This is everything from my summer on one page: what was wrong with omegaUp's background jobs, what I built and why, where each piece stands today, every pull request and issue I opened, and the things I got wrong along the way. It used to be three pages and it is one now, because whoever opens this link should not have to go looking for the rest of it.
Summary
omegaUp is an open source competitive programming platform, used mostly by students and teachers across Latin America. You submit a solution to a problem, it gets judged automatically, and your result feeds into rankings, badges and the suggestions the site makes about what to solve next.
Behind the site a handful of Python scripts run on a schedule overnight. They recompute the rankings, hand out badges, aggregate the feedback students leave on problems, and train the model that decides what to recommend. None of that is visible from the website, and none of it was visible to the people running the website either.
That was the problem. There was no record that a job had run at all, so when one died halfway through, the first sign would be a student noticing days later that a problem they had solved had never counted. Nothing stopped two copies of the same job running at once either, which matters because the ranking job used to delete every row of the rankings table before writing fresh ones. Two copies overlapping does not raise an error. It quietly leaves half a leaderboard and carries on. And an admin who wanted to rerun a job needed shell access to a production server to do it.
While reading omegaUp's production deployment repository I found a commit from December 2021 saying "Make update-ranks only run once a day for now". The "for now" had been holding for five years. In the same afternoon I found that the recommendation model is never trained in production at all, because that job was simply never added to the schedule.
So I built three things that share one foundation. The first is a control plane for the cron
jobs: two new tables that record every job the platform has and every time one of them ran, a
shared runner that wraps each job so it writes itself into that history and takes a lock so two
copies can never overlap, an admin page at /admin/crons that shows all of it, and a
button there that lets an admin rerun a job without touching a server. The second is scheduled
training for the recommendation model, where every run records how good the model it produced was
and a guardrail refuses to publish one that is below a floor or measurably worse than the model it
would replace. The third is a nightly job that looks for problems which silently stopped working,
which needed no dashboard code of its own because it runs through the same shared runner and so
appeared on the admin page by itself.
I opened thirty one pull requests between 26 May and 3 August. Twenty one of my pull requests merged during the coding period, and the whole foundation is on main, including the shared runner and every existing cron job now running through it. The rest are in review or stacked behind something in review, and the detail of exactly what is where is in current status.
In one paragraph: before this summer, if a background job on omegaUp failed at three in the morning, nobody found out until a student complained. Now the run is written down, it cannot collide with itself, an admin can see it and rerun it from a web page without server access, and the two jobs most likely to publish bad data check their own work before they publish it.
About me
I am a third year Integrated B.Tech and M.Tech student in Information Technology at IIITM Gwalior. I have been contributing to omegaUp since August 2025, nine months before this project started, and by the time the coding period began I had twenty one merged changes on the platform including GitHub sign in, human readable contest dates, and the system settings table and its data access layer. That matters mostly because it meant I did not spend the first month of the summer finding my way around the codebase.
How the project changed shape
My original proposal was narrow. It was about cleaning up four Python scripts: five specific SQL problems, the near total absence of unit tests, fourteen bare except blocks, some magic numbers, badge atomicity, and a string replacement involving NOW(). Two of those four scripts had zero Python unit tests between them.
At the midterm I passed, and the feedback was that the work so far did not add up to one thing you could point at. It was many small improvements spread thinly across existing scripts, and what was wanted was something vertical: a single project cutting through the database, the backend, the jobs, the interface and the tests, where the result is a feature somebody can actually see.
Rather than choose the next project alone, I read the production deployment repository so I would be arguing from facts, wrote up fourteen candidate projects, and sent them to my mentor saying I was not attached to any of them and would rather know what the organisation needed. Juan Pablo chose one and named the two that should follow it:
the highest priority is building a Cron Control Plane. Today our cron infrastructure lacks visibility and operational tooling. After that, automated problem health checks and scheduled training for the recommendation model.
So the second half became three projects sharing one foundation, which is how the rest of this page is arranged.
The shared foundation
How the shared runner works
Every job now wraps its main() in one small helper:
with lib.runner.run(parser.prog, args) as cron_run:
with cron_run.phase('update_users_stats'):
update_users_stats(...)
cron_run.set_rows_affected(rows)
- Check the enabled flag. If the registry says this job is switched off, it exits cleanly and records nothing.
- Take a database lock. Two copies of the same job can never run at once.
- Write a running row. So the admin page can see a job that is currently in flight, not only finished ones.
- Time every phase. Each step inside the job is recorded separately.
- Run the guardrail. On the ranking job, the freshly computed ranking is checked before it is published.
- Record the outcome and free the lock. In a finally block, so it always happens.
The lock lives in the database rather than on disk because these jobs can run on different machines, so a lock file on one machine means nothing to a job on another. It also solves the failure I was most worried about: the database ties the lock to the connection, so if a job dies while holding it, the lock releases by itself. A crash cannot jam the door shut forever.
Project one: the cron control plane
What was wrong is the thing described at the top: the jobs ran and left nothing behind. No row saying a job had started, finished, taken a while or fallen over, no protection against two copies overlapping, and no way to rerun one without a shell on a production pod.
What I built is twelve pull requests, arranged as a stack where each one depends on the one
below it. Two tables store the registry of jobs, with their schedule and a flag for whether each
one is switched on, and the history of every run, with the status, the duration, the per step
timings and the error text. A shared runner puts every job into that history and stops them
overlapping. An admin page at /admin/crons shows job health, run durations, per step
timings and error output, and you can click into a single run to see which step was slow or which
one threw. A rerun button lets an admin run a job again without server access.
The rerun button is the piece I thought hardest about. The obvious version is that the button sends a request and the server runs the job. I deliberately did not do that, because the moment a web page can cause a program to run on your servers you have built a door, and doors get picked. Instead the button writes down a request, and a separate trusted worker already running on the inside picks that row up and runs the job, and only jobs on a fixed list it knows about. The web page never runs anything. It leaves a note. Two things fall out of that for free: every rerun is recorded, and the person pressing the button needs no server access, which was the entire point.
The ranking job also gets a guardrail. If the newly computed ranking is empty, has negative scores, or has lost more than half the previously ranked users, it raises and the transaction rolls back. I would rather have yesterday's rankings than today's broken ones. I proved it rather than asserting it: on a forced failure the job exits non zero and the checksum of the rankings table is identical before and after, which is what "the rollback left the data untouched" actually means.
Below is a walkthrough of the dashboard, recorded from a local branch with the whole stack applied.
Project two: scheduled recommendation model training
omegaUp has a model that answers "this student just solved problem X, what should they try next". It was an on demand script. Nobody was on the hook for running it, there was no record of what past training produced, and a bad run would silently overwrite the good model file serving real students.
It is now in the registry with a weekly schedule, which means it is scheduled at all for the first time. Every run records how good the model it produced was, using a MAP score. That is a number between zero and one measuring how often the problem the model suggested next turned out to be one the student actually solved, weighted so a good suggestion near the top of the list counts for more. Alongside it each run also records precision, recall and NDCG over held out users, because one number on its own does not tell you much about a ranking.
A guardrail applies two rules before the model file is written: an absolute floor, and a regression bar against the last published model. The floor alone only catches disasters, because a model that squeaks over the bar while being clearly worse than the one it replaces would sail straight through. I ran it both ways in Docker, watched a weak model get refused with the reason stored, and a good one get published.

Above, the guardrail working: one model published at 0.3419, and one refused at 0.2151 with the reason recorded rather than hidden.
Project three: automated problem health checks
A problem on omegaUp can stop working after it is published, and nothing errors loudly. Students just hit a wall, and the only signal was somebody filing a report, which means the platform found out about broken problems from the people it had already failed. A nightly job now looks for four things.
| Check | Severity | What it means |
|---|---|---|
judge_errors |
error | Five or more recent submissions to one problem came back as a judge or validator error. Students are submitting correct code and getting a system failure. To the student it looks like their fault. |
no_languages |
error | A public problem with no submission languages enabled. Listed, browsable, and impossible to submit to. |
never_solved |
warning | Public, twenty or more submissions, zero accepted. Either the test data is wrong or the statement is misleading. |
deprecated_public |
warning | Retired, but still visible to students. |
Findings are upserted against a unique key, so each one keeps the date it was first detected. That is the difference between "this is broken" and "this has been broken for three weeks". A finding that stops appearing is marked resolved rather than deleted.
This job needed no dashboard code of its own. It runs through the shared runner, so it appeared
on /admin/crons automatically. That is the moment the foundation paid for itself, and
it is the reason the first project was worth building as a foundation rather than as one more
feature.
Work alongside the main project
Two other pieces ran alongside the cron project and both are merged.
The first was a generic way to tell somebody that part of the site is switched off, in #9919. The easy version would have been to write that message into the one page that needed it. I built a component the whole platform can use instead, in every language the site supports. Juan Pablo then pointed out that the way I had wired it in would mean copying a file every time another view needed disabling, so #10025 made the entry point generic and driven by the page payload.

Because the message and its heading come from the translation files rather than from the page that raised it, it arrives in whatever language the reader is already using. The same component in Spanish is below.

The second was rebuilding the navigation menus from a single configuration, in #9968, building on #9871 and #9801. The menus had been hand written markup repeated per menu, so I extracted a reusable item with an icon, a title and a description, then rebuilt every menu from one configuration file with per entry visibility rules.
The clearest illustration of what a review is for is the pair below. First I shipped entries reading "Create zip file" and "I have a zip file", which make sense only if you already know what they do. Juan Pablo asked for something more descriptive, and the second image is the result.



The rest of the menus came from the same configuration file rather than from separate markup, which is what made adding an icon and a description to all of them a single change instead of five.




Current status
This is where everything stood on 23 August 2026. The status column in the tables further down refreshes from GitHub when this page loads, so if something has moved since I wrote this, those tables will say so and this section is the part that will be out of date.
The foundation is merged, and so is every job that now runs through it. That means the registry and run history tables in #9995, the data access layer in #9996, the shared runner itself in #9998, the admin API in #10000, and the three changes that put the existing jobs through the runner: update_ranks in #10003, assign_badges and aggregate_feedback in #10009, and the remaining scripts in #10022. The problem health checks table in #10065 is on main, and so are the reliability fixes to the scripts that were already there: the parameterised coder of the month queries, the typed exception handling, the ranking upsert and the named acl and role constants.
So as of today, every cron job on omegaUp writes itself into a run history, takes a lock that stops a second copy starting, and times each of its phases separately. That part is done and running on main. The runner was the keystone and it landed on 19 August, which unblocked the three changes that put the jobs through it, and those landed over the following three days.
Nine pull requests are open and in review. All nine merge cleanly against main with no failing checks, and #9883 is approved. They are the admin dashboard in #10008, the ranking guardrail in #10011, structured phase logging in #9914, the cron unit test infrastructure in #9883 and the tests built on it for update_ranks in #9889 and assign_badges in #9920, the recommendation model runs table in #10050 and the training job that records and guards itself in #10051, and the nightly problem health check job in #10066.
Five more are drafts, and they are drafts on purpose. Each one sits directly behind a pull request that has not merged yet, so raising it now would put a diff in front of a reviewer that contains an earlier change as well as its own. The rerun button in #10021, the dashboard end to end test in #10023 and the dashboard improvements in #10048 are behind the dashboard. The model quality table in #10052 is behind the model training pair, and the admin view of health findings in #10069 is behind the health check job. Those last two add their own panels to the dashboard page, so all five of them ultimately need #10008, which is why it is the one I would push hardest. Each surfaces as the one below it lands, and every one of them has been verified end to end in Docker.
One pull request from the summer was closed without merging, #9967, because I replaced it with #9968, which rebuilt the menus from one configuration instead of extending the old markup.
What sets the pace now is review time rather than code. That is a real constraint on a volunteer maintained project rather than a complaint. Nothing is half built and nothing is blocked on me.
Pull requests
I opened 31 pull requests between 26 May and 3 August, and 16 of those have merged so far. Counting by merge date instead, 21 of my pull requests landed during the coding period, because five of them had been waiting since before it began. Each one had an issue written before it, so that the reason for a change existed in writing before the change did.
The control plane pull requests are stacked, each built on the one before it, which means GitHub's diff for a later one still carries its unmerged predecessors and shrinks to just its own change the moment the one below it lands. That is deliberate. It is more work for me and far less to read for whoever is reviewing.
| Component | Description | Status | PR |
|---|---|---|---|
| Cron control plane | add cron control plane tables | merged | #9995 |
| add cron control plane dao | merged | #9996 | |
| add cron runner library | merged | #9998 | |
| add cron admin api | merged | #10000 | |
| record update ranks runs | merged | #10003 | |
| add cron admin dashboard | in review | #10008 | |
| record badges and feedback runs | merged | #10009 | |
| add ranking audit guardrail | in review | #10011 | |
| add cron rerun | draft | #10021 | |
| record remaining cron runs | merged | #10022 | |
| add cron dashboard e2e | draft | #10023 | |
| improve cron dashboard with health cards filters and auto refresh | draft | #10048 | |
| Cron tests, queries and reliability | feat(cron): add unit test infrastructure and aggregate_feedback/utils tests | approved | #9883 |
| feat(cron): add update_ranks unit tests | in review | #9889 | |
| fix(cron): parameterize coder_of_the_month queries | merged | #9900 | |
| fix(cron): replace bare except clauses with typed except | merged | #9901 | |
| fix(cron): merge user rank with upsert instead of full reinsert | merged | #9903 | |
| add structured phase logging to crons | in review | #9914 | |
| feat(cron): add assign_badges unit tests | in review | #9920 | |
| refactor(cron): name acl/role magic numbers with constants | merged | #9940 | |
| Recommendation model training | add recommendation model runs table | in review | #10050 |
| record and guard recommendation training | in review | #10051 | |
| add recommendation model dao and dashboard | draft | #10052 | |
| Problem health checks | add problem health checks table | merged | #10065 |
| add problem health check cron | in review | #10066 | |
| show problem health findings for admins | draft | #10069 | |
| Alongside the main project | Extract NavbarItem component for help menu | merged | #9871 |
| add view unavailable component | merged | #9919 | |
| extend navbar item style to remaining submenus | closed | #9967 | |
| build navbar menus from a single configuration | merged | #9968 | |
| make view unavailable entrypoint generic | merged | #10025 |
Issues
Every change above has an issue behind it, written first. The one oddity in this list is #9961, which I opened by accident and closed the same day. It is here because the list is meant to be complete rather than tidy.
| Issue | Title | Status |
|---|---|---|
| #9868 | As a developer, I want a reusable NavbarItem component so that submenu entries are easier to maintain | completed |
| #9869 | As a contributor, I want editable fields in the feature request form so that I can fill in the As a / I want / so that prompts directly | completed |
| #9870 | Refactor navbar help submenu into a reusable NavbarItem component | completed |
| #9882 | Add unit test infrastructure for cron scripts | open |
| #9888 | Add unit tests for update_ranks.py ranking logic | open |
| #9890 | Add unit tests for assign_badges.py | open |
| #9897 | Add a generic "View unavailable" component | completed |
| #9899 | Parameterize string-built queries in coder_of_the_month.py | completed |
| #9902 | Avoid full delete and reinsert of User_Rank in update_ranks | completed |
| #9904 | Prevent and merge duplicate school profiles | open |
| #9913 | Emit structured per-phase logs from cron scripts | open |
| #9939 | Replace acl_id/role_id magic numbers in cron scripts with named constants | completed |
| #9961 | duplicate, opened by mistake and closed the same day | not planned |
| #9962 | Extend the navbar item style with icons and descriptions to remaining submenus | open |
| #9966 | Build the navbar menus from a single configuration with visibility rules | completed |
| #9992 | Cron Control Plane : an observability and operational tooling for background jobs (parent issue) | open |
| #9993 | persistant cron execution history by registry and run tables | completed |
| #9994 | DAOs for the cron tables | completed |
| #9997 | Reusable cron runner that records runs and prevents overlap | completed |
| #9999 | Admin API for cron run history and health | completed |
| #10002 | Record update_ranks executions through the runner | completed |
| #10004 | Admin dashboard page for cron jobs | open |
| #10007 | Record assign_badges and aggregate_feedback through the runner | completed |
| #10010 | Pre-publish guardrail for the ranking job | open |
| #10017 | Make the view unavailable entrypoint generic and driven by the payload | completed |
| #10018 | Let admins safely rerun a cron job | open |
| #10019 | Record the remaining crons through the runner | completed |
| #10020 | End to end test for the cron dashboard | open |
| #10047 | Improve the cron dashboard with health cards, filters, relative times and auto refresh | open |
| #10049 | Make the recommendation model training scheduled, recorded and guarded (parent issue) | open |
| #10064 | Detect problems that silently stopped working (parent issue) | open |
What I learned and what went wrong
The single biggest change is how I write tests. I used to write the test after the fix and move on the moment it went green. Now I break the code on purpose first and check that the test actually fails, because a test that passes against broken code is decoration. It costs about a minute and it caught two real bugs in my own work this summer, both in code I had written minutes earlier and was certain was correct.
The first was in the health checks. The step that closes findings which are no longer detected was reading the clock after writing the new findings rather than before. On a slow run, a finding written moments earlier could be judged as not seen this run and closed immediately, so the job would report a problem as fixed at the instant it discovered it. The fix was to take one timestamp at the start and use it everywhere. I only found it because I wrote the test, and I only trusted the fix because I put the bug back afterwards and watched the test go red.
The second one I nearly shipped without noticing at all. Three of the four health checks fired on their first real run and the judge errors one found nothing. My first reaction was relief. That was wrong. It found nothing because the development database contained no submissions with that kind of error in it, so the check had never been exercised even once. I seeded the failure case and watched it fire before I was willing to trust it. No results and not working look identical from the outside, and I now treat an empty result from a new check as unproven rather than as good news.
A review taught me the same lesson from the other side. Ankit went through the test
infrastructure I had built for the cron scripts and found that my fake database cursor never
advanced, so fetchone returned the first row forever. Every test standing on top of it
was passing for a reason that had nothing to do with the code under test. Being told early and
precisely that my fakes were subtly wrong was worth more to me than an approval would have been,
and it is why I now write a test for the test helper before I trust anything built on it.
The hardest thing I did was withdraw a promise from my own funded proposal. I had committed to rebuilding the rankings in a second table and swapping it in with a rename. When I sat down to research it properly I found five problems. A rename is a schema change, so it forces a commit and destroys the single transaction the phase depends on. Copying a table does not bring its foreign key constraints with it, and constraint names are unique per database, so the swapped table either loses three of them or needs an awkward alternating naming scheme. It needs privileges the production database user may not have. It changes what the coder of the month calculation sees while it runs. And it still rewrites every row every day, so the write load is not reduced, only moved.
Then I found the part that settled it. The problem the swap was meant to solve had already been fixed by somebody else while I was writing the proposal. The phase had become a single transaction, so readers already see the complete old ranking the whole way through, and the empty leaderboard I had written the proposal against no longer existed. I wrote all of that up and posted it rather than quietly building something smaller and hoping nobody compared it against the proposal. It felt bad to do and it was obviously the right call. Nobody is served by a contributor who ships something they know does not work in order to match a document.
The hardest part of the summer had nothing to do with code. At the midterm, not one of my pull requests for the main project had merged. Every one had green checks and every one was sitting in the queue, and there was no version of working harder that would change it. Up to that point every problem I had hit was one I could solve by putting in more hours. My mentors were straightforward that review bandwidth was the constraint and told me to keep building, so I kept building. The first landed on 3 July and sixteen more have landed since.
What I took from that is that on a project maintained by volunteers, how reviewable your work is matters as much as whether it is correct. A large correct change and a small correct change are not worth the same thing to the person who has to read them. I used to think a bigger pull request showed more work. It mostly shows less consideration for the reviewer.
Acting on that had a cost I underestimated. Keeping pull requests small and ordered is right for reviewers and expensive for me. Every time one merges I rebase the next so its diff stands alone, and when the tables merged the migration number changed, so every later database change had to shift with it. At one point a change to a shared file had to be propagated across seventeen branches, because all of them descended from it. I would work the same way again, but I would plan for the rebasing rather than treating it as something that happens to me.
One problem was purely mechanical and taught me something anyway. After the tables merged, one required check insisted that a generated schema file be byte for byte identical to the main schema, and another required check crashed trying to read that same file. Both were mandatory, so satisfying one broke the other. The crash was at one specific table, caused by a full text index clause another contributor had added that the schema parser did not understand. Nobody had hit it because main's copy of the generated file was out of date and did not contain that line yet, so mine was the first change to regenerate it. The fix was one line in the parser. When two automated checks contradict each other, the contradiction is usually pointing at something real rather than at you.
I also learned to be careful about what a green page actually means. On one pull request the PHP job showed as skipping rather than failing, because it declares a dependency on another job that had died, so the PHP tests for that change had never run at all. Nothing was red and nothing had been tested. Skipping is not passing, and I read the job list now instead of the summary.
The last thing is about claims, because there are three I could phrase so they sound better than they are, and I would rather say them plainly. The system can tell you a job ran and failed, but it cannot tell you a job never started, which is the one failure mode run history structurally cannot catch. The lock lives in the database, so it is only as available as the database is, and it is an efficiency lock: the correctness comes from the transaction underneath it rather than from the lock itself. And the model guardrail proves that a new model is not worse than the last one on an offline number, not that it is better for the students using it. All three are fine for what they are, and none of them is more than that.
What is left to do next
The gap I care most about is the first of the three limits above, that the system cannot tell you a job never started. The schedule is already stored in the registry, so the pieces for it are there and it is the next thing I want to build. After that, failure alerting, reusing the notification path the rerun dispatcher already uses, which is the piece that turns history into monitoring. Then retention, because nothing deletes from the run history, which is fine today and will not be in two years.
Smaller things: pagination past the fifty run cap, a duration sparkline per job, and a next scheduled run column computed from the schedule already stored. I also deliberately left model versioning and rollback alone, because the model file layout and the serving side live in omegaUp's production repository and building it blind risked doing it wrong.
Before any of that, the work that is already written has to finish landing, and I intend to be there for it. I was contributing to omegaUp before this project started and I am not stopping at the end of the coding period.
Acknowledgments
Thanks to Juan Pablo for reviews that were specific and for pushing back on the parts that deserved it. The question about storing a name instead of an id made the schema better, and the request to split a large pull request changed how I work rather than just how I structured that one change.
Thanks to Ankit for the review of my test infrastructure that found real bugs in it, including a fetchone that never advanced and so returned the first row forever. Being told early and precisely that my fake database was subtly wrong was worth more than any approval.
And thanks to omegaUp for letting a contributor propose a direction, argue for it with evidence, and then build it.
All my pull requests · All my issues · The epic · Progress tracker