You can't restart the system that's running your work

There's a paradox that shows up whenever you build the tool you work with: to improve it you have to restart it, and restarting it kills the work in flight.
In my case the cost of that paradox was brutal, because the work queue lived in the process's memory. Every server restart — a deploy, an update, a power cut — wiped the entire plan. Not just what was pending: also any trace of what had been running and what state it was left in. And the server is precisely the thing I'm modifying all the time.
The fix isn't "store the queue in a database" and done. What you actually have to decide is what the state of a half-finished task means, and who has the right to change it when the process comes back up.
How it used to work
The original queue was an in-memory list. You ticked the tasks you wanted to run on screen, the order was the order you'd selected them in, and the system processed them one by one.
It worked as long as nothing interrupted it. With two serious problems.
The first is the one above: every time I wanted to deploy an improvement to the orchestrator, I had to choose between waiting for the queue to drain or losing the plan.
The second, quieter one: the first failure aborted the whole queue. An error in the third of ten tasks cancelled the remaining seven. And those seven had nothing to do with the one that failed: they were simply behind it in line.
The redesign
The queue moved to the database, with one design decision I particularly like: position is a real number, not an integer.
That allows inserting between two elements without renumbering the whole list. If you want to slot something between positions 3 and 4, you assign it 3.5. It's an old technique and it solves at the root the problem of reordering a persisted list without writing fifty rows every time someone drags one.
The order is curated by hand, mixing fixes, features and audits into whatever sequence makes sense that day. You reorder by dragging, as long as tasks are still waiting. Whatever is running is pinned and locked, which is the bare minimum: you can't reorder something that already started.
What to do when a task fails
There was a product decision disguised as a technical detail here, and I resolved it by not resolving it: I made it a per-run option.
You can choose for the queue to carry on with the rest, flagging the failed one, or to stop and leave everything else waiting. Both positions are defensible depending on the case. If you're processing ten independent fixes, you want it to continue. If the third failed because the environment broke, you want it to stop before the remaining seven fail for the same reason.
Later I added a stricter rule that turned out to be the most useful in practice: the queue doesn't resume past an unresolved failure. You can continue, but someone has to have looked at that failure and decided what to do about it. Without that, the queue stays put.
It's deliberately inconvenient. The alternative — always continue — lets failures pile up unseen, which is exactly the problem behind the 1,771 unprocessed incidents I described a few entries back.
A single execution mechanism
A side effect of the redesign I hadn't anticipated: the "run everything" and "repair everything" buttons stopped being separate mechanisms.
Before, each was its own pipeline, with its own advancement logic and its own bugs. Now they're simply shortcuts that bulk-enqueue with a preset order and start the queue. One execution mechanism, all of it visible and reorderable.
It's the kind of simplification that appears on its own once the right abstraction finally exists. While the queue was ephemeral, having parallel pipelines seemed reasonable. With a persistent, curatable queue, it makes no sense at all.
Actually surviving a restart
Persisting the queue solves half the problem. The other half is what happens to whatever was running at the moment of the restart.
That task lands in limbo: the database says it's in progress, but the process executing it no longer exists. If nobody does anything, it sits there forever, occupying the head of the queue and blocking everything behind it.
In August I added the two missing pieces: graceful shutdown, so in-flight runs are marked before dying, and startup reconciliation, which checks what's declared as running with no process behind it and resolves it.
And a rule that seems obvious and took me a while to find: re-read the task's status before launching its row. Because between enqueuing and its turn arriving, anything could have happened — the task might have been completed by another path, or postponed. Launching blind against a stale snapshot is a constant source of duplicated work.
Why this sums up the whole project
Of everything I built over these months, the persistent queue is what best captures the thesis: in a system with agents, the state of in-flight work matters as much as the work itself.
An in-memory queue works perfectly until the first restart, and the first restart always arrives at the worst moment. It isn't an optimisation: it's the difference between a tool you use while you're watching and a system that runs when you're not.
Today the queue is the system's only execution path. There's no way to run a task outside it, and that — which felt like an annoying restriction at first — turned out to be what makes everything else observable.
With the queue surviving restarts, the system could run on its own for hours with me nowhere near it. And that raised a question I'd never asked, because I'd never let it run that long: how long to let it think before interrupting. That's what the next entry is about.