Your agent ran out of turns. Your system logs it as an error.

An agent that runs out of budget and an agent that gets it wrong end the same way in your terminal: exit code 1.
For two months my system treated them as the same case, which is wrong, because they call for opposite responses: one resumes, the other gets fixed.
Behind that confusion sits a number almost nobody calibrates and everybody leaves as it came: how many turns to give the agent before cutting it off. With no limit, it can spend hours circling something it won't solve. With too low a limit, it abandons work that was two steps from done.
That number isn't a guess. Only in August, with 30 days of real runs in the database, could I build a table of budgets per task type instead of using the same one for everything.
How it was before: no limit and a hammer
At first there was no turn cap at all. Every step ran with the global default configuration, and the only brake was a watchdog that kills stuck processes after an hour.
A time-based watchdog is a blunt instrument. It can't tell an agent working well on something hard from one stuck in an unproductive loop: both burn an hour the same way. All it guarantees is that nothing runs forever.
That watchdog also had its own operational lesson, which I cover in the final entry: the machine suspends overnight, so suspended time had to be subtracted from the calculation. Otherwise it killed healthy runs over hours that had never actually elapsed.
In June I added a global turn cap as a backstop against runaway runs. Better than nothing, but still a single number for tasks that differ wildly from one another.
The table the data produced
In August the cap became per task type, calibrated by looking at how the last 30 days of runs had actually behaved.
What I like about the result is that the table has a logic you can explain without looking at the numbers:
- The verifier is the most generous, at 160 turns. That makes sense: it reviews someone else's work, and reviewing properly means reading far more than you write.
- The fix implementer follows at 150. A fix usually starts without knowing where the problem is, so a good share of the budget goes into finding it.
- The frontend implementers sit around 90 to 110, depending on whether they work on the admin or the public application.
- The explorer gets 60. It only has to map the terrain and return the touch points: if it needs more than that, something is wrongly framed.
Above all of it sits a global safety ceiling of 200 turns, which no task should ever reach and is there just in case.
The mistake that took me two months to understand
And now the part I consider most important in this entry, because it isn't about the number but about what running out of it means.
For the first two months, an agent that exhausted its turns showed up in the system as "exit code 1". That is: exactly like any other error.
That's a serious diagnostic problem. Running out of budget and failing are completely different things:
- A real failure means something went wrong and you need to investigate what.
- Exhausting the budget means the work was in progress and got interrupted. Nothing is broken: something is unfinished.
Confusing them sent me hunting for bugs that didn't exist, more than once. And the inverse, which is worse: assuming certain failures were "the turn limit again" when they were legitimate errors.
In August I separated them. Running out of budget now has its own message, states clearly what happened, and no longer kills the sequence: the step goes back to the queue instead of dragging the whole task into failure. Which is correct, because the work done up to that point still stands.
Budgets are editable without touching code
A small decision with a large impact: the per-task caps are editable from the interface, with no restart.
The reason is that calibrating them isn't a one-time job. The model changes, project complexity changes, the skills change. A number you adjust by editing code and redeploying is a number that in practice never gets adjusted.
The same criterion, applied to the model
There's a related decision I made at the same time with the same shape: not every step deserves the same model.
The verifier — the most expensive step in the system, that 29% of the bill I described in the entry about measurement — deliberately runs on a cheaper model. It verifies against explicit criteria, and that doesn't require the most capable model available.
Meanwhile, the component making autonomous decisions about what to do next runs on the highest tier with maximum reasoning effort. It's the one that executes least often and has the worst consequences when it gets things wrong.
It's the same idea as the turn budgets, applied along another dimension: budget should follow consequence, not volume.
What I learned
A turn budget looks like a technical parameter and is a product decision: it defines how much autonomy each task gets before someone interrupts it.
And the expensive mistake isn't getting the number wrong — that's corrected with data, and the data shows up on its own if you're measuring. The expensive mistake is failing to distinguish "ran out of budget" from "failed", because the first is resumed and the second is investigated, and confusing them wastes your time in both directions.
What's left to tell
Everything I've covered so far is design decisions: things you think about, plan and implement.
The final entry is about the other half of the work, the half that appears in no architecture diagram and consumed half the time: everything that breaks when you leave the system actually running, for months, on a machine that suspends, restarts and runs out of memory.