Repository navigation
no resume after a provider error mid-turn #913
Description
Activity
@Svector-anu thanks for splitting this out, and sorry it sat for a couple of days.
Confirmed, and I think I can name the mechanism rather than just agreeing with the symptom.
The agent loop is not the part that loses the work. It deliberately preserves what it has on every provider-error return,
internal/agent/loop.goaround lines 468, 476 and 494:result.Messages = copyMessages(messages) return result, err
So the tool results from the interrupted turn, including the files already read, do survive the error and come back to the caller.
The caller drops them.
agentResponseMsgininternal/tui/model.go:643has no messages field at all, and the error path at 5822 builds one from rows, usage events, session events and the error only. Nothing carriesresult.Messagesback into the model, on the error path or the success path.So the next turn does not have the interrupted turn's tool results to resume from, and re-reading from scratch is the model doing the only thing left to it. That matches your report exactly: not just the failed request lost, but the work in progress.
One thing your issue got right that I want to underline, because it changes where the fix goes: this is not a read-tracking bug.
FileTrackeris created once per session (internal/cli/app.go:918) and passed into every turn, so Zero itself does remember which files were read. It is the model's context that gets rebuilt, not Zero's bookkeeping. A fix that only touched the tracker would look right and change nothing.What I have not traced yet, and would want to before anyone writes a patch: exactly how the next turn's history is assembled, and whether partial tool results from a failed turn are persisted to the session at all or only at turn end. If they are only written at turn end, then carrying
result.Messagesthrough is necessary but not sufficient, and the persistence point has to move too.Leaving this open and assigned to me. The repro session id and the event numbers you pinned in #674 are what made it possible to answer this in an hour rather than a day, so thank you for that.
- added a commit that references this issue
on Sep 8, 2026 Labelled as a bug. The fix is #1016, in review; it cites this issue and closes it on merge.
- added 4 commits that reference this issue
on Sep 9, 2026 - added a commit that references this issue
on Sep 12, 2026
no resume after a provider error mid-turn
Split out of #674 (issue 2) per @Vasanthdev2004's request.
After a provider error interrupts a turn, the next turn does not resume from
where it left off — it re-reads previously-read files from scratch to rebuild
context. Work in progress at the time of the error is lost, not just the
single failed request.
Repro context: same session referenced in #674,
zero_20260713163241_1783960361665303000_1— a provider error mid-turn (see #674's events #361/#586/#632/#646/#669/#689)
was followed by the next turn re-reading files it had already read earlier in
the session, rather than resuming from the interrupted state.
Expected: on retry after a provider error, the turn should resume from the
last-known state (files already read, partial tool results, etc.) instead of
rebuilding context from scratch.