Skip to content

no resume after a provider error mid-turn #913

Description

@Svector-anu

no resume after a provider error mid-turn

Split out of #674 (issue 2) per @Vasanthdev2004's request.

After a provider error interrupts a turn, the next turn does not resume from
where it left off — it re-reads previously-read files from scratch to rebuild
context. Work in progress at the time of the error is lost, not just the
single failed request.

Repro context: same session referenced in #674, zero_20260713163241_1783960361665303000_1
— a provider error mid-turn (see #674's events #361/#586/#632/#646/#669/#689)
was followed by the next turn re-reading files it had already read earlier in
the session, rather than resuming from the interrupted state.

Expected: on retry after a provider error, the turn should resume from the
last-known state (files already read, partial tool results, etc.) instead of
rebuilding context from scratch.

Activity

  1. Vasanthdev2004 commented on Aug 19, 2026

    @Vasanthdev2004
    Collaborator

    @Svector-anu thanks for splitting this out, and sorry it sat for a couple of days.

    Confirmed, and I think I can name the mechanism rather than just agreeing with the symptom.

    The agent loop is not the part that loses the work. It deliberately preserves what it has on every provider-error return, internal/agent/loop.go around lines 468, 476 and 494:

    result.Messages = copyMessages(messages)
    return result, err

    So the tool results from the interrupted turn, including the files already read, do survive the error and come back to the caller.

    The caller drops them. agentResponseMsg in internal/tui/model.go:643 has no messages field at all, and the error path at 5822 builds one from rows, usage events, session events and the error only. Nothing carries result.Messages back into the model, on the error path or the success path.

    So the next turn does not have the interrupted turn's tool results to resume from, and re-reading from scratch is the model doing the only thing left to it. That matches your report exactly: not just the failed request lost, but the work in progress.

    One thing your issue got right that I want to underline, because it changes where the fix goes: this is not a read-tracking bug. FileTracker is created once per session (internal/cli/app.go:918) and passed into every turn, so Zero itself does remember which files were read. It is the model's context that gets rebuilt, not Zero's bookkeeping. A fix that only touched the tracker would look right and change nothing.

    What I have not traced yet, and would want to before anyone writes a patch: exactly how the next turn's history is assembled, and whether partial tool results from a failed turn are persisted to the session at all or only at turn end. If they are only written at turn end, then carrying result.Messages through is necessary but not sufficient, and the persistence point has to move too.

    Leaving this open and assigned to me. The repro session id and the event numbers you pinned in #674 are what made it possible to answer this in an hour rather than a day, so thank you for that.

  2. Vasanthdev2004 commented on Sep 8, 2026

    @Vasanthdev2004
    Collaborator

    Labelled as a bug. The fix is #1016, in review; it cites this issue and closes it on merge.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions