Agentic unit test fixing - #325
Merged
Merged
Conversation
Replace the stateless /fix_unittests_issue call with an agent loop: the first failure starts a session via /agent/start, the client executes the tool calls the model asks for (read/grep/ls/edit/write/delete files, run the unit tests) and posts results to /agent/continue until the agent calls submit_fix. The state machine then re-runs the unit tests as before; a repeated failure is fed back into the same session as the submit_fix result, so earlier attempts stay in context instead of being re-diagnosed from scratch. - render_machine/agent_tools.py: tool implementations confined to the build folder (writes) and project root (reads); changed files are tracked in the unit tests running context so they reach the FRID commit as before. - FixUnitTests: start/continue the session, per-attempt turn cap, fresh session when the previous one ended without a submission. - UnitTestsRunningContext carries the session id and the pending submit_fix call; it is recreated per unit-test loop so a session spans one FRID.
- Send raw test output as test_output (run_unit_tests and the submit_fix follow-up) so the server condenses it, instead of head/tail truncation that can drop the root cause. - Make the full test logs readable by read_file/grep and point the agent at them; grep gains context_lines and include. - Seed the first turn with the build folder's file tree and the files changed for the current FRID. - Skip the harness unit-test run when the agent's own run passed and no file changed since. - Start a new session after an abandoned one with previous_session_id so the server can pass on what was tried. - edit_file returns the edited region; repeated read-only calls with no file change in between return a pointer instead of the same output again.
…e conformance fixes After the conformance tests fixer changed implementation code, the unit tests failed and a fresh agent session "repaired" the implementation back to what the stale unit tests asserted. The conformance fixer re-applied its change and the loop repeated (17 times on one FRID in render b6f3ff73), each new session unaware it had already done this. - The agent session state moves into UnitTestsAgentSession. Outside the conformance phase it lives on the unit-tests running context (one session per loop, as before); during the conformance phase it lives on the conformance tests running context, so one session spans every unit-test loop of the phase (RenderContext.unit_tests_agent_session). - A new session gets the conformance tests fixes recorded so far (implementation_code_fixes) as task_params.conformance_tests_fixes, with the files they changed seeded into relevant_files. - When a later unit-test loop continues the session, the pending submit_fix is answered with "your fix was accepted, then the code was changed to fix the conformance tests", only the fixes the session has not seen yet, and the new failure output. - The conformance-context tests are rewritten for the agentic path (they still targeted the stateless /fix_unittests_issue call and errored).
Agent sessions expire on the server (Redis TTL). Continuing an expired one returned 404, which failed the render; FixUnitTests now starts a new session instead. Other errors are raised as before.
Contributor
|
There are 2 security things that I would fix in the follow up PR:
Not a real threat for now but can become when we get a lot of users. We can discuss the proper solution. |
pedjaradenkovic
approved these changes
Oct 1, 2026
pedjaradenkovic
left a comment
Contributor
There was a problem hiding this comment.
Don't forget to increase the min server version with this.
Contributor
read_file, grep and ls_files also accepted anything under the directory codeplain runs from, which exposes the specs, other modules and any .env in the project root; whatever the agent reads goes to the server and the LLM. Reads are now limited to the build folder plus the test logs the agent is pointed to.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Agentic workflow implementation - agentic unit-test fixer
Description:
Replaces the stateless one-shot unit-test fix with an agent session that spans every fix attempt for a functionality. FixUnitTests no longer makes one stateless /fix_unittests_issue call per failed test run. Instead it drives a server-side agent session that can explore the code, edit it, run the tests itself, and remember everything it tried. The old fixer often one-shot a fix, but it had no memory between attempts, so on harder failures it went back and forth between issues.
How FixUnitTests works now
Release note: this needs the server from codeplain-api feat/agentic-unit-test-fix. An older server returns 404 on /agent/start, so release this only after that server is live in production.