Skip to content

Agentic unit test fixing - #325

Merged
VitjanZ merged 5 commits into
mainfrom
feat/agentic-unit-test-fix
Oct 2, 2026
Merged

VitjanZ merged 5 commits into
mainfrom
feat/agentic-unit-test-fix

Conversation

@VitjanZ

@VitjanZ VitjanZ commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator

Agentic workflow implementation - agentic unit-test fixer

Description:
Replaces the stateless one-shot unit-test fix with an agent session that spans every fix attempt for a functionality. FixUnitTests no longer makes one stateless /fix_unittests_issue call per failed test run. Instead it drives a server-side agent session that can explore the code, edit it, run the tests itself, and remember everything it tried. The old fixer often one-shot a fix, but it had no memory between attempts, so on harder failures it went back and forth between issues.

How FixUnitTests works now

  • First failure: the action starts an agent session with the spec, the failing test output and the relevant code (/agent/start).
  • Within an attempt: the action runs the agent's tool calls locally against the build folder (read, search, edit, run the unit tests), sends the results back (/agent/continue), and repeats. It stops when the agent calls submit_fix, or after at most 40 turns.
  • After a submit: the state machine re-runs the unit tests as before. If the agent's own last run already passed and nothing changed since, that re-run is skipped.
  • Next failure: it goes back into the same session as the answer to the agent's submit_fix. The agent sees its previous attempts instead of starting from scratch.
  • Abandoned session (turn budget used up, or an LLM failure): the next attempt starts a new session, with a summary of what the abandoned one tried.

Release note: this needs the server from codeplain-api feat/agentic-unit-test-fix. An older server returns 404 on /agent/start, so release this only after that server is live in production.

Replace the stateless /fix_unittests_issue call with an agent loop: the first
failure starts a session via /agent/start, the client executes the tool calls
the model asks for (read/grep/ls/edit/write/delete files, run the unit tests)
and posts results to /agent/continue until the agent calls submit_fix. The
state machine then re-runs the unit tests as before; a repeated failure is fed
back into the same session as the submit_fix result, so earlier attempts stay
in context instead of being re-diagnosed from scratch.

- render_machine/agent_tools.py: tool implementations confined to the build
  folder (writes) and project root (reads); changed files are tracked in the
  unit tests running context so they reach the FRID commit as before.
- FixUnitTests: start/continue the session, per-attempt turn cap, fresh
  session when the previous one ended without a submission.
- UnitTestsRunningContext carries the session id and the pending submit_fix
  call; it is recreated per unit-test loop so a session spans one FRID.
- Send raw test output as test_output (run_unit_tests and the submit_fix
  follow-up) so the server condenses it, instead of head/tail truncation that
  can drop the root cause.
- Make the full test logs readable by read_file/grep and point the agent at
  them; grep gains context_lines and include.
- Seed the first turn with the build folder's file tree and the files changed
  for the current FRID.
- Skip the harness unit-test run when the agent's own run passed and no file
  changed since.
- Start a new session after an abandoned one with previous_session_id so the
  server can pass on what was tried.
- edit_file returns the edited region; repeated read-only calls with no file
  change in between return a pointer instead of the same output again.
…e conformance fixes

After the conformance tests fixer changed implementation code, the unit
tests failed and a fresh agent session "repaired" the implementation back
to what the stale unit tests asserted. The conformance fixer re-applied
its change and the loop repeated (17 times on one FRID in render
b6f3ff73), each new session unaware it had already done this.

- The agent session state moves into UnitTestsAgentSession. Outside the
  conformance phase it lives on the unit-tests running context (one
  session per loop, as before); during the conformance phase it lives on
  the conformance tests running context, so one session spans every
  unit-test loop of the phase (RenderContext.unit_tests_agent_session).
- A new session gets the conformance tests fixes recorded so far
  (implementation_code_fixes) as task_params.conformance_tests_fixes,
  with the files they changed seeded into relevant_files.
- When a later unit-test loop continues the session, the pending
  submit_fix is answered with "your fix was accepted, then the code was
  changed to fix the conformance tests", only the fixes the session has
  not seen yet, and the new failure output.
- The conformance-context tests are rewritten for the agentic path (they
  still targeted the stateless /fix_unittests_issue call and errored).
Agent sessions expire on the server (Redis TTL). Continuing an expired one
returned 404, which failed the render; FixUnitTests now starts a new session
instead. Other errors are raised as before.
@pedjaradenkovic

Copy link
Copy Markdown
Contributor

There are 2 security things that I would fix in the follow up PR:

  1. Secrets can leave the machine. _readable() allows reads anywhere under the CWD (the project root), not just the build folder. The agent can therefore read_file or grep a .env or credentials file in the project root, and that content goes into the session history and to the LLM provider. I'd restrict reads to the build folder, plus the registered log paths and any linked resources.
Here I would:  Add a denylist inside the build folder. The build folder can also contain secrets, such as a generated .env or a copied key file. Refuse basenames and patterns like .env*, *.pem, *.key, id_rsa*, credentials* and .git/. GREP_EXCLUDED_DIRS already skips .git for grep, but read_file and ls_files don't apply it.
  1. Symlinks. The path checks use normpath/abspath, not realpath. A symlink inside the build folder can point outside it and pass the read check. The same goes for the write check, so writes could also escape the build folder.

Not a real threat for now but can become when we get a lot of users. We can discuss the proper solution.

@pedjaradenkovic pedjaradenkovic left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Don't forget to increase the min server version with this.

@dusano

dusano commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

I agree with #1 that only build folder and folders Predrag mentioned should be allowed.

With regards to #2 I don't see issues with symlinks. If the developer wants to give renderer access to some other location through symlinks, that's their responsibility.

read_file, grep and ls_files also accepted anything under the directory
codeplain runs from, which exposes the specs, other modules and any .env
in the project root; whatever the agent reads goes to the server and the
LLM. Reads are now limited to the build folder plus the test logs the agent
is pointed to.
@VitjanZ
VitjanZ merged commit a87a7d9 into main Oct 2, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants