Skip to content

feat: summarize a project for coding agents in one call with vf context (COR-14197) - #36

Open
Bradenream wants to merge 8 commits into
braden/vf-link/COR-14197from
braden/vf-context/COR-14197
Open

Bradenream wants to merge 8 commits into
braden/vf-link/COR-14197from
braden/vf-context/COR-14197

Conversation

@Bradenream

Copy link
Copy Markdown
Contributor

Stacked on #35 (vf link); this PR's diff is only the vf context work. Review #35 first. Once it merges, this PR retargets to master.

Summary

A coding agent's first job on a Voiceflow project is finding out what it is. With today's CLI that means one call per resource type, each an agent turn. The API has no single read that fits in a context window: the v2 export is not in the SDK, and at 100 KB+ it would flood one anyway.

vf context makes 11 typed reads itself, 4 at a time under a 60 s deadline, and returns an outline:

  • the model, global prompt and instructions (line counts and an excerpt)
  • playbooks, summarized by the description the agent routes on
  • functions and the agent's own tools, named after what they call
  • variables and the knowledge base
  • recent changes, newest first, and the latest conversations
  • the rules of the platform (compile, draft vs published, publish), stated as facts
  • drillDown: the commands for anything the outline leaves out

Lists are capped newest first, text is clipped, and the true totals are under counts. A test holds the worst case under 20 KB of TOON; a large real project comes to about 9 KB.

If a part can't be read, the outline still prints: that part's count is null and warnings names it. Two limits are stated rather than guessed:

  • The API doesn't say who changed a resource.
  • Some resources carry no edit time: instructions, the global prompt, settings and KB documents. When the project record moved after everything the outline can date, unexplainedChange says the change can't be identified with vf.

vf link's agent snippet now starts with vf context.

Measured with cold coding agents

Same question to every run, read-only harness, a real project, 3 runs per variant. Medians; calls and bytes come from the harness log.

vf calls CLI output wall clock answer
today's CLI 34 1.8 MB 271 s accurate
this branch, no hints 5 17 KB 146 s accurate
this branch, linked folder + snippet 1 9 KB 126 s accurate

The runs also shaped the code:

  • An open-ended note sent agents hunting (28 and 17 calls). unexplainedChange now states when a change can't be identified.
  • Agents distrusted an imperative in tool output. The output now states facts only; instructions stay in the snippet for the user's own instructions file.
  • The most recently edited function could fall outside the cap. Capped lists are now newest first.

Test plan

  • gofmt, go vet ./... and go test ./... pass; go.mod is unchanged.
  • internal/outline/outline_test.go covers mapping and caps, empty-vs-unknown counts, plain TOON keys, unexplainedChange, clipping, and the worst-case size budget (18.9 KB of a 20 KB budget).
  • test/context.test.ts has 6 hermetic cases:
    • the full outline and its 11 reads
    • the TOON size
    • a partial failure becoming a warning
    • a missing agent being fatal
    • no project given, with the fix named
    • --dry-run sending nothing
  • Live and read-only on a real project; live end-to-end in a linked folder on a throwaway project (edit, compile, draft test, then vf context reflects it).
  • The full vitest suite passes on master plus the stack, apart from the 4 cases that already fail on master.
  • CI

Fixes COR-14197

…xt (COR-14197)

A coding agent's first job on a Voiceflow project is finding out what it is:
roughly ten calls, plus one per playbook and function, each an agent turn.
The API has no single read an agent can use. The v2 export is not in the
SDK, and at 100KB+ it would flood a context window anyway.

vf context makes eleven typed reads itself, four at a time under a 60s
deadline, and returns an outline:

- the model, global prompt and instructions (line counts and an excerpt)
- playbooks, summarized by the description the agent routes on
- functions, agent tools, variables and the knowledge base
- recent changes, newest first, and the latest conversations
- the working rules agents otherwise learn by breaking them, and the
  drill-down commands for anything the outline leaves out

Lists are capped and text is clipped, with true totals under counts. The
worst case is held under 20 KB of TOON by a test; a large real project comes
to about 9 KB.

The project and agent reads are essential. Any other part that fails becomes
a warning, with that part's count null rather than zero. Two limits are
stated in the output rather than guessed: the API does not say who changed a
resource, and the agent's own instructions carry no timestamp.

vf link's agent snippet now points at vf context. The outline avoids json tag
options, which the TOON encoder prints verbatim.
…'s tools (COR-14197)

Two cold agents answering "what changed most recently?" both made follow-up
calls that the outline should have saved them.

- The capped function list kept the first 20 the API returned, so the most
  recently edited function could fall outside it. Playbooks and functions are
  now sorted newest first before capping.
- Agent tools were counted by type, and an agent had to match their function
  IDs to names by hand. agentTools now lists each tool, named after the
  function it calls, or its description when it calls something else.

To keep the worst case well inside the 20 KB budget with the new list,
summaries are clipped at 120 characters (was 140) and agent tools are capped
at 15. The worst case is now 18.9 KB; a large real project is 9.0 KB.
A cold agent on today's CLI noticed that the project record's updatedAt was
a month later than every resource timestamp it could read, meaning something
had changed that no resource showed. The outline dropped that signal.

project.updatedAt is now in the outline, and the note on untimestamped
instructions says what a later project.updatedAt means.
Go writes a whole-second time as 11:50:00Z and JavaScript as 11:50:00.000Z, so the string comparison added with project.updatedAt failed although the value was right.
… stop searching (COR-14197)

Cold agents asked "what changed most recently?" saw project.updatedAt a month
later than every dated change, read the note that something had changed out
of sight, and went looking: 17 and 28 vf calls, where the same question took
3 before. It cannot be found. The API does not timestamp the instructions,
the global prompt or agent settings, and vf has no history or diff command.

unexplainedChange now states that in the outline, and only when it applies:
the project record moved more than a minute after both the newest dated
change and the last release. The minute absorbs an ordinary edit or publish
touching the project record a moment later.
A cold agent read "Report it as unexplained rather than searching for it" in
vf context's output as an instruction arriving through tool output, and
checked the project record itself before following it. That is the right
instinct: agents should distrust instructions in tool output. So the output
should not give any.

- unexplainedChange and the rules now state facts: what cannot be
  identified, and what compile, draft and publish do.
- The instruction "ask before publishing" stays in the snippet vf link prints
  for the user's own instructions file, which is where instructions belong.
- Another cold agent spent 3 calls ruling out the knowledge base, which also
  carries no edit time. It is now named with the instructions, global prompt
  and agent settings.
Copilot AI balanced review requested due to automatic review settings September 29, 2026 17:59
@linear-code

linear-code Bot commented Sep 29, 2026

Copy link
Copy Markdown

COR-14197

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Partial failures can produce false conclusions, capped ordering is inconsistent, and --usage is not integrated.

Review effort: Balanced
Findings: 4 Medium severity

Open (4)
What changed in this PR

Adds vf context, an agent-focused project summary command with bounded, partial-failure-tolerant output.

Changes:

  • Fetches 11 project resources concurrently and builds a capped outline.
  • Adds unit and integration coverage for output, failures, and size limits.
  • Documents the command and updates linked-project guidance.
File Description
internal/​cli/​context.go Implements the command and concurrent API reads.
internal/​outline/​outline.go Builds the bounded project summary.
internal/​outline/​outline_test.go Tests mapping, clipping, ordering, and size.
internal/​cli/​root.go Registers vf context.
internal/​cli/​link.go Recommends starting with vf context.
test/​context.test.ts Adds end-to-end command tests.
README.md Documents project context retrieval.
docs/​vf.md Adds the command to generated documentation.
docs/​vf_context.md Provides command reference documentation.
Files not reviewed (1)
  • internal/cli/root.go: Generated file

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread internal/cli/context.go
Comment on lines +59 to +60
Args: cobra.NoArgs,
RunE: runContextCmd,
Comment thread internal/outline/outline.go Outdated
Comment thread internal/outline/outline.go
Comment thread internal/outline/outline.go
…read (COR-14197)

Review found that unexplainedChange could state something false. It said a
change "cannot be identified with vf" even when one of the dated reads had
failed, and the part that failed (the environment and its releases,
playbooks, functions, agent tools, variables, MCP servers or tests) may be
exactly what explains the project record's timestamp.

The outline now makes the claim only when every dated part was read. A part
that failed is nil, and one that was read and is empty is an empty slice, so
the check is exact. warnings already names what is missing.
…hem (COR-14197)

Review found two lists that did not keep the newest entries when capped,
although the other capped lists do:

- Variables were sorted by name, so in a project with more than 30 a
  recently edited variable could fall off the list.
- MCP servers were capped in the order the API returned them.

Both are now sorted by updatedAt before the cap, like playbooks, functions and
agent tools. Workflows carry no timestamp, so they keep the agent's routing
order, and the package doc now says so.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants