feat: summarize a project for coding agents in one call with vf context (COR-14197) - #36
Open
Bradenream wants to merge 8 commits into
Open
Bradenream wants to merge 8 commits into
Bradenream wants to merge 8 commits into
Conversation
…xt (COR-14197) A coding agent's first job on a Voiceflow project is finding out what it is: roughly ten calls, plus one per playbook and function, each an agent turn. The API has no single read an agent can use. The v2 export is not in the SDK, and at 100KB+ it would flood a context window anyway. vf context makes eleven typed reads itself, four at a time under a 60s deadline, and returns an outline: - the model, global prompt and instructions (line counts and an excerpt) - playbooks, summarized by the description the agent routes on - functions, agent tools, variables and the knowledge base - recent changes, newest first, and the latest conversations - the working rules agents otherwise learn by breaking them, and the drill-down commands for anything the outline leaves out Lists are capped and text is clipped, with true totals under counts. The worst case is held under 20 KB of TOON by a test; a large real project comes to about 9 KB. The project and agent reads are essential. Any other part that fails becomes a warning, with that part's count null rather than zero. Two limits are stated in the output rather than guessed: the API does not say who changed a resource, and the agent's own instructions carry no timestamp. vf link's agent snippet now points at vf context. The outline avoids json tag options, which the TOON encoder prints verbatim.
…'s tools (COR-14197) Two cold agents answering "what changed most recently?" both made follow-up calls that the outline should have saved them. - The capped function list kept the first 20 the API returned, so the most recently edited function could fall outside it. Playbooks and functions are now sorted newest first before capping. - Agent tools were counted by type, and an agent had to match their function IDs to names by hand. agentTools now lists each tool, named after the function it calls, or its description when it calls something else. To keep the worst case well inside the 20 KB budget with the new list, summaries are clipped at 120 characters (was 140) and agent tools are capped at 15. The worst case is now 18.9 KB; a large real project is 9.0 KB.
A cold agent on today's CLI noticed that the project record's updatedAt was a month later than every resource timestamp it could read, meaning something had changed that no resource showed. The outline dropped that signal. project.updatedAt is now in the outline, and the note on untimestamped instructions says what a later project.updatedAt means.
Go writes a whole-second time as 11:50:00Z and JavaScript as 11:50:00.000Z, so the string comparison added with project.updatedAt failed although the value was right.
… stop searching (COR-14197) Cold agents asked "what changed most recently?" saw project.updatedAt a month later than every dated change, read the note that something had changed out of sight, and went looking: 17 and 28 vf calls, where the same question took 3 before. It cannot be found. The API does not timestamp the instructions, the global prompt or agent settings, and vf has no history or diff command. unexplainedChange now states that in the outline, and only when it applies: the project record moved more than a minute after both the newest dated change and the last release. The minute absorbs an ordinary edit or publish touching the project record a moment later.
A cold agent read "Report it as unexplained rather than searching for it" in vf context's output as an instruction arriving through tool output, and checked the project record itself before following it. That is the right instinct: agents should distrust instructions in tool output. So the output should not give any. - unexplainedChange and the rules now state facts: what cannot be identified, and what compile, draft and publish do. - The instruction "ask before publishing" stays in the snippet vf link prints for the user's own instructions file, which is where instructions belong. - Another cold agent spent 3 calls ruling out the knowledge base, which also carries no edit time. It is now named with the instructions, global prompt and agent settings.
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Partial failures can produce false conclusions, capped ordering is inconsistent, and --usage is not integrated.
Review effort: Balanced
Findings: 4
Open (4)
What changed in this PR
Adds vf context, an agent-focused project summary command with bounded, partial-failure-tolerant output.
Changes:
- Fetches 11 project resources concurrently and builds a capped outline.
- Adds unit and integration coverage for output, failures, and size limits.
- Documents the command and updates linked-project guidance.
| File | Description |
|---|---|
internal/cli/context.go |
Implements the command and concurrent API reads. |
internal/outline/outline.go |
Builds the bounded project summary. |
internal/outline/outline_test.go |
Tests mapping, clipping, ordering, and size. |
internal/cli/root.go |
Registers vf context. |
internal/cli/link.go |
Recommends starting with vf context. |
test/context.test.ts |
Adds end-to-end command tests. |
README.md |
Documents project context retrieval. |
docs/vf.md |
Adds the command to generated documentation. |
docs/vf_context.md |
Provides command reference documentation. |
Files not reviewed (1)
- internal/cli/root.go: Generated file
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
+59
to
+60
| Args: cobra.NoArgs, | ||
| RunE: runContextCmd, |
…read (COR-14197) Review found that unexplainedChange could state something false. It said a change "cannot be identified with vf" even when one of the dated reads had failed, and the part that failed (the environment and its releases, playbooks, functions, agent tools, variables, MCP servers or tests) may be exactly what explains the project record's timestamp. The outline now makes the claim only when every dated part was read. A part that failed is nil, and one that was read and is empty is an empty slice, so the check is exact. warnings already names what is missing.
…hem (COR-14197) Review found two lists that did not keep the newest entries when capped, although the other capped lists do: - Variables were sorted by name, so in a project with more than 30 a recently edited variable could fall off the list. - MCP servers were capped in the order the API returned them. Both are now sorted by updatedAt before the cap, like playbooks, functions and agent tools. Workflows carry no timestamp, so they keep the agent's routing order, and the package doc now says so.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Summary
A coding agent's first job on a Voiceflow project is finding out what it is. With today's CLI that means one call per resource type, each an agent turn. The API has no single read that fits in a context window: the v2 export is not in the SDK, and at 100 KB+ it would flood one anyway.
vf contextmakes 11 typed reads itself, 4 at a time under a 60 s deadline, and returns an outline:drillDown: the commands for anything the outline leaves outLists are capped newest first, text is clipped, and the true totals are under
counts. A test holds the worst case under 20 KB of TOON; a large real project comes to about 9 KB.If a part can't be read, the outline still prints: that part's count is
nullandwarningsnames it. Two limits are stated rather than guessed:unexplainedChangesays the change can't be identified with vf.vf link's agent snippet now starts withvf context.Measured with cold coding agents
Same question to every run, read-only harness, a real project, 3 runs per variant. Medians; calls and bytes come from the harness log.
The runs also shaped the code:
unexplainedChangenow states when a change can't be identified.Test plan
gofmt,go vet ./...andgo test ./...pass;go.modis unchanged.internal/outline/outline_test.gocovers mapping and caps, empty-vs-unknown counts, plain TOON keys,unexplainedChange, clipping, and the worst-case size budget (18.9 KB of a 20 KB budget).test/context.test.tshas 6 hermetic cases:--dry-runsending nothingvf contextreflects it).Fixes COR-14197