Skip to content

feat(eval): one tagged 88-task dataset with runtime tasks, plus bashkit_generate eval - #2664

Merged
chaliy merged 14 commits into
mainfrom
claude/project-thread-hndwnc
Oct 10, 2026
Merged

chaliy merged 14 commits into
mainfrom
claude/project-thread-hndwnc

Conversation

@chaliy

@chaliy chaliy commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

What changed

  • One tagged dataset. The basic, repo and hard tasks are merged into bashkit_bash, 88 tasks in total. Each task is tagged mode=agent|runtime and difficulty=basic|repo|hard, and carries its own max_turns. Subsets run with --tag (just eval-repo, just eval-hard, just eval-runtime).

  • 12 runtime tasks (rt_*). In these, bashkit is the agent's execution runtime:

    • python3 file-processing scripts
    • python3 + bash pipelines
    • sqlite3 reports and migrations
    • state kept across tool calls

    The eval Bash now has CPython and sqlite. A new hidden verify script re-runs what the model built against unseen input.

  • New bashkit_generate eval. It has 15 one-shot tasks, run with just eval-generate. The model replies with a single script, which bashkit runs once and the shared expectations scorer scores.

  • Bashkit fixes found by the reference solutions:

    • A script whose shebang names a builtin (#!/usr/bin/env python3) now runs that builtin.
    • The sqlite builtin writes rollback-journal databases, so CPython's sqlite can open them.
  • Results. The README, the eval README, the site and the knowledge log are updated. The homepage table now lists only runs on the newest run's task count, so 58-task and 88-task scores never mix.

Why

The old dataset only measured bash as an agent tool. Agents also use bashkit as a runtime (python3, sqlite, persistent VFS) and to execute generated scripts. Keeping one tagged dataset avoids splitting it into four tracks; generation gets its own eval because it is scored differently.

Before / After

Before, the 58-task set was saturated, with three models at 58/58.

After, on 2026-10-09:

Model bashkit_bash (88) basic repo hard runtime generate (15)
Claude Opus 5.5 87 58 7 10 12 13
Kimi K3 86 57 8 9 12 15
Claude Sonnet 5.5 85 56 8 9 12 15
Muse Spark 1.3 84 58 5 9 12 15
Gemini 3.8 Flash 74 55 4 6 9 13

GPT-6.1 Sol and GPT-6 Luna are not in this run because the OpenAI key ran out of credit. Their cases can be filled in later with mira run ... --resume <run_id>.

Risk

  • Low. The eval harness changes are internal.
  • The shebang fix touches script dispatch, and the sqlite fix changes the journal mode of new database files. Both are covered by tests.

Checklist

  • Tests added or updated (reference replay, tag coverage, shebang and sqlite tests)
  • Backward compatibility considered (internal; eval names bashkit_repo/bashkit_hard are replaced by tags)

Generated by Claude Code

chaliy added 8 commits October 9, 2026 22:20
One bashkit_bash eval; tasks carry mode (agent|runtime), difficulty
(basic|repo|hard) and an optional per-task max_turns. Reference solutions
live in data/solutions.jsonl and are required for every non-basic task.
The agent-loop eval hides how often a model's first script fails on
bashkit. bashkit_generate gives the model the task plus a fixed sandbox
description, takes ONE bash script from a single reply (no tools, no
feedback), runs it once on build_task_bash and scores it with the shared
expectations scorer.

- generate.rs: subject, documented extraction rule (bash/sh/shell fence,
  else untagged fence, else whole reply), 60s run limit, metrics
  (script_found, extracted, script_bytes, exit_code, timing); no script
  scores 0, provider errors are infra errors
- providers omit tools when none are offered
- 15 tasks (basic/hard) in data/generate-tasks.jsonl with single-script
  references; tests require references to pass and `true` to fail
- just eval-generate, knowledge/operations/eval.md Generate Eval section
Turso writes WAL-mode headers (bytes 18/19 = 2). The Memory backend
persists only the checkpointed main file, so the image is a valid
legacy-mode database; mark it as one. CPython's sqlite3 module in the
WASI guest has no WAL and rejected the file as "file is not a
database", so python3 could not open a db written by sqlite3.

Also record known sqlite3 shell divergences (csv quoting, exit codes,
.import) found by the eval runtime tasks.
…ltin

An executable file with #!/usr/bin/env python3 (or #!/usr/bin/python3,
#!/usr/bin/awk -f) run by path or PATH lookup had its shebang stripped
and its body parsed as bash, so agent-built Python CLIs failed with a
parse error. A shebang naming a registered non-shell builtin now runs
that builtin with the script path as its argument (Linux semantics:
one optional argument; env -S splits). bash/sh, unknown interpreters
and shell-only builtins keep running the content as bash.

Adds CPython integration tests for shebang scripts and for python3
reading/writing a database made by the sqlite builtin.
Every eval Bash now runs real CPython 3.14 (cpython feature) and the
sqlite builtin (sqlite feature, opt-in env set); the tool prompt the
model sees registers the same builtins, so it lists their hints instead
of "python/python3 not available". bashkit-replay uses the same
runtimes.

12 runtime tasks (tag runtime, just eval-runtime): python3 file
processing (CSV rollup, JSONL flatten, sessionization, encodings, TOML
render, a 12k-line log), a bash+python report script, sqlite3 (CSV load
+ report, a reusable migration file, python reading a sqlite3-built db)
and state across calls (a CLI on PATH, an idempotent ingester).

Tasks gain a hidden verify script, run after the agent in a fresh Bash
on the same filesystem; its output (/.eval/verify.out) is scored, so a
built tool is re-run on unseen input. Goldens come from the generator,
references pass in bashkit and against real python 3.13, sqlite3 3.45
and bash 5.2. The untouched-fixture test covers runtime tasks.
@chaliy chaliy self-assigned this Oct 9, 2026
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
bashkit 40956d8 Commit Preview URL

Branch Preview URL
Oct 09 2026, 11:42 PM

chaliy added 6 commits October 9, 2026 22:42
The python3 setup took ~7s in debug and timed out (30s) under coverage
instrumentation. The awk generator writes the byte-identical log
(same md5 in bashkit and real bash) in under 1s.

Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
Instrumented CPython runs ~10x slower; replaying rt_py_large_log's
python3 reference exceeded the 30s shell and python3 limits under
tarpaulin. cfg(tarpaulin) builds get 300s; eval runs keep defaults.

Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
bashkit-eval now enables bashkit's cpython and sqlite features, so
workspace feature unification adds the CPython/sqlite tests to the
bashkit integration binary under coverage, which then ran past 300s.

Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
bashkit-eval enables bashkit's cpython feature; under --workspace
feature unification that pulled CPython tests into the instrumented
bashkit integration binary, which then timed out. Excluding it keeps
coverage on main's feature set; eval tests still run in CI's test job.

Claude-Session: https://claude.ai/code/session_015PWfrXwJG9NsMJqio8M1gW
@chaliy
chaliy merged commit 1518430 into main Oct 10, 2026
49 checks passed
@chaliy
chaliy deleted the claude/project-thread-hndwnc branch October 10, 2026 00:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant