Skip to content

probbit 0.6.0: persona fuzz + prove - #11

Merged
BitmapAsset merged 17 commits into
mainfrom
r38-fuzz-prove
Oct 6, 2026
Merged

BitmapAsset merged 17 commits into
mainfrom
r38-fuzz-prove

Conversation

@BitmapAsset

@BitmapAsset BitmapAsset commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Character testing for personas, and probbit 0.6.0. Draft until CI is green on the head.

What is in it

  • probbit persona fuzz PERSONA --seeds 0-99 --never '{when: {...}, then: {...}}' (or --props FILE): random event scripts over the persona's declared inputs plus an odds-guided beam; each individual's shortest counterexample shrunk by delta debugging and printed with the breaking turn's levels, odds, why, line and the replay / explain commands that reproduce it. Deterministic (Philox keyed by --fuzz-seed; byte-identical on any thread count). Exit 0 nothing found, 1 counterexample, 2 bad input; --json gives probbit_persona_fuzz: 1.
  • probbit persona prove: per rule held by construction, proved for every event sequence (a sound per-cell bound: mood accumulators in a box, finitely many cells of inputs, history and previous levels) or unknown (the earliest undecided cell and the fuzz command to search it). Exit 0 all held or proved, 1 some unknown; --json gives probbit_persona_prove: 1.
  • Soundness tests: every event sequence of tiny random personas agrees with prove; prove never proves what the fuzzer breaks on random personas; habits are held by construction; mutations of the bound are caught.
  • persona lint --props FILE / --never RULE: prove every rule, fuzz the unknown ones, exit 1 when one breaks. lint also lists warnings where a trait's planned level differs from its own most likely level (decoding unchanged).
  • MCP tool probbit_persona_fuzz; Python probbit.persona_fuzz and probbit.persona_prove.
  • The tutor example gets one habit, no_play_when_upset: on the 0.5.0 tutor (kept byte for byte as a test fixture) the fuzzer finds 67 of 100 individuals that break "never playful when upset", 37 of them on the single message {"sentiment":"negative"}; on the fixed tutor prove says held by construction and the fuzzer finds 0 of 100. Goldens regenerated; the playground's embedded copy updated.
  • why names a push by its direction (learner upset -> valence down, humour down) instead of a level name.
  • docs/persona.md §5.6 "Testing a character", §5.5 plan and mode, §5.3 (reference parity stated for 0.5.0); README "Test a character"; docs/agents.md; version 0.6.0; CHANGELOG 0.6.0 - unreleased.

Limits (also in the docs)

Both commands test the stance a host gets, not the words a model writes. "None found" is evidence, not a proof. Numbers are gridded. Rules are habit-shaped. A soft rule with a small margin stays unknown even when no script breaks it.

Test plan

  • cargo test --release --workspace (179 passed locally)
  • python3 python/test_persona.py, python3 python/test_mcp.py
  • sh playground/build.sh && node playground/check.mjs playground/probbit.wasm

probbit persona fuzz PERSONA --seeds 0-99 --never RULE (or --props FILE)
takes a rule in habit syntax and searches, per individual, event
scripts built from the persona's declared inputs for the shortest one
whose stance breaks it: random scripts, then a beam guided by the exact
odds (the next events that put the most odds on a level the rule
forbids), then delta debugging. The report gives the breaking turn's
levels, odds, why and line and the replay / explain commands that
reproduce it; --json gives the probbit_persona_fuzz document.

- Philox (probbit-core) keyed by --fuzz-seed and the individual's seed;
  individuals are searched independently, so the output is
  byte-identical on any number of threads (timing goes to stderr)
- exit 0 nothing found, 1 a counterexample, 2 bad input; a bad rule is
  one error object at its path, as a bad habit is
- habits and rules share one reader (habit_spec) and one allowed-levels
  function; no change to any turn's document
- tests: the 0.5.0 tutor kept byte for byte as a fixture with its fuzz
  report pinned (seed 1, one event {"sentiment":"negative"}, humour
  playful); determinism across threads; habits as rules never break
  (the tutor's, and every habit of 8 random personas); bad rules/flags
Version 0.6.0 in Cargo.toml, the crate manifests, npm/package.json, the
installers' examples and the bench workflow's tag. The golden test is
stdout_matches_the_0_6_0_goldens; its documents are unchanged. CHANGELOG:
0.6.0 - unreleased with persona fuzz.
… or unknown

`probbit persona prove PERSONA (--never RULE | --props FILE) [--seeds] [--threads] [--json]`
answers each rule over the population with one verdict:

- held by construction: the habits in force whenever the rule applies imply it,
  for every previous level. Under on_conflict: yield a habit counts when no turn
  where the rule applies ever drops it: whether a turn's rules conflict, and
  which habit yields, is a function of the habits in force and the previous
  levels, so one scan of every cell decides it.
- proved for every event sequence: a bound per individual. The mood accumulators
  stay in a box; inputs (numbers at their thresholds and the intervals between),
  history values and previous levels fall into finitely many cells; in every
  cell the best allowed stance must beat the best forbidden one (clamped exact
  solves) by more than the box, the number intervals and the rounding can move
  the two scores. Moods outside the rule's component do not count; held traits
  and conflicts are replayed as the turn resolves them; the engine runs the
  turn's own program (with its habit-free twin), and every answer it reads must
  be exact.
- unknown: the earliest cell the bound cannot decide, and the fuzz command that
  searches it.

Human report or --json (probbit_persona_prove: 1). Exit 0 every rule held or
proved, 1 some unknown, 2 bad input. Test: the three verdicts on the 0.5.0 tutor.
prove must never say held or proved where a stance breaks the rule:
- every event sequence of 16 tiny random personas (2 individuals each): up to 3
  events over all 12 input combinations and up to 5 over 4 of them;
- the fuzzer on 16 random personas with numbers, a streak and habits (3
  individuals each): wherever it finds a counterexample, prove did not say held
  or proved;
- a habit's own rule is held by construction (every habit under fallback, the
  top-ranked one under yield).
Each test also requires both sides to occur (rules proved and rules broken), and
dropping the mood box from the bound, or the yield scan from held by
construction, makes them fail.
What a character property is and where it is checked (the stance, never the
words); fuzz's search and its determinism; prove's three verdicts and the bound
behind "proved for every event sequence"; the limits (exact answers, habits as
rules kept by the engine, raw rules not proved, unknown is not broken); exit
codes and JSON formats. Command table rows and exit codes in section 7.
`probbit.persona_fuzz(persona, never=None, props=None, seeds="0-99", **flags)`
and `probbit.persona_prove(...)` run `probbit persona fuzz|prove --json` and
return its document. A rule is a dict or one line of YAML; several are a list or
a props file; seeds a range string or a list; list flags (grid, hours) become
comma lists. A counterexample (fuzz) or an unknown rule (prove) is an answer,
not an error; a bad rule raises ProbbitInputError. Standard library, Python 3.9.
Test: both on the 0.5.0 tutor fixture. docs/persona.md section 7 lists them.
The MCP server gets an eighth tool: a persona (inline document or a path) and a
character property (never: one rule in habit syntax, as an object or one line
of YAML; or props: a list of rules), optional seeds ("0-99", "1,4,9" or a list),
fuzz_seed, scripts, depth, beam, grid, hours, threads -> the document of
`probbit persona fuzz --json` for the same arguments, byte for byte. It runs in
the server's process and keeps nothing; a counterexample is an answer, a bad
persona or rule a tool error with the persona error object. The CLI and the
tool read seeds with one parser. Tests: the MCP answer equals the CLI's on the
0.5.0 tutor fixture, an inline persona, four argument errors; docs/agents.md
and docs/persona.md section 7 list the tool.
`probbit persona lint PERSONA --props FILE` (or `--never RULE`, with
`--seeds` and `--threads`) reports the contradicting habits as before, then runs
prove on every rule and the default fuzz search on the rules the bound leaves
unknown. The lint document gets `props` (each rule's prove entry, plus a `fuzz`
entry when it was unknown) and `broken`; `ok` is false and the exit code 1 when
a contradiction is unresolved or a rule breaks, so CI can gate on it. Without
rules lint is unchanged. fuzz, prove and lint read rules with one function.
Test: on the 0.5.0 tutor the loss rule is held and "never playful when upset"
breaks (exit 1); a proved rule alone passes (exit 0). Help, docs 5.6 and 7,
CHANGELOG.
…_when_upset

`why` said "humour none" next to a playful stance: it named an input's push by the
trait's lowest or highest level. It now names the direction ("learner upset ->
valence down, humour down").

The tutor (persona version 1.1.0) has one habit more, no_play_when_upset
(when: {sentiment: negative}, then: {humour: {at_most: light}}). In 0.5.0 one upset
message made 67 of seeds 0-99 playful, and the workday golden itself had seed 1
playful on turn 11; that turn is now light, `persona prove` says held by
construction and `persona fuzz` finds nothing. Genes are unchanged.

Goldens regenerated with this binary (YAML and JSON forms agree); in the ops
engineer's and the trader assistant's goldens nothing but `why` changes. The 0.5.0 tutor fixture
is untouched; its pinned fuzz report changes in the why line alone. docs 5.3 now
states the reference-implementation parity for the 0.5.0 documents.
…y level

The stance is the joint plan, so on coupled traits a trait's planned level can
differ from its marginal mode (the 0.5.0 tutor's seed 1, upset: humour playful at
0.411 while light had 0.425). `lint` now probes each individual of --seeds (at
rest, a quiet turn, every declared input alone) and lists, per trait, how many
probe turns plan and mode differ on and the earliest one, under `warnings`. A warning
leaves the exit code alone; decoding is unchanged. docs 5.5 gives the options.

Test: the fixed tutor is held by construction and fuzzes clean; on the 0.5.0
tutor `why` reads "humour down" next to the playful stance and lint warns on
seed 1's humour.
README "Test a character": the fuzz report on the 0.5.0 tutor fixture (67 of 100
individuals, seed 1's one-event counterexample) and `held by construction` on the
tutor as it ships now, with the commands that print them. The persona example
shows the new line and why; the reference-parity line is stated for 0.5.0.
CHANGELOG 0.6.0: lint warnings, the tutor fix, the why wording, docs 5.3.
docs/agents.md: persona rules in CI.
The page carries the three example personas inline; check.mjs (the wasm CI job)
requires each to equal examples/persona/<name>.json. Locally: build.sh, then
check.mjs exit 0, all three personas 20/20 workday stances equal to the goldens.
@BitmapAsset
BitmapAsset marked this pull request as ready for review October 6, 2026 03:55
@BitmapAsset
BitmapAsset merged commit 35d4870 into main Oct 6, 2026
8 checks passed
@BitmapAsset
BitmapAsset deleted the r38-fuzz-prove branch October 6, 2026 03:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant