Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,22 @@ All notable changes to this project are documented here. The format follows

## 0.5.0 - unreleased

### Added
- **`probbit persona`, the individuality layer** (docs/persona.md): a persona file (YAML subset or JSON: traits with priors,
moods with inertia, couplings, per-turn inputs, history features, habits as hard rules in the probbit-ir vocabulary) + an
individual's state + a turn's inputs compile to one probbit-ir program, run in process at fixed work; the stance (levels with
exact odds, habits in force and bound, refusals, a why and a stance line of at most 40 estimated tokens) and the next state are
canonical JSON. Subcommands `init`, `turn`, `replay`, `explain`, `diff`, `lint`, `check`, `compile`, `describe`. A strict YAML
subset reader (probbit-cli/src/yaml.rs) and a strict validator (one `{"error": {"code": "persona", ...}}` object, exit 2). Three
example personas with goldens in examples/persona/; the documents equal an independent reference implementation's.
- **MCP tools `probbit_persona_init` and `probbit_persona_turn`** in `probbit mcp` (stateless; the CLI's documents).
- **Python `probbit.persona_init`, `persona_turn`, `persona_replay`** (python/probbit.py; tests python/test_persona.py).
- **probbit-wasm ops 4 (persona init) and 5 (persona turn)**; the playground's "Meet three individuals from one persona".
- `probbit persona ... --timing` reports the 1-minute load average next to its timings (`sys::loadavg`).

### Changed
- `probbit mcp` lists seven tools (the two persona tools after the five before; their order and contracts are unchanged).

- **Renamed: pbit → probbit everywhere** (crates, binary, env vars, MCP tools, Python module, npm package). 'p-bit' remains the
term for the probabilistic bit. No behaviour changes.
- Crates `probbit-core`, `probbit-ir`, `probbit-decide`, `probbit-cli` (binary `probbit`), `probbit-wasm`; Rust paths
Expand Down
53 changes: 51 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -330,6 +330,53 @@ Build the binary once with `cargo build --release -p probbit-cli` (it lands in `
bindings and no MCP server yet; every script in `bench/` is an example of the subprocess pattern. `probbit <command> --help`
lists every flag of a command with its default, and the exit codes.

## The individuality layer: `probbit persona`

Give an agent a temperament that lives outside the model. A **persona** is a small file (YAML subset or JSON): traits with
priors, moods with inertia, soft couplings, per-turn evidence, and habits that are hard rules. A seed makes an individual. Every
turn compiles persona + the individual's state + the turn's inputs into ONE probbit-ir program, answered exactly in process; the
**stance** (a level per trait with exact odds, the habits in force and the ones that changed it, a refusal when the engine cannot
vouch) comes with a short **stance line** a host puts into any model's prompt. The model still writes every word; the persona
decides how they are written, and the same individual answers whatever model the host calls. [docs/persona.md](docs/persona.md)
is the format; [examples/persona/](examples/persona/) has three fictional personas and their goldens.

```sh
probbit persona init examples/persona/tutor.yaml --seed 2 --out pip.json # an individual: genes from the seed
probbit persona turn examples/persona/tutor.yaml --state pip.json --inputs '{"loss": true, "sentiment": "negative"}'
# "status":"ok", "humour":{"level":"none","p":1,...}, "warmth":{"level":"warm","p":0.952032,...}, habits active: no_jokes_on_loss, ...
# "line":"Stance: no jokes, be kind; one emoji at most; gentle; explain step by step; warm and encouraging; suggest what to try next; ask what they think first; casual."
# "why":"learner reports a failure -> humour none, valence down; learner upset -> valence down, humour none" (pip.json is now turn 1)
probbit persona replay examples/persona/tutor.yaml --seed 2 --script examples/persona/workday.json # 20 stances, the same bytes every run
probbit persona lint examples/persona/tutor.yaml # contradicting habits: confused_no_playful vs after_error_check, resolved by priority
```

```python
import probbit # python/probbit.py, stdlib only
state = probbit.persona_init("examples/persona/tutor.yaml", seed=2)
turn = probbit.persona_turn("examples/persona/tutor.yaml", state, {"loss": True})
prompt_tail, state = turn["stance"]["line"], turn["state"] # the line goes after any cached prompt prefix
```

MCP: `probbit_persona_init` and `probbit_persona_turn` in `probbit mcp` (stateless: the persona, inline or a path, and the state go
in; the state comes back). Browser: the playground's "Meet three individuals from one persona" (probbit-wasm ops 4 and 5, no
threads). Every surface gives the CLI's documents byte for byte (tests/persona.rs, python/test_mcp.py, probbit-wasm/tests).

Measured in this run (Apple M4, macOS 26.5.2, 2026-10-02; the three example personas):
- a whole turn in process (compile, engine, decode, digests): median 0.20-0.32 ms, p95 0.22-0.67 ms, of which the engine 0.09-0.18
ms (median); N = 1,000 turns per persona, 1-minute load 3.3-3.5, `nice 10`;
- the stance line: 36-38 estimated tokens on average over the 20 workday turns, at most 40;
- individuals: three seeds of the tutor differ by a mean total-variation distance of 0.12-0.18 between their trait odds on the
workday script; the tutor and the ops engineer (two persona files) by 0.45; the same persona and seed by 0;
- habits: 0 violations over 2,160 turns of 12 random personas (every habit in force re-checked from the file), and each of the 481
habits reported as binding is broken by the habit-free twin;
- determinism: 1,000-turn replays are byte-identical across processes (3 personas x 2 seeds); the documents equal an independent
reference implementation's on 5,460 turns (the 60 golden turns and 5,400 turns of random inputs).

What a persona can NOT do: write, read or check text (text rules stay in the prompt or a checker); see what the host does not tell
it; be a safety gate (keep approvals and permissions in plain code); its odds are its own model's, not measured probabilities that
a user will like the reply; and it cannot make a model follow the line: whether a given model writes in the stance it is given has
to be measured per model (no such measurement has been made here).

## After any judge: `probbit evaluate`

Decision models (TypeSafe's Jev, Cloudflare's Clef and Clef-flash, local System One servers) are judges: content in, a
Expand Down Expand Up @@ -520,8 +567,10 @@ or Windows.
keeps every object's fields equal to the parser's).
- [docs/agents.md](docs/agents.md): calling probbit from a shell, Python, Node, PowerShell, MCP agents and a browser, and
`probbit evaluate` after a decision model.
- [python/](python/): `probbit.py`, a zero-dependency subprocess wrapper (with `evaluate` and a stdlib mock judge), its tests and
three examples.
- [docs/persona.md](docs/persona.md): `probbit persona`, the individuality layer: the persona file, the compilation, the stance and
state documents, the canonical JSON, what a persona can not do; [examples/persona/](examples/persona/): three personas + goldens.
- [python/](python/): `probbit.py`, a zero-dependency subprocess wrapper (with `evaluate`, the persona functions and a stdlib mock
judge), its tests and three examples.
- [playground/](playground/): one static page that runs probbit in a browser (`probbit-wasm`, built by `playground/build.sh`).
- [bench/](bench/): the scripts behind BENCHMARKS §2 and §5 (Python 3; the ILP baselines need `numpy` and `scipy >= 1.9`).
- [CONTRIBUTING.md](CONTRIBUTING.md), [CHANGELOG.md](CHANGELOG.md).
Expand Down
8 changes: 7 additions & 1 deletion docs/agents.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,7 +80,7 @@ mark with `[Console]::InputEncoding = [System.Text.UTF8Encoding]::new($false)` a
## MCP: one line per agent

`probbit mcp` is a [Model Context Protocol](https://modelcontextprotocol.io) server on stdio (JSON-RPC 2.0, one message per
line; stdout carries only protocol messages, logs go to stderr; it exits when stdin closes). Its five tools take the
line; stdout carries only protocol messages, logs go to stderr; it exits when stdin closes). Its seven tools take the
commands' own documents and return the commands' own JSON, byte for byte:

| tool | arguments | returns |
Expand All @@ -90,6 +90,12 @@ commands' own documents and return the commands' own JSON, byte for byte:
| `probbit_stats` | optional `sweeps`, `chains`, `threads` | `probbit stats` |
| `probbit_demo` | optional `tasks`, `seed`, `hard` | `probbit demo` (a router document) |
| `probbit_evaluate` | a System One request plus the optional `probbit` block ([probbit-ir-json.md](probbit-ir-json.md#decision-api-probbit-evaluate)) plus optional `flags` | `probbit evaluate` |
| `probbit_persona_init` | `persona` (the document) or `persona_path`, optional `seed` | `probbit persona init` (the state) |
| `probbit_persona_turn` | `persona` or `persona_path`, `state`, `inputs`, optional `flags` (`timing`, `no_inertia`) | `{stance, state}`: `probbit persona turn`'s stance and the state it writes ([persona.md](persona.md)) |

The two persona tools run in the server's process and keep nothing between calls: pass the returned state back on the next turn and
put `stance.line` into the model's prompt (after any cached prefix). A refused or fallback stance is an answer; a bad persona,
state or input is a tool error with the `{"error": {"code": "persona", ...}}` object.

`flags` are the command's flags without the dashes: `{"budget_ms": 200, "seed": 3, "summary": true}`; `summary: true`
returns the compact answer (README, "First five minutes"). `infeasible` and `refused` / `declined` are answers; bad input
Expand Down
Loading
Loading