Skip to content

Latest commit

 

History

13 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

probbit

probbit

The p-bit processor your agents decide on.
A p-bit is a bit that holds a probability. probbit couples thousands of them on the CPU you already own,
settles a whole plan under hard rules, prints the odds for every part, and refuses when it can't vouch for the answer.

ci license

60 PII leaks → 0 180,540 plans, counted exactly refused → partial → passed
300 agent tasks, six workers, privacy rules and quotas. A rule-ignoring router sends 60 private tasks to cloud models (133 broken rules on probbit demo's queue, 144 on the agent_router example's). Every plan probbit returns breaks 0: the sampler never leaves the rule-abiding set. 12 tasks × 6 workers: every rule-abiding plan enumerated, exact odds for every task, in 7.3-8.2 ms on an Apple M4 (measured). Too little time and it says "I can't vouch for this" (exit 3) instead of guessing. More time: it releases what passed its own error bar and escalates the rest (watch it refuse).
curl -fsSL https://raw.githubusercontent.com/BitmapAsset/probbit/main/install.sh | sh

Linux and macOS, no sudo, the archive's SHA-256 checked first. Windows, npm and the from-source build: Install. One binary under 2 MB, zero third-party dependencies, nothing phones home. Also inside: the individuality layer, probbit persona, a temperament that lives outside the model: swap the model, keep the individual.

Try it in 30 seconds

probbit                                                         # at a terminal: the spec line of this machine and three commands to try
probbit demo                                                    # at a terminal: 300 agent tasks routed live (~4 s): a field of p-bits, the gate, the verdict
probbit demo --tasks 300 | probbit decide --summary --pretty    # the same queue as one JSON answer: verdict, 0 violations, the least certain tasks
probbit demo --tasks 300 --hard | probbit decide --budget-ms 50 # too little time: "refused", exit 3, nothing released
probbit demo | probbit decide --top > answer.json               # 12 tasks, the full answer, with a live monitor while it runs (tier, updates/s, gate)
probbit evaluate --summary --pretty < examples/evaluate/support-12.json   # in a clone: after a decision model, its answers + your rules
probbit run --sweeps 2000 --polish-ms 0 --seed 7 --summary < examples/rabbit-field.json   # in a clone: the white rabbit, 3,840 p-bits: partial, 31 cells escalated
claude mcp add probbit -- probbit mcp                           # hand the processor to an agent: probbit_decide, probbit_run, probbit_evaluate, ...

Visuals go to stderr, only at a terminal; NO_COLOR, --plain or PROBBIT_THEME=plain turn them off. Piped, every command writes exactly one JSON document, the same bytes as 0.2.1 apart from the name, the version and the timings. Piped, probbit demo writes a 12-task queue unless --tasks says otherwise.

probbit is a software processor built from p-bits: bits that are 1 with a probability you set, coupled so that they sample whole configurations together. You give it a program (variables, their allowed values, scores, pairwise couplings, hard caps) and it returns three things:

  1. an answer that obeys every hard rule by construction (the sampler never leaves the feasible set; if it cannot find a feasible start within the budget it returns refused and no plan);
  2. odds for every variable, exact when the program is small or thin enough, otherwise sampled with an error bound;
  3. a verdict: exact (proof-grade under the score model), diagnostics_passed (sampled: every released item passed the versioned diagnostics gate; not a proof), partial (act on the released variables, escalate the rest), refused or infeasible (a proof).

The first front-end is a task router. Give it a queue of things to assign (tasks to workers, tickets to agents, jobs to machines), the score of every option, the rules that must hold (allowed sets, capacities, forced choices) and how much items that belong together should stay together. It samples whole plans on a p-bit substrate, so every plan it returns obeys every rule, and it returns odds per item (exact, or Monte-Carlo estimates of the marginals of the score model you supplied: accuracy under that model, not predictive calibration) with a diagnostics gate that releases the items whose diagnostics pass and escalates the rest. The gate can be fooled: when every chain misses the same mode it releases wrong odds with a tiny error bar (see Known failure modes).

A p-bit is a bit that is 0 or 1 with a tunable probability; coupled p-bits sample joint configurations. probbit is quantum-inspired and entirely classical (see "Is this quantum?" below): it runs on the CPU you already have, needs no network and has zero external crates. The same problem lowers to the QUBO/Ising form a p-bit or annealer chip takes (probbit-ir), which is the hardware hand-off for later.

$ cargo run --release -q -p probbit-cli -- demo --tasks 300 | cargo run --release -q -p probbit-cli -- decide --budget-ms 200
{"engine":"probbit 0.2.0","tasks":300,"workers":6,"affinity":1.2,"verdict":"diagnostics_passed","plan":{"T000":"human",…},
 "violations":0,"odds":{"T000":{"human":0.78,"local-gemma":0.22},…},"released":[…300…],"escalated":[],
 "gate":{"rhat":1.00015,"tv_bound":0.023,"tv_tol":0.05,"frozen_saturated_workers":0,…},"ms":291.6}

diagnostics_passed has a definition, and it is not a proof. (It replaces the 0.1.0 verdict certified.) The answer passed gate/3: a 3-sigma batch-means Monte-Carlo bound (at most 0.05 total variation per released item), split R-hat of the log-weight trace and of every released variable's value indicators, batch-size stability and frozen-resource checks, tuned against exact oracles on small synthetic families. Every sampled answer reports gate.version, gate.assumptions, both MCSE estimates, the worst item's per-chain means and a release_reason per item. These checks cannot see a mode that no chain visits: on the frozen stress corpus (probbit-cli/tests/stress/) 2-3 of 20 seeds per family pass with every released item wrong (BENCHMARKS "Known failure modes"). The tables are in BENCHMARKS.md.

Install

Today, from source with Rust 1.78 or later: cargo install --git https://github.com/BitmapAsset/probbit probbit-cli installs the probbit binary; or clone and cargo build --release -p probbit-cli (it lands in target/release/probbit). Nothing is downloaded after the clone: there are no external crates.

From a release (the installers fetch the latest archive from the Releases page; the commands below were tested against a local server, PORTABILITY.md):

how command
Linux, macOS (no sudo) curl -fsSL https://raw.githubusercontent.com/BitmapAsset/probbit/main/install.sh | sh
Windows PowerShell irm https://raw.githubusercontent.com/BitmapAsset/probbit/main/install.ps1 | iex
npm npm i -g probbit
crates.io cargo install probbit-cli
by hand archives for Linux x86_64 (glibc or static musl), Linux arm64 (static), macOS (Apple silicon and Intel) and Windows x86_64 on the Releases page, each with a .sha256

The installers check the archive's SHA-256 before installing anything; at a terminal, install.sh ends with probbit's own hero screen.

60-second tour

git clone https://github.com/BitmapAsset/probbit && cd probbit
cargo test --release --workspace                                   # 127 tests (3 core, 48 CLI, 11 stress, 18 acceptance, 47 probbit-ir; 1 ignored), ~25 s once built
cargo run --release -p probbit-cli -- demo --tasks 12 | cargo run --release -p probbit-cli -- decide --pretty
cargo run --release --example agent_router                         # the full narrated demo, ~3 s

The demo problem is an AI-agent task router: a queue of tasks and a set of workers (opus, sonnet, luna-pro, local-gemma, codex, human) under hard policy (PII only on-prem or with a human; production-DB migrations never on cheap models), per-hour quotas, and a bonus for keeping one customer's workflow on one worker. A rule-ignoring per-task argmax over synthetic scores (a strawman baseline) produces 60 PII leaks on 300 tasks, with 133 violations on probbit demo's queue and 144 on the agent_router example's. probbit produces 0 violations, per-task odds that match exact enumeration wherever exact enumeration is possible, and escalates the tasks whose odds it cannot pin down in the time budget.

To watch it refuse when its diagnostics fail, make the queue hard (96 % full, strong affinity):

$ probbit demo --tasks 300 --hard | probbit decide --budget-ms 50     # "refused": 0 released, exit 3
$ probbit demo --tasks 300 --hard | probbit decide --budget-ms 200    # "partial": 87-95 of 300 released, the rest escalated
$ probbit demo --tasks 300 --hard | probbit decide --budget-ms 1000   # "diagnostics_passed": 300 released, 0 violations

(5 runs each on an Apple M4.) Every escalated task's odds are in the output; the decision is left to you. The budget is wall-clock, so a slower machine needs a longer one; --sweeps N --polish-sweeps M gives the same output on every machine.

Everything runs with stock stable Rust and no external crates, so cargo build needs no network (a fresh export builds cargo build --release in 7.82 s and every target in 26-28 s on the M4).

Inside the processor

how one decision flows through probbit

The substrate (probbit-core)

A Philox4x32-10 counter RNG (reproducible, stream-splittable: the same seed gives the same decision on any machine) and two p-bit kernels: a bit-sliced multispin kernel (64 replicas per u64, one shared uniform per site, plus a parallel-tempering-in-a-word prefix-mask variant) and an SoA f32 heat-bath.

what runs updates/s on the M4 (BENCHMARKS §1)
structured multispin kernel, 64 replicas bit-sliced per word, 4 threads 1.48e11 replica-site updates/s
heat-bath kernel, one lattice, 1 thread 9.13e8
a general probbit-ir program, 1 / 4 threads 2.28e7 / 7.68e7

The first line counts 64 packed replicas per word (they share one random draw per site, so they are not independent); one problem does not run 64x faster. The first two rows are standalone lattice kernels (probbit-core/examples/kernels): probbit decide and probbit run use only probbit-core's random number generator, so your program runs at the third row's rate, about 40x slower per thread than the one-lattice kernel. probbit stats measures that rate on your machine (a re-run: 2.24e7 / 7.54e7 at 1 / 4 threads).

The instruction set (probbit-ir, JSON v1: docs/probbit-ir-json.md)

  • Variables: categorical, each with its allowed values (allowed, forbid).
  • Weights: unary log-weights per value; pairwise couplings (potts for "same / differ", table for Ising-style couplings).
  • Hard constraints, all lowered to caps so every tier runs them: caps with limit (at most k), min (at least k) or both (exactly k) over any set of (variable, value) pairs; all_different; implies (if x = a then y in a set); tables (allowed or forbidden tuples of 1-3 variables); precedes (values read as ordered slots: slot(y) >= slot(x) + gap); linear (Σ weight x [var = value] <= L: cost budgets, knapsack, bin capacity; examples/knapsack-20.json); clamp forces a value (the what-if); start gives chain 0 a warm start.
  • Instructions: decide (exact tiers, then sampler + gate), exact, sample; every answer carries marginals.
  • Front-ends lower onto it. The assignment router (probbit decide) lowers bit-identically: the acceptance test ir_lowering_bit_identical pins the router engine's gate digests. probbit run executes any program.

The inference compiler

Before any search, probbit compiles the program: exact-zero couplings and caps that can never bind are dropped, constant and separable tables are folded into the per-value weights, and the variables are split into independent components. Each part then goes to the cheapest exact method its structure allows: trees by sum-/max-product, two-value groups with uniform coupling by an occupancy-count dynamic program, groups that share few resources by a frontier dynamic program, small spaces by enumeration. Every answer says which (tier, and compiled counts on probbit run). Measured (BENCHMARKS §5b): a program with 496 exact-zero pairs 115.7 ms sampled -> 0.051 ms exact; 496 constant tables 385.7 ms -> 0.045 ms; a 1,000-variable chain 281.3 ms -> 0.390 ms; the external review's one-group routers ~0.09 ms exact. Whatever no exact method covers is sampled whole (sampling only the residual component is open).

Three tiers: exact, then sample, then gate

  • Exact tiers run first: enumeration (up to 2M feasible plans), a frontier dynamic program for programs whose groups share few resources, minimum-remaining-values enumeration for puzzles, and a components tier: independent parts solved separately, trees by sum-/max-product (a 1,000-variable chain: 0.39 ms exact vs 281 ms sampled), two-value groups with uniform coupling by an occupancy-count dynamic program. Router latency: 7.1-8.0 ms p50 at 12 tasks (enumeration) and 25.1-27.3 ms at 24 tasks (frontier) inside probbit, 9.0-10.1 / 27.2-29.6 ms for the whole process (§2.3: an earlier build and the 0.2.0 re-run, N = 20 each); 3.17 ms p50 on 200-task oracle queues (§2.1). The enumeration is bounded: it declines after max(64 x limit, 1e8) capacity checks without a new plan, so a hard search hands over to the sampler in tens of milliseconds; --op exact is unbounded.
  • The sampler: constraint-preserving Gibbs sampling (per-variable heat-bath inside the feasible set, Metropolis swap moves, an exact two-group joint heat-bath at strong affinity, auto-enabled at affinity ≥ 2; and, with --collective on (the default since 0.2.0), two Metropolis-corrected collective moves per sweep: a global flip of every free two-value variable and a swap of two value labels everywhere, which cross between mirror-image and label-permuted modes; and, with --cluster on (the default in 0.2.0), a Wolff cluster move over the positive couplings, attempted with probability 1/2 per sweep, which flips one cluster at a time: without it two weakly bridged clusters gave 8/8 and 9/9 false whole answers on seeded holdouts, see BENCHMARKS "Known failure modes"; and, with --cycles on (the default in 0.2.0; attempted only on programs with three or more values and at least one cap), n/4 three-cycle rotations per sweep that keep every value's count, which move quota-saturated programs where no single-variable change is feasible: 20/20 refused -> 20/20 whole answers, 0 false, on an untouched seeded holdout), 4 chains from over-dispersed random feasible starts on 4 threads, so no sample ever leaves the feasible set. --budget-ms covers each chain's feasible-start search as well as the sampling, so a program that is slow to start gets less sampling, or refused, rather than a late answer (except at tens of thousands of chains, where building the chains alone outlasts it: 269-318 ms of sampling on 200 at 100,000 chains, medians at loads 3.5-4.2, and no chain sweeps at all: the answer is the starts, which the gate refuses; BENCHMARKS §6).
  • The gate (the part that took the most iteration): per-marginal multi-chain batch-means Monte-Carlo standard error around the pooled mean, at two batch sizes (√n and n^(2/3)); a whole answer passes iff split-R̂ < 1.05, 3·max σ_TV ≤ 0.05, ≥ 8 batches per chain, batch-size stability (σ_long ≤ 1.5·σ), and no frozen saturated worker (a worker that is full in every sample and whose occupant classes never changed in some chain: the chain is reducible there and no within-run statistic can see it). Per-task release additionally needs R̂ < 1.002. Since gate/3 every release, and the whole answer, also needs each released variable's indicator split-R̂ < 1.05 and the occupancy counts of its capacities to agree across chains, both batch passes need ≥ 8 batches, and a variable that no chain ever moved counts as stuck (on partition and cap-free programs too) unless it is proved forced (unit propagation; on partition programs also a capacitated matching). Every threshold was calibrated against exact transfer-matrix oracles and tightened when a fresh set leaked. It was calibrated on small synthetic families only (router chain-of-blocks queues up to 200 tasks, max-cut up to 16 spins, colouring up to 14 vertices, scheduling up to 12 jobs); it is not a proof, and the frozen stress corpus defeats it (Known failure modes).

What the gate has done against exact answers: 1,731 released router tasks, 0 wrong (§2.1; current gate, re-run on 0.2.0; 1,916 on the 0.1.0 gate); 144 max-cut runs, 0 false whole answers, 1,346 released spins, 0 wrong (§3; 0.1.0 gate: with 0.2.0's dual-pass batch rule the same suite gives 61 whole answers and 929 released spins, still 0 wrong: every 500-sweep run now refuses); 525 released colouring vertices, 0 outside tolerance (§4); 428 scheduled jobs released on 39 oracle instances, 0 outside tolerance (§4.3). It also refuses needlessly (on the current gate four router oracle queues whose odds were within 0.012-0.042 of exact, two of them saturated; §2.1), and it refuses every generated sudoku (20/20) and most 3-colourings (9 of 14 on the default set, 59 of 82 over 6 sets), where its chains cannot mix (§4).

Programs that run today

program checked against result classical baseline and who wins
task router (assignment under rules, quotas, affinity) exact odds and optimum (§2.1) released odds 0 wrong, 51% of tasks released (58% on the 0.1.0 gate; §2.1); plan 0.000-0.011 nats from the optimum (median; final build / 0.1.0) on the oracle queues, 0.45-2.0 nats short at 300 tasks (§2.2) an ILP solver proves the single best plan faster on 5 of 6 queues (§2.2): ILP wins the plan; probbit gives the odds and the verdict
max-cut / Ising brute force up to 16 spins (§3) 0 false whole answers simulated annealing finds better cuts at equal time (§3): SA wins the cut
graph colouring exact odds (§4) 525 released, 0 outside tolerance the exact tier itself beats the sampler at these sizes
sudoku the unique solution / backtracking counts (§4) exact per-cell odds equal backtracking backtracking is 15-102x faster per puzzle (§4): backtracking wins
scheduling (unit jobs, windows, slot capacity, precedence) exact odds on 39 small instances (§4.3) 428 released, 0 outside tolerance at 200-500 jobs the anneal from a greedy plan beat restarted greedy + repair at equal time on 4/5 seeds each; an ILP solver (HiGHS) proves the 200-job optimum in 0.68-0.95 s while probbit's plan is 3.3-4.3 nats short (3 programs); at 200 jobs the sampler starts in 0.7 s and refuses at 1 s (§4.3)

What to use it for, and what not to promise: USE-CASES.md.

Controls and monitoring

Every control works on probbit decide and probbit run (and probbit evaluate, which takes every probbit run flag), from a flag; chains, threads, CPU limit, memory limit and priority also from a PROBBIT_* environment variable or probbit.json (flag > environment > file > default). Measured effects (BENCHMARKS §5):

control what it does measured
--threads N worker threads (sampler and plan polish) 300-task router at fixed work: 74.1 / 58.5 / 50.9 ms at 1 / 2 / 4 threads, same answer (5.1); the polish obeys it too (5.8)
--chains N independent chains not benchmarked on its own (more chains = stronger between-chain checks, more work)
--cpu-limit PCT duty-cycles the sampler stays under the cap; +11% CPU per unit of work on the router at 25%, 4.1x on a 2,000-spin ring (5.7, §6)
--priority low nice 10 + macOS background band another job slowed 1.02x instead of 1.78x; costs 4.2x wall when alone (5.4)
--mem-limit-mb N bounds the sample buffers (thinning); default 1024, 0 = unbounded 3 s run: 585.6 MB unbounded, 33.5 MB at 16 MB; peak memory = cap + ~20 MB; 10 s at the default: 1,024 vs 1,625 MB unbounded; costs releases on slowly mixing programs (5.3)
--progress [MS] JSON lines on stderr while sampling about ±5% on realistic programs, up to ~8% on tiny ones (5.2)
--sweeps N, --polish-sweeps M fixed work instead of wall-clock budgets whole output identical across runs and thread counts (test polish_sweeps_makes_the_whole_answer_deterministic)
probbit stats spec sheet: machine, build features, every control and its source, a self-test 23.05 M / 77.65 M updates/s on 1 / 4 threads (5.6; a later single run: 22.40 M / 75.36 M)

Every sampled decision carries a telemetry object: chains, threads, sweeps, updates/s, sample / gate / polish time, process CPU, peak memory, priority, CPU and memory limits (exact-tier answers and infeasible / no-start outputs carry only ms). Linux behaviour of --priority and the CPU telemetry is unmeasured; on Windows --priority low exits 2.

Use it from anything (the JSON contract)

probbit decide reads one problem document on stdin and writes one decision document on stdout (decide, run, evaluate, demo and stats write exactly one JSON document on stdout; probbit ir writes probbit-ir v0 text, probbit version one line, --help the usage; only a probbit demo whose stdout and stderr are both a terminal shows the live view instead). Exit code 0 = a plan was returned (verdict exact, diagnostics_passed or partial), 3 = refused (the whole queue should be escalated; the plan is still in the output as a best effort) or declined (an exact-only mode could not answer), 1 = infeasible (no plan satisfies the rules), 2 = bad input. Agents: probbit mcp serves the same commands over the Model Context Protocol (docs/agents.md).

--budget-ms bounds the sampling phase only. A call also parses, tries the exact tiers, runs the gate and polishes the plan, so its total is longer: the 300-task --hard demo took 206-216 ms in all at --budget-ms 50 (sampling 53-54 ms, the default 50 ms polish, the rest parsing, the exact tiers and the gate; M4, load ~3, N = 3). Every probbit run answer lists each phase in phases, and probbit run --deadline-ms N targets the whole call.

--summary prints the answer without its per-item tables: the verdict, counts (items, released, escalated), the gate, the telemetry and the items to look at first: worst_released and worst_escalated (up to 5 each, largest error bar first; on an exact answer the closest odds), each with its plan value, that value's probability p, its top two odds and its bar. Same verdict, same exit code (docs/probbit-ir-json.md "Summary").

Modelling notes

  • One entity, one variable. Decisions that belong to one entity and are tied together by implies (an action, its tool and its length for each message) freeze single-site moves: no single variable can change without breaking an implication, so the chains cannot move and the gate refuses at every budget (correctly), with a plan that may be far from the best. Make the entity ONE variable whose values are the allowed combinations, each with its combined score. Measured on an external tester's day-planning program (120 variables, 203 caps): refused at --budget-ms 2000 with plan log-weight 65.37; as 20 product variables: diagnostics_passed at --budget-ms 20, plan log-weight 93.25 (27.9 nats better).
  • Plan or odds. plan is one joint plan that obeys every rule at once; odds (marginals on probbit run) are each item's probabilities over all plans. They can disagree: an item's plan value need not be its most likely value, because the plan has to fit every other item (quotas, affinity). Act on the plan when the items must be consistent with each other (assignments under quotas); use the odds to see how sure one item is, or to rank items for a person to review.

Problem:

{
  "workers":  [ {"id": "opus", "cap": 2}, {"id": "human", "cap": 3} ],
  "tasks": [
    { "id": "T001", "group": "acme-wf0", "allowed": ["opus", "human"],
      "scores": {"opus": 2.1, "human": 1.2}, "clamp": null }
  ],
  "affinity": 1.2
}
  • The document is checked strictly (types, unknown fields, duplicate ids, empty domains; finite numbers only): bad input exits 2 with one {"error":{"code","path","message"}} object on stdout (docs/probbit-ir-json.md "Input contract").
  • scores: finite numbers with |x| <= 1e9 (natural-log odds; docs/probbit-ir-json.md "Numeric contract"), one per worker the task may go to (log-odds from your judge, an LLM, a heuristic, a price). A worker missing from scores is not allowed unless listed in allowed (then it enters with score 0).
  • allowed (optional): hard rule. cap: hard per-worker quota. clamp (optional): force this task to one worker (what-if).
  • group (optional): tasks sharing a group get +affinity in log-weight for every pair placed on the same worker.
  • The model is log w(plan) = Σ scores + affinity · #(same-group pairs on one worker), restricted to plans that obey every rule. Odds are the marginals of that distribution; the plan is its (polished) mode.

Decision:

{ "verdict": "partial", "plan": {"T001": "human", …}, "plan_logw": 412.7, "violations": 0,
  "odds": {"T001": {"human": 0.52, "opus": 0.48}, …},
  "released": ["T002", …], "escalated": ["T001"],
  "gate": {"rhat": 1.0004, "tv_bound": 0.061, "tv_tol": 0.05, "frozen_saturated_workers": 0, "min_batches": 14,
           "batch_ratio": 1.21, "samples": 36288, "chains": 4, "budget_ms": 200, "seed": 7}, "ms": 203.9 }
  • exact: odds are exact, logz and top_plans are included, and tier says how. enumerate: the feasible set was small enough (≤ --exact-limit, default 2M plans) to list (n_feasible included; ≈ ≤ 12 tasks × 6 workers, ~7 ms). frontier: a dynamic program over worker loads solved it without listing plans, which works when groups share few workers (frontier_states = its largest layer, capped by --frontier-states, default 4096, 0 = off). On the demo this covers 13-24 tasks in 13-63 ms (13 tasks: same odds and log Z as enumerating all 17.8M plans, which takes 340 ms) and declines in 6.5-9.1 ms (median) on 36-300 tasks, which then go to the sampler. occupancy / forest / components (the inference compiler): independent groups solved one by one; a two-worker group with uniform affinity and whole-worker caps by a count dynamic program over "how many tasks on B" (the external review's 32-task counterexamples: 0.09 ms exact, where the sampler took ~426 ms), cap-free trees by sum-/max-product (components = the per-tier counts).
  • diagnostics_passed: the sampler's whole answer passed the gate's whole-answer test (a 3-sigma Monte-Carlo bound of ±0.05 total variation per task plus the mixing checks, see below). Diagnostics, not a proof: a mode no chain visits is invisible to them (Known failure modes). Its released tasks did NOT all pass the stricter per-item R-hat test (release_reason: whole_answer_gate).
  • partial: the run passed the per-item global checks (R-hat < 1.002, batches, batch stability) and only released tasks passed their own error bar (release_reason: item_gate); act on those, escalate the rest.
  • refused: the sampler cannot vouch for the answer (chains disagree, or a saturated worker is frozen). Escalate.

Options: --budget-ms (default 200, wall-clock for the 4 chains including their feasible-start search; the gate and the polish run after it, so the default answers in ~300 ms; probbit demo also takes --tasks, --seed, --hard), --seed, --exact-limit, --exact-ms (opt-in wall-clock cap on the exact tiers; absent = no cap), --frontier-states, --polish-ms (annealed plan polish after sampling, default 50), --mode auto|exact|sample (exact enumerates with no plan limit, ignoring --exact-limit, so on a big input it may never finish: one router group of 3,000 tasks on 2 workers (2^3000 plans) gave no answer within 90 s. Pair it with --exact-ms), --pretty. Processor controls (same on probbit run): --sweeps, --chains, --threads, --cpu-limit, --mem-limit-mb, --priority, --progress (the five from --chains to --priority also from env PROBBIT_* or probbit.json); probbit stats prints the machine, the effective controls and a measured self-test (docs/probbit-ir-json.md, "Resource controls"). probbit ir prints the problem in probbit-ir v0, the text form that lowers to a one-hot p-bit/QUBO topology (what a hardware p-bit fabric samples).

From Python (python/probbit.py, stdlib only, Python >= 3.9; typed errors, deadlines; python/examples/):

import probbit                                  # python/probbit.py
decision = probbit.decide(problem, budget_ms=200)   # exit 1 / 3 come back as answers; bad input raises probbit.ProbbitInputError
answer = probbit.run(program, deadline_ms=1000)     # a probbit-ir program (docs/probbit-ir-json.md)

From any agent harness, the raw pattern:

import json, subprocess
problem = {"workers": [...], "tasks": [...], "affinity": 1.0}
r = subprocess.run(["probbit", "decide", "--budget-ms", "200"], input=json.dumps(problem), capture_output=True, text=True)
decision = json.loads(r.stdout)          # r.returncode: 0 plan, 3 refused, 1 infeasible, 2 bad input (or --mode exact stopped by --exact-ms)

Build the binary once with cargo build --release -p probbit-cli (it lands in target/release/probbit). There are no native Python bindings (probbit mcp is the MCP server, docs/agents.md); every script in bench/ is an example of the subprocess pattern. probbit <command> --help lists every flag of a command with its default, and the exit codes.

The individuality layer: probbit persona

Give an agent a temperament that lives outside the model. A persona is a small file (YAML subset or JSON): traits with priors, moods with inertia, soft couplings, per-turn evidence, and habits that are hard rules. A seed makes an individual. Every turn compiles persona + the individual's state + the turn's inputs into ONE probbit-ir program, answered exactly in process; the stance (a level per trait with exact odds, the habits in force and the ones that changed it, a refusal when the engine cannot vouch) comes with a short stance line a host puts into any model's prompt. The model still writes every word; the persona decides how they are written, and the same individual answers whatever model the host calls. docs/persona.md is the format; examples/persona/ has three fictional personas and their goldens.

probbit persona init examples/persona/tutor.yaml --seed 2 --out pip.json      # an individual: genes from the seed
probbit persona turn examples/persona/tutor.yaml --state pip.json --inputs '{"loss": true, "sentiment": "negative"}'
#   "status":"ok", "humour":{"level":"none","p":1,...}, "warmth":{"level":"warm","p":0.952032,...}, habits active: no_jokes_on_loss, ...
#   "line":"Stance: no jokes, be kind; one emoji at most; gentle; explain step by step; warm and encouraging; suggest what to try next; ask what they think first; casual."
#   "why":"learner reports a failure -> humour none, valence down; learner upset -> valence down, humour none"     (pip.json is now turn 1)
probbit persona replay examples/persona/tutor.yaml --seed 2 --script examples/persona/workday.json   # 20 stances, the same bytes every run
probbit persona lint examples/persona/tutor.yaml   # contradicting habits: confused_no_playful vs after_error_check, resolved by priority
import probbit                                        # python/probbit.py, stdlib only
state = probbit.persona_init("examples/persona/tutor.yaml", seed=2)
turn = probbit.persona_turn("examples/persona/tutor.yaml", state, {"loss": True})
prompt_tail, state = turn["stance"]["line"], turn["state"]   # the line goes after any cached prompt prefix

MCP: probbit_persona_init and probbit_persona_turn in probbit mcp (stateless: the persona, inline or a path, and the state go in; the state comes back). Browser: the playground's "Meet three individuals from one persona" (probbit-wasm ops 4 and 5, no threads). Every surface gives the CLI's documents byte for byte (tests/persona.rs, python/test_mcp.py, probbit-wasm/tests).

Measured in this run (Apple M4, macOS 26.5.2, 2026-10-02; the three example personas):

  • a whole turn in process (compile, engine, decode, digests): median 0.20-0.32 ms, p95 0.22-0.67 ms, of which the engine 0.09-0.18 ms (median); N = 1,000 turns per persona, 1-minute load 3.3-3.5, nice 10;
  • the stance line: 36-38 estimated tokens on average over the 20 workday turns, at most 40;
  • individuals: three seeds of the tutor differ by a mean total-variation distance of 0.12-0.18 between their trait odds on the workday script; the tutor and the ops engineer (two persona files) by 0.45; the same persona and seed by 0;
  • habits: 0 violations over 2,160 turns of 12 random personas (every habit in force re-checked from the file), and each of the 481 habits reported as binding is broken by the habit-free twin;
  • determinism: 1,000-turn replays are byte-identical across processes (3 personas x 2 seeds); the documents equal an independent reference implementation's on 5,460 turns (the 60 golden turns and 5,400 turns of random inputs).

What a persona can NOT do: write, read or check text (text rules stay in the prompt or a checker); see what the host does not tell it; be a safety gate (keep approvals and permissions in plain code); its odds are its own model's, not measured probabilities that a user will like the reply; and it cannot make a model follow the line: whether a given model writes in the stance it is given has to be measured per model (no such measurement has been made here).

After any judge: probbit evaluate

Decision models (TypeSafe's Jev, Cloudflare's Clef and Clef-flash, local System One servers) are judges: content in, a probability per option for each question out, each question on its own, no rules, no joint answer, no refusal. probbit evaluate is the layer after any of them. It reads the judge's System One request (the shape both vendors publish: model, state, questions keyed by id with types noul, choice, score) plus a probbit block with the judge's answers and your rules over question ids, compiles it to a probbit-ir program (one variable per question, log-weight = ln max(p, 1e-6), the rules as written) and answers in the judge's response shape: the most likely answer set that obeys every rule, odds per option under the rules, and per answer whether probbit moved it from the judge's own answer, plus the whole probbit run document (verdict, gate, violations, released / escalated). Without rules it returns the judge's answers and the same probabilities (150 random requests: largest gap 1.1e-16); with rules only rule-connected questions whose judge answers break a rule move, each group to its most likely rule-abiding assignment (191 groups checked by brute force). docs/probbit-ir-json.md "Decision API".

# shell
probbit evaluate --summary --pretty < examples/evaluate/support-12.json
# Python (python/probbit.py, stdlib only)
PYTHONPATH=python python3 -c 'import json, probbit; print(probbit.evaluate(json.load(open("examples/evaluate/support-12.json")))["verdict"])'
# MCP: tool probbit_evaluate, here over a pipe (in an agent: claude mcp add probbit -- probbit mcp)
{ echo '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"sh","version":"1"}}}'
  python3 -c 'import json,sys; print(json.dumps({"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"probbit_evaluate","arguments":json.load(sys.stdin)}}))' < examples/evaluate/support-12.json; } | probbit mcp
# browser: build the module, then open the page from disk (no server)
sh playground/build.sh && python3 -m webbrowser "file://$PWD/playground/index.html"

On the example (12 questions, 12 rules): exact, 5 answers moved, 1.24 ms inside the answer (Apple M4, median of 7). Python can also ask the judge first (probbit.evaluate(request, judge=<a callable or a System One URL>), stdlib urllib); the browser page runs the CLI's own code compiled to WebAssembly without threads (1.17 MB with the persona ops, 830 KB before them; the 300-task router demo at 3,200 sweeps in 531 ms in Chrome against 346 / 133 ms native at 1 / 4 threads). More in docs/agents.md, "After a judge". The refusal is the point: when the gate does not pass, the verdict is refused (exit 3) and every answer is marked unreleased instead of guessed.

Library API (Rust)

use probbit_decide::*;
let p: Problem = /* t tasks, a workers, h logits, allowed mask, cap, group, lam, clamp */;
let (d, gate) = decide_gated(&p, 2_000_000, 0, Some(200.0), 7, &GATE).unwrap();   // exact or 4-chain sampler + gate
let released = gate.as_ref().map(|g| g.released_tasks(&GATE));                     // per-task release
let anytime = decide_anytime(&p, 500.0, 25.0, 1.0, 7, &GATE, false, None);         // stop as soon as the gate passes
let (logw, plan) = polish_plan(&p, Some(&d.map), 50.0, 7).unwrap();                // annealed best plan

Exact-oracle instance families for your own calibration live in probbit_decide::oracle (build, build_sat, exact_dp, exact_map_logw).

Examples and calibration harnesses (cargo run --release --example <name>)

agent_router (start here) · dispatch (the original ticket-routing narrative, [1]-[9]) · accuracy (T=200 vs exact DP) · calibrate, calib_sat (gate false-release / false-refusal grids vs exact oracles) · scale (T=200..1000) · anytime · hard_affinity · gpair_check (the frozen-split trap) · map_gap (best-seen / polished plan vs the exact optimum) · tempering, chains, focus (refuted or off-by-default accelerators, kept so the negative results are reproducible) · hardware_lowering (probbit-ir → one-hot QUBO) · -p probbit-core --example kernels. The probbit-ir crate has its own oracle examples for max-cut, colouring, sudoku and scheduling (maxcut_oracle, maxcut_vs_sa, colouring_oracle, sudoku_bench, schedule_oracle).

Acceptance tests (probbit-decide/tests/acceptance.rs): TV ≤ 0.01 to exact within 10 ms; 0 violations in ≥ 10⁶ samples; exact = brute force; the refusal gate refuses; diagnostics_passed ⇒ within tolerance on the calibration oracles (not on the stress corpus); IR round trip and lowering exact; the frozen-split trap is refused and the two-group move is exact; anytime answers and polish are correct vs oracles. acc1 has a wall-clock bound and can fail on a heavily loaded machine.

Build note: the default build is portable (no CPU pin): a binary runs on any CPU of its target triple. For the last bit on your own machine, RUSTFLAGS="-C target-cpu=native" cargo build --release (or uncomment the lines in .cargo/config.toml); that binary may crash with an illegal instruction on other CPUs. Measured on an Apple M4 (5 runs): native vs portable 0.975-1.012x on all 16 kernel/sampler metrics, identical answers. On x86-64 (where native can add AVX2/AVX-512) it is unmeasured.

What it does, measured (Apple M4, one machine)

Rows marked earlier build were measured on an earlier version of the engine, before the general instruction set; those calibration reports are not included in this repository. Everything else reproduces from the examples and BENCHMARKS.md.

exact joint answer, 12 tasks × 6 workers 180,540 rule-abiding plans enumerated in 7.3-8.2 ms (8.2 on 0.2.0: +10-12% at 12 tasks and +22% under full enumeration of small groups, BENCHMARKS §6), exact odds; the forced sampler takes 34-37x longer and is only approximately right (BENCHMARKS §6)
300-task router queue, 6 workers, 315 quota slots (95 % full) rule-ignoring argmax over synthetic scores (the agent_router example's queue): 144 violations, 60 PII leaks; probbit: 0 violations; 25 ms: 0-128/300 released (the budget is wall-clock and this one sits at the R-hat threshold: 0, 0, 121 and 126 in four 0.2.0 runs, 128 in an earlier build), 200 ms: 300/300; anytime diagnostics_passed after 211-215 ms (0.2.0 runs; 186 ms in an earlier one)
300-task pod queue with a computable exact answer (89 workers) 200 ms: 291-298 of 300 released, 1 s: 300/300; 0 released tasks wrong across 11 seeds; polished plan within 0.75-3.6 nats of the proven optimum (earlier build; current build, seed 7: 298/300 released, 0 wrong; 296-297/300, 0 wrong on 0.2.0 re-runs)
released tasks that were wrong (vs exact oracles, all families) 0 of 28,007 on a fully out-of-sample calibration set (earlier build; a rule-of-three bound of ≈ 1.1e-4 per task would assume independent tasks, which tasks sharing a run are not). Current build (gate/3; all four re-run on 0.2.0): 0 wrong of 1,731 released router tasks, 929 max-cut spins, 525 colouring vertices and 428 scheduled jobs (BENCHMARKS §2-§4)
whole-answer false passes on exact-oracle sets 1 of 408 (earlier build); the guard added for it gives 0 of 393 in-sample; the one fresh set after it (96 hard runs) passed none. 0 false of 72 max-cut whole answers on the 0.1.0 gate, 0 of 61 on 0.2.0 (BENCHMARKS §3, re-run on 0.2.0). Adversarial stress corpus (0.2.0): 2-3 of 20 seeds per family pass with every released item wrong (Known failure modes)
anytime mode (decide_anytime in the library) median 117 ms to a whole answer vs a 300 ms fixed budget, 0 false in 70 (earlier build)
the trap: 4 chains agree, all wrong (true error 0.51) an earlier gate passed it; the shipped gate refuses (7 frozen saturated workers)
kernel throughput (probbit-core, 64 replicas per u64 word) see cargo run --release -p probbit-core --example kernels

What it does not do, also measured:

  • It is not faster than an ILP/CP solver at finding the single best plan, and it does not prove optimality. Its polished plan is 0.45-2.0 nats short of the optimum at 300 tasks (9.3-9.7 on the hard queue), and an ILP solver (HiGHS) proved that optimum in 52-352 ms, faster than probbit's default path on 5 of 6 queues (BENCHMARKS §2.2); on 200-job schedules HiGHS proves the optimum in 0.68-0.95 s and probbit's plan is 3.3-4.3 nats short (§4.3). If all you want is the argmax plan, use a solver. The value here is per-item odds under your score model plus per-item escalation.
  • If you only need one good feasible plan, greedy + the same anneal polish is better at 300 tasks: at equal time it beat probbit's plan on 8 of 8 runs (BENCHMARKS §2.2). Greedy alone takes 0.05-0.2 ms. The sampler earns its 200-1,000 ms when you need per-item odds (its marginal argmax matches the exact most likely worker on 96-100% of tasks on the oracle queues, a single optimal plan on 71-78%, BENCHMARKS §2.1), a gated release/escalate decision, or what-if marginals.
  • Max-cut: simulated annealing finds better cuts at equal time (11470 vs 11463 and 572 vs 568, §3). Sudoku: plain DFS backtracking does the same exact job 15-102x faster per puzzle (§4); the sampler refuses sudoku, the exact tier answers it.
  • It refuses a lot in hard regimes: exactly 100 % full queues, strong affinity with small dedicated quotas, and 1,000-task queues at 200 ms. On the 16 oracle queues of BENCHMARKS §2.1, 51% of tasks were released (58% on the 0.1.0 gate) and 4 of the 9 queues that released nothing were refused needlessly (3 of 8 on 0.1.0; an earlier build measured a whole-answer false-refusal rate of 0.3-0.5 at 300 tasks).
  • The router document (probbit decide) expresses allowed sets, per-worker quotas, clamps and same-group affinity only. Cost budgets (linear), ordering (precedes), all_different, implications and tuple tables are in the general probbit-ir format (probbit run). Missing in both: overflow / deferral as a rule (every task must go somewhere; model it as an extra "nobody" value with its own score), and objectives beyond per-value scores and pairwise terms.
  • No real traffic yet. The demo's scores are a seeded stub. Calibration is on synthetic exact-oracle families.
  • Nothing beyond the CPU's own compute, and no special hardware. A Metal GPU path was measured and rejected (a decision is ~4k sites, too small to feed a GPU). Unmeasured: x86-64, Linux, Windows, CP-SAT, annealing hardware, real customer data, programs over 100k variables.

Known failure modes

The full list, with the exact commands and numbers, is in BENCHMARKS "Known failure modes" and §6 "Where it loses". In short:

  • The gate can be fooled. With the barrier-crossing moves off (--collective, --cluster, --cycles), adversarial stress families (one-group routers with a slightly better worker, complete-graph ferromagnets, heterogeneous groups) pass it with every released item wrong in 2-3 of 20 seeds (four of the six frozen families; the other two are refused 20/20; re-run on 0.2.0: 2 / 3 / 2 / 0 / 3 / 0): every chain sits in the same wrong mode, which no within-run diagnostic can see. With the default moves the frozen stress corpus gives 0 false and 0 wrong (280 runs, 9,280 released items, re-run twice on 0.2.0); an untouched family can still defeat it.
  • It refuses a lot: saturated and strong-affinity router queues (49% of the oracle-queue tasks are not released on the current gate), every generated sudoku, most 3-colourings, 3-SAT sampling (BENCHMARKS §2.1, §4, §4.8).
  • --budget-ms is not a whole-call deadline. The exact tiers run outside it: the default enumeration walks up to 64 x --exact-limit nodes before it declines (2.7 s on a near-saturated 3,000-task group in 0.2.0, 199.9 s before), and at 100,000 chains the chain builds alone exceed it. --exact-ms and --deadline-ms (probbit run) are the controls.

Hard questions

  • Is this quantum? No. A p-bit sits between a bit and a qubit only as a metaphor: it holds a probability: it is never negative, and nothing interferes with anything. Quantum hardware works with amplitudes that can cancel; probbit has none of that, is not quantum hardware and does not simulate any. What ships here is classical Markov chain Monte Carlo (constraint-preserving Gibbs sampling) on your CPU. What the two share is a job: drawing samples of good joint configurations under couplings and constraints, which is what people rent annealing machines for. probbit's gains on that job are the programming model (rules hold by construction instead of through penalty terms, BENCHMARKS §7), odds you can check, and a gate that refuses when its diagnostics fail (a heuristic with published counterexamples); not more compute than the chip has. It has not been compared with annealing hardware (unmeasured). The honest label is quantum-inspired.
  • Isn't this just Gibbs sampling or simulated annealing with a new name? The sampler is Gibbs sampling, and simulated annealing finds better max-cuts at equal time (§3). What is added is the instruction set with hard rules kept by construction, the exact tiers first, and a gate that refuses.
  • Why not an ILP or CP-SAT solver? For one best plan, use one: HiGHS proved the 300-task optimum in 52-352 ms, faster than probbit on 5 of 6 queues, and probbit's plan was 0.45-2.0 nats short (§2.2); on 200-job schedules HiGHS proved the optimum in 0.68-0.95 s and probbit's plan was 3.3-4.3 nats short (§4.3). An ILP returns no per-item odds, no check on them and no refusal. CP-SAT was not run.
  • What does "diagnostics_passed" mean? (0.1.0 called it "certified".) A 3-sigma Monte-Carlo error bound (at most 0.05 total variation per released item) plus split R-hat, batch-size stability and frozen-resource checks, tuned on exact oracles. It is a diagnostic, not a proof: 0 false out of 61 passed max-cut answers (72 on the 0.1.0 gate) bounds that family's false-pass rate below ~5% (95%, rule of three; §3, re-run on 0.2.0), while adversarial families (one-group routers with a slightly better worker, complete-graph ferromagnets) pass it with every released item wrong in 2-3 of 20 seeds when the barrier-crossing moves are off (0 of 20 with the default moves; BENCHMARKS "Known failure modes").
  • "0 wrong" is easy if you refuse. Where does it lose? It refuses a lot: 51% of tasks released on the router oracle queues (58% on the 0.1.0 gate) and none on saturated or strong-affinity queues (§2.1), every generated sudoku and most 3-colourings (§4). It also loses the single best plan (ILP), the best cut (simulated annealing), puzzle speed (backtracking counts sudoku 15-102x faster per puzzle) and tiny inputs (the exact tier is 34-37x faster than sampling at 12 tasks, §6).

The hardware hand-off

probbit ir prints a router document in a text form that lowers to a binary one-hot QUBO with slack bits, in which the energy of every feasible plan equals −log w exactly (acceptance test ir_round_trip_and_lowering_exact; examples/hardware_lowering). That is the input a p-bit or annealing chip takes, after minor embedding on real hardware (unmeasured). No hardware has been run; embedding cost and hardware speed are unmeasured. The CPU path does not sample the QUBO: on three 34-bit instances (BENCHMARKS §7) the native sampler had 0 infeasible samples and odds within 0.0018-0.0049 TV, while a p-bit sampler on the QUBO lowering either leaks infeasible samples (penalty weight 0.5: 100% infeasible on all three; 2: 82-90%) or loses accuracy as the penalty grows (16: 0% infeasible, but 0.09-0.23 TV). This measures mixing on this CPU, not hardware.

How it was built and checked

Every claim in this README points at a table in BENCHMARKS.md, a use case in USE-CASES.md, or a test. Every gate threshold was set against exact oracles and tightened whenever a fresh set leaked; when a claim broke, the fix shipped with a test that would have caught it. Accelerators that did not help (tempering, extra chains, focused sweeps) and baselines that win (ILP, simulated annealing, backtracking) are kept as runnable examples so the negative results stay reproducible. All timings come from one Apple M4 Mac mini (4 performance + 6 efficiency cores, 16 GB), portable release build, some of it under load; nothing here has been measured on x86-64, Linux or Windows.

Roadmap

  • Cut false refusals without raising false releases. A Rao-Blackwellized marginal estimator was tried and held behind a flag: safe, but it does not lower false refusals, because the worst item's error bar comes from slow mixing, not sampling noise. Next is an exact k-group block move that unfreezes strongly coupled, saturated instances, and a gate statistic that is not set by the slowest item.
  • Overflow (an explicit "nobody" worker) in the router document (cost budgets, windows and precedence exist in probbit-ir).
  • Per-item error bars (σ, R̂) for every item in the JSON output, so integrators can set their own release thresholds (--summary reports the bar of the 5 worst released and escalated items).
  • The sampling phase at 100,000 chains overruns --budget-ms (269-318 ms for 200 in 0.2.0 at loads 3.5-4.2, 354-448 before: the chains' builds alone exceed it, and no chain sweeps: sweeps is 0; the gate takes 84-100 ms there; BENCHMARKS §6). --deadline-ms (probbit run) bounds the whole call.
  • A CP-SAT baseline (ILP baselines: router BENCHMARKS §2.2; scheduling §4.3).
  • Native Python bindings (today: python/probbit.py, a stdlib-only subprocess wrapper with typed errors, python/examples/, and the MCP server probbit mcp, docs/agents.md).
  • Real routing logs: the scorer side has never been validated on real traffic.
  • Hardware backend through probbit-ir when a p-bit fabric is available.

Documentation map

  • BENCHMARKS.md: every number, with machine, load and N.
  • USE-CASES.md: what the measurements support, by problem shape and by industry.
  • docs/probbit-ir-json.md: the probbit-ir JSON v1 wire format, instructions, resource controls, limits.
  • docs/probbit-ir.schema.json: a JSON Schema of that wire format (the page above is normative; a test keeps every object's fields equal to the parser's).
  • docs/agents.md: calling probbit from a shell, Python, Node, PowerShell, MCP agents and a browser, and probbit evaluate after a decision model.
  • docs/persona.md: probbit persona, the individuality layer: the persona file, the compilation, the stance and state documents, the canonical JSON, what a persona can not do; examples/persona/: three personas + goldens.
  • python/: probbit.py, a zero-dependency subprocess wrapper (with evaluate, the persona functions and a stdlib mock judge), its tests and three examples.
  • playground/: one static page that runs probbit in a browser (probbit-wasm, built by playground/build.sh).
  • bench/: the scripts behind BENCHMARKS §2 and §5 (Python 3; the ILP baselines need numpy and scipy >= 1.9).
  • CONTRIBUTING.md, CHANGELOG.md.

License

Apache-2.0.

About

probbit — the p-bit processor your agents decide on: every rule kept by construction, exact odds for every part, and a refusal when it can't vouch. Zero-dependency Rust, CPU only.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages