One set of test assertions, defined as data, so the same test means the same thing in every language that implements it.
This repository holds the definition. It ships no library. Each implementation reads these files and runs this corpus in its own CI, so a library that omits an assertion, names one wrongly, or disagrees about what an assertion means fails its own build.
Write the same test twice, once in Go and once in Python, and the two should pass or fail together. Today they do not.
Go's cmp package, configured the way most test helpers configure it,
says a nil slice equals an empty slice. Python says None does not
equal []. Port a suite from one to the other and both builds go green
while testing two different things. The difference surfaces months later
as a bug that reproduces in one service and not its rewrite.
Six libraries written from a prose specification drift the same way, and the drift stays invisible because each library's own tests pass. Each team tests what it believes the specification says. Nothing tests whether the beliefs agree. Only a shared artifact, read by all of them, can catch that, and it has to be data because prose does not fail a build.
spec/assertions.yaml what each assertion means edited by people
spec/naming.yaml what each language calls it edited by people
spec/assertions.json the same, rendered read by libraries
spec/naming.json the same, rendered read by libraries
spec/conformance.md what converges and what does not
spec/manifest.json a digest of everything an implementation vendors
spec/zones.json the zone list and its offset changes, from tzdata
spec/encoding.md how a corpus case states a value
spec/overlays.md how a language declares it cannot comply
spec/recording.md what a recorded run writes for each call
corpus/*.json the cases, one file per assertion they cover
corpus/prop/*.yaml the property engine's vector inputs edited by people
corpus/prop/*.json the vectors with their outputs read by libraries
corpus/history/*.yaml the history's vector inputs edited by people
corpus/history/*.json the vectors with their outputs read by libraries
corpus/stateful/*.yaml the machines' vector inputs edited by people
corpus/stateful/*.json the vectors with their outputs read by libraries
corpus/files/*.yaml the file assertions' vector inputs edited by people
corpus/files/*.json the vectors with their outputs read by libraries
overlays/*.json one per language, declaring divergences
tools/render.py YAML to JSON, and the vectors' outputs
tools/validate.py the rules, checked
tools/prop/ the property engine's executable reference
tools/history/ the history's and the checkers' executable reference
tools/stateful/ the machines' and the scheduler's executable reference
tools/files/ the trees' and the file assertions' executable reference
tools/spec-sync.sh how an implementation vendors the definition
tools/spec-check.sh how an implementation checks its copy
VERSION 7.0.0
People edit the YAML. make render produces the JSON, which is
committed and is what implementations read: every target language parses
JSON from its standard library, and several would otherwise take a
dependency just to read the definition. CI re-renders and fails if the
result differs from what is committed.
An assertion has a canonical id that no user types. Each language maps that id to a name its users recognise.
# spec/assertions.yaml — what it means
"throws":
arity: 3
summary: >
A callable raises. Yields what was raised.
detail_fields: []# spec/naming.yaml — what a user types
"throws":
go: "Panics"
python: "raises"A single shared vocabulary would read as a translation in most of the
six languages. Splitting the id from the name lets a Python developer
write raises and a Go developer write Panics while both answer to
one definition.
71 assertions: 50 in the root namespace, 4 for golden files and golden trees, 9 for trees of files and single paths, 4 for benchmark ceilings, 1 property check and 3 checks of a recorded history. They cover equality, truth, nullity, length, containment, text, numbers, ordering, errors, raising, cancellation and deadlines, retrying, goroutine and task leaks, allocations, relations between runs of a subject, recorded output, files and directories, performance ceilings, properties over generated inputs, the linearizability of concurrent calls, and the isolation of transactions.
39 of them also have a property form, which runs the assertion on every
input that a property generates. Rendering adds the forms to the prop
package by one rule, so the assertion table states 110 entries.
An assertion earns its place by answering two questions. Does it state something that must be true, and fail when it is not? Does it mean the same thing in every target language? Anything that fails the second is a helper, and helpers live in the libraries.
Conformance is checked three ways, and they catch different things.
The corpus checks meaning. Each case states arguments as typed literals and says whether the assertion passes or fails, and sometimes what the failure must mention:
{
"id": "equal/null-against-empty-list",
"args": [
{ "type": "list", "of": "int", "value": [] },
{ "type": "null" }
],
"expect": "fail",
"detail": {
"want": { "type": "null" },
"got": { "type": "list", "of": "int", "value": [] }
}
}Typed literals cross a language boundary only as data, so the corpus
covers 39 of the 71 assertions. Eighteen of those state their
arguments. The other 21 name a behaviour instead, because what they
take is a callable and no encoding states one. The ten assertions that
read files take a directory or a path, which a case states as a tree
that the runner writes before the call, so their cases are vectors. The
remaining 22 take an error value, a predicate, a callable that no subject
describes, a golden file, a benchmark measurement, a property's body, a
spec or a recorded history, and none of those is a typed literal
either. The property engine
itself is data in and data out, so 430 vectors under corpus/prop/ pin
its decoding, generation, shrinking, coverage test, fuzz bridge, replay
token, run detail and store. They also pin the values each shape
generates, the choices that produce a value, the shape each fixture type
reads as, a passing and a failing run of every property form but
prop-max-allocs, the runs of a form's examples, and the call records of
a property's runs. The history
and the checkers are data in and data out too. 80 vectors under
corpus/history/ pin the events that calls record, the entries that the
history refuses, the verdict, steps and record of a check against each
named spec, and the verdict and record of each isolation check. So are
the steps of a machine. 14 vectors under corpus/stateful/ pin the run of
each machine subject, the minimal steps of each fault, the traces that a
run follows or refuses, and the step at which a replay of a subject whose
refusals change its actions diverges. 57 vectors under corpus/files/
pin the verdict and the record of the ten assertions that read files:
the comparison of two trees, its bound of 64 paths and its digest of a
large file, a golden tree and its update, and the kind, the target, the
content and the mode at one path.
The completeness gate checks membership. Every assertion must be
present under the name the naming table gives it, with the arity the
definition states as far as the language can read it. The gate covers
the 22 assertions that the corpus cannot state, and prop-max-allocs and
prop-max-allocs-with-setup, whose allocation counts no vector can
state. The standard checks a
library's meaning where meaning can be stated, and its membership
everywhere else.
An overlay is where a language declares it cannot comply, with the reason. A divergence nobody wrote down is a bug; one written down is a decision someone can argue with.
Which of these a given difference belongs to, and which differences need
no recording at all, is stated in spec/conformance.md.
A run with DOKIMI_ASSERT_RECORD=1 writes a call record for every
assertion call, pass or fail, into the artifact that the language's test
runner already writes for a run. In Go, that artifact is the event
stream of go test -json. A call record states the assertion, the
contract, the verdict and the detail of a failure. The calls in a
property's cases appear under the property's own call, with the phase
of each case. spec/recording.md fixes the record, and each overlay
names its language's artifact.
An implementation vendors a copy of the definition, so its build fails on its own without reaching the network. What a copy cannot tell you is whether it is current, and the version does not answer that: adding the relaxations changed the definition without changing the version, and by the rule below it should not have.
spec/manifest.json carries a digest of every file an implementation
vendors, so there is something to compare against that tracks the bytes
rather than the meaning. Each implementation runs spec-check in its own
CI. A copy that does not match the manifest beside it fails, always: each
file, the overlay of the copy's language included, and the manifest's
own digest of those files. A copy that differs from this repository
fails only when that change is the one that touched it: falling behind
is allowed and is tracked by an issue, and committing a copy nobody else
has is not. A change that touches the copy also fails when this
repository cannot be read, because the comparison did not happen.
Each implementation opens that issue on itself, on a weekday schedule, by running the same check against this repository's main branch. It reads rather than being told, so nothing here holds a key to five other repositories and there is no token to rotate. A library already current opens nothing, because the check compares digests rather than counting pushes.
spec-sync fetches a pinned ref rather than reading a sibling
directory, so it answers the same way on a laptop and on a runner. Set
SPEC_LOCAL to try a change before pushing it; it says loudly that the
copy it leaves behind is reproducible nowhere else.
The two scripts are themselves vendored from tools/ here and carried
in the manifest, and spec-sync refreshes them along with the
definition. Five copies of a script drift exactly the way five copies
of the definition did, and running one from the network instead would
make an offline check depend on being online.
A change here opens an issue on each of the five, so a definition change becomes five pieces of visible work rather than five silent divergences.
Everything runs through uv, which fetches its own Python. Clone and run the gate; there is nothing else to install.
make install # create the environment
make check # the full pre-merge gate
make render # rebuild the JSON from the YAML
make validate # hold the definition and corpus to their rules
make test # check the validator catches what it claims to
make fmt # format the toolsmake lint-md needs markdownlint, and falls back to npx when it is
not on the path. Every other target needs only uv.
make validate reads the rendered JSON, not the YAML, because that is
what implementations read. It checks that the version files agree, that
every assertion is described, that every language that names one
assertion names all of them, that a qualified name names a member of the
package its assertion declares, that every corpus case names a defined
assertion with a unique id, decodable literals and options its assertion
accepts, and that an overlay extends this version and diverges only from
assertions that exist. It also checks that each overlay states where its
language writes the call records and how it runs the concurrent section of
a machine, that each history vector names a defined spec, that the
vectors of each isolation level report every kind the level forbids, that
each machine subject runs in a vector named for it, that the workspace and
the golden tree of each files vector follow the rules of a tree, and that
each language whose threads run on more than one core limits the
history's recorder. It reports everything it finds in one run.
make test breaks each rule of the validator in a scratch copy and
requires the validator to report it, and breaks a vendored copy each way
that spec-check.sh must refuse. A validator only ever run on a clean
tree would pass just as readily with every rule deleted.
VERSION carries the version of the definition. An overlay names the
version it extends, so an overlay left behind by a change to the
standard fails validation rather than passing quietly.
Adding an assertion is a minor version. Changing what an existing assertion means, or renaming one, is a major version, because it changes whether an existing test still states what its author meant. Renaming a member of the surface table is a major version for the same reason, and so is a corpus case that pins an answer the implementations gave differently.
| Language | Repository | Assertions |
|---|---|---|
| Go | assert-go | 110 of 110 |
| Java | assert-java | 105 of 110 |
| Kotlin | assert-java | 105 of 110 |
| Python | assert-python | 105 of 110 |
| Rust | assert-rust | 110 of 110 |
| TypeScript | assert-typescript | 104 of 110 |
Each count is what the language's overlay declares against version 7.0.0. An implementation that has not synced to it yet has a drift issue open until it does.
Java and Kotlin ship from one repository and are named identically, so a test reads the same in both. Neither states a ceiling on allocation count in a test, a property or a benchmark. The JVM reports bytes allocated per thread and no count of allocations. Python states none of the three, because CPython reports the memory alive at one moment and no running count. For the same reason its ceiling on bytes bounds the peak of an iteration and not the bytes it allocates, which its overlay declares as a limit. TypeScript states none of the four allocation ceilings, because V8 reports allocation only as a heap-usage delta that moves with whether the collector ran. Each gap is in that language's overlay with the measurement behind it.
Go and Rust state all 110 and declare nothing absent. Go checks no
allocation ceiling in a build with the race detector, msan or asan, in
one whose -gcflags turn off optimisation or inlining, or in a test
binary that a mutation run instrumented, because those builds allocate
differently from the one that ships. Go also reads an
int at the platform's width, which is 32 bits on a 32-bit platform. A
property over an int generates other values there. In Rust seven are
partial: the six
allocation ceilings need a counting allocator installed as the test
binary's global allocator, and no-task-leaks sees tasks on a runtime but
not a thread, because nothing in Rust's standard library enumerates
threads. Rust also declines the three relaxations. Its types keep an
absent container and an empty one apart, and its own equality already
treats NaN as unequal to itself. Its assertions compare through
PartialEq, which has no identity to compare by.
Go, Java, Kotlin and Rust declare a limit on the history's recorder because their threads run on more than one core. The recorder's counter synchronizes the clients that record into it. That synchronization can supply a memory barrier that the subject lacks. TypeScript searches the partitions of a history one at a time for any number of workers.
PHP is declared as a target language and the naming table carries no names for it yet, so adding it starts by filling that column.
The argument behind the design is in docs/rfc/0001-the-standardized-assertion-set.md.
MIT. See LICENSE.