Repository navigation
Validate C++ disaster recovery traces with Lean - #8282
Merged
Merged
Conversation
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 4, 2026 18:47
a443db8 to
89872d7
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 4, 2026 22:18
89872d7 to
7721ba2
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 4, 2026 23:31
7721ba2 to
b49f888
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 7, 2026 20:09
b49f888 to
d772ae2
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 8, 2026 09:58
1eaf0c0 to
a5d2879
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
2 times, most recently
from
September 8, 2026 21:00
eedf71b to
5298374
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 9, 2026 14:15
5298374 to
a66bd00
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 9, 2026 14:16
a66bd00 to
62efec2
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 9, 2026 17:15
62efec2 to
af3ec5a
Compare
Amaury Chamayou (achamayou)
force-pushed
the
achamayou-fluffy-parakeet
branch
from
September 9, 2026 17:31
af3ec5a to
bca9e9f
Compare
Amaury Chamayou (achamayou)
added a commit
that referenced
this pull request
Sep 11, 2026
Keep Ubuntu 26.04 runners and pull-request-only Lean validation while retaining the trace-validator job and recovery path filters. Move the recovery restart fix to the upcoming 7.0.16 release notes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Amaury Chamayou (achamayou)
added a commit
that referenced
this pull request
Sep 11, 2026
Include the JavaScript wrapped-value ownership fix merged while recovery-trace validation was running. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Heidi Howard (heidihoward)
approved these changes
Oct 9, 2026
Add a Lean replayer that checks recovery-decision-protocol traces from -DCCF_RECOVERY_TRACE=ON builds against the disaster recovery model. DisasterRecovery/Replay/Records.lean extracts the RDP_TRACE records from node logs and validates their fields. DisasterRecovery/Replay/Reduction.lean orders each node's records by sequence, links receives to their sends, checks retry batches, infers which handler executions took effect, and linearises them into model actions and state observations. DisasterRecovery/Replay.lean replays those through Model.transitionSystem. The disaster-recovery-replay exe waits for the logs to record a complete scenario, then reduces and replays them. The SNP Genoa job builds the replayer and CCF with tracing, and tests/infra/recovery_trace.py replays the quorum, failover and multiple-timeout scenarios of tests/e2e_operations.py. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Tracing is enabled by the CCF_RECOVERY_TRACE environment variable, not a CMake option, so drop the unused -DCCF_RECOVERY_TRACE=ON and set the variable for the SNP Genoa tests instead. The recovery decision protocol scenarios then only run traced in that job. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Records now carry the versions of the sm_state and timeout_sm_state values each execution read, the sm_state version each retry read, and timeout request ids. With them, CCF's optimistic concurrency control determines the order in which each node committed its executions, so replace the per-node search over accounts of which executions took effect: - only the last execution of a message or timeout request can commit; - the version pairs that executions read form one chain, and the next pair shows the one execution that wrote it; - executions that read the gossips or votes follow the insert of the set they read, and inserts off that chain did not commit; - each retry runs where its sm_state version is first read. Any remaining ambiguity fails rather than being resolved by preference. The scenario accepts participants that end in Opening, and is checked after the replay, so that its failures are reported as such. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Two executions of the same message, or of timeout requests that found the same state, can read the same versions and record the same writes. One commits, and the other's re-execution can be rejected before advance(), for example because the gossips already chose a node, so it is not traced again. Both then remain final writers of the pair, and the reduction rejected a correct trace. Such writers have the same model action and recorded fields, and the model replays either alike, since it delivers any queued copy of an envelope. So treat them as one, and replay the lowest sequence. Writers that differ in any recorded field still fail. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
- Reformat and delint tests/infra/recovery_trace_mutations.py: run black with the current unpinned uvx black (26.5.1), which reflows a couple of long lines that an older cached black had left alone, and drop the shebang, which no other tests/infra module has and which ruff's EXE001 (not skipped outside WSL) flags since the file is not executable. - Reduction.lean: when no later retry shows the newest pair's writer's exact sm_state version, require its committed version to exceed both versions of that pair (the pair the writer read), not just the highest version known from earlier committed records. A version between the pair's two versions is impossible, since the writer read that pair, but previously passed. - Add a targeted committed_version_within_newest_pair mutant that lowers such a version into the newest pair and expects the replay to fail; update the README's description of this bound. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Replace the separate CURATED table and its 30 single-purpose mutant functions with a merged, declarative MUTANTS table: each row states a predicate and an edit, interpreted by generic select/edit combinators (simple, fallback, record_mutant) and a handful of tiny field-edit factories (set/del/bump/flip/relabel/ copy/dyn). Bespoke functions are kept only for structural mutants (insert/delete/swap records, consistent TxID relabeling, batch deletion, argument changes, whole-log transforms, and the committed-version-within-newest-pair check). Also simplify the CLI: drop --work-dir/--results-json/--keep-work (unused outside this file) in favour of a temporary directory, and print only unexpected outcomes plus the summary lines. Verified against the current fixtures and replayer: the refactored harness reproduces the pre-refactor harness's per-mutant results exactly (3 baselines, 130 targeted, 520 sweep, 3 allowlisted; no mutant merged or renamed). 1367 -> 1239 lines. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Cut about 75 more lines (1239 -> 1164) by reducing the number of top-level functions and their separator overhead, not just line wrapping: - Add `dyns` (a single-field edit needing full trace state) and `combine` (compose several independent field edits into one) factories, and use them to replace six standalone `_edit_*` functions with inline expressions in the MUTANTS table. - Merge `m_txid_raise_consistent`/`m_txid_lower_consistent` into one `_txid_consistent(state, raise_it)` function. - Replace the `m_delete_middle_unreceived_batch`/ `m_delete_tail_unreceived_batch` wrappers with a `tail=` keyword on `_delete_unreceived_batch`, called directly from the table. - Condense `run_mutant`, `build_jobs` and `main` (the runner, CLI and report) onto fewer lines without changing their behaviour. - Trim docstrings/comments that restated the code. Verified byte-for-byte equivalent outcomes against the previous harness (ccf4b88) on the fixtures: 653/653 results match by scenario/group/name/expected/outcome-class (3 baselines, 130 targeted, 520 sweep). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…ed tables - Drop hand-written mutation descriptions (note(), detail plumbing, and per-mutant description strings) in favour of a single describe(base, mutated) helper computed only when reporting an unexpected outcome; it diffs the mutated scenario's records against the base and reports changed fields, insertions, deletions, and participants/open_kind/order changes. - Replace the single flat MUTANTS table with six group-specific tables (DECISION, CAUSALITY, COMMIT_ORDER, FORMAT, ARGS, BENIGN) of one line per mutant, tied together by a GROUPS dict; hoist any predicate/edit that would not fit on one row to a short named helper above its table. - Replace record_mutant/simple/fallback/find_any with a single one(...) combinator (predicates in order, edit last). - Switch the runner to a Result NamedTuple instead of dicts; fold run_mutant/build_jobs/_bad/allow_sweep/classify/main into a leaner shape that prints two summary lines, the allowlisted classes, and each unexpected outcome's describe() diff plus the replayer's last line. - Tighten several bespoke structural mutants now that they no longer need to build description strings. - Trim the module docstring and drop section-banner comments that only restated the code. Verified equivalent to 69afba4's behaviour: replaying the fixtures with the current disaster-recovery-replay gives 653/653 identical per-mutant results (scenario, group, name, expected, outcome class): 3 baselines, 130 targeted mutants, 520 sweep mutants, the same 3 allowlisted sweep classes. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#8366 now logs recovery-decision-protocol handler executions from each endpoint's locally committed function, with the TxID seqno CCF reported for the transaction and a wrote flag, instead of the sm_state and timeout_sm_state versions a handler read. Only executions whose transaction committed are logged, so the replayer no longer needs to infer which candidate commits from several readers of one pair: each node's records already carry their own commit order. Update the replayer accordingly: - Replay/Records.lean: handler records carry version and wrote instead of pre_version/pre_timeout_version; the global hook's commit record becomes start (version and expected_locations only, once per node); timeouts no longer have a caused_by. - Replay/Reduction.lean: order each node's executions by version, with the write at a version before its readers, and run retries right after the sm_state write they read. The rolled-back, segment, writer, set-chain and committed-record rules are gone, since the trace now shows the commit order directly. - replay/README.md: rewrite the Records and Rules sections and the explanation of why this is the commit order for the new fields and rules, and fix stale references to the dropped allowlist, rules and record kinds in the Mutation test section. - tests/infra/recovery_trace_mutations.py: replace the commit-order mutants for the new rules, and replace the static sweep allowlist with a computed commit_order/same_order check that accepts a sweep pass on version or wrote only when it admits no commit order the original trace does not admit. Fix the same stale allowlist wording in the module docstring. - .github/workflows/README.md: fix the same stale allowlist wording. - Regenerate the fixtures from local virtual-platform recovery traces in the new format (refreshed from a real SNP run in a follow-up commit). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Regenerate the quorum, timeout and multiple-timeout fixtures from the ACI SNP Genoa job's logs-caci-snp-genoa artifact on run 36849042909 (https://github.com/microsoft/CCF/actions/runs/36849042909), the PR's own CI run on the new commit-time replayer format, replacing the local virtual-platform traces. Mutation test against the refreshed fixtures: Curated mutants: 127/127 matched expectation. Sweep mutants: 389/425 caught, 36 allowlisted harmless pass(es). Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Remove two checks from the trace parser that re-encoded protocol rules the model already enforces during replay: that every advance() in Joining requests a restart, and that gossips are only accepted in Gossiping (the model rejects gossips once a node has chosen). Also drop the now-redundant && restart from the IAmOpen check, which still requires pre == .joining && post == .joining && chosen == some source to describe its own Joining write. Also fix a stale 'still builds' in the workflows README now that the Genoa SNP job always builds the replayer. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
#8366 no longer adds trace ids to protocol messages, so receive records carry no caused_by. The reduction now lets each receive take an earlier send of its message from its source that no other receive has taken, as the model's network delivers any queued copy of an envelope. This removes the causes rule and the caused_by parsing. The mutation test no longer sweeps version and wrote, whose effect on commit order the targeted commit-order mutants cover, so every sweep mutant must now fail and the order-preservation allowlist is gone. Its causality mutants are reworked for content matching, and the fixtures drop caused_by. The replay README's fixture refresh is now a short shell loop that produces the same files. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Regenerate the quorum, timeout and multiple-timeout fixtures from the ACI SNP Genoa job's logs-caci-snp-genoa artifact on run 36868572170 (https://github.com/microsoft/CCF/actions/runs/36868572170), the PR's own CI run on the content-matching replayer, using the replay README's grep refresh loop. Mutation test against the refreshed fixtures: Curated mutants: 119/119 matched expectation. Sweep mutants: 343/343 caught. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Build disaster-recovery-replay from its own Lake package in lean/disaster-recovery/replay. The replayer imports no Mathlib module, so the package has no dependencies, and the SNP and Lean CI jobs no longer clone Mathlib to build it. Parse each record kind directly, reading send records' new message and target fields, and drop two unused definitions. Compare model observations with the model types' Repr instances instead of hand-written JSON renderers. Drop five targeted mutants whose mutated traces are identical to sweep mutants, and convert the fixtures' send records to the new fields. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Regenerate the quorum, timeout and multiple-timeout fixtures from the ACI SNP Genoa job's logs-caci-snp-genoa artifact on run 36982801711 (https://github.com/microsoft/CCF/actions/runs/36982801711), the PR's own CI run on the simplified standalone replayer, using the replay README's grep refresh loop. Mutation test against the refreshed fixtures: Curated mutants: 107/107 matched expectation. Sweep mutants: 350/350 caught. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
advance() no longer logs the gossips or votes maps it evaluates, so the replayer no longer parses or compares them: the parser drops the fields from Execution and the presence rules, and Replay.lean and Reduction.lean drop the corresponding state checks. The mutation harness drops the mutants that only exercised the snapshot comparison (open_without_quorum, set_chain_off), rewrites premature_voting and swap_gossip_writes to fail for a reason the replayer still checks without the snapshots, and drops the gossips and votes sweep perturbations along with the now-unused helper they used. Relabeling a gossip's source to a node that also sent its txid is an equivalent mutant now that the snapshot is gone, since every node here gossips the same txid: the sweep's source relabel now only picks a node that never sent it, and skips the field when none exists, so it cannot introduce a new sweep pass. Fixtures: gossips and votes stripped textually from every recorded line, since the parser now rejects the unknown fields. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
advance() no longer logs whether a handler or timeout execution wrote, so the replayer derives it instead of parsing it: a vote or IAmOpen always writes (a vote is a Set::insert, and an IAmOpen puts Joining and the chosen node before advance() reads them), and any other write the replayer can observe changes a phase. A write that changes no phase, such as storing a new gossip or node info, is not observable here, so treating it as a read changes nothing. The mutation harness derives the same rule for its own notion of which records wrote, drops unwrite_read_version (its field is gone), and drops the now-redundant write flag it set in m_receive_before_send. duplicate_write_version picks from the same derived writes, so it now forces two phase-changing writes to share a version rather than two store-only gossip writes, and still fails for the right reason. swap_gossip_writes relied on a node writing a store-only gossip before its advancing one, which the derived rule no longer records as a write, so it never fired. Renamed to swap_writes and broadened to any two of a node's consecutive writes: the sweep never perturbs version, so this is the only test that swapped writes are caught. Fixtures: wrote stripped textually from every handler and timeout record, since the parser now rejects the unknown field. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
trace_safely now logs the line of its caller instead of a label, so a failed trace reads "Failed to trace recovery-decision-protocol at line N". Make the failed_to_trace_line mutant inject that form. The line numbers are illustrative: the replayer only looks for the unchanged prefix. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
With the tracing now merged, remove what earlier trace formats needed, and make the comments say what the code does now. - Records: require `chosen` only on the move to Voting, where nothing else in the execution shows the choice; the other presence rules repeated checks that the replay makes. Only an IAmOpen's `pre` goes uncompared, so check only that. Drop an unused BEq derive, and an empty-log check that reduce already makes. - Reduction: drop the rule of execution items, which was always commit-order, and the pre-state chosen check, which the post-state check covers. - Replay: name every condition of Config.isValid in its error. - Mutation harness: the fixtures hold only canonical RDP_TRACE lines, so drop the handling of other lines, and the strip_non_trace_lines and read_only_before_its_write mutants, which no longer changed anything. Add chosen_not_recorded, the only test of the remaining presence rule. - Docs: describe the CI setup as it is now, and drop a README paragraph that repeated Replay.lean's docstrings. lean.yml's fixtures path was already covered by lean/**. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Check in the mutants that recovery_trace_mutations.py generated, as 417 diffs against the fixtures under each scenario's mutants/fail and mutants/pass, and replace the generator with check-fixtures.sh, which applies each diff and checks that the replayer rejects or accepts it. Mutants that changed the replayer's arguments now edit scenario.json. Those that rewrote whole files, by shuffling lines or reversing keys, or that reordered the log files, are dropped. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
As the Lean proofs are, so that GitHub collapses them in diffs. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
They are traces that the replayer must reject or accept, so name them for that rather than for how they were made: each scenario's mutants/fail and mutants/pass become invalid/ and valid/, and the sweep.* traces, which differ from the recorded ones in one field, become field.*. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
trace.schema.json gives each record kind exactly its fields, and each field the values that src/node/recovery_decision_protocol.cpp writes for it: non-empty location names, non-negative counters, versions from 1, the phase, open kind and message names, TxIDs as view.seqno, a txid only on gossips, a source only on receives, and restart only as true. It has none of the protocol's rules, which the replay checks. check-schema.py checks the records of node logs against it, and the Lean workflow runs it on the fixtures. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The replayer only used it to check, after a successful replay, that the run ended with the kind of opening the e2e test expected. That is a test expectation, not a condition for a valid trace: the model decides the open kind, and the replay checks each recorded one against it. The quorum and timeout tests already assert the open kind from the ledger. Drop --open-kind and the open_kind of each scenario.json. Remove the args.wrong_open_kind_arg invalid traces, and rewrite the args.wrong_participants_arg ones for the shorter scenario.json. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Each branch now fixes the record's kind with const, or enum for the vote and IAmOpen receives, which states the tagged union directly to readers and tools. The schema accepts and rejects the same records as before. JSON Schema has no discriminator, so validators report a record that matches no branch as matching none. check-schema.py reports the errors of the branch for the record's kind instead, so its messages still name the failing field. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The modules in DisasterRecovery/Replay.lean and DisasterRecovery/Replay/ and the Lake package in replay/ had the same name, and the READMEs referred to the modules by paths relative to an unstated directory. - Rename the modules to DisasterRecovery/TraceValidation.lean and DisasterRecovery/TraceValidation/, in namespace DisasterRecovery.TraceValidation. - Rename the package directory to replayer/, and its ReplayMain.lean to Main.lean. - Give paths from lean/disaster-recovery in the READMEs, list the modules in lean/AGENT.md's layout table, and label the delivery and scenario rules in Reduction.lean with the names the README gives them. - Drop the pointer to the deleted mutation generator from the fixture refresh instructions. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Checks that CCF nodes run the recovery decision protocol as the Lean model in
lean/disaster-recoveryspecifies, by replaying through the model the traces that nodes log whenCCF_RECOVERY_TRACEis set. It changes no CCF C++ code and does not change the model.Implementation summary
disaster-recovery-replay, a Lean executable, checks the node logs of one e2e recovery scenario against the model:%%{init: {"theme": "base", "themeVariables": {"background": "#ffffff", "fontFamily": "BlinkMacSystemFont, Segoe UI, Noto Sans, Helvetica, Arial", "primaryColor": "#f6f8fa", "primaryTextColor": "#1f2328", "primaryBorderColor": "#d0d7de", "secondaryColor": "#f6f8fa", "tertiaryColor": "#ffffff", "mainBkg": "#f6f8fa", "nodeBorder": "#d0d7de", "textColor": "#1f2328", "titleColor": "#1f2328", "lineColor": "#656d76", "defaultLinkColor": "#656d76", "clusterBkg": "#ffffff", "clusterBorder": "#d0d7de", "edgeLabelBackground": "#ffffff"}}}%% flowchart TD snp["<b>SNP e2e tests</b><br/>traced with CCF_RECOVERY_TRACE=1"] lean["<b>Lean workflow</b><br/>checked-in valid and invalid traces"] snp -->|"node logs"| extract lean -->|"fixtures"| extract subgraph replayer["disaster-recovery-replay"] direction TB extract["<b>Extract</b> (Records.lean)<br/>well-formed records, no trace failures"] config["<b>Config</b><br/>every start record has the same expected_locations"] subgraph order["Reduction.lean: commit order per node"] direction TB startcheck["<b>Start</b><br/>one start per node, no execution before it"] commitorder["<b>Commit order</b><br/>CCF's commit versions, each write before its readers"] retry["<b>Retry</b><br/>right after the sm_state write it read"] startcheck --> commitorder --> retry end interleave["<b>Interleave</b><br/>each receive takes an earlier send of its message"] replay["<b>Replay</b> (TraceValidation.lean)<br/>actions enabled, observations match"] scenario{"<b>Scenario</b><br/>one opens, others join"} extract --> config --> order --> interleave --> replay --> scenario end wait["<b>Wait</b><br/>up to --wait-ms, 20 s by default"] ok(["<b>Pass</b>"]) bad(["<b>Fail</b><br/>naming the log line and rule"]) order & interleave -.->|"incomplete"| wait wait -.-> extract scenario -->|"yes"| ok scenario -->|"no"| bad replayer -->|"check fails"| bad classDef input fill:#e7f1ff,stroke:#0d6efd,color:#084298 classDef good fill:#d1e7dd,stroke:#198754,color:#0f5132 classDef failure fill:#f8d7da,stroke:#dc3545,color:#842029 class snp,lean input class ok good class bad failureIn
lean/disaster-recovery:DisasterRecovery/TraceValidation/Records.leanparses theRDP_TRACErecords and checks each one's fields. AnyFailed to traceline fails the replay.DisasterRecovery/TraceValidation/Reduction.leanorders each node's executions by the version that CCF reported, with each write before the reads at its version, and each retry right after thesm_statewrite it read. A vote or IAmOpen always writes, and so does any execution that changes a phase. It then interleaves the nodes so that each receive takes an earlier send of the same message from its source.DisasterRecovery/TraceValidation.leanreplays the result throughModel.transitionSystem: each action must be enabled, and each recorded state and message, and each notification that the recorded writes imply, must match the model.replayer/Main.leanwaits up to--wait-ms(20 s by default) until every participant has opened or joined.replayer/trace.schema.jsonis the JSON Schema of the records, format only, andreplayer/check-schema.pychecks node logs against it.replayer/lakefile.tomlbuilds the replayer in its own Lake package, without Mathlib.In CI, the ACI SNP Genoa job, whose tests run traced, replays the quorum, failover and multiple-timeout scenarios of
tests/e2e_operations.pyafter each one runs. The Lean workflow checks recorded SNP traces of those scenarios against the schema.replayer/check-fixtures.shthen replays them and 414 traces stored as diffs against them: the recorded traces and 3 valid ones underfixtures/*/valid/must pass, and 411 invalid ones underfixtures/*/invalid/must fail.To review: the model checks are all in
TraceValidation.lean, andReduction.leanis the only code that orders records.replayer/README.mddocuments the records and the rules. The fixtures are test data, and.gitattributesmarks their diffs as generated.Safety and compatibility
No runtime impact: this PR changes no CCF code, and only adds replay steps to CI. The fixtures only contain test protocol metadata: location names, TxIDs, phases, KV versions and trace counters.