4.7 KiB
4.7 KiB
T2.6 — Crash matrix
| Field | Value |
|---|---|
| Phase | P2 — Durability hard parts |
| Size | L — over 3 days |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
Goal
Systematic kill -9 at every kernel transition, in a loop, verifying
recoverability after each. This is the phase gate.
Facts (inlined — no spec read needed)
- P2 is the phase most likely to be skipped and most expensive to retrofit. The matrix is what proves the three preceding tasks actually hold together.
- What "recoverable" means concretely: on restart, every attempt resolves to a
legal state per T0.2's table; every
Dispatchedintent resolves per T2.2's class table; the fold from LSN 0 reproduces the state exactly (T0.8). - "Zero double-applied effects" is the second, independent assertion. An idempotent tool retried after a crash must leave one effect, not two.
- Lessons this suite exists to catch:
- A green suite says nothing about coverage. A prior idempotency feature generated its own keys and checked for duplicates among them — unreachable for a whole phase, with tests asserting the count was zero.
- A test run that prints nothing cannot distinguish slow from hung. Per-test progress and per-test timeouts from the first commit.
Steps
- Enumerate the crash points: every kernel transition emitted by T1.3/T1.4, plus T2.1's intra-intent windows, plus each of T0.6's four table writes.
- Build the harness as a real process kill (
kill -9on a child), not a panic-and-catch. Shadow paging behaviour under a hard kill is part of what is being tested. - Drive a workload with a scripted stub model (T1.2) and the three effect-class tools (T2.2), so both durability and effect recovery are exercised.
- Seed the crash-point selection from a recorded seed. Print the seed on failure — an unreproducible crash matrix failure is not a finding.
- After each restart assert: state is legal, intents resolve, the cold re-fold matches, effects applied exactly once.
- Emit per-test progress and enforce a per-test timeout so a hang is distinguishable from a slow case.
- Run 500 randomized crash points in CI; keep the exhaustive enumeration as a longer nightly job.
Acceptance
- 500 randomized crash points, zero unrecoverable states, zero double-applied effects.
Phase gate
P2 closes on a green crash matrix.
Verify
Harness: the P2 capstone. Child process under kill -9; the external
side-effect ledger from T2.2; the three effect-class tools; the stub model; a
recorded seed printed on every run.
Integration test — tests/it_crash_matrix.rs:
- Enumerate crash points: every kernel transition, T2.1's two intent windows, T0.6's four table writes. Assert the enumerated count matches an expected constant, so a newly added transition without a crash point fails the test.
- For each of 500 seeded random points: run the workload,
kill -9at the point, restart, then assert all four properties:- every attempt is in a legal state per T0.2's table;
- every
Dispatchedintent resolved per T2.2's class table; assert_refold_identicalpasses;- the external ledger shows exactly one effect per intended effect.
- Print the seed and the crash point on failure; a failure that cannot be replayed is not a finding.
- Per-test progress output and a per-test timeout, so a hang is distinguishable from a slow case.
- Keep the exhaustive enumeration as a nightly job; the 500-point random sample runs in CI.
Command:
cargo test -p durability --features test-hooks --test it_crash_matrix -- --nocapture
False pass:
panic!instead ofkill -9. Destructors run, buffers flush, and the whole matrix passes against an implementation with no durability at all.- Counting effects only among those the harness itself generated. That is the documented prior failure: an idempotency feature generated its own keys and checked for duplicates among them — unreachable for a whole phase, with tests asserting the count was zero. Count against the external ledger.
- Step 1 omitted, so new transitions silently escape the matrix.
- Asserting recoverability as "the process restarted without error". Restarting cleanly into wrong state is the failure being hunted.
Traps
- Simulating a crash with
panic!and unwinding. Destructors run, which is exactly what a real crash does not do. - Counting only effects the harness itself created — the coverage failure quoted above, repeated.
Background (not required to do this task): rust-agentic-sys.md §8, §19 · rust-agentic-task.md