Files
poimen/tasks/T2.6-crash-matrix.md
T

4.7 KiB

T2.6 — Crash matrix

Field Value
Phase P2 — Durability hard parts
Size L — over 3 days
Status Not started
Flags
Spec inlined below
Blocks

Goal

Systematic kill -9 at every kernel transition, in a loop, verifying recoverability after each. This is the phase gate.

Facts (inlined — no spec read needed)

  • P2 is the phase most likely to be skipped and most expensive to retrofit. The matrix is what proves the three preceding tasks actually hold together.
  • What "recoverable" means concretely: on restart, every attempt resolves to a legal state per T0.2's table; every Dispatched intent resolves per T2.2's class table; the fold from LSN 0 reproduces the state exactly (T0.8).
  • "Zero double-applied effects" is the second, independent assertion. An idempotent tool retried after a crash must leave one effect, not two.
  • Lessons this suite exists to catch:
    • A green suite says nothing about coverage. A prior idempotency feature generated its own keys and checked for duplicates among them — unreachable for a whole phase, with tests asserting the count was zero.
    • A test run that prints nothing cannot distinguish slow from hung. Per-test progress and per-test timeouts from the first commit.

Steps

  1. Enumerate the crash points: every kernel transition emitted by T1.3/T1.4, plus T2.1's intra-intent windows, plus each of T0.6's four table writes.
  2. Build the harness as a real process kill (kill -9 on a child), not a panic-and-catch. Shadow paging behaviour under a hard kill is part of what is being tested.
  3. Drive a workload with a scripted stub model (T1.2) and the three effect-class tools (T2.2), so both durability and effect recovery are exercised.
  4. Seed the crash-point selection from a recorded seed. Print the seed on failure — an unreproducible crash matrix failure is not a finding.
  5. After each restart assert: state is legal, intents resolve, the cold re-fold matches, effects applied exactly once.
  6. Emit per-test progress and enforce a per-test timeout so a hang is distinguishable from a slow case.
  7. Run 500 randomized crash points in CI; keep the exhaustive enumeration as a longer nightly job.

Acceptance

  • 500 randomized crash points, zero unrecoverable states, zero double-applied effects.

Phase gate

P2 closes on a green crash matrix.

Verify

Harness: the P2 capstone. Child process under kill -9; the external side-effect ledger from T2.2; the three effect-class tools; the stub model; a recorded seed printed on every run.

Integration testtests/it_crash_matrix.rs:

  1. Enumerate crash points: every kernel transition, T2.1's two intent windows, T0.6's four table writes. Assert the enumerated count matches an expected constant, so a newly added transition without a crash point fails the test.
  2. For each of 500 seeded random points: run the workload, kill -9 at the point, restart, then assert all four properties:
    • every attempt is in a legal state per T0.2's table;
    • every Dispatched intent resolved per T2.2's class table;
    • assert_refold_identical passes;
    • the external ledger shows exactly one effect per intended effect.
  3. Print the seed and the crash point on failure; a failure that cannot be replayed is not a finding.
  4. Per-test progress output and a per-test timeout, so a hang is distinguishable from a slow case.
  5. Keep the exhaustive enumeration as a nightly job; the 500-point random sample runs in CI.

Command: cargo test -p durability --features test-hooks --test it_crash_matrix -- --nocapture

False pass:

  • panic! instead of kill -9. Destructors run, buffers flush, and the whole matrix passes against an implementation with no durability at all.
  • Counting effects only among those the harness itself generated. That is the documented prior failure: an idempotency feature generated its own keys and checked for duplicates among them — unreachable for a whole phase, with tests asserting the count was zero. Count against the external ledger.
  • Step 1 omitted, so new transitions silently escape the matrix.
  • Asserting recoverability as "the process restarted without error". Restarting cleanly into wrong state is the failure being hunted.

Traps

  • Simulating a crash with panic! and unwinding. Destructors run, which is exactly what a real crash does not do.
  • Counting only effects the harness itself created — the coverage failure quoted above, repeated.

Background (not required to do this task): rust-agentic-sys.md §8, §19 · rust-agentic-task.md