4.9 KiB
4.9 KiB
T2.7 — P2 composition gate
| Field | Value |
|---|---|
| Phase | P2 — Durability hard parts |
| Size | L — over 3 days |
| Status | Not started |
| Flags | gate |
| Spec | inlined below |
| Blocks | P3 |
Goal
Prove intents, effect classes, forks, checkpoints and schema evolution survive each other — not just crashes in isolation. T2.6 crashes a simple workload; this crashes a compound one.
Phase gate criterion: kill -9 at every transition leaves recoverable state.
Facts (inlined — no spec read needed)
- Each P2 task was verified against a simple run. The interactions are where the
cost lives, and each is individually plausible-looking:
- a crash during a rewind — a fork event committed, the new branch's LSN 0 not yet written;
- a
Dispatchedintent on a branch that was subsequently forked — recovery must resolve it on the branch that owns it, not the live one; - a checkpoint written before a schema upgrade — the checkpoint holds v1 state, the log above it holds v2 records;
- reduction ordering interacting with a fork — dead-branch text is reducible, the live branch's is not yet.
- Recovery is defined jointly: legal states (T0.2), intents resolved by class (T2.2), byte-identical re-fold (T0.8), and exactly-once effects measured externally.
- The discarded branch is evidence. A recovery path that "cleans up" a partial fork destroys the failure that motivated the rewind.
Steps
- Build the compound workload: a run that retries, dispatches one intent per effect class, rewinds at least once, and crosses a checkpoint boundary.
- Extend T2.6's crash-point enumeration to cover the compound path, including points inside a rewind and inside recovery itself.
- Add the schema dimension: run half the seeds with a v1→v2 upgrade applied between crash and restart.
- Assert the four recovery properties jointly after every restart.
- Run recovery twice per seed — recovery must itself be idempotent, since a crash during recovery is an ordinary event.
- Make this the required CI job gating P3.
Acceptance
- 500 seeded crash points on the compound workload: zero unrecoverable states, zero double-applied effects, byte-identical re-fold, all intents resolved by class.
- Recovery run twice produces the same state as recovery run once.
- Half the seeds pass with a schema upgrade interposed.
Verify
Harness: T2.6's child-process kill -9 rig, the external side-effect ledger,
the three effect-class tools, and a seed printed on every run.
Integration test — tests/it_p2_composition.rs:
- Crash inside a rewind: abort between the fork event commit and the new branch's first record. On restart assert the parent branch is byte-identical to its pre-rewind snapshot and the run is in a legal state — either the fork completed or it did not, never half.
- Intent on a forked branch: dispatch an
Unsafeintent on branch 0, fork to branch 1, crash. Assert recovery marks the branch-0 attemptIndeterminateand does not touch branch 1. - Checkpoint across schema versions: write a checkpoint at v1, append v2 records, crash, restart under the v2 binary. Assert rebuild from that checkpoint equals a rebuild from LSN 0.
- Recovery idempotence: run recovery, snapshot state, run recovery again, assert byte equality and assert the external ledger count is unchanged.
- Effect exactness under compounding: across all seeds, assert the ledger holds exactly one entry per intended effect — including effects issued before a fork.
- Crash during recovery: abort partway through the recovery pass itself; restart; assert convergence.
- Regression: re-run T2.1–T2.6 plus P0/P1 gates in the same job.
Command:
cargo test -p durability --features test-hooks --test it_p2_composition -- --nocapture
False pass:
- Running the compound matrix on the simple workload. T2.6 already covers that; this gate exists for the interactions, and every one of them is green in isolation.
- Step 4 omitted. A recovery pass that is correct once and destructive twice passes every test in P2 and fails the first time a machine crashes while recovering — which is exactly when it runs.
- Step 1 asserting only that the run recovered. Recovering by truncating the parent branch also recovers, and destroys the evidence.
- Measuring effect counts in the framework's log rather than the external ledger.
Traps
- Treating a crash during recovery as out of scope. It is the ordinary case in a
crash loop, and a crash loop is what a nonzero
Indeterminaterate means. - Letting the compound workload drift out of sync with new transitions. Assert the enumerated crash-point count against a constant, as T2.6 does.
Background (not required to do this task): rust-agentic-sys.md §8.4, §8.5, §8.6, §8.7 · rust-agentic-task.md