(tasks) add tasks for harness
This commit is contained in:
@@ -0,0 +1,109 @@
|
||||
# T2.7 — P2 composition gate
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | P2 — Durability hard parts |
|
||||
| Size | L — over 3 days |
|
||||
| Status | Not started |
|
||||
| Flags | gate |
|
||||
| Spec | inlined below |
|
||||
| Blocks | P3 |
|
||||
|
||||
## Goal
|
||||
|
||||
Prove intents, effect classes, forks, checkpoints and schema evolution survive
|
||||
each other — not just crashes in isolation. T2.6 crashes a simple workload; this
|
||||
crashes a **compound** one.
|
||||
|
||||
**Phase gate criterion:** `kill -9` at every transition leaves recoverable state.
|
||||
|
||||
## Facts (inlined — no spec read needed)
|
||||
|
||||
- Each P2 task was verified against a simple run. The interactions are where the
|
||||
cost lives, and each is individually plausible-looking:
|
||||
- a **crash during a rewind** — a fork event committed, the new branch's LSN 0
|
||||
not yet written;
|
||||
- a **`Dispatched` intent on a branch that was subsequently forked** — recovery
|
||||
must resolve it on the branch that owns it, not the live one;
|
||||
- a **checkpoint written before a schema upgrade** — the checkpoint holds v1
|
||||
state, the log above it holds v2 records;
|
||||
- **reduction ordering** interacting with a fork — dead-branch text is
|
||||
reducible, the live branch's is not yet.
|
||||
- Recovery is defined jointly: legal states (T0.2), intents resolved by class
|
||||
(T2.2), byte-identical re-fold (T0.8), and exactly-once effects measured
|
||||
**externally**.
|
||||
- The discarded branch is evidence. A recovery path that "cleans up" a partial
|
||||
fork destroys the failure that motivated the rewind.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Build the compound workload: a run that retries, dispatches one intent per
|
||||
effect class, rewinds at least once, and crosses a checkpoint boundary.
|
||||
2. Extend T2.6's crash-point enumeration to cover the compound path, including
|
||||
points **inside** a rewind and **inside** recovery itself.
|
||||
3. Add the schema dimension: run half the seeds with a v1→v2 upgrade applied
|
||||
between crash and restart.
|
||||
4. Assert the four recovery properties jointly after every restart.
|
||||
5. Run recovery **twice** per seed — recovery must itself be idempotent, since a
|
||||
crash during recovery is an ordinary event.
|
||||
6. Make this the required CI job gating P3.
|
||||
|
||||
## Acceptance
|
||||
|
||||
- 500 seeded crash points on the **compound** workload: zero unrecoverable
|
||||
states, zero double-applied effects, byte-identical re-fold, all intents
|
||||
resolved by class.
|
||||
- Recovery run twice produces the same state as recovery run once.
|
||||
- Half the seeds pass with a schema upgrade interposed.
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** T2.6's child-process `kill -9` rig, the external side-effect ledger,
|
||||
the three effect-class tools, and a seed printed on every run.
|
||||
|
||||
**Integration test** — `tests/it_p2_composition.rs`:
|
||||
1. **Crash inside a rewind:** abort between the fork event commit and the new
|
||||
branch's first record. On restart assert the parent branch is **byte-identical
|
||||
to its pre-rewind snapshot** and the run is in a legal state — either the fork
|
||||
completed or it did not, never half.
|
||||
2. **Intent on a forked branch:** dispatch an `Unsafe` intent on branch 0, fork
|
||||
to branch 1, crash. Assert recovery marks the **branch-0** attempt
|
||||
`Indeterminate` and does not touch branch 1.
|
||||
3. **Checkpoint across schema versions:** write a checkpoint at v1, append v2
|
||||
records, crash, restart under the v2 binary. Assert rebuild from that
|
||||
checkpoint equals a rebuild from LSN 0.
|
||||
4. **Recovery idempotence:** run recovery, snapshot state, run recovery again,
|
||||
assert byte equality and assert the external ledger count is unchanged.
|
||||
5. **Effect exactness under compounding:** across all seeds, assert the ledger
|
||||
holds exactly one entry per intended effect — including effects issued before
|
||||
a fork.
|
||||
6. **Crash during recovery:** abort partway through the recovery pass itself;
|
||||
restart; assert convergence.
|
||||
7. **Regression:** re-run T2.1–T2.6 plus P0/P1 gates in the same job.
|
||||
|
||||
**Command:**
|
||||
`cargo test -p durability --features test-hooks --test it_p2_composition -- --nocapture`
|
||||
|
||||
**False pass:**
|
||||
- Running the compound matrix on the simple workload. T2.6 already covers that;
|
||||
this gate exists for the interactions, and every one of them is green in
|
||||
isolation.
|
||||
- Step 4 omitted. A recovery pass that is correct once and destructive twice
|
||||
passes every test in P2 and fails the first time a machine crashes while
|
||||
recovering — which is exactly when it runs.
|
||||
- Step 1 asserting only that the run recovered. Recovering by truncating the
|
||||
parent branch also recovers, and destroys the evidence.
|
||||
- Measuring effect counts in the framework's log rather than the external ledger.
|
||||
|
||||
## Traps
|
||||
|
||||
- Treating a crash during recovery as out of scope. It is the ordinary case in a
|
||||
crash loop, and a crash loop is what a nonzero `Indeterminate` rate means.
|
||||
- Letting the compound workload drift out of sync with new transitions. Assert
|
||||
the enumerated crash-point count against a constant, as T2.6 does.
|
||||
|
||||
---
|
||||
|
||||
Background (not required to do this task):
|
||||
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §8.4, §8.5, §8.6, §8.7 ·
|
||||
[rust-agentic-task.md](../../../rust-agentic-task.md)
|
||||
Reference in New Issue
Block a user