Files

124 lines
5.5 KiB
Markdown
Raw Permalink Normal View History

2026-08-17 23:05:20 -07:00
# T2.1 — Write-ahead intent
| Field | Value |
|---|---|
| Phase | P2 — Durability hard parts |
| Size | L — over 3 days |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | T2.2, T6.4 |
## Goal
Three intent records around every external effect, so a crash can be classified
rather than guessed at.
## Facts (inlined — no spec read needed)
```
1. append Intent{Pending} ──► commit + fsync
◄── window A: crash here, the call was never issued
2. append Intent{Dispatched} ──► commit + fsync
3. perform the call
◄── window B: crash here, the call may have landed
4. append outcome, Intent{Committed} ──► commit + fsync
```
- **Three records, not two.** Two cannot separate the windows: a crash before the
call and a crash after it both leave a lone `Pending` with no outcome, which
makes every interrupted effect maximally suspicious and pushes recoverable work
into `Indeterminate`. The second fsync buys the distinction, is paid only on
steps with external effects, and is small next to the call it guards.
- Appending after the fact records history but does not make a failed step
resumable — the dangerous window is *before* the record exists.
- Restart classification, from the last committed intent state:
| Last state | Meaning | Resolution |
|---|---|---|
| `Pending` | the call was never issued | retry freely, whatever the effect class |
| `Dispatched` | the call *may* have been issued | by effect class (T2.2) |
| `Committed` | outcome already recorded | nothing to do |
- Intents always commit with `Durability::Immediate`. **Batching them defeats
their only purpose.**
- Tool calls need this more than model calls. A model call is a metered read; a
tool call writes files, pushes commits, and touches the world.
- The intent record is also what decides `Cancelled` vs `Indeterminate` on a
cancelled attempt (T0.2, T1.1).
## Steps
1. Define `IntentState { Pending, Dispatched, Committed }` and an `IntentId`
scoped to `(BranchKey, AttemptNo, step)`. Include the idempotency key or
request id the effect class will need at recovery (T2.2).
2. Wrap every external effect in the four-step sequence above. Each append is its
own `EventLog::commit` with `Immediate` durability — three separate fsyncs, by
design.
3. Make the wrapper the **only** way to reach a tool. If a call site can issue an
effect without an intent, the protocol is decorative.
4. Build the crash harness: a test hook that aborts the process at a named point.
Two named points minimum — after `Pending` commit, and after `Dispatched`
commit but before outcome commit.
5. On restart, scan for the newest intent per `(BranchKey, AttemptNo)` and
classify per the table. Emit the classification as a log event so the decision
itself is recorded.
6. Assert the fsync count on the intent path in a test — a batching "optimization"
introduced later must fail loudly.
## Acceptance
- Crash harness kills the process in both windows: between `Pending` and
`Dispatched`, and between `Dispatched` and outcome-commit.
- Restart classifies the first as "never sent" and the second as "may have been
sent". A two-phase implementation cannot pass this — both crashes leave one
`Pending` with no outcome and are indistinguishable. That is the point of the
test.
## Verify
**Harness:** a real child-process kill (`kill -9`), not a panic. Named abort
points compiled under a `test-hooks` feature. A recording tool that appends to a
file outside the database, so "did the call actually happen" is observable
independently of what the log claims.
**Integration test**`tests/it_intent_windows.rs`:
1. **Window A:** arm the abort after the `Pending` commit, before `Dispatched`.
Kill. Reopen.
- Assert the last intent state is `Pending`.
- Assert restart classifies it **"never sent"**.
- Assert the side-channel file is **empty** — the call really did not happen.
2. **Window B:** arm the abort after the `Dispatched` commit, before the
outcome commit. Kill. Reopen.
- Assert the last intent state is `Dispatched`.
- Assert restart classifies it **"may have been sent"**.
3. Assert the two classifications **differ**. A two-phase implementation makes
them identical, which is the whole point of the test.
4. Count fsyncs on the intent path for one guarded effect; assert exactly 3
commits, so a later batching "optimization" fails loudly.
5. Bypass audit: assert no tool can be invoked except through the intent wrapper —
a compile-level guard if the tool handle is only constructible inside it.
**Command:** `cargo test -p durability --features test-hooks intent_windows -- --test-threads=1`
**False pass:**
- Simulating the crash with `panic!` and catching the unwind. Destructors run and
buffered writes flush — precisely what a real crash does not do.
- Asserting only that both windows "recover". Both recover under a two-phase
implementation too; they just recover **identically**. Step 3 is the test.
- Trusting the log to tell you whether the call happened. Step 1's side-channel
file is the independent witness.
## Traps
- Batching the two pre-call fsyncs into one commit. It compiles, it is faster,
and it deletes the distinction the task exists to create.
- Recording the intent after dispatch "since that is when we know the request
id". Derive the request id before dispatch instead.
---
Background (not required to do this task):
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §8.4 ·
[rust-agentic-task.md](../../../rust-agentic-task.md)