Files
poimen/tasks/T2.5-checkpoints.md
T

3.7 KiB

T2.5 — Checkpoints

Field Value
Phase P2 — Durability hard parts
Size M — 1 to 3 days
Status Not started
Flags
Spec inlined below
Blocks

Goal

Materialized state snapshots tagged with their LSN, so restart folds forward from the newest one instead of from zero.

Facts (inlined — no spec read needed)

  • A checkpoint is a materialized state snapshot tagged with its LSN. Restart folds forward from the newest one at or below the target LSN.
  • An optimization only. Deleting every checkpoint costs startup time and nothing else. That property is the acceptance test, and it is what keeps checkpoints from quietly becoming a second source of truth.
  • Port surface already exists from T0.5: put_checkpoint(key, upto, state) and latest_checkpoint(key, upto) -> Option<Checkpoint>. None means fold from LSN 0.
  • Checkpoints are per BranchKey, like everything else in the log.

Steps

  1. Serialize the materialized state (T0.8) deterministically — the same ordered collections, so a checkpoint written twice from the same state is byte-equal.
  2. Write checkpoints on a policy: every N records or every M seconds per branch. Keep the policy in one place and make it configurable rather than scattered.
  3. rebuild(key, upto): call latest_checkpoint, deserialize if present, then fold records from checkpoint.lsn + 1. On None, fold from 0.
  4. Never let the checkpoint be the only holder of a fact. Validate this by making the None path the default in tests.
  5. Add a prune path — old checkpoints for a branch are deletable at any time without coordination.
  6. Test both directions: with checkpoints and after deleting them all, comparing final state and reporting both startup times.

Acceptance

  • Deleting all checkpoints changes startup time and nothing else — final state identical.

Verify

Harness: a run long enough to trigger the checkpoint policy several times — at least 3 checkpoints, so "newest at or below" is a real choice.

Integration testtests/it_checkpoints.rs:

  1. Run to completion with checkpointing on. Record final state bytes and startup time for a rebuild.
  2. Delete every checkpoint row. Rebuild from LSN 0.
  3. Assert the final state bytes are identical; assert startup time differs (log both, do not assert a threshold — that is flaky).
  4. latest_checkpoint(key, upto) with upto between two checkpoints: assert it returns the lower one, not the newest overall.
  5. Determinism: write a checkpoint twice from the same state; assert byte equality of the two serializations.
  6. Crash between checkpoint write and subsequent commits; assert rebuild still lands on the same state.
  7. Prune: delete an older checkpoint while a newer exists; assert rebuild is unaffected.

Command: cargo test -p durability checkpoints

False pass:

  • A run short enough to produce zero or one checkpoint. Then step 2 deletes nothing and the test is vacuous — assert the checkpoint count is ≥ 3 before deleting.
  • Step 4 omitted: an implementation returning the globally newest checkpoint passes everything else and breaks the moment rewind or point-in-time rebuild needs an earlier LSN.
  • Comparing state with PartialEq rather than bytes.

Traps

  • Writing a checkpoint inside the same transaction as the commit path and then depending on it for correctness. It is a cache; keep it separable.
  • A checkpoint that serializes a HashMap. Byte-identity fails intermittently and reads as flakiness.

Background (not required to do this task): rust-agentic-sys.md §7, §8.6 · rust-agentic-task.md