Files
poimen/tasks/T1.2-stub-model-provider.md
T

78 lines
3.1 KiB
Markdown
Raw Normal View History

2026-08-17 23:05:20 -07:00
# T1.2 — Stub model provider
| Field | Value |
|---|---|
| Phase | P1 — Walking skeleton |
| Size | S — under 1 day |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
## Goal
A deterministic scripted model provider so the walking skeleton runs end to end
with no network and no spend.
## Facts (inlined — no spec read needed)
- Nothing is built before the record is trustworthy: no dashboards and no
learning loop until one full run works end to end **against a stub model**.
- The stub is not a test fixture that gets deleted — it is the fixed point every
later determinism assertion is measured against, including T0.8's re-fold
property and T5.x's grading tests.
- Failure injection here is what T1.4 (retry), T2.1 (crash windows) and T2.6
(crash matrix) drive their scenarios with.
## Steps
1. Define the model port the executor calls: request in, response plus `Usage`
out, `async`, cancellation-aware.
2. Implement the stub over a script: an ordered list of responses, optionally
keyed by step id so a workflow with branches stays readable.
3. Add configurable latency per response, driven by `tokio::time::sleep` so
`tokio::time::pause` in tests makes it instant.
4. Add failure injection: error responses, timeouts, and a hang that outlives the
deadline. Each selectable per script entry.
5. Make every field of the emitted response deterministic — no timestamps from
the wall clock, no generated ids that are not seeded.
6. Test: run the same script twice, serialize both episodes, assert equal bytes.
## Acceptance
- The same script yields byte-identical episodes across runs.
## Verify
**Harness:** the stub itself plus T0.8's `assert_refold_identical`.
**Integration test**`tests/it_stub_determinism.rs`:
1. Run the same script twice, in two fresh databases, in the same process.
2. Serialize both episodes; assert **byte equality**.
3. Run a third time in a **separate process** and compare against a committed
expected serialization — catches anything seeded from process state.
4. With `tokio::time::pause`, assert a script declaring 30s latency completes in
near-zero wall time. If it does not, the stub is sleeping off the tokio clock.
5. Failure injection: assert an error entry produces a failed attempt, a timeout
entry produces `TimedOut`, and a hang entry is cut off by the deadline.
**Command:** `cargo test -p testkit stub`
**False pass:**
- Steps 12 in one process with a lazily-initialized global seed: both runs share
it and agree, while a fresh process differs. Step 3 is the guard.
- Comparing episodes with a `PartialEq` that skips timestamps. That hides exactly
the nondeterminism being hunted — compare the full serialized bytes.
## Traps
- A wall-clock timestamp or a random id in the stub response. It defeats every
byte-identity assertion downstream and looks like a fold bug.
- Sleeping with `std::thread::sleep`, which `tokio::time::pause` cannot skip.
---
Background (not required to do this task):
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §16 ·
[rust-agentic-task.md](../../../rust-agentic-task.md)