(tasks) add tasks for harness
This commit is contained in:
@@ -0,0 +1,77 @@
|
||||
# T1.2 — Stub model provider
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | P1 — Walking skeleton |
|
||||
| Size | S — under 1 day |
|
||||
| Status | Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | — |
|
||||
|
||||
## Goal
|
||||
|
||||
A deterministic scripted model provider so the walking skeleton runs end to end
|
||||
with no network and no spend.
|
||||
|
||||
## Facts (inlined — no spec read needed)
|
||||
|
||||
- Nothing is built before the record is trustworthy: no dashboards and no
|
||||
learning loop until one full run works end to end **against a stub model**.
|
||||
- The stub is not a test fixture that gets deleted — it is the fixed point every
|
||||
later determinism assertion is measured against, including T0.8's re-fold
|
||||
property and T5.x's grading tests.
|
||||
- Failure injection here is what T1.4 (retry), T2.1 (crash windows) and T2.6
|
||||
(crash matrix) drive their scenarios with.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Define the model port the executor calls: request in, response plus `Usage`
|
||||
out, `async`, cancellation-aware.
|
||||
2. Implement the stub over a script: an ordered list of responses, optionally
|
||||
keyed by step id so a workflow with branches stays readable.
|
||||
3. Add configurable latency per response, driven by `tokio::time::sleep` so
|
||||
`tokio::time::pause` in tests makes it instant.
|
||||
4. Add failure injection: error responses, timeouts, and a hang that outlives the
|
||||
deadline. Each selectable per script entry.
|
||||
5. Make every field of the emitted response deterministic — no timestamps from
|
||||
the wall clock, no generated ids that are not seeded.
|
||||
6. Test: run the same script twice, serialize both episodes, assert equal bytes.
|
||||
|
||||
## Acceptance
|
||||
|
||||
- The same script yields byte-identical episodes across runs.
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** the stub itself plus T0.8's `assert_refold_identical`.
|
||||
|
||||
**Integration test** — `tests/it_stub_determinism.rs`:
|
||||
1. Run the same script twice, in two fresh databases, in the same process.
|
||||
2. Serialize both episodes; assert **byte equality**.
|
||||
3. Run a third time in a **separate process** and compare against a committed
|
||||
expected serialization — catches anything seeded from process state.
|
||||
4. With `tokio::time::pause`, assert a script declaring 30s latency completes in
|
||||
near-zero wall time. If it does not, the stub is sleeping off the tokio clock.
|
||||
5. Failure injection: assert an error entry produces a failed attempt, a timeout
|
||||
entry produces `TimedOut`, and a hang entry is cut off by the deadline.
|
||||
|
||||
**Command:** `cargo test -p testkit stub`
|
||||
|
||||
**False pass:**
|
||||
- Steps 1–2 in one process with a lazily-initialized global seed: both runs share
|
||||
it and agree, while a fresh process differs. Step 3 is the guard.
|
||||
- Comparing episodes with a `PartialEq` that skips timestamps. That hides exactly
|
||||
the nondeterminism being hunted — compare the full serialized bytes.
|
||||
|
||||
## Traps
|
||||
|
||||
- A wall-clock timestamp or a random id in the stub response. It defeats every
|
||||
byte-identity assertion downstream and looks like a fold bug.
|
||||
- Sleeping with `std::thread::sleep`, which `tokio::time::pause` cannot skip.
|
||||
|
||||
---
|
||||
|
||||
Background (not required to do this task):
|
||||
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §16 ·
|
||||
[rust-agentic-task.md](../../../rust-agentic-task.md)
|
||||
Reference in New Issue
Block a user