121 lines
5.6 KiB
Markdown
121 lines
5.6 KiB
Markdown
# T5.6 — Sequential test and stopping
|
||
|
||
| Field | Value |
|
||
|---|---|
|
||
| Phase | P5 — Grading |
|
||
| Size | M — 1 to 3 days |
|
||
| Status | Not started |
|
||
| Flags | — |
|
||
| Spec | inlined below |
|
||
| Blocks | T6.3 |
|
||
|
||
## Goal
|
||
|
||
Stop on evidence, not on a fixed sample. Verdicts accumulate into a likelihood
|
||
ratio tested against α/β boundaries.
|
||
|
||
## Facts (inlined — no spec read needed)
|
||
|
||
```
|
||
┌────────────────┬────────────────┐
|
||
▼ ▼ ▼
|
||
accept continue reject
|
||
challenger becomes keep sampling discard, keep
|
||
the current current
|
||
```
|
||
|
||
- **Stopping is a sequential test, not a fixed sample.** Boundaries are set by α,
|
||
β and the smallest win-rate shift worth acting on; the test stops as soon as a
|
||
boundary is crossed. A clearly better challenger is accepted in far fewer
|
||
comparisons than a fixed-n design would spend, and a clearly worse one is
|
||
rejected early instead of running to completion.
|
||
- This is what makes cost adaptive: cheap decisions cost little, close decisions
|
||
cost more, and nothing costs the worst case by default.
|
||
- **Draws are recorded and excluded from the ratio.** A tie carries no evidence
|
||
about which is stronger, so folding it in as half a win manufactures
|
||
information.
|
||
- **But a high draw rate is itself a result** — it says the challenger is not
|
||
meaningfully different — so the test also **rejects on a draw-rate ceiling**
|
||
rather than sampling forever toward a boundary it will never reach.
|
||
- α, β and the minimum detectable win-rate shift are **policy, not physics**: too
|
||
tight and no challenger is ever accepted, too loose and the loop churns the
|
||
current version on noise. Same for the draw ceiling, which interacts with judge
|
||
quality — a weak judge draws more. Expose all four as configuration with
|
||
documented defaults.
|
||
|
||
## Steps
|
||
|
||
1. Define `WinRecord { wins, losses, draws }` and persist it with the decision,
|
||
not only in memory — the decision must survive a restart.
|
||
2. Implement the sequential probability ratio test: maintain the log-likelihood
|
||
ratio over wins and losses under H0 (win rate = 0.5) and H1 (win rate =
|
||
0.5 + δ), with boundaries derived from α and β.
|
||
3. Exclude draws from the ratio; count them separately.
|
||
4. Add the draw-rate ceiling check after each comparison: over the ceiling with a
|
||
minimum sample, reject.
|
||
5. Return a three-valued decision — `Accept`, `Reject`, `Continue` — and let the
|
||
caller decide whether to sample another episode.
|
||
6. Expose α, β, δ and the draw ceiling as configuration. Document the defaults as
|
||
provisional and needing calibration against a workflow whose true improvement
|
||
is known.
|
||
7. Build the mock judges: a 70% winner, a 50/50 coin, and an all-draws judge.
|
||
|
||
## Acceptance
|
||
|
||
- A mock judge with a true 70% win rate crosses the accept boundary in
|
||
**materially fewer comparisons than a fixed-n design at the same α** — assert
|
||
the count, since an implementation that ignores the boundary and runs to n
|
||
still reaches the right answer.
|
||
- A 50/50 judge rejects.
|
||
- A judge returning `Draw` on every comparison terminates on the draw ceiling
|
||
instead of running forever.
|
||
|
||
## Verify
|
||
|
||
**Harness:** seeded mock judges with known true win rates. Every test asserts a
|
||
**comparison count**, not only an outcome — the outcome is reachable by an
|
||
implementation that ignores the boundary entirely.
|
||
|
||
**Integration test** — `tests/it_sequential_test.rs`:
|
||
1. **70% winner:** assert `Accept`, and assert the comparison count is
|
||
**materially below** the fixed-n sample size for the same α. Compute the
|
||
fixed-n figure in the test and assert `count < fixed_n`, with both printed.
|
||
2. **50/50 judge:** assert `Reject`, and assert it terminates — bounded count.
|
||
3. **All-draws judge:** assert termination **on the draw-rate ceiling**, and
|
||
assert the reason is the ceiling, not the α boundary. Without this it loops
|
||
forever.
|
||
4. **30% winner (clearly worse):** assert `Reject` early — assert the count is
|
||
well below the 70% case's, proving early rejection works in both directions.
|
||
5. Draws excluded: feed a sequence of `win, draw, win, draw`; assert the
|
||
likelihood ratio equals that of `win, win`, and that `draws` is counted
|
||
separately in the `WinRecord`.
|
||
6. Persistence: kill the process mid-decision, restart, assert the `WinRecord`
|
||
survived and the decision resumes rather than restarting.
|
||
7. Repeat each case over 100 seeds; assert the accept/reject rates sit within
|
||
α/β. A single seed says nothing about a statistical test.
|
||
|
||
**Command:** `cargo test -p grading sequential -- --nocapture`
|
||
|
||
**False pass:**
|
||
- Asserting only the decision. A fixed-n implementation reaches the **right
|
||
answer** on all of steps 1–4 and delivers none of the cost saving that is the
|
||
entire justification for the design. The count assertion is the test.
|
||
- Step 5 omitted: counting a draw as half a win tightens the boundary on invented
|
||
evidence and still produces plausible decisions.
|
||
- One seed per case. A sequential test is a statistical object; step 7 is what
|
||
makes the α/β claim meaningful.
|
||
|
||
## Traps
|
||
|
||
- Counting a draw as half a win. It tightens the boundary on evidence that does
|
||
not exist.
|
||
- Checking the boundary only at the end. The test then costs exactly what a
|
||
fixed-n design costs, and the acceptance count catches it — which is why the
|
||
count is asserted rather than the outcome.
|
||
|
||
---
|
||
|
||
Background (not required to do this task):
|
||
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.3, §18 ·
|
||
[rust-agentic-task.md](../../../rust-agentic-task.md)
|