Files
poimen/tasks/T5.12-p5-composition-gate.md
T

121 lines
5.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# T5.12 — P5 composition gate
| Field | Value |
|---|---|
| Phase | P5 — Grading |
| Size | L — over 3 days |
| Status | Not started |
| Flags | gate |
| Spec | inlined below |
| Blocks | P6 |
## Goal
Prove the grading strategies are **interchangeable**: the same recorded episodes,
graded under any strategy, flow through the same downstream pipeline unchanged.
**Phase gate criterion:** `PairwiseSequential` decides an accept and a reject
against a mock judge, on one resident model with zero swaps.
## Facts (inlined — no spec read needed)
- **Defaults ship working; every default is a port.** Grading is
strategy-pluggable, and the strategy declares what hardware it needs before it
is allowed to run. That claim is only proven by running **every** strategy
through the same pipeline on the same input.
- The four strategies and their declared costs per episode, on a group of eight:
| Strategy | Calls / episode | Models resident | Produces | Default |
|---|---|---|---|---|
| `DeterministicGrader` | 0 | 0 | `Ranked` | — |
| **`PairwiseSequential`** | **12** | **1, the agent's** | `Relative` | **yes** |
| `TournamentGrader` | 35 | 1 | `Ranked` | opt-in |
| `ReplayTournament` | 35 + N agent runs | 1 | `Ranked` | opt-in |
- The properties no single P5 task owns:
- every strategy's output flows into retention (T4.4), metering and the run
lifecycle **identically**`Ranked`, `Relative`, `Capped` and `Ungraded` all
reach a terminal run state;
- the declared `calls_per_episode` **matches measured** calls, or every
capacity and spend projection built on it is fiction;
- **zero model swaps** hold across the whole phase, not just per strategy;
- the outcome-class bracketing rule has **no exemptions** — assert it once per
strategy, since each has its own pairing path.
- `opt-in` strategies ship **disabled** and gate nothing. This gate must pass
with them off, and separately with them on.
## Steps
1. Fix one corpus of recorded episodes covering: same-outcome groups,
mixed-outcome groups, a novel `TaskId`, a capped episode, and a
budget-exhausted tenant.
2. Grade the corpus under each of the four strategies.
3. Assert the downstream pipeline — score handling, run terminal state, retention
eligibility, metering attribution — is identical in shape across all four.
4. Assert declared-versus-measured cost for each strategy.
5. Run the whole gate twice: opt-in flags **off**, then **on**.
6. Make this the required CI job gating P6.
## Acceptance
- All four strategies grade the same corpus; every episode reaches a terminal run
state with retention eligibility correct under each.
- `PairwiseSequential` decides an **accept** and a **reject** against mock judges.
- **Zero model swaps** across the entire gate.
- Measured `compare` calls per episode fall within each strategy's declared
`calls_per_episode`.
- The gate passes with opt-in flags off **and** on.
## Verify
**Harness:** one committed episode corpus; mock judges (70% winner, 50/50,
all-draws, position-biased, one `CoreViolation`); counters for `compare` calls,
model **loads**, and agent runs, all separate.
**Integration test**`tests/it_p5_composition.rs`:
1. **Strategy interchange:** parameterize over all four strategies. For each,
assert every episode ends in a terminal run state, and that
`is_reducible` (T4.4) agrees with that state. A strategy leaving runs in
`Grading` fails here regardless of its scores.
2. **Cost truth:** assert measured `compare` calls per episode fall inside the
declared `calls_per_episode` band. A declaration that does not match
measurement makes every capacity and spend projection downstream wrong.
3. **Zero swaps:** assert the swap counter is **exactly 0** across the whole
parameterized run, not per strategy.
4. **Bracketing, per strategy:** feed the mixed-outcome group to each strategy;
assert **zero** cross-class `compare` calls in every one. Each strategy has
its own pairing path and the rule has no exemptions.
5. **Score-type handling:** assert `Capped` and `Ungraded` episodes are handled
identically by every strategy's caller — same terminal state, excluded from
aggregates, no numeric coercion anywhere (cross-check T5.3's compile-fail).
6. **Accept and reject:** with the 70% judge assert `Accept`; with the 50/50
judge assert `Reject`. Assert both are reached in fewer comparisons than
fixed-n.
7. **Flag matrix:** run steps 16 with opt-in off (tournament paths must be
unreachable) and on.
8. **Regression:** re-run P0P4 gates and all P5 suites.
**Command:** `cargo test -p grading --test it_p5_composition -- --nocapture`
**False pass:**
- Testing each strategy against its own tailored corpus. The claim is
interchangeability, which only means something on **one** shared input.
- Step 2 omitted: `calls_per_episode` is a declared number nobody measures, and
admission control, capacity planning and spend projection all consume it.
- Step 3 counted per strategy rather than across the run — a swap between two
strategies' evaluations is exactly the failure the one-resident-model default
exists to prevent.
- Step 1 asserting scores only. A strategy that produces correct scores and
leaves runs resting in `Grading` strands storage forever.
## Traps
- Enabling the opt-in flags to make the gate pass. They ship disabled and gate
nothing; the off-run is the one that must be green.
---
Background (not required to do this task):
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11, §14.2 ·
[rust-agentic-task.md](../../../rust-agentic-task.md)