Files
poimen/tasks/T5.3-evaluationstrategy-port-resourceprofile.md
T

159 lines
7.0 KiB
Markdown
Raw Normal View History

2026-08-17 23:05:20 -07:00
# T5.3 — `EvaluationStrategy` port + `ResourceProfile`
| Field | Value |
|---|---|
| Phase | P5 — Grading |
| Size | S — under 1 day |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | T5.4, T5.5 |
## Goal
The grading ports, and the type discipline that stops a capped or ungraded
episode from being averaged into a promotion gate.
## Facts (inlined — no spec read needed)
```rust
#[async_trait]
pub trait EvaluationStrategy: Send + Sync {
fn id(&self) -> StrategyId;
/// Declared before any work is admitted. Validated against capacity limits
/// at load; a strategy whose profile does not fit is rejected by name.
fn resources(&self) -> ResourceProfile;
async fn evaluate(&self, cx: &EvalCtx) -> Result<Vec<Score>>;
}
pub struct ResourceProfile {
/// Models this strategy calls. One entry means it runs on the agent's
/// already-resident model and forces no swap.
pub models: Vec<ModelId>,
/// Largest single-call context. A pairwise judge reads two episodes, so this
/// is roughly twice an episode budget.
pub max_context_tokens: u32,
/// Model calls per episode evaluated, for capacity and spend projection.
pub calls_per_episode: f32,
}
#[async_trait]
pub trait Grader: Send + Sync {
async fn grade(&self, group: &Group, rubric: &RubricDef) -> Result<Vec<Score>>;
}
#[async_trait]
pub trait Judge: Send + Sync {
/// Relative comparison only. Deliberately cannot return an absolute score.
async fn compare(&self, a: &EpisodeView, b: &EpisodeView, r: &RubricDef)
-> Result<Verdict>;
/// Unary, separate from `compare`: a `Core` violation caps an episode on its
/// own terms, not relative to an opponent. Runs before pairing.
async fn screen(&self, e: &EpisodeView, r: &RubricDef) -> Result<Vec<CoreViolation>>;
}
pub enum Verdict { A, B, Draw }
pub struct CoreViolation { pub criterion: RubricCriterionId, pub evidence: BlobRef }
pub enum Score {
Relative { against: RunId, verdict: Verdict, record: WinRecord },
Ranked { strength: f64, interval: (f64, f64), group_size: u32 },
Capped { violations: Vec<CoreViolation> },
Ungraded { reason: UngradedReason },
}
```
- **`compare` returns `Verdict`, never a number.** Absolute scores fail three
ways here: calibration drift (same episode scores differently across weeks and
model versions, and drift is indistinguishable from a variant trend);
weak discrimination (four competent episodes all score 0.8); and saturation —
as workflows improve, pass rate approaches 100% and absolute rubric scores
saturate identically. A tournament cannot saturate; better candidates just make
it harder.
- **`Score` is a sum, not a number with flags.** A capped episode and an ungraded
one are not low scores; they are different kinds of answer. A type that can
represent them as numbers will eventually have them averaged into a promotion
gate by code that meant no harm.
- **The strategy declares what hardware it needs before it is allowed to run.**
Strategies differ by more than an order of magnitude in model calls and VRAM; a
deployment that cannot afford one must be told at load, not by an OOM at 3am.
Strategy catalogue — cost per episode evaluated, group of eight:
| Strategy | Calls / episode | Models resident | Produces | Default |
|---|---|---|---|---|
| `DeterministicGrader` | 0 | 0 | `Ranked` on a computed number | — |
| **`PairwiseSequential`** | **12** | **1, the agent's** | `Relative` | **yes** |
| `TournamentGrader` | 35 | 1 | `Ranked` with intervals | opt-in |
| `ReplayTournament` | 35 **plus N full agent runs** | 1 | `Ranked` across variants | opt-in |
## Steps
1. Define all types above verbatim. `Verdict` has exactly three variants.
2. Give `Score` **no** `Into<f64>`, no `as_number()`, no `unwrap_or(0.0)`
convenience. Aggregation code must match on the variant.
3. Keep `screen` separate and unary. It is the only source of `CoreViolation`,
and it runs before pairing.
4. At strategy registration, call T5.2's residency check against
`resources().models` plus the pinned set. Reject by strategy **name** with the
shortfall.
5. Validate `resources().max_context_tokens` against T5.2's derived ceiling at
the same point.
6. Write the two `trybuild` compile-fail cases described in acceptance.
## Acceptance
- Compile-fail test: a judge cannot return `f64`.
- Compile-fail test: `Score::Capped` and `Score::Ungraded` cannot be coerced to a
number — no `Into<f64>`, no `unwrap_or(0.0)` path through the aggregate.
- A strategy whose `ResourceProfile` names a second model is rejected at load
when the residency invariant (T5.2) does not admit both.
## Verify
**Harness:** `trybuild` for the type discipline; T5.2's capacity checker for the
load-time rejection.
**Integration test**`tests/it_grading_ports.rs` + `tests/compile_fail/`:
1. **Compile-fail:** a `Judge` impl whose `compare` returns `f64`. Assert the
stderr names the `Verdict` return type.
2. **Compile-fail:** `Into::<f64>::into(Score::Capped { .. })`,
`score.unwrap_or(0.0)`, and any `as_number()` call. One case each — a single
case leaves the other routes open.
3. **Audit test:** enumerate every aggregation site and assert each matches on
`Score` variants exhaustively with no `_` arm.
4. Load-time rejection: a strategy whose `ResourceProfile` names a **second**
model, under a `CapacityLimits` that admits only one. Assert rejection **by
strategy name**, at registration, before any evaluation runs.
5. Positive: the same strategy under limits that admit both models is accepted —
otherwise the check is just "reject two models", which is the wrong rule.
6. Context check: a strategy declaring `max_context_tokens` above T5.2's derived
ceiling is rejected with the computed limit in the error.
7. `screen` separation: assert `CoreViolation` can only originate from `screen`
grep plus a test that a `compare` result cannot construct one.
**Command:** `cargo test -p grading ports && cargo test -p grading --test compile_fail`
**False pass:**
- One compile-fail case for `f64`. The leak that matters is on `Score`, not on
`Verdict` — a capped episode reaching an average is the failure this type
discipline exists to prevent, and step 2 is where it is caught.
- Step 5 omitted, so a naive "more than one model is illegal" implementation
passes and blocks a legitimate two-resident-model deployment.
- Step 3 omitted: the types can be sound while one aggregation site does
`if let Some(n) = ...` and silently skips capped episodes instead of failing.
## Traps
- A `Score::value() -> Option<f64>` helper. Every call site then writes
`.unwrap_or(0.0)` and the cap becomes a zero.
- Folding `screen` into `compare`. A judge asked to express a `Core` violation
through a comparison can only rank the offender lower, which is not a cap.
---
Background (not required to do this task):
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.1, §11.2, §14.2 ·
[rust-agentic-task.md](../../../rust-agentic-task.md)