Files
poimen/tasks/T5.3-evaluationstrategy-port-resourceprofile.md
T

159 lines
7.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# T5.3 — `EvaluationStrategy` port + `ResourceProfile`
| Field | Value |
|---|---|
| Phase | P5 — Grading |
| Size | S — under 1 day |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | T5.4, T5.5 |
## Goal
The grading ports, and the type discipline that stops a capped or ungraded
episode from being averaged into a promotion gate.
## Facts (inlined — no spec read needed)
```rust
#[async_trait]
pub trait EvaluationStrategy: Send + Sync {
fn id(&self) -> StrategyId;
/// Declared before any work is admitted. Validated against capacity limits
/// at load; a strategy whose profile does not fit is rejected by name.
fn resources(&self) -> ResourceProfile;
async fn evaluate(&self, cx: &EvalCtx) -> Result<Vec<Score>>;
}
pub struct ResourceProfile {
/// Models this strategy calls. One entry means it runs on the agent's
/// already-resident model and forces no swap.
pub models: Vec<ModelId>,
/// Largest single-call context. A pairwise judge reads two episodes, so this
/// is roughly twice an episode budget.
pub max_context_tokens: u32,
/// Model calls per episode evaluated, for capacity and spend projection.
pub calls_per_episode: f32,
}
#[async_trait]
pub trait Grader: Send + Sync {
async fn grade(&self, group: &Group, rubric: &RubricDef) -> Result<Vec<Score>>;
}
#[async_trait]
pub trait Judge: Send + Sync {
/// Relative comparison only. Deliberately cannot return an absolute score.
async fn compare(&self, a: &EpisodeView, b: &EpisodeView, r: &RubricDef)
-> Result<Verdict>;
/// Unary, separate from `compare`: a `Core` violation caps an episode on its
/// own terms, not relative to an opponent. Runs before pairing.
async fn screen(&self, e: &EpisodeView, r: &RubricDef) -> Result<Vec<CoreViolation>>;
}
pub enum Verdict { A, B, Draw }
pub struct CoreViolation { pub criterion: RubricCriterionId, pub evidence: BlobRef }
pub enum Score {
Relative { against: RunId, verdict: Verdict, record: WinRecord },
Ranked { strength: f64, interval: (f64, f64), group_size: u32 },
Capped { violations: Vec<CoreViolation> },
Ungraded { reason: UngradedReason },
}
```
- **`compare` returns `Verdict`, never a number.** Absolute scores fail three
ways here: calibration drift (same episode scores differently across weeks and
model versions, and drift is indistinguishable from a variant trend);
weak discrimination (four competent episodes all score 0.8); and saturation —
as workflows improve, pass rate approaches 100% and absolute rubric scores
saturate identically. A tournament cannot saturate; better candidates just make
it harder.
- **`Score` is a sum, not a number with flags.** A capped episode and an ungraded
one are not low scores; they are different kinds of answer. A type that can
represent them as numbers will eventually have them averaged into a promotion
gate by code that meant no harm.
- **The strategy declares what hardware it needs before it is allowed to run.**
Strategies differ by more than an order of magnitude in model calls and VRAM; a
deployment that cannot afford one must be told at load, not by an OOM at 3am.
Strategy catalogue — cost per episode evaluated, group of eight:
| Strategy | Calls / episode | Models resident | Produces | Default |
|---|---|---|---|---|
| `DeterministicGrader` | 0 | 0 | `Ranked` on a computed number | — |
| **`PairwiseSequential`** | **12** | **1, the agent's** | `Relative` | **yes** |
| `TournamentGrader` | 35 | 1 | `Ranked` with intervals | opt-in |
| `ReplayTournament` | 35 **plus N full agent runs** | 1 | `Ranked` across variants | opt-in |
## Steps
1. Define all types above verbatim. `Verdict` has exactly three variants.
2. Give `Score` **no** `Into<f64>`, no `as_number()`, no `unwrap_or(0.0)`
convenience. Aggregation code must match on the variant.
3. Keep `screen` separate and unary. It is the only source of `CoreViolation`,
and it runs before pairing.
4. At strategy registration, call T5.2's residency check against
`resources().models` plus the pinned set. Reject by strategy **name** with the
shortfall.
5. Validate `resources().max_context_tokens` against T5.2's derived ceiling at
the same point.
6. Write the two `trybuild` compile-fail cases described in acceptance.
## Acceptance
- Compile-fail test: a judge cannot return `f64`.
- Compile-fail test: `Score::Capped` and `Score::Ungraded` cannot be coerced to a
number — no `Into<f64>`, no `unwrap_or(0.0)` path through the aggregate.
- A strategy whose `ResourceProfile` names a second model is rejected at load
when the residency invariant (T5.2) does not admit both.
## Verify
**Harness:** `trybuild` for the type discipline; T5.2's capacity checker for the
load-time rejection.
**Integration test**`tests/it_grading_ports.rs` + `tests/compile_fail/`:
1. **Compile-fail:** a `Judge` impl whose `compare` returns `f64`. Assert the
stderr names the `Verdict` return type.
2. **Compile-fail:** `Into::<f64>::into(Score::Capped { .. })`,
`score.unwrap_or(0.0)`, and any `as_number()` call. One case each — a single
case leaves the other routes open.
3. **Audit test:** enumerate every aggregation site and assert each matches on
`Score` variants exhaustively with no `_` arm.
4. Load-time rejection: a strategy whose `ResourceProfile` names a **second**
model, under a `CapacityLimits` that admits only one. Assert rejection **by
strategy name**, at registration, before any evaluation runs.
5. Positive: the same strategy under limits that admit both models is accepted —
otherwise the check is just "reject two models", which is the wrong rule.
6. Context check: a strategy declaring `max_context_tokens` above T5.2's derived
ceiling is rejected with the computed limit in the error.
7. `screen` separation: assert `CoreViolation` can only originate from `screen`
grep plus a test that a `compare` result cannot construct one.
**Command:** `cargo test -p grading ports && cargo test -p grading --test compile_fail`
**False pass:**
- One compile-fail case for `f64`. The leak that matters is on `Score`, not on
`Verdict` — a capped episode reaching an average is the failure this type
discipline exists to prevent, and step 2 is where it is caught.
- Step 5 omitted, so a naive "more than one model is illegal" implementation
passes and blocks a legitimate two-resident-model deployment.
- Step 3 omitted: the types can be sound while one aggregation site does
`if let Some(n) = ...` and silently skips capped episodes instead of failing.
## Traps
- A `Score::value() -> Option<f64>` helper. Every call site then writes
`.unwrap_or(0.0)` and the cap becomes a zero.
- Folding `screen` into `compare`. A judge asked to express a `Core` violation
through a comparison can only rank the offender lower, which is not a cap.
---
Background (not required to do this task):
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.1, §11.2, §14.2 ·
[rust-agentic-task.md](../../../rust-agentic-task.md)