Files
poimen/tasks/T5.3-evaluationstrategy-port-resourceprofile.md

7.0 KiB
Raw Permalink Blame History

T5.3 — EvaluationStrategy port + ResourceProfile

Field Value
Phase P5 — Grading
Size S — under 1 day
Status Not started
Flags
Spec inlined below
Blocks T5.4, T5.5

Goal

The grading ports, and the type discipline that stops a capped or ungraded episode from being averaged into a promotion gate.

Facts (inlined — no spec read needed)

#[async_trait]
pub trait EvaluationStrategy: Send + Sync {
    fn id(&self) -> StrategyId;
    /// Declared before any work is admitted. Validated against capacity limits
    /// at load; a strategy whose profile does not fit is rejected by name.
    fn resources(&self) -> ResourceProfile;
    async fn evaluate(&self, cx: &EvalCtx) -> Result<Vec<Score>>;
}

pub struct ResourceProfile {
    /// Models this strategy calls. One entry means it runs on the agent's
    /// already-resident model and forces no swap.
    pub models: Vec<ModelId>,
    /// Largest single-call context. A pairwise judge reads two episodes, so this
    /// is roughly twice an episode budget.
    pub max_context_tokens: u32,
    /// Model calls per episode evaluated, for capacity and spend projection.
    pub calls_per_episode: f32,
}

#[async_trait]
pub trait Grader: Send + Sync {
    async fn grade(&self, group: &Group, rubric: &RubricDef) -> Result<Vec<Score>>;
}

#[async_trait]
pub trait Judge: Send + Sync {
    /// Relative comparison only. Deliberately cannot return an absolute score.
    async fn compare(&self, a: &EpisodeView, b: &EpisodeView, r: &RubricDef)
        -> Result<Verdict>;
    /// Unary, separate from `compare`: a `Core` violation caps an episode on its
    /// own terms, not relative to an opponent. Runs before pairing.
    async fn screen(&self, e: &EpisodeView, r: &RubricDef) -> Result<Vec<CoreViolation>>;
}

pub enum Verdict { A, B, Draw }

pub struct CoreViolation { pub criterion: RubricCriterionId, pub evidence: BlobRef }

pub enum Score {
    Relative { against: RunId, verdict: Verdict, record: WinRecord },
    Ranked { strength: f64, interval: (f64, f64), group_size: u32 },
    Capped { violations: Vec<CoreViolation> },
    Ungraded { reason: UngradedReason },
}
  • compare returns Verdict, never a number. Absolute scores fail three ways here: calibration drift (same episode scores differently across weeks and model versions, and drift is indistinguishable from a variant trend); weak discrimination (four competent episodes all score 0.8); and saturation — as workflows improve, pass rate approaches 100% and absolute rubric scores saturate identically. A tournament cannot saturate; better candidates just make it harder.
  • Score is a sum, not a number with flags. A capped episode and an ungraded one are not low scores; they are different kinds of answer. A type that can represent them as numbers will eventually have them averaged into a promotion gate by code that meant no harm.
  • The strategy declares what hardware it needs before it is allowed to run. Strategies differ by more than an order of magnitude in model calls and VRAM; a deployment that cannot afford one must be told at load, not by an OOM at 3am.

Strategy catalogue — cost per episode evaluated, group of eight:

Strategy Calls / episode Models resident Produces Default
DeterministicGrader 0 0 Ranked on a computed number
PairwiseSequential 12 1, the agent's Relative yes
TournamentGrader 35 1 Ranked with intervals opt-in
ReplayTournament 35 plus N full agent runs 1 Ranked across variants opt-in

Steps

  1. Define all types above verbatim. Verdict has exactly three variants.
  2. Give Score no Into<f64>, no as_number(), no unwrap_or(0.0) convenience. Aggregation code must match on the variant.
  3. Keep screen separate and unary. It is the only source of CoreViolation, and it runs before pairing.
  4. At strategy registration, call T5.2's residency check against resources().models plus the pinned set. Reject by strategy name with the shortfall.
  5. Validate resources().max_context_tokens against T5.2's derived ceiling at the same point.
  6. Write the two trybuild compile-fail cases described in acceptance.

Acceptance

  • Compile-fail test: a judge cannot return f64.
  • Compile-fail test: Score::Capped and Score::Ungraded cannot be coerced to a number — no Into<f64>, no unwrap_or(0.0) path through the aggregate.
  • A strategy whose ResourceProfile names a second model is rejected at load when the residency invariant (T5.2) does not admit both.

Verify

Harness: trybuild for the type discipline; T5.2's capacity checker for the load-time rejection.

Integration testtests/it_grading_ports.rs + tests/compile_fail/:

  1. Compile-fail: a Judge impl whose compare returns f64. Assert the stderr names the Verdict return type.
  2. Compile-fail: Into::<f64>::into(Score::Capped { .. }), score.unwrap_or(0.0), and any as_number() call. One case each — a single case leaves the other routes open.
  3. Audit test: enumerate every aggregation site and assert each matches on Score variants exhaustively with no _ arm.
  4. Load-time rejection: a strategy whose ResourceProfile names a second model, under a CapacityLimits that admits only one. Assert rejection by strategy name, at registration, before any evaluation runs.
  5. Positive: the same strategy under limits that admit both models is accepted — otherwise the check is just "reject two models", which is the wrong rule.
  6. Context check: a strategy declaring max_context_tokens above T5.2's derived ceiling is rejected with the computed limit in the error.
  7. screen separation: assert CoreViolation can only originate from screen — grep plus a test that a compare result cannot construct one.

Command: cargo test -p grading ports && cargo test -p grading --test compile_fail

False pass:

  • One compile-fail case for f64. The leak that matters is on Score, not on Verdict — a capped episode reaching an average is the failure this type discipline exists to prevent, and step 2 is where it is caught.
  • Step 5 omitted, so a naive "more than one model is illegal" implementation passes and blocks a legitimate two-resident-model deployment.
  • Step 3 omitted: the types can be sound while one aggregation site does if let Some(n) = ... and silently skips capped episodes instead of failing.

Traps

  • A Score::value() -> Option<f64> helper. Every call site then writes .unwrap_or(0.0) and the cap becomes a zero.
  • Folding screen into compare. A judge asked to express a Core violation through a comparison can only rank the offender lower, which is not a cap.

Background (not required to do this task): rust-agentic-sys.md §11.1, §11.2, §14.2 · rust-agentic-task.md