Files
poimen/tasks/T5.4-deterministicgrader.md

3.5 KiB
Raw Permalink Blame History

T5.4 — DeterministicGrader

Field Value
Phase P5 — Grading
Size S — under 1 day
Status Not started
Flags
Spec inlined below
Blocks

Goal

Grade on a number the user already computes — latency, cost, test pass count — with zero model calls.

Facts (inlined — no spec read needed)

  • It exists so that adopting the framework does not require adopting LLM-as-judge at all.
  • Cost profile: 0 model calls per episode, 0 models resident, produces Ranked on a computed number.
  • Its ResourceProfile declares an empty models vector. That is the whole point — it forces no residency and passes the capacity check trivially.
  • It still produces Score::Ranked, so the aggregation path downstream is unchanged.
  • Never let a rubric judge what a verifier can check: criteria that can be made mechanical belong here or in a Verifier, not in a rubric line.

Steps

  1. Implement EvaluationStrategy with resources() returning ResourceProfile { models: vec![], max_context_tokens: 0, calls_per_episode: 0.0 }.
  2. Define the metric extractor as a user-supplied function over EpisodeView — the framework supplies latency, cost and usage; the user supplies anything domain-specific such as test pass count.
  3. Rank within the group by the computed number and emit Score::Ranked with group_size.
  4. Ensure the interval is meaningful or explicitly degenerate — do not fabricate a confidence interval around a deterministic measurement.
  5. Test with a workflow graded purely on test-pass-count and latency; assert the model-call counter is exactly zero.

Acceptance

  • A workflow graded purely on test-pass-count and latency, no model calls.
  • Its ResourceProfile declares zero models.

Verify

Harness: a model-call counter installed at the provider boundary, counting all purposes.

Integration testtests/it_deterministic_grader.rs:

  1. Run a workflow graded on test-pass-count and latency.
  2. Assert the model-call counter is exactly 0 across the whole run's grading phase — including any tiebreak path.
  3. Assert resources().models.is_empty() and calls_per_episode == 0.0.
  4. Assert registration passes the capacity check under a CapacityLimits with zero devices — nothing about this strategy should need a GPU.
  5. Determinism: grade the same group twice; assert identical Score::Ranked values including ordering of ties.
  6. Assert the emitted score is Ranked, so downstream aggregation is unchanged from the tournament path.
  7. Assert no fabricated confidence interval — either a real one or an explicitly degenerate one, asserted as such.

Command: cargo test -p grading deterministic

False pass:

  • Counting only judge-purpose model calls. A tiebreak issued under Purpose::Work would slip through — count all purposes.
  • Step 4 omitted: declaring the agent's model in models "since it is resident anyway" passes steps 13, then fails a capacity check it should pass and misreports the spend projection.
  • Step 5 with a group of one, where ordering is trivially stable.

Traps

  • Declaring the agent's model in models "since it is resident anyway". It then fails a capacity check it should pass, and it lies in the spend projection.
  • Calling a model for a tiebreak.

Background (not required to do this task): rust-agentic-sys.md §11.1, §11.7 · rust-agentic-task.md