Files
poimen/tasks/T5.4-deterministicgrader.md
T

89 lines
3.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# T5.4 — `DeterministicGrader`
| Field | Value |
|---|---|
| Phase | P5 — Grading |
| Size | S — under 1 day |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
## Goal
Grade on a number the user already computes — latency, cost, test pass count —
with zero model calls.
## Facts (inlined — no spec read needed)
- It exists so that **adopting the framework does not require adopting
LLM-as-judge at all.**
- Cost profile: 0 model calls per episode, 0 models resident, produces `Ranked`
on a computed number.
- Its `ResourceProfile` declares an **empty** `models` vector. That is the whole
point — it forces no residency and passes the capacity check trivially.
- It still produces `Score::Ranked`, so the aggregation path downstream is
unchanged.
- Never let a rubric judge what a verifier can check: criteria that can be made
mechanical belong here or in a `Verifier`, not in a rubric line.
## Steps
1. Implement `EvaluationStrategy` with `resources()` returning
`ResourceProfile { models: vec![], max_context_tokens: 0, calls_per_episode: 0.0 }`.
2. Define the metric extractor as a user-supplied function over `EpisodeView`
the framework supplies latency, cost and usage; the user supplies anything
domain-specific such as test pass count.
3. Rank within the group by the computed number and emit `Score::Ranked` with
`group_size`.
4. Ensure the interval is meaningful or explicitly degenerate — do not fabricate
a confidence interval around a deterministic measurement.
5. Test with a workflow graded purely on test-pass-count and latency; assert the
model-call counter is exactly zero.
## Acceptance
- A workflow graded purely on test-pass-count and latency, **no model calls**.
- Its `ResourceProfile` declares **zero models**.
## Verify
**Harness:** a model-call counter installed at the provider boundary, counting
**all** purposes.
**Integration test**`tests/it_deterministic_grader.rs`:
1. Run a workflow graded on test-pass-count and latency.
2. Assert the model-call counter is **exactly 0** across the whole run's grading
phase — including any tiebreak path.
3. Assert `resources().models.is_empty()` and `calls_per_episode == 0.0`.
4. Assert registration passes the capacity check under a `CapacityLimits` with
**zero** devices — nothing about this strategy should need a GPU.
5. Determinism: grade the same group twice; assert identical `Score::Ranked`
values including ordering of ties.
6. Assert the emitted score is `Ranked`, so downstream aggregation is unchanged
from the tournament path.
7. Assert no fabricated confidence interval — either a real one or an explicitly
degenerate one, asserted as such.
**Command:** `cargo test -p grading deterministic`
**False pass:**
- Counting only judge-purpose model calls. A tiebreak issued under
`Purpose::Work` would slip through — count all purposes.
- Step 4 omitted: declaring the agent's model in `models` "since it is resident
anyway" passes steps 13, then fails a capacity check it should pass and
misreports the spend projection.
- Step 5 with a group of one, where ordering is trivially stable.
## Traps
- Declaring the agent's model in `models` "since it is resident anyway". It then
fails a capacity check it should pass, and it lies in the spend projection.
- Calling a model for a tiebreak.
---
Background (not required to do this task):
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.1, §11.7 ·
[rust-agentic-task.md](../../../rust-agentic-task.md)