# T5.4 — `DeterministicGrader` | Field | Value | |---|---| | Phase | P5 — Grading | | Size | S — under 1 day | | Status | Not started | | Flags | — | | Spec | inlined below | | Blocks | — | ## Goal Grade on a number the user already computes — latency, cost, test pass count — with zero model calls. ## Facts (inlined — no spec read needed) - It exists so that **adopting the framework does not require adopting LLM-as-judge at all.** - Cost profile: 0 model calls per episode, 0 models resident, produces `Ranked` on a computed number. - Its `ResourceProfile` declares an **empty** `models` vector. That is the whole point — it forces no residency and passes the capacity check trivially. - It still produces `Score::Ranked`, so the aggregation path downstream is unchanged. - Never let a rubric judge what a verifier can check: criteria that can be made mechanical belong here or in a `Verifier`, not in a rubric line. ## Steps 1. Implement `EvaluationStrategy` with `resources()` returning `ResourceProfile { models: vec![], max_context_tokens: 0, calls_per_episode: 0.0 }`. 2. Define the metric extractor as a user-supplied function over `EpisodeView` — the framework supplies latency, cost and usage; the user supplies anything domain-specific such as test pass count. 3. Rank within the group by the computed number and emit `Score::Ranked` with `group_size`. 4. Ensure the interval is meaningful or explicitly degenerate — do not fabricate a confidence interval around a deterministic measurement. 5. Test with a workflow graded purely on test-pass-count and latency; assert the model-call counter is exactly zero. ## Acceptance - A workflow graded purely on test-pass-count and latency, **no model calls**. - Its `ResourceProfile` declares **zero models**. ## Verify **Harness:** a model-call counter installed at the provider boundary, counting **all** purposes. **Integration test** — `tests/it_deterministic_grader.rs`: 1. Run a workflow graded on test-pass-count and latency. 2. Assert the model-call counter is **exactly 0** across the whole run's grading phase — including any tiebreak path. 3. Assert `resources().models.is_empty()` and `calls_per_episode == 0.0`. 4. Assert registration passes the capacity check under a `CapacityLimits` with **zero** devices — nothing about this strategy should need a GPU. 5. Determinism: grade the same group twice; assert identical `Score::Ranked` values including ordering of ties. 6. Assert the emitted score is `Ranked`, so downstream aggregation is unchanged from the tournament path. 7. Assert no fabricated confidence interval — either a real one or an explicitly degenerate one, asserted as such. **Command:** `cargo test -p grading deterministic` **False pass:** - Counting only judge-purpose model calls. A tiebreak issued under `Purpose::Work` would slip through — count all purposes. - Step 4 omitted: declaring the agent's model in `models` "since it is resident anyway" passes steps 1–3, then fails a capacity check it should pass and misreports the spend projection. - Step 5 with a group of one, where ordering is trivially stable. ## Traps - Declaring the agent's model in `models` "since it is resident anyway". It then fails a capacity check it should pass, and it lies in the spend projection. - Calling a model for a tiebreak. --- Background (not required to do this task): [rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.1, §11.7 · [rust-agentic-task.md](../../../rust-agentic-task.md)