Files

89 lines
3.5 KiB
Markdown
Raw Permalink Normal View History

2026-08-17 23:05:20 -07:00
# T5.4 — `DeterministicGrader`
| Field | Value |
|---|---|
| Phase | P5 — Grading |
| Size | S — under 1 day |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
## Goal
Grade on a number the user already computes — latency, cost, test pass count —
with zero model calls.
## Facts (inlined — no spec read needed)
- It exists so that **adopting the framework does not require adopting
LLM-as-judge at all.**
- Cost profile: 0 model calls per episode, 0 models resident, produces `Ranked`
on a computed number.
- Its `ResourceProfile` declares an **empty** `models` vector. That is the whole
point — it forces no residency and passes the capacity check trivially.
- It still produces `Score::Ranked`, so the aggregation path downstream is
unchanged.
- Never let a rubric judge what a verifier can check: criteria that can be made
mechanical belong here or in a `Verifier`, not in a rubric line.
## Steps
1. Implement `EvaluationStrategy` with `resources()` returning
`ResourceProfile { models: vec![], max_context_tokens: 0, calls_per_episode: 0.0 }`.
2. Define the metric extractor as a user-supplied function over `EpisodeView`
the framework supplies latency, cost and usage; the user supplies anything
domain-specific such as test pass count.
3. Rank within the group by the computed number and emit `Score::Ranked` with
`group_size`.
4. Ensure the interval is meaningful or explicitly degenerate — do not fabricate
a confidence interval around a deterministic measurement.
5. Test with a workflow graded purely on test-pass-count and latency; assert the
model-call counter is exactly zero.
## Acceptance
- A workflow graded purely on test-pass-count and latency, **no model calls**.
- Its `ResourceProfile` declares **zero models**.
## Verify
**Harness:** a model-call counter installed at the provider boundary, counting
**all** purposes.
**Integration test**`tests/it_deterministic_grader.rs`:
1. Run a workflow graded on test-pass-count and latency.
2. Assert the model-call counter is **exactly 0** across the whole run's grading
phase — including any tiebreak path.
3. Assert `resources().models.is_empty()` and `calls_per_episode == 0.0`.
4. Assert registration passes the capacity check under a `CapacityLimits` with
**zero** devices — nothing about this strategy should need a GPU.
5. Determinism: grade the same group twice; assert identical `Score::Ranked`
values including ordering of ties.
6. Assert the emitted score is `Ranked`, so downstream aggregation is unchanged
from the tournament path.
7. Assert no fabricated confidence interval — either a real one or an explicitly
degenerate one, asserted as such.
**Command:** `cargo test -p grading deterministic`
**False pass:**
- Counting only judge-purpose model calls. A tiebreak issued under
`Purpose::Work` would slip through — count all purposes.
- Step 4 omitted: declaring the agent's model in `models` "since it is resident
anyway" passes steps 13, then fails a capacity check it should pass and
misreports the spend projection.
- Step 5 with a group of one, where ordering is trivially stable.
## Traps
- Declaring the agent's model in `models` "since it is resident anyway". It then
fails a capacity check it should pass, and it lies in the spend projection.
- Calling a model for a tiebreak.
---
Background (not required to do this task):
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.1, §11.7 ·
[rust-agentic-task.md](../../../rust-agentic-task.md)