4.7 KiB
4.7 KiB
T8.1 — Metering
| Field | Value |
|---|---|
| Phase | P8 — Operability |
| Size | M — 1 to 3 days |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
Goal
Attribute every model call to (TenantId, RunId, Purpose) and enforce separate
ceilings for judging and replay.
Facts (inlined — no spec read needed)
- A framework that spends users' model budget on grading must account for it.
- Every model call — agent, judge, or proposer — is attributed to
(TenantId, RunId, Purpose), wherePurposedistinguishes work from grading from replay. - Per-tenant ceilings on grading spend are enforced at the group boundary, where group size is known and a tournament can be skipped or downsampled before it starts — not mid-tournament.
- Replay is the expensive one and needs its own ceiling separate from judging. N variants × M tasks is N·M full agent runs plus judge calls.
- Grading cost is reported next to grading value. A loop that costs more than the work it grades may still be worth running; it should never be an unpleasant discovery.
- A skipped run must reach terminal
UngradedwithBudgetExhausted, not rest inGrading. A cost ceiling that also strands storage is a bill that arrives twice, landing on the tenants least able to absorb it. - The grading spend ratio ceiling itself is unset — metered here, policy elsewhere.
Steps
- Define
Purpose { Work, Grading, Replay }and thread it through every model call site. No default — the call site knows. - Record usage per call keyed
(TenantId, RunId, Purpose): tokens in/out, model, and cost if a price table is configured. - Define per-tenant ceilings: one for grading, a separate one for replay.
- Check the grading ceiling at the group boundary. Over the ceiling: skip or
downsample, record
UngradedReason::BudgetExhausted, and transition the run to terminalUngraded(T5.11). - Never let a ceiling check touch the agent execution path. Work continues regardless of grading spend.
- Emit a report pairing grading cost with grading value per tenant.
- Test: drive a tenant over its grading ceiling and assert agent throughput is unchanged.
Acceptance
- A tenant exceeding its grading ceiling has tournaments skipped with a recorded reason, and agent work continues unaffected.
- The skipped runs reach terminal
Ungradedrather than resting inGrading.
Verify
Harness: a model provider wrapper recording (TenantId, RunId, Purpose, tokens) for every call, plus a tenant configured with a zero grading ceiling
and a generous replay ceiling (and the reverse).
Integration test — tests/it_metering.rs:
- Run a full cycle: agent work, grading, replay. Assert every recorded call
carries a
Purpose, and that the three purposes are all present. Assert no call is attributed to a default or unknown purpose. - Ceiling separation: with grading ceiling = 0 and replay ceiling high, assert grading is skipped and replay still runs. Then invert and assert the opposite. One shared ceiling fails one of these.
- Skip happens before spend: assert model calls under
Purpose::Gradingis exactly 0 when over the ceiling — the check is at the group boundary, before any call. - Assert agent throughput and latency are unchanged while the tenant is over its grading ceiling; compare against a baseline run.
- Assert skipped runs reach terminal
Ungraded { BudgetExhausted }and that T4.4'sis_reduciblereturns true for them. - Assert the cost-versus-value report is produced per tenant.
- Assert attribution survives a cold re-fold — the usage lives in the log, not only in state.
Command: cargo test -p operability metering
False pass:
- Step 3 asserting only the recorded reason. Enforcing the ceiling after the calls produces the correct reason and spends the money anyway.
- Step 5 omitted: this is the "bill arrives twice" case — the tenant is capped on
spend and then also accumulates storage forever because the run never leaves
Grading. - Step 2 omitted: a single ceiling passes every other assertion here, then replay consumes the whole budget in production.
Traps
- One shared ceiling for judging and replay. Replay dominates and consumes it.
- Checking the ceiling mid-tournament, which pays for half a group and produces nothing.
- Leaving budget-skipped runs in
Grading. The cost ceiling then also strands storage.
Background (not required to do this task): rust-agentic-sys.md §10.3, §11.6, §12.4, §14.1 · rust-agentic-task.md