6.4 KiB
T6.3 — Promotion gates
| Field | Value |
|---|---|
| Phase | P6 — Learning loop |
| Size | M — 1 to 3 days |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
Goal
The default three-rung ladder, closing on T5.6's sequential test, plus automatic
rollback triggered by Score::Capped.
Facts (inlined — no spec read needed)
Default ladder — one challenger, pairwise. Three rungs, because a graduated ramp is a population instrument and there is no population here:
| Rung | Traffic | Entry criterion |
|---|---|---|
| shadow | 0% | registered, validated, sandbox-clean, ResourceProfile fits |
| trial | 5% | no Core violation on any trial episode |
| current | 100% | sequential test crosses the accept boundary; drift check clean |
- The swap at the last rung is deliberate: the challenger takes all traffic at once rather than ramping. A ramp exists to limit blast radius while evidence accumulates, and here the evidence has already accumulated — the sequential test does not cross its boundary until the win rate is established at the configured α. Ramping after that spends traffic to re-learn what the test already concluded.
- What guards the swap instead is the rollback rule, which fires on a single
Score::Cappedand does not wait for a boundary. - Rollback is automatic and unconditional on any
Score::Cappedattributed to the variant, or a verifier pass-rate regression beyond a configured margin.Score::Cappedis the only signal for the first of those; a gate that reads a low number instead is reading something the cap exists to prevent from existing. - Rollback is a traffic change, never a version delete. The failed variant stays in the DAG with its results.
- A
Coreviolation caps rather than subtracts. A weighted sum lets a variant buy past a safety failure with speed, which is exactly what prescriptive rubrics exist to prevent. The cap comes fromJudge::screen, runs before pairing, and yieldsScore::Capped. - A capped episode is excluded from the bracket, not ranked last in it. Left in, it still contributes comparisons that shape everyone else's strength, and a variant with one safety failure and seven strong episodes aggregates to a promotion.
- No rung reads held-out (T6.6). A gate that reads held-out has converted it into a second selection set.
For deployments running the tournament with N challengers, the resourced ladder
is shadow (0%) → canary (5%, beats current on selection replay, BT interval
excludes zero) → ramp (20→50%, no Core violation, cost within budget,
sequential test at α) → current (100%, sustained over N groups, drift clean).
Steps
- Implement the three rungs as an explicit state machine on the challenger record, with each entry criterion as a named predicate.
- Shadow entry: validation (T3.3), sandbox cleanliness (T6.5), and the
ResourceProfilefit (T5.2/T5.3). - Trial entry: no
Coreviolation on any trial episode, read fromScore::Capped. - Current entry: T5.6's sequential test returns
Accept, and T6.7's drift check is clean. Perform the swap atomically through T6.1's CAS. - Rollback path: on any
Score::Cappedattributed to the variant, or a pass-rate regression past the margin, restore the prior version as current in a single CAS — no interval in which neither is live. - Exclude capped episodes from the bracket before comparisons are issued, so no other episode's score is influenced by them.
- Assert no gate predicate reads held-out data.
Acceptance
- A
Judge::screenstub returning oneCoreViolationtriggers rollback without human action; the rolled-back version remains in the DAG with its results intact. - The capped episode contributed zero comparisons, so no other episode's score moved because of it.
- The swap at
currentis atomic — no ramp — and a rollback immediately after it restores the prior version without a gap in which neither is live.
Verify
Harness: a Judge::screen stub returning one CoreViolation on demand; a
traffic router observable at every instant; the mock judges from T5.6.
Integration test — tests/it_promotion_gates.rs:
- Drive a challenger through shadow → trial → current with a 70% mock judge. Assert each rung's entry criterion was evaluated and recorded.
- Rollback: fire one
CoreViolation. Assert rollback happens without human action, and that it triggered onScore::Capped— assert the gate never reads a numeric score by instrumenting the score accessor. - Assert the rolled-back version remains in the DAG with its results intact — read it back after rollback.
- Capped exclusion: grade a group containing the capped episode. Assert it contributed zero comparisons, and assert the other episodes' scores are byte-identical to a control run where the capped episode was absent. Ranking it last would change them.
- Atomic swap: sample the live-version pointer at high frequency across the promotion. Assert it goes 100% old → 100% new with no intermediate percentage — no ramp.
- No gap: roll back immediately after the swap. Assert every sample shows exactly one live version; a sample showing none is a failure.
- Assert no gate predicate reads held-out data (cross-check with T6.6's audit).
Command: cargo test -p loop promotion -- --test-threads=1
False pass:
- Step 4 asserting only "the capped episode has no score". Leaving it in the bracket still lets it shape everyone else's strength, and its own score can be absent while it does. The control-run comparison is the evidence.
- Step 5 asserting the final state only. A ramp also ends at 100%.
- Step 6 sampled too coarsely to observe a gap. Sample from a tight loop or instrument the pointer swap directly.
- A rollback test that triggers on a low score. It passes, and the cap exists precisely so that number does not exist.
Traps
- A gate reading a low numeric score instead of
Score::Capped. The cap exists precisely so that number does not exist. - Ranking a capped episode last rather than excluding it. It still shapes the bracket.
- Deleting a rolled-back version. Its results are attributed to it.
Background (not required to do this task): rust-agentic-sys.md §11.7, §12.3, §12.5 · rust-agentic-task.md