135 lines
6.4 KiB
Markdown
135 lines
6.4 KiB
Markdown
# T6.3 — Promotion gates
|
||
|
||
| Field | Value |
|
||
|---|---|
|
||
| Phase | P6 — Learning loop |
|
||
| Size | M — 1 to 3 days |
|
||
| Status | Not started |
|
||
| Flags | — |
|
||
| Spec | inlined below |
|
||
| Blocks | — |
|
||
|
||
## Goal
|
||
|
||
The default three-rung ladder, closing on T5.6's sequential test, plus automatic
|
||
rollback triggered by `Score::Capped`.
|
||
|
||
## Facts (inlined — no spec read needed)
|
||
|
||
**Default ladder — one challenger, pairwise.** Three rungs, because a graduated
|
||
ramp is a population instrument and there is no population here:
|
||
|
||
| Rung | Traffic | Entry criterion |
|
||
|---|---|---|
|
||
| shadow | 0% | registered, validated, sandbox-clean, `ResourceProfile` fits |
|
||
| trial | 5% | no `Core` violation on any trial episode |
|
||
| current | 100% | sequential test crosses the accept boundary; drift check clean |
|
||
|
||
- **The swap at the last rung is deliberate: the challenger takes all traffic at
|
||
once rather than ramping.** A ramp exists to limit blast radius while evidence
|
||
accumulates, and here the evidence has already accumulated — the sequential
|
||
test does not cross its boundary until the win rate is established at the
|
||
configured α. Ramping after that spends traffic to re-learn what the test
|
||
already concluded.
|
||
- What guards the swap instead is the rollback rule, which fires on a **single**
|
||
`Score::Capped` and does not wait for a boundary.
|
||
- **Rollback is automatic and unconditional** on any `Score::Capped` attributed
|
||
to the variant, or a verifier pass-rate regression beyond a configured margin.
|
||
`Score::Capped` is the only signal for the first of those; a gate that reads a
|
||
low *number* instead is reading something the cap exists to prevent from
|
||
existing.
|
||
- **Rollback is a traffic change, never a version delete.** The failed variant
|
||
stays in the DAG with its results.
|
||
- A `Core` violation **caps** rather than subtracts. A weighted sum lets a
|
||
variant buy past a safety failure with speed, which is exactly what prescriptive
|
||
rubrics exist to prevent. The cap comes from `Judge::screen`, runs before
|
||
pairing, and yields `Score::Capped`.
|
||
- **A capped episode is excluded from the bracket, not ranked last in it.** Left
|
||
in, it still contributes comparisons that shape everyone else's strength, and a
|
||
variant with one safety failure and seven strong episodes aggregates to a
|
||
promotion.
|
||
- **No rung reads held-out** (T6.6). A gate that reads held-out has converted it
|
||
into a second selection set.
|
||
|
||
For deployments running the tournament with N challengers, the resourced ladder
|
||
is shadow (0%) → canary (5%, beats current on selection replay, BT interval
|
||
excludes zero) → ramp (20→50%, no `Core` violation, cost within budget,
|
||
sequential test at α) → current (100%, sustained over N groups, drift clean).
|
||
|
||
## Steps
|
||
|
||
1. Implement the three rungs as an explicit state machine on the challenger
|
||
record, with each entry criterion as a named predicate.
|
||
2. Shadow entry: validation (T3.3), sandbox cleanliness (T6.5), and the
|
||
`ResourceProfile` fit (T5.2/T5.3).
|
||
3. Trial entry: no `Core` violation on any trial episode, read from
|
||
`Score::Capped`.
|
||
4. Current entry: T5.6's sequential test returns `Accept`, and T6.7's drift check
|
||
is clean. Perform the swap **atomically** through T6.1's CAS.
|
||
5. Rollback path: on any `Score::Capped` attributed to the variant, or a pass-rate
|
||
regression past the margin, restore the prior version as current in a single
|
||
CAS — no interval in which neither is live.
|
||
6. Exclude capped episodes from the bracket **before** comparisons are issued, so
|
||
no other episode's score is influenced by them.
|
||
7. Assert no gate predicate reads held-out data.
|
||
|
||
## Acceptance
|
||
|
||
- A `Judge::screen` stub returning one `CoreViolation` triggers rollback **without
|
||
human action**; the rolled-back version remains in the DAG with its results
|
||
intact.
|
||
- The capped episode contributed **zero comparisons**, so no other episode's score
|
||
moved because of it.
|
||
- The swap at `current` is **atomic — no ramp** — and a rollback immediately after
|
||
it restores the prior version **without a gap in which neither is live**.
|
||
|
||
## Verify
|
||
|
||
**Harness:** a `Judge::screen` stub returning one `CoreViolation` on demand; a
|
||
traffic router observable at every instant; the mock judges from T5.6.
|
||
|
||
**Integration test** — `tests/it_promotion_gates.rs`:
|
||
1. Drive a challenger through shadow → trial → current with a 70% mock judge.
|
||
Assert each rung's entry criterion was evaluated and recorded.
|
||
2. **Rollback:** fire one `CoreViolation`. Assert rollback happens **without
|
||
human action**, and that it triggered on `Score::Capped` — assert the gate
|
||
never reads a numeric score by instrumenting the score accessor.
|
||
3. Assert the rolled-back version **remains in the DAG with its results intact**
|
||
— read it back after rollback.
|
||
4. **Capped exclusion:** grade a group containing the capped episode. Assert it
|
||
contributed **zero comparisons**, and assert the other episodes' scores are
|
||
byte-identical to a control run where the capped episode was absent. Ranking
|
||
it last would change them.
|
||
5. **Atomic swap:** sample the live-version pointer at high frequency across the
|
||
promotion. Assert it goes 100% old → 100% new with **no intermediate
|
||
percentage** — no ramp.
|
||
6. **No gap:** roll back immediately after the swap. Assert every sample shows
|
||
exactly one live version; a sample showing none is a failure.
|
||
7. Assert no gate predicate reads held-out data (cross-check with T6.6's audit).
|
||
|
||
**Command:** `cargo test -p loop promotion -- --test-threads=1`
|
||
|
||
**False pass:**
|
||
- Step 4 asserting only "the capped episode has no score". Leaving it in the
|
||
bracket still lets it shape everyone else's strength, and its own score can be
|
||
absent while it does. The **control-run comparison** is the evidence.
|
||
- Step 5 asserting the final state only. A ramp also ends at 100%.
|
||
- Step 6 sampled too coarsely to observe a gap. Sample from a tight loop or
|
||
instrument the pointer swap directly.
|
||
- A rollback test that triggers on a low score. It passes, and the cap exists
|
||
precisely so that number does not exist.
|
||
|
||
## Traps
|
||
|
||
- A gate reading a low numeric score instead of `Score::Capped`. The cap exists
|
||
precisely so that number does not exist.
|
||
- Ranking a capped episode last rather than excluding it. It still shapes the
|
||
bracket.
|
||
- Deleting a rolled-back version. Its results are attributed to it.
|
||
|
||
---
|
||
|
||
Background (not required to do this task):
|
||
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.7, §12.3, §12.5 ·
|
||
[rust-agentic-task.md](../../../rust-agentic-task.md)
|