(tasks) add tasks for harness
This commit is contained in:
@@ -0,0 +1,134 @@
|
||||
# T6.3 — Promotion gates
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | P6 — Learning loop |
|
||||
| Size | M — 1 to 3 days |
|
||||
| Status | Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | — |
|
||||
|
||||
## Goal
|
||||
|
||||
The default three-rung ladder, closing on T5.6's sequential test, plus automatic
|
||||
rollback triggered by `Score::Capped`.
|
||||
|
||||
## Facts (inlined — no spec read needed)
|
||||
|
||||
**Default ladder — one challenger, pairwise.** Three rungs, because a graduated
|
||||
ramp is a population instrument and there is no population here:
|
||||
|
||||
| Rung | Traffic | Entry criterion |
|
||||
|---|---|---|
|
||||
| shadow | 0% | registered, validated, sandbox-clean, `ResourceProfile` fits |
|
||||
| trial | 5% | no `Core` violation on any trial episode |
|
||||
| current | 100% | sequential test crosses the accept boundary; drift check clean |
|
||||
|
||||
- **The swap at the last rung is deliberate: the challenger takes all traffic at
|
||||
once rather than ramping.** A ramp exists to limit blast radius while evidence
|
||||
accumulates, and here the evidence has already accumulated — the sequential
|
||||
test does not cross its boundary until the win rate is established at the
|
||||
configured α. Ramping after that spends traffic to re-learn what the test
|
||||
already concluded.
|
||||
- What guards the swap instead is the rollback rule, which fires on a **single**
|
||||
`Score::Capped` and does not wait for a boundary.
|
||||
- **Rollback is automatic and unconditional** on any `Score::Capped` attributed
|
||||
to the variant, or a verifier pass-rate regression beyond a configured margin.
|
||||
`Score::Capped` is the only signal for the first of those; a gate that reads a
|
||||
low *number* instead is reading something the cap exists to prevent from
|
||||
existing.
|
||||
- **Rollback is a traffic change, never a version delete.** The failed variant
|
||||
stays in the DAG with its results.
|
||||
- A `Core` violation **caps** rather than subtracts. A weighted sum lets a
|
||||
variant buy past a safety failure with speed, which is exactly what prescriptive
|
||||
rubrics exist to prevent. The cap comes from `Judge::screen`, runs before
|
||||
pairing, and yields `Score::Capped`.
|
||||
- **A capped episode is excluded from the bracket, not ranked last in it.** Left
|
||||
in, it still contributes comparisons that shape everyone else's strength, and a
|
||||
variant with one safety failure and seven strong episodes aggregates to a
|
||||
promotion.
|
||||
- **No rung reads held-out** (T6.6). A gate that reads held-out has converted it
|
||||
into a second selection set.
|
||||
|
||||
For deployments running the tournament with N challengers, the resourced ladder
|
||||
is shadow (0%) → canary (5%, beats current on selection replay, BT interval
|
||||
excludes zero) → ramp (20→50%, no `Core` violation, cost within budget,
|
||||
sequential test at α) → current (100%, sustained over N groups, drift clean).
|
||||
|
||||
## Steps
|
||||
|
||||
1. Implement the three rungs as an explicit state machine on the challenger
|
||||
record, with each entry criterion as a named predicate.
|
||||
2. Shadow entry: validation (T3.3), sandbox cleanliness (T6.5), and the
|
||||
`ResourceProfile` fit (T5.2/T5.3).
|
||||
3. Trial entry: no `Core` violation on any trial episode, read from
|
||||
`Score::Capped`.
|
||||
4. Current entry: T5.6's sequential test returns `Accept`, and T6.7's drift check
|
||||
is clean. Perform the swap **atomically** through T6.1's CAS.
|
||||
5. Rollback path: on any `Score::Capped` attributed to the variant, or a pass-rate
|
||||
regression past the margin, restore the prior version as current in a single
|
||||
CAS — no interval in which neither is live.
|
||||
6. Exclude capped episodes from the bracket **before** comparisons are issued, so
|
||||
no other episode's score is influenced by them.
|
||||
7. Assert no gate predicate reads held-out data.
|
||||
|
||||
## Acceptance
|
||||
|
||||
- A `Judge::screen` stub returning one `CoreViolation` triggers rollback **without
|
||||
human action**; the rolled-back version remains in the DAG with its results
|
||||
intact.
|
||||
- The capped episode contributed **zero comparisons**, so no other episode's score
|
||||
moved because of it.
|
||||
- The swap at `current` is **atomic — no ramp** — and a rollback immediately after
|
||||
it restores the prior version **without a gap in which neither is live**.
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** a `Judge::screen` stub returning one `CoreViolation` on demand; a
|
||||
traffic router observable at every instant; the mock judges from T5.6.
|
||||
|
||||
**Integration test** — `tests/it_promotion_gates.rs`:
|
||||
1. Drive a challenger through shadow → trial → current with a 70% mock judge.
|
||||
Assert each rung's entry criterion was evaluated and recorded.
|
||||
2. **Rollback:** fire one `CoreViolation`. Assert rollback happens **without
|
||||
human action**, and that it triggered on `Score::Capped` — assert the gate
|
||||
never reads a numeric score by instrumenting the score accessor.
|
||||
3. Assert the rolled-back version **remains in the DAG with its results intact**
|
||||
— read it back after rollback.
|
||||
4. **Capped exclusion:** grade a group containing the capped episode. Assert it
|
||||
contributed **zero comparisons**, and assert the other episodes' scores are
|
||||
byte-identical to a control run where the capped episode was absent. Ranking
|
||||
it last would change them.
|
||||
5. **Atomic swap:** sample the live-version pointer at high frequency across the
|
||||
promotion. Assert it goes 100% old → 100% new with **no intermediate
|
||||
percentage** — no ramp.
|
||||
6. **No gap:** roll back immediately after the swap. Assert every sample shows
|
||||
exactly one live version; a sample showing none is a failure.
|
||||
7. Assert no gate predicate reads held-out data (cross-check with T6.6's audit).
|
||||
|
||||
**Command:** `cargo test -p loop promotion -- --test-threads=1`
|
||||
|
||||
**False pass:**
|
||||
- Step 4 asserting only "the capped episode has no score". Leaving it in the
|
||||
bracket still lets it shape everyone else's strength, and its own score can be
|
||||
absent while it does. The **control-run comparison** is the evidence.
|
||||
- Step 5 asserting the final state only. A ramp also ends at 100%.
|
||||
- Step 6 sampled too coarsely to observe a gap. Sample from a tight loop or
|
||||
instrument the pointer swap directly.
|
||||
- A rollback test that triggers on a low score. It passes, and the cap exists
|
||||
precisely so that number does not exist.
|
||||
|
||||
## Traps
|
||||
|
||||
- A gate reading a low numeric score instead of `Score::Capped`. The cap exists
|
||||
precisely so that number does not exist.
|
||||
- Ranking a capped episode last rather than excluding it. It still shapes the
|
||||
bracket.
|
||||
- Deleting a rolled-back version. Its results are attributed to it.
|
||||
|
||||
---
|
||||
|
||||
Background (not required to do this task):
|
||||
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §11.7, §12.3, §12.5 ·
|
||||
[rust-agentic-task.md](../../../rust-agentic-task.md)
|
||||
Reference in New Issue
Block a user