# T6.6 — Held-out split | Field | Value | |---|---| | Phase | P6 — Learning loop | | Size | M — 1 to 3 days | | Status | Not started | | Flags | — | | Spec | inlined below | | Blocks | — | ## Goal Partition tasks into selection and held-out sets. Promote on selection, report held-out without optimizing against it. ## Facts (inlined — no spec read needed) - Selecting on a fixed set of recorded tasks overfits to those tasks, **silently**: shadow scores improve while live performance does not. - **No promotion gate takes held-out as an entry criterion.** A gate that reads held-out has converted it into a second selection set and left nothing measuring generalization. - **The proposer reads the selection set only.** If the slow loop consumes held-out failures to generate candidates, the held-out set is contaminated through the generator instead of the selector. **This leak is easy to introduce and invisible once present** — which is why the audit test is the control, not a nicety. - A widening selection-versus-held-out gap is the overfitting alarm. - Under one-state convergence there is no competing variant whose divergence would reveal a stalled loop, so **the held-out report is also the only stall detector there is**. It must be produced and surfaced even when no promotion is pending. - Held-out catches overfitting to *episodes*. It does not catch drift in the *task mix* — that is T6.7. ## Steps 1. Assign each `TaskId` to selection or held-out by a deterministic hash of the id, so the partition is stable and needs no stored membership list. 2. Expose two separate query surfaces — `selection_tasks()` and `heldout_tasks()` — rather than one with a flag. A flag defaults wrong. 3. Give the proposer access to the selection surface only, at the type level if possible. 4. Compute the held-out report on a **schedule**, independent of promotion activity, and surface the selection-versus-held-out gap as a metric. 5. Write the audit test: enumerate every proposer input and every gate evaluation, assert no held-out `TaskId` appears in either. 6. Keep the audit test running in CI permanently. It is the only thing that will notice the leak. ## Acceptance - An audit test asserting **no held-out `TaskId` appears in proposer input**, and none in any promotion-gate evaluation. This leak is invisible once present, so the test is the control. - The held-out report is **emitted on a schedule, not only at a gate** — asserted, since under one-state convergence it is the only stall detector there is. ## Verify **Harness:** an interception layer recording **every** `TaskId` that reaches the proposer and every `TaskId` read during a gate evaluation. The audit is the deliverable — it is the only thing that will ever notice this leak. **Integration test** — `tests/it_heldout_audit.rs`: 1. Partition a corpus of 1000 `TaskId`s. Assert the split is deterministic: recompute it in a second process and assert identical membership. 2. Run a full generate-and-promote cycle with interception on. 3. Assert **no held-out `TaskId`** appears in the recorded proposer input set — intersection with the held-out set is empty. 4. Assert **no held-out `TaskId`** appears in any gate evaluation. 5. Assert the two query surfaces are distinct types or distinct functions — `trybuild` a call that tries to pass held-out tasks into the proposer. 6. **Schedule:** advance simulated time with **no promotion pending**; assert the held-out report is still emitted. Under one-state convergence it is the only stall detector there is. 7. Assert the selection-versus-held-out gap is emitted as a metric. 8. Negative control: deliberately wire a held-out task into the proposer in a test-only build; assert the audit **fails**. An audit that never fails is not a control. **Command:** `cargo test -p loop heldout_audit` **False pass:** - Steps 3–4 passing because the interception layer only sees one of several code paths into the proposer. Assert the interception is at the single chokepoint, or the audit measures nothing. - Step 8 omitted. This is the most important one: an audit test that has never been observed failing may simply be looking in the wrong place, and the leak it guards is invisible once present. - Step 6 omitted, so the report is computed only at gates — exactly when a stalled loop produces none, which is when it is needed. ## Traps - A single task query with an `include_heldout: bool`. Someone passes `true`. - Computing the report only when a promotion is pending, which is exactly when a stalled loop produces none. --- Background (not required to do this task): [rust-agentic-sys.md](../../../rust-agentic-sys.md) §12.1, §12.3, §12.5, §15 · [rust-agentic-task.md](../../../rust-agentic-task.md)