4.7 KiB
4.7 KiB
T6.6 — Held-out split
| Field | Value |
|---|---|
| Phase | P6 — Learning loop |
| Size | M — 1 to 3 days |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
Goal
Partition tasks into selection and held-out sets. Promote on selection, report held-out without optimizing against it.
Facts (inlined — no spec read needed)
- Selecting on a fixed set of recorded tasks overfits to those tasks, silently: shadow scores improve while live performance does not.
- No promotion gate takes held-out as an entry criterion. A gate that reads held-out has converted it into a second selection set and left nothing measuring generalization.
- The proposer reads the selection set only. If the slow loop consumes held-out failures to generate candidates, the held-out set is contaminated through the generator instead of the selector. This leak is easy to introduce and invisible once present — which is why the audit test is the control, not a nicety.
- A widening selection-versus-held-out gap is the overfitting alarm.
- Under one-state convergence there is no competing variant whose divergence would reveal a stalled loop, so the held-out report is also the only stall detector there is. It must be produced and surfaced even when no promotion is pending.
- Held-out catches overfitting to episodes. It does not catch drift in the task mix — that is T6.7.
Steps
- Assign each
TaskIdto selection or held-out by a deterministic hash of the id, so the partition is stable and needs no stored membership list. - Expose two separate query surfaces —
selection_tasks()andheldout_tasks()— rather than one with a flag. A flag defaults wrong. - Give the proposer access to the selection surface only, at the type level if possible.
- Compute the held-out report on a schedule, independent of promotion activity, and surface the selection-versus-held-out gap as a metric.
- Write the audit test: enumerate every proposer input and every gate
evaluation, assert no held-out
TaskIdappears in either. - Keep the audit test running in CI permanently. It is the only thing that will notice the leak.
Acceptance
- An audit test asserting no held-out
TaskIdappears in proposer input, and none in any promotion-gate evaluation. This leak is invisible once present, so the test is the control. - The held-out report is emitted on a schedule, not only at a gate — asserted, since under one-state convergence it is the only stall detector there is.
Verify
Harness: an interception layer recording every TaskId that reaches the
proposer and every TaskId read during a gate evaluation. The audit is the
deliverable — it is the only thing that will ever notice this leak.
Integration test — tests/it_heldout_audit.rs:
- Partition a corpus of 1000
TaskIds. Assert the split is deterministic: recompute it in a second process and assert identical membership. - Run a full generate-and-promote cycle with interception on.
- Assert no held-out
TaskIdappears in the recorded proposer input set — intersection with the held-out set is empty. - Assert no held-out
TaskIdappears in any gate evaluation. - Assert the two query surfaces are distinct types or distinct functions —
trybuilda call that tries to pass held-out tasks into the proposer. - Schedule: advance simulated time with no promotion pending; assert the held-out report is still emitted. Under one-state convergence it is the only stall detector there is.
- Assert the selection-versus-held-out gap is emitted as a metric.
- Negative control: deliberately wire a held-out task into the proposer in a test-only build; assert the audit fails. An audit that never fails is not a control.
Command: cargo test -p loop heldout_audit
False pass:
- Steps 3–4 passing because the interception layer only sees one of several code paths into the proposer. Assert the interception is at the single chokepoint, or the audit measures nothing.
- Step 8 omitted. This is the most important one: an audit test that has never been observed failing may simply be looking in the wrong place, and the leak it guards is invisible once present.
- Step 6 omitted, so the report is computed only at gates — exactly when a stalled loop produces none, which is when it is needed.
Traps
- A single task query with an
include_heldout: bool. Someone passestrue. - Computing the report only when a promotion is pending, which is exactly when a stalled loop produces none.
Background (not required to do this task): rust-agentic-sys.md §12.1, §12.3, §12.5, §15 · rust-agentic-task.md