106 lines
4.7 KiB
Markdown
106 lines
4.7 KiB
Markdown
# T6.6 — Held-out split
|
||||
|
|
|
|||
|
|
| Field | Value |
|
|||
|
|
|---|---|
|
|||
|
|
| Phase | P6 — Learning loop |
|
|||
|
|
| Size | M — 1 to 3 days |
|
|||
|
|
| Status | Not started |
|
|||
|
|
| Flags | — |
|
|||
|
|
| Spec | inlined below |
|
|||
|
|
| Blocks | — |
|
|||
|
|
|
|||
|
|
## Goal
|
|||
|
|
|
|||
|
|
Partition tasks into selection and held-out sets. Promote on selection, report
|
|||
|
|
held-out without optimizing against it.
|
|||
|
|
|
|||
|
|
## Facts (inlined — no spec read needed)
|
|||
|
|
|
|||
|
|
- Selecting on a fixed set of recorded tasks overfits to those tasks, **silently**:
|
|||
|
|
shadow scores improve while live performance does not.
|
|||
|
|
- **No promotion gate takes held-out as an entry criterion.** A gate that reads
|
|||
|
|
held-out has converted it into a second selection set and left nothing
|
|||
|
|
measuring generalization.
|
|||
|
|
- **The proposer reads the selection set only.** If the slow loop consumes
|
|||
|
|
held-out failures to generate candidates, the held-out set is contaminated
|
|||
|
|
through the generator instead of the selector. **This leak is easy to introduce
|
|||
|
|
and invisible once present** — which is why the audit test is the control, not
|
|||
|
|
a nicety.
|
|||
|
|
- A widening selection-versus-held-out gap is the overfitting alarm.
|
|||
|
|
- Under one-state convergence there is no competing variant whose divergence
|
|||
|
|
would reveal a stalled loop, so **the held-out report is also the only stall
|
|||
|
|
detector there is**. It must be produced and surfaced even when no promotion is
|
|||
|
|
pending.
|
|||
|
|
- Held-out catches overfitting to *episodes*. It does not catch drift in the
|
|||
|
|
*task mix* — that is T6.7.
|
|||
|
|
|
|||
|
|
## Steps
|
|||
|
|
|
|||
|
|
1. Assign each `TaskId` to selection or held-out by a deterministic hash of the
|
|||
|
|
id, so the partition is stable and needs no stored membership list.
|
|||
|
|
2. Expose two separate query surfaces — `selection_tasks()` and
|
|||
|
|
`heldout_tasks()` — rather than one with a flag. A flag defaults wrong.
|
|||
|
|
3. Give the proposer access to the selection surface only, at the type level if
|
|||
|
|
possible.
|
|||
|
|
4. Compute the held-out report on a **schedule**, independent of promotion
|
|||
|
|
activity, and surface the selection-versus-held-out gap as a metric.
|
|||
|
|
5. Write the audit test: enumerate every proposer input and every gate
|
|||
|
|
evaluation, assert no held-out `TaskId` appears in either.
|
|||
|
|
6. Keep the audit test running in CI permanently. It is the only thing that will
|
|||
|
|
notice the leak.
|
|||
|
|
|
|||
|
|
## Acceptance
|
|||
|
|
|
|||
|
|
- An audit test asserting **no held-out `TaskId` appears in proposer input**, and
|
|||
|
|
none in any promotion-gate evaluation. This leak is invisible once present, so
|
|||
|
|
the test is the control.
|
|||
|
|
- The held-out report is **emitted on a schedule, not only at a gate** — asserted,
|
|||
|
|
since under one-state convergence it is the only stall detector there is.
|
|||
|
|
|
|||
|
|
## Verify
|
|||
|
|
|
|||
|
|
**Harness:** an interception layer recording **every** `TaskId` that reaches the
|
|||
|
|
proposer and every `TaskId` read during a gate evaluation. The audit is the
|
|||
|
|
deliverable — it is the only thing that will ever notice this leak.
|
|||
|
|
|
|||
|
|
**Integration test** — `tests/it_heldout_audit.rs`:
|
|||
|
|
1. Partition a corpus of 1000 `TaskId`s. Assert the split is deterministic:
|
|||
|
|
recompute it in a second process and assert identical membership.
|
|||
|
|
2. Run a full generate-and-promote cycle with interception on.
|
|||
|
|
3. Assert **no held-out `TaskId`** appears in the recorded proposer input set —
|
|||
|
|
intersection with the held-out set is empty.
|
|||
|
|
4. Assert **no held-out `TaskId`** appears in any gate evaluation.
|
|||
|
|
5. Assert the two query surfaces are distinct types or distinct functions —
|
|||
|
|
`trybuild` a call that tries to pass held-out tasks into the proposer.
|
|||
|
|
6. **Schedule:** advance simulated time with **no promotion pending**; assert the
|
|||
|
|
held-out report is still emitted. Under one-state convergence it is the only
|
|||
|
|
stall detector there is.
|
|||
|
|
7. Assert the selection-versus-held-out gap is emitted as a metric.
|
|||
|
|
8. Negative control: deliberately wire a held-out task into the proposer in a
|
|||
|
|
test-only build; assert the audit **fails**. An audit that never fails is not
|
|||
|
|
a control.
|
|||
|
|
|
|||
|
|
**Command:** `cargo test -p loop heldout_audit`
|
|||
|
|
|
|||
|
|
**False pass:**
|
|||
|
|
- Steps 3–4 passing because the interception layer only sees one of several code
|
|||
|
|
paths into the proposer. Assert the interception is at the single chokepoint,
|
|||
|
|
or the audit measures nothing.
|
|||
|
|
- Step 8 omitted. This is the most important one: an audit test that has never
|
|||
|
|
been observed failing may simply be looking in the wrong place, and the leak it
|
|||
|
|
guards is invisible once present.
|
|||
|
|
- Step 6 omitted, so the report is computed only at gates — exactly when a
|
|||
|
|
stalled loop produces none, which is when it is needed.
|
|||
|
|
|
|||
|
|
## Traps
|
|||
|
|
|
|||
|
|
- A single task query with an `include_heldout: bool`. Someone passes `true`.
|
|||
|
|
- Computing the report only when a promotion is pending, which is exactly when a
|
|||
|
|
stalled loop produces none.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
Background (not required to do this task):
|
|||
|
|
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §12.1, §12.3, §12.5, §15 ·
|
|||
|
|
[rust-agentic-task.md](../../../rust-agentic-task.md)
|