4.7 KiB
4.7 KiB
T6.5 — Sandboxed replay
| Field | Value |
|---|---|
| Phase | P6 — Learning loop |
| Size | L — over 3 days |
| Status | Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | — |
Goal
Re-execute a task under a new workflow version with Unsafe effects denied.
Depends on T2.2's effect classes.
Facts (inlined — no spec read needed)
- "Replay against recorded episodes" means re-running the task under a new workflow version, not pushing a recorded trajectory through new logic. Trajectory replay tells you only where behaviour would first diverge, and everything after divergence is unknown — near-worthless for grading.
- Replay is not free. "0% live traffic" means no user sees the result, not that it costs nothing. N variants × M tasks is N·M full agent runs plus judge calls. Shadow is the most expensive rung, not the cheapest.
- Replay executes real tools. Re-running a workflow that pushes commits
pushes commits. Shadow execution runs in a sandbox with
Unsafeeffects denied, and a workflow whose steps cannot run sandboxed is ineligible for shadow evaluation and must say so at load time rather than at 3am. - Sandbox is layer 3 of the effect-class defence: declaration (T2.2), keyed capability (T8.3), sandbox here.
- Replay needs its own spend ceiling, separate from judging.
Steps
- Implement the sandbox as a capability restriction over the tool registry:
Unsafetools are not resolvable inside a replay context. Denial is structural, not a runtimeif. - Add the load-time eligibility check: walk the workflow's steps, resolve
each
ToolId, and reject the workflow for shadow evaluation if any resolves toUnsafe. The diagnostic names the tool. - Replay execution path: fresh runs against the recorded task input (
TaskId), under the challenger version, inside the sandbox. - Record replay runs with a
Purposedistinguishing them from live work, so metering (T8.1) separates them and their own ceiling applies. - Feed the resulting episodes to the tournament strategy — the group exists because you paid for it, so pairwise grading would waste it.
- Wire the replay ceiling: exceeding it stops replay and reports, rather than silently truncating the variant set.
Acceptance
- A workflow calling an unsafe tool is rejected for shadow evaluation at load time, with a diagnostic naming the tool.
Verify
Harness: the external side-effect ledger from T2.2, plus a workflow that
calls an Unsafe tool and one that does not.
Integration test — tests/it_sandboxed_replay.rs:
- Submit the
Unsafe-calling workflow for shadow evaluation. Assert it is rejected at load time, with a diagnostic naming the tool. - Assert the rejection happens before any run is spawned — check the run counter is 0 and the ledger is empty.
- Positive control: the safe workflow is accepted and replays successfully.
- Defence in depth: bypass the load check in a test-only path and attempt an
Unsafecall inside the sandbox anyway. Assert it is denied at call time too. The load check is the good error; the sandbox is the guarantee. - Assert replay re-executes against the recorded task input, not a recorded trajectory — plant a divergence and assert execution continues past it and produces a complete episode.
- Assert replay runs are recorded with a distinct
Purpose, and that they count against the replay ceiling, not the judging one. - Cost visibility: assert N variants × M tasks produces N·M runs, and that the count is reported — shadow is the most expensive rung and must not look free.
Command: cargo test -p loop sandboxed_replay
False pass:
- Step 1 alone. A load-time check with no runtime enforcement passes it, and any code path that skips validation then executes real effects. Step 4 is the guarantee.
- Step 5 omitted: trajectory replay produces an episode that looks complete and is meaningless after the first divergence, and no assertion about tool denial catches it.
- Charging replay to the judging ceiling — step 6. It passes every functional test and consumes the entire grading budget in production.
Traps
- Trajectory replay. It is cheaper, it looks like replay, and its output cannot be graded past the first divergence.
- Denying
Unsafeat call time instead of load time. The rejection then arrives mid-run, at 3am, after real work has been done. - Charging replay against the judging ceiling. Replay dominates and will consume the whole budget.
Background (not required to do this task): rust-agentic-sys.md §12.4, §13.2, §14.1 · rust-agentic-task.md