Files
poimen/tasks/T6.5-sandboxed-replay.md

4.7 KiB
Raw Permalink Blame History

T6.5 — Sandboxed replay

Field Value
Phase P6 — Learning loop
Size L — over 3 days
Status Not started
Flags
Spec inlined below
Blocks

Goal

Re-execute a task under a new workflow version with Unsafe effects denied. Depends on T2.2's effect classes.

Facts (inlined — no spec read needed)

  • "Replay against recorded episodes" means re-running the task under a new workflow version, not pushing a recorded trajectory through new logic. Trajectory replay tells you only where behaviour would first diverge, and everything after divergence is unknown — near-worthless for grading.
  • Replay is not free. "0% live traffic" means no user sees the result, not that it costs nothing. N variants × M tasks is N·M full agent runs plus judge calls. Shadow is the most expensive rung, not the cheapest.
  • Replay executes real tools. Re-running a workflow that pushes commits pushes commits. Shadow execution runs in a sandbox with Unsafe effects denied, and a workflow whose steps cannot run sandboxed is ineligible for shadow evaluation and must say so at load time rather than at 3am.
  • Sandbox is layer 3 of the effect-class defence: declaration (T2.2), keyed capability (T8.3), sandbox here.
  • Replay needs its own spend ceiling, separate from judging.

Steps

  1. Implement the sandbox as a capability restriction over the tool registry: Unsafe tools are not resolvable inside a replay context. Denial is structural, not a runtime if.
  2. Add the load-time eligibility check: walk the workflow's steps, resolve each ToolId, and reject the workflow for shadow evaluation if any resolves to Unsafe. The diagnostic names the tool.
  3. Replay execution path: fresh runs against the recorded task input (TaskId), under the challenger version, inside the sandbox.
  4. Record replay runs with a Purpose distinguishing them from live work, so metering (T8.1) separates them and their own ceiling applies.
  5. Feed the resulting episodes to the tournament strategy — the group exists because you paid for it, so pairwise grading would waste it.
  6. Wire the replay ceiling: exceeding it stops replay and reports, rather than silently truncating the variant set.

Acceptance

  • A workflow calling an unsafe tool is rejected for shadow evaluation at load time, with a diagnostic naming the tool.

Verify

Harness: the external side-effect ledger from T2.2, plus a workflow that calls an Unsafe tool and one that does not.

Integration testtests/it_sandboxed_replay.rs:

  1. Submit the Unsafe-calling workflow for shadow evaluation. Assert it is rejected at load time, with a diagnostic naming the tool.
  2. Assert the rejection happens before any run is spawned — check the run counter is 0 and the ledger is empty.
  3. Positive control: the safe workflow is accepted and replays successfully.
  4. Defence in depth: bypass the load check in a test-only path and attempt an Unsafe call inside the sandbox anyway. Assert it is denied at call time too. The load check is the good error; the sandbox is the guarantee.
  5. Assert replay re-executes against the recorded task input, not a recorded trajectory — plant a divergence and assert execution continues past it and produces a complete episode.
  6. Assert replay runs are recorded with a distinct Purpose, and that they count against the replay ceiling, not the judging one.
  7. Cost visibility: assert N variants × M tasks produces N·M runs, and that the count is reported — shadow is the most expensive rung and must not look free.

Command: cargo test -p loop sandboxed_replay

False pass:

  • Step 1 alone. A load-time check with no runtime enforcement passes it, and any code path that skips validation then executes real effects. Step 4 is the guarantee.
  • Step 5 omitted: trajectory replay produces an episode that looks complete and is meaningless after the first divergence, and no assertion about tool denial catches it.
  • Charging replay to the judging ceiling — step 6. It passes every functional test and consumes the entire grading budget in production.

Traps

  • Trajectory replay. It is cheaper, it looks like replay, and its output cannot be graded past the first divergence.
  • Denying Unsafe at call time instead of load time. The rejection then arrives mid-run, at 3am, after real work has been done.
  • Charging replay against the judging ceiling. Replay dominates and will consume the whole budget.

Background (not required to do this task): rust-agentic-sys.md §12.4, §13.2, §14.1 · rust-agentic-task.md