# Rust Agentic System — Design A framework for building distributed agents that record what they did, verify it, grade it, and improve from it. Three things distinguish this from an agent library: - **The episode is a first-class durable artifact**, not a log line. It survives crashes, supports rewind, and is the input to every learning mechanism. - **Learning is built in and pluggable.** A default reinforcement loop ships working; the grading schema, the rubric and the judge are all replaceable. - **The workflow is data, not code.** Users define workflows in their own format and version them; the framework executes and grades them without recompiling. Previous revision of this document described a bespoke single-tenant observer embedded in one specific agent. That framing is gone. What survives is the durability model, the state-machine discipline, and the list of mistakes worth not repeating. --- ## 1. Principles **The framework cannot break the agent it runs.** Observation, verification and grading failures degrade the record, never the work. Any code path where a grader can fail into an agent's execution is a defect. **The record is grounded.** Renderers and graders state only facts the episode holds. An invented fact produces a lesson about something that never happened. **Verification decides, grading explains.** Verification returns ground truth. Grading attributes cause and ranks. Collapse them and the system grades its own homework. **Nothing derived is authoritative.** Materialized state is a cache of the log. If it cannot be dropped and rebuilt byte-identically, it has hidden inputs and that is a bug. **Defaults ship working; every default is a port.** A user who wants the built-in behaviour writes no code. A user who wants their own writes an impl, not a fork. --- ## 2. Layering The central structural decision, and the one the previous revision got wrong. ``` ┌──────────────────────────────────────────────────────────┐ │ DOMAIN — user-defined, data, versioned, hot-swappable │ │ workflow definition · rubrics · verifiers · tools │ └──────────────────────────────────────────────────────────┘ │ executed / graded by ┌──────────────────────────────────────────────────────────┐ │ KERNEL — framework-owned, compiled, exhaustively typed │ │ attempt lifecycle · event log · intents · branches │ │ scheduling · tournament · partitioning · tenancy │ └──────────────────────────────────────────────────────────┘ ``` **Kernel states are closed.** An attempt is `Pending → Running → {Succeeded, Failed, TimedOut, Indeterminate}`. That enum is exhaustive, matched at compile time, and users cannot extend it. Everything the durability and learning machinery reasons about lives here. **Domain states are open.** A workflow declares its own step names, ordering and transitions as data. The kernel validates the declaration, then executes it. Adding a domain state is a config change. The previous revision declared one machine and called it normative, then celebrated that adding a state was a compile error at every call site. Correct for a bespoke tool, fatal for a framework — a user defining their own workflow would have to fork and recompile. The fix is not to loosen the kernel. It is to stop conflating the two. ```rust /// Kernel. Closed. Users never extend this. #[derive(Clone, Copy, PartialEq, Eq)] pub enum AttemptState { Pending, Running, Succeeded, Failed, /// Crashed mid-side-effect; recovery could not determine what happened (§8.4). Indeterminate, /// Stopped by decision, with no dispatched intent outstanding. Distinct /// from `Indeterminate`: nothing external is in doubt (§5.1). Cancelled, TimedOut, } /// Domain. Open. Declared by the workflow, validated at load. pub struct StepState(SmolStr); ``` `Indeterminate` is a kernel state because only the kernel knows about intents. It cannot be expressed in a user workflow and must not be collapsed into `Failed` — the two demand different operator responses. --- ## 3. Identity and tenancy Multi-tenant from the first commit. Retrofitting a tenant key through a schema, a partition scheme and a blob store is a rewrite, and the previous revision had no tenant concept at all. ```rust pub struct TenantId(Uuid); pub struct WorkflowId(SmolStr); // logical workflow, stable across versions pub struct WorkflowVersion(Blake3Hash); // content hash of the definition pub struct StepId(SmolStr); // stable across versions — see §4.3 pub struct TaskId(Blake3Hash); // the comparison group key — see §11.5 pub struct RunId(Ulid); // one execution of one workflow pub struct BranchId(u32); // rewind fork — see §8.5 pub struct AttemptNo(u32); pub struct Lsn(u64); // per (run, branch) sequence pub struct GroupEpoch(u32); // comparison-group generation — see §11.6 ``` `Ulid` for `RunId`: lexicographically sortable by creation time, which makes range scans over recent runs a prefix scan rather than a secondary index. **Every key is tenant-prefixed.** Not "most". A single unprefixed table is a cross-tenant read waiting to happen, and it will be found by a customer rather than by us. ```rust pub struct Scoped { pub tenant: TenantId, pub inner: T } ``` Tables key on `Scoped<_>`. The type makes an unscoped access a compile error rather than a review comment. **Blobs are namespaced per tenant even though they are content-addressed.** Global deduplication of prompt and output blobs is tempting — identical system prompts across tenants are common — and it is a leak. A shared blob means one tenant's storage accounting depends on another's, and a hash becomes an oracle for "does anyone else have this content". Deduplicate within a tenant, never across. --- ## 4. Workflow definition The workflow is data. The framework provides a validated intermediate representation and a parser trait; a format is a plugin. ### 4.1 The IR ```rust pub struct WorkflowDef { pub id: WorkflowId, pub schema: SchemaVersion, pub steps: Vec, pub transitions: Vec, pub rubric: RubricDef, pub budget: BudgetDef, } pub struct StepDef { /// Author-assigned, stable across versions. See §4.3. pub id: StepId, pub kind: StepKind, pub tools: ToolSelector, pub verify: Vec, pub retry: RetryPolicy, pub timeout: Duration, } pub enum StepKind { Model { prompt: PromptTemplate, effort: ReasoningEffort }, Tool { tool: ToolId, args: ArgTemplate }, Parallel { branches: Vec, join: JoinPolicy }, Conditional { on: Predicate, then: StepId, otherwise: Option }, SubWorkflow { workflow: WorkflowId, version: VersionSelector }, } ``` `WorkflowVersion` is the Blake3 hash of the canonicalized IR — not of the source text. Two YAML files differing only in key order produce the same version, which is what makes "did this change affect results" answerable. ### 4.2 Formats are plugins ```rust pub trait WorkflowFormat: Send + Sync { fn extensions(&self) -> &[&str]; fn parse(&self, src: &[u8]) -> Result; } ``` Ship YAML and JSON. A user wanting a DSL, Starlark, or a database row implements the trait. Validation and canonicalization live in the kernel and run on the IR, so a new format inherits every check without reimplementing one. ### 4.3 `StepId` stability is the user's contract Credit assignment (§11.7) attributes outcomes to steps across workflow versions. That requires a step identity that survives edits — insert a step at position 2 and every positional index shifts, but `StepId` does not. The framework cannot infer this. It is a documented obligation on the workflow author, enforced by three checks at load time: - `StepId` unique within a version. - On a version bump, report added, removed and retained ids. A version that retains no ids from its parent is almost certainly a renumbering accident and is rejected unless explicitly marked as a rewrite. - `StepId` is opaque to the framework. Never parsed, never ordered, never assumed numeric. ### 4.4 Sub-workflows and version pinning `SubWorkflow` pins by `VersionSelector`: `Exact(hash)` or `Latest`. `Latest` resolves **once, at run spawn**, and the resolved hash is recorded. A run whose sub-workflow version can change mid-execution is a run whose results attribute to nothing. Recursion depth is bounded by the kernel and the cycle is detected at load, not at execution. --- ## 5. Execution model ### 5.1 The two machines ``` KERNEL — per attempt, closed Pending ──► Running ──┬──► Succeeded │ ├──► Failed │ ├──► TimedOut │ ├──► Cancelled ├──► Cancelled └──► Indeterminate └──► TimedOut DOMAIN — per run, declared by WorkflowDef whatever the user wrote, validated as a DAG with explicit loop bounds ``` The legal set, exhaustively — this table is normative and the transition function matches it arm for arm: | From | To | When | |---|---|---| | `Pending` | `Running` | admitted | | `Pending` | `Cancelled` | run cancelled before the attempt started | | `Pending` | `TimedOut` | queue deadline passed before admission | | `Running` | `Succeeded` | completed | | `Running` | `Failed` | completed with an error | | `Running` | `TimedOut` | step deadline passed | | `Running` | `Cancelled` | run cancelled, **no intent was `Dispatched`** (§8.4) | | `Running` | `Indeterminate` | an intent was `Dispatched` and did not resolve | | terminal | — | nothing leaves a terminal state | **`Cancelled` is a kernel state and is not `Indeterminate`.** Cancellation is a decision; indeterminacy is an unknown. Collapsing them was tempting because §6 drops a tool call at its next await point and the tool may have been mid-something — but that "may" is exactly what §8.4's `Dispatched` record answers. If no intent was dispatched, nothing external happened and the attempt is cleanly `Cancelled`. Only a dispatched-and-unresolved intent earns `Indeterminate`. Getting this wrong is not cosmetic: §15 alerts on `Indeterminate` count and expects near-zero, so routing every cancelled attempt there converts the alert into background noise and it stops being read. A run advances through domain steps. Each step execution is one or more kernel attempts. Retry creates **attempt N+1** and never mutates attempt N — this makes "did the retry do better, and why" answerable, and it is what makes log replay idempotent for free. ### 5.2 Run lifecycle ``` spawn │ v Scheduled ─────────────────┐ │ admitted │ v │ Running ⇄ Suspended ──────┤ │ all steps terminal │ cancel v │ Verifying ─────────────────┤ ◄── rests here while verifier │ │ futures are outstanding v │ Verified{pass|fail} │ │ │ v v Grading ──────────────► Cancelled ● │ ◄── may wait for a tournament group to fill (§11.6) │ ├──► Graded ● ──┐ │ ├──► Archived ● └──► Ungraded ● ──┘ ``` **Cancel is reachable from every non-terminal state**, not only from `Scheduled`. A run cancelled mid-step is the ordinary case — it is what a user clicking stop does — and a lifecycle offering cancel only before admission describes a system nobody would ship. Cancelling during `Verifying` or `Grading` is rarer and still legal; the work is done and the record stands, the framework just stops spending on judging it. `Verifying` and `Grading` are states the run **rests in**, not synchronous branches. The previous implementation collapsed `Verifying` by resolving pass/fail inside `verify()`, which blocked async verifiers, mid-run UI, and the snapshot barrier in §10.2 — one collapsed state, three blocked features. **`Ungraded` is terminal and sits beside `Graded`, not below it.** Grading can legitimately end without a score: G = 1 with no group to join (§11.6), a group that closed on timeout without this run, or a tenant over its grading ceiling (§14). Those runs are finished. Without a terminal state saying so they sit in `Grading` forever, and since §10.3 gates retention on grading having ended, they also become permanently irreducible — unbounded storage growth landing precisely on the low-volume tenants §11.6 exists to accommodate, and on the tenants who hit a cost ceiling, which is the worst possible pairing. `Ungraded` carries its `UngradedReason` (§11.1) so the dashboard can state why rather than showing a gap. `Suspended` is new and load-bearing for distribution: a run awaiting human approval or a webhook must release its worker. A run that holds an executor slot across a human decision does not scale past a handful of concurrent runs. ### 5.3 Concurrency shape | Scope | Parallel | Why | |---|---|---| | across runs | unbounded | no shared state | | steps within a run | serial by default | step N+1 reads N's output | | `Parallel` step branches | fan-out/join | declared explicitly in the IR | | attempts within a step | strictly serial | a retry needs the prior failure | | verifiers for one attempt | fan-out/join | independent checks | | tournament group | join | needs the whole group (§11.6) | The serial spine is `(TenantId, RunId)`. Parallelism lives between runs and inside declared fan-out. Nothing else may interleave. --- ## 6. Runtime **Tokio.** This reverses the previous revision, and the reversal is a direct consequence of the goal change. The prior choice was `asupersync` — structured concurrency, capability-secure, cancel-correct, with a deterministic test lab. Genuinely better primitives, and it was correct when this was a component embedded inside one agent that already used it. As a framework that users embed, it fails on one axis that outweighs the rest: **no tokio compatibility means users cannot use the ecosystem.** No `sqlx`, `rdkafka`, `aws-sdk`, `tonic`, `axum`, `reqwest`, or object-store clients. For a distributed framework those are not optional dependencies; they are the distribution layer. The prior revision already conceded this in its "crossing the runtime boundary" section, treating tokio interop as an exception for two clients. Under the framework goal, that boundary is the common case, and a design whose exception path is the main path is the wrong design. What is lost, and how it is recovered: | asupersync gave | Recovered by | |---|---| | Regions — structural task-tree cancellation | `TaskTracker` + `CancellationToken` from `tokio-util`, one tracker per run, enforced by a `RunScope` guard that refuses detached spawns | | Cancel Protocol — work actually stops | `CancellationToken` selected against at every await in kernel code; user tool calls get a hard timeout. A dropped call resolves `Cancelled` if no intent reached `Dispatched`, `Indeterminate` if one did (§5.1) — the intent record, not the drop, is what decides | | `Cx` — explicit capability passing | An explicit `Ctx` struct threaded through every call. Never `task_local!` for anything causal — that is the `AsyncLocalStorage` mistake in different clothing | | The Lab — deterministic schedules | `turmoil` for network partition and latency simulation; `loom` for the lock-free bits; `tokio::time::pause` for time. Weaker than a seeded scheduler; sufficient with discipline | The residual risk is honest: tokio's cancellation is cooperative, so a `select!`-dropped future stops at its next await and not before. Kernel code must never hold a lock or a half-applied state across an await that can be cancelled. This is enforceable by review and by the `RunScope` guard; it is not enforced by the compiler the way Regions did. **An `asupersync` backend stays possible** behind a runtime trait, feature-gated, for embedding in agents that already use it. Not built until someone needs it, and not on the default path. --- ## 7. Storage ports Two deployment modes, one set of ports. ```rust #[async_trait] pub trait EventLog: Send + Sync { /// The whole of §8.3 in one call: append the records, apply the derived /// state, advance the consumer position, enqueue the outbox — all four or /// none. Returns the assigned LSNs. async fn commit(&self, batch: CommitBatch) -> Result>; async fn read(&self, key: BranchKey, from: Lsn, limit: usize) -> Result>; async fn put_checkpoint(&self, key: BranchKey, upto: Lsn, state: &[u8]) -> Result<()>; /// Newest checkpoint at or below `upto`. Restart folds forward from here; /// `None` means fold from LSN 0 (§8.6). async fn latest_checkpoint(&self, key: BranchKey, upto: Lsn) -> Result>; /// Committed-but-unshipped export intents, in `(BranchKey, Lsn)` order. /// Read by the relay (§9.3), never from the execution path. async fn drain_outbox(&self, tenant: TenantId, limit: usize) -> Result>; async fn ack_outbox(&self, shipped: &[(BranchKey, Lsn)]) -> Result<()>; } pub struct CommitBatch { /// Carries the tenant. Never a separate parameter beside a key that /// already holds one — two sources for one fact is one too many. pub key: BranchKey, pub records: Vec, // LSNs assigned by the implementation pub state: Vec, pub position: Option, pub outbox: Vec, } #[async_trait] pub trait BlobStore: Send + Sync { async fn put(&self, tenant: TenantId, content: Bytes) -> Result; async fn get(&self, tenant: TenantId, r: &BlobRef) -> Result>; /// Reduction (§8.6) and tenant deletion (§3) both require this. A store /// that cannot delete cannot honour either, and both are obligations. async fn delete(&self, tenant: TenantId, r: &BlobRef) -> Result<()>; } ``` **`commit` is one method rather than four because atomicity is the contract.** A port exposing `append` alone puts the other three writes of §8.3 outside the transaction, which is the durability guarantee gone — and gone invisibly, since each write individually succeeds. The port must be able to express the strongest thing the implementation promises, or the abstraction quietly weakens it. Same reasoning behind `latest_checkpoint` and `BlobStore::delete`: a checkpoint that can be written and not read is an optimization that cannot be used, and §8.6's retention path is unimplementable without a delete. Everything is `async`. The previous revision declared the hot-path store synchronous because the local implementation was a B-tree, then documented in the same file that a network-backed implementation could not honour the signature. That is a port finished while already known to be unimplementable. An async signature over a local call costs a negligible poll; a sync signature over a network call is impossible. The port-completeness failure above is the same mistake one level up: a signature that cannot express what the caller needs is not finished either. | Mode | Log + state | Blobs | Coordination | |---|---|---|---| | **Embedded** — single binary, no services | `redb` | `redb` table | in-process | | **Distributed** — multi-node, multi-tenant | Postgres | S3-compatible | Postgres advisory locks, or Redis if leases dominate | `redb` remains the right embedded engine: pure Rust, ACID, MVCC, stable file format, no server. It uses copy-on-write shadow paging rather than a WAL, so a torn write cannot corrupt the file — it simply does not take effect. Commits must be `Durability::Immediate`; the enum is `#[non_exhaustive]`, so set it explicitly rather than relying on the default. Avoid `sled` — years at 0.34 beta with known space amplification. **Embedded mode is a first-class product**, not a test harness. A user must be able to `cargo add` this, run an agent, and get durability and grading with zero infrastructure. That constraint is what keeps the ports honest. --- ## 8. Durability ### 8.1 Two requirements, one solved by the engine **Crash-atomicity** — a crash must not leave half-written state. `redb`'s shadow paging and Postgres transactions both handle this. **History** — rewind, resume-from-failure, and "what did this look like at step 3" need the *sequence* of transitions. Neither engine keeps one. This is ours. The log is a table on top of the engine, not a competitor to it. Because the engine's transactions are atomic, appending to the log and applying the state happen together or not at all: no torn records, no redo/undo pass, no checkpoint-consistency problem. ### 8.2 The record ```rust pub struct LogRecord { pub key: BranchKey, // (TenantId, RunId, BranchId) pub lsn: Lsn, /// Wire-format version of `event`. Never removed, never reused. See §8.7. pub schema: SchemaVersion, pub at: Timestamp, pub event: WorkEvent, } pub struct BranchKey { pub tenant: TenantId, pub run: RunId, pub branch: BranchId } ``` `BranchId` is in the key, not implied. The previous revision keyed the log on `(RunId, Lsn)` while a later section claimed state was "keyed by branch as well as attempt" — a contradiction that made forking unimplementable as written. **LSNs are per branch, not global.** A global counter serializes every run through one atomic. The ordering contract is per-run total order with nothing promised across runs, so a per-branch sequence is exactly as strong as the contract requires, contention-free, and keeps the log partitionable by run. ### 8.3 The commit protocol ```rust let txn = db.begin_write()?; { let mut log = txn.open_table(EVENT_LOG)?; let mut state = txn.open_table(RUN_STATE)?; let mut position = txn.open_table(CONSUMER_POSITION)?; let mut outbox = txn.open_table(OUTBOX)?; log.insert((branch_key, lsn), &record)?; // append: the durable fact state.insert((branch_key, attempt_no), &attempt)?; // apply: the derived view position.insert(Scoped::new(tenant, stream), pos)?; // advance: where to resume outbox.insert((branch_key, lsn), &intent)?; // relay separately (§13.3) } txn.commit()?; // all four, or none ``` **Every key here carries `BranchKey` or a `Scoped<_>` (§3).** The outbox is the one that invites the mistake: an outbox keyed on `Lsn` alone reads naturally and is wrong, because LSNs are per branch (§8.2), so a bare LSN collides across every branch of every run of every tenant. The same key shape also gives the relay a defined order — per `BranchKey`, ascending `Lsn` — which is the only ordering §9.2 promises. Commit per event, not per run. A projection that accumulates in memory and writes at run end loses the whole run on a crash. Episodes are small and append-mostly; one fsync per transition is cheap next to model latency. ### 8.4 Write-ahead intent Appending after the fact records history. It does not make a failed step resumable, because the dangerous window is *before* the record exists. ``` 1. append Intent{Pending} ──► commit + fsync ◄── window A: crash here, the call was never issued 2. append Intent{Dispatched} ──► commit + fsync 3. perform the call ◄── window B: crash here, the call may have landed 4. append outcome, Intent{Committed} ──► commit + fsync ``` Three records, not two. Two records cannot separate the windows: a crash before the call and a crash after it both leave a lone `Pending` with no outcome, which makes every interrupted effect maximally suspicious and pushes recoverable work into `Indeterminate`. The second fsync buys the distinction. It is paid only on steps with external effects and is small next to the call it guards. On restart, the last committed intent state classifies the crash: | Last state | Meaning | Resolution | |---|---|---| | `Pending` | the call was never issued | retry freely, whatever the effect class | | `Dispatched` | the call *may* have been issued | by effect class, below | | `Committed` | outcome already recorded | nothing to do | For a `Dispatched` intent, resolution is by declared effect class: | Class | Recovery | |---|---| | `Idempotent` | retry with the same idempotency key; the provider deduplicates | | `Queryable` | ask the provider whether the request id landed, then complete or retry | | `Unsafe` | **never auto-retry.** Attempt becomes `Indeterminate`, operator is notified | The third row is the honest one. Some effects cannot be made safe by any protocol. The intent log's value is converting an invisible unknown into a recorded one — `Indeterminate` is a fact a grader and an operator can both use; a silently retried payment is not. Intents always commit `Immediate`. Batching them defeats their only purpose. Tool calls need this more than model calls. A model call is a metered read; a tool call writes files, pushes commits, and touches the world. ### 8.5 Rewind is a fork ``` lsn 0 ─ 1 ─ 2 ─ 3 ─ 4 ─ 5 ─ 6(failed) branch 0, retained └─ 0 ─ 1 ─ 2 ─ ... branch 1, forked at (0, 3) ``` A rewind allocates a new `BranchId` and starts its LSNs at zero, recording the fork point. Nothing is removed. Same rule as attempt N+1 never mutating attempt N, for the same reason: **the discarded branch is the evidence**. Truncating it destroys the failure that motivated the rewind, which is what the learning loop exists to consume. - Queries default to the live branch; grading may read all of them. - Only the live branch is exported. - Rewinding past a `Committed` intent with a non-idempotent effect is a compensation problem, not a replay problem. The log records what happened; it cannot un-happen it. Flag rather than pretend. ### 8.6 Retention: reduce, then tier Three mechanisms, escalating. Note the vocabulary: **`Archived`** is a run state; **reduction** is the token-budget operation; **tiering** is the move to cold storage. The previous revision called two of these "compaction" and the collision was guaranteed to confuse implementers. **Checkpoints.** A materialized state snapshot tagged with its LSN. Restart folds forward from the newest one. An optimization only — deleting every checkpoint costs startup time and nothing else. **Reduction at a token ceiling.** Retention is measured in **tokens**, because tokens are the currency of every downstream consumer: what a replay costs, what fits in a judge's window, what an export is billed at. Default ceiling 200k per run; per-tenant configurable. On self-hosted weights that ceiling is **not free to choose** — it is bounded by `max_context_tokens / 2` from §14.2, because a pairwise judge reads two episodes into one KV cache and two 200k episodes do not fit on any single device. Where the derived bound is tighter than the configured ceiling, the bound wins and the judge reads a further-reduced view. A retention number set without reference to the hardware that must read it is a number that will be discovered wrong at the first judge call. Loss order is fixed: | Kept | Reduced | Dropped | |---|---|---| | every transition record | blob bodies → summary blob | raw text on dead branches | | context partitions, tool info, usage | dead-branch attempts → summary | | | verifier results, grades | | | **Reduction never rewrites a blob and never edits the log.** Blobs are content-addressed (§3), so replacing a body under its existing ref makes the ref a lie; and repointing the log at a new ref is the history rewrite §8.7 forbids. Reduction instead writes the summary as a *new* blob, appends a `Reduced{original: BlobRef, summary: BlobRef}` event to the log, and only then deletes the original body. The reduction is a later fact about an earlier record, not a change to it. `BlobStore::get` on the original ref returns `None`, and the fold knows why and what stands in its place — so drop-and-re-fold from LSN 0 still yields byte-identical state, which it would not if the mapping lived only in the projection. The transition sequence always survives. What reduces is *text*, because it dominates token count and is the only part with a cheap lossy representation. A reduced episode can still be graded, attributed and structurally rewound — it just cannot be replayed verbatim. **Tiering.** Terminal, reduced runs move to cold storage. Moved, not copied, with local rows deleted only after the remote commit acknowledges. ### 8.7 Log schema evolution An append-only log plus an evolving event enum is a trap the previous revision walked straight past. Two years of records, one `WorkEvent` variant renamed, and the "drop derived state and re-fold" property is silently gone. Rules, from record one: - Every record carries `SchemaVersion`. Written always, even at v1. - `WorkEvent` is `#[non_exhaustive]`; decode is version-dispatched. - **Variants are never removed or repurposed.** Deprecated variants stay decodable forever. Storage is cheap; an undecodable log is not. - Migrations are **upcasters** — `fn upcast(vN) -> vN+1` — applied on read, never by rewriting history. Rewriting an append-only log is a contradiction. - A round-trip test per version, asserting that a stored fixture of every historical version still folds to the expected state. This test is the whole guarantee; without it the rules are aspirational. --- ## 9. Distribution ### 9.1 Partitioning ``` ingest ──► execute ──► verify ──► tournament ──► aggregate ──► decide │ │ │ │ │ │ (T,Run) (T,Run) (T,Run) (T,Task,Class, (T,Variant) (T,Workflow) Epoch) ▲ ▲ ▲ shuffle 1 shuffle 2 single writer ``` Three keys, two shuffles, one single-writer stage. Everything up to verification keys on `(TenantId, RunId)` and is embarrassingly parallel. Grading is a **join** — a comparison group must be co-located — and it keys on `(TenantId, TaskId, VerifierOutcome, GroupEpoch)`: the outcome class because §11.4 brackets only within one, and the epoch because a closed group never reopens for a late arrival (§11.6). The decide stage is single-writer per `(TenantId, WorkflowId)`. Two schedulers adjusting traffic allocation concurrently produce an allocation neither holds. A compare-and-swap on a generation counter is sufficient; no lock service needed at this size. ### 9.2 Ordering and delivery Per-run total order, nothing promised across runs. Downstream consumers must therefore partition by run key, and the framework's broker adapters set the partition key from `(TenantId, RunId)` — never from a correlation id, which collapses unrelated runs onto one partition while splitting single runs across several. Delivery is at-least-once. Exactly-once is achieved at the fold, not in transport: `(BranchKey, Lsn)` is the natural idempotency key, so a redelivered record is a no-op insert. ### 9.3 Outbox The framework never calls a broker from the execution path. Export intent is written in the same transaction as the state (§8.3); a separate relay reads committed intents and ships them. This makes export restartable, keeps a broker outage from stalling a run, and is the only pattern that survives a crash between "state committed" and "event published". ### 9.4 Leases and work distribution Runs are claimed by workers under a lease with a TTL. A worker that dies has its runs reclaimed after expiry. Lease renewal is a heartbeat on the run record, and `Suspended` runs (§5.2) release their lease entirely rather than heartbeating through a human's lunch break. Fencing tokens on every lease. A partitioned worker that resumes must not write under an expired claim, and a monotonic fence in the run record makes that a rejected write rather than a silent double-execution. --- ## 10. Verification Verification returns ground truth. It is a port with a fail-closed contract. ```rust #[async_trait] pub trait Verifier: Send + Sync { fn id(&self) -> VerifierId; /// Any error, panic or timeout MUST resolve to `Fail`. A verifier that /// throws or hangs can never report `Pass`. async fn verify(&self, cx: &VerifierCtx) -> VerifierOutcome; } ``` ### 10.1 Verifiers need the inputs, not just the ids A verifier seeing only identifiers can answer "did it work". Answering "did the agent have what it needed" requires the context and the prompt. ```rust pub struct VerifierCtx { pub run: RunId, pub step: StepId, pub attempt: AttemptNo, /// Frozen on entry to `Verifying`. Never mutates while a verifier holds it. pub episode: EpisodeView, /// Lazy. Verifiers needing no text never pay for it. pub blobs: Arc, pub deadline: Instant, } pub struct AttemptView { pub step: StepId, pub attempt: AttemptNo, pub state: AttemptState, pub workflow_version: WorkflowVersion, pub context: Option, // identifiers, small, inline pub prompt: Option, // large, by reference pub output: Option, pub tools: Option, pub usage: Option, } ``` Context partitions hold *identifiers* — packed, available-but-not-packed, dropped — and are small enough to inline. Prompts and outputs are large and go by reference. Laziness matters: a verifier that shells out and checks an exit code needs none of this, and making every verifier carry prompt text penalizes the common case while blowing broker payload limits. Include **failed attempts**. "Retried three times because context was missing X" is the learning signal; shipping only the winning attempt discards it. ### 10.2 The snapshot barrier Verifiers read after work completes, so state must stop moving beneath them. Entry to `Verifying` freezes the view. This is why `Verifying` must be a real resting state and not a synchronous branch. Worth naming what the barrier actually defends against, because the obvious answer is wrong. It is **not** a concurrent retry: `Verifying` is entered only when every step is terminal (§5.2), so no attempt can still be running. The real mutators are the ones that arrive from outside the run's own execution: - a **rewind** (§8.5) forking a new `BranchId` while verifiers hold a view of the old one; - a **cancel** (§5.2), now legal from `Verifying`; - **recovery** resolving a `Dispatched` intent left by an earlier crash, which writes an outcome into an attempt a verifier is already reading. Each of these is a write to the run while verifiers are mid-flight, and each is rare enough to be missed in testing and ordinary enough to happen in production. ### 10.3 Retention ordering Reduction (§8.6) must not outrun verification or grading. Eligibility is **`Graded`, `Ungraded` or `Archived`** — the condition is *grading has terminated*, not *grading succeeded*. Never a step-level finish timestamp: a step can finish, be reduced, and then run-level verification finds nothing. `Ungraded` belongs in that set for a reason worth stating plainly, since the tighter-looking `Graded`-only rule is the one that gets written. A run that never gets a score — G = 1, a group that closed without it, a tenant over its grading ceiling — is finished, and gating retention on `Graded` alone leaves it irreducible forever. The tenants that hit this are the low-volume ones and the cost-capped ones: the two populations least able to absorb unbounded storage, and the two least likely to have anyone watching for it. --- ## 11. Grading and the learning loop The default loop ships working. Every component is a port. ### 11.1 Ports Grading is **strategy-pluggable, and the strategy declares what hardware it needs before it is allowed to run.** That second half is not a detail: the strategies below differ by more than an order of magnitude in model calls and in VRAM, and a deployment that cannot afford one must be told at load time rather than by an OOM at 3am. ```rust #[async_trait] pub trait EvaluationStrategy: Send + Sync { fn id(&self) -> StrategyId; /// Declared before any work is admitted. Validated against §14.2's limits /// at load; a strategy whose profile does not fit is rejected by name. fn resources(&self) -> ResourceProfile; async fn evaluate(&self, cx: &EvalCtx) -> Result>; } pub struct ResourceProfile { /// Models this strategy calls. One entry means it runs on the agent's /// already-resident model and forces no swap (§14.2). pub models: Vec, /// Largest single-call context the strategy will request. A pairwise judge /// reads two episodes, so this is roughly twice an episode budget. pub max_context_tokens: u32, /// Model calls per episode evaluated, for capacity planning and for §14's /// spend projection. pub calls_per_episode: f32, } #[async_trait] pub trait Grader: Send + Sync { /// Produce comparable scores for a group of episodes. async fn grade(&self, group: &Group, rubric: &RubricDef) -> Result>; } #[async_trait] pub trait Judge: Send + Sync { /// Relative comparison only. Deliberately cannot return an absolute score. async fn compare(&self, a: &EpisodeView, b: &EpisodeView, r: &RubricDef) -> Result; /// Unary, and separate from `compare` for a structural reason: a `Core` /// violation caps an episode on its own terms, not relative to an opponent /// (§11.7). Runs before pairing. async fn screen(&self, e: &EpisodeView, r: &RubricDef) -> Result>; } pub enum Verdict { A, B, Draw } pub struct CoreViolation { pub criterion: RubricCriterionId, pub evidence: BlobRef, } pub enum Score { /// Default path (§11.3): one verdict against the current reference, plus /// the running record the sequential test consumes. Relative { against: RunId, verdict: Verdict, record: WinRecord }, /// Bradley-Terry strength as a delta from control, with its interval and /// the group size that produced it (§11.2, §11.4). Only the tournament /// strategy produces this. Ranked { strength: f64, interval: (f64, f64), group_size: u32 }, /// A `Core` violation caps the episode. Carries the violations and no /// number, so there is nothing for an aggregate to average past. Capped { violations: Vec }, /// No comparison was possible (§11.6). A reason, never a neutral score. Ungraded { reason: UngradedReason }, } ``` `Judge::compare` returning `Verdict` rather than `f64` is the schema decision that matters most, and §11.2 is why. `Score` is deliberately a sum rather than a number with flags. A capped episode and an ungraded one are not low scores; they are different kinds of answer, and a type that can represent them as numbers will eventually have them averaged into a promotion gate by code that meant no harm. **The strategy catalogue.** Cost is per episode evaluated, on a group of eight. | Strategy | Model calls / episode | Models resident | Produces | Default | |---|---|---|---|---| | `DeterministicGrader` | 0 | 0 | `Ranked` on a computed number | — | | **`PairwiseSequential`** (§11.3) | **1–2** | **1, the agent's** | `Relative` | **yes** | | `TournamentGrader` (§11.4) | 3–5 | 1 | `Ranked` with intervals | opt-in | | `ReplayTournament` (§12.4) | 3–5 **plus N full agent runs** | 1 | `Ranked` across variants | opt-in | The default is `PairwiseSequential` because it is the only one whose cost does not scale with how much you want to know. The tournament's `(G/2)·log₂(G)` comparisons buy a full ranking with composable strengths; that is genuinely more information, and a deployment that can afford it should turn it on. Most cannot, and a framework whose default path assumes a grading budget larger than the work being graded will simply be run with grading disabled — which is the outcome this whole section exists to avoid. `DeterministicGrader` remains for users whose quality signal is a number they already compute — latency, cost, test pass count. It exists so that adopting the framework does not require adopting LLM-as-judge at all. ### 11.2 Absolute scores do not work here Three failure modes, all of which this system would hit: **Calibration drift.** A judge asked for 0..1 returns different numbers for the same episode across weeks and model versions. Drift is indistinguishable from a variant trend, so promotion decisions fire on grader noise. **Weak discrimination.** Four competent episodes all score 0.8. No gradient, no selection pressure, and the loop reports "nothing beats control" because the grader cannot resolve them — not because they are equivalent. **Saturation.** The one that kills the loop outright. As workflows improve, pass rate approaches 100% and pass/fail carries zero information; absolute rubric scores saturate identically. A tournament cannot saturate — better candidates just make it harder. Group-normalized relative scores also give something absolute scores cannot: **cross-task comparability** — but only through a shared anchor, and that qualification is load-bearing. A Bradley-Terry fit identifies strengths only up to an additive constant *within one connected comparison graph*. Two groups on different tasks are disjoint graphs, so their strengths sit on unlinked scales, and averaging them directly commits the same error this section accuses point tallies of, one layer further in. **The anchor is control.** §12.1 gives control a traffic share in every allocation, so every group contains at least one control episode; the fit pins control to zero and every other strength is read as a delta from it. A variant's aggregate is then a mean of like-for-like deltas rather than a mean of incomparable scales — which is what makes the per-variant stage in §9.1 sound rather than approximate. A group that happens to contain no control episode is not aggregatable: it still grades its own members and is still worth reading, it just does not feed the aggregate, and it is recorded as such rather than folded in on the assumption that scales match. ### 11.3 Pairwise sequential (default) The system holds **one current version and at most one challenger**, and grading answers one question: has the challenger accumulated enough evidence to replace the current one? Not "rank these eight", not "what is each episode worth" — a single accept/reject that converges toward one state. ``` current version ──► episode ──┐ ├──► Judge::compare ──► verdict challenger ──► episode ──┘ │ (same TaskId) ▼ accumulate into WinRecord │ ┌────────────────┼────────────────┐ ▼ ▼ ▼ accept continue reject challenger becomes keep sampling discard, keep the current current ``` One comparison per episode. Against the tournament's `(G/2)·log₂(G)` pairs doubled for both orderings, that is 24 judge calls dropping to 8 at G = 8, and the saving grows with G rather than shrinking. **The reference is the current version's recorded episode on the same `TaskId`.** It is already on disk — no re-run, no group to fill, no `GroupEpoch` timeout. Where the task has never been seen before there is nothing to compare against, and that case degrades per §11.6 rather than being papered over. **Stopping is a sequential test, not a fixed sample.** Verdicts accumulate into a likelihood ratio against boundaries set by α, β and the smallest win-rate shift worth acting on; the test stops as soon as a boundary is crossed. A clearly better challenger is accepted in far fewer comparisons than a fixed-n design would spend, and a clearly worse one is rejected early instead of running to completion. This is the mechanism that makes the cost adaptive: cheap decisions cost little, close decisions cost more, and nothing costs the worst case by default. **Draws are recorded and excluded from the ratio.** A tie carries no evidence either way about which is stronger, so folding it in as half a win manufactures information. But a *high draw rate* is itself a result — it says the challenger is not meaningfully different — so the test also rejects on a draw-rate ceiling rather than sampling forever toward a boundary it will never reach. **Order alternates rather than doubling.** §11.4's both-orderings rule pays 2× on every comparison to cancel position bias. Here the challenger takes position A on even-numbered comparisons and position B on odd ones: bias cancels across the sequence instead of within each pair, at no extra cost. Order-consistency is still measured, on a sampled fraction of comparisons, and reported as the grader's error bar exactly as before. **A fixed reference is a cacheable prefix.** The same reference episode leads every comparison in a decision, so it can be cached across calls. The tournament cannot do this — shuffling into brackets makes every pair a novel combination by design, which is the point of the shuffle and the reason nothing caches. What this gives up, stated rather than discovered later: **parallel exploration and composable strengths.** One challenger at a time is hill-climbing, which is slower to find improvements and can settle in a local optimum with nothing in the loop able to report that it has. And a `Relative` score answers "better than the current version on this task" — it is not a strength that composes across tasks the way §11.2's anchored Bradley-Terry deltas do. Deployments that can afford the tournament get real information for the money; this is the right default, not the better mechanism. ### 11.4 Tournament grading (opt-in) Not the default path — §11.3 is — and not a fallback either. This is the strategy to enable when episodes are **already co-present at no extra cost**, or when a deployment can afford full rankings. Two cases qualify naturally: - **Attempt tournaments** (§11.5). The attempts of one step are on disk the moment a retry happens. No agent runs to pay for, and this is the only source of per-step credit the system has. - **Replay** (§12.4), where N variants are executed against one task deliberately. The episodes exist because you paid for them; grading them pairwise would waste the group you bought. Everything below is unchanged in substance from when it was the default. What changed is the claim: it is more information per episode, at three to five times the model calls, and that trade is now the user's to make explicitly rather than one the framework makes for them. ``` [ G comparable episodes for one task, one outcome class ] │ ▼ [ shuffle into brackets ] ◄── shuffling also cancels position bias │ ▼ ┌─────────────────────────────────┐ │ Swiss pairing, log₂(G) rounds │ ◄── rubric-guided Judge, relative only └────────────────┬────────────────┘ │ ▼ [ Bradley-Terry fit over all comparisons ] │ ▼ [ strength per episode + confidence interval ] ``` Swiss rather than round-robin: O(G log G) instead of O(G²). Eight episodes is twelve comparisons rather than twenty-eight. Swiss rather than single elimination because we want a full ranking, not a champion — eliminated candidates still carry signal. **Bradley-Terry, not point-tally z-scores.** Accumulating tournament points and normalizing to mean 0 / sd 1 within a group is the obvious approach and it is statistically wrong for our aggregate: small groups produce extreme z-scores, so a variant that appears in many small groups wins on variance rather than quality. A Bradley-Terry fit over the pairwise outcomes yields a strength parameter with a real confidence interval, which composes correctly across groups of different sizes and feeds the sample gate in §12.3 directly. Three details of that fit are decisions, not implementation freedom. Textbook Bradley-Terry does none of them, and each failure looks like a result rather than a bug. **Draws need a draw model.** Plain BT is binary and has no tie term, so the draws this section deliberately permits have nowhere to go. Dropping them discards the judge's most confident statements; splitting each half-and-half fabricates two comparisons that never happened and tightens the interval on invented evidence. Use the **Davidson extension** — one additional tie parameter fit alongside the strengths. **Small groups separate.** At G = 4..8 an episode that wins every comparison drives the unpenalized maximum-likelihood estimate to infinite strength. That is the extreme-score failure §11.2 rejects, arriving through the fit instead of through z-scores. A weakly-informative prior on the strengths — equivalently, a penalized likelihood — is **required, not tuning**. It is the mechanism that turns "won all three of its comparisons" into a wide interval rather than an unbounded one, and without it the reassuring sentence about small groups producing wide intervals is simply false. **A strength alone means nothing.** The fit is identified only up to an additive constant (§11.2), so an interval on a single raw strength is an interval on an arbitrary origin. Pin control to zero and report every strength as a delta against it. The §12.3 canary gate says "BT interval excludes zero" — zero is control, and it is only zero because it was pinned there. **Draws are permitted.** A judge forced to separate two equivalent episodes invents a distinction, and the optimizer will chase the invention. Draws cost gradient; forced choices cost correctness. **Both orderings are judged.** Pairwise judges have position bias. Order-consistency is recorded per comparison, and the disagreement rate is the grader's own error bar — it belongs in the report next to the scores, and an inconsistent judge should widen the sample gate rather than silently promote. ### 11.5 Groups: where they come from A tournament needs a comparison group, and production runs are one-shot on tasks that mostly never repeat. This is the binding constraint on the whole learning loop. `TaskId` is the group key: a content hash of the task input **before any workflow touches it**. Hashing the prompt does not work — the workflow changes the prompt by construction, which is the entire point of a variant. | Source | Group | Cost | Needs | |---|---|---|---| | **Attempts** | attempts 1..N of one step, identical context | free, already recorded | nothing | | **Recurring tasks** | runs sharing a `TaskId` over time | free, slow to fill | `TaskId` | | **Replay** | one task re-executed under N variants | N full runs | blobs, sandbox | Attempt tournaments first. Retries are already on disk and attempt N+1 never mutates attempt N, so the attempts of one step are a group on identical context — the cheapest per-step credit signal available. A failure and its successful retry are **not** a judge comparison. §11.8 forbids that pairing and the verifier has already ordered it; asking a judge which is better asks it to re-decide what verification decided. The pair is consumed *structurally* instead: what differed between attempt N and N+1 — context partition, tool selection, prompt — is attributed to the `StepId` as the change that turned a fail into a pass. The judge sees only same-outcome attempts, where the question it answers is "which failure got further" — which no verifier can answer. **`TaskId` cannot be backfilled.** A run recorded without one is permanently ungroupable, which is why it is required at spawn with no `Default` and no `From`. ### 11.6 Low-volume degradation Most of this problem is a tournament problem, and §11.3 does not have it: a pairwise comparison needs one partner, and the current version's recorded episode on that `TaskId` is already on disk. Volume stops mattering the moment a task recurs even once. What survives is the genuinely irreducible case — **a `TaskId` never seen before**. There is nothing to compare against, because nothing else has done this task. That is not a degradation to engineer around; it is the first observation of a new task, and it becomes the reference for the next one. 1. **Novel `TaskId`** — no comparison. Verifier outcome and deterministic dimensions (cost, latency, tool efficiency) still recorded; no relative score. Reported as `Score::Ungraded { reason: NoReference }`, never as a neutral score. The episode is retained as the reference for that `TaskId`. 2. **Attempt tournaments** — available to any workflow that retries, regardless of volume, and unaffected by either of the above. 3. **Grading budget exhausted** — a tenant over its §14 ceiling reports `Ungraded { BudgetExhausted }`. Not a quality signal; a spend signal. The remainder of this section applies **only when the tournament strategy is enabled** (§11.4), where a group must genuinely fill: 4. **Group completeness trigger** — a group closes on quorum *or* on a timeout, grading whatever arrived, with group size attached to the confidence interval. **A closed group is immutable, and the next episode starts a new one.** This is the question the trigger raises and does not answer on its own: a group closed on timeout at G = 3, then a fourth episode with the same `TaskId` arrives an hour later. Re-opening and re-fitting is the wrong answer — strengths from that group have already been published, aggregated, and possibly acted on by a promotion gate, and a fit that silently changes underneath a decision already made is worse than a small group. So the group key carries a generation: `(TenantId, TaskId, VerifierOutcome, GroupEpoch)`. Closure increments the epoch; late arrivals accumulate into the next one. The cost is honest and should be stated rather than discovered — a low-volume tenant with a long inter-arrival time gets a run of G = 1 epochs, each reported `Ungraded { InsufficientGroup }`. That is a real signal about their volume, and the fix is a longer timeout, which is a tenant-level setting and a tradeoff between waiting and grading, not a bug in the trigger. A tenant whose loop never engages must see that in the dashboard as a stated reason. Silent no-op is the worst outcome: it looks like a working loop that finds no improvements. ### 11.7 Rubrics and credit assignment ```rust pub enum RubricLayer { /// Mandatory. A violation caps the result regardless of everything else. Core, /// Anti-gaming. Written explicitly against known exploits. Prescriptive, /// Context-specific, user-authored, weighed rather than binding. Contextual, } ``` `Core` violations **cap** rather than subtract. A weighted sum lets a variant buy past a safety failure with speed, which is the exact failure prescriptive rubrics exist to prevent. The cap needs somewhere to live, and `Verdict` is the wrong place — a `Core` violation is a fact about one episode, not about a pair, and a judge asked to express it through a comparison can only rank the offender lower. It comes from `Judge::screen` (§11.1) instead, which runs before pairing and yields `Score::Capped`. A capped episode is **excluded from the bracket, not ranked last in it**: left in, it still contributes comparisons that shape everyone else's strength, and a variant with one safety failure and seven strong episodes aggregates to a promotion. Never let a rubric judge what a verifier can check. Every criterion that can be made mechanical should be a `Verifier`, not a rubric line — deterministic, cheap, and not subject to judge drift. Per-step credit attributes a group's outcome to `StepId`s, which is why §4.3's stability contract is load-bearing rather than cosmetic. ### 11.8 Grading never overrides the verifier Episodes are bracketed **within** a verifier outcome class, never across. A verified pass beats a verified fail by definition and that pairing is never shown to a judge. Ranking failures against each other is not wasted work: "failed at step 2" versus "failed at step 7 after recovering twice" is exactly the signal a pass rate cannot see. This binds attempt groups too (§11.5), which is where the rule is easiest to break: a step's failed attempt and its successful retry sit side by side on disk and look like a free comparison. They are a free *credit* signal and not a comparison at all. The bracketing rule has no exemptions — if a pairing crosses an outcome class, it is evidence for attribution, never input to a judge. --- ## 12. Optimization loops Two loops at different clock speeds. The fast loop **selects** among existing workflow versions; the slow loop **generates** new ones. ### 12.1 Shape ``` ┌──── GENERATE (slow, human-gated by default) ─────┐ │ failure evidence ──► propose ──► challenger │ │ ▲ │ │ └────────┼──────────────────────────────┼──────────┘ │ │ register — at most one ┌────────┼──── SELECT (fast) ───────────┼──────────┐ │ │ ┌── allocation ────────────┘ │ │ │ │ current 95% · challenger 5% │ │ │ └────┬─────── │ │ │ │ spawn — pin WorkflowVersion │ │ │ ▼ │ │ │ run ──► verify ──► compare vs │ │ │ current (§11.3) │ │ │ │ │ │ └──────────────────────────────┤ │ │ ▼ │ │ sequential test boundary │ │ │ │ │ │ accept │ │ reject │ │ ▼ ▼ │ │ challenger discard, │ │ becomes current keep current │ └───────────────────────────────────────────────────┘ ``` **One current version, at most one challenger.** The loop converges toward a single state rather than maintaining a population. This is the change that makes everything else affordable: no N-way traffic split, no per-variant aggregation across groups, no allocation state to contend over, and one comparison per episode instead of a bracket. The cost is exploration. A single challenger at a time is hill-climbing — it finds improvements more slowly than a population would, and it can sit in a local optimum indefinitely with nothing in the loop able to say so. §12.5's held-out report is the only instrument that will notice, which makes it more important here than it was under the population design, not less. **Multi-variant selection remains available** for deployments that can afford it: enable the tournament strategy (§11.4), allow N challengers, and §12.3's full rung ladder applies with Bradley-Terry aggregation across groups. The machinery is the same; what changes is how many versions are live at once and which statistical object closes the decision. The slow loop fires when the fast loop **rejects a challenger without finding a replacement** — a trigger, not a timer. ### 12.2 Versions form a DAG ``` v1 ────┬───► v2 ───┐ control │ └──► v4 (merge) └───► v3 ───────┘ ``` Content-addressed, parent-pointered, never edited. Editing a version in place destroys every result already attributed to it. ### 12.3 Promotion gates Every rung needs a criterion. The previous revision drew the ladder and stated a rule for only the first rung. **Default ladder — one challenger, pairwise (§11.3).** Three rungs, because a graduated ramp is a population instrument and there is no population here: | Rung | Traffic | Entry criterion | |---|---|---| | shadow | 0% | registered, validated, sandbox-clean, `ResourceProfile` fits (§14.2) | | trial | 5% | no `Core` violation on any trial episode | | current | 100% | sequential test crosses the accept boundary; drift check clean | The swap at the last rung is deliberate and worth naming: the challenger takes all traffic at once rather than ramping. A ramp exists to limit blast radius while evidence accumulates, and here the evidence has already accumulated — the sequential test does not cross its boundary until the win rate is established at the configured α. Ramping after that spends traffic to re-learn what the test already concluded. What guards the swap instead is the rollback rule below, which fires on a single `Score::Capped` and does not wait for a boundary. **Resourced ladder — N challengers, tournament (§11.4).** For deployments that enabled the tournament strategy and can carry multiple live versions: | Rung | Traffic | Entry criterion | |---|---|---| | shadow | 0% | registered, validated, sandbox-clean, profile fits | | canary | 5% | beats current on **selection** replay, BT interval excludes zero | | ramp | 20→50% | no `Core` violation, cost within budget, sequential test at α | | current | 100% | sustained over N groups, drift check clean | Rollback is automatic and unconditional on any `Score::Capped` attributed to the variant, or a verifier pass-rate regression beyond a configured margin. `Score::Capped` is the only signal for the first of those (§11.1); a gate that reads a low *number* instead is reading something the cap exists to prevent from existing. Rollback is a traffic change, never a version delete — the failed variant stays in the DAG with its results. ### 12.4 Replay is re-execution "Replay against recorded episodes" means **re-running the task under a new workflow version**, not pushing a recorded trajectory through new logic. Trajectory replay tells you only where behaviour would first diverge, and everything after divergence is unknown — near-worthless for grading. Two consequences the prior revision missed: **Replay is not free.** "0% live traffic" means no user sees the result, not that it costs nothing. N variants × M tasks is N·M full agent runs plus judge calls. Shadow is the most expensive rung, not the cheapest. **Replay executes real tools.** Re-running a workflow that pushes commits pushes commits. Shadow execution runs in a **sandbox with `Unsafe` effects denied**, and a workflow whose steps cannot run sandboxed is ineligible for shadow evaluation and must say so at load time rather than at 3am. ### 12.5 Held-out set, and the leak Selecting on a fixed set of recorded tasks overfits to those tasks, silently: shadow scores improve while live performance does not. Partition tasks into a selection set and a held-out set. Promote on selection, report held-out without optimizing against it, and treat a widening gap as the overfitting alarm. Concretely: **no rung in §12.3 takes held-out as an entry criterion.** A gate that reads held-out has converted it into a second selection set and left nothing measuring generalization. **The proposer reads the selection set only.** The slow loop consumes failure evidence to generate candidates; if it reads held-out failures, the held-out set is contaminated through the generator instead of the selector. This leak is easy to introduce and invisible once present. Held-out catches overfitting to *episodes*. It does not catch drift in the *task mix* — that needs a separate distribution check on incoming `TaskId` characteristics over time. --- ## 13. Extensions, tools, and trust ### 13.1 Capability-gated hostcalls Tools register through capability-gated hostcalls. The gate is where effect class is declared, and a tool that cannot state whether it is safe to retry does not register. ```rust pub struct ToolRegistration { pub id: ToolId, /// No Default. The author is the only party who knows this. pub effects: EffectClass, pub caps: CapabilitySet, pub timeout: Duration, } /// Each class carries what its §8.4 recovery path actually needs. A bare /// discriminant would let a tool claim `Idempotent` while withholding the one /// thing that makes the claim actionable. pub enum EffectClass { Idempotent { key: KeyDerivation }, Queryable { lookup: RequestIdLookup }, Unsafe, } ``` Tool calls appear in the episode as **first-class steps**. The hostcall boundary already knows the identity, arguments and capabilities; recording them as opaque invocations discards information the runtime is holding anyway. ### 13.2 Declaration is not enforcement A framework cannot trust a user-supplied `EffectClass`. A tool declared `Idempotent` that is not will be retried after an indeterminate crash, and the damage is the user's data. Three layers, in order of strength: 1. **Declaration** — required, recorded, auditable. 2. **Keyed capability** — `Idempotent` is a claim about *retry*, not about abstaining from writes. Denying an `Idempotent` tool network and filesystem writes would deny the recovery path in §8.4, which retries the write under an idempotency key. So the restriction is on the *shape* of the write, not its existence: an `Idempotent` tool declares how its key derives from its arguments (§13.1), the kernel derives and supplies that key on every call, and a write issued without it is denied. A tool that cannot derive a stable key cannot be `Idempotent` — which is the same claim as before, now refused at registration rather than discovered after a double-charge. 3. **Sandbox** — shadow and replay execution deny `Unsafe` effects outright. Layer 2 is the one that makes layer 1 more than paperwork. ### 13.3 The relay Export runs from the outbox, out of process, with its own retry and its own failure domain. A user's broker being down is not an agent outage. --- ## 14. Metering, cost, and capacity ### 14.1 Metering A framework that spends users' model budget on grading must account for it. - Every model call — agent, judge, or proposer — is attributed to `(TenantId, RunId, Purpose)` where purpose distinguishes work from grading from replay. - Per-tenant ceilings on grading spend, enforced at the group boundary in §9.1 where group size is known and a tournament can be skipped or downsampled before it starts. - Grading cost is reported next to grading value. A loop that costs more than the work it grades may still be worth running; it should never be an unpleasant discovery. Replay is the expensive one (§12.4) and needs its own ceiling separate from judging. ### 14.2 Model capacity is a hard constraint, not a budget Metering counts tokens after the fact. On self-hosted weights the binding limit arrives earlier and harder: **VRAM**, and the cost of moving weights in and out of it. A model that is not resident cannot be inferred against, and making it resident means evicting something else and paying a load measured in tens of seconds. A design that treats "call the judge model" as equivalent in cost to "call the agent model" is wrong by two orders of magnitude on this hardware. The limits are declared, not discovered: ```rust pub struct CapacityLimits { /// Per-device VRAM this framework may use. Not the card's total — leave /// headroom for anything else sharing the device. pub vram_bytes_per_device: u64, pub devices: u32, /// Hard cap on simultaneously resident models across all devices. pub max_resident_models: u32, /// Never evictable. The agent's model belongs here; if it can be evicted, /// grading can stall agent work, which §1 forbids outright. pub pinned: Vec, } pub struct ModelProfile { pub id: ModelId, pub weights_bytes: u64, /// KV cache cost per token at the deployed dtype and parallelism. The term /// that decides how long a context may be — see below. pub kv_bytes_per_token: u64, /// Devices this model spans under tensor parallelism. pub devices_required: u32, } ``` **Residency invariant**, checked at load and before any admission: ``` sum(weights of resident models) + peak_concurrent_kv ≤ vram_bytes_per_device × devices resident_model_count ≤ max_resident_models ``` **One resident model is the default configuration.** The agent's model is pinned; `Judge` and the §12 proposer run on that same model. §11.3's `ResourceProfile` declares a single `ModelId` precisely so that the default grading path adds no resident model and forces no swap. A strategy naming a second model is legal and is rejected at load unless the invariant still holds with both resident — never by swapping between them per call, which is the failure mode this section exists to prevent. **Context length is a VRAM quantity, and this bites hardest on the judge.** A pairwise comparison reads two episodes, so its context is roughly twice an episode budget. Working the invariant backwards gives the ceiling: ``` max_context_tokens = (vram_bytes_per_device × devices − sum(weights)) / kv_bytes_per_token ``` Put numbers on it, because the result is not marginal. At GQA fp16, per-token KV runs roughly 0.13 MB for an 8B-class model and 0.33 MB for a 70B-class one. Two episodes at §8.6's 200k-token retention ceiling is 400k tokens of context: **52 GB of KV cache at 8B, 131 GB at 70B**, before weights. Neither fits an 80 GB device. A judge reading two full-ceiling episodes is not expensive — it is impossible. Three consequences, all forced rather than chosen: - **The retention ceiling must be derived from this, not set beside it.** §8.6's 200k default is a token-budget number that was picked without reference to any device. The reduced episode is what the judge reads, so the reduction target is `max_context_tokens / 2`, and where that is smaller than the retention ceiling, the judge reads a further-reduced view. - **Admission control, not backpressure.** Work whose `ResourceProfile` does not fit the current residency is refused at admission with the limit named. Queuing it would stall behind an eviction that §1 does not permit. - **Grading yields to agent work.** When both contend for the same resident model, agent inference wins and grading queues. A framework that lets a judge call delay the work it is judging has inverted its own first principle. Distributed GPUs change the arithmetic, not the rule. `devices_required` expresses tensor parallelism across cards; `max_resident_models` is a fleet-wide count, so two nodes each holding the agent model are two resident instances, not one. Residency is per device, and a model resident on node A does not make node B's runs admissible. --- ## 15. Observability The framework observes agents; it must also be observable. - Kernel state transitions as metrics, tagged by tenant and workflow version. - Lag on every stage boundary in §9.1. Tournament and reduction backlogs are the two that grow silently. - `Indeterminate` attempt count as a first-class alert. It should be near zero; a nonzero rate means either a crash loop or a misdeclared effect class. This threshold only holds because `Cancelled` is a separate state (§5.1) — route cancellations here and the alert has a noisy floor, which is the same as not having it. - `Ungraded` run count by reason (§11.6), separately from `Graded`. A tenant whose runs are mostly `InsufficientGroup` has a loop that is not engaging, and that reads as healthy on any dashboard that only counts failures. - Judge order-inconsistency rate (§11.4) as a grader-health metric, measured on a sampled fraction of comparisons under §11.3's alternating-order scheme. - **Judge-versus-verifier agreement on the calibration set** (§16). Under one resident model the judge is the agent's model, and this is the only instrument that can see the grader drifting toward its own bias. - **Model residency and swap count** (§14.2). A nonzero swap rate on a single-model deployment means something is requesting a non-resident model, and the load cost will dominate everything else in the trace. - Admission refusals by reason, separating "capacity" from "budget". They look identical in a queue-depth graph and have opposite fixes. - Held-out versus selection gap (§12.5) as the overfitting alarm. - Trace context propagated through `Ctx`, never through task-locals. --- ## 16. Deliberately not built - Anything before the record is trustworthy. No dashboards and no learning loop until one full run works end to end against a stub model. - Grading that decides. It attributes; the verifier decides. - Automatic workflow mutation without human approval. The generator is gated by default; a user may ungate it once their loop is calibrated against verifier ground truth, and that is their decision to make explicitly. - Self-critique using the same model family being graded, **ungated**. §14.2's one-resident-model default means the judge normally *is* the agent's model, so this is no longer a prohibition but a precondition: it is admissible only with the bootstrap the prior revision named — validation against verifiable tasks. Concretely, a calibration set of tasks with known verifier ground truth is replayed through the judge on a schedule, and judge-versus-verifier agreement is tracked as a health metric (§15). It costs no agent runs, since the episodes are already recorded. Without it the grader's bias is unmeasured and the optimizer will find it — and under one-state convergence (§12.1) there is no competing variant whose divergence would make that visible. - Cross-tenant blob deduplication (§3). - Semantic retrieval over episodes. No consumer yet. --- ## 17. Decisions | Question | Decision | Where | |---|---|---| | Framework or application | **Framework.** Kernel/domain split; workflows are data | §2, §4 | | Async runtime | **Tokio.** Reverses the prior asupersync choice; ecosystem access is decisive for a distributed framework | §6 | | Storage | **Two modes**, one port set: `redb` embedded, Postgres + object store distributed | §7 | | Port signatures | **Async everywhere**, including local implementations | §7 | | Durability | Engine gives atomicity; **the log gives history**. Append + apply + advance in one transaction | §8 | | Side effects | **Write-ahead intent in three phases** (pending / dispatched / committed), effect-class recovery, `Indeterminate` as a real state | §8.4 | | Rewind | **Fork**, never truncate. `BranchId` in the log key | §8.5 | | Log evolution | **Versioned records, upcasters on read.** Variants never removed | §8.7 | | Grading | **Relative only**, and **strategy-pluggable**. Default `PairwiseSequential`: one current, one challenger, one comparison per episode, sequential-test stop. `Judge` returns a verdict not a number | §11.1, §11.3 | | Tournament | **Opt-in, not default.** Swiss + Bradley-Terry with a Davidson tie term and a prior, strengths as deltas from a pinned anchor. Enabled where episodes are co-present anyway — attempts, replay | §11.4 | | Optimization target | **Converge to one state.** One current version plus at most one challenger; N-variant population selection is the resourced option | §12.1, §12.3 | | Model capacity | **One resident model by default**, agent's model pinned, judge and proposer share it. VRAM residency invariant checked at load; context ceiling derived from KV cost, not chosen | §14.2 | | Self-critique | **Admissible under calibration.** Same-model judging is the default consequence of one resident model, gated on judge-versus-verifier agreement tracking | §16 | | Tenancy | **Tenant key on every row and every blob namespace**, from commit one | §3 | | Multi-agent orchestration | **Deferred.** The log already provides durable state and resumability; revisit when cross-agent coordination is real | §18 | --- ## 18. Open questions - **What defines `TaskId` for a given user.** Ticket id, input fixture, or a hash of the pre-workflow goal. Depends on where work enters their system, so the framework provides the type and a default hasher and lets it be overridden. Must be settled before any run is recorded, since it cannot be backfilled. - **Minimum useful group size.** G ≥ 2 runs, but a two-episode tournament is one comparison and carries little. Where the useful floor sits is empirical. - **Group timeout default.** §11.6 closes a group on quorum or timeout and increments the epoch. The timeout trades grading latency against group size, and the right default depends on tenant arrival rate — which the framework can measure but has no data for yet. Only bites when the tournament is enabled. - **Sequential test boundaries.** §11.3 stops on α, β and a minimum detectable win-rate shift. All three are policy, not physics: too tight and no challenger is ever accepted, too loose and the loop churns the current version on noise. Needs calibration against a workflow whose true improvement is known. - **Draw-rate ceiling.** §11.3 rejects a challenger that draws too often, since the test would otherwise never terminate. Where the ceiling sits is empirical and interacts with judge quality — a weak judge draws more. - **Local optima under one-state convergence.** §12.1 accepts hill-climbing. Nothing currently detects a loop that has stalled in a local optimum versus one correctly reporting no improvement exists. The held-out gap (§12.5) is the nearest instrument and was not designed for this. - **`kv_bytes_per_token` measurement.** §14.2's context ceiling depends on it, and it varies with dtype, quantization, attention implementation and parallelism. Measured per deployment or read from a profile the operator supplies — the framework should refuse a guess. - **Grading spend ratio.** §14 meters it; nobody has set the ceiling. - **Reduction summarizer.** §8.6 reduces blob bodies to "a summary" without saying what produces it. A model call makes reduction non-deterministic, which interacts badly with replay. Extractive or structural reduction may suffice. - **Effect classification for built-in tools.** An audit, not a decision. Until done, default `Unsafe` and never auto-retry. - **Postgres schema for the log at scale.** Partitioning by tenant and time, index strategy for branch scans, and whether the outbox is a table or a logical replication slot. - **Cancellation rigour under tokio.** §6 accepts cooperative cancellation as a residual risk. Whether a lint, a wrapper type, or a `loom` harness is the right enforcement is unresolved. --- ## 19. Lessons carried forward Each of these cost real time in the prior TypeScript implementation. Most are now structural; the rest stay written down. **Structurally handled:** half a concurrency guarantee is worse than none (awaitable delivery without cancellation); unbounded state (retention declared at construction); ordering under concurrency (per-run scope, not a global queue). **Still on us:** - *A green suite says nothing about coverage.* An idempotency feature generated its own keys and checked for duplicates among them — unreachable for a whole phase, with tests asserting the count was zero. - *Suspect the guards before the mechanism.* A "batching" failure was the depth guard, because two unrelated limits shared a default value. - *A test run that prints nothing cannot distinguish slow from hung.* Per-test progress and per-test timeouts from the first commit. - *Collapsing a state removes the seam that needed it.* Resolving pass/fail inside `verify()` blocked three separate features. - *A placeholder that type-checks is invisible.* A hardcoded `"current"` version hash compiled, passed tests, and made every result unattributable. A newtype with no `Default` refuses to compile instead. - *A port finished while known to be unimplementable is not finished.* The synchronous dedupe store documented, in its own doc comment, that a networked implementation could not honour its signature. --- ## 20. Build order Detailed, dependency-ordered tasks: [rust-agentic-task.md](rust-agentic-task.md). The walking skeleton is one full run of one workflow against a stub model, in embedded mode, with no network — recorded, verified, and re-derivable from the log. Nothing after it is worth starting until it runs.