1701 lines
81 KiB
Markdown
1701 lines
81 KiB
Markdown
# Rust Agentic System — Design
|
||||
|
|
|
|||
|
|
A framework for building distributed agents that record what they did, verify
|
|||
|
|
it, grade it, and improve from it.
|
|||
|
|
|
|||
|
|
Three things distinguish this from an agent library:
|
|||
|
|
|
|||
|
|
- **The episode is a first-class durable artifact**, not a log line. It survives
|
|||
|
|
crashes, supports rewind, and is the input to every learning mechanism.
|
|||
|
|
- **Learning is built in and pluggable.** A default reinforcement loop ships
|
|||
|
|
working; the grading schema, the rubric and the judge are all replaceable.
|
|||
|
|
- **The workflow is data, not code.** Users define workflows in their own format
|
|||
|
|
and version them; the framework executes and grades them without recompiling.
|
|||
|
|
|
|||
|
|
Previous revision of this document described a bespoke single-tenant observer
|
|||
|
|
embedded in one specific agent. That framing is gone. What survives is the
|
|||
|
|
durability model, the state-machine discipline, and the list of mistakes worth
|
|||
|
|
not repeating.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 1. Principles
|
|||
|
|
|
|||
|
|
**The framework cannot break the agent it runs.** Observation, verification and
|
|||
|
|
grading failures degrade the record, never the work. Any code path where a
|
|||
|
|
grader can fail into an agent's execution is a defect.
|
|||
|
|
|
|||
|
|
**The record is grounded.** Renderers and graders state only facts the episode
|
|||
|
|
holds. An invented fact produces a lesson about something that never happened.
|
|||
|
|
|
|||
|
|
**Verification decides, grading explains.** Verification returns ground truth.
|
|||
|
|
Grading attributes cause and ranks. Collapse them and the system grades its own
|
|||
|
|
homework.
|
|||
|
|
|
|||
|
|
**Nothing derived is authoritative.** Materialized state is a cache of the log.
|
|||
|
|
If it cannot be dropped and rebuilt byte-identically, it has hidden inputs and
|
|||
|
|
that is a bug.
|
|||
|
|
|
|||
|
|
**Defaults ship working; every default is a port.** A user who wants the
|
|||
|
|
built-in behaviour writes no code. A user who wants their own writes an impl,
|
|||
|
|
not a fork.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 2. Layering
|
|||
|
|
|
|||
|
|
The central structural decision, and the one the previous revision got wrong.
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
┌──────────────────────────────────────────────────────────┐
|
|||
|
|
│ DOMAIN — user-defined, data, versioned, hot-swappable │
|
|||
|
|
│ workflow definition · rubrics · verifiers · tools │
|
|||
|
|
└──────────────────────────────────────────────────────────┘
|
|||
|
|
│ executed / graded by
|
|||
|
|
┌──────────────────────────────────────────────────────────┐
|
|||
|
|
│ KERNEL — framework-owned, compiled, exhaustively typed │
|
|||
|
|
│ attempt lifecycle · event log · intents · branches │
|
|||
|
|
│ scheduling · tournament · partitioning · tenancy │
|
|||
|
|
└──────────────────────────────────────────────────────────┘
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**Kernel states are closed.** An attempt is `Pending → Running → {Succeeded,
|
|||
|
|
Failed, TimedOut, Indeterminate}`. That enum is exhaustive, matched at compile
|
|||
|
|
time, and users cannot extend it. Everything the durability and learning
|
|||
|
|
machinery reasons about lives here.
|
|||
|
|
|
|||
|
|
**Domain states are open.** A workflow declares its own step names, ordering and
|
|||
|
|
transitions as data. The kernel validates the declaration, then executes it.
|
|||
|
|
Adding a domain state is a config change.
|
|||
|
|
|
|||
|
|
The previous revision declared one machine and called it normative, then
|
|||
|
|
celebrated that adding a state was a compile error at every call site. Correct
|
|||
|
|
for a bespoke tool, fatal for a framework — a user defining their own workflow
|
|||
|
|
would have to fork and recompile. The fix is not to loosen the kernel. It is to
|
|||
|
|
stop conflating the two.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
/// Kernel. Closed. Users never extend this.
|
|||
|
|
#[derive(Clone, Copy, PartialEq, Eq)]
|
|||
|
|
pub enum AttemptState {
|
|||
|
|
Pending,
|
|||
|
|
Running,
|
|||
|
|
Succeeded,
|
|||
|
|
Failed,
|
|||
|
|
/// Crashed mid-side-effect; recovery could not determine what happened (§8.4).
|
|||
|
|
Indeterminate,
|
|||
|
|
/// Stopped by decision, with no dispatched intent outstanding. Distinct
|
|||
|
|
/// from `Indeterminate`: nothing external is in doubt (§5.1).
|
|||
|
|
Cancelled,
|
|||
|
|
TimedOut,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
/// Domain. Open. Declared by the workflow, validated at load.
|
|||
|
|
pub struct StepState(SmolStr);
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`Indeterminate` is a kernel state because only the kernel knows about intents.
|
|||
|
|
It cannot be expressed in a user workflow and must not be collapsed into
|
|||
|
|
`Failed` — the two demand different operator responses.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 3. Identity and tenancy
|
|||
|
|
|
|||
|
|
Multi-tenant from the first commit. Retrofitting a tenant key through a schema,
|
|||
|
|
a partition scheme and a blob store is a rewrite, and the previous revision had
|
|||
|
|
no tenant concept at all.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub struct TenantId(Uuid);
|
|||
|
|
pub struct WorkflowId(SmolStr); // logical workflow, stable across versions
|
|||
|
|
pub struct WorkflowVersion(Blake3Hash); // content hash of the definition
|
|||
|
|
pub struct StepId(SmolStr); // stable across versions — see §4.3
|
|||
|
|
pub struct TaskId(Blake3Hash); // the comparison group key — see §11.5
|
|||
|
|
pub struct RunId(Ulid); // one execution of one workflow
|
|||
|
|
pub struct BranchId(u32); // rewind fork — see §8.5
|
|||
|
|
pub struct AttemptNo(u32);
|
|||
|
|
pub struct Lsn(u64); // per (run, branch) sequence
|
|||
|
|
pub struct GroupEpoch(u32); // comparison-group generation — see §11.6
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`Ulid` for `RunId`: lexicographically sortable by creation time, which makes
|
|||
|
|
range scans over recent runs a prefix scan rather than a secondary index.
|
|||
|
|
|
|||
|
|
**Every key is tenant-prefixed.** Not "most". A single unprefixed table is a
|
|||
|
|
cross-tenant read waiting to happen, and it will be found by a customer rather
|
|||
|
|
than by us.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub struct Scoped<T> { pub tenant: TenantId, pub inner: T }
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Tables key on `Scoped<_>`. The type makes an unscoped access a compile error
|
|||
|
|
rather than a review comment.
|
|||
|
|
|
|||
|
|
**Blobs are namespaced per tenant even though they are content-addressed.**
|
|||
|
|
Global deduplication of prompt and output blobs is tempting — identical system
|
|||
|
|
prompts across tenants are common — and it is a leak. A shared blob means one
|
|||
|
|
tenant's storage accounting depends on another's, and a hash becomes an oracle
|
|||
|
|
for "does anyone else have this content". Deduplicate within a tenant, never
|
|||
|
|
across.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 4. Workflow definition
|
|||
|
|
|
|||
|
|
The workflow is data. The framework provides a validated intermediate
|
|||
|
|
representation and a parser trait; a format is a plugin.
|
|||
|
|
|
|||
|
|
### 4.1 The IR
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub struct WorkflowDef {
|
|||
|
|
pub id: WorkflowId,
|
|||
|
|
pub schema: SchemaVersion,
|
|||
|
|
pub steps: Vec<StepDef>,
|
|||
|
|
pub transitions: Vec<Transition>,
|
|||
|
|
pub rubric: RubricDef,
|
|||
|
|
pub budget: BudgetDef,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub struct StepDef {
|
|||
|
|
/// Author-assigned, stable across versions. See §4.3.
|
|||
|
|
pub id: StepId,
|
|||
|
|
pub kind: StepKind,
|
|||
|
|
pub tools: ToolSelector,
|
|||
|
|
pub verify: Vec<VerifierRef>,
|
|||
|
|
pub retry: RetryPolicy,
|
|||
|
|
pub timeout: Duration,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub enum StepKind {
|
|||
|
|
Model { prompt: PromptTemplate, effort: ReasoningEffort },
|
|||
|
|
Tool { tool: ToolId, args: ArgTemplate },
|
|||
|
|
Parallel { branches: Vec<StepId>, join: JoinPolicy },
|
|||
|
|
Conditional { on: Predicate, then: StepId, otherwise: Option<StepId> },
|
|||
|
|
SubWorkflow { workflow: WorkflowId, version: VersionSelector },
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`WorkflowVersion` is the Blake3 hash of the canonicalized IR — not of the source
|
|||
|
|
text. Two YAML files differing only in key order produce the same version, which
|
|||
|
|
is what makes "did this change affect results" answerable.
|
|||
|
|
|
|||
|
|
### 4.2 Formats are plugins
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub trait WorkflowFormat: Send + Sync {
|
|||
|
|
fn extensions(&self) -> &[&str];
|
|||
|
|
fn parse(&self, src: &[u8]) -> Result<WorkflowDef, ParseError>;
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Ship YAML and JSON. A user wanting a DSL, Starlark, or a database row implements
|
|||
|
|
the trait. Validation and canonicalization live in the kernel and run on the IR,
|
|||
|
|
so a new format inherits every check without reimplementing one.
|
|||
|
|
|
|||
|
|
### 4.3 `StepId` stability is the user's contract
|
|||
|
|
|
|||
|
|
Credit assignment (§11.7) attributes outcomes to steps across workflow versions.
|
|||
|
|
That requires a step identity that survives edits — insert a step at position 2
|
|||
|
|
and every positional index shifts, but `StepId` does not.
|
|||
|
|
|
|||
|
|
The framework cannot infer this. It is a documented obligation on the workflow
|
|||
|
|
author, enforced by three checks at load time:
|
|||
|
|
|
|||
|
|
- `StepId` unique within a version.
|
|||
|
|
- On a version bump, report added, removed and retained ids. A version that
|
|||
|
|
retains no ids from its parent is almost certainly a renumbering accident and
|
|||
|
|
is rejected unless explicitly marked as a rewrite.
|
|||
|
|
- `StepId` is opaque to the framework. Never parsed, never ordered, never
|
|||
|
|
assumed numeric.
|
|||
|
|
|
|||
|
|
### 4.4 Sub-workflows and version pinning
|
|||
|
|
|
|||
|
|
`SubWorkflow` pins by `VersionSelector`: `Exact(hash)` or `Latest`. `Latest`
|
|||
|
|
resolves **once, at run spawn**, and the resolved hash is recorded. A run whose
|
|||
|
|
sub-workflow version can change mid-execution is a run whose results attribute
|
|||
|
|
to nothing.
|
|||
|
|
|
|||
|
|
Recursion depth is bounded by the kernel and the cycle is detected at load, not
|
|||
|
|
at execution.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 5. Execution model
|
|||
|
|
|
|||
|
|
### 5.1 The two machines
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
KERNEL — per attempt, closed
|
|||
|
|
Pending ──► Running ──┬──► Succeeded
|
|||
|
|
│ ├──► Failed
|
|||
|
|
│ ├──► TimedOut
|
|||
|
|
│ ├──► Cancelled
|
|||
|
|
├──► Cancelled └──► Indeterminate
|
|||
|
|
└──► TimedOut
|
|||
|
|
|
|||
|
|
DOMAIN — per run, declared by WorkflowDef
|
|||
|
|
whatever the user wrote, validated as a DAG with explicit loop bounds
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
The legal set, exhaustively — this table is normative and the transition
|
|||
|
|
function matches it arm for arm:
|
|||
|
|
|
|||
|
|
| From | To | When |
|
|||
|
|
|---|---|---|
|
|||
|
|
| `Pending` | `Running` | admitted |
|
|||
|
|
| `Pending` | `Cancelled` | run cancelled before the attempt started |
|
|||
|
|
| `Pending` | `TimedOut` | queue deadline passed before admission |
|
|||
|
|
| `Running` | `Succeeded` | completed |
|
|||
|
|
| `Running` | `Failed` | completed with an error |
|
|||
|
|
| `Running` | `TimedOut` | step deadline passed |
|
|||
|
|
| `Running` | `Cancelled` | run cancelled, **no intent was `Dispatched`** (§8.4) |
|
|||
|
|
| `Running` | `Indeterminate` | an intent was `Dispatched` and did not resolve |
|
|||
|
|
| terminal | — | nothing leaves a terminal state |
|
|||
|
|
|
|||
|
|
**`Cancelled` is a kernel state and is not `Indeterminate`.** Cancellation is a
|
|||
|
|
decision; indeterminacy is an unknown. Collapsing them was tempting because
|
|||
|
|
§6 drops a tool call at its next await point and the tool may have been
|
|||
|
|
mid-something — but that "may" is exactly what §8.4's `Dispatched` record
|
|||
|
|
answers. If no intent was dispatched, nothing external happened and the attempt
|
|||
|
|
is cleanly `Cancelled`. Only a dispatched-and-unresolved intent earns
|
|||
|
|
`Indeterminate`. Getting this wrong is not cosmetic: §15 alerts on
|
|||
|
|
`Indeterminate` count and expects near-zero, so routing every cancelled attempt
|
|||
|
|
there converts the alert into background noise and it stops being read.
|
|||
|
|
|
|||
|
|
A run advances through domain steps. Each step execution is one or more kernel
|
|||
|
|
attempts. Retry creates **attempt N+1** and never mutates attempt N — this makes
|
|||
|
|
"did the retry do better, and why" answerable, and it is what makes log replay
|
|||
|
|
idempotent for free.
|
|||
|
|
|
|||
|
|
### 5.2 Run lifecycle
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
spawn
|
|||
|
|
│
|
|||
|
|
v
|
|||
|
|
Scheduled ─────────────────┐
|
|||
|
|
│ admitted │
|
|||
|
|
v │
|
|||
|
|
Running ⇄ Suspended ──────┤
|
|||
|
|
│ all steps terminal │ cancel
|
|||
|
|
v │
|
|||
|
|
Verifying ─────────────────┤ ◄── rests here while verifier
|
|||
|
|
│ │ futures are outstanding
|
|||
|
|
v │
|
|||
|
|
Verified{pass|fail} │
|
|||
|
|
│ │
|
|||
|
|
v v
|
|||
|
|
Grading ──────────────► Cancelled ●
|
|||
|
|
│ ◄── may wait for a tournament group to fill (§11.6)
|
|||
|
|
│
|
|||
|
|
├──► Graded ● ──┐
|
|||
|
|
│ ├──► Archived ●
|
|||
|
|
└──► Ungraded ● ──┘
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**Cancel is reachable from every non-terminal state**, not only from
|
|||
|
|
`Scheduled`. A run cancelled mid-step is the ordinary case — it is what a user
|
|||
|
|
clicking stop does — and a lifecycle offering cancel only before admission
|
|||
|
|
describes a system nobody would ship. Cancelling during `Verifying` or `Grading`
|
|||
|
|
is rarer and still legal; the work is done and the record stands, the framework
|
|||
|
|
just stops spending on judging it.
|
|||
|
|
|
|||
|
|
`Verifying` and `Grading` are states the run **rests in**, not synchronous
|
|||
|
|
branches. The previous implementation collapsed `Verifying` by resolving
|
|||
|
|
pass/fail inside `verify()`, which blocked async verifiers, mid-run UI, and the
|
|||
|
|
snapshot barrier in §10.2 — one collapsed state, three blocked features.
|
|||
|
|
|
|||
|
|
**`Ungraded` is terminal and sits beside `Graded`, not below it.** Grading can
|
|||
|
|
legitimately end without a score: G = 1 with no group to join (§11.6), a group
|
|||
|
|
that closed on timeout without this run, or a tenant over its grading ceiling
|
|||
|
|
(§14). Those runs are finished. Without a terminal state saying so they sit in
|
|||
|
|
`Grading` forever, and since §10.3 gates retention on grading having ended, they
|
|||
|
|
also become permanently irreducible — unbounded storage growth landing precisely
|
|||
|
|
on the low-volume tenants §11.6 exists to accommodate, and on the tenants who
|
|||
|
|
hit a cost ceiling, which is the worst possible pairing. `Ungraded` carries its
|
|||
|
|
`UngradedReason` (§11.1) so the dashboard can state why rather than showing a
|
|||
|
|
gap.
|
|||
|
|
|
|||
|
|
`Suspended` is new and load-bearing for distribution: a run awaiting human
|
|||
|
|
approval or a webhook must release its worker. A run that holds an executor slot
|
|||
|
|
across a human decision does not scale past a handful of concurrent runs.
|
|||
|
|
|
|||
|
|
### 5.3 Concurrency shape
|
|||
|
|
|
|||
|
|
| Scope | Parallel | Why |
|
|||
|
|
|---|---|---|
|
|||
|
|
| across runs | unbounded | no shared state |
|
|||
|
|
| steps within a run | serial by default | step N+1 reads N's output |
|
|||
|
|
| `Parallel` step branches | fan-out/join | declared explicitly in the IR |
|
|||
|
|
| attempts within a step | strictly serial | a retry needs the prior failure |
|
|||
|
|
| verifiers for one attempt | fan-out/join | independent checks |
|
|||
|
|
| tournament group | join | needs the whole group (§11.6) |
|
|||
|
|
|
|||
|
|
The serial spine is `(TenantId, RunId)`. Parallelism lives between runs and
|
|||
|
|
inside declared fan-out. Nothing else may interleave.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 6. Runtime
|
|||
|
|
|
|||
|
|
**Tokio.** This reverses the previous revision, and the reversal is a direct
|
|||
|
|
consequence of the goal change.
|
|||
|
|
|
|||
|
|
The prior choice was `asupersync` — structured concurrency, capability-secure,
|
|||
|
|
cancel-correct, with a deterministic test lab. Genuinely better primitives, and
|
|||
|
|
it was correct when this was a component embedded inside one agent that already
|
|||
|
|
used it. As a framework that users embed, it fails on one axis that outweighs
|
|||
|
|
the rest: **no tokio compatibility means users cannot use the ecosystem.** No
|
|||
|
|
`sqlx`, `rdkafka`, `aws-sdk`, `tonic`, `axum`, `reqwest`, or object-store
|
|||
|
|
clients. For a distributed framework those are not optional dependencies; they
|
|||
|
|
are the distribution layer.
|
|||
|
|
|
|||
|
|
The prior revision already conceded this in its "crossing the runtime boundary"
|
|||
|
|
section, treating tokio interop as an exception for two clients. Under the
|
|||
|
|
framework goal, that boundary is the common case, and a design whose exception
|
|||
|
|
path is the main path is the wrong design.
|
|||
|
|
|
|||
|
|
What is lost, and how it is recovered:
|
|||
|
|
|
|||
|
|
| asupersync gave | Recovered by |
|
|||
|
|
|---|---|
|
|||
|
|
| Regions — structural task-tree cancellation | `TaskTracker` + `CancellationToken` from `tokio-util`, one tracker per run, enforced by a `RunScope` guard that refuses detached spawns |
|
|||
|
|
| Cancel Protocol — work actually stops | `CancellationToken` selected against at every await in kernel code; user tool calls get a hard timeout. A dropped call resolves `Cancelled` if no intent reached `Dispatched`, `Indeterminate` if one did (§5.1) — the intent record, not the drop, is what decides |
|
|||
|
|
| `Cx` — explicit capability passing | An explicit `Ctx` struct threaded through every call. Never `task_local!` for anything causal — that is the `AsyncLocalStorage` mistake in different clothing |
|
|||
|
|
| The Lab — deterministic schedules | `turmoil` for network partition and latency simulation; `loom` for the lock-free bits; `tokio::time::pause` for time. Weaker than a seeded scheduler; sufficient with discipline |
|
|||
|
|
|
|||
|
|
The residual risk is honest: tokio's cancellation is cooperative, so a
|
|||
|
|
`select!`-dropped future stops at its next await and not before. Kernel code
|
|||
|
|
must never hold a lock or a half-applied state across an await that can be
|
|||
|
|
cancelled. This is enforceable by review and by the `RunScope` guard; it is not
|
|||
|
|
enforced by the compiler the way Regions did.
|
|||
|
|
|
|||
|
|
**An `asupersync` backend stays possible** behind a runtime trait, feature-gated,
|
|||
|
|
for embedding in agents that already use it. Not built until someone needs it,
|
|||
|
|
and not on the default path.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 7. Storage ports
|
|||
|
|
|
|||
|
|
Two deployment modes, one set of ports.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
#[async_trait]
|
|||
|
|
pub trait EventLog: Send + Sync {
|
|||
|
|
/// The whole of §8.3 in one call: append the records, apply the derived
|
|||
|
|
/// state, advance the consumer position, enqueue the outbox — all four or
|
|||
|
|
/// none. Returns the assigned LSNs.
|
|||
|
|
async fn commit(&self, batch: CommitBatch) -> Result<Vec<Lsn>>;
|
|||
|
|
|
|||
|
|
async fn read(&self, key: BranchKey, from: Lsn, limit: usize) -> Result<Vec<LogRecord>>;
|
|||
|
|
|
|||
|
|
async fn put_checkpoint(&self, key: BranchKey, upto: Lsn, state: &[u8]) -> Result<()>;
|
|||
|
|
/// Newest checkpoint at or below `upto`. Restart folds forward from here;
|
|||
|
|
/// `None` means fold from LSN 0 (§8.6).
|
|||
|
|
async fn latest_checkpoint(&self, key: BranchKey, upto: Lsn)
|
|||
|
|
-> Result<Option<Checkpoint>>;
|
|||
|
|
|
|||
|
|
/// Committed-but-unshipped export intents, in `(BranchKey, Lsn)` order.
|
|||
|
|
/// Read by the relay (§9.3), never from the execution path.
|
|||
|
|
async fn drain_outbox(&self, tenant: TenantId, limit: usize) -> Result<Vec<OutboxEntry>>;
|
|||
|
|
async fn ack_outbox(&self, shipped: &[(BranchKey, Lsn)]) -> Result<()>;
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub struct CommitBatch {
|
|||
|
|
/// Carries the tenant. Never a separate parameter beside a key that
|
|||
|
|
/// already holds one — two sources for one fact is one too many.
|
|||
|
|
pub key: BranchKey,
|
|||
|
|
pub records: Vec<WorkEvent>, // LSNs assigned by the implementation
|
|||
|
|
pub state: Vec<StateDelta>,
|
|||
|
|
pub position: Option<ConsumerPosition>,
|
|||
|
|
pub outbox: Vec<ExportIntent>,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
#[async_trait]
|
|||
|
|
pub trait BlobStore: Send + Sync {
|
|||
|
|
async fn put(&self, tenant: TenantId, content: Bytes) -> Result<BlobRef>;
|
|||
|
|
async fn get(&self, tenant: TenantId, r: &BlobRef) -> Result<Option<Bytes>>;
|
|||
|
|
/// Reduction (§8.6) and tenant deletion (§3) both require this. A store
|
|||
|
|
/// that cannot delete cannot honour either, and both are obligations.
|
|||
|
|
async fn delete(&self, tenant: TenantId, r: &BlobRef) -> Result<()>;
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**`commit` is one method rather than four because atomicity is the contract.**
|
|||
|
|
A port exposing `append` alone puts the other three writes of §8.3 outside the
|
|||
|
|
transaction, which is the durability guarantee gone — and gone invisibly, since
|
|||
|
|
each write individually succeeds. The port must be able to express the strongest
|
|||
|
|
thing the implementation promises, or the abstraction quietly weakens it. Same
|
|||
|
|
reasoning behind `latest_checkpoint` and `BlobStore::delete`: a checkpoint that
|
|||
|
|
can be written and not read is an optimization that cannot be used, and §8.6's
|
|||
|
|
retention path is unimplementable without a delete.
|
|||
|
|
|
|||
|
|
Everything is `async`. The previous revision declared the hot-path store
|
|||
|
|
synchronous because the local implementation was a B-tree, then documented in
|
|||
|
|
the same file that a network-backed implementation could not honour the
|
|||
|
|
signature. That is a port finished while already known to be unimplementable.
|
|||
|
|
An async signature over a local call costs a negligible poll; a sync signature
|
|||
|
|
over a network call is impossible. The port-completeness failure above is the
|
|||
|
|
same mistake one level up: a signature that cannot express what the caller needs
|
|||
|
|
is not finished either.
|
|||
|
|
|
|||
|
|
| Mode | Log + state | Blobs | Coordination |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| **Embedded** — single binary, no services | `redb` | `redb` table | in-process |
|
|||
|
|
| **Distributed** — multi-node, multi-tenant | Postgres | S3-compatible | Postgres advisory locks, or Redis if leases dominate |
|
|||
|
|
|
|||
|
|
`redb` remains the right embedded engine: pure Rust, ACID, MVCC, stable file
|
|||
|
|
format, no server. It uses copy-on-write shadow paging rather than a WAL, so a
|
|||
|
|
torn write cannot corrupt the file — it simply does not take effect. Commits
|
|||
|
|
must be `Durability::Immediate`; the enum is `#[non_exhaustive]`, so set it
|
|||
|
|
explicitly rather than relying on the default.
|
|||
|
|
|
|||
|
|
Avoid `sled` — years at 0.34 beta with known space amplification.
|
|||
|
|
|
|||
|
|
**Embedded mode is a first-class product**, not a test harness. A user must be
|
|||
|
|
able to `cargo add` this, run an agent, and get durability and grading with zero
|
|||
|
|
infrastructure. That constraint is what keeps the ports honest.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 8. Durability
|
|||
|
|
|
|||
|
|
### 8.1 Two requirements, one solved by the engine
|
|||
|
|
|
|||
|
|
**Crash-atomicity** — a crash must not leave half-written state. `redb`'s shadow
|
|||
|
|
paging and Postgres transactions both handle this.
|
|||
|
|
|
|||
|
|
**History** — rewind, resume-from-failure, and "what did this look like at step
|
|||
|
|
3" need the *sequence* of transitions. Neither engine keeps one. This is ours.
|
|||
|
|
|
|||
|
|
The log is a table on top of the engine, not a competitor to it. Because the
|
|||
|
|
engine's transactions are atomic, appending to the log and applying the state
|
|||
|
|
happen together or not at all: no torn records, no redo/undo pass, no
|
|||
|
|
checkpoint-consistency problem.
|
|||
|
|
|
|||
|
|
### 8.2 The record
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub struct LogRecord {
|
|||
|
|
pub key: BranchKey, // (TenantId, RunId, BranchId)
|
|||
|
|
pub lsn: Lsn,
|
|||
|
|
/// Wire-format version of `event`. Never removed, never reused. See §8.7.
|
|||
|
|
pub schema: SchemaVersion,
|
|||
|
|
pub at: Timestamp,
|
|||
|
|
pub event: WorkEvent,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub struct BranchKey { pub tenant: TenantId, pub run: RunId, pub branch: BranchId }
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`BranchId` is in the key, not implied. The previous revision keyed the log on
|
|||
|
|
`(RunId, Lsn)` while a later section claimed state was "keyed by branch as well
|
|||
|
|
as attempt" — a contradiction that made forking unimplementable as written.
|
|||
|
|
|
|||
|
|
**LSNs are per branch, not global.** A global counter serializes every run
|
|||
|
|
through one atomic. The ordering contract is per-run total order with nothing
|
|||
|
|
promised across runs, so a per-branch sequence is exactly as strong as the
|
|||
|
|
contract requires, contention-free, and keeps the log partitionable by run.
|
|||
|
|
|
|||
|
|
### 8.3 The commit protocol
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
let txn = db.begin_write()?;
|
|||
|
|
{
|
|||
|
|
let mut log = txn.open_table(EVENT_LOG)?;
|
|||
|
|
let mut state = txn.open_table(RUN_STATE)?;
|
|||
|
|
let mut position = txn.open_table(CONSUMER_POSITION)?;
|
|||
|
|
let mut outbox = txn.open_table(OUTBOX)?;
|
|||
|
|
|
|||
|
|
log.insert((branch_key, lsn), &record)?; // append: the durable fact
|
|||
|
|
state.insert((branch_key, attempt_no), &attempt)?; // apply: the derived view
|
|||
|
|
position.insert(Scoped::new(tenant, stream), pos)?; // advance: where to resume
|
|||
|
|
outbox.insert((branch_key, lsn), &intent)?; // relay separately (§13.3)
|
|||
|
|
}
|
|||
|
|
txn.commit()?; // all four, or none
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**Every key here carries `BranchKey` or a `Scoped<_>` (§3).** The outbox is the
|
|||
|
|
one that invites the mistake: an outbox keyed on `Lsn` alone reads naturally and
|
|||
|
|
is wrong, because LSNs are per branch (§8.2), so a bare LSN collides across every
|
|||
|
|
branch of every run of every tenant. The same key shape also gives the relay a
|
|||
|
|
defined order — per `BranchKey`, ascending `Lsn` — which is the only ordering
|
|||
|
|
§9.2 promises.
|
|||
|
|
|
|||
|
|
Commit per event, not per run. A projection that accumulates in memory and
|
|||
|
|
writes at run end loses the whole run on a crash. Episodes are small and
|
|||
|
|
append-mostly; one fsync per transition is cheap next to model latency.
|
|||
|
|
|
|||
|
|
### 8.4 Write-ahead intent
|
|||
|
|
|
|||
|
|
Appending after the fact records history. It does not make a failed step
|
|||
|
|
resumable, because the dangerous window is *before* the record exists.
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
1. append Intent{Pending} ──► commit + fsync
|
|||
|
|
◄── window A: crash here, the call was never issued
|
|||
|
|
2. append Intent{Dispatched} ──► commit + fsync
|
|||
|
|
3. perform the call
|
|||
|
|
◄── window B: crash here, the call may have landed
|
|||
|
|
4. append outcome, Intent{Committed} ──► commit + fsync
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Three records, not two. Two records cannot separate the windows: a crash before
|
|||
|
|
the call and a crash after it both leave a lone `Pending` with no outcome, which
|
|||
|
|
makes every interrupted effect maximally suspicious and pushes recoverable work
|
|||
|
|
into `Indeterminate`. The second fsync buys the distinction. It is paid only on
|
|||
|
|
steps with external effects and is small next to the call it guards.
|
|||
|
|
|
|||
|
|
On restart, the last committed intent state classifies the crash:
|
|||
|
|
|
|||
|
|
| Last state | Meaning | Resolution |
|
|||
|
|
|---|---|---|
|
|||
|
|
| `Pending` | the call was never issued | retry freely, whatever the effect class |
|
|||
|
|
| `Dispatched` | the call *may* have been issued | by effect class, below |
|
|||
|
|
| `Committed` | outcome already recorded | nothing to do |
|
|||
|
|
|
|||
|
|
For a `Dispatched` intent, resolution is by declared effect class:
|
|||
|
|
|
|||
|
|
| Class | Recovery |
|
|||
|
|
|---|---|
|
|||
|
|
| `Idempotent` | retry with the same idempotency key; the provider deduplicates |
|
|||
|
|
| `Queryable` | ask the provider whether the request id landed, then complete or retry |
|
|||
|
|
| `Unsafe` | **never auto-retry.** Attempt becomes `Indeterminate`, operator is notified |
|
|||
|
|
|
|||
|
|
The third row is the honest one. Some effects cannot be made safe by any
|
|||
|
|
protocol. The intent log's value is converting an invisible unknown into a
|
|||
|
|
recorded one — `Indeterminate` is a fact a grader and an operator can both use;
|
|||
|
|
a silently retried payment is not.
|
|||
|
|
|
|||
|
|
Intents always commit `Immediate`. Batching them defeats their only purpose.
|
|||
|
|
|
|||
|
|
Tool calls need this more than model calls. A model call is a metered read; a
|
|||
|
|
tool call writes files, pushes commits, and touches the world.
|
|||
|
|
|
|||
|
|
### 8.5 Rewind is a fork
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
lsn 0 ─ 1 ─ 2 ─ 3 ─ 4 ─ 5 ─ 6(failed) branch 0, retained
|
|||
|
|
└─ 0 ─ 1 ─ 2 ─ ... branch 1, forked at (0, 3)
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
A rewind allocates a new `BranchId` and starts its LSNs at zero, recording the
|
|||
|
|
fork point. Nothing is removed. Same rule as attempt N+1 never mutating attempt
|
|||
|
|
N, for the same reason: **the discarded branch is the evidence**. Truncating it
|
|||
|
|
destroys the failure that motivated the rewind, which is what the learning loop
|
|||
|
|
exists to consume.
|
|||
|
|
|
|||
|
|
- Queries default to the live branch; grading may read all of them.
|
|||
|
|
- Only the live branch is exported.
|
|||
|
|
- Rewinding past a `Committed` intent with a non-idempotent effect is a
|
|||
|
|
compensation problem, not a replay problem. The log records what happened; it
|
|||
|
|
cannot un-happen it. Flag rather than pretend.
|
|||
|
|
|
|||
|
|
### 8.6 Retention: reduce, then tier
|
|||
|
|
|
|||
|
|
Three mechanisms, escalating. Note the vocabulary: **`Archived`** is a run
|
|||
|
|
state; **reduction** is the token-budget operation; **tiering** is the move to
|
|||
|
|
cold storage. The previous revision called two of these "compaction" and the
|
|||
|
|
collision was guaranteed to confuse implementers.
|
|||
|
|
|
|||
|
|
**Checkpoints.** A materialized state snapshot tagged with its LSN. Restart
|
|||
|
|
folds forward from the newest one. An optimization only — deleting every
|
|||
|
|
checkpoint costs startup time and nothing else.
|
|||
|
|
|
|||
|
|
**Reduction at a token ceiling.** Retention is measured in **tokens**, because
|
|||
|
|
tokens are the currency of every downstream consumer: what a replay costs, what
|
|||
|
|
fits in a judge's window, what an export is billed at. Default ceiling 200k per
|
|||
|
|
run; per-tenant configurable.
|
|||
|
|
|
|||
|
|
On self-hosted weights that ceiling is **not free to choose** — it is bounded by
|
|||
|
|
`max_context_tokens / 2` from §14.2, because a pairwise judge reads two episodes
|
|||
|
|
into one KV cache and two 200k episodes do not fit on any single device. Where
|
|||
|
|
the derived bound is tighter than the configured ceiling, the bound wins and the
|
|||
|
|
judge reads a further-reduced view. A retention number set without reference to
|
|||
|
|
the hardware that must read it is a number that will be discovered wrong at the
|
|||
|
|
first judge call.
|
|||
|
|
|
|||
|
|
Loss order is fixed:
|
|||
|
|
|
|||
|
|
| Kept | Reduced | Dropped |
|
|||
|
|
|---|---|---|
|
|||
|
|
| every transition record | blob bodies → summary blob | raw text on dead branches |
|
|||
|
|
| context partitions, tool info, usage | dead-branch attempts → summary | |
|
|||
|
|
| verifier results, grades | | |
|
|||
|
|
|
|||
|
|
**Reduction never rewrites a blob and never edits the log.** Blobs are
|
|||
|
|
content-addressed (§3), so replacing a body under its existing ref makes the ref
|
|||
|
|
a lie; and repointing the log at a new ref is the history rewrite §8.7 forbids.
|
|||
|
|
Reduction instead writes the summary as a *new* blob, appends a
|
|||
|
|
`Reduced{original: BlobRef, summary: BlobRef}` event to the log, and only then
|
|||
|
|
deletes the original body. The reduction is a later fact about an earlier record,
|
|||
|
|
not a change to it. `BlobStore::get` on the original ref returns `None`, and the
|
|||
|
|
fold knows why and what stands in its place — so drop-and-re-fold from LSN 0
|
|||
|
|
still yields byte-identical state, which it would not if the mapping lived only
|
|||
|
|
in the projection.
|
|||
|
|
|
|||
|
|
The transition sequence always survives. What reduces is *text*, because it
|
|||
|
|
dominates token count and is the only part with a cheap lossy representation. A
|
|||
|
|
reduced episode can still be graded, attributed and structurally rewound — it
|
|||
|
|
just cannot be replayed verbatim.
|
|||
|
|
|
|||
|
|
**Tiering.** Terminal, reduced runs move to cold storage. Moved, not copied,
|
|||
|
|
with local rows deleted only after the remote commit acknowledges.
|
|||
|
|
|
|||
|
|
### 8.7 Log schema evolution
|
|||
|
|
|
|||
|
|
An append-only log plus an evolving event enum is a trap the previous revision
|
|||
|
|
walked straight past. Two years of records, one `WorkEvent` variant renamed, and
|
|||
|
|
the "drop derived state and re-fold" property is silently gone.
|
|||
|
|
|
|||
|
|
Rules, from record one:
|
|||
|
|
|
|||
|
|
- Every record carries `SchemaVersion`. Written always, even at v1.
|
|||
|
|
- `WorkEvent` is `#[non_exhaustive]`; decode is version-dispatched.
|
|||
|
|
- **Variants are never removed or repurposed.** Deprecated variants stay
|
|||
|
|
decodable forever. Storage is cheap; an undecodable log is not.
|
|||
|
|
- Migrations are **upcasters** — `fn upcast(vN) -> vN+1` — applied on read, never
|
|||
|
|
by rewriting history. Rewriting an append-only log is a contradiction.
|
|||
|
|
- A round-trip test per version, asserting that a stored fixture of every
|
|||
|
|
historical version still folds to the expected state. This test is the whole
|
|||
|
|
guarantee; without it the rules are aspirational.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 9. Distribution
|
|||
|
|
|
|||
|
|
### 9.1 Partitioning
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
ingest ──► execute ──► verify ──► tournament ──► aggregate ──► decide
|
|||
|
|
│ │ │ │ │ │
|
|||
|
|
(T,Run) (T,Run) (T,Run) (T,Task,Class, (T,Variant) (T,Workflow)
|
|||
|
|
Epoch)
|
|||
|
|
▲ ▲ ▲
|
|||
|
|
shuffle 1 shuffle 2 single writer
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Three keys, two shuffles, one single-writer stage. Everything up to verification
|
|||
|
|
keys on `(TenantId, RunId)` and is embarrassingly parallel. Grading is a **join**
|
|||
|
|
— a comparison group must be co-located — and it keys on
|
|||
|
|
`(TenantId, TaskId, VerifierOutcome, GroupEpoch)`: the outcome class because
|
|||
|
|
§11.4 brackets only within one, and the epoch because a closed group never
|
|||
|
|
reopens for a late arrival (§11.6).
|
|||
|
|
|
|||
|
|
The decide stage is single-writer per `(TenantId, WorkflowId)`. Two schedulers
|
|||
|
|
adjusting traffic allocation concurrently produce an allocation neither holds.
|
|||
|
|
A compare-and-swap on a generation counter is sufficient; no lock service
|
|||
|
|
needed at this size.
|
|||
|
|
|
|||
|
|
### 9.2 Ordering and delivery
|
|||
|
|
|
|||
|
|
Per-run total order, nothing promised across runs. Downstream consumers must
|
|||
|
|
therefore partition by run key, and the framework's broker adapters set the
|
|||
|
|
partition key from `(TenantId, RunId)` — never from a correlation id, which
|
|||
|
|
collapses unrelated runs onto one partition while splitting single runs across
|
|||
|
|
several.
|
|||
|
|
|
|||
|
|
Delivery is at-least-once. Exactly-once is achieved at the fold, not in
|
|||
|
|
transport: `(BranchKey, Lsn)` is the natural idempotency key, so a redelivered
|
|||
|
|
record is a no-op insert.
|
|||
|
|
|
|||
|
|
### 9.3 Outbox
|
|||
|
|
|
|||
|
|
The framework never calls a broker from the execution path. Export intent is
|
|||
|
|
written in the same transaction as the state (§8.3); a separate relay reads
|
|||
|
|
committed intents and ships them. This makes export restartable, keeps a broker
|
|||
|
|
outage from stalling a run, and is the only pattern that survives a crash
|
|||
|
|
between "state committed" and "event published".
|
|||
|
|
|
|||
|
|
### 9.4 Leases and work distribution
|
|||
|
|
|
|||
|
|
Runs are claimed by workers under a lease with a TTL. A worker that dies has its
|
|||
|
|
runs reclaimed after expiry. Lease renewal is a heartbeat on the run record, and
|
|||
|
|
`Suspended` runs (§5.2) release their lease entirely rather than heartbeating
|
|||
|
|
through a human's lunch break.
|
|||
|
|
|
|||
|
|
Fencing tokens on every lease. A partitioned worker that resumes must not write
|
|||
|
|
under an expired claim, and a monotonic fence in the run record makes that a
|
|||
|
|
rejected write rather than a silent double-execution.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 10. Verification
|
|||
|
|
|
|||
|
|
Verification returns ground truth. It is a port with a fail-closed contract.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
#[async_trait]
|
|||
|
|
pub trait Verifier: Send + Sync {
|
|||
|
|
fn id(&self) -> VerifierId;
|
|||
|
|
/// Any error, panic or timeout MUST resolve to `Fail`. A verifier that
|
|||
|
|
/// throws or hangs can never report `Pass`.
|
|||
|
|
async fn verify(&self, cx: &VerifierCtx) -> VerifierOutcome;
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
### 10.1 Verifiers need the inputs, not just the ids
|
|||
|
|
|
|||
|
|
A verifier seeing only identifiers can answer "did it work". Answering "did the
|
|||
|
|
agent have what it needed" requires the context and the prompt.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub struct VerifierCtx {
|
|||
|
|
pub run: RunId,
|
|||
|
|
pub step: StepId,
|
|||
|
|
pub attempt: AttemptNo,
|
|||
|
|
/// Frozen on entry to `Verifying`. Never mutates while a verifier holds it.
|
|||
|
|
pub episode: EpisodeView,
|
|||
|
|
/// Lazy. Verifiers needing no text never pay for it.
|
|||
|
|
pub blobs: Arc<dyn BlobStore>,
|
|||
|
|
pub deadline: Instant,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub struct AttemptView {
|
|||
|
|
pub step: StepId,
|
|||
|
|
pub attempt: AttemptNo,
|
|||
|
|
pub state: AttemptState,
|
|||
|
|
pub workflow_version: WorkflowVersion,
|
|||
|
|
pub context: Option<ContextPartition>, // identifiers, small, inline
|
|||
|
|
pub prompt: Option<BlobRef>, // large, by reference
|
|||
|
|
pub output: Option<BlobRef>,
|
|||
|
|
pub tools: Option<ToolInfo>,
|
|||
|
|
pub usage: Option<Usage>,
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Context partitions hold *identifiers* — packed, available-but-not-packed,
|
|||
|
|
dropped — and are small enough to inline. Prompts and outputs are large and go
|
|||
|
|
by reference. Laziness matters: a verifier that shells out and checks an exit
|
|||
|
|
code needs none of this, and making every verifier carry prompt text penalizes
|
|||
|
|
the common case while blowing broker payload limits.
|
|||
|
|
|
|||
|
|
Include **failed attempts**. "Retried three times because context was missing X"
|
|||
|
|
is the learning signal; shipping only the winning attempt discards it.
|
|||
|
|
|
|||
|
|
### 10.2 The snapshot barrier
|
|||
|
|
|
|||
|
|
Verifiers read after work completes, so state must stop moving beneath them.
|
|||
|
|
Entry to `Verifying` freezes the view. This is why `Verifying` must be a real
|
|||
|
|
resting state and not a synchronous branch.
|
|||
|
|
|
|||
|
|
Worth naming what the barrier actually defends against, because the obvious
|
|||
|
|
answer is wrong. It is **not** a concurrent retry: `Verifying` is entered only
|
|||
|
|
when every step is terminal (§5.2), so no attempt can still be running. The real
|
|||
|
|
mutators are the ones that arrive from outside the run's own execution:
|
|||
|
|
|
|||
|
|
- a **rewind** (§8.5) forking a new `BranchId` while verifiers hold a view of
|
|||
|
|
the old one;
|
|||
|
|
- a **cancel** (§5.2), now legal from `Verifying`;
|
|||
|
|
- **recovery** resolving a `Dispatched` intent left by an earlier crash, which
|
|||
|
|
writes an outcome into an attempt a verifier is already reading.
|
|||
|
|
|
|||
|
|
Each of these is a write to the run while verifiers are mid-flight, and each is
|
|||
|
|
rare enough to be missed in testing and ordinary enough to happen in production.
|
|||
|
|
|
|||
|
|
### 10.3 Retention ordering
|
|||
|
|
|
|||
|
|
Reduction (§8.6) must not outrun verification or grading. Eligibility is
|
|||
|
|
**`Graded`, `Ungraded` or `Archived`** — the condition is *grading has
|
|||
|
|
terminated*, not *grading succeeded*. Never a step-level finish timestamp: a
|
|||
|
|
step can finish, be reduced, and then run-level verification finds nothing.
|
|||
|
|
|
|||
|
|
`Ungraded` belongs in that set for a reason worth stating plainly, since the
|
|||
|
|
tighter-looking `Graded`-only rule is the one that gets written. A run that
|
|||
|
|
never gets a score — G = 1, a group that closed without it, a tenant over its
|
|||
|
|
grading ceiling — is finished, and gating retention on `Graded` alone leaves it
|
|||
|
|
irreducible forever. The tenants that hit this are the low-volume ones and the
|
|||
|
|
cost-capped ones: the two populations least able to absorb unbounded storage,
|
|||
|
|
and the two least likely to have anyone watching for it.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 11. Grading and the learning loop
|
|||
|
|
|
|||
|
|
The default loop ships working. Every component is a port.
|
|||
|
|
|
|||
|
|
### 11.1 Ports
|
|||
|
|
|
|||
|
|
Grading is **strategy-pluggable, and the strategy declares what hardware it
|
|||
|
|
needs before it is allowed to run.** That second half is not a detail: the
|
|||
|
|
strategies below differ by more than an order of magnitude in model calls and in
|
|||
|
|
VRAM, and a deployment that cannot afford one must be told at load time rather
|
|||
|
|
than by an OOM at 3am.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
#[async_trait]
|
|||
|
|
pub trait EvaluationStrategy: Send + Sync {
|
|||
|
|
fn id(&self) -> StrategyId;
|
|||
|
|
|
|||
|
|
/// Declared before any work is admitted. Validated against §14.2's limits
|
|||
|
|
/// at load; a strategy whose profile does not fit is rejected by name.
|
|||
|
|
fn resources(&self) -> ResourceProfile;
|
|||
|
|
|
|||
|
|
async fn evaluate(&self, cx: &EvalCtx) -> Result<Vec<Score>>;
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub struct ResourceProfile {
|
|||
|
|
/// Models this strategy calls. One entry means it runs on the agent's
|
|||
|
|
/// already-resident model and forces no swap (§14.2).
|
|||
|
|
pub models: Vec<ModelId>,
|
|||
|
|
/// Largest single-call context the strategy will request. A pairwise judge
|
|||
|
|
/// reads two episodes, so this is roughly twice an episode budget.
|
|||
|
|
pub max_context_tokens: u32,
|
|||
|
|
/// Model calls per episode evaluated, for capacity planning and for §14's
|
|||
|
|
/// spend projection.
|
|||
|
|
pub calls_per_episode: f32,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
#[async_trait]
|
|||
|
|
pub trait Grader: Send + Sync {
|
|||
|
|
/// Produce comparable scores for a group of episodes.
|
|||
|
|
async fn grade(&self, group: &Group, rubric: &RubricDef) -> Result<Vec<Score>>;
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
#[async_trait]
|
|||
|
|
pub trait Judge: Send + Sync {
|
|||
|
|
/// Relative comparison only. Deliberately cannot return an absolute score.
|
|||
|
|
async fn compare(&self, a: &EpisodeView, b: &EpisodeView, r: &RubricDef)
|
|||
|
|
-> Result<Verdict>;
|
|||
|
|
|
|||
|
|
/// Unary, and separate from `compare` for a structural reason: a `Core`
|
|||
|
|
/// violation caps an episode on its own terms, not relative to an opponent
|
|||
|
|
/// (§11.7). Runs before pairing.
|
|||
|
|
async fn screen(&self, e: &EpisodeView, r: &RubricDef) -> Result<Vec<CoreViolation>>;
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub enum Verdict { A, B, Draw }
|
|||
|
|
|
|||
|
|
pub struct CoreViolation {
|
|||
|
|
pub criterion: RubricCriterionId,
|
|||
|
|
pub evidence: BlobRef,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub enum Score {
|
|||
|
|
/// Default path (§11.3): one verdict against the current reference, plus
|
|||
|
|
/// the running record the sequential test consumes.
|
|||
|
|
Relative { against: RunId, verdict: Verdict, record: WinRecord },
|
|||
|
|
/// Bradley-Terry strength as a delta from control, with its interval and
|
|||
|
|
/// the group size that produced it (§11.2, §11.4). Only the tournament
|
|||
|
|
/// strategy produces this.
|
|||
|
|
Ranked { strength: f64, interval: (f64, f64), group_size: u32 },
|
|||
|
|
/// A `Core` violation caps the episode. Carries the violations and no
|
|||
|
|
/// number, so there is nothing for an aggregate to average past.
|
|||
|
|
Capped { violations: Vec<CoreViolation> },
|
|||
|
|
/// No comparison was possible (§11.6). A reason, never a neutral score.
|
|||
|
|
Ungraded { reason: UngradedReason },
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`Judge::compare` returning `Verdict` rather than `f64` is the schema decision
|
|||
|
|
that matters most, and §11.2 is why.
|
|||
|
|
|
|||
|
|
`Score` is deliberately a sum rather than a number with flags. A capped episode
|
|||
|
|
and an ungraded one are not low scores; they are different kinds of answer, and
|
|||
|
|
a type that can represent them as numbers will eventually have them averaged
|
|||
|
|
into a promotion gate by code that meant no harm.
|
|||
|
|
|
|||
|
|
**The strategy catalogue.** Cost is per episode evaluated, on a group of eight.
|
|||
|
|
|
|||
|
|
| Strategy | Model calls / episode | Models resident | Produces | Default |
|
|||
|
|
|---|---|---|---|---|
|
|||
|
|
| `DeterministicGrader` | 0 | 0 | `Ranked` on a computed number | — |
|
|||
|
|
| **`PairwiseSequential`** (§11.3) | **1–2** | **1, the agent's** | `Relative` | **yes** |
|
|||
|
|
| `TournamentGrader` (§11.4) | 3–5 | 1 | `Ranked` with intervals | opt-in |
|
|||
|
|
| `ReplayTournament` (§12.4) | 3–5 **plus N full agent runs** | 1 | `Ranked` across variants | opt-in |
|
|||
|
|
|
|||
|
|
The default is `PairwiseSequential` because it is the only one whose cost does
|
|||
|
|
not scale with how much you want to know. The tournament's `(G/2)·log₂(G)`
|
|||
|
|
comparisons buy a full ranking with composable strengths; that is genuinely more
|
|||
|
|
information, and a deployment that can afford it should turn it on. Most cannot,
|
|||
|
|
and a framework whose default path assumes a grading budget larger than the work
|
|||
|
|
being graded will simply be run with grading disabled — which is the outcome
|
|||
|
|
this whole section exists to avoid.
|
|||
|
|
|
|||
|
|
`DeterministicGrader` remains for users whose quality signal is a number they
|
|||
|
|
already compute — latency, cost, test pass count. It exists so that adopting the
|
|||
|
|
framework does not require adopting LLM-as-judge at all.
|
|||
|
|
|
|||
|
|
### 11.2 Absolute scores do not work here
|
|||
|
|
|
|||
|
|
Three failure modes, all of which this system would hit:
|
|||
|
|
|
|||
|
|
**Calibration drift.** A judge asked for 0..1 returns different numbers for the
|
|||
|
|
same episode across weeks and model versions. Drift is indistinguishable from a
|
|||
|
|
variant trend, so promotion decisions fire on grader noise.
|
|||
|
|
|
|||
|
|
**Weak discrimination.** Four competent episodes all score 0.8. No gradient, no
|
|||
|
|
selection pressure, and the loop reports "nothing beats control" because the
|
|||
|
|
grader cannot resolve them — not because they are equivalent.
|
|||
|
|
|
|||
|
|
**Saturation.** The one that kills the loop outright. As workflows improve, pass
|
|||
|
|
rate approaches 100% and pass/fail carries zero information; absolute rubric
|
|||
|
|
scores saturate identically. A tournament cannot saturate — better candidates
|
|||
|
|
just make it harder.
|
|||
|
|
|
|||
|
|
Group-normalized relative scores also give something absolute scores cannot:
|
|||
|
|
**cross-task comparability** — but only through a shared anchor, and that
|
|||
|
|
qualification is load-bearing. A Bradley-Terry fit identifies strengths only up
|
|||
|
|
to an additive constant *within one connected comparison graph*. Two groups on
|
|||
|
|
different tasks are disjoint graphs, so their strengths sit on unlinked scales,
|
|||
|
|
and averaging them directly commits the same error this section accuses point
|
|||
|
|
tallies of, one layer further in.
|
|||
|
|
|
|||
|
|
**The anchor is control.** §12.1 gives control a traffic share in every
|
|||
|
|
allocation, so every group contains at least one control episode; the fit pins
|
|||
|
|
control to zero and every other strength is read as a delta from it. A variant's
|
|||
|
|
aggregate is then a mean of like-for-like deltas rather than a mean of
|
|||
|
|
incomparable scales — which is what makes the per-variant stage in §9.1 sound
|
|||
|
|
rather than approximate. A group that happens to contain no control episode is
|
|||
|
|
not aggregatable: it still grades its own members and is still worth reading, it
|
|||
|
|
just does not feed the aggregate, and it is recorded as such rather than folded
|
|||
|
|
in on the assumption that scales match.
|
|||
|
|
|
|||
|
|
### 11.3 Pairwise sequential (default)
|
|||
|
|
|
|||
|
|
The system holds **one current version and at most one challenger**, and grading
|
|||
|
|
answers one question: has the challenger accumulated enough evidence to replace
|
|||
|
|
the current one? Not "rank these eight", not "what is each episode worth" — a
|
|||
|
|
single accept/reject that converges toward one state.
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
current version ──► episode ──┐
|
|||
|
|
├──► Judge::compare ──► verdict
|
|||
|
|
challenger ──► episode ──┘ │
|
|||
|
|
(same TaskId) ▼
|
|||
|
|
accumulate into WinRecord
|
|||
|
|
│
|
|||
|
|
┌────────────────┼────────────────┐
|
|||
|
|
▼ ▼ ▼
|
|||
|
|
accept continue reject
|
|||
|
|
challenger becomes keep sampling discard, keep
|
|||
|
|
the current current
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
One comparison per episode. Against the tournament's `(G/2)·log₂(G)` pairs
|
|||
|
|
doubled for both orderings, that is 24 judge calls dropping to 8 at G = 8, and
|
|||
|
|
the saving grows with G rather than shrinking.
|
|||
|
|
|
|||
|
|
**The reference is the current version's recorded episode on the same `TaskId`.**
|
|||
|
|
It is already on disk — no re-run, no group to fill, no `GroupEpoch` timeout.
|
|||
|
|
Where the task has never been seen before there is nothing to compare against,
|
|||
|
|
and that case degrades per §11.6 rather than being papered over.
|
|||
|
|
|
|||
|
|
**Stopping is a sequential test, not a fixed sample.** Verdicts accumulate into a
|
|||
|
|
likelihood ratio against boundaries set by α, β and the smallest win-rate shift
|
|||
|
|
worth acting on; the test stops as soon as a boundary is crossed. A clearly
|
|||
|
|
better challenger is accepted in far fewer comparisons than a fixed-n design
|
|||
|
|
would spend, and a clearly worse one is rejected early instead of running to
|
|||
|
|
completion. This is the mechanism that makes the cost adaptive: cheap decisions
|
|||
|
|
cost little, close decisions cost more, and nothing costs the worst case by
|
|||
|
|
default.
|
|||
|
|
|
|||
|
|
**Draws are recorded and excluded from the ratio.** A tie carries no evidence
|
|||
|
|
either way about which is stronger, so folding it in as half a win manufactures
|
|||
|
|
information. But a *high draw rate* is itself a result — it says the challenger
|
|||
|
|
is not meaningfully different — so the test also rejects on a draw-rate ceiling
|
|||
|
|
rather than sampling forever toward a boundary it will never reach.
|
|||
|
|
|
|||
|
|
**Order alternates rather than doubling.** §11.4's both-orderings rule pays 2× on
|
|||
|
|
every comparison to cancel position bias. Here the challenger takes position A on
|
|||
|
|
even-numbered comparisons and position B on odd ones: bias cancels across the
|
|||
|
|
sequence instead of within each pair, at no extra cost. Order-consistency is
|
|||
|
|
still measured, on a sampled fraction of comparisons, and reported as the
|
|||
|
|
grader's error bar exactly as before.
|
|||
|
|
|
|||
|
|
**A fixed reference is a cacheable prefix.** The same reference episode leads
|
|||
|
|
every comparison in a decision, so it can be cached across calls. The tournament
|
|||
|
|
cannot do this — shuffling into brackets makes every pair a novel combination by
|
|||
|
|
design, which is the point of the shuffle and the reason nothing caches.
|
|||
|
|
|
|||
|
|
What this gives up, stated rather than discovered later: **parallel exploration
|
|||
|
|
and composable strengths.** One challenger at a time is hill-climbing, which is
|
|||
|
|
slower to find improvements and can settle in a local optimum with nothing in the
|
|||
|
|
loop able to report that it has. And a `Relative` score answers "better than the
|
|||
|
|
current version on this task" — it is not a strength that composes across tasks
|
|||
|
|
the way §11.2's anchored Bradley-Terry deltas do. Deployments that can afford the
|
|||
|
|
tournament get real information for the money; this is the right default, not the
|
|||
|
|
better mechanism.
|
|||
|
|
|
|||
|
|
### 11.4 Tournament grading (opt-in)
|
|||
|
|
|
|||
|
|
Not the default path — §11.3 is — and not a fallback either. This is the
|
|||
|
|
strategy to enable when episodes are **already co-present at no extra cost**, or
|
|||
|
|
when a deployment can afford full rankings. Two cases qualify naturally:
|
|||
|
|
|
|||
|
|
- **Attempt tournaments** (§11.5). The attempts of one step are on disk the
|
|||
|
|
moment a retry happens. No agent runs to pay for, and this is the only source
|
|||
|
|
of per-step credit the system has.
|
|||
|
|
- **Replay** (§12.4), where N variants are executed against one task
|
|||
|
|
deliberately. The episodes exist because you paid for them; grading them
|
|||
|
|
pairwise would waste the group you bought.
|
|||
|
|
|
|||
|
|
Everything below is unchanged in substance from when it was the default. What
|
|||
|
|
changed is the claim: it is more information per episode, at three to five times
|
|||
|
|
the model calls, and that trade is now the user's to make explicitly rather than
|
|||
|
|
one the framework makes for them.
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
[ G comparable episodes for one task, one outcome class ]
|
|||
|
|
│
|
|||
|
|
▼
|
|||
|
|
[ shuffle into brackets ] ◄── shuffling also cancels position bias
|
|||
|
|
│
|
|||
|
|
▼
|
|||
|
|
┌─────────────────────────────────┐
|
|||
|
|
│ Swiss pairing, log₂(G) rounds │ ◄── rubric-guided Judge, relative only
|
|||
|
|
└────────────────┬────────────────┘
|
|||
|
|
│
|
|||
|
|
▼
|
|||
|
|
[ Bradley-Terry fit over all comparisons ]
|
|||
|
|
│
|
|||
|
|
▼
|
|||
|
|
[ strength per episode + confidence interval ]
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Swiss rather than round-robin: O(G log G) instead of O(G²). Eight episodes is
|
|||
|
|
twelve comparisons rather than twenty-eight. Swiss rather than single
|
|||
|
|
elimination because we want a full ranking, not a champion — eliminated
|
|||
|
|
candidates still carry signal.
|
|||
|
|
|
|||
|
|
**Bradley-Terry, not point-tally z-scores.** Accumulating tournament points and
|
|||
|
|
normalizing to mean 0 / sd 1 within a group is the obvious approach and it is
|
|||
|
|
statistically wrong for our aggregate: small groups produce extreme z-scores, so
|
|||
|
|
a variant that appears in many small groups wins on variance rather than
|
|||
|
|
quality. A Bradley-Terry fit over the pairwise outcomes yields a strength
|
|||
|
|
parameter with a real confidence interval, which composes correctly across
|
|||
|
|
groups of different sizes and feeds the sample gate in §12.3 directly.
|
|||
|
|
|
|||
|
|
Three details of that fit are decisions, not implementation freedom. Textbook
|
|||
|
|
Bradley-Terry does none of them, and each failure looks like a result rather
|
|||
|
|
than a bug.
|
|||
|
|
|
|||
|
|
**Draws need a draw model.** Plain BT is binary and has no tie term, so the
|
|||
|
|
draws this section deliberately permits have nowhere to go. Dropping them
|
|||
|
|
discards the judge's most confident statements; splitting each half-and-half
|
|||
|
|
fabricates two comparisons that never happened and tightens the interval on
|
|||
|
|
invented evidence. Use the **Davidson extension** — one additional tie parameter
|
|||
|
|
fit alongside the strengths.
|
|||
|
|
|
|||
|
|
**Small groups separate.** At G = 4..8 an episode that wins every comparison
|
|||
|
|
drives the unpenalized maximum-likelihood estimate to infinite strength. That is
|
|||
|
|
the extreme-score failure §11.2 rejects, arriving through the fit instead of
|
|||
|
|
through z-scores. A weakly-informative prior on the strengths — equivalently, a
|
|||
|
|
penalized likelihood — is **required, not tuning**. It is the mechanism that
|
|||
|
|
turns "won all three of its comparisons" into a wide interval rather than an
|
|||
|
|
unbounded one, and without it the reassuring sentence about small groups
|
|||
|
|
producing wide intervals is simply false.
|
|||
|
|
|
|||
|
|
**A strength alone means nothing.** The fit is identified only up to an additive
|
|||
|
|
constant (§11.2), so an interval on a single raw strength is an interval on an
|
|||
|
|
arbitrary origin. Pin control to zero and report every strength as a delta
|
|||
|
|
against it. The §12.3 canary gate says "BT interval excludes zero" — zero is
|
|||
|
|
control, and it is only zero because it was pinned there.
|
|||
|
|
|
|||
|
|
**Draws are permitted.** A judge forced to separate two equivalent episodes
|
|||
|
|
invents a distinction, and the optimizer will chase the invention. Draws cost
|
|||
|
|
gradient; forced choices cost correctness.
|
|||
|
|
|
|||
|
|
**Both orderings are judged.** Pairwise judges have position bias.
|
|||
|
|
Order-consistency is recorded per comparison, and the disagreement rate is the
|
|||
|
|
grader's own error bar — it belongs in the report next to the scores, and an
|
|||
|
|
inconsistent judge should widen the sample gate rather than silently promote.
|
|||
|
|
|
|||
|
|
### 11.5 Groups: where they come from
|
|||
|
|
|
|||
|
|
A tournament needs a comparison group, and production runs are one-shot on tasks
|
|||
|
|
that mostly never repeat. This is the binding constraint on the whole learning
|
|||
|
|
loop.
|
|||
|
|
|
|||
|
|
`TaskId` is the group key: a content hash of the task input **before any
|
|||
|
|
workflow touches it**. Hashing the prompt does not work — the workflow changes
|
|||
|
|
the prompt by construction, which is the entire point of a variant.
|
|||
|
|
|
|||
|
|
| Source | Group | Cost | Needs |
|
|||
|
|
|---|---|---|---|
|
|||
|
|
| **Attempts** | attempts 1..N of one step, identical context | free, already recorded | nothing |
|
|||
|
|
| **Recurring tasks** | runs sharing a `TaskId` over time | free, slow to fill | `TaskId` |
|
|||
|
|
| **Replay** | one task re-executed under N variants | N full runs | blobs, sandbox |
|
|||
|
|
|
|||
|
|
Attempt tournaments first. Retries are already on disk and attempt N+1 never
|
|||
|
|
mutates attempt N, so the attempts of one step are a group on identical context —
|
|||
|
|
the cheapest per-step credit signal available.
|
|||
|
|
|
|||
|
|
A failure and its successful retry are **not** a judge comparison. §11.8 forbids
|
|||
|
|
that pairing and the verifier has already ordered it; asking a judge which is
|
|||
|
|
better asks it to re-decide what verification decided. The pair is consumed
|
|||
|
|
*structurally* instead: what differed between attempt N and N+1 — context
|
|||
|
|
partition, tool selection, prompt — is attributed to the `StepId` as the change
|
|||
|
|
that turned a fail into a pass. The judge sees only same-outcome attempts, where
|
|||
|
|
the question it answers is "which failure got further" — which no verifier can
|
|||
|
|
answer.
|
|||
|
|
|
|||
|
|
**`TaskId` cannot be backfilled.** A run recorded without one is permanently
|
|||
|
|
ungroupable, which is why it is required at spawn with no `Default` and no
|
|||
|
|
`From<RunId>`.
|
|||
|
|
|
|||
|
|
### 11.6 Low-volume degradation
|
|||
|
|
|
|||
|
|
Most of this problem is a tournament problem, and §11.3 does not have it: a
|
|||
|
|
pairwise comparison needs one partner, and the current version's recorded
|
|||
|
|
episode on that `TaskId` is already on disk. Volume stops mattering the moment a
|
|||
|
|
task recurs even once.
|
|||
|
|
|
|||
|
|
What survives is the genuinely irreducible case — **a `TaskId` never seen
|
|||
|
|
before**. There is nothing to compare against, because nothing else has done this
|
|||
|
|
task. That is not a degradation to engineer around; it is the first observation
|
|||
|
|
of a new task, and it becomes the reference for the next one.
|
|||
|
|
|
|||
|
|
1. **Novel `TaskId`** — no comparison. Verifier outcome and deterministic
|
|||
|
|
dimensions (cost, latency, tool efficiency) still recorded; no relative score.
|
|||
|
|
Reported as `Score::Ungraded { reason: NoReference }`, never as a neutral
|
|||
|
|
score. The episode is retained as the reference for that `TaskId`.
|
|||
|
|
2. **Attempt tournaments** — available to any workflow that retries, regardless
|
|||
|
|
of volume, and unaffected by either of the above.
|
|||
|
|
3. **Grading budget exhausted** — a tenant over its §14 ceiling reports
|
|||
|
|
`Ungraded { BudgetExhausted }`. Not a quality signal; a spend signal.
|
|||
|
|
|
|||
|
|
The remainder of this section applies **only when the tournament strategy is
|
|||
|
|
enabled** (§11.4), where a group must genuinely fill:
|
|||
|
|
|
|||
|
|
4. **Group completeness trigger** — a group closes on quorum *or* on a timeout,
|
|||
|
|
grading whatever arrived, with group size attached to the confidence interval.
|
|||
|
|
|
|||
|
|
**A closed group is immutable, and the next episode starts a new one.** This is
|
|||
|
|
the question the trigger raises and does not answer on its own: a group closed
|
|||
|
|
on timeout at G = 3, then a fourth episode with the same `TaskId` arrives an hour
|
|||
|
|
later. Re-opening and re-fitting is the wrong answer — strengths from that group
|
|||
|
|
have already been published, aggregated, and possibly acted on by a promotion
|
|||
|
|
gate, and a fit that silently changes underneath a decision already made is
|
|||
|
|
worse than a small group.
|
|||
|
|
|
|||
|
|
So the group key carries a generation: `(TenantId, TaskId, VerifierOutcome,
|
|||
|
|
GroupEpoch)`. Closure increments the epoch; late arrivals accumulate into the
|
|||
|
|
next one. The cost is honest and should be stated rather than discovered — a
|
|||
|
|
low-volume tenant with a long inter-arrival time gets a run of G = 1 epochs, each
|
|||
|
|
reported `Ungraded { InsufficientGroup }`. That is a real signal about their
|
|||
|
|
volume, and the fix is a longer timeout, which is a tenant-level setting and a
|
|||
|
|
tradeoff between waiting and grading, not a bug in the trigger.
|
|||
|
|
|
|||
|
|
A tenant whose loop never engages must see that in the dashboard as a stated
|
|||
|
|
reason. Silent no-op is the worst outcome: it looks like a working loop that
|
|||
|
|
finds no improvements.
|
|||
|
|
|
|||
|
|
### 11.7 Rubrics and credit assignment
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub enum RubricLayer {
|
|||
|
|
/// Mandatory. A violation caps the result regardless of everything else.
|
|||
|
|
Core,
|
|||
|
|
/// Anti-gaming. Written explicitly against known exploits.
|
|||
|
|
Prescriptive,
|
|||
|
|
/// Context-specific, user-authored, weighed rather than binding.
|
|||
|
|
Contextual,
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`Core` violations **cap** rather than subtract. A weighted sum lets a variant buy
|
|||
|
|
past a safety failure with speed, which is the exact failure prescriptive
|
|||
|
|
rubrics exist to prevent.
|
|||
|
|
|
|||
|
|
The cap needs somewhere to live, and `Verdict` is the wrong place — a `Core`
|
|||
|
|
violation is a fact about one episode, not about a pair, and a judge asked to
|
|||
|
|
express it through a comparison can only rank the offender lower. It comes from
|
|||
|
|
`Judge::screen` (§11.1) instead, which runs before pairing and yields
|
|||
|
|
`Score::Capped`. A capped episode is **excluded from the bracket, not ranked
|
|||
|
|
last in it**: left in, it still contributes comparisons that shape everyone
|
|||
|
|
else's strength, and a variant with one safety failure and seven strong episodes
|
|||
|
|
aggregates to a promotion.
|
|||
|
|
|
|||
|
|
Never let a rubric judge what a verifier can check. Every criterion that can be
|
|||
|
|
made mechanical should be a `Verifier`, not a rubric line — deterministic,
|
|||
|
|
cheap, and not subject to judge drift.
|
|||
|
|
|
|||
|
|
Per-step credit attributes a group's outcome to `StepId`s, which is why §4.3's
|
|||
|
|
stability contract is load-bearing rather than cosmetic.
|
|||
|
|
|
|||
|
|
### 11.8 Grading never overrides the verifier
|
|||
|
|
|
|||
|
|
Episodes are bracketed **within** a verifier outcome class, never across. A
|
|||
|
|
verified pass beats a verified fail by definition and that pairing is never shown
|
|||
|
|
to a judge. Ranking failures against each other is not wasted work: "failed at
|
|||
|
|
step 2" versus "failed at step 7 after recovering twice" is exactly the signal a
|
|||
|
|
pass rate cannot see.
|
|||
|
|
|
|||
|
|
This binds attempt groups too (§11.5), which is where the rule is easiest to
|
|||
|
|
break: a step's failed attempt and its successful retry sit side by side on disk
|
|||
|
|
and look like a free comparison. They are a free *credit* signal and not a
|
|||
|
|
comparison at all. The bracketing rule has no exemptions — if a pairing crosses
|
|||
|
|
an outcome class, it is evidence for attribution, never input to a judge.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 12. Optimization loops
|
|||
|
|
|
|||
|
|
Two loops at different clock speeds. The fast loop **selects** among existing
|
|||
|
|
workflow versions; the slow loop **generates** new ones.
|
|||
|
|
|
|||
|
|
### 12.1 Shape
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
┌──── GENERATE (slow, human-gated by default) ─────┐
|
|||
|
|
│ failure evidence ──► propose ──► challenger │
|
|||
|
|
│ ▲ │ │
|
|||
|
|
└────────┼──────────────────────────────┼──────────┘
|
|||
|
|
│ │ register — at most one
|
|||
|
|
┌────────┼──── SELECT (fast) ───────────┼──────────┐
|
|||
|
|
│ │ ┌── allocation ────────────┘ │
|
|||
|
|
│ │ │ current 95% · challenger 5% │
|
|||
|
|
│ │ └────┬─────── │
|
|||
|
|
│ │ │ spawn — pin WorkflowVersion │
|
|||
|
|
│ │ ▼ │
|
|||
|
|
│ │ run ──► verify ──► compare vs │
|
|||
|
|
│ │ current (§11.3) │
|
|||
|
|
│ │ │ │
|
|||
|
|
│ └──────────────────────────────┤ │
|
|||
|
|
│ ▼ │
|
|||
|
|
│ sequential test boundary │
|
|||
|
|
│ │ │ │
|
|||
|
|
│ accept │ │ reject │
|
|||
|
|
│ ▼ ▼ │
|
|||
|
|
│ challenger discard, │
|
|||
|
|
│ becomes current keep current │
|
|||
|
|
└───────────────────────────────────────────────────┘
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**One current version, at most one challenger.** The loop converges toward a
|
|||
|
|
single state rather than maintaining a population. This is the change that makes
|
|||
|
|
everything else affordable: no N-way traffic split, no per-variant aggregation
|
|||
|
|
across groups, no allocation state to contend over, and one comparison per
|
|||
|
|
episode instead of a bracket.
|
|||
|
|
|
|||
|
|
The cost is exploration. A single challenger at a time is hill-climbing — it
|
|||
|
|
finds improvements more slowly than a population would, and it can sit in a local
|
|||
|
|
optimum indefinitely with nothing in the loop able to say so. §12.5's held-out
|
|||
|
|
report is the only instrument that will notice, which makes it more important
|
|||
|
|
here than it was under the population design, not less.
|
|||
|
|
|
|||
|
|
**Multi-variant selection remains available** for deployments that can afford it:
|
|||
|
|
enable the tournament strategy (§11.4), allow N challengers, and §12.3's full
|
|||
|
|
rung ladder applies with Bradley-Terry aggregation across groups. The machinery
|
|||
|
|
is the same; what changes is how many versions are live at once and which
|
|||
|
|
statistical object closes the decision.
|
|||
|
|
|
|||
|
|
The slow loop fires when the fast loop **rejects a challenger without finding a
|
|||
|
|
replacement** — a trigger, not a timer.
|
|||
|
|
|
|||
|
|
### 12.2 Versions form a DAG
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
v1 ────┬───► v2 ───┐
|
|||
|
|
control │ └──► v4 (merge)
|
|||
|
|
└───► v3 ───────┘
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Content-addressed, parent-pointered, never edited. Editing a version in place
|
|||
|
|
destroys every result already attributed to it.
|
|||
|
|
|
|||
|
|
### 12.3 Promotion gates
|
|||
|
|
|
|||
|
|
Every rung needs a criterion. The previous revision drew the ladder and stated a
|
|||
|
|
rule for only the first rung.
|
|||
|
|
|
|||
|
|
**Default ladder — one challenger, pairwise (§11.3).** Three rungs, because a
|
|||
|
|
graduated ramp is a population instrument and there is no population here:
|
|||
|
|
|
|||
|
|
| Rung | Traffic | Entry criterion |
|
|||
|
|
|---|---|---|
|
|||
|
|
| shadow | 0% | registered, validated, sandbox-clean, `ResourceProfile` fits (§14.2) |
|
|||
|
|
| trial | 5% | no `Core` violation on any trial episode |
|
|||
|
|
| current | 100% | sequential test crosses the accept boundary; drift check clean |
|
|||
|
|
|
|||
|
|
The swap at the last rung is deliberate and worth naming: the challenger takes
|
|||
|
|
all traffic at once rather than ramping. A ramp exists to limit blast radius
|
|||
|
|
while evidence accumulates, and here the evidence has already accumulated — the
|
|||
|
|
sequential test does not cross its boundary until the win rate is established at
|
|||
|
|
the configured α. Ramping after that spends traffic to re-learn what the test
|
|||
|
|
already concluded. What guards the swap instead is the rollback rule below, which
|
|||
|
|
fires on a single `Score::Capped` and does not wait for a boundary.
|
|||
|
|
|
|||
|
|
**Resourced ladder — N challengers, tournament (§11.4).** For deployments that
|
|||
|
|
enabled the tournament strategy and can carry multiple live versions:
|
|||
|
|
|
|||
|
|
| Rung | Traffic | Entry criterion |
|
|||
|
|
|---|---|---|
|
|||
|
|
| shadow | 0% | registered, validated, sandbox-clean, profile fits |
|
|||
|
|
| canary | 5% | beats current on **selection** replay, BT interval excludes zero |
|
|||
|
|
| ramp | 20→50% | no `Core` violation, cost within budget, sequential test at α |
|
|||
|
|
| current | 100% | sustained over N groups, drift check clean |
|
|||
|
|
|
|||
|
|
Rollback is automatic and unconditional on any `Score::Capped` attributed to the
|
|||
|
|
variant, or a verifier pass-rate regression beyond a configured margin.
|
|||
|
|
`Score::Capped` is the only signal for the first of those (§11.1); a gate that
|
|||
|
|
reads a low *number* instead is reading something the cap exists to prevent from
|
|||
|
|
existing. Rollback is a traffic
|
|||
|
|
change, never a version delete — the failed variant stays in the DAG with its
|
|||
|
|
results.
|
|||
|
|
|
|||
|
|
### 12.4 Replay is re-execution
|
|||
|
|
|
|||
|
|
"Replay against recorded episodes" means **re-running the task under a new
|
|||
|
|
workflow version**, not pushing a recorded trajectory through new logic.
|
|||
|
|
Trajectory replay tells you only where behaviour would first diverge, and
|
|||
|
|
everything after divergence is unknown — near-worthless for grading.
|
|||
|
|
|
|||
|
|
Two consequences the prior revision missed:
|
|||
|
|
|
|||
|
|
**Replay is not free.** "0% live traffic" means no user sees the result, not
|
|||
|
|
that it costs nothing. N variants × M tasks is N·M full agent runs plus judge
|
|||
|
|
calls. Shadow is the most expensive rung, not the cheapest.
|
|||
|
|
|
|||
|
|
**Replay executes real tools.** Re-running a workflow that pushes commits pushes
|
|||
|
|
commits. Shadow execution runs in a **sandbox with `Unsafe` effects denied**, and
|
|||
|
|
a workflow whose steps cannot run sandboxed is ineligible for shadow evaluation
|
|||
|
|
and must say so at load time rather than at 3am.
|
|||
|
|
|
|||
|
|
### 12.5 Held-out set, and the leak
|
|||
|
|
|
|||
|
|
Selecting on a fixed set of recorded tasks overfits to those tasks, silently:
|
|||
|
|
shadow scores improve while live performance does not.
|
|||
|
|
|
|||
|
|
Partition tasks into a selection set and a held-out set. Promote on selection,
|
|||
|
|
report held-out without optimizing against it, and treat a widening gap as the
|
|||
|
|
overfitting alarm. Concretely: **no rung in §12.3 takes held-out as an entry
|
|||
|
|
criterion.** A gate that reads held-out has converted it into a second selection
|
|||
|
|
set and left nothing measuring generalization.
|
|||
|
|
|
|||
|
|
**The proposer reads the selection set only.** The slow loop consumes failure
|
|||
|
|
evidence to generate candidates; if it reads held-out failures, the held-out set
|
|||
|
|
is contaminated through the generator instead of the selector. This leak is
|
|||
|
|
easy to introduce and invisible once present.
|
|||
|
|
|
|||
|
|
Held-out catches overfitting to *episodes*. It does not catch drift in the *task
|
|||
|
|
mix* — that needs a separate distribution check on incoming `TaskId`
|
|||
|
|
characteristics over time.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 13. Extensions, tools, and trust
|
|||
|
|
|
|||
|
|
### 13.1 Capability-gated hostcalls
|
|||
|
|
|
|||
|
|
Tools register through capability-gated hostcalls. The gate is where effect
|
|||
|
|
class is declared, and a tool that cannot state whether it is safe to retry does
|
|||
|
|
not register.
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub struct ToolRegistration {
|
|||
|
|
pub id: ToolId,
|
|||
|
|
/// No Default. The author is the only party who knows this.
|
|||
|
|
pub effects: EffectClass,
|
|||
|
|
pub caps: CapabilitySet,
|
|||
|
|
pub timeout: Duration,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
/// Each class carries what its §8.4 recovery path actually needs. A bare
|
|||
|
|
/// discriminant would let a tool claim `Idempotent` while withholding the one
|
|||
|
|
/// thing that makes the claim actionable.
|
|||
|
|
pub enum EffectClass {
|
|||
|
|
Idempotent { key: KeyDerivation },
|
|||
|
|
Queryable { lookup: RequestIdLookup },
|
|||
|
|
Unsafe,
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Tool calls appear in the episode as **first-class steps**. The hostcall boundary
|
|||
|
|
already knows the identity, arguments and capabilities; recording them as opaque
|
|||
|
|
invocations discards information the runtime is holding anyway.
|
|||
|
|
|
|||
|
|
### 13.2 Declaration is not enforcement
|
|||
|
|
|
|||
|
|
A framework cannot trust a user-supplied `EffectClass`. A tool declared
|
|||
|
|
`Idempotent` that is not will be retried after an indeterminate crash, and the
|
|||
|
|
damage is the user's data.
|
|||
|
|
|
|||
|
|
Three layers, in order of strength:
|
|||
|
|
|
|||
|
|
1. **Declaration** — required, recorded, auditable.
|
|||
|
|
2. **Keyed capability** — `Idempotent` is a claim about *retry*, not about
|
|||
|
|
abstaining from writes. Denying an `Idempotent` tool network and filesystem
|
|||
|
|
writes would deny the recovery path in §8.4, which retries the write under an
|
|||
|
|
idempotency key. So the restriction is on the *shape* of the write, not its
|
|||
|
|
existence: an `Idempotent` tool declares how its key derives from its
|
|||
|
|
arguments (§13.1), the kernel derives and supplies that key on every call,
|
|||
|
|
and a write issued without it is denied. A tool that cannot derive a stable
|
|||
|
|
key cannot be `Idempotent` — which is the same claim as before, now refused
|
|||
|
|
at registration rather than discovered after a double-charge.
|
|||
|
|
3. **Sandbox** — shadow and replay execution deny `Unsafe` effects outright.
|
|||
|
|
|
|||
|
|
Layer 2 is the one that makes layer 1 more than paperwork.
|
|||
|
|
|
|||
|
|
### 13.3 The relay
|
|||
|
|
|
|||
|
|
Export runs from the outbox, out of process, with its own retry and its own
|
|||
|
|
failure domain. A user's broker being down is not an agent outage.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 14. Metering, cost, and capacity
|
|||
|
|
|
|||
|
|
### 14.1 Metering
|
|||
|
|
|
|||
|
|
A framework that spends users' model budget on grading must account for it.
|
|||
|
|
|
|||
|
|
- Every model call — agent, judge, or proposer — is attributed to
|
|||
|
|
`(TenantId, RunId, Purpose)` where purpose distinguishes work from grading
|
|||
|
|
from replay.
|
|||
|
|
- Per-tenant ceilings on grading spend, enforced at the group boundary in §9.1
|
|||
|
|
where group size is known and a tournament can be skipped or downsampled
|
|||
|
|
before it starts.
|
|||
|
|
- Grading cost is reported next to grading value. A loop that costs more than
|
|||
|
|
the work it grades may still be worth running; it should never be an
|
|||
|
|
unpleasant discovery.
|
|||
|
|
|
|||
|
|
Replay is the expensive one (§12.4) and needs its own ceiling separate from
|
|||
|
|
judging.
|
|||
|
|
|
|||
|
|
### 14.2 Model capacity is a hard constraint, not a budget
|
|||
|
|
|
|||
|
|
Metering counts tokens after the fact. On self-hosted weights the binding limit
|
|||
|
|
arrives earlier and harder: **VRAM**, and the cost of moving weights in and out
|
|||
|
|
of it. A model that is not resident cannot be inferred against, and making it
|
|||
|
|
resident means evicting something else and paying a load measured in tens of
|
|||
|
|
seconds. A design that treats "call the judge model" as equivalent in cost to
|
|||
|
|
"call the agent model" is wrong by two orders of magnitude on this hardware.
|
|||
|
|
|
|||
|
|
The limits are declared, not discovered:
|
|||
|
|
|
|||
|
|
```rust
|
|||
|
|
pub struct CapacityLimits {
|
|||
|
|
/// Per-device VRAM this framework may use. Not the card's total — leave
|
|||
|
|
/// headroom for anything else sharing the device.
|
|||
|
|
pub vram_bytes_per_device: u64,
|
|||
|
|
pub devices: u32,
|
|||
|
|
/// Hard cap on simultaneously resident models across all devices.
|
|||
|
|
pub max_resident_models: u32,
|
|||
|
|
/// Never evictable. The agent's model belongs here; if it can be evicted,
|
|||
|
|
/// grading can stall agent work, which §1 forbids outright.
|
|||
|
|
pub pinned: Vec<ModelId>,
|
|||
|
|
}
|
|||
|
|
|
|||
|
|
pub struct ModelProfile {
|
|||
|
|
pub id: ModelId,
|
|||
|
|
pub weights_bytes: u64,
|
|||
|
|
/// KV cache cost per token at the deployed dtype and parallelism. The term
|
|||
|
|
/// that decides how long a context may be — see below.
|
|||
|
|
pub kv_bytes_per_token: u64,
|
|||
|
|
/// Devices this model spans under tensor parallelism.
|
|||
|
|
pub devices_required: u32,
|
|||
|
|
}
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**Residency invariant**, checked at load and before any admission:
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
sum(weights of resident models) + peak_concurrent_kv ≤ vram_bytes_per_device × devices
|
|||
|
|
resident_model_count ≤ max_resident_models
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
**One resident model is the default configuration.** The agent's model is pinned;
|
|||
|
|
`Judge` and the §12 proposer run on that same model. §11.3's `ResourceProfile`
|
|||
|
|
declares a single `ModelId` precisely so that the default grading path adds no
|
|||
|
|
resident model and forces no swap. A strategy naming a second model is legal and
|
|||
|
|
is rejected at load unless the invariant still holds with both resident — never
|
|||
|
|
by swapping between them per call, which is the failure mode this section
|
|||
|
|
exists to prevent.
|
|||
|
|
|
|||
|
|
**Context length is a VRAM quantity, and this bites hardest on the judge.** A
|
|||
|
|
pairwise comparison reads two episodes, so its context is roughly twice an
|
|||
|
|
episode budget. Working the invariant backwards gives the ceiling:
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
max_context_tokens = (vram_bytes_per_device × devices − sum(weights)) / kv_bytes_per_token
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
Put numbers on it, because the result is not marginal. At GQA fp16, per-token KV
|
|||
|
|
runs roughly 0.13 MB for an 8B-class model and 0.33 MB for a 70B-class one. Two
|
|||
|
|
episodes at §8.6's 200k-token retention ceiling is 400k tokens of context:
|
|||
|
|
**52 GB of KV cache at 8B, 131 GB at 70B**, before weights. Neither fits an 80 GB
|
|||
|
|
device. A judge reading two full-ceiling episodes is not expensive — it is
|
|||
|
|
impossible.
|
|||
|
|
|
|||
|
|
Three consequences, all forced rather than chosen:
|
|||
|
|
|
|||
|
|
- **The retention ceiling must be derived from this, not set beside it.** §8.6's
|
|||
|
|
200k default is a token-budget number that was picked without reference to any
|
|||
|
|
device. The reduced episode is what the judge reads, so the reduction target
|
|||
|
|
is `max_context_tokens / 2`, and where that is smaller than the retention
|
|||
|
|
ceiling, the judge reads a further-reduced view.
|
|||
|
|
- **Admission control, not backpressure.** Work whose `ResourceProfile` does not
|
|||
|
|
fit the current residency is refused at admission with the limit named. Queuing
|
|||
|
|
it would stall behind an eviction that §1 does not permit.
|
|||
|
|
- **Grading yields to agent work.** When both contend for the same resident
|
|||
|
|
model, agent inference wins and grading queues. A framework that lets a judge
|
|||
|
|
call delay the work it is judging has inverted its own first principle.
|
|||
|
|
|
|||
|
|
Distributed GPUs change the arithmetic, not the rule. `devices_required`
|
|||
|
|
expresses tensor parallelism across cards; `max_resident_models` is a
|
|||
|
|
fleet-wide count, so two nodes each holding the agent model are two resident
|
|||
|
|
instances, not one. Residency is per device, and a model resident on node A does
|
|||
|
|
not make node B's runs admissible.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 15. Observability
|
|||
|
|
|
|||
|
|
The framework observes agents; it must also be observable.
|
|||
|
|
|
|||
|
|
- Kernel state transitions as metrics, tagged by tenant and workflow version.
|
|||
|
|
- Lag on every stage boundary in §9.1. Tournament and reduction backlogs are the
|
|||
|
|
two that grow silently.
|
|||
|
|
- `Indeterminate` attempt count as a first-class alert. It should be near zero;
|
|||
|
|
a nonzero rate means either a crash loop or a misdeclared effect class. This
|
|||
|
|
threshold only holds because `Cancelled` is a separate state (§5.1) — route
|
|||
|
|
cancellations here and the alert has a noisy floor, which is the same as not
|
|||
|
|
having it.
|
|||
|
|
- `Ungraded` run count by reason (§11.6), separately from `Graded`. A tenant
|
|||
|
|
whose runs are mostly `InsufficientGroup` has a loop that is not engaging, and
|
|||
|
|
that reads as healthy on any dashboard that only counts failures.
|
|||
|
|
- Judge order-inconsistency rate (§11.4) as a grader-health metric, measured on
|
|||
|
|
a sampled fraction of comparisons under §11.3's alternating-order scheme.
|
|||
|
|
- **Judge-versus-verifier agreement on the calibration set** (§16). Under one
|
|||
|
|
resident model the judge is the agent's model, and this is the only instrument
|
|||
|
|
that can see the grader drifting toward its own bias.
|
|||
|
|
- **Model residency and swap count** (§14.2). A nonzero swap rate on a
|
|||
|
|
single-model deployment means something is requesting a non-resident model, and
|
|||
|
|
the load cost will dominate everything else in the trace.
|
|||
|
|
- Admission refusals by reason, separating "capacity" from "budget". They look
|
|||
|
|
identical in a queue-depth graph and have opposite fixes.
|
|||
|
|
- Held-out versus selection gap (§12.5) as the overfitting alarm.
|
|||
|
|
- Trace context propagated through `Ctx`, never through task-locals.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 16. Deliberately not built
|
|||
|
|
|
|||
|
|
- Anything before the record is trustworthy. No dashboards and no learning loop
|
|||
|
|
until one full run works end to end against a stub model.
|
|||
|
|
- Grading that decides. It attributes; the verifier decides.
|
|||
|
|
- Automatic workflow mutation without human approval. The generator is gated by
|
|||
|
|
default; a user may ungate it once their loop is calibrated against verifier
|
|||
|
|
ground truth, and that is their decision to make explicitly.
|
|||
|
|
- Self-critique using the same model family being graded, **ungated**. §14.2's
|
|||
|
|
one-resident-model default means the judge normally *is* the agent's model, so
|
|||
|
|
this is no longer a prohibition but a precondition: it is admissible only with
|
|||
|
|
the bootstrap the prior revision named — validation against verifiable tasks.
|
|||
|
|
Concretely, a calibration set of tasks with known verifier ground truth is
|
|||
|
|
replayed through the judge on a schedule, and judge-versus-verifier agreement
|
|||
|
|
is tracked as a health metric (§15). It costs no agent runs, since the episodes
|
|||
|
|
are already recorded. Without it the grader's bias is unmeasured and the
|
|||
|
|
optimizer will find it — and under one-state convergence (§12.1) there is no
|
|||
|
|
competing variant whose divergence would make that visible.
|
|||
|
|
- Cross-tenant blob deduplication (§3).
|
|||
|
|
- Semantic retrieval over episodes. No consumer yet.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 17. Decisions
|
|||
|
|
|
|||
|
|
| Question | Decision | Where |
|
|||
|
|
|---|---|---|
|
|||
|
|
| Framework or application | **Framework.** Kernel/domain split; workflows are data | §2, §4 |
|
|||
|
|
| Async runtime | **Tokio.** Reverses the prior asupersync choice; ecosystem access is decisive for a distributed framework | §6 |
|
|||
|
|
| Storage | **Two modes**, one port set: `redb` embedded, Postgres + object store distributed | §7 |
|
|||
|
|
| Port signatures | **Async everywhere**, including local implementations | §7 |
|
|||
|
|
| Durability | Engine gives atomicity; **the log gives history**. Append + apply + advance in one transaction | §8 |
|
|||
|
|
| Side effects | **Write-ahead intent in three phases** (pending / dispatched / committed), effect-class recovery, `Indeterminate` as a real state | §8.4 |
|
|||
|
|
| Rewind | **Fork**, never truncate. `BranchId` in the log key | §8.5 |
|
|||
|
|
| Log evolution | **Versioned records, upcasters on read.** Variants never removed | §8.7 |
|
|||
|
|
| Grading | **Relative only**, and **strategy-pluggable**. Default `PairwiseSequential`: one current, one challenger, one comparison per episode, sequential-test stop. `Judge` returns a verdict not a number | §11.1, §11.3 |
|
|||
|
|
| Tournament | **Opt-in, not default.** Swiss + Bradley-Terry with a Davidson tie term and a prior, strengths as deltas from a pinned anchor. Enabled where episodes are co-present anyway — attempts, replay | §11.4 |
|
|||
|
|
| Optimization target | **Converge to one state.** One current version plus at most one challenger; N-variant population selection is the resourced option | §12.1, §12.3 |
|
|||
|
|
| Model capacity | **One resident model by default**, agent's model pinned, judge and proposer share it. VRAM residency invariant checked at load; context ceiling derived from KV cost, not chosen | §14.2 |
|
|||
|
|
| Self-critique | **Admissible under calibration.** Same-model judging is the default consequence of one resident model, gated on judge-versus-verifier agreement tracking | §16 |
|
|||
|
|
| Tenancy | **Tenant key on every row and every blob namespace**, from commit one | §3 |
|
|||
|
|
| Multi-agent orchestration | **Deferred.** The log already provides durable state and resumability; revisit when cross-agent coordination is real | §18 |
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 18. Open questions
|
|||
|
|
|
|||
|
|
- **What defines `TaskId` for a given user.** Ticket id, input fixture, or a hash
|
|||
|
|
of the pre-workflow goal. Depends on where work enters their system, so the
|
|||
|
|
framework provides the type and a default hasher and lets it be overridden.
|
|||
|
|
Must be settled before any run is recorded, since it cannot be backfilled.
|
|||
|
|
- **Minimum useful group size.** G ≥ 2 runs, but a two-episode tournament is one
|
|||
|
|
comparison and carries little. Where the useful floor sits is empirical.
|
|||
|
|
- **Group timeout default.** §11.6 closes a group on quorum or timeout and
|
|||
|
|
increments the epoch. The timeout trades grading latency against group size,
|
|||
|
|
and the right default depends on tenant arrival rate — which the framework can
|
|||
|
|
measure but has no data for yet. Only bites when the tournament is enabled.
|
|||
|
|
- **Sequential test boundaries.** §11.3 stops on α, β and a minimum detectable
|
|||
|
|
win-rate shift. All three are policy, not physics: too tight and no challenger
|
|||
|
|
is ever accepted, too loose and the loop churns the current version on noise.
|
|||
|
|
Needs calibration against a workflow whose true improvement is known.
|
|||
|
|
- **Draw-rate ceiling.** §11.3 rejects a challenger that draws too often, since
|
|||
|
|
the test would otherwise never terminate. Where the ceiling sits is empirical
|
|||
|
|
and interacts with judge quality — a weak judge draws more.
|
|||
|
|
- **Local optima under one-state convergence.** §12.1 accepts hill-climbing.
|
|||
|
|
Nothing currently detects a loop that has stalled in a local optimum versus one
|
|||
|
|
correctly reporting no improvement exists. The held-out gap (§12.5) is the
|
|||
|
|
nearest instrument and was not designed for this.
|
|||
|
|
- **`kv_bytes_per_token` measurement.** §14.2's context ceiling depends on it,
|
|||
|
|
and it varies with dtype, quantization, attention implementation and
|
|||
|
|
parallelism. Measured per deployment or read from a profile the operator
|
|||
|
|
supplies — the framework should refuse a guess.
|
|||
|
|
- **Grading spend ratio.** §14 meters it; nobody has set the ceiling.
|
|||
|
|
- **Reduction summarizer.** §8.6 reduces blob bodies to "a summary" without
|
|||
|
|
saying what produces it. A model call makes reduction non-deterministic, which
|
|||
|
|
interacts badly with replay. Extractive or structural reduction may suffice.
|
|||
|
|
- **Effect classification for built-in tools.** An audit, not a decision. Until
|
|||
|
|
done, default `Unsafe` and never auto-retry.
|
|||
|
|
- **Postgres schema for the log at scale.** Partitioning by tenant and time,
|
|||
|
|
index strategy for branch scans, and whether the outbox is a table or a
|
|||
|
|
logical replication slot.
|
|||
|
|
- **Cancellation rigour under tokio.** §6 accepts cooperative cancellation as a
|
|||
|
|
residual risk. Whether a lint, a wrapper type, or a `loom` harness is the right
|
|||
|
|
enforcement is unresolved.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 19. Lessons carried forward
|
|||
|
|
|
|||
|
|
Each of these cost real time in the prior TypeScript implementation. Most are
|
|||
|
|
now structural; the rest stay written down.
|
|||
|
|
|
|||
|
|
**Structurally handled:** half a concurrency guarantee is worse than none
|
|||
|
|
(awaitable delivery without cancellation); unbounded state (retention declared
|
|||
|
|
at construction); ordering under concurrency (per-run scope, not a global queue).
|
|||
|
|
|
|||
|
|
**Still on us:**
|
|||
|
|
|
|||
|
|
- *A green suite says nothing about coverage.* An idempotency feature generated
|
|||
|
|
its own keys and checked for duplicates among them — unreachable for a whole
|
|||
|
|
phase, with tests asserting the count was zero.
|
|||
|
|
- *Suspect the guards before the mechanism.* A "batching" failure was the depth
|
|||
|
|
guard, because two unrelated limits shared a default value.
|
|||
|
|
- *A test run that prints nothing cannot distinguish slow from hung.* Per-test
|
|||
|
|
progress and per-test timeouts from the first commit.
|
|||
|
|
- *Collapsing a state removes the seam that needed it.* Resolving pass/fail
|
|||
|
|
inside `verify()` blocked three separate features.
|
|||
|
|
- *A placeholder that type-checks is invisible.* A hardcoded `"current"` version
|
|||
|
|
hash compiled, passed tests, and made every result unattributable. A newtype
|
|||
|
|
with no `Default` refuses to compile instead.
|
|||
|
|
- *A port finished while known to be unimplementable is not finished.* The
|
|||
|
|
synchronous dedupe store documented, in its own doc comment, that a networked
|
|||
|
|
implementation could not honour its signature.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
## 20. Build order
|
|||
|
|
|
|||
|
|
Detailed, dependency-ordered tasks: [rust-agentic-task.md](rust-agentic-task.md).
|
|||
|
|
|
|||
|
|
The walking skeleton is one full run of one workflow against a stub model, in
|
|||
|
|
embedded mode, with no network — recorded, verified, and re-derivable from the
|
|||
|
|
log. Nothing after it is worth starting until it runs.
|