# T5.2 — `CapacityLimits` and the residency invariant | Field | Value | |---|---| | Phase | P5 — Grading | | Size | M — 1 to 3 days | | Status | Not started | | Flags | — | | Spec | inlined below | | Blocks | T5.3 | ## Goal Declared VRAM limits, the residency invariant, and a `max_context_tokens` that is **derived** from them rather than configured beside them. ## Facts (inlined — no spec read needed) ```rust pub struct CapacityLimits { /// Per-device VRAM this framework may use. Not the card's total — leave /// headroom for anything else sharing the device. pub vram_bytes_per_device: u64, pub devices: u32, /// Hard cap on simultaneously resident models across all devices. pub max_resident_models: u32, /// Never evictable. The agent's model belongs here; if it can be evicted, /// grading can stall agent work, which the first principle forbids. pub pinned: Vec, } pub struct ModelProfile { pub id: ModelId, pub weights_bytes: u64, /// KV cache cost per token at the deployed dtype and parallelism. pub kv_bytes_per_token: u64, /// Devices this model spans under tensor parallelism. pub devices_required: u32, } ``` Invariant, checked at load **and** before any admission: ``` sum(weights of resident models) + peak_concurrent_kv ≤ vram_bytes_per_device × devices resident_model_count ≤ max_resident_models ``` Derived ceiling — never a separate config knob: ``` max_context_tokens = (vram_bytes_per_device × devices − sum(weights)) / kv_bytes_per_token ``` - Metering counts tokens after the fact; on self-hosted weights the binding limit arrives earlier and harder. A model that is not resident cannot be inferred against, and making it resident means evicting something else and paying a load measured in tens of seconds. - **One resident model is the default configuration.** The agent's model is pinned; the judge and the proposer run on that same model. A strategy naming a second model is legal and is rejected at load unless the invariant holds with both resident — **never** by swapping between them per call. - The worked case, because the result is not marginal: at GQA fp16, per-token KV runs roughly 0.13 MB for an 8B-class model and 0.33 MB for a 70B-class one. Two episodes at a 200k-token retention ceiling is 400k tokens of context — **52 GB of KV cache at 8B, 131 GB at 70B**, before weights. Neither fits an 80 GB device. A judge reading two full-ceiling episodes is not expensive, it is impossible. - Therefore the retention ceiling is derived from this, not set beside it: the reduction target is `max_context_tokens / 2`, and where that is smaller than the configured ceiling, **the bound wins**. - Distributed GPUs change the arithmetic, not the rule. `devices_required` expresses tensor parallelism; `max_resident_models` is a fleet-wide count, so two nodes each holding the agent model are two resident instances, not one. - `kv_bytes_per_token` varies with dtype, quantization, attention implementation and parallelism. It is measured per deployment or read from an operator-supplied profile — **the framework refuses a guess.** ## Steps 1. Define `CapacityLimits` and `ModelProfile` as above. No `Default` for `kv_bytes_per_token` — an absent value is a load error, not a guess. 2. Write `check_residency(&CapacityLimits, &[ModelProfile]) -> Result<(), CapacityError>` implementing both inequalities. `CapacityError` names the model and the shortfall in bytes. 3. Write `derive_max_context_tokens(...)` from the formula above. Expose it as a function, not a settable field. 4. Call the check at load, and again before admission (T8.6 enforces the runtime half). 5. Refuse a context request over the derived ceiling with the computed limit in the error message. 6. Encode the worked case as a test: 70B-class profile, 0.33 MB/token, 80 GB device, two 200k-token episodes → refused. ## Acceptance - A config whose resident set exceeds VRAM is rejected **at load**, naming the model and the shortfall — not at first inference. - A judge context request exceeding the derived ceiling is refused with the computed limit in the error. - The worked case asserts the arithmetic: 70B-class, 0.33 MB/token, 80 GB device, two 200k-token episodes must be refused. ## Verify **Harness:** pure arithmetic — no GPU required. Model profiles as fixtures with declared `weights_bytes` and `kv_bytes_per_token`. **Integration test** — `tests/it_capacity_invariant.rs`: 1. **The worked case, hardcoded:** 70B-class profile, `kv_bytes_per_token = 0.33 MB`, one 80 GB device. Two 200k-token episodes = 400k tokens = 131 GB of KV before weights. Assert the request is **refused** and that the error carries the computed limit. 2. Same arithmetic at 8B / 0.13 MB per token: 52 GB. Assert also refused on an 80 GB device once weights are counted. 3. Over-capacity config at load: resident set exceeding VRAM → rejected **at load**, error naming **the model and the byte shortfall**. Assert the message contains both. 4. Assert `max_context_tokens` is only reachable as a **function**, not a settable field — `trybuild` if it is a private field with no setter. 5. Multi-device: `devices_required = 4` under tensor parallelism; assert the invariant uses `vram × devices` and that `max_resident_models` is counted fleet-wide, so the same model on two nodes counts as two. 6. Missing `kv_bytes_per_token` → load error, **not** a default. Assert no `Default` impl exists. 7. Pinned model: assert no code path can evict a pinned model to satisfy the invariant — the eviction candidate list excludes pinned entries. **Command:** `cargo test -p capacity invariant` **False pass:** - Asserting only that "some error" is returned. The acceptance criterion is that the **arithmetic** is right — pin the expected byte figures in the test, so an off-by-a-factor error in the KV formula is caught rather than rounded away. - Step 3 checked at first inference rather than at load. Both produce an error; only one produces it before 3am. - Testing with a permissive config where everything fits. Every implementation passes that. ## Traps - A `max_context_tokens` config field. It will be set to something plausible and discovered wrong at the first judge call. - Treating "call the judge model" as equivalent in cost to "call the agent model". On this hardware that is wrong by two orders of magnitude. - Allowing an eviction of a pinned model to satisfy the invariant. --- Background (not required to do this task): [rust-agentic-sys.md](../../../rust-agentic-sys.md) §8.6, §14.2, §18 · [rust-agentic-task.md](../../../rust-agentic-task.md)