(tasks) add tasks for harness
This commit is contained in:
@@ -0,0 +1,105 @@
|
||||
# T6.4 — Per-variant aggregation
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | P6 — Learning loop |
|
||||
| Size | M — 1 to 3 days |
|
||||
| Status | Not started |
|
||||
| Flags | opt-in |
|
||||
| Spec | inlined below |
|
||||
| Blocks | — |
|
||||
|
||||
## Goal
|
||||
|
||||
Compose Bradley-Terry strengths across groups of differing size, as deltas from
|
||||
the anchor present in each group. Only reachable with the tournament enabled and
|
||||
N challengers allowed.
|
||||
|
||||
## Facts (inlined — no spec read needed)
|
||||
|
||||
- Group-normalized relative scores give **cross-task comparability** — but only
|
||||
through a shared anchor, and that qualification is load-bearing.
|
||||
- A BT fit identifies strengths only up to an additive constant **within one
|
||||
connected comparison graph**. Two groups on different tasks are disjoint
|
||||
graphs, so their strengths sit on unlinked scales. Averaging them directly
|
||||
commits the same error that point tallies are accused of, one layer further in.
|
||||
- **The anchor is control.** Control gets a traffic share in every allocation, so
|
||||
every group contains at least one control episode; the fit pins control to zero
|
||||
and every other strength is read as a delta from it. A variant's aggregate is
|
||||
then a mean of like-for-like deltas rather than a mean of incomparable scales.
|
||||
- **A group with no control episode is not aggregatable.** It still grades its own
|
||||
members and is worth reading; it just does not feed the aggregate, and it is
|
||||
**recorded as such** rather than folded in on the assumption that scales match.
|
||||
- Why not point-tally z-scores: small groups produce extreme z-scores, so a
|
||||
variant appearing in many small groups wins on variance rather than quality. BT
|
||||
intervals compose correctly across group sizes and feed the sample gate
|
||||
directly.
|
||||
|
||||
## Steps
|
||||
|
||||
1. Consume `Score::Ranked { strength, interval, group_size }` from T5.9. Strength
|
||||
is already a delta from control because control was pinned at fit time.
|
||||
2. Partition by variant. Weight each group's contribution by its interval width
|
||||
(inverse-variance), not by group size alone.
|
||||
3. **Reject any group lacking a control episode** from the aggregate, and count
|
||||
the exclusions in a reported total — exclusion must be visible, not silent.
|
||||
4. Never combine raw strengths from independently fitted groups without the
|
||||
anchor. Make that impossible in the type, if the type can carry "anchored".
|
||||
5. Emit the aggregate with its own interval and the number of contributing groups.
|
||||
6. Keep it behind the same off-by-default flag as T5.8/T5.9.
|
||||
|
||||
## Acceptance
|
||||
|
||||
- A variant appearing only in small groups **does not outrank** one with a
|
||||
tighter interval at equal mean.
|
||||
- Anchor guard: two groups fit independently with different origins are **not
|
||||
averaged** — a group lacking the anchor is excluded from the aggregate and
|
||||
**counted in a reported total**, so exclusion is visible rather than silent.
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** synthetic `Score::Ranked` inputs with known true strengths and
|
||||
deliberately varied group sizes — no judge needed.
|
||||
|
||||
**Integration test** — `tests/it_variant_aggregation.rs`:
|
||||
1. Variant X appears only in small groups (wide intervals); variant Y appears in
|
||||
large groups (tight intervals). Set their **means equal**. Assert Y does not
|
||||
lose to X — and specifically assert X does **not** outrank Y, which is the
|
||||
variance failure being guarded against.
|
||||
2. Assert weighting is by interval width, not group count: construct a case where
|
||||
the two disagree and assert the interval-weighted answer wins.
|
||||
3. **Anchor guard:** feed two groups fit independently with different origins.
|
||||
Assert they are **not averaged**.
|
||||
4. Assert a group lacking the anchor is **excluded** and that the exclusion is
|
||||
**counted in a reported total** — read the total back and assert it is
|
||||
non-zero. Silent exclusion is the failure.
|
||||
5. Assert the aggregate carries its own interval and the contributing group
|
||||
count.
|
||||
6. Assert this path is unreachable with the tournament flag off.
|
||||
|
||||
**Command:** `cargo test -p loop variant_aggregation`
|
||||
|
||||
**False pass:**
|
||||
- Step 1 with unequal means, where the right variant wins for the wrong reason.
|
||||
Equal means is what isolates the variance effect.
|
||||
- Step 4 asserting only that the anchorless group was dropped. Dropping silently
|
||||
looks identical to having no such group, so a systematic allocation bug that
|
||||
strips control from every group reads as a quiet dataset. The **reported
|
||||
count** is the instrument.
|
||||
- Averaging raw strengths and passing because the synthetic groups happened to
|
||||
share an origin. Construct step 3's groups with deliberately different origins.
|
||||
|
||||
## Traps
|
||||
|
||||
- Averaging strengths across groups without checking for the anchor. It compiles,
|
||||
it produces plausible numbers, and the numbers mean nothing.
|
||||
- Weighting by group size. That is the variance failure BT was chosen to avoid,
|
||||
reintroduced at the aggregation layer.
|
||||
- Dropping anchorless groups silently. Then a systematic allocation bug looks
|
||||
like a quiet dataset.
|
||||
|
||||
---
|
||||
|
||||
Background (not required to do this task):
|
||||
[rust-agentic-sys.md](../../../rust-agentic-sys.md) §9.1, §11.2, §11.4, §12.1 ·
|
||||
[rust-agentic-task.md](../../../rust-agentic-task.md)
|
||||
Reference in New Issue
Block a user