6.2 KiB
T5.9 — Bradley-Terry fit
| Field | Value |
|---|---|
| Phase | P5 — Grading |
| Size | L — over 3 days |
| Status | Not started |
| Flags | opt-in |
| Spec | inlined below |
| Blocks | T6.4 |
Goal
Strength parameter plus confidence interval over the pairwise outcomes. Ships disabled — it gates nothing.
Facts (inlined — no spec read needed)
Bradley-Terry, not point-tally z-scores. Accumulating tournament points and normalizing to mean 0 / sd 1 within a group is the obvious approach and it is statistically wrong for the aggregate: small groups produce extreme z-scores, so a variant appearing in many small groups wins on variance rather than quality. A BT fit yields a strength with a real confidence interval, which composes across groups of different sizes.
Three decisions, none optional. Textbook BT does none of them, and each failure looks like a result rather than a bug:
- Davidson tie term. Plain BT is binary and has no tie parameter, so the permitted draws have nowhere to go. Dropping them discards the judge's most confident statements; splitting each half-and-half fabricates two comparisons that never happened and tightens the interval on invented evidence. Fit one additional tie parameter alongside the strengths.
- Weakly-informative prior — required, not tuning. At G = 4..8 an episode that wins every comparison drives the unpenalized maximum-likelihood estimate to infinite strength. The prior (equivalently, a penalized likelihood) is what turns "won all three of its comparisons" into a wide interval rather than an unbounded one. Without it, the reassuring claim that small groups produce wide intervals is simply false.
- Control pinned to zero. The fit is identified only up to an additive constant, so an interval on a single raw strength is an interval on an arbitrary origin. Pin control to zero and report every strength as a delta against it. Downstream gates say "BT interval excludes zero" — zero is control, and it is only zero because it was pinned there.
- A group containing no control episode is not aggregatable. It still grades its own members and is worth reading; it just does not feed the aggregate, and it is recorded as such rather than folded in on the assumption that scales match.
Steps
- Take the comparison list from T5.8:
(episode_a, episode_b, Verdict). - Implement the Davidson model: strengths plus one tie parameter, fit by penalized maximum likelihood.
- Add the weakly-informative prior on strengths. Make its strength a named constant with a comment stating it is required for convergence, not a tuning knob — otherwise it gets removed as "unnecessary regularization".
- Constrain control's strength to zero during the fit rather than subtracting it afterwards.
- Compute intervals (profile likelihood or bootstrap over comparisons) and emit
Score::Ranked { strength, interval, group_size }. - Reject a group with no control episode as non-aggregatable, with a typed error — do not fit against an arbitrary origin.
- Build the synthetic generator: known strengths, configurable draw rate.
Acceptance
- On synthetic data with known strengths, recovered ordering is correct and the delta-from-control interval covers the true delta at the stated rate.
Three targeted cases beyond that:
- A group where one episode wins every comparison converges, with a wide interval. Without the prior this diverges — assert the bound, not just that the fit returns.
- A synthetic generator producing 30% draws recovers the true strengths; a fit that drops draws does not, and the test asserts the gap.
- A group with no control episode is rejected as non-aggregatable rather than fit against an arbitrary origin.
Verify
Harness: a synthetic generator with known true strengths and a configurable draw rate. Every claim here is checkable against ground truth, so none of it needs a real judge.
Integration test — tests/it_bradley_terry.rs:
- Generate comparisons from known strengths. Assert the recovered ordering is correct.
- Coverage test: repeat over 200 seeded datasets; assert the delta-from-control interval covers the true delta at the stated rate (e.g. ~95% for a 95% interval). One dataset proves nothing about an interval.
- Clean sweep: a group where one episode wins every comparison. Assert the fit converges and assert a numeric upper bound on the interval width — not merely that the call returned. Without the prior this diverges.
- Draws at 30%: assert recovered strengths are close to truth. Then run a draw-dropping fit on the same data and assert the gap — the test must show the Davidson term is doing work, not just that the fit runs.
- Half-split control: also fit with draws split half-and-half; assert its interval is narrower than truth warrants (over-confident), demonstrating why splitting is rejected.
- No control: a group with no control episode is rejected as non-aggregatable with a typed error — not fit against an arbitrary origin.
- Assert control's fitted strength is exactly 0 and every other strength is reported as a delta.
Command: cargo test -p grading bradley_terry -- --nocapture
False pass:
- Step 3 asserting only that the fit returns a value. An unpenalized fit returns a very large finite number on many datasets and looks fine. Assert the bound.
- Step 4 without the comparison fit from step 5. "The fit recovers strengths" passes on easy data whether or not draws are modelled — the gap is the evidence.
- Step 2 with one dataset. Interval coverage is a frequency claim; it needs repetition.
- Reporting raw strengths anywhere: an interval on a raw strength is an interval on an arbitrary origin, and it will look perfectly reasonable.
Traps
- Splitting draws half-and-half. It is the standard workaround and it invents evidence.
- Removing the prior after seeing it barely move the estimate on well-separated data. It exists for the clean-sweep case.
- Reporting raw strengths anywhere outside the fit.
Background (not required to do this task): rust-agentic-sys.md §11.2, §11.4 · rust-agentic-task.md