Files
poimen/tasks/T5.7-order-alternation-and-sampled-consistency.md

4.2 KiB
Raw Permalink Blame History

T5.7 — Order alternation and sampled consistency

Field Value
Phase P5 — Grading
Size S — under 1 day
Status Not started
Flags
Spec inlined below
Blocks

Goal

Cancel judge position bias across the sequence instead of within each pair, and measure order-consistency on a sample rather than by doubling every call.

Facts (inlined — no spec read needed)

  • Pairwise judges have position bias. The tournament path cancels it by judging both orderings, paying 2× on every comparison.
  • The default path does not pay that: the challenger takes position A on even-numbered comparisons and position B on odd ones. Bias cancels across the sequence, at no extra cost.
  • Order-consistency is still measured — on a sampled fraction of comparisons — and reported as the grader's own error bar.
  • That disagreement rate is a first-class grader-health metric. An inconsistent judge should widen the sample gate rather than silently promote.
  • Total calls must stay near 1× the comparison count. A doubling implementation produces the right consistency number and the wrong cost, which is why the call count is asserted rather than the metric.

Steps

  1. Alternate position by comparison index: even → challenger is A, odd → challenger is B. Record the position used on each comparison.
  2. Normalize the verdict back to challenger-relative terms before folding into the WinRecord (T5.6) — otherwise alternation inverts half the record.
  3. Sample a configurable fraction of comparisons for a second, order-swapped call. Default the fraction low; it is an error bar, not the measurement.
  4. Compute the disagreement rate over sampled pairs and emit it as a metric tagged by judge and rubric.
  5. Count total compare calls and assert the ratio to comparison count in the test.
  6. Build a deliberately position-biased mock judge — always picks A — as the detection case.

Acceptance

  • A deliberately position-biased mock judge is detected and its disagreement rate reported.
  • Total compare calls stay within 1.01.2× the comparison count. A doubling implementation fails this, which is the point.

Verify

Harness: a mock judge that always picks position A regardless of content — maximal position bias, so detection is unambiguous.

Integration testtests/it_order_alternation.rs:

  1. Run 100 comparisons against the always-A judge.
  2. Assert the disagreement rate is detected and reported near 100% on the sampled subset.
  3. Assert total compare calls fall in 1.01.2× the comparison count. A doubling implementation lands at 2.0× and fails here, which is the point.
  4. Assert position alternation: challenger in A on even indices, B on odd. Read the recorded position per comparison, do not infer it.
  5. Normalization: with the always-A judge, assert the resulting WinRecord is near 50/50 rather than 100% challenger wins. A missing verdict normalization after the swap inverts half the record and produces exactly 100%, which looks like a strong result.
  6. Unbiased judge control: a content-driven mock judge yields a low disagreement rate — otherwise the detector fires on everything.
  7. Assert the disagreement rate is emitted as a metric tagged by judge and rubric.

Command: cargo test -p grading order_alternation

False pass:

  • Step 3 omitted. Judging both orderings on every comparison produces a perfect consistency measurement at double the cost — correct-looking and exactly the design this task rejects.
  • Step 5 omitted: forgetting to normalize verdicts after alternating is the most likely implementation error, and it makes a biased judge look like a decisive one.
  • Step 6 omitted, so a detector that always reports high bias passes.

Traps

  • Forgetting to normalize the verdict after swapping positions. The record then averages toward 50% regardless of the truth.
  • Sampling "every comparison, it is cheap" — that is the doubling implementation with extra steps.

Background (not required to do this task): rust-agentic-sys.md §11.3, §11.4, §15 · rust-agentic-task.md