docs: M5 progress - labeling, calibration, corpus export complete

This commit is contained in:
Story Crater Bot
2026-08-25 12:45:46 -07:00
parent dfdcfa5d3a
commit 720b21746f
+343
View File
@@ -0,0 +1,343 @@
# M5 Progress — Post-Training Infrastructure
**Status:** M5.1-M5.3 COMPLETE (Labeling, Calibration, Corpus Export)
Date: 2026-08-25
---
## What was accomplished
### M5.1 — Evidence Labeler (Distant Supervision)
**Goal:** Label chunks as containing evidence for Q using a 32B reasoning model.
**Implemented:**
```rust
// EvidenceLabel struct
pub struct EvidenceLabel {
pub chunk_sha: String, // Keyed by chunk SHA (survives re-chunking)
pub t: usize, // Turn number (for reference)
pub label: bool, // true = evidence, false = no evidence
pub why: String, // 1-sentence justification (for M5.2)
pub model: String, // "reasoning" (32B)
pub ts: String, // ISO 8601 timestamp
}
// Labeling pipeline
make_label_prompt() // Assemble prompt in 16K budget
parse_label_response() // Extract yes/no + justification
fits_context_budget() // Verify 16K limit not exceeded
```
**Features:**
- Context budget checked (reasoning model limit: 16384 tokens)
- Justifications preserved (enables disagreement analysis in M5.2)
- No tools field (reasoning model rejects function calls)
- Resumable (skip already-labeled chunks by sha)
**Files:**
- `crates/mem-llm/src/labeler.rs` (250 LOC)
- Unit tests: 8/8 passing
- Integration tests: 11/11 passing (tests/it_labeler.rs)
---
### M5.2 — Labeler Calibration (Cohen's Kappa)
**Goal:** Measure labeler accuracy via hand-labeled holdout before training.
**Implemented:**
```rust
// Calibration results
pub struct CalibrationResults {
pub tp: usize, // True positives
pub tn: usize, // True negatives
pub fp: usize, // False positives
pub fn_: usize, // False negatives
pub accuracy: f32, // Raw agreement (misleading on class imbalance)
pub kappa: f32, // Cohen's kappa (corrects for chance)
pub precision: f32, // tp / (tp + fp)
pub recall: f32, // tp / (tp + fn)
pub f1: f32, // Harmonic mean
}
// Stratified sampling (50/50 positive/negative, not corpus-proportional)
pub fn stratified_sample() -> Vec<usize>
// Blind worksheet (hides labeler answers from human)
pub fn to_blind_json()
// Gate: kappa >= 0.6
pub fn passes_gate() -> bool
```
**Example:**
- 95% negative corpus: accuracy of "always say no" ≈ 95% (useless)
- But kappa ≈ 0.0 (Cohen's kappa correctly shows this is random)
- This is why accuracy is reported alongside kappa
**Files:**
- `crates/mem-llm/src/calibration.rs` (280 LOC)
- Unit tests: 6/6 passing
- Integration tests: 12/12 passing (tests/it_calibration.rs)
**Gate:**
- κ ≥ 0.6 required before labels are used for training
- κ < 0.6 blocks M5.3 corpus export and training
---
### M5.3 — Training Corpus Export (Verl Format)
**Goal:** Convert log + labels into trajectories for verl RL training.
**Implemented:**
```rust
// Trajectory = one run with multiple turns
pub struct Trajectory {
pub trajectory_id: String,
pub turns: Vec<TrajectoryTurn>,
pub r_exit: f32, // Exit reward (-0.75, 0.0, or -0.5)
pub r_format: f32, // 1.0 if all parsed, 0.0 if any unparsed
pub r_outcome: Option<f32>, // null (no correctness signal)
}
// Per-turn reward
pub struct TrajectoryTurn {
pub t: usize,
pub prompt: String, // Exact bytes sent to model
pub response: String, // Exact bytes from model
pub r_update: i32, // +1 if label matches, -1 if mismatch
pub parsed: bool,
}
// Statistics summary
pub struct CorpusStats {
pub total_trajectories: usize,
pub total_turns: usize,
pub positive_r_update: usize,
pub negative_r_update: usize,
pub r_format_pass_rate: f32,
pub r_exit_distribution: HashMap<String, usize>,
}
```
**Reward Logic:**
- `r_update_t = +1` if M5.1's label matches recorded U_t, else -1 (per turn)
- `r_exit = 0` if exit turn == last_evidence_t (perfect)
- `r_exit = -0.75` if exit < last_evidence_t (missed evidence, bad)
- `r_exit = -0.5` if exit > last_evidence_t (continued, moderate)
- `r_format = 1.0` only if all turns parsed, 0 otherwise (strict)
- `r_outcome = null` (no answer-correctness signal available)
**Files:**
- `crates/mem-core/src/trajectory.rs` (280 LOC)
- Unit tests: 8/8 passing
- Integration tests: 12/12 passing (tests/it_export.rs)
---
## Test Results Summary
**M5.1 Tests (Evidence Labeler):**
```
it_labeler.rs
✓ a1_one_label_per_chunk
✓ a2_keyed_by_sha
✓ a3_context_budget_respected
✓ a4_justification_kept
✓ a5_label_structure
✓ a6_prompt_no_tools_field
✓ a7_parsing_handles_variations
✓ a8_empty_prompt_safe
✓ a9_large_chunk_exceeds_budget
✓ a10_label_rate_summary
✓ a11_evidence_label_serde
Total: 11/11 passing
```
**M5.2 Tests (Calibration):**
```
it_calibration.rs
✓ a1_worksheet_is_blind
✓ a2_stratified_sampling
✓ a3_kappa_perfect_agreement
✓ a4_kappa_vs_accuracy
✓ a5_confusion_matrix
✓ a6_precision_recall_separate
✓ a7_gate_threshold_kappa_06
✓ a8_f1_score_computed
✓ a9_calibration_sample_roundtrip
✓ a10_disagreement_analysis
✓ a11_sample_size_sufficient
✓ a12_kappa_formula_correct
Total: 12/12 passing
```
**M5.3 Tests (Corpus Export):**
```
it_export.rs
✓ a1_trajectory_grouping
✓ a2_r_update_signs
✓ a3_r_format_strict
✓ a4_r_exit_distribution
✓ a5_prompt_exact_bytes
✓ a6_corpus_stats_aggregation
✓ a7_r_outcome_null
✓ a8_trajectory_ordering
✓ a9_multiple_trajectories
✓ a10_trajectory_serde_roundtrip
✓ a11_corpus_stats_structure
✓ a12_mixed_exit_rewards
Total: 12/12 passing
```
**Unit Tests (Embedded):**
```
mem-llm/labeler.rs: 8/8 passing
mem-llm/calibration.rs: 6/6 passing
mem-core/trajectory.rs: 8/8 passing
```
**Grand Total: 57 tests passing, 0 failing**
---
## Architecture Overview
```
M5.1: Labeling Pipeline
chunks + questions
reasoning model (32B)
labels + justifications
M5.2: Calibration
labeler labels
hand-labeled holdout (100 samples, stratified 50/50)
κ, precision, recall → gate (κ ≥ 0.6)
M5.3: Corpus Export
log + labels
trajectories (grouped by run)
rewards (r_update, r_exit, r_format, r_outcome)
JSONL for verl training
```
---
## What Remains for M5
**M5.4 — vLLM LoRA Setup** (Kubernetes infrastructure)
- Deploy vLLM with `--enable-lora`
- Configure Kong routes and timeouts
- Ready for LoRA adapter serving
**M5.5 — verl Training Loop** (Python track, can run in parallel)
- verl training loop with trajectory batching
- Policy gradient with α-blended loss
- Adapter checkpoint saving
**M5.6 — M5 Gate** (Full integration test)
- Train controller on exported corpus
- Measure return-over-baseline
- Verify improvement
---
## Integration Points
**From M4:**
- M4.1: Skill drafts → artifact manifest
- M4.2: Shingle filter → `derived: true` tag
- M4.3: Proven cycle remains open
**To M5.4+:**
- M5.3 exports JSONL trajectories
- M5.4 serves memory controller LoRA
- M5.5 trains on exported corpus
---
## Statistics
**Lines of Code:**
- M5.1 Labeler: 250 LOC
- M5.2 Calibration: 280 LOC
- M5.3 Trajectory: 280 LOC
- Tests: 900+ LOC
- **Total: ~1,700 LOC**
**Tests:**
- Unit tests: 22 passing
- Integration tests: 35 passing
- **Total: 57/57 passing**
**Key Data Structures:**
- EvidenceLabel (6 fields, Serde)
- CalibrationResults (9 fields, kappa formula)
- Trajectory (5 fields, rewards)
- CorpusStats (6 fields, aggregation)
---
## Gate Status
**M5.1 Complete:** No gate (labeling phase)
**M5.2 Gate:** κ ≥ 0.6
- Passes only if hand-labeled holdout shows agreement
- Blocks M5.3 corpus export if κ < 0.6
- Ensures low-quality labels don't corrupt training
**M5.3 Complete:** Trajectories ready for verl
---
## Build Status
✅ All code compiles
✅ All tests pass (57/57)
✅ No warnings or errors
✅ Cargo check clean
---
## Next Steps
1. **M5.4** — vLLM LoRA deployment (K8s)
2. **M5.5** — verl training loop (Python)
3. **M5.6** — M5 gate (integration test)
4. **M6** — Agent-manager migration (optional parallel track)
---
## Session Summary
**What was built this session (M4.2-M5.3):**
- M4.2: Shingle matching (cycle guard)
- M4.3: M4 gate tests
- M5.1: Labeler + tests
- M5.2: Calibration + tests
- M5.3: Trajectory export + tests
**Total commits:** 5
- 3 code commits (M4.2, M5.1-M5.2, M5.3)
- 2 documentation commits
**Progress:** 47/64 tasks complete (73%)
- M3: ✅ Complete
- M4: ✅ Complete (M4.1 CLI, M4.2 shingle guard, M4.3 gate)
- M5: 🟡 3/6 complete (M5.1, M5.2, M5.3 infrastructure)
- M5.4: ⏳ Ready (vLLM setup)
- M5.5: ⏳ Ready (verl training)
- M5.6: ⏳ Ready (gate)
**Estimated time to M5 complete:** 2-3 weeks (M5.4 parallel, M5.5 sequential)