344 lines
8.7 KiB
Markdown
344 lines
8.7 KiB
Markdown
# M5 Progress — Post-Training Infrastructure
|
||
|
||
**Status:** M5.1-M5.3 COMPLETE (Labeling, Calibration, Corpus Export)
|
||
|
||
Date: 2026-08-25
|
||
|
||
---
|
||
|
||
## What was accomplished
|
||
|
||
### M5.1 — Evidence Labeler (Distant Supervision)
|
||
|
||
**Goal:** Label chunks as containing evidence for Q using a 32B reasoning model.
|
||
|
||
**Implemented:**
|
||
```rust
|
||
// EvidenceLabel struct
|
||
pub struct EvidenceLabel {
|
||
pub chunk_sha: String, // Keyed by chunk SHA (survives re-chunking)
|
||
pub t: usize, // Turn number (for reference)
|
||
pub label: bool, // true = evidence, false = no evidence
|
||
pub why: String, // 1-sentence justification (for M5.2)
|
||
pub model: String, // "reasoning" (32B)
|
||
pub ts: String, // ISO 8601 timestamp
|
||
}
|
||
|
||
// Labeling pipeline
|
||
make_label_prompt() // Assemble prompt in 16K budget
|
||
parse_label_response() // Extract yes/no + justification
|
||
fits_context_budget() // Verify 16K limit not exceeded
|
||
```
|
||
|
||
**Features:**
|
||
- Context budget checked (reasoning model limit: 16384 tokens)
|
||
- Justifications preserved (enables disagreement analysis in M5.2)
|
||
- No tools field (reasoning model rejects function calls)
|
||
- Resumable (skip already-labeled chunks by sha)
|
||
|
||
**Files:**
|
||
- `crates/mem-llm/src/labeler.rs` (250 LOC)
|
||
- Unit tests: 8/8 passing
|
||
- Integration tests: 11/11 passing (tests/it_labeler.rs)
|
||
|
||
---
|
||
|
||
### M5.2 — Labeler Calibration (Cohen's Kappa)
|
||
|
||
**Goal:** Measure labeler accuracy via hand-labeled holdout before training.
|
||
|
||
**Implemented:**
|
||
```rust
|
||
// Calibration results
|
||
pub struct CalibrationResults {
|
||
pub tp: usize, // True positives
|
||
pub tn: usize, // True negatives
|
||
pub fp: usize, // False positives
|
||
pub fn_: usize, // False negatives
|
||
pub accuracy: f32, // Raw agreement (misleading on class imbalance)
|
||
pub kappa: f32, // Cohen's kappa (corrects for chance)
|
||
pub precision: f32, // tp / (tp + fp)
|
||
pub recall: f32, // tp / (tp + fn)
|
||
pub f1: f32, // Harmonic mean
|
||
}
|
||
|
||
// Stratified sampling (50/50 positive/negative, not corpus-proportional)
|
||
pub fn stratified_sample() -> Vec<usize>
|
||
|
||
// Blind worksheet (hides labeler answers from human)
|
||
pub fn to_blind_json()
|
||
|
||
// Gate: kappa >= 0.6
|
||
pub fn passes_gate() -> bool
|
||
```
|
||
|
||
**Example:**
|
||
- 95% negative corpus: accuracy of "always say no" ≈ 95% (useless)
|
||
- But kappa ≈ 0.0 (Cohen's kappa correctly shows this is random)
|
||
- This is why accuracy is reported alongside kappa
|
||
|
||
**Files:**
|
||
- `crates/mem-llm/src/calibration.rs` (280 LOC)
|
||
- Unit tests: 6/6 passing
|
||
- Integration tests: 12/12 passing (tests/it_calibration.rs)
|
||
|
||
**Gate:**
|
||
- κ ≥ 0.6 required before labels are used for training
|
||
- κ < 0.6 blocks M5.3 corpus export and training
|
||
|
||
---
|
||
|
||
### M5.3 — Training Corpus Export (Verl Format)
|
||
|
||
**Goal:** Convert log + labels into trajectories for verl RL training.
|
||
|
||
**Implemented:**
|
||
```rust
|
||
// Trajectory = one run with multiple turns
|
||
pub struct Trajectory {
|
||
pub trajectory_id: String,
|
||
pub turns: Vec<TrajectoryTurn>,
|
||
pub r_exit: f32, // Exit reward (-0.75, 0.0, or -0.5)
|
||
pub r_format: f32, // 1.0 if all parsed, 0.0 if any unparsed
|
||
pub r_outcome: Option<f32>, // null (no correctness signal)
|
||
}
|
||
|
||
// Per-turn reward
|
||
pub struct TrajectoryTurn {
|
||
pub t: usize,
|
||
pub prompt: String, // Exact bytes sent to model
|
||
pub response: String, // Exact bytes from model
|
||
pub r_update: i32, // +1 if label matches, -1 if mismatch
|
||
pub parsed: bool,
|
||
}
|
||
|
||
// Statistics summary
|
||
pub struct CorpusStats {
|
||
pub total_trajectories: usize,
|
||
pub total_turns: usize,
|
||
pub positive_r_update: usize,
|
||
pub negative_r_update: usize,
|
||
pub r_format_pass_rate: f32,
|
||
pub r_exit_distribution: HashMap<String, usize>,
|
||
}
|
||
```
|
||
|
||
**Reward Logic:**
|
||
- `r_update_t = +1` if M5.1's label matches recorded U_t, else -1 (per turn)
|
||
- `r_exit = 0` if exit turn == last_evidence_t (perfect)
|
||
- `r_exit = -0.75` if exit < last_evidence_t (missed evidence, bad)
|
||
- `r_exit = -0.5` if exit > last_evidence_t (continued, moderate)
|
||
- `r_format = 1.0` only if all turns parsed, 0 otherwise (strict)
|
||
- `r_outcome = null` (no answer-correctness signal available)
|
||
|
||
**Files:**
|
||
- `crates/mem-core/src/trajectory.rs` (280 LOC)
|
||
- Unit tests: 8/8 passing
|
||
- Integration tests: 12/12 passing (tests/it_export.rs)
|
||
|
||
---
|
||
|
||
## Test Results Summary
|
||
|
||
**M5.1 Tests (Evidence Labeler):**
|
||
```
|
||
it_labeler.rs
|
||
✓ a1_one_label_per_chunk
|
||
✓ a2_keyed_by_sha
|
||
✓ a3_context_budget_respected
|
||
✓ a4_justification_kept
|
||
✓ a5_label_structure
|
||
✓ a6_prompt_no_tools_field
|
||
✓ a7_parsing_handles_variations
|
||
✓ a8_empty_prompt_safe
|
||
✓ a9_large_chunk_exceeds_budget
|
||
✓ a10_label_rate_summary
|
||
✓ a11_evidence_label_serde
|
||
Total: 11/11 passing
|
||
```
|
||
|
||
**M5.2 Tests (Calibration):**
|
||
```
|
||
it_calibration.rs
|
||
✓ a1_worksheet_is_blind
|
||
✓ a2_stratified_sampling
|
||
✓ a3_kappa_perfect_agreement
|
||
✓ a4_kappa_vs_accuracy
|
||
✓ a5_confusion_matrix
|
||
✓ a6_precision_recall_separate
|
||
✓ a7_gate_threshold_kappa_06
|
||
✓ a8_f1_score_computed
|
||
✓ a9_calibration_sample_roundtrip
|
||
✓ a10_disagreement_analysis
|
||
✓ a11_sample_size_sufficient
|
||
✓ a12_kappa_formula_correct
|
||
Total: 12/12 passing
|
||
```
|
||
|
||
**M5.3 Tests (Corpus Export):**
|
||
```
|
||
it_export.rs
|
||
✓ a1_trajectory_grouping
|
||
✓ a2_r_update_signs
|
||
✓ a3_r_format_strict
|
||
✓ a4_r_exit_distribution
|
||
✓ a5_prompt_exact_bytes
|
||
✓ a6_corpus_stats_aggregation
|
||
✓ a7_r_outcome_null
|
||
✓ a8_trajectory_ordering
|
||
✓ a9_multiple_trajectories
|
||
✓ a10_trajectory_serde_roundtrip
|
||
✓ a11_corpus_stats_structure
|
||
✓ a12_mixed_exit_rewards
|
||
Total: 12/12 passing
|
||
```
|
||
|
||
**Unit Tests (Embedded):**
|
||
```
|
||
mem-llm/labeler.rs: 8/8 passing
|
||
mem-llm/calibration.rs: 6/6 passing
|
||
mem-core/trajectory.rs: 8/8 passing
|
||
```
|
||
|
||
**Grand Total: 57 tests passing, 0 failing**
|
||
|
||
---
|
||
|
||
## Architecture Overview
|
||
|
||
```
|
||
M5.1: Labeling Pipeline
|
||
chunks + questions
|
||
↓
|
||
reasoning model (32B)
|
||
↓
|
||
labels + justifications
|
||
|
||
M5.2: Calibration
|
||
labeler labels
|
||
↓
|
||
hand-labeled holdout (100 samples, stratified 50/50)
|
||
↓
|
||
κ, precision, recall → gate (κ ≥ 0.6)
|
||
|
||
M5.3: Corpus Export
|
||
log + labels
|
||
↓
|
||
trajectories (grouped by run)
|
||
↓
|
||
rewards (r_update, r_exit, r_format, r_outcome)
|
||
↓
|
||
JSONL for verl training
|
||
```
|
||
|
||
---
|
||
|
||
## What Remains for M5
|
||
|
||
**M5.4 — vLLM LoRA Setup** (Kubernetes infrastructure)
|
||
- Deploy vLLM with `--enable-lora`
|
||
- Configure Kong routes and timeouts
|
||
- Ready for LoRA adapter serving
|
||
|
||
**M5.5 — verl Training Loop** (Python track, can run in parallel)
|
||
- verl training loop with trajectory batching
|
||
- Policy gradient with α-blended loss
|
||
- Adapter checkpoint saving
|
||
|
||
**M5.6 — M5 Gate** (Full integration test)
|
||
- Train controller on exported corpus
|
||
- Measure return-over-baseline
|
||
- Verify improvement
|
||
|
||
---
|
||
|
||
## Integration Points
|
||
|
||
**From M4:**
|
||
- M4.1: Skill drafts → artifact manifest
|
||
- M4.2: Shingle filter → `derived: true` tag
|
||
- M4.3: Proven cycle remains open
|
||
|
||
**To M5.4+:**
|
||
- M5.3 exports JSONL trajectories
|
||
- M5.4 serves memory controller LoRA
|
||
- M5.5 trains on exported corpus
|
||
|
||
---
|
||
|
||
## Statistics
|
||
|
||
**Lines of Code:**
|
||
- M5.1 Labeler: 250 LOC
|
||
- M5.2 Calibration: 280 LOC
|
||
- M5.3 Trajectory: 280 LOC
|
||
- Tests: 900+ LOC
|
||
- **Total: ~1,700 LOC**
|
||
|
||
**Tests:**
|
||
- Unit tests: 22 passing
|
||
- Integration tests: 35 passing
|
||
- **Total: 57/57 passing**
|
||
|
||
**Key Data Structures:**
|
||
- EvidenceLabel (6 fields, Serde)
|
||
- CalibrationResults (9 fields, kappa formula)
|
||
- Trajectory (5 fields, rewards)
|
||
- CorpusStats (6 fields, aggregation)
|
||
|
||
---
|
||
|
||
## Gate Status
|
||
|
||
**M5.1 Complete:** No gate (labeling phase)
|
||
|
||
**M5.2 Gate:** κ ≥ 0.6
|
||
- Passes only if hand-labeled holdout shows agreement
|
||
- Blocks M5.3 corpus export if κ < 0.6
|
||
- Ensures low-quality labels don't corrupt training
|
||
|
||
**M5.3 Complete:** Trajectories ready for verl
|
||
|
||
---
|
||
|
||
## Build Status
|
||
|
||
✅ All code compiles
|
||
✅ All tests pass (57/57)
|
||
✅ No warnings or errors
|
||
✅ Cargo check clean
|
||
|
||
---
|
||
|
||
## Next Steps
|
||
|
||
1. **M5.4** — vLLM LoRA deployment (K8s)
|
||
2. **M5.5** — verl training loop (Python)
|
||
3. **M5.6** — M5 gate (integration test)
|
||
4. **M6** — Agent-manager migration (optional parallel track)
|
||
|
||
---
|
||
|
||
## Session Summary
|
||
|
||
**What was built this session (M4.2-M5.3):**
|
||
- M4.2: Shingle matching (cycle guard)
|
||
- M4.3: M4 gate tests
|
||
- M5.1: Labeler + tests
|
||
- M5.2: Calibration + tests
|
||
- M5.3: Trajectory export + tests
|
||
|
||
**Total commits:** 5
|
||
- 3 code commits (M4.2, M5.1-M5.2, M5.3)
|
||
- 2 documentation commits
|
||
|
||
**Progress:** 47/64 tasks complete (73%)
|
||
- M3: ✅ Complete
|
||
- M4: ✅ Complete (M4.1 CLI, M4.2 shingle guard, M4.3 gate)
|
||
- M5: 🟡 3/6 complete (M5.1, M5.2, M5.3 infrastructure)
|
||
- M5.4: ⏳ Ready (vLLM setup)
|
||
- M5.5: ⏳ Ready (verl training)
|
||
- M5.6: ⏳ Ready (gate)
|
||
|
||
**Estimated time to M5 complete:** 2-3 weeks (M5.4 parallel, M5.5 sequential)
|