Files
poimen-memory/M5-PROGRESS.md
T

344 lines
8.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M5 Progress — Post-Training Infrastructure
**Status:** M5.1-M5.3 COMPLETE (Labeling, Calibration, Corpus Export)
Date: 2026-08-25
---
## What was accomplished
### M5.1 — Evidence Labeler (Distant Supervision)
**Goal:** Label chunks as containing evidence for Q using a 32B reasoning model.
**Implemented:**
```rust
// EvidenceLabel struct
pub struct EvidenceLabel {
pub chunk_sha: String, // Keyed by chunk SHA (survives re-chunking)
pub t: usize, // Turn number (for reference)
pub label: bool, // true = evidence, false = no evidence
pub why: String, // 1-sentence justification (for M5.2)
pub model: String, // "reasoning" (32B)
pub ts: String, // ISO 8601 timestamp
}
// Labeling pipeline
make_label_prompt() // Assemble prompt in 16K budget
parse_label_response() // Extract yes/no + justification
fits_context_budget() // Verify 16K limit not exceeded
```
**Features:**
- Context budget checked (reasoning model limit: 16384 tokens)
- Justifications preserved (enables disagreement analysis in M5.2)
- No tools field (reasoning model rejects function calls)
- Resumable (skip already-labeled chunks by sha)
**Files:**
- `crates/mem-llm/src/labeler.rs` (250 LOC)
- Unit tests: 8/8 passing
- Integration tests: 11/11 passing (tests/it_labeler.rs)
---
### M5.2 — Labeler Calibration (Cohen's Kappa)
**Goal:** Measure labeler accuracy via hand-labeled holdout before training.
**Implemented:**
```rust
// Calibration results
pub struct CalibrationResults {
pub tp: usize, // True positives
pub tn: usize, // True negatives
pub fp: usize, // False positives
pub fn_: usize, // False negatives
pub accuracy: f32, // Raw agreement (misleading on class imbalance)
pub kappa: f32, // Cohen's kappa (corrects for chance)
pub precision: f32, // tp / (tp + fp)
pub recall: f32, // tp / (tp + fn)
pub f1: f32, // Harmonic mean
}
// Stratified sampling (50/50 positive/negative, not corpus-proportional)
pub fn stratified_sample() -> Vec<usize>
// Blind worksheet (hides labeler answers from human)
pub fn to_blind_json()
// Gate: kappa >= 0.6
pub fn passes_gate() -> bool
```
**Example:**
- 95% negative corpus: accuracy of "always say no" ≈ 95% (useless)
- But kappa ≈ 0.0 (Cohen's kappa correctly shows this is random)
- This is why accuracy is reported alongside kappa
**Files:**
- `crates/mem-llm/src/calibration.rs` (280 LOC)
- Unit tests: 6/6 passing
- Integration tests: 12/12 passing (tests/it_calibration.rs)
**Gate:**
- κ ≥ 0.6 required before labels are used for training
- κ < 0.6 blocks M5.3 corpus export and training
---
### M5.3 — Training Corpus Export (Verl Format)
**Goal:** Convert log + labels into trajectories for verl RL training.
**Implemented:**
```rust
// Trajectory = one run with multiple turns
pub struct Trajectory {
pub trajectory_id: String,
pub turns: Vec<TrajectoryTurn>,
pub r_exit: f32, // Exit reward (-0.75, 0.0, or -0.5)
pub r_format: f32, // 1.0 if all parsed, 0.0 if any unparsed
pub r_outcome: Option<f32>, // null (no correctness signal)
}
// Per-turn reward
pub struct TrajectoryTurn {
pub t: usize,
pub prompt: String, // Exact bytes sent to model
pub response: String, // Exact bytes from model
pub r_update: i32, // +1 if label matches, -1 if mismatch
pub parsed: bool,
}
// Statistics summary
pub struct CorpusStats {
pub total_trajectories: usize,
pub total_turns: usize,
pub positive_r_update: usize,
pub negative_r_update: usize,
pub r_format_pass_rate: f32,
pub r_exit_distribution: HashMap<String, usize>,
}
```
**Reward Logic:**
- `r_update_t = +1` if M5.1's label matches recorded U_t, else -1 (per turn)
- `r_exit = 0` if exit turn == last_evidence_t (perfect)
- `r_exit = -0.75` if exit < last_evidence_t (missed evidence, bad)
- `r_exit = -0.5` if exit > last_evidence_t (continued, moderate)
- `r_format = 1.0` only if all turns parsed, 0 otherwise (strict)
- `r_outcome = null` (no answer-correctness signal available)
**Files:**
- `crates/mem-core/src/trajectory.rs` (280 LOC)
- Unit tests: 8/8 passing
- Integration tests: 12/12 passing (tests/it_export.rs)
---
## Test Results Summary
**M5.1 Tests (Evidence Labeler):**
```
it_labeler.rs
✓ a1_one_label_per_chunk
✓ a2_keyed_by_sha
✓ a3_context_budget_respected
✓ a4_justification_kept
✓ a5_label_structure
✓ a6_prompt_no_tools_field
✓ a7_parsing_handles_variations
✓ a8_empty_prompt_safe
✓ a9_large_chunk_exceeds_budget
✓ a10_label_rate_summary
✓ a11_evidence_label_serde
Total: 11/11 passing
```
**M5.2 Tests (Calibration):**
```
it_calibration.rs
✓ a1_worksheet_is_blind
✓ a2_stratified_sampling
✓ a3_kappa_perfect_agreement
✓ a4_kappa_vs_accuracy
✓ a5_confusion_matrix
✓ a6_precision_recall_separate
✓ a7_gate_threshold_kappa_06
✓ a8_f1_score_computed
✓ a9_calibration_sample_roundtrip
✓ a10_disagreement_analysis
✓ a11_sample_size_sufficient
✓ a12_kappa_formula_correct
Total: 12/12 passing
```
**M5.3 Tests (Corpus Export):**
```
it_export.rs
✓ a1_trajectory_grouping
✓ a2_r_update_signs
✓ a3_r_format_strict
✓ a4_r_exit_distribution
✓ a5_prompt_exact_bytes
✓ a6_corpus_stats_aggregation
✓ a7_r_outcome_null
✓ a8_trajectory_ordering
✓ a9_multiple_trajectories
✓ a10_trajectory_serde_roundtrip
✓ a11_corpus_stats_structure
✓ a12_mixed_exit_rewards
Total: 12/12 passing
```
**Unit Tests (Embedded):**
```
mem-llm/labeler.rs: 8/8 passing
mem-llm/calibration.rs: 6/6 passing
mem-core/trajectory.rs: 8/8 passing
```
**Grand Total: 57 tests passing, 0 failing**
---
## Architecture Overview
```
M5.1: Labeling Pipeline
chunks + questions
reasoning model (32B)
labels + justifications
M5.2: Calibration
labeler labels
hand-labeled holdout (100 samples, stratified 50/50)
κ, precision, recall → gate (κ ≥ 0.6)
M5.3: Corpus Export
log + labels
trajectories (grouped by run)
rewards (r_update, r_exit, r_format, r_outcome)
JSONL for verl training
```
---
## What Remains for M5
**M5.4 — vLLM LoRA Setup** (Kubernetes infrastructure)
- Deploy vLLM with `--enable-lora`
- Configure Kong routes and timeouts
- Ready for LoRA adapter serving
**M5.5 — verl Training Loop** (Python track, can run in parallel)
- verl training loop with trajectory batching
- Policy gradient with α-blended loss
- Adapter checkpoint saving
**M5.6 — M5 Gate** (Full integration test)
- Train controller on exported corpus
- Measure return-over-baseline
- Verify improvement
---
## Integration Points
**From M4:**
- M4.1: Skill drafts → artifact manifest
- M4.2: Shingle filter → `derived: true` tag
- M4.3: Proven cycle remains open
**To M5.4+:**
- M5.3 exports JSONL trajectories
- M5.4 serves memory controller LoRA
- M5.5 trains on exported corpus
---
## Statistics
**Lines of Code:**
- M5.1 Labeler: 250 LOC
- M5.2 Calibration: 280 LOC
- M5.3 Trajectory: 280 LOC
- Tests: 900+ LOC
- **Total: ~1,700 LOC**
**Tests:**
- Unit tests: 22 passing
- Integration tests: 35 passing
- **Total: 57/57 passing**
**Key Data Structures:**
- EvidenceLabel (6 fields, Serde)
- CalibrationResults (9 fields, kappa formula)
- Trajectory (5 fields, rewards)
- CorpusStats (6 fields, aggregation)
---
## Gate Status
**M5.1 Complete:** No gate (labeling phase)
**M5.2 Gate:** κ ≥ 0.6
- Passes only if hand-labeled holdout shows agreement
- Blocks M5.3 corpus export if κ < 0.6
- Ensures low-quality labels don't corrupt training
**M5.3 Complete:** Trajectories ready for verl
---
## Build Status
✅ All code compiles
✅ All tests pass (57/57)
✅ No warnings or errors
✅ Cargo check clean
---
## Next Steps
1. **M5.4** — vLLM LoRA deployment (K8s)
2. **M5.5** — verl training loop (Python)
3. **M5.6** — M5 gate (integration test)
4. **M6** — Agent-manager migration (optional parallel track)
---
## Session Summary
**What was built this session (M4.2-M5.3):**
- M4.2: Shingle matching (cycle guard)
- M4.3: M4 gate tests
- M5.1: Labeler + tests
- M5.2: Calibration + tests
- M5.3: Trajectory export + tests
**Total commits:** 5
- 3 code commits (M4.2, M5.1-M5.2, M5.3)
- 2 documentation commits
**Progress:** 47/64 tasks complete (73%)
- M3: ✅ Complete
- M4: ✅ Complete (M4.1 CLI, M4.2 shingle guard, M4.3 gate)
- M5: 🟡 3/6 complete (M5.1, M5.2, M5.3 infrastructure)
- M5.4: ⏳ Ready (vLLM setup)
- M5.5: ⏳ Ready (verl training)
- M5.6: ⏳ Ready (gate)
**Estimated time to M5 complete:** 2-3 weeks (M5.4 parallel, M5.5 sequential)