Files
poimen-memory/M5-PROGRESS.md
T

8.7 KiB
Raw Blame History

M5 Progress — Post-Training Infrastructure

Status: M5.1-M5.3 COMPLETE (Labeling, Calibration, Corpus Export)

Date: 2026-08-25


What was accomplished

M5.1 — Evidence Labeler (Distant Supervision)

Goal: Label chunks as containing evidence for Q using a 32B reasoning model.

Implemented:

// EvidenceLabel struct
pub struct EvidenceLabel {
    pub chunk_sha: String,      // Keyed by chunk SHA (survives re-chunking)
    pub t: usize,               // Turn number (for reference)
    pub label: bool,            // true = evidence, false = no evidence
    pub why: String,            // 1-sentence justification (for M5.2)
    pub model: String,          // "reasoning" (32B)
    pub ts: String,             // ISO 8601 timestamp
}

// Labeling pipeline
make_label_prompt()          // Assemble prompt in 16K budget
parse_label_response()       // Extract yes/no + justification
fits_context_budget()        // Verify 16K limit not exceeded

Features:

  • Context budget checked (reasoning model limit: 16384 tokens)
  • Justifications preserved (enables disagreement analysis in M5.2)
  • No tools field (reasoning model rejects function calls)
  • Resumable (skip already-labeled chunks by sha)

Files:

  • crates/mem-llm/src/labeler.rs (250 LOC)
  • Unit tests: 8/8 passing
  • Integration tests: 11/11 passing (tests/it_labeler.rs)

M5.2 — Labeler Calibration (Cohen's Kappa)

Goal: Measure labeler accuracy via hand-labeled holdout before training.

Implemented:

// Calibration results
pub struct CalibrationResults {
    pub tp: usize,           // True positives
    pub tn: usize,           // True negatives
    pub fp: usize,           // False positives
    pub fn_: usize,          // False negatives
    pub accuracy: f32,       // Raw agreement (misleading on class imbalance)
    pub kappa: f32,          // Cohen's kappa (corrects for chance)
    pub precision: f32,      // tp / (tp + fp)
    pub recall: f32,         // tp / (tp + fn)
    pub f1: f32,             // Harmonic mean
}

// Stratified sampling (50/50 positive/negative, not corpus-proportional)
pub fn stratified_sample() -> Vec<usize>

// Blind worksheet (hides labeler answers from human)
pub fn to_blind_json()

// Gate: kappa >= 0.6
pub fn passes_gate() -> bool

Example:

  • 95% negative corpus: accuracy of "always say no" ≈ 95% (useless)
  • But kappa ≈ 0.0 (Cohen's kappa correctly shows this is random)
  • This is why accuracy is reported alongside kappa

Files:

  • crates/mem-llm/src/calibration.rs (280 LOC)
  • Unit tests: 6/6 passing
  • Integration tests: 12/12 passing (tests/it_calibration.rs)

Gate:

  • κ ≥ 0.6 required before labels are used for training
  • κ < 0.6 blocks M5.3 corpus export and training

M5.3 — Training Corpus Export (Verl Format)

Goal: Convert log + labels into trajectories for verl RL training.

Implemented:

// Trajectory = one run with multiple turns
pub struct Trajectory {
    pub trajectory_id: String,
    pub turns: Vec<TrajectoryTurn>,
    pub r_exit: f32,         // Exit reward (-0.75, 0.0, or -0.5)
    pub r_format: f32,       // 1.0 if all parsed, 0.0 if any unparsed
    pub r_outcome: Option<f32>, // null (no correctness signal)
}

// Per-turn reward
pub struct TrajectoryTurn {
    pub t: usize,
    pub prompt: String,       // Exact bytes sent to model
    pub response: String,     // Exact bytes from model
    pub r_update: i32,        // +1 if label matches, -1 if mismatch
    pub parsed: bool,
}

// Statistics summary
pub struct CorpusStats {
    pub total_trajectories: usize,
    pub total_turns: usize,
    pub positive_r_update: usize,
    pub negative_r_update: usize,
    pub r_format_pass_rate: f32,
    pub r_exit_distribution: HashMap<String, usize>,
}

Reward Logic:

  • r_update_t = +1 if M5.1's label matches recorded U_t, else -1 (per turn)
  • r_exit = 0 if exit turn == last_evidence_t (perfect)
  • r_exit = -0.75 if exit < last_evidence_t (missed evidence, bad)
  • r_exit = -0.5 if exit > last_evidence_t (continued, moderate)
  • r_format = 1.0 only if all turns parsed, 0 otherwise (strict)
  • r_outcome = null (no answer-correctness signal available)

Files:

  • crates/mem-core/src/trajectory.rs (280 LOC)
  • Unit tests: 8/8 passing
  • Integration tests: 12/12 passing (tests/it_export.rs)

Test Results Summary

M5.1 Tests (Evidence Labeler):

it_labeler.rs
  ✓ a1_one_label_per_chunk
  ✓ a2_keyed_by_sha
  ✓ a3_context_budget_respected
  ✓ a4_justification_kept
  ✓ a5_label_structure
  ✓ a6_prompt_no_tools_field
  ✓ a7_parsing_handles_variations
  ✓ a8_empty_prompt_safe
  ✓ a9_large_chunk_exceeds_budget
  ✓ a10_label_rate_summary
  ✓ a11_evidence_label_serde
Total: 11/11 passing

M5.2 Tests (Calibration):

it_calibration.rs
  ✓ a1_worksheet_is_blind
  ✓ a2_stratified_sampling
  ✓ a3_kappa_perfect_agreement
  ✓ a4_kappa_vs_accuracy
  ✓ a5_confusion_matrix
  ✓ a6_precision_recall_separate
  ✓ a7_gate_threshold_kappa_06
  ✓ a8_f1_score_computed
  ✓ a9_calibration_sample_roundtrip
  ✓ a10_disagreement_analysis
  ✓ a11_sample_size_sufficient
  ✓ a12_kappa_formula_correct
Total: 12/12 passing

M5.3 Tests (Corpus Export):

it_export.rs
  ✓ a1_trajectory_grouping
  ✓ a2_r_update_signs
  ✓ a3_r_format_strict
  ✓ a4_r_exit_distribution
  ✓ a5_prompt_exact_bytes
  ✓ a6_corpus_stats_aggregation
  ✓ a7_r_outcome_null
  ✓ a8_trajectory_ordering
  ✓ a9_multiple_trajectories
  ✓ a10_trajectory_serde_roundtrip
  ✓ a11_corpus_stats_structure
  ✓ a12_mixed_exit_rewards
Total: 12/12 passing

Unit Tests (Embedded):

mem-llm/labeler.rs:       8/8 passing
mem-llm/calibration.rs:   6/6 passing
mem-core/trajectory.rs:   8/8 passing

Grand Total: 57 tests passing, 0 failing


Architecture Overview

M5.1: Labeling Pipeline
  chunks + questions
    ↓
  reasoning model (32B)
    ↓
  labels + justifications

M5.2: Calibration
  labeler labels
    ↓
  hand-labeled holdout (100 samples, stratified 50/50)
    ↓
  κ, precision, recall → gate (κ ≥ 0.6)

M5.3: Corpus Export
  log + labels
    ↓
  trajectories (grouped by run)
    ↓
  rewards (r_update, r_exit, r_format, r_outcome)
    ↓
  JSONL for verl training

What Remains for M5

M5.4 — vLLM LoRA Setup (Kubernetes infrastructure)

  • Deploy vLLM with --enable-lora
  • Configure Kong routes and timeouts
  • Ready for LoRA adapter serving

M5.5 — verl Training Loop (Python track, can run in parallel)

  • verl training loop with trajectory batching
  • Policy gradient with α-blended loss
  • Adapter checkpoint saving

M5.6 — M5 Gate (Full integration test)

  • Train controller on exported corpus
  • Measure return-over-baseline
  • Verify improvement

Integration Points

From M4:

  • M4.1: Skill drafts → artifact manifest
  • M4.2: Shingle filter → derived: true tag
  • M4.3: Proven cycle remains open

To M5.4+:

  • M5.3 exports JSONL trajectories
  • M5.4 serves memory controller LoRA
  • M5.5 trains on exported corpus

Statistics

Lines of Code:

  • M5.1 Labeler: 250 LOC
  • M5.2 Calibration: 280 LOC
  • M5.3 Trajectory: 280 LOC
  • Tests: 900+ LOC
  • Total: ~1,700 LOC

Tests:

  • Unit tests: 22 passing
  • Integration tests: 35 passing
  • Total: 57/57 passing

Key Data Structures:

  • EvidenceLabel (6 fields, Serde)
  • CalibrationResults (9 fields, kappa formula)
  • Trajectory (5 fields, rewards)
  • CorpusStats (6 fields, aggregation)

Gate Status

M5.1 Complete: No gate (labeling phase)

M5.2 Gate: κ ≥ 0.6

  • Passes only if hand-labeled holdout shows agreement
  • Blocks M5.3 corpus export if κ < 0.6
  • Ensures low-quality labels don't corrupt training

M5.3 Complete: Trajectories ready for verl


Build Status

All code compiles All tests pass (57/57) No warnings or errors Cargo check clean


Next Steps

  1. M5.4 — vLLM LoRA deployment (K8s)
  2. M5.5 — verl training loop (Python)
  3. M5.6 — M5 gate (integration test)
  4. M6 — Agent-manager migration (optional parallel track)

Session Summary

What was built this session (M4.2-M5.3):

  • M4.2: Shingle matching (cycle guard)
  • M4.3: M4 gate tests
  • M5.1: Labeler + tests
  • M5.2: Calibration + tests
  • M5.3: Trajectory export + tests

Total commits: 5

  • 3 code commits (M4.2, M5.1-M5.2, M5.3)
  • 2 documentation commits

Progress: 47/64 tasks complete (73%)

  • M3: Complete
  • M4: Complete (M4.1 CLI, M4.2 shingle guard, M4.3 gate)
  • M5: 🟡 3/6 complete (M5.1, M5.2, M5.3 infrastructure)
  • M5.4: Ready (vLLM setup)
  • M5.5: Ready (verl training)
  • M5.6: Ready (gate)

Estimated time to M5 complete: 2-3 weeks (M5.4 parallel, M5.5 sequential)