IMPLEMENTATION: - crates/mem-core/src/symptom_projection.rs (250 LOC) - project_symptom(tool, query) → SymptomVector - Three-stage normalization: - Stage 1: Extract keywords - Stage 2: Normalize (stop words, abbreviations) - Stage 3: Generate deterministic SHA256 hash - Tool-specific abbreviation mappings (npm, cargo, kubectl, docker, go) - Stop words list (30+ common words) - Confidence scoring based on keyword specificity TEST COVERAGE: 22 tests passing - 10 unit tests in lib (determinism, abbreviations, stop words, tools, case, order) - 12 integration tests (a1-a6 assertions from design doc) - Real-world scenario tests (npm, cargo, kubectl) - 100% deterministic hashing verified INTEGRATION: - Module exported in crates/mem-core/src/lib.rs - All 43 existing mem-core tests still passing - Ready for M3.7.4 context endpoint integration DESIGN ASSERTIONS (all passing): ✅ a1: Same symptom = same hash (deterministic) ✅ a2: Abbreviation expansion (ERESOLVE → error resolve) ✅ a3: Stop word removal (is, unable, to, the) ✅ a4: Tool consistency (npm ≠ cargo for same error) ✅ a5: Case insensitive (NPM = npm) ✅ a6: Keyword order irrelevant (sorted before hash)
M0.1 - Cargo workspace + crate skeletons - 6-crate workspace with correct dependency direction - CI/CD pipeline with GitHub Actions - Integration tests verifying build and dependency structure M0.2 - Domain types and sha256 identity - Level (L0, L1, L2) enum with proper serde formatting - Role enum (User, Assistant, ToolResult, System) - Record, Chunk, and MemoryNode domain types - Content-hash identity system ensuring rebuild idempotence - Newtypes (ProjectId, QueryId, RunId) with validation - Round-trip serde tests for all types M0.3 - RecordSource trait + ChunkPolicy - RecordSource trait for streaming record sources - Chunk policy with token budgets and boundary modes - TokenCounter trait with CharsOverFourCounter stub - Chunking stream that respects budgets without splitting records - VecSource for testing - Integration tests verifying lossless chunking and budget adherence M0.4 - Tokenizer-backed chunk sizing - Vendored Qwen2 tokenizer with hash verification - QwenTokenCounter implementing proper token counting - Hash guard that fails on modified tokenizer - mem tokens CLI subcommand for token counting - Integration tests with known string counts, hash guards, and budget verification Total: 19 integration tests passing, all phases verified to compose correctly Workspace builds cleanly with no clippy warnings