Adds normalized shingle matching to prevent feedback loops where emitted skills are re-ingested as evidence: Files created: crates/mem-core/src/shingle.rs (250 LOC) - Shingle: normalized n-gram wrapper - ShingleConfig: configurable threshold (default 0.80) and size (default 4) - normalize(): removes markdown, code fences, collapses whitespace - get_shingles(): overlapping token n-grams - jaccard_similarity(): Jaccard index for text comparison - matches_artifact(): detect if record matches any artifact above threshold tests/it_derived_filter.rs (11 tests, all passing) - a1: Verbatim artifact copies detected - a2: Reformatted copies (whitespace/markdown) detected - a3: Mere mentions of skill names NOT excluded (false positive guard) - a4: Unrelated text NOT excluded - a5: Multiple artifacts handled correctly - a6: Threshold configurable - a7: Similarity score returned - a8: No artifacts is safe (empty list) - a9: Empty text is safe - a10: Case-insensitive matching - a11: Partial coverage detection Files modified: crates/mem-core/src/lib.rs - Add shingle module - Export ShingleConfig, jaccard_similarity, matches_artifact Architecture: During ingest: compare record against vault/.artifacts.jsonl If overlap >= threshold: tag derived=true, exclude from evidence Log exclusion event for auditability Threshold tuning: - 0.80: strict, catches verbatim + reformatted - 0.70: moderate, catches variants - 0.60: permissive, catches substantial overlap Default 0.80 prevents false positives (mentioning skill != using skill text) Tests: ✓ 11/11 passing ✓ Unit tests in shingle module: 11/11 passing ✓ Integration tests: 11/11 passing Blocks: M4.3 gate (needs ingest integration) Depends: M4.1 ✓ (skill draft)
M0.1 - Cargo workspace + crate skeletons - 6-crate workspace with correct dependency direction - CI/CD pipeline with GitHub Actions - Integration tests verifying build and dependency structure M0.2 - Domain types and sha256 identity - Level (L0, L1, L2) enum with proper serde formatting - Role enum (User, Assistant, ToolResult, System) - Record, Chunk, and MemoryNode domain types - Content-hash identity system ensuring rebuild idempotence - Newtypes (ProjectId, QueryId, RunId) with validation - Round-trip serde tests for all types M0.3 - RecordSource trait + ChunkPolicy - RecordSource trait for streaming record sources - Chunk policy with token budgets and boundary modes - TokenCounter trait with CharsOverFourCounter stub - Chunking stream that respects budgets without splitting records - VecSource for testing - Integration tests verifying lossless chunking and budget adherence M0.4 - Tokenizer-backed chunk sizing - Vendored Qwen2 tokenizer with hash verification - QwenTokenCounter implementing proper token counting - Hash guard that fails on modified tokenizer - mem tokens CLI subcommand for token counting - Integration tests with known string counts, hash guards, and budget verification Total: 19 integration tests passing, all phases verified to compose correctly Workspace builds cleanly with no clippy warnings