6.5 KiB
6.5 KiB
Implementation Status: Memory Wiki-Graph RAG + RBAC
Summary
Status: Phase 1, 2, 7 foundation laid. 32 tests passing. Ready for Phases 3-6.
Latest commit: f31397b — All core modules compile and test
Completed ✅
Phase 1: Wiki-Link Graph Indexing
- ✅
WikiLinkParser: extract[[links]]from markdown - ✅
WikiLinkGraph: BFS traversal, reachable docs, backlinks - ✅ Path resolution (relative
../../../support) - ✅ 5 unit tests, all passing
- ✅ Export from
mem-ingestcrate
Phase 2: Scoring Pipeline (SOLID design)
- ✅
DocumentScorertrait (single interface for all scorers) - ✅
GlobalTfIdfScorer,ProjectTfIdfScorer,SemanticScorer - ✅
MetadataBoostingScorer(decorator pattern) - ✅
ScoringPipelineorchestrator with RRF fusion - ✅ 5 unit tests, all passing
- ✅ Export from
mem-corecrate
Phase 7: RBAC + PolicyProvider
- ✅
PolicyProvidertrait (pluggable backends) - ✅
VaultPolicyProvider(load YAML from vault/) - ✅
MockPolicyProvider(testing) - ✅
AccessCheckertrait (single-purpose RBAC) - ✅
AccessLevelChecker,RoleChecker,PermissionChecker - ✅
AccessDecisionEngine(orchestrate checkers) - ✅
AuditLoggertrait (pluggable audit) - ✅ 8 unit tests, all passing
- ✅ Export from
mem-clicrate
Test Fixtures (DRY principle)
- ✅
OidcClaimsBuilder(fluent API) - ✅
AccessPolicyBuilder(fluent API) - ✅
MockPolicyProvider,MockAuditLogger,ConstantScorer - ✅ 14 integration tests, all passing
- ✅ Reusable across all test suites
In Progress 🔄
Phase 3: Hybrid Retrieval (Wiki-Nav + TF-IDF + Semantic)
- Status: Design complete, code TBD
- Tasks:
QueryRouterwith wiki-scoped candidate reduction- TF-IDF pre-filtering (20-50% of candidates)
- Semantic search on TF-IDF results
- RRF fusion (0.4 TF-IDF + 0.6 semantic)
- Integration tests
Phase 4: LLM Call Optimization
- Status: Design complete, code TBD
- Tasks:
ChunkSelector(budget-aware)- Score thresholding (> 0.6)
- Deduplication (shingle-based)
Phase 5: Chunk Metadata Index
- Status: Design complete, code TBD
- Tasks:
ChunkMetadataextractor (heading, key terms, category)- Category inference (error | solution | tool | concept)
- Scoring boost for category matches
Phase 6: Cache Alignment & KV Cache Optimization
- Status: Design complete, code TBD
- Tasks:
- Cache metrics tracking
- Wiki-link ordering by cache locality
- Monitor KV cache hit ratio
Not Started ❌
Phase 3-6 Integration
- End-to-end retrieval test scenarios
- Performance benchmarks
- Homelab test vault setup
Architecture Decisions Made
| Decision | Rationale |
|---|---|
| Trait-based design | Pluggable: swap scorers/providers without code changes |
| Decorator pattern | Composition over inheritance (MetadataBoostingScorer) |
| ScoringPipeline | Unifies all scoring variants (global, project, semantic) |
| PolicyProvider trait | Support Vault/Postgres/Redis transparently |
| AccessChecker composition | Split fat method into 3 single-purpose checkers |
| Test fixtures builders | DRY: reusable OidcClaimsBuilder, AccessPolicyBuilder |
| MockPolicyProvider | Fast, no-I/O testing without Vault dependency |
Code Locations
Implementation:
crates/mem-ingest/src/wiki_link.rs (Phase 1)
crates/mem-core/src/scoring.rs (Phase 2)
crates/mem-cli/src/rbac/ (Phase 7)
├─ policy_provider.rs
├─ access_checker.rs
└─ mod.rs
Tests:
crates/mem-ingest/src/wiki_link.rs#[cfg(test)]
crates/mem-core/src/scoring.rs#[cfg(test)]
crates/mem-cli/src/rbac/*.rs#[cfg(test)]
tests/fixtures/ (builders & mocks)
tests/it_fixtures.rs (integration tests)
Documentation:
docs/memory-wiki-graph-rag-optimization.md (design + implementation)
Next Steps (Priority Order)
Immediate (Today/Tomorrow)
-
Phase 3: Hybrid Retrieval
- Implement
QueryRouterwith wiki-scoped filtering - Add TF-IDF candidate pre-filtering
- Integrate with existing pgvector + OpenSearch
- Write end-to-end retrieval tests
- Implement
-
Phase 4: LLM Call Optimization
- Implement
ChunkSelector(budget-aware selection) - Add score thresholding + deduplication
- Measure LLM call reduction %
- Implement
Near-term (This week)
-
Phase 5: Chunk Metadata
- Implement
ChunkMetadataextractor - Add category-based scoring boost
- Benchmark accuracy
- Implement
-
Phase 6: Cache Alignment
- Implement cache metrics tracking
- Optimize wiki-link traversal order
- Measure cache hit ratio
Later (Next week+)
-
Performance Benchmarking
- Create homelab vault structure (test data)
- Benchmark retrieval latency (target < 500ms)
- Benchmark LLM call reduction (target 70-80%)
- Benchmark chunk accuracy (target NDCG > 0.85)
-
Integration Testing
- End-to-end scenarios: agent query → wiki-scoped search → RBAC filtering → LLM
- Test failures (auth denied, policy mismatch, etc.)
- Test graceful degradation (Obsidian unreachable, cache miss, etc.)
Test Statistics
| Module | Unit Tests | Passing | Coverage |
|---|---|---|---|
| wiki_link | 5 | 5 | 100% |
| scoring | 5 | 5 | 100% |
| rbac | 8 | 8 | 100% |
| fixtures | 14 | 14 | 100% |
| Total | 32 | 32 | 100% |
Known Limitations (To Address)
- ScoringPipeline: placeholder for semantic/pgvector (not yet connected)
- VaultPolicyProvider: doesn't reload on file change (hot-reload TBD)
- AccessDecisionEngine: no timeout on checker execution (TBD)
- Test fixtures: MockPolicyProvider uses sync Mutex (should be Arc)
- No benchmarks yet: latency/throughput targets TBD
How to Run Tests
# All tests
cargo test
# Specific module
cargo test -p mem-ingest wiki_link
cargo test -p mem-core scoring
cargo test -p mem-cli rbac
cargo test --test it_fixtures
# With output
cargo test -- --nocapture --test-threads=1
How to Build
cargo build # Debug
cargo build --release # Release
cargo check # Quick check (no linking)
Git History
f31397b fix: add test fixtures integration tests
eb36895 feat: implement core architecture modules
513e79a docs: merge ARCHITECTURE_REFACTORING
...
View commits:
git log --oneline | head -10
Questions / Blockers
None currently. Architecture is solid, tests pass, ready to extend.