Deleted 31 completed task files: - M0.x: 8 tasks (cargo, domain types, recordsource, tokenizer, adapters, gate) - M1.x: 8 tasks (llm-chat, standing-query, prompt template, parser, loop, log, e2e, gate) - M3.x: 4 tasks (l2-synthesis, rerank, mem-query, gate) - M3.5.x: 8 tasks (http-server, ingest, query, federation, skills, projects, rate-limiting, gate) - M3.6.1: DocCorpusSource (heading-boundary chunking) - M4.1-2: skill-draft, derived-filter Updated INDEX.md: - Removed M0 & M1 phase sections (archived in git history) - Updated progress table: 65 active tasks (42✅ + 2🟡 + 21⬜) - Updated status: M0/M1 complete, M3/M3.5 gates passing, M4.1-2 done - Noted M3.5.10 JWT auth implementation complete (awaiting image rollout) - Cleaned up broken links to deleted task files Total test count: 239 passing, 2 ignored (up from 196 at M3.4) Ready for M4.3 gate composition, M5 post-training, M7 source connectors.
12 KiB
12 KiB
API Review: Query APIs vs Retrieval APIs
Current API Endpoints (9 routes)
Overview Table
| # | Endpoint | Method | Category | Auth | Rate Limit | Purpose |
|---|---|---|---|---|---|---|
| 1 | /health |
GET | System | ❌ No | ❌ No | Health check |
| 2 | /memory/ingest |
POST | Write | ✅ (write) | 100/hr | Queue ingest job |
| 3 | /memory/ingest/{id} |
GET | Status | ✅ (read) | ✅ Query | Check ingest progress |
| 4 | /memory/query |
GET | QUERY API ⭐ | ✅ (read) | 1000/hr | Semantic search |
| 5 | /memory/projects |
GET | Metadata | ✅ (read) | 100/hr | List projects |
| 6 | /memory/skills |
GET | Metadata | ✅ (read) | ✅ Query | List skills |
| 7 | /memory/vault/generate |
POST | Write | ✅ (write) | 100/hr | Generate vault |
| 8 | /memory/vault |
GET | RETRIEVAL API ⭐ | ✅ (read) | ✅ Query | Browse vault root |
| 9 | /memory/vault/{project} |
GET | RETRIEVAL API ⭐ | ✅ (read) | ✅ Query | Browse project |
| 10 | /memory/vault/{project}/{file} |
GET | RETRIEVAL API ⭐ | ✅ (read) | ✅ Query | Read file |
Distinction: Query APIs vs Retrieval APIs
Query APIs (Search/Semantic)
🔍 /memory/query — Semantic Search (PRIMARY)
GET /memory/query?query=<text>&project=<proj>&limit=<k>
Authorization: Bearer <JWT>
Query Parameters:
- query (required): Search string (e.g., "fix kubernetes port 8080")
- project (required): Filter by project
- limit (optional): Top-K results (default: 5)
Returns:
{
"query": "fix kubernetes port 8080",
"project": "poimen",
"results": [
{
"id": "chunk-123",
"text": "kubectl port-forward...",
"source": "runbooks/port-forward.md",
"level": "L1",
"breadcrumb": ["runbooks", "kubernetes"],
"score": 0.92,
"sem_component": 0.88,
"lex_component": 0.96 // [NEW] when hybrid enabled
},
...
]
}
Current Implementation:
- ✅ JWT auth enforced
- ✅ Rate limited: 1000/hr
- ✅ Capability check:
memory:read - ⏳ MISSING: Hybrid search integration
- ⏳ MISSING: Query routing decision tree
- ⏳ MISSING: OpenSearch lexical search
Flow:
GET /memory/query
│
├─ Validate JWT / ApiKey
├─ Check capability: memory:read
├─ Check rate limit: 1000/hr per user
│
├─ (Current) QueryWorker.query()
│ └─ Only semantic via pgvector
│
└─ (NEEDED) HybridQueryWorker.query()
├─ Parallel pgvector search (top-50)
├─ Parallel OpenSearch BM25 (top-50) + JWT forward
├─ Normalize scores to [0.0, 1.0]
├─ Fusion: 0.6*sem + 0.4*lex
└─ Return top-10 results with breakdown
Retrieval APIs (Browse/Read)
📂 /memory/vault — List Vault Root
GET /memory/vault
Authorization: Bearer <JWT>
Returns:
{
"projects": [
{
"name": "poimen",
"path": "vault/poimen",
"file_count": 42,
"updated_at": "2024-01-15T10:30:00Z"
},
...
]
}
Purpose: Browse vault directory structure (top-level)
📂 /memory/vault/{project} — List Project Files
GET /memory/vault/poimen
Authorization: Bearer <JWT>
Returns:
{
"project": "poimen",
"files": [
{
"name": "runbooks",
"type": "directory",
"path": "vault/poimen/runbooks",
"file_count": 15
},
{
"name": "deployment.md",
"type": "file",
"path": "vault/poimen/deployment.md",
"size_bytes": 4096,
"updated_at": "2024-01-15T10:30:00Z"
},
...
]
}
Purpose: Browse files in a project (file tree view)
📄 /memory/vault/{project}/{file} — Read File Content
GET /memory/vault/poimen/deployment.md
Authorization: Bearer <JWT>
Returns:
{
"project": "poimen",
"file": "deployment.md",
"path": "vault/poimen/deployment.md",
"content": "# Deployment Guide\n\n## Overview\n...",
"size_bytes": 4096,
"updated_at": "2024-01-15T10:30:00Z",
"sections": [
{
"title": "Overview",
"level": 1,
"content": "..."
},
...
]
}
Purpose: Read/view full file content (no AI processing)
Current Issues & Gaps
❌ Issue 1: /memory/query is Semantic-Only
Problem:
- Current
/memory/querycallsQueryWorker.query()which ONLY does pgvector - No access to OpenSearch (lexical search)
- No fusion/reranking logic
- No score breakdown for debugging
Solution:
// Instead of:
state.query_worker.query(&project, &question, Some(limit)).await
// Should be:
state.hybrid_query_worker.hybrid_query(
&project,
&question,
limit,
SearchMethod::Hybrid, // or Semantic, Lexical
HybridWeights { semantic: 0.6, lexical: 0.4 },
claims,
&token // Forward JWT to OpenSearch
).await
❌ Issue 2: No Query Routing
Problem:
- All queries use hybrid (once implemented)
- No decision tree for:
- Short queries (< 3 tokens) → lexical better
- Special syntax (#tag) → lexical better
- LLM unavailable → fallback to lexical
Solution:
pub fn route_query(query: &str, method_override: Option<SearchMethod>) -> SearchMethod {
// If user explicitly requested a method, use it
if let Some(method) = method_override {
return method;
}
// Otherwise, auto-route based on query characteristics
let token_count = query.split_whitespace().count();
if token_count < 3 {
SearchMethod::Lexical // Short queries
} else if query.contains('#') || query.contains('@') {
SearchMethod::Lexical // Special syntax
} else {
SearchMethod::Hybrid // Normal queries
}
}
❌ Issue 3: Retrieval APIs Don't Integrate with Search
Problem:
/memory/vault/*only reads files from PVC- No integration with indexed chunks
- User can't "click through" from search results to source
- No breadcrumb context
Opportunity:
Query Result (from /memory/query):
- source: "runbooks/kubernetes/networking.md"
- breadcrumb: ["runbooks", "kubernetes"]
- section: "Port Forwarding"
User Clicks "View Full Document"
→ /memory/vault/poimen/runbooks/kubernetes/networking.md
→ Server highlights the relevant section
→ Shows context (adjacent sections)
Recommended API Enhancements
Phase 1: Enhance /memory/query (Now)
GET /memory/query?query=<text>&project=<proj>&limit=<k>&method=<hybrid|semantic|lexical>
Authorization: Bearer <JWT>
Returns:
{
"query": "fix kubernetes port",
"project": "poimen",
"search_method": "hybrid",
"results": [
{
"id": "chunk-123",
"text": "...",
"source": "...",
"score": 0.92,
"breakdown": {
"semantic": 0.88,
"lexical": 0.96,
"semantic_weight": 0.6,
"lexical_weight": 0.4,
"reason": "High semantic + lexical match"
}
}
],
"metrics": {
"retrieval_time_ms": 245,
"semantic_engine": "pgvector",
"lexical_engine": "opensearch",
"total_docs_searched": 100
}
}
Implementation:
- Update
query_handlerto acceptmethodparameter - Implement hybrid retrieval in
HybridQueryWorker - Forward JWT token to OpenSearch
- Return score breakdown
Phase 2: Add Context-Aware Retrieval (Week 2)
New Endpoint: /memory/query/with-context
GET /memory/query/with-context?query=<text>&project=<proj>&context_chunks=2
Authorization: Bearer <JWT>
Returns:
{
"query": "...",
"results": [
{
"id": "chunk-123",
"text": "...",
"score": 0.92,
"context": {
"previous_chunk": { "id": "chunk-122", "text": "..." },
"next_chunk": { "id": "chunk-124", "text": "..." },
"section_title": "Port Forwarding",
"breadcrumb_full": ["runbooks", "kubernetes", "troubleshooting"]
}
}
]
}
Purpose:
- Return adjacent chunks for full context
- Help LLM (agent) understand query context
- Support follow-up questions
Phase 3: Add Faceted Search (Week 3)
Enhancement to /memory/query:
GET /memory/query?query=<text>&project=<proj>&filters=<json>
Authorization: Bearer <JWT>
Query:
- filters: {"level": "L1", "source_pattern": "runbooks/*", "updated_since": "2024-01-01"}
Returns:
{
"query": "...",
"filters_applied": { "level": "L1", ... },
"results": [...],
"facets": {
"levels": { "L1": 15, "L2": 8 },
"sources": { "runbooks": 12, "docs": 11 },
"dates": { "2024-01": 14, "2024-02": 9 }
}
}
Phase 4: Add Relevance Feedback (Week 4)
New Endpoint: /memory/query/feedback
POST /memory/query/feedback
Authorization: Bearer <JWT>
Body:
{
"query_id": "q-123",
"query_text": "fix kubernetes port",
"result_id": "chunk-123",
"relevant": true, // or false
"rating": 4, // 1-5 stars
"feedback_text": "This was exactly what I needed"
}
Returns:
{
"query_id": "q-123",
"status": "recorded",
"message": "Thank you for feedback"
}
Purpose:
- Train weight tuning (0.6/0.4 optimal?)
- Detect broken results
- Improve future rankings
Architecture Changes Needed
Current State (Semantic-Only)
/memory/query
│
└─ QueryWorker.query()
└─ pgvector (only)
Target State (Hybrid)
/memory/query
│
├─ Route query (decision tree)
│
├─ If HYBRID:
│ └─ HybridQueryWorker.hybrid_query()
│ ├─ Parallel pgvector (top-50)
│ ├─ Parallel OpenSearch + JWT (top-50)
│ ├─ Normalize scores
│ ├─ Fusion: 0.6*sem + 0.4*lex
│ └─ Return top-10 + breakdown
│
├─ If SEMANTIC:
│ └─ QueryWorker.query() [fallback]
│
└─ If LEXICAL:
└─ OpenSearchWorker.query() + JWT [new]
Implementation Checklist
✅ Already Implemented
/health— Health check/memory/ingest— Queue ingest/memory/ingest/{id}— Check status/memory/query— Semantic search (pgvector only)/memory/projects— List projects/memory/skills— List skills/memory/vault/generate— Generate vault/memory/vault— Browse vault root/memory/vault/{project}— Browse project/memory/vault/{project}/{file}— Read file- JWT auth (validates Authentik tokens)
- Rate limiting (per user, per endpoint)
- Capability checks (memory:read, memory:write)
- Idempotency (ingest deduplication)
📋 TODO: Hybrid Query Support
- Implement
HybridQueryWorker - Add OpenSearch client integration
- Implement score normalization
- Implement fusion strategy (weighted linear)
- Add query routing decision tree
- Add
?method=parameter to/memory/query - Return score breakdown in response
- Test with multiple weights (A/B testing)
- Add metrics endpoint for latency tracking
📋 TODO: Enhanced Retrieval
- Add
/memory/query/with-contextendpoint - Implement adjacent chunk retrieval
- Add faceted search to
/memory/query - Add
/memory/query/feedbackfor relevance feedback - Link search results to vault files
Summary
Query APIs (Search/AI)
/memory/query⭐ Primary semantic search- Searches indexed chunks (pgvector)
- Should become hybrid (pgvector + OpenSearch)
- Returns: scored, ranked results + breakdown
- Used by: LLM agents, frontend search UI
- Target: 0-500ms latency
Retrieval APIs (Browse/Read)
-
/memory/vault⭐ Browse vault tree- Lists projects and files from PVC
- Returns: file structure (no content)
- Used by: Web UI vault browser
- Target: 100-200ms latency
-
/memory/vault/{project}⭐ Browse project- Lists files in a project
- Returns: file tree with metadata
- Used by: Web UI file picker
-
/memory/vault/{project}/{file}⭐ Read file- Reads full file content from PVC
- Returns: markdown + sections
- Used by: Web UI file viewer
- Target: 50-100ms latency
Gap
- Query APIs and Retrieval APIs are disconnected
- Search results should link to vault files
- No way to "view full context" after search
- Solution: Add
/memory/query/with-contextendpoint