Deleted 31 completed task files: - M0.x: 8 tasks (cargo, domain types, recordsource, tokenizer, adapters, gate) - M1.x: 8 tasks (llm-chat, standing-query, prompt template, parser, loop, log, e2e, gate) - M3.x: 4 tasks (l2-synthesis, rerank, mem-query, gate) - M3.5.x: 8 tasks (http-server, ingest, query, federation, skills, projects, rate-limiting, gate) - M3.6.1: DocCorpusSource (heading-boundary chunking) - M4.1-2: skill-draft, derived-filter Updated INDEX.md: - Removed M0 & M1 phase sections (archived in git history) - Updated progress table: 65 active tasks (42✅ + 2🟡 + 21⬜) - Updated status: M0/M1 complete, M3/M3.5 gates passing, M4.1-2 done - Noted M3.5.10 JWT auth implementation complete (awaiting image rollout) - Cleaned up broken links to deleted task files Total test count: 239 passing, 2 ignored (up from 196 at M3.4) Ready for M4.3 gate composition, M5 post-training, M7 source connectors.
502 lines
12 KiB
Markdown
502 lines
12 KiB
Markdown
# API Review: Query APIs vs Retrieval APIs
|
|
|
|
## Current API Endpoints (9 routes)
|
|
|
|
### Overview Table
|
|
|
|
| # | Endpoint | Method | Category | Auth | Rate Limit | Purpose |
|
|
|---|----------|--------|----------|------|-----------|---------|
|
|
| 1 | `/health` | GET | System | ❌ No | ❌ No | Health check |
|
|
| 2 | `/memory/ingest` | POST | **Write** | ✅ (write) | 100/hr | Queue ingest job |
|
|
| 3 | `/memory/ingest/{id}` | GET | **Status** | ✅ (read) | ✅ Query | Check ingest progress |
|
|
| 4 | `/memory/query` | GET | **QUERY API** ⭐ | ✅ (read) | 1000/hr | Semantic search |
|
|
| 5 | `/memory/projects` | GET | Metadata | ✅ (read) | 100/hr | List projects |
|
|
| 6 | `/memory/skills` | GET | Metadata | ✅ (read) | ✅ Query | List skills |
|
|
| 7 | `/memory/vault/generate` | POST | **Write** | ✅ (write) | 100/hr | Generate vault |
|
|
| 8 | `/memory/vault` | GET | **RETRIEVAL API** ⭐ | ✅ (read) | ✅ Query | Browse vault root |
|
|
| 9 | `/memory/vault/{project}` | GET | **RETRIEVAL API** ⭐ | ✅ (read) | ✅ Query | Browse project |
|
|
| 10 | `/memory/vault/{project}/{file}` | GET | **RETRIEVAL API** ⭐ | ✅ (read) | ✅ Query | Read file |
|
|
|
|
---
|
|
|
|
## Distinction: Query APIs vs Retrieval APIs
|
|
|
|
### Query APIs (Search/Semantic)
|
|
|
|
#### 🔍 `/memory/query` — Semantic Search (PRIMARY)
|
|
|
|
```
|
|
GET /memory/query?query=<text>&project=<proj>&limit=<k>
|
|
Authorization: Bearer <JWT>
|
|
|
|
Query Parameters:
|
|
- query (required): Search string (e.g., "fix kubernetes port 8080")
|
|
- project (required): Filter by project
|
|
- limit (optional): Top-K results (default: 5)
|
|
|
|
Returns:
|
|
{
|
|
"query": "fix kubernetes port 8080",
|
|
"project": "poimen",
|
|
"results": [
|
|
{
|
|
"id": "chunk-123",
|
|
"text": "kubectl port-forward...",
|
|
"source": "runbooks/port-forward.md",
|
|
"level": "L1",
|
|
"breadcrumb": ["runbooks", "kubernetes"],
|
|
"score": 0.92,
|
|
"sem_component": 0.88,
|
|
"lex_component": 0.96 // [NEW] when hybrid enabled
|
|
},
|
|
...
|
|
]
|
|
}
|
|
```
|
|
|
|
**Current Implementation:**
|
|
- ✅ JWT auth enforced
|
|
- ✅ Rate limited: 1000/hr
|
|
- ✅ Capability check: `memory:read`
|
|
- ⏳ **MISSING**: Hybrid search integration
|
|
- ⏳ **MISSING**: Query routing decision tree
|
|
- ⏳ **MISSING**: OpenSearch lexical search
|
|
|
|
**Flow:**
|
|
```
|
|
GET /memory/query
|
|
│
|
|
├─ Validate JWT / ApiKey
|
|
├─ Check capability: memory:read
|
|
├─ Check rate limit: 1000/hr per user
|
|
│
|
|
├─ (Current) QueryWorker.query()
|
|
│ └─ Only semantic via pgvector
|
|
│
|
|
└─ (NEEDED) HybridQueryWorker.query()
|
|
├─ Parallel pgvector search (top-50)
|
|
├─ Parallel OpenSearch BM25 (top-50) + JWT forward
|
|
├─ Normalize scores to [0.0, 1.0]
|
|
├─ Fusion: 0.6*sem + 0.4*lex
|
|
└─ Return top-10 results with breakdown
|
|
```
|
|
|
|
---
|
|
|
|
### Retrieval APIs (Browse/Read)
|
|
|
|
#### 📂 `/memory/vault` — List Vault Root
|
|
|
|
```
|
|
GET /memory/vault
|
|
Authorization: Bearer <JWT>
|
|
|
|
Returns:
|
|
{
|
|
"projects": [
|
|
{
|
|
"name": "poimen",
|
|
"path": "vault/poimen",
|
|
"file_count": 42,
|
|
"updated_at": "2024-01-15T10:30:00Z"
|
|
},
|
|
...
|
|
]
|
|
}
|
|
```
|
|
|
|
**Purpose:** Browse vault directory structure (top-level)
|
|
|
|
---
|
|
|
|
#### 📂 `/memory/vault/{project}` — List Project Files
|
|
|
|
```
|
|
GET /memory/vault/poimen
|
|
Authorization: Bearer <JWT>
|
|
|
|
Returns:
|
|
{
|
|
"project": "poimen",
|
|
"files": [
|
|
{
|
|
"name": "runbooks",
|
|
"type": "directory",
|
|
"path": "vault/poimen/runbooks",
|
|
"file_count": 15
|
|
},
|
|
{
|
|
"name": "deployment.md",
|
|
"type": "file",
|
|
"path": "vault/poimen/deployment.md",
|
|
"size_bytes": 4096,
|
|
"updated_at": "2024-01-15T10:30:00Z"
|
|
},
|
|
...
|
|
]
|
|
}
|
|
```
|
|
|
|
**Purpose:** Browse files in a project (file tree view)
|
|
|
|
---
|
|
|
|
#### 📄 `/memory/vault/{project}/{file}` — Read File Content
|
|
|
|
```
|
|
GET /memory/vault/poimen/deployment.md
|
|
Authorization: Bearer <JWT>
|
|
|
|
Returns:
|
|
{
|
|
"project": "poimen",
|
|
"file": "deployment.md",
|
|
"path": "vault/poimen/deployment.md",
|
|
"content": "# Deployment Guide\n\n## Overview\n...",
|
|
"size_bytes": 4096,
|
|
"updated_at": "2024-01-15T10:30:00Z",
|
|
"sections": [
|
|
{
|
|
"title": "Overview",
|
|
"level": 1,
|
|
"content": "..."
|
|
},
|
|
...
|
|
]
|
|
}
|
|
```
|
|
|
|
**Purpose:** Read/view full file content (no AI processing)
|
|
|
|
---
|
|
|
|
## Current Issues & Gaps
|
|
|
|
### ❌ Issue 1: `/memory/query` is Semantic-Only
|
|
|
|
**Problem:**
|
|
- Current `/memory/query` calls `QueryWorker.query()` which ONLY does pgvector
|
|
- No access to OpenSearch (lexical search)
|
|
- No fusion/reranking logic
|
|
- No score breakdown for debugging
|
|
|
|
**Solution:**
|
|
```rust
|
|
// Instead of:
|
|
state.query_worker.query(&project, &question, Some(limit)).await
|
|
|
|
// Should be:
|
|
state.hybrid_query_worker.hybrid_query(
|
|
&project,
|
|
&question,
|
|
limit,
|
|
SearchMethod::Hybrid, // or Semantic, Lexical
|
|
HybridWeights { semantic: 0.6, lexical: 0.4 },
|
|
claims,
|
|
&token // Forward JWT to OpenSearch
|
|
).await
|
|
```
|
|
|
|
---
|
|
|
|
### ❌ Issue 2: No Query Routing
|
|
|
|
**Problem:**
|
|
- All queries use hybrid (once implemented)
|
|
- No decision tree for:
|
|
- Short queries (< 3 tokens) → lexical better
|
|
- Special syntax (#tag) → lexical better
|
|
- LLM unavailable → fallback to lexical
|
|
|
|
**Solution:**
|
|
```rust
|
|
pub fn route_query(query: &str, method_override: Option<SearchMethod>) -> SearchMethod {
|
|
// If user explicitly requested a method, use it
|
|
if let Some(method) = method_override {
|
|
return method;
|
|
}
|
|
|
|
// Otherwise, auto-route based on query characteristics
|
|
let token_count = query.split_whitespace().count();
|
|
|
|
if token_count < 3 {
|
|
SearchMethod::Lexical // Short queries
|
|
} else if query.contains('#') || query.contains('@') {
|
|
SearchMethod::Lexical // Special syntax
|
|
} else {
|
|
SearchMethod::Hybrid // Normal queries
|
|
}
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
### ❌ Issue 3: Retrieval APIs Don't Integrate with Search
|
|
|
|
**Problem:**
|
|
- `/memory/vault/*` only reads files from PVC
|
|
- No integration with indexed chunks
|
|
- User can't "click through" from search results to source
|
|
- No breadcrumb context
|
|
|
|
**Opportunity:**
|
|
```
|
|
Query Result (from /memory/query):
|
|
- source: "runbooks/kubernetes/networking.md"
|
|
- breadcrumb: ["runbooks", "kubernetes"]
|
|
- section: "Port Forwarding"
|
|
|
|
User Clicks "View Full Document"
|
|
→ /memory/vault/poimen/runbooks/kubernetes/networking.md
|
|
→ Server highlights the relevant section
|
|
→ Shows context (adjacent sections)
|
|
```
|
|
|
|
---
|
|
|
|
## Recommended API Enhancements
|
|
|
|
### Phase 1: Enhance `/memory/query` (Now)
|
|
|
|
```
|
|
GET /memory/query?query=<text>&project=<proj>&limit=<k>&method=<hybrid|semantic|lexical>
|
|
Authorization: Bearer <JWT>
|
|
|
|
Returns:
|
|
{
|
|
"query": "fix kubernetes port",
|
|
"project": "poimen",
|
|
"search_method": "hybrid",
|
|
"results": [
|
|
{
|
|
"id": "chunk-123",
|
|
"text": "...",
|
|
"source": "...",
|
|
"score": 0.92,
|
|
"breakdown": {
|
|
"semantic": 0.88,
|
|
"lexical": 0.96,
|
|
"semantic_weight": 0.6,
|
|
"lexical_weight": 0.4,
|
|
"reason": "High semantic + lexical match"
|
|
}
|
|
}
|
|
],
|
|
"metrics": {
|
|
"retrieval_time_ms": 245,
|
|
"semantic_engine": "pgvector",
|
|
"lexical_engine": "opensearch",
|
|
"total_docs_searched": 100
|
|
}
|
|
}
|
|
```
|
|
|
|
**Implementation:**
|
|
1. Update `query_handler` to accept `method` parameter
|
|
2. Implement hybrid retrieval in `HybridQueryWorker`
|
|
3. Forward JWT token to OpenSearch
|
|
4. Return score breakdown
|
|
|
|
---
|
|
|
|
### Phase 2: Add Context-Aware Retrieval (Week 2)
|
|
|
|
**New Endpoint:** `/memory/query/with-context`
|
|
|
|
```
|
|
GET /memory/query/with-context?query=<text>&project=<proj>&context_chunks=2
|
|
Authorization: Bearer <JWT>
|
|
|
|
Returns:
|
|
{
|
|
"query": "...",
|
|
"results": [
|
|
{
|
|
"id": "chunk-123",
|
|
"text": "...",
|
|
"score": 0.92,
|
|
"context": {
|
|
"previous_chunk": { "id": "chunk-122", "text": "..." },
|
|
"next_chunk": { "id": "chunk-124", "text": "..." },
|
|
"section_title": "Port Forwarding",
|
|
"breadcrumb_full": ["runbooks", "kubernetes", "troubleshooting"]
|
|
}
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
**Purpose:**
|
|
- Return adjacent chunks for full context
|
|
- Help LLM (agent) understand query context
|
|
- Support follow-up questions
|
|
|
|
---
|
|
|
|
### Phase 3: Add Faceted Search (Week 3)
|
|
|
|
**Enhancement to `/memory/query`:**
|
|
|
|
```
|
|
GET /memory/query?query=<text>&project=<proj>&filters=<json>
|
|
Authorization: Bearer <JWT>
|
|
|
|
Query:
|
|
- filters: {"level": "L1", "source_pattern": "runbooks/*", "updated_since": "2024-01-01"}
|
|
|
|
Returns:
|
|
{
|
|
"query": "...",
|
|
"filters_applied": { "level": "L1", ... },
|
|
"results": [...],
|
|
"facets": {
|
|
"levels": { "L1": 15, "L2": 8 },
|
|
"sources": { "runbooks": 12, "docs": 11 },
|
|
"dates": { "2024-01": 14, "2024-02": 9 }
|
|
}
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
### Phase 4: Add Relevance Feedback (Week 4)
|
|
|
|
**New Endpoint:** `/memory/query/feedback`
|
|
|
|
```
|
|
POST /memory/query/feedback
|
|
Authorization: Bearer <JWT>
|
|
|
|
Body:
|
|
{
|
|
"query_id": "q-123",
|
|
"query_text": "fix kubernetes port",
|
|
"result_id": "chunk-123",
|
|
"relevant": true, // or false
|
|
"rating": 4, // 1-5 stars
|
|
"feedback_text": "This was exactly what I needed"
|
|
}
|
|
|
|
Returns:
|
|
{
|
|
"query_id": "q-123",
|
|
"status": "recorded",
|
|
"message": "Thank you for feedback"
|
|
}
|
|
```
|
|
|
|
**Purpose:**
|
|
- Train weight tuning (0.6/0.4 optimal?)
|
|
- Detect broken results
|
|
- Improve future rankings
|
|
|
|
---
|
|
|
|
## Architecture Changes Needed
|
|
|
|
### Current State (Semantic-Only)
|
|
|
|
```
|
|
/memory/query
|
|
│
|
|
└─ QueryWorker.query()
|
|
└─ pgvector (only)
|
|
```
|
|
|
|
### Target State (Hybrid)
|
|
|
|
```
|
|
/memory/query
|
|
│
|
|
├─ Route query (decision tree)
|
|
│
|
|
├─ If HYBRID:
|
|
│ └─ HybridQueryWorker.hybrid_query()
|
|
│ ├─ Parallel pgvector (top-50)
|
|
│ ├─ Parallel OpenSearch + JWT (top-50)
|
|
│ ├─ Normalize scores
|
|
│ ├─ Fusion: 0.6*sem + 0.4*lex
|
|
│ └─ Return top-10 + breakdown
|
|
│
|
|
├─ If SEMANTIC:
|
|
│ └─ QueryWorker.query() [fallback]
|
|
│
|
|
└─ If LEXICAL:
|
|
└─ OpenSearchWorker.query() + JWT [new]
|
|
```
|
|
|
|
---
|
|
|
|
## Implementation Checklist
|
|
|
|
### ✅ Already Implemented
|
|
- [x] `/health` — Health check
|
|
- [x] `/memory/ingest` — Queue ingest
|
|
- [x] `/memory/ingest/{id}` — Check status
|
|
- [x] `/memory/query` — Semantic search (pgvector only)
|
|
- [x] `/memory/projects` — List projects
|
|
- [x] `/memory/skills` — List skills
|
|
- [x] `/memory/vault/generate` — Generate vault
|
|
- [x] `/memory/vault` — Browse vault root
|
|
- [x] `/memory/vault/{project}` — Browse project
|
|
- [x] `/memory/vault/{project}/{file}` — Read file
|
|
- [x] JWT auth (validates Authentik tokens)
|
|
- [x] Rate limiting (per user, per endpoint)
|
|
- [x] Capability checks (memory:read, memory:write)
|
|
- [x] Idempotency (ingest deduplication)
|
|
|
|
### 📋 TODO: Hybrid Query Support
|
|
- [ ] Implement `HybridQueryWorker`
|
|
- [ ] Add OpenSearch client integration
|
|
- [ ] Implement score normalization
|
|
- [ ] Implement fusion strategy (weighted linear)
|
|
- [ ] Add query routing decision tree
|
|
- [ ] Add `?method=` parameter to `/memory/query`
|
|
- [ ] Return score breakdown in response
|
|
- [ ] Test with multiple weights (A/B testing)
|
|
- [ ] Add metrics endpoint for latency tracking
|
|
|
|
### 📋 TODO: Enhanced Retrieval
|
|
- [ ] Add `/memory/query/with-context` endpoint
|
|
- [ ] Implement adjacent chunk retrieval
|
|
- [ ] Add faceted search to `/memory/query`
|
|
- [ ] Add `/memory/query/feedback` for relevance feedback
|
|
- [ ] Link search results to vault files
|
|
|
|
---
|
|
|
|
## Summary
|
|
|
|
### Query APIs (Search/AI)
|
|
- **`/memory/query`** ⭐ Primary semantic search
|
|
- Searches indexed chunks (pgvector)
|
|
- Should become hybrid (pgvector + OpenSearch)
|
|
- Returns: scored, ranked results + breakdown
|
|
- Used by: LLM agents, frontend search UI
|
|
- **Target**: 0-500ms latency
|
|
|
|
### Retrieval APIs (Browse/Read)
|
|
- **`/memory/vault`** ⭐ Browse vault tree
|
|
- Lists projects and files from PVC
|
|
- Returns: file structure (no content)
|
|
- Used by: Web UI vault browser
|
|
- **Target**: 100-200ms latency
|
|
|
|
- **`/memory/vault/{project}`** ⭐ Browse project
|
|
- Lists files in a project
|
|
- Returns: file tree with metadata
|
|
- Used by: Web UI file picker
|
|
|
|
- **`/memory/vault/{project}/{file}`** ⭐ Read file
|
|
- Reads full file content from PVC
|
|
- Returns: markdown + sections
|
|
- Used by: Web UI file viewer
|
|
- **Target**: 50-100ms latency
|
|
|
|
### Gap
|
|
- Query APIs and Retrieval APIs are **disconnected**
|
|
- Search results should link to vault files
|
|
- No way to "view full context" after search
|
|
- **Solution**: Add `/memory/query/with-context` endpoint
|
|
|