Files
poimen-memory/docs/API_REVIEW.md
T
Story Crater Bot 56bee1915e chore: Archive completed task files (M0, M1, M3, M3.5, M4.1-2, M3.6.1)
Deleted 31 completed task files:
- M0.x: 8 tasks (cargo, domain types, recordsource, tokenizer, adapters, gate)
- M1.x: 8 tasks (llm-chat, standing-query, prompt template, parser, loop, log, e2e, gate)
- M3.x: 4 tasks (l2-synthesis, rerank, mem-query, gate)
- M3.5.x: 8 tasks (http-server, ingest, query, federation, skills, projects, rate-limiting, gate)
- M3.6.1: DocCorpusSource (heading-boundary chunking)
- M4.1-2: skill-draft, derived-filter

Updated INDEX.md:
- Removed M0 & M1 phase sections (archived in git history)
- Updated progress table: 65 active tasks (42 + 2🟡 + 21)
- Updated status: M0/M1 complete, M3/M3.5 gates passing, M4.1-2 done
- Noted M3.5.10 JWT auth implementation complete (awaiting image rollout)
- Cleaned up broken links to deleted task files

Total test count: 239 passing, 2 ignored (up from 196 at M3.4)
Ready for M4.3 gate composition, M5 post-training, M7 source connectors.
2026-08-27 20:25:05 -07:00

12 KiB

API Review: Query APIs vs Retrieval APIs

Current API Endpoints (9 routes)

Overview Table

# Endpoint Method Category Auth Rate Limit Purpose
1 /health GET System No No Health check
2 /memory/ingest POST Write (write) 100/hr Queue ingest job
3 /memory/ingest/{id} GET Status (read) Query Check ingest progress
4 /memory/query GET QUERY API (read) 1000/hr Semantic search
5 /memory/projects GET Metadata (read) 100/hr List projects
6 /memory/skills GET Metadata (read) Query List skills
7 /memory/vault/generate POST Write (write) 100/hr Generate vault
8 /memory/vault GET RETRIEVAL API (read) Query Browse vault root
9 /memory/vault/{project} GET RETRIEVAL API (read) Query Browse project
10 /memory/vault/{project}/{file} GET RETRIEVAL API (read) Query Read file

Distinction: Query APIs vs Retrieval APIs

Query APIs (Search/Semantic)

🔍 /memory/query — Semantic Search (PRIMARY)

GET /memory/query?query=<text>&project=<proj>&limit=<k>
Authorization: Bearer <JWT>

Query Parameters:
  - query (required): Search string (e.g., "fix kubernetes port 8080")
  - project (required): Filter by project
  - limit (optional): Top-K results (default: 5)

Returns:
{
  "query": "fix kubernetes port 8080",
  "project": "poimen",
  "results": [
    {
      "id": "chunk-123",
      "text": "kubectl port-forward...",
      "source": "runbooks/port-forward.md",
      "level": "L1",
      "breadcrumb": ["runbooks", "kubernetes"],
      "score": 0.92,
      "sem_component": 0.88,
      "lex_component": 0.96  // [NEW] when hybrid enabled
    },
    ...
  ]
}

Current Implementation:

  • JWT auth enforced
  • Rate limited: 1000/hr
  • Capability check: memory:read
  • MISSING: Hybrid search integration
  • MISSING: Query routing decision tree
  • MISSING: OpenSearch lexical search

Flow:

GET /memory/query
  │
  ├─ Validate JWT / ApiKey
  ├─ Check capability: memory:read
  ├─ Check rate limit: 1000/hr per user
  │
  ├─ (Current) QueryWorker.query()
  │  └─ Only semantic via pgvector
  │
  └─ (NEEDED) HybridQueryWorker.query()
     ├─ Parallel pgvector search (top-50)
     ├─ Parallel OpenSearch BM25 (top-50) + JWT forward
     ├─ Normalize scores to [0.0, 1.0]
     ├─ Fusion: 0.6*sem + 0.4*lex
     └─ Return top-10 results with breakdown

Retrieval APIs (Browse/Read)

📂 /memory/vault — List Vault Root

GET /memory/vault
Authorization: Bearer <JWT>

Returns:
{
  "projects": [
    {
      "name": "poimen",
      "path": "vault/poimen",
      "file_count": 42,
      "updated_at": "2024-01-15T10:30:00Z"
    },
    ...
  ]
}

Purpose: Browse vault directory structure (top-level)


📂 /memory/vault/{project} — List Project Files

GET /memory/vault/poimen
Authorization: Bearer <JWT>

Returns:
{
  "project": "poimen",
  "files": [
    {
      "name": "runbooks",
      "type": "directory",
      "path": "vault/poimen/runbooks",
      "file_count": 15
    },
    {
      "name": "deployment.md",
      "type": "file",
      "path": "vault/poimen/deployment.md",
      "size_bytes": 4096,
      "updated_at": "2024-01-15T10:30:00Z"
    },
    ...
  ]
}

Purpose: Browse files in a project (file tree view)


📄 /memory/vault/{project}/{file} — Read File Content

GET /memory/vault/poimen/deployment.md
Authorization: Bearer <JWT>

Returns:
{
  "project": "poimen",
  "file": "deployment.md",
  "path": "vault/poimen/deployment.md",
  "content": "# Deployment Guide\n\n## Overview\n...",
  "size_bytes": 4096,
  "updated_at": "2024-01-15T10:30:00Z",
  "sections": [
    {
      "title": "Overview",
      "level": 1,
      "content": "..."
    },
    ...
  ]
}

Purpose: Read/view full file content (no AI processing)


Current Issues & Gaps

Issue 1: /memory/query is Semantic-Only

Problem:

  • Current /memory/query calls QueryWorker.query() which ONLY does pgvector
  • No access to OpenSearch (lexical search)
  • No fusion/reranking logic
  • No score breakdown for debugging

Solution:

// Instead of:
state.query_worker.query(&project, &question, Some(limit)).await

// Should be:
state.hybrid_query_worker.hybrid_query(
    &project,
    &question,
    limit,
    SearchMethod::Hybrid,  // or Semantic, Lexical
    HybridWeights { semantic: 0.6, lexical: 0.4 },
    claims,
    &token  // Forward JWT to OpenSearch
).await

Issue 2: No Query Routing

Problem:

  • All queries use hybrid (once implemented)
  • No decision tree for:
    • Short queries (< 3 tokens) → lexical better
    • Special syntax (#tag) → lexical better
    • LLM unavailable → fallback to lexical

Solution:

pub fn route_query(query: &str, method_override: Option<SearchMethod>) -> SearchMethod {
    // If user explicitly requested a method, use it
    if let Some(method) = method_override {
        return method;
    }
    
    // Otherwise, auto-route based on query characteristics
    let token_count = query.split_whitespace().count();
    
    if token_count < 3 {
        SearchMethod::Lexical  // Short queries
    } else if query.contains('#') || query.contains('@') {
        SearchMethod::Lexical  // Special syntax
    } else {
        SearchMethod::Hybrid   // Normal queries
    }
}

Problem:

  • /memory/vault/* only reads files from PVC
  • No integration with indexed chunks
  • User can't "click through" from search results to source
  • No breadcrumb context

Opportunity:

Query Result (from /memory/query):
  - source: "runbooks/kubernetes/networking.md"
  - breadcrumb: ["runbooks", "kubernetes"]
  - section: "Port Forwarding"
  
User Clicks "View Full Document"
  → /memory/vault/poimen/runbooks/kubernetes/networking.md
  → Server highlights the relevant section
  → Shows context (adjacent sections)

Phase 1: Enhance /memory/query (Now)

GET /memory/query?query=<text>&project=<proj>&limit=<k>&method=<hybrid|semantic|lexical>
Authorization: Bearer <JWT>

Returns:
{
  "query": "fix kubernetes port",
  "project": "poimen",
  "search_method": "hybrid",
  "results": [
    {
      "id": "chunk-123",
      "text": "...",
      "source": "...",
      "score": 0.92,
      "breakdown": {
        "semantic": 0.88,
        "lexical": 0.96,
        "semantic_weight": 0.6,
        "lexical_weight": 0.4,
        "reason": "High semantic + lexical match"
      }
    }
  ],
  "metrics": {
    "retrieval_time_ms": 245,
    "semantic_engine": "pgvector",
    "lexical_engine": "opensearch",
    "total_docs_searched": 100
  }
}

Implementation:

  1. Update query_handler to accept method parameter
  2. Implement hybrid retrieval in HybridQueryWorker
  3. Forward JWT token to OpenSearch
  4. Return score breakdown

Phase 2: Add Context-Aware Retrieval (Week 2)

New Endpoint: /memory/query/with-context

GET /memory/query/with-context?query=<text>&project=<proj>&context_chunks=2
Authorization: Bearer <JWT>

Returns:
{
  "query": "...",
  "results": [
    {
      "id": "chunk-123",
      "text": "...",
      "score": 0.92,
      "context": {
        "previous_chunk": { "id": "chunk-122", "text": "..." },
        "next_chunk": { "id": "chunk-124", "text": "..." },
        "section_title": "Port Forwarding",
        "breadcrumb_full": ["runbooks", "kubernetes", "troubleshooting"]
      }
    }
  ]
}

Purpose:

  • Return adjacent chunks for full context
  • Help LLM (agent) understand query context
  • Support follow-up questions

Phase 3: Add Faceted Search (Week 3)

Enhancement to /memory/query:

GET /memory/query?query=<text>&project=<proj>&filters=<json>
Authorization: Bearer <JWT>

Query:
  - filters: {"level": "L1", "source_pattern": "runbooks/*", "updated_since": "2024-01-01"}

Returns:
{
  "query": "...",
  "filters_applied": { "level": "L1", ... },
  "results": [...],
  "facets": {
    "levels": { "L1": 15, "L2": 8 },
    "sources": { "runbooks": 12, "docs": 11 },
    "dates": { "2024-01": 14, "2024-02": 9 }
  }
}

Phase 4: Add Relevance Feedback (Week 4)

New Endpoint: /memory/query/feedback

POST /memory/query/feedback
Authorization: Bearer <JWT>

Body:
{
  "query_id": "q-123",
  "query_text": "fix kubernetes port",
  "result_id": "chunk-123",
  "relevant": true,  // or false
  "rating": 4,       // 1-5 stars
  "feedback_text": "This was exactly what I needed"
}

Returns:
{
  "query_id": "q-123",
  "status": "recorded",
  "message": "Thank you for feedback"
}

Purpose:

  • Train weight tuning (0.6/0.4 optimal?)
  • Detect broken results
  • Improve future rankings

Architecture Changes Needed

Current State (Semantic-Only)

/memory/query
  │
  └─ QueryWorker.query()
     └─ pgvector (only)

Target State (Hybrid)

/memory/query
  │
  ├─ Route query (decision tree)
  │
  ├─ If HYBRID:
  │  └─ HybridQueryWorker.hybrid_query()
  │     ├─ Parallel pgvector (top-50)
  │     ├─ Parallel OpenSearch + JWT (top-50)
  │     ├─ Normalize scores
  │     ├─ Fusion: 0.6*sem + 0.4*lex
  │     └─ Return top-10 + breakdown
  │
  ├─ If SEMANTIC:
  │  └─ QueryWorker.query() [fallback]
  │
  └─ If LEXICAL:
     └─ OpenSearchWorker.query() + JWT [new]

Implementation Checklist

Already Implemented

  • /health — Health check
  • /memory/ingest — Queue ingest
  • /memory/ingest/{id} — Check status
  • /memory/query — Semantic search (pgvector only)
  • /memory/projects — List projects
  • /memory/skills — List skills
  • /memory/vault/generate — Generate vault
  • /memory/vault — Browse vault root
  • /memory/vault/{project} — Browse project
  • /memory/vault/{project}/{file} — Read file
  • JWT auth (validates Authentik tokens)
  • Rate limiting (per user, per endpoint)
  • Capability checks (memory:read, memory:write)
  • Idempotency (ingest deduplication)

📋 TODO: Hybrid Query Support

  • Implement HybridQueryWorker
  • Add OpenSearch client integration
  • Implement score normalization
  • Implement fusion strategy (weighted linear)
  • Add query routing decision tree
  • Add ?method= parameter to /memory/query
  • Return score breakdown in response
  • Test with multiple weights (A/B testing)
  • Add metrics endpoint for latency tracking

📋 TODO: Enhanced Retrieval

  • Add /memory/query/with-context endpoint
  • Implement adjacent chunk retrieval
  • Add faceted search to /memory/query
  • Add /memory/query/feedback for relevance feedback
  • Link search results to vault files

Summary

Query APIs (Search/AI)

  • /memory/query Primary semantic search
    • Searches indexed chunks (pgvector)
    • Should become hybrid (pgvector + OpenSearch)
    • Returns: scored, ranked results + breakdown
    • Used by: LLM agents, frontend search UI
    • Target: 0-500ms latency

Retrieval APIs (Browse/Read)

  • /memory/vault Browse vault tree

    • Lists projects and files from PVC
    • Returns: file structure (no content)
    • Used by: Web UI vault browser
    • Target: 100-200ms latency
  • /memory/vault/{project} Browse project

    • Lists files in a project
    • Returns: file tree with metadata
    • Used by: Web UI file picker
  • /memory/vault/{project}/{file} Read file

    • Reads full file content from PVC
    • Returns: markdown + sections
    • Used by: Web UI file viewer
    • Target: 50-100ms latency

Gap

  • Query APIs and Retrieval APIs are disconnected
  • Search results should link to vault files
  • No way to "view full context" after search
  • Solution: Add /memory/query/with-context endpoint