50 KiB
Memory UI Flow - Complete Workflow
Table of Contents
- Read Flow
- Search Flow (Semantic + Lexical Hybrid)
- Edit Flow (GRM Workflow)
- Agent Context Flow
- System Architecture
- OpenSearch + JWT Authentication
- Pod Infrastructure
Read Flow
Browse vault documents from the web UI.
┌─────────────────────────────────────────────────────┐
│ User: memory.riotpiao.com │
│ (Browser, JWT token in localStorage) │
└────────────┬────────────────────────────────────────┘
│
│ GET /memory/vault?project=poimen
│ Authorization: Bearer <JWT>
│
↓
┌─────────────────────────────────────────────────────┐
│ Memory Service Pod │
│ ├─ Load vault files from PVC │
│ ├─ Build file tree (directory structure) │
│ └─ Return JSON response │
└────────────┬────────────────────────────────────────┘
│
│ {files: [{path, title, updated_at}...]}
│
↓
┌─────────────────────────────────────────────────────┐
│ UI: Render Vault Browser │
│ ├─ Project selector (dropdown) │
│ ├─ File tree (collapsible folders) │
│ ├─ Breadcrumb navigation │
│ └─ Preview panel (markdown rendering) │
└─────────────────────────────────────────────────────┘
Search Flow (Semantic + Lexical Hybrid)
Hybrid retrieval pipeline combining:
- Semantic path: pgvector embeddings (query understanding)
- Lexical path: OpenSearch BM25 (exact term matching)
- Fusion: Weighted linear combination (60% semantic, 40% lexical)
- Result limit: Top 50 from each engine, merge to top 10 final
Complete Retrieval Pipeline
User Query: "fix kubernetes port 8080 conflict"
│
↓
┌────────────────────────────────────────────────────────┐
│ Stage 1: Query Normalization │
├────────────────────────────────────────────────────────┤
│ ├─ Tokenize: ["fix", "kubernetes", ...] │
│ ├─ Extract entities: {port: "8080"} │
│ ├─ Generate embedding (LLM) │
│ └─ Create QueryContext │
└────────────────────────┬───────────────────────────────┘
│
↓
┌────────────────────────────────────────────────────────┐
│ Stage 2: Parallel Retrieval (Both Engines) │
├────────────────────────────────────────────────────────┤
│ │
│ ┌─ SEMANTIC (pgvector) │
│ │ ├─ Query embedding <-> vector │
│ │ ├─ Filter: project_id, level, date │
│ │ ├─ ORDER BY cosine_similarity DESC │
│ │ └─ Return: Top 50 with scores │
│ │ │
│ └─ LEXICAL (OpenSearch + JWT) │
│ ├─ multi_match on [content^2, breadcrumb] │
│ ├─ BM25 ranking with fuzziness │
│ ├─ Filter: project_id, level, date │
│ └─ Return: Top 50 with raw scores │
│ │
└────────────────┬──────────────────────────┬────────────┘
│ │
sem_results: [(doc1, 0.92), ...] lex_results: [(doc1, 8.5), ...]
│ │
└──────────┬───────────────┘
│
↓
┌────────────────────────────────────────────────────────┐
│ Stage 3: Score Normalization │
├────────────────────────────────────────────────────────┤
│ ├─ Normalize semantic: [0.0 ... 1.0] │
│ │ (already 0-1 from cosine) │
│ │ │
│ ├─ Normalize lexical: [0.0 ... 1.0] │
│ │ (min-max: (score-min)/(max-min)) │
│ │ │
│ └─ Result: Both in [0.0, 1.0] range │
│ │
└────────────────┬──────────────────────────────────────┘
│
sem_norm: [(doc1, 1.0), ...] lex_norm: [(doc1, 0.98), ...]
│
↓
┌────────────────────────────────────────────────────────┐
│ Stage 4: Fusion (Weighted Linear) │
├────────────────────────────────────────────────────────┤
│ ├─ Merge all doc IDs from both results │
│ ├─ For each doc: score = 0.6*sem + 0.4*lex │
│ │ │
│ │ doc1: 0.6*1.0 + 0.4*0.98 = 0.992 │
│ │ doc2: 0.6*0.96 + 0.4*0.0 = 0.576 │
│ │ doc3: 0.6*0.0 + 0.4*0.88 = 0.352 │
│ │ │
│ ├─ Sort descending: [doc1, doc2, doc3] │
│ └─ Take top-10 │
│ │
└────────────────┬──────────────────────────────────────┘
│
↓
┌────────────────────────────────────────────────────────┐
│ Stage 5: Score Breakdown (Transparency) │
├────────────────────────────────────────────────────────┤
│ doc1: { │
│ "score": 0.992, │
│ "sem_component": 1.0, │
│ "lex_component": 0.98, │
│ "sem_weight": 0.6, │
│ "lex_weight": 0.4, │
│ "reason": "Exact semantic match + strong lexical" │
│ } │
│ │
└────────────────┬──────────────────────────────────────┘
│
↓
┌────────────────────────────────────────────────────────┐
│ Final Results (Top-10) │
├────────────────────────────────────────────────────────┤
│ 1. port-forward.md (0.992) │
│ - Semantic: 1.0 | Lexical: 0.98 │
│ │
│ 2. troubleshooting.md (0.576) │
│ - Semantic: 0.96 | Lexical: 0.0 │
│ │
│ 3. k8s-basics.md (0.352) │
│ - Semantic: 0.0 | Lexical: 0.88 │
│ │
└────────────────────────────────────────────────────────┘
Query Routing Decision Tree
Query Received
│
├─ Token count < 3?
│ ├─ YES → Use LEXICAL_ONLY
│ │ (short queries: "fix port" → better BM25)
│ │
│ └─ NO → Continue...
│
├─ Contains special syntax (#tag, @mention)?
│ ├─ YES → Use LEXICAL_WITH_FILTERS
│ │ (preserve exact matches)
│ │
│ └─ NO → Continue...
│
├─ Can we embed the query?
│ ├─ YES → Use HYBRID
│ │ (both engines)
│ │
│ └─ NO → Use LEXICAL_ONLY
│ (LLM unavailable, fallback)
│
└─ Execute chosen strategy
Index Optimization
PostgreSQL (pgvector) - Semantic Path:
- Index:
ivfflat (embedding vector_cosine_ops) WITH (lists=100) - Filter:
project_id, level IN ('L0','L1','L2'), created_at > now-1y - Retrieve: Top 50 chunks, then merge with lexical
- Lookup: O(log n) pre-filter + O(1) embedding distance
OpenSearch (BM25) - Lexical Path:
- Analyzer:
standard(lowercase, stop words) - Fields:
content^2(2x boost) +breadcrumb+source - Tokenizer: Standard + n-gram for typo tolerance
- Retrieve: Top 50 results via BM25, then merge with semantic
- Lookup: O(n) inverted index scan + TF-IDF ranking
Merging Strategy:
- Normalize both score ranges to [0.0, 1.0]
- Weighted sum:
0.6 * semantic + 0.4 * lexical - Sort by final score
- Return top-10 to user
Accuracy Metrics (A/B Testing):
- MRR (Mean Reciprocal Rank): Position of first correct result
- NDCG@10 (Normalized Discounted Cumulative Gain): Quality of top-10 ranking
- Precision@K: Relevant results in top-K
- Recall@K: Coverage of all relevant results in top-K
Edit Flow (GRM Workflow)
Full Git Review Merge workflow: create branch → MR → human approval → auto-sync vault.
┌──────────────────────────────────────────────────────┐
│ User: Clicks "Edit" on document │
│ Example: runbook-deploy.md │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ UI: Switch to Edit Mode │
│ ├─ Load document content from Memory Service │
│ ├─ Show markdown editor (CodeMirror) │
│ ├─ Disable Save button (drafts only) │
│ └─ Show "Submit for Review" button │
└────────────┬─────────────────────────────────────────┘
│
├─ User makes edits (e.g., update deploy steps)
│
↓
┌──────────────────────────────────────────────────────┐
│ User: Click "Submit for Review" │
└────────────┬─────────────────────────────────────────┘
│
│
╔═══════╩═══════════════════════════════════════════╗
║ STEP 1: CREATE BRANCH ║
╚═══════╤═══════════════════════════════════════════╝
│
↓
┌──────────────────────────────────────────────────────┐
│ Frontend: POST /memory/grc/draft │
│ { │
│ "document_path": "vault/runbooks/deploy.md", │
│ "content": "<new markdown content>", │
│ "message": "Update deploy steps", │
│ "user": "[email protected]" │
│ } │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ Memory Service Pod: GRC Handler │
│ ├─ Generate branch name: edit/rock/deploy-<ts> │
│ ├─ Call Forgejo API (create branch) │
│ ├─ Commit changes to branch │
│ └─ Return PR URL + branch name │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ Forgejo Git Service │
│ ├─ Create branch: edit/rock/deploy-<ts> │
│ ├─ From: main │
│ ├─ Commit: "Update deploy steps" │
│ └─ Trigger CI checks (markdown lint) │
└────────────┬─────────────────────────────────────────┘
│
↓
╔═══════╩═══════════════════════════════════════════╗
║ STEP 2: AUTO-CREATE MERGE REQUEST ║
╚═══════╤═══════════════════════════════════════════╝
│
↓
┌──────────────────────────────────────────────────────┐
│ UI Feedback │
│ ✅ "Draft saved - Merge Request created" │
│ └─ Show clickable MR link │
└────────────┬─────────────────────────────────────────┘
│
│
╔═══════╩═══════════════════════════════════════════╗
║ STEP 3: HUMAN REVIEW (in Forgejo) ║
╚═══════╤═══════════════════════════════════════════╝
│
↓
┌──────────────────────────────────────────────────────┐
│ Reviewer (e.g., lead engineer) │
│ ├─ Open MR in Forgejo web UI │
│ ├─ Review diff (before/after) │
│ ├─ Comment/suggest edits │
│ ├─ Approve or request changes │
│ └─ Click "Merge to main" │
└────────────┬─────────────────────────────────────────┘
│
↓
╔═══════╩═══════════════════════════════════════════╗
║ STEP 4: AUTO-SYNC TO VAULT ║
╚═══════╤═══════════════════════════════════════════╝
│
↓
┌──────────────────────────────────────────────────────┐
│ Forgejo: Merge Complete │
│ ├─ Branch merged to main │
│ ├─ Trigger webhook: pull_request_merged │
│ └─ Payload: {pr_id, merged_at, branch} │
└────────────┬─────────────────────────────────────────┘
│
│ Webhook trigger
│
↓
┌──────────────────────────────────────────────────────┐
│ ArgoCD Application │
│ ├─ Webhook receiver │
│ ├─ Trigger sync of poimen-memory-app │
│ └─ Pull latest from git (main) │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ Git-Sync Sidecar Pod (poimen namespace) │
│ ├─ Receive ArgoCD sync signal │
│ ├─ `git pull origin main` in vault/ │
│ ├─ File appears in PVC │
│ └─ Update complete │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ Memory Service Pod: Indexing Job │
│ ├─ Detect vault file change │
│ ├─ Tokenize + embed new content │
│ ├─ Insert into pgvector index │
│ └─ Document now searchable │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ UI Notification │
│ ✅ "Document published!" │
│ ├─ Document now visible to all │
│ ├─ Embeddings indexed │
│ └─ Available in search │
└──────────────────────────────────────────────────────┘
Agent Context Flow
Real-time agent execution with memory retrieval tracking.
┌──────────────────────────────────────────────────────┐
│ User: Navigate to "Agent Workspace" tab │
│ (Shows live agent execution) │
└────────────┬─────────────────────────────────────────┘
│
│ Establish connection
│
↓
┌──────────────────────────────────────────────────────┐
│ Frontend: WebSocket /memory/agents/stream │
│ (Fallback: HTTP polling) │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ Memory Service Pod │
│ ├─ Tail agent execution log │
│ ├─ Emit events: │
│ │ - agent_started │
│ │ - memory_retrieved {query, chunks, scores} │
│ │ - tool_invoked {tool_name, args} │
│ │ - tool_result {result} │
│ │ - agent_decision {reasoning} │
│ │ - agent_complete │
│ └─ Stream as JSON events │
└────────────┬─────────────────────────────────────────┘
│
↓
┌──────────────────────────────────────────────────────┐
│ UI: Real-time Dashboard │
│ ├─ Timeline of agent actions (bottom-up) │
│ ├─ Memory chunks used (with similarity scores) │
│ ├─ Tool calls + outputs (expandable) │
│ ├─ Decision tree (branching logic) │
│ └─ Knowledge graph overlay (related docs) │
└──────────────────────────────────────────────────────┘
System Architecture
Complete deployment topology with all components.
┌────────────────────────────────────────────────────────────────────────┐
│ EXTERNAL: User → memory.riotpiao.com (DNS A record) │
└────────────┬───────────────────────────────────────────────────────────┘
│
│ HTTPS
│
↓
┌────────────────────────────────────────────────────────────────────────┐
│ K8s Ingress Controller (nginx-ingress) │
│ ├─ TLS termination (memory.riotpiao.com) │
│ ├─ Route to frontend Service (port 80) │
│ └─ Route to memory Service (port 8080) │
└────┬───────────────────────────────────┬───────────────────────────────┘
│ │
↓ (frontend) ↓ (API)
┌──────────────────────────┐ ┌────────────────────────────────────┐
│ Frontend Service │ │ Memory Service (8080) │
│ (port 80) │ │ ├─ LoadBalancer type │
└────┬─────────────────────┘ └────┬───────────────────────────────┘
│ │
↓ ↓
┌──────────────────────────┐ ┌────────────────────────────────────┐
│ Frontend Pod (React SPA)│ │ Memory Pod 1 (poimen-memory-*) │
│ ├─ React app │ │ ├─ HTTP server (actix-web) │
│ ├─ Vite build │ │ ├─ JWT validation │
│ ├─ Static files │ │ ├─ GRC handler (Forgejo API) │
│ └─ API client │ │ ├─ Vault browser │
└──────────────────────────┘ │ ├─ Query (embedding) handler │
│ ├─ Skills handler │
┌──────────────────────────┼─ Projects handler │
│ │ └─ Volume: /data/vault (PVC) │
│ │ │
│ └────┬───────────────────────────────┘
│ │
│ Memory Pod 2 (HA replica) │
│ (identical to Pod 1) │
│ │
├───────────────────────────────┤
│ │
↓ ↓
┌─────────────────────────────────────────────────────────────┐
│ PostgreSQL StatefulSet (memory-db-0, memory-db-1) │
│ ├─ Primary: memory-db-0 (PVC: 20Gi) │
│ ├─ Replica: memory-db-1 (PVC: 20Gi) │
│ ├─ Service: memory-db (headless) │
│ ├─ Tables: │
│ │ ├─ chunks (id, text, project_id, embedding, source) │
│ │ ├─ skills (id, name, metadata) │
│ │ ├─ projects (id, name) │
│ │ └─ agent_logs (id, agent_id, action, timestamp) │
│ └─ Extension: pgvector (vector similarity) │
└─────────────────────────────────────────────────────────────┘
│
│ INDEX: embedding <-> vector[]
│
└─ Used by: /memory/query (similarity search)
┌─────────────────────────────────────────────────────────────┐
│ Storage: PVC (poimen-memory-vault, 10Gi, Longhorn) │
│ ├─ Mount path: /data/vault │
│ ├─ Content: │
│ │ ├─ vault/skills/ │
│ │ ├─ vault/runbooks/ │
│ │ ├─ vault/evidence/ │
│ │ └─ .git/ (full git history) │
│ └─ Sync: git-sync sidecar (on file changes) │
└─────────────────────────────────────────────────────────────┘
│
│ ArgoCD monitors + syncs
│
↓
┌─────────────────────────────────────────────────────────────┐
│ Git-Sync Sidecar (runs in Memory Pod) │
│ ├─ Watches: https://forgejo.riotpiao.com/.../memory.git │
│ ├─ Branch: main │
│ ├─ Sync interval: 30s │
│ ├─ On merge: pulls to /data/vault │
│ └─ Triggers indexing │
└─────────────────────────────────────────────────────────────┘
│
│ Webhooks
│
↓
┌─────────────────────────────────────────────────────────────┐
│ External: Forgejo + ArgoCD │
│ ├─ Forgejo webhook: pr_merged → ArgoCD │
│ ├─ ArgoCD watches: poimen-memory-app (in git) │
│ ├─ Auto-sync enabled (prune + selfHeal) │
│ └─ Revision tracking │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ External: Authentik (OIDC) │
│ ├─ Issuer: https://authentik.riotpiao.com/.../ │
│ ├─ JWKS: .../jwks/ │
│ ├─ OAuth2 App: poimen-memory │
│ └─ Used by: JWT validation in Memory Service │
└─────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────┐
│ External: LLM Service (Embeddings) │
│ ├─ Provider: Vertex AI / Hugging Face / etc │
│ ├─ Used by: /memory/query (tokenize + embed) │
│ └─ Cached results (1hr TTL) │
└─────────────────────────────────────────────────────────────┘
OpenSearch + JWT Authentication
JWT Flow with OpenSearch
Frontend
│ Authorization: Bearer <JWT>
│ (Authentik-signed token)
│
↓
Memory Service
├─ Extract JWT from header
├─ Validate signature (Authentik JWKS)
├─ Verify expiry + issuer + audience
└─ Extract claims (sub, roles, permissions)
│
├─ Can query pgvector (no auth needed)
│
└─ Forward JWT to OpenSearch
│ Authorization: Bearer <same-JWT>
│
↓
OpenSearch
├─ Receive JWT in Authorization header
├─ JWT realm validates signature
├─ Extract roles from JWT claims
├─ Map to internal roles (read_vault, write_vault)
└─ Check index permissions
│
├─ Query allowed → return results
└─ Write denied → 403 Forbidden
Hybrid Search: Semantic + Lexical
Memory Service executes parallel searches:
POST /memory/query + JWT
│
├─ Path 1 (Semantic): pgvector
│ ├─ LLM embedding
│ ├─ Cosine similarity
│ └─ Score: 0.88 (understanding)
│
├─ Path 2 (Lexical): OpenSearch + JWT
│ ├─ Tokenize query
│ ├─ BM25 ranking
│ └─ Score: 0.96 (exact terms)
│
└─ Rerank (Weighted: 60% semantic + 40% lexical)
└─ Final score: 0.92
Returns: Combined results sorted by hybrid score
OpenSearch JWT Realm Configuration
opensearch_security:
authc:
realms:
jwt_realm:
type: jwt
order: 1
# Token extraction
token_name: Authorization
token_extractor: "Bearer " # Strip "Bearer " prefix
# JWKS from Authentik (auto-refresh hourly)
jwks_uri: "https://authentik.riotpiao.com/application/o/poimen-memory/jwks/"
jwks_refresh_interval_ms: 3600000
# Issuer validation
issuer: "https://authentik.riotpiao.com/application/o/poimen-memory/"
# Extract claims
roles_key: "roles" # From JWT claim
subject_key: "sub" # User identifier
Roles Mapping (JWT → OpenSearch)
JWT claims example:
{
"iss": "https://authentik.riotpiao.com/application/o/poimen-memory/",
"sub": "[email protected]",
"roles": ["read_vault", "write_vault"],
"permissions": ["memory:read", "memory:write"]
}
OpenSearch role definitions:
read_vault:
cluster_permissions: ["cluster:monitor/health"]
index_permissions:
- index_patterns: ["vault-*"]
allowed_actions: ["indices:data/read/search"]
write_vault:
cluster_permissions: ["cluster:monitor/health"]
index_permissions:
- index_patterns: ["vault-*"]
allowed_actions: ["indices:data/write/index", "indices:data/read/search"]
Network Security (K8s NetworkPolicy)
# Only Memory Service can access OpenSearch
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: opensearch-access
namespace: poimen
spec:
podSelector:
matchLabels:
app: opensearch
policyTypes:
- Ingress
ingress:
- from:
- podSelector:
matchLabels:
app.kubernetes.io/name: poimen-memory
ports:
- protocol: TCP
port: 9200
M3.7.7 → M3.7.8: Failure Diagnosis Pipeline
M3.7.7 Complete: Signature Extraction (✅ 18 tests passing)
Purpose: Reduce failure logs to canonical signatures that are byte-identical across runs.
Failure Log (50KB, noisy):
2026-08-21T10:02:11.482Z
/home/runner/work/Poimen/memory/...
npm ERR! code ERESOLVE unable to resolve dependency tree
npm ERR! [... 100 lines of consequence errors ...]
│
↓ [M3.7.7: extract() + normalise()]
│
Signature:
tool: "npm"
raw: "npm ERR! code ERESOLVE unable to resolve dependency tree"
normalised: "npm error eresolve dependency tree"
sig_sha: "abc123def456..." ← deterministic hash
rule: "npm_error_line"
Key invariants (all 18 tests verify):
- ✅ Same failure from 2 different runs → identical sig_sha
- ✅ Different failures → different sig_sha
- ✅ Removes timestamps, paths, SHAs, line:col, durations, addresses
- ✅ Picks first (root) error, not consequence lines
- ✅ Unknown tools fallback gracefully
- ✅ Tool part of signature identity (npm vs cargo errors differ)
M3.7.8 In Progress: Symptom Projection
Purpose: Transform user queries into normalized symptom vectors that can match extracted signatures.
The Problem:
User reports: "npm can't find tslib module"
Canonical sig: "npm ERR! code ERESOLVE unable to resolve dependency tree"
These don't hash the same way, so tier 1 (exact signature match) fails.
Symptom projection bridges this gap.
Design (3-stage normalization pipeline):
Stage 1: Extract Symptom Keywords
Input query: "npm error: unable to resolve dependency tree"
│
├─ Tool detection: "npm" (from query or context)
├─ Error pattern extraction: ["unable", "resolve", "dependency"]
├─ Stop word removal: [unable, to, the, of, ...]
└─ Normalize abbreviations: ERESOLVE → "resolve error"
ERR → "error"
RC → "return code"
OOM → "out of memory"
EACCES → "permission denied"
ENOENT → "not found"
│
↓
Keywords: ["npm", "error", "resolve", "dependency", "tree"]
Stage 2: Normalize to Canonical Form
Keywords: ["npm", "error", "resolve", "dependency", "tree"]
│
├─ Remove stop words (a, the, is, can, may, etc.)
├─ Lowercase & stemming (resolved → resolve)
├─ Expand shorthands:
│ - ERESOLVE → "resolve error"
│ - ERR → "error"
│ - EOF → "end of file"
│ - EACCES → "permission denied"
│
└─ Tool mapping (npm, cargo, kubectl, etc.)
(different tools handle same error differently)
│
↓
Normalized: "npm error resolve dependency tree"
Stage 3: Generate Symptom Hash
Normalized: "npm error resolve dependency tree"
│
├─ Sort keywords alphabetically (for consistency)
│ → "dependency error npm resolve tree"
│
├─ Join with spaces
├─ SHA256 hash (same algorithm as M3.7.7)
│ sym_sha = SHA256("npm" + "\n" + normalized)
│
└─ Store: SymptomVector { tool, normalised, sym_sha }
│
├─ sym_sha: "xyz789abc..." ← for lookup
├─ normalised: "dependency error npm resolve tree"
└─ keywords: ["dependency", "error", "npm", "resolve", "tree"]
Key insight: If extracted signature and user query normalize to the same sym_sha, they match (tier 1 exact hit). If not, fall through to tier 2 (hybrid search).
M3.7.4: Three-Tier Context Endpoint Integration
GET /memory/context?query=npm+ERR+ERESOLVE&tool=npm
│
│ Normalize with M3.7.8: SymptomVector
│
├─ TIER 1 (Exact): sym_sha lookup
│ ├─ Query: SELECT * FROM lessons WHERE sym_sha = ?
│ ├─ Hit: Return past solution (cached in memory)
│ └─ Miss: Continue to Tier 2
│
├─ TIER 2 (Semantic/Hybrid): M8 search
│ ├─ Query: POST /memory/query (hybrid: 60% pgvector + 40% OpenSearch)
│ ├─ Results: Top-10 chunks ranked by similarity
│ └─ Confidence: med-high (vector match)
│
└─ TIER 3 (Reference): M3.6 documentation
├─ Query: Reference corpus (kubectl docs, npm docs)
├─ Results: General guidance (not specific solution)
└─ Confidence: low (generic info)
Response:
{
"tier": 1,
"confidence": "high",
"results": [
{
"source": "lesson",
"title": "Fix npm ERESOLVE errors",
"content": "npm ERR! code ERESOLVE...",
"solution": "Run npm ci instead of npm install"
}
]
}
M3.7.8 Implementation Plan
Files to create:
-
crates/mem-core/src/symptom_projection.rs(250 LOC)project_symptom(tool: &str, query: &str) -> SymptomVectornormalize_query(text: &str) -> Stringabbrev_expand(word: &str) -> Stringstop_words()— predefined list- Abbreviation mappings per tool
-
tests/it_symptom_projection.rs(400 LOC, 6 assertions)a1_same_symptom_same_hash— query variants → same sym_shaa2_abbrev_expansion— ERESOLVE, ERR, OOM normalize correctlya3_stop_word_removal— "unable to resolve" → "resolve"a4_tool_consistency— tool name included in sym_shaa5_case_insensitive— "NPM" = "npm"a6_keyword_order_irrelevant— sorted before hashing
Test fixtures:
- 6 query examples (npm, cargo, kubectl) with expected normalized form
- Abbreviation expansions per tool
Integration with M3.7.4 endpoint:
- M3.7.4 calls
symptom_projection::project_symptom(tool, user_query) - Returns
SymptomVector { sym_sha, normalised, keywords } - Looks up sym_sha in lesson cache
- If miss: delegates to M8 hybrid search
Pod Infrastructure
Complete pod inventory deployed in poimen namespace.
Production Pods
| Pod Name | Role | Replicas | PVC | Purpose |
|---|---|---|---|---|
| poimen-memory-* | API Server | 2 | 10Gi vault | HTTP server, JWT auth, GRC, hybrid search |
| memory-db-0 | PostgreSQL Primary | 1 | 20Gi | pgvector semantic search index |
| memory-db-1 | PostgreSQL Replica | 1 | 20Gi | High availability, read replicas |
| opensearch-0 | OpenSearch Primary | 1 | 30Gi | Lexical (BM25) search, JWT realm |
| opensearch-1 | OpenSearch Replica | 1 | 30Gi | HA cluster node, JWT validation |
| frontend-* | React SPA | 1+ | — | Web UI (memory.riotpiao.com) |
Supporting Infrastructure (External)
| Component | Role | Location |
|---|---|---|
| Git-Sync Sidecar | Auto-pull vault | Embedded in memory pod |
| ArgoCD Application | CD orchestration | argocd namespace |
| Ingress Controller | Reverse proxy | ingress-nginx namespace |
| Longhorn | Storage provider | Storage layer |
Total Pod Count: 8 Production Pods
Namespace: poimen
├─ poimen-memory (ReplicaSet) × 2 pods ...................... (2)
├─ memory-db-0 (StatefulSet) ............................... (1)
├─ memory-db-1 (StatefulSet) ............................... (1)
├─ opensearch-0 (StatefulSet) .............................. (1) [NEW]
├─ opensearch-1 (StatefulSet) .............................. (1) [NEW]
└─ frontend (Deployment) × 1-2 pods ......................... (1-2)
Namespace: argocd
└─ argocd-server, argocd-repo-server, etc .................. (5+)
Namespace: ingress-nginx
└─ nginx-ingress-controller ................................ (1)
TOTAL: 14-16 pods (8 core + 6-8 supporting)
Pod Responsibilities
Memory Service Pod (×2, HA)
- Listen: 0.0.0.0:8080
- Endpoints:
GET /memory/vault— Read vault filesPOST /memory/query— Hybrid search (semantic + lexical)GET /memory/skills— List skillsPOST /memory/grc/draft— Create branch + MRGET /memory/grc/status— Check MR statusGET /memory/agents/logs— Stream agent events
- Auth: JWT (Authentik)
- Hybrid Search Logic:
- Validates JWT (Authentik JWKS)
- Queries pgvector (semantic in parallel)
- Queries OpenSearch with JWT (lexical in parallel)
- Reranks results (weighted: 60% semantic, 40% lexical)
- Returns combined results
- Connections:
- PostgreSQL (pgvector semantic search)
- OpenSearch (lexical search with JWT)
- Forgejo API (GRC)
- LLM service (embeddings)
- PVC (vault files)
PostgreSQL Pod (×2, Primary + Replica)
- Listen: 5432
- Service: memory-db (headless for StatefulSet)
- Storage: 20Gi per pod (PVC)
- Replication: Streaming replication (primary → replica)
- Extensions: pgvector
- Data:
- chunks table (with vector index for semantic search)
- skills table
- projects table
- agent_logs table
- Role: Semantic search engine (embeddings)
OpenSearch Pod (×2, Primary + Replica)
- Listen: 9200 (HTTP), 9300 (cluster communication)
- Service: opensearch (headless for cluster), opensearch-internal (for queries)
- Storage: 30Gi per pod (PVC)
- Cluster: poimen-memory (2-node minimum)
- Security:
- JWT realm enabled (validates Authentik tokens)
- JWKS endpoint: https://authentik.riotpiao.com/application/o/poimen-memory/jwks/
- Role mapping: Extract roles from JWT claims
- Index permissions: read_vault, write_vault roles
- Indices:
- vault-* (BM25 text search with TF-IDF scoring)
- Role: Lexical search engine (exact terms + TF-IDF)
Frontend Pod (×1-2)
- Listen: 80
- Serve: React SPA static files
- Endpoints:
/— App shell/api/*— Proxy to Memory Service (8080)
- Auth: JWT (localStorage)
- Build: Vite (production bundle)
Git-Sync Sidecar (embedded in Memory Pod)
- Runs: As a init container + background process
- Watch: Forgejo main branch
- Sync interval: 30 seconds
- Action on merge:
git pull→ trigger re-index
Traffic Flow Diagram
Internet (Users)
│
│ HTTPS
│
↓
┌─────────────────────────┐
│ Ingress Controller │
│ (nginx) │
└──┬────────────┬─────────┘
│ │
│ port 80 │ port 8080
│ │
↓ ↓
┌──────────┐ ┌──────────────────┐
│ Frontend │ │ Memory Service │
│ (React) │ │ (Rust + Actix) │
└────┬─────┘ └──┬───────────┬────┘
│ │ │
│ ┌──┘ └──┐
│ │ │
↓ ↓ ↓
┌───────────────────┐ ┌──────────────┐
│ PostgreSQL │ │ Vault PVC │
│ (pgvector index) │ │ (git files) │
└───────────────────┘ └──────────────┘
Deployment Checklist
Core Services
- Memory Service (2 pods) deployed in
poimennamespace - PostgreSQL StatefulSet (2 pods) deployed with pgvector
- OpenSearch StatefulSet (2 pods) deployed with JWT realm
- PVC: poimen-memory-vault (10Gi) attached
- PVC: opensearch data (30Gi per pod) attached
Security & Auth
- JWT auth enabled (Authentik integration at Memory Service)
- OpenSearch JWT realm configured (validates Authentik tokens)
- K8s NetworkPolicy (only Memory Service → OpenSearch)
- Role mapping (JWT claims → OpenSearch roles)
Integration
- Git-sync sidecar configured (auto-pull on merge)
- Ingress configured (memory.riotpiao.com)
- Hybrid search endpoints (/memory/query?search_method=hybrid)
Design Documentation
- Retrieval pipeline architecture (4-stage: normalize → parallel → fusion → ranking)
- Index optimization (pgvector ivfflat + OpenSearch BM25 tuning)
- Score fusion strategy (weighted linear + RRF alternative)
- Query routing decision tree (short queries → lexical, normal → hybrid)
- Accuracy metrics (MRR, NDCG@10, Precision@K, Recall@K)
- Weight tuning strategy (A/B testing framework)
- Reference: See
docs/HYBRID_SEARCH_DESIGN.md(19KB, comprehensive design)
Testing & Deployment
- Test fixture setup (query + expected results dataset)
- NDCG/MRR baseline measurements (semantic vs lexical)
- Weight tuning experiments (0.5/0.5, 0.6/0.4, 0.7/0.3, 0.4/0.6)
- Performance benchmarks (latency: parallel vs serial)
- Frontend pod deployment
- End-to-end hybrid search tests
- Gradual rollout (Phase 1: lexical-only → Phase 2: hybrid 10% → Phase 3: 100%)
- GRC endpoints tested
- Agent logging endpoints tested