feat: simplify queue naming, remove stale docs, add Queue CRDs
- Queue name now just 'poimen-chunks' (no project suffix) - Delete outdated CI/DESIGN docs (CLAUDE.md is source of truth) - Add k8s/infra/queue.yaml: poimen-chunks + DLQ (Ready) - Update test to expect new queue name format
This commit is contained in:
@@ -1,224 +0,0 @@
|
||||
# CI/CD for All Poimen Repos — Standardized Pattern
|
||||
|
||||
## Overview
|
||||
|
||||
All Poimen repos should follow the same CI/CD pattern for consistency and maintainability.
|
||||
|
||||
**Pattern**: Test locally → Build image → Push to registry → ArgoCD deploys
|
||||
|
||||
**Based on**: homelab-frontend (proven production pattern)
|
||||
|
||||
---
|
||||
|
||||
## Repos & Status
|
||||
|
||||
### Repos That Need Docker Deployment
|
||||
|
||||
| Repo | Status | Dockerfile | Notes |
|
||||
|------|--------|-----------|-------|
|
||||
| **poimen-memory** | ✅ Ready | Yes | This repo - see `.forgejo/workflows/build.yaml` |
|
||||
| **poimen** | ⏳ TBD | Yes (assumed) | Orchestrator - needs deployment |
|
||||
| **poimen-workflows** | ⏳ TBD | Maybe | Check if containerized |
|
||||
|
||||
### Repos That Don't Need Docker
|
||||
|
||||
| Repo | Status | Type | Notes |
|
||||
|------|--------|------|-------|
|
||||
| **homelab** | ✅ Done | K8s manifests | Validates with yamllint + kubeval |
|
||||
|
||||
---
|
||||
|
||||
## Implementation Checklist for Each Repo
|
||||
|
||||
### Step 0: Prerequisites
|
||||
- [ ] Repo has a `Dockerfile`
|
||||
- [ ] Repo has a `.forgejo/` or `.gitea/` directory
|
||||
- [ ] Docker builds successfully: `docker build -t test:latest .`
|
||||
- [ ] Tests pass: `cargo test` / `npm test` / etc
|
||||
|
||||
### Step 1: Create Workflow File
|
||||
```bash
|
||||
# Copy from poimen-memory:
|
||||
cp ~/workplace/Poimen/memory/.forgejo/workflows/build.yaml \
|
||||
~/workplace/Poimen/<repo>/.forgejo/workflows/build.yaml
|
||||
|
||||
# Edit if needed:
|
||||
# - Change IMAGE_NAME from "rock/poimen-memory" to "rock/<your-repo>"
|
||||
# - Adjust test command if not Rust (cargo test)
|
||||
```
|
||||
|
||||
### Step 2: Set Repository Secret
|
||||
```
|
||||
https://git.riotpiao.com/rock/<repo>/settings/secrets
|
||||
|
||||
Add:
|
||||
- Name: REGISTRY_PAT
|
||||
- Value: <org-token-or-personal-token>
|
||||
```
|
||||
|
||||
### Step 3: Commit & Push
|
||||
```bash
|
||||
git add .forgejo/workflows/build.yaml
|
||||
git commit -m "Add CI/CD: auto-build and push to registry"
|
||||
git push origin main
|
||||
```
|
||||
|
||||
### Step 4: Create ArgoCD Application
|
||||
```bash
|
||||
# Create k8s/argocd/<repo>-app.yaml
|
||||
|
||||
apiVersion: argoproj.io/v1alpha1
|
||||
kind: Application
|
||||
metadata:
|
||||
name: poimen-<repo>-app
|
||||
namespace: argocd
|
||||
spec:
|
||||
project: homelab
|
||||
source:
|
||||
repoURL: https://forgejo.riotpiao.com/rock/poimen-<repo>.git
|
||||
targetRevision: main
|
||||
path: k8s/app # adjust if different
|
||||
destination:
|
||||
server: https://kubernetes.default.svc
|
||||
namespace: poimen
|
||||
syncPolicy:
|
||||
automated:
|
||||
prune: true
|
||||
selfHeal: true
|
||||
```
|
||||
|
||||
### Step 5: Apply Application
|
||||
```bash
|
||||
kubectl apply -f k8s/argocd/<repo>-app.yaml
|
||||
```
|
||||
|
||||
### Done! ✅
|
||||
- Every push to main triggers:
|
||||
1. Test suite
|
||||
2. Docker build
|
||||
3. Push to `forgejo.riotpiao.com/rock/<repo>:latest`
|
||||
4. ArgoCD auto-deploys
|
||||
|
||||
---
|
||||
|
||||
## File Reference
|
||||
|
||||
### Workflow Comparison
|
||||
|
||||
**poimen-memory** (current):
|
||||
```yaml
|
||||
runs-on: golang
|
||||
container:
|
||||
image: docker:27-cli
|
||||
volumes:
|
||||
- /docker-certs/client:/docker-certs/client:ro
|
||||
env:
|
||||
DOCKER_HOST: tcp://localhost:2376
|
||||
DOCKER_TLS_VERIFY: "1"
|
||||
DOCKER_CERT_PATH: /docker-certs/client
|
||||
```
|
||||
|
||||
**Why this setup:**
|
||||
- Runs on `golang` runner (has Docker daemon)
|
||||
- Uses Docker CLI in container with DinD (Docker-in-Docker)
|
||||
- TLS certs mounted for secure daemon access
|
||||
- Allows building AND pushing in same job
|
||||
|
||||
### Test Job
|
||||
|
||||
Adjust for your language:
|
||||
|
||||
**Rust** (poimen-memory):
|
||||
```yaml
|
||||
runs-on: rust
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- run: cargo test --all
|
||||
```
|
||||
|
||||
**Go**:
|
||||
```yaml
|
||||
runs-on: golang
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- run: go test ./...
|
||||
```
|
||||
|
||||
**Node.js**:
|
||||
```yaml
|
||||
runs-on: docker
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- run: npm install && npm test
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Organization-Wide Setup
|
||||
|
||||
### One-Time: Set Organization Secret
|
||||
Instead of per-repo secrets, Forgejo supports organization secrets.
|
||||
|
||||
**If available**, set `REGISTRY_PAT` at org level:
|
||||
```
|
||||
https://git.riotpiao.com/rock/settings/secrets
|
||||
```
|
||||
|
||||
Then all repos automatically inherit it (no per-repo setup needed).
|
||||
|
||||
**Check**: Try accessing org secrets settings
|
||||
- If available: set once, use everywhere
|
||||
- If not: set per-repo (5 minutes per repo)
|
||||
|
||||
---
|
||||
|
||||
## Monitoring & Troubleshooting
|
||||
|
||||
### Build Failures
|
||||
|
||||
**Check logs:**
|
||||
```
|
||||
https://git.riotpiao.com/rock/<repo>/actions
|
||||
```
|
||||
|
||||
**Common issues:**
|
||||
- Test failures → Fix tests locally
|
||||
- Docker build error → Check Dockerfile syntax
|
||||
- Push fails → Verify REGISTRY_PAT token
|
||||
|
||||
### Deployment Issues
|
||||
|
||||
**Watch ArgoCD:**
|
||||
```bash
|
||||
kubectl get application -n argocd poimen-<repo>-app -w
|
||||
kubectl logs -n argocd argocd-application-controller | grep poimen
|
||||
```
|
||||
|
||||
**Check pods:**
|
||||
```bash
|
||||
kubectl get pods -n poimen -l app.kubernetes.io/name=poimen-<repo> -w
|
||||
kubectl describe pod -n poimen <pod-name>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
**Effort**: ~10 minutes per repo (once)
|
||||
|
||||
**Benefit**:
|
||||
- Zero-touch deployments
|
||||
- Every commit automatically tested & deployed
|
||||
- Consistent across organization
|
||||
- No manual image pushes ever
|
||||
|
||||
**Best practice**: Use org-level secret if available (1 setup, unlimited repos)
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
1. **poimen-memory**: ✅ Done (this repo)
|
||||
2. **poimen**: Set up workflow + secret
|
||||
3. **poimen-workflows**: Set up workflow + secret
|
||||
4. **Document in**: homelab-poimen-standard.md (org wiki)
|
||||
-147
@@ -1,147 +0,0 @@
|
||||
# CI/CD Setup — Registry Push Configuration
|
||||
|
||||
## One-Time Setup
|
||||
|
||||
The CI/CD pipeline automatically builds and pushes Docker images when you push to `main`.
|
||||
|
||||
### 1. Create or Get Registry Token
|
||||
|
||||
**Option A: Use Organization Token** (Recommended)
|
||||
```bash
|
||||
# Ask Rock for the existing 'rock' organization PAT
|
||||
# It should already have write:package permissions
|
||||
```
|
||||
|
||||
**Option B: Create Personal Token**
|
||||
```bash
|
||||
# In browser: https://git.riotpiao.com/user/settings/tokens
|
||||
# 1. Click "Generate New Token"
|
||||
# 2. Name: "Docker Registry"
|
||||
# 3. Scope: Check `write:package`
|
||||
# 4. Generate and copy the token
|
||||
```
|
||||
|
||||
### 2. Add Repository Secret
|
||||
|
||||
Go to: **https://git.riotpiao.com/rock/poimen-memory/settings/secrets**
|
||||
|
||||
Add secret:
|
||||
- **Name**: `REGISTRY_PAT`
|
||||
- **Value**: `<token-from-step-1>`
|
||||
- **Save**
|
||||
|
||||
### 3. Verify Setup
|
||||
|
||||
```bash
|
||||
# Push a commit (any change will do)
|
||||
cd ~/workplace/Poimen/memory
|
||||
git commit --allow-empty -m "Trigger CI build"
|
||||
git push origin main
|
||||
|
||||
# Check Actions tab
|
||||
# https://git.riotpiao.com/rock/poimen-memory/actions
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## How It Works
|
||||
|
||||
```
|
||||
Push to main
|
||||
↓
|
||||
Forgejo Actions triggered
|
||||
↓
|
||||
Test: cargo test --all
|
||||
↓ (only if tests pass)
|
||||
Build: docker build -t forgejo.riotpiao.com/rock/poimen-memory:latest .
|
||||
↓
|
||||
Push: docker push (using REGISTRY_PAT secret)
|
||||
↓
|
||||
ArgoCD detects new image
|
||||
↓
|
||||
Auto-deploy to poimen namespace
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Check Status
|
||||
|
||||
**Web UI** — See build progress:
|
||||
```
|
||||
https://git.riotpiao.com/rock/poimen-memory/actions
|
||||
```
|
||||
|
||||
**CLI** — Watch deployment:
|
||||
```bash
|
||||
kubectl get application -n argocd poimen-memory-app -w
|
||||
kubectl get pods -n poimen -l app.kubernetes.io/name=poimen-memory -w
|
||||
```
|
||||
|
||||
**Verify Image** — Check registry:
|
||||
```bash
|
||||
docker pull forgejo.riotpiao.com/rock/poimen-memory:latest
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Once Image is Ready
|
||||
|
||||
```bash
|
||||
# Port forward to local
|
||||
kubectl port-forward -n poimen svc/poimen-memory 8080:80 &
|
||||
|
||||
# Test
|
||||
curl http://localhost:8080/health
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Secret Not Found Error
|
||||
- Go to: https://git.riotpiao.com/rock/poimen-memory/settings/secrets
|
||||
- Verify `REGISTRY_PAT` is set
|
||||
|
||||
### Login Failed
|
||||
- Token might be expired or revoked
|
||||
- Create a new token and update the secret
|
||||
|
||||
### Build Failed
|
||||
- Check Actions logs for the error
|
||||
- Usually: tests failed
|
||||
- Fix locally: `cargo test --all`
|
||||
|
||||
### Image Exists But Pods Not Running
|
||||
- Check pod events: `kubectl describe pod -n poimen <pod-name>`
|
||||
- Usually: image pull policy issue or pod crashed
|
||||
- Check logs: `kubectl logs -n poimen deployment/poimen-memory`
|
||||
|
||||
---
|
||||
|
||||
## Apply to Other Repos
|
||||
|
||||
The same setup works for all Poimen repos:
|
||||
|
||||
```bash
|
||||
# For poimen, poimen-workflows, etc:
|
||||
# 1. Create .forgejo/workflows/build.yaml (copy from template below)
|
||||
# 2. Add REGISTRY_PAT secret
|
||||
# 3. Push and watch it deploy
|
||||
```
|
||||
|
||||
**Template**: See `.forgejo/workflows/TEMPLATE.md` in this repo
|
||||
|
||||
---
|
||||
|
||||
## Pattern Overview
|
||||
|
||||
**Based on**: homelab-frontend (proven production pattern)
|
||||
- Uses `REGISTRY_PAT` secret ✓
|
||||
- Docker login + push ✓
|
||||
- Tags: commit SHA + latest ✓
|
||||
- ArgoCD watches tags ✓
|
||||
|
||||
**Consistency**: All Poimen repos use same pattern
|
||||
- Same secret name: `REGISTRY_PAT`
|
||||
- Same workflow structure
|
||||
- Same deployment process
|
||||
@@ -1,387 +0,0 @@
|
||||
# Query Optimization & Hybrid Search Design — Complete
|
||||
|
||||
## What Was Built
|
||||
|
||||
### ✅ 1. Query Optimization Engine (`query_optimizer.rs` - 450 LOC)
|
||||
|
||||
**6-stage pipeline for understanding queries:**
|
||||
|
||||
1. **Normalization** — Lowercase, trim whitespace
|
||||
2. **Tokenization** — Break into words
|
||||
3. **Entity Extraction** — Find years, quoted phrases, tags
|
||||
4. **Characteristic Analysis** — Detect dates, negation, special syntax
|
||||
5. **Question Classification** — Procedural vs Factual vs Troubleshooting, etc
|
||||
6. **Search Strategy Routing** — Choose optimal retrieval method
|
||||
|
||||
**Output:** `QueryContext` + `SearchStrategy` + `Confidence`
|
||||
|
||||
```rust
|
||||
pub enum SearchStrategy {
|
||||
Hybrid, // Both pgvector + OpenSearch (best accuracy)
|
||||
SemanticOnly, // pgvector only (fallback)
|
||||
LexicalOnly, // OpenSearch only (fallback)
|
||||
LexicalFirst, // OpenSearch narrow → pgvector rerank (fastest)
|
||||
}
|
||||
```
|
||||
|
||||
**Key Features:**
|
||||
- ✅ RRF (Reciprocal Rank Fusion) algorithm — no parameter tuning
|
||||
- ✅ Cascading strategy support — multi-stage retrieval
|
||||
- ✅ 15+ unit tests
|
||||
- ✅ Zero external dependencies (pure logic)
|
||||
|
||||
---
|
||||
|
||||
### ✅ 2. Hybrid Query Worker (`hybrid_query_worker.rs` - 380 LOC)
|
||||
|
||||
**Orchestrates parallel retrieval across engines:**
|
||||
|
||||
- **Stage 1**: Route query using QueryOptimizer
|
||||
- **Stage 2**: Generate embedding (LLM)
|
||||
- **Stage 3**: Execute parallel queries
|
||||
- pgvector semantic (top-50)
|
||||
- OpenSearch lexical (top-50) with JWT auth
|
||||
- **Stage 4**: Fuse results using RRF
|
||||
- **Stage 5**: Build rich response with score breakdown
|
||||
|
||||
**Output:** `HybridQueryResponse` with:
|
||||
- Top-10 results
|
||||
- Score breakdown (semantic + lexical components)
|
||||
- Metrics (latency, engine counts, fusion method)
|
||||
- Retrieval engine used
|
||||
|
||||
**Strategies Supported:**
|
||||
- HYBRID: Parallel pgvector + OpenSearch → RRF fusion
|
||||
- CASCADING: OpenSearch narrow (200) → pgvector rerank (10)
|
||||
- SEMANTIC: pgvector only (fallback)
|
||||
- LEXICAL: OpenSearch only (fallback)
|
||||
|
||||
---
|
||||
|
||||
### ✅ 3. Design Documentation (5 comprehensive documents)
|
||||
|
||||
#### **QUERY_OPTIMIZATION_ENGINE.md** (500+ LOC)
|
||||
- **Executive summary** — Why Approach A (Parallel RRF)
|
||||
- **Architecture overview** — Complete data flow
|
||||
- **6-stage pipeline** — Detailed implementation of QueryOptimizer
|
||||
- **Question classification** — Type detection + routing examples
|
||||
- **Search strategy routing** — Decision tree with confidence scores
|
||||
- **RRF algorithm** — Why RRF > Weighted Linear, formula, Rust code
|
||||
- **Response format** — API contract with score breakdown
|
||||
- **Integration path** — How to update /memory/query endpoint
|
||||
- **4-phase implementation plan** — Week 1-4 deliverables
|
||||
- **Testing checklist** — Unit + integration + A/B testing
|
||||
- **Configuration reference** — Env vars + tuning parameters
|
||||
|
||||
#### **HYBRID_SEARCH_DESIGN.md** (760+ LOC)
|
||||
- 5-stage retrieval pipeline (normalize → parallel → normalize → fuse → rank)
|
||||
- Index optimization for pgvector (HNSW, filtering, queries)
|
||||
- Index optimization for OpenSearch (BM25, field boosts, analyzers)
|
||||
- Accuracy metrics (MRR, NDCG@10, Precision@K, Recall@K)
|
||||
- Query routing decision tree
|
||||
- Weight tuning strategy (A/B testing framework)
|
||||
- Indexing pipeline (write side)
|
||||
- Testing strategy with fixtures
|
||||
|
||||
#### **API_REVIEW.md** (400+ LOC)
|
||||
- 10 endpoints reviewed (health, ingest, query, vault-*, etc)
|
||||
- Distinction: Query APIs vs Retrieval APIs
|
||||
- Current implementation gaps
|
||||
- Recommended Phase 1-4 enhancements
|
||||
- Architecture changes needed
|
||||
- Implementation checklist
|
||||
|
||||
#### **IMPLEMENTATION_NOTES.md** (280+ LOC)
|
||||
- Compilation status (non-blocking API mismatches noted)
|
||||
- VectorStore API corrections
|
||||
- OpenSearchClient API fixes
|
||||
- Phase 2 checklist (5-day implementation)
|
||||
- Code diff preview
|
||||
- Design validation matrix
|
||||
|
||||
#### **memory-flow.md** (updated - 833 LOC)
|
||||
- Complete retrieval pipeline diagram (5 stages)
|
||||
- Query routing decision tree
|
||||
- Index optimization details
|
||||
- Pod infrastructure (now 8 core pods)
|
||||
- Deployment checklist reorganized
|
||||
|
||||
---
|
||||
|
||||
## Architecture Decision: Approach A (Parallel RRF)
|
||||
|
||||
### Why This Approach?
|
||||
|
||||
| Criterion | Score | Reasoning |
|
||||
|-----------|-------|-----------|
|
||||
| **Accuracy** | ⭐⭐⭐⭐⭐ | Semantic + Lexical covers all cases |
|
||||
| **Fault Tolerance** | ⭐⭐⭐⭐⭐ | Fallback to semantic if OpenSearch down |
|
||||
| **No False Negatives** | ⭐⭐⭐⭐⭐ | Semantic catches synonyms lexical misses |
|
||||
| **Debugging** | ⭐⭐⭐⭐⭐ | Clear score breakdown for transparency |
|
||||
| **Decoupled** | ⭐⭐⭐⭐⭐ | Embedding model changes don't break system |
|
||||
| **Latency** | ⭐⭐⭐ | 150-250ms (parallel) vs 60-100ms (single engine) |
|
||||
| **Complexity** | ⭐⭐⭐ | Moderate RRF logic + parallel orchestration |
|
||||
|
||||
**Mission-critical for agent reasoning:** Agents make decisions based on retrieved context. Missing docs = wrong decisions.
|
||||
|
||||
---
|
||||
|
||||
## Key Components
|
||||
|
||||
### 1. QueryOptimizer (Pure Logic)
|
||||
|
||||
```rust
|
||||
optimizer.optimize_query("How do I fix kubernetes port 8080?")
|
||||
→ QueryContext {
|
||||
raw_query: "How do I fix kubernetes port 8080?",
|
||||
normalized: "how do i fix kubernetes port 8080?",
|
||||
tokens: ["how", "do", "i", "fix", "kubernetes", "port", "8080"],
|
||||
entities: {},
|
||||
token_count: 7,
|
||||
has_special_syntax: false,
|
||||
has_date_filters: false,
|
||||
has_negation: false,
|
||||
question_type: Procedural,
|
||||
search_strategy: Hybrid,
|
||||
confidence: 0.95,
|
||||
}
|
||||
```
|
||||
|
||||
### 2. HybridQueryWorker (Parallel Orchestration)
|
||||
|
||||
```rust
|
||||
worker.query("poimen", "How do I fix kubernetes port 8080?", 10, &jwt)
|
||||
→ HybridQueryResponse {
|
||||
query: "How do I fix kubernetes port 8080?",
|
||||
project: "poimen",
|
||||
search_strategy: "Hybrid",
|
||||
strategy_confidence: 0.95,
|
||||
results: [
|
||||
{
|
||||
id: "chunk-123",
|
||||
rank: 1,
|
||||
final_score: 0.0328,
|
||||
semantic_score: 0.95,
|
||||
lexical_score: 8.5,
|
||||
fusion_method: "rrf",
|
||||
text: "kubectl port-forward service port:8080...",
|
||||
source: "runbooks/kubernetes/networking.md",
|
||||
score_breakdown: {
|
||||
semantic_rank: 1,
|
||||
lexical_rank: 1,
|
||||
rrf_components: {...}
|
||||
}
|
||||
},
|
||||
...
|
||||
],
|
||||
metrics: {
|
||||
total_time_ms: 245,
|
||||
semantic_time_ms: 120,
|
||||
lexical_time_ms: 118,
|
||||
fusion_time_ms: 7,
|
||||
semantic_results_count: 50,
|
||||
lexical_results_count: 50,
|
||||
final_results_count: 10
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
### 3. RRF Algorithm (No Parameter Tuning)
|
||||
|
||||
```rust
|
||||
// Input: two ranked lists
|
||||
semantic: [(doc1, 0.95), (doc2, 0.88), (doc3, 0.82)]
|
||||
lexical: [(doc1, 8.5), (doc4, 7.2), (doc2, 6.8)]
|
||||
|
||||
// RRF formula: 1 / (k + rank) where k=60
|
||||
doc1: 1/(60+1) + 1/(60+1) = 0.0328 ← Top result
|
||||
doc2: 1/(60+2) + 1/(60+3) = 0.0317
|
||||
doc4: 1/(60+2) = 0.0159
|
||||
doc3: 1/(60+3) = 0.0158
|
||||
|
||||
// Output: [doc1, doc2, doc4, doc3] (merged + ranked)
|
||||
```
|
||||
|
||||
**Why RRF?**
|
||||
- ✅ No parameter tuning (k=60 is academic standard)
|
||||
- ✅ Robust to score distribution differences
|
||||
- ✅ Works if embedding model changes
|
||||
- ✅ Academic consensus for multi-engine fusion
|
||||
- ❌ Loses score magnitudes (but transparency provided)
|
||||
|
||||
---
|
||||
|
||||
## Implementation Phases
|
||||
|
||||
### Phase 1: ✅ COMPLETE (This Session)
|
||||
|
||||
**Deliverables:**
|
||||
- ✅ QueryOptimizer (450 LOC, 15+ tests)
|
||||
- ✅ HybridQueryWorker (380 LOC, stub with API fixes noted)
|
||||
- ✅ RRF Fusion algorithm (no parameter tuning)
|
||||
- ✅ Complete design documentation (2000+ LOC)
|
||||
- ✅ Implementation notes + API corrections
|
||||
|
||||
**Time: 4 hours of design + coding**
|
||||
|
||||
### Phase 2: TODO (Week 2, 3-4 days)
|
||||
|
||||
**Tasks:**
|
||||
- [ ] Fix VectorStore API calls (15 min)
|
||||
- [ ] Make OpenSearchClient::lexical_search public (5 min)
|
||||
- [ ] Integrate HybridQueryWorker into /memory/query handler
|
||||
- [ ] Add fallback strategy (hybrid → semantic → error)
|
||||
- [ ] Update response format (include metrics + score breakdown)
|
||||
- [ ] Write 10+ integration tests
|
||||
- [ ] Measure latency (hybrid vs semantic vs cascading)
|
||||
|
||||
### Phase 3: TODO (Week 3, 2-3 days)
|
||||
|
||||
**Performance Optimization:**
|
||||
- [ ] Benchmark all search strategies
|
||||
- [ ] Optimize pgvector index (HNSW tuning)
|
||||
- [ ] Optimize OpenSearch queries (field boosts)
|
||||
- [ ] Add query result caching (1hr TTL)
|
||||
- [ ] Profile parallel execution
|
||||
|
||||
### Phase 4: TODO (Week 4, 2-3 days)
|
||||
|
||||
**Testing & Validation:**
|
||||
- [ ] Create test fixture dataset (50+ queries with ground truth)
|
||||
- [ ] Measure NDCG@10, MRR, Precision@K
|
||||
- [ ] A/B test: Hybrid vs Semantic-only
|
||||
- [ ] A/B test: RRF vs Weighted Linear (0.6/0.4)
|
||||
- [ ] Experiment with different question types
|
||||
- [ ] Finalize configuration (env vars + defaults)
|
||||
|
||||
---
|
||||
|
||||
## Files & Statistics
|
||||
|
||||
### Code Files (830 LOC)
|
||||
|
||||
```
|
||||
crates/mem-cli/src/
|
||||
├─ query_optimizer.rs (450 LOC, 15 tests)
|
||||
│ ├─ QueryOptimizer (6-stage pipeline)
|
||||
│ ├─ QueryContext (data structure)
|
||||
│ ├─ QuestionType enum (6 types)
|
||||
│ ├─ SearchStrategy enum (4 strategies)
|
||||
│ ├─ RRFConfig (tuning parameters)
|
||||
│ └─ RRFFusion (RRF algorithm)
|
||||
│
|
||||
├─ hybrid_query_worker.rs (380 LOC, stub)
|
||||
│ ├─ HybridQueryWorker (orchestrator)
|
||||
│ ├─ retrieve_hybrid() (parallel)
|
||||
│ ├─ retrieve_cascading() (2-stage)
|
||||
│ ├─ fuse_results() (RRF)
|
||||
│ └─ HybridQueryResponse (response type)
|
||||
│
|
||||
└─ lib.rs
|
||||
├─ pub mod query_optimizer
|
||||
└─ pub mod hybrid_query_worker
|
||||
```
|
||||
|
||||
### Design Documents (2100+ LOC)
|
||||
|
||||
```
|
||||
docs/
|
||||
├─ QUERY_OPTIMIZATION_ENGINE.md (500+ LOC)
|
||||
│ ├─ Executive Summary
|
||||
│ ├─ 6-Stage Pipeline Detailed
|
||||
│ ├─ Question Classification
|
||||
│ ├─ RRF Algorithm Explained
|
||||
│ ├─ 4-Phase Implementation Plan
|
||||
│ └─ Testing Checklist
|
||||
│
|
||||
├─ HYBRID_SEARCH_DESIGN.md (760+ LOC)
|
||||
│ ├─ 5-Stage Retrieval Pipeline
|
||||
│ ├─ Index Optimization (pgvector + OpenSearch)
|
||||
│ ├─ Accuracy Metrics
|
||||
│ └─ Weight Tuning Strategy
|
||||
│
|
||||
├─ API_REVIEW.md (400+ LOC)
|
||||
│ ├─ 10 Endpoints Reviewed
|
||||
│ ├─ Query vs Retrieval APIs
|
||||
│ ├─ Current Gaps
|
||||
│ └─ Phase 1-4 Enhancements
|
||||
│
|
||||
├─ IMPLEMENTATION_NOTES.md (280+ LOC)
|
||||
│ ├─ Compilation Status
|
||||
│ ├─ API Corrections
|
||||
│ └─ Phase 2 Checklist
|
||||
│
|
||||
└─ memory-flow.md (updated, 833 LOC)
|
||||
├─ 5-Stage Hybrid Retrieval Pipeline
|
||||
├─ Query Routing Decision Tree
|
||||
└─ Pod Infrastructure (8 core)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Next Steps
|
||||
|
||||
### Immediate (End of Session)
|
||||
|
||||
✅ Review & approve design
|
||||
✅ Commit code to repository
|
||||
✅ Document in CLAUDE.md
|
||||
|
||||
### Week 2 (Phase 2 Implementation)
|
||||
|
||||
- [ ] Fix compilation errors (API mismatches)
|
||||
- [ ] Integrate into /memory/query handler
|
||||
- [ ] Add hybrid search tests
|
||||
- [ ] Deploy to staging
|
||||
|
||||
### Metrics to Track
|
||||
|
||||
| Metric | Target | Notes |
|
||||
|--------|--------|-------|
|
||||
| Hybrid latency | 150-250ms | Parallel pgvector + OpenSearch |
|
||||
| Cascading latency | 100-180ms | Lexical narrow → semantic rerank |
|
||||
| NDCG@10 | ≥0.85 | Ranking quality |
|
||||
| MRR | ≥0.8 | First correct result position |
|
||||
| Precision@5 | ≥0.8 | Correct results in top-5 |
|
||||
| Zero false negatives | 100% | Semantic catches synonyms |
|
||||
|
||||
---
|
||||
|
||||
## Key Decisions
|
||||
|
||||
✅ **Approach A: Parallel RRF** — Highest accuracy, fault tolerant
|
||||
✅ **RRF over Weighted Linear** — No parameter tuning, robust
|
||||
✅ **6-stage QueryOptimizer** — Understand query before retrieval
|
||||
✅ **4 Search Strategies** — Hybrid/Semantic/Lexical/Cascading
|
||||
✅ **JWT forwarding to OpenSearch** — Consistent auth
|
||||
✅ **Fallback strategy** — Hybrid → Semantic → Error
|
||||
✅ **Score breakdown in API** — Transparency + debugging
|
||||
|
||||
---
|
||||
|
||||
## Success Criteria (Phase 1)
|
||||
|
||||
✅ Design document complete and reviewed
|
||||
✅ Code compiles (after API fixes)
|
||||
✅ 15+ unit tests passing
|
||||
✅ Architecture decisions documented
|
||||
✅ Phase 2 implementation plan clear
|
||||
✅ No architectural changes needed
|
||||
|
||||
**All criteria met.** 🎉
|
||||
|
||||
---
|
||||
|
||||
## Summary
|
||||
|
||||
We've designed and implemented a **production-grade Query Optimization Engine** for Poimen Memory:
|
||||
|
||||
1. **QueryOptimizer** — 6-stage pipeline that understands queries
|
||||
2. **HybridQueryWorker** — Parallel retrieval + RRF fusion
|
||||
3. **4 Search Strategies** — Optimize for different query types
|
||||
4. **Comprehensive Documentation** — 2100+ LOC covering architecture to testing
|
||||
|
||||
**Approach:** Parallel RRF (Approach A) — highest accuracy for mission-critical agent reasoning.
|
||||
|
||||
**Status:** Ready for Phase 2 implementation (3-4 day integration + testing).
|
||||
|
||||
@@ -1,75 +0,0 @@
|
||||
# Quick Start — One-Time CI/CD Setup (3 minutes)
|
||||
|
||||
## 🚀 Setup
|
||||
|
||||
### 1. Add Secret to Repository
|
||||
```bash
|
||||
# Go to: https://git.riotpiao.com/rock/poimen-memory/settings/secrets
|
||||
# Add: Name=REGISTRY_PAT, Value=bdf6a1d2317c28a332447083c61bb463d24defb7
|
||||
# Save
|
||||
|
||||
# (This token is encrypted & managed in homelab via SOPS)
|
||||
# See: CI-SETUP-WITH-KSOPS.md for details
|
||||
```
|
||||
|
||||
### 3. Push to Trigger Build
|
||||
```bash
|
||||
cd ~/workplace/Poimen/memory
|
||||
git commit --allow-empty -m "Trigger CI"
|
||||
git push
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## ✨ That's It!
|
||||
|
||||
After these 3 steps, every push automatically:
|
||||
- ✅ Runs all tests
|
||||
- ✅ Builds Docker image
|
||||
- ✅ Pushes to `forgejo.riotpiao.com/rock/poimen-memory:latest`
|
||||
- ✅ ArgoCD deploys to K8s
|
||||
|
||||
---
|
||||
|
||||
## 📊 Monitor
|
||||
|
||||
```bash
|
||||
# Watch build
|
||||
https://git.riotpiao.com/rock/poimen-memory/actions
|
||||
|
||||
# Watch deployment
|
||||
kubectl get pods -n poimen -l app.kubernetes.io/name=poimen-memory -w
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🧪 Test When Ready
|
||||
|
||||
```bash
|
||||
# Port forward
|
||||
kubectl port-forward -n poimen svc/poimen-memory 8080:80 &
|
||||
|
||||
# Health check
|
||||
curl http://localhost:8080/health
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 🔄 Apply to Other Repos
|
||||
|
||||
Same pattern for `poimen`, `poimen-workflows`, etc:
|
||||
|
||||
1. Add `REGISTRY_PAT` secret
|
||||
2. Copy `.forgejo/workflows/build.yaml` from this repo
|
||||
3. Push
|
||||
|
||||
See `CI-SETUP.md` for details.
|
||||
|
||||
---
|
||||
|
||||
## Pattern Details
|
||||
|
||||
- **Based on**: homelab-frontend (proven pattern)
|
||||
- **Runner**: docker:27-cli (supports buildx)
|
||||
- **Tags**: commit SHA + "latest"
|
||||
- **No manual steps**: Fully automated
|
||||
@@ -217,8 +217,8 @@ impl GatewayQueueAdapter {
|
||||
}
|
||||
}
|
||||
|
||||
fn queue_name(&self, project: &str) -> String {
|
||||
format!("{}-{}", self.default_queue_prefix, project)
|
||||
fn queue_name(&self, _project: &str) -> String {
|
||||
self.default_queue_prefix.clone()
|
||||
}
|
||||
}
|
||||
|
||||
@@ -505,7 +505,7 @@ mod tests {
|
||||
"test-token".to_string(),
|
||||
);
|
||||
|
||||
assert_eq!(adapter.queue_name("myproject"), "poimen-chunks-myproject");
|
||||
assert_eq!(adapter.queue_name("myproject"), "poimen-chunks");
|
||||
}
|
||||
|
||||
#[test]
|
||||
|
||||
@@ -0,0 +1,32 @@
|
||||
apiVersion: kmsvc.io/v1
|
||||
kind: Queue
|
||||
metadata:
|
||||
name: poimen-chunks
|
||||
namespace: sqs
|
||||
spec:
|
||||
fifoQueue: false
|
||||
visibilityTimeoutSeconds: 30
|
||||
messageRetentionPeriodSeconds: 86400
|
||||
maxReceiveCount: 3
|
||||
deadLetterTargetQueue: poimen-chunks-dlq
|
||||
delaySeconds: 0
|
||||
partitionsPerShard: 2
|
||||
minShards: 1
|
||||
maxShards: 4
|
||||
shardSplitThresholdBytesPerSec: 1048576
|
||||
shardSplitCooldownSeconds: 300
|
||||
---
|
||||
apiVersion: kmsvc.io/v1
|
||||
kind: Queue
|
||||
metadata:
|
||||
name: poimen-chunks-dlq
|
||||
namespace: sqs
|
||||
spec:
|
||||
fifoQueue: false
|
||||
isDLQ: true
|
||||
visibilityTimeoutSeconds: 30
|
||||
messageRetentionPeriodSeconds: 604800
|
||||
maxReceiveCount: 3
|
||||
partitionsPerShard: 1
|
||||
minShards: 1
|
||||
maxShards: 1
|
||||
Reference in New Issue
Block a user