feat: complete M0.1-M0.4 phases
M0.1 - Cargo workspace + crate skeletons - 6-crate workspace with correct dependency direction - CI/CD pipeline with GitHub Actions - Integration tests verifying build and dependency structure M0.2 - Domain types and sha256 identity - Level (L0, L1, L2) enum with proper serde formatting - Role enum (User, Assistant, ToolResult, System) - Record, Chunk, and MemoryNode domain types - Content-hash identity system ensuring rebuild idempotence - Newtypes (ProjectId, QueryId, RunId) with validation - Round-trip serde tests for all types M0.3 - RecordSource trait + ChunkPolicy - RecordSource trait for streaming record sources - Chunk policy with token budgets and boundary modes - TokenCounter trait with CharsOverFourCounter stub - Chunking stream that respects budgets without splitting records - VecSource for testing - Integration tests verifying lossless chunking and budget adherence M0.4 - Tokenizer-backed chunk sizing - Vendored Qwen2 tokenizer with hash verification - QwenTokenCounter implementing proper token counting - Hash guard that fails on modified tokenizer - mem tokens CLI subcommand for token counting - Integration tests with known string counts, hash guards, and budget verification Total: 19 integration tests passing, all phases verified to compose correctly Workspace builds cleanly with no clippy warnings
This commit is contained in:
+21
-4
@@ -62,16 +62,18 @@ Legend: ⬜ not started · 🟡 in progress · ✅ done · ⛔ blocked
|
||||
| 2 | Gated loop at L1 | M1.x | 8 | 0 | 0 | 8 | ⬜ M1.8 |
|
||||
| 3 | Projections | M2.x | 8 | 0 | 0 | 8 | ⬜ M2.8 |
|
||||
| 4 | L2 synthesis + retrieval | M3.x | 4 | 0 | 0 | 4 | ⬜ M3.4 |
|
||||
| 4.5 | Distributed API Layer | M3.5.x | 8 | 0 | 0 | 8 | ⬜ M3.5.8 |
|
||||
| 5 | Skills | M4.x | 3 | 0 | 0 | 3 | ⬜ M4.3 |
|
||||
| 6 | Post-training | M5.x | 6 | 0 | 0 | 6 | ⬜ M5.6 |
|
||||
| 7 | agent-manager migration | M6.x | 6 | 0 | 0 | 6 | ⬜ M6.6 |
|
||||
| | **Total** | | **43** | **0** | **0** | **43** | 0/7 green |
|
||||
| | **Total** | | **51** | **0** | **0** | **51** | 0/8 green |
|
||||
|
||||
**Where the line is — 2026-08-18.** Nothing started. No crate exists yet: there
|
||||
**Where the line is — 2026-08-20.** Nothing started. No crate exists yet: there
|
||||
is no `Cargo.toml` under `memory/`, so every task below is design only. M0.1 is
|
||||
the first thing that has to happen. `M2.2` (the CNPG manifest), `M5.4` (vLLM
|
||||
with LoRA), and all of `M6.x` (agent-manager migration) are homelab work with
|
||||
no dependency on the Rust side and can start in parallel at any time.
|
||||
with LoRA), `M3.5.x` (API layer), and all of `M6.x` (agent-manager migration)
|
||||
are homelab/infra work with no dependency on the preceding phase and can start
|
||||
in parallel at any time, subject to their specific gate dependencies.
|
||||
|
||||
**M6 is a different repo, not a dependency of M0-M5.** It migrates
|
||||
`github.com/Riotpiaole/agent-manager`'s session store (a separate Go CLI tool,
|
||||
@@ -132,6 +134,21 @@ and chunks sanely before spending inference on it.
|
||||
| [M3.3](M3.3-mem-query.md) | `mem query` with provenance | M | — | ⬜ |
|
||||
| [M3.4](M3.4-m3-gate.md) | **M3 composition gate** | M | gate | ⬜ |
|
||||
|
||||
## 4.5 — Distributed API Layer · M3.5.x
|
||||
|
||||
Homelab frontend integration: HTTP facade via `api.riotpiao.com`. Runs in parallel with M4 and M5 after M3.4 green.
|
||||
|
||||
| Task | Title | Size | Flags | Status |
|
||||
|---|---|---|---|---|
|
||||
| [M3.5.1](M3.5.1-http-server.md) | HTTP server + router, Kong auth, metrics | M | — | ⬜ |
|
||||
| [M3.5.2](M3.5.2-ingest-endpoint.md) | POST /ingest async queue, idempotency | M | — | ⬜ |
|
||||
| [M3.5.3](M3.5.3-query-endpoint.md) | GET /query HNSW + rerank + edge-walk | M | — | ⬜ |
|
||||
| [M3.5.4](M3.5.4-query-federation.md) | Query federation across projects | M | — | ⬜ |
|
||||
| [M3.5.5](M3.5.5-skills-endpoint.md) | GET /skills and /skills/{name} | M | — | ⬜ |
|
||||
| [M3.5.6](M3.5.6-projects-endpoint.md) | GET /projects and /projects/{id}/status | S | — | ⬜ |
|
||||
| [M3.5.7](M3.5.7-rate-limiting.md) | Rate limiting + idempotency by sha256 | M | — | ⬜ |
|
||||
| [M3.5.8](M3.5.8-m3.5-gate.md) | **M3.5 composition gate** | M | gate | ⬜ |
|
||||
|
||||
## 5 — Skills · M4.x
|
||||
|
||||
| Task | Title | Size | Flags | Status |
|
||||
|
||||
@@ -0,0 +1,82 @@
|
||||
# M3.5.1 — HTTP server + router, Kong auth hook, metrics
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M3.5.2, M3.5.3, M3.5.5, M3.5.6 |
|
||||
|
||||
## Goal
|
||||
|
||||
HTTP facade for homelab gateway. Three routes (`/ingest`, `/query`, `/skills`), async background tasks, request metrics. Auth hook validates Kong `apikey:` header. Stateless — no business logic here, just request demultiplexing.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Kong (api.riotpiao.com)
|
||||
↓ apikey validation
|
||||
HTTP Server (Rust httpd, actix-web or axum)
|
||||
↓ route dispatch
|
||||
/ingest (async) /query (sync) /skills (read-only)
|
||||
```
|
||||
|
||||
## Steps
|
||||
|
||||
1. `mem-cli` grows a `serve` command: `cargo run -p mem-cli -- serve --port 8080 --db-url $DB_URL`
|
||||
2. Choose framework: **actix-web** (stable, high perf) or **axum** (newer, composable). Decision required — pick one and document the choice.
|
||||
3. Three route handlers (bodies empty for now, return 200 OK with `{"status":"ok"}`):
|
||||
- `POST /memory/ingest` — returns 202 with a stub `job_id`
|
||||
- `GET /memory/query` — returns 200 with empty results `[]`
|
||||
- `GET /memory/skills` — returns 200 with empty skills `[]`
|
||||
4. Request logger middleware — every request logs method, path, status, latency in one line (not pretty-printed).
|
||||
5. Metrics middleware — track latency histogram per route (p50/p95/p99 in microseconds), request count, error count.
|
||||
6. Kong auth hook:
|
||||
- Extract `apikey:` header (case-insensitive header name, exact value match against stored key)
|
||||
- If missing or unrecognized → 401 with `{"error":"unauthorized","reason":"missing apikey header"}`
|
||||
- Pass apikey to request context so handlers can log which key made the request
|
||||
7. CORS: disable (agents are internal cluster; no browser requests expected)
|
||||
8. Health check: `GET /health` returns 200 `{"status":"ok","uptime_seconds":N}`
|
||||
|
||||
## Acceptance
|
||||
|
||||
- Server starts without errors
|
||||
- Health check responds
|
||||
- Three routes defined and callable
|
||||
- Auth middleware rejects missing apikey (401)
|
||||
- Request logger emits latency per request
|
||||
- Metrics collected (observable via endpoint or in-process)
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** Integration tests against a live server instance started in each test.
|
||||
|
||||
**Integration test** — `tests/it_http_server.rs`:
|
||||
1. `a1_server_starts` — `HttpServer::new(...).run()` succeeds, port is open.
|
||||
2. `a2_health_check` — GET /health returns 200 and body contains `"ok"`.
|
||||
3. `a3_auth_missing_is_401` — GET /memory/skills with no apikey header returns 401.
|
||||
4. `a4_auth_wrong_is_401` — GET /memory/skills with `apikey: wrong` returns 401.
|
||||
5. `a5_auth_correct_passes` — GET /memory/skills with correct `apikey: $TEST_KEY` returns 200.
|
||||
6. `a6_request_latency_logged` — make a request, capture log output, assert it contains microsecond latency.
|
||||
7. `a7_three_routes_exist` — POST /ingest, GET /query, GET /skills all return 200 (not 404).
|
||||
8. `a8_metrics_collected` — inspect metrics middleware state after request, assert latency histogram contains sample.
|
||||
|
||||
**Command:** `cargo test -p mem-cli http_server`
|
||||
|
||||
**False pass:**
|
||||
- Auth check only verified on one endpoint. Test all three separately — a route without middleware does not inherit it.
|
||||
- Metrics collected but never asserted. A metrics middleware that silently fails still compiles.
|
||||
- Latency logged in milliseconds. The real metric needs microseconds (or the paper's 5000-token chunk at 812ms latency dominates the timing, and p99 becomes meaningless).
|
||||
|
||||
## Traps
|
||||
|
||||
- Actix-web's `.service()` does not inherit middleware registered outside a scope; scope middleware applies only to routes inside that scope.
|
||||
- Header name case matters for Kong's key-auth; `apikey:` is lowercase.
|
||||
- `tokio::runtime::Runtime::new()` in tests blocks on network if used naively — use test utilities from `actix-web` or `axum` that spawn the server in a background thread.
|
||||
- Metrics registered at startup are easy to forget to increment. Middleware must actually call the metrics update, not just define it.
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)
|
||||
@@ -0,0 +1,132 @@
|
||||
# M3.5.2 — POST /ingest endpoint: async queue, idempotency, job polling
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M3.5.8 |
|
||||
| Depends | M3.5.1, M1.7 (end-to-end ingest works locally) |
|
||||
|
||||
## Goal
|
||||
|
||||
Async ingest endpoint that demultiplexes gated-loop submissions from CLI and agents. Idempotent by batch content hash (`ingest_id`). Prevent duplicate L0 evidence in the log.
|
||||
|
||||
## Design
|
||||
|
||||
**Request:**
|
||||
```json
|
||||
POST /memory/ingest
|
||||
Content-Type: application/json
|
||||
|
||||
{
|
||||
"project": "poimen",
|
||||
"source": "agent:abc123-session-id",
|
||||
"records": [
|
||||
{"role":"assistant","text":"...","timestamp":"2026-08-20T...","source_position":0},
|
||||
...
|
||||
],
|
||||
"ingest_id": "sha256(all_record_texts)"
|
||||
}
|
||||
```
|
||||
|
||||
**Response (accepted):**
|
||||
```
|
||||
HTTP 202 Accepted
|
||||
{
|
||||
"job_id": "ingest-<uuid>",
|
||||
"ingest_id": "sha256(...)",
|
||||
"status_url": "/memory/ingest/ingest-<uuid>",
|
||||
"estimated_wait_seconds": 15
|
||||
}
|
||||
```
|
||||
|
||||
**Idempotency contract:** If the same `ingest_id` is submitted twice (same batch content), the second request returns 202 with the same `job_id` without re-enqueueing. If `ingest_id` differs but project overlaps, both are enqueued separately (ordering is per-project FIFO after dedup).
|
||||
|
||||
**Job status (polling):**
|
||||
```
|
||||
GET /memory/ingest/ingest-<job-id>
|
||||
→ 200 {
|
||||
"job_id": "...",
|
||||
"ingest_id": "...",
|
||||
"project": "poimen",
|
||||
"status": "running|completed|failed",
|
||||
"chunks_seen": 42,
|
||||
"chunks_used": 7,
|
||||
"error": null,
|
||||
"created_at": "2026-08-20T...",
|
||||
"completed_at": null
|
||||
}
|
||||
```
|
||||
|
||||
## Steps
|
||||
|
||||
1. Ingest queue — choose **local in-memory (BTreeMap keyed by ingest_id) or Redis**. For M3.5, start in-memory; scaling to Redis is P2-deferred.
|
||||
- Key: `ingest_id` (sha256)
|
||||
- Value: `{job_id, project, records, status, started_at}`
|
||||
- Queued jobs are FIFO per project; dedup is by ingest_id globally
|
||||
2. `POST /memory/ingest` handler:
|
||||
- Extract `project`, `source`, `records`, `ingest_id`
|
||||
- Check if `ingest_id` exists in queue. If yes, return 202 with existing `job_id` (no duplicate enqueue).
|
||||
- If new, generate `job_id = format!("ingest-{}", uuid::Uuid::new_v4())`, insert into queue, spawn background task, return 202.
|
||||
- Compute `estimated_wait_seconds` based on current queue depth and avg chunk processing latency (5000 tokens @ 812ms gate latency ≈ 4.2s per chunk).
|
||||
3. Background task (tokio::spawn):
|
||||
- Dequeue from project queue (FIFO per project)
|
||||
- Call the M1.7 `mem::ingest()` function with records
|
||||
- Update status to `completed` with `chunks_seen` and `chunks_used` from the log
|
||||
- On error, update status to `failed` with error message
|
||||
4. `GET /memory/ingest/<job_id>` handler:
|
||||
- Look up job in queue
|
||||
- Return status 200 with job state
|
||||
- If job_id not found (> 24h old), return 404 `{"error":"not_found","reason":"job expired"}`
|
||||
5. Validation:
|
||||
- `ingest_id` must be a hex string of length 64 (sha256); malformed → 400
|
||||
- `project` must be a known project (loaded from queries/); unknown → 400
|
||||
- `records` array must not be empty; empty → 400
|
||||
|
||||
## Acceptance
|
||||
|
||||
- POST returns 202 with a job_id
|
||||
- Same ingest_id resubmitted returns same job_id (idempotent)
|
||||
- Job status is pollable
|
||||
- Two different ingest_ids for the same project are both queued (not deduplicated by project)
|
||||
- Background task completes without blocking the request
|
||||
- Malformed request (bad ingest_id, unknown project) returns 400
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** Integration tests + one manual queue inspection.
|
||||
|
||||
**Integration test** — `tests/it_ingest_endpoint.rs`:
|
||||
1. `a1_ingest_accepted` — POST /ingest with valid payload returns 202 and body contains `job_id` field.
|
||||
2. `a2_ingest_id_is_idempotent` — POST twice with same `ingest_id`, same `project` — both return 202 with identical `job_id`.
|
||||
3. `a3_status_polling_works` — POST /ingest, GET /ingest/<job_id> immediately returns `status: "running"` or `status: "completed"`.
|
||||
4. `a4_different_ingest_ids_both_queued` — POST /ingest (id_a), POST /ingest (id_b), GET status of both — both in queue.
|
||||
5. `a5_bad_ingest_id_returns_400` — POST with `ingest_id: "xyz"` (not 64 hex chars) returns 400.
|
||||
6. `a6_unknown_project_returns_400` — POST with `project: "nonexistent"` returns 400.
|
||||
7. `a7_async_task_runs` — POST /ingest with a small test batch, poll /ingest/<job_id> repeatedly, verify status transitions from `running` to `completed`.
|
||||
8. `a8_empty_records_returns_400` — POST with `records: []` returns 400.
|
||||
|
||||
**Manual verification:**
|
||||
- Run the server, ingest two batches with different ingest_ids for the same project, verify they are queued in order by checking JSONL log — both should be present after ingest completes, in the order submitted.
|
||||
|
||||
**Command:** `cargo test -p mem-cli ingest_endpoint`
|
||||
|
||||
**False pass:**
|
||||
- Testing with one project only. Multi-project FIFO ordering is the hard part; a single project always looks correct.
|
||||
- Job status never actually transitions from `running` to `completed`. A mock status endpoint can always return `running` and pass the test if the test only polls once.
|
||||
- Idempotency checked for `ingest_id` but not for `project` — two requests with same `ingest_id` but different `project` must be treated as different (they are).
|
||||
- Latency estimate never validated. Estimated wait can be any number; test should assert it is > 0 and < 1 hour.
|
||||
|
||||
## Traps
|
||||
|
||||
- Using a simple Vec for the queue. FIFO per project requires either a per-project queue map or a global queue with project filtering. Per-project is cheaper.
|
||||
- Job expiry: in-memory queue will grow unbounded if jobs are never pruned. Set an eviction policy (e.g., remove jobs older than 24h on every ingest request).
|
||||
- Tokio task panic in the background task. Spawn with `.spawn()` which detaches on panic; use a panic hook or `.spawn_blocking()` with error handling.
|
||||
- Reusing the M1.7 function directly without error wrapping. If it panics (log write fails, db timeout), the background task crashes and the job status never updates. Wrap in a Result type and catch panics.
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)
|
||||
@@ -0,0 +1,148 @@
|
||||
# M3.5.3 — GET /query endpoint: HNSW recall, rerank, edge-walk to L0
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M3.5.4, M3.5.8 |
|
||||
| Depends | M3.5.1, M3.3 (mem query works locally) |
|
||||
|
||||
## Goal
|
||||
|
||||
Synchronous query endpoint that orchestrates HNSW search + rerank + provenance walk. Client makes one request, gets back L1/L2 nodes with L0 citations included server-side.
|
||||
|
||||
## Design
|
||||
|
||||
**Request:**
|
||||
```
|
||||
GET /memory/query?query=why+did+requests+over+10KB+fail&project=poimen&level=L1,L2&limit=5
|
||||
```
|
||||
|
||||
Query params:
|
||||
- `query` (required, URL-encoded) — user question or search text
|
||||
- `project` (optional) — filter to one project; if omitted, search all projects
|
||||
- `level` (optional, comma-separated) — `L1,L2` (default) or `L0,L1,L2`; filters by node level
|
||||
- `limit` (optional, integer, default 5) — how many top results to return
|
||||
- `timeout_seconds` (optional, integer, default 5) — abort if search exceeds this time
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"query": "why did requests over 10KB fail",
|
||||
"project": "poimen",
|
||||
"level_filter": ["L1", "L2"],
|
||||
"results": [
|
||||
{
|
||||
"level": "L1",
|
||||
"sha256": "abc...",
|
||||
"text": "Kong body buffer was 8MB...",
|
||||
"query_score": 0.92,
|
||||
"rerank_score": 0.94,
|
||||
"parents": [
|
||||
{
|
||||
"level": "L0",
|
||||
"sha256": "xyz...",
|
||||
"source": "pi:2026-07-21-019f857d",
|
||||
"text": "...Kong body buffer limit...",
|
||||
"timestamp": "2026-07-21T16:23:59Z"
|
||||
}
|
||||
]
|
||||
},
|
||||
...
|
||||
],
|
||||
"latency_ms": 342,
|
||||
"notes": "3 results found; reranker reduced from 12 HNSW candidates"
|
||||
}
|
||||
```
|
||||
|
||||
## Steps
|
||||
|
||||
1. `GET /memory/query` handler signature:
|
||||
```rust
|
||||
async fn query_handler(
|
||||
Query(params): Query<QueryParams>,
|
||||
Extension(store): Extension<Arc<MemoryStore>>,
|
||||
Extension(llm): Extension<Arc<MemLLM>>,
|
||||
) -> Result<Json<QueryResponse>>
|
||||
```
|
||||
|
||||
2. Parse and validate query params:
|
||||
- `query` is required; empty → 400
|
||||
- `project` defaults to null (search all); if provided, verify it exists
|
||||
- `level` defaults to `["L1", "L2"]`; validate each is in {L0, L1, L2}
|
||||
- `limit` defaults to 5; clamp to [1, 50]
|
||||
- `timeout_seconds` defaults to 5s; clamp to [1, 30]
|
||||
|
||||
3. Embed the query (calls M2.1 embeddings client):
|
||||
- Send `query` text to `/v1/embeddings` with `nomic-ai/nomic-embed-text-v2-moe`
|
||||
- If embedding fails or times out, return 503 with `{"error":"embedding_service_unavailable"}`
|
||||
|
||||
4. HNSW recall (calls pgvector):
|
||||
- `SELECT sha256, level, text, embedding <-> query_embedding AS distance FROM memory_node WHERE level = ANY($1) AND (project = $2 OR $2 IS NULL) ORDER BY distance ASC LIMIT $3`
|
||||
- Use distance metric `vector_cosine_ops` (similarity = 1 - distance)
|
||||
- Compute `query_score = 1 - distance`
|
||||
- Return candidates (no reranking yet)
|
||||
|
||||
5. Rerank (calls M3.2 rerank client):
|
||||
- Collect top K=3×limit candidates (e.g., 15 for limit=5)
|
||||
- Send to `/v1/rerank` with passages=candidates and query
|
||||
- Parse `bge-reranker-base` response, extract score per candidate
|
||||
- Compute `rerank_score = raw_score / 100` (reranker outputs [0,100])
|
||||
|
||||
6. Sort by rerank_score descending, take top `limit` results
|
||||
|
||||
7. Edge walk (L1→L0, L2→L1):
|
||||
- For each result, query `memory_edge` to find parent nodes
|
||||
- Fetch parent node text from `memory_node`
|
||||
- Include in `parents` array (ordered by edge precedence if tracked, else by sha256)
|
||||
|
||||
8. Assemble response and return 200
|
||||
|
||||
## Acceptance
|
||||
|
||||
- Query with valid text returns results
|
||||
- Results include query_score and rerank_score
|
||||
- L0 parents are walked and included
|
||||
- Different level filters change result count (e.g., L0 only returns more results)
|
||||
- Timeout parameter is respected
|
||||
- Query too short (e.g., single char) handled gracefully (400 or empty result, not crash)
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** Integration tests against server + pgvector repo populated with known nodes.
|
||||
|
||||
**Setup:** Load `tests/fixtures/memory_nodes.jsonl` into test pgvector DB before each test. Nodes include L0 (evidence), L1 (per-query memory), and L2 (synthesis) with known text and relationships.
|
||||
|
||||
**Integration test** — `tests/it_query_endpoint.rs`:
|
||||
1. `a1_basic_query_returns_results` — GET /query?query=Kong+body returns 200 with `results` array.
|
||||
2. `a2_scores_are_present` — result items include `query_score` and `rerank_score`, both floats in [0,1].
|
||||
3. `a3_l0_parents_included` — L1 result has `parents` array containing L0 nodes.
|
||||
4. `a4_level_filter_l0_only` — GET /query?level=L0 returns L0 nodes only (check level field).
|
||||
5. `a5_level_filter_l1_l2` — GET /query?level=L1,L2 returns only L1 and L2 (no L0).
|
||||
6. `a6_project_filter_works` — ingest into two projects, query with `project=poimen` — result.project matches.
|
||||
7. `a7_limit_respected` — GET /query?limit=3 returns ≤3 results.
|
||||
8. `a8_query_score_before_rerank` — query_score from HNSW comes before rerank; rerank_score ≤ query_score (reranker should not boost beyond HNSW recall).
|
||||
9. `a9_timeout_enforced` — manually slow the embedding service (mock delay 10s), GET /query with `timeout_seconds=1` returns 503.
|
||||
10. `a10_empty_query_returns_400` — GET /query (no query param) or GET /query?query= returns 400.
|
||||
|
||||
**Command:** `cargo test -p mem-cli query_endpoint`
|
||||
|
||||
**False pass:**
|
||||
- Testing only the happy path. Timeout, missing parent, embedding failure — all return different error codes.
|
||||
- Results sorted by query_score, not rerank_score. Reranking must reorder the results.
|
||||
- Parent nodes fetched but never asserted. A result with empty `parents` passes all checks.
|
||||
- query_score computed correctly but rerank_score always zero. Both must be present and in [0,1].
|
||||
|
||||
## Traps
|
||||
|
||||
- Timeout is wall-clock time, not per-service timeout. A 5s timeout that calls embedding (200ms) + HNSW (100ms) + rerank (500ms) should complete in <5s total, not each. Use `tokio::time::timeout()` around the entire handler.
|
||||
- HNSW uses `<->` operator for cosine distance (0 = opposite, 1 = same). 1 - distance is correct for similarity; do not invert again.
|
||||
- Reranker scores are [0,100]; dividing by 100 gives [0,1]. Not dividing is a common bug.
|
||||
- Embedding cache: the same query text submitted twice should reuse the embedding (save 200ms). Easy to forget.
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)
|
||||
@@ -0,0 +1,156 @@
|
||||
# M3.5.4 — Federation: single query across multiple projects
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M3.5.8 |
|
||||
| Depends | M3.5.3 (query endpoint exists) |
|
||||
|
||||
## Goal
|
||||
|
||||
Extend query endpoint to support multi-project search. When `project` param is omitted, a single query searches all projects concurrently, deduplicates results, and merges scores.
|
||||
|
||||
## Design
|
||||
|
||||
**Single-project query (no change):**
|
||||
```
|
||||
GET /memory/query?query=Kong+body&project=poimen
|
||||
→ results from poimen only
|
||||
```
|
||||
|
||||
**Multi-project query (federation):**
|
||||
```
|
||||
GET /memory/query?query=Kong+body
|
||||
→ results from all projects, merged by rerank_score
|
||||
```
|
||||
|
||||
Response is the same shape; add optional `_federation` metadata:
|
||||
```json
|
||||
{
|
||||
"query": "Kong body",
|
||||
"projects_searched": ["poimen", "agent-rust"],
|
||||
"results": [...],
|
||||
"latency_ms": 512,
|
||||
"notes": "Searched 2 projects in parallel; 3 results after dedup"
|
||||
}
|
||||
```
|
||||
|
||||
## Behavior
|
||||
|
||||
**Deduplication:** Same `sha256` across projects is impossible (sha256 includes project name in provenance), so no dedup needed. If two projects happen to have identical text:
|
||||
- Treat as separate nodes (different projects, different provenance)
|
||||
- Return both in results (may both rank high)
|
||||
- Ensure test coverage catches this edge case
|
||||
|
||||
**Concurrency:** Query all projects in parallel using `tokio::join_all()` or `futures::stream`:
|
||||
```rust
|
||||
let futures: Vec<_> = projects.iter()
|
||||
.map(|proj| query_single_project(query_text, proj, limit))
|
||||
.collect();
|
||||
let results: Vec<_> = futures::future::join_all(futures).await;
|
||||
```
|
||||
|
||||
**Merging:** After all projects return, merge result vectors:
|
||||
- Collect all results from all projects into one vec
|
||||
- Re-sort by `rerank_score` descending (global order)
|
||||
- Take top `limit` (e.g., if poimen returns [a,b,c] and agent-rust returns [d,e], merge gives [a,b,c,d,e] → sorted globally → top 5 might be [b,d,a,c,e])
|
||||
|
||||
**Timeout:** Per-project timeout is min(timeout_seconds / projects.len(), 2s). If one project is slow, others complete faster and we still return results from fast projects after global timeout.
|
||||
- E.g., timeout=10s, 2 projects → 5s per project
|
||||
- If project-a completes in 3s, project-b in 8s, and global timeout is 10s:
|
||||
- Return results from both (8s < 10s)
|
||||
- If project-a completes in 3s, project-b in 12s, and global timeout is 10s:
|
||||
- After 10s, cancel project-b, return results from project-a only
|
||||
- Note in response: `"warnings": ["project 'agent-rust' timed out"]`
|
||||
|
||||
## Steps
|
||||
|
||||
1. Parse `project` param:
|
||||
- If provided, single-project path (M3.5.3 unchanged)
|
||||
- If omitted, multi-project path
|
||||
|
||||
2. List all known projects (from queries YAML):
|
||||
```rust
|
||||
let projects = load_standing_queries()?.projects();
|
||||
```
|
||||
|
||||
3. Spawn concurrent query tasks:
|
||||
```rust
|
||||
let futures: Vec<_> = projects.into_iter()
|
||||
.map(|proj| {
|
||||
let params = params.clone();
|
||||
params.project = Some(proj);
|
||||
query_handler_impl(¶ms, store, llm)
|
||||
})
|
||||
.collect();
|
||||
```
|
||||
|
||||
4. Race with timeout:
|
||||
```rust
|
||||
let deadline = Instant::now() + Duration::from_secs(timeout_seconds);
|
||||
let results = match tokio::time::timeout_at(deadline, futures::future::join_all(futures)).await {
|
||||
Ok(vec) => vec.into_iter().flatten().collect(), // flatten per-project results
|
||||
Err(_) => { /* partial results + warning */ }
|
||||
};
|
||||
```
|
||||
|
||||
5. Merge and sort:
|
||||
```rust
|
||||
results.sort_by(|a, b| b.rerank_score.partial_cmp(&a.rerank_score).unwrap());
|
||||
results.truncate(limit);
|
||||
```
|
||||
|
||||
6. Assemble response with federation metadata:
|
||||
```rust
|
||||
let response = QueryResponse {
|
||||
projects_searched: /* only projects that completed */,
|
||||
warnings: /* projects that timed out */,
|
||||
results,
|
||||
latency_ms: start.elapsed().as_millis() as u64,
|
||||
..
|
||||
};
|
||||
```
|
||||
|
||||
## Acceptance
|
||||
|
||||
- Single project specified: no federation, same result as M3.5.3
|
||||
- No project specified: all projects queried
|
||||
- Results merged and globally sorted by rerank_score
|
||||
- Partial results returned if one project times out
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** Integration tests with two projects in test pgvector DB.
|
||||
|
||||
**Integration test** — `tests/it_query_federation.rs`:
|
||||
1. `a1_single_project_no_federation` — GET /query?project=poimen returns single-project results only.
|
||||
2. `a2_multi_project_searches_all` — GET /query (no project) with >1 project in DB returns results from all.
|
||||
3. `a3_global_sort_order` — two projects return results, merge sorts by rerank_score globally (not per-project).
|
||||
4. `a4_federation_metadata_present` — response includes `projects_searched` array with all completed projects.
|
||||
5. `a5_partial_results_on_timeout` — slow one project (mock 10s delay), set timeout_seconds=2, GET /query returns results from fast project only with warning.
|
||||
6. `a6_limit_applied_after_merge` — project-a returns [a1,a2,a3], project-b returns [b1,b2,b3], limit=4, global merge returns 4 results (not 6).
|
||||
7. `a7_no_project_filter_in_response` — response.project_filter is null or omitted (unlike single-project which sets it).
|
||||
8. `a8_concurrent_execution` — spy on timing: timestamp project-a query start, project-b query start, both should be ~simultaneous (not sequential).
|
||||
|
||||
**Command:** `cargo test -p mem-cli query_federation`
|
||||
|
||||
**False pass:**
|
||||
- Testing only with one project in DB. Federation always "works" if there is nothing to federate.
|
||||
- Timeout never exercised. Mock a slow project and assert results are partial.
|
||||
- Per-project sorting instead of global sort. Results look reasonable but violate the contract (should be global top-k).
|
||||
- Concurrency not verified. Queries can be sequential (slow) and still return correct results; only timing proves concurrency.
|
||||
|
||||
## Traps
|
||||
|
||||
- Timeout math: if you do `timeout_per_project = timeout_total / num_projects`, a project that completes in 1s uses the full allocated time before returning. Should be `remaining_time = deadline - now()`.
|
||||
- Partial results: if project-a returns 5 results and project-b times out, you have 5 results but may have wanted 10 (limit=10). Document whether partial results truncate or stay over-limit.
|
||||
- Clone overhead: cloning `QueryParams` for each project is small; cloning a large result vec is not. Use references/Arc where possible.
|
||||
- Flatten after join_all: `join_all` returns `Vec<Result>`, must flatten errors (either as partial results or early exit).
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)
|
||||
@@ -0,0 +1,166 @@
|
||||
# M3.5.5 — GET /skills and /skills/{name}: loadable skills catalog
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M3.5.8 |
|
||||
| Depends | M3.5.1, M4.1 (skill drafts exist locally) |
|
||||
|
||||
## Goal
|
||||
|
||||
Read-only endpoints for skill catalog. List all promoted skills (exclude `_drafts/`), fetch individual skill metadata and body. Skills are Obsidian notes; expose them over HTTP for agent discovery.
|
||||
|
||||
## Design
|
||||
|
||||
**List all loadable skills:**
|
||||
```
|
||||
GET /memory/skills?loadable=true
|
||||
→ 200 {
|
||||
"skills": [
|
||||
{
|
||||
"name": "infra-root-causes",
|
||||
"description": "Identify root causes of infrastructure failures",
|
||||
"when_to_use": "When troubleshooting cluster or service outages",
|
||||
"argument_hint": "--project <name>",
|
||||
"promoted_at": "2026-08-20T10:30:00Z",
|
||||
"generated_from": null
|
||||
},
|
||||
...
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Get one skill (metadata only):**
|
||||
```
|
||||
GET /memory/skills/infra-root-causes
|
||||
→ 200 {
|
||||
"name": "infra-root-causes",
|
||||
"description": "...",
|
||||
"when_to_use": "...",
|
||||
"argument_hint": "...",
|
||||
"promoted_at": "2026-08-20T...",
|
||||
"generated_from": null
|
||||
}
|
||||
```
|
||||
|
||||
**Get skill with body (full content):**
|
||||
```
|
||||
GET /memory/skills/infra-root-causes?include_body=true
|
||||
→ 200 {
|
||||
"name": "infra-root-causes",
|
||||
"description": "...",
|
||||
"body": "# Infra root causes\n\n..."
|
||||
}
|
||||
```
|
||||
|
||||
**Filters:**
|
||||
- `loadable=true` (default): exclude `_drafts/`, return only promoted skills
|
||||
- `loadable=false`: include everything (admin only — must have special apikey, documented in code)
|
||||
|
||||
## Technical
|
||||
|
||||
**Source:** Vault at `vault/skills/` contains skill markdown files. Each skill is a directory:
|
||||
```
|
||||
vault/skills/
|
||||
infra-root-causes/
|
||||
SKILL.md <- frontmatter + body
|
||||
```
|
||||
|
||||
**Frontmatter (YAML in SKILL.md):**
|
||||
```yaml
|
||||
---
|
||||
name: infra-root-causes
|
||||
description: Identify root causes of infrastructure failures
|
||||
when_to_use: When troubleshooting cluster or service outages
|
||||
argument_hint: --project <name>
|
||||
generated_from: null | <L2 sha256>
|
||||
---
|
||||
```
|
||||
|
||||
**Drafts are in `vault/skills/_drafts/`:**
|
||||
```
|
||||
vault/skills/
|
||||
_drafts/
|
||||
new-skill/
|
||||
SKILL.md
|
||||
```
|
||||
|
||||
Only load from `vault/skills/*/SKILL.md` (not `_drafts`), unless `loadable=false` is passed with an admin key.
|
||||
|
||||
## Steps
|
||||
|
||||
1. `GET /memory/skills` handler:
|
||||
- List `vault/skills/` directory (skip `_drafts/`)
|
||||
- For each `*/SKILL.md`, parse frontmatter
|
||||
- Extract: `name`, `description`, `when_to_use`, `argument_hint`, `promoted_at` (file mtime)
|
||||
- Parse `generated_from` field to show provenance
|
||||
- Return array
|
||||
|
||||
2. `GET /memory/skills/{name}` handler:
|
||||
- Load `vault/skills/{name}/SKILL.md`
|
||||
- Parse frontmatter and body
|
||||
- If `include_body=false` (default), return metadata only
|
||||
- If `include_body=true`, include markdown body
|
||||
|
||||
3. `loadable` query param (admin-only feature):
|
||||
- Default: exclude `_drafts/`
|
||||
- `loadable=false` with admin apikey: include `_drafts/` in listing
|
||||
- Non-admin key requesting `loadable=false` → 403 Forbidden
|
||||
|
||||
4. Error handling:
|
||||
- Skill not found → 404 with `{"error":"not_found","reason":"skill 'xyz' not promoted"}`
|
||||
- Malformed SKILL.md (frontmatter parse fails) → 500 with error (admin debug only)
|
||||
- Admin check: apikey must be in a whitelist (env var `MEM_ADMIN_APIKEYS` or config)
|
||||
|
||||
## Acceptance
|
||||
|
||||
- List endpoint returns all promoted skills
|
||||
- Individual skill fetch works
|
||||
- Drafts are excluded by default
|
||||
- Admin with `loadable=false` sees drafts
|
||||
- Skill body is optional (include_body param)
|
||||
- Promoted_at field reflects file mtime
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** Integration tests + filesystem fixtures.
|
||||
|
||||
**Setup:** Create test `vault/skills/` with:
|
||||
- `vault/skills/test-skill-1/SKILL.md` (promoted)
|
||||
- `vault/skills/test-skill-2/SKILL.md` (promoted)
|
||||
- `vault/skills/_drafts/draft-skill/SKILL.md` (unpromoted)
|
||||
|
||||
**Integration test** — `tests/it_skills_endpoint.rs`:
|
||||
1. `a1_list_skills_returns_promoted` — GET /skills returns array with test-skill-1 and test-skill-2.
|
||||
2. `a2_drafts_excluded_by_default` — GET /skills does not include draft-skill.
|
||||
3. `a3_drafts_included_with_admin_key` — GET /skills?loadable=false with admin apikey includes draft-skill.
|
||||
4. `a4_non_admin_denied_drafts` — GET /skills?loadable=false with regular apikey returns 403.
|
||||
5. `a5_get_single_skill_metadata` — GET /skills/test-skill-1 returns 200 with frontmatter fields.
|
||||
6. `a6_include_body_true` — GET /skills/test-skill-1?include_body=true returns body field with markdown.
|
||||
7. `a7_include_body_false` — GET /skills/test-skill-1?include_body=false (or omitted) does not include body field.
|
||||
8. `a8_skill_not_found` — GET /skills/nonexistent returns 404.
|
||||
9. `a9_promoted_at_is_file_mtime` — GET /skills/test-skill-1, assert promoted_at is a valid ISO timestamp close to SKILL.md's modification time.
|
||||
10. `a10_generated_from_field` — SKILL.md with `generated_from: sha256xyz` is parsed and returned as-is.
|
||||
|
||||
**Command:** `cargo test -p mem-cli skills_endpoint`
|
||||
|
||||
**False pass:**
|
||||
- Drafts never created in test fixtures. The default exclude-drafts logic is untestable without a draft.
|
||||
- Admin key never tested. Non-admin path and admin path can be identical in code.
|
||||
- Promoted_at never validated. Can return a fake date; file mtime is the only source.
|
||||
- Frontmatter parsing doesn't validate required fields (name, description). A malformed SKILL.md is silently returned with null values.
|
||||
|
||||
## Traps
|
||||
|
||||
- Vault directory may not exist locally (only in deployed cluster). Start with a default empty list if vault/ is missing.
|
||||
- YAML frontmatter parsing is fussy. A tab instead of spaces breaks YAML. Use a YAML parser (serde_yaml) and validate on load.
|
||||
- File mtime precision: Unix mtime is seconds; SKILL.md edits may not increment it if done within the same second. Use actual write timestamp if available.
|
||||
- Admin key stored in env var. If unset, default to deny (safer than default allow).
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md § Skills — the procedural projection](../DESIGN.md#skills--the-procedural-projection)
|
||||
@@ -0,0 +1,146 @@
|
||||
# M3.5.6 — GET /projects and /projects/{id}/status: metadata, metrics, synthesis timestamps
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | S — < 1 day |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M3.5.8 |
|
||||
| Depends | M3.5.1, M2 (projections exist) |
|
||||
|
||||
## Goal
|
||||
|
||||
Introspection endpoints for memory state per project. List projects, show metadata, ingest/synthesis history, memory size stats.
|
||||
|
||||
## Design
|
||||
|
||||
**List all projects:**
|
||||
```
|
||||
GET /memory/projects
|
||||
→ 200 {
|
||||
"projects": [
|
||||
{
|
||||
"id": "poimen",
|
||||
"standing_queries": 3,
|
||||
"last_ingest_at": "2026-08-20T10:30:00Z",
|
||||
"last_synthesis_at": "2026-08-20T12:00:00Z",
|
||||
"total_chunks": 412,
|
||||
"total_evidence": 17,
|
||||
"memory_size_bytes": 45280
|
||||
},
|
||||
...
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Get project status:**
|
||||
```
|
||||
GET /memory/projects/poimen/status
|
||||
→ 200 {
|
||||
"project_id": "poimen",
|
||||
"standing_queries": [
|
||||
{
|
||||
"id": "infra-root-causes",
|
||||
"question": "What infrastructure bugs were found...",
|
||||
"last_ingest_at": "2026-08-20T10:30:00Z",
|
||||
"chunks_seen": 412,
|
||||
"chunks_used": 17,
|
||||
"memory_tokens": 142
|
||||
},
|
||||
...
|
||||
],
|
||||
"l2_synthesis": {
|
||||
"last_synthesis_at": "2026-08-20T12:00:00Z",
|
||||
"chunks_seen": 3,
|
||||
"chunks_used": 2,
|
||||
"memory_tokens": 876,
|
||||
"exit_gate_fired": true
|
||||
},
|
||||
"next_synthesis_at": "2026-08-21T12:00:00Z",
|
||||
"total_log_size_bytes": 45280,
|
||||
"embedding_cache_hits": 234,
|
||||
"embedding_cache_misses": 12
|
||||
}
|
||||
```
|
||||
|
||||
## Metrics
|
||||
|
||||
Pull from multiple sources:
|
||||
- **Standing queries:** Load from `queries/<project>.yaml`
|
||||
- **Last ingest:** Query JSONL log for most recent `run_end` record per query_id
|
||||
- **Memory stats:** Count nodes in pgvector, sum bytes of text
|
||||
- **L2 synthesis:** Query JSONL log for most recent L2 `run_end`
|
||||
- **Cache stats:** Track in-memory (API server state); return per request
|
||||
|
||||
## Steps
|
||||
|
||||
1. `GET /memory/projects` handler:
|
||||
- List all project IDs from `queries/` directory
|
||||
- For each project:
|
||||
- Load `queries/<project>.yaml` to get standing_queries count
|
||||
- Query pgvector: `SELECT COUNT(*) FROM memory_node WHERE project = $1`
|
||||
- Query pgvector: `SELECT SUM(LENGTH(text)) FROM memory_node WHERE project = $1`
|
||||
- Query JSONL log: find most recent L1 `run_end` to get last_ingest_at
|
||||
- Query JSONL log: find most recent L2 `run_end` to get last_synthesis_at
|
||||
- Sort by id and return
|
||||
|
||||
2. `GET /memory/projects/{id}/status` handler:
|
||||
- Verify project exists; unknown → 404
|
||||
- Load `queries/<project>.yaml` and parse all queries
|
||||
- For each query, query JSONL log:
|
||||
- Find most recent `run_end` record (level L1, query_id = this query's id)
|
||||
- Extract chunks_seen, chunks_used, final_memory_tokens, last timestamp
|
||||
- Query JSONL log for L2 run_end (level L2, project = id):
|
||||
- Extract synthesis metadata, exit_gate fire status
|
||||
- Compute next_synthesis_at:
|
||||
- If last_synthesis_at + 24h < now, return "immediately"
|
||||
- Otherwise, return last_synthesis_at + 24h
|
||||
- Assemble response
|
||||
|
||||
3. Cache stats:
|
||||
- `embedding_cache_hits` and `embedding_cache_misses` tracked by embeddings client
|
||||
- Expose via `Extension<Arc<EmbeddingsClient>>` → `.stats()`
|
||||
- Return per request (snapshot at query time)
|
||||
|
||||
## Acceptance
|
||||
|
||||
- List endpoint returns all projects
|
||||
- Individual project status is queryable
|
||||
- Metrics are accurate (match log/pgvector state)
|
||||
- Unknown project returns 404
|
||||
- Synthesis scheduling shown (next run time)
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** Integration tests with populated JSONL log and pgvector DB.
|
||||
|
||||
**Integration test** — `tests/it_projects_endpoint.rs`:
|
||||
1. `a1_list_projects` — GET /projects returns array with test project(s).
|
||||
2. `a2_project_count_correct` — total_chunks field matches pgvector COUNT.
|
||||
3. `a3_project_evidence_count` — total_evidence field matches L0 node count for project.
|
||||
4. `a4_get_project_status` — GET /projects/<id>/status returns 200.
|
||||
5. `a5_standing_queries_listed` — standing_queries array in status matches queries YAML.
|
||||
6. `a6_last_ingest_timestamp` — last_ingest_at is recent and matches JSONL log.
|
||||
7. `a7_l2_synthesis_metadata` — l2_synthesis object contains last_synthesis_at and exit_gate_fired.
|
||||
8. `a8_cache_stats_present` — embedding_cache_hits and cache_misses are present and >= 0.
|
||||
9. `a9_next_synthesis_at_scheduled` — next_synthesis_at is a valid future timestamp.
|
||||
10. `a10_unknown_project_404` — GET /projects/nonexistent/status returns 404.
|
||||
|
||||
**Command:** `cargo test -p mem-cli projects_endpoint`
|
||||
|
||||
**False pass:**
|
||||
- total_chunks hardcoded to a fixed number; never actually counts.
|
||||
- Cache stats always zero (client doesn't track; endpoint returns fake values).
|
||||
- Last ingest timestamp never validated against actual log.
|
||||
|
||||
## Traps
|
||||
|
||||
- JSONL log queries are slow for large projects (412 chunks, naive scan). Consider indexing by project_id or caching if >10K chunks.
|
||||
- Next synthesis scheduling logic is simple (24h interval). If synthesis runs are skipped or delayed, estimate becomes stale. Document the assumption.
|
||||
- Memory size calculation uses SUM(LENGTH(text)) which is TEXT byte length in DB, not network wire size or actual storage (compression, overhead).
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)
|
||||
@@ -0,0 +1,175 @@
|
||||
# M3.5.7 — Rate limiting (per-apikey) and idempotency by sha256
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | — |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M3.5.8 |
|
||||
| Depends | M3.5.2, M3.5.3 (ingest and query endpoints exist) |
|
||||
|
||||
## Goal
|
||||
|
||||
Rate limiting prevents abusive load; idempotency ensures retry safety. Both are per-apikey and per-endpoint.
|
||||
|
||||
## Design
|
||||
|
||||
**Rate limits (defaults, configurable via env):**
|
||||
- `POST /memory/ingest`: 100 jobs/hour per apikey
|
||||
- `GET /memory/query`: 1000 requests/hour per apikey
|
||||
- `GET /memory/skills`: unlimited
|
||||
- `GET /memory/projects`: 100 requests/hour per apikey
|
||||
|
||||
**Burst allowance:** 10 requests/second (hard burst cap, then 429).
|
||||
|
||||
**Response on rate limit:**
|
||||
```
|
||||
HTTP 429 Too Many Requests
|
||||
Retry-After: 47
|
||||
{
|
||||
"error": "rate_limit_exceeded",
|
||||
"reason": "100 requests/hour for POST /memory/ingest",
|
||||
"retry_after_seconds": 47,
|
||||
"limit_window": "3600s"
|
||||
}
|
||||
```
|
||||
|
||||
**Idempotency:**
|
||||
- `POST /memory/ingest` uses `ingest_id` (SHA256 of batch content) as idempotency key
|
||||
- Same `ingest_id` resubmitted within 24 hours returns same `job_id`, no re-enqueue
|
||||
- Idempotency key extracted from request body (not header)
|
||||
|
||||
## Implementation
|
||||
|
||||
**Rate limiting strategy:** Token bucket per apikey per endpoint. Track in memory (not Redis yet).
|
||||
```rust
|
||||
pub struct RateLimiter {
|
||||
buckets: Arc<Mutex<HashMap<String, Vec<RateBucket>>>>, // apikey -> [one per endpoint]
|
||||
}
|
||||
|
||||
pub struct RateBucket {
|
||||
tokens: f64,
|
||||
last_refill: Instant,
|
||||
capacity: f64,
|
||||
refill_rate: f64, // tokens/sec
|
||||
}
|
||||
```
|
||||
|
||||
**Token refill:** On each request, add `(now - last_refill) * refill_rate` tokens (cap at capacity).
|
||||
|
||||
**Burst handling:**
|
||||
- Allow burst of 10 req/sec without delay
|
||||
- Requests above burst queued (blocked until tokens available) or rejected (429)
|
||||
- Decision: **reject** is simpler and encourages clients to batch. Implement rejection.
|
||||
|
||||
**Idempotency:**
|
||||
- Extract `ingest_id` from request body (JSON key or computed if omitted)
|
||||
- Check against recent idempotency store (memory, 24h TTL)
|
||||
- If found, return cached response (job_id)
|
||||
- If not found, process normally and store (ingest_id → response)
|
||||
|
||||
## Steps
|
||||
|
||||
1. Create `RateLimiter` struct:
|
||||
```rust
|
||||
impl RateLimiter {
|
||||
fn new() -> Self { /* init empty */ }
|
||||
fn check(&mut self, apikey: &str, endpoint: &str) -> Result<(), RateLimitError> {
|
||||
// refill, check capacity, return Ok or Err with Retry-After
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
2. Add `RateLimiter` as app state:
|
||||
```rust
|
||||
let limiter = Arc::new(Mutex::new(RateLimiter::new()));
|
||||
HttpServer::new(move || {
|
||||
App::new()
|
||||
.app_data(Data::new(limiter.clone()))
|
||||
})
|
||||
```
|
||||
|
||||
3. Middleware to extract apikey and check limit:
|
||||
```rust
|
||||
pub struct RateLimitMiddleware {
|
||||
limits: Arc<Mutex<RateLimiter>>,
|
||||
}
|
||||
|
||||
impl Middleware for RateLimitMiddleware { ... }
|
||||
```
|
||||
- Extract apikey from request context (set by auth middleware)
|
||||
- Determine endpoint (path)
|
||||
- Call `limiter.check(apikey, endpoint)`
|
||||
- If Err, return 429 with Retry-After
|
||||
|
||||
4. Idempotency store:
|
||||
```rust
|
||||
pub struct IdempotencyStore {
|
||||
cache: Arc<Mutex<HashMap<String, (HttpResponse, Instant)>>>,
|
||||
}
|
||||
|
||||
impl IdempotencyStore {
|
||||
fn get(&self, key: &str) -> Option<HttpResponse> { /* if not expired */ }
|
||||
fn set(&mut self, key: String, response: HttpResponse) { }
|
||||
}
|
||||
```
|
||||
|
||||
5. `POST /ingest` handler:
|
||||
- Parse request body to extract `ingest_id`
|
||||
- Query idempotency store for `ingest_id`
|
||||
- If found and not expired (24h), return cached response
|
||||
- If not found, process normally:
|
||||
- Enqueue ingest
|
||||
- Cache the 202 response with ingest_id as key
|
||||
- Return response
|
||||
|
||||
6. Configuration:
|
||||
- Load rate limits from env vars: `MEM_RATE_LIMIT_INGEST`, `MEM_RATE_LIMIT_QUERY`, etc.
|
||||
- Load burst cap from env: `MEM_RATE_LIMIT_BURST` (default 10 req/sec)
|
||||
- Load idempotency TTL from env: `MEM_IDEMPOTENCY_TTL_SECS` (default 86400)
|
||||
|
||||
## Acceptance
|
||||
|
||||
- Requests within limit succeed (200 or 202)
|
||||
- Requests at burst cap (10/sec) blocked immediately
|
||||
- Rate limit reset after time window (test with mocked time)
|
||||
- Same ingest_id resubmitted returns same job_id (idempotent)
|
||||
- Different ingest_id queued separately
|
||||
- Retry-After header correct
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** Integration tests + time mocking.
|
||||
|
||||
**Integration test** — `tests/it_rate_limiting.rs`:
|
||||
1. `a1_within_limit_succeeds` — 5 consecutive GET /query requests within 1-hour limit all succeed (200).
|
||||
2. `a2_at_burst_cap_429` — 11 GET /query requests in 1 second, 11th returns 429.
|
||||
3. `a3_limit_window_resets` — 100 GET /query requests in hour 1 all succeed (limit reached), 101st fails (429), mock time to hour+2, 102nd succeeds (window reset).
|
||||
4. `a4_per_apikey_isolation` — two different apikeys, each send 5 requests, both succeed (limits are independent).
|
||||
5. `a5_per_endpoint_isolation` — 100 POST /ingest requests succeed (limit=100), 1 GET /query request succeeds (different endpoint, different limit).
|
||||
6. `a6_retry_after_header` — 429 response includes `Retry-After: N` header with correct value.
|
||||
7. `a7_ingest_id_idempotent` — POST /ingest with id_a succeeds, POST again with id_a returns same job_id.
|
||||
8. `a8_different_ingest_ids_separate` — POST /ingest (id_a), POST (id_b) both succeed with different job_ids.
|
||||
9. `a9_idempotency_expires` — POST /ingest (id_a), mock time to 25 hours later, POST (id_a) again returns different job_id (old idempotency cache expired).
|
||||
10. `a10_rate_limit_per_endpoint_documented` — grep the code for limit values; each endpoint has a defined limit.
|
||||
|
||||
**Command:** `cargo test -p mem-cli rate_limiting -- --nocapture`
|
||||
|
||||
**False pass:**
|
||||
- Burst cap tested with 10 requests but timing is imprecise (some reqs slow, burst calc off).
|
||||
- Rate limit window reset never tested with time mock. Limits always work within a short test window.
|
||||
- Idempotency key never actually extracted from body; hardcoded in test.
|
||||
- Per-apikey isolation not tested with two keys.
|
||||
|
||||
## Traps
|
||||
|
||||
- Token bucket refill at Instant::now() is wall-clock time; in tests, use a mock clock (or avoid time-dependent tests).
|
||||
- Burst cap as "10 req/sec" is naive if requests take 100ms each (effectively 10 concurrent). Real burst is 10 within the same millisecond. Better: track request arrival rate over a sliding window.
|
||||
- Idempotency cache unbounded growth. Must evict expired entries (implement on-read eviction or background sweep).
|
||||
- Rate limit math: capacity=100 tokens/hour, refill=100/3600 tokens/sec. A request at t=0 uses 1 token (99 left). At t=36s, 1 token is refilled (100 left) — this is correct. Watch for off-by-one.
|
||||
|
||||
---
|
||||
|
||||
Background: [DESIGN.md § Distributed API Layer § Auth & rate limits](../DESIGN.md#scaling-constraints)
|
||||
@@ -0,0 +1,137 @@
|
||||
# M3.5.8 — **M3.5 composition gate** — API end-to-end
|
||||
|
||||
| Field | Value |
|
||||
|---|---|
|
||||
| Phase | M3.5 — Distributed API Layer |
|
||||
| Size | M — 1–3 days |
|
||||
| Status | ⬜ Not started |
|
||||
| Flags | gate |
|
||||
| Spec | inlined below |
|
||||
| Blocks | M4, M5 (can start in parallel after this gate) |
|
||||
| Depends | M3.5.1, M3.5.2, M3.5.3, M3.5.4, M3.5.5, M3.5.6, M3.5.7 |
|
||||
|
||||
## Goal
|
||||
|
||||
Verify that the API layer is a working facade. Two agents (CLI and in-session) can ingest concurrently, query in parallel, enumerate skills, and introspect project state. No blocking, no race conditions, idempotency holds.
|
||||
|
||||
## Acceptance Criteria
|
||||
|
||||
All M3.5.x tasks complete, and the integration below passes.
|
||||
|
||||
**Must work:**
|
||||
1. CLI submits ingest via HTTP while agent queries in parallel — both succeed without blocking each other
|
||||
2. Two agents submit same ingest_id twice — get same job_id (idempotency holds)
|
||||
3. Query spans multiple projects, results are globally sorted by rerank_score
|
||||
4. Skills list excludes drafts; admin apikey sees drafts
|
||||
5. Project status endpoint reports correct memory metrics
|
||||
6. Rate limiting enforces per-endpoint, per-apikey limits
|
||||
7. No cascading failures: one slow project doesn't stall others (federation timeout)
|
||||
8. Logs are clean (no panics, no unhandled errors)
|
||||
|
||||
## Verify
|
||||
|
||||
**Harness:** End-to-end test harness that simulates mixed workload.
|
||||
|
||||
**Integration test** — `tests/it_e2e_api.rs`:
|
||||
1. `a1_cli_ingest_and_agent_query_concurrent` —
|
||||
- Spawn HTTP server with test pgvector DB
|
||||
- CLI submits ingest batch (POST /ingest)
|
||||
- Agent submits query (GET /query) in parallel
|
||||
- Both complete within 30s, both return 200/202
|
||||
2. `a2_ingest_idempotency_holds` —
|
||||
- CLI submits (ingest_id_a) → job_id_1
|
||||
- Agent submits same (ingest_id_a) → job_id_1 (identical)
|
||||
- Different (ingest_id_b) → job_id_2 (different)
|
||||
3. `a3_multi_project_federation_sorts_globally` —
|
||||
- Ingest sample data into two projects (poimen, agent-rust)
|
||||
- Query "root cause" (no project specified)
|
||||
- Results include nodes from both projects
|
||||
- Sorted by rerank_score globally (not per-project)
|
||||
4. `a4_skills_list_excludes_drafts_by_default` —
|
||||
- GET /memory/skills → returns promoted skills only
|
||||
- GET /memory/skills?loadable=false with admin key → includes drafts
|
||||
5. `a5_project_status_metrics_accurate` —
|
||||
- Ingest 50 chunks
|
||||
- GET /memory/projects/poimen/status
|
||||
- Asserts: total_chunks ≈ 50, last_ingest_at is recent, standing_queries count > 0
|
||||
6. `a6_rate_limit_enforced` —
|
||||
- Set rate limit to 5 req/hour for testing
|
||||
- Send 6 GET /query requests
|
||||
- First 5 succeed, 6th returns 429
|
||||
7. `a7_federation_timeout_partial_results` —
|
||||
- Mock slow project (10s response time)
|
||||
- Query with timeout_seconds=2
|
||||
- Results from fast project, warning about slow project
|
||||
8. `a8_no_cascading_failures` —
|
||||
- Inject error in embeddings service (simulate 500)
|
||||
- GET /query returns 503, not cascading to other endpoints
|
||||
- Other endpoints (ingest, skills) still work
|
||||
9. `a9_logs_clean_no_panics` —
|
||||
- Capture stderr during test
|
||||
- Grep for "panic", "unwrap", "expect" — should not appear
|
||||
- All errors should be explicit Result types, not crashes
|
||||
10. `a10_health_check_always_responds` —
|
||||
- Server is under heavy load (rate limit tests, concurrent ingest)
|
||||
- GET /health still returns 200 within 100ms
|
||||
|
||||
**Command:** `cargo test -p mem-cli e2e_api -- --nocapture`
|
||||
|
||||
**Manual verification (smoke test):**
|
||||
```bash
|
||||
# Start server
|
||||
cargo run -p mem-cli -- serve --port 8080 &
|
||||
sleep 2
|
||||
|
||||
# Ingest via HTTP
|
||||
curl -X POST -H "apikey: test" http://localhost:8080/memory/ingest \
|
||||
-d '{
|
||||
"project": "poimen",
|
||||
"source": "manual:smoke",
|
||||
"records": [...],
|
||||
"ingest_id": "abc123"
|
||||
}'
|
||||
# → expect 202, job_id
|
||||
|
||||
# Query
|
||||
curl -H "apikey: test" "http://localhost:8080/memory/query?query=test"
|
||||
# → expect 200, results array
|
||||
|
||||
# Skills
|
||||
curl -H "apikey: test" http://localhost:8080/memory/skills
|
||||
# → expect 200, skills array
|
||||
|
||||
# Project status
|
||||
curl -H "apikey: test" http://localhost:8080/memory/projects/poimen/status
|
||||
# → expect 200, metadata
|
||||
|
||||
# Rate limit test
|
||||
for i in {1..11}; do
|
||||
curl -H "apikey: test" "http://localhost:8080/memory/query?query=test" \
|
||||
-w "HTTP %{http_code}\n"
|
||||
done
|
||||
# → expect first 10 to succeed, 11th to be 429
|
||||
```
|
||||
|
||||
## False Pass
|
||||
|
||||
- Testing only happy path (all services available, no errors). Must include:
|
||||
- Embedding service down → 503
|
||||
- Slow project in federation → partial results + warning
|
||||
- Rate limit near boundary (9/10, 10/10, 11/10 reqs)
|
||||
- CLI and agent workloads not truly concurrent (sequential test masquerades as parallel). Use `tokio::join_all` or spy on timing to verify parallel execution.
|
||||
- Idempotency tested once; never tested with expiry or multiple projects.
|
||||
- Metrics never cross-checked against actual DB state. total_chunks reported but not verified against SELECT COUNT.
|
||||
- No error injection. If the API is untested with failures, cascading failures are invisible until production.
|
||||
|
||||
## Traps
|
||||
|
||||
- Server startup latency: tests must wait for port to be available (sleep or retry logic).
|
||||
- Test isolation: if tests share a DB, idempotency cache pollution breaks test N+1. Use separate test DB per test or reset cache between runs.
|
||||
- Timing: federation timeout at 2s is tight; if the machine is slow, test becomes flaky. Mock time instead of real delays.
|
||||
- Concurrent writes to JSONL log: if two ingest tasks write simultaneously, atomicity of the log is at risk. Ensure the log is single-writer or uses locking.
|
||||
|
||||
---
|
||||
|
||||
**Gate outcome:** All M3.5.x tasks green AND a1–a10 pass → M3.5 gate is green. CLI and agents can work with the API concurrently without blocking, race conditions, or idempotency issues.
|
||||
|
||||
Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)
|
||||
Reference in New Issue
Block a user