138 lines
5.7 KiB
Markdown
138 lines
5.7 KiB
Markdown
# M3.5.8 — **M3.5 composition gate** — API end-to-end
|
||
|
||
| Field | Value |
|
||
|---|---|
|
||
| Phase | M3.5 — Distributed API Layer |
|
||
| Size | M — 1–3 days |
|
||
| Status | ✅ Done |
|
||
| Flags | gate |
|
||
| Spec | inlined below |
|
||
| Blocks | M4, M5 (can start in parallel after this gate) |
|
||
| Depends | M3.5.1, M3.5.2, M3.5.3, M3.5.4, M3.5.5, M3.5.6, M3.5.7 |
|
||
|
||
## Goal
|
||
|
||
Verify that the API layer is a working facade. Two agents (CLI and in-session) can ingest concurrently, query in parallel, enumerate skills, and introspect project state. No blocking, no race conditions, idempotency holds.
|
||
|
||
## Acceptance Criteria
|
||
|
||
All M3.5.x tasks complete, and the integration below passes.
|
||
|
||
**Must work:**
|
||
1. CLI submits ingest via HTTP while agent queries in parallel — both succeed without blocking each other
|
||
2. Two agents submit same ingest_id twice — get same job_id (idempotency holds)
|
||
3. Query spans multiple projects, results are globally sorted by rerank_score
|
||
4. Skills list excludes drafts; admin apikey sees drafts
|
||
5. Project status endpoint reports correct memory metrics
|
||
6. Rate limiting enforces per-endpoint, per-apikey limits
|
||
7. No cascading failures: one slow project doesn't stall others (federation timeout)
|
||
8. Logs are clean (no panics, no unhandled errors)
|
||
|
||
## Verify
|
||
|
||
**Harness:** End-to-end test harness that simulates mixed workload.
|
||
|
||
**Integration test** — `tests/it_e2e_api.rs`:
|
||
1. `a1_cli_ingest_and_agent_query_concurrent` —
|
||
- Spawn HTTP server with test pgvector DB
|
||
- CLI submits ingest batch (POST /ingest)
|
||
- Agent submits query (GET /query) in parallel
|
||
- Both complete within 30s, both return 200/202
|
||
2. `a2_ingest_idempotency_holds` —
|
||
- CLI submits (ingest_id_a) → job_id_1
|
||
- Agent submits same (ingest_id_a) → job_id_1 (identical)
|
||
- Different (ingest_id_b) → job_id_2 (different)
|
||
3. `a3_multi_project_federation_sorts_globally` —
|
||
- Ingest sample data into two projects (poimen, agent-rust)
|
||
- Query "root cause" (no project specified)
|
||
- Results include nodes from both projects
|
||
- Sorted by rerank_score globally (not per-project)
|
||
4. `a4_skills_list_excludes_drafts_by_default` —
|
||
- GET /memory/skills → returns promoted skills only
|
||
- GET /memory/skills?loadable=false with admin key → includes drafts
|
||
5. `a5_project_status_metrics_accurate` —
|
||
- Ingest 50 chunks
|
||
- GET /memory/projects/poimen/status
|
||
- Asserts: total_chunks ≈ 50, last_ingest_at is recent, standing_queries count > 0
|
||
6. `a6_rate_limit_enforced` —
|
||
- Set rate limit to 5 req/hour for testing
|
||
- Send 6 GET /query requests
|
||
- First 5 succeed, 6th returns 429
|
||
7. `a7_federation_timeout_partial_results` —
|
||
- Mock slow project (10s response time)
|
||
- Query with timeout_seconds=2
|
||
- Results from fast project, warning about slow project
|
||
8. `a8_no_cascading_failures` —
|
||
- Inject error in embeddings service (simulate 500)
|
||
- GET /query returns 503, not cascading to other endpoints
|
||
- Other endpoints (ingest, skills) still work
|
||
9. `a9_logs_clean_no_panics` —
|
||
- Capture stderr during test
|
||
- Grep for "panic", "unwrap", "expect" — should not appear
|
||
- All errors should be explicit Result types, not crashes
|
||
10. `a10_health_check_always_responds` —
|
||
- Server is under heavy load (rate limit tests, concurrent ingest)
|
||
- GET /health still returns 200 within 100ms
|
||
|
||
**Command:** `cargo test -p mem-cli e2e_api -- --nocapture`
|
||
|
||
**Manual verification (smoke test):**
|
||
```bash
|
||
# Start server
|
||
cargo run -p mem-cli -- serve --port 8080 &
|
||
sleep 2
|
||
|
||
# Ingest via HTTP
|
||
curl -X POST -H "apikey: test" http://localhost:8080/memory/ingest \
|
||
-d '{
|
||
"project": "poimen",
|
||
"source": "manual:smoke",
|
||
"records": [...],
|
||
"ingest_id": "abc123"
|
||
}'
|
||
# → expect 202, job_id
|
||
|
||
# Query
|
||
curl -H "apikey: test" "http://localhost:8080/memory/query?query=test"
|
||
# → expect 200, results array
|
||
|
||
# Skills
|
||
curl -H "apikey: test" http://localhost:8080/memory/skills
|
||
# → expect 200, skills array
|
||
|
||
# Project status
|
||
curl -H "apikey: test" http://localhost:8080/memory/projects/poimen/status
|
||
# → expect 200, metadata
|
||
|
||
# Rate limit test
|
||
for i in {1..11}; do
|
||
curl -H "apikey: test" "http://localhost:8080/memory/query?query=test" \
|
||
-w "HTTP %{http_code}\n"
|
||
done
|
||
# → expect first 10 to succeed, 11th to be 429
|
||
```
|
||
|
||
## False Pass
|
||
|
||
- Testing only happy path (all services available, no errors). Must include:
|
||
- Embedding service down → 503
|
||
- Slow project in federation → partial results + warning
|
||
- Rate limit near boundary (9/10, 10/10, 11/10 reqs)
|
||
- CLI and agent workloads not truly concurrent (sequential test masquerades as parallel). Use `tokio::join_all` or spy on timing to verify parallel execution.
|
||
- Idempotency tested once; never tested with expiry or multiple projects.
|
||
- Metrics never cross-checked against actual DB state. total_chunks reported but not verified against SELECT COUNT.
|
||
- No error injection. If the API is untested with failures, cascading failures are invisible until production.
|
||
|
||
## Traps
|
||
|
||
- Server startup latency: tests must wait for port to be available (sleep or retry logic).
|
||
- Test isolation: if tests share a DB, idempotency cache pollution breaks test N+1. Use separate test DB per test or reset cache between runs.
|
||
- Timing: federation timeout at 2s is tight; if the machine is slow, test becomes flaky. Mock time instead of real delays.
|
||
- Concurrent writes to JSONL log: if two ingest tasks write simultaneously, atomicity of the log is at risk. Ensure the log is single-writer or uses locking.
|
||
|
||
---
|
||
|
||
**Gate outcome:** All M3.5.x tasks green AND a1–a10 pass → M3.5 gate is green. CLI and agents can work with the API concurrently without blocking, race conditions, or idempotency issues.
|
||
|
||
Background: [DESIGN.md § Distributed API Layer](../DESIGN.md#distributed-api-layer-homelab-frontend)
|