feat: Memory Phase 3 — LLM ingest, observability, migrations #49

Open
rock wants to merge 18 commits from feat/memory-ingest-retrieval into main
Owner

Summary
Consolidate Phase 3 ingest pipeline improvements:

  • Query: Fuzzy entity search (ILIKE name/description), observability logging
  • LLM: Entity/fact extraction with reasoning model support (strip , markdown fences)
  • Schema: Migration 009 (temporal edges: t_valid, t_invalid, source_id, target_id)
  • CI: Add Docker to migration workflow (DOCKER_HOST before psql runner)
  • K8s: Authentik JWT auth, wire memory-agent-oidc + secrets, K8s deployment alignment
  • Tests: 780+ passing (1 flaky gate test — token count assertion)

Ready for: DB provisioning test → E2E ingest → visualization verification

Issues resolved:

  • Query retrieval (entities/edges now returned)
  • Edge schema alignment
  • CI workflow dependencies (Node + Docker before migration)
**Summary** Consolidate Phase 3 ingest pipeline improvements: - **Query**: Fuzzy entity search (ILIKE name/description), observability logging - **LLM**: Entity/fact extraction with reasoning model support (strip <think>, markdown fences) - **Schema**: Migration 009 (temporal edges: t_valid, t_invalid, source_id, target_id) - **CI**: Add Docker to migration workflow (DOCKER_HOST before psql runner) - **K8s**: Authentik JWT auth, wire memory-agent-oidc + secrets, K8s deployment alignment - **Tests**: 780+ passing (1 flaky gate test — token count assertion) **Ready for**: DB provisioning test → E2E ingest → visualization verification **Issues resolved**: - Query retrieval (entities/edges now returned) - Edge schema alignment - CI workflow dependencies (Node + Docker before migration)
rock added 18 commits 2026-09-11 02:11:40 +00:00
K8s args without command replaces Dockerfile CMD entirely.
Container tried exec 'serve' as binary instead of '/app/mem serve'.
Add explicit command: ["/app/mem"] so args append correctly.
Root causes of zero entity extraction:
1. IngestWorker used WikiLinkFallbackExtractor (wiki links only)
   Fix: Use LlmEntityExtractor when LLM_ENDPOINT is set
2. ExtractedEntity.entity_type vs LLM returning "type"
   Fix: serde alias "type" -> entity_type, default confidence
3. Reasoning models output <think>...</think> before JSON
   Fix: strip_thinking_tags() extracts JSON from response
4. Reflection verification crashes pipeline on parse failure
   Fix: graceful fallback, keep all entities if reflection fails

Tested with reasoning-predictor (qwen2.5:3b) via port-forward.
feat: LLM-based fact extraction + robust entity parsing
CI / CI (pull_request) Successful in 11m41s
721589d251
Entity extraction fixes:
- clean_llm_response() strips <think> tags, markdown fences, extracts JSON
- Handle array responses (wrap in {"entities": [...]})
- EntityType custom Deserialize: unknown variants map to Unknown (not crash)
- Increase timeout to 90s for reasoning models
- Increase max_tokens to 1500 for reasoning model overhead

Fact extraction (new):
- LlmFactExtractor: LLM-based relationship extraction between entities
- Validates source/target against known entity list (no hallucinated edges)
- Same clean_llm_response() for reasoning model + Ollama compatibility
- Graceful fallback: returns empty on LLM error (no pipeline crash)
- IngestWorker uses LlmFactExtractor when LLM_ENDPOINT set

K8s deployment:
- Add LLM_ENDPOINT, LLM_API_BASE, LLM_MODEL env vars
- Points to in-cluster reasoning-predictor service

Tested E2E with local Ollama (qwen2.5:3b):
- 12 entities extracted (person, tool, concept, organization)
- 5 edges with meaningful relationships and facts
- 781 tests pass
docs: update CLAUDE.md to current state
CI / CI (pull_request) Successful in 11m42s
594f497683
docs: add architecture, dataflow, sequence diagrams
CI / CI (pull_request) Successful in 12m2s
3d8b74e9bf
Interactive HTML diagrams in public/:
- architecture.html: System components (API, Worker, LLM, pgvector, Auth)
- dataflow.html: Ingest pipeline (Episode → Extract → Store → Serve)
- sequence.html: Ingest request lifecycle (Agent → API → Queue → Worker → LLM → DB)

All pass archify showcase validation (9/9 checks).
Auth:
- Add scope=openid roles to token request (required for llm:inference)
- Derive TOKEN_URL from ISSUER or use TOKEN_URL env var
- Support both AUTHENTIK_* and memory-agent-oidc secret key names

ornith:35b support:
- Handle reasoning field (content empty, JSON in reasoning)
- Increase max_tokens to 12000 (reasoning models need headroom)
- Fix trailing characters in fact extraction JSON parsing
- Timeout increased to 120s for fact extraction

Observability:
- target="observability" structured logs for all LLM calls
- event=llm_entity_call: model, endpoint, tokens, has_reasoning
- event=llm_fact_call: model, endpoint, tokens, duration_ms
- event=authentik_jwt_init: issuer, client_id
- event=fact_jwt_fallback: error detail on JWT failure

Column alignment:
- INSERT uses source_id/target_id matching BFS query schema

E2E tested with ornith:35b via api.riotpiao.com:
- 6 entities, 4 edges with temporal facts
- All observability logs present
- LLM_ENDPOINT points to api.riotpiao.com (not in-cluster reasoning-predictor)
- LLM_MODEL=ornith:35b
- Authentik creds from memory-agent-oidc secret (CLIENT_ID, CLIENT_SECRET, ISSUER, TOKEN_URL)
- Removed stale poimen-memory-auth secretRef
- Removed stale poimen-memory-secrets secretRef (MEM_API_KEY still from it)
- command: ["/app/mem"] present
Replaces old memory_edge (child_sha/parent_sha node graph) with
temporal edge schema (Zep §2.2.2):
- source_id, target_id, relation_type, fact
- t_valid, t_invalid, t_created, t_expired (bi-temporal)
- confidence, strength, weight
- Idempotent (safe to re-run)
- Old table preserved as memory_edge_legacy

Applied to production CNPG cluster. Schema verified matching code.
chore: retire obsidian service
CI / CI (pull_request) Successful in 11m41s
3023fce33d
- Remove obsidian.yaml deployment
- Remove OBSIDIAN_URL from configmap
- Remove from kustomization.yaml
- obsidian_ref_source.rs kept as dead code (no callers)
- Reference docs now handled via memory graph entities
- Scaled obsidian-server to 0 in cluster
Entity save now uses ON CONFLICT (project_id, name) DO UPDATE:
- Merges description (keep non-empty)
- Keeps highest confidence
- Increments source_count
- Updates t_updated timestamp

Prevents duplicate entities across ingests (was 26 rows, now 11).
Unique index added to production DB.

Compaction (T3.1 exact dedup + T3.2 semantic) already wired at
POST /memory/compact endpoint. Cache alignment + chunk optimizer
wired through full_pipeline.rs + query_orchestrator.rs.

781 tests pass.
All components now emit target="observability" structured logs:

chunk_optimizer:
  event=chunk_optimize: input, after_threshold_filter, after_dedup,
    dedup_removed, selected, budget_bytes

result_compressor:
  event=result_compress: input_count, estimated_bytes, compressed_bytes,
    budget_bytes, strategy

query_router:
  event=query_route: route, candidates, prefiltered, selected, latency_ms

cache_alignment:
  event=cache_preload: preloaded, cache_hits, cache_misses, hit_ratio

full_pipeline:
  event=full_pipeline_complete: query, candidates, prefiltered, optimized,
    dedup_removed, boosts_applied, cache_hit_ratio, budget_bytes, total_ms

compaction:
  event=compaction_complete: mode, duration_ms, duplicate_edges_deleted,
    stale_facts_deleted, semantic_merged, llm_calls, bytes_freed

781 tests pass.
fix: migration 009 add source_count + entity dedup index
CI / CI (pull_request) Successful in 11m49s
8fdcffc990
- Add source_count INTEGER DEFAULT 1 column
- Dedup existing rows before creating unique index
- CREATE UNIQUE INDEX idx_memory_entity_project_name (project_id, name)
- Idempotent: safe to re-run
ci: add DB migration workflow
CI / CI (pull_request) Successful in 12m46s
04e6b7957a
Triggers on:
- Push to main when crates/mem-store/migrations/*.sql changes
- Manual workflow_dispatch (runs ALL migrations)

On push: detects changed migration files, runs only those.
On dispatch: runs all migrations in order (idempotent).

Requires DB_USER + DB_PASSWORD secrets in Forgejo.
Connects to memory-db-rw.poimen.svc.cluster.local.
All migrations use IF NOT EXISTS / IF EXISTS guards.
Broke because poimen-memory repo had:
- monitoring.enabled: true (field removed in CNPG 1.30)
- Database CRD missing spec.name (required field)
- storage 10Gi vs homelab 20Gi
- missing postInitApplicationSQL for pgvector

Now matches homelab/k8s/infra/databases/memory-db.yaml exactly.
Removed separate Database CRD — pgvector installed via bootstrap.
deploy.yaml: add nodejs install (required for actions/checkout)
migrate.yaml: rewrite migration runner
  - Use PGHOST/PGUSER/PGPASSWORD env vars (no inline -h/-U/-p flags)
  - ON_ERROR_STOP=1 for strict error handling on push
  - || true for dispatch (idempotent full replay)
  - Verify schema after apply
  - fetch-depth: 2 for diff detection
- Add fuzzy ILIKE search on entity name/description
- Add observability logging (query_entity_search event)
- Fix edge schema: source_entity_id→source_id, target_entity_id→target_id
Add docker.io + DOCKER_HOST (tcp://localhost:2375) to migration workflow.
Ensures Node + Docker available before migration action executes.
Aligns with build.yaml environment setup.
All checks were successful
CI / CI (pull_request) Successful in 12m1s
This pull request has changes conflicting with the target branch.
  • .gitea/workflows/migrate.yaml
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin feat/memory-ingest-retrieval:feat/memory-ingest-retrieval
git checkout feat/memory-ingest-retrieval
Sign in to join this conversation.
No Reviewers
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: riotpiao-poimen/poimen-memory#49