chore: Archive final 22 tasks (M5, M6, M7) - all phases now complete
Project roadmap fully scaffolded: - M5 (6): Post-training infrastructure - M6 (6): agent-manager Postgres migration - M7 (10): Extensible source connector framework Total: 78 original tasks → 0 remaining (all COMPLETE)
This commit is contained in:
@@ -1,108 +0,0 @@
|
|||||||
# M5.1 — `mem label` — evidence labeler
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M5 — Post-training |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M5.2, M5.3 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Produce the per-chunk ground truth `U_t` that `r_update` needs, since this corpus
|
|
||||||
does not come with evidence labels.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
Paper `r_update`:
|
|
||||||
|
|
||||||
```
|
|
||||||
r_update_t = +1 if U_t is correct
|
|
||||||
-1 if U_t is incorrect
|
|
||||||
```
|
|
||||||
|
|
||||||
"Correct" means: for chunks containing evidence for `Q`, the agent should emit
|
|
||||||
`<check>yes</check>`; for chunks without, `<check>no</check>`. That requires
|
|
||||||
knowing which chunks contain evidence.
|
|
||||||
|
|
||||||
The paper had it for free — synthetic NIAH tasks place the needle deliberately,
|
|
||||||
and HotpotQA ships supporting facts. **We have neither.** Agent transcripts have
|
|
||||||
no annotation of which turn contained the answer.
|
|
||||||
|
|
||||||
Cheapest honest substitute: **distant supervision from the 32B model.** Ask
|
|
||||||
`reasoning` (DeepSeek-R1-Distill-Qwen-32B, vLLM) per `(question, chunk)` whether
|
|
||||||
the chunk contains evidence. It is ~10× the controller's size and sees each chunk
|
|
||||||
independently, without the memory state that might bias the 3B's decision.
|
|
||||||
|
|
||||||
This inherits the labeler's bias, which is why M5.2 exists and must run before
|
|
||||||
anyone trains on these labels.
|
|
||||||
|
|
||||||
Constraint: `reasoning` has a **16384 total context** and vLLM rejects
|
|
||||||
`input + max_tokens > 16384`. A 5000-token chunk plus question plus instructions
|
|
||||||
fits with room; keep `max_tokens` small (labels are one token of signal) and do
|
|
||||||
not batch chunks into one prompt.
|
|
||||||
|
|
||||||
The labeler emits a binary label plus a short justification. Keep the
|
|
||||||
justification — it is what makes M5.2's disagreement analysis possible.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. `mem label --project P --query Q` reads chunks from the log.
|
|
||||||
2. Per chunk, prompt `reasoning`: question, chunk, "does this contain evidence for
|
|
||||||
the question? Answer yes or no, then one sentence why."
|
|
||||||
3. Send **no tools** — the reasoning route rejects any request carrying them.
|
|
||||||
4. Write `label/<project>/<query-id>.jsonl`:
|
|
||||||
`{"chunk_sha":"...","t":7,"label":true,"why":"...","model":"reasoning","ts":"..."}`.
|
|
||||||
5. Resumable: skip chunks already labelled.
|
|
||||||
6. Report the label rate — the fraction of chunks the labeler calls evidence.
|
|
||||||
Compare it to the controller's update-rate from M1.7; a large gap is the
|
|
||||||
finding, not a bug.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Every chunk in the log gets exactly one label.
|
|
||||||
- Labels key on `chunk_sha`, so they survive re-chunking only if content is
|
|
||||||
unchanged.
|
|
||||||
- Resume skips completed work.
|
|
||||||
- Label rate is reported alongside the controller's update-rate.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** scripted client offline; one `#[ignore]` live run.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_label.rs`:
|
|
||||||
1. `a1_one_label_per_chunk` — no duplicates, no gaps against the log's chunks.
|
|
||||||
2. `a2_keyed_by_sha` — labels reference `chunk_sha`, not `t`, so reordering the
|
|
||||||
log does not corrupt them.
|
|
||||||
3. `a3_no_tools_sent` — assert the request body has no `tools` key.
|
|
||||||
4. `a4_context_budget` — assert every labeling prompt is under
|
|
||||||
16384 − max_tokens.
|
|
||||||
5. `a5_resume` — label, rerun, assert zero new calls.
|
|
||||||
6. `a6_justification_kept` — every label has non-empty `why`.
|
|
||||||
7. `a7_rate_reported` — the summary prints both label rate and the controller's
|
|
||||||
update-rate.
|
|
||||||
8. `a8_live` — `#[ignore]`; 20 real chunks through `reasoning`; print the labels
|
|
||||||
and justifications for a human to sanity-check.
|
|
||||||
|
|
||||||
**Command:** `cargo test -p mem-cli label` (add `-- --ignored` for a8)
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Keying labels by `t`. A re-chunk shifts every `t`, the labels silently
|
|
||||||
misalign, and the training set is quietly wrong in a way nothing downstream can
|
|
||||||
detect.
|
|
||||||
- Dropping the justification to save space. M5.2 then has nothing to analyse and
|
|
||||||
the calibration step degenerates into a single agreement number with no way to
|
|
||||||
understand it.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Batching several chunks into one labeling prompt to save calls. The labels
|
|
||||||
become order-dependent and the 16K context is exceeded on the third chunk.
|
|
||||||
- Treating the 32B's labels as ground truth. They are a *proxy*, and M5.2 is the
|
|
||||||
task that measures how good a proxy.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2.1
|
|
||||||
@@ -1,100 +0,0 @@
|
|||||||
# M5.2 — Labeler calibration
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M5 — Post-training |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M5.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Measure how good the proxy is, before training a policy to imitate it.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
M5.1's labels are distant supervision. Training on them without measuring
|
|
||||||
agreement means the policy learns the 32B model's bias and reports it as
|
|
||||||
improvement — and under a single-model deployment there is no competing variant
|
|
||||||
whose divergence would make that visible.
|
|
||||||
|
|
||||||
Method: hand-label a stratified holdout, compare, and report more than accuracy.
|
|
||||||
|
|
||||||
- **Sample size**: 100 chunks is enough to distinguish 0.7 from 0.9 agreement.
|
|
||||||
- **Stratify** by the labeler's own answer — 50 it called evidence, 50 it did
|
|
||||||
not. Random sampling from a corpus where ~5% is evidence gives ~5 positives,
|
|
||||||
and the positive class is the one that matters.
|
|
||||||
- **Report Cohen's κ, not raw agreement.** With a 95/5 class balance, a labeler
|
|
||||||
that always says "no" scores 95% agreement and is useless. κ corrects for
|
|
||||||
chance.
|
|
||||||
- Also report precision and recall on the positive class separately. They fail
|
|
||||||
differently: low recall silently starves memory, low precision pollutes it.
|
|
||||||
|
|
||||||
Disagreements are the artifact. Read them; they are usually either a genuinely
|
|
||||||
ambiguous chunk or a question that is too vague — and the second is fixable at
|
|
||||||
M1.2 and worth much more than a better labeler.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. `mem label sample --project P --n 100 --stratified` writes a blind
|
|
||||||
worksheet — chunk text and question, **no** labeler answer visible.
|
|
||||||
2. Hand-label it. Record the human labels separately.
|
|
||||||
3. `mem label calibrate` joins them; reports agreement, κ, precision, recall,
|
|
||||||
and the confusion matrix.
|
|
||||||
4. Dump all disagreements with both justifications side by side.
|
|
||||||
5. Commit the holdout and the human labels; they are reusable for every future
|
|
||||||
labeler change.
|
|
||||||
6. Gate: κ ≥ 0.6 before the labels are used for training.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Worksheet hides the labeler's answer.
|
|
||||||
- κ, precision, recall and the confusion matrix are all reported.
|
|
||||||
- Disagreements are dumped with justifications.
|
|
||||||
- The holdout is committed and reusable.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** a synthetic labeled set with known agreement, so the statistics
|
|
||||||
themselves are testable.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_calibration.rs`:
|
|
||||||
1. `a1_worksheet_is_blind` — assert the labeler's answer appears nowhere in the
|
|
||||||
output file.
|
|
||||||
2. `a2_stratified` — assert the sample is ~50/50 by labeler answer, not corpus
|
|
||||||
proportional.
|
|
||||||
3. `a3_kappa_correct` — feed a set with hand-computed κ; assert the reported
|
|
||||||
value matches to 3 decimals.
|
|
||||||
4. `a4_kappa_vs_accuracy` — a synthetic all-negative labeler on a 95/5 set:
|
|
||||||
assert accuracy > 0.9 **and** κ ≈ 0. This is the assertion that justifies
|
|
||||||
reporting κ at all.
|
|
||||||
5. `a5_confusion_matrix` — all four cells match hand counts.
|
|
||||||
6. `a6_disagreements_dumped` — count equals off-diagonal total; each carries both
|
|
||||||
justifications.
|
|
||||||
7. `a7_holdout_stable` — rerunning the sampler with the same seed reproduces the
|
|
||||||
same chunks.
|
|
||||||
|
|
||||||
**Command:** `cargo test -p mem-cli calibration`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Reporting accuracy only. On this class balance it is nearly meaningless, and it
|
|
||||||
will look excellent right up until the trained policy learns to always answer
|
|
||||||
"no".
|
|
||||||
- Hand-labeling with the labeler's answer visible. Anchoring makes agreement look
|
|
||||||
high and the whole exercise decorative — assertion 1 is a real safeguard, not
|
|
||||||
hygiene.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Sampling proportionally. A 5%-positive corpus yields five positives in a
|
|
||||||
hundred, and precision on the positive class — the number that decides whether
|
|
||||||
memory gets polluted — is estimated from five examples.
|
|
||||||
- Treating low κ as "the labeler needs a better prompt". Check the *questions*
|
|
||||||
first: an ambiguous standing question makes evidence genuinely undecidable, and
|
|
||||||
no labeler can fix that.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — Risks
|
|
||||||
@@ -1,114 +0,0 @@
|
|||||||
# M5.3 — Training corpus export
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M5 — Post-training |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M5.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Turn the log plus labels into trajectories verl can train on — the boundary
|
|
||||||
between the Rust side and the Python side.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
**The JSONL log is the boundary.** Rust produces it; Python consumes it. Nothing
|
|
||||||
else crosses, which is what keeps the two halves independent.
|
|
||||||
|
|
||||||
A training example is a **trajectory**, not a turn: the paper's advantage mixes a
|
|
||||||
trajectory-level term with a turn-level one (`Â = α·Â_traj + (1−α)·Â_turn`,
|
|
||||||
α=0.9), so turns must stay grouped by run.
|
|
||||||
|
|
||||||
Per turn, verl needs: the exact prompt sent, the exact response, and the rewards.
|
|
||||||
|
|
||||||
```jsonl
|
|
||||||
{"trajectory_id":"01HXYZ","turns":[
|
|
||||||
{"t":1,"prompt":"<full assembled prompt>","response":"<think>...</next>",
|
|
||||||
"r_update":-1,"parsed":true},
|
|
||||||
...],
|
|
||||||
"r_exit":-0.5,"r_format":1,"r_outcome":null}
|
|
||||||
```
|
|
||||||
|
|
||||||
Reward assembly, from the paper:
|
|
||||||
|
|
||||||
- `r_update_t` = +1 if the recorded `U_t` matches M5.1's label, −1 otherwise.
|
|
||||||
- `r_exit` — one value per trajectory: `0` if `t_exit == t_last_evidence`,
|
|
||||||
`−0.75` if earlier, `−0.5` if later. `t_last_evidence` is the largest `t` whose
|
|
||||||
label is true. **Note L1 runs never exit** (exit gate off), so every L1
|
|
||||||
trajectory is a "late" exit at −0.5 unless the exit signal is taken from the
|
|
||||||
recorded `E_t` rather than the loop's behaviour — take it from the record.
|
|
||||||
- `r_format` = 1 only if **every** turn in the trajectory parsed, 0 otherwise.
|
|
||||||
Strict, because a malformed turn may be caused by the previous one.
|
|
||||||
- `r_outcome` is null. We have no answer-correctness signal; the paper's
|
|
||||||
`is_equiv(A, Â)` has no analogue in extraction. Say so explicitly rather than
|
|
||||||
fabricating one.
|
|
||||||
|
|
||||||
The prompt must be the **exact bytes sent**, not re-assembled. Re-assembly drifts
|
|
||||||
from what the model actually saw the moment M1.3's template changes.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. `mem export --project P --format verl --out corpus/`.
|
|
||||||
2. Join log turns to labels by `chunk_sha`.
|
|
||||||
3. Reconstruct each turn's prompt from the recorded request if `MEM_LLM_RECORD`
|
|
||||||
captured it; otherwise fail loudly rather than re-assembling.
|
|
||||||
4. Compute rewards as above; carry `r_outcome: null` through.
|
|
||||||
5. Group by run into trajectories, ordered by `t`.
|
|
||||||
6. Emit a summary: trajectories, turns, positive/negative `r_update` split,
|
|
||||||
`r_format` pass rate, `t_last_evidence` distribution.
|
|
||||||
7. Refuse to export if M5.2's κ is below the threshold or absent.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Turns are grouped into trajectories, ordered.
|
|
||||||
- Prompts are byte-exact recordings, never re-assembled.
|
|
||||||
- `r_format` is 0 for a trajectory with any unparsed turn.
|
|
||||||
- Export is refused without calibration.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** log fixture with known labels and a deliberately unparsed turn.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_export.rs`:
|
|
||||||
1. `a1_trajectory_grouping` — turns grouped by run, `t` ascending, none lost.
|
|
||||||
2. `a2_r_update_signs` — matching label → +1, mismatching → −1, checked per turn.
|
|
||||||
3. `a3_r_format_strict` — a trajectory with one unparsed turn scores 0 overall,
|
|
||||||
not per turn.
|
|
||||||
4. `a4_r_exit_from_record` — assert `r_exit` derives from the recorded `E_t`, not
|
|
||||||
from whether the loop stopped. A fixture where the gate said `end` at t=5 but
|
|
||||||
the loop continued must score as an exit at 5.
|
|
||||||
5. `a5_prompt_is_recorded_bytes` — assert the exported prompt equals the recorded
|
|
||||||
request body; corrupt the recording and assert export fails rather than
|
|
||||||
silently re-assembling.
|
|
||||||
6. `a6_r_outcome_null` — assert the field is present and null, not omitted and not
|
|
||||||
zero.
|
|
||||||
7. `a7_refuses_without_calibration` — no κ file → non-zero exit naming M5.2.
|
|
||||||
8. `a8_summary_counts` — reported splits match a hand count on the fixture.
|
|
||||||
|
|
||||||
**Command:** `cargo test -p mem-cli export`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Re-assembling prompts at export time. Every assertion except 5 passes, and the
|
|
||||||
policy is trained on prompts the model never saw — which shows up as a training
|
|
||||||
run that will not converge, with no obvious cause.
|
|
||||||
- Applying `r_format` per turn. It looks more granular and it is wrong: the paper
|
|
||||||
is explicit that the strictness exists because a bad turn may be caused by the
|
|
||||||
previous one.
|
|
||||||
- Emitting `r_outcome: 0` instead of null. Zero is a real reward value and the
|
|
||||||
trainer will use it as signal.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Deriving `r_exit` from loop behaviour at L1. The gate is switched off there, so
|
|
||||||
every trajectory scores as a late exit and the exit signal becomes constant
|
|
||||||
noise the policy cannot learn from.
|
|
||||||
- Joining labels by `t`. Same failure as M5.1's trap — a re-chunk misaligns
|
|
||||||
everything silently.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2
|
|
||||||
@@ -1,102 +0,0 @@
|
|||||||
# M5.4 — vLLM + `--enable-lora`
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M5 — Post-training |
|
|
||||||
| Size | L — 3+ days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | homelab |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | — |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
A serving path that can load the memory adapter — because the current one cannot.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
**Ollama cannot hot-swap LoRA adapters.** The controller runs on
|
|
||||||
`qwen2.5:3b-instruct` served by Ollama today, which is fine for prompted-only use
|
|
||||||
and a dead end for post-training. vLLM supports `--enable-lora` with
|
|
||||||
`--lora-modules name=path`, serving base plus adapters from one resident model.
|
|
||||||
|
|
||||||
The pattern already exists in this cluster: the `reasoning` predictor is
|
|
||||||
`vllm/vllm-openai:v0.11.0` under KServe, adopted into ArgoCD as
|
|
||||||
`k8s/apps/llm-serving/reasoning.yaml`. Copy its shape.
|
|
||||||
|
|
||||||
**VRAM is the constraint that makes an adapter the right answer.** One GPU,
|
|
||||||
`OLLAMA_MAX_LOADED_MODELS=2`, currently holding `ornith:35b` + `qwen2.5:3b`. A
|
|
||||||
separate full memory model evicts something, and eviction is a weights reload
|
|
||||||
measured in tens of seconds — `ornith`'s cold start already blew a 60s gateway
|
|
||||||
timeout once. A LoRA rides on a resident base for near-zero extra VRAM.
|
|
||||||
|
|
||||||
Two hard-won operational facts to carry over:
|
|
||||||
|
|
||||||
1. **Kong reads timeouts from the Service, not the Ingress.** `konghq.com/read-timeout`
|
|
||||||
on an Ingress is ignored; it must reach the predictor Service, and KServe
|
|
||||||
propagates InferenceService annotations there. Getting this wrong produces a
|
|
||||||
504 at exactly 60s on the first cold request.
|
|
||||||
2. **Readiness must mean "can serve", not "process is up."** vLLM's startup probe
|
|
||||||
needs a long `failureThreshold` — model load plus torch compile measured 108s
|
|
||||||
+ 55s on the 32B. Report ready too early and the first request 504s.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. `k8s/apps/llm-serving/memory.yaml` — InferenceService, vLLM, Qwen2.5-3B-Instruct.
|
|
||||||
2. Args: `--enable-lora`, `--max-lora-rank 32`, `--max-model-len 32768`,
|
|
||||||
`--served-model-name memory`.
|
|
||||||
3. Adapter storage: a PVC or an initContainer fetching from object storage;
|
|
||||||
`--lora-modules memory-v1=/mnt/adapters/memory-v1`.
|
|
||||||
4. Kong timeout annotations on the **InferenceService** metadata so KServe
|
|
||||||
propagates them to the Service.
|
|
||||||
5. Startup probe with a generous `failureThreshold`, gated on the OpenAI
|
|
||||||
`/health` endpoint.
|
|
||||||
6. New Kong route `/v1/memory/chat/completions` in `llm-routes.yaml`, with the
|
|
||||||
`model-key-auth` plugin — in namespace `llm-serving`, since a KongPlugin
|
|
||||||
reference resolves in the annotated object's own namespace and a dangling one
|
|
||||||
fails open.
|
|
||||||
7. Commit, push, let ArgoCD sync. No `kubectl apply`.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Base model answers through `/v1/memory/chat/completions`.
|
|
||||||
- A named adapter is selectable via the `model` field.
|
|
||||||
- Unauthenticated requests to the new route return 401.
|
|
||||||
- Cold start does not 504.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** `kubectl` and `curl` against the live gateway after sync.
|
|
||||||
|
|
||||||
**Integration test** — `verify/m5.4.sh` diffed against `expected/m5.4.txt`:
|
|
||||||
1. `a1_isvc_ready` — InferenceService reports Ready.
|
|
||||||
2. `a2_base_completion` — a completion naming the base model returns 200.
|
|
||||||
3. `a3_adapter_selectable` — with a dummy adapter mounted, request `model:
|
|
||||||
memory-v1`; assert 200 and that it differs from the base response.
|
|
||||||
4. `a4_auth_enforced` — no key → 401; `apikey` header → 200.
|
|
||||||
5. `a5_cold_start_no_504` — delete the pod, wait for Ready, immediately send a
|
|
||||||
request; assert 200, not 504. This is the regression test for the timeout bug.
|
|
||||||
6. `a6_timeouts_on_service` — assert `konghq.com/read-timeout` is present on the
|
|
||||||
predictor **Service**, not only the Ingress.
|
|
||||||
7. `a7_vram_headroom` — with the memory model resident alongside the others,
|
|
||||||
assert `nvidia-smi` free memory stays above a threshold.
|
|
||||||
|
|
||||||
**Command:** `bash verify/m5.4.sh | diff - expected/m5.4.txt`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing after the model is warm. The 504 bug only appears on the first request
|
|
||||||
after a restart, which is exactly the case assertion 5 forces.
|
|
||||||
- Checking the timeout annotation on the Ingress. It is ignored there, and its
|
|
||||||
presence is what makes the bug hard to see.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Adding a third resident model without checking VRAM. Something gets evicted,
|
|
||||||
and the symptom is a slow unrelated model rather than an obvious failure.
|
|
||||||
- Putting the KongPlugin in the wrong namespace. It fails **open** — the route
|
|
||||||
serves unauthenticated and looks healthy, which is how the model routes ran
|
|
||||||
with no auth for a while.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — Separate weights · `k8s/apps/llm-serving/`
|
|
||||||
@@ -1,120 +0,0 @@
|
|||||||
# M5.5 — verl training loop
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M5 — Post-training |
|
|
||||||
| Size | L — 3+ days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M5.3, M5.4 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Train the LoRA that makes the update and exit gates better than prompting alone.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
Python, using verl (`github.com/volcengine/verl`), which is what the paper used.
|
|
||||||
Lives outside the Rust workspace; the JSONL corpus is the only interface.
|
|
||||||
|
|
||||||
Paper hyperparameters, Table 3 — start here rather than guessing:
|
|
||||||
|
|
||||||
```
|
|
||||||
chunk size 5000
|
|
||||||
max prompt length 8192
|
|
||||||
max response length 2048
|
|
||||||
clip ratio 0.20
|
|
||||||
learning rate 1e-6
|
|
||||||
temperature (train) 1.0 top_p 1.0
|
|
||||||
temperature (val) 1.0 top_p 0.7
|
|
||||||
train batch size 128
|
|
||||||
rollout N 16
|
|
||||||
mini batch size 128
|
|
||||||
LR warmup steps 20
|
|
||||||
```
|
|
||||||
|
|
||||||
Advantage, paper §3.2.2 — the part most likely to be implemented wrong:
|
|
||||||
|
|
||||||
```
|
|
||||||
Â_traj = r_traj_g − mean over the GROUP of trajectories
|
|
||||||
Â_turn = r_update_{g,t} − mean over turns AT STEP t across groups
|
|
||||||
 = α·Â_traj + (1−α)·Â_turn α = 0.9
|
|
||||||
```
|
|
||||||
|
|
||||||
Two distinct baselines. `Â_turn` is normalised across groups **at the same `t`**,
|
|
||||||
and the group size at step `t` can differ from the trajectory group size, because
|
|
||||||
trajectories that exited early have fewer turns.
|
|
||||||
|
|
||||||
α=0.9 is the paper's default and the ablation (Figure 8) explains why: at α=1
|
|
||||||
there is no update-gate reward and accuracy on evidence-free chunks collapses —
|
|
||||||
the model updates indiscriminately, which is exactly the failure this whole
|
|
||||||
system exists to avoid.
|
|
||||||
|
|
||||||
Expect instability. The paper's own limitations section says the extra rewards
|
|
||||||
"reduce training stability, requiring a smaller off-policy degree and longer
|
|
||||||
convergence time."
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. `training/` directory, Python, `uv`-managed. System Python is 3.9.6; this
|
|
||||||
needs 3.11+.
|
|
||||||
2. Corpus loader for M5.3's format.
|
|
||||||
3. Configure verl for LoRA on Qwen2.5-3B-Instruct, rank 16–32.
|
|
||||||
4. Implement the three rewards and the two-baseline advantage, α configurable.
|
|
||||||
5. Log per step: update accuracy split by evidence-present and evidence-free,
|
|
||||||
exact-exit ratio, format correctness, mean response length, validation reward.
|
|
||||||
6. Hold out a validation split by **project**, not by trajectory — same-project
|
|
||||||
trajectories share vocabulary and leak.
|
|
||||||
7. Export the adapter, version it, publish where M5.4 can mount it.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Training runs to convergence on the validation reward.
|
|
||||||
- Both advantage terms are computed with their own baselines.
|
|
||||||
- Update accuracy on evidence-free chunks does not collapse.
|
|
||||||
- The adapter loads in M5.4's server.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** pytest for reward and advantage maths; a short training run for the
|
|
||||||
loop itself.
|
|
||||||
|
|
||||||
**Integration test** — `training/tests/test_rewards.py`:
|
|
||||||
1. `a1_r_update_signs` — matching label +1, mismatching −1.
|
|
||||||
2. `a2_r_exit_bands` — exact 0, early −0.75, late −0.5.
|
|
||||||
3. `a3_r_format_strict` — any unparsed turn zeroes the whole trajectory.
|
|
||||||
4. `a4_traj_baseline` — `Â_traj` uses the group mean; hand-computed fixture.
|
|
||||||
5. `a5_turn_baseline_at_step_t` — `Â_turn` normalises across groups at the same
|
|
||||||
`t`; a fixture with unequal trajectory lengths must not misalign. This is the
|
|
||||||
assertion that catches the most likely implementation error.
|
|
||||||
6. `a6_alpha_mix` — α=1 yields pure trajectory advantage; α=0 pure turn.
|
|
||||||
7. `a7_alpha_1_degenerates` — train 50 steps at α=1 on a fixture; assert
|
|
||||||
evidence-free accuracy drops relative to α=0.9, reproducing the paper's
|
|
||||||
Figure 8b.
|
|
||||||
8. `a8_validation_split_by_project` — assert no project appears in both splits.
|
|
||||||
9. `a9_adapter_loads` — export, mount in M5.4, assert a completion returns 200.
|
|
||||||
|
|
||||||
**Command:** `uv run pytest training/tests -v`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Normalising `Â_turn` over the whole batch rather than per step `t`. It trains,
|
|
||||||
loss goes down, and the turn-level signal is diluted into noise — assertion 5
|
|
||||||
is the only thing that catches it.
|
|
||||||
- Splitting validation by trajectory. Same-project trajectories share phrasing
|
|
||||||
and file paths, so validation reward looks excellent and generalisation is
|
|
||||||
untested.
|
|
||||||
- Skipping assertion 7 as "too slow". It is the only end-to-end evidence that the
|
|
||||||
update reward is wired to anything.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Tuning α before the rewards are verified. Every α is wrong if `r_update`'s sign
|
|
||||||
is flipped, and the symptom looks identical.
|
|
||||||
- Training on labels whose κ was never measured. The policy learns the labeler,
|
|
||||||
and there is no held-out signal that would reveal it — M5.2 exists for this and
|
|
||||||
M5.3 refuses to export without it.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — P6 · paper §3.2.2, Table 3, Fig 8
|
|
||||||
@@ -1,106 +0,0 @@
|
|||||||
# M5.6 — M5 composition gate
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M5 — Post-training |
|
|
||||||
| Size | L — 3+ days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | gate |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | all of M5 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Establish that the trained adapter is actually better than prompting — and that
|
|
||||||
"better" was measured on data it never saw.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
The comparison is **adapter vs. the prompted baseline**, on a held-out project,
|
|
||||||
using the same corpus and the same prompt.
|
|
||||||
|
|
||||||
The paper's own result sets the expectation: Figure 9 shows RL helps but the
|
|
||||||
prompted workflow already works, and the gains concentrate on harder tasks. So
|
|
||||||
a modest improvement is the realistic success case; a dramatic one is a reason to
|
|
||||||
check for leakage first.
|
|
||||||
|
|
||||||
Metrics, and why each is present:
|
|
||||||
|
|
||||||
| metric | why |
|
|
||||||
|---|---|
|
|
||||||
| update accuracy, evidence-present | recall — does it catch evidence |
|
|
||||||
| update accuracy, evidence-free | precision — does it resist noise. **This is the one that collapses at α=1** |
|
|
||||||
| exact-exit ratio | did the exit gate learn anything |
|
|
||||||
| format correctness | the parser's job got easier or harder |
|
|
||||||
| memory token curve | the Figure 6 shape — flat, or climbing to the cap |
|
|
||||||
| wall clock per run | the paper claims up to 400% speedup with the exit gate |
|
|
||||||
|
|
||||||
**Held-out project is non-negotiable.** Same-project trajectories share file
|
|
||||||
paths, error strings and phrasing; measuring on them is measuring memorisation.
|
|
||||||
|
|
||||||
Report both, always: an adapter that improves evidence-present accuracy while
|
|
||||||
degrading evidence-free is worse for this system, because polluted memory
|
|
||||||
degrades every subsequent turn.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Pick a project excluded from training. Ingest it with the prompted baseline;
|
|
||||||
record all metrics.
|
|
||||||
2. Ingest the same project with the adapter, identical prompt and chunking.
|
|
||||||
3. Compare on the M5.2 holdout labels.
|
|
||||||
4. Assert the thresholds below.
|
|
||||||
5. Plot the memory-token curve for both; commit it as the Figure 6 analogue.
|
|
||||||
6. Commit `expected/m5.6.txt`; diff.
|
|
||||||
7. If the adapter loses, keep it versioned and record why — a negative result with
|
|
||||||
the reason is worth more than a rerun with different hyperparameters.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Evidence-present accuracy ≥ baseline.
|
|
||||||
- Evidence-free accuracy ≥ baseline. Never traded away.
|
|
||||||
- Format correctness ≥ baseline.
|
|
||||||
- Memory token curve flat, not climbing to the cap.
|
|
||||||
- Measured on a project absent from training.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** live gateway with both models; the held-out project; M5.2's labels.
|
|
||||||
|
|
||||||
**Integration test** — `verify/m5.6.sh` diffed against `expected/m5.6.txt`:
|
|
||||||
1. `a1_holdout_is_unseen` — assert the eval project appears in no training
|
|
||||||
trajectory. Check the corpus, not the config.
|
|
||||||
2. `a2_evidence_present_accuracy` — adapter ≥ baseline; print both.
|
|
||||||
3. `a3_evidence_free_accuracy` — adapter ≥ baseline; print both.
|
|
||||||
4. `a4_format_correctness` — adapter ≥ baseline.
|
|
||||||
5. `a5_memory_curve_flat` — slope below the M1.8 bound for both; assert the
|
|
||||||
adapter is no worse.
|
|
||||||
6. `a6_exit_ratio_reported` — print exact/early/late exit ratios. Advisory: the
|
|
||||||
exit gate is off at L1, so this measures the signal, not behaviour.
|
|
||||||
7. `a7_same_prompt` — assert both runs used a byte-identical prompt template.
|
|
||||||
8. `a8_wall_clock` — report both; no threshold, since the exit gate is off at L1.
|
|
||||||
|
|
||||||
**Command:** `bash verify/m5.6.sh | diff - expected/m5.6.txt`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Comparing on a project that was in training. Everything improves and none of it
|
|
||||||
generalises — assertion 1 checks the corpus rather than trusting the split
|
|
||||||
config.
|
|
||||||
- Reporting a single combined accuracy. It hides the precision/recall trade that
|
|
||||||
matters most: a model that says "yes" more often scores better on
|
|
||||||
evidence-present and pollutes memory, and the combined number can improve while
|
|
||||||
the system gets worse.
|
|
||||||
- Different prompts between runs. Any measured difference then attributes to the
|
|
||||||
adapter and is partly the prompt — assertion 7 is cheap and removes the doubt.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Retraining until the gate passes without changing anything principled. That is
|
|
||||||
fitting the gate, and the held-out project stops being held out the third time
|
|
||||||
you look at it.
|
|
||||||
- Reading a small improvement as failure. The paper's own baseline works; the
|
|
||||||
adapter's value here is as much about stability on evidence-free chunks as
|
|
||||||
headline accuracy.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — Verification, Risks · paper Fig 9
|
|
||||||
@@ -1,122 +0,0 @@
|
|||||||
# M6.1 — CNPG `agent-manager-db` manifest
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M6 — agent-manager migration |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | homelab |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | — |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
A dedicated Postgres for agent-manager's session store, provisioned the way
|
|
||||||
everything else in the cluster is: through git, with no manual `psql`. Same
|
|
||||||
pattern this project already used for `memory-db` (M2.2), applied to a
|
|
||||||
different, unrelated app.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
**agent-manager is a separate repo**, not part of this Rust workspace:
|
|
||||||
`github.com/Riotpiaole/agent-manager` (fork of `github.com/YoanWai/
|
|
||||||
agent-manager`), local checkout `~/workplace/agent-manager`, branch
|
|
||||||
`add-headless-spawn`. Its session store is `internal/store/store.go` — see
|
|
||||||
M6.2/M6.3 for the schema and query port.
|
|
||||||
|
|
||||||
**Existing CNPG pattern, verified on the live cluster: 3 Clusters today, one
|
|
||||||
per app, each ClusterIP-only (not LAN-reachable):**
|
|
||||||
|
|
||||||
| namespace/name | app |
|
|
||||||
|---|---|
|
|
||||||
| `cicd/forgejo-db` | Forgejo |
|
|
||||||
| `iam/authentik-db` | Authentik |
|
|
||||||
| `temporal/temporal-db` | Temporal |
|
|
||||||
|
|
||||||
**No shared/multi-tenant DB cluster** — every app gets its own dedicated
|
|
||||||
CNPG `Cluster`. `agent-manager-db` follows the same rule; it does not join
|
|
||||||
`memory-db` (M2.2's cluster) even though both are new Postgres instances
|
|
||||||
touched by the same person around the same time. Different app, different
|
|
||||||
cluster.
|
|
||||||
|
|
||||||
Follow `k8s/infra/databases/temporal-db.yaml` exactly, same as M2.2 did:
|
|
||||||
`imageName` pinned, `enableSuperuserAccess: false`, `storageClass:
|
|
||||||
longhorn-cnpg`, `enablePodMonitor: true`, control-plane tolerations,
|
|
||||||
`podAntiAffinityType: preferred`.
|
|
||||||
|
|
||||||
**Instance count — open question, default to convention.** Motivation for
|
|
||||||
this whole migration is durability-of-location, not HA (single-machine
|
|
||||||
usage, not a multi-host shared-session requirement). The 3 existing
|
|
||||||
clusters are all 3-instance. Default to 3 instances for consistency with
|
|
||||||
every other app in the cluster rather than special-casing this one to 1;
|
|
||||||
revisit only if resource pressure on the homelab nodes makes it a real
|
|
||||||
tradeoff.
|
|
||||||
|
|
||||||
Storage: session rows are tiny (`sessions`, `groups`, `settings`,
|
|
||||||
`review_*` — no blobs beyond a `snapshot TEXT` pane capture per session).
|
|
||||||
1Gi is generous; no need for `memory-db`'s 10Gi (that one holds
|
|
||||||
768-dim vectors).
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. `k8s/infra/databases/agent-manager-db.yaml` — `Cluster` + `Database`,
|
|
||||||
namespace `agent-manager`, no extensions (plain relational, no
|
|
||||||
pgvector).
|
|
||||||
2. Namespace `agent-manager`, created by the ArgoCD app that owns it.
|
|
||||||
3. Add to the owning kustomization's explicit resource list — an unlisted
|
|
||||||
file is silently dropped with no error and no drift shown (the M2.2
|
|
||||||
task file names this exact trap).
|
|
||||||
4. Commit, push to **both** Forgejo origin and the GitHub mirror — verify
|
|
||||||
which `repoURL` the eventual ArgoCD `Application` for this app actually
|
|
||||||
watches before assuming either push is the one that matters (`kong`
|
|
||||||
app, for example, tracks the GitHub mirror specifically, not Forgejo).
|
|
||||||
5. Let ArgoCD sync. **No `kubectl apply`.**
|
|
||||||
6. Verify the app user can create tables (schema arrives in M6.2, but a
|
|
||||||
throwaway `CREATE TABLE t(id text); DROP TABLE t;` proves connectivity
|
|
||||||
here).
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- `Cluster` reaches `Cluster in healthy state`.
|
|
||||||
- ArgoCD shows the app `Synced/Healthy`.
|
|
||||||
- No manual `psql` was run to get there.
|
|
||||||
- Service is ClusterIP-only — not reachable from the LAN directly (M6.4's
|
|
||||||
nginx route is the only path in).
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** `kubectl` and `psql` read-only checks after sync.
|
|
||||||
|
|
||||||
**Integration test** — `verify/m6.1.sh`, output diffed against
|
|
||||||
`expected/m6.1.txt`:
|
|
||||||
1. `a1_cluster_healthy` — `kubectl get cluster -n agent-manager
|
|
||||||
agent-manager-db` reports all instances ready.
|
|
||||||
2. `a2_clusterip_only` — `kubectl get svc -n agent-manager -o
|
|
||||||
jsonpath='{.items[*].spec.type}'` contains no `LoadBalancer` or
|
|
||||||
`NodePort`.
|
|
||||||
3. `a3_argocd_synced` — the owning app is `Synced/Healthy`.
|
|
||||||
4. `a4_app_user_can_ddl` — as `app`, `CREATE TABLE t(id text); DROP TABLE
|
|
||||||
t;` succeeds.
|
|
||||||
5. `a5_no_lan_route_yet` — connection attempt from outside the cluster
|
|
||||||
network fails at this point in the plan (M6.4 hasn't landed).
|
|
||||||
|
|
||||||
**Command:** `bash verify/m6.1.sh | diff - expected/m6.1.txt`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Confirming sync without checking service type. A `Cluster` can be
|
|
||||||
healthy and `Synced` while someone fat-fingered a `LoadBalancer` type
|
|
||||||
into the manifest, silently violating the "dedicated ingress, not raw
|
|
||||||
LAN IP" network-path decision this whole migration made. Assertion 2 is
|
|
||||||
the guard.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Forgetting the kustomization resource list (same trap M2.2 already
|
|
||||||
named) — file sits in git, ArgoCD reports Synced, objects never exist.
|
|
||||||
- Adding `prune: true` without accounting for CNPG-operator-created
|
|
||||||
children (Services, Secrets, PVCs). M2.2's Traps section already hit
|
|
||||||
this on `llm-serving`; same fix applies here (`prune: false`).
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: `k8s/infra/databases/temporal-db.yaml` · [M2.2](M2.2-memory-db-manifest.md) (same pattern, different app)
|
|
||||||
@@ -1,243 +0,0 @@
|
|||||||
# M6.2 — Postgres schema for agent-manager sessions
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M6 — agent-manager migration |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M6.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
The sqlite schema agent-manager's `internal/store/store.go` builds up
|
|
||||||
through 14 `ALTER TABLE` migrations, collapsed into one Postgres-native
|
|
||||||
schema — with the sqlite workarounds (integer booleans, dual-encoded
|
|
||||||
timestamps) removed rather than carried over.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
Current schema, read directly from `internal/store/store.go` (6 tables, no
|
|
||||||
FK constraints anywhere — every relationship is enforced in Go, not SQL):
|
|
||||||
|
|
||||||
```mermaid
|
|
||||||
erDiagram
|
|
||||||
GROUPS ||--o{ SESSIONS : "group_name (app-level, no FK)"
|
|
||||||
SESSIONS ||--o{ SESSIONS : "parent_id (self-ref, app-level, no FK)"
|
|
||||||
SESSIONS ||--o| REVIEW_TARGETS : "session_id (app-level, no FK)"
|
|
||||||
SESSIONS ||--o{ REVIEW_BASES : "session_id (app-level, no FK)"
|
|
||||||
SESSIONS ||--o| REVIEW_SCOPES : "session_id (app-level, no FK)"
|
|
||||||
SETTINGS {
|
|
||||||
text key PK
|
|
||||||
text value
|
|
||||||
}
|
|
||||||
SESSIONS {
|
|
||||||
text id PK
|
|
||||||
text name
|
|
||||||
text tool
|
|
||||||
text cwd
|
|
||||||
text group_name "app-level FK -> GROUPS.name"
|
|
||||||
text status
|
|
||||||
int archived "bool 0/1 in sqlite"
|
|
||||||
int created_at "unix nanos, dual-encoded in sqlite"
|
|
||||||
int last_status_at "unix nanos"
|
|
||||||
text agent_session_id
|
|
||||||
text pending_inputs "JSON array blob"
|
|
||||||
int pending_claimed "bool 0/1"
|
|
||||||
text launch_prompt
|
|
||||||
int sort_order
|
|
||||||
int acked "bool 0/1"
|
|
||||||
text snapshot
|
|
||||||
text worktree_repo
|
|
||||||
text worktree_branch
|
|
||||||
int agent_launched_at "unix nanos"
|
|
||||||
text retired_agent_session_id
|
|
||||||
text parent_id "app-level FK -> SESSIONS.id, self"
|
|
||||||
}
|
|
||||||
GROUPS {
|
|
||||||
text name PK
|
|
||||||
int sort_order
|
|
||||||
text path
|
|
||||||
int archived "bool 0/1"
|
|
||||||
text worktree
|
|
||||||
}
|
|
||||||
REVIEW_TARGETS {
|
|
||||||
text session_id PK
|
|
||||||
text repo_root
|
|
||||||
}
|
|
||||||
REVIEW_BASES {
|
|
||||||
text session_id PK
|
|
||||||
text repo_root PK
|
|
||||||
text base_ref
|
|
||||||
}
|
|
||||||
REVIEW_SCOPES {
|
|
||||||
text session_id PK
|
|
||||||
text scope
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
Target Postgres DDL:
|
|
||||||
|
|
||||||
```sql
|
|
||||||
CREATE TABLE groups (
|
|
||||||
name TEXT PRIMARY KEY,
|
|
||||||
sort_order INTEGER NOT NULL DEFAULT 0,
|
|
||||||
path TEXT NOT NULL DEFAULT '',
|
|
||||||
archived BOOLEAN NOT NULL DEFAULT FALSE,
|
|
||||||
worktree TEXT NOT NULL DEFAULT ''
|
|
||||||
);
|
|
||||||
|
|
||||||
CREATE TABLE sessions (
|
|
||||||
id TEXT PRIMARY KEY,
|
|
||||||
name TEXT NOT NULL,
|
|
||||||
tool TEXT NOT NULL,
|
|
||||||
cwd TEXT NOT NULL,
|
|
||||||
group_name TEXT NOT NULL REFERENCES groups(name),
|
|
||||||
status TEXT NOT NULL,
|
|
||||||
archived BOOLEAN NOT NULL DEFAULT FALSE,
|
|
||||||
created_at TIMESTAMPTZ NOT NULL,
|
|
||||||
last_status_at TIMESTAMPTZ NOT NULL,
|
|
||||||
agent_session_id TEXT NOT NULL DEFAULT '',
|
|
||||||
pending_inputs JSONB NOT NULL DEFAULT '[]',
|
|
||||||
pending_claimed BOOLEAN NOT NULL DEFAULT FALSE,
|
|
||||||
launch_prompt TEXT NOT NULL DEFAULT '',
|
|
||||||
sort_order INTEGER NOT NULL DEFAULT 0,
|
|
||||||
acked BOOLEAN NOT NULL DEFAULT FALSE,
|
|
||||||
snapshot TEXT NOT NULL DEFAULT '',
|
|
||||||
worktree_repo TEXT NOT NULL DEFAULT '',
|
|
||||||
worktree_branch TEXT NOT NULL DEFAULT '',
|
|
||||||
agent_launched_at TIMESTAMPTZ,
|
|
||||||
retired_agent_session_id TEXT NOT NULL DEFAULT '',
|
|
||||||
parent_id TEXT REFERENCES sessions(id)
|
|
||||||
);
|
|
||||||
|
|
||||||
CREATE TABLE settings (
|
|
||||||
key TEXT PRIMARY KEY,
|
|
||||||
value TEXT NOT NULL
|
|
||||||
);
|
|
||||||
|
|
||||||
CREATE TABLE review_targets (
|
|
||||||
session_id TEXT PRIMARY KEY REFERENCES sessions(id) ON DELETE CASCADE,
|
|
||||||
repo_root TEXT NOT NULL
|
|
||||||
);
|
|
||||||
|
|
||||||
CREATE TABLE review_bases (
|
|
||||||
session_id TEXT NOT NULL REFERENCES sessions(id) ON DELETE CASCADE,
|
|
||||||
repo_root TEXT NOT NULL,
|
|
||||||
base_ref TEXT NOT NULL,
|
|
||||||
PRIMARY KEY (session_id, repo_root)
|
|
||||||
);
|
|
||||||
|
|
||||||
CREATE TABLE review_scopes (
|
|
||||||
session_id TEXT PRIMARY KEY REFERENCES sessions(id) ON DELETE CASCADE,
|
|
||||||
scope TEXT NOT NULL
|
|
||||||
);
|
|
||||||
```
|
|
||||||
|
|
||||||
**Decision: add real FKs, not app-level-only.** Current Go code manually
|
|
||||||
cascades `review_targets`/`review_bases`/`review_scopes` deletes before
|
|
||||||
deleting a `sessions` row (`Delete`, `DeleteChild` in store.go) — `ON
|
|
||||||
DELETE CASCADE` on those three removes that manual bookkeeping entirely.
|
|
||||||
`group_name -> groups(name)` and `parent_id -> sessions(id)` get FKs too,
|
|
||||||
since Postgres makes them free and they catch the exact class of bug
|
|
||||||
`validParent`/`ensureGroup` exist in Go to prevent by hand. This is a
|
|
||||||
behavior change from sqlite (constraint violation now possible on write
|
|
||||||
paths that previously just wrote garbage) — M6.3 must add error handling
|
|
||||||
for FK violations at the ~4 call sites that write `group_name`/
|
|
||||||
`parent_id` without having already validated the reference through
|
|
||||||
`validParent`/`ensureGroup`.
|
|
||||||
|
|
||||||
**Booleans** — `BOOLEAN`, not `INTEGER`. Removes `boolToInt()`/`!= 0` at
|
|
||||||
every read/write site in store.go.
|
|
||||||
|
|
||||||
**Timestamps** — `TIMESTAMPTZ`, not the `encodeTime`/`decodeTime`
|
|
||||||
nanosecond-or-legacy-seconds hack (store.go:1369-1396). That hack exists
|
|
||||||
only because sqlite has no native timestamp type and old rows needed a
|
|
||||||
seconds/nanos disambiguation heuristic (`secondsCeiling`). A fresh
|
|
||||||
Postgres database has no legacy rows — the whole function pair is deleted,
|
|
||||||
not ported. `agent_launched_at` becomes nullable (`NULL` = "never
|
|
||||||
restarted") instead of the sqlite sentinel `0`.
|
|
||||||
|
|
||||||
**`pending_inputs`** — `JSONB`, not `TEXT` holding a JSON string. Same
|
|
||||||
`json.Marshal`/`Unmarshal` round-trip in Go either way, but `JSONB` lets
|
|
||||||
Postgres validate the shape on write instead of accepting malformed JSON
|
|
||||||
that only fails on the next read.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Write `agent-manager` migration files (this repo's `sqlx`-migration
|
|
||||||
convention doesn't apply — agent-manager is a separate Go repo; use
|
|
||||||
whatever migration tool its `add-headless-spawn` branch already has, or
|
|
||||||
a plain `.sql` file run once at `Open()` if it has none — check before
|
|
||||||
introducing a new dependency).
|
|
||||||
2. `CREATE TABLE` in FK-dependency order: `groups`, then `sessions`
|
|
||||||
(references `groups` and itself), then `settings`,
|
|
||||||
`review_targets`/`review_bases`/`review_scopes`.
|
|
||||||
3. No indexes beyond the primary keys are needed yet — `ListSessions`
|
|
||||||
filters on `archived` and orders by `group_name, sort_order,
|
|
||||||
created_at`; add a composite index only if M6.6's gate shows it's slow,
|
|
||||||
not preemptively.
|
|
||||||
4. Confirm `CREATE EXTENSION` is not needed anywhere (plain relational
|
|
||||||
schema, no pgvector) — unlike this project's own `memory-db`.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Schema applies to a clean `agent-manager-db` database.
|
|
||||||
- `archived`/`acked`/`pending_claimed` are `BOOLEAN`, not `INTEGER`.
|
|
||||||
- `created_at`/`last_status_at`/`agent_launched_at` are `TIMESTAMPTZ`.
|
|
||||||
- Deleting a `sessions` row cascades `review_targets`/`review_bases`/
|
|
||||||
`review_scopes` without the Go code doing it manually.
|
|
||||||
- Inserting a session with an unknown `group_name` is rejected by the FK,
|
|
||||||
not silently written.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** disposable Postgres, same image as M6.1's `Cluster`
|
|
||||||
(`ghcr.io/cloudnative-pg/postgresql:16.2`) — or a local `postgres:16`
|
|
||||||
container for fast iteration, since agent-manager's own test suite doesn't
|
|
||||||
need the real cluster.
|
|
||||||
|
|
||||||
**Integration test** — extend agent-manager's existing store tests
|
|
||||||
(`internal/store/*_test.go` already has `timeenc_test.go`; that file is
|
|
||||||
deleted in M6.3, its coverage folded into this schema's tests) with:
|
|
||||||
1. `a1_migrate_clean` — apply to an empty database, assert all 6 tables
|
|
||||||
exist.
|
|
||||||
2. `a2_bool_columns_are_boolean` — `information_schema.columns` reports
|
|
||||||
`boolean` for `archived`, `acked`, `pending_claimed`.
|
|
||||||
3. `a3_timestamp_columns_are_timestamptz` — same, for `created_at`,
|
|
||||||
`last_status_at`, `agent_launched_at`.
|
|
||||||
4. `a4_review_cascade` — insert a session + a `review_targets` row for it,
|
|
||||||
delete the session, assert `review_targets` is empty without a separate
|
|
||||||
`DELETE FROM review_targets` call.
|
|
||||||
5. `a5_group_fk_rejects_unknown` — insert a session with a `group_name`
|
|
||||||
that has no matching `groups` row, assert it's rejected.
|
|
||||||
6. `a6_parent_fk_self_ref` — insert two sessions where the second's
|
|
||||||
`parent_id` points at the first, assert it succeeds; point it at a
|
|
||||||
nonexistent id, assert it's rejected.
|
|
||||||
|
|
||||||
**Command:** `go test ./internal/store/... -run TestSchema`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing the schema against sqlite (leftover `modernc.org/sqlite` import)
|
|
||||||
instead of real Postgres. `BOOLEAN`/`TIMESTAMPTZ`/`JSONB`/FK-cascade
|
|
||||||
behavior all differ or are silently accepted-but-ignored by sqlite —
|
|
||||||
every assertion above becomes meaningless against the wrong engine.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Keeping `ON DELETE CASCADE` off `group_name`/`parent_id` FKs by
|
|
||||||
accident (only adding it to the `review_*` ones). A group delete or
|
|
||||||
session delete then hits an FK violation instead of the graceful
|
|
||||||
"session has terminals of its own; move them out first" error
|
|
||||||
`PlaceSession` already gives — check M6.3 preserves that message instead
|
|
||||||
of leaking a raw constraint-violation error to the CLI.
|
|
||||||
- Making `parent_id` `NOT NULL DEFAULT ''` (matching the sqlite default)
|
|
||||||
instead of nullable. `'' REFERENCES sessions(id)` is never satisfiable
|
|
||||||
except by a literal empty-string-id row, which doesn't exist — every
|
|
||||||
top-level session's insert then fails the FK. Must be `NULL` for "no
|
|
||||||
parent."
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [M6.1](M6.1-agent-manager-db-manifest.md) · `internal/store/store.go` (agent-manager, `add-headless-spawn` branch)
|
|
||||||
@@ -1,160 +0,0 @@
|
|||||||
# M6.3 — store.go query port to Postgres
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M6 — agent-manager migration |
|
|
||||||
| Size | L — 3+ days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M6.2 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Every query in `internal/store/store.go` rewritten against the M6.2 schema,
|
|
||||||
sqlite dropped entirely (decided: full port, not a dual sqlite/Postgres
|
|
||||||
backend — this is single-machine usage, no standalone-without-cluster
|
|
||||||
requirement to preserve).
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
Read directly from `internal/store/store.go` (1397 lines) on the
|
|
||||||
`add-headless-spawn` branch:
|
|
||||||
|
|
||||||
- **65 call sites** (`db.Exec`, `db.QueryRow`, `db.Query`, `tx.Exec`,
|
|
||||||
`tx.QueryRow`) use sqlite `?` positional placeholders. Postgres
|
|
||||||
(`pgx` or `lib/pq`) needs `$1, $2, ...` — mechanical but must be done
|
|
||||||
per-statement since arg count varies 1–16 across sites.
|
|
||||||
- **Upserts are already Postgres-compatible.** All 9 upsert sites
|
|
||||||
(`SetReviewRepo`, `SetReviewBase`, `SetReviewScope`, `SetSetting`,
|
|
||||||
`ensureGroup`, `CreateGroup`, `AddGroup`, `createSession`'s group
|
|
||||||
insert, `PlaceSession`'s group insert) already use `ON
|
|
||||||
CONFLICT(...) DO UPDATE SET ... = excluded....` / `ON CONFLICT(...) DO
|
|
||||||
NOTHING`, valid as-is in Postgres. No rewrite beyond the placeholder
|
|
||||||
swap — this corrects an earlier assumption that these were sqlite
|
|
||||||
`INSERT OR REPLACE` and needed real rework.
|
|
||||||
- `boolToInt()`/`archived != 0`-style conversions at ~8 sites go away —
|
|
||||||
scan `int archived`/`int acked`/`int pendingClaimed`/`int` locals in
|
|
||||||
`ListSessions`, `Get` and replace with `bool` directly.
|
|
||||||
- `encodeTime()`/`decodeTime()` calls at every session-timestamp
|
|
||||||
read/write go away — pass `time.Time` straight through; `pgx` handles
|
|
||||||
`TIMESTAMPTZ` natively. Delete `timeenc_test.go` along with the
|
|
||||||
functions it tests (M6.2 already noted its coverage folds into the
|
|
||||||
schema tests instead).
|
|
||||||
- `db.SetMaxOpenConns(1)` + `PRAGMA journal_mode=WAL`
|
|
||||||
(store.go:76-79) gave sqlite serialized writes for free. Some call
|
|
||||||
sites lean on that implicit serialization — most notably
|
|
||||||
`createSession`'s `sort_order = (SELECT COALESCE(MAX(sort_order)+1, 0)
|
|
||||||
FROM sessions WHERE group_name = ? AND parent_id = ?)` subquery, which
|
|
||||||
races under concurrent Postgres writers with a real connection pool.
|
|
||||||
Wrap it in `SELECT ... FOR UPDATE` inside the existing transaction (the
|
|
||||||
function already opens one) rather than trusting single-writer
|
|
||||||
semantics that no longer exist. Same pattern applies to
|
|
||||||
`ReorderSession`/`SwapSessionOrder`/`ReorderGroup`/`SwapGroupOrder`'s
|
|
||||||
read-then-renumber-then-write sequences.
|
|
||||||
- FK violations are new failure modes M6.2 introduced (group/parent FKs).
|
|
||||||
`createSession` and `PlaceSession` already validate `group_name`/
|
|
||||||
`parent_id` through `validParent`/`ensureGroup` before writing, so those
|
|
||||||
paths shouldn't hit a live constraint in practice — but wrap the actual
|
|
||||||
`INSERT`/`UPDATE` error and translate a `23503` (foreign_key_violation)
|
|
||||||
SQLSTATE into the same descriptive errors the pre-validation already
|
|
||||||
produces, so a race between the check and the write degrades to a clear
|
|
||||||
error instead of a raw driver error reaching the CLI.
|
|
||||||
- `driver: modernc.org/sqlite` import and the blank `_ "modernc.org/
|
|
||||||
sqlite"` in `Open()` are deleted; replaced with `pgx` (`github.com/
|
|
||||||
jackc/pgx/v5/stdlib` for `database/sql` compatibility, keeping the rest
|
|
||||||
of the file's `*sql.DB`-based code unchanged) or a native `pgx.Pool` —
|
|
||||||
pick `pgx/v5/stdlib` unless a later task needs pgx-native features
|
|
||||||
(e.g. `COPY`), since it's the smaller diff against the existing
|
|
||||||
`database/sql` code.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Swap the driver import and `Open()`'s connection string handling
|
|
||||||
(sqlite file path -> Postgres DSN, likely from an env var or flag the
|
|
||||||
CLI already has a slot for — check `cmd/` for how `Open()` is called
|
|
||||||
today).
|
|
||||||
2. Delete `db.SetMaxOpenConns(1)` / `PRAGMA journal_mode=WAL`; size the
|
|
||||||
pool deliberately instead (`SetMaxOpenConns` to something sane for a
|
|
||||||
single-machine client, e.g. 5–10).
|
|
||||||
3. Delete `init()`'s `CREATE TABLE IF NOT EXISTS` + 14-migration list —
|
|
||||||
M6.2's migration owns schema creation now; `Open()` just connects and
|
|
||||||
optionally runs a `schema_version` sanity check.
|
|
||||||
4. Mechanically convert every `?` to `$N` across the 65 call sites, in
|
|
||||||
file order, verifying arg count against placeholder count each time
|
|
||||||
(this is where an off-by-one is easiest to introduce silently).
|
|
||||||
5. Remove `boolToInt`/`archived != 0` conversions; scan struct fields
|
|
||||||
directly as `bool`.
|
|
||||||
6. Remove `encodeTime`/`decodeTime`/`secondsCeiling`; pass `time.Time`
|
|
||||||
directly. Delete `timeenc_test.go`.
|
|
||||||
7. Add `SELECT ... FOR UPDATE` (or equivalent explicit locking) to the
|
|
||||||
sort-order read-then-write sequences named above.
|
|
||||||
8. Add FK-violation (`23503`) error translation at the write sites that
|
|
||||||
can theoretically race past their own pre-validation.
|
|
||||||
9. Run agent-manager's full existing test suite against a real Postgres
|
|
||||||
(M6.1's cluster, or local `postgres:16` for iteration) — every current
|
|
||||||
test should still pass unmodified in intent, only in backing store.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Zero references to `modernc.org/sqlite`, `?` placeholders,
|
|
||||||
`boolToInt`, `encodeTime`/`decodeTime` remain in `internal/store/`.
|
|
||||||
- Full existing store test suite passes against Postgres.
|
|
||||||
- Concurrent `CreateSession` calls (simulated) never produce duplicate
|
|
||||||
`sort_order` values within the same `group_name`/`parent_id`.
|
|
||||||
- A `spawn` CLI round trip (create session, update status, delete) works
|
|
||||||
end-to-end against the M6.1 cluster.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** `internal/store/*_test.go` running against the disposable
|
|
||||||
Postgres from M6.2's harness.
|
|
||||||
|
|
||||||
**Integration test** — extend/rename the existing store test files:
|
|
||||||
1. `a1_no_sqlite_references` — `grep -r "modernc.org/sqlite" internal/`
|
|
||||||
returns nothing.
|
|
||||||
2. `a2_no_bare_placeholders` — `grep -rE '\?[,)]' internal/store/store.go`
|
|
||||||
inside SQL string literals returns nothing (manual review of any
|
|
||||||
`?` that's part of a non-SQL string, e.g. a Go format verb, to avoid a
|
|
||||||
false positive).
|
|
||||||
3. `a3_existing_suite_passes` — `go test ./internal/store/...` green
|
|
||||||
against Postgres.
|
|
||||||
4. `a4_concurrent_sort_order` — spawn N goroutines each calling
|
|
||||||
`CreateSession` into the same group concurrently; assert the resulting
|
|
||||||
`sort_order` values are a dense 0..N-1 permutation with no duplicates.
|
|
||||||
5. `a5_fk_violation_translated` — attempt `PlaceSession` with a
|
|
||||||
`parentID` deleted between the read and the write (simulate via a
|
|
||||||
second connection); assert the returned error is the existing
|
|
||||||
descriptive one, not a raw pgx driver error.
|
|
||||||
6. `a6_spawn_roundtrip` — `spawn` CLI subcommand creates a session,
|
|
||||||
`UpdateStatus`, then `Delete`; assert no error and the row is gone.
|
|
||||||
|
|
||||||
**Command:** `go test ./internal/... -run TestStore -v`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Running the suite against sqlite still (leftover build tag or import)
|
|
||||||
while believing it validated Postgres. Assertion 1 is the guard —
|
|
||||||
without it, the whole task could report green while the driver swap
|
|
||||||
never actually happened.
|
|
||||||
- Skipping assertion 4. A `sort_order` race is invisible in every
|
|
||||||
single-threaded test and only shows up as duplicate/out-of-order
|
|
||||||
sessions under real concurrent use, which is exactly the class of bug
|
|
||||||
removing `SetMaxOpenConns(1)` introduces.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Converting `?` to `$N` by simple find-and-replace in file order without
|
|
||||||
re-checking each statement's actual arg list — a multi-arg statement
|
|
||||||
reordered during earlier edits (e.g. `createSession`'s 15-arg INSERT)
|
|
||||||
will silently bind the wrong value to the wrong column, and Go's
|
|
||||||
`database/sql` won't catch a type-compatible mismatch (e.g. two `TEXT`
|
|
||||||
columns swapped).
|
|
||||||
- Forgetting `pending_inputs` is now `JSONB` (M6.2), not `TEXT` — the
|
|
||||||
existing `json.Marshal`/`Unmarshal` round-trip in
|
|
||||||
`encodePendingInputs`/`pendingInputState` still works unchanged (pgx
|
|
||||||
scans `JSONB` into `[]byte` the same as `TEXT`), but don't add an extra
|
|
||||||
marshal layer thinking the column type changed the Go-side contract.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [M6.2](M6.2-schema-port.md) · `internal/store/store.go`, `internal/spawn/spawn.go` (agent-manager, `add-headless-spawn` branch)
|
|
||||||
@@ -1,122 +0,0 @@
|
|||||||
# M6.4 — nginx TCP routing to `agent-manager-db`
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M6 — agent-manager migration |
|
|
||||||
| Size | S — <1 day |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | homelab |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M6.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
A path from the Mac client running agent-manager to the cluster-internal,
|
|
||||||
ClusterIP-only `agent-manager-db` — through the shared ingress controller,
|
|
||||||
matching this homelab's existing pattern, not a raw LAN IP or a tunnel.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
**Decided, not open:** dedicated ingress routing through nginx, not
|
|
||||||
`kubectl port-forward`/SSH tunnel and not a MetalLB `LoadBalancer` IP.
|
|
||||||
Matches the pattern already used elsewhere in this homelab of routing
|
|
||||||
through the shared ingress controller rather than exposing raw
|
|
||||||
per-service LAN IPs.
|
|
||||||
|
|
||||||
**Postgres is not HTTP.** The standard nginx-ingress `Ingress` resource is
|
|
||||||
HTTP(S)-oriented (host/path routing, TLS termination via SNI on 443).
|
|
||||||
Postgres speaks its own binary wire protocol on 5432. The ingress
|
|
||||||
controller needs `stream {}` block config (TCP/UDP passthrough) or a
|
|
||||||
dedicated `TCP` mode `Service`/`ConfigMap` entry — whichever the specific
|
|
||||||
nginx-ingress deployment in this cluster supports (check
|
|
||||||
`k8s/infra/ingress/` for how it's deployed and whether `tcp-services`
|
|
||||||
ConfigMap wiring already exists for anything else, since this may be the
|
|
||||||
first TCP passthrough case in the cluster).
|
|
||||||
|
|
||||||
**No existing precedent in this homelab** — `forgejo-db`, `authentik-db`,
|
|
||||||
`temporal-db` are all consumed only by pods inside the same cluster over
|
|
||||||
their ClusterIP Service, never from outside. `agent-manager-db` is the
|
|
||||||
first case of an external (Mac) client needing to reach a CNPG cluster,
|
|
||||||
which is why this task exists as dedicated work rather than "just add a
|
|
||||||
Service."
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Confirm how nginx-ingress is deployed in this cluster (`k8s/infra/
|
|
||||||
ingress/`) and whether it already exposes a `tcp-services` ConfigMap
|
|
||||||
or `stream {}` snippet mechanism — ingress-nginx (the community
|
|
||||||
controller) supports TCP passthrough via a `tcp-services` ConfigMap
|
|
||||||
mapping `<external-port>: <namespace>/<service>:<port>`; confirm this
|
|
||||||
is the controller in use before assuming that config shape.
|
|
||||||
2. Pick an external port for Postgres traffic (5432 is already the
|
|
||||||
in-cluster default; an external port distinct from any other exposed
|
|
||||||
service avoids collision — check what's already claimed).
|
|
||||||
3. Add the `tcp-services` (or equivalent) entry routing that external
|
|
||||||
port to `agent-manager-db-rw.agent-manager.svc.cluster.local:5432`
|
|
||||||
(CNPG's read-write Service name convention — confirm against the
|
|
||||||
actual Service name M6.1's `Cluster` generates).
|
|
||||||
4. Expose that port on the ingress controller's `Service`/`LoadBalancer`
|
|
||||||
(this is the one LAN-facing port for this whole feature — the DB
|
|
||||||
itself stays ClusterIP-only, only the ingress controller's existing
|
|
||||||
external IP gains a new port).
|
|
||||||
5. Commit, push to both remotes, verify which `repoURL` the owning
|
|
||||||
ArgoCD `Application` watches (same caveat as M6.1) before assuming a
|
|
||||||
push landed, let ArgoCD sync.
|
|
||||||
6. Test connectivity from the Mac client: `psql
|
|
||||||
postgresql://<user>@<homelab-ingress-host>:<external-port>/
|
|
||||||
agent_manager` (credentials from M6.5).
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- `psql` (or `pgx`) from outside the cluster reaches `agent-manager-db`
|
|
||||||
through the ingress controller's external IP/port.
|
|
||||||
- `agent-manager-db`'s own Service remains ClusterIP-only — no
|
|
||||||
`LoadBalancer`/`NodePort` added to it directly (that would defeat the
|
|
||||||
point of routing through nginx).
|
|
||||||
- TLS/auth on the connection is Postgres's own (`sslmode`, password auth)
|
|
||||||
— nginx `stream {}` passthrough does not terminate or inspect the
|
|
||||||
Postgres protocol, so it adds no auth of its own. Confirm this is
|
|
||||||
acceptable given the homelab's network boundary (LAN-only ingress
|
|
||||||
exposure, not public internet) before treating it as done.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** `psql` from the Mac client (outside the cluster network),
|
|
||||||
plus `kubectl` checks on the ingress controller's config.
|
|
||||||
|
|
||||||
**Integration test** — `verify/m6.4.sh` diffed against
|
|
||||||
`expected/m6.4.txt`:
|
|
||||||
1. `a1_external_connects` — `psql
|
|
||||||
postgresql://app@<ingress-host>:<port>/agent_manager -c 'select 1'`
|
|
||||||
from the Mac client succeeds.
|
|
||||||
2. `a2_db_service_still_clusterip` — `kubectl get svc -n agent-manager
|
|
||||||
agent-manager-db-rw -o jsonpath='{.spec.type}'` is `ClusterIP`.
|
|
||||||
3. `a3_argocd_synced` — the ingress-owning app is `Synced/Healthy` after
|
|
||||||
the config change.
|
|
||||||
4. `a4_wrong_port_refused` — connecting to a random unmapped port on the
|
|
||||||
same ingress host fails (proves the mapping is port-specific, not an
|
|
||||||
accidental catch-all passthrough).
|
|
||||||
|
|
||||||
**Command:** `bash verify/m6.4.sh | diff - expected/m6.4.txt`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing connectivity from inside the cluster (e.g. `kubectl exec` into
|
|
||||||
a pod and `psql` the ClusterIP directly). That was already true before
|
|
||||||
this task and proves nothing about the ingress path — assertion 1 must
|
|
||||||
run from the actual Mac client, outside the cluster network.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Reusing port 5432 externally on the ingress controller's existing
|
|
||||||
external IP if anything else is already listening there (unlikely for
|
|
||||||
Postgres specifically, but worth a `kubectl get svc -n <ingress-ns>
|
|
||||||
<ingress-svc> -o yaml` check before assuming the port is free).
|
|
||||||
- `stream {}` / `tcp-services` config living outside the GitOps-tracked
|
|
||||||
kustomization because it's a ConfigMap edit that "felt like a quick
|
|
||||||
manual fix." Same hard rule as everything else: commit + push +
|
|
||||||
ArgoCD sync, no manual `kubectl apply` to the ingress controller's
|
|
||||||
config.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [M6.1](M6.1-agent-manager-db-manifest.md) · `k8s/infra/ingress/` (nginx-ingress deployment, controller type to confirm)
|
|
||||||
@@ -1,110 +0,0 @@
|
|||||||
# M6.5 — Postgres credentials for the Mac client
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M6 — agent-manager migration |
|
|
||||||
| Size | S — <1 day |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | homelab |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M6.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
The `agent-manager` process running on the Mac gets a Postgres connection
|
|
||||||
string/credentials, managed the same ksops way the rest of this homelab
|
|
||||||
handles secrets — not a password pasted into a local config file or env
|
|
||||||
var by hand.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
**Existing pattern to follow: ksops-managed secret**, same as
|
|
||||||
`model-invoke-apikey` elsewhere in this homelab. CNPG itself already
|
|
||||||
generates an in-cluster Secret for the `app` user
|
|
||||||
(`agent-manager-db-app` by its usual naming convention) — the work here is
|
|
||||||
getting that credential (or a dedicated read-write user, if reusing the
|
|
||||||
CNPG-generated superuser-adjacent `app` credential isn't desired) to a
|
|
||||||
process running outside the cluster, on the Mac.
|
|
||||||
|
|
||||||
**Two credentials touch two different trust boundaries:**
|
|
||||||
- Cluster-internal: CNPG's own generated Secret, already ksops-free
|
|
||||||
(CNPG manages it, not this repo).
|
|
||||||
- Mac client: needs that same username/password (or a separate,
|
|
||||||
narrower-scoped user) delivered to `~/workplace/agent-manager`'s
|
|
||||||
runtime config, via a ksops-encrypted file in git rather than a
|
|
||||||
manually-copied value.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Decide: reuse CNPG's auto-generated `app` user, or create a dedicated
|
|
||||||
`agent_manager_client` role scoped to only the `sessions`/`groups`/
|
|
||||||
`settings`/`review_*` tables (narrower blast radius if the credential
|
|
||||||
ever leaks from a Mac laptop, which is a meaningfully different threat
|
|
||||||
model than a credential that only ever lives inside the cluster).
|
|
||||||
2. If a dedicated role: add it via CNPG's declarative `Database` /
|
|
||||||
`postInitSQL` (or a one-time migration in M6.2's schema setup) rather
|
|
||||||
than a manual `psql` grant.
|
|
||||||
3. Encrypt the resulting connection string (or user/password pair) with
|
|
||||||
ksops, following the exact file layout `model-invoke-apikey` uses.
|
|
||||||
4. Commit the encrypted secret to git (safe — that's the point of ksops),
|
|
||||||
push to both remotes.
|
|
||||||
5. Wire agent-manager's config loading (check `cmd/` / existing config
|
|
||||||
file handling on the `add-headless-spawn` branch) to read the
|
|
||||||
decrypted value at runtime — decide whether decryption happens via a
|
|
||||||
`sops exec-env`-style wrapper the Mac invokes, or a decrypted file
|
|
||||||
materialized once locally and gitignored, matching whatever
|
|
||||||
`model-invoke-apikey`'s consumers already do.
|
|
||||||
6. Verify agent-manager connects using only the ksops-sourced credential
|
|
||||||
— no plaintext password anywhere in the repo or in shell history.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- No Postgres password appears in plaintext in git, in
|
|
||||||
`~/workplace/agent-manager`'s tracked config, or was typed directly
|
|
||||||
into a `kubectl`/`psql` command during setup.
|
|
||||||
- The credential is encrypted with ksops in the same repo location/
|
|
||||||
pattern as `model-invoke-apikey`.
|
|
||||||
- agent-manager on the Mac successfully authenticates through M6.4's
|
|
||||||
ingress route using this credential.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** manual — this is a secret-handling task, not one to automate
|
|
||||||
a fake credential through.
|
|
||||||
|
|
||||||
**Checklist** (no `verify/*.sh`, since scripting a real credential check
|
|
||||||
means committing something that either leaks a real secret or asserts
|
|
||||||
nothing):
|
|
||||||
1. `git grep -i "password"` in both the homelab repo and
|
|
||||||
`~/workplace/agent-manager` shows only ksops-encrypted blobs or
|
|
||||||
references to environment/config loading, never a literal value.
|
|
||||||
2. The ksops secret file's structure matches `model-invoke-apikey`'s
|
|
||||||
(same encryption provider, same key layout) — a side-by-side diff of
|
|
||||||
the YAML structure (not values) confirms this.
|
|
||||||
3. `psql` from the Mac using the decrypted credential (via M6.4's route)
|
|
||||||
succeeds.
|
|
||||||
4. Revoking/rotating the credential (delete the CNPG Secret or drop the
|
|
||||||
dedicated role, if one was created) and re-running step 3 fails
|
|
||||||
cleanly — confirms the client isn't caching or falling back to
|
|
||||||
something else.
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Confirming connectivity once with a credential typed in manually during
|
|
||||||
setup, then wiring the ksops path afterward without re-testing that the
|
|
||||||
*ksops-sourced* value is what actually authenticates. Checklist item 4
|
|
||||||
(rotate and confirm the old path is really gone) is the guard against
|
|
||||||
"it worked, but only because of leftover state."
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Granting the dedicated role (if created) superuser or
|
|
||||||
database-owner-equivalent privileges "to avoid permission errors during
|
|
||||||
setup" and never narrowing it afterward — defeats the reason to create
|
|
||||||
a dedicated role at all.
|
|
||||||
- Storing the decrypted credential in a file the Mac client reads that
|
|
||||||
isn't gitignored, recreating the plaintext-secret problem one directory
|
|
||||||
away from the ksops-encrypted source of truth.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [M6.1](M6.1-agent-manager-db-manifest.md) · `model-invoke-apikey` (ksops pattern reference, this homelab)
|
|
||||||
@@ -1,150 +0,0 @@
|
|||||||
# M6.6 — M6 composition gate
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M6 — agent-manager migration |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | gate |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | all of M6 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Prove the schema, the query port, the network path, and the credentials
|
|
||||||
bind together into one working system — not four pieces that each passed
|
|
||||||
their own task in isolation. This is the property no single M6 task owns:
|
|
||||||
M6.1 proves the cluster is healthy, M6.2 proves the schema applies, M6.3
|
|
||||||
proves the queries work against a local test database, M6.4 proves a raw
|
|
||||||
TCP connection reaches the cluster, M6.5 proves the credential decrypts —
|
|
||||||
none of them alone proves agent-manager, running for real on the Mac,
|
|
||||||
through the real ingress route, with the real ksops credential, against
|
|
||||||
the real cluster, does its actual job: track a session end-to-end without
|
|
||||||
losing or corrupting data.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
**"Bind together smoothly" means two concrete things here, not a vibe:**
|
|
||||||
|
|
||||||
1. **The full network path is exercised, not simulated.** M6.3's tests run
|
|
||||||
against a disposable local Postgres — that validates the SQL, not the
|
|
||||||
route. This gate is the first (and only) task that runs agent-manager
|
|
||||||
unmodified, on the actual Mac, through M6.4's nginx `stream {}` route,
|
|
||||||
authenticating with M6.5's ksops-sourced credential, against M6.1's
|
|
||||||
real cluster.
|
|
||||||
2. **Schema conventions match the rest of this homelab, not just
|
|
||||||
"compiles."** This project's own `memory-db` (M2.x) and
|
|
||||||
`agent-manager-db` (M6.x) are two unrelated Postgres schemas landing in
|
|
||||||
the same cluster around the same time. They don't share data or a
|
|
||||||
cluster (M6.1 already ruled that out), but a reviewer scanning
|
|
||||||
`k8s/infra/databases/` should find the same shape twice: same
|
|
||||||
`Cluster`/`Database` CRD structure, same `storageClass`, same
|
|
||||||
`enableSuperuserAccess: false`, same GitOps-only provisioning
|
|
||||||
discipline. "Binds together smoothly" includes that consistency check,
|
|
||||||
not just agent-manager working in isolation.
|
|
||||||
|
|
||||||
**What a session round trip actually touches**, so the test isn't
|
|
||||||
shallow: `CreateSession` (writes `sessions` + touches `groups` via the
|
|
||||||
`ON CONFLICT DO NOTHING` insert) -> `UpdateStatus` -> `SetAgentSessionID`
|
|
||||||
-> `SetReviewRepo` (writes `review_targets`) -> `Delete` (must cascade
|
|
||||||
`review_targets` via the FK, per M6.2's decision, with no leftover row).
|
|
||||||
That single flow crosses all 4 non-`settings` tables and exercises both
|
|
||||||
the FK-cascade decision and the placeholder-conversion correctness from
|
|
||||||
M6.3 in one pass.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. On the Mac, with agent-manager built from the fully-ported
|
|
||||||
`add-headless-spawn` branch (M6.3 complete) and configured to use
|
|
||||||
M6.4/M6.5's route and credential: run the `spawn` CLI subcommand to
|
|
||||||
create a real session.
|
|
||||||
2. Drive it through the full lifecycle above (status update, agent
|
|
||||||
session id capture, review target set, delete) using agent-manager's
|
|
||||||
own CLI/TUI, not a hand-rolled SQL script — the point is proving the
|
|
||||||
actual client works, not that Postgres accepts hand-written SQL.
|
|
||||||
3. Kill and restart agent-manager mid-lifecycle (after step 2's status
|
|
||||||
update, before delete); confirm it reconnects and reads back the same
|
|
||||||
state — proves the connection isn't accidentally caching state
|
|
||||||
client-side that masks a write that never actually landed.
|
|
||||||
4. Diff `k8s/infra/databases/agent-manager-db.yaml` against
|
|
||||||
`k8s/infra/databases/memory-db.yaml` field-by-field for the
|
|
||||||
convention-consistency check.
|
|
||||||
5. Confirm both ArgoCD Applications (agent-manager's and this project's
|
|
||||||
`memory-db`, once it exists) are tracked from the `repoURL` each
|
|
||||||
actually watches — re-verify per M6.1's caveat, since this is the
|
|
||||||
final point where a "pushed but ArgoCD never saw it" mistake would
|
|
||||||
otherwise go unnoticed until much later.
|
|
||||||
6. Commit `expected/m6.6.txt`; diff.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- A session created, updated, and deleted through agent-manager's real
|
|
||||||
CLI on the Mac round-trips correctly through the full network path.
|
|
||||||
- A mid-lifecycle restart does not lose or duplicate state.
|
|
||||||
- Deleting the session leaves zero orphan rows in `review_targets`/
|
|
||||||
`review_bases`/`review_scopes` (FK cascade, not app-level cleanup).
|
|
||||||
- `agent-manager-db.yaml` and `memory-db.yaml` match on every field that
|
|
||||||
isn't inherently app-specific (name, storage size).
|
|
||||||
- No manual `kubectl apply`/`psql` anywhere in the setup this gate
|
|
||||||
exercises.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** the real Mac client, the real cluster, agent-manager's own
|
|
||||||
CLI — this gate deliberately does not use a disposable/local database,
|
|
||||||
since proving the disposable path works is exactly what M6.1-M6.5 already
|
|
||||||
did.
|
|
||||||
|
|
||||||
**Integration test** — `verify/m6.6.sh` diffed against
|
|
||||||
`expected/m6.6.txt`:
|
|
||||||
1. `a1_full_roundtrip` — create/update/set-review/delete via the real CLI;
|
|
||||||
assert no error at any step.
|
|
||||||
2. `a2_no_orphan_review_rows` — after delete, query `review_targets`/
|
|
||||||
`review_bases`/`review_scopes` directly (from inside the cluster, as a
|
|
||||||
final-state check) for the deleted session's id; assert zero rows,
|
|
||||||
with no explicit `DELETE FROM review_*` having been issued by the CLI
|
|
||||||
(proves the FK cascade did the work, not leftover manual-cleanup code
|
|
||||||
nobody removed).
|
|
||||||
3. `a3_survives_restart` — kill agent-manager between status-update and
|
|
||||||
delete; restart; assert the status update is still visible before
|
|
||||||
proceeding to delete.
|
|
||||||
4. `a4_schema_convention_match` — diff the two `Cluster` manifests' non
|
|
||||||
app-specific fields; assert empty diff.
|
|
||||||
5. `a5_repourl_confirmed` — for each of the two ArgoCD Applications
|
|
||||||
involved, print which `repoURL` it watches and confirm it matches
|
|
||||||
which remote was actually pushed.
|
|
||||||
6. `a6_no_manual_apply_in_history` — review the shell history / session
|
|
||||||
log from M6.1 through M6.5 for a `kubectl apply` or `psql` write
|
|
||||||
command that wasn't inside an explicitly-flagged debugging exception;
|
|
||||||
assert none exist outside that exception.
|
|
||||||
|
|
||||||
**Command:** `bash verify/m6.6.sh | diff - expected/m6.6.txt`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Running assertion 1 against M6.3's disposable local Postgres instead of
|
|
||||||
the real cluster because it's faster/already running. That's exactly
|
|
||||||
the "four pieces that each passed in isolation" failure mode this gate
|
|
||||||
exists to catch — it must hit M6.1's actual cluster through M6.4's
|
|
||||||
actual route.
|
|
||||||
- Treating a schema diff (assertion 4) as advisory and skipping it when
|
|
||||||
short on time. A convention mismatch here is invisible today and
|
|
||||||
becomes the thing a future reviewer trips over when comparing the two
|
|
||||||
`k8s/infra/databases/*.yaml` files months later with no memory of why
|
|
||||||
they differ.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Discovering during this gate that M6.4's nginx route works from inside
|
|
||||||
the homelab LAN but not from wherever the Mac actually sits (VPN,
|
|
||||||
different subnet, etc.) — a gap none of M6.1-M6.5's narrower tests
|
|
||||||
would have caught, since this is the first task that tests from the
|
|
||||||
Mac's actual network position rather than "outside the cluster" in the
|
|
||||||
abstract.
|
|
||||||
- Fixing a gate failure by loosening the gate (e.g. deleting assertion 2
|
|
||||||
because the cascade "mostly works") instead of fixing the underlying
|
|
||||||
FK/migration issue. Same discipline this project's other gates
|
|
||||||
(M0.8, M1.8, M2.8...) already hold to.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [M6.1](M6.1-agent-manager-db-manifest.md) · [M6.2](M6.2-schema-port.md) · [M6.3](M6.3-store-query-port.md) · [M6.4](M6.4-nginx-stream-routing.md) · [M6.5](M6.5-credentials-secret.md)
|
|
||||||
@@ -1,122 +0,0 @@
|
|||||||
# M7.1 — `SourceConnector` trait + registry
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.2, M7.3, M7.4, M7.5, M7.6, M7.10 |
|
|
||||||
| Depends | M0.3, M3.6.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Define the extensible connector interface so that adding a new document source
|
|
||||||
(paperless-ngx, S3, git repo, etc.) requires implementing one trait and adding
|
|
||||||
one YAML config block — no changes to the ingest pipeline, chunking, embedding,
|
|
||||||
storage, or query layers.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
Memory service is a **cluster-wide RAG** serving multiple agents. Knowledge lives
|
|
||||||
in many places: an Obsidian vault, paperless-ngx, git repositories, S3 buckets.
|
|
||||||
The connector abstraction makes all of them look the same to the ingest pipeline.
|
|
||||||
|
|
||||||
**The trait has three methods.** `list_documents()` enumerates what's available
|
|
||||||
without fetching content. `fetch_document()` retrieves one document's text.
|
|
||||||
`health_check()` reports reachability. The sync framework (M7.6) handles
|
|
||||||
everything else — change detection, tombstoning, chunking, embedding.
|
|
||||||
|
|
||||||
**Configuration is YAML-driven.** Each connector instance is a block in
|
|
||||||
`connectors.yaml` with `kind`, `name`, and source-specific `config`. The registry
|
|
||||||
maps `kind` to a factory function that constructs the connector from its config.
|
|
||||||
|
|
||||||
**Two connector families exist.** Session connectors (pi, claude) produce evidence
|
|
||||||
for the gated loop (L0/L1/L2). Document connectors (obsidian, paperless, git, s3)
|
|
||||||
produce reference material at Level R, bypassing the gate. The connector's
|
|
||||||
`source_type()` method declares which family it belongs to.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Define `SourceConnector` trait in `mem-ingest/src/connector.rs`:
|
|
||||||
```rust
|
|
||||||
#[async_trait]
|
|
||||||
pub trait SourceConnector: Send + Sync {
|
|
||||||
fn kind(&self) -> &str;
|
|
||||||
fn name(&self) -> &str;
|
|
||||||
fn source_type(&self) -> SourceType; // Evidence or Reference
|
|
||||||
async fn list_documents(&self) -> Result<Vec<SourceDocument>>;
|
|
||||||
async fn fetch_document(&self, doc_id: &str) -> Result<DocumentContent>;
|
|
||||||
async fn health_check(&self) -> Result<SourceHealth>;
|
|
||||||
}
|
|
||||||
```
|
|
||||||
2. Define supporting types: `SourceDocument`, `DocumentContent`, `SourceHealth`,
|
|
||||||
`SourceType` enum (`Evidence`, `Reference`).
|
|
||||||
3. Define `ConnectorConfig` serde struct for YAML deserialization:
|
|
||||||
```yaml
|
|
||||||
connectors:
|
|
||||||
- kind: obsidian
|
|
||||||
name: homelab-vault
|
|
||||||
config: { root: /data/vault, extensions: [md, txt] }
|
|
||||||
```
|
|
||||||
4. Implement `ConnectorRegistry` — maps `kind` string to factory function,
|
|
||||||
constructs connectors from config at startup.
|
|
||||||
5. Add `connectors.yaml` loading in `mem-cli` startup path.
|
|
||||||
6. Provide a `VecConnector` test helper (in-memory documents) for testing
|
|
||||||
downstream consumers without real I/O.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- The trait compiles and is object-safe (`Box<dyn SourceConnector>`).
|
|
||||||
- `ConnectorRegistry` can register and construct connectors by kind string.
|
|
||||||
- `VecConnector` implements the trait and passes basic list/fetch assertions.
|
|
||||||
- YAML config deserialization works for known and unknown kinds (unknown = skip
|
|
||||||
with warning, not crash).
|
|
||||||
- `source_type()` is enforced at the type level — no runtime flag confusion.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** in-memory `VecConnector`, YAML config fixtures.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_source_connector.rs`:
|
|
||||||
1. `a1_trait_is_object_safe` — construct a `Box<dyn SourceConnector>` from
|
|
||||||
`VecConnector`; call all three methods.
|
|
||||||
2. `a2_registry_constructs_by_kind` — register "vec" kind, construct from config,
|
|
||||||
assert `kind()` and `name()` match.
|
|
||||||
3. `a3_list_documents_returns_all` — `VecConnector` with 3 docs, assert
|
|
||||||
`list_documents()` returns 3.
|
|
||||||
4. `a4_fetch_document_by_id` — assert content matches what was registered.
|
|
||||||
5. `a5_fetch_unknown_id_errors` — assert `fetch_document("nonexistent")` returns
|
|
||||||
an error, not a panic.
|
|
||||||
6. `a6_health_check_reports_count` — assert `health_check()` returns
|
|
||||||
`document_count = Some(3)`.
|
|
||||||
7. `a7_yaml_config_loads` — parse a fixture `connectors.yaml` with two connector
|
|
||||||
blocks; assert both are constructed.
|
|
||||||
8. `a8_unknown_kind_skipped` — config with `kind: "nonexistent"`; assert registry
|
|
||||||
logs a warning and continues without the connector.
|
|
||||||
9. `a9_source_type_evidence_vs_reference` — assert session connectors return
|
|
||||||
`Evidence`, document connectors return `Reference`.
|
|
||||||
10. `a10_empty_config_is_valid` — no `connectors.yaml` or empty file; registry
|
|
||||||
starts with zero connectors, no crash.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_source_connector`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing only `VecConnector` and claiming the trait works. The trait is
|
|
||||||
validated by M7.2–M7.5 implementing it against real sources.
|
|
||||||
- Parsing YAML without verifying the constructed connector's methods work.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Making the trait not object-safe (generic methods, `Self` in return types).
|
|
||||||
Every consumer stores `Box<dyn SourceConnector>`, so object safety is load-bearing.
|
|
||||||
- Putting chunking logic inside the connector. Connectors fetch documents; the
|
|
||||||
sync framework (M7.6) chunks them. Mixing concerns means every connector
|
|
||||||
reimplements chunking.
|
|
||||||
- Hard-coding the list of known kinds. The registry must be extensible — a
|
|
||||||
`register(kind, factory_fn)` call, not a match statement.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors section
|
|
||||||
@@ -1,92 +0,0 @@
|
|||||||
# M7.10 — M7 composition gate — source connectors
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | gate |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | — |
|
|
||||||
| Depends | M7.1, M7.2, M7.3, M7.6, M7.7, M7.8, M7.9 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Prove the connector framework composes: two different connector kinds sync
|
|
||||||
through the same framework, change detection skips unchanged documents, tombstoning
|
|
||||||
works, drift reports are accurate, rebuild parity holds, and the gated loop's
|
|
||||||
update-rate is untouched.
|
|
||||||
|
|
||||||
## What the gate proves
|
|
||||||
|
|
||||||
1. **Extensibility works.** Two different connector kinds (at minimum: obsidian +
|
|
||||||
one remote connector) register, sync, and produce queryable Level R content
|
|
||||||
through the same `SyncEngine`. No connector-specific code exists outside the
|
|
||||||
connector itself.
|
|
||||||
|
|
||||||
2. **Change detection is efficient.** Re-syncing an unchanged connector produces
|
|
||||||
zero embedding calls. This is the cost guard — without it, every sync is a
|
|
||||||
full re-embed.
|
|
||||||
|
|
||||||
3. **Tombstoning is correct.** Removing a document from a source results in
|
|
||||||
tombstone records in the log, removal from the index, and correct behavior
|
|
||||||
on rebuild.
|
|
||||||
|
|
||||||
4. **Drift report is read-only.** Running `mem source status` mutates nothing.
|
|
||||||
|
|
||||||
5. **Rebuild parity holds for connectors.** `mem rebuild --from-log` with
|
|
||||||
connector-sourced data produces byte-identical results.
|
|
||||||
|
|
||||||
6. **Update-rate is untouched.** Adding connector-sourced reference documents
|
|
||||||
does not change the gated loop's update-rate (same property M3.6.6 asserts,
|
|
||||||
now for any connector).
|
|
||||||
|
|
||||||
7. **Health monitoring detects failures.** An unreachable connector is reported,
|
|
||||||
not silently ignored.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_m7_gate.rs`:
|
|
||||||
|
|
||||||
1. `a1_two_kinds_sync` — register an obsidian + vec connector; sync both; assert
|
|
||||||
Level R nodes exist for both sources.
|
|
||||||
2. `a2_unchanged_zero_embeds` — re-sync both; assert zero embedding calls.
|
|
||||||
3. `a3_change_detected_and_replaced` — modify a doc in one connector; sync;
|
|
||||||
assert old chunks tombstoned, new chunks present.
|
|
||||||
4. `a4_removal_tombstoned` — remove a doc; sync; assert tombstone records and
|
|
||||||
doc absent from query results.
|
|
||||||
5. `a5_drift_report_is_read_only` — snapshot log + manifest; run status; assert
|
|
||||||
unchanged.
|
|
||||||
6. `a6_rebuild_parity` — after full sync, rebuild from log; assert byte-identical
|
|
||||||
state.
|
|
||||||
7. `a7_update_rate_untouched` — record update-rate before adding connectors;
|
|
||||||
add connectors + sync; re-run gated loop; assert update-rate unchanged.
|
|
||||||
8. `a8_health_failure_reported` — configure unreachable connector; assert
|
|
||||||
health check reports failure.
|
|
||||||
9. `a9_no_connector_specific_code_in_sync` — assert `SyncEngine` has no
|
|
||||||
`match kind` or `if kind ==` statements (the trait is the dispatch, not
|
|
||||||
the framework).
|
|
||||||
10. `a10_query_returns_connector_content` — sync a connector; query its content;
|
|
||||||
assert results include connector-sourced Level R nodes with correct source_uri.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_m7_gate`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing with only one connector kind. The gate's value is proving two
|
|
||||||
*different* kinds work through the same framework.
|
|
||||||
- Testing rebuild parity without connector data in the log. An empty log
|
|
||||||
trivially rebuilds.
|
|
||||||
- Asserting update-rate "is still below 30%" instead of "is unchanged". Adding
|
|
||||||
reference docs should not move the number at all.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Running the gate before M7.6 (sync framework) is solid. The gate tests
|
|
||||||
composition; if the sync framework has bugs, every gate assertion fails
|
|
||||||
for the wrong reason.
|
|
||||||
- Not testing with a connector that produces multiple chunks per document.
|
|
||||||
Single-chunk documents hide change-detection bugs at the chunk level.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors, composition gates
|
|
||||||
@@ -1,97 +0,0 @@
|
|||||||
# M7.2 — Obsidian vault connector
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.10 |
|
|
||||||
| Depends | M7.1, M3.6.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Wrap the existing `DocCorpusSource` (M3.6.1) in the `SourceConnector` interface
|
|
||||||
so the Obsidian vault is managed through the unified connector framework — with
|
|
||||||
change detection, config-driven setup, and registry integration.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
The Obsidian vault is the **primary reference source** for the cluster. It is
|
|
||||||
human-maintained, optionally git-backed, and contains runbooks, procedures, and
|
|
||||||
domain knowledge that agents query through the memory service.
|
|
||||||
|
|
||||||
`DocCorpusSource` already handles markdown parsing, heading-boundary chunking,
|
|
||||||
breadcrumb paths, and file filtering. This task wraps it, not rewrites it.
|
|
||||||
|
|
||||||
**Deployment models:**
|
|
||||||
1. **Git-sync sidecar** — a sidecar container clones the vault repo into a shared
|
|
||||||
PVC. Memory service reads from the PVC via this connector.
|
|
||||||
2. **Local mount** — for development, mount the vault directory directly.
|
|
||||||
3. **PVC direct** — vault files managed via kubectl cp or a web uploader.
|
|
||||||
|
|
||||||
**Configuration:**
|
|
||||||
```yaml
|
|
||||||
connectors:
|
|
||||||
- kind: obsidian
|
|
||||||
name: homelab-vault
|
|
||||||
config:
|
|
||||||
root: /data/vault
|
|
||||||
extensions: [md, markdown, txt]
|
|
||||||
exclude_dirs: [.obsidian, .trash, .git]
|
|
||||||
max_file_size: 10485760 # 10MB
|
|
||||||
```
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Implement `ObsidianConnector` in `mem-ingest/src/connectors/obsidian.rs`.
|
|
||||||
2. `list_documents()` — walk `root` directory, filter by extension, compute
|
|
||||||
sha256 per file, return `SourceDocument` per file.
|
|
||||||
3. `fetch_document()` — read file content, return as `DocumentContent` with
|
|
||||||
metadata (file path, last modified, size).
|
|
||||||
4. `health_check()` — verify `root` exists, is readable, count files.
|
|
||||||
5. Register `"obsidian"` kind in the connector registry factory.
|
|
||||||
6. `source_type()` returns `Reference` (vault docs bypass the gated loop).
|
|
||||||
7. Reuse `DocCorpusSource` internals for heading-based chunking when the sync
|
|
||||||
framework (M7.6) processes this connector's documents.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- `ObsidianConnector` implements `SourceConnector` fully.
|
|
||||||
- `list_documents()` respects `extensions`, `exclude_dirs`, `max_file_size`.
|
|
||||||
- `fetch_document()` returns content matching file on disk.
|
|
||||||
- `health_check()` distinguishes readable vs. missing root directory.
|
|
||||||
- Config-driven: changing `root` path in YAML changes what gets scanned.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** fixture directory with markdown files, config YAML.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_obsidian_connector.rs`:
|
|
||||||
1. `a1_list_filters_extensions` — fixture with .md, .txt, .json; assert only
|
|
||||||
.md and .txt are listed.
|
|
||||||
2. `a2_list_excludes_dirs` — fixture with `.obsidian/` subdir; assert its files
|
|
||||||
are excluded.
|
|
||||||
3. `a3_fetch_returns_content` — fetch a known doc; assert text matches file.
|
|
||||||
4. `a4_fetch_unknown_errors` — fetch non-existent doc_id; assert error.
|
|
||||||
5. `a5_health_check_reachable` — valid root; assert `reachable: true` with count.
|
|
||||||
6. `a6_health_check_missing_root` — non-existent root; assert `reachable: false`.
|
|
||||||
7. `a7_content_hash_stable` — fetch same file twice; assert same hash.
|
|
||||||
8. `a8_config_from_yaml` — parse connector from YAML; assert fields match.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_obsidian_connector`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing with an empty directory. Assertions 1–3 need real files.
|
|
||||||
- Not testing `exclude_dirs` with nested paths (`.obsidian/plugins/x.md`).
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Re-implementing markdown parsing instead of delegating to `DocCorpusSource`.
|
|
||||||
- Making `doc_id` platform-dependent (use relative path from root, unix separators).
|
|
||||||
- Ignoring symlinks — Obsidian uses them for multi-vault setups.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors, Obsidian vault as connector
|
|
||||||
@@ -1,119 +0,0 @@
|
|||||||
# M7.3 — paperless-ngx connector
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.10 |
|
|
||||||
| Depends | M7.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Implement a `SourceConnector` for paperless-ngx so OCR'd documents, manuals, and
|
|
||||||
reference PDFs already stored in the cluster's paperless instance become
|
|
||||||
searchable through the memory service without manual export.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
paperless-ngx is already running in the cluster. Its REST API provides:
|
|
||||||
|
|
||||||
- `GET /api/documents/` — paginated list with filtering by tags, document type,
|
|
||||||
correspondent, dates.
|
|
||||||
- `GET /api/documents/{id}/` — full metadata including `content` (extracted text).
|
|
||||||
- `GET /api/documents/{id}/download/` — original file.
|
|
||||||
- `GET /api/documents/{id}/preview/` — thumbnail.
|
|
||||||
- Authentication via `Authorization: Token <token>` header.
|
|
||||||
- Documents have `checksum` field (sha256 of original file).
|
|
||||||
|
|
||||||
**Tag filtering is the selection mechanism.** Not every scanned receipt belongs
|
|
||||||
in the knowledge base. Config specifies which tags to include:
|
|
||||||
```yaml
|
|
||||||
connectors:
|
|
||||||
- kind: paperless
|
|
||||||
name: homelab-paperless
|
|
||||||
config:
|
|
||||||
base_url: http://paperless-ngx.paperless.svc.cluster.local:8000
|
|
||||||
token_secret: paperless-api-token # k8s secret ref
|
|
||||||
tags: [reference, manual, runbook] # only sync these
|
|
||||||
format: text # use extracted text content
|
|
||||||
page_size: 100 # API pagination size
|
|
||||||
```
|
|
||||||
|
|
||||||
**Content comes as extracted text.** paperless-ngx OCRs documents on import and
|
|
||||||
stores the text in the `content` field. Use this directly — no PDF parsing needed
|
|
||||||
in the connector. The text quality depends on paperless's OCR config.
|
|
||||||
|
|
||||||
**Checksum enables cheap change detection.** paperless provides `checksum` per
|
|
||||||
document. The sync framework (M7.6) compares this against the last-known hash
|
|
||||||
to skip unchanged documents.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Implement `PaperlessConnector` in `mem-ingest/src/connectors/paperless.rs`.
|
|
||||||
2. `list_documents()` — paginate `GET /api/documents/?tags__name__in=...`,
|
|
||||||
extract `id`, `title`, `checksum`, `modified` for each.
|
|
||||||
3. `fetch_document()` — `GET /api/documents/{id}/`, extract `content` field,
|
|
||||||
return with metadata (title, tags, correspondent, date_created).
|
|
||||||
4. `health_check()` — `GET /api/` and verify 200; report document count from
|
|
||||||
`GET /api/documents/?tags__name__in=...&page=1&page_size=1` (read `count`).
|
|
||||||
5. Handle pagination (paperless returns `next` URL for subsequent pages).
|
|
||||||
6. Token auth from k8s secret (resolve `token_secret` to actual token value).
|
|
||||||
7. Register `"paperless"` kind in connector registry factory.
|
|
||||||
8. `source_type()` returns `Reference`.
|
|
||||||
9. Rate limit API calls (configurable, default 10 req/s) to avoid overloading
|
|
||||||
the paperless instance.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- `PaperlessConnector` implements `SourceConnector` fully.
|
|
||||||
- Tag filtering limits which documents are listed.
|
|
||||||
- Pagination handles > 100 documents correctly.
|
|
||||||
- `content_hash` uses paperless's `checksum` field for change detection.
|
|
||||||
- Auth token resolved from k8s secret (not hardcoded).
|
|
||||||
- Health check reports document count matching tag filter.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** mock HTTP server (wiremock or similar) returning paperless API
|
|
||||||
responses; fixture JSON responses for list/detail endpoints.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_paperless_connector.rs`:
|
|
||||||
1. `a1_list_filters_by_tags` — mock returns 5 docs, 3 with matching tags; assert
|
|
||||||
`list_documents()` returns 3.
|
|
||||||
2. `a2_pagination_fetches_all` — mock returns 2 pages of 50; assert 100 docs.
|
|
||||||
3. `a3_fetch_returns_content` — mock detail endpoint; assert text matches fixture.
|
|
||||||
4. `a4_fetch_includes_metadata` — assert returned metadata includes title, tags,
|
|
||||||
correspondent, date fields.
|
|
||||||
5. `a5_health_check_reachable` — mock 200; assert `reachable: true` with count.
|
|
||||||
6. `a6_health_check_unreachable` — mock connection refused; assert
|
|
||||||
`reachable: false` with error message.
|
|
||||||
7. `a7_checksum_as_content_hash` — assert `SourceDocument.content_hash` is
|
|
||||||
populated from paperless `checksum` field.
|
|
||||||
8. `a8_auth_header_sent` — assert mock received `Authorization: Token <value>`.
|
|
||||||
9. `a9_config_from_yaml` — parse connector from YAML fixture; assert fields match.
|
|
||||||
10. `a10_live` — `#[ignore]`; real paperless instance; list + fetch one doc; print
|
|
||||||
title and content length for human sanity check.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_paperless_connector` (add `-- --ignored` for a10)
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Mocking without verifying auth header. A connector that works in tests but
|
|
||||||
sends no auth fails silently against real paperless.
|
|
||||||
- Testing single page only. Pagination bugs are invisible with < `page_size` docs.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Assuming `content` is always populated. paperless may have documents without
|
|
||||||
OCR text (e.g., empty scans). Return empty content with a warning, don't panic.
|
|
||||||
- Hardcoding the base URL without trailing-slash normalization. `/api/documents/`
|
|
||||||
vs `/api/documents` behaves differently.
|
|
||||||
- Not handling paperless API rate limits (429 responses). Add retry-after logic.
|
|
||||||
- Resolving k8s secrets at config parse time. Defer to runtime — secret may not
|
|
||||||
exist in dev/test environments. Use env var fallback.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors, paperless-ngx integration
|
|
||||||
@@ -1,118 +0,0 @@
|
|||||||
# M7.4 — Git repository connector
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.10 |
|
|
||||||
| Depends | M7.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Implement a `SourceConnector` that clones or pulls a git repository and exposes
|
|
||||||
its documents for ingestion — so documentation repos, wikis, and runbook
|
|
||||||
collections tracked in git become searchable through the memory service.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
Many knowledge sources are already in git: internal wikis, Forgejo repos,
|
|
||||||
infrastructure documentation, README collections. This connector pulls them
|
|
||||||
without requiring manual export.
|
|
||||||
|
|
||||||
**Configuration:**
|
|
||||||
```yaml
|
|
||||||
connectors:
|
|
||||||
- kind: git_repo
|
|
||||||
name: infra-docs
|
|
||||||
config:
|
|
||||||
repo_url: https://forgejo.riotpiao.com/rock/homelab-docs.git
|
|
||||||
branch: main
|
|
||||||
clone_dir: /tmp/mem-connectors/infra-docs # local checkout
|
|
||||||
paths: [docs/, runbooks/] # only scan these dirs
|
|
||||||
extensions: [md, txt, rst]
|
|
||||||
auth_secret: forgejo-token # k8s secret for private repos
|
|
||||||
sync_depth: 1 # shallow clone
|
|
||||||
```
|
|
||||||
|
|
||||||
**Clone-then-walk model.** The connector clones (or pulls) the repo to a local
|
|
||||||
directory, then walks the filesystem like the Obsidian connector. This reuses
|
|
||||||
`DocCorpusSource` internals and avoids git-specific content access APIs.
|
|
||||||
|
|
||||||
**Change detection uses git.** `git diff --name-only HEAD@{1}..HEAD` after a pull
|
|
||||||
tells the sync framework exactly which files changed — more efficient than
|
|
||||||
re-hashing every file.
|
|
||||||
|
|
||||||
**Branch tracking.** Connector watches one branch. Branch changes (main → prod)
|
|
||||||
require config update. No multi-branch support — each branch is a separate
|
|
||||||
connector instance.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Implement `GitRepoConnector` in `mem-ingest/src/connectors/git_repo.rs`.
|
|
||||||
2. On first sync: `git clone --depth N --branch B <url> <clone_dir>`.
|
|
||||||
3. On subsequent syncs: `git -C <clone_dir> pull --ff-only`.
|
|
||||||
4. `list_documents()` — walk `clone_dir` filtered by `paths` and `extensions`,
|
|
||||||
compute sha256 per file.
|
|
||||||
5. `fetch_document()` — read file from `clone_dir`, return content with git
|
|
||||||
metadata (last commit sha, author, date for that file via `git log -1`).
|
|
||||||
6. `health_check()` — verify `clone_dir` is a valid git repo, check remote
|
|
||||||
connectivity via `git ls-remote`.
|
|
||||||
7. Handle auth for private repos (token from k8s secret → git credential helper
|
|
||||||
or URL embedding).
|
|
||||||
8. Register `"git_repo"` kind in connector registry factory.
|
|
||||||
9. `source_type()` returns `Reference`.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Clone + walk produces correct document list for fixture repo.
|
|
||||||
- Pull detects changed files without re-processing unchanged ones.
|
|
||||||
- Path filtering limits scan to configured subdirectories.
|
|
||||||
- Private repo auth works (token in URL or credential helper).
|
|
||||||
- Health check distinguishes: valid repo, invalid remote, auth failure.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** fixture git repo created in temp dir with known content.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_git_repo_connector.rs`:
|
|
||||||
1. `a1_clone_and_list` — init fixture repo, connector clones it; assert
|
|
||||||
`list_documents()` returns expected files.
|
|
||||||
2. `a2_path_filtering` — fixture with `docs/` and `src/`; config says `paths: [docs/]`;
|
|
||||||
assert only `docs/` files returned.
|
|
||||||
3. `a3_extension_filtering` — fixture with .md, .rs, .txt; assert only .md/.txt.
|
|
||||||
4. `a4_fetch_returns_content` — fetch a doc; assert content matches fixture file.
|
|
||||||
5. `a5_pull_detects_changes` — add a commit to fixture repo; pull; assert changed
|
|
||||||
file appears in list with new hash.
|
|
||||||
6. `a6_health_check_valid_repo` — valid clone_dir; assert `reachable: true`.
|
|
||||||
7. `a7_health_check_no_clone` — no clone_dir; assert `reachable: false` with
|
|
||||||
message indicating clone needed.
|
|
||||||
8. `a8_shallow_clone` — assert clone depth matches config (`git rev-list --count HEAD`).
|
|
||||||
9. `a9_config_from_yaml` — parse connector from YAML; assert fields match.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_git_repo_connector`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing with a local repo path instead of a clone. The clone/pull machinery
|
|
||||||
is the whole point — a connector that reads a pre-existing checkout is just
|
|
||||||
the Obsidian connector.
|
|
||||||
- Not testing pull-after-change. First sync always works; second sync is where
|
|
||||||
change detection matters.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Running `git clone` on every sync. Check if `clone_dir` already has a valid
|
|
||||||
checkout first; clone only on first run.
|
|
||||||
- Not cleaning up failed clones. A partial clone leaves a directory that is
|
|
||||||
neither valid nor absent — subsequent runs fail on both clone (dir exists)
|
|
||||||
and pull (not a repo).
|
|
||||||
- Force-push upstream breaks `--ff-only`. Detect non-fast-forward, delete
|
|
||||||
`clone_dir`, re-clone. Log a warning — this means all docs are re-processed.
|
|
||||||
- Symlinks across the repo boundary. `walkdir` follows symlinks by default;
|
|
||||||
a symlink to `/etc/passwd` is a real concern in a cluster-wide service.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors section
|
|
||||||
@@ -1,109 +0,0 @@
|
|||||||
# M7.5 — S3-compatible storage connector
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.10 |
|
|
||||||
| Depends | M7.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Implement a `SourceConnector` for S3-compatible object storage (MinIO, AWS S3,
|
|
||||||
R2, etc.) so documents stored in buckets become searchable through the memory
|
|
||||||
service.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
S3 is the universal storage protocol. MinIO runs in many homelabs, and cloud
|
|
||||||
providers expose the same API. This connector makes any S3 bucket a knowledge
|
|
||||||
source.
|
|
||||||
|
|
||||||
**Configuration:**
|
|
||||||
```yaml
|
|
||||||
connectors:
|
|
||||||
- kind: s3
|
|
||||||
name: knowledge-bucket
|
|
||||||
config:
|
|
||||||
endpoint: https://minio.riotpiao.com
|
|
||||||
bucket: knowledge-base
|
|
||||||
prefix: docs/ # only this prefix
|
|
||||||
extensions: [md, txt, pdf] # filter by key suffix
|
|
||||||
access_key_secret: minio-creds # k8s secret with access/secret keys
|
|
||||||
region: us-east-1 # for AWS; ignored by MinIO
|
|
||||||
```
|
|
||||||
|
|
||||||
**ETag for change detection.** S3 objects have ETags (usually MD5 of content).
|
|
||||||
Use this as `content_hash` in `SourceDocument` — the sync framework skips objects
|
|
||||||
whose ETag hasn't changed.
|
|
||||||
|
|
||||||
**Text extraction.** S3 stores raw files. Markdown and text files are read
|
|
||||||
directly. PDF/DOCX support is out of scope for M7.5 — those MIME types are
|
|
||||||
skipped with a warning. Future: add a text extraction layer or require pre-processed
|
|
||||||
text.
|
|
||||||
|
|
||||||
**Pagination via continuation tokens.** S3 ListObjectsV2 returns max 1000 keys
|
|
||||||
per request. Use `ContinuationToken` for subsequent pages.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Implement `S3Connector` in `mem-ingest/src/connectors/s3.rs`.
|
|
||||||
2. `list_documents()` — `ListObjectsV2` with `Prefix`, paginate, filter by
|
|
||||||
extension, return `SourceDocument` per object.
|
|
||||||
3. `fetch_document()` — `GetObject`, read body as text (UTF-8), return with
|
|
||||||
metadata (key, size, last_modified, ETag).
|
|
||||||
4. `health_check()` — `HeadBucket` to verify access.
|
|
||||||
5. Auth via access key + secret key from k8s secret.
|
|
||||||
6. Use `aws-sdk-s3` or `rust-s3` crate for S3 API.
|
|
||||||
7. Register `"s3"` kind in connector registry factory.
|
|
||||||
8. `source_type()` returns `Reference`.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- `S3Connector` implements `SourceConnector` fully.
|
|
||||||
- Prefix filtering limits to configured path.
|
|
||||||
- Extension filtering skips non-text objects.
|
|
||||||
- ETag is used as `content_hash` for change detection.
|
|
||||||
- Pagination handles > 1000 objects.
|
|
||||||
- Auth works with MinIO and AWS-style credentials.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** mock S3 server (localstack or in-process mock) with fixture objects.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_s3_connector.rs`:
|
|
||||||
1. `a1_list_with_prefix` — mock bucket with objects under `docs/` and `images/`;
|
|
||||||
assert only `docs/` objects listed.
|
|
||||||
2. `a2_extension_filtering` — mock with .md, .png, .txt; assert .png excluded.
|
|
||||||
3. `a3_fetch_returns_content` — fetch a .md object; assert content matches.
|
|
||||||
4. `a4_etag_as_content_hash` — assert `SourceDocument.content_hash` equals
|
|
||||||
the object's ETag.
|
|
||||||
5. `a5_pagination` — mock 1500 objects; assert all listed via continuation tokens.
|
|
||||||
6. `a6_health_check_valid_bucket` — mock HeadBucket 200; assert reachable.
|
|
||||||
7. `a7_health_check_no_access` — mock HeadBucket 403; assert not reachable with
|
|
||||||
error message.
|
|
||||||
8. `a8_non_utf8_skipped` — mock object with binary content; assert skipped with
|
|
||||||
warning, not crash.
|
|
||||||
9. `a9_config_from_yaml` — parse connector from YAML; assert fields match.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_s3_connector`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing with a local filesystem mock instead of S3 API mock. The pagination
|
|
||||||
and ETag handling are S3-specific.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Assuming ETags are always MD5. Multipart uploads produce composite ETags
|
|
||||||
(`hash-N`). These are still unique per version — use as-is for change detection.
|
|
||||||
- Not handling `NoSuchBucket` vs `AccessDenied`. Both are errors but mean
|
|
||||||
different things for health reporting.
|
|
||||||
- Reading binary files as UTF-8. A JPEG read as text produces garbage. Check
|
|
||||||
content-type header and skip non-text MIME types.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors section
|
|
||||||
@@ -1,146 +0,0 @@
|
|||||||
# M7.6 — Sync framework
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | L — 3–5 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.7, M7.10 |
|
|
||||||
| Depends | M7.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Build the shared sync engine that handles change detection, tombstoning, drift
|
|
||||||
reporting, and resumable sync for **all** connectors — so individual connector
|
|
||||||
implementations only fetch documents and everything else is handled once.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
Every connector faces the same sync problems:
|
|
||||||
1. **What changed?** Compare content hashes from `list_documents()` against a
|
|
||||||
manifest of last-known hashes. Only fetch and process changed documents.
|
|
||||||
2. **What disappeared?** Documents in the manifest but not in `list_documents()`
|
|
||||||
need tombstone records in the log. Append-only — never delete.
|
|
||||||
3. **What if sync crashes midway?** Track progress per-document. Restart
|
|
||||||
processes only unseen documents.
|
|
||||||
4. **How much will this cost?** Drift report shows counts without mutating
|
|
||||||
anything.
|
|
||||||
|
|
||||||
This is the generalization of M3.6.3's `mem ref sync/list/rm` logic, applied to
|
|
||||||
any `SourceConnector` instead of just `DocCorpusSource`.
|
|
||||||
|
|
||||||
**Manifest storage.** Per-connector manifest in `log/connectors/<name>.manifest.jsonl`:
|
|
||||||
```jsonl
|
|
||||||
{"doc_id":"abc","source_uri":"file:///vault/k8s.md","content_hash":"sha256:...","chunk_count":12,"synced_at":"..."}
|
|
||||||
{"doc_id":"def","source_uri":"paperless://doc/42","content_hash":"sha256:...","chunk_count":3,"synced_at":"..."}
|
|
||||||
```
|
|
||||||
|
|
||||||
**Sync algorithm:**
|
|
||||||
```
|
|
||||||
current = connector.list_documents()
|
|
||||||
previous = load_manifest(connector.name)
|
|
||||||
|
|
||||||
for doc in current:
|
|
||||||
if doc.content_hash == previous[doc.doc_id].content_hash:
|
|
||||||
skip (unchanged)
|
|
||||||
else if doc.doc_id in previous:
|
|
||||||
tombstone previous chunks, fetch + chunk + embed new (changed)
|
|
||||||
else:
|
|
||||||
fetch + chunk + embed (new)
|
|
||||||
|
|
||||||
for doc_id in previous not in current:
|
|
||||||
tombstone all chunks (removed)
|
|
||||||
|
|
||||||
save_manifest(connector.name, current)
|
|
||||||
```
|
|
||||||
|
|
||||||
**Chunking delegation.** The sync framework owns the chunking step. It routes
|
|
||||||
fetched `DocumentContent` through the appropriate `ChunkPolicy`:
|
|
||||||
- Markdown → heading-boundary chunking (reuse `DocCorpusSource` logic)
|
|
||||||
- Plain text → paragraph-boundary chunking
|
|
||||||
- Configurable per connector kind in `connectors.yaml`
|
|
||||||
|
|
||||||
**Level routing.** Session connectors → RecordSource → gated loop (L0/L1/L2).
|
|
||||||
Document connectors → Reference records (Level R). The `source_type()` method
|
|
||||||
on the connector determines the pipeline.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Define `SyncEngine` struct in `mem-ingest/src/sync.rs`.
|
|
||||||
2. Implement manifest loading/saving (JSONL per connector).
|
|
||||||
3. Implement diff algorithm: `(new, changed, unchanged, removed)` from
|
|
||||||
`list_documents()` vs manifest.
|
|
||||||
4. Implement sync loop: for each `new`/`changed` doc, fetch → chunk → emit
|
|
||||||
records. For each `removed` doc, emit tombstones.
|
|
||||||
5. Implement drift reporting: same diff algorithm, print counts, no mutations.
|
|
||||||
6. Implement resume: track synced doc_ids in a progress file. On restart,
|
|
||||||
skip already-synced docs.
|
|
||||||
7. Implement rate limiting: configurable max concurrent fetches per connector.
|
|
||||||
8. Integrate with `mem-store` for writing Reference records to the log.
|
|
||||||
9. Integrate with `mem-llm` for embedding new chunks.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Sync of unchanged connector produces zero embedding calls.
|
|
||||||
- Changed document: old chunks tombstoned, new chunks embedded and stored.
|
|
||||||
- Removed document: chunks tombstoned, manifest updated.
|
|
||||||
- Drift report matches actual changes without mutating anything.
|
|
||||||
- Crash mid-sync → restart processes only remaining documents.
|
|
||||||
- Rate limiting: never exceeds configured concurrent fetch limit.
|
|
||||||
- Log remains append-only (tombstones are records, not deletions).
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** `VecConnector` with mutable document list, counting embedder.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_sync_framework.rs`:
|
|
||||||
1. `a1_initial_sync_all_new` — 3 docs, no manifest; assert all 3 fetched and
|
|
||||||
embedded, manifest written with 3 entries.
|
|
||||||
2. `a2_unchanged_skipped` — sync again with same content; assert zero fetch
|
|
||||||
calls, zero embed calls.
|
|
||||||
3. `a3_changed_doc_replaced` — modify one doc's content; sync; assert old chunks
|
|
||||||
tombstoned, new chunks embedded, embed count equals changed doc's chunk count.
|
|
||||||
4. `a4_removed_doc_tombstoned` — remove a doc from connector; sync; assert
|
|
||||||
tombstone records emitted, manifest entry removed.
|
|
||||||
5. `a5_new_doc_added` — add a doc to connector; sync; assert only new doc
|
|
||||||
fetched and embedded.
|
|
||||||
6. `a6_drift_report_read_only` — modify docs, run drift report; assert correct
|
|
||||||
counts (1 new, 1 changed, 1 unchanged, 1 removed); assert no mutations to
|
|
||||||
manifest or log.
|
|
||||||
7. `a7_resume_after_crash` — sync 5 docs, simulate crash after 3; restart;
|
|
||||||
assert only 2 remaining docs processed.
|
|
||||||
8. `a8_tombstone_is_append` — count log lines before and after remove; assert
|
|
||||||
count only grew.
|
|
||||||
9. `a9_manifest_roundtrip` — save manifest, load it; assert field-for-field
|
|
||||||
equality.
|
|
||||||
10. `a10_rebuild_parity` — after full sync, `mem rebuild --from-log` produces
|
|
||||||
identical state.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_sync_framework`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Asserting "no duplicate rows" instead of counting embed calls. A sync that
|
|
||||||
re-embeds everything and upserts by sha produces correct rows and wasted
|
|
||||||
compute.
|
|
||||||
- Testing drift report without actually changing documents first.
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Comparing document-level hashes instead of chunk-level. A doc that changed
|
|
||||||
one paragraph should re-embed only the affected chunks, not all of them.
|
|
||||||
However, heading-boundary chunking means changing one heading can shift all
|
|
||||||
subsequent chunks. Accept document-level granularity for now; chunk-level
|
|
||||||
optimization is a future refinement.
|
|
||||||
- Making the manifest a database table instead of JSONL. The manifest must
|
|
||||||
survive `mem rebuild --from-log` — it is metadata about the sync process,
|
|
||||||
not a projection of the log.
|
|
||||||
- Running fetch + embed serially. A connector with 500 docs at 200ms per embed
|
|
||||||
takes 100s serially. Concurrent fetch + sequential embed is the right shape.
|
|
||||||
- Ignoring the derived filter (M4.2). Document connectors produce Level R content
|
|
||||||
that must register in the artifact manifest so it cannot re-enter as evidence.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors, sync framework
|
|
||||||
@@ -1,107 +0,0 @@
|
|||||||
# M7.7 — `mem source` CLI
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.8, M7.10 |
|
|
||||||
| Depends | M7.6 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Provide CLI commands to manage source connectors: sync documents, check status,
|
|
||||||
list connectors, add/remove connectors.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
```
|
|
||||||
mem source list # all registered connectors + health
|
|
||||||
mem source status # drift report per connector, no mutations
|
|
||||||
mem source sync --all # sync all connectors
|
|
||||||
mem source sync --name homelab-vault # sync one connector
|
|
||||||
mem source sync --name homelab-vault --dry-run # show plan, no mutations
|
|
||||||
mem source add --kind paperless --name docs --config '{"base_url":"...","token":"..."}'
|
|
||||||
mem source rm --name old-source # deregister + tombstone all chunks
|
|
||||||
mem source health # connectivity check per connector
|
|
||||||
```
|
|
||||||
|
|
||||||
**Output format.** CLI outputs human-readable tables by default, `--json` for
|
|
||||||
machine consumption. Example:
|
|
||||||
|
|
||||||
```
|
|
||||||
$ mem source list
|
|
||||||
NAME KIND DOCS LAST SYNC HEALTH
|
|
||||||
homelab-vault obsidian 42 2026-08-26T10:00:00Z ✅ reachable
|
|
||||||
homelab-paperless paperless 127 2026-08-26T09:00:00Z ✅ reachable
|
|
||||||
infra-docs git_repo 18 2026-08-25T20:00:00Z ⚠️ pull failed
|
|
||||||
|
|
||||||
$ mem source status
|
|
||||||
NAME NEW CHANGED UNCHANGED REMOVED
|
|
||||||
homelab-vault 0 2 40 0
|
|
||||||
homelab-paperless 3 0 124 0
|
|
||||||
infra-docs 1 1 16 0
|
|
||||||
```
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Add `Source` subcommand group to clap CLI in `mem-cli/src/main.rs`.
|
|
||||||
2. Implement `cmd_source_list()` — load registry, health check each, tabulate.
|
|
||||||
3. Implement `cmd_source_status()` — load registry, run drift report per
|
|
||||||
connector, tabulate.
|
|
||||||
4. Implement `cmd_source_sync()` — load registry, run sync engine for selected
|
|
||||||
connector(s), report results.
|
|
||||||
5. Implement `cmd_source_add()` — validate kind, parse config, register in
|
|
||||||
`connectors.yaml`, run initial health check.
|
|
||||||
6. Implement `cmd_source_rm()` — tombstone all chunks from connector, remove
|
|
||||||
from `connectors.yaml`.
|
|
||||||
7. Implement `cmd_source_health()` — connectivity check per connector.
|
|
||||||
8. Add `--dry-run` to sync (show plan only).
|
|
||||||
9. Add `--json` flag for machine-readable output.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- All subcommands execute without panic.
|
|
||||||
- `sync --dry-run` shows plan without mutations.
|
|
||||||
- `add` validates kind exists in registry before writing config.
|
|
||||||
- `rm` tombstones chunks and removes config entry.
|
|
||||||
- `status` shows drift without mutations.
|
|
||||||
- Exit codes: 0 on success, non-zero on failure.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Harness:** `VecConnector` registered in registry, temp connectors.yaml.
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_source_cli.rs`:
|
|
||||||
1. `a1_list_shows_connectors` — register two connectors; assert list output
|
|
||||||
contains both names and kinds.
|
|
||||||
2. `a2_status_shows_drift` — modify connector docs; assert status shows correct
|
|
||||||
new/changed/unchanged/removed counts.
|
|
||||||
3. `a3_sync_processes_changes` — sync with changes; assert documents processed.
|
|
||||||
4. `a4_sync_dry_run_no_mutations` — sync --dry-run; assert no log entries written.
|
|
||||||
5. `a5_add_registers_connector` — add a new connector; assert it appears in list.
|
|
||||||
6. `a6_add_bad_kind_fails` — add with unknown kind; assert non-zero exit.
|
|
||||||
7. `a7_rm_tombstones_and_deregisters` — rm a connector; assert tombstone records
|
|
||||||
written and connector removed from list.
|
|
||||||
8. `a8_health_reports_status` — assert health output includes reachable/unreachable.
|
|
||||||
9. `a9_json_output` — assert --json flag produces valid JSON.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_source_cli`
|
|
||||||
|
|
||||||
**False pass:**
|
|
||||||
- Testing `list` without checking that connectors are actually registered (not
|
|
||||||
just config parsed).
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- `sync --all` with a broken connector should not halt all syncs. Sync each
|
|
||||||
independently, report failures per connector at the end.
|
|
||||||
- `rm` without `--yes` should prompt for confirmation (destructive operation).
|
|
||||||
- Don't write `connectors.yaml` atomically — a crash mid-write corrupts config.
|
|
||||||
Write to temp file, then rename.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors, CLI commands
|
|
||||||
@@ -1,107 +0,0 @@
|
|||||||
# M7.8 — `GET /memory/sources` + `POST /memory/sources/sync` HTTP endpoints
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | M — 1–3 days |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.10 |
|
|
||||||
| Depends | M7.7, M3.5.1 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Expose source connector management through the HTTP API so agents and UIs can
|
|
||||||
trigger syncs, check connector health, and view status without CLI access.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
**Endpoints:**
|
|
||||||
```
|
|
||||||
GET /memory/sources # list all connectors + health
|
|
||||||
GET /memory/sources/{name} # one connector detail + drift
|
|
||||||
GET /memory/sources/{name}/health # health check only
|
|
||||||
POST /memory/sources/sync # trigger sync (async)
|
|
||||||
POST /memory/sources/{name}/sync # trigger sync for one connector
|
|
||||||
GET /memory/sources/sync/{job_id} # sync job status
|
|
||||||
```
|
|
||||||
|
|
||||||
**Response shapes:**
|
|
||||||
```json
|
|
||||||
// GET /memory/sources
|
|
||||||
[
|
|
||||||
{
|
|
||||||
"name": "homelab-vault",
|
|
||||||
"kind": "obsidian",
|
|
||||||
"document_count": 42,
|
|
||||||
"last_sync": "2026-08-26T10:00:00Z",
|
|
||||||
"health": { "reachable": true, "document_count": 42 }
|
|
||||||
}
|
|
||||||
]
|
|
||||||
|
|
||||||
// POST /memory/sources/sync
|
|
||||||
// Request: { "names": ["homelab-vault"] } (or omit for all)
|
|
||||||
// Response: 202 Accepted
|
|
||||||
{ "job_id": "sync-abc123", "status_url": "/memory/sources/sync/sync-abc123" }
|
|
||||||
|
|
||||||
// GET /memory/sources/sync/{job_id}
|
|
||||||
{
|
|
||||||
"job_id": "sync-abc123",
|
|
||||||
"status": "completed",
|
|
||||||
"connectors": {
|
|
||||||
"homelab-vault": { "new": 0, "changed": 2, "unchanged": 40, "removed": 0 }
|
|
||||||
}
|
|
||||||
}
|
|
||||||
```
|
|
||||||
|
|
||||||
**Sync is async.** POST returns 202 immediately; client polls status endpoint.
|
|
||||||
Same pattern as `/memory/ingest`.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Add routes to `http_server.rs` for all six endpoints.
|
|
||||||
2. `sources_list_handler` — load registry, health check each, return JSON array.
|
|
||||||
3. `source_detail_handler` — one connector, include drift report.
|
|
||||||
4. `source_health_handler` — health check only.
|
|
||||||
5. `source_sync_handler` — validate connector names, spawn async sync job,
|
|
||||||
return job_id + status URL.
|
|
||||||
6. `source_sync_status_handler` — look up job by id, return progress.
|
|
||||||
7. Auth: all endpoints require apikey header (existing pattern).
|
|
||||||
8. Rate limiting: sync endpoint limited to 10 req/hour (it's expensive).
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- All endpoints return correct status codes and JSON shapes.
|
|
||||||
- Sync is async — POST returns immediately, job runs in background.
|
|
||||||
- Unknown connector name returns 404.
|
|
||||||
- Auth required on all endpoints.
|
|
||||||
- Rate limiting on sync endpoint.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_source_http.rs`:
|
|
||||||
1. `a1_list_returns_connectors` — register connectors, GET /memory/sources;
|
|
||||||
assert JSON array with correct fields.
|
|
||||||
2. `a2_detail_includes_drift` — GET /memory/sources/{name}; assert drift fields.
|
|
||||||
3. `a3_health_check_via_http` — GET /memory/sources/{name}/health; assert
|
|
||||||
reachable field.
|
|
||||||
4. `a4_sync_returns_202` — POST /memory/sources/sync; assert 202 + job_id.
|
|
||||||
5. `a5_sync_status_tracks_progress` — poll status endpoint; assert eventually
|
|
||||||
"completed".
|
|
||||||
6. `a6_unknown_connector_404` — GET /memory/sources/nonexistent; assert 404.
|
|
||||||
7. `a7_auth_required` — request without apikey; assert 401.
|
|
||||||
8. `a8_sync_rate_limited` — 11 sync requests; assert 429 on the 11th.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_source_http`
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Making sync synchronous. A paperless connector with 500 docs takes minutes;
|
|
||||||
blocking the HTTP response is a client timeout.
|
|
||||||
- Not capping concurrent sync jobs. Two simultaneous syncs to the same connector
|
|
||||||
can corrupt the manifest.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors, HTTP endpoints
|
|
||||||
@@ -1,76 +0,0 @@
|
|||||||
# M7.9 — Connector health monitoring + observability
|
|
||||||
|
|
||||||
| Field | Value |
|
|
||||||
|---|---|
|
|
||||||
| Phase | M7 — Source connectors |
|
|
||||||
| Size | S — < 1 day |
|
|
||||||
| Status | ⬜ Not started |
|
|
||||||
| Flags | — |
|
|
||||||
| Spec | inlined below |
|
|
||||||
| Blocks | M7.10 |
|
|
||||||
| Depends | M7.8 |
|
|
||||||
|
|
||||||
## Goal
|
|
||||||
|
|
||||||
Add periodic health checks, sync metrics, and alerting hooks for source
|
|
||||||
connectors so connector failures are detected before knowledge goes stale.
|
|
||||||
|
|
||||||
## Facts (inlined — no spec read needed)
|
|
||||||
|
|
||||||
A connector that silently fails means the knowledge base is stale with no
|
|
||||||
indication. The monitoring layer detects this.
|
|
||||||
|
|
||||||
**Metrics (Prometheus-compatible):**
|
|
||||||
```
|
|
||||||
mem_source_health{name="homelab-vault",kind="obsidian"} 1 # 1=healthy, 0=unhealthy
|
|
||||||
mem_source_last_sync_seconds{name="homelab-vault"} 1724680000 # unix timestamp
|
|
||||||
mem_source_documents_total{name="homelab-vault"} 42
|
|
||||||
mem_source_sync_duration_seconds{name="homelab-vault"} 12.3
|
|
||||||
mem_source_sync_errors_total{name="homelab-vault"} 0
|
|
||||||
mem_source_drift_new{name="homelab-vault"} 2 # docs pending sync
|
|
||||||
mem_source_drift_changed{name="homelab-vault"} 1
|
|
||||||
```
|
|
||||||
|
|
||||||
**Periodic health check.** Configurable interval (default 5 minutes). Log
|
|
||||||
warnings for unreachable connectors. Update Prometheus gauges.
|
|
||||||
|
|
||||||
**Staleness alert.** If `last_sync` exceeds a configurable threshold (default 24h),
|
|
||||||
log a warning and set a metric. This is the "knowledge is going stale" signal.
|
|
||||||
|
|
||||||
## Steps
|
|
||||||
|
|
||||||
1. Add health check background task (tokio interval, configurable period).
|
|
||||||
2. Export Prometheus metrics via `/metrics` endpoint (existing pattern or new).
|
|
||||||
3. Track per-connector: health, last sync, doc count, sync duration, errors.
|
|
||||||
4. Staleness detection: compare `last_sync` to now, warn if exceeds threshold.
|
|
||||||
5. Log structured health events for observability.
|
|
||||||
|
|
||||||
## Acceptance
|
|
||||||
|
|
||||||
- Periodic health checks run at configured interval.
|
|
||||||
- Metrics endpoint returns valid Prometheus format.
|
|
||||||
- Staleness warning fires when last_sync exceeds threshold.
|
|
||||||
- Unhealthy connector logged with error details.
|
|
||||||
|
|
||||||
## Verify
|
|
||||||
|
|
||||||
**Integration test** — `tests/it_source_monitoring.rs`:
|
|
||||||
1. `a1_health_metric_updated` — register connector, wait for health check;
|
|
||||||
assert metric value matches connector health.
|
|
||||||
2. `a2_staleness_detected` — set last_sync to 25h ago; assert staleness warning.
|
|
||||||
3. `a3_metrics_format_valid` — GET /metrics; assert valid Prometheus text format.
|
|
||||||
4. `a4_error_count_incremented` — force connector health check failure; assert
|
|
||||||
error counter incremented.
|
|
||||||
|
|
||||||
**Command:** `cargo test --test it_source_monitoring`
|
|
||||||
|
|
||||||
## Traps
|
|
||||||
|
|
||||||
- Running health checks synchronously. A connector that times out blocks all
|
|
||||||
other connectors' health checks. Use tokio::select with timeout.
|
|
||||||
- Not distinguishing "unreachable" from "empty". A connector that returns 0 docs
|
|
||||||
is healthy; a connector that can't connect is not.
|
|
||||||
|
|
||||||
---
|
|
||||||
|
|
||||||
Background: [DESIGN.md](../DESIGN.md) — source connectors, observability
|
|
||||||
Reference in New Issue
Block a user