Files
poimen-memory/tasks/M3.7.8-symptom-projection.md
T

142 lines
6.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M3.7.8 — Symptom projection: make an answer findable from an error
| Field | Value |
|---|---|
| Phase | M3.7 — Tool context |
| Size | M — 13 days |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M3.7.6 |
| Depends | M3.7.7, M2.4, M1.5 |
## Goal
Give every memory a second vector describing the failures it would explain, so a
stack trace can find an answer written in prose.
## Facts (inlined — no spec read needed)
The asymmetry this exists to fix:
```
L1 memory (how it is written):
"Requests over 10KB failed because Kong buffered the whole body before
proxying; resolved with proxy-body-size: 0 on the ingress."
Query (how it arrives):
"413 Request Entity Too Large" + a curl trace
```
Same incident. Embedded with the same model and compared by cosine, they are
mediocre neighbours — one is an explanation, the other is a symptom. This is the
main reason retrieval that looks correct in a unit test disappoints in use.
**Fix at write time, not read time.** When the gated loop accepts a memory,
generate a short *symptom projection* — the errors, messages and observable
behaviour this memory would explain — and embed that as a second vector:
```
symptom projection for the memory above:
"413 Request Entity Too Large; large POST bodies rejected at the ingress;
uploads over 10KB fail while small ones succeed; nginx/Kong body buffer limit"
```
The alternative, HyDE, generates a hypothetical answer per *query* and puts an
LLM call on every lookup. Writes are rare here — the gate keeps acceptance under
30% by design — and lookups should be fast, so paying once at write is the right
side of that trade.
**It is a projection, so it obeys the projection rules.** Regenerated by
`mem rebuild --from-log`, never authoritative, and byte-identical on replay —
which means the generation call must be deterministic: temperature 0, pinned
prompt, and the model id recorded in the log record so a model change is visible
as a rebuild difference rather than silent drift.
**Only L1 and L2 get one.** L0 is raw evidence already phrased as symptoms; R is
documentation and its headings already read like problems. Generating projections
for those doubles the index for no gain.
**Empty is allowed.** Not every memory explains a failure — an architectural
decision has no symptoms. The controller returns nothing, no vector is written,
and the memory remains findable by its text vector alone. A model that invents
symptoms for a memory that has none pollutes the symptom index with plausible
nonsense, which is worse than a smaller index.
**Signature linking.** Where the L0 evidence behind an accepted memory contains a
parseable failure (M3.7.7), write a `failure_signature` row pointing at the L1.
That is what turns the next occurrence into an exact-match tier-1 hit instead of
a vector search.
## Steps
1. Prompt template `prompts/symptom.tmpl` — memory text in, symptom lines out,
explicit "return nothing if this describes no failure".
2. Hook into the gated loop after a memory is accepted; L1 and L2 only.
3. Temperature 0; record model id and prompt sha on the log record.
4. Embed and write `memory_vector(kind='symptom')`.
5. Extract signatures from the memory's L0 parents; write `failure_signature`
rows keyed to the memory.
6. `mem rebuild --from-log` regenerates projections; assert stability.
7. `mem symptoms show <sha>` prints the projection for inspection.
## Acceptance
- An accepted L1 that explains a failure gains a symptom vector.
- A memory describing a decision gains none.
- Symptom text retrieves its memory from a raw error string that shares no
vocabulary with the memory text.
- Rebuild regenerates projections byte-identically.
- Signature rows link to the right memory.
## Verify
**Harness:** fixture memories — five that explain failures, three that do not —
plus raw error strings for the five, deliberately worded with no vocabulary
overlap with the memory text.
**Integration test**`tests/it_symptom_projection.rs`:
1. `a1_projection_generated` — the five failure memories each gain a
`kind='symptom'` vector.
2. `a2_no_projection_for_non_failure` — the three others gain none.
3. `a3_retrieval_by_symptom` — searching `kind='symptom'` with each raw error
string returns its memory first. This is the assertion the task exists for.
4. `a4_beats_text_vector` — the same query against `kind='text'` ranks the
correct memory *lower*; assert the symptom search strictly improves rank.
Without this the projection could be doing nothing.
5. `a5_deterministic` — generate twice; assert byte-identical projections.
6. `a6_rebuild_stable` — drop and rebuild; assert projections and their
embeddings match the originals.
7. `a7_signature_linked` — a memory whose evidence contains a parseable error has
a `failure_signature` row pointing at it, with the tool set.
8. `a8_l0_and_r_skipped` — assert no L0 or R node has a symptom vector.
9. `a9_model_id_recorded` — the log record names the model and prompt sha.
**Command:** `cargo test -p mem-core symptom && cargo test -p mem-cli rebuild`
**False pass:**
- Test queries that reuse the memory's own wording. The text vector already finds
those, assertion 3 passes, and the projection is never exercised. The error
strings must share no meaningful vocabulary — that constraint is the test.
- Asserting 3 without 4. If the text vector already ranked it first, assertion 3
is satisfied by a projection that is empty or useless.
- Skipping determinism because output "looks stable". A default temperature makes
it stable for ten runs and different on the eleventh, and the symptom is a
rebuild diff nobody can explain.
## Traps
- Generating projections for rejected chunks. The gate rejected them; embedding
their symptoms puts evidence-free content in the index through a side door.
- Letting the projection restate the memory. If the model paraphrases the answer
instead of naming the symptoms, the second vector duplicates the first and
assertion 4 fails — which is the correct outcome, but the cause is the prompt,
not the plumbing.
- Treating an empty projection as an error and retrying. It is the right answer
for most non-incident memories, and a retry loop turns it into invented
symptoms.
---
Background: [DESIGN.md](../DESIGN.md) — tool context, write path