Files
poimen-memory/tasks/M3.2-rerank-client.md
T

89 lines
3.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# M3.2 — Rerank client
| Field | Value |
|---|---|
| Phase | M3 — L2 synthesis and retrieval |
| Size | S — under 1 day |
| Status | ⬜ Not started |
| Flags | — |
| Spec | inlined below |
| Blocks | M1.1 |
## Goal
Reorder vector-recall candidates by actual relevance, because embedding distance
is a coarse filter.
## Facts (inlined — no spec read needed)
```
POST /v1/rerank
{"query":"what is rust","texts":["rust is a language","bananas"]}
-> [{"index":0,"score":0.98176736},{"index":1,"score":0.00008833366}]
```
`BAAI/bge-reranker-base` via TEI. Note the **response shape is a bare array**, not
an OpenAI-style `{"data": [...]}` envelope — this route does not follow the chat
convention.
The discrimination is real: 0.98 vs 0.00009 on that probe, four orders of
magnitude. Embedding cosine on the same pair would be far closer, which is why
recall-then-rerank beats recall alone.
`index` refers to the position in the submitted `texts` array; results come back
**sorted by score**, so the index is the only way to map back. Do not assume
order.
Batch limits apply as with embeddings — keep candidate lists modest (top-50 from
recall is plenty).
## Steps
1. `RerankClient::rerank(query, texts) -> Result<Vec<Scored>>` in `mem-llm`.
2. Parse the bare array; map `index` back to the caller's items.
3. Preserve the caller's item type: take `&[T]`, return `Vec<(T, f32)>` so the
caller does not re-associate by position.
4. Same `apikey` header, retry and timeout policy as M1.1.
5. Empty input returns empty without a request.
## Acceptance
- Results map correctly back to input items via `index`.
- A more relevant text scores above a less relevant one on a live call.
- Empty input makes no request.
## Verify
**Harness:** `wiremock` offline, one `#[ignore]` live test.
**Integration test**`tests/it_rerank.rs`:
1. `a1_bare_array_parsed` — mock returns `[{"index":1,...},{"index":0,...}]`;
assert parsing succeeds.
2. `a2_index_mapping` — with the out-of-order mock above, assert the returned
items correspond to inputs 1 and 0 respectively, not 0 and 1.
3. `a3_empty_no_request` — empty texts; assert zero requests.
4. `a4_apikey_sent` — header present.
5. `a5_live_discriminates``#[ignore]`; real gateway, query "what is rust"
against `["rust is a language","bananas"]`; assert the first scores at least
10× the second.
**Command:** `cargo test -p mem-llm rerank` (add `-- --ignored` for a5)
**False pass:**
- Assuming the response preserves input order. A mock that returns results in
input order passes a naive test and the live endpoint returns them sorted,
which silently mislabels every result. Assertion 2 must use an out-of-order
mock.
- Expecting an OpenAI envelope. It will fail immediately against the live route,
but a mock written to match the wrong shape hides that until integration.
## Traps
- Re-associating results by position instead of by `index`. The scores are right
and attached to the wrong documents — a bug that looks like poor retrieval
quality rather than a mapping error.
---
Background: [DESIGN.md](../DESIGN.md) — Verified facts, retrieval