89 lines
3.1 KiB
Markdown
89 lines
3.1 KiB
Markdown
# M3.2 — Rerank client
|
||||
|
|
|
|||
|
|
| Field | Value |
|
|||
|
|
|---|---|
|
|||
|
|
| Phase | M3 — L2 synthesis and retrieval |
|
|||
|
|
| Size | S — under 1 day |
|
|||
|
|
| Status | ⬜ Not started |
|
|||
|
|
| Flags | — |
|
|||
|
|
| Spec | inlined below |
|
|||
|
|
| Blocks | M1.1 |
|
|||
|
|
|
|||
|
|
## Goal
|
|||
|
|
|
|||
|
|
Reorder vector-recall candidates by actual relevance, because embedding distance
|
|||
|
|
is a coarse filter.
|
|||
|
|
|
|||
|
|
## Facts (inlined — no spec read needed)
|
|||
|
|
|
|||
|
|
```
|
|||
|
|
POST /v1/rerank
|
|||
|
|
{"query":"what is rust","texts":["rust is a language","bananas"]}
|
|||
|
|
-> [{"index":0,"score":0.98176736},{"index":1,"score":0.00008833366}]
|
|||
|
|
```
|
|||
|
|
|
|||
|
|
`BAAI/bge-reranker-base` via TEI. Note the **response shape is a bare array**, not
|
|||
|
|
an OpenAI-style `{"data": [...]}` envelope — this route does not follow the chat
|
|||
|
|
convention.
|
|||
|
|
|
|||
|
|
The discrimination is real: 0.98 vs 0.00009 on that probe, four orders of
|
|||
|
|
magnitude. Embedding cosine on the same pair would be far closer, which is why
|
|||
|
|
recall-then-rerank beats recall alone.
|
|||
|
|
|
|||
|
|
`index` refers to the position in the submitted `texts` array; results come back
|
|||
|
|
**sorted by score**, so the index is the only way to map back. Do not assume
|
|||
|
|
order.
|
|||
|
|
|
|||
|
|
Batch limits apply as with embeddings — keep candidate lists modest (top-50 from
|
|||
|
|
recall is plenty).
|
|||
|
|
|
|||
|
|
## Steps
|
|||
|
|
|
|||
|
|
1. `RerankClient::rerank(query, texts) -> Result<Vec<Scored>>` in `mem-llm`.
|
|||
|
|
2. Parse the bare array; map `index` back to the caller's items.
|
|||
|
|
3. Preserve the caller's item type: take `&[T]`, return `Vec<(T, f32)>` so the
|
|||
|
|
caller does not re-associate by position.
|
|||
|
|
4. Same `apikey` header, retry and timeout policy as M1.1.
|
|||
|
|
5. Empty input returns empty without a request.
|
|||
|
|
|
|||
|
|
## Acceptance
|
|||
|
|
|
|||
|
|
- Results map correctly back to input items via `index`.
|
|||
|
|
- A more relevant text scores above a less relevant one on a live call.
|
|||
|
|
- Empty input makes no request.
|
|||
|
|
|
|||
|
|
## Verify
|
|||
|
|
|
|||
|
|
**Harness:** `wiremock` offline, one `#[ignore]` live test.
|
|||
|
|
|
|||
|
|
**Integration test** — `tests/it_rerank.rs`:
|
|||
|
|
1. `a1_bare_array_parsed` — mock returns `[{"index":1,...},{"index":0,...}]`;
|
|||
|
|
assert parsing succeeds.
|
|||
|
|
2. `a2_index_mapping` — with the out-of-order mock above, assert the returned
|
|||
|
|
items correspond to inputs 1 and 0 respectively, not 0 and 1.
|
|||
|
|
3. `a3_empty_no_request` — empty texts; assert zero requests.
|
|||
|
|
4. `a4_apikey_sent` — header present.
|
|||
|
|
5. `a5_live_discriminates` — `#[ignore]`; real gateway, query "what is rust"
|
|||
|
|
against `["rust is a language","bananas"]`; assert the first scores at least
|
|||
|
|
10× the second.
|
|||
|
|
|
|||
|
|
**Command:** `cargo test -p mem-llm rerank` (add `-- --ignored` for a5)
|
|||
|
|
|
|||
|
|
**False pass:**
|
|||
|
|
- Assuming the response preserves input order. A mock that returns results in
|
|||
|
|
input order passes a naive test and the live endpoint returns them sorted,
|
|||
|
|
which silently mislabels every result. Assertion 2 must use an out-of-order
|
|||
|
|
mock.
|
|||
|
|
- Expecting an OpenAI envelope. It will fail immediately against the live route,
|
|||
|
|
but a mock written to match the wrong shape hides that until integration.
|
|||
|
|
|
|||
|
|
## Traps
|
|||
|
|
|
|||
|
|
- Re-associating results by position instead of by `index`. The scores are right
|
|||
|
|
and attached to the wrong documents — a bug that looks like poor retrieval
|
|||
|
|
quality rather than a mapping error.
|
|||
|
|
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
Background: [DESIGN.md](../DESIGN.md) — Verified facts, retrieval
|