mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 02:18:29 +02:00
v1
This commit is contained in:
@@ -0,0 +1,47 @@
|
||||
# Deep review: "Route Retrieval by Evidence" article
|
||||
|
||||
You are doing a rigorous, research-backed review of a draft technical article for the Qdrant website. The bar is a gatekeeping CTO (Andrey): accurate, measurable, lean, in the right channel, honest about tradeoffs, and aware of prior art. Work in two phases — **deep external research first, then strong prioritized feedback**. Do not rewrite the article; produce findings and concrete recommended fixes, and only edit if the user approves.
|
||||
|
||||
Before anything else, read **`article-research/REVIEW_CONTEXT.md`** — it has every artifact path, the environment, and the verified numbers so you can check the article against known-good values instead of re-deriving them.
|
||||
|
||||
## The artifact
|
||||
- Article: `/Users/dylanc/Documents/GitHub/landing_page/qdrant-landing/content/articles/route-retrieval-by-evidence.md` (git branch `signal-driven-retrieval-article`).
|
||||
- Figures: `qdrant-landing/static/articles_data/route-retrieval-by-evidence/` (loop diagram, signal-separation boxplots, gate histogram, routing distribution, cost/quality frontier).
|
||||
- Live Hugo preview: `http://localhost:1313/articles/route-retrieval-by-evidence/`.
|
||||
- **Read the whole article first.** Then verify it against the source rather than trusting it.
|
||||
|
||||
## What the article argues (orientation)
|
||||
A "self-correcting retrieval loop": retrieve once with hybrid search (dense `bge-base` + sparse miniCOIL, RRF), judge the result with a **cheap signal** computed from the result with no extra model call, and escalate to a more expensive action (ColBERT rerank, or query decomposition / IRCoT) only when the signal predicts weak retrieval. Thesis: **route by evidence, not by question shape.** The differentiated core is the *signals* plus an AUC-based method to pick which ones predict weak retrieval on your data, then a threshold gate. Benchmarked on a MuSiQue workload recast as single-hop / multi-hop / unanswerable, and generalized to non-QA use cases.
|
||||
|
||||
## Source material and environment
|
||||
See `REVIEW_CONTEXT.md`. In short: the workshop repo at `/Users/dylanc/Documents/GitHub/self-correcting-loops-workshop/` is the source of truth (lab notebook + `artifacts/*.json` scorecards + `data/`); the live Qdrant on `:6333` has the full `musique` and `musique_colbert` collections so you can re-run the signal benchmark; the workshop `.venv` has the deps; API keys are in `landing_page/.env`.
|
||||
|
||||
## Phase 1 — deep research (do this first, with sub-agents)
|
||||
Fan out cheap-model research sub-agents (one topic each; keep conclusions, drop raw dumps). The point is to find **prior art we may have missed, concepts we skipped, and claims that don't hold up against the literature.** Investigate at least:
|
||||
|
||||
1. **Adaptive / corrective / self-correcting RAG — prior art (highest priority).** Map and read: **CRAG (Corrective RAG)**, **Self-RAG**, **Adaptive-RAG (Jeong et al.)**, **FLARE**, **Self-Ask**, **IRCoT**, least-to-most. For each: what it routes on, and how close it is to "cheap signal + gate." Be skeptical of any novelty claim — CRAG uses a lightweight *retrieval evaluator* to grade retrieval and trigger correction (very close to our signal+gate); Adaptive-RAG routes by a *query-complexity classifier* (exactly the "question shape" approach we position against). Does the article engage this prior art honestly, mis-credit it, or reinvent it without citation? This is the likeliest fatal gap.
|
||||
2. **Query Performance Prediction (QPP).** Post-retrieval predictors (NQC, UQC, Clarity, WIG, SMV, UEF, query feedback, coherence-based) and dense/neural QPP (BERT-QPP, ADG-QPP). Verify the article's claim that the variance signals are "Normalized Query Commitment (NQC)" — check whether raw score variance is actually **UQC (unnormalized)** and NQC requires normalization by a collection score. Identify any cheap, strong predictor we omitted that a reviewer would expect.
|
||||
3. **Selective prediction / abstention / confidence in retrieval and RAG** — how the field frames "know when retrieval failed," to test our "good retrieval = needed evidence in the consumed window" framing against standard IR metrics (recall@k, nDCG) and calibration work.
|
||||
4. **Product-fact grounding** against canonical Qdrant docs: RRF (and `k`), DBSF, ColBERT / late interaction / MaxSim, miniCOIL. Confirm nothing is misstated.
|
||||
5. **Multi-hop QA + decomposition** (IRCoT and relatives) to check our decomposition description is accurate and fairly credited.
|
||||
|
||||
Prefer **web-only** research agents and **cheap models** for this fan-out. Reserve full-strength reasoning for synthesis and critique.
|
||||
|
||||
## Phase 2 — verify accuracy (check the numbers, don't trust them)
|
||||
Extract every quantitative and technical claim from the article and verify it against `REVIEW_CONTEXT.md` and the artifacts; re-run the live signal benchmark on the full `musique` collection if anything looks off (CP2 logic in the workshop notebook; the collection is ready). Confirm the frontier table, signal AUCs, correlations (dense×score ≈ 0.54, max_score×score ≈ 0.92), gate operating points (78%/49% dense alone, 80%/57% OR), and the coverage-saturation stat. Verify every technical claim from first principles. Flag anything stated as universal that is corpus-specific, and any place the single-dataset evidence is overreached.
|
||||
|
||||
## Phase 3 — strong feedback (then, and only then)
|
||||
Deliver, in this order:
|
||||
1. **Research findings** — what the field says, the prior art that matters, the concepts we missed, with sources.
|
||||
2. **Inaccuracies / corrections** (highest priority) — each with the evidence and the concrete fix.
|
||||
3. **Missed concepts / prior art to engage** — what the article must cite or position against to survive expert review.
|
||||
4. **Improvements** — framing, depth, lean (cut anything off-topic), structure.
|
||||
5. **Hygiene** — links (the `#query-decomposition-...` anchor used 3× lives in a separate unmerged PR: a known merge-order dependency, not a master defect), figure/asset sizes, channel fit (it's an `articles/` piece modeled on `miniCOIL`/`bm42`), and the missing `social_preview_image`/`preview` asset.
|
||||
|
||||
Andrey register: terse, specific, no flattery; lead each finding with the problem, name the fix, cite the source. Conclude the piece is solid where it is.
|
||||
|
||||
## Guardrails
|
||||
- **Do not rewrite the article.** Produce the review and recommended changes; edit only on approval.
|
||||
- Invoke the `qdrant-landing-page` and `qdrant-messaging` skills before proposing any copy (voice: "vector search engine" not "vector database"; no em dashes; problem-first; numbers over adjectives).
|
||||
- **macOS file-access caveat:** the repo and workshop live under `~/Documents`; a burst of parallel local sub-agents reading that path can trigger `Operation not permitted` (TCC) mid-session and block reads / venv launches. Keep research agents **web-only**; avoid large parallel bursts of local-file agents; if reads or `.venv` launches start failing with EPERM, stop and ask the user to restart the terminal.
|
||||
- Don't touch other branches or the separate hybrid-queries PR.
|
||||
@@ -0,0 +1,63 @@
|
||||
# Review context: verified facts for "Route Retrieval by Evidence"
|
||||
|
||||
Reference sheet so the review session can check the article against known-good numbers
|
||||
without re-deriving everything. All numbers below were computed this session from the
|
||||
workshop artifacts and the live Qdrant. Re-run to confirm; expect ~1-2 label drift from
|
||||
floating-point/ANN variation (the AUCs and the qualitative picture are stable).
|
||||
|
||||
## Paths
|
||||
- Article: `qdrant-landing/content/articles/route-retrieval-by-evidence.md` (branch `signal-driven-retrieval-article`)
|
||||
- Figures: `qdrant-landing/static/articles_data/route-retrieval-by-evidence/` (loop diagram, signal-separation, gate-histogram, routing-distribution, cost-quality-frontier)
|
||||
- Workshop (source of truth): `/Users/dylanc/Documents/GitHub/self-correcting-loops-workshop/`
|
||||
- `notebooks/lab.ipynb` (CP2 = the signal benchmark), `artifacts/headline_final_v25.json`, `artifacts/targeted_stop_v25.json`, `artifacts/mixed_manifest.json`, `data/corpus.jsonl`, `data/questions_mixed.jsonl`
|
||||
- Research copies: `article-research/` (notebook + both scorecards)
|
||||
- Live Qdrant: `http://localhost:6333` — `musique` (22,808 docs, `dense`+`minicoil`), `musique_colbert` (`dense`+`colbert`)
|
||||
- Python with deps: `/Users/dylanc/Documents/GitHub/self-correcting-loops-workshop/.venv/bin/python` (qdrant-client, fastembed, litellm, sklearn)
|
||||
- Keys: `ANTHROPIC_API_KEY`, `OPENAI_API_KEY` in `landing_page/.env`
|
||||
- Hugo preview: `http://localhost:1313/articles/route-retrieval-by-evidence/`
|
||||
|
||||
## Stack / config (workshop)
|
||||
- Dense: `BAAI/bge-base-en-v1.5` (768-d, cosine) | Sparse: `Qdrant/minicoil-v1` (IDF modifier)
|
||||
- ColBERT action: `answerdotai/answerai-colbert-small-v1` (96-d/token, MaxSim multivector)
|
||||
- Fusion: RRF, server-side. RETRIEVE_N=50, TOP_K=10, ANSWER_K=3 (the answer window)
|
||||
- Agent: Claude Sonnet 4.6 (decompose+answer); STOP autorater: Claude Haiku 4.5
|
||||
- Note: workshop `config.py` comments RRF "k=60"; Qdrant's RRF default is k=2. Article does NOT cite k. Verify if any version ever does.
|
||||
|
||||
## Headline frontier (`headline_final_v25.json`, overall; 321 test Qs = 180 answerable + 141 unanswerable)
|
||||
| policy | recall@3 | full_gold@3 | mrr_first | llm_calls | latency_s |
|
||||
|---|---|---|---|---|---|
|
||||
| always_answer | 0.8167 | 0.700 | 0.8798 | 0 | 0.15 |
|
||||
| always_rerank | 0.8079 | 0.6944 | 0.9051 | 0 | 0.15 |
|
||||
| always_colbert | 0.8023 | 0.6833 | 0.8921 | 0 | 0.30 |
|
||||
| always_decompose | 0.8773 | 0.7944 | 0.9006 | 1.883 | 3.296 |
|
||||
| ladder | 0.8523 | 0.7611 | 0.9131 | 0.778 | 1.498 |
|
||||
|
||||
CIs (`ci_vs`): ladder vs always_answer — recall@3 +0.0356 [0.0032, 0.0676], full_gold@3 +0.0611 [0.0111, 0.1111], em +0.05 [0.0111, 0.0889]. ladder vs always_decompose — recall@3 -0.025 [-0.0509, -0.0028], full_gold@3 -0.0333 [-0.0667, -0.0056], em -0.0111 [-0.0444, 0.0222] (not significant).
|
||||
|
||||
tier_dist: single_hop {1:70, 2:44, 3:6}; multi_hop {1:9, 2:4, 3:47}; unanswerable {1:19, 2:22, 3:100}.
|
||||
|
||||
## Signal benchmark (live, full corpus, 150 calibration Qs, ~100 good / ~50 weak)
|
||||
AUCs (direction-agnostic): dense_variance **0.74**, score_variance **0.70**, max_score **0.66**, evidence_coverage **0.59**, retriever_divergence **0.53**. Bar = 0.65; correlation drop bar = 0.85.
|
||||
- KEPT: dense_variance, score_variance. Dropped: max_score (correlation), coverage + divergence (below bar).
|
||||
- Floors (Youden): DV ≈ 0.039, SV ≈ 0.238 (article does not hardcode these).
|
||||
|
||||
Correlations (Pearson |r|): dense_variance × score_variance = **0.54**; score_variance × max_score = **0.92** (why max_score is dropped); dense × max_score = 0.51; coverage/divergence ≈ 0 with all.
|
||||
|
||||
Gate operating points: dense_variance alone catches **78%** / escalates **49%**; OR of both floors catches **80%** / escalates **57%**. Fired-set Jaccard(dense, score) = 0.52; 12 escalations unique to score, 29 unique to dense.
|
||||
|
||||
evidence_coverage saturation: good mean 0.88, **87% score exactly 1.0**; weak mean 0.82, **69% score exactly 1.0**. A `coverage < 0.99` threshold catches only 16/51 weak (recall 0.31) with 13 false alarms — confirms the AUC is the number that matters.
|
||||
|
||||
## STOP (`targeted_stop_v25.json`)
|
||||
| variant | selective_accuracy | abstain_unans | false_stop_ans |
|
||||
|---|---|---|---|
|
||||
| baseline_hybrid_gentle | 0.5296 | 0.5461 | 0.2111 |
|
||||
| ladder_gentle | 0.4984 | 0.4113 | 0.15 |
|
||||
| ladder_autorater_all | 0.6293 | 0.8652 | 0.4111 |
|
||||
| ladder_targeted | 0.6168 | 0.8085 | 0.3667 |
|
||||
|
||||
## Known open items (already identified, not yet resolved)
|
||||
1. The article links `#query-decomposition-for-multi-hop-questions` (×3) — that anchor lives in a SEPARATE committed-but-unmerged PR (branch `self-correcting-retrieval`, commit `2e2371c30`). Resolves once that PR merges; not a master defect.
|
||||
2. `social_preview_image` + `preview/` assets are referenced in frontmatter but not created (design/publisher item).
|
||||
3. The article calls the variance signals "Normalized Query Commitment (NQC)" — verify against the QPP literature; raw score variance is closer to UQC (unnormalized). Likely imprecise.
|
||||
4. Prior art not yet cited: CRAG, Self-RAG, Adaptive-RAG, FLARE. The "route by evidence not shape" thesis is in direct dialogue with these; positioning gap is the top review risk.
|
||||
5. All numbers come from ONE MuSiQue workload (single dataset). Generalization to non-QA use cases in the article is reasoned, not measured.
|
||||
@@ -0,0 +1,311 @@
|
||||
{
|
||||
"n_test": 321,
|
||||
"n_answerable": 180,
|
||||
"n_unanswerable": 141,
|
||||
"touched_once": false,
|
||||
"test_reuse": {
|
||||
"v2": {
|
||||
"multi_in_prior_ans": 16,
|
||||
"unans_in_prior_unans": 60,
|
||||
"single_src_in_prior_ans": 23
|
||||
},
|
||||
"v1": {
|
||||
"multi_in_prior_ans": 10,
|
||||
"unans_in_prior_unans": 60,
|
||||
"single_src_in_prior_ans": 23
|
||||
}
|
||||
},
|
||||
"lead_metrics": [
|
||||
"recall@3",
|
||||
"full_gold@3"
|
||||
],
|
||||
"note_mrr": "mrr_first = reciprocal rank of the FIRST gold doc; lenient on multi-hop (rewards any one support passage). NOT the headline metric.",
|
||||
"overall": {
|
||||
"always_answer": {
|
||||
"recall@1": 0.6755,
|
||||
"recall@3": 0.8167,
|
||||
"full_gold@3": 0.7,
|
||||
"mrr_first": 0.8798,
|
||||
"qdrant_calls": 1,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.15
|
||||
},
|
||||
"always_colbert": {
|
||||
"recall@1": 0.7153,
|
||||
"recall@3": 0.8023,
|
||||
"full_gold@3": 0.6833,
|
||||
"mrr_first": 0.8921,
|
||||
"qdrant_calls": 2,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.3
|
||||
},
|
||||
"always_rerank": {
|
||||
"recall@1": 0.7204,
|
||||
"recall@3": 0.8079,
|
||||
"full_gold@3": 0.6944,
|
||||
"mrr_first": 0.9051,
|
||||
"qdrant_calls": 1,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.15
|
||||
},
|
||||
"always_decompose": {
|
||||
"recall@1": 0.6898,
|
||||
"recall@3": 0.8773,
|
||||
"full_gold@3": 0.7944,
|
||||
"mrr_first": 0.9006,
|
||||
"qdrant_calls": 3.139,
|
||||
"llm_calls": 1.883,
|
||||
"avg_latency_s": 3.296
|
||||
},
|
||||
"ladder": {
|
||||
"recall@1": 0.7231,
|
||||
"recall@3": 0.8523,
|
||||
"full_gold@3": 0.7611,
|
||||
"mrr_first": 0.9131,
|
||||
"qdrant_calls": 2.206,
|
||||
"llm_calls": 0.778,
|
||||
"avg_latency_s": 1.498
|
||||
}
|
||||
},
|
||||
"by_type": {
|
||||
"single_hop": {
|
||||
"always_answer": {
|
||||
"recall@1": 0.85,
|
||||
"recall@3": 0.975,
|
||||
"full_gold@3": 0.975,
|
||||
"mrr_first": 0.9056,
|
||||
"qdrant_calls": 1,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.15
|
||||
},
|
||||
"always_colbert": {
|
||||
"recall@1": 0.9,
|
||||
"recall@3": 0.9417,
|
||||
"full_gold@3": 0.9417,
|
||||
"mrr_first": 0.9181,
|
||||
"qdrant_calls": 2,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.3
|
||||
},
|
||||
"always_rerank": {
|
||||
"recall@1": 0.9083,
|
||||
"recall@3": 0.95,
|
||||
"full_gold@3": 0.95,
|
||||
"mrr_first": 0.9329,
|
||||
"qdrant_calls": 1,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.15
|
||||
},
|
||||
"always_decompose": {
|
||||
"recall@1": 0.85,
|
||||
"recall@3": 0.975,
|
||||
"full_gold@3": 0.975,
|
||||
"mrr_first": 0.9076,
|
||||
"qdrant_calls": 2.658,
|
||||
"llm_calls": 1.517,
|
||||
"avg_latency_s": 2.674
|
||||
},
|
||||
"ladder": {
|
||||
"recall@1": 0.9,
|
||||
"recall@3": 0.9583,
|
||||
"full_gold@3": 0.9583,
|
||||
"mrr_first": 0.9264,
|
||||
"qdrant_calls": 1.517,
|
||||
"llm_calls": 0.117,
|
||||
"avg_latency_s": 0.402
|
||||
}
|
||||
},
|
||||
"multi_hop": {
|
||||
"always_answer": {
|
||||
"recall@1": 0.3264,
|
||||
"recall@3": 0.5,
|
||||
"full_gold@3": 0.15,
|
||||
"mrr_first": 0.8283,
|
||||
"qdrant_calls": 1,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.15
|
||||
},
|
||||
"always_colbert": {
|
||||
"recall@1": 0.3458,
|
||||
"recall@3": 0.5236,
|
||||
"full_gold@3": 0.1667,
|
||||
"mrr_first": 0.8403,
|
||||
"qdrant_calls": 2,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.3
|
||||
},
|
||||
"always_rerank": {
|
||||
"recall@1": 0.3444,
|
||||
"recall@3": 0.5236,
|
||||
"full_gold@3": 0.1833,
|
||||
"mrr_first": 0.8496,
|
||||
"qdrant_calls": 1,
|
||||
"llm_calls": 0,
|
||||
"avg_latency_s": 0.15
|
||||
},
|
||||
"always_decompose": {
|
||||
"recall@1": 0.3694,
|
||||
"recall@3": 0.6819,
|
||||
"full_gold@3": 0.4333,
|
||||
"mrr_first": 0.8867,
|
||||
"qdrant_calls": 4.1,
|
||||
"llm_calls": 2.617,
|
||||
"avg_latency_s": 4.54
|
||||
},
|
||||
"ladder": {
|
||||
"recall@1": 0.3694,
|
||||
"recall@3": 0.6403,
|
||||
"full_gold@3": 0.3667,
|
||||
"mrr_first": 0.8867,
|
||||
"qdrant_calls": 3.583,
|
||||
"llm_calls": 2.1,
|
||||
"avg_latency_s": 3.688
|
||||
}
|
||||
}
|
||||
},
|
||||
"answers": {
|
||||
"always_answer": {
|
||||
"overall": {
|
||||
"n": 180,
|
||||
"em": 0.5167,
|
||||
"f1": 0.5971
|
||||
},
|
||||
"single_hop": {
|
||||
"n": 120,
|
||||
"em": 0.675,
|
||||
"f1": 0.7716
|
||||
},
|
||||
"multi_hop": {
|
||||
"n": 60,
|
||||
"em": 0.2,
|
||||
"f1": 0.2482
|
||||
}
|
||||
},
|
||||
"always_decompose": {
|
||||
"overall": {
|
||||
"n": 180,
|
||||
"em": 0.5778,
|
||||
"f1": 0.6575
|
||||
},
|
||||
"single_hop": {
|
||||
"n": 120,
|
||||
"em": 0.6833,
|
||||
"f1": 0.7789
|
||||
},
|
||||
"multi_hop": {
|
||||
"n": 60,
|
||||
"em": 0.3667,
|
||||
"f1": 0.4148
|
||||
}
|
||||
},
|
||||
"ladder": {
|
||||
"overall": {
|
||||
"n": 180,
|
||||
"em": 0.5667,
|
||||
"f1": 0.6321
|
||||
},
|
||||
"single_hop": {
|
||||
"n": 120,
|
||||
"em": 0.6917,
|
||||
"f1": 0.77
|
||||
},
|
||||
"multi_hop": {
|
||||
"n": 60,
|
||||
"em": 0.3167,
|
||||
"f1": 0.3564
|
||||
}
|
||||
}
|
||||
},
|
||||
"selective_accuracy": {
|
||||
"always_answer": {
|
||||
"selective_accuracy": 0.5296,
|
||||
"abstain_rate_unans": 0.5461,
|
||||
"answerable_em": 0.5167
|
||||
},
|
||||
"ladder": {
|
||||
"selective_accuracy": 0.4984,
|
||||
"abstain_rate_unans": 0.4113,
|
||||
"answerable_em": 0.5667
|
||||
},
|
||||
"always_decompose": {
|
||||
"selective_accuracy": 0.4829,
|
||||
"abstain_rate_unans": 0.3617,
|
||||
"answerable_em": 0.5778
|
||||
}
|
||||
},
|
||||
"tier_dist": {
|
||||
"single_hop": {
|
||||
"1": 70,
|
||||
"2": 44,
|
||||
"3": 6
|
||||
},
|
||||
"multi_hop": {
|
||||
"1": 9,
|
||||
"2": 4,
|
||||
"3": 47
|
||||
},
|
||||
"unanswerable": {
|
||||
"1": 19,
|
||||
"2": 22,
|
||||
"3": 100
|
||||
}
|
||||
},
|
||||
"ci_vs": {
|
||||
"always_answer": {
|
||||
"recall@3": {
|
||||
"lift": 0.0356,
|
||||
"ci95": [
|
||||
0.0032,
|
||||
0.0676
|
||||
]
|
||||
},
|
||||
"full_gold@3": {
|
||||
"lift": 0.0611,
|
||||
"ci95": [
|
||||
0.0111,
|
||||
0.1111
|
||||
]
|
||||
},
|
||||
"em": {
|
||||
"lift": 0.05,
|
||||
"ci95": [
|
||||
0.0111,
|
||||
0.0889
|
||||
]
|
||||
}
|
||||
},
|
||||
"always_decompose": {
|
||||
"recall@3": {
|
||||
"lift": -0.025,
|
||||
"ci95": [
|
||||
-0.0509,
|
||||
-0.0028
|
||||
]
|
||||
},
|
||||
"full_gold@3": {
|
||||
"lift": -0.0333,
|
||||
"ci95": [
|
||||
-0.0667,
|
||||
-0.0056
|
||||
]
|
||||
},
|
||||
"em": {
|
||||
"lift": -0.0111,
|
||||
"ci95": [
|
||||
-0.0444,
|
||||
0.0222
|
||||
]
|
||||
}
|
||||
}
|
||||
},
|
||||
"gate": [
|
||||
"dense_variance",
|
||||
"score_variance"
|
||||
],
|
||||
"answer_k": 3,
|
||||
"latency_model": {
|
||||
"description": "Estimated routing latency before final answer generation; final answer call is shared and excluded.",
|
||||
"qdrant_call_s": 0.15,
|
||||
"llm_subquery_call_s": 1.5
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,325 @@
|
||||
{
|
||||
"cells": [
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "# Self-Correcting Agentic Retrieval Loops\n\nMost retrieval agents run the same pipeline on every query. That overspends on easy\nquestions and underserves hard ones. This notebook builds a **self-correcting loop**\ninstead: retrieve once with hybrid search, judge the result with a **cheap signal**,\nand escalate to a more expensive action only when the signal says retrieval was weak.\n\nThe lesson is a **method**, not a fixed recipe. The signals that work here are specific\nto this corpus; the way you find them transfers to yours:\n\n1. Define what \"good retrieval\" means for your task.\n2. Engineer cheap candidate signals from the retrieval result.\n3. Benchmark which signals predict weak evidence on your data.\n4. Turn the winning signal into a gate, and route each query to the cheapest action that fixes it.\n5. Measure quality and cost together.\n\n**What \"good retrieval\" means here.** The agent answers from a **top-3 answer context**:\nit reads only the first three retrieved passages. So good retrieval means the supporting\nevidence lands in that top-3 window. A cheap signal tells you when it didn't, before you\nspend an LLM call finding out.\n\n> This notebook is the runnable companion to the [Self-Correcting Retrieval Loops tutorial](https://qdrant.tech/documentation/tutorials-search-engineering/self-correcting-retrieval-loops/).\n> It builds two small collections from a curated MuSiQue slice so the whole loop runs in\n> a few minutes. The headline cost and quality numbers near the end are **reported from a\n> run on the full workload**, not recomputed on the slice."
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "## Setup\n\nInstall the client, the local embedders, and the agent's LLM client."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "%pip install -q \"qdrant-client[fastembed]\" litellm scikit-learn pandas numpy matplotlib"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "This notebook uses a **Qdrant Cloud** cluster to hold the collections and generates all\nembeddings **locally with FastEmbed**. The signals read raw, per-retriever score views\n(`bge-base` dense cosines, miniCOIL ranks, ColBERT MaxSim), so the exact models matter:\nthey are what the cited results were measured on. Create a free cluster at\n[cloud.qdrant.io](https://cloud.qdrant.io/), then set `QDRANT_URL`, `QDRANT_API_KEY`, and\n`ANTHROPIC_API_KEY` as Colab secrets (the key icon in the left sidebar)."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "import os, json, re, string, statistics, urllib.request\nfrom pathlib import Path\nfrom dataclasses import dataclass\n\nimport numpy as np\nimport pandas as pd\nimport litellm\nimport matplotlib.pyplot as plt\nfrom qdrant_client import QdrantClient, models\nfrom fastembed import TextEmbedding, SparseTextEmbedding, LateInteractionTextEmbedding\nfrom sklearn.metrics import roc_auc_score\n\n# Credentials. On Colab, set these as secrets (key icon). Get the values from cloud.qdrant.io.\ntry:\n from google.colab import userdata\n QDRANT_URL = userdata.get(\"QDRANT_URL\")\n QDRANT_API_KEY = userdata.get(\"QDRANT_API_KEY\")\n os.environ[\"ANTHROPIC_API_KEY\"] = userdata.get(\"ANTHROPIC_API_KEY\")\nexcept Exception:\n QDRANT_URL = os.environ.get(\"QDRANT_URL\", \"http://localhost:6333\")\n QDRANT_API_KEY = os.environ.get(\"QDRANT_API_KEY\")\n\nclient = QdrantClient(url=QDRANT_URL, api_key=QDRANT_API_KEY, timeout=120)\n\n# Collection schema, model ids, and retrieval sizes (the whole config in one place).\nCOLLECTION = \"musique\" # dense (bge) + miniCOIL sparse\nCOLBERT_COLLECTION = \"musique_colbert\" # dense (reused) + ColBERT multivector\nDENSE_MODEL = \"BAAI/bge-base-en-v1.5\" # 768-d dense, cosine\nMINICOIL_MODEL = \"Qdrant/minicoil-v1\" # word-sense-aware sparse, IDF modifier\nCOLBERT_MODEL = \"answerdotai/answerai-colbert-small-v1\"\nDENSE_VEC, MINICOIL_VEC, COLBERT_VEC = \"dense\", \"minicoil\", \"colbert\"\nRETRIEVE_N = 50 # per-retriever prefetch depth before fusion\nTOP_K = 10 # signal / pool window\nANSWER_K = 3 # focused passages the LLM reads to answer\nAGENT_MODEL = \"anthropic/claude-sonnet-4-6\" # decompose + answer\nFAST_MODEL = \"anthropic/claude-haiku-4-5\" # the fast sufficiency autorater (STOP)\n\n# The query-side embedders (local, ONNX). Warm them so later timings are clean.\ndense_model = TextEmbedding(DENSE_MODEL)\nminicoil_model = SparseTextEmbedding(MINICOIL_MODEL)\ncolbert_model = LateInteractionTextEmbedding(COLBERT_MODEL)\nfor warm in (dense_model, minicoil_model):\n next(iter(warm.query_embed(\"warm up\")))\nnext(iter(colbert_model.query_embed(\"warm up\")))\nprint(\"models ready\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### The curated corpus slice\n\nThis pulls a few hundred MuSiQue passages and a mixed set of questions (single-hop,\nmulti-hop, and unanswerable). The questions carry their split label, so the calibration\nset used to benchmark signals is available directly."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "# The slice ships with this notebook. Locally these files already sit next to it; on Colab\n# they download once.\nDATA_REPO = \"https://raw.githubusercontent.com/qdrant/examples/master/self-correcting-retrieval-loops\" # TODO: confirm path when the examples PR merges\nfor fn in (\"slice_corpus.jsonl\", \"slice_questions.jsonl\",\n \"headline_final_v25.json\", \"targeted_stop_v25.json\"):\n if not Path(fn).exists():\n urllib.request.urlretrieve(f\"{DATA_REPO}/{fn}\", fn)\n\ncorpus = [json.loads(l) for l in open(\"slice_corpus.jsonl\")]\nquestions = [json.loads(l) for l in open(\"slice_questions.jsonl\")]\nby_id = {q[\"id\"]: q for q in questions}\n\ndef doc_embed_text(d):\n title, text = (d.get(\"title\") or \"\").strip(), (d.get(\"text\") or \"\").strip()\n return f\"{title}. {text}\" if title else text\n\nprint(f\"{len(corpus)} passages, {len(questions)} questions \"\n f\"({sum(q['split']=='calibration' for q in questions)} calibration)\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### Build the two collections\n\nOne collection holds the **hybrid baseline** (dense + miniCOIL sparse). A second holds the\n**ColBERT** multivector used by one of the corrective actions. Both reuse the same dense\nvectors. The named-vector and multivector mechanics are covered in the\n[multivector tutorial](https://qdrant.tech/documentation/tutorials-search-engineering/using-multivector-representations/);\nhere we just build them."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def to_sparse(emb):\n return models.SparseVector(indices=emb.indices.tolist(), values=emb.values.tolist())\n\ndef batched_upsert(coll, points, batch_size=128):\n for i in range(0, len(points), batch_size):\n client.upsert(coll, points[i:i + batch_size], wait=True)\n\ndef build_collections():\n texts = [doc_embed_text(d) for d in corpus]\n dense_vecs = [v.tolist() for v in dense_model.embed(texts, batch_size=64)]\n minicoil_vecs = list(minicoil_model.embed(texts, batch_size=64))\n colbert_vecs = [[r.tolist() for r in v] for v in colbert_model.embed(texts, batch_size=32)]\n\n if client.collection_exists(COLLECTION):\n client.delete_collection(COLLECTION)\n client.create_collection(\n COLLECTION,\n vectors_config={DENSE_VEC: models.VectorParams(size=768, distance=models.Distance.COSINE)},\n sparse_vectors_config={MINICOIL_VEC: models.SparseVectorParams(modifier=models.Modifier.IDF)},\n )\n if client.collection_exists(COLBERT_COLLECTION):\n client.delete_collection(COLBERT_COLLECTION)\n client.create_collection(\n COLBERT_COLLECTION,\n vectors_config={\n DENSE_VEC: models.VectorParams(size=768, distance=models.Distance.COSINE),\n COLBERT_VEC: models.VectorParams(\n size=96, distance=models.Distance.COSINE,\n multivector_config=models.MultiVectorConfig(comparator=models.MultiVectorComparator.MAX_SIM),\n hnsw_config=models.HnswConfigDiff(m=0)), # rescorer only; reached via the dense prefetch\n },\n )\n base, colb = [], []\n for i, d in enumerate(corpus):\n payload = {\"title\": d[\"title\"], \"text\": d[\"text\"]}\n base.append(models.PointStruct(id=d[\"doc_id\"],\n vector={DENSE_VEC: dense_vecs[i], MINICOIL_VEC: to_sparse(minicoil_vecs[i])}, payload=payload))\n colb.append(models.PointStruct(id=d[\"doc_id\"],\n vector={DENSE_VEC: dense_vecs[i], COLBERT_VEC: colbert_vecs[i]}, payload=payload))\n batched_upsert(COLLECTION, base)\n batched_upsert(COLBERT_COLLECTION, colb)\n\nwant = len(corpus)\nhave = client.count(COLLECTION).count if client.collection_exists(COLLECTION) else -1\nif have != want:\n build_collections()\nprint(f\"'{COLLECTION}': {client.count(COLLECTION).count} | \"\n f\"'{COLBERT_COLLECTION}': {client.count(COLBERT_COLLECTION).count}\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### The LLM call\n\nThe agent uses an LLM for three jobs: generating a decomposition sub-question, writing the\nfinal answer, and deciding when the evidence is enough. All three go through one helper set\nto deterministic output, with retries for transient API errors."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def ask_llm(system, user, max_tokens=256, model=AGENT_MODEL, temperature=0.0):\n litellm.suppress_debug_info = True\n resp = litellm.completion(\n model=model,\n messages=[{\"role\": \"system\", \"content\": system}, {\"role\": \"user\", \"content\": user}],\n max_tokens=max_tokens, temperature=temperature, timeout=45, num_retries=3,\n )\n return (resp.choices[0].message.content or \"\").strip()"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "## The hybrid baseline, and where it breaks\n\nThe workload has three kinds of questions:\n\n| kind | what has to happen |\n|---|---|\n| single-hop | one retrieved passage contains the answer |\n| multi-hop | the first passage reveals the next thing to look up |\n| unanswerable | the corpus lacks the evidence, so the agent should abstain |\n\nThe baseline is one **hybrid** retrieval. Qdrant gets candidates two ways (dense vectors for\nsemantic matches, miniCOIL sparse for lexical matches) and fuses the ranks server-side with\n[Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/search/hybrid-queries/). No\ncross-encoder in the baseline: the signals read the fusion scores directly."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def embed(text):\n dense = next(iter(dense_model.query_embed(text))).tolist()\n minicoil = next(iter(minicoil_model.query_embed(text)))\n return dense, minicoil\n\ndef hybrid_search(question, limit=TOP_K, enc=None):\n dense, minicoil = enc or embed(question)\n return client.query_points(\n COLLECTION,\n prefetch=[\n models.Prefetch(query=dense, using=DENSE_VEC, limit=RETRIEVE_N),\n models.Prefetch(query=to_sparse(minicoil), using=MINICOIL_VEC, limit=RETRIEVE_N),\n ],\n query=models.FusionQuery(fusion=models.Fusion.RRF),\n limit=limit, with_payload=True,\n ).points\n\n@dataclass\nclass Passage:\n doc_id: str\n title: str\n text: str\n score: float\n\ndef to_passages(points):\n return [Passage(p.id, p.payload[\"title\"], p.payload[\"text\"], p.score) for p in points]\n\ndef show_hits(hits, gold_ids, k=3, snippet=95):\n gold = set(gold_ids)\n for rank, h in enumerate(hits[:k], 1):\n did = h.id if hasattr(h, \"id\") else h.doc_id\n title = h.payload[\"title\"] if hasattr(h, \"payload\") else h.title\n text = h.payload[\"text\"] if hasattr(h, \"payload\") else h.text\n print(f\" [{'GOLD' if did in gold else ' '}] #{rank} {title}\")\n print(f\" {' '.join((text or '').split())[:snippet]}...\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "**Single-hop: the cheap path is enough.** One passage carries the answer, and hybrid retrieval\nputs it in the top-3 answer context."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "single = by_id[\"2hop__101521_42157__h0\"]\nprint(f\"Q: {single['question']}\\n\")\nshow_hits(hybrid_search(single[\"question\"]), single[\"gold_doc_ids\"])"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "**Multi-hop: the baseline breaks.** The first retrieval can only find the bridge entity. The\npassage that answers the question is never retrieved from the question as written,\nso no amount of reordering recovers it. This is a recall problem, not a ranking problem."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "multi = by_id[\"2hop__615262_131886\"]\ngold = set(multi[\"gold_doc_ids\"])\nprint(f\"Q: {multi['question']}\\n\")\nprint(\"baseline retrieve, question as written (top-3):\")\nshow_hits(hybrid_search(multi[\"question\"]), gold)\nprint(f\"\\ngold passages in the top-3: {len({p.id for p in hybrid_search(multi['question'])[:3]} & gold)} of {len(gold)}\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "## Signals: predicting weak retrieval cheaply\n\nA **signal** is a low-cost metric computed from the retrieval result that predicts whether the\nevidence is weak, before the agent spends anything to find out. A useful signal has to do two\nthings: separate good retrievals from weak ones on calibration data, and be cheap enough to run\non every query.\n\nThe plan: on calibration data we know whether all the gold landed in the top-3. We engineer\ncandidate signals, score each by how well it separates good from weak (its AUC), keep the ones\nthat clear a bar, and drop near-duplicates.\n\n### The candidates, by family\n\nWe test five candidates. The split that matters is whether a signal reads the already-fused\nhybrid result or asks Qdrant for one extra raw retriever view.\n\n| family | signal | what it reads | flags weak when |\n|---|---|---|---|\n| height | `max_score` | top-1 fused score | low |\n| spread (fused) | `score_variance` | spread of the fused top-k scores | low (flat ranking) |\n| spread (raw) | `dense_variance` | spread of the raw dense cosines | low (dense can't separate its hits) |\n| coverage | `evidence_coverage` | question entities present in the top-k text | low (text misses them) |\n| agreement | `retriever_divergence` | dense vs miniCOIL top-k overlap | high (they disagree) |\n\nRRF combines ranks well but flattens score shape, so we test both the fused score spread (free)\nand the raw dense cosine spread (one extra Qdrant query, no extra model call)."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def dense_ranking(question, enc=None):\n # the RAW dense cosine ranking, pre-fusion: [(doc_id, cosine), ...]\n dense, _ = enc or embed(question)\n pts = client.query_points(COLLECTION, query=dense, using=DENSE_VEC, limit=TOP_K, with_payload=False).points\n return [(p.id, p.score) for p in pts]\n\ndef minicoil_ranking(question, enc=None):\n # the RAW miniCOIL ranking, pre-fusion: [doc_id, ...]\n _, minicoil = enc or embed(question)\n pts = client.query_points(COLLECTION, query=to_sparse(minicoil), using=MINICOIL_VEC,\n limit=TOP_K, with_payload=False).points\n return [p.id for p in pts]\n\n_QUESTION_STOP = {\"what\", \"who\", \"whom\", \"whose\", \"where\", \"when\", \"which\", \"why\", \"how\",\n \"is\", \"was\", \"are\", \"were\", \"did\", \"do\", \"does\", \"the\", \"a\", \"an\", \"name\",\n \"in\", \"of\", \"on\", \"at\", \"to\", \"for\", \"by\", \"as\", \"that\", \"this\"}\n_NAME_CONNECTORS = {\"of\", \"the\", \"and\", \"de\", \"von\", \"van\", \"del\", \"la\", \"el\", \"da\", \"di\", \"&\"}\n\ndef question_entities(question):\n # the capitalized spans + 4-digit years in the question, e.g. {\"pulp fiction\", \"1994\"}\n toks = (question or \"\").split()\n ents, cur = [], []\n for i, tok in enumerate(toks):\n w = tok.strip(string.punctuation)\n if not w:\n if cur: ents.append(\" \".join(cur)); cur = []\n continue\n if w[0].isupper() and not (i == 0 and w.lower() in _QUESTION_STOP):\n cur.append(w)\n elif cur and w.lower() in _NAME_CONNECTORS and i + 1 < len(toks) \\\n and toks[i + 1].strip(string.punctuation)[:1].isupper():\n cur.append(w)\n elif cur:\n ents.append(\" \".join(cur)); cur = []\n if cur: ents.append(\" \".join(cur))\n out = {e.lower() for e in ents if len(e) >= 2}\n out |= set(re.findall(r\"\\b\\d{4}\\b\", question or \"\"))\n return out"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "Now the five candidate signals. `retrieve_signals` runs the shared reads once (one encode, the\nfused hybrid result, and the two raw single-retriever rankings); each signal then reads one\nvalue from those results."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def retrieve_signals(question, enc=None):\n enc = enc or embed(question)\n return hybrid_search(question, enc=enc), dense_ranking(question, enc=enc), minicoil_ranking(question, enc=enc)\n\ndef max_score(fused): return fused[0].score\ndef score_variance(fused): return statistics.pstdev([p.score for p in fused])\ndef dense_variance(dense): return statistics.pstdev([s for _, s in dense])\n\ndef evidence_coverage(question, fused):\n ents = question_entities(question)\n if not ents:\n return 1.0\n blob = \" \".join(f\"{p.payload['title']} {p.payload['text']}\" for p in fused).lower()\n return sum(1 for e in ents if e in blob) / len(ents)\n\ndef retriever_divergence(dense, sparse_ids):\n dense_ids = [i for i, _ in dense]\n return 1.0 - len(set(dense_ids) & set(sparse_ids)) / max(len(dense_ids), len(sparse_ids))"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### Benchmark the candidates, live\n\nWe score each signal on the calibration questions. The label is `full_gold@3`: did **all** the\nsupporting passages land in the top-3 answer context? AUC reads as a separation score: 0.5 is\nchance, 1.0 is perfect. This runs against Qdrant only, no LLM calls."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "cal_questions = [q for q in questions\n if q[\"split\"] == \"calibration\" and q.get(\"answerable\") and q.get(\"gold_doc_ids\")]\n\ndef feature_row(q):\n fused, dense, sparse_ids = retrieve_signals(q[\"question\"])\n gold = set(q[\"gold_doc_ids\"])\n return {\n \"full_gold_label\": 1 if gold.issubset({p.id for p in fused[:ANSWER_K]}) else 0,\n \"dense_variance\": dense_variance(dense), \"score_variance\": score_variance(fused),\n \"max_score\": max_score(fused), \"evidence_coverage\": evidence_coverage(q[\"question\"], fused),\n \"retriever_divergence\": retriever_divergence(dense, sparse_ids),\n }\n\ncalibration = [feature_row(q) for q in cal_questions]\nlabels = [r[\"full_gold_label\"] for r in calibration]\nprint(f\"benchmarked {len(calibration)} calibration questions: \"\n f\"{sum(labels)} good / {len(labels) - sum(labels)} weak retrievals\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "Keep signals with separation at least **0.65**, then drop any that is a near-duplicate\n(absolute correlation above **0.85**) of a stronger kept signal."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "AUC_BAR, CORR_MAX = 0.65, 0.85\nSIGNALS = [\"dense_variance\", \"score_variance\", \"max_score\", \"evidence_coverage\", \"retriever_divergence\"]\n\ndef signal_auc(name):\n raw = roc_auc_score(labels, [r[name] for r in calibration])\n return max(raw, 1 - raw) # separation strength, regardless of direction\n\ndef abs_corr(a, b):\n return abs(float(np.corrcoef([r[a] for r in calibration], [r[b] for r in calibration])[0, 1]))\n\naucs = {name: signal_auc(name) for name in SIGNALS}\nkept = []\nfor name in sorted((n for n in aucs if aucs[n] >= AUC_BAR), key=lambda n: -aucs[n]):\n if all(abs_corr(name, k) <= CORR_MAX for k in kept):\n kept.append(name)\nkept = set(kept)\n\npd.DataFrame([{\"signal\": n, \"AUC\": round(aucs[n], 3), \"verdict\": \"kept\" if n in kept else \"dropped\"}\n for n in sorted(aucs, key=lambda n: -aucs[n])])"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "The same benchmark as a picture: when the good and weak boxes pull apart, the signal separates;\nwhen they overlap, it doesn't."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def plot_signal_separation(features, aucs, kept):\n good = lambda col: [r[col] for r in features if r[\"full_gold_label\"] == 1]\n weak = lambda col: [r[col] for r in features if r[\"full_gold_label\"] == 0]\n order = sorted(aucs, key=lambda s: -aucs[s])\n fig, axes = plt.subplots(1, len(order), figsize=(14, 3.4))\n for ax, name in zip(axes, order):\n b = ax.boxplot([good(name), weak(name)], tick_labels=[\"good\", \"weak\"],\n widths=0.6, patch_artist=True, showfliers=False)\n b[\"boxes\"][0].set(facecolor=\"#1f9d55\", alpha=0.55)\n b[\"boxes\"][1].set(facecolor=\"#d64545\", alpha=0.55)\n ax.set_title(f\"{name}\\nAUC {aucs[name]:.2f} ({'kept' if name in kept else 'dropped'})\", fontsize=9)\n ax.tick_params(labelsize=8)\n fig.suptitle(\"Each signal on good vs weak retrievals: separation = predictive power\", fontsize=11)\n fig.tight_layout(rect=[0, 0, 1, 0.92]); plt.show()\n\nplot_signal_separation(calibration, aucs, kept)"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "Two signals survive: `dense_variance` (spread of the raw dense cosines) and `score_variance`\n(spread of the fused RRF scores). Low spread means the retriever couldn't separate its top hits,\nwhich is what weak retrieval looks like. The other three either don't separate well enough or\nduplicate a stronger signal.\n\n**This is the part that transfers, not the winners.** On a different corpus the coverage or\ndivergence signal might win instead. Engineer candidates, benchmark them against labeled\nretrieval outcomes on your data, and keep what separates. In production, pick signals on one\nsplit, tune the threshold on another, and keep a test split held out."
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### The gate: one signal, one floor\n\nThe gate turns a signal into a decision: answer from the current result, or escalate. Each\nsignal gets a floor; below it, retrieval is treated as weak. Raising the floor catches more weak\nretrievals but escalates more queries. `youden_floor` picks the threshold that best balances\ncatching weak retrievals against false alarms."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "weak = [r[\"full_gold_label\"] == 0 for r in calibration] # True = retrieval was weak\n\ndef youden_floor(values):\n best = None\n for thr in sorted(set(values)):\n pred = [v < thr for v in values]\n tp = sum(p and w for p, w in zip(pred, weak)); fp = sum(p and not w for p, w in zip(pred, weak))\n fn = sum((not p) and w for p, w in zip(pred, weak)); tn = sum((not p) and not w for p, w in zip(pred, weak))\n tpr = tp / (tp + fn) if (tp + fn) else 0.0\n fpr = fp / (fp + tn) if (fp + tn) else 0.0\n if best is None or (tpr - fpr) > best[1]:\n best = (thr, tpr - fpr)\n return best[0]\n\nDV_FLOOR = youden_floor([r[\"dense_variance\"] for r in calibration])\nSV_FLOOR = youden_floor([r[\"score_variance\"] for r in calibration])\nprint(f\"calibrated floors: dense_variance < {DV_FLOOR:.5f} score_variance < {SV_FLOOR:.5f}\\n\")\n\nfires = [r[\"dense_variance\"] < DV_FLOOR for r in calibration]\ntp = sum(f and w for f, w in zip(fires, weak)); fp = sum(f and not w for f, w in zip(fires, weak))\nfn = sum((not f) and w for f, w in zip(fires, weak))\nprint(\"at the dense_variance floor:\")\nprint(f\" precision = {tp / (tp + fp):.3f} (of the queries we escalate, how many were truly weak)\")\nprint(f\" recall = {tp / (tp + fn):.3f} (of the truly weak queries, how many we catch)\")\nprint(f\" escalation rate = {sum(fires) / len(fires):.2f} (fraction of queries sent past the cheap path)\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "The histogram shows where the floor sits and what it buys: how much of the weak (red) mass falls\nbelow the line (caught) versus how much of all traffic falls below it (escalated)."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def plot_gate(features, floor, signal=\"dense_variance\", desc=\"spread of the raw dense scores\"):\n good = [r[signal] for r in features if r[\"full_gold_label\"] == 1]\n weakv = [r[signal] for r in features if r[\"full_gold_label\"] == 0]\n vals = np.array([r[signal] for r in features])\n is_weak = np.array([r[\"full_gold_label\"] == 0 for r in features])\n below = vals < floor\n recall = (below & is_weak).sum() / max(is_weak.sum(), 1)\n esc = below.mean()\n fig, ax = plt.subplots(figsize=(8.5, 4.2))\n bins = np.linspace(0, vals.max(), 26)\n ax.hist(good, bins=bins, alpha=0.6, color=\"#1f9d55\", label=\"good (full gold present)\")\n ax.hist(weakv, bins=bins, alpha=0.6, color=\"#d64545\", label=\"weak (full gold missing)\")\n ax.axvspan(0, floor, color=\"black\", alpha=0.06)\n ax.axvline(floor, color=\"black\", ls=\"--\", lw=1.6, label=f\"gate floor = {floor:.3f}\")\n ax.set_title(f\"Low spread predicts weak retrieval\\n{signal} floor {floor:.3f} -> \"\n f\"catches {recall:.0%} of weak retrievals, escalates {esc:.0%} of all queries\", fontsize=10)\n ax.set_xlabel(f\"{signal} ({desc})\"); ax.set_ylabel(\"queries\")\n ax.legend(fontsize=8); fig.tight_layout(); plt.show()\n\nplot_gate(calibration, DV_FLOOR)"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "Wire the gate the loop will call. It reads `score_variance` from the fused result and\n`dense_variance` from the dense-only ranking; if either falls below its floor, the retrieval is\nweak."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def retrieval_is_weak(fused, dense):\n return score_variance(fused) < SV_FLOOR or dense_variance(dense) < DV_FLOOR"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "> **Targeting recall instead of balance.** The Youden floor balances catches against false alarms.\n> If a missed weak retrieval costs more than an extra escalation, raise the floor to hit a recall\n> target (say, catch 90% of weak retrievals) at the price of escalating more. The loop below keeps\n> the balanced floors."
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "## The corrective loop\n\nThe gate says *when* retrieval is weak. The loop decides *what to do* about it, matching the fix\nto the failure mode:\n\n- **Weak single-hop lookup:** the right passage exists but ranks too low. Fix the ranking.\n- **Weak multi-hop query:** the next passage isn't reachable from the question as written. Recover\n the missing evidence.\n\n### Action: ColBERT late interaction (single-hop precision)\n\nColBERT scores query tokens against passage tokens, so it can promote a specific passage that the\npooled dense vector blurred away. We prefetch a dense candidate pool, then rescore with the ColBERT\nmultivector (MaxSim). The mechanics of late interaction and multivectors are covered in the\n[ColBERT how-to](https://qdrant.tech/documentation/fastembed/fastembed-colbert/) and the\n[multivector tutorial](https://qdrant.tech/documentation/tutorials-search-engineering/using-multivector-representations/);\nhere it is just the corrective action the gate routes to."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "def colbert_rerank(question, limit=TOP_K, enc=None):\n dense, _ = enc or embed(question)\n colbert_vecs = [v.tolist() for v in next(iter(colbert_model.query_embed(question)))]\n return client.query_points(\n COLBERT_COLLECTION,\n prefetch=[models.Prefetch(query=dense, using=DENSE_VEC, limit=RETRIEVE_N)],\n query=colbert_vecs, using=COLBERT_VEC, limit=limit, with_payload=True,\n ).points\n\ntier2 = by_id[\"2hop__82744_23140__h0\"]\ngold = tier2[\"gold_doc_ids\"]\nprint(f\"Q: {tier2['question']}\\n\")\nprint(\"hybrid retrieve, the answer passage buried:\")\nshow_hits(hybrid_search(tier2[\"question\"]), gold)\nprint(\"\\nColBERT late interaction, the answer passage promoted:\")\nshow_hits(colbert_rerank(tier2[\"question\"]), gold)"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "Hybrid retrieval matched the entity in the question and buried the passage with the actual answer.\nColBERT promoted it to the top. Across the full validation set ColBERT adds little on average, but\nit earns its keep on the weak single-hop retrievals the gate routes to it. Cross-encoder rerankers\nare a viable alternative; ColBERT produced the best results on this corpus.\n\n### Action: decompose (IRCoT) for multi-hop recall\n\nA reranker can only reorder passages already retrieved. When the needed evidence is missing\nentirely, the agent decomposes: read the evidence so far, ask the next still-missing sub-question,\nretrieve again, and repeat until it has enough. This retrieve-read-ask loop is **IRCoT**.\nDecomposition as a standalone technique is covered in\n[Hybrid and Multi-Stage Queries](https://qdrant.tech/documentation/search/hybrid-queries/#query-decomposition-for-multi-hop-questions)."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "IRCOT_SYSTEM = (\n \"You are running an iterative retrieve-and-reason loop to answer a multi-hop \"\n \"question. Given the main question, the evidence retrieved so far, and the \"\n \"sub-questions already asked, output the NEXT single sub-question whose answer is \"\n \"still MISSING and is needed to answer the main question. Make it self-contained: \"\n \"name entities explicitly, resolving any bridge entity from the evidence so far. \"\n \"If the evidence already contains everything needed to answer the main question, \"\n \"reply with exactly: ENOUGH. Output ONLY the sub-question text or ENOUGH - no prose.\"\n)\n\ndef evidence_digest(pools, max_docs=6, max_chars=160):\n seen, lines = set(), []\n for pool in pools:\n for c in pool:\n if c.doc_id in seen:\n continue\n seen.add(c.doc_id)\n lines.append(f\"- {c.title}: {(c.text or '')[:max_chars]}\")\n if len(lines) >= max_docs:\n return \"\\n\".join(lines)\n return \"\\n\".join(lines) if lines else \"(none)\"\n\ndef next_subquery(question, pools, sub_queries):\n user = (\n f\"Main question: {question}\\n\\n\"\n f\"Evidence so far:\\n{evidence_digest(pools)}\\n\\n\"\n \"Sub-questions already asked:\\n\" + (\"\\n\".join(f\"- {s}\" for s in sub_queries) or \"(none)\") +\n \"\\n\\nNext sub-question (or ENOUGH):\"\n )\n t = (ask_llm(IRCOT_SYSTEM, user, max_tokens=80) or \"\").strip()\n if not t or t.upper().startswith(\"ENOUGH\"):\n return None\n return re.sub(r\"^[\\-\\d\\.\\)\\s]+\", \"\", t.splitlines()[0]).strip() or None\n\ndef union_pool(pools, k=TOP_K):\n # merge the per-hop pools, keep the max score per doc; this is what lets a later hop's\n # passage surface into the final answer context.\n best = {}\n for pool in pools:\n for c in pool:\n if c.doc_id not in best or c.score > best[c.doc_id].score:\n best[c.doc_id] = c\n return sorted(best.values(), key=lambda c: c.score, reverse=True)[:k]\n\ndef decompose(question, max_hops=4, hop0=None):\n pools = [hop0 if hop0 is not None else to_passages(hybrid_search(question))]\n sub_queries = []\n for _ in range(max_hops - 1):\n next_q = next_subquery(question, pools, sub_queries)\n if next_q is None:\n break\n sub_queries.append(next_q)\n pools.append(to_passages(hybrid_search(next_q)))\n return union_pool(pools, TOP_K), sub_queries\n\nmulti = by_id[\"2hop__615262_131886\"]\ngold = multi[\"gold_doc_ids\"]\nprint(f\"Q: {multi['question']}\\n\")\nprint(\"hybrid retrieve (single pass), only the first hop is reachable:\")\nshow_hits(hybrid_search(multi[\"question\"]), gold)\npool, sub_queries = decompose(multi[\"question\"])\nprint(\"\\ndecompose reads hop 1, then asks the still-missing sub-question:\")\nfor sq in sub_queries:\n print(f\" -> {sq}\")\nprint(\"\\nunioned evidence, the second supporting passage now in context:\")\nshow_hits(pool, gold, k=4)"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### Assemble the loop: `solve()`\n\n`solve()` retrieves once, reads the gate, and escalates only as far as needed. The gate decides\n*whether* to escalate; a cheap question-shape check (`looks_multi_hop`) decides *how*. The answer\nstep reads only the top-3 passages and returns `INSUFFICIENT_CONTEXT` when the evidence is thin."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "ANSWER_SYSTEM = (\n \"You answer a question using ONLY the numbered context passages provided. \"\n \"Reply with ONLY the final answer on a single line: a name, date, number, or short \"\n \"noun phrase, usually one to six words. Do NOT show reasoning or steps, do NOT \"\n \"restate the question. If the passages do not contain the information needed to \"\n \"answer, reply with exactly: INSUFFICIENT_CONTEXT\"\n)\n\ndef generate_answer(question, passages):\n if not passages:\n return \"INSUFFICIENT_CONTEXT\"\n ctx = \"\\n\".join(f\"[{i}] {c.title}. {(c.text or '')[:700]}\" for i, c in enumerate(passages[:ANSWER_K], 1))\n t = (ask_llm(ANSWER_SYSTEM, f\"Context:\\n{ctx}\\n\\nQuestion: {question}\\nAnswer (answer only):\",\n max_tokens=150) or \"\").strip()\n if \"INSUFFICIENT_CONTEXT\" in t.upper() or not t:\n return \"INSUFFICIENT_CONTEXT\"\n last = [ln.strip() for ln in t.splitlines() if ln.strip()][-1]\n return re.sub(r\"^(answer|the answer is|final answer)[:\\-\\s]+\", \"\", last, flags=re.I).strip()\n\ndef looks_multi_hop(question):\n # cheap, gold-free router: >= 2 named entities or a long question -> likely a missing hop\n return len(question_entities(question)) >= 2 or len(question.split()) >= 12\n\ndef solve(question):\n enc = embed(question)\n fused = hybrid_search(question, enc=enc)\n dense = dense_ranking(question, enc=enc)\n if not retrieval_is_weak(fused, dense):\n return to_passages(fused), \"cheap path: confident, answer from the hybrid top-3\"\n if looks_multi_hop(question):\n pool, sub_queries = decompose(question, hop0=to_passages(fused))\n return pool, f\"decompose: weak + multi-hop ({len(sub_queries)} sub-question(s))\"\n return to_passages(colbert_rerank(question, enc=enc)), \"ColBERT: weak single-hop, rerank\""
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "Three queries, three paths from one agent."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "routing_demos = [\n (\"confident single-hop\", \"2hop__101521_42157__h0\"),\n (\"weak single-hop\", \"2hop__130545_45439__h0\"),\n (\"multi-hop\", \"2hop__615262_131886\"),\n]\nfor label, qid in routing_demos:\n q = by_id[qid]\n pool, route = solve(q[\"question\"])\n answer = generate_answer(q[\"question\"], pool)\n found = len({p.doc_id for p in pool[:3]} & set(q[\"gold_doc_ids\"]))\n print(f\"[{label}] {q['question']}\")\n print(f\" route: {route}\")\n print(f\" answer: {answer[:72]}\")\n print(f\" gold in answer context: {found}/{len(q['gold_doc_ids'])}\\n\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### STOP: deciding whether to answer at all\n\nRouting asks \"should I spend more retrieval?\" STOP asks \"should I answer at all?\" The default\n**gentle stop** is built into `generate_answer`: it answers from the context or returns\n`INSUFFICIENT_CONTEXT`, with no extra model call. The stricter option is a separate **sufficiency\njudge** (a fast, cheap model) that reads the same top-3 context and decides whether every needed\nfact is present. It catches more unanswerables but can over-refuse answerable questions."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "SUFFICIENCY_SYSTEM = (\n \"Judge whether the provided context contains enough information to answer the question. \"\n \"Use ONLY the context. Reply with exactly one word: SUFFICIENT or INSUFFICIENT.\"\n)\n\ndef sufficiency_judge(question, passages):\n if not passages:\n return False\n context = \"\\n\".join(f\"[{i}] {c.title}. {(c.text or '')[:600]}\" for i, c in enumerate(passages, 1))\n user = (f\"Question:\\n{question}\\n\\nContext:\\n{context}\\n\\n\"\n \"Can the question be answered completely from this context? \"\n \"Reply with exactly one word: SUFFICIENT or INSUFFICIENT.\")\n try:\n t = ask_llm(SUFFICIENCY_SYSTEM, user, max_tokens=20, model=FAST_MODEL)\n except Exception:\n return True # on API failure, keep the gentle-stop behavior\n words = re.findall(r\"[A-Z]+\", (t or \"\").upper())\n return \"INSUFFICIENT\" not in words"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "One unanswerable question, traced end to end: the first retrieval, the gate, the decompose\nsub-questions, the final answer context, and both STOP decisions."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "unanswerable = by_id[\"2hop__108098_170204\"]\nquestion = unanswerable[\"question\"]\nprint(f\"Q: {question}\\n\")\n\nenc = embed(question)\nfused = hybrid_search(question, enc=enc)\ndense = dense_ranking(question, enc=enc)\nis_weak = retrieval_is_weak(fused, dense)\n\nprint(\"baseline hybrid retrieve:\")\nshow_hits(fused, [], k=ANSWER_K)\nprint(f\"\\ngate: dense_variance={dense_variance(dense):.5f} (floor {DV_FLOOR:.5f}), \"\n f\"score_variance={score_variance(fused):.5f} (floor {SV_FLOOR:.5f})\")\nprint(f\"decision: {'weak -> escalate' if is_weak else 'strong -> answer now'}\")\n\npool, route = solve(question)\nprint(f\"route: {route}\")\ngentle = generate_answer(question, pool)\nsufficient = sufficiency_judge(question, pool[:ANSWER_K])\nprint(f\"\\ngentle stop = {'abstain' if gentle == 'INSUFFICIENT_CONTEXT' else 'answer'} ({gentle})\")\nprint(f\"sufficiency judge = {'answer' if sufficient else 'abstain'}\")"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "## Results on the full workload\n\nThe numbers below are **reported from a run on the full workload** (321 held-out questions:\nsingle-hop, multi-hop, and unanswerable), not recomputed on the slice above. The slice is a\nrunnable demo; these are what the method buys at scale.\n\nThe answer context is top-3, so the retrieval metrics focus on that small window:\n\n| metric | what it answers |\n|---|---|\n| recall@3 | did at least one supporting passage reach the answer context? |\n| full_gold@3 | did every required passage reach the answer context? |\n| MRR | how high did the first supporting passage rank? |\n| LLM calls/query | how many LLM calls did the route spend? |\n| latency | routing time before the final answer, end to end |"
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "headline = json.loads(Path(\"headline_final_v25.json\").read_text())[\"overall\"]\nrows = []\nfor name in (\"always_answer\", \"always_colbert\", \"always_decompose\", \"ladder\"):\n m = headline[name]\n rows.append({\n \"policy\": name.replace(\"_\", \"-\"),\n \"recall@3\": m[\"recall@3\"], \"full_gold@3\": m[\"full_gold@3\"], \"MRR\": m[\"mrr_first\"],\n \"LLM calls/query\": m[\"llm_calls\"], \"routing latency (s)\": m[\"avg_latency_s\"],\n })\npd.set_option(\"display.precision\", 3)\npd.DataFrame(rows)"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "The adaptive loop (`ladder`) keeps most of the quality of always-decompose at a fraction of the\nLLM calls and latency: it only pays for decomposition on the queries the gate flags. Always-answer\nis cheapest but leaves multi-hop quality on the table; always-decompose is the quality ceiling but\nspends an LLM call on every query, including the ones that didn't need it.\n\n### STOP, side by side\n\n\"Full workload handled\" means the agent either answers correctly or correctly refuses an\nunanswerable question."
|
||||
},
|
||||
{
|
||||
"cell_type": "code",
|
||||
"metadata": {},
|
||||
"execution_count": null,
|
||||
"outputs": [],
|
||||
"source": "stop = json.loads(Path(\"targeted_stop_v25.json\").read_text())[\"variants\"]\nrows = [\n (\"hybrid baseline + gentle\", \"baseline_hybrid_gentle\"),\n (\"loop + gentle (default)\", \"ladder_gentle\"),\n (\"loop + LLM sufficiency check\", \"ladder_autorater_all\"),\n]\npd.DataFrame([{\n \"setup\": label,\n \"catches unanswerable\": stop[key][\"abstain_unans\"],\n \"over-refuses answerable\": stop[key][\"false_stop_ans\"],\n \"full workload handled\": stop[key][\"selective_accuracy\"],\n} for label, key in rows])"
|
||||
},
|
||||
{
|
||||
"cell_type": "markdown",
|
||||
"metadata": {},
|
||||
"source": "### Adapt this to your corpus\n\n1. **Set the answer context.** Decide how many passages the LLM reads. That defines \"good retrieval.\"\n2. **Label retrieval outcomes** on a calibration split: did the needed evidence reach that window?\n3. **Engineer and benchmark cheap signals.** Keep the ones that separate weak from healthy retrievals on your data.\n4. **Match a fix to each failure mode.** Add the cheapest corrective action that addresses it.\n5. **Measure quality and cost together.** Keep the loop only where the quality justifies the spend.\n6. **Choose STOP deliberately.** Gentle by default; stricter when wrong answers are expensive.\n\nA single signal plus one floor was enough here because the kept signals were redundant. When your\nsignals carry independent information, fuse them with a small classifier instead of one threshold.\nThe method is the same: define good retrieval, find signals that predict it on your data, and route\nto the cheapest sufficient fix."
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
"kernelspec": {
|
||||
"display_name": "Python 3",
|
||||
"language": "python",
|
||||
"name": "python3"
|
||||
},
|
||||
"language_info": {
|
||||
"name": "python",
|
||||
"version": "3.12"
|
||||
}
|
||||
},
|
||||
"nbformat": 4,
|
||||
"nbformat_minor": 5
|
||||
}
|
||||
@@ -0,0 +1,36 @@
|
||||
{
|
||||
"n": 321,
|
||||
"n_answerable": 180,
|
||||
"n_unanswerable": 141,
|
||||
"tier_dist": {
|
||||
"1": 98,
|
||||
"2": 70,
|
||||
"3": 153
|
||||
},
|
||||
"variants": {
|
||||
"baseline_hybrid_gentle": {
|
||||
"selective_accuracy": 0.5296,
|
||||
"answerable_em": 0.5167,
|
||||
"abstain_unans": 0.5461,
|
||||
"false_stop_ans": 0.2111
|
||||
},
|
||||
"ladder_gentle": {
|
||||
"selective_accuracy": 0.4984,
|
||||
"answerable_em": 0.5667,
|
||||
"abstain_unans": 0.4113,
|
||||
"false_stop_ans": 0.15
|
||||
},
|
||||
"ladder_autorater_all": {
|
||||
"selective_accuracy": 0.6293,
|
||||
"answerable_em": 0.4444,
|
||||
"abstain_unans": 0.8652,
|
||||
"false_stop_ans": 0.4111
|
||||
},
|
||||
"ladder_targeted": {
|
||||
"selective_accuracy": 0.6168,
|
||||
"answerable_em": 0.4667,
|
||||
"abstain_unans": 0.8085,
|
||||
"false_stop_ans": 0.3667
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,203 @@
|
||||
---
|
||||
title: "Route Retrieval by Evidence: Cheap Signals for Self-Correcting RAG"
|
||||
short_description: "Cheap signals that flag weak retrieval, so you escalate only when needed."
|
||||
description: "Route agentic retrieval by evidence, not question shape: cheap signals predict weak retrieval and a gate escalates to ColBERT or decomposition only when needed."
|
||||
preview_dir: /articles_data/route-retrieval-by-evidence/preview
|
||||
social_preview_image: /articles_data/route-retrieval-by-evidence/preview/social_preview.jpg
|
||||
weight: -200
|
||||
author: Dylan Couzon
|
||||
date: 2026-06-17T00:00:00+03:00
|
||||
draft: false
|
||||
keywords:
|
||||
- signal-driven retrieval
|
||||
- agentic RAG
|
||||
- self-correcting retrieval
|
||||
- query routing
|
||||
- hybrid search
|
||||
category: rag-and-genai
|
||||
---
|
||||
|
||||
Most retrieval systems do the same thing on every query, whether they feed a RAG prompt, an agent loop, or a search results page. They run one [hybrid search](/documentation/search/hybrid-queries/) and use the top results, or they rerank everything, or they rewrite and re-retrieve on every query. Each is the wrong default for some fraction of traffic: the single pass under-serves the hard queries, and the expensive path wastes compute and latency on the easy ones. Worse, a single pass fails *silently*. When the relevant document never reaches the top results, the system still returns something, ranked from whatever it retrieved.
|
||||
|
||||
A self-correcting loop fixes the default by making it conditional. Retrieve once, judge the result with a **cheap signal**, and escalate to a more expensive action only when the signal says the evidence is weak. The interesting question is not which expensive action to run. [Cross-encoders](/documentation/fastembed/fastembed-rerankers/), [ColBERT late interaction](/articles/late-interaction-models/), query rewriting, and [decomposition](/documentation/search/hybrid-queries/#query-decomposition-for-multi-hop-questions) are all well understood. The interesting question is *when* to spend them, and how to decide cheaply, on every query, without an extra model call.
|
||||
|
||||
This article is about that decision, which is the same whatever the corpus: the signals that predict weak retrieval, how to find the ones that work on your data, and how to turn a signal into a routing gate. To put real numbers on it we lean on one concrete benchmark, a [workshop we built](https://github.com/qdrant-labs/self-correcting-loops-workshop) on a mixed question-answering workload, where routing by signal kept most of the quality of always-escalating at a fraction of the cost. The benchmark is question answering; the method is not.
|
||||
|
||||
{{< figure src="/articles_data/route-retrieval-by-evidence/self-correcting-loop.png" alt="The self-correcting retrieval loop: a query goes through hybrid retrieval to a cheap signal and gate, which either answers from the top-3 or escalates to ColBERT or decomposition before answering." caption="The loop. One retrieval, a cheap signal that decides whether the evidence is weak, and a more expensive action only when it is." width="100%" >}}
|
||||
|
||||
This sits in a line of work on corrective and adaptive retrieval, and the difference is specific. [CRAG](https://arxiv.org/abs/2401.15884) also grades the retrieval and corrects when it looks weak, but it trains a dedicated evaluator model to do the grading. [Adaptive-RAG](https://arxiv.org/abs/2403.14403) routes on a query-complexity classifier, deciding from the shape of the question before retrieving anything, which is the approach this article argues against. [Self-RAG](https://arxiv.org/abs/2310.11511) bakes the decision into a fine-tuned model, and [FLARE](https://arxiv.org/abs/2305.06983) gates on the generator's token confidence. The gate here reads a cheap statistic of the result already retrieved: no extra model call, no model to train, no query classifier, just a labeled calibration split to choose the signals and set the floors. Whether the retrieved evidence is enough is the same question [sufficient-context work](https://arxiv.org/abs/2411.06037) asks with an LLM judge, answered here from a free signal instead.
|
||||
|
||||
## What "good retrieval" means
|
||||
|
||||
Before you can detect weak retrieval, you have to define it, and the definition has to be operational. Every retrieval system has a window that gets used: the few passages an LLM reads, the first page a user scans, the candidate pool a reranker reorders. Good retrieval means the items you need land inside that window. Everything below it is invisible, no matter how good the recall looks deeper down the list.
|
||||
|
||||
In the benchmark the window is the **top-3** passages an LLM reads to answer, so a retrieval is *good* when every supporting passage the question needs lands in the top-3, and *weak* when at least one is missing. Pick the window that matches your system and the rest of the method is unchanged. This is the hook that makes a signal meaningful: a signal is useful precisely when it predicts that the needed items did not make the window.
|
||||
|
||||
The benchmark workload is [MuSiQue](https://github.com/StonyBrookNLP/musique), built from three failure modes that show up well beyond question answering:
|
||||
|
||||
| Failure mode | In the benchmark | Where else it appears |
|
||||
|---|---|---|
|
||||
| easy | single-hop: one passage answers | the bulk of queries in a typical workload |
|
||||
| needs a follow-up | multi-hop: the next passage isn't reachable from the query as written | faceted product search, queries whose filter depends on a first lookup |
|
||||
| not in the corpus | unanswerable: the evidence isn't there | anything that should return "no good match" instead of the least-bad result |
|
||||
|
||||
The baseline is a single **hybrid** retrieval: a [dense retriever](/documentation/search/search/) (`bge-base`) and a sparse one ([miniCOIL](/documentation/fastembed/fastembed-minicoil/)) fused server-side with [Reciprocal Rank Fusion (RRF)](/documentation/search/hybrid-queries/#reciprocal-rank-fusion-rrf). It handles the easy cases. It breaks when the needed evidence is not reachable from the query as written, and it has no way to know it failed.
|
||||
|
||||
## Signals: predicting weak retrieval for free
|
||||
|
||||
A **signal** is a number you compute from the retrieval result that predicts whether the evidence is weak. A good signal meets two criteria, and the second is what makes this practical:
|
||||
|
||||
1. It separates good retrievals from weak ones on calibration data.
|
||||
2. It is cheap enough to run on every query.
|
||||
|
||||
"Cheap" is the constraint that rules out the obvious approaches. You could ask an LLM "is this enough to answer?" on every query, but that adds a model call and latency to the cheap path you were trying to protect. The signals here read only what the retriever already returned, or at most ask Qdrant for one extra ranking. No extra model call, no LLM in the hot path.
|
||||
|
||||
### The candidate families
|
||||
|
||||
There is no single signal that works everywhere, so the right move is to test several and keep what separates on your data. We test five candidates across five families. They are a starter set, not the universe: predicting whether a retrieval succeeded is a studied problem, [query performance prediction](https://arxiv.org/abs/2504.01101), with a couple of dozen predictors in the literature. The two spread signals here are score-variance predictors of the kind that work calls Unnormalized Query Commitment (UQC), the unnormalized counterpart of the better-known Normalized Query Commitment (NQC), which divides the same spread by a collection-wide score. The split that matters is whether a signal reads the already-fused hybrid result or asks for one extra raw retriever view.
|
||||
|
||||
| Family | Signal | What it reads | Flags weak when |
|
||||
|---|---|---|---|
|
||||
| height | `max_score` | top-1 fused score | low |
|
||||
| spread (fused) | `score_variance` | spread of the fused top-k scores | low (flat ranking) |
|
||||
| spread (raw) | `dense_variance` | spread of the raw dense cosines | low (dense can't separate its hits) |
|
||||
| coverage | `evidence_coverage` | question entities present in the top-k text | low (text misses them) |
|
||||
| agreement | `retriever_divergence` | dense vs miniCOIL top-k overlap | high (the retrievers disagree) |
|
||||
|
||||
The intuition behind each is worth stating, because it is what transfers to a new corpus:
|
||||
|
||||
- **Height** asks whether the top hit scored high in absolute terms. It is most useful when scores are calibrated; RRF ranks rather than calibrates, so on a fused result the top score tends to echo the spread signals.
|
||||
- **Spread** asks whether the retriever could pull its top results apart. When a retriever is confident, the scores fan out; when it is lost, they bunch together. We measure spread on two substrates: the fused RRF scores (free) and the raw dense cosines (one extra query, no model call). RRF combines ranks well but flattens score shape, which is why the raw dense spread can carry information the fused spread loses.
|
||||
- **Coverage** asks whether the entities named in the question appear in the retrieved text. Cheap to compute, strong on entity-lookup data.
|
||||
- **Agreement** asks whether the dense and sparse retrievers returned the same documents. When they diverge sharply, one of them is usually lost, which happens on jargon or out-of-vocabulary terms.
|
||||
|
||||
### How the signals are computed
|
||||
|
||||
Each signal is a few lines. The two spread signals take the population standard deviation of a score list (the `_variance` names are loose; the code returns the standard deviation):
|
||||
|
||||
```python
|
||||
import statistics
|
||||
|
||||
def score_variance(fused_points): # spread of the fused RRF scores
|
||||
return statistics.pstdev([p.score for p in fused_points])
|
||||
|
||||
def dense_variance(dense_ranking): # spread of the raw dense cosines
|
||||
return statistics.pstdev([score for _, score in dense_ranking])
|
||||
```
|
||||
|
||||
`score_variance` reads the hybrid result you already have. `dense_variance` needs the raw dense ranking, which is one extra Qdrant query that reuses the embedding the hybrid query already computed, so there is no extra model call:
|
||||
|
||||
```python
|
||||
def dense_ranking(query_vector):
|
||||
points = client.query_points(
|
||||
collection_name="{collection_name}",
|
||||
query=query_vector,
|
||||
using="dense",
|
||||
limit=10,
|
||||
with_payload=False,
|
||||
).points
|
||||
return [(p.id, p.score) for p in points]
|
||||
```
|
||||
|
||||
Coverage and agreement are equally cheap: coverage counts how many of the question's capitalized spans and years appear in the top-k text, and agreement is one minus the overlap between the dense and sparse top-k id sets.
|
||||
|
||||
## Finding the winners on your data
|
||||
|
||||
This is the part that matters, and the part that does not transfer as a fixed answer. The signals that win are corpus-specific. The *method* for finding them is not.
|
||||
|
||||
On a calibration split you know the ground truth: whether all the gold passages landed in the top-3. So you can score each candidate by how well its value separates good retrievals from weak ones, using the area under the ROC curve (AUC). Read it as a separation score, where 0.5 is a coin flip and 1.0 is perfect. Because some signals fire high for weak retrieval and others fire low, we take the stronger of the signal and its inverse:
|
||||
|
||||
```python
|
||||
from sklearn.metrics import roc_auc_score
|
||||
|
||||
def separation(values, weak_labels):
|
||||
auc = roc_auc_score(weak_labels, values)
|
||||
return max(auc, 1 - auc) # separation, regardless of direction
|
||||
```
|
||||
|
||||
Then two rules: keep any signal that separates above a bar (0.65 here), and drop any survivor that is a near-duplicate of a stronger one (absolute correlation above 0.85). The correlation rule matters because redundant signals add cost without adding information.
|
||||
|
||||
{{< figure src="/articles_data/route-retrieval-by-evidence/signal-separation.png" alt="Box plots of each candidate signal on good versus weak retrievals. dense_variance and score_variance separate the two; the others overlap." caption="Each candidate on good versus weak retrievals over the calibration split. When the boxes pull apart, the signal separates; when they overlap, it does not." width="100%" >}}
|
||||
|
||||
On this corpus, two signals clear the bar and survive the correlation check:
|
||||
|
||||
| Signal | AUC | Verdict |
|
||||
|---|---|---|
|
||||
| `dense_variance` | 0.74 | kept |
|
||||
| `score_variance` | 0.70 | kept |
|
||||
| `max_score` | 0.66 | dropped (redundant) |
|
||||
| `evidence_coverage` | 0.59 | dropped (below the bar) |
|
||||
| `retriever_divergence` | 0.53 | dropped (below the bar) |
|
||||
|
||||
Both winners are spread signals: when the retriever cannot separate its top results, the evidence is usually weak. `max_score` clears the separation bar too, but it correlated 0.92 with `score_variance` on this split, so the correlation rule drops it as a near-duplicate. `evidence_coverage` fails for a concrete reason worth seeing: the question's entities turn up somewhere in the top-k for most retrievals, weak or not (87% of good and 69% of weak score a perfect 1.0), so the signal saturates and a threshold has nothing to cut on. The AUC, not any single threshold, is what tells you that. Agreement is weak here too, though on a jargon-heavy corpus it can be the strongest signal you have, and on entity-lookup data coverage often wins. Run the benchmark on your data and keep what separates: pick signals on one split, tune thresholds on a second, and keep a test split held out.
|
||||
|
||||
## The gate: a floor on each signal
|
||||
|
||||
A signal becomes a decision through a threshold. Each kept signal gets a floor; if either falls below it, the retrieval is treated as weak and the query escalates:
|
||||
|
||||
```python
|
||||
def retrieval_is_weak(fused_points, dense_ranking):
|
||||
return (score_variance(fused_points) < SV_FLOOR
|
||||
or dense_variance(dense_ranking) < DV_FLOOR)
|
||||
```
|
||||
|
||||
The floor is a tradeoff, not a constant. Raise it and you catch more weak retrievals but escalate more queries; lower it and you escalate less but miss more. We set it at the threshold that maximizes catches minus false alarms on the calibration set (the Youden point on the ROC curve), though you can instead target a recall (say, catch 90% of weak retrievals) when a missed weak retrieval costs more than an extra escalation.
|
||||
|
||||
{{< figure src="/articles_data/route-retrieval-by-evidence/gate-histogram.png" alt="Histogram of dense_variance for good and weak retrievals, with the gate floor marked. Most weak retrievals fall below the floor." caption="The chosen signal's distribution for good and weak retrievals, with the calibrated floor. The shaded region escalates." width="85%" >}}
|
||||
|
||||
The two kept signals differ in cost as well as strength. `score_variance` reads the fused result you already have, so it is free: at its balanced floor it catches 63% of the weak retrievals while escalating 37% of queries. `dense_variance` is stronger, catching 78% while escalating 49%, but it spends one extra dense query on every request, a retrieval round-trip added to the cheap path you are protecting. The free signal alone already catches most weak retrievals; the extra query buys about 15 more points of catch. That is the shape of the win either way: most weak cases caught, and the cheap path still serves the rest.
|
||||
|
||||
The gate ORs the two floors because the signals are complementary, not redundant. On the calibration split `dense_variance` and `score_variance` correlate only 0.54 and flag overlapping but different weak retrievals, with about half the escalations unique to one signal. Together they catch 80% against 78% for `dense_variance` alone, at a slightly higher escalation rate. Whether two signals are redundant is a property of the corpus, not a fixed fact, which is why the correlation check belongs in the benchmark rather than a rule of thumb. With more signals that each carry independent information, hand-tuning a floor per signal stops scaling, and a small classifier over the signal vector is the better gate.
|
||||
|
||||
## What routing by evidence buys
|
||||
|
||||
The gate decides whether to escalate; what you escalate *to* is open. A reranker, a query rewrite, a larger embedding model, decomposition, relaxing a filter, or handing off to a person are all valid escalations, and the gate does not care which you pick. That division is the point of routing by evidence: the signal makes the decision that costs you, whether to escalate at all; choosing which escalation to run is a cheaper, lower-stakes call, and a query-shape heuristic is fine there. The benchmark picks between two with a light, label-free check on the query (its entity count and length): a likely multi-hop query goes to **decomposition**, which [asks the next still-missing sub-question](/documentation/search/hybrid-queries/#query-decomposition-for-multi-hop-questions), retrieves again, and repeats, interleaving retrieval with reasoning ([Self-Ask](https://arxiv.org/abs/2210.03350), [IRCoT](https://arxiv.org/abs/2212.10509)); anything else goes to **ColBERT** late interaction, which re-scores with token-level precision to promote a buried result and adds no LLM call. That split is why the gate escalates about 70% of the test queries but the loop still averages only 0.78 LLM calls: only the roughly 48% routed to decomposition spend a call, and the ColBERT reranks add none. Both actions are linked rather than re-explained, because the point of the loop is the routing, not the actions.
|
||||
|
||||
Because the gate routes by evidence, each question type lands where it needs to without anyone hard-coding a rule per type:
|
||||
|
||||
{{< figure src="/articles_data/route-retrieval-by-evidence/routing-distribution.png" alt="Stacked bars showing how single-hop, multi-hop, and unanswerable questions are routed across the three tiers." caption="How the gate routes each question type. Single-hop questions mostly stay cheap; multi-hop questions mostly decompose; unanswerable questions spread across tiers as the signal finds them weak." width="85%" >}}
|
||||
|
||||
Unanswerable queries show both the reach of routing by evidence and its limit: the gate flags most of them as weak, but escalation cannot conjure evidence that is not in the corpus, so the win there is detecting the weak retrieval, not repairing it. Deciding when to abstain on those instead of answering from thin evidence is a separate step, the STOP autorater described in the [workshop](https://github.com/qdrant-labs/self-correcting-loops-workshop); this article measures retrieval quality on the answerable set and makes no abstention claim.
|
||||
|
||||
The loop sits on the cost/quality frontier. Quality is measured over the 180 answerable questions, since the 141 unanswerable ones have no gold passage to recover, while cost and routing cover all 321. The lead metrics are top-3 metrics, because the top-3 is the window the model reads: `recall@3` asks whether at least one supporting passage reached it, and `full_gold@3` asks whether every required passage did, a strict set recall at the window.
|
||||
|
||||
| Policy | recall@3 | full_gold@3 | MRR | LLM calls/query | Routing latency |
|
||||
|---|---|---|---|---|---|
|
||||
| always answer | 0.82 | 0.70 | 0.88 | 0.0 | 0.15s |
|
||||
| always rerank | 0.81 | 0.69 | 0.91 | 0.0 | 0.15s |
|
||||
| always ColBERT | 0.80 | 0.68 | 0.89 | 0.0 | 0.30s |
|
||||
| always decompose | 0.88 | 0.79 | 0.90 | 1.88 | 3.30s |
|
||||
| **signal-gated ladder** | **0.85** | **0.76** | **0.91** | **0.78** | **1.50s** |
|
||||
|
||||
{{< figure src="/articles_data/route-retrieval-by-evidence/cost-quality-frontier.png" alt="Cost versus quality across the five retrieval policies. The signal-gated ladder sits on the efficient frontier, near always-decompose but at far lower cost." caption="The same five policies as a frontier: quality (all gold passages in the top-3) against cost (LLM calls per query). The signal-gated ladder comes within a few points of always-decompose at roughly 40% of its LLM calls." width="90%" >}}
|
||||
|
||||
Two honest readings of this table. First, the win is efficiency, not a new quality ceiling. Always-decompose is the most accurate policy, and the ladder does not beat it: it reaches `recall@3` of 0.85 against 0.88, and `full_gold@3` of 0.76 against 0.79. What it changes is the bill. It spends 0.78 LLM calls per query against 1.88, and 1.5 seconds of routing latency against 3.3, because it only decomposes the queries the gate flags. Against the cheap always-answer baseline, the lift is real and significant: `recall@3` improves by 3.6 points (95% CI 0.3 to 6.8) and `full_gold@3` by 6.1 points (95% CI 1.1 to 11.1).
|
||||
|
||||
Second, the ladder posts the best MRR (mean reciprocal rank) at 0.91, just ahead of always-rerank at 0.905 and always-decompose at 0.901. The gap is small and carries no confidence interval, and MRR is the lenient metric here, rewarding any single supporting passage, so read it as a sign that routing the right queries to ColBERT holds rank quality steady, not as a standalone win.
|
||||
|
||||
## This is not only for question answering
|
||||
|
||||
The benchmark is a QA workload, but nothing in the method is: the signals are statistics of a ranking and the gate is a threshold on them, so the pattern fits any retrieval system with a cheap default and a more expensive fallback. The cases below are not measured here. They are extrapolations from the mechanism, and they are worth stating only as hypotheses to run on your own data, not as results:
|
||||
|
||||
- **Product and site search.** A query that matches nothing relevant produces a flat, low-spread result. The same spread signal that flags weak QA retrieval can trigger query relaxation or a broader recall pass instead of returning the least-bad items.
|
||||
- **Support and knowledge-base RAG.** When the first retrieval over a help center is weak, escalate to a query rewrite or ask the user to disambiguate, rather than answering from thin context.
|
||||
- **Code and technical search.** Identifier- and jargon-heavy queries are exactly where dense and lexical retrievers disagree, so the agreement signal earns its keep and routes those queries to a lexical-heavy pass.
|
||||
- **Heterogeneous or multi-tenant corpora.** Some queries land in a region of the corpus the default index serves poorly; a weak-retrieval signal can route them to a different index or filter.
|
||||
|
||||
In each case the three ingredients are the same: a window you can define, cheap signals that predict when the needed items missed it, and a more expensive action worth spending only when they did.
|
||||
|
||||
## Adapt this to your corpus
|
||||
|
||||
The specific winners here are not portable, but the recipe is:
|
||||
|
||||
1. **Set the window.** Decide what your system consumes (the passages an LLM reads, the results a user sees) and how many. That window defines good retrieval.
|
||||
2. **Label retrieval outcomes** on a calibration split: did the needed evidence reach the window?
|
||||
3. **Engineer and benchmark cheap signals.** Keep the ones that separate weak from healthy retrievals on your data; drop the redundant ones.
|
||||
4. **Turn the winner into a gate.** One floor if your signals are redundant, a small classifier if they carry independent information.
|
||||
5. **Match a fix to each failure mode** and measure quality and cost together. Keep the loop only where the quality justifies the spend.
|
||||
|
||||
A caveat worth stating plainly: every number here is from one workload. The mixed single-hop, multi-hop, and unanswerable set is representative of the failure modes that make routing worthwhile, but it is a single corpus, and the AUCs, the floors, and the winning signals would differ on yours. The benchmark code is open-source, so the honest way to use this is to run it on your data rather than to adopt these thresholds.
|
||||
|
||||
The reusable idea is small and sturdy. Retrieval quality is observable from cheap statistics of the result, before you spend anything to act on it. Measure it, set a floor, and let the evidence, not the question, decide which queries are worth the expensive path.
|
||||
|
||||
*The full method, including the corrective actions and the evaluation harness, is in the [self-correcting retrieval loops workshop](https://github.com/qdrant-labs/self-correcting-loops-workshop). For the building blocks the loop escalates to, see [late interaction models](/articles/late-interaction-models/), [hybrid search](/articles/hybrid-search/), and [query decomposition](/documentation/search/hybrid-queries/#query-decomposition-for-multi-hop-questions).*
|
||||
BIN
Binary file not shown.
|
After Width: | Height: | Size: 61 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 50 KiB |
BIN
Binary file not shown.
|
After Width: | Height: | Size: 54 KiB |
BIN
Binary file not shown.
|
After Width: | Height: | Size: 54 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 69 KiB |
Reference in New Issue
Block a user