diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md index 3169aeda8..6277ecca5 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-golden-set.md @@ -43,7 +43,11 @@ You are helping build an evaluation dataset for a search system. Generate 3 realistic search queries for the document below. Each query should be what a real user would type to find it. Phrase queries naturally, not as paraphrases of the document. -Return only the queries, one per line. No numbering or explanation. + +Return exactly 3 lines, one query per line. No numbering, no bullets, no preamble. Example: +how does X work +best way to configure Y +what is Z used for Document: {document_text} diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md index 952d9bb17..4e13eec44 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality-pipeline-output.md @@ -23,6 +23,8 @@ For orientation on the four layers of retrieval evaluation and where this tutori **1. Prepare the evaluation data.** Each entry needs a `query_id`, a `query_text` (for prompting the generator), a `query_vector` (for retrieval), and `labels`. For `context_precision` only, also include a `ground_truth` reference answer. +If your queries came from synthetic generation, they don't carry ground-truth answers natively. A simple workaround: make one more LLM pass per query, constrained to the source document, asking for a one-to-two-sentence reference answer. Skip this step if you're only scoring `faithfulness` and `answer_relevancy` (both are reference-free). + ```python # Example of an evaluation-ready entry. { @@ -53,6 +55,8 @@ Question: """ ``` +The prompt above is a starting point; tune it for your domain: answer style, refusal behavior, whether outside knowledge is allowed, and output format. + **3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample`. `SingleTurnSample` is Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference. The example uses Anthropic, but any LLM provider works (OpenAI, Cohere, a local model). Only the `generate_answer` body changes: ```python diff --git a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md index 337ae537a..dd2adff77 100644 --- a/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md +++ b/qdrant-landing/content/documentation/tutorials-search-engineering/retrieval-quality.md @@ -55,7 +55,7 @@ Tune until you hit the point that matches your quality and cost targets. The Web UI is the fastest way to check precision interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute precision@k. -This helper takes a list of query vectors and returns the average precision@k. Use a representative sample of query vectors from your workload as your test set. +This helper takes a list of query vectors and returns the average precision@k. Use a representative sample of query vectors from your workload (typically 20–50, embedded with the same model your collection uses) as your test set. ```python from qdrant_client import QdrantClient, models