mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-30 00:18:32 +02:00
Adding more details to the instructions
This commit is contained in:
+5
-1
@@ -43,7 +43,11 @@ You are helping build an evaluation dataset for a search system.
|
||||
Generate 3 realistic search queries for the document below.
|
||||
Each query should be what a real user would type to find it.
|
||||
Phrase queries naturally, not as paraphrases of the document.
|
||||
Return only the queries, one per line. No numbering or explanation.
|
||||
|
||||
Return exactly 3 lines, one query per line. No numbering, no bullets, no preamble. Example:
|
||||
how does X work
|
||||
best way to configure Y
|
||||
what is Z used for
|
||||
|
||||
Document:
|
||||
{document_text}
|
||||
|
||||
+4
@@ -23,6 +23,8 @@ For orientation on the four layers of retrieval evaluation and where this tutori
|
||||
|
||||
**1. Prepare the evaluation data.** Each entry needs a `query_id`, a `query_text` (for prompting the generator), a `query_vector` (for retrieval), and `labels`. For `context_precision` only, also include a `ground_truth` reference answer.
|
||||
|
||||
If your queries came from synthetic generation, they don't carry ground-truth answers natively. A simple workaround: make one more LLM pass per query, constrained to the source document, asking for a one-to-two-sentence reference answer. Skip this step if you're only scoring `faithfulness` and `answer_relevancy` (both are reference-free).
|
||||
|
||||
```python
|
||||
# Example of an evaluation-ready entry.
|
||||
{
|
||||
@@ -53,6 +55,8 @@ Question:
|
||||
"""
|
||||
```
|
||||
|
||||
The prompt above is a starting point; tune it for your domain: answer style, refusal behavior, whether outside knowledge is allowed, and output format.
|
||||
|
||||
**3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample`. `SingleTurnSample` is Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference. The example uses Anthropic, but any LLM provider works (OpenAI, Cohere, a local model). Only the `generate_answer` body changes:
|
||||
|
||||
```python
|
||||
|
||||
+1
-1
@@ -55,7 +55,7 @@ Tune until you hit the point that matches your quality and cost targets.
|
||||
|
||||
The Web UI is the fastest way to check precision interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute precision@k.
|
||||
|
||||
This helper takes a list of query vectors and returns the average precision@k. Use a representative sample of query vectors from your workload as your test set.
|
||||
This helper takes a list of query vectors and returns the average precision@k. Use a representative sample of query vectors from your workload (typically 20–50, embedded with the same model your collection uses) as your test set.
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
Reference in New Issue
Block a user