Adding more details to the instructions

This commit is contained in:
Dylan Couzon
2026-04-24 00:38:27 -04:00
parent d14bbd3fc1
commit 3e378a3647
3 changed files with 10 additions and 2 deletions
@@ -43,7 +43,11 @@ You are helping build an evaluation dataset for a search system.
Generate 3 realistic search queries for the document below. Generate 3 realistic search queries for the document below.
Each query should be what a real user would type to find it. Each query should be what a real user would type to find it.
Phrase queries naturally, not as paraphrases of the document. Phrase queries naturally, not as paraphrases of the document.
Return only the queries, one per line. No numbering or explanation.
Return exactly 3 lines, one query per line. No numbering, no bullets, no preamble. Example:
how does X work
best way to configure Y
what is Z used for
Document: Document:
{document_text} {document_text}
@@ -23,6 +23,8 @@ For orientation on the four layers of retrieval evaluation and where this tutori
**1. Prepare the evaluation data.** Each entry needs a `query_id`, a `query_text` (for prompting the generator), a `query_vector` (for retrieval), and `labels`. For `context_precision` only, also include a `ground_truth` reference answer. **1. Prepare the evaluation data.** Each entry needs a `query_id`, a `query_text` (for prompting the generator), a `query_vector` (for retrieval), and `labels`. For `context_precision` only, also include a `ground_truth` reference answer.
If your queries came from synthetic generation, they don't carry ground-truth answers natively. A simple workaround: make one more LLM pass per query, constrained to the source document, asking for a one-to-two-sentence reference answer. Skip this step if you're only scoring `faithfulness` and `answer_relevancy` (both are reference-free).
```python ```python
# Example of an evaluation-ready entry. # Example of an evaluation-ready entry.
{ {
@@ -53,6 +55,8 @@ Question:
""" """
``` ```
The prompt above is a starting point; tune it for your domain: answer style, refusal behavior, whether outside knowledge is allowed, and output format.
**3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample`. `SingleTurnSample` is Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference. The example uses Anthropic, but any LLM provider works (OpenAI, Cohere, a local model). Only the `generate_answer` body changes: **3. Run retrieval and generation.** For each entry, retrieve the top-k chunks, pass them through the generator, and record a `SingleTurnSample`. `SingleTurnSample` is Ragas's data class for one evaluation record: question, retrieved context, generated answer, and optional reference. The example uses Anthropic, but any LLM provider works (OpenAI, Cohere, a local model). Only the `generate_answer` body changes:
```python ```python
@@ -55,7 +55,7 @@ Tune until you hit the point that matches your quality and cost targets.
The Web UI is the fastest way to check precision interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute precision@k. The Web UI is the fastest way to check precision interactively. For continuous integration or scripted regression tests, the Qdrant client exposes the same exact-search mode via `search_params=models.SearchParams(exact=True)`. Compare the ANN and exact top-k sets yourself and compute precision@k.
This helper takes a list of query vectors and returns the average precision@k. Use a representative sample of query vectors from your workload as your test set. This helper takes a list of query vectors and returns the average precision@k. Use a representative sample of query vectors from your workload (typically 20–50, embedded with the same model your collection uses) as your test set.
```python ```python
from qdrant_client import QdrantClient, models from qdrant_client import QdrantClient, models