rename dense_summary to dense_abstract and explain chunking choice

The arxiv data has abstracts, not summaries. Renaming the named
vector and prose throughout removes the ambiguity flagged on the PR.
Adds a short paragraph to the Dataset section explaining that
abstracts fit any embedding model's context window, so chunking is
included to mirror the pipeline shape you'd use on full bodies.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Dylan Couzon
2026-05-11 15:03:52 -04:00
co-authored by Claude Opus 4.7
parent 972568940d
commit 433cae62a5
@@ -5,7 +5,7 @@ aliases:
- /documentation/tutorials/multi-representation-search/
---
# Multi-Representation Search Across Titles, Summaries, and Chunks
# Multi-Representation Search Across Titles, Abstracts, and Chunks
| Time: 45 min | Level: Intermediate | Output: [GitHub](https://github.com/qdrant/examples/blob/master/multi-representation-search/multi-representation-search.ipynb) | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://githubtocolab.com/qdrant/examples/blob/master/multi-representation-search/multi-representation-search.ipynb) |
| --- | ----------- | ----------- | ----------- |
@@ -30,7 +30,9 @@ This tutorial uses <a href="/documentation/inference/#qdrant-cloud-inference">Qd
## Dataset
You'll work with 20 000 arXiv papers from the [`gfissore/arxiv-abstracts-2021`](https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021) Hugging Face dataset, filtered to ML/CS categories and to papers from 2018 onward, since earlier ML papers predate most of the topics queries care about. Each paper has a title, an abstract, and category tags, which gives you four natural representations once the abstract is split into chunks: title, full abstract as a summary, abstract sentences as chunks, and categories as tags.
You'll work with 20 000 arXiv papers from the [`gfissore/arxiv-abstracts-2021`](https://huggingface.co/datasets/gfissore/arxiv-abstracts-2021) Hugging Face dataset, filtered to ML/CS categories and to papers from 2018 onward, since earlier ML papers predate most of the topics queries care about. Each paper has a title, an abstract, and category tags, which gives you four natural representations once the abstract is split into chunks: title, full abstract, abstract sentences as chunks, and categories as tags.
arXiv abstracts are short enough to fit any dense embedding model's context window, so chunking isn't strictly required for this dataset. We chunk here because the same pipeline shape (chunk-level retrieval, document-level grouping) is what you'd use on full paper bodies in production, where context limits force the issue. The abstract stands in for what would be a longer body field in your own corpus.
```python
from datasets import load_dataset
@@ -68,7 +70,7 @@ for i in range(len(dataset) - 1, -1, -1):
## Collection Schema
Design the collection before writing any queries. The point granularity is the chunk: every chunk of every paper becomes one point. Title and summary embeddings are stored on every chunk so the Query API can fuse across them in a single request without an extra lookup.
Design the collection before writing any queries. The point granularity is the chunk: every chunk of every paper becomes one point. Title and abstract embeddings are stored on every chunk so the Query API can fuse across them in a single request without an extra lookup.
```python
from qdrant_client import QdrantClient, models
@@ -85,7 +87,7 @@ client.create_collection(
vectors_config={
"dense_chunk": models.VectorParams(size=384, distance=models.Distance.COSINE),
"dense_title": models.VectorParams(size=384, distance=models.Distance.COSINE),
"dense_summary": models.VectorParams(size=384, distance=models.Distance.COSINE),
"dense_abstract": models.VectorParams(size=384, distance=models.Distance.COSINE),
},
sparse_vectors_config={
"sparse_keywords": models.SparseVectorParams(
@@ -99,12 +101,12 @@ Each vector covers a different signal:
- `dense_chunk`: the content workhorse. Chunks are short enough that a single embedding represents them faithfully.
- `dense_title`: a few tokens that name the topic. A title hit is a strong signal even when no chunk matches.
- `dense_summary`: between title and chunk in length and specificity. Catches queries about the contribution rather than a passage.
- `dense_abstract`: between title and chunk in length and specificity. Catches queries about the contribution rather than a single passage.
- `sparse_keywords`: BM25 over title and tags concatenated. BM25 pays off on short structured fields where exact lexical matches matter.
Title and summary vectors are duplicated across every chunk of the same document. That trades storage for query simplicity: one collection, one Query API call, no `lookup_from`. If your titles or summaries are heavy, or you have millions of chunks per document, store them in a separate collection and reach them with `lookup_from`. For the typical case (a few dozen chunks per document, embeddings under a kilobyte each), inline storage is the simpler choice.
Title and abstract vectors are duplicated across every chunk of the same document. That trades storage for query simplicity: one collection, one Query API call, no `lookup_from`. If your titles or abstracts are heavy, or you have millions of chunks per document, store them in a separate collection and reach them with `lookup_from`. For the typical case (a few dozen chunks per document, embeddings under a kilobyte each), inline storage is the simpler choice.
Ingestion produces one point per chunk and reuses the title and summary embeddings.
Ingestion produces one point per chunk and reuses the title and abstract embeddings.
```python
DENSE_MODEL = "sentence-transformers/all-minilm-l6-v2"
@@ -120,10 +122,10 @@ points = []
for paper in papers:
chunks = chunk_sentences(paper["abstract"])
# Title, summary, and sparse docs are reused across every chunk of this paper; only the chunk text varies.
# Title, abstract, and sparse docs are reused across every chunk of this paper; only the chunk text varies.
# Cloud Inference embeds each Document on the server, so you don't need a client-side embedding library.
title_doc = models.Document(text=paper["title"], model=DENSE_MODEL)
summary_doc = models.Document(text=paper["abstract"], model=DENSE_MODEL)
abstract_doc = models.Document(text=paper["abstract"], model=DENSE_MODEL)
sparse_doc = models.Document(
text=paper["title"] + " " + " ".join(paper["categories"]),
model=BM25_MODEL,
@@ -135,7 +137,7 @@ for paper in papers:
vector={
"dense_chunk": models.Document(text=chunk, model=DENSE_MODEL),
"dense_title": title_doc,
"dense_summary": summary_doc,
"dense_abstract": abstract_doc,
"sparse_keywords": sparse_doc,
},
payload={
@@ -150,7 +152,7 @@ for paper in papers:
client.upload_points(collection_name="arxiv_multi_repr", points=points, batch_size=64)
```
After the upload completes, opening any point in the Qdrant Cloud UI shows all four named vectors attached to one chunk. `dense_chunk` carries the chunk's own embedding, while `dense_title`, `dense_summary`, and `sparse_keywords` are the same across every chunk of this paper.
After the upload completes, opening any point in the Qdrant Cloud UI shows all four named vectors attached to one chunk. `dense_chunk` carries the chunk's own embedding, while `dense_title`, `dense_abstract`, and `sparse_keywords` are the same across every chunk of this paper.
![A point in the arxiv_multi_repr collection showing all four named vectors](/documentation/tutorials/multi-representation-search/point.png)
@@ -188,7 +190,7 @@ A representation only earns its own prefetch if it carries signal independent of
- `dense_title` carries the topical naming. For a query like "diffusion models for high-resolution image synthesis", a paper titled "High-Resolution Image Synthesis with Latent Diffusion Models" surfaces from the title prefetch even when its abstract phrases the contribution differently. The chunk prefetch alone misses it.
- `sparse_keywords` carries lexical hits on title and tags that the dense embedding has averaged out: rare entity names, jargon, controlled-vocabulary tags.
Adding `dense_summary` as a fourth prefetch is worth it only if summaries surface what chunks don't. If summaries paraphrase the chunks, the extra prefetch adds latency without lift.
Adding `dense_abstract` as a fourth prefetch is worth it only if the abstract surfaces what chunks don't. If the abstract paraphrases the chunks, the extra prefetch adds latency without lift.
### How to Fuse
@@ -236,9 +238,9 @@ For the step-by-step build-up that produced this design (dense baseline, plus sp
## Where This Pattern Doesn't Fit
Multi-representation retrieval pays off when items are long and structured. It's overkill, and sometimes worse than a single dense vector, when items are short and homogeneous. Tweets, product names, and one-line forum titles don't have a meaningful title-versus-summary-versus-body distinction; splitting them into representations adds query latency without adding signal. For those cases, a single well-chosen dense model and a sparse fallback are usually enough.
Multi-representation retrieval pays off when items are long and structured. It's overkill, and sometimes worse than a single dense vector, when items are short and homogeneous. Tweets, product names, and one-line forum titles don't have a meaningful title-versus-body distinction; splitting them into representations adds query latency without adding signal. For those cases, a single well-chosen dense model and a sparse fallback are usually enough.
The pattern also assumes you can identify the representations cleanly. If your corpus has inconsistent metadata (some documents have summaries, others don't), the missing-representation case becomes its own design problem: empty vectors, fallback strategies, or a separate index per representation availability.
The pattern also assumes you can identify the representations cleanly. If your corpus has inconsistent metadata (some documents have abstracts, others don't), the missing-representation case becomes its own design problem: empty vectors, fallback strategies, or a separate index per representation availability.
## Open Ends