Merge branch 'master' into course/multi-vector-search

This commit is contained in:
kanungle
2026-03-14 04:26:04 -05:00
committed by GitHub
1923 changed files with 40829 additions and 3830 deletions
@@ -1,12 +1,12 @@
---
values:
- id: 0
title: 10 Million Downloads
title: 250+ Million Downloads
icon:
src: /img/about-us/rocket-logo.svg
alt: Rocket logo
- id: 1
title: 23K GitHub Stars
title: 29K GitHub Stars
icon:
src: /img/about-us/github-logo.svg
alt: github-logo
@@ -19,13 +19,13 @@ values:
alt: rust-logo
link: /articles/why-rust
- id: 3
title: 7.5k Members
title: 9k Members
icon:
src: /img/about-us/discord-logo.svg
alt: discord-logo
link: /community
- id: 4
title: 75+ employees
title: 100+ employees
icon:
src: /img/about-us/sparks-logo.svg
alt: sparks-logo
@@ -20,5 +20,21 @@ photoCards:
img: /img/carousel/qdrant-people-9.jpg
- id: 3
img: /img/carousel/qdrant-people-10.jpg
- id: 4
img: /img/carousel/qdrant-people-11.jpg
- id: 5
img: /img/carousel/qdrant-people-12.jpg
- id: 6
img: /img/carousel/qdrant-people-13.jpg
- id: 7
img: /img/carousel/qdrant-people-14.jpg
- id: 8
img: /img/carousel/qdrant-people-15.jpg
- id: 9
img: /img/carousel/qdrant-people-16.jpg
- id: 10
img: /img/carousel/qdrant-people-17.jpg
- id: 11
img: /img/carousel/qdrant-people-18.jpg
sitemapExclude: true
---
+14 -9
View File
@@ -1,19 +1,24 @@
---
title: Our Investors
link1:
title: Read Our Series A Story
url: /blog/series-a-funding-round/
link2:
title: Read About Our Seed Funding
url: /articles/seed-round/
links:
- title: We're Celebrating Series B
url: /blog/series-b-announcement/
- title: Read Our Series A Story
url: /blog/series-a-funding-round/
- title: Read About Our Seed Funding
url: /articles/seed-round/
investors:
- id: 0
logo: '/img/investors/spark-capital.svg'
logo: '/img/investors/avp.svg'
- id: 1
logo: '/img/investors/unusual-ventures.svg'
logo: '/img/investors/bosch-ventures.svg'
- id: 2
logo: '/img/investors/42cap.svg'
logo: '/img/investors/spark-capital.svg'
- id: 3
logo: '/img/investors/unusual-ventures.svg'
- id: 4
logo: '/img/investors/ibb-ventures.svg'
- id: 5
logo: '/img/investors/42cap.svg'
sitemapExclude: true
---
+1 -1
View File
@@ -2,7 +2,7 @@
cards:
- id: 0
title: Our Vision
content: To be the global leader in vector search technology, driving the future of AI-powered data retrieval and analysis.
content: To build the most efficient, scalable vector search engine, available as open source and managed cloud services, giving engineers explicit control over how they index, search, and retrieve high-dimensional data.
- id: 1
title: Our Mission
content: To deliver cutting-edge vector search software, available as open-source and scalable cloud services, empowering developers and businesses to unlock the full potential of high-dimensional data.
+2 -2
View File
@@ -11,11 +11,11 @@ extraImage1:
alt: Team
src: /img/about-us/team2-1x.jpg
srcLarge: /img/about-us/team2-2x.jpg
extraContent2: As the project gained traction, the decision was made to formally establish Qdrant and continue developing the vector search engine into its current form.<br/><br/>Today, Qdrant is the backbone of the most ambitious AI applications, powering everything from groundbreaking startups to enterprise-scale deployments with the best open-source vector database and enterprise-ready solutions.
extraContent2: As the project gained traction, the decision was made to formally establish Qdrant and continue developing the vector search engine into its current form.<br/><br/>Today, Qdrant is the backbone of the most ambitious AI applications, powering everything from groundbreaking startups to enterprise-scale deployments with the best open-source vector search and enterprise-ready solutions.
extraImage2:
alt: Team
src: /img/about-us/team3-1x.jpg
srcLarge: /img/about-us/team3-2x.jpg
subContent: "Our team has grown to 75+ experts across 20+ countries, but our mission remains unchanged: building the most scalable, high-performance vector search engine to fuel the future of AI and machine learning."
subContent: "Our team has grown to 100+ experts across 20+ countries, but our mission remains unchanged: building the most scalable, high-performance vector search engine to fuel the future of AI and machine learning."
sitemapExclude: true
---
@@ -37,7 +37,7 @@ features:
description: Qdrant enhances search speeds and control and context understanding through filtering on any nested entry in our payload. Unique architecture allows Qdrant to avoid expensive pre-filtering and post-filtering stages, making search faster and accurate.
link:
text: Learn More
url: /articles/filtrable-hnsw/
url: /articles/filterable-hnsw/
sitemapExclude: true
---
@@ -39,7 +39,7 @@ features:
description: Qdrant’s real-time, advanced vector search enables AI agents to act instantly on live data, which is crucial for time-sensitive, autonomous decision-making.
link:
text: HNSW
url: /articles/filtrable-hnsw/
url: /articles/filterable-hnsw/
- id: 3
icon:
src: /icons/outline/server-rack-blue.svg
@@ -5,7 +5,7 @@ tag:
src: /icons/outline/training-purple.svg
alt: Training
title: Building AI Agents for personalized recommendations with Qdrant and n8n
description: Learn in this video how to build an AI-powered recommendation system using Qdrant and n8n. It demonstrates how an AI agent retrieves data from Qdrant's vector database and leverages a large language model (LLM) to generate personalized recommendations based on user inputs.
description: Learn in this video how to build an AI-powered recommendation system using Qdrant and n8n. It demonstrates how an AI agent retrieves data from Qdrant's vector search engine and leverages a large language model (LLM) to generate personalized recommendations based on user inputs.
link:
text: Watch Now
url: https://www.youtube.com/watch?v=O5mT8M7rqQQ
@@ -145,7 +145,7 @@ In many vector search solutions, filtering is approached in two ways: **pre-filt
| ❌ | **Pre-filtering** | Has the linear complexity of computing the vector mask and becomes a bottleneck for large datasets. |
| ❌ | **Post-filtering** | The problem with **post-filtering** is tied to vector search "*everything fits and doesn't at the same time*" nature: imagine a low-cardinality filter that leaves only a few matching elements in the database. If none of them are similar enough to the query to appear in the top-X retrieved results, they'll all be filtered out. |
Qdrant [**took filtering in vector search further**](/articles/vector-search-filtering/), recognizing the limitations of pre-filtering & post-filtering strategies. We developed an adaptation of HNSW — [**filterable HNSW**](/articles/filtrable-hnsw/) — that also enables **in-place filtering** during graph traversal. To make this possible, we condition HNSW index construction on possible filtering conditions reflected by [**payload indexes**](/documentation/concepts/indexing/#payload-index) (inverted indexes built on vectors' [**metadata**](/documentation/concepts/payload/)).
Qdrant [**took filtering in vector search further**](/articles/vector-search-filtering/), recognizing the limitations of pre-filtering & post-filtering strategies. We developed an adaptation of HNSW — [**filterable HNSW**](/articles/filterable-hnsw/) — that also enables **in-place filtering** during graph traversal. To make this possible, we condition HNSW index construction on possible filtering conditions reflected by [**payload indexes**](/documentation/concepts/indexing/#payload-index) (inverted indexes built on vectors' [**metadata**](/documentation/concepts/payload/)).
**Qdrant was designed with a vector index being a central component of the system.** That made it possible to organize optimizers, payload indexes and other components around the vector index, unlocking the possibility of building a filterable HNSW.
@@ -1,17 +1,17 @@
---
title: Filtrable HNSW
title: Filterable HNSW
short_description: How to make ANN search with custom filtering?
description: How to make ANN search with custom filtering? Search in selected subsets without loosing the results.
# external_link: https://blog.vasnetsov.com/posts/categorical-hnsw/
social_preview_image: /articles_data/filtrable-hnsw/social_preview.jpg
preview_dir: /articles_data/filtrable-hnsw/preview
small_preview_image: /articles_data/filtrable-hnsw/global-network.svg
social_preview_image: /articles_data/filterable-hnsw/social_preview.jpg
preview_dir: /articles_data/filterable-hnsw/preview
small_preview_image: /articles_data/filterable-hnsw/global-network.svg
weight: 60
date: 2019-11-24T22:44:08+03:00
author: Andrei Vasnetsov
author_link: https://blog.vasnetsov.com/
category: qdrant-internals
# aliases: [ /articles/filtrable-hnsw/ ]
aliases: [ /articles/filtrable-hnsw/ ]
---
If you need to find some similar objects in vector space, provided e.g. by embeddings or matching NN, you can choose among a variety of libraries: Annoy, FAISS or NMSLib.
@@ -37,7 +37,7 @@ We need to build a navigation graph among all indexed points so that the greedy
This graph is constructed by sequentially adding points that are connected by a fixed number of edges to previously added points.
In the resulting graph, the number of edges at each point does not exceed a given threshold $m$ and always contains the nearest considered points.
![NSW](/articles_data/filtrable-hnsw/NSW.png)
![NSW](/articles_data/filterable-hnsw/NSW.png)
### How can we modify it?
@@ -55,9 +55,9 @@ Therefore, the theoretical conclusions obtained in the [Percolation theory](http
This statement also confirmed by experiments:
{{< figure src=/articles_data/filtrable-hnsw/exp_connectivity_glove_m0.png caption="Dependency of connectivity to the number of edges" >}}
{{< figure src=/articles_data/filterable-hnsw/exp_connectivity_glove_m0.png caption="Dependency of connectivity to the number of edges" >}}
{{< figure src=/articles_data/filtrable-hnsw/exp_connectivity_glove_num_elements.png caption="Dependency of connectivity to the number of point (no dependency)." >}}
{{< figure src=/articles_data/filterable-hnsw/exp_connectivity_glove_num_elements.png caption="Dependency of connectivity to the number of point (no dependency)." >}}
There is a clear threshold when the search begins to fail.
@@ -79,13 +79,13 @@ In this case, the total number of edges will increase by no more than 2 times, r
Second case is a little harder. A connection may be lost between two categories if they lie in different clusters.
![category clusters](/articles_data/filtrable-hnsw/hnsw_graph_category.png)
![category clusters](/articles_data/filterable-hnsw/hnsw_graph_category.png)
The idea here is to build same navigation graph but not between nodes, but between categories.
Distance between two categories might be defined as distance between category entry points (or, for precision, as the average distance between a random sample). Now we can estimate expected graph connectivity by number of excluded categories, not nodes.
It still does not guarantee that two random categories will be connected, but allows us to switch to multiple searches in each category if connectivity threshold passed. In some cases, multiple searches can be even faster if you take advantage of parallel processing.
{{< figure src=/articles_data/filtrable-hnsw/exp_random_groups.png caption="Dependency of connectivity to the random categories included in search" >}}
{{< figure src=/articles_data/filterable-hnsw/exp_random_groups.png caption="Dependency of connectivity to the random categories included in search" >}}
Third case might be resolved in a same way it is resolved in classical databases.
Depending on labeled subsets size ration we can go for one of the following scenarios:
@@ -100,7 +100,7 @@ Next we also connect neighboring buckets to achieve graph connectivity. We still
Geographical case is a lot like a numerical one.
Usual geographical search involves [geohash](https://en.wikipedia.org/wiki/Geohash), which matches any geo-point to a fixes length identifier.
![Geohash example](/articles_data/filtrable-hnsw/geohash.png)
![Geohash example](/articles_data/filterable-hnsw/geohash.png)
We can use this identifiers as categories and additionally make connections between neighboring geohashes.
It will ensure that any selected geographical region will also contain connected HNSW graph.
@@ -102,7 +102,7 @@ Let's take a look at a non-exhaustive list of data structures and potential impr
| Tenant Isolation | Vector Storage | Defragmented Vector Storage | Faster access to on-disk data |
For more info on payload-aware connections in HNSW, read our [previous article](/articles/filtrable-hnsw/).
For more info on payload-aware connections in HNSW, read our [previous article](/articles/filterable-hnsw/).
This time around, we will focus on the latest additions to Qdrant:
- **the immutable hash map with perfect hashing**
@@ -214,7 +214,7 @@ But let's first see how much RAM we need to serve 1 million vectors and then we
### Vectors and HNSW graph stored using MMAP
In the third experiment, we tested how well our system performs when vectors and [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) graph are stored using the memory-mapped files.
In the third experiment, we tested how well our system performs when vectors and [HNSW](https://qdrant.tech/articles/filterable-hnsw/) graph are stored using the memory-mapped files.
Create collection with:
```http
@@ -61,7 +61,7 @@ negative examples.
## HNSW ANN example and strategy
Let’s start with an example to help you understand the [HNSW graph](/articles/filtrable-hnsw/). Assume you want
Let’s start with an example to help you understand the [HNSW graph](/articles/filterable-hnsw/). Assume you want
to travel to a small city on another continent:
1. You start from your hometown and take a bus to the local airport.
@@ -145,7 +145,7 @@ significantly lower. However, if the best negative score is higher than the best
further away from the negatives. That procedure effectively **pulls the traversal procedure away from the negative examples**.
If you want to know more about the internals of HNSW, you can check out the article about the
[Filtrable HNSW](/articles/filtrable-hnsw/) that covers the topic thoroughly.
[Filterable HNSW](/articles/filterable-hnsw/) that covers the topic thoroughly.
## Food Discovery demo
@@ -0,0 +1,466 @@
---
title: "Relevance Feedback in Qdrant"
short_description: "The story behind the vector search-native relevance feedback feature, available since 1.17.0, which increases the relevance of search results universally, cheaply, and at scale."
description: "The story behind the vector search-native relevance feedback feature, available since 1.17.0, which increases the relevance of search results universally, cheaply, and at scale."
social_preview_image: /articles_data/relevance-feedback/preview/social_preview.jpg
preview_dir: /articles_data/relevance-feedback/preview
weight: -170
author: Evgeniya Sukhodolskaya
date: 2026-02-20T00:00:00+03:00
draft: false
keywords:
- search relevance
- reranking
- query rewriting
- relevance
category: machine-learning
---
A year ago, we dropped a statement-bomb in the “[Relevance Feedback in Information Retrieval](https://qdrant.tech/articles/search-feedback-loop/)” article and then went silent.
We claimed that even though the information retrieval research field has proposed many useful mechanisms for increasing the relevance of search results, none of them made it to the neural search industry, simply because these approaches are not scalable.
Certainly, there are methods widely used to improve the relevance of retrieved results: query rewriting, for example. Yet none of the vector search solutions out there have tried to use the possibilities that come with full access to the vector search index: traversing it in the direction of relevance, instead of guessing where to shoot the query to the vector space.
To break this vector-search-native tools silence, we’re introducing [Relevance Feedback Query](https://qdrant.tech/blog/qdrant-1.17.x/#relevance-feedback-query), a universal, cheap and scalable method for improving the relevance recall of your search results.
## Vector Search Optimization
In vector search, like in any field, you’re always trading off speed, cost, and quality on the way to production. Everyone needs instruments to tune that balance for their use case, like [vector quantization](https://qdrant.tech/articles/what-is-vector-quantization/) to lower costs or [reranking](https://qdrant.tech/documentation/search-precision/reranking-semantic-search/) to increase relevance.
For those instruments to work (and to stick in industry), they have to be **universal**, **native to the stack**, and **reasonable**.
> “Reasonable” means optimizing one side of the speed–cost–quality triangle **without** tanking the other two.
At scale, neural search is usually limited to fairly simple dense encoders (embeddings up to a few thousand dimensions), as running large models to sift through billions of data points is, well, unreasonable.
So most search quality-boosting tools aim to adjust/align/guide these simple retrievers. The guidance can come from the search end-user or a domain expert: rescoring via business-logic-based rules, reranking with a learned-to-rank model, and using **relevance feedback** mechanisms.
### What is Relevance Feedback
**Relevance feedback** distills signals from current search results into the next retrieval iteration to surface more relevant documents.
It can be provided by a human or a model, be binary (relevant–irrelevant) or granular, rescoring top documents by their relative relevance to the query.
{{< figure src="/articles_data/relevance-feedback/relevance_feedback_types.png" alt="Diagram showing a query flowing into a retrieval system, which returns top-ranked documents. These documents are then processed in three ways: pseudo-relevance feedback, binary user or classifier feedback, and model-driven re-scored feedback. Arrows illustrate how the feedback signals update document rankings" caption="Feedback types">}}
Based on this feedback, one of the retrieval components is adjusted: the query or the scoring method between the query and documents. The next retrieval iteration is done with this aligned component.
> For a detailed taxonomy with various methods from the field, check out our article “[Relevance Feedback in Information Retrieval](https://qdrant.tech/articles/search-feedback-loop/)”.
Relevance-feedback-based methods are a standard in full-text search, some of them proposed more than 50 years ago.
Yet it feels like, when it comes to modern vector search, users have to reinvent the relevance feedback wheel due to the lack of universal interfaces: prompt a search agent to rewrite queries in endless, costly loops, do heuristics-based vector math and fine-tune models per request on the client side... Simply put, they have to struggle.
### Tools of a Vector Search Engine
The goal is to fill this gap and let search end-users (doesn't matter if humans or agents) get a proper relevance feedback interface to use in our vector search engine.
So what makes a method a production-ready for vector search?
#### ...Should Be Cheap
Since additional retrieval iterations already add to the search latency, we can't afford to spend too much time or money on gathering feedback. That means:
**Very few documents for feedback**
The set of documents we pass to a **feedback model** should be extremely small. It is rarely affordable to run an LLM, use a cross encoder, or fine tune a machine learning model on hundreds of documents per request.
> A **feedback model** here is anything – agent, ML model, formula, … – that’s capable of providing (numerical) information on how relevant the retrieved document is to the query.
We should already improve recall of the next retrieval iteration with that small sample. We can’t afford having the whole dataset as a feedback context for an agent.
**Automated**
You might have noticed that we wrote "a feedback **model**". Humans are known to be the laziest at providing feedback, so the search feedback loop must run on its own.
**No labeling required**
Models which require many labels or domain expert input are hard to adopt. The feedback tool should be easy to train, and data -- self-supervised.
#### ...Should Be Universal
**For every type of data**
Vector search is attractive because it is agnostic to the data type: images, video, audio, molecules… Not only text. Vector search native tools should be the same, operating on vectors without minding their nature. That is why query rewriting as reformulating text won’t suffice.
**For every type of signal**
Some relevance feedback methods depend on clean feedback signals: strictly relevant or irrelevant documents.
In real life, especially when we need relevance boosting, results returned to a request could be, for example, *meh* and *more meh*. Ideally, the method should be able to use the noisy signals.
## Our Idea
Usually neither LLMs nor subject matter experts providing feedback know the dataset’s exact shape or the vector space it spans. For that they would need to process a large subset of stored vectors.
What they can do is help define a relevance direction from the examples they have seen. A retriever could then use these direction hints on the next retrieval iteration.
### Relevance Direction
{{< figure src="/articles_data/relevance-feedback/lost_astronaut.png" alt="An astronaut wearing a red spacesuit and backpack stands in a dense, shadowy forest with tall trees and blue-tinted light filtering through, looking toward a brighter glow in the distance." >}}
Imagine a hiker lost in a forest with a weak phone signal. They still manage to send a couple of quick photos to a friend who knows orienteering.
The friend does not have a map of the whole forest or a live view of every tree. Still, they try to help and reply:
- *“Better go away from anything looking like this plant. It might signal that a swamp is near.”*
- *“I see moss on this tree. Moss grows on the north side. If you keep north you might find a way out.”*
Now the hiker has extra knowledge to choose a path, comparing the landscape around them to the photos and applying the direction hints from their experienced friend.
Read this as vector search.
1. The forest is a **vector space**.
2. The hiker is a **retriever model.**
3. The brilliant friend is a **feedback model**.
4. The photos are the hiker's **context**, **limited** to the amount that is affordable to use for feedback.
And the hiker’s new “where-should-I-go” decisions based on the acquired information are **feedback-based** (forest path) **scoring**.
### Using Entire Vector Space
So the idea is to use feedback to define a more relevant direction (**closer** to this, **further** from that) within the vector space spanned by the dataset.
That implies warping the notion of "closer" and "further", the distance (or similarity) **scoring metric** used during the retrieval.
Tweaking the relevance scoring formula based on the feedback model’s signal is nothing revolutionary. What changes is that this feedback-based scoring is used during the traversal of the **entire** dataset used by the vector search engine, not just a subset of results.
### Sketch of the Method
So, stepping away from analogies (into an even deeper forest), the algorithm should consist of:
1. Initial retrieval.
2. Getting a small amount of feedback (within the **context limit**) on the results of this initial retrieval.
> A **context limit** is how many top documents from the initial retrieval the feedback model scores. We want to keep it as small as possible to **save time and resources**.
A good choice for feedback format is granular judgments, aka pointwise relevance scores, between the query and the retrieved documents. They are more helpful when the retrieved documents are not strictly relevant or irrelevant.
3. Extracting relevance signals from this feedback and propagating them into a new similarity (distance) scoring formula.
4. Using this new feedback-based similarity scoring formula in the next retrieval iteration **on the whole dataset of documents** to traverse the vector index in the direction of more relevance.
Simply put, we're no longer using cosine to score similarities, but rather a formula adjusted for feedback.
## Dissecting Feedback
Let's talk about how a very small amount of feedback can be used most effectively, what relevance signals we can extract from it, and how exactly.
### Context Pairs
If we show only one document to the feedback model, it is hard for it to say whether it is already a good match. Perhaps there are no relevant matches in the dataset at all.
With two documents, this model could already judge which one is closer to what the user expects. Feedback then could be a **context pair** of (more) **positive** and (more) **negative** examples-documents.
> A **context pair (positive, negative)** is two documents from the top context limit results of the initial retrieval. The positive received a higher relevance-to-query score from the feedback model, and the negative received a lower score.
{{< figure src="/articles_data/relevance-feedback/context_pair.png" alt="Diagram titled Re-score with Feedback Model showing five documents labeled Doc 1 to Doc 5, each with two horizontal bars for Retriever Score in blue and Feedback Score in green. Doc 2 has the highest feedback score and Doc 5 has the lowest feedback score. Doc 2 and Doc 4 show higher feedback scores than retriever scores, while Doc 3 and Doc 5 show lower feedback scores. On the right, a Context Pair section shows Doc 2 labeled Positive and Doc 5 labeled Negative, with curved arrows linking these examples back to the document list" caption="How context pairs are formed">}}
### Context Pair's Confidence
Two documents nearly indistinguishable from the perspective of a feedback model give far less information than a context pair with one clearly more relevant document.
When documents in the context pair seem different to the feedback model, it is more **confident** in guiding the retriever.
{{< figure src="/articles_data/relevance-feedback/confidence_of_context_pair.png" alt="Diagram titled Re-score with Feedback Model showing five documents labeled Doc 1 to Doc 5, each with two horizontal bars for Retriever Score in blue and Feedback Score in green. Doc 2 has the highest feedback score and Doc 5 has the lowest feedback score. Doc 2 and Doc 4 show higher feedback scores than retriever scores, while Doc 3 and Doc 5 show lower feedback scores. On the right, a Context Pair section shows Doc 2 labeled Positive and Doc 5 labeled Negative, with curved arrows linking these examples back to the document list" caption="Context pair's confidence">}}
### Direction's Delta
Then, if the retriever favors a document closer to the negative example by some **delta**, the retriever should most probably adjust its search direction in the vector space.
{{< figure src="/articles_data/relevance-feedback/delta_as_distance.png" alt="The image shows a candidate document represented as a dark dot in the center, with dashed arrows indicating its distances to a circled positive example above and a circled negative example below. A ‘-delta’ segment on the arrow to the positive example highlights how much more the retriever currently favors the negative example. The diagram illustrates that the retriever should adjust its direction toward candidates closer to the positive example, as shown by an additional dashed arrow." caption="The point in vector space is closer (more similar) to the negative element of a context pair by a delta." width="80%">}}
## Feedback-Based Scoring
With the context pair(s) at our expense, we can try the following feedback-based scoring during retrieval:
1. Let the retriever still have a say in what is relevant to the query, count in its **score** (to query).
2. Yet reward candidates that are closer to the positive element of the context pair based on **delta**.
3. Especially when the feedback model had high **confidence** in this pair.
{{< figure src="/articles_data/relevance-feedback/scoring.png" alt="Diagram titled Scoring with context pairs showing how candidates from a collection are scored using three components: similarity to the query, similarity to a positive example Doc 2, and similarity to a negative example Doc 5. For each candidate, horizontal bars show score to query in blue, score to positive in green, and score to negative in red. A dashed delta indicates the difference between the positive and negative scores, which rewards candidates closer to the positive example and farther from the negative one. Larger deltas represent stronger separation in relevance, especially when the context pair is known with high confidence." caption="Feedback-based scoring using candidate’s similarity to the query and the context pair.">}}
So, we need to combine signals (**score, delta, confidence**) in a reasonable, simple scoring formula with a few parameters.
Usually what helps in coming up with a reasonable formula is looking at edge cases.
### The Math of Edge Cases
**Direction’s delta is zero**
If the retriever sees the positive and the negative documents (from the context pair) as equally similar to the candidate document, there’s no direction-establishing signal.
**Context pair's confidence is zero.**
If, from the feedback model’s perspective, both documents in a context pair are identical to the query, there’s, once again, no information on direction of relevance.
In both cases the final formula should rely on the retriever's judgment, as opposed to a situation where feedback signals are strong.
### Naive Formula
One of the options to express this behavior is through a weighted sum of signals, ensuring confidence and delta don't collapse into a single joint term by exponentiating one of them.
Applying the math of edge cases, we came up with this three-parameter (**a**, **b** and **c**) formula.
$$
F = a \cdot \text{score} + \sum_{p=1}^{\text{# pairs}} \text{confidence}_{p}^{b} \cdot c \cdot \text{delta}_p
$$
It computes the score between the query and a candidate document on the retrieval-with-relevance-feedback iteration.
*As you see, the amount of context pairs used in the scoring formula can be more than one but one is also an option*.
> Is this formula set in stone? Absolutely not! We tried three others in experiments, and this one was the simplest that worked. In future releases we plan to allow providing custom formulas.
#### What Goes Into It
$\text{score}_\text{retriever}(\text{query}, \text{candidate document})$
|||
|---|---|
| What does it mean? | A similarity score between the query and the candidate embedding generated by the retriever model. F.e., cosine similarity. |
| When is it calculated? | On the second step of retrieval, during search for more relevant candidates in the vector space. |
| Example | If $\text{cosine}(\text{query}, \text{candidate document}) = 0.83$, then $\text{score} = 0.83$. |
$\text{confidence}_\text{feedback}(\text{context pair}_p)$
|||
|---|---|
| What does it mean? | A difference in relevance to the query for the two documents forming a context pair, as scored by the feedback model. |
| When is it calculated? | At the moment feedback is collected, right after the initial retrieval. |
| Example | We have $\text{context pair}_p$ out of $\text{doc}_1$ and $\text{doc}_2$.<br/>A feedback model scores them $0.99$ (more relevant, so “positive”) and $0.70$ (less relevant, so “negative”) respectively against the query.<br/>$\text{confidence}$ of the $\text{context pair}_p = 0.99 - 0.70 = 0.29$. |
$\text{delta}_\text{retriever}(\text{context pair}_p, \text{candidate document})$
|||
|---|---|
| What does it mean? | The difference between the candidate’s similarity (e.g., cosine) to the positive document and to the negative document from the context pair. <br> All the similarity scores are calculated on embeddings generated by the retriever. |
| When is it calculated? | On the second step of retrieval, during search for more relevant candidates in the vector space. |
| Example | Given the context pair ($\text{doc}_1$, $\text{doc}_2$) mentioned above:<br/>If $\text{cosine}(\text{doc}_1, \text{candidate})$ and $\text{cosine}(\text{doc}_2, \text{candidate})$ are respectively $0.78$ and $0.40$.<br/>$\text{delta} = 0.78 - 0.40 = 0.38$. |
Now the question is: **Where do we take a, b, and c from?**
> We'll need to obtain formula parameters (a, b and c) through some simple training process, because score, confidence, and delta value distributions will likely vary by dataset, retriever, and feedback model.
### Training Objective
With the right a, b and c, adjusted to dataset, retriever and feedback scores distribution, the naive formula should rank documents better than our simple retriever.
That leads us to **minimizing Pairwise Ranking Loss** as the training objective.
Since the formula has only three parameters, **the amount of training data needed is very small**. A few hundred domain-relevant queries (per dataset) are more than enough.
For each query, **the goal is for the feedback-based scoring formula to rank documents as well as the feedback model would**.
Hence, to form a training dataset:
1. We retrieve the top X (`limit`) documents per query.
X is chosen as large as affordable, given the cost of obtaining a golden ranking on these X documents from the feedback model.
2. We use feedback from the top K (`context limit`) results to mine context pairs for the formula, with K being much smaller than X.
3. The remaining X − K documents form the training pool, on which our naive formula learns to rank better (similarly to how the feedback model would).
## Experiments
Before building a new interface, we needed to confirm our instinct: that a relevance feedback scoring can surface relevant documents that initial retrieval misses.
### Setup
Each triplet **retriever**, **feedback model**, **dataset** forms one experiment.
<details>
<summary><b>Datasets</b></summary>
<p>A subset of <a href="https://github.com/beir-cellar/beir">BEIR (Informational Retrieval benchmark)</a>:</p>
<ul>
<li>MSMARCO (8.84mln documents)</li>
<li>SCIDOCS (25K documents)</li>
<li>Quora (523K documents)</li>
<li>FiQA-2018 (57K documents)</li>
<li>NFCorpus (3.6K documents)</li>
</ul>
<p>— documents and queries, no qrels needed.</p>
</details>
<details>
<summary><b>Retrievers</b></summary>
<p>As retrievers we chose:</p>
<ul>
<li><a href="https://huggingface.co/jinaai/jina-embeddings-v2-base-en">jina-embeddings-v2-base-en</a> (768 output dimensions)</li>
<li><a href="https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1">mxbai-embed-large-v1</a> (1024 output dimensions)</li>
<li><a href="https://huggingface.co/michaelfeil/Qwen3-Embedding-0.6B-auto">Qwen3-Embedding-0.6B</a> (1024 output dimensions)</li>
</ul>
</details>
<details>
<summary><b>Feedback models</b></summary>
<p>The cheapest feedback type suitable for testing the hypothesis that came to mind was <b>using embedding models (bi-encoders) with higher dimensionality than those used for retrieval</b>.</p>
<p>The feedback from these models is a similarity score between query and document embeddings (so, simply put, cosine similarity).</p>
<p>We chose as feedback models:</p>
<ul>
<li><a href="https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1">mxbai-embed-large-v1</a> (1024 output dimensions)</li>
<li><a href="https://huggingface.co/michaelfeil/Qwen3-Embedding-0.6B-auto">Qwen3-Embedding-0.6B</a> (1024 output dimensions)</li>
<li><a href="https://huggingface.co/michaelfeil/Qwen3-Embedding-4B-auto">Qwen3-Embedding-4B</a> (2560 output dimensions)</li>
<li><a href="https://huggingface.co/colbert-ir/colbertv2.0">colBERTv2.0</a> (multivectors of 128 dimensions each)</li>
</ul>
</details>
With [Qdrant’s Cloud Inference](https://qdrant.tech/cloud-inference/) at our disposal, we ran embedding inference once for all datasets on both retrieval and feedback models, stored the embeddings in Qdrant, and then experimented with different formulas.
This way, we could use in our experiments one model interchangeably as a retriever and a feedback model.
### Metric
We want to check if relevance feedback-based retrieval increases relevance recall. That means surfacing documents more relevant to the query than the vanilla retriever did.
**How did we emulate "we-want-to-surface" quality in experiments?**
* For each query, the retriever returned a ranked list of documents.
* We obtained a ground-truth relevance scoring of this list using the feedback model.
* We mined context pairs for the feedback-based scoring formula only from the top-K retrieved results.
This top-K emulated the context limit in production, aka the initial retrieval results available for feedback.
* From the feedback model's scores within the top-K, we took the highest score as a **threshold**.
Any document in the list **outside the top-K** whose feedback model score **exceeded this threshold** was considered a **desired result** for "the next retrieval iteration".
{{< figure src="/articles_data/relevance-feedback/goal.png" alt="Diagram contrasting retriever's ranking with ranking made by a feedback model. The top section shows the retriever model’s initial ranking, where only the highest scoring items on the left are included in the default top K. The bottom section shows the golden ground-truth relevance scoring on the same cadidates from a feedback model, with some different items now receiving higher scores. Orange bars labeled Desired results appear below the original top K threshold line, illustrating documents that were scored too low by the retriever. The figure demonstrates more relevant documents that were not included in the original top K results that we'd want to surface." caption="More relevant documents we'd like to surface">}}
**What we wanted to measure**
On the next retrieval iteration, can our feedback-based formula pull more relevant documents into the top N than the vanilla retriever did?
*Relevance is judged by the feedback model's ground-truth scores.*
For that, we came up with the **abovethreshold@K** metric.
#### Abovethreshold@N
A custom metric to compare vanilla retriever versus relevance feedback-based scoring.
**How we computed it (per query):**
We looked at N positions after the top-K documents used for mining context pairs (positions K+1 through K+N).
* **Vanilla retriever:** counted how many "we-want-to-surface" documents already appeared in these positions.
* **Relevance feedback-based scoring:** rescored all remaining documents (outside top-K) using a trained naive formula, reranked them, and counted how many "we-want-to-surface" documents appeared in the first N positions of the new ranking.
{{< figure src="/articles_data/relevance-feedback/metric.png" alt="Diagram showing the abovethreshold at 10 metric comparing vanilla second retrieval with feedback-based second retrieval. The threshold is the highest feedback model score within the original top K. Documents outside the top K whose feedback score exceeds this threshold are considered desired results. In the vanilla ranking, zero out of two desired documents appear in the next 10 positions, while after feedback-based rescoring, two out of two appear, showing improved recall of relevant documents." caption="Abovethreshold@N metric">}}
To get one number per test set, we summed abovethreshold@N counts across all queries per method and calculated the **relative gain**: `(feedback_count - vanilla_count) / vanilla_count`.
### Parameters
<details>
<summary><b>Training data size</b></summary>
| Dataset | Train queries | Test queries |
|------------|---------------:|-------------:|
| NFCorpus | 223 | 100 |
| FiQa-2018 | 500 | 148 |
| SCIDOCS | 500 | 500 |
| MSMARCO | 4000 | 1000 |
| Quora | 6000 | 1000 |
Queries for training are split in 50% train, 50% validation.
</details>
<details>
<summary><b>Training parameters</b></summary>
| Parameter | Value |
|-------------------------------------------------------------|------:|
| Golden feedback model scoring limit (documents per query) | 100 |
| Context Limit, aka top K (to mine context pairs) | 5 |
| Context pairs used (out of all, sorted by the confidence) | top-1 |
| Learning rate | 0.005 |
| Epochs with early stopping and patience 200 | 2000 |
</details>
<details>
<summary><b>Testing parameters</b></summary>
| Parameter | Value |
|-------------------------------------------------------------|------:|
| Context Limit, aka top K (to mine context pairs) | 3 |
| Context pairs used (out of all, sorted by the confidence) | all |
| Evaluation window size (N in abovethreshold@N) | 10 |
</details>
### Results
Rescoring humongous datasets like MSMARCO on the user side for every query would not have been fun, especially while iterating on different formulas and hyperparameters. We faced the same problem that limits many relevance feedback researchers and practitioners, testing approaches only on a subset of all documents.
Our Relevance Feedback Query API is a remedy against this limitation, but first we needed experiment results to justify its implementation. A chicken-and-egg problem. **Hen**ce, to break the loop, we simulated feedback-based scoring on a hundred documents per query.
Out of all retriever–feedback model pairs, three leaders emerged with the following relative gain in **abovethreshold@10** compared to the vanilla retriever:
| | Qwen3-0.6B → colBERTv2.0 | Qwen3-0.6B → Qwen3-4B | mxbai-large-v1 → colBERTv2.0 |
| ----- | ----- | ----- | ----- |
| **NFCorpus** | +10.34% | +10.61% | **+21.57%** |
| **FiQA-2018** | +6.45% | +10.94% | **+12.24%** |
| **SCIDOCS** | **+38.72%** | +0.69% | +9.55% |
| **MSMARCO** | **+23.23%** | +16.73% | +2.40% |
| **Quora** | **+5.04%** | +2.67% | 0.00% |
[jina-embeddings-v2-base-en](https://huggingface.co/jinaai/jina-embeddings-v2-base-en), being a smaller and less expressive retriever, did not benefit from the feedback signal. Results varied significantly across queries, with no consistent improvement from adding feedback-based scoring.
| | jina-v2-base → mxbai-large-v1 | jina-v2-base → Qwen3-0.6B | jina-v2-base → Qwen3-4B |
| ----- | ----- | ----- | ----- |
| **NFCorpus** | **+4.62%** | −3.85% | +2.86% |
| **FiQA-2018** | −3.90% | −1.59% | **+3.97%** |
| **SCIDOCS** | **+4.55%** | +1.62% | +1.82% |
| **MSMARCO** | +2.57% | +2.23% | **+2.82%** |
| **Quora** | 0.00% | −1.37% | 0.00% |
#### Takeaways
* Relevance feedback scoring is most effective when the feedback model (reasonably) disagrees with the retriever's ordering within the top-K (context limit). "Reasonably" meaning the feedback aligns with the user's actual notion of relevance.
* The retriever's expressiveness limits how much it can benefit from feedback. The retriever operates in a lower-dimensional space and can't capture all the distinctions the feedback model makes. Past a certain point, a more sophisticated feedback model won't help.
* Initially we used only the single highest-confidence context pair in the feedback-based scoring formula, both during training and at inference. We then discovered that at inference time, incorporating signals from additional (all) context pairs improves results.
## Relevance Feedback Query
The results were convincing enough to justify implementing Relevance Feedback Query. And [here it is](https://qdrant.tech/documentation/concepts/search-relevance/#relevance-feedback), ready for your retrieval pipelines!
### When to Use It
Use Relevance Feedback Query once a basic retrieval pipeline is in place and you're looking for additional techniques to boost result relevance, such as [Maximal Marginal Relevance (MMR)](https://qdrant.tech/documentation/concepts/search-relevance/#maximal-marginal-relevance-mmr), [Reranking](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries), or [Score Boosting](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting).
**It's here not to replace but to complement other search relevance tools.**
For example, Relevance Feedback Query can be a great aid for search agents, letting you propagate the agent's understanding of your use case directly to the vector search index.
### How to Use It
For ease of use, as one shouldn't need a machine learning degree to use a new feature, we published a [Python package that customizes Naive Formula weights for your dataset, retriever, and feedback model](https://pypi.org/project/qdrant-relevance-feedback/).
What you need is a Qdrant collection, an idea of which feedback model you'd like to use to guide your retriever, and, optionally, a small set of use case-specific queries (50–300).
We provide `QdrantRetriever` (using our [Cloud Inference](https://qdrant.tech/documentation/cloud/inference/) or [FastEmbed](https://qdrant.tech/documentation/fastembed/) locally) and `FastembedFeedback` with bi-encoders, late interaction models, and cross-encoders as feedback model options.
That said, you can define your own retrievers and feedback models, whatever works best for your data modality and use case.
> A feedback model can be anything: a bi-encoder, a late interaction model, a cross-encoder, or an LLM.
If no training queries are supplied, the package will train the formula directly on documents sampled from your Qdrant collection.
> **Warning:** If your use case doesn't involve document-to-document semantic similarity search, training on sampled documents alone may completely cancel the effect of relevance feedback scoring on real data.
> It's far more effective to use real queries.
Once you've obtained the weights, simply plug them into your [Qdrant Client of choice](https://qdrant.tech/documentation/concepts/search-relevance/#relevance-feedback).
### Evaluating Your Gains
Additionally, the [Relevance Feedback Parameters package](https://pypi.org/project/qdrant-relevance-feedback/) provides an `Evaluator` module with two metrics: **relative gain** based on the **abovethreshold@N** metric from the "Experiments" section above, and a metric more recognizable to people in search -- **Discounted Cumulative Gain (DCG) Win Rate**.
- **Discounted Cumulative Gain (DCG) Win Rate**
For each query, we compute DCG@N for both compared methods (vanilla and relevance feedback-based retrieval) against ground truth relevancy scores from a feedback model. The method with the higher DCG@N gets a "win".
With this module you can immediately check the potential benefit of adding Relevance Feedback-based retrieval to your pipelines (or reproduce the benchmarks above).
If you don't see any **relative gain** on your test queries, this particular "feedback model, retriever, dataset" triplet might just not work together; for example, if the feedback model fully agrees with the retriever's ranking in all the cases, making the relevance feedback signal redundant.
If you're unsure which feedback model to try next or have questions about Relevance Feedback Query in general, reach out in our [Discord community](https://qdrant.to/discord).
## Conclusion
We've released [a new relevance feedback tool in Qdrant 1.17.0](https://qdrant.tech/blog/qdrant-1.17.x/#relevance-feedback-query) to help increase the relevance of vector search results. **It is built for scale, cheap, customizable, and universal.**
**It is cheap to use**, as the time and resources spent on obtaining relevance feedback are minimal.
**It is easy to adapt** to your use case: dataset, retriever, and feedback model (any, from a bi-encoder to an LLM).
**It is universal** -- data type agnostic (texts, images, code, molecules, you name it), as it works directly on embeddings.
**It applies to the whole vector space**, not just a subset of documents, a key distinction from approaches limited to reranking retrieved results.
It's a tool for production.
*If you'd like advice on Relevance Feedback Query usage or have ideas on how to enhance the method, reach out in our [Discord community](https://qdrant.to/discord). Follow [our LinkedIn](https://www.linkedin.com/company/qdrant/) for upcoming tutorials on various applications of Relevance Feedback Query.*
@@ -50,7 +50,7 @@ Our plan for the current [open-source roadmap](https://github.com/qdrant/qdrant/
Qdrant started more than two years ago with the mission of building a vector database powered by a well-thought-out tech stack. Using Rust as the system programming language and technical architecture decision during the development of the engine made Qdrant the leading and one of the most popular vector database solutions.
Our unique custom modification of the [HNSW algorithm](/articles/filtrable-hnsw/) for Approximate Nearest Neighbor Search (ANN) allows querying the result with a state-of-the-art speed and applying filters without compromising on results. Cloud-native support for distributed deployment and replications makes the engine suitable for high-throughput applications with real-time latency requirements. Rust brings stability, efficiency, and the possibility to make optimization on a very low level. In general, we always aim for the best possible results in [performance](/benchmarks/), code quality, and feature set.
Our unique custom modification of the [HNSW algorithm](/articles/filterable-hnsw/) for Approximate Nearest Neighbor Search (ANN) allows querying the result with a state-of-the-art speed and applying filters without compromising on results. Cloud-native support for distributed deployment and replications makes the engine suitable for high-throughput applications with real-time latency requirements. Rust brings stability, efficiency, and the possibility to make optimization on a very low level. In general, we always aim for the best possible results in [performance](/benchmarks/), code quality, and feature set.
Most importantly, we want to say a big thank you to our [open-source community](https://qdrant.to/discord), our adopters, our contributors, and our customers. Your active participation in the development of our products has helped make Qdrant the best vector database on the market. I cannot imagine how we could do what we’re doing without the community or without being open-source and having the TRUST of the engineers. Thanks to all of you!
@@ -0,0 +1,173 @@
---
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 1: Why Sparse Embeddings Beat BM25"
short_description: "Dense embeddings blur exact matches. Sparse embeddings keep the details that matter in e-commerce search."
description: "Part 1 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Learn how sparse embeddings outperform BM25 and dense models for product search, how SPLADE works, and why Qdrant's native sparse vector support matters."
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-1/preview
social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-1/preview/social_preview.jpg
weight: -200
author: Thierry Damiba
author_link: https://github.com/thierrydamiba
date: 2026-03-09T00:00:00.000Z
category: practicle-examples
---
*This is Part 1 of a 5-part series on fine-tuning sparse embeddings for e-commerce search. We'll go from "why bother?" to a production system that beats BM25 by 29%.*
**Series:**
- Part 1: Why Sparse Embeddings Beat BM25 (here)
- [Part 2: Training on Modal](/articles/sparse-embeddings-ecommerce-part-2/)
- [Part 3: Evaluation & Hard Negatives](/articles/sparse-embeddings-ecommerce-part-3/)
- [Part 4: Specialization vs Generalization](/articles/sparse-embeddings-ecommerce-part-4/)
- [Part 5: From Research to Product](/articles/sparse-embeddings-ecommerce-part-5/)
---
Search "iPhone 15 Pro Max 256GB" on a dense embedding system and it happily returns the 128GB model. The semantic similarity is high - it's the same phone! But the customer specified 256GB for a reason. In e-commerce, the details aren't noise. They're the whole point.
![Dense embedding search returns the wrong iPhone storage variant](/articles_data/sparse-embeddings-ecommerce-part-1/wrong-iphone-result.png)
This is the gap that sparse embeddings fill. And with fine-tuning, they fill it dramatically well - we achieved a **29% improvement over BM25** on Amazon's ESCI dataset, one of the largest public e-commerce search benchmarks.
In this series, we'll build the entire system: data loading, GPU training on Modal, evaluation with Qdrant, and hard negative mining. The [full code is on GitHub](https://github.com/thierrypdamiba/finetune-ecommerce-search) and the [fine-tuned models are on HuggingFace](https://huggingface.co/thierrydamiba/splade-ecommerce-esci). If you want to skip the walkthrough and fine-tune on your own data, the [`sparse-finetune`](https://github.com/qdrant/sparse-finetune) CLI runs the entire pipeline with one command. But first, let's understand why sparse embeddings are the right tool for e-commerce search.
## The Problem with Dense Embeddings in E-Commerce
Dense embeddings (the kind you get from [OpenAI](https://platform.openai.com/docs/guides/embeddings), [Cohere](https://docs.cohere.com/docs/embeddings), or a fine-tuned [sentence transformer](https://www.sbert.net/)) compress text into a fixed-size vector, typically 384 to 1536 dimensions, all non-zero. They're excellent at capturing semantic meaning. "Running shoes" and "jogging sneakers" land close together in the embedding space.
But this strength becomes a weakness in e-commerce:
**Exact matches get blurred.** When every dimension carries a non-zero value, the model prioritizes broad semantic similarity over exact term matching. SKU numbers, model names, specific sizes - these critical differentiators get averaged into the same neighborhood as similar-but-wrong products.
**Retrieval is approximate.** At scale, dense vectors require Approximate Nearest Neighbor (ANN) indexes like HNSW. The "approximate" part means you're trading recall for speed. For search, where missing a relevant product means a lost sale, this tradeoff hurts.
**Results are opaque.** Why did product X rank above product Y? With dense embeddings, you can't say. The 768-dimensional vector offers no interpretability. When a merchandising team asks why a product isn't showing up, you're stuck.
## Enter Sparse Embeddings
[Sparse embeddings](https://qdrant.tech/articles/sparse-vectors/) take a fundamentally different approach. Instead of compressing text into a small, dense vector, they project it onto a large vocabulary space - typically 30,000+ dimensions (one per token in the vocabulary). But only 100-300 of those dimensions are non-zero.
| | Dense | Sparse |
|---|---|---|
| **Vector size** | 384-1536 dims | ~30,000 dims |
| **Non-zero values** | All dimensions carry a value | 100-300 terms |
| **Index type** | ANN (HNSW) | Inverted index |
| **Exact matching** | Weak | Strong |
| **Interpretability** | Black box | Per-term weights |
![Visualization comparing dense and sparse vector representations](/articles_data/sparse-embeddings-ecommerce-part-1/dense-vs-sparse-viz.png)
Both approaches encode text into vectors, but sparse embeddings preserve individual term signals that dense models compress away.
The key difference: each dimension in a sparse vector corresponds to an actual word in the vocabulary. You can inspect the vector and see exactly which terms the model considers important and how much weight it gives each one.
## SPLADE: Learned Sparse Representations
SPLADE (Sparse Lexical and Expansion) is the model architecture that makes this work. It passes text through a transformer with a [masked language model](https://huggingface.co/docs/transformers/tasks/masked_language_modeling) (MLM) head, then applies max pooling and log saturation to produce sparse weights:
For an input like `"noise canceling headphones"`, SPLADE encodes it in four steps:
1. **Tokenize and encode** the input through DistilBERT with a masked language model (MLM) head
2. **Max pool** across all token positions to get a single score per vocabulary term
3. **Apply log saturation** — `log(1 + ReLU(x))` — a learned version of BM25's saturation curve that prevents any single term from dominating
4. **Output a sparse vector** with ~200 non-zero values out of 30,522 vocabulary dimensions
![The SPLADE encoding pipeline from input text to sparse vector](/articles_data/sparse-embeddings-ecommerce-part-1/splade-pipeline.png)
| Token | Weight |
|---|---|
| headphones | 2.3 |
| noise | 1.9 |
| canceling | 1.7 |
| audio | 1.2 |
| wireless | 0.8 |
| sound | 0.6 |
The **log saturation** step is important. Without it, a single high-confidence term could dominate the score. The log compression keeps results balanced - "headphones" matters more than "audio", but not 10x more.
The model learns three things simultaneously:
1. **Weight important terms** higher (product names, key attributes)
2. **Expand queries** with related terms (implicit synonym expansion)
3. **Suppress noise** (common words get near-zero weights)
### Query Expansion: Why SPLADE Beats BM25
This expansion is what separates SPLADE from traditional keyword search. BM25 can only match terms that literally appear in both the query and the document. SPLADE adds related terms that the model learned from training data:
**Query: "summer dress"**
Original terms: `dress` (2.5), `summer` (2.1)
Expanded by SPLADE: `sundress` (1.8), `floral` (0.9), `lightweight` (0.7), `cotton` (0.6)
The model adds "sundress", "floral", and "cotton" - terms that appear in product titles even when "summer" doesn't. This matches products like *"Floral Sundress for Women - Lightweight Cotton"* that BM25 would miss entirely.
No manual synonym file. No query rewriting rules. The model learned these associations from seeing millions of query-product pairs.
## Why Qdrant for Sparse Vectors?
Not every vector database treats sparse vectors as a first-class citizen. Qdrant does, and the difference matters in practice.
**Weighted sparse vectors.** SPLADE emits arbitrary learned weights; Qdrant stores them natively. If you're running a BM25 baseline, you can still add IDF at query time:
```python
sparse_vectors_config={
"bm25": models.SparseVectorParams(
modifier=models.Modifier.IDF, # Apply IDF at query time
)
}
```
**Hybrid in one request.** Combine sparse precision with dense semantics via native [RRF/prefetch](https://qdrant.tech/documentation/concepts/hybrid-queries/) - no external reranker:
```python
client.query_points(
collection_name="products",
prefetch=[
models.Prefetch(query=sparse_vector, using="sparse", limit=100),
models.Prefetch(query=dense_vector, using="dense", limit=100),
],
query=models.FusionQuery(fusion=models.Fusion.RRF),
limit=10,
)
```
**Production-ready scaling.** Rust + [SIMD-optimized inverted index](https://qdrant.tech/articles/sparse-vectors/) with an on-disk option keeps RAM low even with 200+ active terms per doc across millions of products.
**No ANN approximation.** Sparse retrieval uses an [inverted index](https://qdrant.tech/articles/sparse-vectors/), the same data structure powering BM25. Results are exact - no recall tradeoffs from approximate nearest neighbor search.
## The Stack
Our training pipeline combines three components:
**[Modal](https://modal.com/)** (GPU Training)
- A100 GPUs on demand
- Persistent volumes for checkpoints
- Detached runs for long training
**[Sentence Transformers v5](https://www.sbert.net/)** (Training Framework)
- SparseEncoder architecture
- SpladeLoss with regularization
- Built-in training utilities
**[Qdrant](https://qdrant.tech/)** (Sparse Vector Store)
- Native sparse vector support
- Inverted index
- Hybrid search ready
Modal gives us serverless A100 GPUs - no idle hardware, no queue management. Sentence Transformers v5 introduced the `SparseEncoder` class that makes SPLADE training straightforward. And Qdrant handles storage, indexing, and retrieval with native sparse vector support.
## What We'll Build
Over the next four articles, we'll walk through the full pipeline:
- [**Part 2: Training on Modal**](/articles/sparse-embeddings-ecommerce-part-2/) - Loading the Amazon ESCI dataset, creating the SPLADE model, configuring loss functions with sparsity regularization, and running GPU training with persistent checkpoints.
- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE hard negative mining loops, and analyzing what fine-tuning actually changes in the model.
- [**Part 4: Specialization vs Generalization**](/articles/sparse-embeddings-ecommerce-part-4/) - Cross-domain evaluation on Wayfair and Home Depot data, multi-domain training, when to specialize vs generalize, and production deployment guidance.
- [**Part 5: From Research to Product**](/articles/sparse-embeddings-ecommerce-part-5/) - An open-source CLI and web dashboard that runs the entire fine-tuning pipeline with a single command.
The end result: a fine-tuned SPLADE model that achieves **nDCG@10 of 0.388** on Amazon ESCI, compared to **0.301** for BM25 and **0.324** for off-the-shelf SPLADE. That 29% improvement over BM25 translates to meaningfully better search results for real e-commerce queries. You can try the models directly from HuggingFace: [splade-ecommerce-esci](https://huggingface.co/thierrydamiba/splade-ecommerce-esci) (best in-domain) and [splade-ecommerce-multidomain](https://huggingface.co/thierrydamiba/splade-ecommerce-multidomain) (better generalization).
@@ -0,0 +1,375 @@
---
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 2: Training SPLADE on Modal"
short_description: "Train a SPLADE model on Amazon's ESCI dataset using Modal's serverless GPUs and Sentence Transformers."
description: "Part 2 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Build a training pipeline on Modal with persistent checkpoints, SpladeLoss, and hyperparameter sweeps."
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-2/preview
social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-2/preview/social_preview.jpg
weight: -199
author: Thierry Damiba
author_link: https://github.com/thierrydamiba
date: 2026-03-09T00:00:00.000Z
category: practicle-examples
---
*This is Part 2 of a 5-part series on fine-tuning sparse embeddings for e-commerce search. In [Part 1](/articles/sparse-embeddings-ecommerce-part-1/), we covered why sparse embeddings beat BM25 for e-commerce. Now we build the training pipeline.*
**Series:**
- [Part 1: Why Sparse Embeddings Beat BM25](/articles/sparse-embeddings-ecommerce-part-1/)
- Part 2: Training SPLADE on Modal (here)
- [Part 3: Evaluation & Hard Negatives](/articles/sparse-embeddings-ecommerce-part-3/)
- [Part 4: Specialization vs Generalization](/articles/sparse-embeddings-ecommerce-part-4/)
- [Part 5: From Research to Product](/articles/sparse-embeddings-ecommerce-part-5/)
---
In the last article we made the case for sparse embeddings in e-commerce search. Now we write the code. All source code is available in the [GitHub repo](https://github.com/thierrypdamiba/finetune-ecommerce-search), and you can try the [fine-tuned models on HuggingFace](https://huggingface.co/thierrydamiba/splade-ecommerce-esci). Want to skip straight to fine-tuning on your own data? See the [`sparse-finetune`](https://github.com/qdrant/sparse-finetune) CLI. By the end of this piece, you'll have a SPLADE model trained on Amazon's ESCI dataset, running on Modal's serverless GPUs, with checkpoints saved to persistent storage.
## The Dataset: Amazon ESCI
We use Amazon's [ESCI dataset](https://github.com/amazon-science/esci-data) (Shopping Queries Dataset), released for KDD Cup 2022. It's one of the most realistic e-commerce search benchmarks available:
- **1.2M+ query-product pairs** with human-annotated relevance labels
- **Four relevance grades**: Exact (E), Substitute (S), Complement (C), Irrelevant (I)
- **Rich product metadata**: titles, descriptions, bullet points, brands
The graded relevance is what makes ESCI interesting:
![ESCI relevance gradient from Exact to Irrelevant](/articles_data/sparse-embeddings-ecommerce-part-2/esci-relevance-gradient.png)
For training, we use Exact and Substitute pairs as positives. This teaches the model that both the exact product and reasonable alternatives are relevant, matching how real shoppers think.
### Loading the Data
```python
from datasets import load_dataset
from src.data.text_builder import build_product_text
def load_esci_training_data(max_samples=None):
"""Load ESCI dataset as anchor-positive pairs for contrastive training."""
dataset = load_dataset("tasksource/esci", split="train")
pairs = []
for row in dataset:
if row["relevance_label"] not in ("E", "S"):
continue
query = row["query"]
product_text = build_product_text(
title=row["product_title"],
brand=row.get("product_brand", ""),
description=row.get("product_description", ""),
bullets=row.get("product_bullet_point", []),
)
pairs.append({"anchor": query, "positive": product_text})
if max_samples and len(pairs) >= max_samples:
break
return pairs
```
### Product Text Formatting
How you format product text matters for sparse embeddings. Unlike dense models that capture broad semantic meaning, SPLADE is lexically grounded: the specific tokens in your text determine which vocabulary dimensions activate:
```python
def build_product_text(title, brand="", description="", bullets=None, max_length=512):
"""Consistent product text formatting for SPLADE."""
parts = []
# Brand in brackets makes it a distinct signal
if brand:
parts.append(f"[{brand}]")
parts.append(title)
# Pipe separators help the model distinguish sections
if description:
parts.append(f"| {description[:200]}")
if bullets:
parts.append(f"| {' | '.join(bullets[:3])}")
text = " ".join(parts)
return text[:max_length]
# Example output:
# "[Sony] WH-1000XM5 Wireless Headphones | Industry-leading noise
# cancellation | 30hr battery | Hi-Res Audio"
```
The bracket notation for brands, pipe separators between sections, and character limits are deliberate. They preserve lexical signals that SPLADE can learn from: brand names, product attributes, and key features remain as distinct tokens rather than blurring into a wall of text.
![Training stack: Modal for GPU compute, Sentence Transformers for training, Qdrant for evaluation](/articles_data/sparse-embeddings-ecommerce-part-2/training-stack.png)
## Setting Up the Modal App
Modal gives us serverless GPUs. No provisioning, no idle hardware, pay-per-second billing. Here's the app configuration:
```python
import modal
app = modal.App("esci-sparse-encoder")
# Persistent storage for checkpoints and datasets
checkpoint_volume = modal.Volume.from_name(
"esci-sparse-checkpoints", create_if_missing=True
)
dataset_volume = modal.Volume.from_name(
"esci-datasets", create_if_missing=True
)
# Docker image with dependencies
image = (
modal.Image.debian_slim(python_version="3.11")
.pip_install(
"sentence-transformers>=5.0.0",
"torch>=2.2.0",
"transformers>=4.45.0",
"datasets>=2.20.0",
"qdrant-client>=1.12.0",
"accelerate>=0.30.0",
)
)
```
Two things matter here:
**Persistent volumes.** Training runs can take hours. If your SSH connection drops or a container restarts, you don't want to lose checkpoints. Modal volumes persist data across runs. Mount them at a path and write to them like a local filesystem.
**Detached runs.** For long training jobs, launch with `--detach` and walk away:
```bash
# Start training and disconnect
uv run modal run --detach modal_app.py --mode train
# Come back later, check your checkpoints
uv run modal volume ls esci-sparse-checkpoints /checkpoints/
```
No S3 uploads, no checkpoint management code, no lost training runs.
## Creating the SPLADE Model
Sentence Transformers v5 introduced `SparseEncoder`, making SPLADE training straightforward. The model has two components:
1. **MLMTransformer**: A transformer with a masked language model head that outputs logits over the full vocabulary
2. **SpladePooling**: Max-pools the token-level logits and applies ReLU + log saturation
```python
from sentence_transformers import SparseEncoder
from sentence_transformers.sparse_encoder.models import (
MLMTransformer,
SpladePooling,
)
def create_sparse_encoder(base_model="distilbert/distilbert-base-uncased"):
"""Create a SPLADE model from a base transformer."""
# MLM transformer outputs logits over vocabulary
mlm = MLMTransformer(base_model)
# SPLADE pooling: max over tokens, ReLU activation
pooling = SpladePooling(pooling_strategy="max")
return SparseEncoder(modules=[mlm, pooling])
```
We start from DistilBERT rather than a pre-trained SPLADE checkpoint (like `naver/splade-v3`). This is a deliberate choice. We want to measure how much domain-specific fine-tuning helps when starting from a general language model, not from a model already trained on web search data.
## The Training Function
Here's the core training logic, decorated as a Modal function:
```python
@app.function(
image=image,
gpu="A100",
volumes={
"/checkpoints": checkpoint_volume,
"/datasets": dataset_volume,
},
timeout=3600 * 6,
)
def train_sparse_encoder(config: dict):
from sentence_transformers import SparseEncoder
from sentence_transformers.sparse_encoder import SparseEncoderTrainer
from sentence_transformers.training_args import SparseEncoderTrainingArguments
from sentence_transformers.losses import SpladeLoss, SparseMultipleNegativesRankingLoss
# Create model
model = create_sparse_encoder(config["base_model"])
# Load ESCI dataset (anchor-positive pairs)
train_dataset = load_esci_training_data(
max_samples=config.get("max_samples")
)
# SPLADE loss combines contrastive learning with sparsity regularization
loss = SpladeLoss(
model=model,
loss=SparseMultipleNegativesRankingLoss(model=model),
query_regularizer_weight=float(config.get("query_regularizer_weight", 5e-5)),
document_regularizer_weight=float(config.get("document_regularizer_weight", 3e-5)),
)
# Training arguments
args = SparseEncoderTrainingArguments(
output_dir=f"/checkpoints/{config['run_name']}",
num_train_epochs=config.get("num_epochs", 1),
per_device_train_batch_size=config.get("batch_size", 32),
learning_rate=float(config.get("learning_rate", 2e-5)),
warmup_ratio=0.1,
fp16=True,
save_steps=1000,
logging_steps=100,
)
# Train
trainer = SparseEncoderTrainer(
model=model,
args=args,
train_dataset=train_dataset,
loss=loss,
)
trainer.train()
# Save final model
model.save_pretrained(f"/checkpoints/{config['run_name']}/final")
return f"/checkpoints/{config['run_name']}/final"
```
### Understanding SpladeLoss
`SpladeLoss` wraps two objectives:
**Contrastive loss** (`SparseMultipleNegativesRankingLoss`): Given a batch of (query, product) pairs, treat other products in the batch as negatives. Push relevant query-product pairs together, push irrelevant ones apart. This is the same in-batch negative approach used for dense embedding training, and it works because most random products are irrelevant to a given query.
**Sparsity regularization**: Penalizes dense outputs to maintain efficiency. Without it, the model would activate all 30,000 vocabulary dimensions for every input. That's technically optimal for matching but useless for retrieval speed and storage.
The regularization weights control this tradeoff:
| Parameter | Value | Effect |
|---|---|---|
| `query_regularizer_weight` | 5e-5 | Higher = sparser queries |
| `document_regularizer_weight` | 3e-5 | Higher = sparser documents |
The sweet spot is 100-300 active terms per vector. Too high regularization produces nearly empty vectors (fast but low recall). Too low produces thousands of terms (slow, huge index).
Document regularization is lower than query regularization because product descriptions need more terms to capture all relevant attributes. A product listing for headphones should activate terms like "audio", "wireless", "bluetooth", "noise", "canceling" - more than the 3-4 words in a typical query.
### Configuration via YAML
We keep hyperparameters in YAML files for easy experimentation:
```yaml
# configs/splade_standard.yaml
run_name: splade_standard
base_model: distilbert/distilbert-base-uncased
architecture: splade
batch_size: 32
learning_rate: 2e-5
num_epochs: 1
query_regularizer_weight: 5e-5
document_regularizer_weight: 3e-5
max_samples: 100000
```
100K samples trains in about 6 minutes on an A100 and costs less than $1 on Modal. The full 1.2M dataset with multiple epochs takes a few hours, still cheap compared to reserved GPU instances.
## Parallel Hyperparameter Sweeps
One of Modal's strengths is embarrassingly parallel workloads. Hyperparameter sweeps are a natural fit. `spawn()` launches one GPU per configuration:
```python
@app.function(gpu="A100")
def train_single_experiment(config: dict):
"""Train one configuration."""
model = create_sparse_encoder(config["base_model"])
# ... training code ...
return {"config": config, "ndcg": evaluate(model)}
@app.local_entrypoint()
def run_hyperparameter_sweep():
"""Launch all experiments in parallel."""
configs = [
{"learning_rate": 1e-5, "regularizer_weight": 3e-5},
{"learning_rate": 2e-5, "regularizer_weight": 3e-5},
{"learning_rate": 2e-5, "regularizer_weight": 5e-5},
{"learning_rate": 5e-5, "regularizer_weight": 5e-5},
# ... more configurations ...
]
# Launch all experiments simultaneously
handles = [train_single_experiment.spawn(c) for c in configs]
# Collect results as they complete
results = [h.get() for h in handles]
best = max(results, key=lambda r: r["ndcg"])
print(f"Best config: {best}")
```
A 24-experiment sweep finishes in the time of a single training run. Each experiment gets its own A100. You pay only for the compute time actually used, not for idle GPUs waiting in a queue.
## What NOT to Do: The Inference-Free SPLADE Trap
We tried replacing the query-side transformer with a static embedding lookup to save latency. The idea is appealing: queries are short, so why run a full transformer?
```python
# DON'T DO THIS (for e-commerce)
router = Router.for_query_document(
query_modules=[
SparseStaticEmbedding(tokenizer=mlm.tokenizer) # Fast but weak
],
document_modules=[
mlm,
SpladePooling(pooling_strategy="max"),
],
)
```
The results were disastrous:
| Architecture | nDCG@10 |
|---|---|
| Standard SPLADE (contextual) | **0.389** |
| Inference-Free (static) | 0.065 |
That's 6x worse without contextual encoding.
The static embedding completely failed because e-commerce queries are highly contextual. "Apple" means different things in "apple iphone" vs "apple fruit". The static embedding can't disambiguate. It looks up "apple" and returns the same vector regardless of context.
The transformer is the bottleneck at ~15ms per query, but 15ms is perfectly acceptable for search. Don't prematurely optimize away the component that makes the model work.
![Modal detached training with persistent volumes](/articles_data/sparse-embeddings-ecommerce-part-2/modal-detached-training.png)
## Running Training
With everything in place, launch training:
```bash
# Quick test run (100K samples)
uv run modal run modal_app.py \
--config-path configs/splade_standard.yaml \
--mode train
# Full dataset, detached
uv run modal run --detach modal_app.py \
--config-path configs/splade_standard.yaml \
--mode train
```
The model checkpoint gets saved to the persistent volume at `/checkpoints/splade_standard/final`. We've also published the trained model on HuggingFace as [splade-ecommerce-esci](https://huggingface.co/thierrydamiba/splade-ecommerce-esci) so you can skip training and use it directly. In the next article, we'll load this model, index products into Qdrant, and run retrieval benchmarks to see exactly how much we've improved over BM25.
## Key Takeaways
- **ESCI's graded relevance** (Exact, Substitute, Complement, Irrelevant) teaches the model nuanced matching, not just binary relevant/not-relevant.
- **Product text formatting matters** for sparse models. Keep lexical signals distinct with structured formatting.
- **SpladeLoss balances two objectives**: contrastive learning for relevance and regularization for sparsity. The regularization weights are the main knob to tune.
- **Modal's persistent volumes** solve the checkpoint management problem. Detached runs survive SSH drops.
- **Don't skip the query transformer.** The 15ms of latency buys you a 6x quality improvement over static embeddings.
---
*Next: [Part 3 - Evaluation, Hard Negatives, and Results](/articles/sparse-embeddings-ecommerce-part-3/)*
@@ -0,0 +1,261 @@
---
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 3: Evaluation and Hard Negatives"
short_description: "Evaluate fine-tuned SPLADE with Qdrant and boost results with hard negative mining."
description: "Part 3 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Index products in Qdrant, run retrieval benchmarks, and implement ANCE hard negative mining for a 28% improvement over BM25."
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-3/preview
social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-3/preview/social_preview.jpg
weight: -198
author: Thierry Damiba
author_link: https://github.com/thierrydamiba
date: 2026-03-09T00:00:00.000Z
category: practicle-examples
---
*This is Part 3 of a 5-part series on fine-tuning sparse embeddings for e-commerce search. In [Part 2](/articles/sparse-embeddings-ecommerce-part-2/), we trained a SPLADE model on Modal. Now we evaluate it and push further with hard negative mining.*
**Series:**
- [Part 1: Why Sparse Embeddings Beat BM25](/articles/sparse-embeddings-ecommerce-part-1/)
- [Part 2: Training SPLADE on Modal](/articles/sparse-embeddings-ecommerce-part-2/)
- Part 3: Evaluation & Hard Negatives (here)
- [Part 4: Specialization vs Generalization](/articles/sparse-embeddings-ecommerce-part-4/)
- [Part 5: From Research to Product](/articles/sparse-embeddings-ecommerce-part-5/)
---
We have a trained SPLADE model sitting on a Modal volume (or grab it from [HuggingFace](https://huggingface.co/thierrydamiba/splade-ecommerce-esci)). Now comes the question that matters: is it actually better? In this article, we'll index products into Qdrant, run retrieval benchmarks, implement hard negative mining, and dig into what the model learned. Full evaluation code is in the [GitHub repo](https://github.com/thierrypdamiba/finetune-ecommerce-search). To run this entire pipeline on your own data, see the [`sparse-finetune`](https://github.com/qdrant/sparse-finetune) CLI.
## Indexing Products in Qdrant
Before we can evaluate, we need products in a searchable index. Qdrant's sparse vector support makes this straightforward:
```python
from qdrant_client import QdrantClient, models
def index_products(model, products, collection_name="ecommerce_splade"):
client = QdrantClient(url=QDRANT_URL, api_key=QDRANT_API_KEY)
# Create collection with sparse vector config
client.create_collection(
collection_name=collection_name,
vectors_config={},
sparse_vectors_config={
"text": models.SparseVectorParams(
index=models.SparseIndexParams(on_disk=True)
)
},
)
# Encode and index in batches
for batch in chunked(products, batch_size=32):
texts = [p["text"] for p in batch]
embeddings = model.encode(texts)
points = []
for product, emb in zip(batch, embeddings):
indices = emb["indices"].tolist()
values = emb["values"].tolist()
points.append(models.PointStruct(
id=product["id"],
vector={
"text": models.SparseVector(indices=indices, values=values)
},
payload={"title": product["title"], "brand": product["brand"]},
))
upsert_with_retry(client, collection_name, points)
```
A few production details:
- **`on_disk=True`** keeps the inverted index on disk instead of RAM. SPLADE vectors average 200 active terms, and across millions of products, this adds up. Requires SSD for acceptable latency.
- **`wait=False`** on upserts (inside `upsert_with_retry`) lets you pipeline batches without blocking. Call with `wait=True` on the final batch.
- **Retry with exponential backoff** for cloud databases. Network hiccups happen in production.
```python
def upsert_with_retry(client, collection_name, points, max_retries=5):
"""Upsert with exponential backoff."""
for attempt in range(max_retries):
try:
client.upsert(collection_name=collection_name, points=points, wait=False)
return
except Exception as e:
if attempt == max_retries - 1:
raise
wait_time = (2 ** attempt) + (attempt * 0.5)
time.sleep(wait_time)
```
## Retrieval Metrics
We evaluate with standard information retrieval metrics on 2,000 test queries against 10,000 products:
- **nDCG@k** — Ranking quality with position bias; top results matter more
- **MRR@k** — How high the first relevant result appears
- **Recall@k** — What fraction of relevant products appear in top-k
- **Precision@k** — What fraction of top-k results are relevant
**nDCG@10** (Normalized Discounted Cumulative Gain) is the primary metric. It rewards putting highly relevant products (Exact matches) at the top and penalizes relevant results that appear lower in the ranking. A perfect score is 1.0; random ranking on this dataset gives roughly 0.1.
### Searching the Index
```python
def search_products(query, model, client, collection_name="ecommerce_splade", limit=10):
query_embedding = model.encode(query)
results = client.query_points(
collection_name=collection_name,
query=models.SparseVector(
indices=query_embedding["indices"].tolist(),
values=query_embedding["values"].tolist(),
),
using="text",
limit=limit,
)
return [
{"id": r.id, "score": r.score, "title": r.payload["title"]}
for r in results.points
]
```
Five lines from query string to ranked products. The sparse vector lookup in Qdrant's inverted index is sub-millisecond, even with millions of products. The bottleneck is the 10-20ms query encoding through the transformer.
## The Results
Here's what we found, evaluated on 2,000 test queries:
| Model | nDCG@10 | MRR@10 | vs BM25 |
|---|---|---|---|
| BM25 (baseline) | 0.305 | 0.313 | - |
| SPLADE (off-the-shelf) | 0.326 | 0.339 | +7.2% |
| **SPLADE (fine-tuned)** | **0.389** | **0.387** | **+27.5%** |
The fine-tuned model beats BM25 by nearly 28%. More telling: it beats the off-the-shelf SPLADE by 19%. The off-the-shelf model was trained on MS MARCO (web search queries), not e-commerce. That 19% gap is the value of domain-specific training.
### What About Hybrid Search?
![Hybrid search fusion combining sparse and dense retrieval](/articles_data/sparse-embeddings-ecommerce-part-3/hybrid-search-fusion.png)
A natural question: can we combine sparse and dense vectors for even better results? We tested this with Qdrant's native Reciprocal Rank Fusion:
```python
client.query_points(
collection_name="products",
prefetch=[
models.Prefetch(query=sparse_vector, using="sparse", limit=100),
models.Prefetch(query=dense_vector, using="dense", limit=100),
],
query=models.FusionQuery(fusion=models.Fusion.RRF),
limit=10,
)
```
With the **off-the-shelf SPLADE**, hybrid helps: +1.3% over sparse alone. Both signals are moderate strength, and combining them catches products that either one misses.
With the **fine-tuned SPLADE**, hybrid actually hurts: SPLADE-only scored 0.413 vs hybrid at 0.405. The fine-tuned sparse model is strong enough that adding a generic dense signal dilutes the ranking. The dense model retrieves semantically similar but irrelevant products that drag down nDCG.
This is a useful finding. Hybrid search isn't always better. It depends on the relative strength of your signals. If your sparse model is domain-tuned and your dense model is generic, the dense component can actively harm results.
## Hard Negative Mining with ANCE
![The ANCE hard negative mining loop](/articles_data/sparse-embeddings-ecommerce-part-3/ance-loop.png)
The training in Part 2 used in-batch negatives: other products in the same batch serve as negatives for a given query. This works but has a limitation: random products are easy negatives. The model doesn't learn to distinguish between genuinely confusable products.
[ANCE](https://www.sbert.net/examples/training/quora_duplicate_questions/README.html) (Approximate Nearest Neighbor Negative Contrastive Estimation) fixes this by mining hard negatives from the current model's own retrieval results:
1. **Index** products into Qdrant with the current model
2. **Retrieve** top-K products for each query
3. **Filter** to non-relevant products — these are the hard negatives
4. **Train** on (query, positive, hard_negatives) triplets
5. **Repeat** with the updated model
Each round mines harder negatives as the model improves.
The idea: if the current model retrieves a product for a query but that product isn't relevant, it's a hard negative. The model thought it was relevant, so training on it teaches the model where its mistakes are.
### Mining Implementation
```python
from src.qdrant.mining import SparseQdrantMiner
# Index products with current model
index_sparse_vectors(client, collection_name, model, products)
# Mine hard negatives
miner = SparseQdrantMiner(client, model, collection_name)
hard_neg_examples = miner.mine_for_training(
queries=queries_with_positives,
top_k=20, # Consider top-20 results
num_negatives=3, # Keep 3 hardest negatives per query
)
# hard_neg_examples now contains:
# [{"anchor": "wireless earbuds",
# "positive": "Sony WF-1000XM5 Earbuds...",
# "negative": ["Generic Bluetooth Earbuds...", ...]}, ...]
```
Sparse retrieval keeps mining cheap, with sub-millisecond per query in Qdrant. For 100K queries, the mining step takes seconds, not minutes. Payload filters exclude known positives so you don't accidentally treat a relevant product as a negative.
### When to Use ANCE
ANCE adds complexity. You need to:
1. Index products with the current model
2. Run retrieval for all training queries
3. Filter and format the results
4. Retrain with the augmented dataset
5. Optionally repeat
This gives an additional 5-10% improvement on top of basic training. Whether that's worth the engineering effort depends on your use case. For a product search system serving millions of queries, 5% nDCG improvement translates to meaningfully better user experience and conversion rates.
## What Fine-Tuning Actually Changes
Looking at the model's outputs before and after fine-tuning reveals what it learned:
**Query expansion improves:**
- "laptop" → adds "notebook", "computer", "macbook"
- "wireless earbuds" → adds "bluetooth", "airpods", "tws"
**Term weighting sharpens:**
- Brand names get higher weights (users searching "Sony headphones" want Sony)
- Generic terms get lower weights ("good", "best", "cheap")
**Domain vocabulary emerges:**
- E-commerce terms like "refurbished", "renewed", "bundle" get meaningful weights
- Web-search-specific terms get downweighted
This domain adaptation explains both the strong in-domain results and, as we'll see in Part 4, the tradeoffs when applying the model to other domains.
## Production Latency
A common concern: isn't running a transformer on every query slow?
| Step | Latency | Note |
|---|---|---|
| Query encoding (SPLADE) | 10-20ms | Bottleneck |
| Sparse retrieval (Qdrant) | <1ms | Negligible |
| **Total** | **10-20ms** | Real-time |
The retrieval itself is negligible. Qdrant's Rust + [SIMD-optimized inverted index](https://qdrant.tech/articles/sparse-vectors/) scans millions of posting lists in sub-millisecond time. All the latency is in the encoder, which runs once per query regardless of catalog size.
Optimization strategies if 15ms isn't fast enough:
- **Batch queries**: Encode multiple queries together (autocomplete, related searches)
- **Distillation**: Train a smaller encoder (TinyBERT, MiniLM) to mimic SPLADE's outputs
- **Caching**: Popular queries can be cached at the sparse vector level
- **GPU inference**: 5-10x speedup on high-traffic systems
For most e-commerce applications, 15ms is fine, especially when it delivers 28% better relevance.
## Key Takeaways
- **Fine-tuned SPLADE beats BM25 by 28% and off-the-shelf SPLADE by 19%.** Domain-specific training matters, even for sparse models.
- **Hybrid search isn't always better.** A strong domain-tuned sparse model can outperform sparse+dense fusion when the dense component is generic.
- **Hard negative mining (ANCE) adds 5-10%** on top of basic training. Qdrant's sparse retrieval makes the mining step cheap.
- **Production latency is 10-20ms total.** Transformer encoding is the bottleneck, not retrieval.
- **The model learns domain-specific patterns**: query expansion, term weighting, and e-commerce vocabulary all improve with fine-tuning.
---
*Next: [Part 4 - Specialization vs Generalization](/articles/sparse-embeddings-ecommerce-part-4/)*
@@ -0,0 +1,186 @@
---
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 4: Specialization vs Generalization"
short_description: "When to fine-tune sparse embeddings and how far to specialize before generalization suffers."
description: "Part 4 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Test cross-domain generalization, train a multi-domain model, and decide when to specialize vs generalize."
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-4/preview
social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-4/preview/social_preview.jpg
weight: -197
author: Thierry Damiba
author_link: https://github.com/thierrydamiba
date: 2026-03-09T00:00:00.000Z
category: practicle-examples
---
*This is Part 4 of a 5-part series on fine-tuning sparse embeddings for e-commerce search. In [Part 3](/articles/sparse-embeddings-ecommerce-part-3/), we evaluated our model and implemented hard negative mining. Now we test how well it generalizes.*
**Series:**
- [Part 1: Why Sparse Embeddings Beat BM25](/articles/sparse-embeddings-ecommerce-part-1/)
- [Part 2: Training SPLADE on Modal](/articles/sparse-embeddings-ecommerce-part-2/)
- [Part 3: Evaluation & Hard Negatives](/articles/sparse-embeddings-ecommerce-part-3/)
- Part 4: Specialization vs Generalization (here)
- [Part 5: From Research to Product](/articles/sparse-embeddings-ecommerce-part-5/)
---
We've built a SPLADE model that beats BM25 by 28% on Amazon ESCI. But here's the question that determines whether this is a lab result or a production strategy: does it work on data it wasn't trained on? Full code is on [GitHub](https://github.com/thierrypdamiba/finetune-ecommerce-search), you can try the [fine-tuned models on HuggingFace](https://huggingface.co/thierrydamiba/splade-ecommerce-esci), or fine-tune on your own catalog with the [`sparse-finetune`](https://github.com/qdrant/sparse-finetune) CLI.
In this final article, we test cross-domain generalization, train a multi-domain model, and lay out a decision framework for when to specialize vs generalize.
## Cross-Domain Evaluation
![Cross-domain nDCG comparison across datasets](/articles_data/sparse-embeddings-ecommerce-part-4/cross-domain-ndcg.png)
We took our Amazon ESCI-trained model and tested it on three additional datasets:
- **WANDS** (Wayfair): Furniture and home goods search
- **Home Depot**: Hardware and home improvement search
- **MS MARCO**: General web search (the "out of distribution" control)
| Dataset | BM25 | SPLADE (OTS) | SPLADE (tuned) | vs BM25 |
|---|---|---|---|---|
| ESCI (Amazon) | 0.305 | 0.326 | **0.389** | +27.5% |
| WANDS (Wayfair) | 0.329 | 0.341 | **0.355** | +7.9% |
| Home Depot | 0.349 | **0.391** | 0.384* | +10.0% |
| MS MARCO (web) | 0.915 | 0.982 | 0.751 | -17.9% |
*On Home Depot, the off-the-shelf model edges out the fine-tuned one (0.391 vs 0.384).
Three patterns emerge:
**In-domain (ESCI): +28% over BM25.** The model was trained on this data. No surprise it does well.
**Cross-domain e-commerce: +8-10% over BM25.** The Amazon-trained model still helps on Wayfair and Home Depot. E-commerce search shares enough structure (brand matching, attribute weighting, product vocabulary) that the patterns transfer. But notice the gap to off-the-shelf SPLADE narrows. On Home Depot, the off-the-shelf model actually wins (0.391 vs 0.384).
**Out-of-domain (MS MARCO): -18% vs BM25.** This is catastrophic forgetting in action. The model overfitted to e-commerce patterns. "Apple" became a brand, not a fruit. "Prime" became a shipping speed, not a math concept. The general IR capabilities of the original DistilBERT were overwritten during fine-tuning.
## Why Generalization Degrades
![Transfer decay curve showing performance drop across domains](/articles_data/sparse-embeddings-ecommerce-part-4/transfer-decay-curve.png)
The cross-domain results reveal a fundamental tradeoff. Fine-tuning teaches the model:
- **Amazon-specific query patterns:** short, product-focused queries with brand names and model numbers
- **Amazon-specific vocabulary:** "renewed" (refurbished), "subscribe & save", "prime eligible"
- **Amazon-specific relevance signals:** what Amazon shoppers consider a good match vs a substitute
Wayfair customers search differently ("mid-century modern coffee table" vs "coffee table"). Home Depot customers use industry terminology ("3/8 inch drive socket set"). The Amazon-trained model helps on these datasets because e-commerce is e-commerce, but it's not optimal.
MS MARCO is the extreme case. Web search queries like "what is the capital of France" or "how to tie a tie" are nothing like e-commerce queries. The model's learned biases actively hurt.
## Multi-Domain Training
![Domain coverage Venn diagram showing overlap between e-commerce datasets](/articles_data/sparse-embeddings-ecommerce-part-4/domain-coverage-venn.png)
To address the generalization problem, we trained a **multi-domain SPLADE model** on combined data from ESCI, WANDS, and Home Depot: roughly 50K training pairs from each dataset, 150K total.
The hypothesis: exposure to diverse e-commerce catalogs should improve cross-domain transfer while maintaining reasonable in-domain performance.
| Dataset | ESCI-only | Multi-domain | Difference |
|---|---|---|---|
| ESCI | **0.389** | 0.372 | -4.4% |
| WANDS | 0.355 | **0.366** | +3.1% |
| Home Depot | 0.384 | **0.410** | +6.8% |
| MS MARCO | 0.751 | **0.829** | +10.4% |
Multi-domain training does exactly what you'd expect:
- **ESCI drops 4%**: Less specialization means less Amazon-specific optimization. The model can't memorize Amazon's vocabulary as deeply when it's also learning Wayfair and Home Depot patterns.
- **WANDS and Home Depot gain 3-7%**: Direct benefit from training data. The model now understands furniture terminology and hardware vocabulary.
- **MS MARCO recovers 10%**: More diverse training data prevents the catastrophic forgetting we saw with ESCI-only training. The model retains more general language understanding.
### Setting Up Multi-Domain Training
The multi-domain loader normalizes labels across datasets:
```yaml
# configs/splade_multidomain.yaml
run_name: splade_multidomain
base_model: distilbert/distilbert-base-uncased
architecture: splade
batch_size: 32
learning_rate: 2e-5
num_epochs: 1
datasets:
- name: esci
max_samples: 50000
- name: wands
max_samples: 50000
- name: homedepot
max_samples: 50000
```
Label normalization is the key challenge. ESCI uses character labels (E, S, C, I), WANDS uses numeric scores (0, 1, 2), and Home Depot uses relevance ratings. The multi-domain loader maps everything to a common format: positive (relevant) and negative (irrelevant) pairs for contrastive training.
## Decision Framework
![When to use specialist vs generalist models](/articles_data/sparse-embeddings-ecommerce-part-4/specialist-vs-generalist.png)
After running all these experiments, here's when to use each approach:
| Scenario | Recommended approach |
|---|---|
| Single retailer, lots of training data | **Domain-specific fine-tuning** — maximum performance on your catalog |
| Multi-retailer or marketplace | **Multi-domain training** — better generalization across catalogs |
| New domain, limited data | **Off-the-shelf SPLADE** — strong baseline without training data |
| Hybrid (e-commerce + general search) | **Multi-domain training** — preserves general IR capabilities |
**Single retailer with abundant data.** If you're building search for Amazon, Wayfair, or any single retailer with click logs, domain-specific fine-tuning wins. The 4% you lose on other domains doesn't matter if you only serve one catalog.
**Marketplace or multi-retailer.** If you're building a platform that serves multiple retailers (Shopify search, a price comparison engine), multi-domain training provides better balance. You sacrifice some peak performance for consistency across catalogs.
**Cold start.** New to a domain with no training data? Off-the-shelf SPLADE (like `naver/splade-v3`) is a strong baseline. It beats BM25 on most e-commerce datasets without any fine-tuning. Start here, collect click data, then fine-tune.
## The Case for Fine-Tuning
Why fine-tune when off-the-shelf models already beat BM25?
**Domain knowledge matters.** Generic models don't know that "AirPods Max" is a specific product, that "prime" means fast shipping, or that "organic" is a critical filter in grocery. Fine-tuning on your catalog teaches the model your vocabulary and your customers' search patterns.
**You control the training data.** Click logs, add-to-cart signals, and purchase data are unique to your business. Fine-tuning converts this proprietary data into a model that understands your domain better than any general-purpose model can.
**The model is portable.** Your fine-tuned model runs wherever you need it: Modal, your own GPUs, CPU inference, or any cloud provider. Deploy it however makes sense for your infrastructure.
**Performance compounds.** As we saw, domain-specific training delivers +28% over BM25. That's not a marginal improvement. It's the difference between showing a customer the right product on the first page or burying it on the third.
## The Data Flywheel
Fine-tuning isn't a one-time investment. It's the start of a compounding loop:
1. **Better model** leads to better rankings
2. **Better rankings** lead to more clicks
3. **More clicks** produce better training data
4. **Better training data** produces an even better model
5. Repeat
**Phase 1: Bootstrap.** Use product metadata and relevance labels (or the ESCI dataset as a proxy). Train the initial model. This is what we've done in this series.
**Phase 2: Implicit feedback.** Log queries with clicked products (positive pairs). Log impressions without clicks (negative signals). Track add-to-cart and purchase events (high-confidence positives).
**Phase 3: Continuous improvement.** Retrain periodically on accumulated click data. A/B test new models against production. Monitor nDCG on held-out queries.
The 28% improvement we demonstrated is the starting point. Each iteration incorporates what customers actually searched for and clicked on, data your competitors can't access.
## What's Next
We've covered the full pipeline: from understanding why sparse embeddings work for e-commerce, through training on Modal and evaluating with Qdrant, to the specialization-generalization tradeoff.
Extensions worth exploring:
- **Cross-encoder reranking**: Add a second-stage ranker for the top-k results from SPLADE. This is the standard two-stage retrieval architecture in production systems.
- **Larger base models**: ModernBERT or DeBERTa instead of DistilBERT. More parameters, better representations, slower inference.
- **Full dataset training**: We used 100K samples from ESCI. The full 1.2M with multiple epochs would likely improve results further.
- **Curriculum learning**: Start with general data, gradually specialize to your domain. This can mitigate catastrophic forgetting while still achieving strong in-domain performance.
The [code is open source](https://github.com/thierrypdamiba/finetune-ecommerce-search). The [pre-trained models are on HuggingFace](https://huggingface.co/thierrydamiba/splade-ecommerce-esci) (including a [multi-domain variant](https://huggingface.co/thierrydamiba/splade-ecommerce-multidomain)). Training runs on Modal for under $1. Qdrant handles the [sparse vectors](https://qdrant.tech/articles/sparse-vectors/), indexing, and retrieval out of the box. The barrier to building better e-commerce search has never been lower.
We also packaged this entire pipeline into an open-source toolkit with a CLI and web dashboard. See [Part 5: From Research to Product](/articles/sparse-embeddings-ecommerce-part-5/) for how to fine-tune a SPLADE model on your own catalog with a single command.
---
## Series Summary
- **[Part 1: Why sparse embeddings for e-commerce](/articles/sparse-embeddings-ecommerce-part-1/)** - SPLADE combines keyword precision with learned expansion
- **[Part 2: Training pipeline on Modal](/articles/sparse-embeddings-ecommerce-part-2/)** - 6 min training, <$1, persistent checkpoints
- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)** - +28% vs BM25, +19% vs off-the-shelf SPLADE
- **[Part 4: Specialization vs generalization](/articles/sparse-embeddings-ecommerce-part-4/)** - Domain-specific wins for single retailers; multi-domain for platforms
- **[Part 5: From research to product](/articles/sparse-embeddings-ecommerce-part-5/)** - CLI + dashboard that runs the full pipeline
@@ -0,0 +1,222 @@
---
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 5: From Research to Product"
short_description: "One command to fine-tune SPLADE for your catalog. No ML pipeline assembly required."
description: "Part 5 of the sparse embeddings series. We packaged the entire training pipeline from Parts 1-4 into an open-source CLI and web dashboard that fine-tunes SPLADE models for any product catalog in minutes."
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-5/preview
social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-5/preview/social_preview.jpg
weight: -196
author: Thierry Damiba
author_link: https://github.com/thierrydamiba
date: 2026-03-09T00:00:00.000Z
category: practicle-examples
---
*This is Part 5 of a series on fine-tuning sparse embeddings for e-commerce search. Parts [1](/articles/sparse-embeddings-ecommerce-part-1/)–[4](/articles/sparse-embeddings-ecommerce-part-4/) built the pipeline from scratch. This article packages it into a tool anyone can use.*
**Series:**
- [Part 1: Why Sparse Embeddings Beat BM25](/articles/sparse-embeddings-ecommerce-part-1/)
- [Part 2: Training SPLADE on Modal](/articles/sparse-embeddings-ecommerce-part-2/)
- [Part 3: Evaluation & Hard Negatives](/articles/sparse-embeddings-ecommerce-part-3/)
- [Part 4: Specialization vs Generalization](/articles/sparse-embeddings-ecommerce-part-4/)
- Part 5: From Research to Product (here)
---
In Parts 1 through 4, we built a SPLADE fine-tuning pipeline piece by piece: data loading, Modal GPU training, Qdrant evaluation, ANCE hard negative mining, cross-domain experiments. The code worked. The results were strong: 28% over BM25 on Amazon ESCI.
Using it required reading four articles, cloning a repo, understanding the training loop internals, wiring up Modal volumes, and configuring Qdrant connections manually. That's fine for a series walkthrough. It's not fine for someone who has a product catalog and wants a better search model by end of day.
So we packaged everything into [`qdrant-sparse-finetune`](https://github.com/qdrant/sparse-finetune): an open-source CLI and web dashboard that runs the entire pipeline (synthetic query generation, SPLADE training with ANCE, evaluation, and HuggingFace publishing) with a single command.
## The Problem We're Solving
![From research repo to production CLI](/articles_data/sparse-embeddings-ecommerce-part-5/research-to-production-pipeline.png)
The series repo ([finetune-ecommerce-search](https://github.com/thierrypdamiba/finetune-ecommerce-search)) is research code. It demonstrates how sparse embedding fine-tuning works. Actually using it on your data means you need to:
1. Format your product data to match the expected schema
2. Either provide labeled queries or set up an LLM API for synthetic generation
3. Configure Modal volumes and GPU settings
4. Wire up Qdrant credentials for indexing and mining
5. Run training, manually trigger ANCE iterations
6. Evaluate, interpret metrics
7. Publish to HuggingFace if you want to share the model
Each step has its own configuration, its own failure modes, and its own set of assumptions about how the previous step ran. It's a pipeline with no orchestration.
`qdrant-sparse-finetune` handles all of that:
```bash
pip install git+https://github.com/qdrant/sparse-finetune.git
qdrant-finetune setup
qdrant-finetune pipeline --data products.csv --gpu modal
```
Three commands. The setup wizard configures Qdrant, your LLM provider, and your GPU backend (Modal or Vultr). The pipeline command runs everything end-to-end. When training finishes, it asks if you want to publish to HuggingFace and gives you the link.
## Two Interfaces, Same Pipeline
### The CLI
For developers and automation:
```bash
# End-to-end on Modal (serverless A10G)
qdrant-finetune pipeline --data products.csv --gpu modal
# End-to-end on Vultr (dedicated cloud GPU)
qdrant-finetune pipeline --data products.csv --gpu vultr
# Or step by step
qdrant-finetune generate-queries --data products.csv --synth-model gpt-4o-mini
qdrant-finetune train --data products.csv --queries queries.jsonl --gpu modal
qdrant-finetune evaluate --model output/finetune/final --queries test_queries.jsonl
qdrant-finetune publish --model output/finetune/final --repo your-name/your-model
```
The `pipeline` command chains all four steps. If you already have labeled queries, pass `--queries` to skip generation. If you pass `--repo`, it publishes automatically. If you don't, it asks after training completes:
```
Training complete! Publish to HuggingFace Hub? [Y/n]: y
Repo name (e.g. your-username/my-splade-model): acme/product-search-splade
Private repo? [y/N]: n
Published → https://huggingface.co/acme/product-search-splade
```
No buried flags. The most valuable output, a shareable model URL, is the natural endpoint of the workflow.
### The Dashboard
For teams who prefer a visual interface:
```bash
qdrant-finetune studio
```
This launches a web dashboard with tabs for each stage of the pipeline:
**Train.** Configure the base model, ANCE iterations, batch size, and GPU backend. Submit a job and watch live logs with a loss chart that updates as training progresses.
**Evaluate.** Point at a trained model and test queries. Get metric cards for nDCG@10, MRR@10, Recall, and Precision.
**[Collections](https://qdrant.tech/documentation/concepts/collections/).** Browse your Qdrant collections, check point counts, and run test searches against indexed products. Useful for sanity-checking that indexing worked before evaluation.
**Publish.** Enter a model path and HuggingFace repo name. Click publish.
**Jobs.** Full history of every training, evaluation, and publish job. Expand any job to see its configuration, logs, and results. Jobs are tracked whether they were launched from the CLI or the dashboard.
**Settings.** Manage API keys for Qdrant, OpenAI, Anthropic, OpenRouter, Vultr, and HuggingFace. All stored in `.env`, all masked in the UI.
The dashboard is the recommended starting point. You can see everything the tool can do without reading documentation, and the job history gives you a record of every experiment.
## What Changed From the Series Code
![Production architecture: CLI, dashboard, and GPU backends](/articles_data/sparse-embeddings-ecommerce-part-5/production-architecture.png)
The pipeline from Parts 2-4 is the same underneath. The toolkit wraps it with:
**Automatic data handling.** Pass a CSV, JSON, JSONL, or HuggingFace dataset path. The loader auto-detects columns (`title`, `description`, `name`, `product_id`) and builds product text using the same bracket-brand, pipe-separated format from Part 2. No manual formatting.
**Synthetic query generation.** If you don't have labeled queries, the toolkit generates them using any [litellm](https://docs.litellm.ai/docs/providers)-supported model. Pass `--synth-model gpt-4o-mini` or `--synth-model ollama/llama3` for local generation. It reads your product catalog and generates realistic search queries with relevance labels.
```bash
# OpenAI
qdrant-finetune generate-queries --data products.csv --synth-model gpt-4o-mini
# Local Ollama
qdrant-finetune generate-queries --data products.csv --synth-model ollama/llama3
# Anthropic
qdrant-finetune generate-queries --data products.csv --synth-model anthropic/claude-sonnet-4-20250514
```
**Multi-backend GPU support.** Modal and Vultr are both supported. The `--gpu modal` flag handles volume creation, data upload, and job launching. `--gpu vultr` does the same for Vultr Cloud GPU. Both download the trained model back to your local machine when done.
**Interactive publishing.** The original series code required you to manually load the model and call `push_to_hub`. The toolkit prompts you after training completes, asks for the repo name, handles authentication, and prints the HuggingFace URL.
**Job tracking.** Every operation (train, evaluate, publish) is logged with configuration, timestamps, status, and output. The dashboard surfaces this as a job history. The CLI stores it in a local SQLite database.
## One-Line Python API
For notebooks and scripts, the same pipeline is available as a function call:
```python
from qdrant_finetune import finetune
model_path = finetune("products.csv")
```
That single line:
1. Loads your product data
2. Generates synthetic queries via LLM
3. Creates a SPLADE encoder
4. Runs 3 rounds of ANCE hard negative mining against Qdrant
5. Saves the fine-tuned model
For more control:
```python
from qdrant_finetune import Trainer, FinetuneConfig
config = FinetuneConfig(
base_model="distilbert/distilbert-base-uncased",
ance_iterations=5,
batch_size=64,
mining_top_k=20,
num_negatives=3,
)
trainer = Trainer(config)
trainer.fit(data="products.csv", queries="queries.csv")
trainer.index(data="products.csv", collection_name="my_products")
metrics = trainer.evaluate(queries="test_queries.csv")
print(metrics) # {'ndcg@10': 0.389, 'mrr@10': 0.387, ...}
```
Same ANCE loop from Part 3, same evaluation metrics. The `Trainer` class wraps the training loop, Qdrant indexing, hard negative mining, and evaluation into a coherent API.
## Getting Started
```bash
# Install from GitHub
pip install git+https://github.com/qdrant/sparse-finetune.git
# Configure (interactive wizard)
qdrant-finetune setup
# Run the full pipeline
qdrant-finetune pipeline --data your-products.csv --gpu modal # or --gpu vultr
```
The setup wizard validates your Qdrant connection, configures your LLM provider for synthetic queries, and sets up your GPU backend: Modal for serverless pay-per-second GPUs, or Vultr for dedicated cloud instances. Everything gets written to a `.env` file.
Or skip the CLI and launch the dashboard:
```bash
qdrant-finetune studio
```
The [source code is on GitHub](https://github.com/qdrant/sparse-finetune). File issues, submit PRs, or fork it for your own use case.
## What's Actually Different
The series taught the concepts. The toolkit removes the friction. Specifically:
- **No pipeline assembly.** You don't wire up Modal → training → Qdrant → mining → retraining. One command does it.
- **No data formatting.** Auto-detection handles CSV columns. The text builder applies the formatting conventions from Part 2 automatically.
- **No query labeling requirement.** Synthetic generation means you can start with just a product CSV. No click logs, no relevance labels needed.
- **No credential juggling.** Setup once, stored in `.env`, used everywhere.
- **No manual publishing.** Interactive prompt after training. HuggingFace link as the final output.
The 28% improvement over BM25 from Part 3 isn't locked behind a research repo anymore. It's a `pip install` away.
---
## Series Summary
- **[Part 1: Why sparse embeddings for e-commerce](/articles/sparse-embeddings-ecommerce-part-1/)**: SPLADE combines keyword precision with learned expansion
- **[Part 2: Training pipeline on Modal](/articles/sparse-embeddings-ecommerce-part-2/)**: A100 training with persistent checkpoints
- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)**: +28% vs BM25, ANCE mining with Qdrant
- **[Part 4: Specialization vs generalization](/articles/sparse-embeddings-ecommerce-part-4/)**: Domain-specific vs multi-domain tradeoffs
- **[Part 5: From research to product](/articles/sparse-embeddings-ecommerce-part-5/)**: CLI + dashboard that runs the full pipeline
@@ -68,7 +68,7 @@ Many users configure complex filters but may not be aware of the need to create
> As a result, every query scans thousands of vectors and their bare payloads before discarding the majority that failed the filter condition. This leads to soaring CPU usage and long response times, especially under higher traffic loads.
Filtering after retrieving thousands of vectors can get expensive. If you don't filter with your queries, then Qdrant will evaluate more vectors than you need. This will make the entire system slower and more resource intensive. Because of this, we have developed out own version of HNSW - [**The Filterable Vector Index**](https://qdrant.tech/articles/filtrable-hnsw/).
Filtering after retrieving thousands of vectors can get expensive. If you don't filter with your queries, then Qdrant will evaluate more vectors than you need. This will make the entire system slower and more resource intensive. Because of this, we have developed out own version of HNSW - [**The Filterable Vector Index**](https://qdrant.tech/articles/filterable-hnsw/).
Unlike some other engines, Qdrant lets you make the optimal choice of which fields to index for your use case rather than creating indexes for every field by default.
@@ -82,7 +82,7 @@ To ensure our catalog is accurate, we can use a dissimilarity search to highligh
To do this, we only need to search for the most dissimilar items using the
embedding of the category title itself as a query.
This can be too broad, so, by combining it with filters —a [Qdrant superpower](/articles/filtrable-hnsw/)—, we can narrow down the search to a specific category.
This can be too broad, so, by combining it with filters —a [Qdrant superpower](/articles/filterable-hnsw/)—, we can narrow down the search to a specific category.
{{< figure src=/articles_data/vector-similarity-beyond-search/mislabelling.png caption="Mislabeling Detection" >}}
@@ -98,7 +98,7 @@ For that reason, when comparing two similar sentences, their embeddings will tur
<img src="/articles_data/what-is-a-vector-database/two-similar-vectors.png" alt="Comparison of the embeddings of 2 similar sentences" width="500">
That’s the beauty of embeddings. Tthe complexity of the data is distilled into something that can be compared across a multi-dimensional space.
That’s the beauty of embeddings. The complexity of the data is distilled into something that can be compared across a multi-dimensional space.
### 3. The Payload: Adding Context with Metadata
@@ -112,7 +112,7 @@ For example, if you’re searching for a picture of a dog, the vector helps the
<img src="/articles_data/what-is-a-vector-database/filtering-example.png" alt="Filtering Example" width="500">
The payload can help you narrow down those results by ignoring vectors that doesn't match your query vector filtering criteria. If you want the full picture of how filtering works in Qdrant, check out our [Complete Guide to Filtering.](https://qdrant.tech/articles/vector-search-filtering/)
The payload can help you narrow down those results by ignoring vectors that don't match your query vector filtering criteria. If you want the full picture of how filtering works in Qdrant, check out our [Complete Guide to Filtering.](https://qdrant.tech/articles/vector-search-filtering/)
## The Architecture of a Vector Database
@@ -184,7 +184,7 @@ In Qdrant, indexing is modular. You can configure indexes for **both vectors and
<img src="/articles_data/what-is-a-vector-database/hnsw-search.png" alt="Searching Data with the HNSW algorithm" width="300">
You need to build the payload index for **each field** you'd like to search. The magic here is in the combination: HNSW finds similar vectors, and the payload index makes sure only the ones that fit your criteria come through. Learn more about Qdrant's [Filtrable HNSW](https://qdrant.tech/articles/filtrable-hnsw/) and why it was built like this.
You need to build the payload index for **each field** you'd like to search. The magic here is in the combination: HNSW finds similar vectors, and the payload index makes sure only the ones that fit your criteria come through. Learn more about Qdrant's [Filterable HNSW](https://qdrant.tech/articles/filterable-hnsw/) and why it was built like this.
> Combining [full-text search](https://qdrant.tech/documentation/concepts/indexing/#full-text-index) with vector-based search gives you even more versatility. You can simultaneously search for conceptually similar documents while ensuring specific keywords are present, all within the same query.
@@ -290,7 +290,7 @@ Sparse vectors are ideal for tasks like **keyword search** or **metadata filteri
Sometimes context alone isn’t enough. Sometimes you need precision, too. Dense vectors are fantastic when you need to retrieve results based on the context or meaning behind the data. Sparse vectors are useful when you also need **keyword or specific attribute matching**.
> With hybrid search you don’t have to choose one over the othe and use both to get searches that are more **relevant** and **filtered**.
> With hybrid search you don’t have to choose one over the other and use both to get searches that are more **relevant** and **filtered**.
To achieve this balance, Qdrant uses **normalization** and **fusion** techniques to blend results from multiple search methods. One common approach is **Reciprocal Rank Fusion (RRF)**, where results from different methods are merged, giving higher importance to items ranked highly by both methods. This ensures that the best candidates, whether identified through dense or sparse vectors, appear at the top of the results.
@@ -324,9 +324,9 @@ This is just a simple example and there's so much more you can do with it. See o
![vector-database-architecture](/articles_data/what-is-a-vector-database/vector-database-2.jpeg)
As your vector dataset grow larger, so do the computational demands of searching through it.
As your vector dataset grows larger, so do the computational demands of searching through it.
Quantized vectors are much smaller and easier to compare. With methods like [**Binary Quantization**](https://qdrant.tech/articles/binary-quantization/), you can see **search speeds improve by up to 40x while memory usage decreases by 32x**. Improvements that can be decicive when dealing with large datasets or needing low-latency results.
Quantized vectors are much smaller and easier to compare. With methods like [**Binary Quantization**](https://qdrant.tech/articles/binary-quantization/), you can see **search speeds improve by up to 40x while memory usage decreases by 32x**. Improvements that can be decisive when dealing with large datasets or needing low-latency results.
It works by converting high-dimensional vectors, which typically use `4 bytes` per dimension, into binary representations, using just `1 bit` per dimension. Values above zero become "1", and everything else becomes "0".
@@ -473,7 +473,7 @@ In more advanced setups, Qdrant uses **JWT (JSON Web Tokens)** to enforce **Role
RBAC defines roles and assigns permissions, while JWT securely encodes these roles into tokens. Each request is validated against the user's JWT, ensuring they can only access or modify data based on their assigned permissions.
You can easily setup you access tokens and secure access to sensitive data through the **Qdrant Web UI:**
You can easily setup your access tokens and secure access to sensitive data through the **Qdrant Web UI:**
<img src="/articles_data/what-is-a-vector-database/jwt-web-ui.png" alt="Qdrant Web UI for generating a new access token." width="1000">
+3 -2
View File
@@ -1,8 +1,9 @@
---
title: Vector Database Benchmarks
description: The first comparative benchmark and benchmarking framework for vector search engines and vector databases.
title: Vector Search Benchmarks
description: The first comparative benchmark and benchmarking framework for vector search engines.
keywords:
- vector databases comparative benchmark
- vector search comparative benchmark
- ANN Benchmark
- Qdrant vs Milvus
- Qdrant vs Weaviate
@@ -9,7 +9,7 @@ weight: 10
## Are we biased?
Probably, yes. Even if we try to be objective, we are not experts in using all the existing vector databases.
Probably, yes. Even if we try to be objective, we are not experts in using all the existing vector search engines.
We build Qdrant and know the most about it.
Due to that, we could have missed some important tweaks in different vector search engines.
@@ -22,7 +22,7 @@ There are several factors considered while deciding on which database to use.
Of course, some of them support a different subset of functionalities, and those might be a key factor to make the decision.
But in general, we all care about the search precision, speed, and resources required to achieve it.
There is one important thing - **the speed of the vector databases should to be compared only if they achieve the same precision**. Otherwise, they could maximize the speed factors by providing inaccurate results, which everybody would rather avoid. Thus, our benchmark results are compared only at a specific search precision threshold.
There is one important thing - **the speed of the vector search engines should to be compared only if they achieve the same precision**. Otherwise, they could maximize the speed factors by providing inaccurate results, which everybody would rather avoid. Thus, our benchmark results are compared only at a specific search precision threshold.
## How we select hardware?
@@ -58,8 +58,8 @@ Those may use some different protocols under the hood, but at the end of the day
## What about closed-source SaaS platforms?
There are some vector databases available as SaaS only so that we couldn’t test them on the same machine as the rest of the systems.
That makes the comparison unfair. That’s why we purely focused on testing the Open Source vector databases, so everybody may reproduce the benchmarks easily.
There are some vector search engines available as SaaS only so that we couldn’t test them on the same machine as the rest of the systems.
That makes the comparison unfair. That’s why we purely focused on testing the Open Source vector search engines, so everybody may reproduce the benchmarks easily.
This is not the final list, and we’ll continue benchmarking as many different engines as possible.
@@ -5,7 +5,7 @@ title: How vector search should be benchmarked?
weight: 1
---
# Benchmarking Vector Databases
# Benchmarking Vector Search
At Qdrant, performance is the top-most priority. We always make sure that we use system resources efficiently so you get the **fastest and most accurate results at the cheapest cloud costs**. So all of our decisions from [choosing Rust](/articles/why-rust/), [io optimisations](/articles/io_uring/), [serverless support](/articles/serverless/), [binary quantization](/articles/binary-quantization/), to our [fastembed library](/articles/fastembed/) are all based on our principle. In this article, we will compare how Qdrant performs against the other vector search engines.
@@ -31,4 +31,4 @@ On top of it, there is also a problem with search accuracy.
It appears if too many vectors are filtered out, so the HNSW graph becomes disconnected.
Qdrant uses a different approach, not requiring pre- or post-filtering while addressing the accuracy problem.
Read more about the Qdrant approach in our [Filtrable HNSW](/articles/filtrable-hnsw/) article.
Read more about the Qdrant approach in our [Filterable HNSW](/articles/filterable-hnsw/) article.
@@ -3,7 +3,7 @@ draft: false
id: 1
title: Single node benchmarks
description: |
We benchmarked several vector databases using various configurations of them on different datasets to check how the results may vary. Those datasets may have different vector dimensionality but also vary in terms of the distance function being used. We also tried to capture the difference we can expect while using some different configuration parameters, for both the engine itself and the search operation separately. </br> </br> <b> Updated: January/June 2024 </b>
We benchmarked several vector search engines using various configurations of them on different datasets to check how the results may vary. Those datasets may have different vector dimensionality but also vary in terms of the distance function being used. We also tried to capture the difference we can expect while using some different configuration parameters, for both the engine itself and the search operation separately. </br> </br> <b> Updated: January/June 2024 </b>
single_node_title: Single node benchmarks
single_node_data: /benchmarks/results-1-100-thread-2024-06-15.json
preview_image: /benchmarks/benchmark-1.png
+137
View File
@@ -0,0 +1,137 @@
---
draft: false
title: "Qdrant 2025 Recap: Powering the Agentic Era"
short_description: "A 2025 recap of Qdrant’s biggest product launches, customer wins, and technical milestones."
description: "A comprehensive recap of Qdrant’s 2025 highlights, including major product releases, enterprise deployments, AI workloads at scale, and the teams building with Qdrant in production."
preview_image: /blog/2025-recap/2025-hero-image.png
social_preview_image: /blog/2025-recap/2025-hero-image.png
date: 2025-12-17
author: "Daniel Azoulai"
featured: true
tags:
- qdrant
- company update
- product recap
- vector search
- enterprise ai
- open source
- 2025
---
![Infographic](/blog/2025-recap/2025-infographic.png)
This year was a defining year for Qdrant. Not because of a single feature or launch, but because of a clear shift in what the platform enables. As AI systems moved from static assistants to autonomous, multi-step agents, the demands placed on retrieval changed fundamentally. Speed alone was no longer enough. Production systems now require precise relevance control, predictable performance at scale, and the flexibility to run wherever data and users live.
We focused on meeting those requirements head-on. Rather than shipping disconnected features, we invested in deep, system-level improvements that strengthen Qdrant as long-term infrastructure. The result is a retrieval engine designed for real-world AI workloads: agentic, multimodal, hybrid, cost-efficient, and enterprise-ready.
Keep reading for a rundown.
## Product Focus: Built for Production, Designed for Agents
Across customers, partners, and open-source users, the same patterns kept surfacing: relevance breaks down under complex queries, costs explode at scale, multi-tenancy becomes fragile, and deployment constraints slow teams down.
In response, our 2025 roadmap centered on four tightly connected capability areas:
• Advanced Retrieval to move beyond basic vector similarity
• Performance & Resource Optimization to control cost without sacrificing speed
• Enterprise Scaling & Isolation to support shared, mission-critical infrastructure
• Deployment Flexibility to run in cloud, hybrid, or even edge environments
![Features](/blog/2025-recap/2025-features.png)
### Advanced Retrieval
In 2025, we focused on giving teams explicit control over retrieval quality as applications moved beyond basic semantic search. Our new capabilities make relevance more explainable, tunable, and aligned with real user intent, especially in agentic and hybrid search workflows.
**Related enhancements:**
• [Score-Boosting Reranking](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting) allowing the blending of vector similarity with business signals
• [Full-Text Filtering](https://qdrant.tech/documentation/concepts/filtering/) which brought native multilingual tokenization, stemming, and phrase matching
• [ACORN algorithm](https://qdrant.tech/documentation/concepts/search/#acorn-search-algorithm) for higher-quality filtered HNSW queries
• [Maximal Marginal Relevance (MMR)](https://qdrant.tech/blog/mmr-diversity-aware-reranking/) to balance relevance and diversity
• ASCII folding for improved multilingual recall
### Performance & Resource Optimization
To support large, cost-sensitive workloads, we targeted the biggest performance bottlenecks in production systems. New improvements help teams scale indexing and querying without over-provisioning memory or compute.
**Related enhancements:**
• [GPU-Accelerated HNSW Indexing](https://qdrant.tech/documentation/guides/running-with-gpu/) unlocks up to an order-of-magnitude faster ingestion
• [Inline Storage](https://qdrant.tech/documentation/guides/optimize/#inline-storage-in-hnsw-index) embedded quantized vectors directly into the graph to dramatically improve disk-based search performance
• [Custom storage engine](https://qdrant.tech/articles/gridstore-key-value-storage/) optimized for predictable low-latency access
• [Incremental HNSW indexing](https://qdrant.tech/documentation/database-tutorials/bulk-upload/?q=incremental+hnsw#choose-an-indexing-strategy) for upsert-heavy workloads
• HNSW graph compression to reduce memory footprint
• Expanded [Quantization](https://qdrant.tech/documentation/guides/quantization/#15-bit-and-2-bit-quantization]) options, including 1.5-bit, 2-bit, and asymmetric quantization
### Enterprise Scaling & Isolation
As Qdrant became shared infrastructure inside larger organizations, we focused on multitenancy, governance, and enterprise needs.
**Related enhancements:**
• [Tiered Multitenancy](https://qdrant.tech/documentation/guides/multitenancy/#tiered-multitenancy) enables efficient support for both small and large tenants within a single system
• [Single Sign-On (SSO) and role-based access control (RBAC)](https://qdrant.tech/enterprise-solutions/)
• Granular database API keys
• [Terraform-enabled Cloud API](https://qdrant.tech/enterprise-solutions/) for automation and governance
• Conditional updates for safe concurrent workflows and embedding migrations
### Deployment Flexibility & New Frontiers
We also expanded where and how Qdrant can run to match modern AI architectures. [Qdrant Cloud Inference](https://qdrant.tech/documentation/cloud/inference/) unified embedding generation and vector search into a single managed workflow, simplifying hybrid and multimodal pipelines. [Qdrant Edge](https://qdrant.tech/edge/) extended retrieval directly onto devices, enabling low-latency, deterministic search without a server dependency.
• Native support for dense, sparse, and image embeddings
• Hybrid retrieval pipelines without external inference infrastructure
• Consistent APIs across cloud, hybrid, and edge deployments
## Enabling Retrieval for the AI Era
Our customers validated our direction as we invested in more capable retrieval for AI. Below are just some of the companies that are use Qdrant.
![Logos](/blog/2025-recap/customer-logos.png)
• [Tripadvisor](https://qdrant.tech/blog/case-study-tripadvisor/) activated a dataset of over one billion reviews to power its AI Trip Planner, driving 2-3x more revenue from users engaged with the new generative experience.
• [OpenTable](https://qdrant.tech/blog/case-study-opentable/) reinvented dining discovery by building its AI Concierge on Qdrant, utilizing sparse embeddings to precisely filter over 60,000 restaurants for natural language queries.
• [HubSpot](https://qdrant.tech/blog/case-study-hubspot/) selected Qdrant to scale Breeze AI, its flagship intelligent assistant, ensuring highly personalized, context-aware responses without compromising on speed or reliability.
## A Thriving Community
In 2025, the Qdrant community achieved high-velocity. Our ecosystem grew from a strong base of early adopters into a global community of tens of thousands of engineers, researchers, and builders shaping the future of AI retrieval together.
### Community Engagement
Working with our community, we were able to create spaces together where practitioners could learn, share, and build.
[Vector Space Day 2025](https://qdrant.tech/blog/vector-space-day-2025-recap/) marked our first global Qdrant conference. Hosted at the Colosseum Theater in Berlin, the event brought together more than 400 in-person attendees, alongside hundreds more participating in a virtual hackathon. Talks and discussions spanned RAG, agentic memory, and distributed systems, with speakers from LlamaIndex, Vultr, and Google DeepMind.
To help developers bridge the gap between “Hello World” and production, we launched [Qdrant Essentials](https://qdrant.tech/course/essentials/). This comprehensive educational program covers vector search fundamentals, quantization strategies, and hybrid retrieval best practices. Thousands of developers have already learned from the course.
We also re-launched [Qdrant Stars](https://qdrant.tech/stars/), our ambassador program recognizing community members who create tutorials, speak at meetups, and mentor new users. Contributors Leaders like Pavan Kumar Mantha and Tarun Jain became the backbone of local Qdrant communities, driving meetups from Hyderabad to San Francisco.
### Momentum by the Numbers
The scale of community engagement in 2025 reflects accelerating adoption and collaboration:
• [GitHub](https://github.com/qdrant/qdrant) surpassed **27,000 stars**
• Added 35 integrations, including an [official n8n node](https://github.com/qdrant/n8n-nodes-qdrant).
• Our Github Issues tab evolved into an active collaboration space, with community contributions such as FastEmbed enhancements landing directly in core workflows
• [Discord](https://discord.com/invite/qdrant) grew to **8,000+ members**, serving as both a support hub and a place to share projects and wins
• On [LinkedIn](https://www.linkedin.com/company/qdrant/posts/?feedView=all), Qdrant appears in over **50 technical deep dives per day**
• Sponsored or supported **50+ AI and data events** worldwide, including ODSC West and /function1
![Awards](/blog/2025-recap/2025-award.png)
### Looking Ahead: The 2026 Roadmap
The progress in 2025 was shaped by real feedback and real use cases from the community. Building on that momentum, our 2026 roadmap doubles down on efficiency, agent-native retrieval, and enterprise-scale operability.
• **Efficiency & Scale**: 4-bit quantization, read-write segregation, block storage integration
• **Advanced Agent Retrieval**: relevance feedback, expanded inference capabilities
• **Robust Enterprise Deployment**: fully scalable multitenancy, faster horizontal scaling, read-only replicas
If you’re building the next generation of intelligent applications, or the infrastructure that supports them, Qdrant is ready. [Explore open roles](https://join.com/companies/qdrant) on our team or [start a free instance on Qdrant Cloud](https://cloud.qdrant.io/login) today.
![team](/blog/2025-recap/2025-team.png)
Thanks for joining our mission\!
@@ -0,0 +1,95 @@
---
draft: false
title: "How Anima Health scaled clinical document intelligence with Qdrant"
short_description: "Anima Health scaled privacy-first clinical intelligence with Qdrant."
description: "Discover how Anima Health used Qdrant to power vector search for clinical document coding, privacy-first retrieval, and agentic workflows in UK primary care."
preview_image: /blog/case-study-anima-health/social_preview_partnership-anima-health.png
social_preview_image: /blog/case-study-anima-health/social_preview_partnership-anima-health.png
date: 2026-01-28
author: "Daniel Azoulai"
featured: true
tags:
- Anima Health
- vector search
- healthcare
- clinical document intelligence
- retrieval-augmented generation
- agentic ai
- privacy
- case study
---
![Anima Health scaled privacy-first clinical intelligence with Qdrant](/blog/case-study-anima-health/anima-bento.png)
Primary care systems across the UK are under intense strain. General practitioners (GPs) balance their time with patient demand, understaffing, administrative burden vs. delivering care. <a href="https://animahealth.com/" target="_blank">Anima Health</a> set out to address this challenge by building a clinical operating system designed to make primary care more efficient, more informed, and more humane for both clinicians and patients.
At the heart of Anima’s platform is the ability to process large volumes of unstructured clinical data, including documents, test results, referral letters, and notes, while maintaining strict privacy guarantees. To achieve this at scale, Anima relies on Qdrant as a core infrastructure component for vector search, similarity analysis, and agentic AI workflows.
### The challenge: Under-capacity clinics and unstructured data overload
Anima focuses on GP practices in the UK, where under-capacity is the defining operational constraint. Clinics must triage and treat more patients than ever, with too few clinicians and limited administrative resources.
A major bottleneck lies in unstructured clinical documents. These can include PDFs, handwritten notes, blood test results, and referral letters. Important information often arrives late, is poorly indexed, or requires manual review by clinicians.
>“The main bottleneck is GPs. We are completely understaffed across the UK, and there is no end in sight. Optimizing GP time is essential.”
-Colin Cooke, Lead AI Engineer, Anima Health
This overload creates two compounding problems: First, clinicians lack timely access to information that could inform better decisions, and second, highly trained medical staff spend a disproportionate amount of time on administrative work rather than patient care.
### The solution: Vector-powered clinical intelligence with Qdrant
From the earliest stages of the product, Anima identified vector search as foundational infrastructure. Qdrant became the backbone of several key workflows, most notably clinical document coding.
Clinical documents are analyzed using large language models (LLMs) to extract meaning from unstructured text. However, LLMs alone cannot reliably handle medical ontologies such as SNOMED (Systematized Nomenclature of Medicine) codes, which are numeric identifiers with precise clinical meaning and downstream implications. These codes must be exact.
Anima represents SNOMED codes inside Qdrant as vector embeddings, enriched with metadata. During document processing, Qdrant is used as a retrieval layer inside an agentic pipeline. It narrows the search space, surfaces candidate codes, and enables high-confidence recommendations that clinicians can review and approve.
>“LLMs are great at understanding unstructured data, but they cannot free recall SNOMED codes. Those numeric IDs matter, and getting them wrong has real consequences.”
-Colin Cooke, Lead AI Engineer, Anima Health
Beyond coding, Anima uses Qdrant to understand documents at scale. By working with embedded representations of documents rather than raw text, the system can identify patterns across documents while preserving patient privacy. These signals influence downstream workflows without exposing sensitive content to models or operators.
>“Embeddings let us do a lot with clinical data without ever coming close to violating patient privacy. That is something we lean on heavily.”
-Colin Cooke, Lead AI Engineer, Anima Health
### Why Qdrant: Deployment control, cost predictability, and flexibility
Several factors made Qdrant a strong fit for healthcare workloads.
[Deployment flexibility](https://qdrant.tech/documentation/guides/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance.
Cost predictability also played a critical role. With a fixed infrastructure cost for vector search, Anima could use retrieval across multiple passes in their pipelines. This unlocked higher-quality results without eroding margins.
>“Knowing that retrieval is reliable and low cost changed how we build. We do not think twice about using vector search as part of our pipelines.”
-Colin Cooke, Lead AI Engineer, Anima Health
Finally, Qdrant’s vector-native capabilities mattered. [Payload-based filtering](https://qdrant.tech/documentation/concepts/payload/) allows Anima to scope searches precisely across different electronic health record systems. [Multivector support](https://qdrant.tech/documentation/concepts/payload/) enables experimentation with multiple embedding strategies and providers, reducing long-term lock-in and easing future transitions.
### Results: Scalable, privacy-first AI in production
Since moving into production, Qdrant has scaled quietly alongside Anima’s growth. Despite rapid increases in workload, the vector layer required minimal operational attention, even as it became deeply embedded across multiple pipelines.
>“We experienced significant growth, and Qdrant was never something I had to think about. It just continued to work.”
-Colin Cooke, Lead AI Engineer, Anima Health
Clinicians benefit from faster document processing, better prioritization, and reduced administrative overhead. Patients benefit from more timely and informed care decisions. Internally, Anima gained the confidence to build increasingly agentic workflows on top of a reliable retrieval foundation.
## What’s next for Anima Health
As Anima Health continues to scale, the team is looking beyond individual workflows toward more fully agentic clinical systems. These systems are designed to operate with greater autonomy while remaining tightly governed by clinical and regulatory constraints.
A key focus area is expanding how medical knowledge and patient context are represented and retrieved. Today, vector search already plays a central role in document understanding and classification. Going forward, Anima sees embeddings as a way to model longer-term patient state and clinical context that can evolve over time.
>“As we move toward more agentic systems, having reliable tools between the agent and our medical knowledge is essential. Qdrant is becoming one of those core tools.”
-Colin Cooke, Lead AI Engineer, Anima Health
Another area of exploration is temporal relevance. Clinical information can become stale at different rates depending on its nature. Anima is interested in systems that can reason about how recent or uncertain a piece of information is, and adjust retrieval and decision-making accordingly.
The team is also investing in flexibility across their AI stack. This includes experimenting with multiple embedding models in parallel and transitioning between providers without disrupting production systems. Qdrant’s capabilities make this kind of controlled experimentation possible, allowing Anima to adopt better models as they emerge.
>“I do not want to be locked into a single embedding provider. Multivector support gives us the freedom to test and transition as models improve.”
-Colin Cooke, Lead AI Engineer, Anima Health
Ultimately, Anima’s roadmap points toward AI systems that can support clinicians continuously rather than reactively. By grounding these systems in robust retrieval, strict privacy boundaries, and predictable infrastructure, Anima aims to scale clinical intelligence without sacrificing trust.
As healthcare organizations look to adopt more advanced AI workflows, Anima’s approach shows how vector search can serve not just as a technical optimization, but as a foundation for safe, future-proof clinical AI.
@@ -0,0 +1,116 @@
---
draft: false
title: "How Bazaarvoice scaled AI-powered product insights with Qdrant"
short_description: "Bazaarvoice scaled AI product insights across billions of reviews."
description: "Discover how Bazaarvoice reduced vector storage by ~100x and unlocked new AI-powered shopping and insights experiences with Qdrant."
preview_image: /blog/case-study-bazaarvoice//bazaarvoice-case-study-preview.png
social_preview_image: /blog/case-study-bazaarvoice//bazaarvoice-case-study-preview.png
date: 2026-02-10
author: "Daniel Azoulai"
featured: true
tags:
- Bazaarvoice
- vector search
- ecommerce
- ai search
- product insights
- cost optimization
- case study
---
![Bazaarvoice overview](/blog/case-study-bazaarvoice/bazaarvoice-bento.png)
## Turning billions of reviews into real-time, actionable intelligence
Bazaarvoice powers ratings and reviews across the global ecommerce ecosystem, connecting brands, retailers, and consumers through authentic product feedback. From brand-owned storefronts to major retailers, Bazaarvoice sources, verifies, and amplifies reviews at a scale few companies ever reach.
As large language models (LLMs) became production-ready, Bazaarvoice saw an opportunity to enhance the experiences of their clients' shoppers. The company wanted to help shoppers ask questions directly on product detail pages using natural language and help brands extract meaningful insights from vast volumes of unstructured customer feedback.
Delivering those experiences required a new foundation. Bazaarvoice needed vector search that could scale to billions of embeddings, support strict tenant isolation, and enable flexible query patterns. That requirement ultimately led the team to Qdrant.
## The challenge: Vector search at scale without operational overload
Bazaarvoice’s core data set is dominated by text reviews, with some images and videos layered in. What makes this data uniquely challenging is not only its size, but its structure. Reviews are syndicated across brands and retailers, meaning a single product can accumulate feedback from many different sources.
This created two fundamental constraints: First, the system had to scale to billions of vectors while remaining cost-efficient. Second, queries had to be scoped dynamically. Most searches are limited to a specific client, product, or category. Searching the entire corpus every time would waste compute, memory, and time.
To move quickly, Bazaarvoice initially implemented vector search using PostgreSQL with the pgvector extension. While this allowed the team to ship early versions of AI-powered features, it was never intended as a long-term solution.
>“We only did Postgres to get a product out the door quickly. We knew when we did it, it was not the right long-term choice.”
— Dr. Lou Kratz, Senior Principal Engineer, Bazaarvoice
As usage grew, the drawbacks became unavoidable. Every new client or product required manual partition creation. Cross-product queries forced Postgres to iterate across thousands of partitions, causing query latency to spike. At the same time, storage and RAM requirements climbed into the multi-terabyte range.
Bazaarvoice needed to build a vector search that could handle selective search at scale without forcing the team to engineer and maintain a parallel partitioning system.
## Why Bazaarvoice chose Qdrant
The team evaluated several approaches, including traditional search engines and other vector databases, but ruled out systems that relied on global approximate nearest neighbor search followed by post-filtering.
With over a billion vectors, post-filtering meant searching far more data than necessary.
>“Why would we search a billion vectors when we only need to search ten thousand?”
— Dr. Lou Kratz, Senior Principal Engineer, Bazaarvoice
Qdrant stood out for a few key reasons:
* [Multitenancy](https://qdrant.tech/documentation/guides/multitenancy/) with payload-based partitioning, allowing searches to be scoped by client, product, or category at query time
* [Quantization](https://qdrant.tech/documentation/guides/quantization/), enabling dramatic reductions in storage and RAM requirements which translates directly to cost
* [Hybrid cloud deployment](https://qdrant.tech/hybrid-cloud/), running inside Bazaarvoice’s VPC on Kubernetes
* Operational simplicity, eliminating manual partition management entirely
From the outset, the team defined clear success criteria: sub-100 millisecond query latency, approximately 98 percent nearest neighbor recall, and hands-off scaling as new clients and products were added continuously.
## Migrating billions of vectors under real-world constraints
The migration from PostgreSQL to Qdrant involved moving between four and five terabytes of data while continuing to ingest new reviews through streaming pipelines. At the same time, message retention limits created a narrow window to complete the migration safely.
>“We had to move about 4 to 5 TBs of data while new data kept coming in. There was a real time crunch.”
— Abhijeet Dhupia, Senior Machine Learning Engineer, Bazaarvoice
During the migration, all data was disk-backed to avoid exhausting RAM. Despite this unoptimized configuration, teams immediately noticed performance improvements.
>“Even with everything on disk, the feedback was that it was pretty fast compared to where we were earlier.”
— Abhijeet Dhupia, Senior Machine Learning Engineer, Bazaarvoice
Today, Bazaarvoice runs 2.7 billion review vectors in Qdrant and continues to tune the system as usage grows.
## Results: \~99% lower vector storage footprint with fast, accurate queries
The most immediate impact came from storage and infrastructure efficiency.
Bazaarvoice reduced its vector storage footprint by approximately 100x, compressing what previously required 4 to 5 terabytes in PostgreSQL down to a few hundred gigabytes in Qdrant. Quantization made it possible to keep billions of vectors accessible without the RAM requirements that full-resolution embeddings would have imposed.
>“Everyone talks about speed and accuracy. Storage is the story nobody talks about, and it’s the most important one at this scale.”
— Dr. Lou Kratz, Senior Principal Engineer, Bazaarvoice
Performance improved at the same time. Even during migration, with disk-based collections, the system consistently delivered sub-100 millisecond query latency while maintaining approximately 98 percent nearest neighbor accuracy. That balance allowed Bazaarvoice to dramatically lower infrastructure costs without sacrificing user experience.
Just as important, Qdrant removed a major source of engineering friction. Under the previous architecture, every new client or product would have required new Postgres partitions. At full scale, this would have meant maintaining close to one million partitions.
With Qdrant’s flexible indexing, cross-product and cross-category queries now work at query time without performance degradation.
>“In Postgres, you’d end up iterating over ten thousand partitions and the query would just go kaput. That simply doesn’t happen anymore.”
— Dr. Lou Kratz, Senior Principal Engineer, Bazaarvoice
## Unlocking New Product Development
By solving storage, cost, and operational complexity in one move, Qdrant unlocked entirely new product development at Bazaarvoice.
Within a single year, the team shipped two new AI-powered products on top of the same vector infrastructure. The AI Shopping Assistant went live in 2026, allowing shoppers to ask natural language questions directly on product pages. AI Insights, currently in prerelease, enables brands to explore sentiment and themes across entire product lines using conversational queries.
>“We’re shipping the AI products we actually bought this for now. And we’re doing it from a single place.”
— Dr. Lou Kratz, Senior Principal Engineer, Bazaarvoice
From a developer perspective, the experience remained consistent across environments. The open-source container made local development and CI/CD straightforward, while the hybrid cloud deployment satisfied security and compliance requirements.
>“It does one thing and it does it really well. It feels simple, even though I know it’s not.”
— Dr. Lou Kratz, Senior Principal Engineer, Bazaarvoice
## What’s next for Bazaarvoice
With Qdrant established as the central vector layer, Bazaarvoice plans to migrate additional workloads, including review summaries, into the same system. The team is also optimizing collections to further improve latency as AI-driven features move from prerelease into broader adoption.
For Bazaarvoice, the transition to Qdrant was not just a database migration. It was a structural shift that made large-scale, AI-powered commerce experiences faster to build, cheaper to run, and easier to evolve.
@@ -1,9 +1,9 @@
---
title: "Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Database"
title: "Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Search"
draft: false
slug: case-study-dailymotion # Change this slug to your page slug if needed
short_description: Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Database
description: Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Database
short_description: Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Search
description: Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Search
preview_image: /case-studies/dailymotion/preview-dailymotion.png # Change this
# social_preview_image: /blog/Article-Image.png # Optional image used for link previews
@@ -21,7 +21,7 @@ weight: 0 # Change this weight to change order of posts
# For more guidance, see https://github.com/qdrant/landing_page?tab=readme-ov-file#blog
---
## Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Database
## Dailymotion's Journey to Crafting the Ultimate Content-Driven Video Recommendation Engine with Qdrant Vector Search
In today's digital age, the consumption of video content has become ubiquitous, with an overwhelming abundance of options available at our fingertips. However, amidst this vast sea of videos, the challenge lies not in finding content, but in discovering the content that truly resonates with individual preferences and interests and yet is diverse enough to not throw users into their own filter bubble. As viewers, we seek meaningful and relevant videos that enrich our experiences, provoke thought, and spark inspiration.
Dailymotion is not just another video application; it's a beacon of curated content in an ocean of options. With a steadfast commitment to providing users with meaningful and ethical viewing experiences, Dailymotion stands as the bastion of videos that truly matter.
@@ -75,7 +75,7 @@ Title , Tags , Description , Transcript (generated by [OpenAI whisper](https://o
![quote-from-Samuel](/case-studies/dailymotion/Dailymotion-Quote.jpg)
Looking at the complexity, scale and adaptability of the desired solution, the team decided to leverage Qdrant’s vector database to implement a content-based video recommendation that undoubtedly offered several advantages over other methods:
Looking at the complexity, scale and adaptability of the desired solution, the team decided to leverage Qdrant’s vector search to implement a content-based video recommendation that undoubtedly offered several advantages over other methods:
**1. Efficiency in High-Dimensional Data Handling:**
@@ -0,0 +1,78 @@
---
draft: false
title: "Building real-time multimodal similarity search in Flipkart Trust & Safety with Qdrant"
short_description: "Tackling fraud and abuse with scalable similarity search."
description: "Tackling fraud and abuse with scalable similarity search."
preview_image: /blog/case-study-flipkart/social_preview_partnership-flipkart.png
social_preview_image: /blog/case-study/social_preview_partnership-flipkart.png
date: 2026-01-09
author: "Daniel Azoulai"
featured: true
tags:
- Flipkart
- vector search
- multimodal search
- fraud detection
- real-time search
- trust and safety
- case study
---
### Tackling fraud and abuse with scalable similarity search
At Flipkart, the Trust & Safety team is focused on detecting and preventing platform abuse and fraud. A critical part of this work involves running large-scale similarity searches across customer and seller-submitted data, particularly images. This allows the team to identify patterns associated with fraudulent activity, such as repeat returns or duplicate seller claims, before they cause downstream harm.
*“Platform integrity is a constant challenge. To stay ahead of fraudulent actors, we needed a system that could compare multimodal data in real time, not just in long-running batch jobs.”*
— Sourabh Sarkar, SDE-III, Trust & Safety at Flipkart
### Limitations of prior batch-based methods
The team’s earlier approach to similarity search used HBase with Locality-Sensitive Hashing (LSH). While workable for batch analysis, this system was slow and could not keep up with the demands of real-time fraud prevention. In some cases, finding similar images in historical data could take up to nine hours.
Additionally, Flipkart’s embedding models produce high-dimensional vectors (2048 dimensions), which added pressure on indexing performance and made efficient real-time querying more difficult.
### Evaluating open-source options and selecting Qdrant
To address these challenges, the team evaluated multiple open-source vector databases through a proof-of-concept. They chose Qdrant because it provided:
• **Deployment flexibility** with official Debian packaging, which fit well with Flipkart’s internal infrastructure
• **Efficient HNSW indexing** capable of handling simultaneous reads and writes
• **Support for high-dimensional embeddings**, critical for the models in production
### Building a multi-tenant similarity service
The Trust & Safety team then built a new multi-tenant similarity service. This platform now supports several important use cases:
• **Fraud detection:** Real-time image similarity checks to identify potentially abusive behavior
• **Address clustering:** Grouping unstructured customer addresses to improve last-mile delivery routing
• **Retrieval-augmented generation (RAG):** Serving as the retrieval layer for internal GenAI initiatives
*“What used to take hours in our old batch workflows can now be done in under a minute. That change has been crucial in stopping fraud before it impacts customers.”*
— Sourabh Sarkar, SDE-III, Trust & Safety at Flipkart
### Detection time reduced from 9 hours to 1 minute
The shift from batch processing to real-time search has significantly reduced detection time, from nine hours to under one minute. This improvement enables much earlier intervention against fraud.
From a developer perspective, integration with the Java gRPC SDK and Prometheus metrics endpoint simplified adoption and monitoring. The team also built custom adapters and backup scripts, ensuring the service could be reused by multiple teams without duplicating effort.
### Looking ahead: Expanding beyond fraud detection
The Trust & Safety team continues to broaden its capabilities. Upcoming projects include:
• Expanding retrieval use cases for company-wide RAG systems
• Standardizing on a Kubernetes-based Qdrant deployment as the embedding store across different groups at Flipkart
• Exploring integrations with agentic AI frameworks to further automate detection and prevention workflows
*“We see vector databases becoming a key part of modern AI infrastructure. It’s not only for fraud detection, but also as a foundation for new AI systems we’re experimenting with.”*
— Sourabh Sarkar, SDE-III, Trust & Safety at Flipkart
@@ -0,0 +1,89 @@
---
draft: false
title: "How GlassDollar improved high-recall sourcing by migrating from Elasticsearch to Qdrant"
short_description: "GlassDollar migrated from Elasticsearch to Qdrant to scale semantic retrieval, cut costs, and improve end-to-end RAG accuracy."
description: "Discover how GlassDollar reduced infrastructure costs by ~40% and increased user engagement 3x by improving end-to-end accuracy with Qdrant."
preview_image: /blog/case-study-glassdollar/social_preview_partnership-glassdollar.jpg
social_preview_image: /blog/case-study-glassdollar/social_preview_partnership-glassdollar.jpg
date: 2026-03-04
author: "Daniel Azoulai"
featured: true
tags:
- GlassDollar
- vector search
- semantic search
- rag
- query expansion
- nodejs
- cost reduction
- case study
---
![GlassDollar overview](/blog/case-study-glassdollar/glassdollar-bento-box.png)
<a href="https://www.glassdollar.com/" target="_blank">GlassDollar</a> helps enterprises such as Siemens, Mahle, and A2A discover, compare, and run proof-of-concepts with innovative startups. The platform combines an chatbot-like experience for discovering innovate companies with tools to manager innovation projects from start to finish.
For GlassDollar, search is not a feature. It is the core mechanism that turns an enterprise problem statement into a shortlist of relevant companies, ranked and contextualized for decision-making.
## Scaling to 10 Million Documents Pushed Search to its Limits
GlassDollar’s dataset spans a few million companies. Each company can map to multiple documents, including product descriptions and technical summaries.
Early on, the team implemented vector search with Elasticsearch using OpenAI embeddings. It was a sensible default for a full-stack team that looked for a mature ecosystem and easy-to-find resources.
As the platform grew, retrieval became a bottleneck. GlassDollar wanted to index more companies, more documents per company, and more sources of text per profile. But performance slowed as the indexed corpus expanded, and it became difficult to see a path to scaling 10x or more without compromising the user experience.
The team also found themselves carrying extra complexity. To achieve acceptable quality, they maintained keyword-based logic alongside semantic retrieval, which increased operational overhead and made iteration harder.
## The Product Evolved from Fast Search to High-Recall Sourcing
GlassDollar’s early interface looked like a familiar search box, so latency targets were strict. Over time, the product shifted closer to the real-world sourcing process enterprises already used. Users described what they needed in a sentence or a paragraph and cared most about whether the right companies appeared, even if results took longer.
That change forced a new definition of success.
Instead of optimizing around a single response time target, GlassDollar focused on retrieval quality. Recall became the guiding metric, since missing the best companies was the fastest way to lose user trust.
>“At first, we were thinking in terms of speed. Later we understood that quality comes first. We measure recall now. If the best companies are not on the screen, nothing else matters.” Kamen Kanev, GlassDollar
## Accuracy Meant End-to-end Results, not Just Retriever Scores
For GlassDollar, accuracy was measured at the workflow level: did the system surface the right companies for a given enterprise need, and did users act on the results. That meant optimizing the full architecture, including [query expansion](https://qdrant.tech/documentation/concepts/hybrid-queries/), retrieval, ranking, and contextual ranking, rather than chasing marginal gains in any single component. Faster retrieval mattered most because it enabled more queries to improve query expansion, which raised recall and improved the final shortlist quality.
## Contextual Embeddings Improved Matching Between User Intent and Company Descriptions
One persistent challenge was representing long company descriptions in a way that matched short, intent-driven queries. Embedding an entire document often diluted meaning and reduced match quality. In practice, the team saw that how text was sliced and embedded mattered as much as the embedding model itself.
GlassDollar adopted a contextual chunking approach inspired by recent research. Instead of embedding isolated sentences, they produced segments that preserved key context across chunks. If a company positioned itself as an automotive solution provider early in its description, that context continued to appear in subsequent embedded segments. This helped retrieval when users searched for outcomes or categories that were implied rather than explicitly listed.
## Query Expansion Raised Recall but Required Faster Vector Retrieval
To push recall higher, GlassDollar leaned into query expansion. A single user prompt could generate multiple related queries, each retrieving candidates from a different angle. Results were then combined and passed into ranking and contextual ranking stages.
This approach improved coverage, but it amplified the importance of retrieval performance. If each search was slow, query expansion became impractical. The team needed a vector search that could handle more queries per request without degrading speed or cost.
## GlassDollar Migrated from Elasticsearch to Qdrant to Scale Retrieval, Reducing Costs by 40%.
To evaluate alternatives, GlassDollar built a benchmark script with a golden dataset: a set of representative queries paired with the companies they expected to retrieve. The team tested Qdrant quickly with partial indexing and minimal configuration to validate whether recall was in the right range.
The results were close enough to their Elasticsearch baseline to justify a full migration. Once Qdrant was in production, GlassDollar saw immediate gains in retrieval speed and the ability to run more expanded queries within the same time budget. They also reduced system complexity by removing keyword-specific compensations while maintaining overall quality.
Cost improved as well. After migration, infrastructure costs dropped to roughly 60 percent of the previous setup, even before the team fully indexed all planned documents and sources.
## Node.js and TypeScript Kept the System Accessible to Full-stack Engineers
A key requirement for GlassDollar was staying productive in a Node.js and TypeScript-first environment. The team used the [Qdrant Node.js SDK](https://github.com/qdrant/qdrant-js) alongside LLM APIs to build retrieval, query expansion, and ranking workflows directly inside their existing backend services.
This mattered for hiring and velocity. It enabled engineers who already shipped product in JavaScript and TypeScript to implement modern RAG pipelines without maintaining a separate Python-only stack.
## User Engagement Rripled After Retrieval Quality Improved
GlassDollar tracked success through product behavior. One signal was how often users saved (bookmarked) companies during sourcing workflows. Before the search improvements, many users either relied on manual sourcing or browsed results without committing them into a shortlist.
After migrating to Qdrant and scaling their recall-focused retrieval strategy, save activity increased sharply. In Q1, GlassDollar saw a 3x increase in bookmarks compared to previous months, driven by more engagement from existing users rather than only new user growth.
## Next Steps Focus on Repeatable Feedback Loops for Accuracy Gains
With retrieval scalability in place, GlassDollar’s roadmap centers on continuous accuracy improvements across the full RAG architecture. They will continue optimizing query expansion strategy and improving reraking models to continually improve.
The goal remains consistent: deliver the most accurate matchmaking between corporate pain points and startups through a search built for high-recall, decision-ready results.
@@ -0,0 +1,97 @@
---
draft: false
title: "How Kakao Built an AI-Powered Internal Service Desk with Qdrant"
short_description: "Kakao built an AI-powered internal service desk with Qdrant."
description: "Discover how Kakao’s Connectivity Platform team built an AI-powered internal Service Desk using Qdrant to enable hybrid search, scale securely on Kubernetes, and improve employee productivity."
preview_image: /blog/case-study-kakao/social_preview_partnership-kakao.png
social_preview_image: /blog/case-study-kakao/social_preview_partnership-kakao.png
date: 2026-01-27
author: "David Koh - Kakao Connectivity Platform"
featured: true
tags:
- Kakao
- vector search
- hybrid search
- retrieval-augmented generation
- internal knowledge search
- case study
---
<a href="https://www.kakaocorp.com/" target="_blank">Kakao</a> is one of South Korea's leading technology companies, best known for KakaoTalk, the country's dominant messaging platform with over 48 million monthly active users. Beyond messaging, Kakao operates a broad ecosystem of services including maps, mobility, fintech, and enterprise solutions.
## Helping employees find answers faster without sacrificing precision or control
Kakao’s Connectivity Platform team set out to solve a familiar internal problem: employees across the organization needed a faster, more reliable way to get answers about internal systems, APIs, and operational procedures. The result was **Service Desk Agent**, an AI-powered internal service desk designed to answer questions in natural language using Kakao’s internal documentation and historical inquiry data.
Built as a Retrieval-Augmented Generation (RAG) system on top of LangGraph, Service Desk Agent acts as a conversational interface to Kakao’s internal knowledge. This helps employees resolve issues quickly, while reducing repetitive work for their support staff.
## The challenge: searching complex internal knowledge at scale
From the beginning, the team faced a search problem that couldn’t be solved with a single retrieval approach.
Service Desk Agent needed to work across two very different types of data:
1. Long-form technical documentation, including project guides and API specifications, where understanding context matters. But, exact system names and proper nouns still need to be matched.
2. Historical Q\&A and incident data, which often includes precise error messages, commands, and configuration details.
Pure keyword search struggled with semantic questions. Pure vector search struggled with exact terms and proper nouns. Neither approach alone was sufficient.
At the same time, Kakao had strict infrastructure and operational requirements. All data needed to remain within internal infrastructure, the system had to be deployable on Kubernetes, and the solution needed a permissive open-source license with strong official documentation for operations like upgrades, backup, and recovery.
This was a greenfield project; Kakao wasn’t replacing an existing vector database. Instead, the team evaluated several options, including Milvus, Weaviate, Qdrant, and Elasticsearch, to determine the best foundation for RAG-based internal AI services.
## Why Kakao chose Qdrant
After evaluating multiple vector databases, the Connectivity Platform team selected Qdrant as their first vector search solution.
The decision came down to a combination of search quality, performance, and operational fit.
Qdrant’s hybrid search capabilities were a key factor. By supporting both dense vectors for semantic search and sparse vectors for keyword-based retrieval—combined using [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/#reciprocal-rank-fusion-rrf), the team could address both conceptual questions and exact-match queries in a single system. Named Vectors made it possible to manage multiple vector types within the same collection.
Performance was another major consideration. Qdrant’s [Rust-based architecture](https://qdrant.tech/articles/why-rust/), efficient [HNSW implementation](https://qdrant.tech/course/essentials/day-2/what-is-hnsw/), and support for [scalar quantization (INT8)](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization) provided low-latency search while optimizing memory usage. This was crucial for an internal service expected to scale over time.
From an operational standpoint, Qdrant fit naturally into Kakao’s environment. Its single-binary design simplified deployment, it ran reliably on Kubernetes, and it allowed Kakao to retain full control over data by self-hosting within internal infrastructure.
Finally, the team highlighted the developer experience: a well-designed Python SDK with async support, an intuitive Web UI for debugging and exploration, and detailed official documentation covering both usage and performance tuning.
## How Service Desk Agent uses Qdrant
Qdrant sits at the core of Service Desk Agent’s RAG architecture, acting as the system’s primary vector store.
The team integrated Qdrant using the asynchronous Python client (`AsyncQdrantClient`) to handle high query concurrency, with secure HTTPS communication for all data access.
Collections were designed around data sources, with separate collections for internal technical documentation, historical inquiry data, and a semantic cache used to speed up repeated queries. Metadata filtering allows the system to narrow search scope by service or time period, while maintaining fast response times.
Each collection stores both dense and sparse vectors using [Named Vectors](https://qdrant.tech/documentation/concepts/vectors/#named-vectors). Hybrid search results are merged using RRF to produce more accurate answers across different query types.
An automated indexing pipeline handles document ingestion end-to-end. This ranges from cleansing and chunking, to embedding generation, to batch upserts into Qdrant.
The entire system runs on Kakao’s internal Kubernetes cluster, using a replication-based Qdrant setup for high availability and rolling updates for zero-downtime deployments. The system integrates with distributed locking and state management solutions within the broader application architecture.
![indexing-pipeline](/blog/case-study-kakao/indexing-pipeline.png)
![search-pipeline-diagram](/blog/case-study-kakao/search-pipeline-diagram.png)
## The result: faster answers, lower support load
Today, Service Desk Agent supports approximately 1 million vectors, with the dataset continuing to grow as more internal knowledge is indexed. The system operates across multiple collections and uses 3072-dimensional embeddings from OpenAI’s model.
By combining hybrid search, metadata filtering, and semantic caching, the team was able to improve both search quality and response times. Scalar quantization reduced memory usage, while tuned HNSW parameters helped maintain low latency under load.
From a business perspective, the impact has been clear:
* Reduced workload for support staff, as more inquiries are resolved automatically before tickets are created.
* Improved employee satisfaction, with faster, self-service access to internal knowledge.
* Higher development productivity, enabled by Qdrant’s SDK, documentation, and ease of integration during rapid prototyping.
Before Service Desk Agent, employees searched across multiple internal knowledge bases manually; a process that could take several minutes depending on query complexity. Now, end-to-end response time averages under 30 seconds, from query submission to complete answer delivery.
## What’s next
Kakao plans to continue expanding Service Desk Agent by indexing more of its internal knowledge base. The team is also exploring deeper GraphRAG integration, potential multimodal search capabilities, and sharding strategies to support future scale and performance needs.
Within the Connectivity Platform team, Qdrant has become the foundation for RAG-based internal AI services. It provides the hybrid search, flexibility, and operational stability required to support AI-driven knowledge access at scale.
*This post authored by David Koh — Kakao Connectivity Platform*
@@ -0,0 +1,94 @@
---
draft: false
title: "How My AskAI Built Self-Improving Support Agents"
short_description: "My AskAI scaled reliable support agents on Qdrant Cloud."
description: "Discover how My AskAI built a self-improving customer support agent platform with Qdrant Cloud, enabling scalable retrieval, hybrid search iteration, and faster operations."
preview_image: /blog/case-study-my-askai/social_preview_partnership-my-askai.png
social_preview_image: /blog/case-study-my-askai/social_preview_partnership-my-askai.png
date: 2026-02-25
author: "Daniel Azoulai"
featured: true
tags:
- My AskAI
- vector search
- customer support
- rag
- hybrid search
- llm agents
- case study
---
![My AskAI overview](/blog/case-study-my-askai/my-askai-bento-box.png)
[My AskAI](https://myaskai.com) built a managed platform for AI customer support agents that plug directly into existing helpdesk tools like [Intercom](https://myaskai.com/ai-agent-integration/intercom) and [Zendesk](https://myaskai.com/ai-agent-integration/zendesk-tickets). The goal was to make AI behave like a reliable coworker, not a brittle chatbot. In production, My AskAI's agents are designed to resolve a large portion of inbound support requests automatically, then hand over to a human when the agent cannot answer confidently. My AskAI positions this as [deflecting around 75 percent of support requests](https://myaskai.com/blog/my-askai-edel-optics-case-study-2026) and sustaining a resolution rate in the low to mid 70s, depending on the time window and workload mix.
As My AskAI narrowed its focus, the team discovered that customer support was not just a use case; it was **the use case**. Support data is messy, unstructured, and constantly changing. Success required strong retrieval, predictable latency, and an infrastructure layer that could scale without forcing the team to become full-time database operators. That combination ultimately led My AskAI to standardize on [Qdrant Cloud](http://cloud.qdrant.io) as the vector search backbone of its platform.
## Customer Support Created Unique Retrieval Requirements
My AskAI started as a broader "chat with your data" product. Users could connect sources like Google Drive, PDFs, and web content, then ask questions in natural language. As usage grew, My AskAI looked for the accounts that were most engaged and willing to pay for the product's value.
That analysis showed a clear pattern.
"When we double clicked on our users, the stickiest users and the ones generating the most revenue, we saw that pretty much all of them were using it in a customer support use case," explains Alex Rainey, co-founder of My AskAI. "So we thought, there's something interesting happening here. Let's focus on this niche."
These teams cared about accuracy, speed, and guardrails because every answer represented their brand. In practice, that meant My AskAI needed to index and retrieve from large sets of unstructured documents and historical support interactions. Keyword search alone was not enough. Customers often described issues differently than the help article headings, and support tickets were full of partial context and inconsistent phrasing. My AskAI needed semantic retrieval that could match intent, not just exact terms.
## Embeddings and RAG Made the Product Viable
My AskAI's earliest attempts at working with customer data hit a wall quickly. Before embedding models were widely available, the team tried fine-tuning early language models on customer support tickets to generate answers to new questions.
>"We tried to do some fine-tuning, which was obviously the wrong way to approach the problem. Fine-tuning on customer support tickets to answer new questions just didn't make sense. It was very early days."
Everything changed when OpenAI released its embedding model. Instead of hoping users would guess the right keyword, My AskAI could retrieve relevant passages by semantic similarity and pass them into an LLM as context. This became the basis of My AskAI's customer support workflow.
>"That was a transformational moment for us. Now we could have hundreds of help articles ingested in the system, and a user can ask a question, and we can answer that really specifically and cheaply and quickly."
Over time, the team also learned that semantic search was strong but not universally sufficient, especially when tickets contained product names, error codes, or specific identifiers that benefit from lexical matching. That realization led My AskAI toward experimentation with [hybrid search](https://qdrant.tech/documentation/concepts/hybrid-queries/) as a way to blend semantic similarity with keyword signals, while keeping the operational footprint small.
## Why My AskAI Chose Qdrant: Scalability, Integrations, and Developer Experience
As My AskAI's usage scaled, the team evaluated vector search choices through a longer-term lens. Switching vector search infrastructure later can be painful, so the decision had to hold up as the product grew 10x.
My AskAI ultimately selected Qdrant for three reasons.
First, cost and scalability needed to be predictable as workloads increased. The team modeled their projected growth and found that their previous provider's pricing would become prohibitively expensive at scale.
Second, the platform needed to integrate smoothly into an ecosystem of connectors and support tools. My AskAI was building integrations with third-party connector tools for sources like Google Drive, Notion, Intercom, and Zendesk, and Qdrant was natively supported as a plugin in that ecosystem.
Third, the developer experience had to be straightforward so the team could scale and manage instances without constant engineering effort.
The team also leaned on community signal to validate the decision. "A lot of developers were building in public, speaking about exactly what infra they were using under the hood. I just kept seeing Qdrant coming up. Engineers at our third-party connector tool spoke extremely highly of Qdrant for scalability and latency."
The migration itself was smooth, in part because My AskAI used the transition to build a V2 of their product focused purely on customer support. That gave them a clean start with Qdrant and a modernized stack, rather than requiring a large-scale data migration.
## What Changed After Migrating
Once on Qdrant Cloud, My AskAI leaned into a workflow where scaling and day-to-day operations were simple. The team relied on the dashboard for performance visibility and managed scaling with a few UI actions, while still using the API for collection setup and configuration when needed.
>"Scaling horizontally or vertically is like two clicks away. It's not really a concern we have. If we notice anything, we get the alert, and we can jump in and scale things up or down as we need to."
The ideal infrastructure, as Alex puts it, is the kind you don't have to think about. "I didn't want to have to think about it."
My AskAI also began running customer-specific proofs of concept for [hybrid search](https://qdrant.tech/documentation/concepts/hybrid-queries/), aiming to find the right blend that improved retrieval in the edge cases where semantic-only results were not enough. Before Qdrant, managing hybrid search had required spinning up separate infrastructure on AWS and handling reranking externally. With Qdrant, the team could enable hybrid search per collection and iterate without managing additional systems.
>"Just being able to turn on hybrid search is super useful. It removes that headache and pushes management of hybrid search down to the vendor."
Just as importantly, My AskAI highlighted the hands-on, highly technical support experience as part of why the platform felt dependable for production workloads.
"Qdrant support is always phenomenal. Super fast, super technical, very hands on. We've always had great experiences whenever we needed them."
## What's Next: Self-Learning Support Agents Powered by Clustered Knowledge
My AskAI's most important next step is self-learning.
When an AI agent escalates a conversation to a human, My AskAI now monitors what the human does next and treats it as a [learning opportunity](https://myaskai.com/features/self-learning). The system compares the human response to what the AI would have done, identifies gaps, and stores new learnings. Those learnings are then clustered and compiled into what My AskAI calls a self-learning article, a continuously updated body of knowledge derived from real support outcomes.
"Most of the time it's because companies just don't keep their help articles super up to date. It's not easy to do that," Alex explains. "These learnings get stored, get clustered, and then we push those into what we call a self-learning article, a collection of new knowledge that doesn't exist in public help articles but exists in the human agent replies."
Those generated articles are re-indexed in Qdrant so the next time the question appears, the agent can answer correctly, sometimes as soon as the next day.
"The next day the question can be answered, because the human agent gave us the answer and that knowledge has now been indexed."
@@ -46,7 +46,7 @@ With the emergence of pure-play, native vector search engines, Nyris conducted e
As part of their selection process, Nyris evaluated several critical factors to ensure they chose the best vector search engine solution:
- **Accuracy and Speed**: These were primary considerations. Nyris needed to understand the performance differences between the [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) graph-based approach and brute-force search. In particular, they examined edge cases that required numerous filters, sometimes necessitating a switch to brute-force search. Even in these scenarios, Qdrant demonstrated impressive speed and reliability, meeting Nyris's stringent performance requirements.
- **Accuracy and Speed**: These were primary considerations. Nyris needed to understand the performance differences between the [HNSW](https://qdrant.tech/articles/filterable-hnsw/) graph-based approach and brute-force search. In particular, they examined edge cases that required numerous filters, sometimes necessitating a switch to brute-force search. Even in these scenarios, Qdrant demonstrated impressive speed and reliability, meeting Nyris's stringent performance requirements.
- **Insert Speed**: Nyris assessed how quickly data could be inserted into the database, including the performance during simultaneous data ingests and query requests. Qdrant excelled in this area, providing the necessary efficiency for their operations.
- **Total Cost of Ownership**: Nyris analyzed the infrastructure costs and licensing fees associated with each solution. Qdrant offered a competitive total cost of ownership, making it an economically viable option.
- **Data Sovereignty**: The ability to deploy Qdrant in their own clusters was a key aspect for Nyris, ensuring they maintained control over their data and complied with relevant data sovereignty requirements.
@@ -17,7 +17,7 @@ tags:
---
A problem we've noticed while monitoring the [Qdrant Discord Community](https://discord.gg/d4MPnX3s) is that due to the extensive list of expressions that the [score boosting](https://qdrant.tech/documentation/concepts/hybrid-queries/#score-boosting) functionality provides, there's room for confusion on how it's supposed to be applied. And that might block you from moving the business logic behind relevance scoring into the Qdrant search engine. We don't want that!
A problem we've noticed while monitoring the [Qdrant Discord Community](https://discord.gg/d4MPnX3s) is that due to the extensive list of expressions that the [score boosting](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting) functionality provides, there's room for confusion on how it's supposed to be applied. And that might block you from moving the business logic behind relevance scoring into the Qdrant search engine. We don't want that!
In this blog, we'd like to de-spooky-fy the **decay functions** part of the score boosting, or, more precisely: `LinDecayExpression`, `ExpDecayExpression`, and `GaussDecayExpression` -- frequent guests on the Discord *#ask-for-help* channel.
@@ -170,7 +170,7 @@ But here's the problem: That 36 might not be a "high" score at all. Maybe your d
Now let's see how using decay functions looks in Qdrant.
We'll provide HTTP request examples, but you can use decay functions [analogously in the Python, TypeScript, Rust, Java, C#, and Go clients](https://qdrant.tech/documentation/concepts/hybrid-queries/#time-based-score-boosting).
We'll provide HTTP request examples, but you can use decay functions [analogously in the Python, TypeScript, Rust, Java, C#, and Go clients](/documentation/concepts/search-relevance/#time-based-score-boosting).
**Note #6.**
Payload variables used within the formula benefit from having [payload indexes](https://qdrant.tech/documentation/concepts/indexing/#payload-index). So, we require you to set up a payload index for any variable used in a formula.
@@ -248,7 +248,7 @@ We truly hope this write-up helped untangle things a bit. Now the only thing lef
Use the snippets in the article as a starting point and experiment with the relevance score boosting in [Qdrant Cloud](https://qdrant.tech/). We offer a free-forever 1GB cluster: enough to test, tweak, and see how the decay functions behave on your data.
And if you feel like diving deeper into decay functions or score boosting in general, check out our [documentation](https://qdrant.tech/documentation/concepts/hybrid-queries/?q=Query+Points+API#score-boosting), which includes a decay-on-distance example and plenty more to learn from.
And if you feel like diving deeper into decay functions or score boosting in general, check out our [documentation](/documentation/concepts/search-relevance/#score-boosting), which includes a decay-on-distance example and plenty more to learn from.
### Tell Us What You're Building
@@ -94,4 +94,4 @@ With this combination, you can simplify infrastructure management, implement sec
## Come Build with Us\!
[Contact Sales](https://qdrant.tech/contact-us/) to enable enterprise features for your team, or [start prototyping with a free Qdrant cluster](https://cloud.qdrant.io/).
[Contact Us](https://qdrant.tech/contact-us/) to enable enterprise features for your team, or [start prototyping with a free Qdrant cluster](https://cloud.qdrant.io/).
@@ -0,0 +1,92 @@
---
title: "Convolve 4.0 - IIT Hackathon Winners"
draft: false
slug: iit-hack-winners
short_description: "Discover the winners of Qdrant’s Convolve 4.0 - IIT Hackathon, where developers built innovative vector search applications using Multi-Agent Intelligence Systems - from e-commerce to healthcare, trends, and more."
preview_image: blog/iit-hack-2026/convolve.png
social_preview_image: blog/iit-hack-2026/convolve.png
date: 2026-02-27
author: Manas Chopra
featured: true
tags:
- news
- blog
---
Builders from across India came together for Convolve 4.0 - A Pan IIT AI/ML Hackathon, hosted by IIT Madras, to develop impactful, real-world AI systems across critical domains. The hackathon focused on multi-agent systems, retrieval-augmented generation (RAG), vector memory, multimodal intelligence, and production-ready AI pipelines.
Participants built solutions spanning healthcare and medical reasoning, disaster response and crisis intelligence, climate and risk monitoring, legal-tech and governance systems, misinformation detection, public safety, infrastructure auditing, education, and civic-tech platforms. Many projects emphasized persistent AI memory, domain-aware guardrails, spatial intelligence, and collaborative agent architectures designed for long-term, scalable deployment.
With a prize pool of around ₹2 lakh, Convolve 4.0 highlighted how advanced AI/ML systems—powered by intelligent retrieval, multimodal data, and agentic workflows - can address complex societal and national-scale challenges.
👉 Full hackathon details: [Hackathon page](https://unstop.com/hackathons/convolve-40-a-pan-iit-aiml-hackathon-open-to-all-iit-guwahati-1609886)
Let’s dive into the top projects and what we liked about each of them.
---
## 🏆 Overall Winners
### 🥇 1st Place: Masthishq (Krishna Koushik Padigala)
**Total Prize: `₹1,25,000 Cash and Qdrant Goodie box`**
**What it is:** Masthishq is a multimodal AI agent that acts as a cognitive prothesis. By leveraging Computer Vision (FaceNet, YOLO), Vector search (Qdrant), and Large Language Models (Llama 3 via Groq), the system provides real-time, context-aware assistance. It identifies people and objects in the patient's environment, retrieves associated long-term memories ("This is Jill, your sister"), and engages in empathetic, looped conversations to soothe the patient. The system includes a Patient App featuring an animated avatar for accessibility and a Caregiver Dashboard for secure memory management.
Repo: [Masthishq Repo](https://github.com/krishk2/Masthishq/)
### 📸 Project Showcase
![Masthishq Patient Interface](/blog/iit-hack-2026/ss-1.png)
---
### 🥈 2nd Place: SignalWeave - A Temporal AI Memory System for Emerging Trendsk (T Mohamed Yaser)
**Total Prize: `₹75,000 Cash and Qdrant Goodie box`**
**What it is:** SignalWeave is a temporal AI memory system for detecting emerging trends from weak signals - small, scattered mentions that typically go unnoticed. It continuously ingests content, converts each signal into vector embeddings, clusters related signals, and accumulates them over time. Instead of treating data as snapshots, it promotes trends only after they gather enough evidence, enabling early detection through temporal merging and growth scoring.
Qdrant serves as the system’s persistent vector memory layer. It stores signal embeddings and evolving cluster centroids, enables fast cosine similarity search for merging and retrieval, and preserves historical context across runs, making temporal trend accumulation possible.
**Demo App:** [SignalWeave](https://signalweave.vercel.app/)
**Repo:** [SignalWeave Repo](https://github.com/Yaser-123/signalweave)
### 📸 Project Showcase
![SignalWeave Trend Detection](/blog/iit-hack-2026/ss-2.png)
---
### 🥉 3rd Place: Demeter (Debarghya Das)
**Total Prize: `₹50,000 Cash and Qdrant Goodie box`**
**What it is:** Demeter is an autonomous multi-agent system for hydroponic farm management that creates a digital twin of the growing environment and makes expert-level decisions without constant human supervision. It ingests multimodal data - crop images and sensor readings (pH, EC, temperature, humidity) - and fuses them into a unified vector representation called a Farm Memory Unit (FMU), which is stored in Qdrant for long-term memory and retrieval.
Qdrant acts as the system’s persistent cognitive backbone, storing fused plant states, agronomic knowledge for RAG, and plant biography histories. Specialized agents (Water, Atmospheric, Doctor, Supervisor) retrieve historical cases, similar environmental states, and scientific references from Qdrant to make grounded decisions. A Contextual Bandit then selects the safest optimal strategy, while feedback from a Judge agent updates Qdrant with outcome data - allowing Demeter to continuously learn which interventions work best over time.
**Repo:** [Demeter Repo](https://github.com/Deb044/Demeter)
### 📸 Project Showcase
![Demeter Dashboard](/blog/iit-hack-2026/ss-3.png)
---
## Why These Projects Matter
These winners stand out because they treat memory and retrieval as foundational infrastructure - not add-ons. Instead of one-off AI outputs, they built systems that remember, reason over history, and improve over time.
- Masthishq uses multimodal vector memory to reconnect dementia patients with people and past context.
- SignalWeave accumulates weak signals over time, using persistent clustering and hybrid search to detect emerging trends early.
- Demeter stores multimodal farm states and historical interventions to power safe, learning-driven hydroponic automation.
Across all three, Qdrant acts as the long-term memory layer, enabling grounded decisions, reuse of past knowledge, and continuous improvement.
As always, many strong projects competed, but these winners highlight what’s possible when AI systems are designed to truly remember.
## What’s Next
A huge thank you to everyone who built, mentored, or joined. You’ve proven that with Qdrant, the possibilities are limitless.
👉 Join the community: [qdrant.tech/community](https://qdrant.tech/community)
@@ -125,7 +125,7 @@ final_results = reranked[:5]
Not every clause is created equal. Legal professionals often care more about specific provisions, jurisdictions, or case types, for example.
Qdrant's [Score Boosting Reranker](https://qdrant.tech/documentation/concepts/hybrid-queries/#score-boosting) lets you integrate domain-specific logic (e.g., jurisdiction or recent cases) directly into search rankings, ensuring results align precisely with legal business rules.
Qdrant's [Score Boosting Reranker](/documentation/concepts/search-relevance/#score-boosting) lets you integrate domain-specific logic (e.g., jurisdiction or recent cases) directly into search rankings, ensuring results align precisely with legal business rules.
```json
POST /collections/legal-docs/points/query
+5 -3
View File
@@ -4,6 +4,8 @@ draft: false
slug: qdrant-1.15.x
short_description: "Smarter Quantization, Healing Indexes, and Multilingual Text Filtering"
description: "Qdrant v1.15 release presents new Quantization Features, advanced Full-Text filtering and a bunch of performance optimizations"
preview_image: /blog/qdrant-1.15.x/social_preview.jpg
social_preview_image: /blog/qdrant-1.15.x/social_preview.jpg
date: 2025-07-18T00:00:00-08:00
author: Derrick Mwiti
featured: true
@@ -211,7 +213,7 @@ The above will match:
## MMR Reranking
We introduce [Maximal Marginal Relevance (MMR)](/documentation/concepts/hybrid-queries/#maximal-marginal-relevance-mmr) reranking to balance relevance and diversity.
We introduce [Maximal Marginal Relevance (MMR)](/documentation/concepts/search-relevance/#maximal-marginal-relevance-mmr) reranking to balance relevance and diversity.
MMR works by selecting the results iteratively, by picking the item with the best combination of similarity to the query and dissimilarity to the already selected items.
It prevents your top-k results from being redundant and helps surface varied but relevant answers, particularly in dense datasets with overlapping entries.
@@ -223,7 +225,7 @@ It prevents your top-k results from being redundant and helps surface varied but
Let’s say you’re building a knowledge assistant or semantic document explorer in which a single query can return multiple highly similar queries.
For instance, searching “climate change” in a scientific paper database might return several similar paragraphs.
You can diversify the results with [Maximal Marginal Relevance (MMR)](/documentation/concepts/hybrid-queries/#maximal-marginal-relevance-mmr).
You can diversify the results with [Maximal Marginal Relevance (MMR)](/documentation/concepts/search-relevance/#maximal-marginal-relevance-mmr).
Instead of returning the top-k results based on pure similarity, MMR helps select a diverse subset of high-quality results.
This gives more coverage and avoids redundant results, which is helpful in dense content domains such as academic papers, product catalogs, or search assistants.
@@ -281,7 +283,7 @@ This modification, in combinations with [incremental HNSW indexing](/blog/qdrant
### HNSW Graph connectivity estimation
Qdrant builds [addtitional HNSW links](/articles/filtrable-hnsw/) to ensure that filtered searches are performed fast and accurate.
Qdrant builds [addtitional HNSW links](/articles/filterable-hnsw/) to ensure that filtered searches are performed fast and accurate.
It does, however, introduce an overhead for indexing complexity, especially when the number of payload indexes is large.
With v1.15, Qdrant introduces an optimization, which quickly estimates graph connectivity before creating additional links.
+4 -2
View File
@@ -4,6 +4,8 @@ draft: false
slug: qdrant-1.16.x
short_description: "v1.16 of Qdrant focuses on tiered multitenancy with tenant promotion and disk-efficient vector search."
description: "v1.16 of Qdrant focuses on tiered multitenancy with tenant promotion, disk-efficient vector search with inline storage, and improved filtered vector search with ACORN."
preview_image: /blog/qdrant-1.16.x/social_preview.jpg
social_preview_image: /blog/qdrant-1.16.x/social_preview.jpg
date: 2025-11-19T00:00:00-08:00
author: Abdon Pijpelink
featured: true
@@ -50,9 +52,9 @@ To use Tiered Multitenancy, after [setting up a collection with a shared fallbac
![Section 2](/blog/qdrant-1.16.x/section-2.png)
To enhance the scalability and speed of vector search, Qdrant employs a graph-based index structure known as [HNSW (Hierarchical Navigable Small World)](/documentation/concepts/indexing/#vector-index). While traditional HNSW is primarily designed for unfiltered searches, Qdrant has addressed this limitation by implementing a [filterable HSNW index](/articles/filtrable-hnsw/). This innovative approach extends the HNSW graph with additional edges that correspond to indexed payload values. This enables Qdrant to maintain search quality even with high filtering selectivity, without introducing any runtime overhead during the search process.
To enhance the scalability and speed of vector search, Qdrant employs a graph-based index structure known as [HNSW (Hierarchical Navigable Small World)](/documentation/concepts/indexing/#vector-index). While traditional HNSW is primarily designed for unfiltered searches, Qdrant has addressed this limitation by implementing a [filterable HSNW index](/articles/filterable-hnsw/). This innovative approach extends the HNSW graph with additional edges that correspond to indexed payload values. This enables Qdrant to maintain search quality even with high filtering selectivity, without introducing any runtime overhead during the search process.
Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen when you use a combination of high cardinality filters, leading to the HNSW graph becoming [disconnected](/documentation/concepts/indexing/#filtrable-index). It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where filtering criteria are not known in advance.
Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen when you use a combination of high cardinality filters, leading to the HNSW graph becoming [disconnected](/documentation/concepts/indexing/#filterable-index). It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where filtering criteria are not known in advance.
To address these limitations, in version 1.16 we are introducing support for [ACORN](/documentation/concepts/search/#acorn-search-algorithm), based on the ACORN-1 algorithm described in the paper [ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data](https://arxiv.org/abs/2403.04871). With ACORN enabled, Qdrant not only traverses direct neighbors (the first hop) in the HNSW graph but also examines neighbors of neighbors (the second hop) if the direct neighbors have been filtered out. This enhancement improves search accuracy at the expense of performance, especially when multiple low-selectivity filters are applied.
@@ -0,0 +1,143 @@
---
title: "Qdrant 1.17 - Relevance Feedback & Search Latency Improvements"
draft: false
slug: qdrant-1.17.x
short_description: "Version 1.17 of Qdrant features a new Relevance Feedback Query and search latency improvements."
description: "Version 1.17 of Qdrant features a new Relevance Feedback Query, search latency improvements, and better operational observability."
preview_image: /blog/qdrant-1.17.x/social_preview.jpg
social_preview_image: /blog/qdrant-1.17.x/social_preview.jpg
date: 2026-02-20T00:00:00-08:00
author: Abdon Pijpelink
featured: true
tags:
- vector search
- relevance feedback
- search performance
- observability
---
[**Qdrant 1.17.0 is out!**](https://github.com/qdrant/qdrant/releases/tag/v1.17.0) Let’s look at the main features for this version:
**Relevance Feedback Query:** Improve the quality of search results by incorporating information about their relevance.
**Search Latency Improvements:** Manage search latency with new tools, such as an update queue and delayed fan-outs, as well as many internal search performance improvements.
**Greater Operational Observability:** Better insights into operational metrics and faster troubleshooting with a new cluster-wide telemetry API and segment optimization monitoring.
## Relevance Feedback Query
![Section 1](/blog/qdrant-1.17.x/section-1.png)
Crafting queries is hard: users often struggle to precisely formulate search queries. At the same time, judging the relevance of a given search result is often much easier. Retrieval systems can leverage this [relevance feedback](/articles/search-feedback-loop/) to iteratively refine results toward user intent.
This release introduces a new [Relevance Feedback Query](/documentation/concepts/search-relevance/#relevance-feedback) as a scalable, vector‑native approach to incorporating relevance feedback. The Relevance Feedback Query uses a small amount of model‑generated feedback to guide the retriever through the entire vector space, effectively nudging search toward “more relevant” results without requiring expensive loops, expensive retrievers, or human labeling. This enables the engine to traverse billions of vectors with improved recall without having to retrain models.
This method works by collecting lightweight feedback on just a few top results, creating “context pairs” of more‑ and less‑relevant examples. These pairs define a signal that adjusts the scoring function during the next retrieval pass. Instead of rewriting queries or rescoring large batches of documents, Qdrant modifies how similarity is computed. Experiments demonstrate substantial gains, especially when pairing expressive retrievers with strong feedback models. For the methodology and experiments behind this feature, see our article [Relevance Feedback in Qdrant](/articles/relevance-feedback). To get started, refer to the [documentation](/documentation/concepts/search-relevance/#relevance-feedback).
<figure>
<img src="/blog/qdrant-1.17.x/relevance-feedback-overview.png">
<figcaption>
Feedback-based scoring combines a candidate’s similarity to a query with its relative distances (delta) to the positive and negative items in context pairs.
</figcaption>
</figure>
## Search Latency Improvements
![Section 2](/blog/qdrant-1.17.x/section-2.png)
This release includes several changes that reduce search latency. To improve query response times in environments with high write loads, Qdrant can now be configured to avoid creating large unoptimized segments. Additionally, delayed fan-outs help reduce tail latency by querying a second replica if the first does not respond within a configurable latency threshold.
### Search Latency Under Write Load
A common pattern with vector search engines like Qdrant involves bulk uploads. For example, periodically refreshing data from an external source of truth using nightly batch updates. Newly ingested data needs to be indexed, which is a resource-intensive operation. When the data ingestion rate exceeds the indexing rate, this can lead to issues such as:
- Back-pressure and rejected update operations due to a full update queue.
- Slow queries over data that has not yet been indexed.
This release addresses these issues by changing how data is ingested. Shards still process data through the familiar stages: WAL persistence, queued updates, application to unoptimized segments, and eventual full indexing, but two new features reshape how systems behave under heavy write load.
A new [update queue](/documentation/guides/low-latency-search/#query-indexed-data-only) tracks up to one million pending changes. When the queue fills, back pressure slows incoming writes, preventing runaway load and helping clusters stay stable even during large batch operations or recovery after downtime.
For applications that demand consistently low-latency search, indexed‑only mode ensures queries touch only fully indexed segments. A side-effect of using indexed-only queries was that they could temporarily hide the newest updates, before they were indexed. A new [`prevent_unoptimized` optimizer setting](/documentation/guides/low-latency-search/#query-indexed-data-only) solves this by throttling updates to match the indexing rate, reducing the creation of large unoptimized segments.
Together, these features give developers tighter control over write throughput, indexing behavior, and search performance, especially in high‑volume environments.
### Reduced Tail Latency with Delayed Fan-Outs
By default, a search operation queries a single replica of each shard within a collection. If one of these replicas responds slowly due to load or network issues, this negatively impacts the overall search latency. This phenomenon, where a single slow replica increases the 95th or 99th percentile latency of the entire system, is known as “tail latency.” High tail latency can noticeably degrade the user experience.
To mitigate tail latency for read operations, this release introduces a new [delayed fan-out](/documentation/guides/low-latency-search/#use-delayed-fan-outs) feature. With delayed fan-outs, if the initial request to a replica exceeds a configurable latency threshold, an additional read request is sent to another replica, and Qdrant will use the first available response. Delayed fan-outs help your application provide a consistent, low latency experience to end-users.
## Greater Operational Observability
![Section 3](/blog/qdrant-1.17.x/section-3.png)
We are continuously working to enhance the operational observability of Qdrant clusters. In this release, we introduce two new features: a new cluster-wide telemetry API and segment optimization monitoring.
### Cluster-Wide Telemetry
Qdrant’s API exposes a `/telemetry` endpoint which provides information about the current state of a peer in a cluster, including the number of vectors, shards, and other useful information. However, obtaining a complete view of the entire cluster using this endpoint is not straightforward, requiring querying each peer and piecing together a complete view yourself.
In version 1.17, we’re introducing a new [`/cluster/telemetry` endpoint](/documentation/guides/monitoring/#cluster-wide-telemetry). This API provides information about all peers in a cluster, offering insights into cluster-wide operations such as leader elections, resharding, and shard transfers.
### Segment Optimization Monitoring
Optimization is a background process where Qdrant removes data marked for deletion, merges segments, and creates indexes. To improve visibility into this process, this release introduces [segment optimization monitoring capabilities](/documentation/concepts/optimizer/#optimization-monitoring).
A new `/collections/{collection_name}/optimizations` API endpoint provides cluster-wide information about the current optimization status, as well as detailed information for current and past optimization operations. Because the output of the API can be verbose, we’ve added a new Optimizations tab to the Collections interface in the Web UI that makes it easier to analyze the data. Here, you can find an overview of the current optimization status, a timeline of current and past optimization operations, and a breakdown of the tasks in a specific cycle and their durations.
<figure>
<img width="75%" src="/blog/qdrant-1.17.x/optimizer-web-ui.png">
<figcaption>
The new user interface in the Web UI provides an overview of the current cluster-wide optimization status and a timeline of current and past optimization cycles.
</figcaption>
</figure>
## Redesigned Web UI Point Search
![Section 5](/blog/qdrant-1.17.x/section-4.png)
[Web UI](/documentation/web-ui/) is Qdrant’s user interface for managing deployments and collections. It enables you to create and manage collections, run API calls, import sample datasets, and learn about Qdrant's API through interactive tutorials.
Many people have been asking about point filtering in web UI. And now it's back, better than ever. In this release, we have redesigned the point search interface in the Web UI to make exploring your data and discovering relevant points easier and more intuitive. The new two-field layout enables searching for points similar to another point, filtering by payload values, and finding points by ID.
<figure>
<img src="/blog/qdrant-1.17.x/web-ui-search.png">
<figcaption>
The redesigned point search interface in the Web UI provides a way to find points similar to another point and filter on payload values.
</figcaption>
</figure>
## Honorable Mentions
![Section 5](/blog/qdrant-1.17.x/section-5.png)
As an open source project, we welcome contributions from the Qdrant community. This release features two contributions from community members:
- Not all payload field indexes are used in combination with dense vector queries. With this release, you can [specify whether individual payload field indexes should be reflected in the HNSW index](/documentation/concepts/indexing/#disable-the-creation-of-extra-edges-for-payload-fields).
- A new API endpoint is available to [list all user-defined shard keys](/documentation/guides/distributed_deployment/#user-defined-sharding).
Additionally, this release adds the following features:
- Upserts now support an [update mode](/documentation/concepts/points/#update-mode) for insert-only or update-only operations.
- To speed up the recovery of the replicas after they’ve been down, shards will [increase the size of their write-ahead log](https://github.com/qdrant/qdrant/pull/7834) when they detect that one of their remote replicas is unavailable.
- Reciprocal Rank Fusion (RRF) combines multiple query results into one list, but its default equal weighting can let weaker rankers dilute stronger ones. [Weighted RRF](/documentation/concepts/hybrid-queries/#reciprocal-rank-fusion-rrf) in Qdrant 1.17 addresses this by letting you assign weights to individual queries.
- A new [user interface in the Web UI enables resharding collections](https://github.com/qdrant/qdrant-web-ui/pull/341) on Qdrant Cloud.
- Qdrant now supports [audit logging](/documentation/guides/security/#audit-logging) to track all API operations that require authentication or authorization.
- [External provider API keys for inference requests](/documentation/concepts/inference/#external-embedding-model-providers) can now be provided in the request header.
For a full list of all changes in version 1.17, please refer to the [change log](https://github.com/qdrant/qdrant/releases/tag/v1.17.0).
## Upgrading to Version 1.17
![Section 6](/blog/qdrant-1.17.x/section-6.png)
On Qdrant Cloud, navigate to the Cluster Details screen and select Version 1.17 from the dropdown menu. The upgrade process may take a few moments.
We recommend upgrading versions one by one. On Qdrant Cloud, this is done automatically when you select the target version. If you are self-hosting, upgrade to the latest patch version of each intermediate minor version first, for example 1.15.5->1.16.3->1.17.0.
## Engage
![Section 7](/blog/qdrant-1.17.x/section-7.png)
We would love to hear your thoughts on this release. If you have any questions or feedback, join our [Discord](https://discord.gg/qdrant) or create an issue on [GitHub](https://github.com/qdrant/qdrant/issues).
@@ -0,0 +1,68 @@
---
title: "Qdrant Academy Expands with Official Certification"
draft: false
slug: qdrant-certification-launch
short_description: "Qdrant Academy launched with first course, Qdrant Essentials"
description: "Master the art of production-grade retrieval with Qdrant Academy’s new certification. Earn credentials, score exclusive swag, and level up your engineering skills."
preview_image: /blog/qdrant-certification-launch/hero-graphic.png
social_preview_image: /blog/qdrant-certification-launch/hero-graphic.png
date: 2026-01-28
author: Neil Kanungo
featured: true
tags:
- Community
- Academy
---
Since we first announced **[Qdrant Academy](https://qdrant.tech/course/)**, our mission has been to provide developers with more than just documentation. We wanted to build a structured path to mastering vector search. As the AI search landscape matures, the distinction between a simple storage layer and a high-performance vector search engine has become the defining factor in production-grade RAG and recommendation systems.
Today, we are thrilled to take the next step in that mission. It’s time to move from learning to proving your expertise with the launch of our first official certification.
### Introducing the "Qdrant Essentials" Certification
The [Qdrant Essentials course](https://qdrant.tech/course/essentials/) has already helped thousands of developers understand the "why" behind high-dimensional search. Now, you can officially validate that knowledge.
By completing the course and passing the final exam at **[train.qdrant.dev](https://train.qdrant.dev)**, you’ll earn a digital credential that proves you can architect search systems that are as efficient as they are accurate.
#### What the Essentials Track Covers:
* **Engine Architecture:** Deep dives into HNSW, distance metrics, and collection structures.
* **Precision Filtering:** Mastering payload-based filtering without sacrificing search speed.
* **Hybrid Search:** Implementing a mix of dense and sparse vectors for superior retrieval.
* **Production Optimization:** Utilizing quantization and rescoring to scale your engine efficiently.
### Why Get Certified?
In a field as fast-moving as AI, "knowing a bit of Python" isn't enough. Moving from a prototype to a production-ready system requires specialized engineering judgment. Becoming **#QdrantCertified** can be a game-changer for your career:
* **Verified Expertise:** It proves you understand the critical trade-offs—like balancing latency vs. accuracy—that separate a hobbyist project from enterprise infrastructure.
* **Career Differentiation:** As companies hunt for RAG and Agentic AI experts, this badge signals that you can handle high-scale vector search, reducing your onboarding time and making you an immediate asset.
* **Standardized Knowledge:** You aren't just learning from assorted tutorials; you’re learning the industry standard for high-performance retrieval directly from the creators of Qdrant.
* **Engineering Authority:** Gain the confidence to lead internal AI workshops or architect your company's next-gen search platform using verified best practices.
### Get Certified. Get Swag.
We want to see those certificates! To celebrate the launch of our certification platform, we’re sending out some exclusive gear to our early achievers.
> **The first 30 people** to post their Qdrant Essentials certification to LinkedIn with the hashtag **#QdrantCertified** will receive a free Qdrant swag pack.
It’s simple: Learn, pass the exam at [train.qdrant.dev](https://train.qdrant.dev), and share your success with the community to claim your prize.
### More Courses Launching Soon
The "Essentials" course is just the foundation. Qdrant Academy is expanding rapidly to support developers at every stage of their journey:
#### The 2-Hour Beginner Launchpad
Coming soon, we are launching a **2-hour Basic Course**. This is designed for those who need a high-impact, low-time-commitment introduction to the world of vector search. You’ll go from "What is an embedding?" to "I have a running search engine" in a quick yet comprehensive Qdrant intro.
#### Advanced Retrieval Topics
For the power users, our upcoming **Multivectors Course** will tackle the cutting edge of retrieval. It will focus on Late Interaction models (like ColBERT), and will cover sophisticated retrieval with MUVERA. You’ll learn how to handle token-level embeddings to achieve incredible retrieval precision for complex datasets.
## Ready to Level Up?
Come grow with Qdrant, and prove your knowledge with Qdrant Certifications:
1. **Learn:** Head over to the [Qdrant Essentials course](https://qdrant.tech/course/essentials/).
2. **Certify:** Take the exam and claim your badge at **[train.qdrant.dev](https://train.qdrant.dev)**.
3. **Win:** Post it on LinkedIn with **#QdrantCertified** and grab your swag.
As always, happy coding!
@@ -30,5 +30,5 @@ What this means for you:
Get started by [signing up for a Qdrant Cloud account](https://cloud.qdrant.io). And learn more about Qdrant Cloud in our [docs](/documentation/cloud/).
<video autoplay="true" loop="true" width="100%" controls><source src="/blog/qdrant-cloud-on-azure/azure-cluster-deployment-short.mp4" type="video/mp4"></video>
<video autoplay="true" loop="true" width="100%" controls><source src="https://storage.googleapis.com/qdrant-landing/files/blog/azure-cluster-deployment-short.mp4" type="video/mp4"></video>
@@ -0,0 +1,61 @@
---
draft: false
title: "New DeepLearning.AI Course on Multi-Vector Image Retrieval with ColPali and MUVERA"
short_description: "Free course on advanced image retrieval using multi-vector techniques."
description: "Join Qdrant and DeepLearning.AI for a free course on multi-vector image retrieval, featuring ColPali, ColBERT, and MUVERA for production-ready visual RAG systems."
preview_image: /blog/qdrant-deeplearning-ai-multi-vector-image-retrieval/preview.png
social_preview_image: /blog/qdrant-deeplearning-ai-multi-vector-image-retrieval/preview.png
date: 2025-12-11T17:00:00Z
author: Kacper Łukawski
featured: true
tags:
- DeepLearning.AI
- Multi-Vector Search
- ColPali
- ColBERT
- Image Retrieval
- Vector Search
- Retrieval-Augmented Generation
- MUVERA
---
We're thrilled to announce our latest collaboration with DeepLearning.AI: [Multi-Vector Image Retrieval](https://www.deeplearning.ai/short-courses/multi-vector-image-retrieval/). Building on the success of our previous course on retrieval optimization, this intermediate-level course takes you deeper into advanced search techniques that are transforming how AI systems understand and retrieve visual content.
Led once again by Qdrant's Kacper Łukawski, Senior Developer Advocate, this free course is designed for AI builders working with multi-modal data who want to implement cutting-edge image retrieval in their applications.
## Why This Collaboration Matters
Our continued partnership with DeepLearning.AI reflects our commitment to advancing the field of vector search through education. While our first course introduced developers to the fundamentals of retrieval optimization, this intermediate course tackles one of the most challenging problems in modern AI: fine-grained matching between text queries and visual content.
Multi-vector approaches represent a significant leap forward from traditional single-vector embeddings. By encoding images as multiple vectors, one for each visual patch, these techniques enable far more precise and nuanced search capabilities. This is particularly powerful for applications requiring detailed visual understanding, from e-commerce product search to document analysis.
<iframe width="560" height="315" src="https://www.youtube.com/embed/5lR0V1PUZ10?si=IdTEKlUHVbGzgJD-" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
## What You'll Learn
This course provides comprehensive coverage of multi-vector retrieval techniques and their practical implementation:
- Understand the fundamentals of multi-vector embeddings and how patch-level representations dramatically improve search accuracy compared to single-vector approaches.
- Remind ColBERT for text retrieval and learn how its late-interaction architecture enables fine-grained semantic matching.
- Master ColPali, a vision language model that generates multi-vector representations for images, enabling detailed text-to-image search.
- Learn optimization techniques including quantization and pooling to minimize memory requirements while maintaining search quality.
- Discover MUVERA's approach to converting high-dimensional multi-vector embeddings into efficient representations for production systems.
- Build complete multi-modal RAG pipelines that combine ColPali retrieval with HNSW search algorithms for real-world applications.
## Who Should Enroll
This course is designed for AI builders working with multi-modal data who want to implement advanced image retrieval systems. You should have foundational familiarity with Python and vector embeddings to get the most from this intermediate-level content.
Whether you're building visual search applications, enhancing document retrieval systems, or developing multi-modal AI assistants, this course provides the practical knowledge you need to implement state-of-the-art retrieval techniques.
### At a Glance:
- **Speaker**: Kacper Łukawski, Senior Developer Advocate at Qdrant
- **Level**: Intermediate
- **Cost**: Free
- **Location**: Online
- **Duration**: 1 hour 33 minutes
## How to Enroll
[Enroll via the DeepLearning.AI website](https://www.deeplearning.ai/short-courses/multi-vector-image-retrieval/?utm_campaign=qdrant-launch&utm_medium=qdrant&utm_source=partner-promo) and start building advanced multi-vector retrieval systems today.
@@ -0,0 +1,64 @@
---
title: "Qdrant Meets Google Gemini Embedding 2"
draft: false
slug: qdrant-gemini-embedding-2
short_description: "Qdrant with Gemini Embedding 2 enables multimodal search across text, images, video, audio, and PDFs all in one vector space. Read more to learn how to get started."
preview_image: /blog/qdrant-gemini-embedding-2/gemini-2-hero.png
social_preview_image: /blog/qdrant-gemini-embedding-2/gemini-2-hero.png
date: 2026-03-10
author: Neil Kanungo
featured: true
tags:
- Community
- Embeddings
---
Until now, building a search system that understands both what a document says *and* what its images, videos, or audio convey could be complex and limited. It meant stitching together multiple embedding models and writing code to reconcile results across modalities. That pipeline complexity may have been the single biggest barrier to production-grade multimodal retrieval.
Today, Google launched [**Gemini Embedding 2**](https://ai.google.dev/gemini-api/docs/embeddings) in Public Preview, the first fully multimodal embedding model in the Gemini family, and Qdrant supports it from day one. This post explains what the model offers, why Qdrant is a natural fit, and how to get started.
### What is Gemini Embedding 2?
Gemini Embedding 2 is Google's first embedding model capable of mapping **text, images, video, audio, *and* PDF documents** (including interleaved combinations of these) into a single, unified vector space. Built on the Gemini architecture with support for over 100 languages, it's now available through the Gemini API.
### What makes it significant?
**Natively multimodal.** Rather than converting images to captions or audio to transcripts before embedding, Gemini Embedding 2 processes each modality directly. This eliminates lossy intermediate steps and preserves the rich semantic information that lives in visual composition, audio tone, or document layout. Information that text-only pipelines discard.
**Flexible output dimensions via Matryoshka Representation Learning (MRL).** The model supports output dimensions ranging from 128 to 3072, with recommended sizes of 756, 1536, and the default 3072\. MRL means the most important semantic information is concentrated in the first dimensions of the vector. You can truncate embeddings to a smaller size for faster search and lower storage costs, then use the full-size embedding only when you need maximum precision.
**Broad format support with clear limits.** The model handles PNG and JPEG images (up to 6 per request), MP4/MOV video (up to 128 seconds), MP3/WAV audio (up to 80 seconds), and PDFs (up to 6 pages), with a maximum input of 8192 tokens.
**Strong benchmark performance.** Gemini Embedding 2 improves over its predecessor and ranks in the top 5 on the MTEB Multilingual leaderboard for text. Across most other modalities, it achieves state-of-the-art results among proprietary models.
### Why Qdrant is Right for This
Qdrant is built to handle exactly the kind of workloads that a multimodal embedding model like Gemini Embedding 2 enables. Here's where the fit is particularly strong:
**A Unified Collection for All Your Modalities.** Because Gemini Embedding 2 maps every modality into the same vector space, you can store text embeddings, image embeddings, video embeddings, audio embeddings, and PDF embeddings all in a single Qdrant collection. No separate indexes, no reconciliation logic. A text query retrieves relevant video clips. An image query surfaces matching documents. The vectors speak the same language.
**Named Vectors for Hybrid Strategies.** For more sophisticated architectures, Qdrant's named vectors let you store multiple vector representations per point within the same collection. You might store a Gemini Embedding 2 vector alongside a sparse BM25 vector for hybrid search, or keep separate embeddings for different aspects of the same data point. Named vectors keep everything in one place while giving you full control over which vector to query against.
**MRL-Friendly Architecture.** Gemini Embedding 2's flexible dimensionality pairs beautifully with Qdrant's support for multi-stage retrieval. You can build a two-pass search pipeline: first, use lower-dimensional embeddings (like 756 dimensions) for fast candidate retrieval over your entire corpus, then rescore the top candidates with the full 3072-dimensional embeddings for maximum precision. Qdrant's named vectors make this pattern straightforward to implement within a single collection.
**Production-Ready at Scale.** Qdrant is purpose-built in Rust for speed and reliability. Features like built-in quantization (scalar, product, and binary) can dramatically reduce memory usage for large 3072-dimensional vectors. Payload filtering lets you combine vector similarity with structured metadata queries — for example, searching for the most relevant video clips *from the last 30 days* or *in a specific language*. And Qdrant Cloud provides managed horizontal scaling for production deployments.
### How this Benefits You
The combination of a truly multimodal embedding model and a high-performance vector database opens up use cases that were previously impractical to build:
**Multimodal RAG.** Build retrieval-augmented generation systems that pull context from text documents, product images, tutorial videos, and customer support audio recordings, all from a single query against a single Qdrant collection.
**Cross-modal semantic search.** A user describes what they're looking for in text, and the system retrieves matching images, videos, or PDF pages. Or they upload a photo and find relevant documents. The shared vector space makes this work without any custom bridging logic.
**Unified content recommendation.** Recommend a mix of articles, videos, podcasts, and documents based on a user's interaction history, without maintaining separate recommendation engines per content type.
**Multilingual document intelligence.** With 100+ language support, embed and search across documents in any language. Combine this with PDF embedding to build document retrieval systems that work across languages and formats without OCR or translation pipelines.
### Start Building Today
Qdrant enables you to use Gemini Embedding 2 today. Check out the full documentation with working examples in Python and JavaScript at our [**Gemini integration guide**](/documentation/embeddings/gemini/).
Gemini Embedding 2 is compatible with both Cloud and Open Source deployments. To get started with Qdrant Cloud, [sign up for a free tier today](https://cloud.qdrant.io/signup).
Multimodal search just got a lot simpler. With Gemini Embedding 2 handling the embeddings and Qdrant handling the storage and retrieval, you can build with the next wave of embedding models and vector search capabilities today.
@@ -27,7 +27,7 @@ The rise of generative AI in the last few years has shone a spotlight on vector
## What sets Qdrant apart?
To meet the needs of the next generation of AI applications, Qdrant has always been built with four keys in mind: efficiency, scalability, performance, and flexibility. Our goal is to give our users unmatched speed and reliability, even when they are building massive-scale AI applications requiring the handling of billions of vectors. We did so by building Qdrant on Rust for performance, memory safety, and scale. Additionally, [our custom HNSW search algorithm](/articles/filtrable-hnsw/) and unique [filtering](/documentation/concepts/filtering/) capabilities consistently lead to [highest RPS](/benchmarks/), minimal latency, and high control with accuracy when running large-scale, high-dimensional operations.
To meet the needs of the next generation of AI applications, Qdrant has always been built with four keys in mind: efficiency, scalability, performance, and flexibility. Our goal is to give our users unmatched speed and reliability, even when they are building massive-scale AI applications requiring the handling of billions of vectors. We did so by building Qdrant on Rust for performance, memory safety, and scale. Additionally, [our custom HNSW search algorithm](/articles/filterable-hnsw/) and unique [filtering](/documentation/concepts/filtering/) capabilities consistently lead to [highest RPS](/benchmarks/), minimal latency, and high control with accuracy when running large-scale, high-dimensional operations.
Beyond performance, we provide our users with the most flexibility in cost savings and deployment options. A combination of cutting-edge efficiency features, like [built-in compression options](/documentation/guides/quantization/), [multitenancy](/documentation/guides/multiple-partitions/) and the ability to [offload data to disk](/documentation/concepts/storage/), dramatically reduce memory consumption. Committed to privacy and security, crucial for modern AI applications, Qdrant now also offers on-premise and hybrid SaaS solutions, meeting diverse enterprise needs in a data-sensitive world. This approach, coupled with our open-source foundation, builds trust and reliability with engineers and developers, making Qdrant a game-changer in the vector database domain.
@@ -0,0 +1,53 @@
---
draft: false
title: "We Raised $50M to Build Composable Vector Search as Core Infrastructure"
short_description: "Qdrant raises $50M in Series B funding to scale composable vector search from edge devices to supercomputers."
description: "Qdrant announces $50M in Series B funding led by AVP to build composable vector search as foundational infrastructure for production AI — from agentic workflows to edge devices."
preview_image: /blog/series-b-announcement/series-b-funding.jpg
social_preview_image: /blog/series-b-announcement/series-b-funding.jpg
date: 2026-03-12
author: "Andre Zayarni"
featured: true
tags:
- series b
- funding
- composable vector search
- infrastructure
- agentic ai
- edge computing
- rust
- qdrant
---
Today we're announcing $50 million in Series B funding, led by AVP, with participation from Bosch Ventures, Unusual Ventures, Spark Capital, and 42CAP.
### Retrieval Is on the Critical Path of Every AI System
Every serious AI workload — RAG, agents, multimodal search — depends on retrieving the right information, at the right time, under real constraints. Teams prototype with whatever is convenient, then hit walls in production: indexes that stall under writes, filtering applied after search instead of during it, tail latencies that spike under load. These aren't configuration problems. They're architectural ones. And they're why we started Qdrant.
Our design philosophy: we are building fundamental infrastructure. The kind that lasts years, maybe decades. Think Linux kernel, not SaaS wrapper. That conviction is why Qdrant is built from the ground up in Rust — because infrastructure on the critical path of production AI cannot afford garbage collection pauses, memory safety issues, or the performance unpredictability of managed runtimes. It's why we control the stack down to assembly, rebuilding any layer that doesn't meet the standard. And it's why Qdrant delivers predictable low tail latency at billion-scale, running consistently from [edge devices](https://github.com/qdrant/qdrant-edge-demo) to bare-metal supercomputers like [Aurora at Argonne National Laboratory](https://arxiv.org/abs/2509.12384).
We’re proud of the enterprises running Qdrant in production, including Canva, Bazaarvoice, HubSpot, Roche, Bosch, and OpenTable. We surpassed 250 million downloads across all our packages and 29,000 GitHub stars.
"Qdrant's technical architecture and performance capabilities have proven to be exactly what we need as we scale our AI-powered features across the platform," said Colin Chauvet, Director of Engineering at Canva. "They are an ideal partner as we standardize our vector search infrastructure to serve millions of users worldwide."
### Fixed Pipelines Break. Composable Primitives Don't.
Most search systems give you a fixed pipeline. Data goes in, queries come out, and the system decides how retrieval happens. Composable vector search inverts that. Features such as dense vectors, sparse vectors, metadata filters, multi-vector representations, and custom scoring are primitives you combine at query time, not features hidden behind an opaque API.
This matters because different workloads need fundamentally different retrieval. A [multimodal system at Tripadvisor](https://qdrant.tech/blog/case-study-tripadvisor/) retrieving across billions of signals looks nothing like an [agentic workflow at Lyzr](https://qdrant.tech/blog/case-study-lyzr/) working with 100s of agents operating at low latency. A composable engine adapts to the problem. A fixed pipeline forces the problem to adapt to the tool.
### Agents and Edge Devices Need the Same Thing: Fast, Flexible Retrieval Everywhere
Agentic AI can turn retrieval into a tight inner loop; imagine thousands of steps per workflow, where latency compounds and an occasional slow result cascades through the entire chain. Agents can't declare their retrieval strategy upfront. They shift from dense to hybrid search, tighten filters, and re-weight scores based on what prior steps returned. This is composability at query time, not configuration time.
Meanwhile, AI is moving to where decisions are made. Not everything can round-trip to the cloud: on-device assistants, field diagnostics, and industrial systems running semi-offline. [Qdrant Edge](https://qdrant.tech/documentation/edge/) will bring the full composable retrieval stack to resource-constrained devices with efficient cloud sync. One retrieval architecture from the data center to the device.
### Thanks to Our Community and the Engineers Who Shape This Engine
Qdrant is shaped by engineers running it under real pressure, and we want to thank them directly. The over 29,000 GitHub stars reflect reach, but it is the code contributions from production systems that move our engine forward. In v1.16, [@eltu added ASCII folding to improve multilingual full-text retrieval without upstream preprocessing](https://github.com/qdrant/qdrant/pull/7408). More recently, [@TY0909 contributed field-level control over HNSW graph construction](https://github.com/qdrant/qdrant/pull/7887), reducing indexing cost and memory overhead in large hybrid deployments. These contributions come from real operational pressure and directly shape how Qdrant behaves under load.
Models get the attention, but retrieval is what makes them useful in production. We believe retrieval will become core infrastructure for AI, not a feature bolted onto something else. And we're building Qdrant to be the engine that lasts.
If you're building with Qdrant or contributing back: thank you. And if you want to work on fundamental infrastructure for AI, [we're hiring](https://join.com/companies/qdrant).
@@ -0,0 +1,115 @@
---
title: "Sketch & Search: Google Deepmind x Qdrant x Freepik Hackathon Winners"
draft: false
slug: sketch-n-search-winners
short_description: "Discover the winners of Qdrant’s Sketch & Search Hackathon in collaboration with Google Deepmind and Freepik, where developers built innovative vector search applications beyond chatbots - from robotics to gaming, e-commerce, and more."
preview_image: blog/sketch-n-search-2025/sketch.png
social_preview_image: blog/sketch-n-search-2025/sketch.png
date: 2026-02-03
author: Manas Chopra
featured: true
tags:
- news
- blog
---
Builders from around the world came together for Sketch & Search, a global hackathon powered by Google DeepMind, Freepik, and Qdrant, to explore the future of AI-driven creative pipelines.
Teams were challenged to go beyond single-prompt generation and build end-to-end systems combining generative models, visual creation, and vector search. Submissions showcased consistent characters and style memory, image-as-prompt and image-to-video workflows, intelligent asset discovery, recommendations, and built-in brand-safe guardrails.
The hackathon kicked off in San Francisco on November 22, 2025, followed by a two-week virtual build window and a live demo day where winners were announced. Projects were judged on creative quality, effective search and similarity, UX tradeoffs, guardrails, and real-world applicability.
With over $25,000 in prizes, including cash awards and Gemini API credits, Sketch & Search highlighted how generation, search, and creativity come together to power production-ready applications.
👉 Full hackathon details: [Hackathon page](https://luma.com/2kt11r0m)
Let’s dive into the top projects and what we liked about each of them.
---
## 🏆 Overall Winners
### 🥇 1st Place: Prometheus (Harry Kabodha)
**Bonus Prize: Best use of Gemini**
**Total Prize: `$1500 USD Cash & 500 Gift Card` \+ `$10000 Gemini API Credits`**
<iframe width="560" height="315"
src="https://www.youtube.com/embed/IwCNnaGZLuM?rel=0"
title="YouTube video player"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
**What it is:** Prometheus turns a PDB file (molecule format) into a cinematic, structure-faithful “mechanism trailer” for scientists and biotech storytellers. Gemini interprets the protein (domains, active site, flexible regions) and writes shot prompts grounded by Mol* renders. Qdrant searches your prior targets to reuse the best prompt/camera templates, then stores what works to build an institutional “visual memory.” Nano Banana generates consistent keyframes, and Freepik stitches them into a short animation ready for decks, talks, and internal knowledge bases.
**Stack:** Gemini/Flash, Nano Banana, Qdrant, Freepik
Repo: [Prometheus](https://github.com/resilienthike/Prometheus)
---
### 🥈 2nd Place: Roast My Snack (Jay Ozer)
**Total Prize: `$500 Cash & 300 Gift Card`**
<iframe width="560" height="315"
src="https://www.loom.com/embed/b1568b443acb4c1db7f3a2af903e2d14"
title="YouTube video player"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
**What it is:** Roast My Snack transforms snack photos into 4-panel comics where Dr. Hawley - a Gen-Z molar with Adult Swim energy - roasts your snack's "aesthetic threat level" to your smile. The insight: Telling teens "sugar causes cavities" doesn't work. But "that snack is gonna turn your smile yellow"? That lands.
We use vanity as a force for good. Built with Gemini Vision, Qdrant semantic search, and clinic-validated dental science from Poppy Kids Pediatric Dentistry. Age-adaptive roasts (Spicy for tweens, Savage for teens), transparent risk scoring, and Instagram-ready exports. Dental education kids actually want to share.
**Stack:** Gemini/Flash, Nano Banana, Qdrant, Freepik
**Repo:** [Roast My Snack](https://github.com/jayozer/snackswap_comics)
---
### 🥉 3rd Place: AutoScape (Tommy Purcell & Rae Jin)
**Bonus Prize: Best use of Nano Banana**
**Total Prize: `$250 Cash & 150 Gift Card` \+ `$10,000 Gemini API Credits`**
<iframe width="560" height="315"
src="https://www.youtube.com/embed/giOQIApFRzE?rel=0"
title="YouTube video player"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
allowfullscreen>
</iframe>
**What it is:** AutoScape is an AI-powered landscape design platform that takes you from a single photo to a build-ready outdoor project. But it's more than a generic image generation tool. AutoScape is a full platform for homeowners, designers, contractors, and public-sector teams working on residential, commercial, or civic spaces. AutoScape produces photorealistic designs, materials lists, and labor cost estimates with product links, exportable to Excel or Google Sheets. Designs can be shared for collaboration. Unlike generic AI tools, AutoScape uses a Qdrant-powered RAG system and curated plant database grounded in real materials and pricing. A public gallery enables inspiration and reuse globally.
**Stack:** Gemini/Flash, Nano Banana, Qdrant, Freepik
**Repo:** [AutoScape](https://github.com/tommypurcell/AutoScape)
---
## Why These Projects Matter
These winners stand out because they treat retrieval as an engineering primitive - not a bolt-on feature. Instead of generating “one-off” outputs, they build systems that can *remember*, *reuse*, and *stay grounded* as inputs and requirements change:
- Prometheus turns structured scientific data (PDB + Mol* renders) into a repeatable video pipeline, and uses similarity search to reuse proven camera/prompt templates across targets.
- Roast My Snack couples vision + semantic search with an explicit scoring layer (age-adaptive tone, transparent risk signals) so the output is controllable and consistent - not just funny.
- AutoScape anchors image-to-design generation in a curated plant/material catalog with retrieval, so designs come with real constraints (availability, pricing) and can be iterated, shared, and audited.
Across all three, Qdrant enables the “long-term memory” layer: storing embeddings of past assets, prompts, and outcomes so future runs can start from what already worked - faster iteration, better consistency, and less prompt thrash.
As with any hackathon, there were tons of amazing submissions, and we wish we could showcase them all, but we hope this sample showcases the strong power of the Qdrant Community when tasked with new challenges and creative approaches.
## What’s Next
We’ll be publishing deep dives and hosting winner interviews over the coming weeks and months. Keep an eye out for these by [subscribing to our newsletter](https://qdrant.tech/community/)\!
A huge thank you to everyone who built, mentored, or joined. You’ve proven that with Qdrant, the possibilities are limitless.
👉 Join the community: [qdrant.tech/community](https://qdrant.tech/community)
@@ -0,0 +1,186 @@
---
title: "Two Approaches to Helping AI Agents Use Your API (And Why You Need Both)"
draft: false
slug: skill-md-meet-repl
description: "Two emerging patterns for agent-assisted development: static knowledge files and dynamic tool access. How they complement each other using Qdrant as a case study."
short_description: "Mintlify's skill.md and Armin Ronacher's REPL-first MCP solve different failure modes. Together, they define how agents should interact with developer tools."
preview_image: /blog/skill-md-meets-repl/repl-skill.png
social_preview_image: /blog/skill-md-meets-repl/repl-skill.png
date: 2026-01-28T00:00:00-08:00
author: Thierry Damiba
featured: true
tags:
- agents
- blog
---
AI coding agents fail in predictable ways when working with APIs. Two recent approaches from Mintlify and Armin Ronacher attack different failure modes. Understanding both reveals something useful about how agents should interact with developer tools.
## Two Failure Modes
When an agent writes code against your API, it can fail because:
1. **It doesn't know what it doesn't know.** The agent uses a deprecated method, misconfigures a parameter, or violates a constraint that isn't obvious from type signatures. This is the "known unknowns" problem: things the API maintainer knows but the agent doesn't.
2. **It can't discover what exists.** The agent doesn't know what collections exist, what the payload schema looks like, or what data is actually in the system. This is the "unknown unknowns" problem: things specific to the user's environment that no amount of documentation covers.
Most agent failures trace back to one of these. Mintlify's SKILL.md aproach addresses the first. Armin Ronacher's REPL-first MCP addresses the second.
## What SKILL.md Gives You
[SKILL.md](https://github.com/AgenticSkills/skills) is an emerging open standard for shipping knowledge to agents before they write code. The idea has roots in the [Cloudflare RFC](https://blog.cloudflare.com/ai-agents-open-standard), the [agentskills proposal](https://agentskills.org), and Vercel's skills CLI. [Mintlify's blog post](https://mintlify.com/blog/skill-md) by [Michael Ryaboy](https://www.linkedin.com/in/michael-ryaboy-software-engineer) showed how to apply it in practice. Decision tables for component selection, explicit gotchas sections, and auto-generating skill files from existing docs. A skill.md isn't documentation. It's a briefing. Decision tables, not tutorials. Gotchas, not explanations.
For Qdrant, a skill.md might include:
<div style="max-width: 640px; margin: 2rem auto; border-radius: 12px; overflow: hidden; font-family: 'JetBrains Mono', 'Fira Code', monospace; font-size: 14px; box-shadow: 0 4px 24px rgba(0,0,0,0.12);">
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-bottom: 2px solid #dc3545;">Decision Table</div>
<div style="background: #16213e; padding: 0;">
<table style="width: 100%; border-collapse: collapse; color: #e0e0e0; font-size: 14px;">
<thead>
<tr style="border-bottom: 1px solid #2a2a4e;">
<th style="text-align: left; padding: 12px 20px; color: #8890a8; font-weight: 400;">Want to...</th>
<th style="text-align: left; padding: 12px 20px; color: #8890a8; font-weight: 400;">Do</th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: 1px solid #2a2a4e;">
<td style="padding: 10px 20px; color: #e0e0f0;">Search</td>
<td style="padding: 10px 20px;"><code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">client.query_points(col, query=vec, limit=10)</code></td>
</tr>
<tr style="border-bottom: 1px solid #2a2a4e;">
<td style="padding: 10px 20px; color: #e0e0f0;">Filter</td>
<td style="padding: 10px 20px;">Add <code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">query_filter=Filter(must=[...])</code></td>
</tr>
<tr style="border-bottom: 1px solid #2a2a4e;">
<td style="padding: 10px 20px; color: #e0e0f0;">Hybrid</td>
<td style="padding: 10px 20px;"><code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">prefetch</code> dense+sparse, <code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">FusionQuery(Fusion.RRF)</code></td>
</tr>
<tr>
<td style="padding: 10px 20px; color: #e0e0f0;">Multi-tenant</td>
<td style="padding: 10px 20px;"><code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">create_payload_index(field, is_tenant=True)</code></td>
</tr>
</tbody>
</table>
</div>
<div style="background: #1a1a2e; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-top: 1px solid #2a2a4e; border-bottom: 2px solid #dc3545; color: #e0e0e0;">Gotchas</div>
<div style="background: #16213e; padding: 16px 20px; line-height: 1.8; color: #c0c0d8;">
<span style="color: #dc3545;">&#x2716;</span> <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">query_points</code> not <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">search</code> - all search variants deprecated<br>
<span style="color: #dc3545;">&#x2716;</span> Never one collection per user - use <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">is_tenant=True</code> index<br>
<span style="color: #dc3545;">&#x2716;</span> BM25 requires <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">Modifier.IDF</code> - sparse search needs it
</div>
</div>
This prevents the agent from using `client.search()` (deprecated), creating a collection per user (anti-pattern), or misconfiguring sparse vectors (common mistake). The guidance for all of these exists across Qdrant's tutorials, docs, and community discussions. However, finding it requires existing Qdrant context because you need to already know enough to ask the right questions. Skills package that accumulated product intuition so agents don't need to build it from scratch.
## What REPL-First MCP Gives You
[Armin Ronacher](https://lucumr.pocoo.org/), creator of Flask and now building [Earendil](https://earendil.dev/), proposed a [different approach](https://lucumr.pocoo.org/2025/1/22/what-i-want-for-ai-tools/). Instead of 30 narrow MCP tools, give the agent a Python shell with the SDK pre-configured:
```python
# Agent can just run this
collections = client.get_collections()
print([c.name for c in collections.collections])
# Then inspect the actual schema
info = client.get_collection("products")
print(info.config.params.vectors)
```
The agent discovers what exists by asking the system directly. No tool for "list collections." No tool for "get schema." Just Python. The REPL handles the unknown unknowns: what's actually in your Qdrant instance right now.
## Why Neither Alone Works
**skill.md without REPL:** The agent knows *how* to use `query_points` but not *what* to query. It guesses collection names. It assumes payload fields. It writes syntactically correct code that fails at runtime.
**REPL without skill.md:** The agent can discover what exists but still uses deprecated methods. It creates collections with wrong configurations. It makes the same mistakes it would have made without the REPL, just with more information about the data.
Together, the agent workflow looks like this:
<div style="max-width: 640px; margin: 2rem auto; border-radius: 12px; overflow: hidden; font-family: 'JetBrains Mono', 'Fira Code', monospace; font-size: 14px; box-shadow: 0 4px 24px rgba(0,0,0,0.12);">
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-bottom: 2px solid #dc3545;">Agent Workflow</div>
<div style="background: #16213e; padding: 24px 28px;">
<div style="margin-bottom: 20px;">
<div style="color: #dc3545; font-weight: 700; margin-bottom: 6px;">1. Read skill.md</div>
<div style="padding-left: 20px; line-height: 1.7;">
<code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">"Use query_points, not search"</code><br>
<code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">"Never one collection per user"</code><br>
<code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">"BM25 needs Modifier.IDF"</code>
</div>
</div>
<div style="margin-bottom: 20px;">
<div style="color: #dc3545; font-weight: 700; margin-bottom: 6px;">2. Use REPL</div>
<div style="padding-left: 20px; line-height: 1.7;">
<code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">client.get_collections()</code><br>
<code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">client.scroll("products", limit=1)</code><br>
<span style="color: #8890a8;"># Now knows: collection exists, payload has "category" field</span>
</div>
</div>
<div>
<div style="color: #dc3545; font-weight: 700; margin-bottom: 6px;">3. Write code</div>
<div style="padding-left: 20px; line-height: 1.7;">
<span style="color: #e0e0f0;">Correct method + correct collection + correct filter fields</span>
</div>
</div>
</div>
</div>
The skill.md prevents known mistakes. The REPL handles environment-specific discovery. Both failure modes addressed.
## Implementation
The skill.md is just a file you drop into your project. The REPL is an MCP tool. A minimal implementation:
```python
@server.call_tool()
async def handle_tool_call(name: str, arguments: dict):
if name == "qdrant-repl":
code = arguments["code"]
# client, models pre-configured in repl_globals
try:
result = eval(compile(code, "<repl>", "eval"), repl_globals)
return repr(result) if result else "(ok)"
except SyntaxError:
exec(compile(code, "<repl>", "exec"), repl_globals)
return "(executed)"
```
State persists between calls. The agent builds up context incrementally.
## The Broader Pattern
This isn't specific to Qdrant. Any API with:
- Deprecated methods or migration paths → needs skill.md
- User-specific state (databases, collections, schemas) → needs REPL
The two approaches complement because they address orthogonal problems. Static knowledge for static mistakes. Dynamic access for dynamic discovery.
Most developer tools need both.
## This Doesn't Make It Easy
Adding skill.md and a REPL doesn't mean agents suddenly work flawlessly. They still hallucinate. They still misunderstand requirements. They still write code that technically runs but doesn't do what you wanted.
What these tools do is eliminate *unnecessary* failures. The agent won't fail because it used a deprecated method. That's a solved problem with skill.md. It won't fail because it guessed a collection name. The REPL lets it check. But it can still fail because it misunderstood what you meant by "similar products" or because the embedding model you're using doesn't capture the semantics you care about.
You might ask: why not just pull context from docs automatically? Tools like [mcp-code-snippets](https://github.com/qdrant/mcp-code-snippets), Qdrant's MCP server for searching documentation and code examples, solve a real problem, but they solve a different one. Auto-generated context gives you API surface area. A SKILL.md gives you judgment. "Don't create one collection per user" isn't obvious from any API reference. "BM25 needs Modifier.IDF" is documented, but you have to know to look for it. Skills distill that intuition into a format agents can use immediately. The two approaches aren't in competition. Use both.
The goal isn't perfect agents. The goal is agents that fail for interesting reasons instead of boring ones. Deprecated API calls are boring failures. Wrong collection names are boring failures. Skill.md and REPL handle the boring stuff so you can focus on the hard problems.
## Try It
The [Qdrant Python skill](https://github.com/qdrant/skills/tree/main/skills/qdrant-python) is just 94 lines, and there's also a [Rust skill](https://github.com/qdrant/skills/tree/main/skills/qdrant-rust) for the gRPC client. Both are minimal, adaptable examples for your own setup. The repo also includes the `AGENTS.md` format alongside `SKILL.md` for flexible configuration.
Drop it into your project:
```bash
mkdir -p .claude/skills/qdrant
curl -o .claude/skills/qdrant/SKILL.md \
https://raw.githubusercontent.com/qdrant/skills/main/skills/qdrant-python/SKILL.md
```
For the REPL side, configure the [Qdrant MCP server with REPL](https://github.com/thierrypdamiba/mcp-server-qdrant-repl). It's a fork of the official [Qdrant MCP server](https://github.com/qdrant/mcp-server-qdrant) that adds a `qdrant-repl` tool giving agents a stateful Python shell with a pre-configured `QdrantClient`. For docs lookup, [mcp-code-snippets](https://github.com/qdrant/mcp-code-snippets) gives agents semantic search over Qdrant documentation and code examples. The REPL fork layers exploratory access on top. The agent gets the briefing, the docs, and the toolkit.
Your agents will stop using deprecated methods. They'll stop guessing collection names. They'll write correct queries on the first attempt. Not because they got smarter, but because they got the right information at the right time.
We'd love to hear how you're using Qdrant with AI agents. What's working, what's breaking, what patterns you've found. If you want to stay ahead of the curve on this stuff, [sign up for our newsletter](https://qdrant.tech/subscribe/). And thanks to you, our readers, for pushing us to make the developer experience better for everyone.
@@ -33,7 +33,7 @@ For example, you might search for hotels in Paris with specific criteria:
> In this blog, we'll show you how we built [**The Hotel Search Demo**](https://hotel-search-recipe.superlinked.io/).
**Figure 1:** Superlinked generates vectors of different modalities which are indexed and served by Qdrant for fast, accurate hotel search.
![superlinked-hotel-search](/blog/superlinked-multimodal-search/frontend.gif)
![superlinked-hotel-search](https://storage.googleapis.com/qdrant-landing/files/blog/frontend.gif)
What makes this app particularly powerful is how it breaks down your natural language query into precise parameters. As you type your question at the top, you can observe the query parameters dynamically update in the left sidebar.
@@ -122,11 +122,11 @@ colors:
code: "FFFFFF"
typography:
title: Typography
description: Main typography is Satoshi, this is employed for both UI and marketing purposes. Headlines are set in Bold (600), while text is rendered in Medium (500).
description: Our main typography is Mona Sans, this is employed for both UI and marketing purposes. Headlines are set in Semibold (500), while text is rendered in Regular (400).
example: AaBb
specimen: "ABCDEFGHIJKLMNOPQRSTUVWXYZ<br>abcdefghijklmnopqrstuvwxyz<br>0123456789 !@#$%^&*()"
link:
url: https://api.fontshare.com/v2/fonts/download/satoshi
url: https://fonts.google.com/specimen/Mona+Sans
text: Download
trademarks:
title: Trademarks
@@ -2,20 +2,20 @@
title: FAQs
questions:
- id: 0
question: Is Qdrant Cloud Inference available on free accounts?
answer: Inference is only available on paid Qdrant Cloud clusters.
question: Is Qdrant Cloud Inference available on free clusters?
answer: Yes, free models and external model providers can be used in free Qdrant Cloud clusters. Paid models require a paid cluster.
- id: 1
question: What kinds of data can I embed?
answer: You can embed both text and image data using the current available models.
- id: 2
question: Where are the embeddings generated?
answer: Embeddings are generated inside the network of your cluster, which removes external API overhead.
answer: For Qdrant hosted models, embeddings are generated inside the network of your cluster, which removes external API overhead. If you use external model providers, embeddings are generated by this provider.
- id: 3
question: How much does it cost?
answer: Inference is billed per token, and costs depend on the model. Each month, Qdrant Paid Cloud users get up to 5 million tokens free, depending on the model, and unlimited tokens for BM25.
answer: Inference is billed per token, and costs depend on the model. Each month, Qdrant Paid Cloud users get up to 5 million tokens free, depending on the model. Several models are offered for free completely, with no token limits. For more details, refer to the Inference section on your cluster detail page.
- id: 4
question: How do I get started?
answer: If you're on a paid plan, Cloud Inference is enabled by default.
answer: For new clusters, Cloud Inference is enabled by default. For older clusters that were created before the release of Cloud Inference, you can enable it from the cluster detail page in the Qdrant Cloud Console.
- id: 5
question: Will there be options for other embedding models?
answer: We plan to add models incrementally based on customer feedback.
@@ -1,13 +1,13 @@
---
title: "Join the</br>Qdrant Community"
description: Connect with over 30,000 community members, get access to educational resources, and stay up to date on all news and discussions about Qdrant and the vector database space.
description: Connect with over 30,000 community members, get access to educational resources, and stay up to date on all news and discussions about Qdrant and the vector search space.
image:
src: /img/community-hero.png
alt: Community
button:
text: Join our Discord
url: https://discord.gg/qdrant
about: Get access to educational resources, and stay up to date on all news and discussions about Qdrant and the vector database space.
about: Get access to educational resources, and stay up to date on all news and discussions about Qdrant and the vector search space.
sitemapExclude: true
---
@@ -6,4 +6,17 @@ weight: 100
# Qdrant Essentials Certification
Coming soon! [Click here](https://forms.gle/QPSfdMjs3QpUCtGT9) to be notified when certifications become available.
Congratulations! You’ve officially navigated the complexities of the **Qdrant Essentials** course. You didn’t just learn how to store data; you learned how to architect a high-performance **Vector Search Engine**.
You’ve moved past simple "Hello World" tutorials and dove deep into HNSW indexing, hybrid search, and production-grade optimization. That effort deserves more than just a "finished" status. It deserves professional recognition.
## 🏆 Get #QdrantCertified
Your expertise is now production-ready. It’s time to validate those skills with our official certification.
**Head over to [train.qdrant.dev](https://train.qdrant.dev) to take the exam.**
Passing this exam proves you aren't just a user; you are a **Search Engineer** capable of:
* **Designing** scalable retrieval systems.
* **Optimizing** for both memory and latency.
* **Mastering** the nuances of the Qdrant architecture.
@@ -135,7 +135,7 @@ collection_info = client.get_collection(collection_name)
print("Collection info:", collection_info)
```
Expected output: Detailed collection information showing `points_count=2`, vector configuration, and [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) settings.
Expected output: Detailed collection information showing `points_count=2`, vector configuration, and [HNSW](https://qdrant.tech/articles/filterable-hnsw/) settings.
## Step 8: Run Your First Similarity Search
@@ -1,16 +1,16 @@
---
title: "Qdrant Cloud Setup"
title: "Qdrant Setup"
description: Set up your Qdrant Cloud cluster in minutes. Learn to create collections, manage data, access the Web UI, and connect securely from Python.
weight: 2
---
{{< date >}} Day 0 {{< /date >}}
# Qdrant Cloud Setup
# Qdrant Setup
<div class="video">
<iframe
src="https://www.youtube.com/embed/PLTlJyrSkng?si=y9fNtxNS34PdcKBk"
src="https://www.youtube.com/embed/9JBlgNBQoOY?si=7t3LAvMsUUtlUMN7&rel=0"
frameborder="0"
allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share"
referrerpolicy="strict-origin-when-cross-origin"
@@ -54,7 +54,7 @@ This is where chunking comes in. The goal is to have chunks
By breaking a document into focused chunks, each chunk gets its own vector that accurately represents a specific idea. This allows the search to be far more precise.
**Example:** Consider a multi-page Document like the [Qdrant Collection Configuration Guide of day 7](/course/essentials/day-7/collection-configuration-guide/) covering everything from HNSW to sharding and quantization.
**Example:** Consider a multi-page Document like the [Qdrant Collection Configuration Guide of Day 7](/course/essentials/day-7/collection-configuration-guide/) covering everything from HNSW to sharding and quantization.
If a user asks: *"What does the m parameter do?"*
@@ -492,7 +492,7 @@ You can read more about grouping [here](/documentation/concepts/hybrid-queries/?
- Original content with source attribution
- Section context for better understanding
- Direct links to full documents
- Creation timestamps for [freshness](/documentation/concepts/hybrid-queries/#time-based-score-boosting)
- Creation timestamps for [freshness](/documentation/concepts/search-relevance/#time-based-score-boosting)
**5. Permission Control**
```python
@@ -234,7 +234,7 @@ Query: 'alien invasion'
## Step 6: Advanced Features
Note: If you are already familiar Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/concepts/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [day 2](/content/course/essentials/day-2/_index.md) of this course.
Note: If you are already familiar Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/concepts/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [Day 2](/content/course/essentials/day-2/_index.md) of this course.
### Filtering by Metadata
@@ -349,4 +349,4 @@ encoder_large = SentenceTransformer("all-mpnet-base-v2") # Larger, potentially
encoder_fast = SentenceTransformer("all-MiniLM-L12-v2") # Different size/speed tradeoff
```
**Ready for Day 2?** Tomorrow you'll learn how Qdrant makes vector search lightning-fast through [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) indexing and how to optimize for production workloads.
**Ready for Day 2?** Tomorrow you'll learn how Qdrant makes vector search lightning-fast through [HNSW](https://qdrant.tech/articles/filterable-hnsw/) indexing and how to optimize for production workloads.
@@ -9,7 +9,7 @@ weight: 30
# Indexing and Performance
Master [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) indexing and practical tuning for fast retrieval.
Master [HNSW](https://qdrant.tech/articles/filterable-hnsw/) indexing and practical tuning for fast retrieval.
---
@@ -8,7 +8,7 @@ weight: 4
# Demo: HNSW Performance Tuning
Learn how to improve vector search speed with [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) tuning and payload indexing on a real 100K dataset.
Learn how to improve vector search speed with [HNSW](https://qdrant.tech/articles/filterable-hnsw/) tuning and payload indexing on a real 100K dataset.
**Follow along in Colab:** <a href="https://colab.research.google.com/github/qdrant/examples/blob/master/course/day_2/hnsw_performance_tuning.ipynb">
<img src="https://colab.research.google.com/assets/colab-badge.svg" style="display:inline; margin:0;" alt="Open In Colab"/>
@@ -8,7 +8,7 @@ weight: 3
# Combining Vector Search and Filtering
We've talked about how Qdrant uses the [HNSW](/documentation/concepts/indexing/#filtrable-index) graph to efficiently search dense vectors. But in real-world applications, you'll often want to constrain your search using filters. This creates unique challenges for graph traversal that Qdrant solves elegantly.
We've talked about how Qdrant uses the [HNSW](/documentation/concepts/indexing/#filterable-index) graph to efficiently search dense vectors. But in real-world applications, you'll often want to constrain your search using filters. This creates unique challenges for graph traversal that Qdrant solves elegantly.
<div class="video">
<iframe
@@ -199,4 +199,4 @@ See more in [the docs](/documentation/concepts/filtering/).
In the next section, we'll define a collection with structured payloads, configure payload indexing, and evaluate how different HNSW parameters impact filtered search performance.
Learn more: [Filterable HNSW Article](https://qdrant.tech/articles/filtrable-hnsw/)
Learn more: [Filterable HNSW Article](https://qdrant.tech/articles/filterable-hnsw/)
@@ -8,7 +8,7 @@ weight: 5
# Project: HNSW Performance Benchmarking
Now that you've seen how [HNSW](https://qdrant.tech/articles/filtrable-hnsw/) parameters and payload indexes affect performance with the DBpedia dataset, it's time to optimize for your own domain and use case.
Now that you've seen how [HNSW](https://qdrant.tech/articles/filterable-hnsw/) parameters and payload indexes affect performance with the DBpedia dataset, it's time to optimize for your own domain and use case.
## Your Mission
@@ -25,7 +25,7 @@ At this point, you've learned how vector search retrieves the nearest vectors to
You might wonder if Qdrant calculates the distance to every single vector in your collection for each query. This method, known as brute force search, technically works but with millions or billions of vectors this is too slow per query.
Fortunately, Qdrant speeds things up with **[HNSW — Hierarchical Navigable Small World](https://qdrant.tech/articles/filtrable-hnsw/)**.
Fortunately, Qdrant speeds things up with **[HNSW — Hierarchical Navigable Small World](https://qdrant.tech/articles/filterable-hnsw/)**.
### The Library Analogy
@@ -3,7 +3,7 @@ caseStudy:
logo:
src: /img/customers-case-studies/customer-logo4.svg
alt: Logo Dailymotion
title: Recommendation Engine with Qdrant Vector Database
title: Recommendation Engine with Qdrant Vector Search
description: Dailymotion leverages Qdrant to optimize its <b>video recommendation engine</b>, managing over 420 million videos and processing 13 million recommendations daily. With this, Dailymotion was able to <b>reduced content processing times from hours to minutes</b> and <b>increased user interactions and click-through rates by more than 3x.</b>
link:
text: Read Case Study
+1 -1
View File
@@ -26,7 +26,7 @@ cards:
title: Categorization Demo -<br> E-Commerce Products
paragraphs:
- id: 0
content: Discover the power of vector databases in e-commerce through our demo. Simply input a product name and watch as our multi-language model intelligently categorizes it. The dots you see represent product clusters, highlighting our system's efficient categorization.
content: Discover the power of vector search in e-commerce through our demo. Simply input a product name and watch as our multi-language model intelligently categorizes it. The dots you see represent product clusters, highlighting our system's efficient categorization.
link:
text: View Demo
url: https://qdrant.to/extreme-classification-demo
@@ -6,7 +6,7 @@ breadcrumb: false
content:
- partial: "documentation/banners/banner-a"
title: Qdrant Documentation
description: Qdrant is an AI-native vector database and a semantic search engine. You can use it to extract meaningful information from unstructured data.
description: Qdrant is an AI-native vector search and a semantic search engine. You can use it to extract meaningful information from unstructured data.
linkDescription: <a href="https://github.com/qdrant/qdrant_demo/" target="_blank">Clone this repo now</a> and build a search engine in five minutes.
cloudButton:
text: Cloud Quickstart
@@ -72,7 +72,7 @@ THIS CONTENT IS GOING TO BE IGNORED FOR NOW
# Documentation
Qdrant is an AI-native vector database and a semantic search engine. You can use it to extract meaningful information from unstructured data. Want to see how it works? [Clone this repo now](https://github.com/qdrant/qdrant_demo/) and build a search engine in five minutes.
Qdrant is an AI-native vector search and a semantic search engine. You can use it to extract meaningful information from unstructured data. Want to see how it works? [Clone this repo now](https://github.com/qdrant/qdrant_demo/) and build a search engine in five minutes.
|||
|-:|:-|
@@ -88,7 +88,7 @@ Qdrant is an AI-native vector database and a semantic search engine. You can use
## Qdrant's most popular features:
||||
|:-|:-|:-|
|[Filtrable HNSW](/documentation/filtering/) </br> Single-stage payload filtering | [Recommendations & Context Search](/documentation/concepts/explore/#explore-the-data) </br> Exploratory advanced search| [Pure-Vector Hybrid Search](/documentation/hybrid-queries/)</br>Full text and semantic search in one|
|[Filterable HNSW](/documentation/filtering/) </br> Single-stage payload filtering | [Recommendations & Context Search](/documentation/concepts/explore/#explore-the-data) </br> Exploratory advanced search| [Pure-Vector Hybrid Search](/documentation/hybrid-queries/)</br>Full text and semantic search in one|
|[Multitenancy](/documentation/guides/multiple-partitions/) </br> Payload-based partitioning|[Custom Sharding](/documentation/guides/distributed_deployment/#sharding) </br> For data isolation and distribution|[Role Based Access Control](/documentation/guides/security/?q=jwt#granular-access-control-with-jwt)</br>Secure JWT-based access |
|[Quantization](/documentation/guides/quantization/) </br> Compress data for drastic speedups|[Multivector Support](/documentation/concepts/vectors/?q=multivect#multivectors) </br> For ColBERT late interaction |[Built-in IDF](/documentation/concepts/indexing/?q=inverse+docu#idf-modifier) </br> Advanced similarity calculation|
@@ -1,19 +0,0 @@
---
title: Advanced Retrieval
weight: 17
# If the index.md file is empty, the link to the section will be hidden from the sidebar
is_empty: false
aliases:
- how-to
- tutorials
partition: qdrant
---
# Advanced Tutorials
| |
|----------------------------------------------------------|
| [Use Collaborative Filtering to Build a Movie Recommendation System with Qdrant](/documentation/advanced-tutorials/collaborative-filtering/) |
| [Build a Text/Image Multimodal Search System with Qdrant and FastEmbed](/documentation/advanced-tutorials/multimodal-search-fastembed/) |
| [Navigate Your Codebase with Semantic Search and Qdrant](/documentation/advanced-tutorials/code-search/) |
| [Ensure optimal large-scale PDF Retrieval with Qdrant and ColPali/ColQwen](/documentation/advanced-tutorials/pdf-retrieval-at-scale/) |
@@ -1,20 +0,0 @@
---
title: Vector Search Basics
aliases:
- /documentation/tutorials/
- how-to
- tutorials
weight: 16
# If the index.md file is empty, the link to the section will be hidden from the sidebar
is_empty: false
partition: qdrant
---
# Beginner Tutorials
| |
|----------------------------------------------------|
| [Build Your First Semantic Search Engine in 5 Minutes](/documentation/beginner-tutorials/search-beginners/) |
| [Build a Neural Search Service with Sentence Transformers and Qdrant](/documentation/beginner-tutorials/neural-search/) |
| [Build a Hybrid Search Service with FastEmbed and Qdrant](/documentation/beginner-tutorials/hybrid-search-fastembed/) |
| [Measure and Improve Retrieval Quality in Semantic Search](/documentation/beginner-tutorials/retrieval-quality/) |
@@ -31,7 +31,7 @@ content:
src: /img/dev-portal-build/rag.png
alt: RAG
title: RAG
description: Build end-to-end prototype chatbots. Learn how Qdrant integrates with popular RAG frameworks like LangChain and Llamaindex.
description: Build end-to-end prototype chatbots. Learn how Qdrant integrates with popular RAG frameworks like LangChain and LlamaIndex.
link:
url: /documentation/frameworks/langchain/
text: Read More
@@ -14,8 +14,8 @@ Qdrant Cloud offers an optional premium tier for customers who require additiona
* **Shorter Response Times**: Premium customers receive priority support and can expect faster response times, with shorter SLAs.
* **99.9% Uptime SLA**: We guarantee 99.9% uptime for your Qdrant Cloud clusters (compared to 99.5% in standard).
* **Single Sign-On (SSO)**: Premium customers can use their existing SSO provider to manage access to Qdrant Cloud.
* **VPC Private Links**: Premium customers can connect their Qdrant Cloud clusters to their VPCs using private links (AWS only).
* **Storage encryption with shared keys**: Premium customers can encrypt their data at rest using their own keys (AWS only).
* **VPC Private Links**: Premium customers can connect their Qdrant Cloud clusters to their VPCs using private links.
* **Storage encryption with shared keys**: Premium customers can encrypt their data at rest using their own keys.
Please refer to the [Qdrant Cloud SLA](https://qdrant.to/sla/) for a detailed definition on uptime and support SLAs.
@@ -19,10 +19,13 @@ You can pay for your Qdrant Cloud database clusters either with a credit card or
Your payment method is charged at the beginning of each month for the previous month's usage. There is no difference in pricing between the different payment methods.
If you choose to pay through a marketplace, the Qdrant Cloud usage costs are added as $0.01 Resource Usage Units to your existing billing for your cloud provider services. E.g. if you create a Qdrant Cluster that costs $85 in a month, 8,500 Resource Usage Units for Qdrant Cloud will be added to your cloud provider bill. A detailed breakdown of your usage is available in the Qdrant Cloud Console.
If you choose to pay through a marketplace, the Qdrant Cloud usage costs are added as \\$0.01 Resource Usage Units to your existing billing for your cloud provider services.
E.g. if you create a Qdrant Cluster that costs \\$85 in a month, 8,500 Resource Usage Units for Qdrant Cloud will be added to your cloud provider bill. A detailed breakdown of your usage is available in the Qdrant Cloud Console.
Note: Even if you pay using a marketplace subscription, your database clusters will still be deployed into Qdrant-owned infrastructure. The setup and management of Qdrant database clusters will also still be done via the Qdrant Cloud Console UI.
Note: If you choose to pay for your Qdrant Cloud clusters through a cloud provider marketplace, Qdrant Cluster creation is limited to paid clusters in regions of this cloud provider only. This is due to compliance requirements of the cloud provider. If you wish to run clusters in other regions or on other cloud providers, you can create additional accounts using a different payment method.
If you wish to deploy Qdrant database clusters into your own environment from Qdrant Cloud then we recommend our [Hybrid Cloud](/documentation/hybrid-cloud/) solution.
![Payment Options](/documentation/cloud/payment-options.png)
@@ -9,122 +9,711 @@ aliases:
- cloud/quickstart-cloud/
- /documentation/quickstart-cloud/
---
# How to Get Started With Qdrant Cloud
<p align="center"><iframe width="560" height="315" src="https://www.youtube.com/embed/3hrQP3hh69Y?si=hypr-vyKywhjoOTQ" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></p>
<p style="text-align: center;">You can try vector search on Qdrant Cloud in three steps.
</br> Instructions are below, but the video is faster:</p>
# Quick Start with Qdrant Cloud
## Setup a Qdrant Cloud Cluster
<p align="center"><iframe width="560" height="315" src="https://www.youtube.com/embed/xvWIssi_cjQ?si=CLhFrUDpQlNog9mz&rel=0" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe></p>
Learn how to set up Qdrant Cloud and perform your first semantic search in just a few minutes. We'll use a sample dataset of menu items pre-embedded with the `BAAI/bge-small-en-v1.5` model.
## 1. Create a Cloud Cluster
1. Register for a [Cloud account](https://cloud.qdrant.io/signup) with your email, Google or Github credentials.
2. Go to **Clusters** and follow the onboarding instructions under **Create First Cluster**.
2. Go to **Clusters** and click **Create First Cluster**.
3. Copy your **API key** when prompted - you'll need it to connect. Store it somewhere safe as it won't be displayed again.
![create a cluster](/docs/gettingstarted/gui-quickstart/create-cluster.png)
For detailed cluster setup instructions, see the [Cloud documentation](/documentation/cloud-intro/).
3. When you create it, you will receive an API key. You will need to copy it and store it somewhere self. It will not be displayed again. If you loose it, you can always create a new one on the **Cluster Detail Page** later.
## 2. Install the Qdrant Client
![get api key](/docs/gettingstarted/gui-quickstart/api-key.png)
## Access the Cluster UI
1. Click on **Cluster UI** on the **Cluster Detail Page** to access the cluster UI dashboard.
2. Paste your new API key here. You can revoke and create new API keys in the **API Keys** tab on your **Cluster Detail Page**.
3. The key will grant you access to your Qdrant instance. Now you can see the cluster Dashboard.
![access the dashboard](/docs/gettingstarted/gui-quickstart/access-dashboard.png)
## Authenticate via SDKs
Now that you have your cluster and key, you can use our official SDKs to access Qdrant Cloud from within your application.
Once you have a cluster, the fastest way to get started is to use our official SDKs which provide a convenient interface for working with Qdrant in your preferred programming language.
```bash
curl \
-X GET https://xyz-example.eu-central.aws.cloud.qdrant.io:6333 \
--header 'api-key: <your-api-key>'
# Alternatively, you can use the `Authorization` header with the `Bearer` prefix
curl \
-X GET https://xyz-example.eu-central.aws.cloud.qdrant.io:6333 \
--header 'Authorization: Bearer <your-api-key>'
pip install qdrant-client fastembed # for Python projects
# cargo add qdrant-client fastembed # for Rust projects
# npm install @qdrant/js-client-rest fastembed # for Node.js projects
```
## 3. Connect to Qdrant Cloud
Import the qdrant client and create a connection to your Qdrant Cloud cluster using your cluster URL and API key.
```python
from qdrant_client import QdrantClient
qdrant_client = QdrantClient(
host="xyz-example.eu-central.aws.cloud.qdrant.io",
api_key="<your-api-key>",
# connect to Qdrant Cloud
client = QdrantClient(
url="https://xyz-example.eu-central.aws.cloud.qdrant.io",
api_key="your-api-key",
)
```
```rust
use qdrant_client::Qdrant;
// Connect to Qdrant Cloud
let client = Qdrant::from_url("https://xyz-example.eu-central.aws.cloud.qdrant.io:6334")
.api_key("your-api-key")
.build()?;
```
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({
host: "xyz-example.eu-central.aws.cloud.qdrant.io",
apiKey: "<your-api-key>",
url: "https://xyz-example.eu-central.aws.cloud.qdrant.io",
apiKey: "your-api-key",
});
```
```bash
# test the connection
curl -X GET \
'http://<your-qdrant-host>:6333/collections' \
--header 'api-key: <api-key-value>'
```
## 4. Create our collection
We will use some sample menu items to demonstrate how to create a collection and add data to it. Each menu item has a name, description, price, and category. First, we need to create a collection in Qdrant to store our menu items.
```python
from qdrant_client.models import Distance, VectorParams
# create collection
client.create_collection(
collection_name="items",
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
)
```
```rust
use qdrant_client::Qdrant;
use qdrant_client::qdrant::{CreateCollectionBuilder, Distance, VectorParamsBuilder};
let client = Qdrant::from_url("https://xyz-example.eu-central.aws.cloud.qdrant.io:6334")
.api_key("<your-api-key>")
.build()?;
// create collection
client.create_collection(
CreateCollectionBuilder::new("items")
.vectors_config(VectorParamsBuilder::new(384, Distance::Cosine)),
)
.await?;
```
```java
import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
QdrantClient client =
new QdrantClient(
QdrantGrpcClient.newBuilder(
"xyz-example.eu-central.aws.cloud.qdrant.io",
6334,
true)
.withApiKey("<your-api-key>")
.build());
```typescript
await client.createCollection("items", {
vectors: { size: 384, distance: "Cosine" },
});
```
```csharp
using Qdrant.Client;
var client = new QdrantClient(
host: "xyz-example.eu-central.aws.cloud.qdrant.io",
https: true,
apiKey: "<your-api-key>"
);
```bash
curl -X PUT \
'http://<your-qdrant-host>:6333/collections/items' \
--header 'api-key: <api-key-value>' \
--header 'Content-Type: application/json' \
--data-raw '{
"vectors": {
"size": 384,
"distance": "Cosine"
}
}'
```
```go
import "github.com/qdrant/go-client/qdrant"
## 5. Populate the collection
Next, we will populate the collection with menu items. Each item will be represented as a point in the collection, with its vector embedding and associated metadata.
client, err := qdrant.NewClient(&qdrant.Config{
Host: "xyz-example.eu-central.aws.cloud.qdrant.io",
Port: 6334,
APIKey: "<your-api-key>",
UseTLS: true,
})
```python
from qdrant_client.models import PointStruct
from fastembed import TextEmbedding
# load the embedding model
model = TextEmbedding('BAAI/bge-small-en-v1.5')
menu_items = [
("Pad Thai with Tofu", "Stir-fried rice noodles with tofu bean sprouts scallions and crushed peanuts in traditional tamarind sauce", "$13.95", "Noodles"),
("Grilled Salmon Fillet", "Wild-caught Atlantic salmon grilled with lemon butter and fresh herbs served with seasonal vegetables", "$24.50", "Seafood Entrees"),
("Mushroom Risotto", "Creamy arborio rice with mixed mushrooms parmesan truffle oil and fresh thyme", "$16.75", "Vegetarian"),
("Bibimbap Bowl", "Korean rice bowl with seasoned vegetables fried egg gochujang sauce and choice of protein", "$14.50", "Korean Bowls"),
("Falafel Wrap", "Crispy chickpea fritters with hummus tahini cucumber tomato and pickled vegetables in warm pita", "$11.25", "Mediterranean"),
("Shrimp Tacos", "Three soft tacos with grilled shrimp cabbage slaw chipotle aioli and fresh lime", "$13.00", "Tacos"),
("Vegetable Curry", "Mixed vegetables in aromatic coconut curry sauce with jasmine rice and naan bread", "$12.95", "Indian Curries"),
("Tuna Poke Bowl", "Fresh ahi tuna with avocado edamame cucumber seaweed salad over sushi rice with spicy mayo", "$16.50", "Poke Bowls"),
("Margherita Pizza", "Fresh mozzarella san marzano tomatoes basil and extra virgin olive oil on wood-fired crust", "$14.00", "Pizza"),
("Chicken Tikka Masala", "Tandoori chicken in creamy tomato sauce with aromatic spices served with basmati rice", "$15.95", "Indian Entrees"),
("Greek Salad", "Romaine lettuce tomatoes cucumbers kalamata olives feta cheese red onion with lemon oregano dressing", "$10.50", "Salads"),
("Lobster Roll", "Fresh Maine lobster meat with light mayo on toasted buttery roll served with chips", "$22.00", "Seafood Sandwiches"),
("Quinoa Buddha Bowl", "Organic quinoa with roasted chickpeas kale sweet potato tahini dressing and hemp seeds", "$13.50", "Healthy Bowls"),
("Beef Pho", "Traditional Vietnamese beef noodle soup with rice noodles fresh herbs bean sprouts and lime", "$12.75", "Noodle Soups"),
("Eggplant Parmesan", "Breaded eggplant layered with marinara mozzarella and parmesan served with pasta", "$15.25", "Italian Entrees"),
("Crab Cakes", "Maryland-style lump crab cakes with remoulade sauce and mixed greens", "$18.50", "Seafood Appetizers"),
("Tofu Stir Fry", "Crispy tofu with broccoli bell peppers snap peas in garlic ginger sauce over steamed rice", "$12.50", "Vegetarian Entrees"),
("Salmon Sushi Platter", "12 pieces of fresh salmon nigiri and sashimi with wasabi pickled ginger and soy sauce", "$19.95", "Sushi"),
("Caprese Sandwich", "Fresh mozzarella tomatoes basil pesto balsamic glaze on ciabatta bread", "$11.75", "Sandwiches"),
("Tom Yum Soup", "Spicy and sour Thai soup with shrimp lemongrass galangal mushrooms and kaffir lime leaves", "$11.50", "Soups"),
("Lentil Dal", "Red lentils simmered with turmeric cumin coriander served with rice and naan", "$11.95", "Vegan Entrees"),
("Fish and Chips", "Beer-battered cod with crispy fries malt vinegar and tartar sauce", "$16.00", "British Classics"),
("Veggie Burger", "House-made black bean and quinoa patty with avocado sprouts tomato on brioche bun", "$13.25", "Burgers"),
("Miso Ramen", "Rich miso broth with ramen noodles soft-boiled egg bamboo shoots nori and scallions", "$14.50", "Ramen"),
("Stuffed Bell Peppers", "Roasted bell peppers filled with rice vegetables herbs and melted cheese", "$13.75", "Vegetarian Entrees"),
("Scallop Risotto", "Pan-seared sea scallops over creamy parmesan risotto with white wine and lemon", "$26.50", "Seafood Specials"),
("Spring Rolls", "Fresh rice paper rolls with vegetables tofu rice noodles herbs and peanut dipping sauce", "$8.95", "Appetizers"),
("Oyster Po Boy", "Fried oysters with lettuce tomato pickles and remoulade on french bread", "$15.50", "Sandwiches"),
("Portobello Mushroom Steak", "Grilled portobello cap marinated in balsamic with roasted vegetables and quinoa", "$14.95", "Vegan Entrees"),
("Coconut Shrimp", "Jumbo shrimp breaded in shredded coconut served with sweet chili sauce", "$14.25", "Seafood Appetizers")
]
# embedding generator
points = []
embeddings = model.embed([f"{item[0]} {item[1]}" for item in menu_items])
for i, embedding in enumerate(embeddings):
vector = embedding.tolist()
point = PointStruct(
id=i,
vector=vector,
payload={
"item_name": menu_items[i][0],
"description": menu_items[i][1],
"price": menu_items[i][2],
"category": menu_items[i][3],
}
)
points.append(point)
# upsert points to collection
client.upsert(
collection_name="items",
points=points,
)
```
## Try the Tutorial Sandbox
```rust
use fastembed::{EmbeddingModel, InitOptions, TextEmbedding};
use qdrant_client::qdrant::{PointStruct, UpsertPointsBuilder};
use qdrant_client::{Qdrant, Payload};
use serde_json::json;
1. Open the interactive **Tutorial**. Here, you can test basic Qdrant API requests.
2. Using the **Quickstart** instructions, create a collection, add vectors and run a search.
3. The output will show you some basic semantic search results.
// load the embedding model
let mut model = TextEmbedding::try_new(
InitOptions::new(EmbeddingModel::BGESmallENV15).with_show_download_progress(true),
)
.expect("Failed to load embedding model");
![interactive-tutorial](/docs/gettingstarted/gui-quickstart/interactive-tutorial.png)
// generate embeddings and prepare points
let menu_items = vec![
(
"Pad Thai with Tofu",
"Stir-fried rice noodles with tofu bean sprouts scallions and crushed peanuts in traditional tamarind sauce",
"$13.95",
"Noodles",
),
(
"Grilled Salmon Fillet",
"Wild-caught Atlantic salmon grilled with lemon butter and fresh herbs served with seasonal vegetables",
"$24.50",
"Seafood Entrees",
),
(
"Mushroom Risotto",
"Creamy arborio rice with mixed mushrooms parmesan truffle oil and fresh thyme",
"$16.75",
"Vegetarian",
),
(
"Bibimbap Bowl",
"Korean rice bowl with seasoned vegetables fried egg gochujang sauce and choice of protein",
"$14.50",
"Korean Bowls",
),
(
"Falafel Wrap",
"Crispy chickpea fritters with hummus tahini cucumber tomato and pickled vegetables in warm pita",
"$11.25",
"Mediterranean",
),
(
"Shrimp Tacos",
"Three soft tacos with grilled shrimp cabbage slaw chipotle aioli and fresh lime",
"$13.00",
"Tacos",
),
(
"Vegetable Curry",
"Mixed vegetables in aromatic coconut curry sauce with jasmine rice and naan bread",
"$12.95",
"Indian Curries",
),
(
"Tuna Poke Bowl",
"Fresh ahi tuna with avocado edamame cucumber seaweed salad over sushi rice with spicy mayo",
"$16.50",
"Poke Bowls",
),
(
"Margherita Pizza",
"Fresh mozzarella san marzano tomatoes basil and extra virgin olive oil on wood-fired crust",
"$14.00",
"Pizza",
),
(
"Chicken Tikka Masala",
"Tandoori chicken in creamy tomato sauce with aromatic spices served with basmati rice",
"$15.95",
"Indian Entrees",
),
(
"Greek Salad",
"Romaine lettuce tomatoes cucumbers kalamata olives feta cheese red onion with lemon oregano dressing",
"$10.50",
"Salads",
),
(
"Lobster Roll",
"Fresh Maine lobster meat with light mayo on toasted buttery roll served with chips",
"$22.00",
"Seafood Sandwiches",
),
(
"Quinoa Buddha Bowl",
"Organic quinoa with roasted chickpeas kale sweet potato tahini dressing and hemp seeds",
"$13.50",
"Healthy Bowls",
),
(
"Beef Pho",
"Traditional Vietnamese beef noodle soup with rice noodles fresh herbs bean sprouts and lime",
"$12.75",
"Noodle Soups",
),
(
"Eggplant Parmesan",
"Breaded eggplant layered with marinara mozzarella and parmesan served with pasta",
"$15.25",
"Italian Entrees",
),
(
"Crab Cakes",
"Maryland-style lump crab cakes with remoulade sauce and mixed greens",
"$18.50",
"Seafood Appetizers",
),
(
"Tofu Stir Fry",
"Crispy tofu with broccoli bell peppers snap peas in garlic ginger sauce over steamed rice",
"$12.50",
"Vegetarian Entrees",
),
(
"Salmon Sushi Platter",
"12 pieces of fresh salmon nigiri and sashimi with wasabi pickled ginger and soy sauce",
"$19.95",
"Sushi",
),
(
"Caprese Sandwich",
"Fresh mozzarella tomatoes basil pesto balsamic glaze on ciabatta bread",
"$11.75",
"Sandwiches",
),
(
"Tom Yum Soup",
"Spicy and sour Thai soup with shrimp lemongrass galangal mushrooms and kaffir lime leaves",
"$11.50",
"Soups",
),
(
"Lentil Dal",
"Red lentils simmered with turmeric cumin coriander served with rice and naan",
"$11.95",
"Vegan Entrees",
),
(
"Fish and Chips",
"Beer-battered cod with crispy fries malt vinegar and tartar sauce",
"$16.00",
"British Classics",
),
(
"Veggie Burger",
"House-made black bean and quinoa patty with avocado sprouts tomato on brioche bun",
"$13.25",
"Burgers",
),
(
"Miso Ramen",
"Rich miso broth with ramen noodles soft-boiled egg bamboo shoots nori and scallions",
"$14.50",
"Ramen",
),
(
"Stuffed Bell Peppers",
"Roasted bell peppers filled with rice vegetables herbs and melted cheese",
"$13.75",
"Vegetarian Entrees",
),
(
"Scallop Risotto",
"Pan-seared sea scallops over creamy parmesan risotto with white wine and lemon",
"$26.50",
"Seafood Specials",
),
(
"Spring Rolls",
"Fresh rice paper rolls with vegetables tofu rice noodles herbs and peanut dipping sauce",
"$8.95",
"Appetizers",
),
(
"Oyster Po Boy",
"Fried oysters with lettuce tomato pickles and remoulade on french bread",
"$15.50",
"Sandwiches",
),
(
"Portobello Mushroom Steak",
"Grilled portobello cap marinated in balsamic with roasted vegetables and quinoa",
"$14.95",
"Vegan Entrees",
),
(
"Coconut Shrimp",
"Jumbo shrimp breaded in shredded coconut served with sweet chili sauce",
"$14.25",
"Seafood Appetizers",
),
];
let embeddings = model
.embed(
menu_items
.iter()
.map(|item| format!("{} {}", item.0, item.1))
.collect::<Vec<_>>(),
None,
)
.expect("Failed to generate embeddings");
let points = embeddings
.into_iter()
.enumerate()
.map(|(idx, embedding)| {
PointStruct::new(
idx as u64,
embedding,
Payload::try_from(json!({
"item_name": menu_items[idx].0,
"description": menu_items[idx].1,
"price": menu_items[idx].2,
"category": menu_items[idx].3,
}))
.unwrap(),
)
})
.collect::<Vec<_>>();
let _ = client
.upsert_points(UpsertPointsBuilder::new("items", points).wait(true))
.await;
```
```typescript
import { TextEmbedding, EmbeddingModel } from 'fastembed';
// load the embedding model
const model = await FlagEmbedding.init({
model: EmbeddingModel.BGESmallENV15,
});
let menuItems = [
[
"Pad Thai with Tofu",
"Stir-fried rice noodles with tofu bean sprouts scallions and crushed peanuts in traditional tamarind sauce",
"$13.95",
"Noodles",
],
[
"Grilled Salmon Fillet",
"Wild-caught Atlantic salmon grilled with lemon butter and fresh herbs served with seasonal vegetables",
"$24.50",
"Seafood Entrees",
],
[
"Mushroom Risotto",
"Creamy arborio rice with mixed mushrooms parmesan truffle oil and fresh thyme",
"$16.75",
"Vegetarian",
],
[
"Bibimbap Bowl",
"Korean rice bowl with seasoned vegetables fried egg gochujang sauce and choice of protein",
"$14.50",
"Korean Bowls",
],
[
"Falafel Wrap",
"Crispy chickpea fritters with hummus tahini cucumber tomato and pickled vegetables in warm pita",
"$11.25",
"Mediterranean",
],
[
"Shrimp Tacos",
"Three soft tacos with grilled shrimp cabbage slaw chipotle aioli and fresh lime",
"$13.00",
"Tacos",
],
[
"Vegetable Curry",
"Mixed vegetables in aromatic coconut curry sauce with jasmine rice and naan bread",
"$12.95",
"Indian Curries",
],
[
"Tuna Poke Bowl",
"Fresh ahi tuna with avocado edamame cucumber seaweed salad over sushi rice with spicy mayo",
"$16.50",
"Poke Bowls",
],
[
"Margherita Pizza",
"Fresh mozzarella san marzano tomatoes basil and extra virgin olive oil on wood-fired crust",
"$14.00",
"Pizza",
],
[
"Chicken Tikka Masala",
"Tandoori chicken in creamy tomato sauce with aromatic spices served with basmati rice",
"$15.95",
"Indian Entrees",
],
[
"Greek Salad",
"Romaine lettuce tomatoes cucumbers kalamata olives feta cheese red onion with lemon oregano dressing",
"$10.50",
"Salads",
],
[
"Lobster Roll",
"Fresh Maine lobster meat with light mayo on toasted buttery roll served with chips",
"$22.00",
"Seafood Sandwiches",
],
[
"Quinoa Buddha Bowl",
"Organic quinoa with roasted chickpeas kale sweet potato tahini dressing and hemp seeds",
"$13.50",
"Healthy Bowls",
],
[
"Beef Pho",
"Traditional Vietnamese beef noodle soup with rice noodles fresh herbs bean sprouts and lime",
"$12.75",
"Noodle Soups",
],
[
"Eggplant Parmesan",
"Breaded eggplant layered with marinara mozzarella and parmesan served with pasta",
"$15.25",
"Italian Entrees",
],
[
"Crab Cakes",
"Maryland-style lump crab cakes with remoulade sauce and mixed greens",
"$18.50",
"Seafood Appetizers",
],
[
"Tofu Stir Fry",
"Crispy tofu with broccoli bell peppers snap peas in garlic ginger sauce over steamed rice",
"$12.50",
"Vegetarian Entrees",
],
[
"Salmon Sushi Platter",
"12 pieces of fresh salmon nigiri and sashimi with wasabi pickled ginger and soy sauce",
"$19.95",
"Sushi",
],
[
"Caprese Sandwich",
"Fresh mozzarella tomatoes basil pesto balsamic glaze on ciabatta bread",
"$11.75",
"Sandwiches",
],
[
"Tom Yum Soup",
"Spicy and sour Thai soup with shrimp lemongrass galangal mushrooms and kaffir lime leaves",
"$11.50",
"Soups",
],
[
"Lentil Dal",
"Red lentils simmered with turmeric cumin coriander served with rice and naan",
"$11.95",
"Vegan Entrees",
],
[
"Fish and Chips",
"Beer-battered cod with crispy fries malt vinegar and tartar sauce",
"$16.00",
"British Classics",
],
[
"Veggie Burger",
"House-made black bean and quinoa patty with avocado sprouts tomato on brioche bun",
"$13.25",
"Burgers",
],
[
"Miso Ramen",
"Rich miso broth with ramen noodles soft-boiled egg bamboo shoots nori and scallions",
"$14.50",
"Ramen",
],
[
"Stuffed Bell Peppers",
"Roasted bell peppers filled with rice vegetables herbs and melted cheese",
"$13.75",
"Vegetarian Entrees",
],
[
"Scallop Risotto",
"Pan-seared sea scallops over creamy parmesan risotto with white wine and lemon",
"$26.50",
"Seafood Specials",
],
[
"Spring Rolls",
"Fresh rice paper rolls with vegetables tofu rice noodles herbs and peanut dipping sauce",
"$8.95",
"Appetizers",
],
[
"Oyster Po Boy",
"Fried oysters with lettuce tomato pickles and remoulade on french bread",
"$15.50",
"Sandwiches",
],
[
"Portobello Mushroom Steak",
"Grilled portobello cap marinated in balsamic with roasted vegetables and quinoa",
"$14.95",
"Vegan Entrees",
],
[
"Coconut Shrimp",
"Jumbo shrimp breaded in shredded coconut served with sweet chili sauce",
"$14.25",
"Seafood Appetizers",
],
] as const;
// generate embeddings and prepare points
const points: any[] = [];
let idx = 0;
const embeddings = model.embed(menuItems.map(item => `${item[0]} ${item[1]}`));
for await (const embedding of embeddings) {
points.push({
id: idx,
vector: Array.from(embedding[0]),
payload: {
item_name: menuItems[idx][0],
description: menuItems[idx][1],
price: menuItems[idx][2],
category: menuItems[idx][3],
},
});
idx++;
}
// upsert points to collection
await client.upsert("items", { points });
```
## 6. Search the Menu Items
Now we can search the menu item dataset! We'll use the same `BAAI/bge-small-en-v1.5` model to embed our query text, then find the best dishes matching that embedding.
```python
# generate query embedding
query_text = "vegetarian dishes"
query_vector = next(iter(model.embed(query_text)))
# search for similar menu items
results = client.query_points(
collection_name="items",
query=query_vector,
with_payload=True,
limit=5
)
# print results
for result in results.points:
print(f"Item: {result.payload.get('item_name', 'N/A')}")
print(f"Score: {result.score}")
print(f"Description: {result.payload['description'][:150]}...")
print(f"Price: {result.payload.get('price', 'N/A')}")
print("---")
```
```rust
// generate query embedding
let query_text = "vegetarian dishes";
let query_embeddings = model
.embed(vec![query_text], None)
.expect("Failed to generate embeddings");
let query_vector = query_embeddings[0].clone();
let results = client
.query(
QueryPointsBuilder::new("items")
.query(query_vector)
.with_payload(true)
.limit(5),
)
.await
.expect("Query failed");
let na_str = "N/A".to_string();
for result in results.result {
let payload = result.payload;
println!("Item: {}", payload.get("item_name")
.and_then(|v| v.as_str()).unwrap_or(&na_str));
println!("Score: {}", result.score);
println!("Description: {}", payload.get("description")
.and_then(|v| v.as_str()).unwrap_or(&na_str));
println!("Price: {}", payload.get("price")
.and_then(|v| v.as_str()).unwrap_or(&na_str));
println!("---");
}
```
```typescript
// generate query embedding
const queryText = "vegetarian dishes";
const queryEmbedding = (await model.embed([queryText]).next()).value!
// search for similar items
const results = await client.query("items", {
query: Array.from(queryEmbedding[0]),
with_payload: true,
limit: 5,
});
// print results
for (const result of results.points) {
console.log(`Item: ${result.payload?.item_name || 'N/A'}`);
console.log(`Score: ${result.score}`);
console.log(`Description: ${result.payload?.description || 'N/A'}`);
console.log(`Price: ${result.payload?.price || 'N/A'}`);
console.log('---');
}
```
## That's Vector Search!
You can stay in the sandbox and continue trying our different API calls.</br>
When ready, use the Console and our complete REST API to try other operations.
You've just performed semantic search on real menu item data. The query "vegetarian dishes" returned similar menu items based on meaning, not just keyword matching.
## What's Next?
Now that you have a Qdrant Cloud cluster up and running, you should [test remote access](/documentation/cloud/authentication/#test-cluster-access) with a Qdrant Client.
For more about Qdrant Cloud, check our [dedicated documentation](/documentation/cloud-intro/).
- Explore [filtering](/documentation/concepts/filtering/) to combine semantic search with structured queries
- Learn about [collections](/documentation/concepts/collections/) and advanced configuration options
- Check out more [examples and tutorials](/documentation/tutorials-overview/)
@@ -9,9 +9,12 @@ weight: 2
## Inviting Users to an Account
Users can be invited via the **User Management** section, where they are assigned the **Base role** by default. Additionally, users have the option to select a specific role when inviting another user. The **Base role** is a predefined role with minimal permissions, granting users access to the platform while restricting them to viewing only their own profile.
Account users can be managed via the **User Management** section. Start by selecting a role from the dropdown and then type the name of the user you wish to manage. For users who are not in your account you will have the option to invite them. For users already in your account you can add them to the role, or see if the role has already been assigned to them.
![image.png](/documentation/cloud/role-based-access-control/user-invitation.png)
![image.png](/documentation/cloud/role-based-access-control/user-addition.png)
### Accepting an Invitation
@@ -39,6 +42,12 @@ Authorized users can give or take away roles from users in **User Management**.
![image.png](/documentation/cloud/role-based-access-control/update-user-role-edit-dialog.png)
## Making a User the Owner of an Account
Only account owners are allowed to transfer ownership of an account, this can be done via the **User Management** page. There can only be one account owner per account.
![image.png](/documentation/cloud/role-based-access-control/make-account-owner.png)
## Removing a User from an Account
Users can be removed from an account by clicking on their name in either **User Management** (via Actions). This option is only available after they've accepted the invitation to join, ensuring that only active users can be removed.
@@ -32,3 +32,7 @@ Next to the cluster endpoint which loadbalances requests across all healthy Qdra
You can finde the node specific endpoints on the cluster detail page in the Qdrant Cloud Console.
![Cluster node endpoints](/documentation/cloud/cloud-node-endpoints.png)
## Restricting Cluster Access by IP Range
You can restrict access to your cluster by specifying allowed IP ranges. This ensures that only clients connecting from the specified IP ranges can access the cluster. For more information, see [Client IP Restrictions](/documentation/cloud/configure-cluster/#client-ip-restrictions).
@@ -19,7 +19,272 @@ Logs of the database cluster are available in the Qdrant Cloud Console in the **
## Alerts
You will receive automatic alerts via email before your cluster reaches the currently configured memory or storage limits, including recommendations for scaling your cluster.
The account owner will receive automatic alerts via email if your cluster has any of the following issues:
{{< accordion >}}
- title: Memory Overutilized
content: |
**Why am I getting this alert?**
Your cluster is using more than 80% of its memory allocation for over 5 minutes.
**What does this mean for me?**
If your usage continues to grow beyond the allocation then your cluster will fail due to resource pressure, and you will see disruption.
**What can I do to resolve this?**
You have the option to scale vertically to increase existing node capacity, or horizontally to spread the load more evenly, increase capacity, and reduce overall pressure.
Alternatively, you can delete data from your cluster to reduce the amount of resources required.
**Where can I learn more about this alert?**
You can learn about high-availability and production readiness [here](/documentation/cloud/create-cluster/?q=high#creating-a-production-ready-cluster).
You can learn about vertical scaling [here](/documentation/cloud/cluster-scaling/#vertical-scaling).
You can learn about horizontal scaling [here](/documentation/cloud/cluster-scaling/#horizontal-scaling).
- title: Disk Space Overutilized
content: |
**Why am I getting this alert?**
Your cluster is using more than 80% of its disk space allocation for over 5 minutes.
**What does this mean for me?**
If your usage continues to grow beyond the allocation then your cluster will fail due to resource pressure, and you will see disruption.
**What can I do to resolve this?**
You have the option to scale vertically to increase existing node capacity, or horizontally to spread the load more evenly, increase capacity, and reduce overall pressure.
Alternatively, you can delete data from your cluster to reduce the amount of resources required.
**Where can I learn more about this alert?**
You can learn about high-availability and production readiness [here](/documentation/cloud/create-cluster/?q=high#creating-a-production-ready-cluster).
You can learn about vertical scaling [here](/documentation/cloud/cluster-scaling/#vertical-scaling).
You can learn about horizontal scaling [here](/documentation/cloud/cluster-scaling/#horizontal-scaling).
- title: A Node Ran Out of Memory
content: |
**Why am I getting this alert?**
Nodes in your cluster ran out of RAM.
**What does this mean for me?**
One or more Qdrant nodes tried to allocate more RAM than available while storing data or serving requests, which resulted in the operating system stopping the Qdrant process.
If your cluster is highly available you may avoid total downtime but the situation is still unstable. Without high availability you should expect disruption until your cluster has been scaled up.
**What can I do to resolve this?**
Ensuring your cluster is highly available will mitigate the worst case scenario, but there is still significant operational risk if untreated.
You have the option to scale vertically to increase existing node capacity, or horizontally to spread the load more evenly and reduce overall pressure.
Alternatively, you can delete data from your cluster to reduce the amount of RAM required.
**Where can I learn more about this alert?**
You can learn about high-availability and production readiness [here](/documentation/cloud/create-cluster/?q=high#creating-a-production-ready-cluster).
You can learn about vertical scaling [here](/documentation/cloud/cluster-scaling/#vertical-scaling).
You can learn about horizontal scaling [here](/documentation/cloud/cluster-scaling/#horizontal-scaling).
- title: A Node Ran Out of Disk Space
content: |
**Why am I getting this alert?**
One or more nodes in the cluster have run out of disk.
**What does this mean for me?**
Nodes that run out of disk will be unable to store new vectors.
This can lead to disruption if not resolved.
**What can I do to resolve this?**
Scaling vertically to increase disk space on existing nodes.
Scaling horizontally by adding new nodes so that shards are spread out more. Re-sharding may be necessary if there are not enough shards to distribute to new nodes.
Alternatively, you can delete data from your cluster to reduce the amount of disk required.
**Where can I learn more about this alert?**
You can learn more about disk capacity [here](/documentation/guides/capacity-planning/#scaling-disk-space-in-qdrant-cloud).
You can learn about vertical scaling [here](/documentation/cloud/cluster-scaling/#vertical-scaling).
You can learn about horizontal scaling [here](/documentation/cloud/cluster-scaling/#horizontal-scaling).
You can learn about re-sharding [here](/documentation/cloud/cluster-scaling/#resharding).
- title: Cluster Has Too Many Collections
content: |
**Why am I getting this alert?**
Your cluster has >500 collections, suggesting an anti-pattern for Qdrant.
**What does this mean for me?**
A large amount of collections brings significant resource overhead. If not addressed this will degrade resilience and even cause outages in the long term.
**What can I do to resolve this?**
A single collection with payload index partitioning is usually optimal compared to many small individual tenant collections.*
It is also possible to split collections across clusters, the [Qdrant Migration CLI](/documentation/database-tutorials/migration/) can help you with this.
**Where can I learn more about this alert?**
You can learn more about how to set up multi-tenancy with a Qdrant collection [here](/documentation/guides/multiple-partitions/).
- title: Cluster is Unhealthy
content: |
**Why am I getting this alert?**
Your cluster’s nodes have been marked as unhealthy for over 5 minutes.
**What does this mean for me?**
One or more nodes in your Qdrant Cluster has been unhealthy for longer than 5 mins indicating there is a serious issue that needs attention.
**What can I do to resolve this?**
There are many reasons for a cluster’s workloads to become unhealthy. We send proactive alerts for common scenarios and recommend checking for other alerts. It is also important to validate any available monitoring statistics for both Qdrant and your application.
We recommend checking any recent changes to client code, or configuration in your environment, looking for increases in search or write requests, to ensure no recent changes are the cause.
- title: Cluster Version is Not Covered by SLA
content: |
**Why am I getting this alert?**
Qdrant Cloud only supports the latest and previous 3 minor versions of Qdrant.
**What does this mean for me?**
Support requests against your cluster will not be covered by the SLA.
**What can I do to resolve this?**
You can upgrade you cluster version by visiting the cluster details page.
**Where can I learn more about this alert?**
Learn more about updating your cluster [here](/documentation/cloud/cluster-upgrades/)
Learn more about the Qdrant SLA and version policy [here](https://cloud.qdrant.io/sla#3-supported-versions)
- title: Database API Key is About to Expire
content: |
**Why am I getting this alert?**
A Database Key is expiring at soon.
**What does this mean for me?**
Requests using an expired key won’t be successful, this could lead to fail queries and application failures.
**What can I do to resolve this?**
If you are still using the key, you can create a new Database API Key and update your application to use the new one.
**Where can I learn more about this alert?**
Learn about the SDKs [here](/documentation/interfaces/).
Learn more about JWT Keys and permissions [here](/documentation/guides/security/?q=jwt#granular-access-control-with-jwt).
- title: A Node is CPU Throttled
content: |
**Why am I getting this alert?**
One or more cluster workloads have been CPU throttled for over 5 mins.
**What does this mean for me?**
CPU usage is constantly saturated which forces the operating system to throttle your Qdrant database nodes. This means that Qdrant will respond much slower to searches and writes.
The reason for this is usually a high write load which overloads the cluster when creating or updating indexes. A very high search load with inefficient or complex queries can contribute.
**What can I do to resolve this?**
You can reduce the amount of writes to your database. e.g. by performing the writes in batches during times when you have less search traffic, or performing less writes in parallel.
Indexing data properly will optimize query results and improve search performance reducing potential load.
While hybrid and multi-stage queries are CPU intensive techniques, they can be used to optimize and reduce the quantity of inefficient operations. On the other hand if you are using these techniques too much, you may consider simplifying some operations to be less intensive.
You can re-configure Qdrant Optimizers to reduce the load on the cluster.
It is also possible to scale your cluster horizontally or vertically for higher CPU capacity.
**Where can I learn more about this alert?**
Learn about how to optimizer Qdrant for performance and how to configure indexing [here](/documentation/concepts/indexing/) and [here](/documentation/guides/optimize/).
Learn about optimizers [here](/documentation/concepts/optimizer/).
Learn more about hybrid search [here](/documentation/concepts/hybrid-queries/).
- title: Node CPU Usage is Not Distributed Equally
content: |
**Why am I getting this alert?**
CPU usage is not consistent across nodes in your cluster.
**What does this mean for me?**
Hotspotting, where a subset of nodes handle more load than the rest, causes slow search performance and potentially node failures.*
For uneven CPU usage it typically means data distribution across nodes is uneven.
**What can I do to resolve this?**
Resharding and shard rebalancing are the primary techniques for ensuring data is evenly distributed and that requests are not concentrated on a single node.
**Where can I learn more about this alert?**
Learn more about cloud rebalancing [here](/documentation/cloud/configure-cluster/#shard-rebalancing).
- title: Node RAM or Disk Space Usage is Not Distributed Equally
content: |
**Why am I getting this alert?**
RAM/Storage usage is not consistent across nodes in your cluster.
**What does this mean for me?**
Hotspotting, where a subset of nodes handle more load than the rest, causes slow search performance and potentially node failures.*
For RAM/Storage usage it typically means data distribution across nodes is uneven.
**What can I do to resolve this?**
Resharding and rebalancing shards are both tactics to redistribute the load across your cluster evenly.
Resharding allows you to scale the number of shards in a collection up or down without recreating the collection.
This would help when your node count has been scaled up so you can reshard a collection to split it more evenly with a rebalance.
Rebalancing is the process of redistributing shards across nodes which is useful if you add a new node and need to fill the capacity. In Qdrant Cloud, rebalancing happens automatically when a cluster is scaled horizontally.
**Where can I learn more about this alert?**
Learn more about distributed deployments and resharding [here](/documentation/guides/distributed_deployment/#resharding).
Learn more about cloud rebalancing [here](/documentation/cloud/configure-cluster/#shard-rebalancing).
{{< /accordion >}}
## Qdrant Database Metrics and Telemetry
@@ -9,6 +9,8 @@ As soon as a new Qdrant version is available. Qdrant Cloud will show you an upda
To update to a new version, go to the Cluster Details page, choose the new version from the version dropdown and click **Update**.
If you are several versions behind, multiple updates might be required to reach the latest version. In this case, Qdrant Cloud will automatically perform the required intermediate updates to ensure a supported update path. You should still ensure that your client applications and used SKDs are compatible with the target version.
![Cluster Updates](/documentation/cloud/cluster-upgrades.png)
If you have a multi-node cluster and if your collections have a replication factor of at least **2**, the update process will be zero-downtime and done in a rolling fashion. You will be able to use your database cluster normally.
@@ -16,3 +18,5 @@ If you have a multi-node cluster and if your collections have a replication fact
If you have a single-node cluster or a collection with a replication factor of **1**, the update process will require a short downtime period to restart your cluster with the new version.
See also [Restart Mode](/documentation/cloud/configure-cluster/#restart-mode) for more details.
We advise taking a [backup](/documentation/cloud/backups/) before updating to allow for rollbacks.
@@ -16,6 +16,8 @@ In adition the cloud platform automatically configures the following settings fo
* The cluster mode is automatically enabled to allow distributed deployments and horizontal scaling.
* The maximum amount of payload indexes per collection is set to 100. Larger numbers of payload indexes lead to performance degradation (starting with Qdrant v1.16.0).
![Cluster node endpoints](/documentation/cloud/cloud-advanced-configuration.png)
## Collection Defaults
You can set default values for the configuration of new collections in your cluster. These defaults will be used when creating a new collection, unless you override them in the collection creation request.
@@ -44,6 +46,8 @@ Enables async scorer which uses io_uring when rescoring. See [Qdrant under the h
If configured, only the chosen IP ranges will be allowed to access the cluster. This is useful for securing your cluster and ensuring that only clients coming from trusted networks can connect to it.
![Cluster node endpoints](/documentation/cloud/cloud-ip-restrictions.png)
## Restart Mode
The cloud platform will automatically choose the optimal restart mode during version upgrades or maintenance for your cluster. If you have a multi-node cluster and one or more collections with a replication factor of at least 2, the cloud platform will use the rolling restart mode. This means that nodes in the cluster will be restarted one at a time, ensuring that the cluster remains available during the restart process.
@@ -52,6 +56,8 @@ If you have a multi-node cluster, but all collections have a replication factor
It is possible to override your cluster's default restart mode in the advanced configuration section of the Cluster Details page.
![Cluster node endpoints](/documentation/cloud/cloud-restart-mode.png)
## Shard Rebalancing
When you scale your cluster horizontally, the cloud platform will automatically rebalance shards across all nodes in the cluster, ensuring that data is evenly distributed. This is done to ensure that all nodes are utilized and that the performance of the cluster is optimal.
@@ -64,8 +70,22 @@ Qdrant Cloud offers three strategies for shard rebalancing:
You can deactivate automatic shard rebalancing by deselecting the `rebalancing_strategy` option. This is useful if you want to manually control the shard distribution across nodes.
![Cluster node endpoints](/documentation/cloud/cloud-shard-rebalancing.png)
## Rename a Cluster
You can rename a Qdrant Cluster by clicking the pencil icon next to the cluster name on the Cluster Details page.
You can rename a Qdrant Cluster from the cluster's detail page.
![Cluster Actions](/documentation/cloud/cloud-cluster-actions.png)
Renaming a cluster does not affect its functionality or configuration. The cluster's unique ID and cluster URLs will remain the same.
![Rename Cluster Dialog](/documentation/cloud/cloud-rename-cluster.png)
## Adding Labels to a Cluster
You can add labels to a Qdrant Cluster from the cluster's detail page. Labels are key-value pairs that help you organize and manage your clusters.
![Cluster Labels](/documentation/cloud/cloud-cluster-labels.png)
@@ -82,6 +82,8 @@ This page shows you how to use the Qdrant Cloud Console to create a custom Qdran
> Each node is automatically attached with a disk, that has enough space to store data with Qdrant's default collection configuration.
1. Select additional disk space for your deployment.
> Depending on your collection configuration, you may need more disk space per RAM. For example, if you configure `on_disk: true` and only use RAM for caching.
1. Choose the speed tier for your disk. (AWS only)
> Higher speed tiers provide better performance, especially for write-heavy workloads, or configurations with a low RAM cache ratio.
1. Review your cluster configuration and pricing.
1. When you're ready, select **Create**. It takes some time to provision your cluster.
@@ -99,6 +101,10 @@ To create a production-ready cluster, you need to ensure the following:
Your cluster should have at least 3 nodes, and each collection should have a replication factor of at least 2. This ensures that is one node fails, or is restarted due to maintenance, a version upgrade, or a scaling operation, that the cluster remains fully operational. You can ensure this by checking the **High Availability** checkbox when creating a cluster.
**Disk Speed (AWS only)**
We recommend the **Balanced** tier for disks >= 32 GiB, and the **Performance** tier for disks >= 256 GiB.
**Backup and Disaster Recovery**
You should create a backup schedule for your cluster. This ensures that you can restore your data in case of a disaster. You can configure backups in the **Backups** section of the cluster detail page. See [**Backups**](/documentation/cloud/backups/) for more information.
@@ -10,7 +10,7 @@ weight: 81
Qdrant Managed Cloud allows you to use inference directly in the cloud, without the need to set up and maintain your own inference infrastructure. You can use [embedding models hosted on Qdrant Cloud](#cloud-inference), or use [externally hosted models](#use-external-models).
<aside role="alert">
Inference is currently executed within a US region, even if the Qdrant Cloud cluster is hosted in another region.
Inference is executed within the EU for Qdrant clusters in EU regions and in the US for Qdrant Clusters in all other regions. Free models are hosted on US region only.
</aside>
![Cluster Cluster UI](/documentation/cloud/cloud-inference.png)
@@ -29,20 +29,25 @@ Clusters on Qdrant Managed Cloud can access embedding models that are hosted on
You can see the list of supported models in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. The list includes models for text, both to produce dense and sparse vectors, as well as multi-modal models for images.
### Free Embedding Models
Several embedding models can be used for free with Qdrant Cloud Inference, also in combination with clusters on the Qdrant Cloud free tier. Free models are identified by the "Cost: Free" label in the Inference tab of the Cluster Detail page.
### Billing
Inference is billed based on the number of tokens processed by the model. The cost is calculated per 1,000,000 tokens. The price depends on the model and is displayed on the Inference tab of the Cluster Detail page. You also can see the current usage of each model there.
Usage of non-free embedding models is billed based on the number of tokens processed by the model. The cost is calculated per 1,000,000 tokens. The price depends on the model and is displayed on the Inference tab of the Cluster Detail page. You also can see the current usage of each model there.
## Use External Models
Qdrant Cloud can act as a proxy for the APIs of three external embedding model providers:
Qdrant Cloud can act as a proxy for the following external embedding providers:
- OpenAI
- Cohere
- Jina AI
- OpenRouter
This enables you to access any of the embedding models provided by these providers through the Qdrant API.
### Billing
To use an external provider's embedding model, you need an API key from that provider. Billing is managed directly through the external provider, based on API key usage. Refer to each external embedding model provider's website for pricing details.
To use an external provider's embedding model, you need an API key from that provider. Billing is managed directly through the external provider, based on API key usage. Refer to each external embedding model provider's website for pricing details.
@@ -27,6 +27,10 @@ A [Payload](/documentation/concepts/payload/) describes information that you can
[Search](/documentation/concepts/search/) describes _similarity search_, which set up related objects close to each other in vector space.
## Search Relevance
[Search Relevance](/documentation/concepts/search-relevance/) describes techniques to improve the ranking of search results by considering factors beyond vector similarity scores.
## Explore
[Explore](/documentation/concepts/explore/) includes several APIs for exploring data in your collections.
@@ -170,7 +170,7 @@ The following parameters can be updated:
* `hnsw_config` - see [indexing](/documentation/concepts/indexing/#vector-index) for details.
* `quantization_config` - see [quantization](/documentation/guides/quantization/#setting-up-quantization-in-qdrant) for details.
* `vectors_config` - vector-specific configuration, including individual `hnsw_config`, `quantization_config` and `on_disk` settings.
* `params` - other collection parameters, including `write_consistency_factor` and `on_disk_payload`.
* `params` - other collection parameters, including `read_fan_out_delay_ms`, `write_consistency_factor` and `on_disk_payload`.
* `strict_mode_config` - see [strict mode](/documentation/guides/administration/#strict-mode) for details.
Full API specification is available in [schema definitions](https://api.qdrant.tech/api-reference/collections/update-collection).
@@ -40,6 +40,9 @@ Suppose we have a set of points with the following payload:
### Must
When using `must`, the clause becomes `true` only if every condition listed inside `must` is satisfied.
In this sense, `must` is equivalent to the operator `AND`.
Example:
{{< code-snippet path="/documentation/headless/snippets/scroll-points/with-must-filter/" >}}
@@ -50,11 +53,11 @@ Filtered points would be:
[{ "id": 2, "city": "London", "color": "red" }]
```
When using `must`, the clause becomes `true` only if every condition listed inside `must` is satisfied.
In this sense, `must` is equivalent to the operator `AND`.
### Should
When using `should`, the clause becomes `true` if at least one condition listed inside `should` is satisfied.
In this sense, `should` is equivalent to the operator `OR`.
Example:
{{< code-snippet path="/documentation/headless/snippets/scroll-points/with-should-filter/" >}}
@@ -70,11 +73,11 @@ Filtered points would be:
]
```
When using `should`, the clause becomes `true` if at least one condition listed inside `should` is satisfied.
In this sense, `should` is equivalent to the operator `OR`.
### Must Not
When using `must_not`, the clause becomes `true` if none of the conditions listed inside `must_not` is satisfied.
In this sense, `must_not` is equivalent to the expression `(NOT A) AND (NOT B) AND (NOT C)`.
Example:
{{< code-snippet path="/documentation/headless/snippets/scroll-points/with-must-not-filter/" >}}
@@ -88,9 +91,6 @@ Filtered points would be:
]
```
When using `must_not`, the clause becomes `true` if none of the conditions listed inside `must_not` is satisfied.
In this sense, `must_not` is equivalent to the expression `(NOT A) AND (NOT B) AND (NOT C)`.
### Clauses combination
It is also possible to use several clauses simultaneously:
@@ -37,26 +37,45 @@ For example, in text search, it is often useful to combine dense and sparse vect
Qdrant has a few ways of fusing the results from different queries: `rrf` and `dbsf`
### Reciprocal Rank Fusion (RRF)
<a href=https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf target="_blank">
RRF</a> considers the positions of results within each query, and boosts the ones that appear closer to the top in multiple sets of results.
The formula is simple, but needs access to the rank of each result in each query.
RRF</a> considers the positions of results within each query and boosts those that appear closer to the top in multiple sets of results. The score of a document is calculated using its rank in each result set:
$$ score(d\in D) = \sum_{r_d\in R(d)} \frac{1}{k + \frac{r_d + 1}{w_r} - 1} $$
Where:
- $D$ the set of points across all results
- $R(d)$ is the set of rankings for a particular document
- $k$ is a constant (set to 2 by default)
- $r$ is an ordered set of results from one source
- $r_d$ is the rank of document $d$ in ranking $r$
- $w_r$ is the weight of ranking $r$ (set to 1 by default)
Because $w_r$ defaults to 1, without setting explicit weights, the formula can be simplified to the original RRF function:
$$ score(d\in D) = \sum_{r_d\in R(d)} \frac{1}{k + r_d} $$
Where $D$ the set of points across all results, $R(d)$ is the set of rankings for a particular document, and $k$ is a constant (set to 2 by default).
Here is an example of RRF for a query containing two prefetches against different named vectors configured to hold sparse and dense vectors, respectively.
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rrf/" >}}
#### Parametrized RRF
#### Setting RRF Constant k
_Available as of v1.16.0_
To change the value of constant $k$ in the formula, use the dedicated `rrf` query variant.
To change the value of constant $k$ in the formula, use the dedicated `rrf` query.
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rrf-k/" >}}
#### Weighted RRF
_Available as of v1.17.0_
By default, each query is assigned an equal weight. In reality, some queries are stronger, more discriminative, or more domain-specific than others. For example, a semantic search model understands meaning better than a simple keyword matcher. Assigning equal weight to both can cause the weaker model to negatively influence results, leading to a suboptimal search experience. To address this, you can assign greater weight to rankers that perform well.
The `rrf` query allows you to configure relative weights for each of the prefetches. For example, if you have two prefetches and assign a weight of 3.0 to the first and 1.0 to the second, a document ranked third in the first query scores the same as a document ranked first in the second query. In the case of non-overlapping result sets, these weights return three results from the first set for every one result from the second set.
Weights should be provided as an array of numbers, where each weight is applied to the corresponding prefetch in the order they are defined. The number of weights must match the number of prefetches.
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rrf-weights/" >}}
### Distribution-Based Score Fusion (DBSF)
@@ -101,144 +120,6 @@ It is possible to combine all the above techniques in a single query:
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rescoring-multistage/" >}}
### Maximal Marginal Relevance (MMR)
_Available as of v1.15.0_
A useful algorithm to improve the diversity of the results is [Maximal Marginal Relevance (MMR)](https://www.cs.cmu.edu/~jgc/publication/The_Use_MMR_Diversity_Based_LTMIR_1998.pdf). It excels when the dataset has many redundant or very similar points for a query.
MMR selects candidates iteratively, starting with the most relevant point (higher similarity to the query). For each next point, it selects the one that hasn't been chosen yet which has the best combination of relevance and higher separation to the already selected points.
$$
MMR = \arg \max_{D_i \in R\setminus S}[\lambda sim(D_i, Q) - (1 - \lambda)\max_{D_j \in S}sim(D_i, D_j)]
$$
<figcaption align="center">Where $R$ is the candidates set, $S$ is the selected set, $Q$ is the query vector, $sim$ is the similarity function, and $\lambda = 1 - diversity$.</figcaption>
<br>
This is implemented in Qdrant as a parameter of a nearest neighbors query. You define the vector to get the nearest candidates, and a `diversity` parameter which controls the balance between relevance (0.0) and diversity (1.0).
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-mmr/" >}}
**Caveat:** Since MMR ranks one point at a time, the scores produced by MMR in Qdrant refer to the similarity to the query vector. This means that the response will not be ordered by score, but rather by the order of selection of MMR.
## Score boosting
_Available as of v1.14.0_
When introducing vector search to specific applications, sometimes business logic needs to be considered for ranking the final list of results.
A quick example is [our own documentation search bar](https://github.com/qdrant/page-search).
It has vectors for every part of the documentation site. If one were to perform a search by "just" using the vectors, all kinds of elements would be equally considered good results.
However, when searching for documentation, we can establish a hierarchy of importance:
`title > content > snippets`
One way to solve this is to weight the results based on the kind of element.
For example, we can assign a higher weight to titles and content, and keep snippets unboosted.
Pseudocode would be something like:
`score = score + (is_title * 0.5) + (is_content * 0.25)`
Query API can rescore points with custom formulas. They can be based on:
- Dynamic payload values
- Conditions
- Scores of prefetches
To express the formula, the syntax uses objects to identify each element.
Taking the documentation example, the request would look like this:
{{< code-snippet path="/documentation/headless/snippets/query-points/score-boost-tags/" >}}
There are multiple expressions available, check the [API docs for specific details](https://api.qdrant.tech/v-1-14-x/api-reference/search/query-points#request.body.query.Query%20Interface.Query.Formula%20Query.formula).
- **constant** - A floating point number. e.g. `0.5`.
- `"$score"` - Reference to the score of the point in the prefetch. This is the same as `"$score[0]"`.
- `"$score[0]"`, `"$score[1]"`, `"$score[2]"`, ... - When using multiple prefetches, you can reference specific prefetch with the index within the array of prefetches.
- **payload key** - Any plain string will refer to a payload key. This uses the jsonpath format used in every other place, e.g. `key` or `key.subkey`. It will try to extract a number from the given key.
- **condition** - A filtering condition. If the condition is met, it becomes `1.0`, otherwise `0.0`.
- **mult** - Multiply an array of expressions.
- **sum** - Sum an array of expressions.
- **div** - Divide an expression by another expression.
- **abs** - Absolute value of an expression.
- **pow** - Raise an expression to the power of another expression.
- **sqrt** - Square root of an expression.
- **log10** - Base 10 logarithm of an expression.
- **ln** - Natural logarithm of an expression.
- **exp** - Exponential function of an expression (`e^x`).
- **geo distance** - Haversine distance between two geographic points. Values need to be `{ "lat": 0.0, "lon": 0.0 }` objects.
- **decay** - Apply a decay function to an expression, which clamps the output between 0 and 1. Available decay functions are **linear**, **exponential**, and **gaussian**. [See more](#boost-points-closer-to-user).
- **datetime** - Parse a datetime string (see formats [here](/documentation/concepts/payload/#datetime)), and use it as a POSIX timestamp, in seconds.
- **datetime key** - Specify that a payload key contains a datetime string to be parsed into POSIX seconds.
It is possible to define a default for when the variable (either from payload or prefetch score) is not found. This is given in the form of a mapping from variable to value.
If there is no variable, and no defined default, a default value of `0.0` is used.
<aside role="status">
**Considerations when using formula queries:**
- Formula queries can only be used as a rescoring step.
- Formula results are always sorted in descending order (bigger is better). **For euclidean scores, make sure to negate them** to sort closest to farthest.
- If a score or variable is not available, and there is no default value, it will return an error.
- If a value is not a number (or the expected type), it will return an error.
- To leverage payload indices, single-value arrays are considered the same as the inner value. For example: `[0.2]` is the same as `0.2`, but `[0.2, 0.7]` will be interpreted as `[0.2, 0.7]`
- Multiplication and division are lazily evaluated, meaning that if a 0 is encountered, the rest of operations don't execute (e.g. `0.0 * condition` won't check the condition).
- Payload variables used within the formula also benefit from having payload indices. Please try to always have a payload index set up for the variables used in the formula for better performance.
</aside>
### Boost points closer to user
Another example. Combine the score with how close the result is to a user.
Considering each point has an associated geo location, we can calculate the distance between the point and the request's location.
Assuming we have cosine scores in the prefetch, we can use a helper function to clamp the geographical distance between 0 and 1, by using a decay function. Once clamped, we can sum the score and the distance together. Pseudocode:
`score = score + gauss_decay(distance)`
In this case we use a **gauss_decay** function.
{{< code-snippet path="/documentation/headless/snippets/query-points/score-boost-closer-to-user/" >}}
### Time-based score boosting
Or combine the score with the information on how "fresh" the result is. It's applicable to (news) articles and in general many other different types of searches (think of the "newest" filter you use in applications).
To implement time-based score boosting, you'll need each point to have a datetime field in its payload, e.g., when the item was uploaded or last updated. Then we can calculate the time difference in seconds between this payload value and the current time, our `target`.
With an exponential decay function, perfect for use cases with time, as freshness is a very quickly lost quality, we can convert this time difference into a value between 0 and 1, then add it to the original score to prioritise fresh results.
`score = score + exp_decay(current_time - point_time)`
That's how it will look for an application where, after 1 day, results start being only half-relevant (so get a score of 0.5):
{{< code-snippet path="/documentation/headless/snippets/query-points/score-boost-time/" >}}
For all decay functions, there are these parameters available
| Parameter | Default | Description |
| ---------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `x` | N/A | The value to decay |
| `target` | 0.0 | The value at which the decay will be at its peak. For distances it is usually set at 0.0, but can be set to any value. |
| `scale` | 1.0 | The value at which the decay function will be equal to `midpoint`. This is in terms of `x` units, for example, if `x` is in meters, `scale` of 5000 means 5km. Must be a non-zero positive number |
| `midpoint` | 0.5 | Output is `midpoint` when `x` equals `target` ± `scale`. Must be in the range (0.0, 1.0), exclusive |
![Decay functions.](/docs/decay-function.png)
The [formulas for each decay function](https://www.desmos.com/calculator/idv5hknwb1) are as follows:
<br>
| Decay Function | Color | Range | Formula |
|----------------|-------|-------|---------|
| **`lin_decay`** | green | `[0, 1]` | $\text{lin_decay}(x) = \max\left(0,\ -\frac{(1-m_{idpoint})}{s_{cale}}\cdot {abs}(x-t_{arget})+1\right)$ |
| **`exp_decay`** | red | `(0, 1]` | $\text{exp_decay}(x) = \exp\left(\frac{\ln(m_{idpoint})}{s_{cale}}\cdot {abs}(x-t_{arget})\right)$ |
| **`gauss_decay`** | purple | `(0, 1]` | $\text{gauss_decay}(x) = \exp\left(\frac{\ln(m_{idpoint})}{s_{cale}^{2}}\cdot (x-t_{arget})^{2}\right)$ |
## Grouping
_Available as of v1.11.0_
@@ -45,6 +45,8 @@ Payload index may occupy some additional memory, so it is recommended to only us
If you need to filter by many fields and the memory limits do not allow for indexing all of them, it is recommended to choose the field that limits the search result the most.
As a rule, the more different values a payload value has, the more efficiently the index will be used.
<aside role="alert">It's highly recommended to create all payload indices immediately after collection creation. Creating them later may block updates for some time. HNSW graphs will also only benefit from <a href="#filterable-index">additional optimizations</a> (extra edges) when they are generated after payload index creation.</aside>
### Parameterized index
*Available as of v1.8.0*
@@ -177,7 +179,7 @@ The choice of tokenizer affects how queries match the indexed text, supporting d
Available tokenizers are:
* `word` - splits the string into words, separated by spaces, punctuation marks, and special characters.
* `word` (default) - splits the string into words, separated by spaces, punctuation marks, and special characters.
* `whitespace` - splits the string into words, separated by spaces.
* `prefix` - splits the string into words, separated by spaces, punctuation marks, and special characters, and then creates a prefix index for each word. For example: `hello` will be indexed as `h`, `he`, `hel`, `hell`, `hello`.
* `multilingual` - a special type of tokenizer based on multiple packages like [charabia](https://github.com/meilisearch/charabia) and [vaporetto](https://github.com/daac-tools/vaporetto) to deliver fast and accurate tokenization for a large variety of languages. It allows proper tokenization and lemmatization for multiple languages, including those with non-Latin alphabets and non-space delimiters. See the [charabia documentation](https://github.com/meilisearch/charabia) for a full list of supported languages and normalization options. Note: For the Japanese language, Qdrant relies on the `vaporetto` project, which has much less overhead compared to `charabia`, while maintaining comparable performance.
@@ -210,7 +212,7 @@ When configuring a full-text index in Qdrant, you can specify a stemmer to be us
Qdrant provides an implementation of [Snowball stemmer](https://snowballstem.org/), a widely used and performant variant for some of the most popular languages.
For the list of supported languages, please visit the [rust-stemmers repository](https://github.com/qdrant/rust-stemmers).
Here is an example of full-text Index configuration with Snowball stemmer:
For full-text indices, stemming is not enabled by default. To enable it, configure the `snowball` stemmer with the desired language:
{{< code-snippet path="/documentation/headless/snippets/create-payload-index/stemmer-full-text/" >}}
@@ -222,8 +224,7 @@ In Qdrant, you can specify a list of stopwords to be ignored during full-text in
You can configure stopwords based on predefined languages, as well as extend existing stopword lists with custom words.
Here is an example of configuring a full-text index with custom stopwords:
For full-text indices, stopword removal is not enabled by default. To enable it, configure the `stopwords` parameter with the desired languages and any custom stopwords:
{{< code-snippet path="/documentation/headless/snippets/create-payload-index/stopwords-full-text/" >}}
@@ -287,6 +288,45 @@ The HNSW parameters can also be configured on a collection and named vector
level by setting [`hnsw_config`](/documentation/concepts/indexing/#vector-index) to fine-tune search
performance.
### Filterable HNSW Index
Separately, a payload index and a vector index cannot completely address the challenges of filtered search.
In the case of high-selectivity (weak) filters, you can use the HNSW index as it is.
In the case of low-selectivity (strict) filters, you can use the payload index and do a complete rescore.
However, for cases in the middle, this approach does not work well.
On one hand, we cannot apply a full scan on too many vectors.
On the other hand, the HNSW graph starts to fall apart when using filters that are too strict.
![HNSW fail](/docs/precision_by_m.png)
<!-- ![hnsw graph](/docs/graph.gif) -->
Qdrant solves this problem by extending the HNSW graph with additional edges based on indexed payload values.
Extra edges allow you to efficiently search for nearby vectors using the HNSW index and apply filters as you search in the graph.
You can find more information on this approach in our [article](/articles/filterable-hnsw/).
#### The ACORN Search Algorithm
*Available as of v1.16.0*
In some cases, the additional edges built for Qdrant's filterable HNSW may not be sufficient.
These extra edges are added for each payload index separately, but not for every possible combination of payload indices.
As a result, a combination of two or more strict filters might still lead to disconnected graph components.
The same can happen when there are a large number of soft-deleted points in the graph.
In such cases, use the [ACORN Search Algorithm](/documentation/concepts/search/#acorn-search-algorithm).
When using ACORN, during graph traversal, it explores not just direct neighbors (first hop), but also neighbors of neighbors (second hop) when direct neighbors are filtered out. This improves search accuracy at the cost of performance.
#### Disable the Creation of Extra Edges for Payload Fields
*Available as of v1.17.0*
Not all payload indices may be intended for use with dense vector search. For example, when a collection contains both dense and sparse vectors, some payload fields may only be used to filter sparse vector searches. Since sparse vector search does not use the HNSW index, it is unnecessary to build extra edges in the HNSW graph for these fields. Creating extra edges adds indexing latency and increases the size of the HNSW graph, which consumes memory as well as disk space, so you may want to disable it for fields that do not require it.
You can disable the creation of extra edges for an indexed payload field by setting `enable_hnsw` to `false` when configuring a payload index:
{{< code-snippet path="/documentation/headless/snippets/create-payload-index/disable-hnsw/" >}}
## Sparse Vector Index
*Available as of v1.7.0*
@@ -339,28 +379,3 @@ Where:
- `N` is the total number of documents in the collection.
- `n` is the number of documents containing non-zero values for the given vector element.
## Filtrable Index
Separately, a payload index and a vector index cannot solve the problem of search using the filter completely.
In the case of high-selectivity (weak) filters, you can use the HNSW index as it is.
In the case of low-selectivity (strict) filters, you can use the payload index and complete rescore.
However, for cases in the middle, this approach does not work well.
On the one hand, we cannot apply a full scan on too many vectors.
On the other hand, the HNSW graph starts to fall apart when using too strict filters.
![HNSW fail](/docs/precision_by_m.png)
<!-- ![hnsw graph](/docs/graph.gif) -->
Qdrant solves this problem by extending the HNSW graph with additional edges based on the stored payload values.
Extra edges allow you to efficiently search for nearby vectors using the HNSW index and apply filters as you search in the graph.
You can find more information on this approach in our [article](/articles/filtrable-hnsw/).
However, in some cases, these additional edges might not be enough.
These extra edges are added per each payload index separately, but not per each possible combination of them.
So, a combination of two or more strict filters still might lead to disconnected graph components.
The same may happen when having a large number of soft-deleted points in the graph.
In such cases, the [ACORN Search Algorithm](/documentation/concepts/search/#acorn-search-algorithm) can be used.
@@ -9,7 +9,7 @@ aliases:
Inference is the process of using a machine learning model to create vector embeddings from text, images, or other data types. While you can create embeddings on the client side, you can also let Qdrant generate them while storing or querying data.
![Inference.](/docs/inference.png)
![Inference](/docs/inference.png)
There are several advantages to generating embeddings with Qdrant:
@@ -175,14 +175,17 @@ This flexibility allows you to develop and test your applications locally or in
## External Embedding Model Providers
Qdrant Cloud can act as a proxy for the APIs of three external embedding model providers:
Qdrant Cloud can act as a proxy for the APIs of external embedding model providers:
- OpenAI
- Cohere
- Jina AI
- OpenRouter
This enables you to access any of the embedding models provided by these providers through the Qdrant API.
![Inference with an external embedding model provider](/docs/inference-external-provider.png)
To use an external provider's embedding model, you need an API key from that provider. For example, to access OpenAI models, you need an OpenAI API key. Qdrant does not store or cache your API keys; they must be provided with each inference request.
When using an external embedding model, ensure that your collection has been configured for vectors with the correct dimensionality. Refer to the model's documentation for details on the output dimensions.
@@ -199,7 +202,7 @@ When using a model from an external provider, refer to the model's documentation
When you prepend a model name with `openai/`, the embedding request is automatically routed to the [OpenAI Embeddings API](https://platform.openai.com/docs/guides/embeddings).
For example, to use OpenAI's `text-embedding-3-large` model when ingesting data, prepend the model name with `openai/` and provide your OpenAI API key in the `options` object. Any OpenAI-specific API parameters can be passed using the `options` object. This example uses the OpenAI-specific API `dimensions` parameter to reduce the dimensionality to 512:
For example, to use OpenAI's `text-embedding-3-large` model when ingesting data, prepend the model name with `openai/`. Provide your OpenAI API key in the request header, or in the request body in the `options` object. Any OpenAI-specific API parameters can be passed using the `options` object. This example uses the OpenAI-specific API `dimensions` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/openai-upsert/" >}}
@@ -215,7 +218,7 @@ Note that, because Qdrant does not store or cache your OpenAI API key, you need
When you prepend a model name with `cohere/`, the embedding request is automatically routed to the [Cohere Embed API](https://docs.cohere.com/reference/embed).
For example, to use Cohere's multimodal `embed-v4.0` model when ingesting data, prepend the model name with `cohere/` and provide your Cohere API key in the `options` object. This example uses the Cohere-specific API `output_dimension` parameter to reduce the dimensionality to 512:
For example, to use Cohere's multimodal `embed-v4.0` model when ingesting data, prepend the model name with `cohere/`. Provide your Cohere API key in the request header, or in the request body in the `options` object. This example uses the Cohere-specific API `output_dimension` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/cohere-upsert/" >}}
@@ -231,7 +234,7 @@ Note that, because Qdrant does not store or cache your Cohere API key, you need
When you prepend a model name with `jinaai/`, the embedding request is automatically routed to the [Jina AI Embedding API](https://jina.ai/embeddings/).
For example, to use Jina AI's multimodal `jina-clip-v2` model when ingesting data, prepend the model name with `jinaai/` and provide your Jina AI API key in the `options` object. This example uses the Jina AI-specific API `dimensions` parameter to reduce the dimensionality to 512:
For example, to use Jina AI's multimodal `jina-clip-v2` model when ingesting data, prepend the model name with `jinaai/`. Provide your Jina AI API key in the request header, or in the request body in the `options` object. This example uses the Jina AI-specific API `dimensions` parameter to reduce the dimensionality to 512:
{{< code-snippet path="/documentation/headless/snippets/inference/jinaai-upsert/" >}}
@@ -241,6 +244,20 @@ At query time, you can use the same model by prepending the model name with `jin
Note that, because Qdrant does not store or cache your Jina AI API key, you need to provide it with each inference request
### OpenRouter
OpenRouter is a platform that provides [several embedding models](https://openrouter.ai/models?fmt=cards&output_modalities=embeddings). To use one of the models provided by the [OpenRouter Embeddings API](https://openrouter.ai/docs/api/reference/embeddings), prepend the model name with `openrouter/`.
For example, to use the `mistralai/mistral-embed-2312` model when ingesting data, prepend the model name with `openrouter/`. Provide your OpenRouter API key in the request header, or in the request body in the `options` object.
{{< code-snippet path="/documentation/headless/snippets/inference/openrouter-upsert/" >}}
At query time, you can use the same model by prepending the model name with `openrouter/` and providing your OpenRouter API key in the `options` object:
{{< code-snippet path="/documentation/headless/snippets/inference/openrouter-query/" >}}
Note that, because Qdrant does not store or cache your OpenRouter API key, you need to provide it with each inference request.
## Multiple Inference Operations
You can run multiple inference operations within a single request, even when models are hosted in different locations. This example generates three different named vectors for a single point: image embeddings using `jina-clip-v2` hosted by Jina AI, text embeddings using `all-minilm-l6-v2` hosted by Qdrant Cloud, and BM25 embeddings using the `bm25` model executed locally by the Qdrant cluster:
@@ -12,7 +12,7 @@ It is much more efficient to apply changes in batches than perform each change i
Storage optimization in Qdrant occurs at the segment level (see [storage](/documentation/concepts/storage/)).
In this case, the segment to be optimized remains readable for the time of the rebuild.
![Segment optimization](/docs/optimization.svg)
![Segment optimization](/articles_data/immutable-data-structures/optimization.png)
The availability is achieved by wrapping the segment into a proxy that transparently handles data changes.
Changed data is placed in the copy-on-write segment, which has priority for retrieval and subsequent updates.
@@ -119,3 +119,34 @@ storage:
In addition to the configuration file, you can also set optimizer parameters separately for each [collection](/documentation/concepts/collections/).
Dynamic parameter updates may be useful, for example, for more efficient initial loading of points. You can disable indexing during the upload process with these settings and enable it immediately after it is finished. As a result, you will not waste extra computation resources on rebuilding the index.
## Optimization Monitoring
*Available as of v1.17.0*
The `/collections/{collection_name}/optimizations` API endpoint returns information about the optimization of a specific collection, including:
- A summary of optimization activity, with the number of queued optimizations, queued segments, queued points, and idle segments (segments that need no optimization).
- Details about any currently running optimization, including:
- the specific optimizer
- its status
- the segments involved
- its progress
Optionally, you can use the `with` query parameter with one or more of the following comma-separated values to retrieve additional information:
- `queued`, to return a list of queued optimizations
- `completed`, to return a list of completed optimizations
- `idle_segments`, to return a list of idle segments
For example:
{{< code-snippet path="/documentation/headless/snippets/optimizations/" >}}
### Web UI
The same information is also accessible via the **Optimizations** tab within the **Collections** interface in [the Web UI](/documentation/web-ui/). For a specific collection, this tab provides an overview of the current optimization status and a timeline of current and past optimization cycles:
![The Optimizations tab in Web UI shows progress and a timeline of optimization cycles](/docs/web-ui-optimizations-progress-timeline.png)
Selecting a specific optimization cycle from the timeline provides detailed information about the tasks performed during that cycle, including their durations:
![The Optimizations tab in Web UI provides access to detailed information about optimization tasks and their durations](/docs/web-ui-optimizations-tree.png)
@@ -164,9 +164,29 @@ In this case, it means that points with the same id will be overwritten when re-
Idempotence property is useful if you use, for example, a message queue that doesn't provide an exactly-once guarantee.
Even with such a system, Qdrant ensures data consistency.
### Update Mode
_Available as of v1.17.0_
By default, an upsert operation inserts a point if it does not exist, or updates it if it does. To change this behavior, use the `update_mode` parameter:
- `upsert` (default): Insert a point if it does not exist, or update it if it does.
- `insert_only`: Insert a point only if it does not already exist. If a point with the same ID exists, the operation is ignored.
- `update_only`: Update a point only if it already exists. Points that do not exist are not inserted.
For example, to use `insert_only` mode:
{{< code-snippet path="/documentation/headless/snippets/insert-points/update-mode/" >}}
`insert_only` mode is especially useful when [migrating from one embedding model to another](/documentation/database-tutorials/embedding-model-migration/), where conflicts between regular updates and background re-embedding tasks need to be resolved.
{{< figure src="/docs/embedding-model-migration.png" caption="Embedding model migration in blue-green deployment" width="80%" >}}
`update_only` mode is useful with [conditional updates](#conditional-updates). Because upserts default to inserts for non-existing points, a conditional update without an explicit `update_mode` will insert a new point even if the condition is not met, which is not the intended behavior in most cases.
### Named vectors
[_Available as of v0.10.0_](#create-vector-name)
_Available as of v0.10.0_
If the collection was created with multiple vectors, each vector data can be provided using the vector's name:
@@ -291,6 +311,8 @@ All update operations (including point insertion, vector updates, payload update
{{< code-snippet path="/documentation/headless/snippets/insert-points/with-condition/" >}}
<aside role="alert">By default, a conditional update on a non-existent point behaves as a regular upsert, inserting the point regardless of the filter. This is undesirable in most cases. To ensure that only existing points that meet the condition are updated, <a href="#update-mode">set <code>update_mode</code></a> to <code>update_only</code>.</aside>
While conditional payload modification and deletion covers the use-case of mass data modification, conditional point insertion and vector updates are particularly useful for implementing optimistic concurrency control in distributed systems.
A common scenario for such mechanism is when multiple clients try to update the same point independently.
@@ -309,10 +331,6 @@ If Client B tries to write back its changes later, the condition would fail (as
Instead of `version`, applications can use timestamps (assuming synchronized clocks) or any other monotonically increasing value that fits their data model.
This mechanism is especially useful in the scenarios of embedding model migration, where we need to resolve conflicts between regular application updates and background re-embedding tasks.
{{< figure src="/docs/embedding-model-migration.png" caption="Embedding model migration in blue-green deployment" width="80%" >}}
## Retrieve points
There is a method for retrieving points by their ids.
@@ -0,0 +1,217 @@
---
title: Search Relevance
weight: 52
---
# Search Relevance
By default, Qdrant ranks search results based on vector similarity scores. However, you may wish to consider additional factors when ranking results. Qdrant offers several tools to help you accomplish this.
## Score Boosting
_Available as of v1.14.0_
When introducing vector search to specific applications, sometimes business logic needs to be considered for ranking the final list of results.
A quick example is [our own documentation search bar](https://github.com/qdrant/page-search).
It has vectors for every part of the documentation site. If one were to perform a search by "just" using the vectors, all kinds of elements would be equally considered good results.
However, when searching for documentation, we can establish a hierarchy of importance:
`title > content > snippets`
One way to solve this is to weight the results based on the kind of element.
For example, we can assign a higher weight to titles and content and keep snippets unboosted.
Pseudocode would be something like:
`score = score + (is_title * 0.5) + (is_content * 0.25)`
The Query API can rescore points with custom formulas based on:
- Dynamic payload values
- Conditions
- Scores of prefetches
To express the formula, the syntax uses objects to identify each element.
Taking the documentation example, the request would look like this:
{{< code-snippet path="/documentation/headless/snippets/query-points/score-boost-tags/" >}}
There are multiple expressions available. Check the [API docs for specific details](https://api.qdrant.tech/v-1-14-x/api-reference/search/query-points#request.body.query.Query%20Interface.Query.Formula%20Query.formula).
- **constant** - A floating point number. e.g. `0.5`.
- `"$score"` - Reference to the score of the point in the prefetch. This is the same as `"$score[0]"`.
- `"$score[0]"`, `"$score[1]"`, `"$score[2]"`, ... - When using multiple prefetches, you can reference specific prefetch with the index within the array of prefetches.
- **payload key** - Any plain string will refer to a payload key. This uses the jsonpath format used in every other place, e.g. `key` or `key.subkey`. It will try to extract a number from the given key.
- **condition** - A filtering condition. If the condition is met, it becomes `1.0`, otherwise `0.0`.
- **mult** - Multiply an array of expressions.
- **sum** - Sum an array of expressions.
- **div** - Divide an expression by another expression.
- **abs** - Absolute value of an expression.
- **pow** - Raise an expression to the power of another expression.
- **sqrt** - Square root of an expression.
- **log10** - Base 10 logarithm of an expression.
- **ln** - Natural logarithm of an expression.
- **exp** - Exponential function of an expression (`e^x`).
- **geo distance** - Haversine distance between two geographic points. Values need to be `{ "lat": 0.0, "lon": 0.0 }` objects.
- **decay** - Apply a decay function to an expression, which clamps the output between 0 and 1. Available decay functions are **linear**, **exponential**, and **gaussian**. [See more](#decay-functions).
- **datetime** - Parse a datetime string (see formats [here](/documentation/concepts/payload/#datetime)), and use it as a POSIX timestamp in seconds.
- **datetime key** - Specify that a payload key contains a datetime string to be parsed into POSIX seconds.
It is possible to define a default for when the variable (either from payload or prefetch score) is not found. This is given in the form of a mapping from variable to value.
If there is no variable and no defined default, a default value of `0.0` is used.
<aside role="status">
**Considerations when using formula queries:**
- Formula queries can only be used as a rescoring step.
- Formula results are always sorted in descending order (bigger is better). **For Euclidean scores, make sure to negate them** to sort closest to farthest.
- If a score or variable is not available and there is no default value, it will return an error.
- If a value is not a number (or the expected type), it will return an error.
- To leverage payload indices, single-value arrays are considered the same as the inner value. For example, `[0.2]` is the same as `0.2`, but `[0.2, 0.7]` will be interpreted as `[0.2, 0.7]`
- Multiplication and division are lazily evaluated, meaning that if a 0 is encountered, the rest of the operations don't execute (for example, `0.0 * condition` won't check the condition).
- Payload variables used within the formula also benefit from having payload indices. Please try to always have a payload index set up for the variables used in the formula for better performance.
</aside>
### Decay Functions
Decay functions enable you to modify the score based on how far a value is from a target using a linear, exponential, or Gaussian decay function. For all decay functions, these are the available parameters:
| Parameter | Default | Description |
| ---------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `x` | N/A | The value to decay |
| `target` | 0.0 | The value at which the decay will be at its peak. For distances, it is usually set at 0.0, but can be set to any value. |
| `scale` | 1.0 | The value at which the decay function will be equal to `midpoint`. This is in terms of `x` units. For example, if `x` is in meters, `scale` of 5000 means 5km. Must be a non-zero positive number. |
| `midpoint` | 0.5 | Output is `midpoint` when `x` equals `target` ± `scale`. Must be in the range (0.0, 1.0), exclusive. |
![Decay functions.](/docs/decay-function.png)
The [formula for each decay function](https://www.desmos.com/calculator/idv5hknwb1) is as follows:
<br>
| Decay Function | Range | Formula |
|----------------|-------|-------|---------|
| **`lin_decay`** | `[0, 1]` | $\text{lin_decay}(x) = \max\left(0,\ -\frac{(1-m_{idpoint})}{s_{cale}}\cdot {abs}(x-t_{arget})+1\right)$ |
| **`exp_decay`** | `(0, 1]` | $\text{exp_decay}(x) = \exp\left(\frac{\ln(m_{idpoint})}{s_{cale}}\cdot {abs}(x-t_{arget})\right)$ |
| **`gauss_decay`** | `(0, 1]` | $\text{gauss_decay}(x) = \exp\left(\frac{\ln(m_{idpoint})}{s_{cale}^{2}}\cdot (x-t_{arget})^{2}\right)$ |
#### Boost Points Closer to User
An example of decay functions is to combine the score with how close a result is to a user.
Considering each point has an associated geo location, we can calculate the distance between the point and the request's location.
Assuming we have cosine scores in the prefetch, we can use a helper function to clamp the geographical distance between 0 and 1, by using a decay function. Once clamped, we can sum the score and the distance together. Pseudocode:
`score = score + gauss_decay(distance)`
In this case, we use a **gauss_decay** function.
{{< code-snippet path="/documentation/headless/snippets/query-points/score-boost-closer-to-user/" >}}
#### Time-Based Score Boosting
Or combine the score with the information on how "fresh" the result is. It's applicable to (news) articles and, in general, many other different types of searches (think of the "newest" filter you use in applications).
To implement time-based score boosting, you'll need each point to have a datetime field in its payload, e.g., when the item was uploaded or last updated. Then we can calculate the time difference in seconds between this payload value and the current time, our `target`.
With an exponential decay function, perfect for use cases with time, as freshness is a very quickly lost quality, we can convert this time difference into a value between 0 and 1, then add it to the original score to prioritize fresh results.
`score = score + exp_decay(current_time - point_time)`
That's how it will look for an application where, after 1 day, results start being only half-relevant (so get a score of 0.5):
{{< code-snippet path="/documentation/headless/snippets/query-points/score-boost-time/" >}}
## Maximal Marginal Relevance (MMR)
_Available as of v1.15.0_
[Maximal Marginal Relevance (MMR)](https://www.cs.cmu.edu/~jgc/publication/The_Use_MMR_Diversity_Based_LTMIR_1998.pdf) is an algorithm to improve the diversity of the results. It excels when the dataset has many redundant or very similar points for a query.
MMR selects candidates iteratively, starting with the most relevant point (higher similarity to the query). For each next point, it selects the one that hasn't been chosen yet which has the best combination of relevance and higher separation to the already selected points.
$$
MMR = \arg \max_{D_i \in R\setminus S}[\lambda sim(D_i, Q) - (1 - \lambda)\max_{D_j \in S}sim(D_i, D_j)]
$$
<figcaption align="center">Where $R$ is the candidates set, $S$ is the selected set, $Q$ is the query vector, $sim$ is the similarity function, and $\lambda = 1 - diversity$.</figcaption>
<br>
This is implemented in Qdrant as a parameter of a nearest neighbors query. You define the vector to get the nearest candidates, and a `diversity` parameter which controls the balance between relevance (0.0) and diversity (1.0).
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-mmr/" >}}
**Caveat:** Since MMR ranks one point at a time, the scores produced by MMR in Qdrant refer to the similarity to the query vector. This means that the response will not be ordered by score, but rather by the order of selection of MMR.
## Relevance Feedback
*Available as of 1.17*
Relevance feedback distills signals from current search results into the next retrieval iteration to surface more relevant documents.
Qdrant provides a subtype of relevance feedback-based retrieval, where feedback is given by any model (relevance oracle) in a granular fashion: it rescores top retrieved results by their relative relevance to the query. A detailed overview of relevance feedback methods can be found in [Relevance Feedback in Information Retrieval](/articles/search-feedback-loop).
To use relevance feedback-based retrieval, two components are required:
1. A collection of vectors to search through.
2. An oracle to determine the relevance of search results.
---
The idea behind using relevance feedback-based retrieval is the following:
1. Run a basic nearest neighbors search. Let's call its results **Retriever Similarity** and the algorithm behind this search -- **retriever**.
2. Use any **feedback model** to assign a relevance score to the top X search results (X is not expected to be large, 3-5 is a good option). Let's call these scores **Feedback Score**.
3. Through analyzing the Feedback Score for the top results, determine if the feedback model agrees with the retriever, or if retrieval can be improved.
4. If it can be improved, use feedback to modify retrieval (vector space traversal) to account for the discrepancies between the feedback model and the retriever.
For example, in this set of retrieved results:
| Point ID | Retriever Similarity | Feedback Score |
| --- | --- | --- |
| 111 | 0.89 | 0.68 |
| 222 | 0.81 | 0.72 |
| 333 | 0.77 | 0.61 |
The feedback model considers the second result with ID 222 to be the most relevant, which is a discrepancy with retriever's ranking. Hence, this feedback can potentially help make the next iteration of retrieval better.
---
To leverage the feedback in search across the entire collection, Qdrant provides a query interface that requires:
1. The original query (`target`), which can be a point ID, an inference object, or a raw vector.
2. A short list of initial retrieval results and their relevance score (`feedback`). Each feedback item consists of:
- `example`, which can be point ID, an inference object, or a raw vector used by the retriever.
- `score`, the feedback score.
3. A definition of the formula that modifies retrieval based on the feedback (`strategy`).
{{< code-snippet path="/documentation/headless/snippets/query-points-explore/relevance-feedback-naive/" >}}
Internally, Qdrant combines the feedback list into pairs, based on the relevance scores, and then uses these pairs in a formula that modifies vector space traversal during retrieval (changes the strategy of retrieval). This relevance feedback-based retrieval considers not only the similarity of candidates to the query but also to each feedback pair. For a more detailed description of how it works, refer to the article [Relevance Feedback in Qdrant](/articles/relevance-feedback).
The `a`, `b`, and `c` parameters of the [`naive` strategy](#naive-strategy) need to be customized for each triplet of retriever, feedback model, and collection. To get these 3 weights adapted to your setup, use [our open source Python package](https://pypi.org/project/qdrant-relevance-feedback/).
<aside role="alert">When using point IDs for <code>target</code> or <code>example</code>, these points are excluded from the search results. To include them, convert them to raw vectors first and use the raw vectors in the query.</aside>
For a hands-on tutorial on determining the parameters, using them with the Relevance Feedback Query, and evaluating the results, check out [Relevance Feedback in Qdrant](/documentation/tutorials-search-engineering/using-relevance-feedback/).
### Naive Strategy
For now, `naive` is the only available strategy.
<details>
<summary>Naive Strategy</summary>
$$
score = a * sim(query, candidate) + \sum_{pair \in pairs}{(confidence_{pair})^b * c * delta_{pair}} \\\\
$$
\begin{align}
\text{where} \\\\
confidence_{pair} &= relevance_{positive} - relevance_{negative} \\\\
delta_{pair} &= sim(positive, candidate) - sim(negative, candidate) \\\\
\end{align}
</details>
@@ -30,7 +30,7 @@ Depending on the `query` parameter, Qdrant might prefer different strategies for
| [Discovery Search](/documentation/concepts/explore/#discovery-api) | Guide the search using context as a one-shot training set |
| [Scroll](/documentation/concepts/points/#scroll-points) | Get all points with optional filtering |
| [Grouping](/documentation/concepts/search/#grouping-api) | Group results by a certain field |
| [Order By](/documentation/concepts/hybrid-queries/#re-ranking-with-stored-values) | Order points by payload key |
| [Order By](/documentation/concepts/points/#order-points-by-payload-key) | Order points by payload key |
| [Hybrid Search](/documentation/concepts/hybrid-queries/#hybrid-search) | Combine multiple queries to get better results |
| [Multi-Stage Search](/documentation/concepts/hybrid-queries/#multi-stage-queries) | Optimize performance for large embeddings |
| [Random Sampling](#random-sampling) | Get random points from the collection |
@@ -175,7 +175,7 @@ Accessing array elements by index is currently not supported.
*Available as of v1.16.0*
For filtered vector search, you are recommended to create a [payload index](/documentation/concepts/indexing/#payload-index) for the fields you want to filter by.
During the search, Qdrant will use a combined [filterable index](/documentation/concepts/indexing/#filtrable-index).
During the search, Qdrant will use a combined [filterable index](/documentation/concepts/indexing/#filterable-index).
However, when combining multiple strict payload filters, this mechanism might not provide sufficient accuracy.
In such cases, you can use the ACORN search algorithm.
@@ -1,6 +1,6 @@
---
title: Data Management
weight: 18
weight: 11
partition: build
---
@@ -16,5 +16,6 @@ partition: build
| [Confluent](/documentation/data-management/confluent/) | Fully-managed data streaming platform with a cloud-native Apache Kafka engine. |
| [DLT](/documentation/data-management/dlt/) | Python library to simplify data loading processes between several sources and destinations. |
| [Fluvio](/documentation/data-management/fluvio/) | Rust-based platform for high speed, real-time data processing. |
| [POMA](/documentation/data-management/poma/) | Python library for data ingestion, and structured chunking from various sources. |
| [Spark](/documentation/data-management/spark/) | A unified analytics engine for large-scale data processing. |
| [Unstructured](/documentation/data-management/unstructured/) | Python library with components for ingesting and pre-processing data from numerous sources. |
@@ -0,0 +1,194 @@
---
title: POMA
---
# POMA + Qdrant: Structure-Preserving Retrieval
| Time: 15 min | Level: Beginner/Intermediate | [Complete Notebook](https://colab.research.google.com/github/poma-ai/.github/blob/main/notebooks/qdrant/poma_meets_qdrant.ipynb) | [Notebook Source](https://github.com/poma-ai/.github/blob/main/notebooks/qdrant/poma_meets_qdrant.ipynb) |
| ------------ | ---------------------------- | -------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------- |
## Overview
- **POMA**, as *document chunking engine*, is built around simplicity for operators: process files into structure-aware chunksets and send them to Qdrant with minimal boilerplate and a patented chunking approach.
- **Qdrant** as your preferred *vector search engine*.
Together, they combine individual simplicity into one streamlined workflow.
This guide walks through the current [POMA AI](https://www.poma-ai.com/) for Qdrant SDK flow: process documents, upsert chunksets, retrieve structure-preserving cheatsheets, and understand where convenience defaults end and advanced knobs begin.
---
## Prerequisites
- Python 3.10+
- A POMA API key
- A Qdrant cluster URL + API key (for cloud)
---
## 1. Get API Keys
### POMA API key
1. Open https://app.poma-ai.com/
2. Register or sign in.
3. Open **API Keys** in the left navigation.
4. Copy your key and export it as `POMA_API_KEY`.
### Qdrant cluster API key
During Qdrant cluster creation, or when creating fine-grained API keys, see further details [here](https://qdrant.tech/documentation/cloud/authentication/).
### Credentials
Set your environment variables:
```bash
POMA_API_KEY="your_poma_api_key"
QDRANT_URL="https://<cluster>.<region>.qdrant.io"
QDRANT_API_KEY="your_qdrant_api_key"
```
---
## 2. Install Dependencies
```bash
pip install "poma[qdrant]"
```
---
## 3. Imports
```python
import os
from poma import Poma
from qdrant_client.http import models as qmodels
from poma.integrations.qdrant.qdrant_poma import PomaQdrant
```
Use your own local document path in the next step (for example `"./docs/your_file.pdf"`).
The linked Colab notebook includes downloadable sample files for quick testing.
---
## 4. Chunk a File with POMA
```python
client = Poma(os.environ["POMA_API_KEY"])
job = client.start_chunk_file("./docs/your_file.pdf")
chunk_data = client.get_chunk_result(
job["job_id"],
show_progress=True,
download_dir="./",
filename="your_file.poma",
)
```
### POMA-specific knobs in `get_chunk_result(...)`
- `show_progress`: prints job status updates while processing.
- `download_dir` + `filename`: save the returned archive as a `.poma` file while still returning parsed `chunk_data`.
- If both `download_dir` and `filename` are omitted, result is returned in-memory only (no `.poma` archive written).
`chunk_data` contains the structured output (`chunks` and `chunksets`) used by `upsert_poma_points(...)`.
If you already have a `.poma` archive, pass the path directly later:
```python
chunk_data = "your_file.poma"
```
---
## 5. Upsert Chunksets into Qdrant
```python
QDRANT_COLLECTION_NAME = "cloud_hybrid"
DENSE_MODEL = "sentence-transformers/all-minilm-l6-v2"
SPARSE_MODEL = "Qdrant/bm25"
DENSE_OPTIONS = {"dimensions": 384}
poma_qdrant = PomaQdrant(
url=os.environ["QDRANT_URL"],
api_key=os.environ["QDRANT_API_KEY"],
cloud_inference=True,
timeout=120,
collection_name=QDRANT_COLLECTION_NAME,
dense_model=DENSE_MODEL,
sparse_model=SPARSE_MODEL,
dense_size=384,
dense_options=DENSE_OPTIONS,
auto_create_collection=True,
)
poma_qdrant.upsert_poma_points(chunk_data)
```
`dense_size` is required when `auto_create_collection=True`.
---
## 6. Retrieve Structure-Preserving Cheatsheets
```python
cheatsheets = poma_qdrant.get_cheatsheets(
query="Whats the positional embeddings frequency?",
limit=10,
)
for i, cs in enumerate(cheatsheets, 1):
print(f"\n=== Cheatsheet {i} ===")
print(f"file_id: {cs['file_id']}")
print("content:")
print(cs["content"])
```
---
## 7. Advanced Query Control (Optional)
Use Qdrant prefetch + RRF fusion explicitly and still return POMA cheatsheets.
```python
query_text = "Whats the positional embeddings frequency?"
query_obj = qmodels.RrfQuery(rrf=qmodels.Rrf(k=60))
prefetch = [
qmodels.Prefetch(
query=qmodels.Document(
text=query_text,
model=DENSE_MODEL,
options=DENSE_OPTIONS,
),
using="dense",
limit=100,
),
qmodels.Prefetch(
query=qmodels.Document(
text=query_text,
model=SPARSE_MODEL,
),
using="sparse",
limit=100,
),
]
cheatsheets = poma_qdrant.get_cheatsheets(
query_obj=query_obj,
prefetch=prefetch,
collection_name=QDRANT_COLLECTION_NAME,
limit=10,
chunk_data=chunk_data,
)
```
---
## Further details:
- [POMA docs hub](https://www.poma-ai.com/docs/)
- [POMA on GitHub](https://github.com/poma-ai)
@@ -1,22 +0,0 @@
---
title: Using the Database
weight: 18
# If the index.md file is empty, the link to the section will be hidden from the sidebar
is_empty: false
aliases:
- how-to
- tutorials
partition: qdrant
---
# Database Tutorials
| |
|--------------------------------------------|
| [Bulk Upload Vectors to a Qdrant Collection](/documentation/database-tutorials/bulk-upload/) |
| [Large Scale Search](/documentation/database-tutorials/large-scale-search/) |
| [Backup and Restore Qdrant Collections Using Snapshots](/documentation/database-tutorials/create-snapshot/) |
| [Load and Search Hugging Face Datasets with Qdrant](/documentation/database-tutorials/huggingface-datasets/) |
| [Using Qdrant’s Async API for Efficient Python Applications](/documentation/database-tutorials/async-api/) |
| [Qdrant Migration Guide](/documentation/database-tutorials/migration/) |
| [Static Embeddings. Should you pay attention?](/documentation/database-tutorials/static-embeddings/) |
@@ -1,6 +1,6 @@
---
title: Practice Datasets
weight: 29
weight: 28
partition: build
---
@@ -1,11 +1,11 @@
---
#Delimiter files are used to separate the list of documentation pages into sections.
title: "Essentials"
title: "Integration Guides"
type: delimiter
weight: 1 # Change this weight to change order of sections
weight: 20 # Change this weight to change order of sections
partition: build
sitemapExclude: True
_build:
publishResources: false
render: never
partition: build
---
@@ -2,7 +2,7 @@
#Delimiter files are used to separate the list of documentation pages into sections.
title: "Integrations"
type: delimiter
weight: 17 # Change this weight to change order of sections
weight: 10 # Change this weight to change order of sections
sitemapExclude: True
_build:
publishResources: false
@@ -0,0 +1,43 @@
---
title: "Qdrant Edge"
weight: 13
partition: qdrant
---
<aside role="status">Qdrant Edge is in beta. The API and functionality may change in future releases.</aside>
# What Is Qdrant Edge?
Qdrant Edge is a lightweight, embedded vector search engine for AI on devices like robots, kiosks, home assistants, and mobile phones. Designed for real-time vector search on edge devices with limited computational resources, Qdrant Edge allows applications to use Qdrant's functionality even with intermittent or no internet connectivity.
Qdrant Edge does not run as a separate process. Instead, it runs inside an application process. Data is stored and queried locally on the device, ensuring low-latency access and enhanced privacy since data does not need to be transmitted to an external server. That said, Qdrant Edge provides APIs to [synchronize data with a Qdrant server](/documentation/edge/edge-data-synchronization-patterns/). This enables you to offload heavy computations such as indexing to more powerful server instances, back up and restore data, and centrally aggregate data from multiple edge devices.
## Qdrant Edge Shard
Qdrant Edge is built around the concept of an **Edge Shard**: a self-contained storage unit that can operate independently on edge devices. Each Edge Shard manages its own data, including vector and payload storage, and can perform local search and retrieval operations.
![Qdrant Edge Shards operate on edge devices](/documentation/edge/qdrant-edge.png)
To work with a Qdrant Edge Shard from a Python application, use the [Python Bindings for Qdrant Edge](https://pypi.org/project/qdrant-edge-py/) package. This package provides an `EdgeShard` class with methods to manage data, query it, and restore snapshots:
- `update`: Updates the data.
- `query`: Queries the data.
- `scroll`: Returns all points.
- `count`: Returns the number of points.
- `retrieve`: Retrieves points with the given IDs.
- `flush`: Flushes the data to ensure that all writes have been persisted to disk.
- `close`: Cleanly destroys the shard instance, ensuring the data is flushed. The data is persisted on disk and can be used to create another shard.
- `info`: Returns metadata information about the shard.
- `unpack_snapshot`: Unpacks a snapshot on disk.
- `snapshot_manifest`: Returns the current shard’s snapshot manifest.
- `update_from_snapshot`: Applies a snapshot to the shard.
## Using Qdrant Edge
To get started with Qdrant Edge, refer to the [Qdrant Edge Quickstart Guide](/documentation/edge/edge-quickstart/).
## More Examples
More examples and advanced usage of Qdrant Edge API can be found in the [GitHub repository](https://github.com/qdrant/qdrant/tree/master/lib/edge/python/examples).

Some files were not shown because too many files have changed in this diff Show More