Merge branch 'master' into blog-vsd-adding-hackathon

This commit is contained in:
kanungle
2025-09-17 11:43:13 -07:00
committed by GitHub
190 changed files with 3690 additions and 813 deletions
@@ -60,9 +60,9 @@ With Qdrant handling retrieval, Alhena no longer needed to customize infrastruct
## Hitting production-grade performance targets
Latency was a critical metric for Alhena. With FAISS, vector search on catalogs with 100,000+ items often took three seconds or more. That delayed the start of agent response streaming, making the AI feel sluggish and hurting the user experience. Pinecone helped on large indexes, but introduced latency on small ones, and couldn’t handle hybrid filtering needs.
Latency was a critical metric for Alhena. With FAISS, vector search on catalogs with 100,000+ items went far above their latency budget. That delayed the start of agent response streaming, making the AI feel sluggish and hurting the user experience. Pinecone helped on large indexes, but introduced latency on small ones, and couldn’t handle hybrid filtering needs.
Qdrant reduced retrieval latency on the same datasets to approximately 300 milliseconds. That enabled Alhena to meet its internal P95 SLA of 3.5 seconds from query to first token, even after accounting for hallucination detection, policy enforcement, and contextual rewriting.
Qdrant reduced retrieval latency by up to 90% on the same datasets. That enabled Alhena to meet its internal P95 SLA from query to first token, even after accounting for hallucination detection, policy enforcement, and contextual rewriting.
*“We track every millisecond. Qdrant helped us cut vector retrieval time by 90 percent at scale. That’s what made it possible to stay under our latency SLA.”*
— Kang-Chi Ho
@@ -0,0 +1,66 @@
---
draft: false
title: "How Fieldy AI Achieved Reliable AI Memory with Qdrant"
short_description: "Fieldy AI migrated to Qdrant to deliver fault-tolerant, real-time memory recall while reducing infrastructure costs by two-thirds."
description: "Discover how Fieldy AI built a fault-tolerant AI memory platform with Qdrant, achieving 100% reliable real-time recall, seamless hybrid search, and significant cost savings at scale."
preview_image: /blog/case-study-fieldy/social_preview_partnership-fieldy.jpg
social_preview_image: /blog/case-study-fieldy/social_preview_partnership-fieldy.jpg
date: 2025-09-04
author: "Daniel Azoulai"
featured: false
tags:
- Fieldy
- vector search
- hybrid search
- AI memory
- cost reduction
- reliability
- case study
---
## Fieldy AI’s migration to Qdrant: Building a fault-tolerant AI memory platform
![How Fieldy AI Achieved Reliable AI Memory with Qdrant](/blog/case-study-fieldy/case-study-fieldy-bento-dark.jpg)
### Capturing and retrieving a lifetime of conversations
<a href="https://fieldy.ai/" target="_blank">Fieldy</a> is a hands-free wearable AI note taker that continuously records, transcribes, and organizes real-world conversations into your personal, searchable memory. The system’s goal is simple in concept but demanding in execution: capture every relevant spoken interaction, transcribe it with high accuracy, and make it instantly retrievable. This requires a robust ingestion pipeline, a scalable [vector search](https://qdrant.tech/documentation/overview/) layer, and a retrieval process capable of handling growing volumes of multimodal data without introducing latency or errors.
From the start, the engineering team treated transcription reliability as the primary design constraint. If a conversation is not captured in the moment, it cannot be reconstructed later. This applies equally to Bluetooth transfer from the AI wearable pendant to the app, HTTPS uploads to the backend, speech-to-text transcription, embedding generation, and [vector database ingestion](https://qdrant.tech/documentation/database-tutorials/bulk-upload/). Every component had to meet this standard.
### How Fieldy AI differentiates in a crowded market
Fieldy’s multilingual, real-time transcription and instant recall make it a trusted voice recorder for professionals, healthcare providers, and anyone needing memory support, including those with ADHD. The engineering team is built to iterate quickly on quality and features in this fast-moving space. They maintain a direct feedback loop with users, enabling rapid prioritization of features that matter most in real usage. Finally, Fieldy’s multilingual transcription enables high-accuracy transcription in over 100 languages, using custom speech-to-text pipelines and continuous evaluation.
![Product screenshot](/blog/case-study-fieldy/fieldy-device-image.jpg)
*Fieldy's device*
### Reliability challenges with the initial vector database
Fieldy’s original architecture used Weaviate for vector storage and retrieval. Both the hosted and self-hosted deployments experienced persistent operational issues, most notably 5xx errors affecting roughly 10% of requests. These failures occurred both during ingestion and search, undermining the product’s promise of complete and accessible memory. Attempts to address the problem, including migrating to self-hosted infrastructure, provided only temporary relief before the errors returned months later.
For the engineering team, these failures had two serious implications. First, missing data meant a permanent loss of user trust. Second, time spent debugging database issues directly slowed feature development. The team’s requirements for a replacement were clear: the new [vector database](https://qdrant.tech/documentation/overview/) had to eliminate persistent query failures, ingest and search tens of millions of vectors with low latency, and run with minimal operational intervention.
### Selecting and deploying Qdrant
After evaluating alternatives, Fieldy selected [Qdrant](http://qdrant.tech) for its stability, straightforward configuration, and suitability for self-hosted deployment. They opted to run Qdrant in the same environment as their backend services, ensuring low-latency access and avoiding the cross-region connectivity issues that had contributed to failures in the previous architecture.
The migration process took less than a week. The team reused their existing vector schema to avoid redesigning their indexing logic during the transition. Embeddings, generated using Cohere’s multilingual v3 model, were batched locally before ingestion, replacing the integrated embedding calls previously handled within Weaviate. Following [Qdrant best practices](https://qdrant.tech/documentation/guides/optimize/), they disabled indexing during bulk import to maximize throughput, re-enabling it only after the migration was complete. The backend API calls were updated to use Qdrant’s gRPC interface, and [hybrid search](https://qdrant.tech/articles/hybrid-search/) with BM25 was combined with dense vector retrieval via [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/#hybrid-search) for relevance scoring.
### Architecture after migration
In Fieldy’s current architecture, the AI transcription device streams audio to the mobile app over Bluetooth. The app sends audio to the backend via HTTPS, where it is processed by a speech-to-text model to produce a transcript. This transcript and embeddings are stored in Qdrant.
When a user submits a query in the chat interface, the backend’s retrieval agent performs a [hybrid search](https://qdrant.tech/articles/hybrid-search/) in Qdrant. The HNSW index is used for dense vector similarity, BM25 handles term matching, and results are merged with Reciprocal Rank Fusion. Conversation metadata is fetched from Firestore for context assembly before the results are returned to the user.
### Results in production
Since the migration, Fieldy has delivered 100% reliable real-time AI recall and seamless memory search, ensuring the AI note taker’s performance matches its promise. Hosting Qdrant alongside the backend has reduced latency and virtually eliminated the network errors previously seen in cross-region deployments.
Fieldy also achieved a two-thirds reduction in infrastructure costs after moving to Qdrant, while scaling to handle tens of millions of embeddings without operational incidents. Post-migration interventions have been limited to planned storage increases.
### Next steps for retrieval quality
With reliability and cost efficiency achieved, Fieldy’s engineering focus is shifting toward retrieval quality. Planned improvements include adding [location filtering](https://qdrant.tech/documentation/concepts/filtering/#geo) and [datetime filtering](https://qdrant.tech/documentation/concepts/filtering/#datetime-range) within Qdrant to refine result sets, experimenting with late chunking strategies, and testing parallel hybrid searches to increase recall on complex multi-faceted queries. The team is also exploring embedding summaries alongside raw transcript segments to improve retrieval performance on high-level or thematic searches.
@@ -41,7 +41,7 @@ GoodData transitioned to a Retrieval-Augmented Generation (RAG) strategy, requir
GoodData leveraged Qdrant’s official Helm chart, deploying smoothly in Kubernetes and efficiently managing near-real-time embedding updates, crucial for multilingual semantic layers.
![Old Approach](/blog/case-study-gooddata/gooddata-diagram-1.png)
![Old Approach](/blog/case-study-gooddata/gooddata-diagram-1.1.png)
### Real-Time Performance and Scalability Gains
@@ -0,0 +1,66 @@
---
draft: false
title: "How OpenTable Reinvented Restaurant Discovery with Qdrant"
short_description: "OpenTable transformed restaurant search with Concierge, an AI-powered assistant built on Qdrant."
description: "Discover how OpenTable built Concierge, a generative AI dining assistant powered by Qdrant, achieving accurate retrieval, global scalability, and operational stability while redefining how diners discover restaurants."
preview_image: /blog/case-study-opentable/social-preview-opentable.jpg
social_preview_image: /blog/case-study-opentable/social-preview-opentable.jpg
date: 2025-09-02
author: "Daniel Azoulai"
featured: true
tags:
- OpenTable
- vector search
- generative AI
- restaurant discovery
- sparse embeddings
- filtering
- case study
---
## **Reinventing Restaurant Discovery: How OpenTable built Concierge, an AI Dining Assistant**
### Recognizing that AI would redefine restaurant discovery
When generative AI tools entered the mainstream, OpenTable knew diners would change how they find and choose restaurants. People were beginning to expect conversational, intelligent and context-aware assistants, rather than static search boxes.
Patrick Lombardo, Staff ML Engineer at OpenTable, recalls that the team wanted to move quickly. “We knew early on that generative AI was going to change user expectations. Concierge was an opportunity for us to transform the way that diners discover restaurants while building the tooling and infrastructure that will support future AI-powered experiences.”
That stepping stone is [Concierge](https://www.opentable.com/blog/concierge-ai-dining-assistant/), an AI-powered assistant designed to answer restaurant-related questions in natural language using OpenTable’s data.
![Concierge screenshot](/blog/case-study-opentable/opentable-concierge-screenshot.png)
### Setting clear priorities for accuracy, domain focus, and speed
For Concierge to succeed, the assistant needed to respond to the vast majority of user questions and every answer had to reflect reality. Incorrect menu items or outdated offerings could erode user and restaurant trust.
“The primary goal was answerability. We wanted to make sure the model could answer most questions. The second most important was accuracy, so that when the model gave an answer it was correct.” Puyuan Liu, Machine Learning Scientist, OpenTable
Beyond the application logic, the team needed a vector database that could handle sparse embeddings for keyword expansions and fine-grained filtering. Queries often narrowed results to a single restaurant out of more than 60,000, which placed heavy demands on filtering performance.
### Qdrant chosen due to sparse embeddings, filtering, and deployment.
Qdrant emerged as the preferred option for several reasons that aligned directly with OpenTable’s priorities.
First, its handling of sparse embeddings was a key differentiator. Concierge’s retrieval often required filtering down to a single restaurant, making the collection effectively sparse. Many vector databases see HNSW graph quality degrade under such conditions, but Qdrant’s optimizations avoided that performance drop.
Second, Qdrant delivered reliable high-precision filtering. In production, each query might target reviews, metadata, and other structured restaurant data all at once. Qdrant handled this with predictable performance, which was essential for hitting their latency budget.
Third, Qdrant Cloud provided a deployment path that was simpler than self-hosting.
Patrick Lombardo summed it up: "Creating a Qdrant Cloud cluster was one of the easiest parts of the project. It just worked."
The production launch was global from the start, allowing Concierge to answer questions about restaurants in many regions without separate deployments.
### Achieving stability and setting the stage for future innovation
Concierge met its latency target and maintained high answerability without extensive post-launch tuning. Operationally, Qdrant became one of the most stable components in the stack. Ant White, Principal Software Engineer at OpenTable, explained, “Since running it in production, it is a frictionless part of the stack. ”
### Key takeaways from the Concierge rollout
Concierge is continuing to pave the way for OpenTable’s AI transformation. While it is a positive new user experience itself, it also gave OpenTable a safe environment to refine retrieval infrastructure before integrating it into core search. Sparse embeddings proved highly effective when combined with heavy filtering, keeping retrieval precise and efficient.
Operational stability turned out to be a significant advantage. With Qdrant handling retrieval without incident, the team was free to focus on improving model performance and user experience rather than troubleshooting the database layer.
With Concierge as a foundation, OpenTable is now positioned to continue to deliver richer, faster, and more intelligent dining experiences, from conversational search to visual dish discovery.
@@ -0,0 +1,60 @@
---
draft: false
title: "How Tavus used Qdrant Edge to create conversational AI "
short_description: "Tavus built its Conversational Video Interface with Qdrant edge retrieval to achieve subsecond, human-grade conversations."
description: "Tavus used Qdrant edge retrieval to cut latency and deliver natural, human-grade conversational AI."
preview_image: /blog/case-study-tavus/social_preview_partnership-tavus.jpg
social_preview_image: /blog/case-study-tavus/social_preview_partnership-tavus.jpg
date: 2025-09-12
author: "Daniel Azoulai"
featured: true
tags:
- Tavus
- vector search
- conversational AI
- retrieval augmented generation
- latency optimization
- case study
- Edge
---
![Tavus Overview](/blog/case-study-tavus/tavus-bento-box-dark.jpg)
## How Tavus delivered human-grade conversational AI with edge retrieval on Qdrant
Tavus is a human–computer research lab building CVI, the <a href="https://www.tavus.io/" target="_blank">Conversational Video Interface</a>. CVI presents a face-to-face AI that reads tone, gesture, and on-screen context in real time, allowing for humans to interface with powerful, functional AI like never before. The team’s north star was simple to say and hard to ship: conversations should feel natural. That meant tracking conversational dynamics like utterance-to-utterance timing, back-channeling, and turn-taking while grounding replies in a customer’s private knowledge.
Early iterations of CVI focused on live conversation quality, but not retrieval. Customers who needed document grounding or recall brought their own RAG layer, which added latency and inconsistency. Tavus wanted to internalize RAG so they could guarantee performance, simplify onboarding, and keep the experience cohesive.
*“I read your docs, had a clear idea of what to do, implemented it, and it just worked. The simplicity and performance were there from day one.”*
Mert Gerdan, ML Engineer, Tavus
## Why network hops threatened subsecond conversational flow
Human conversation tolerates very little lag. Literature on conversational systems shows that in highly engaging exchanges, the optimal time from one speaker finishing to the other starting is about 200ms. For CVI, even 500 to 600ms best case end to end felt tight once you include understanding, planning, text to speech, facial rendering, and streaming.
Adding network hops for retrieval threatened to push utterance-to-utterance into an unacceptable range. On top of that, customers needed multimodal grounding that spans video, audio, and screen share, as well as per-conversation isolation for security and correctness. The retrieval layer had to be fast, local, and simple.
## How per-conversation edge vector stores removed latency
Tavus implemented a self-hosted central Qdrant and spun up per-conversation edge collections that were colocated with the conversational worker. Each conversation generated embeddings on a local GPU and queried its local Qdrant store, which removed the network latency during retrieval.
### Design choices that balanced speed and quality
The first choice was to keep both embedding and approximate nearest neighbor lookup local to the node where the conversation runs. Because the data never left the machine, the system avoided serialization and transit delays that would otherwise dominate the latency budget. Most of the usage is in a few collections, and they filter based on conversation id. This made reasoning about context scope easier and created a clean boundary for privacy, auditing, and lifecycle management.
With the network hop removed, the team chose simplicity over tuning. They did not need quantization or aggressive compression to chase a few extra milliseconds, so they focused on retrieval quality and multimodal accuracy instead. That simplicity enabled a second architectural move: retrieval on every utterance. Each turn could fire an embedding and vector search without imposing a perceptible pause. Finally, Tavus layered speculative execution on top, predicting likely continuations so the agent could prepare responses while the user was still speaking. Together, these choices produced a pipeline that felt natural without sacrificing correctness.
### Business impact from faster retrieval and smoother launches
By eliminating the network hop, Tavus reduced retrieval to roughly 20 to 25ms at the edge. End-to-end utterance-to-utterance timing now landed near 500 to 600ms in the best case, leaving enough headroom to add artificial delays for deeper topics where a slower cadence feels more human. Because retrieval was cheap in terms of perceived latency, the team could ground every turn and keep answers accurate even as conversations grew complex.
The operational picture improved as well. Within the first three weeks, Tavus indexed about 3 to 3.5 million points, with each point representing around 1,500 characters. The launch was uneventful in a good way. Support queues stayed quiet, and customers were able to bring private knowledge into CVI without standing up their own RAG stacks. Developer velocity benefited from a smaller, clearer deployment model that was easier to reason about and extend.
*“We wanted companies to experience CVI’s quality without building a RAG system themselves. With Qdrant at the edge, retrieval became effectively invisible to the user.”*
Mert Gerdan, ML Engineer, Tavus
## What the team learned about architecture and speed
Tavus validated that architecture beats micro-optimizations. Removing the network hop and colocating compute with data produced a larger latency win than squeezing a few milliseconds with quantization. With per-conversation edge stores in place, the team could optimize for quality and safety, such as richer multimodal features and better turn-taking, without sacrificing speed.
@@ -0,0 +1,258 @@
---
title: "Untangling Relevance Score Boosting and Decay Functions"
draft: false
slug: decay-functions # Change this slug to your page slug if needed
short_description: Why's and how's of decay functions in Qdrant's relevance score boosting. # Change this
description: Understanding decay functions for relevance score boosting. # Change this
preview_image: /blog/decay-functions/preview/preview.jpg # Change this
social_preview_image: /blog/decay-functions/preview/social_preview.jpg # Optional image used for link previews
title_preview_image: /blog/decay-functions/preview/title.jpg # Optional image used for blog post title
date: 2025-09-01T14:55:45+02:00
author: Evgeniya Sukhodolskaya
featured: true
tags:
- features
- tech
- blog
---
A problem we've noticed while monitoring the [Qdrant Discord Community](https://discord.gg/d4MPnX3s) is that due to the extensive list of expressions that the [score boosting](https://qdrant.tech/documentation/concepts/hybrid-queries/#score-boosting) functionality provides, there's room for confusion on how it's supposed to be applied. And that might block you from moving the business logic behind relevance scoring into the Qdrant search engine. We don't want that!
In this blog, we'd like to de-spooky-fy the **decay functions** part of the score boosting, or, more precisely: `LinDecayExpression`, `ExpDecayExpression`, and `GaussDecayExpression` -- frequent guests on the Discord *#ask-for-help* channel.
## Purpose of Decay Functions
Decay functions help turn numeric properties of your dataset items (like sizes or ratings) into values between 1.0 (most relevant) and 0.0 (not relevant). This makes it possible for those properties to meaningfully influence the final relevance score.
Decay functions are useful when a change in some numeric property of an item should *smoothly* and *proportionally* affect its relevance score.
Think of it like this:
- News articles become less relevant over time, so relevance decays with days passed from the publication date.
- A further restaurant is less relevant for food ordering, so relevance decays as the distance to the user increases.
- A better reputation makes a movie more relevant, so relevance decays as the number of positive IMDb reviews decreases.
### Three Options
Qdrant has added three decay functions to the score boosting functionality, each one capturing a different way in which relevance can decay.
{{< figure src="/blog/decay-functions/decay_1.png" alt="Three decay functions used in the score boosting." caption="**Image 1.** Three decay functions used in the score boosting.<br>An interactive version of this graph is available [here](https://www.desmos.com/calculator/idv5hknwb1)." width="100%" >}}
**Linear**
Relevance changes at a constant rate with the variable. Each change in the value has the same impact.
→ For example, the discount percentage: the more, the merrier!
**Gaussian**
Relevance decays smoothly and gradually. Small deviations from the ideal are forgivable, but the further you go, the less relevant it becomes.
→ Perfect for things like product price; small differences are usually fine for users, but when the price gap gets big, interest drops off fast.
**Exponential**
Relevance drops sharply with even small changes. Deviation from the ideal is punished quickly.
→ Delivery time is a great candidate here. If the item takes too long to arrive, users instantly lose interest.
**All three decay functions in Qdrant are *symmetrical*. They assign 1.0 relevance to a certain value of a variable and decay toward 0.0 as the variable deviates from this target.**
This symmetry is useful when there's a clear "ideal" value, and anything more or less than that is equally off. For example:
- A user has a target price in mind. Anything more expensive is less relevant (obviously), but anything cheaper might also raise quality concerns. (Yes, not always, but many think so.)
- You're searching for a 30-minute exercise video. Both 25 and 35 minutes are okay-ish, but 5 minutes or 1 hour are clearly not what you had in mind.
## Decay Function Parameters
To use decay functions for score boosting, you need to figure out what values to provide for their parameters.
At first glance, it might seem like there are many of them: `x`, `target`, `midpoint`, `scale`...
Let's demystify these. And let's start with a quick win: two of them, `x` and `target`, we've already used many times in the examples above.
### `x` parameter
This is just the variable you want to transform with a decay function: score, time, distance, age, number of reviews, price, etc.
You can think of it as the *input value*. It may come from a payload field of an item or the embedding similarity score. Basically, it’s the x-axis of the decay function (and y is the output, the relevance score, decaying from 1.0 to 0.0).
{{< figure src="/blog/decay-functions/decay_2.png" alt="x (x-axis) and target (point on x-axis) of a decay function." caption="**Image 2.** x (x-axis) and target (point on x-axis) of a decay function.<br>An interactive version of this graph is available [here](https://www.desmos.com/calculator/idv5hknwb1)." width="100%" >}}
### `target` parameter
This is the value your `x` variable needs to match for the item to be considered 100% relevant, to get a 1.0 relevance score from the decay function.
By default, `target` is 0.0, which makes sense in many scenarios. For example:
- 0.0 meters for delivery distance: the closer, the better.
- 0.0 seconds since publishing: the fresher, the better.
But of course, `target` can be anything else, depending on your use case: desired and most relevant price, age, size, score, etc.
*So, in short, `x` is the current value, `target` is the relevance best-case value.*
### `scale & midpoint` parameters
Now we're left with `midpoint` and `scale`. What's their purpose? To control how the decay function of your choice will, well... function. Its shape needs to match your definition of relevance and the nature of the variable you're transforming.
{{< figure src="/blog/decay-functions/decay_3.png" alt="scale (segment on x-axis) and midpoint (point on y-axis), defining the shape of decay functions." caption="**Image 3.** scale (segment on x-axis) and midpoint (point on y-axis), defining the shape of decay functions.<br>An interactive version of this graph is available [here](https://www.desmos.com/calculator/idv5hknwb1)." width="100%" >}}
*Together, `scale` and `midpoint` define the slope of the function, how quickly or smoothly relevance decays. It reads: "To what `scale` should `x` change to reach a `midpoint` value of relevance."*
A decay function's shape is defined by two key points:
- (`target`, 1.0) --- the ideal use case
- (`target ± scale`, `midpoint`) --- how relevance drops after the `x` variable changes by `scale` from the ideal `target` value.
The choice of `scale` and `midpoint` defines a certain behavior for each type of decay function.
- For **Gaussian decay**, relevance drops slowly and smoothly from 1.0 toward `midpoint`, as `x` changes by `scale`. After that, the decay accelerates.
- For **Exponential decay**, it's the opposite: fast decay at first till `midpoint`, then slower.
- For **Linear decay**, `midpoint` and `scale` define the point at which the relevance score hits 0.0, as it's the only decay function that actually reaches zero.
**Note #1.** `midpoint` defaults to 0.5, but can be anything in the (0.0, 1.0) range. For linear decay, it's also valid to set `midpoint` to 0.0.
**Note #2.** `scale` defaults to 1.0, but it can be anything that reflects the relationship between your variable `x` and how you define relevance. Only *you* know what makes sense here.
**Note #3.** We expect `scale` to be a **positive** value; it just makes calculations for us simpler.
**Note #4.** Exponential and Gaussian decay functions never reach 0.0. The relevance score approaches zero but stays positive. Only Linear decay can reach exactly 0.0, and it's the only one where setting `midpoint` to 0.0 is valid.
### How to Pick Parameters: Examples
**Example #1**
**Use case:** A user is searching for educational videos in German about techno club culture to practice language comprehension. They've chosen 5 minutes as the ideal video length.
**Decay:** Gaussian
**`x`**: Video length in minutes, stored in the video's payload
**`target`**: 5 (*minutes*)
**`scale`**: 4 (*minutes*)
**`midpoint`**: 0.5
**Explanation:**
We assume that, out of all videos relevant by content, the user will tolerate deviations in video length by up to ±4 minutes, so 1-minute to 9-minute videos. Relevance should therefore decay *smoothly and slowly* from 1.0 to 0.5 in a Gaussian fashion.
Anything longer than 9 minutes or shorter than 1 minute quickly becomes less relevant, even if the content still matches.
**Example #2**
**Use case:** A promo code aggregator app boosts freshness to always show the latest promo codes for products and events.
**Decay:** Exponential
**`x`**: datetime of promo code upload, stored in the payload
**`target`**: Current datetime (moment of search)
**`scale`**: 604800 (1 week in seconds)
**`midpoint`**: 0.1
**Explanation:**
Out of all promo codes for different products/events, users will strongly prefer ones uploaded *just now*, as they’re most likely to work. But that relevance drops quickly over time: within a week, it reaches a midpoint of 0.1. After that, if a promo code is still active, it’s a gamble anyway: might work, might not. So old-but-not-expired codes are roughly equally irrelevant.
**Note #5.** For Qdrant [datetime](https://qdrant.tech/documentation/concepts/payload/#datetime) payloads, `scale` should always be provided in seconds!
### I Don't Know All the Parameters in Advance
As you can see, using decay functions in Qdrant's score boosting means you'll have to know the parameters in advance.
What we've seen in our Discord Community quite a few times is that people try to apply decay functions to normalize similarity scores from [prefetches](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries), usually as a way to fuse results from different types of similarity searches.
The common question is:
> How can I use a decay function for score normalization if I don't know the scale in advance?
The answer is: **you can't dynamically set the `midpoint` and `scale` parameters, hence you can't dynamically normalize scores, and you probably shouldn't even try**:
Say you're prefetching a subset of items scored by a late interaction model like ColBERT, where a higher score means higher similarity. You get scores like 36, 22, and 1. Separately, you also have cosine similarity scores from some dense vector search, and you'd like to fuse both sets of results.
If you normalize ColBERT scores dynamically, based on just this subset, 36 will become 1.0, and everything else scales accordingly.
But here's the problem: That 36 might not be a "high" score at all. Maybe your dataset just didn't contain any good matches. The normalization step will strip away that context, and when you fuse the scores, you'll create a false sense of high relevance.
*If you're planning to use a decay function for score normalization, you need to know the expected parameter values beforehand. If you don't know the range of your input variable (`x`), you won't be able to use a decay function reliably.*
## Code Snippets
Now let's see how using decay functions looks in Qdrant.
We'll provide HTTP request examples, but you can use decay functions [analogously in the Python, TypeScript, Rust, Java, C#, and Go clients](https://qdrant.tech/documentation/concepts/hybrid-queries/#time-based-score-boosting).
**Note #6.**
Payload variables used within the formula benefit from having [payload indexes](https://qdrant.tech/documentation/concepts/indexing/#payload-index). So, we require you to set up a payload index for any variable used in a formula.
Let's take our "educational videos in the German language" example and see how it takes shape in Qdrant:
```http
POST collections/video/points/query
{
"prefetch": {
"query": <video_description_embedding>,
"limit": 10 // limit of prefetched results
},
"query": {
"formula": {
"sum": [
"$score", // so the final score = score + gauss_decay(duration)
{
"gauss_decay": {
"target": 5,
"scale": 4,
"midpoint": 0.5,
"x": "duration" // payload key
}
}
]
}
}
}
```
And here’s the “fresh promo codes” example, so you’ve got a better grip on how to use Qdrant’s score boosting:
```http
POST collections/promocodes/points/query
{
"prefetch": {
"query": <promocode_description_embedding>,
"limit": 10 // limit of prefetched results
},
"query": {
"formula": {
"sum": [
"$score", // so the final score = score + exp_decay(search_time - upload_time)
{
"exp_decay": {
"x": {
"datetime_key": "upload_time" // payload key
},
"target": {
"datetime": "2025-08-04T00:00:00Z" // time of the search
},
"scale": 604800, // 1 week in seconds
"midpoint": 0.1
}
}
]
}
}
}
```
**Note #7.**
`datetime_key` and `datetime` are used to distinguish between payload keys that hold a datetime-type string and directly provided datetime strings.
## To Sum Up
So, we've covered quite a bit, including:
- What decay functions are and why they matter;
- How `target`, `x`, `scale`, and `midpoint` shape decay behavior;
- And how to use decay functions in Qdrant's score boosting.
We truly hope this write-up helped untangle things a bit. Now the only thing left for you is to get your hands dirty and experiment!
Use the snippets in the article as a starting point and experiment with the relevance score boosting in [Qdrant Cloud](https://qdrant.tech/). We offer a free-forever 1GB cluster: enough to test, tweak, and see how the decay functions behave on your data.
And if you feel like diving deeper into decay functions or score boosting in general, check out our [documentation](https://qdrant.tech/documentation/concepts/hybrid-queries/?q=Query+Points+API#score-boosting), which includes a decay-on-distance example and plenty more to learn from.
### Tell Us What You're Building
We'd love to hear what you're experimenting with! What are you considering "relevant"? Which features of the last releases are you enjoying? What's missing?
If you're still unsure about anything, feel free to ask us [on Discord](https://discord.gg/d4MPnX3s) or [connect with me on LinkedIn](https://www.linkedin.com/in/evgeniya-sukhodolskaya/). We're always happy to explain more, and we'd love to know what to write about next!
@@ -0,0 +1,404 @@
---
draft: false
title: Balancing Relevance and Diversity with MMR Search
slug: mmr-diversity-aware-reranking
short_description: Discover how Qdrant's Maximum Marginal Relevance (MMR) balances relevance with diversity in fashion search to avoid echo chambers of similar results.
description: Learn how to implement Maximum Marginal Relevance (MMR) with Qdrant to create diverse search results in fashion discovery. This guide shows how to balance relevance and diversity using the DeepFashion dataset and CLIP embeddings.
preview_image: /blog/mmr-diversity-aware-reranking/preview/preview.webp
social_preview_image: /blog/mmr-diversity-aware-reranking/preview/social_preview.jpg
title_preview_image: /blog/mmr-diversity-aware-reranking/preview/title.webp
date: 2025-09-04
author: Thierry Damiba
featured: true
categories:
- Tutorial
- Vector Search
tags:
- MMR
- Fashion Search
- Diversity
- CLIP
- Vector Search
---
Variety is the spice of life! Yet often, with search engines, users find that the results are too similar to get value. You search for a black jacket on your favorite shopping site, and you get 5 black full zip bomber jackets. Search for a black dress and you get 5 strapless dresses. Traditional vector search focuses on returning the most relevant items, which creates an echo chamber of similar results.
![Similar black dresses](/blog/mmr-diversity-aware-reranking/mmr-food-diversity.webp)
*Problem: A search for "black dress" returns only strapless dresses*
Qdrant's native Maximum Marginal Relevance (MMR) fixes this by balancing similar results with diverse results. Instead of showing variations of the same item, MMR makes sure each result adds something novel to your search.
While MMR applies to any domain and modality, today we'll explore it through fashion search using the DeepFashion dataset. This visual approach makes the diversity benefits immediately obvious, but the same principles work whether you're searching documents, building recommendation engines, or retrieving context for AI systems.
For a different perspective on MMR with text-based movie recommendations, check out [Tarun Jain's implementation guide](https://python.plainenglish.io/understanding-maximal-marginal-relevance-mmr-a0a7a8df0a1a).
## What is Maximum Marginal Relevance?
![Standard Search returns all "bomber jacket"](/blog/mmr-diversity-aware-reranking/standard-search-results.webp)
*Note the diversity of food dishes on the right, with MMR*
MMR solves the redundancy problem by reranking search results based on two criteria:
1. **Relevance to your query** (how well does this match what you're looking for?)
2. **Diversity from already selected items** (how different is this from what you already have?)
The algorithm picks the most relevant item first, then for each subsequent item, it balances relevance against similarity to already-selected results. A lambda parameter controls this balance:
- **λ = 1.0:** Pure relevance (regular vector search)
- **λ = 0.5:** Balanced approach
- **λ = 0.0:** Pure diversity
*MMR is implemented in Qdrant as a parameter of a nearest neighbor query. Code examples can be found below.*
## Why Vector Search for Fashion Discovery?
![Fashion search demonstration](/blog/mmr-diversity-aware-reranking/fashion-search-demo.webp)
*A search for "black dress" returns only strapless dresses, showing the need for diversity in search results*
Fashion search is perfect for demonstrating MMR because visual similarity doesn't always match shopping intent. When someone searches for "black jacket," they might want to explore:
- 🏃 Bomber jacket (sporty)
- 👔 Blazer (professional)
- 🧥 Leather jacket (edgy)
- 👖 Denim jacket (casual)
Your typical search might show four black jackets that look nearly identical. MMR gives you one or two bombers plus diverse alternatives, helping users discover a wider variety of styles.
Let's take a look at a practical example.
## Setting Up the Environment
We'll need several libraries for this fashion discovery project:
```bash
pip install qdrant-client # Vector Search Engine
pip install fastembed # Fast, lightweight embedding generation with CLIP
pip install datasets # For DeepFashion data access
```
## Exploring the DeepFashion Dataset
The [DeepFashion dataset](https://huggingface.co/datasets/SaffalPoosh/deepFashion-with-masks) contains over 40,000 clothing images across different categories and styles. It includes rich metadata, such as category, color, and style attributes, that make it perfect for testing diversity algorithms.
The dataset includes realistic fashion photography: items worn by models, flat lay product shots, and detailed images. The variety of items with similar names makes this dataset perfect for testing whether MMR can distinguish between visually similar items that serve different fashion purposes.
---
## Step 1: Loading and Processing Fashion Data
```python
from datasets import load_dataset
from fastembed import ImageEmbedding, TextEmbedding
from qdrant_client import QdrantClient, models
import uuid
def load_fashion_data(sample_size=10):
"""Load fashion items from DeepFashion-with-masks dataset"""
dataset = load_dataset("SaffalPoosh/deepFashion-with-masks", split="train")
sample_dataset = dataset.shuffle(seed=42).select(range(sample_size))
fashion_items = []
for i, item in enumerate(sample_dataset):
# Extract all available metadata
metadata = {}
for key, value in item.items():
if key not in ["images", "mask", "mask_overlay"] and value is not None:
metadata[key] = value
fashion_items.append({
"image": item["images"], # PIL Image object
**metadata
})
return fashion_items
fashion_items = load_fashion_data(sample_size=100)
```
---
## Step 2: Creating Fashion Embeddings
We'll use CLIP to create embeddings that understand both visual similarity and semantic meaning in fashion:
```python
# Create image embeddings for all fashion items
image_model = ImageEmbedding(model_name="Qdrant/clip-ViT-B-32-vision")
embeddings = list(image_model.embed([item["image"] for item in fashion_items]))
```
---
## Step 3: Setting Up Qdrant for Visual Fashion Search
Create the client and set up your collection
```python
# Initialize Qdrant client, get your credentials at https://qdrant.tech
client = QdrantClient(
host="your-qdrant url",
api_key="<your-qdrant-api-key>")
collection_name = "fashion_discovery"
# Create collection optimized for CLIP embeddings
client.recreate_collection(
collection_name=collection_name,
vectors_config=models.VectorParams(
size=512, # CLIP embedding dimension
distance=models.Distance.COSINE
)
)
# Upload fashion items with embeddings
points = []
for i, (item, embedding) in enumerate(zip(fashion_items, embeddings)):
# Remove PIL image from payload
payload = {k: v for k, v in item.items() if k != "image"}
payload["item_id"] = i
point = models.PointStruct(
id=str(uuid.uuid4()),
vector=embedding.tolist(),
payload=payload
)
points.append(point)
client.upsert(collection_name=collection_name, points=points)
```
---
## Step 4: Implementing Fashion Search Functions
Now let's create search functions to compare standard similarity versus MMR diversity. We'll use each function with the query: "black jacket"
### Standard Search
Standard search gives you all similar styles: bomber jacket.
```python
# Standard fashion search — often returns similar looking items
def fashion_search_standard(query_text, limit=5):
text_model = TextEmbedding(model_name="Qdrant/clip-ViT-B-32-text")
query_embedding = list(text_model.embed([query_text]))[0]
results = client.search(
collection_name=collection_name,
query_vector=query_embedding.tolist(),
limit=limit,
with_payload=True
)
return results
query_text = "black jacket"
# Standard search
standard_results = fashion_search_standard(query_text)
print("\nSTANDARD FASHION SEARCH: 'black jacket'")
for i, point in enumerate(standard_results, 1):
payload = point.payload
print(f"{i}️⃣ Score: {point.score:.4f} | {payload.get('item_description', 'Fashion Item')}")
print(f" Style: {payload.get('style', 'N/A')} | Color: {payload.get('color', 'N/A')}")
print()
```
![Standard Search returns all "bomber jacket"](/blog/mmr-diversity-aware-reranking/black-dress-problem.webp)
*Standard Search returns all "bomber jacket" - notice the lack of diversity in results*
```
STANDARD SEARCH RESULTS: "black jacket"
1️⃣ Black Bomber Jacket with Orange lining
Style: casual | Color: black | Type: bomber jacket
Score: 0.9234
`
2️⃣ Slim-fit Bomber Jacket with Zipper pocket
Style: modern | Color: black | Type: bomber jacket
Score: 0.9156
3️⃣ Lightweight Bomber Jacket with ribbed cuffs
Style: minimalist | Color: black | Type: bomber jacket
Score: 0.9089
4️⃣ Quilted Bomber Jacket with padded lining
Style: classic | Color: black | Type: bomber jacket
Score: 0.9012
5️⃣ Streamlined bomber jacket with sleek zipper
Style: contemporary | Color: black | Type: bomber jacket
Score: 0.8967
```
### MMR Search
MMR gives you diverse styles: a hoodie, a windbreaker, and a blazer.
```python
# MMR fashion search — balances relevance with style diversity
def fashion_search_mmr(query_text, limit=5, diversity=0.5):
text_model = TextEmbedding(model_name="Qdrant/clip-ViT-B-32-text")
query_embedding = list(text_model.embed([query_text]))[0]
results = client.query_points(
collection_name=collection_name,
query=models.NearestQuery(
nearest=query_embedding.tolist(),
mmr=models.Mmr(
diversity=diversity, # 0.0 - relevance; 1.0 - diversity
candidates_limit=100 # num of candidates to preselect
)
),
limit=limit,
with_payload=True
)
return results
query_text = "black jacket"
# MMR search with diversity=0.5
mmr_results = fashion_search_mmr(query_text, diversity=0.5)
print("\nMMR FASHION SEARCH: 'black jacket' (diversity=0.5)")
for i, point in enumerate(mmr_results.points, 1):
payload = point.payload
print(f"{i}️⃣ Score: {point.score:.4f} | {payload.get('item_description', 'Fashion Item')}")
print(f" Style: {payload.get('style', 'N/A')} | Color: {payload.get('color', 'N/A')}")
print()
```
![MMR Search returns different styles](/blog/mmr-diversity-aware-reranking/mmr-search-results.webp)
*MMR Search returns different styles - bomber jacket, hoodie, windbreaker, and blazer for better diversity*
```
MMR SEARCH RESULTS: "black jacket" (diversity=0.5)
1️⃣ Black Bomber Jacket with Orange lining
Style: casual | Color: black | Type: bomber jacket
Score: 0.9234 | Selected: Most relevant
2️⃣ Black zip-up hoodie with drawstring
Style: relaxed | Color: black | Type: hoodie
Score: 0.8456 | Selected: Different style (hoodie vs bomber)
3️⃣ Lightweight Bomber jacket with ribbed cuffs
Style: minimalist | Color: black | Type: bomber jacket
Score: 0.9089 | Selected: Different aesthetic (minimalist)
4️⃣ Black Windbreaker with high collar
Style: sporty | Color: black | Type: windbreaker
Score: 0.8234 | Selected: Different function (windbreaker)
5️⃣ Black tailored blazer with notch lapel
Style: classic | Color: black | Type: coat
Score: 0.7891 | Selected: Different formality (blazer)
```
Keep in mind that MMR selects results one by one. The scores you see in Qdrant are representative of the similarity between each item and the original query. The final ranking won't be sorted by score; it will be sorted by the order in which the MMR algorithm selects each item.
---
## Step 5: Advanced Fashion Filtering with MMR
Combine MMR with Qdrant's metadata-filtering for targeted fashion discovery. In this example, we will filter on two categories. We will search for casual outerwear, but filter on casual and sporty. This allows us to diversify the results while making sure that only outdoor and sporty examples show up.
```python
def filtered_fashion_search(query_text, metadata_filter=None, limit=5, diversity=0.4):
text_model = TextEmbedding(model_name="Qdrant/clip-ViT-B-32-text")
query_embedding = list(text_model.embed([query_text]))[0]
# Build filter for metadata fields
filter_conditions = []
if metadata_filter:
for key, value in metadata_filter.items():
filter_conditions.append(
models.FieldCondition(
key=key,
match=models.MatchValue(value=value)
)
)
query_filter = models.Filter(must=filter_conditions) if filter_conditions else None
results = client.query_points(
collection_name=collection_name,
query=models.NearestQuery(
nearest=query_embedding.tolist(),
mmr=models.Mmr(
diversity=diversity, # 0.0 - relevance; 1.0 - diversity
candidates_limit=100 # num of candidates to preselect
)
),
query_filter=query_filter,
limit=limit,
with_payload=True
)
return results
```
### Practical Example: Filter for women's blouses
```python
# Example: Filter for women's blouses
results = filtered_fashion_search(
"professional attire",
metadata_filter={"gender": "WOMEN", "cloth_type": "Blouses_Shirts"},
diversity=0.5)
print("FILTERED SEARCH (Women's Professional Attire):")
for point in results.points:
payload = point.payload
print(f"Score: {point.score:.4f}")
print(f"Gender: {payload.get('gender', 'N/A')}")
print(f"Cloth Type: {payload.get('cloth_type', 'N/A')}")
print()
```
**Sample Output:**
```
FILTERED SEARCH (Women's Professional Attire):
1️⃣ White Cotton Button-Down Shirt
Score: 0.8934 | Gender: WOMEN | Type: Blouses_Shirts
Style: classic professional
2️⃣ Navy Silk Blouse with Bow Tie
Score: 0.8567 | Gender: WOMEN | Type: Blouses_Shirts
Style: elegant business
3️⃣ Striped Long-Sleeve Shirt
Score: 0.8234 | Gender: WOMEN | Type: Blouses_Shirts
Style: modern casual-professional
```
---
## What We Achieved
1. **Visual embeddings with CLIP** turned fashion images into searchable vectors
2. **MMR reranking** eliminated duplicate-looking recommendations
3. **Diversity control** let us tune exploration vs relevance
4. **Style filtering** combined semantic search with structured metadata for targeted discovery
The result is a fashion search that actually helps users discover new styles instead of showing variations of the same item.
---
## Where to Go Next
This pipeline opens up several possibilities:
- **Visual similarity with style diversity:** Upload a photo and find similar items in different styles
- **Outfit completion:** Given one item, find diverse pieces that create complete outfits
- **Seasonal recommendations:** Balance color preferences with seasonal appropriateness
- **Personal styling AI:** Learn user preferences and recommend diverse items within their taste profile
---
## Try It Yourself
MMR transforms fashion search from "here are similar items" to "here are diverse options you might love." The native Qdrant implementation handles everything under the hood while giving you fine control over the relevance-diversity balance.
Start with diversity=0.5 and then adjust based on whether you want more exploration (lower diversity) or precision (higher diversity).
Try it with your own fashion dataset on [Qdrant Cloud's free tier](https://cloud.qdrant.io/).
If you enjoyed this article, give me a follow on [LinkedIn](https://www.linkedin.com/in/thierrydamiba/) or [Twitter](https://twitter.com/thierrydamiba) to stay up to date with more guides involving vector search and retrieval optimization.
@@ -166,13 +166,13 @@ Are you interested in becoming a Qdrant Star?
We're on the lookout for individuals who are passionate about vector search technology and looking to make an impact in the AI community.
If you have a strong understanding of vector search technologies, enjoy creating content, speaking at conferences, and actively engage with our community. If this sounds like you, don't hesitate to apply. We look forward to potentially welcoming you as our next Qdrant Star. [Apply here!](https://forms.gle/q4fkwudDsy16xAZk8)
If you have a strong understanding of vector search technologies, enjoy creating content, speaking at conferences, and actively engage with our community. If this sounds like you, don't hesitate to apply. We look forward to potentially welcoming you as our next Qdrant Star. [Apply here!](https://forms.gle/vTuy8Fe9RFdt4SiB9)
Share your journey with vector search technologies and how you plan to contribute further.
#### Nominate a Qdrant Star
Do you know someone who could be our next Qdrant Star? Please submit your nomination through our [nomination form](https://forms.gle/n4zv7JRkvnp28qv17), explaining why they're a great fit. Your recommendation could help us find the next standout ambassador.
Do you know someone who could be our next Qdrant Star? Please submit your nomination through our [nomination form](hhttps://forms.gle/jsEJ9zjdaxqk7F5b9), explaining why they're a great fit. Your recommendation could help us find the next standout ambassador.
#### Learn More
@@ -0,0 +1,73 @@
---
title: "Announcing the Vector Space Day 2025 Speaker Lineup"
draft: false
slug: vector-space-day-lineup-2025
short_description: "We are just days away from Vector Space Day in Berlin, and the full speaker lineup is here! "
description: "We are just days away from Vector Space Day in Berlin, and the full speaker lineup is here! This year’s program spans keynotes, deep-dive technical sessions, and lightning talks, covering everything from benchmarking search engines to scalable AI memory and multimodal embeddings."
preview_image: /blog/vector-space-day-2025-lineup/lineup-hero.jpg
social_preview_image: /blog/vector-space-day-2025-lineup/lineup-hero.jpg
date: 2025-09-15
author: Qdrant
featured: true
tags:
- news
- blog
---
# Announcing the Vector Space Day 2025 Speaker Lineup
We are just days away from [Vector Space Day](https://luma.com/p7w9uqtz) in Berlin, and the full speaker lineup is here\! This year’s program spans keynotes, deep-dive technical sessions, and lightning talks, covering everything from benchmarking search engines to scalable AI memory and multimodal embeddings. Here’s what to expect.
## Opening Keynotes
The day begins with perspectives from across the ecosystem:
* **Andre Zayarni, Andrey Vasnetsov,** and **Neil Kanungo** sharing Qdrant’s vision for the future of vector search and how devs can engage with the Qdrant Community.
* **Robert Eichenseer (Microsoft), Kevin Cochrane (Vultr),** and **Inaam Syed (AWS)** offering insights on how cloud, infrastructure, and developer communities are reshaping AI systems.
## Breakout Sessions
#### Track A: Milky Way \- Architectures, Infrastructure and Multimodal Retrieval
* **AskNews** \- *Building a News Sleuth for the Deep Research Paradigm:* How high-performance hybrid retrieval can support investigative journalism and geopolitical risk monitoring.
* **Delivery Hero** \- *How to Cheat at Benchmarking Search Engines:* Lessons from building reproducible benchmarking harnesses and public leaderboards.
* **Neo4j** \- *Hands-On GraphRAG:* Practical guidance on combining knowledge graphs with RAG for more explainable retrieval.
* **Superlinked** \- *Beyond Text-Only:* How mixture of encoders unlocks advanced retrieval using Google DeepMind’s latest embeddings.
* **Jina AI** \- *Vision-Language Models for Embedding:* Training insights for multimodal embeddings that span text, diagrams, and UI screenshots.
* **TwelveLabs** \- *Practical Multimodal Embeddings:* Real workflows for cross-modal video search and recommendations.
* **Baseten** \- *High Throughput, Low Latency Embedding Pipelines:* Patterns and open-source tools for production-ready embedding inference.
* **Google** **DeepMind** \- *Vector Search with Gemini and EmbeddingGemma:* Deploying cutting-edge embeddings with the right indexing strategies.
#### Track B: Andromeda \- AI Workflows, Agents and Applications
* **Linkup** \- *Beyond Web Search:* Infrastructure for AI-native agents that need structured, real-time web intelligence.
* **Cognee** \- *Building Scalable AI Memory:* Abstractions that sync graphs and vectors for durable, multi-backend AI memory.
* **n8n** \- *Evaluate Your Qdrant-RAG Agents:* A live no-code session on agent evaluation using n8n’s native tools.
* **Arize AI** \- *Self-Improving Evaluations:* Feedback loops and tracing for reliable agentic RAG in production.
* **LlamaIndex** \- *Vector Databases for Workflow Engineering:* Using Qdrant to orchestrate context-aware AI pipelines.
* **deepset** \- *Agent-Powered Retrieval with Haystack and Qdrant:* When retrieval agents outperform or overcomplicate pipelines.
* **GoodData** \- *Scaling Real-Time RAG for Analytics:* Lessons from streaming BI artifacts into Qdrant for natural-language analytics.
* **Equal** \- *Redefining Long-Term Memory:* Streaming-driven ingestion architectures that give agents enterprise-grade responsiveness.
## Lightning Talks
The afternoon features rapid-fire sessions from innovators including:
* **bakdata** \- Streaming pipelines with Kafka and Qdrant.
* **KI Reply** \- GDPR-compliant retrieval with graph-augmented RAG.
* **iCompetence** \- Personalized product discovery with multimodal vectors.
* **Raiva** **Technologies** \- Voice-first multimodal search with Qdrant.
* **Superlinked** \- Stories from the AI Search Frontier.
#### 👉 [**View Full Agenda Details**](https://try.qdrant.tech/hubfs/VSD-2025-program.pdf)
## Hackathon Awards and After Party
We will celebrate the winners of the [Think Outside the Bot Hackathon](https://try.qdrant.tech/hackathon-2025), followed by closing remarks and an after party with live DJ and networking.
## Don’t Miss Out
Vector Space Day 2025 takes place in Berlin on September 26, 2025\. Space is limited, and registration is filling fast\!
#### [**Register today**](https://luma.com/p7w9uqtz)
@@ -5,7 +5,7 @@ slug: vector-space-day-2025
short_description: "We’re hosting our first-ever full-day in-person Vector Space Day this September in Berlin, and you’re invited."
description: "From building scalable RAG pipelines to enabling real-time AI memory and next-gen context engineering, we’re covering the full spectrum of modern vector-native search."
preview_image: /blog/vector-space-day-2025/Vector-Space-Day-Hero.jpg
social_preview_image: /blog/vector-space-day-2025/partners_6-aug.png.jpg
social_preview_image: /blog/vector-space-day-2025/partners-20.08.png
date: 2025-07-14
author: Qdrant
featured: false
@@ -61,7 +61,7 @@ Missed the deadline? Go ahead and send it in anyway. Proposals submitted after t
### Partners
We’ll be joined by leading organizations including AWS, Microsoft, Vultr, Jina, DeepSet, LlamaIndex, TwelveLabs, n8n, Neo4j, MistralAI, DataTalks.Club, and the MLOps Community, and more.
We’ll be joined by leading organizations including AWS, Microsoft, Vultr, Jina, DeepSet, LlamaIndex, TwelveLabs, n8n, Neo4j, Superlinked, Linkup, DataTalks.Club, and the MLOps Community, and more.
These partners represent a cross-section of the most influential players in AI infrastructure and applied research, and we’re proud to collaborate with them to bring this event to life.
@@ -69,6 +69,10 @@ Their involvement underscores the growing momentum behind vector search and retr
![Partners](/blog/vector-space-day-2025/Vector-Space-Day-Partners-sept17.jpg)
​Vielen Dank an unseren Medienpartner. Thanks to our media partner.
![Media](/blog/vector-space-day-2025/media-partner.png)
### Get Your Ticket
General admission: €50
@@ -77,7 +81,9 @@ General admission: €50
Space is limited.
### Global Hackathon - Submissions Closed
### Global Hackathon — Submissions Closed
![Hackathon](/blog/vector-space-day-2025/hackathon-26aug.png)
In the lead-up to Vector Space Day, we're hosting **Think Outside the Bot**, a global, virtual hackathon challenging devs to reimagine what's possible with vector search. Forget the classical RAG chatbot! Explore multi-modal applications, intelligent recommendations, and advanced vector search that go far beyond conversational interfaces.
@@ -91,7 +97,6 @@ In the lead-up to Vector Space Day, we're hosting **Think Outside the Bot**, a g
[**Submissions are closed but you can learn more.**](https://try.qdrant.tech/hackathon-2025)
### Need your manager’s approval to attend?
We’ve got you covered. Download this ready-to-send request letter to help explain why attending Vector Search Day is a valuable use of your time (and budget). [Download now](https://docs.google.com/document/d/1EivCVK47XEFXAhyoo8QaCBX0Op6uicUODAxTGXhZxrs/edit?usp=sharing).