diff --git a/.github/workflows/internal-dead-links.yml b/.github/workflows/internal-dead-links.yml index 699e30260..0f8763f71 100644 --- a/.github/workflows/internal-dead-links.yml +++ b/.github/workflows/internal-dead-links.yml @@ -31,7 +31,7 @@ jobs: id: lychee uses: lycheeverse/lychee-action@v1.8.0 with: - args: --max-redirects 0 --exclude '.*' --include 'http://localhost:1313/.*' --base http://localhost:1314/ qdrant-landing/public/ + args: --max-redirects 0 --exclude '.*' --include '^http://localhost:1314/[^%]+$' --base http://localhost:1314/ qdrant-landing/public/ fail: true env: GITHUB_TOKEN: ${{secrets.GITHUB_TOKEN}} diff --git a/qdrant-landing/content/about-us/about-us-resources.md b/qdrant-landing/content/about-us/about-us-resources.md index b93754ddb..6f3854075 100644 --- a/qdrant-landing/content/about-us/about-us-resources.md +++ b/qdrant-landing/content/about-us/about-us-resources.md @@ -2,6 +2,6 @@ title: Brand Resources seeOpenRoles: text: View Brand Resources - url: /brand-resources + url: /brand-resources/ image: /img/about-us/style-guide.svg --- diff --git a/qdrant-landing/content/about-us/about-us-values.md b/qdrant-landing/content/about-us/about-us-values.md index eb4022ed6..dc7512d75 100644 --- a/qdrant-landing/content/about-us/about-us-values.md +++ b/qdrant-landing/content/about-us/about-us-values.md @@ -17,13 +17,13 @@ values: icon: src: /img/about-us/rust-logo.svg alt: rust-logo - link: /articles/why-rust + link: /articles/why-rust/ - id: 3 title: 9k Members icon: src: /img/about-us/discord-logo.svg alt: discord-logo - link: /community + link: /community/ - id: 4 title: 100+ employees icon: diff --git a/qdrant-landing/content/advanced-search/advanced-search-features.md b/qdrant-landing/content/advanced-search/advanced-search-features.md index 16a6507df..237abba7e 100644 --- a/qdrant-landing/content/advanced-search/advanced-search-features.md +++ b/qdrant-landing/content/advanced-search/advanced-search-features.md @@ -10,7 +10,7 @@ features: description: Qdrant optimizes similarity search, identifying the closest database items to any query vector for applications like recommendation systems, RAG and image retrieval, enhancing accuracy and user experience. link: text: Learn More - url: /documentation/concepts/search/ + url: /documentation/search/search/ - id: 1 icon: src: /icons/outline/search-text-blue.svg diff --git a/qdrant-landing/content/ai-agents/ai-agents-features.md b/qdrant-landing/content/ai-agents/ai-agents-features.md index 48c377597..a452c17df 100644 --- a/qdrant-landing/content/ai-agents/ai-agents-features.md +++ b/qdrant-landing/content/ai-agents/ai-agents-features.md @@ -48,7 +48,7 @@ features: description: Qdrant’s architecture is optimized for high-throughput embedding processing, minimizing CPU load and preventing performance bottlenecks. This enables AI agents in Agentic RAG workflows to execute complex, multi-step tasks efficiently, ensuring smooth operation even at scale. link: text: Distributed Deployment - url: /documentation/guides/distributed_deployment/ + url: /documentation/operations/distributed_deployment/ - id: 4 icon: src: /icons/outline/speedometer-blue.svg diff --git a/qdrant-landing/content/ai-agents/ai-agents-integrations.md b/qdrant-landing/content/ai-agents/ai-agents-integrations.md index e25fb98b4..93e47217b 100644 --- a/qdrant-landing/content/ai-agents/ai-agents-integrations.md +++ b/qdrant-landing/content/ai-agents/ai-agents-integrations.md @@ -14,7 +14,7 @@ integrations: alt: Open AI logo title: Swarm description: Decentralized platform enabling collaboration among AI agents for task completion. - url: /documentation/frameworks/swarm/ + url: /documentation/frameworks/ - id: 2 icon: src: /img/integrations/integration-crew-ai.svg diff --git a/qdrant-landing/content/ai-agents/ai-agents-use-cases.md b/qdrant-landing/content/ai-agents/ai-agents-use-cases.md index 665b2d40f..cfd40aee9 100644 --- a/qdrant-landing/content/ai-agents/ai-agents-use-cases.md +++ b/qdrant-landing/content/ai-agents/ai-agents-use-cases.md @@ -20,7 +20,7 @@ features: description: Learn how to build OpenAI Swarm agents using Qdrant for fast, scalable vector search and real-time actions. link: text: Read the Docs - url: /documentation/frameworks/swarm/ + url: /documentation/frameworks/ - id: 2 image: src: /img/ai-agents-use-cases/ai-scheduler.svg diff --git a/qdrant-landing/content/articles/agentic-builders-guide.md b/qdrant-landing/content/articles/agentic-builders-guide.md index b440769a5..cdf2927f2 100644 --- a/qdrant-landing/content/articles/agentic-builders-guide.md +++ b/qdrant-landing/content/articles/agentic-builders-guide.md @@ -38,7 +38,7 @@ Reliable agentic workflows require a clear plan executed with precise tools. A v * [**Real-time memory layer:**](https://qdrant.tech/blog/case-study-fieldy/) fast access to prior steps, actions, and knowledge * [**Multimodal support**](https://qdrant.tech/blog/case-study-mixpeek/): text, image, videos, audio, and code * [**Hybrid search**](https://qdrant.tech/articles/hybrid-search/)**:** combining dense \+ sparse vectors -* [**Advanced filtering**](https://qdrant.tech/documentation/concepts/filtering/)**:** semantic \+ metadata \+ keyword constraints +* [**Advanced filtering**](https://qdrant.tech/documentation/search/filtering/)**:** semantic \+ metadata \+ keyword constraints * [**Millisecond vector retrieval**](https://qdrant.tech/articles/vector-search-production/)**:** Fast retrieval at \>billion vector scale Throughout this article, we’ll use TripAdvisor’s [TripBuilder](https://www.tripadvisor.com/TripBuilder) to illustrate how each of the four concepts above is critical for building and deploying agentic search at scale. @@ -55,7 +55,7 @@ When your agent has to design the itinerary for the trip to Berlin, TripBuilder ## Why Speed Matters in Agentic Retrieval -Agents may run thousands of searches to answer a single complex question. Sometimes those searches are run in parallel, but sometimes they’re run sequentially. For those sequential searches, each millisecond saved compounds. That’s when Qdrant's [Rust-based](https://qdrant.tech/articles/why-rust/) [HNSW indexing](https://qdrant.tech/documentation/concepts/indexing/#hnsw-graph-index), which enables millisecond-level retrieval at scale, becomes a differentiator. +Agents may run thousands of searches to answer a single complex question. Sometimes those searches are run in parallel, but sometimes they’re run sequentially. For those sequential searches, each millisecond saved compounds. That’s when Qdrant's [Rust-based](https://qdrant.tech/articles/why-rust/) [HNSW indexing](https://qdrant.tech/documentation/manage-data/indexing/#hnsw-graph-index), which enables millisecond-level retrieval at scale, becomes a differentiator. Let’s imagine that you see a 75ms retrieval performance gain with Qdrant. On its own, that might not make a demonstrable difference in the search experience. But if an agent performs four sequential searches, that 300ms starts to have a user experience impact. With that delay, a user can quickly get bored or distracted, leading to frustration. With agentic AI, the retrieval speed optimizations are compounded. @@ -81,7 +81,7 @@ Combining semantic search, metadata, and keyword [filters](https://qdrant.tech/a ## Real-Time Memory Layer for Agents -Users expect agents to remember the details of their conversation. Your agent's memory must be updated every time it gets new information, not just at the start of each session. Qdrant supports [real-time upserts](https://qdrant.tech/documentation/concepts/points/#upsert-points), giving your agent access to the freshest data for short-term memory and long-term memory. +Users expect agents to remember the details of their conversation. Your agent's memory must be updated every time it gets new information, not just at the start of each session. Qdrant supports [real-time upserts](https://qdrant.tech/documentation/manage-data/points/#upsert-points), giving your agent access to the freshest data for short-term memory and long-term memory. Not all information is timeless. For an agent, knowing what's recent is as important as knowing what's relevant. This is where [decay functions](https://qdrant.tech/blog/decay-functions/) come in, acting as a "recency boost" during a search. @@ -109,13 +109,13 @@ Once you’ve built a fast, accurate, secure, and scalable agent, how do you kno Your agent’s ability to complete complex tasks is only as good as the context it can retrieve. It is crucial to closely and continuously monitor the agent’s performance with grounding checks to spot and prevent hallucinations, recall@k to ensure relevance in search, and MMR parameters for diversity. -#### [**Performance-Cost Tradeoff**](https://qdrant.tech/documentation/guides/optimize/) +#### [**Performance-Cost Tradeoff**](https://qdrant.tech/documentation/operations/optimize/) In production agents, efficiency is one of, if not the, most important metric to track. First, in enterprise environments you must meet strict latency budgets. Evaluations also let you track the cost per task by tracking token usage and end-to-end compute time. Finally, you can track the effectiveness of your memory layer by monitoring cache hit rates for your memory banks. ![Tradeoff Triangle](/articles_data/agentic-builders-guide/tradeoff-triangle.png) -#### [**Guardrails & Fallbacks**](https://qdrant.tech/documentation/guides/security/) +#### [**Guardrails & Fallbacks**](https://qdrant.tech/documentation/operations/security/) Just like humans, agents aren’t perfect. A production agentic system should expect and anticipate failures and have guardrails to handle them gracefully. You can also use a human in the loop when confidence scores are below a chosen threshold or the query touches on a high-stakes or sensitive topic. To handle a wide range of queries, your agent should use hybrid search. Hybrid search combines the power of semantic search for understanding meaning with the precision of keyword search for exact matches, giving you the best of both worlds. @@ -125,13 +125,13 @@ The same agent that speeds through a toy dataset with 10,000 points will become We’ll talk about three concepts you can take advantage of to improve your scale, but if you want even more information on how to scale, check out this [article](https://qdrant.tech/documentation/database-tutorials/large-scale-search/) on large scale search. -As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/guides/distributed_deployment/), creating copies of your shards across the cluster. +As your dataset and traffic grow, Qdrant Cloud offers a suite of features to ensure your system can scale effectively. Horizontal scaling is achieved through [sharding](https://qdrant.tech/articles/multitenancy/), which splits your collection across multiple nodes to distribute the load and improve performance. For high availability and fault tolerance, Qdrant supports [replication](https://qdrant.tech/documentation/operations/distributed_deployment/), creating copies of your shards across the cluster. -Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/guides/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search. +Qdrant provides robust tools for resource and cost optimization. Vector [quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses your data, significantly reducing its memory footprint and speeding up search. ![Quantization](/articles_data/agentic-builders-guide/quantization.png) -To further manage costs as your dataset expands, [on-disk storage](https://qdrant.tech/documentation/concepts/storage/) allows you to keep the full vectors on more affordable SSDs while the necessary index data remains in RAM. This hybrid approach enables searches over billions of vectors without the high cost of keeping all data in memory. +To further manage costs as your dataset expands, [on-disk storage](https://qdrant.tech/documentation/manage-data/storage/) allows you to keep the full vectors on more affordable SSDs while the necessary index data remains in RAM. This hybrid approach enables searches over billions of vectors without the high cost of keeping all data in memory. To enhance the relevance and precision of search queries, Qdrant natively supports hybrid search, which combines traditional keyword-based search with semantic vector search. By doing so, your application can find documents that match exact terms, such as product codes or names, while also discovering documents that are semantically similar in meaning. This ensures that you can find the most relevant information, even in massive and complex datasets. @@ -143,11 +143,11 @@ Workflows that are effective for a single user or a handful of users in a develo Qdrant provides production-grade [authorization and authentication](https://qdrant.tech/documentation/cloud/authentication/) via API keys, multitenancy, and Role-Based Access Control (RBAC) to make sure your agent doesn’t go rogue. -Authentication for agents is handled by [API keys](https://qdrant.tech/documentation/guides/security/#api-keys). With Qdrant, API keys are more than just a password. They act as smart credentials that also carry the details of the authorization rules that are enforced once the agent’s identity is confirmed. These keys can be dynamically created as temporary credentials for each user session, making access limited to the session and secure. +Authentication for agents is handled by [API keys](https://qdrant.tech/documentation/operations/security/#api-keys). With Qdrant, API keys are more than just a password. They act as smart credentials that also carry the details of the authorization rules that are enforced once the agent’s identity is confirmed. These keys can be dynamically created as temporary credentials for each user session, making access limited to the session and secure. Note: Qdrant also supports concurrent queries, so your search won’t slow down as more users are writing queries simultaneously. -Authorization is handled by [RBAC](https://qdrant.tech/articles/data-privacy/) and [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/), which work hand-in-hand to define and enforce permissions specific to the agent. RBAC is a set of rules that defines the allowed permissions inlcuding read-only, read-write, and admin controls. It answers the question, “What is this agent allowed to do?” For instance, can it only search for hotels (read-only), or can it also add, update, and delete them (read-write)? +Authorization is handled by [RBAC](https://qdrant.tech/articles/data-privacy/) and [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/), which work hand-in-hand to define and enforce permissions specific to the agent. RBAC is a set of rules that defines the allowed permissions including read-only, read-write, and admin controls. It answers the question, “What is this agent allowed to do?” For instance, can it only search for hotels (read-only), or can it also add, update, and delete them (read-write)? ![Multi-tenancy](/articles_data/agentic-builders-guide/multi-tenancy.png) diff --git a/qdrant-landing/content/articles/agentic-rag.md b/qdrant-landing/content/articles/agentic-rag.md index a6e77c5ec..057ecc64d 100644 --- a/qdrant-landing/content/articles/agentic-rag.md +++ b/qdrant-landing/content/articles/agentic-rag.md @@ -212,7 +212,7 @@ Some of the key concepts of CrewAI include: Qdrant comes into play, as it might be used as a long-term memory layer.** CrewAI provides a rich set of tools integrated into the framework. That may be a huge advantage for those who want to -combine RAG with e.g. code execution, or image generation. The ecosystem is rich, however brining your own tools is +combine RAG with e.g. code execution, or image generation. The ecosystem is rich, however bringing your own tools is not a big deal, as CrewAI is designed to be extensible. A simple agentic RAG application implemented in CrewAI could look like this: diff --git a/qdrant-landing/content/articles/binary-quantization-openai.md b/qdrant-landing/content/articles/binary-quantization-openai.md index 5c23fc902..0e0556640 100644 --- a/qdrant-landing/content/articles/binary-quantization-openai.md +++ b/qdrant-landing/content/articles/binary-quantization-openai.md @@ -216,6 +216,6 @@ We recommend the following best practices for leveraging Binary Quantization to Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/login) in our cloud hosted service. -The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/guides/quantization/). +The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials-develop/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/manage-data/quantization/). Want to discuss these findings and learn more about Binary Quantization? [Join our Discord community.](https://discord.gg/qdrant) diff --git a/qdrant-landing/content/articles/binary-quantization.md b/qdrant-landing/content/articles/binary-quantization.md index 8f7830909..60d164ae6 100644 --- a/qdrant-landing/content/articles/binary-quantization.md +++ b/qdrant-landing/content/articles/binary-quantization.md @@ -231,6 +231,6 @@ If you determine that binary quantization is appropriate for your datasets and q Binary quantization is exceptional if you need to work with large volumes of data under high recall expectations. You can try this feature either by spinning up a [Qdrant container image](https://hub.docker.com/r/qdrant/qdrant) locally or, having us create one for you through a [free account](https://cloud.qdrant.io/signup) in our cloud hosted service. -The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/guides/quantization/). +The article gives examples of data sets and configuration you can use to get going. Our documentation covers [adding large datasets to Qdrant](/documentation/tutorials-develop/bulk-upload/) to your Qdrant instance as well as [more quantization methods](/documentation/manage-data/quantization/). If you have any feedback, drop us a note on Twitter or LinkedIn to tell us about your results. [Join our lively Discord Server](https://discord.gg/Qy6HCJK9Dc) if you want to discuss BQ with like-minded people! diff --git a/qdrant-landing/content/articles/cars-recognition.md b/qdrant-landing/content/articles/cars-recognition.md index c567b6686..6b5d6d092 100644 --- a/qdrant-landing/content/articles/cars-recognition.md +++ b/qdrant-landing/content/articles/cars-recognition.md @@ -34,7 +34,7 @@ However, similarity learning comes with its own difficulties such as: Quaterion is a fine tuning framework built to tackle such problems in similarity learning. It uses [PyTorch Lightning](https://www.pytorchlightning.ai/) -as a backend, which is advertized with the motto, "spend more time on research, less on engineering." +as a backend, which is advertised with the motto, "spend more time on research, less on engineering." This is also true for Quaterion, and it includes: 1. Trainable and servable model classes, diff --git a/qdrant-landing/content/articles/data-privacy.md b/qdrant-landing/content/articles/data-privacy.md index 70afd8374..5447e0165 100644 --- a/qdrant-landing/content/articles/data-privacy.md +++ b/qdrant-landing/content/articles/data-privacy.md @@ -55,7 +55,7 @@ On Qdrant Cloud, you can create API keys using the [Cloud Dashboard](https://qdr For on-premise or local deployments, you'll need to configure API key authentication. This involves specifying a key in either the Qdrant configuration file or as an environment variable. This ensures that all requests to the server must include a valid API key sent in the header. -When using the simple API key-based authentication, you should also turn on TLS encryption. Otherwise, you are exposing the connection to sniffing and MitM attacks. To secure your connection using TLS, you would need to create a certificate and private key, and then [enable TLS](/documentation/guides/security/#tls) in the configuration. +When using the simple API key-based authentication, you should also turn on TLS encryption. Otherwise, you are exposing the connection to sniffing and MitM attacks. To secure your connection using TLS, you would need to create a certificate and private key, and then [enable TLS](/documentation/operations/security/#tls) in the configuration. API authentication, coupled with TLS encryption, offers a first layer of security for your Qdrant instance. However, to enable more granular access control, the recommended approach is to leverage JSON Web Tokens (JWTs). diff --git a/qdrant-landing/content/articles/dedicated-service.md b/qdrant-landing/content/articles/dedicated-service.md index 3f30a02e9..c045617d6 100644 --- a/qdrant-landing/content/articles/dedicated-service.md +++ b/qdrant-landing/content/articles/dedicated-service.md @@ -107,7 +107,7 @@ In practice, this means that your main database becomes burdened with high memor Fortunately, the data synchronization problem is not new and definitely not unique to vector search. There are many well-known solutions, starting with message queues and ending with specialized ETL tools. -For example, we recently released our [integration with Airbyte](/documentation/integrations/airbyte/), allowing you to synchronize data from various sources into Qdrant incrementally. +For example, we recently released our [integration with Airbyte](/documentation/data-management/airbyte/), allowing you to synchronize data from various sources into Qdrant incrementally. ###### You have to pay for a vector service uptime and data transfer of both solutions. @@ -115,7 +115,7 @@ In the open-source world, you pay for the resources you use, not the number of d Resources depend more on the optimal solution for each use case. As a result, running a dedicated vector search engine can be even cheaper, as it allows optimization specifically for vector search use cases. -For instance, Qdrant implements a number of [quantization techniques](/documentation/guides/quantization/) that can significantly reduce the memory footprint of embeddings. +For instance, Qdrant implements a number of [quantization techniques](/documentation/manage-data/quantization/) that can significantly reduce the memory footprint of embeddings. In terms of data transfer costs, on most cloud providers, network use within a region is usually free. As long as you put the original source data and the vector store in the same region, there are no added data transfer costs. diff --git a/qdrant-landing/content/articles/dedicated-vector-search.md b/qdrant-landing/content/articles/dedicated-vector-search.md index 237e548ed..aa2fe22ad 100644 --- a/qdrant-landing/content/articles/dedicated-vector-search.md +++ b/qdrant-landing/content/articles/dedicated-vector-search.md @@ -22,9 +22,9 @@ In this article, we will describe the unique challenges vector search poses and ## Vectors ![vectors](/articles_data/dedicated-vector-search/image1.jpg) -Let's look at the central concept of vector databases — [**vectors**](/documentation/concepts/vectors/). +Let's look at the central concept of vector databases — [**vectors**](/documentation/manage-data/vectors/). -Vectors (also known as embeddings) are high-dimensional representations of various data points — texts, images, videos, etc. Many state-of-the-art (SOTA) embedding models generate representations of over 1,500 dimensions. When it comes to state-of-the-art PDF retrieval, the representations can reach [**over 100,000 dimensions per page**](/documentation/advanced-tutorials/pdf-retrieval-at-scale/). +Vectors (also known as embeddings) are high-dimensional representations of various data points — texts, images, videos, etc. Many state-of-the-art (SOTA) embedding models generate representations of over 1,500 dimensions. When it comes to state-of-the-art PDF retrieval, the representations can reach [**over 100,000 dimensions per page**](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/). This brings us to the first challenge of vector search — vectors are heavy. @@ -52,7 +52,7 @@ However, vectors have positive properties as well. One of the most important is Embedding models are designed to produce vectors of a fixed size. We have to use it to our advantage. -For fast search, vectors need to be instantly accessible. Whether in [**RAM or disk**](/documentation/concepts/storage/), vectors should be stored in a format that allows quick access and comparison. This is essential, as vector comparison is a very hot operation in vector search workloads. It is often performed thousands of times per search query, so even a small overhead can lead to a significant slowdown. +For fast search, vectors need to be instantly accessible. Whether in [**RAM or disk**](/documentation/manage-data/storage/), vectors should be stored in a format that allows quick access and comparison. This is essential, as vector comparison is a very hot operation in vector search workloads. It is often performed thousands of times per search query, so even a small overhead can lead to a significant slowdown. For dedicated storage, vectors' fixed size comes as a blessing. Knowing how much space one data point needs, we don't have to deal with the usual overhead of locating data — the location of elements in storage is straightforward to calculate. @@ -103,7 +103,7 @@ A strictly consistent transactional approach also loses its attractiveness when ## Vector Index ![vector-index](/articles_data/dedicated-vector-search/image3.jpg) -[**Vector search**](/documentation/concepts/search/) relies on high-dimensional vector mathematics, making it computationally heavy at scale. A brute-force similarity search would require comparing a query against every vector in the database. In a database with 100 million 1536-dimensional vectors, performing 100 million comparisons per one query is unfeasible for production scenarios. Instead of a brute-force approach, vector databases have specialized approximate nearest neighbour (ANN) indexes that balance search precision and speed. These indexes require carefully designed architectures to make their maintenance in production feasible. +[**Vector search**](/documentation/search/search/) relies on high-dimensional vector mathematics, making it computationally heavy at scale. A brute-force similarity search would require comparing a query against every vector in the database. In a database with 100 million 1536-dimensional vectors, performing 100 million comparisons per one query is unfeasible for production scenarios. Instead of a brute-force approach, vector databases have specialized approximate nearest neighbour (ANN) indexes that balance search precision and speed. These indexes require carefully designed architectures to make their maintenance in production feasible. {{< figure src=/articles_data/dedicated-vector-search/hnsw.png caption="HNSW Index" width=80% >}} @@ -111,7 +111,7 @@ One of the most popular vector indexes is **HNSW (Hierarchical Navigable Small W ### Index Complexity -[**HNSW**](/documentation/concepts/indexing/) is structured as a multi-layered graph. With a new data point inserted, the algorithm must compare it to existing nodes across several layers to index it. As the number of vectors grows, these comparisons will noticeably slow down the construction process, making updates increasingly time-consuming. The indexing operation can quickly become the bottleneck in the system, slowing down search requests. +[**HNSW**](/documentation/manage-data/indexing/) is structured as a multi-layered graph. With a new data point inserted, the algorithm must compare it to existing nodes across several layers to index it. As the number of vectors grows, these comparisons will noticeably slow down the construction process, making updates increasingly time-consuming. The indexing operation can quickly become the bottleneck in the system, slowing down search requests. Building an HNSW monolith means limiting the scalability of your solution — its size has to be capped, as its construction time scales **non-linearly** with the number of elements. To keep the construction process feasible and ensure it doesn't affect the search time, we came up with a layered architecture that breaks down all data management into small units called **segments**. @@ -128,7 +128,7 @@ With index maintenance divided between segments, Qdrant can ensure high performa | | | |---------------------|-------------| | **Mutable Segments** | These are used for quickly ingesting new data and handling changes (updates) to existing data. | -| **Immutable Segments** | Once a mutable segment reaches a certain size, an optimization process converts it into an immutable segment, constructing an HNSW index – you could [**read about these optimizers here**](/documentation/concepts/optimizer/#optimizer) in detail. This immutability trick allowed us, for example, to ensure effective [**tenant isolation**](/documentation/concepts/indexing/#tenant-index). | +| **Immutable Segments** | Once a mutable segment reaches a certain size, an optimization process converts it into an immutable segment, constructing an HNSW index – you could [**read about these optimizers here**](/documentation/operations/optimizer/#optimizer) in detail. This immutability trick allowed us, for example, to ensure effective [**tenant isolation**](/documentation/manage-data/indexing/#tenant-index). | Immutable segments are an implementation detail transparent for users — they can delete vectors at any time, while additions and updates are applied to a mutable segment instead. This combination of mutability and immutability allows search and indexing to smoothly run simultaneously, even under heavy loads. This approach minimizes the performance impact of indexing time and allows on-the-fly configuration changes on a collection level (such as enabling or disabling data quantization) without downtimes. @@ -145,7 +145,7 @@ In many vector search solutions, filtering is approached in two ways: **pre-filt | ❌ | **Pre-filtering** | Has the linear complexity of computing the vector mask and becomes a bottleneck for large datasets. | | ❌ | **Post-filtering** | The problem with **post-filtering** is tied to vector search "*everything fits and doesn't at the same time*" nature: imagine a low-cardinality filter that leaves only a few matching elements in the database. If none of them are similar enough to the query to appear in the top-X retrieved results, they'll all be filtered out. | -Qdrant [**took filtering in vector search further**](/articles/vector-search-filtering/), recognizing the limitations of pre-filtering & post-filtering strategies. We developed an adaptation of HNSW — [**filterable HNSW**](/articles/filterable-hnsw/) — that also enables **in-place filtering** during graph traversal. To make this possible, we condition HNSW index construction on possible filtering conditions reflected by [**payload indexes**](/documentation/concepts/indexing/#payload-index) (inverted indexes built on vectors' [**metadata**](/documentation/concepts/payload/)). +Qdrant [**took filtering in vector search further**](/articles/vector-search-filtering/), recognizing the limitations of pre-filtering & post-filtering strategies. We developed an adaptation of HNSW — [**filterable HNSW**](/articles/filterable-hnsw/) — that also enables **in-place filtering** during graph traversal. To make this possible, we condition HNSW index construction on possible filtering conditions reflected by [**payload indexes**](/documentation/manage-data/indexing/#payload-index) (inverted indexes built on vectors' [**metadata**](/documentation/manage-data/payload/)). **Qdrant was designed with a vector index being a central component of the system.** That made it possible to organize optimizers, payload indexes and other components around the vector index, unlocking the possibility of building a filterable HNSW. @@ -165,7 +165,7 @@ The strength of vector search lies in its ability to facilitate [**discovery**]( ### Recommendations -Vector search is perfect for [**recommendations**](/documentation/concepts/explore/#recommendation-api). Imagine browsing for a new book or movie. Instead of searching for an exact match, you might look for stories that capture a certain mood or theme but differ in key aspects from what you already know. For example, you may [**want a film featuring wizards without the familiar feel of the "Harry Potter" series**](https://www.youtube.com/watch?v=O5mT8M7rqQQ). This flexibility is possible because vector search is not tied to the binary "match/not match" concept but operates on distances in a vector space. +Vector search is perfect for [**recommendations**](/documentation/search/explore/#recommendation-api). Imagine browsing for a new book or movie. Instead of searching for an exact match, you might look for stories that capture a certain mood or theme but differ in key aspects from what you already know. For example, you may [**want a film featuring wizards without the familiar feel of the "Harry Potter" series**](https://www.youtube.com/watch?v=O5mT8M7rqQQ). This flexibility is possible because vector search is not tied to the binary "match/not match" concept but operates on distances in a vector space. ### Big Unstructured Data Analysis @@ -193,17 +193,17 @@ Consider some of the advanced features implemented in Qdrant: GPU acceleration in Qdrant is a custom solution developed by an enthusiast from our core team. It's vendor-free and natively supports all Qdrant's unique architectural features, from FIlterable HNSW to multivectors. -- [**Multivectors**](/documentation/concepts/vectors/?q=multivectors#multivectors) +- [**Multivectors**](/documentation/manage-data/vectors/?q=multivectors#multivectors) Some modern embedding models produce an entire matrix (a list of vectors) as output rather than a single vector. Qdrant supports multivectors natively. - This feature is critical when using state-of-the-art retrieval models such as [**ColBERT**](/documentation/fastembed/fastembed-colbert/), ColPali, or ColQwen. For instance, ColPali and ColQwen produce multivector outputs, and supporting them natively is crucial for [**state-of-the-art (SOTA) PDF-retrieval**](/documentation/advanced-tutorials/pdf-retrieval-at-scale/). + This feature is critical when using state-of-the-art retrieval models such as [**ColBERT**](/documentation/fastembed/fastembed-colbert/), ColPali, or ColQwen. For instance, ColPali and ColQwen produce multivector outputs, and supporting them natively is crucial for [**state-of-the-art (SOTA) PDF-retrieval**](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/). In addition to that, we continuously look for improvements in: | | | |----------------------------------|-------------| -| **Memory Efficiency & Compression** | Techniques such as [**quantization**](/documentation/guides/quantization/) and [**HNSW compression**](/blog/qdrant-1.13.x/#hnsw-graph-compression) to reduce storage requirements | -| **Retrieval Algorithms** | Support for the latest retrieval algorithms, including [**sparse neural retrieval**](/articles/modern-sparse-neural-retrieval/), [**hybrid search**](/documentation/concepts/hybrid-queries/) methods, and [**re-rankers**](/documentation/fastembed/fastembed-rerankers/). | +| **Memory Efficiency & Compression** | Techniques such as [**quantization**](/documentation/manage-data/quantization/) and [**HNSW compression**](/blog/qdrant-1.13.x/#hnsw-graph-compression) to reduce storage requirements | +| **Retrieval Algorithms** | Support for the latest retrieval algorithms, including [**sparse neural retrieval**](/articles/modern-sparse-neural-retrieval/), [**hybrid search**](/documentation/search/hybrid-queries/) methods, and [**re-rankers**](/documentation/fastembed/fastembed-rerankers/). | | **Vector Data Analysis & Visualization** | Tools like the [**distance matrix API**](/blog/qdrant-1.12.x/#distance-matrix-api-for-data-insights) provide insights into vectorized data, and a [**Web UI**](/blog/qdrant-1.11.x/#web-ui-search-quality-tool) allows for intuitive exploration of data. | | **Search Speed & Scalability** | Includes optimizations for [**multi-tenant environments**](/articles/multitenancy/) to ensure efficient and scalable search. | diff --git a/qdrant-landing/content/articles/discovery-search.md b/qdrant-landing/content/articles/discovery-search.md index 1cdc07422..91db4532a 100644 --- a/qdrant-landing/content/articles/discovery-search.md +++ b/qdrant-landing/content/articles/discovery-search.md @@ -38,7 +38,7 @@ This is where a __vector _context___ can help. We define _context_ as a list of ![Discovery search visualization](/articles_data/discovery-search/discovery-search.png) -While positive and negative vectors might suggest the use of the recommendation interface, in the case of _context_ they require to be paired up in a positive-negative fashion. This is inspired from the machine-learning concept of _triplet loss_, where you have three vectors: an anchor, a positive, and a negative. Triplet loss is an evaluation of how much the anchor is closer to the positive than to the negative vector, so that learning happens by "moving" the positive and negative points to try to get a better evaluation. However, during discovery, we consider the positive and negative vectors as static points, and we search through the whole dataset for the "anchors", or result candidates, which fit this characteristic better. +While positive and negative vectors might suggest the use of the recommendation interface, in the case of _context_ they require to be paired up in a positive-negative fashion. This is inspired from the machine-learning concept of _triplet loss_, where you have three vectors: an anchor, a positive, and a negative. Triplet loss is an evaluation of how much the anchor is closer to the positive than to the negative vector, so that learning happens by "moving" the positive and negative points to try to get a better evaluation. However, during discovery, we consider the positive and negative vectors as static points, and we search through the whole dataset for the "anchors", or result candidates, which fit this characteristic better. ![Triplet loss](/articles_data/discovery-search/triplet-loss.png) @@ -100,4 +100,4 @@ This way you can give refreshing recommendations, while still being in control b - Discovery search is a powerful tool for controlled exploration in vector spaces. Context, consisting of positive and negative vectors constrain the search space, while a target guides the search. - Real-world applications include multimodal search, diverse recommendations, and context-driven exploration. -- Ready to learn more about the math behind it and how to use it? Check out the [documentation](/documentation/concepts/explore/#discovery-api) \ No newline at end of file +- Ready to learn more about the math behind it and how to use it? Check out the [documentation](/documentation/search/explore/#discovery-api) \ No newline at end of file diff --git a/qdrant-landing/content/articles/distance-based-exploration.md b/qdrant-landing/content/articles/distance-based-exploration.md index c3036e050..24ca789d5 100644 --- a/qdrant-landing/content/articles/distance-based-exploration.md +++ b/qdrant-landing/content/articles/distance-based-exploration.md @@ -11,7 +11,7 @@ draft: false keywords: - clusterization - dimensionality reduction - - vizualization + - visualization category: data-exploration --- @@ -25,8 +25,8 @@ Examining data points individually is not always the best way to grasp the struc As numbers in a table obtain meaning when plotted on a graph, visualising distances (similar/dissimilar) between unstructured data items can reveal hidden structures and patterns. -{{< figure src="/articles_data/distance-based-exploration/data-on-chart.png" alt="Data visualization" caption="Vizualized chart, very intuitive" >}} -There are many tools to investigate data similarity, and Qdrant's [1.12 release](https://qdrant.tech/blog/qdrant-1.12.x/) made it much easier to start this investigation. With the new [Distance Matrix API](/documentation/concepts/explore/#distance-matrix), Qdrant handles the most computationally expensive part of the process—calculating the distances between data points. +{{< figure src="/articles_data/distance-based-exploration/data-on-chart.png" alt="Data visualization" caption="Visualized chart, very intuitive" >}} +There are many tools to investigate data similarity, and Qdrant's [1.12 release](https://qdrant.tech/blog/qdrant-1.12.x/) made it much easier to start this investigation. With the new [Distance Matrix API](/documentation/search/explore/#distance-matrix), Qdrant handles the most computationally expensive part of the process—calculating the distances between data points. In many implementations, the distance matrix calculation was part of the clustering or visualization processes, requiring either brute-force computation or building a temporary index. With Qdrant, however, the data is already indexed, and the distance matrix can be computed relatively cheaply. @@ -62,7 +62,7 @@ PUT /collections/midlib/snapshots/recover ```
-We also need to prepare our python enviroment: +We also need to prepare our python environment: ```bash pip install umap-learn seaborn matplotlib qdrant-client @@ -77,7 +77,7 @@ from qdrant_client import QdrantClient from umap import UMAP # Python implementation for sparse matrices from scipy.sparse import csr_matrix -# For vizualization +# For visualization import seaborn as sns ``` diff --git a/qdrant-landing/content/articles/faq-question-answering.md b/qdrant-landing/content/articles/faq-question-answering.md index ea255751c..ce1f0c19d 100644 --- a/qdrant-landing/content/articles/faq-question-answering.md +++ b/qdrant-landing/content/articles/faq-question-answering.md @@ -55,7 +55,7 @@ As embeddings are vectors, one can apply a simple function to calculate the simi So with similarity learning, all we need to do is provide pairs of correct questions and answers. And then, the model will learn to distinguish proper answers by the similarity of embeddings. ->If you want to learn more about similarity learning and applications, check out this [article](/documentation/tutorials/neural-search/) which might be an asset. +>If you want to learn more about similarity learning and applications, check out this [article](/documentation/tutorials-search-engineering/neural-search/) which might be an asset. ## Let's build diff --git a/qdrant-landing/content/articles/fastembed.md b/qdrant-landing/content/articles/fastembed.md index 1532f6d11..2496fbb67 100644 --- a/qdrant-landing/content/articles/fastembed.md +++ b/qdrant-landing/content/articles/fastembed.md @@ -238,7 +238,7 @@ If you're curious about how FastEmbed and Qdrant can make your search tasks a br 1. **Cloud**: Get started with a free plan on the [Qdrant Cloud](https://qdrant.to/cloud?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article). -2. **Docker Container**: If you're the DIY type, you can set everything up on your own machine. Here's a quick guide to help you out: [Quick Start with Docker](/documentation/quick-start/?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article). +2. **Docker Container**: If you're the DIY type, you can set everything up on your own machine. Here's a quick guide to help you out: [Quick Start with Docker](/documentation/quickstart/?utm_source=qdrant&utm_medium=website&utm_campaign=fastembed&utm_content=article). So, go ahead, take it for a test drive. We're excited to hear what you think! diff --git a/qdrant-landing/content/articles/filterable-hnsw.md b/qdrant-landing/content/articles/filterable-hnsw.md index 8254ca2b3..894d8c61f 100644 --- a/qdrant-landing/content/articles/filterable-hnsw.md +++ b/qdrant-landing/content/articles/filterable-hnsw.md @@ -1,7 +1,7 @@ --- title: Filterable HNSW short_description: How to make ANN search with custom filtering? -description: How to make ANN search with custom filtering? Search in selected subsets without loosing the results. +description: How to make ANN search with custom filtering? Search in selected subsets without losing the results. # external_link: https://blog.vasnetsov.com/posts/categorical-hnsw/ social_preview_image: /articles_data/filterable-hnsw/social_preview.jpg preview_dir: /articles_data/filterable-hnsw/preview diff --git a/qdrant-landing/content/articles/food-discovery-demo.md b/qdrant-landing/content/articles/food-discovery-demo.md index 3e919558e..5f35f52fe 100644 --- a/qdrant-landing/content/articles/food-discovery-demo.md +++ b/qdrant-landing/content/articles/food-discovery-demo.md @@ -27,7 +27,7 @@ Otherwise, read on to learn more about the demo and how it works! In general, our application consists of three parts: a [FastAPI](https://fastapi.tiangolo.com/) backend, a [React](https://react.dev/) frontend, and a [Qdrant](/) instance. The architecture diagram below shows how these components interact with each other: -![Archtecture diagram](/articles_data/food-discovery-demo/architecture-diagram.png) +![Architecture diagram](/articles_data/food-discovery-demo/architecture-diagram.png) ## Why did we use a CLIP model? @@ -92,7 +92,7 @@ in the vector space. ![Random points selection](/articles_data/food-discovery-demo/textual-search.png) This is implemented as [a group search query to Qdrant](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L44). -We didn't use a simple search, but performed grouping by the restaurant to get more diverse results. [Search groups](/documentation/concepts/search/#search-groups) +We didn't use a simple search, but performed grouping by the restaurant to get more diverse results. [Search groups](/documentation/search/search/#search-groups) is a mechanism similar to `GROUP BY` clause in SQL, and it's useful when you want to get a specific number of result per group (in our case just one). ```python @@ -121,7 +121,7 @@ and the demo will update the search results accordingly. #### Negative feedback only -Qdrant [Recommendation API](/documentation/concepts/search/#recommendation-api) needs at least one positive example to work. However, in our demo +Qdrant [Recommendation API](/documentation/search/search/#recommendation-api) needs at least one positive example to work. However, in our demo we want to be able to provide only negative examples. This is because we want to be able to say “I don’t like this dish” without having to like anything first. To achieve this, we use a trick. We negate the vectors of the disliked dishes and use their mean as a query. This way, the disliked dishes will be pushed away from the search results. **This works because the cosine distance is based on the angle between two vectors, and the angle between a vector and its negation is 180 degrees.** @@ -129,8 +129,8 @@ from the search results. **This works because the cosine distance is based on th ![CLIP model](/articles_data/food-discovery-demo/negated-vector.png) Food Discovery Demo [implements that trick](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L122) -by calling Qdrant twice. Initially, we use the [Scroll API](/documentation/concepts/points/#scroll-points) to find disliked items, -and then calculate a negated mean of all their vectors. That allows using the [Search Groups API](/documentation/concepts/search/#search-groups) +by calling Qdrant twice. Initially, we use the [Scroll API](/documentation/manage-data/points/#scroll-points) to find disliked items, +and then calculate a negated mean of all their vectors. That allows using the [Search Groups API](/documentation/search/search/#search-groups) to find the nearest neighbors of the negated mean vector. ```python @@ -163,7 +163,7 @@ response = client.search_groups( #### Positive and negative feedback -Since the [Recommendation API](/documentation/concepts/search/#recommendation-api) requires at least one positive example, we can use it only when +Since the [Recommendation API](/documentation/search/search/#recommendation-api) requires at least one positive example, we can use it only when the user has liked at least one dish. We could theoretically use the same trick as above and negate the disliked dishes, but it would be a bit weird, as Qdrant has that feature already built-in, and we can call it just once to do the job. It's always better to perform the search server-side. Thus, in this case [we just call the Qdrant server with a list of positive and negative examples](https://github.com/qdrant/demo-food-discovery/blob/6b49e11cfbd6412637d527cdd62fe9b9f74ac699/backend/discovery.py#L166), @@ -185,7 +185,7 @@ From the user perspective nothing changes comparing to the previous case. Last but not least, location plays an important role in the food discovery process. You are definitely looking for something you can find nearby, not on the other side of the globe. Therefore, your current location can be toggled as a filtering condition. You can enable it by clicking on “Find near me” icon -in the top right. This way you can find the best pizza in your neighborhood, not in the whole world. Qdrant [geo radius filter](/documentation/concepts/filtering/#geo-radius) is a perfect choice for this. It lets you +in the top right. This way you can find the best pizza in your neighborhood, not in the whole world. Qdrant [geo radius filter](/documentation/search/filtering/#geo-radius) is a perfect choice for this. It lets you filter the results by distance from a given point. ```python @@ -208,7 +208,7 @@ query_filter = models.Filter( ) ``` -Such a filter needs [a payload index](/documentation/concepts/indexing/#payload-index) to work efficiently, and it was created on a collection +Such a filter needs [a payload index](/documentation/manage-data/indexing/#payload-index) to work efficiently, and it was created on a collection we used to create the snapshot. When you import it into your instance, the index will be already there. ## Using the demo @@ -224,7 +224,7 @@ docker-compose up -d ``` The demo will be available at `http://localhost:8001`, but you won't be able to search anything until you [import the snapshot into your Qdrant -instance](/documentation/concepts/snapshots/#recover-via-api). If you don't want to bother with hosting a local one, you can use the [Qdrant +instance](/documentation/operations/snapshots/#recover-via-api). If you don't want to bother with hosting a local one, you can use the [Qdrant Cloud](https://cloud.qdrant.io/) cluster. 4 GB RAM is enough to load all the 2 million entries. ## Fork and reuse diff --git a/qdrant-landing/content/articles/geo-polygon-filter-gsoc.md b/qdrant-landing/content/articles/geo-polygon-filter-gsoc.md index 65aa17777..4afa54d35 100644 --- a/qdrant-landing/content/articles/geo-polygon-filter-gsoc.md +++ b/qdrant-landing/content/articles/geo-polygon-filter-gsoc.md @@ -75,4 +75,4 @@ Being selected for Google Summer of Code 2023 and collaborating with Arnaud and Without a doubt, I'm eager to continue growing alongside this community and contribute to new features and enhancements that elevate the product. I've also become an advocate for Qdrant, introducing this project to numerous coworkers and friends in the tech industry. I'm excited to witness new users and contributors emerge from within my own network! -If you want to try out my work, read the [documentation](/documentation/concepts/filtering/#geo-polygon) and then, either sign up for a free [cloud account](https://cloud.qdrant.io) or download the [Docker image](https://hub.docker.com/r/qdrant/qdrant). I look forward to seeing how people are using my work in their own applications! +If you want to try out my work, read the [documentation](/documentation/search/filtering/#geo-polygon) and then, either sign up for a free [cloud account](https://cloud.qdrant.io) or download the [Docker image](https://hub.docker.com/r/qdrant/qdrant). I look forward to seeing how people are using my work in their own applications! diff --git a/qdrant-landing/content/articles/gridstore-key-value-storage.md b/qdrant-landing/content/articles/gridstore-key-value-storage.md index f5ce9d933..572be9b30 100644 --- a/qdrant-landing/content/articles/gridstore-key-value-storage.md +++ b/qdrant-landing/content/articles/gridstore-key-value-storage.md @@ -97,7 +97,7 @@ Real-world systems don’t operate in a vacuum. Failures happen: software bugs c If one component is updated but another isn’t, the entire system could become inconsistent. Worse, if an operation is only partially written to disk, it could lead to orphaned data, unusable space, or even data corruption. ### Stability Through Idempotency: Recovering With WAL -To guard against these risks, Qdrant relies on a [**Write-Ahead Log (WAL)**](/documentation/concepts/storage/). Before committing an operation, Qdrant ensures that it is at least recorded in the WAL. If a crash happens before all updates are flushed, the system can safely replay operations from the log. +To guard against these risks, Qdrant relies on a [**Write-Ahead Log (WAL)**](/documentation/manage-data/storage/). Before committing an operation, Qdrant ensures that it is at least recorded in the WAL. If a crash happens before all updates are flushed, the system can safely replay operations from the log. This recovery mechanism introduces another essential property: [**idempotence**](https://en.wikipedia.org/wiki/Idempotence). @@ -167,7 +167,7 @@ For even sharper debugging, Property-Based Testing adds automated test generatio Designing for crash resilience is one thing, and proving it works under stress is another. To push Qdrant’s data integrity to the limit, we built [**Crasher**](https://github.com/qdrant/crasher), a test bench that brutally kills and restarts Qdrant while it handles a heavy update workload. -Crasher runs a loop that continuously writes data, then randomly crashes Qdrant. On each restart, Qdrant replays its [**Write-Ahead Log (WAL)**](/documentation/concepts/storage/), and we verify if data integrity holds. Possible anomalies include: +Crasher runs a loop that continuously writes data, then randomly crashes Qdrant. On each restart, Qdrant replays its [**Write-Ahead Log (WAL)**](/documentation/manage-data/storage/), and we verify if data integrity holds. Possible anomalies include: - Missing data (points, vectors, or payloads) - Corrupt payload values @@ -194,7 +194,7 @@ This shows a clear boost in performance. As we can see, the investment in Gridst ### End-to-End Benchmarking -Now, let’s test the impact on a real Qdrant instance. So far, we’ve only integrated Gridstore for [**payloads**](/documentation/concepts/payload/) and [**sparse vectors**](/documentation/concepts/vectors/#sparse-vectors), but even this partial switch should show noticeable improvements. +Now, let’s test the impact on a real Qdrant instance. So far, we’ve only integrated Gridstore for [**payloads**](/documentation/manage-data/payload/) and [**sparse vectors**](/documentation/manage-data/vectors/#sparse-vectors), but even this partial switch should show noticeable improvements. For benchmarking, we used our in-house [**bfb tool**](https://github.com/qdrant/bfb) to generate a workload. Our configuration: diff --git a/qdrant-landing/content/articles/hybrid-search.md b/qdrant-landing/content/articles/hybrid-search.md index ed5a234c3..b62d7f9c9 100644 --- a/qdrant-landing/content/articles/hybrid-search.md +++ b/qdrant-landing/content/articles/hybrid-search.md @@ -21,7 +21,7 @@ required to combine the results from different methods on your end. **Qdrant 1.10 introduces a new Query API that lets you build a search system by combining different search methods to improve retrieval quality**. Everything is now done on the server side, and you can focus on building the best search experience for your users. In this article, we will show you how to utilize the new [Query -API](/documentation/concepts/search/#query-api) to build a hybrid search system. +API](/documentation/search/search/#query-api) to build a hybrid search system. ## Introducing the new Query API diff --git a/qdrant-landing/content/articles/immutable-data-structures.md b/qdrant-landing/content/articles/immutable-data-structures.md index e5b86d272..9dda36ffb 100644 --- a/qdrant-landing/content/articles/immutable-data-structures.md +++ b/qdrant-landing/content/articles/immutable-data-structures.md @@ -169,12 +169,12 @@ If we knew some additional information about the data, we could combine all rele {{< figure src="/articles_data/immutable-data-structures/defragmentation.png" alt="Defragmentation" caption="Defragmentation" width="70%" >}} -This additional information is available to Qdrant via the [payload index](/documentation/concepts/indexing/#payload-index). +This additional information is available to Qdrant via the [payload index](/documentation/manage-data/indexing/#payload-index). By specifying the payload index, which is going to be used for filtering most of the time, we can put all vectors with the same payload together. This way, reading a single page will also read nearby vectors, which will be used in the search. -This approach is especially efficient for [multi-tenant systems](/documentation/guides/multiple-partitions/), where only a small subset of vectors is actively used for search. +This approach is especially efficient for [multi-tenant systems](/documentation/manage-data/multitenancy/), where only a small subset of vectors is actively used for search. The capacity of such a deployment is typically defined by the size of the hot subset, which is much smaller than the total number of vectors. > Grouping relevant vectors together allows us to optimize the size of the hot subset by avoiding caching of irrelevant data. @@ -200,7 +200,7 @@ All benchmarks are made with minimal RAM allocation to demonstrate disk cache ef As you can see, the biggest impact is on the small tenant size, where defragmentation allows us to achieve **100x more RPS**. Of course, the real-world impact of defragmentation depends on the specific workload and the size of the hot subset, but enabling this feature can significantly improve the performance of Qdrant. -Please find more details on how to enable defragmentation in the [indexing documentation](/documentation/concepts/indexing/#tenant-index). +Please find more details on how to enable defragmentation in the [indexing documentation](/documentation/manage-data/indexing/#tenant-index). ## Updating Immutable Data Structures diff --git a/qdrant-landing/content/articles/miniCOIL.md b/qdrant-landing/content/articles/miniCOIL.md index d62fde1bb..b6eabf722 100644 --- a/qdrant-landing/content/articles/miniCOIL.md +++ b/qdrant-landing/content/articles/miniCOIL.md @@ -227,7 +227,7 @@ Here are the specific characteristics of the miniCOIL model we trained based on | **Input Dense Encoder** | [`jina-embeddings-v2-small-en`](https://huggingface.co/jinaai/jina-embeddings-v2-small-en) (512 dimensions) | | **miniCOIL Vectors Size** | 4 dimensions | | **miniCOIL Vocabulary** | List of 30,000 of the most common English words, cleaned of stop words and words shorter than 3 letters, [taken from here](https://github.com/arstgit/high-frequency-vocabulary/tree/master). Words are stemmed to align miniCOIL with our BM25 implementation. | -| **Training Data** | 40 million sentences — a random subset of the [OpenWebText dataset](https://paperswithcode.com/dataset/openwebtext). To make triplet sampling convenient, we uploaded sentences and their [`mxbai-embed-large-v1`](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) embeddings to Qdrant and built a [full-text payload index](https://qdrant.tech/documentation/concepts/indexing/#full-text-index) on sentences with a tokenizer of type `word`. | +| **Training Data** | 40 million sentences — a random subset of the [OpenWebText dataset](https://paperswithcode.com/dataset/openwebtext). To make triplet sampling convenient, we uploaded sentences and their [`mxbai-embed-large-v1`](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) embeddings to Qdrant and built a [full-text payload index](https://qdrant.tech/documentation/manage-data/indexing/#full-text-index) on sentences with a tokenizer of type `word`. | | **Training Data per Word** | We sample 8000 sentences per word and form triplets with a margin of at least **0.1**.
Additionally, we apply **augmentation** — take a sentence and cut out the target word plus its 1–3 neighbours. We reuse the same similarity score between original and augmented sentences for simplicity. | | **Training Parameters** | **Epochs**: 60
**Optimizer**: Adam with a learning rate of 1e-4
**Validation set**: 20% | @@ -237,7 +237,7 @@ We included this `minicoil-v1` version in the [v0.7.0 release of our FastEmbed l You can check an example of `minicoil-v1` usage with FastEmbed in the [HuggingFace card](https://huggingface.co/Qdrant/minicoil-v1). ## Results diff --git a/qdrant-landing/content/articles/modern-sparse-neural-retrieval.md b/qdrant-landing/content/articles/modern-sparse-neural-retrieval.md index 02bd6a1ee..76266e647 100644 --- a/qdrant-landing/content/articles/modern-sparse-neural-retrieval.md +++ b/qdrant-landing/content/articles/modern-sparse-neural-retrieval.md @@ -234,7 +234,7 @@ internal document expansion idea, which made the retrieval quality noticeably be - The SPARTA model is not sparse enough by construction, so authors of the SPLADE family of models introduced explicit **sparsity regularisation**, preventing the model from producing too many non-zero values. -- The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specifity of Information Retrieval problem, +- The SPARTA model mostly uses the BERT model as-is, without any additional neural network to capture the specificity of Information Retrieval problem, so SPLADE models introduce a trainable neural network on top of BERT with a specific architecture choice to make it perfectly fit the task. - SPLADE family of models, finally, uses **knowledge distillation**, which is learning from a bigger (and therefore much slower, not-so-fit for production tasks) model how to predict good representations. @@ -397,8 +397,8 @@ from qdrant_client import QdrantClient, models qdrant_client = QdrantClient(":memory:") # Qdrant is running from RAM. ``` -Now, let's create a [collection](https://qdrant.tech/documentation/concepts/collections/) in which could upload our sparse SPLADE++ embeddings. \ -For that, we will use the [sparse vectors](https://qdrant.tech/documentation/concepts/vectors/#sparse-vectors) representation supported in Qdrant. +Now, let's create a [collection](https://qdrant.tech/documentation/manage-data/collections/) in which could upload our sparse SPLADE++ embeddings. \ +For that, we will use the [sparse vectors](https://qdrant.tech/documentation/manage-data/vectors/#sparse-vectors) representation supported in Qdrant. ```python qdrant_client.create_collection( diff --git a/qdrant-landing/content/articles/multitenancy.md b/qdrant-landing/content/articles/multitenancy.md index d84136634..ce509f721 100644 --- a/qdrant-landing/content/articles/multitenancy.md +++ b/qdrant-landing/content/articles/multitenancy.md @@ -19,14 +19,14 @@ category: vector-search-manuals # Scaling Your Machine Learning Setup: The Power of Multitenancy and Custom Sharding in Qdrant -We are seeing the topics of [multitenancy](/documentation/guides/multiple-partitions/) and [distributed deployment](/documentation/guides/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup. +We are seeing the topics of [multitenancy](/documentation/manage-data/multitenancy/) and [distributed deployment](/documentation/operations/distributed_deployment/#sharding) pop-up daily on our [Discord support channel](https://qdrant.to/discord). This tells us that many of you are looking to scale Qdrant along with the rest of your machine learning setup. Whether you are building a bank fraud-detection system, [RAG](https://qdrant.tech/articles/what-is-rag-in-ai/) for e-commerce, or services for the federal government - you will need to leverage a multitenant architecture to scale your product. In the world of SaaS and enterprise apps, this setup is the norm. It will considerably increase your application's performance and lower your hosting costs. ## Multitenancy & custom sharding with Qdrant -We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/guides/multiple-partitions/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/guides/distributed_deployment/#user-defined-sharding). +We have developed two major features just for this. __You can now scale a single Qdrant cluster and support all of your customers worldwide.__ Under [multitenancy](/documentation/manage-data/multitenancy/), each customer's data is completely isolated and only accessible by them. At times, if this data is location-sensitive, Qdrant also gives you the option to divide your cluster by region or other criteria that further secure your customer's access. This is called [custom sharding](/documentation/operations/distributed_deployment/#user-defined-sharding). Combining these two will result in an efficiently-partitioned architecture that further leverages the convenience of a single Qdrant cluster. This article will briefly explain the benefits and show how you can get started using both features. @@ -41,7 +41,7 @@ Qdrant is built to excel in a single collection with a vast number of tenants. Y ## Sharding your database -With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/guides/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node. +With Qdrant, you can also specify a shard for each vector individually. This feature is useful if you want to [control where your data is kept in the cluster](/documentation/operations/distributed_deployment/#sharding). For example, one set of vectors can be assigned to one shard on its own node, while another set can be on a completely different node. During vector search, your operations will be able to hit only the subset of shards they actually need. In massive-scale deployments, __this can significantly improve the performance of operations that do not require the whole collection to be scanned__. @@ -49,7 +49,7 @@ This works in the other direction as well. Whenever you search for something, yo ### Common use cases -A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/guides/distributed_deployment/#moving-shards). +A clear use-case for this feature is managing a multitenant collection, where each tenant (let it be a user or organization) is assumed to be segregated, so they can have their data stored in separate shards. Sharding solves the problem of region-based data placement, whereby certain data needs to be kept within specific locations. To do this, however, you will need to [move your shards between nodes](/documentation/operations/distributed_deployment/#moving-shards). **Figure 2:** Users can both upsert and query shards that are relevant to them, all within the same collection. Regional sharding can help avoid cross-continental traffic. ![Qdrant Multitenancy](/articles_data/multitenancy/shards.png) @@ -79,7 +79,7 @@ client.create_shard_key("{tenant_data}", "germany") ``` In this example, your cluster is divided between Germany and Canada. Canadian and German law differ when it comes to international data transfer. Let's say you are creating a RAG application that supports the healthcare industry. Your Canadian customer data will have to be clearly separated for compliance purposes from your German customer. -Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/guides/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech). +Even though it is part of the same collection, data from each shard is isolated from other shards and can be retrieved as such. For additional examples on shards and retrieval, consult [Distributed Deployments](/documentation/operations/distributed_deployment/) documentation and [Qdrant Client specification](https://python-client.qdrant.tech). ## Configure a multitenant setup for users @@ -182,7 +182,7 @@ client.create_payload_index( ## Explore multitenancy and custom sharding in Qdrant for scalable solutions -Qdrant is ready to support a massive-scale architecture for your machine learning project. If you want to see whether our [vector database](https://qdrant.tech/) is right for you, try the [quickstart tutorial](/documentation/quick-start/) or read our [docs and tutorials](/documentation/). +Qdrant is ready to support a massive-scale architecture for your machine learning project. If you want to see whether our [vector database](https://qdrant.tech/) is right for you, try the [quickstart tutorial](/documentation/quickstart/) or read our [docs and tutorials](/documentation/). To spin up a free instance of Qdrant, sign up for [Qdrant Cloud](https://qdrant.to/cloud) - no strings attached. diff --git a/qdrant-landing/content/articles/muvera-embeddings.md b/qdrant-landing/content/articles/muvera-embeddings.md index f26839c9b..c33f16077 100644 --- a/qdrant-landing/content/articles/muvera-embeddings.md +++ b/qdrant-landing/content/articles/muvera-embeddings.md @@ -14,7 +14,7 @@ category: vector-search-manuals Multi-vector representations are superior to single-vector embeddings in many benchmarks. It might be tempting to use them right away, but there is a catch: they are slower to search. Traditional vector search structures like -[HNSW](/documentation/concepts/indexing/#vector-index) are optimized for retrieving the nearest neighbors of a single +[HNSW](/documentation/manage-data/indexing/#vector-index) are optimized for retrieving the nearest neighbors of a single query vector using simple metrics such as cosine similarity. These indexes are not suitable for multi-vector retrieval strategies, such as MaxSim, where a query and document are each represented by multiple vectors and the final score is computed as the maximum similarity over all cross-pairings. MaxSim is inherently asymmetric and non-metric, so HNSW @@ -66,7 +66,7 @@ done by computing the dot product of the input vector with each hyperplane norma result. Since each of our regions can be represented as a binary string of length `k_sim` (where each bit indicates which side of a hyperplane the vector is on), we can interpret this binary string as an integer to get a cluster ID. -![SimHash cluster assignement](/articles_data/muvera-embeddings/simhash-cluster-assignment.png) +![SimHash cluster assignment](/articles_data/muvera-embeddings/simhash-cluster-assignment.png) ### Fixed Dimensional Encoding (FDE) creation diff --git a/qdrant-landing/content/articles/new-recommendation-api.md b/qdrant-landing/content/articles/new-recommendation-api.md index 29b3afdc8..d6bf2270b 100644 --- a/qdrant-landing/content/articles/new-recommendation-api.md +++ b/qdrant-landing/content/articles/new-recommendation-api.md @@ -17,15 +17,15 @@ does exist, and recommendation systems are a great example. Recommendations migh to find items close to positive and far from negative examples. This use of vector databases has many applications, including recommendation systems for e-commerce, content, or even dating apps. -Qdrant has provided the [Recommendation API](/documentation/concepts/search/#recommendation-api) for a while, and with the latest release, [Qdrant 1.6](https://github.com/qdrant/qdrant/releases/tag/v1.6.0), +Qdrant has provided the [Recommendation API](/documentation/search/search/#recommendation-api) for a while, and with the latest release, [Qdrant 1.6](https://github.com/qdrant/qdrant/releases/tag/v1.6.0), we're glad to give you more flexibility and control over the Recommendation API. Here, we'll discuss some internals and show how they may be used in practice. ### Recap of the old recommendations API -The previous [Recommendation API](/documentation/concepts/search/#recommendation-api) in Qdrant came with some limitations. First of all, it was required to pass vector IDs for +The previous [Recommendation API](/documentation/search/search/#recommendation-api) in Qdrant came with some limitations. First of all, it was required to pass vector IDs for both positive and negative example points. If you wanted to use vector embeddings directly, you had to either create a new point -in a collection or mimic the behaviour of the Recommendation API by using the [Search API](/documentation/concepts/search/#search-api). +in a collection or mimic the behaviour of the Recommendation API by using the [Search API](/documentation/search/search/#search-api). Moreover, in the previous releases of Qdrant, you were always asked to provide at least one positive example. This requirement was based on the algorithm used to combine multiple samples into a single query vector. It was a simple, yet effective approach. However, if the only information you had was that your user dislikes some items, you couldn't use it directly. @@ -149,7 +149,7 @@ If you want to know more about the internals of HNSW, you can check out the arti ## Food Discovery demo -Our [Food Discovery demo](/articles/food-discovery-demo/) is an application built on top of the new [Recommendation API](/documentation/concepts/search/#recommendation-api). +Our [Food Discovery demo](/articles/food-discovery-demo/) is an application built on top of the new [Recommendation API](/documentation/search/search/#recommendation-api). It allows you to find a meal based on liked and disliked photos. There are some updates, enabled by the new Qdrant release: * **Ability to include multiple textual queries in the recommendation request.** Previously, we only allowed passing a single diff --git a/qdrant-landing/content/articles/product-quantization.md b/qdrant-landing/content/articles/product-quantization.md index d1315fbb5..3fd45d341 100644 --- a/qdrant-landing/content/articles/product-quantization.md +++ b/qdrant-landing/content/articles/product-quantization.md @@ -224,7 +224,7 @@ In circumstances that do not align with the above, Scalar Quantization should be ## Using Qdrant for Product Quantization -If you’re already a Qdrant user, we have, documentation on [Product Quantization](/documentation/guides/quantization/#setting-up-product-quantization) that will help you to set and configure the new quantization for your data and achieve even +If you’re already a Qdrant user, we have, documentation on [Product Quantization](/documentation/manage-data/quantization/#setting-up-product-quantization) that will help you to set and configure the new quantization for your data and achieve even up to 64x memory reduction. Ready to experience the power of Product Quantization? [Sign up now](https://cloud.qdrant.io/signup) for a free Qdrant demo and optimize your data management today! \ No newline at end of file diff --git a/qdrant-landing/content/articles/qdrant-0-10-release.md b/qdrant-landing/content/articles/qdrant-0-10-release.md index ae65636b3..09ee523ea 100644 --- a/qdrant-landing/content/articles/qdrant-0-10-release.md +++ b/qdrant-landing/content/articles/qdrant-0-10-release.md @@ -27,7 +27,7 @@ set up your collections. Previously, you had to send multiple requests to the Qdrant API to perform multiple non-related tasks. However, this can cause significant network overhead and slow down the process, especially if you have a poor connection speed. -Fortunately, the [new batch search feature](/documentation/concepts/search/#batch-search-api) allows +Fortunately, the [new batch search feature](/documentation/search/search/#batch-search-api) allows you to avoid this issue. With just one API call, Qdrant will handle multiple search requests in the most efficient way possible. This means that you can perform multiple tasks simultaneously without having to worry about network overhead or slow performance. @@ -44,6 +44,6 @@ both ARM and non-ARM architectures using similar setups to understand the potent Qdrant is a vector database that allows you to quickly search for the nearest neighbors. However, you may need to apply additional filters on top of the semantic search. Up until version 0.10, Qdrant only supported keyword filters. With the -release of Qdrant 0.10, [you can now use full-text filters](/documentation/concepts/filtering/#full-text-match) +release of Qdrant 0.10, [you can now use full-text filters](/documentation/search/filtering/#full-text-match) as well. This new filter type can be used on its own or in combination with other filter types to provide even more flexibility in your searches. diff --git a/qdrant-landing/content/articles/qdrant-1.2.x.md b/qdrant-landing/content/articles/qdrant-1.2.x.md index 27663f635..ee14c353b 100644 --- a/qdrant-landing/content/articles/qdrant-1.2.x.md +++ b/qdrant-landing/content/articles/qdrant-1.2.x.md @@ -39,7 +39,7 @@ why we have introduced the [Scalar Quantization](/articles/scalar-quantization/) which makes it possible to reduce the memory requirements by up to four times. Today, we are bringing a new quantization mechanism to life. A separate article on [Product -Quantization](/documentation/quantization/#product-quantization) will describe that feature in more +Quantization](/documentation/manage-data/quantization/#product-quantization) will describe that feature in more detail. In a nutshell, you can **reduce the memory requirements by up to 64 times**! ### Optional named vectors @@ -83,7 +83,7 @@ returned. Unlike some other vector databases, Qdrant accepts any arbitrary JSON payload, including arrays, objects, and arrays of objects. You can also [filter the search results using nested -keys](/documentation/filtering/#nested-key), even though arrays (using the `[]` syntax). +keys](/documentation/search/filtering/#nested-key), even though arrays (using the `[]` syntax). Before Qdrant 1.2 it was impossible to express some more complex conditions for the nested structures. For example, let's assume we have the following payload: @@ -195,18 +195,18 @@ Out-of-Memory errors. Qdrant 1.2 enters recovery mode, if enabled, when it detects a failure on startup. That makes the service halt the loading of collection data and commence operations in a partial state. This state allows for removing collections but doesn't support search or update functions. -**Recovery mode [has to be enabled by user](/documentation/administration/#recovery-mode).** +**Recovery mode [has to be enabled by user](/documentation/operations/administration/#recovery-mode).** ### Appendable mmap For a long time, segments using mmap storage were `non-appendable` and could only be constructed by the optimizer. Dynamically adding vectors to the mmap file is fairly complicated and thus not implemented in Qdrant, but we did our best to implement it in the recent release. If you want -to read more about segments, check out our docs on [vector storage](/documentation/storage/#vector-storage). +to read more about segments, check out our docs on [vector storage](/documentation/manage-data/storage/#vector-storage). ## Security -There are two major changes in terms of [security](/documentation/security/): +There are two major changes in terms of [security](/documentation/operations/security/): 1. **API-key support** - basic authentication with a static API key to prevent unwanted access. Previously API keys were only supported in [Qdrant Cloud](https://cloud.qdrant.io/). diff --git a/qdrant-landing/content/articles/qdrant-1.3.x.md b/qdrant-landing/content/articles/qdrant-1.3.x.md index bc770bc90..502916d04 100644 --- a/qdrant-landing/content/articles/qdrant-1.3.x.md +++ b/qdrant-landing/content/articles/qdrant-1.3.x.md @@ -33,7 +33,7 @@ Your feedback is valuable to us, and are always tying to include some of your fe ## New features -### Asychronous I/O interface +### Asynchronous I/O interface Going forward, we will support the `io_uring` asychnronous interface for storage devices on Linux-based systems. Since its introduction, `io_uring` has been proven to speed up slow-disk deployments as it decouples kernel work from the IO process. @@ -59,7 +59,7 @@ Please keep in mind that this feature is experimental and that the interface may ### Oversampling for quantization -We are introducing [oversampling](/documentation/guides/quantization/#oversampling) as a new way to help you improve the accuracy and performance of similarity search algorithms. With this method, you are able to significantly compress high-dimensional vectors in memory and then compensate the accuracy loss by re-scoring additional points with the original vectors. +We are introducing [oversampling](/documentation/manage-data/quantization/#oversampling) as a new way to help you improve the accuracy and performance of similarity search algorithms. With this method, you are able to significantly compress high-dimensional vectors in memory and then compensate the accuracy loss by re-scoring additional points with the original vectors. You will experience much faster performance with quantization due to parallel disk usage when reading vectors. Much better IO means that you can keep quantized vectors in RAM, so the pre-selection will be even faster. Finally, once pre-selection is done, you can use parallel IO to retrieve original vectors, which is significantly faster than traversing HNSW on slow disks. @@ -180,7 +180,7 @@ client.search_groups( We are excited to announce a more user-friendly way to organize and work with your collections inside of Qdrant. Our dashboard's design is simple, but very intuitive and easy to access. -Try it out now! If you have Docker running, you can [quickstart Qdrant](/documentation/quick-start/) and access the Dashboard locally from [http://localhost:6333/dashboard](http://localhost:6333/dashboard). You should see this simple access point to Qdrant: +Try it out now! If you have Docker running, you can [quickstart Qdrant](/documentation/quickstart/) and access the Dashboard locally from [http://localhost:6333/dashboard](http://localhost:6333/dashboard). You should see this simple access point to Qdrant: ![Qdrant Web UI](/articles_data/qdrant-1.3.x/web-ui.png) @@ -201,7 +201,7 @@ Internally, `is_empty` was not using the index when it was called, so it had to ### Faster read access with mmap -If you used mmap, you most likely found that segments were always created with cold caches. The first request to the database needed to request the disk, which made startup slower despite plenty of RAM being available. We have implemeneted a way to ask the kernel to "heat up" the disk cache and make initialization much faster. +If you used mmap, you most likely found that segments were always created with cold caches. The first request to the database needed to request the disk, which made startup slower despite plenty of RAM being available. We have implemented a way to ask the kernel to "heat up" the disk cache and make initialization much faster. The function is expected to be used on startup and after segment optimization and reloading of newly indexed segment. So far this is only implemented for "immutable" memmaps. diff --git a/qdrant-landing/content/articles/qdrant-1.7.x.md b/qdrant-landing/content/articles/qdrant-1.7.x.md index 3a1efaa93..fe9d611d9 100644 --- a/qdrant-landing/content/articles/qdrant-1.7.x.md +++ b/qdrant-landing/content/articles/qdrant-1.7.x.md @@ -51,11 +51,11 @@ Things have changed since then, as so many of you wanted a single tool for spars If you're coming across the topic of sparse vectors for the first time, our [Brief History of Search](/documentation/overview/vector-search/) explains the difference between sparse and dense vectors. -Check out the [sparse vectors article](/articles/sparse-vectors/) and [sparse vectors index docs](/documentation/concepts/indexing/#sparse-vector-index) for more details on what this new index means for Qdrant users. +Check out the [sparse vectors article](/articles/sparse-vectors/) and [sparse vectors index docs](/documentation/manage-data/indexing/#sparse-vector-index) for more details on what this new index means for Qdrant users. ### Discovery API -The recently launched [Discovery API](/documentation/concepts/explore/#discovery-api) extends the range of scenarios for leveraging vectors. While its interface mirrors the [Recommendation API](/documentation/concepts/explore/#recommendation-api), it focuses on refining the search parameters for greater precision. +The recently launched [Discovery API](/documentation/search/explore/#discovery-api) extends the range of scenarios for leveraging vectors. While its interface mirrors the [Recommendation API](/documentation/search/explore/#recommendation-api), it focuses on refining the search parameters for greater precision. The concept of 'context' refers to a collection of positive-negative pairs that define zones within a space. Each pair effectively divides the space into positive or negative segments. This concept guides the search operation to prioritize points based on their inclusion within positive zones or their avoidance of negative zones. Essentially, the search algorithm favors points that fall within multiple positive zones or steer clear of negative ones. The Discovery API can be used in two ways - either with or without the target point. The first case is called a **discovery search**, while the second is called a **context search**. @@ -66,7 +66,7 @@ The Discovery API can be used in two ways - either with or without the target po ![Discovery search visualization](/articles_data/qdrant-1.7.x/discovery-search.png) -Please refer to the [Discovery API documentation on discovery search](/documentation/concepts/explore/#discovery-search) for more details and the internal mechanics of the operation. +Please refer to the [Discovery API documentation on discovery search](/documentation/search/explore/#discovery-search) for more details and the internal mechanics of the operation. #### Context search @@ -92,7 +92,7 @@ POST /collections/my_collection/points/search } ``` -If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/guides/distributed_deployment/#sharding). +If you want to know more about the user-defined sharding, please refer to the [sharding documentation](/documentation/operations/distributed_deployment/#sharding). ### Snapshot-based shard transfer @@ -101,7 +101,7 @@ That's a really more in depth technical improvement for the distributed mode use Moving shards is required for dynamical scaling of the cluster. Your data can migrate between nodes, and the way you move it is crucial for the performance of the whole system. The good old `stream_records` method (still the default one) transmits all the records between the machines and indexes them on the target node. In the case of moving the shard, it's necessary to recreate the HNSW index each time. However, with the introduction of the new `snapshot` approach, the snapshot itself, inclusive of all data and potentially quantized content, is transferred to the target node. This comprehensive snapshot includes the entire index, enabling the target node to seamlessly load it and promptly begin handling requests without the need for index recreation. -There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/guides/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future. +There are multiple scenarios in which you may prefer one over the other. Please check out the docs of the [shard transfer method](/documentation/operations/distributed_deployment/#shard-transfer-method) for more details and head-to-head comparison. As for now, the old `stream_records` method is still the default one, but we may decide to change it in the future. ## Minor improvements diff --git a/qdrant-landing/content/articles/qdrant-1.8.x.md b/qdrant-landing/content/articles/qdrant-1.8.x.md index 67a1f8279..baa56b644 100644 --- a/qdrant-landing/content/articles/qdrant-1.8.x.md +++ b/qdrant-landing/content/articles/qdrant-1.8.x.md @@ -27,7 +27,7 @@ This time around, we have focused on Qdrant's internals. Our goal was to optimiz - **Faster [sparse vectors](https://qdrant.tech/articles/sparse-vectors/):** [Hybrid search](https://qdrant.tech/articles/hybrid-search/) is up to 16x faster now! - **CPU resource management:** You can allocate CPU threads for faster indexing. -- **Better indexing performance:** We optimized text [indexing](https://qdrant.tech/documentation/concepts/indexing/) on the backend. +- **Better indexing performance:** We optimized text [indexing](https://qdrant.tech/documentation/manage-data/indexing/) on the backend. ## Faster search with sparse vectors @@ -55,7 +55,7 @@ The colors within both scatter plots show the frequency of results. The red dots This performance increase can have a dramatic effect on hybrid search implementations. [Read more about how to set this up.](/articles/sparse-vectors/) -FYI, sparse vectors were released in [Qdrant v.1.7.0](/articles/qdrant-1.7.x/#sparse-vectors). They are stored using a different index, so first [check out the documentation](/documentation/concepts/indexing/#sparse-vector-index) if you want to try an implementation. +FYI, sparse vectors were released in [Qdrant v.1.7.0](/articles/qdrant-1.7.x/#sparse-vectors). They are stored using a different index, so first [check out the documentation](/documentation/manage-data/indexing/#sparse-vector-index) if you want to try an implementation. ## CPU resource management @@ -65,7 +65,7 @@ This isn't mandatory, as Qdrant is by default tuned to strike the right balance This version introduces a `optimizer_cpu_budget` parameter to control the maximum number of CPUs used for indexing. -> Read more about `config.yaml` in the [configuration file](/documentation/guides/configuration/). +> Read more about `config.yaml` in the [configuration file](/documentation/operations/configuration/). ```yaml # CPU budget, how many CPUs (threads) to allocate for an optimization job. @@ -99,8 +99,8 @@ This approach ensures stability in the [vector search](https://qdrant.tech/docum Beyond these enhancements, [Qdrant v1.8.0](https://github.com/qdrant/qdrant/releases/tag/v1.8.0) adds and improves on several smaller features: -1. **Order points by payload:** In addition to searching for semantic results, you might want to retrieve results by specific metadata (such as price). You can now use Scroll API to [order points by payload key](/documentation/concepts/points/#order-points-by-payload-key). -2. **Datetime support:** We have implemented [datetime support for the payload index](/documentation/concepts/filtering/#datetime-range). Prior to this, if you wanted to search for a specific datetime range, you would have had to convert dates to UNIX timestamps. ([PR#3320](https://github.com/qdrant/qdrant/issues/3320)) +1. **Order points by payload:** In addition to searching for semantic results, you might want to retrieve results by specific metadata (such as price). You can now use Scroll API to [order points by payload key](/documentation/manage-data/points/#order-points-by-payload-key). +2. **Datetime support:** We have implemented [datetime support for the payload index](/documentation/search/filtering/#datetime-range). Prior to this, if you wanted to search for a specific datetime range, you would have had to convert dates to UNIX timestamps. ([PR#3320](https://github.com/qdrant/qdrant/issues/3320)) 3. **Check collection existence:** You can check whether a collection exists via the `/exists` endpoint to the `/collections/{collection_name}`. You will get a true/false response. ([PR#3472](https://github.com/qdrant/qdrant/pull/3472)). 4. **Find points** whose payloads match more than the minimal amount of conditions. We included the `min_should` match feature for a condition to be `true` ([PR#3331](https://github.com/qdrant/qdrant/pull/3466/)). 5. **Modify nested fields:** We have improved the `set_payload` API, adding the ability to update nested fields ([PR#3548](https://github.com/qdrant/qdrant/pull/3548)). diff --git a/qdrant-landing/content/articles/rapid-rag-optimization-with-qdrant-and-quotient.md b/qdrant-landing/content/articles/rapid-rag-optimization-with-qdrant-and-quotient.md index 100480d27..35977160e 100755 --- a/qdrant-landing/content/articles/rapid-rag-optimization-with-qdrant-and-quotient.md +++ b/qdrant-landing/content/articles/rapid-rag-optimization-with-qdrant-and-quotient.md @@ -467,7 +467,7 @@ We will reprocess the data with the updated parameters above: ```python ## for iteration 2 - lets modify chunk configuration -## We will start with creating seperate collection to store vectors +## We will start with creating separate collection to store vectors chunk_size = 1024 chunk_overlap = 128 diff --git a/qdrant-landing/content/articles/relevance-feedback.md b/qdrant-landing/content/articles/relevance-feedback.md index d28e3c295..45c9c079f 100644 --- a/qdrant-landing/content/articles/relevance-feedback.md +++ b/qdrant-landing/content/articles/relevance-feedback.md @@ -408,11 +408,11 @@ Out of all retriever–feedback model pairs, three leaders emerged with the foll ## Relevance Feedback Query -The results were convincing enough to justify implementing Relevance Feedback Query. And [here it is](https://qdrant.tech/documentation/concepts/search-relevance/#relevance-feedback), ready for your retrieval pipelines! +The results were convincing enough to justify implementing Relevance Feedback Query. And [here it is](https://qdrant.tech/documentation/search/search-relevance/#relevance-feedback), ready for your retrieval pipelines! ### When to Use It -Use Relevance Feedback Query once a basic retrieval pipeline is in place and you're looking for additional techniques to boost result relevance, such as [Maximal Marginal Relevance (MMR)](https://qdrant.tech/documentation/concepts/search-relevance/#maximal-marginal-relevance-mmr), [Reranking](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries), or [Score Boosting](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting). +Use Relevance Feedback Query once a basic retrieval pipeline is in place and you're looking for additional techniques to boost result relevance, such as [Maximal Marginal Relevance (MMR)](https://qdrant.tech/documentation/search/search-relevance/#maximal-marginal-relevance-mmr), [Reranking](https://qdrant.tech/documentation/search/hybrid-queries/#multi-stage-queries), or [Score Boosting](https://qdrant.tech/documentation/search/search-relevance/#score-boosting). **It's here not to replace but to complement other search relevance tools.** For example, Relevance Feedback Query can be a great aid for search agents, letting you propagate the agent's understanding of your use case directly to the vector search index. @@ -434,7 +434,7 @@ If no training queries are supplied, the package will train the formula directly > **Warning:** If your use case doesn't involve document-to-document semantic similarity search, training on sampled documents alone may completely cancel the effect of relevance feedback scoring on real data. > It's far more effective to use real queries. -Once you've obtained the weights, simply plug them into your [Qdrant Client of choice](https://qdrant.tech/documentation/concepts/search-relevance/#relevance-feedback). +Once you've obtained the weights, simply plug them into your [Qdrant Client of choice](https://qdrant.tech/documentation/search/search-relevance/#relevance-feedback). ### Evaluating Your Gains diff --git a/qdrant-landing/content/articles/scalar-quantization.md b/qdrant-landing/content/articles/scalar-quantization.md index 9c8d7f327..d0cde7e17 100644 --- a/qdrant-landing/content/articles/scalar-quantization.md +++ b/qdrant-landing/content/articles/scalar-quantization.md @@ -290,6 +290,6 @@ expensive setup if you can agree to a small decrease in the search precision. ### Accessing best practices -Qdrant documentation on [Scalar Quantization](/documentation/quantization/#setting-up-quantization-in-qdrant) +Qdrant documentation on [Scalar Quantization](/documentation/manage-data/quantization/#setting-up-quantization-in-qdrant) is a great resource describing different scenarios and strategies to achieve up to 4x lower memory footprint and even up to 2x performance increase. diff --git a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md index 4338fd66d..45fb0beac 100644 --- a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md +++ b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md @@ -118,7 +118,7 @@ sparse_vectors_config={ } ``` -**Hybrid in one request.** Combine sparse precision with dense semantics via native [RRF/prefetch](https://qdrant.tech/documentation/concepts/hybrid-queries/) - no external reranker: +**Hybrid in one request.** Combine sparse precision with dense semantics via native [RRF/prefetch](https://qdrant.tech/documentation/search/hybrid-queries/) - no external reranker: ```python client.query_points( diff --git a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md index d452bb860..904e245fc 100644 --- a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md +++ b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md @@ -100,7 +100,7 @@ This launches a web dashboard with tabs for each stage of the pipeline: **Evaluate.** Point at a trained model and test queries. Get metric cards for nDCG@10, MRR@10, Recall, and Precision. -**[Collections](https://qdrant.tech/documentation/concepts/collections/).** Browse your Qdrant collections, check point counts, and run test searches against indexed products. Useful for sanity-checking that indexing worked before evaluation. +**[Collections](https://qdrant.tech/documentation/manage-data/collections/).** Browse your Qdrant collections, check point counts, and run test searches against indexed products. Useful for sanity-checking that indexing worked before evaluation. **Publish.** Enter a model path and HuggingFace repo name. Click publish. diff --git a/qdrant-landing/content/articles/vector-search-filtering.md b/qdrant-landing/content/articles/vector-search-filtering.md index dbb3e8059..7b4ccc773 100644 --- a/qdrant-landing/content/articles/vector-search-filtering.md +++ b/qdrant-landing/content/articles/vector-search-filtering.md @@ -83,7 +83,7 @@ Most people use default settings and build vector search apps that aren't proper #### Remember to run all tutorial code in Qdrant's Dashboard -The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/quickstart-cloud/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses. +The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/cloud-quickstart/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses. ![qdrant-filtering-tutorial](/articles_data/vector-search-filtering/qdrant-filtering-tutorial.png) @@ -93,13 +93,13 @@ Qdrant follows a specific method of searching and filtering through dense vector Let's take a look at this **3-stage diagram**. In this case, we are trying to find the nearest neighbour to the query vector **(green)**. Your search journey starts at the bottom **(orange)**. -By default, Qdrant connects all your data points within the [**vector index**](/documentation/concepts/indexing/). After you [**introduce filters**](/documentation/concepts/filtering/), some data points become disconnected. Vector search can't cross the grayed out area and it won't reach the nearest neighbor. +By default, Qdrant connects all your data points within the [**vector index**](/documentation/manage-data/indexing/). After you [**introduce filters**](/documentation/search/filtering/), some data points become disconnected. Vector search can't cross the grayed out area and it won't reach the nearest neighbor. How can we bridge this gap? **Figure 1:** How Qdrant maintains a filterable vector index. ![filterable-vector-index](/articles_data/vector-search-filtering/filterable-vector-index.png) -[**Filterable vector index**](/documentation/concepts/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points. +[**Filterable vector index**](/documentation/manage-data/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points. ### Qdrant's approach vs traditional filtering methods @@ -201,7 +201,7 @@ As you can see, Qdrant's filtering method has a greater chance of capturing all This specific example uses the `range` condition for filtering. Qdrant, however, offers many other possible ways to structure a filter -**For detailed usage examples, [filtering](/documentation/concepts/filtering/) docs are the best resource.** +**For detailed usage examples, [filtering](/documentation/search/filtering/) docs are the best resource.** ### Scrolling instead of searching @@ -240,7 +240,7 @@ POST /collections/online_store/points/scroll ``` The response contains a batch of points that match the criteria and a reference (offset or next page token) to retrieve the next set of points. -> [**Scrolling**](/documentation/concepts/points/#scroll-points) is designed to be efficient. It minimizes the load on the server and reduces memory consumption on the client side by returning only manageable chunks of data at a time. +> [**Scrolling**](/documentation/manage-data/points/#scroll-points) is designed to be efficient. It minimizes the load on the server and reduces memory consumption on the client side by returning only manageable chunks of data at a time. #### Available filtering conditions @@ -254,7 +254,7 @@ The response contains a batch of points that match the criteria and a reference | **Full Text Match** | Search in text fields. | **Is Empty** | Filter empty fields. | | **Has ID** | Filter by unique ID. | **Is Null** | Filter null values. | -> All clauses and conditions are outlined in Qdrant's [filtering](/documentation/concepts/filtering/) documentation. +> All clauses and conditions are outlined in Qdrant's [filtering](/documentation/search/filtering/) documentation. #### Filtering clauses to remember @@ -559,10 +559,10 @@ You can use filters to retrieve data points without knowing their `id`. You can | Action | Description | Action | Description | |--------|-------------|--------|-------------| -| [Delete Points](/documentation/concepts/points/#delete-points) | Deletes all points matching the filter. | [Set Payload](/documentation/concepts/payload/#set-payload) | Adds payload fields to all points matching the filter. | -| [Scroll Points](/documentation/concepts/points/#scroll-points) | Lists all points matching the filter. | [Update Payload](/documentation/concepts/payload/#overwrite-payload) | Updates payload fields for points matching the filter. | -| [Order Points](/documentation/concepts/points/#order-points-by-payload-key) | Lists all points, sorted by the filter. | [Delete Payload](/documentation/concepts/payload/#delete-payload-keys) | Deletes fields for points matching the filter. | -| [Count Points](/documentation/concepts/points/#counting-points) | Totals the points matching the filter. | | | +| [Delete Points](/documentation/manage-data/points/#delete-points) | Deletes all points matching the filter. | [Set Payload](/documentation/manage-data/payload/#set-payload) | Adds payload fields to all points matching the filter. | +| [Scroll Points](/documentation/manage-data/points/#scroll-points) | Lists all points matching the filter. | [Update Payload](/documentation/manage-data/payload/#overwrite-payload) | Updates payload fields for points matching the filter. | +| [Order Points](/documentation/manage-data/points/#order-points-by-payload-key) | Lists all points, sorted by the filter. | [Delete Payload](/documentation/manage-data/payload/#delete-payload-keys) | Deletes fields for points matching the filter. | +| [Count Points](/documentation/manage-data/points/#counting-points) | Totals the points matching the filter. | | | ## Filtering with the payload index @@ -577,7 +577,7 @@ Just how the vector index organizes vectors, the payload index will structure yo ![payload-index-vector-search](/articles_data/vector-search-filtering/payload-index-vector-search.png) -On its own, semantic searching over terabytes of data can take up lots of RAM. [**Filtering**](/documentation/concepts/filtering/) and [**Indexing**](/documentation/concepts/indexing/) are two easy strategies to reduce your compute usage and still get the best results. Remember, this is only a guide. For an exhaustive list of filtering options, you should read the [filtering documentation](/documentation/concepts/filtering/). +On its own, semantic searching over terabytes of data can take up lots of RAM. [**Filtering**](/documentation/search/filtering/) and [**Indexing**](/documentation/manage-data/indexing/) are two easy strategies to reduce your compute usage and still get the best results. Remember, this is only a guide. For an exhaustive list of filtering options, you should read the [filtering documentation](/documentation/search/filtering/). Here is how you can create a single index for a metadata field "category": @@ -660,11 +660,11 @@ If your users are often filtering by **laptop** when looking up a product **cate | Index Type | Description | |---------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------| -| [Full-text Index](/documentation/concepts/indexing/#full-text-index) | Enables efficient text search in large datasets. | -| [Tenant Index](/documentation/concepts/indexing/#tenant-index) | For data isolation and retrieval efficiency in multi-tenant architectures. | -| [Principal Index](/documentation/concepts/indexing/#principal-index) | Manages data based on primary entities like users or accounts. | -|[On-Disk Index](/documentation/concepts/indexing/#on-disk-payload-index) | Stores indexes on disk to manage large datasets without memory usage. | -| [Parameterized Index](/documentation/concepts/indexing/#parameterized-index) | Allows for dynamic querying, where the index can adapt based on different parameters or conditions provided by the user. Useful for numeric data like prices or timestamps. | +| [Full-text Index](/documentation/manage-data/indexing/#full-text-index) | Enables efficient text search in large datasets. | +| [Tenant Index](/documentation/manage-data/indexing/#tenant-index) | For data isolation and retrieval efficiency in multi-tenant architectures. | +| [Principal Index](/documentation/manage-data/indexing/#principal-index) | Manages data based on primary entities like users or accounts. | +|[On-Disk Index](/documentation/manage-data/indexing/#on-disk-payload-index) | Stores indexes on disk to manage large datasets without memory usage. | +| [Parameterized Index](/documentation/manage-data/indexing/#parameterized-index) | Allows for dynamic querying, where the index can adapt based on different parameters or conditions provided by the user. Useful for numeric data like prices or timestamps. | ### Indexing payloads in multitenant setups @@ -688,7 +688,7 @@ PUT /collections/{collection_name}/index ``` Additionally, we offer a way of organizing data efficiently by means of the tenant index. This is another variant of the payload index that makes tenant data more accessible. This time, the request will specify the field as a tenant. This means that you can mark various customer types and user id’s as `is_tenant: true`. -Read more about setting up [tenant defragmentation](/documentation/concepts/indexing/?q=tenant#tenant-index) in multitenant environments, +Read more about setting up [tenant defragmentation](/documentation/manage-data/indexing/?q=tenant#tenant-index) in multitenant environments, ## Key takeaways in filtering and indexing ![best-practices](/articles_data/vector-search-filtering/best-practices.png) @@ -744,7 +744,7 @@ As a conclusion to this guide, let's look at some real-life use cases where filt #### Before you go - all the code is in Qdrant's Dashboard -The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/quickstart-cloud/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses. +The easiest way to reach that "Hello World" moment is to [**try filtering in a live cluster**](/documentation/cloud-quickstart/). Our interactive tutorial will show you how to create a cluster, add data and try some filtering clauses. **It's all in your free cluster!** diff --git a/qdrant-landing/content/articles/vector-search-production.md b/qdrant-landing/content/articles/vector-search-production.md index 3480a30d1..3ff10e4d5 100644 --- a/qdrant-landing/content/articles/vector-search-production.md +++ b/qdrant-landing/content/articles/vector-search-production.md @@ -52,17 +52,17 @@ This article will help you successfully deploy and maintain vector search system ### Ensure your hot dataset fits in RAM for low-latency queries. -If not, then you'll have to [**offload data 'on_disk'**](/documentation/concepts/storage/#configuring-memmap-storage). If this parameter is enabled, Qdrant caches your most frequently accessed vectors loaded into RAM, and the rest is memory-mapped onto the disk. +If not, then you'll have to [**offload data 'on_disk'**](/documentation/manage-data/storage/#configuring-memmap-storage). If this parameter is enabled, Qdrant caches your most frequently accessed vectors loaded into RAM, and the rest is memory-mapped onto the disk. This ensures minimal disk access during queries, significantly reducing latency and boosting overall performance. By monitoring query patterns and usage metrics, you can identify which subsets of your data deserve dedicated in-memory storage, reserving disk access only for colder, less frequently queried vectors. || |-| -|**Read More:** [**Storage Documentation**](https://qdrant.tech/documentation/concepts/storage/)| +|**Read More:** [**Storage Documentation**](https://qdrant.tech/documentation/manage-data/storage/)| ### Index Your Important Metadata to Avoid Costly Queries -✅ You should always [**create payload indexes**](https://qdrant.tech/documentation/concepts/indexing/#payload-index) for all fields used in filters or sorting. +✅ You should always [**create payload indexes**](https://qdrant.tech/documentation/manage-data/indexing/#payload-index) for all fields used in filters or sorting. Many users configure complex filters but may not be aware of the need to create corresponding payload indexes. @@ -72,11 +72,11 @@ Filtering after retrieving thousands of vectors can get expensive. If you don't Unlike some other engines, Qdrant lets you make the optimal choice of which fields to index for your use case rather than creating indexes for every field by default. -> **Note:** Don't forget to use the correct [**payload index type**](https://qdrant.tech/documentation/concepts/indexing/#payload-index). If there are numeric values, the you must use a numeric index. If you represent numbers in strings ("123"), a numeric index will not work. +> **Note:** Don't forget to use the correct [**payload index type**](https://qdrant.tech/documentation/manage-data/indexing/#payload-index). If there are numeric values, the you must use a numeric index. If you represent numbers in strings ("123"), a numeric index will not work. || |-| -|**Read More:** [**Filtering Documentation**](https://qdrant.tech/documentation/concepts/filtering/)| +|**Read More:** [**Filtering Documentation**](https://qdrant.tech/documentation/search/filtering/)| ### Don't Forget to Tune HNSW Search Parameters @@ -88,7 +88,7 @@ Sometimes users don't properly balance HNSW search parameters. Setting the HNSW ❓ **Use Case:** A customer ran advanced similarity searches across their vast dataset of nearly 800 million vectors. Initially, they found that queries took anywhere from 10 to 20 seconds, especially when combining multiple filters and metadata fields. -> ✅ How can they retain accuracy and keep things fast? [**The answer is optimization.**](https://qdrant.tech/documentation/guides/optimize/) +> ✅ How can they retain accuracy and keep things fast? [**The answer is optimization.**](https://qdrant.tech/documentation/operations/optimize/) **Figure 1:** Qdrant is highly configurable. You can configure it for speed, precision or resource use. ![qdrant resource tradeoffs](/docs/tradeoff.png) @@ -99,8 +99,8 @@ This strategy balanced memory usage with performance: only the compact vectors n || |-| -|**Read More:** [**Optimization Guide**](https://qdrant.tech/documentation/guides/optimize/)|#optimizing-qdrant-performance-three-scenarios -|**Read More:** [**HNSW Documentation**](https://qdrant.tech/documentation/concepts/indexing/#vector-index)| +|**Read More:** [**Optimization Guide**](https://qdrant.tech/documentation/operations/optimize/)|#optimizing-qdrant-performance-three-scenarios +|**Read More:** [**HNSW Documentation**](https://qdrant.tech/documentation/manage-data/indexing/#vector-index)| ### Compress Your Data with Quantization Strategies @@ -108,21 +108,21 @@ This strategy balanced memory usage with performance: only the compact vectors n |:-:| |**"We're using too much memory for our massive dataset."**| -Many users skip [**quantization**](https://qdrant.tech/documentation/guides/quantization/), causing their index to consume excessive RAM and produce uneven performance. Some users hesitate to compromise precision, but this is not always the case. +Many users skip [**quantization**](https://qdrant.tech/documentation/manage-data/quantization/), causing their index to consume excessive RAM and produce uneven performance. Some users hesitate to compromise precision, but this is not always the case. If your workload can tolerate a moderate drop in embedding precision, data compression offers a powerful way to shrink vector size and slash memory usage. By converting high-dimensional floating-point values into lower-bit formats (such as 8-bit scalar or even a single bit-sized representations), you can keep far more vectors in RAM while reducing disk footprint. -> ✅ [**You should evaluate and apply quantization**](https://qdrant.tech/documentation/guides/quantization/#how-to-choose-the-right-quantization-method) if your use case allows. Quantization seriously improves performance and reduces storage costs. +> ✅ [**You should evaluate and apply quantization**](https://qdrant.tech/documentation/manage-data/quantization/#how-to-choose-the-right-quantization-method) if your use case allows. Quantization seriously improves performance and reduces storage costs. This not only speeds up query throughput for large-scale datasets, but also cuts hardware costs and storage overhead. While Scalar Quantization is a midrange compression alternative, Binary quantization is more drastic, so be sure to test your accuracy requirements for each thoroughly. -When using [**quantization**](https://qdrant.tech/documentation/guides/quantization/), you can store only the compressed vectors in memory while leaving the original floating-point versions on disk for reference. This approach dramatically lowers RAM consumption—since quantized vectors take far less space—yet still allows you to retrieve full-precision vectors if needed for downstream tasks like re-ranking. +When using [**quantization**](https://qdrant.tech/documentation/manage-data/quantization/), you can store only the compressed vectors in memory while leaving the original floating-point versions on disk for reference. This approach dramatically lowers RAM consumption—since quantized vectors take far less space—yet still allows you to retrieve full-precision vectors if needed for downstream tasks like re-ranking. >**Sidenote:** You can always enable `async_io` scorer when the linux kernel supports it and if you have `on_disk` vectors. || |-| -|**Read More:** [**Quantization Documentation**](https://qdrant.tech/documentation/guides/quantization/)| +|**Read More:** [**Quantization Documentation**](https://qdrant.tech/documentation/manage-data/quantization/)| ## 2. How do I Ingest and Index Large Amounts of Data? ![vector-search-production](/articles_data/vector-search-production/vector-search-production-2.jpg) @@ -141,7 +141,7 @@ Once all records are inserted, you can rebuild the index in a single pass. Consi || |-| -|**Read More:** [**Configuring the Vector Index**](https://qdrant.tech/documentation/concepts/indexing/#vector-index)| +|**Read More:** [**Configuring the Vector Index**](https://qdrant.tech/documentation/manage-data/indexing/#vector-index)| ### Other Solutions to Alleviate Indexing Bottleneck @@ -155,7 +155,7 @@ Once all records are inserted, you can rebuild the index in a single pass. Consi || |-| -|**Read More:** [**Configuration Documentation**](https://qdrant.tech/documentation/guides/configuration/)| +|**Read More:** [**Configuration Documentation**](https://qdrant.tech/documentation/operations/configuration/)| ### When Indexing Falls Behind Ingestion ![vector-search-production](/articles_data/vector-search-production/vector-search-production-3.jpg) @@ -166,7 +166,7 @@ By default, searches include unindexed data. However, a large number of unindexe If the maximum number of indexed points remains consistently low, this is likely not an issue. If you anticipate periods with many unindexed points, you should take measures to prevent search disruptions in production. -One option is to [**set `indexed_only=true` in search requests**](https://qdrant.tech/documentation/concepts/search/#search-api). This will ensure fast searches by only considering indexed data, at the expense of eventual consistency (new data becomes searchable only after indexing). +One option is to [**set `indexed_only=true` in search requests**](https://qdrant.tech/documentation/search/search/#search-api). This will ensure fast searches by only considering indexed data, at the expense of eventual consistency (new data becomes searchable only after indexing). Alternatively, you can perform [**bulk vector uploads**](https://qdrant.tech/documentation/database-tutorials/bulk-upload/) during low-traffic periods to allow indexing to complete before increased traffic. @@ -174,7 +174,7 @@ Alternatively, you can perform [**bulk vector uploads**](https://qdrant.tech/doc || |-| -|**Read More:** [**Indexing Documentation**](https://qdrant.tech/documentation/concepts/indexing/)| +|**Read More:** [**Indexing Documentation**](https://qdrant.tech/documentation/manage-data/indexing/)| ### How to Arrange Metadata and Schema for Consistency @@ -184,7 +184,7 @@ Alternatively, you can perform [**bulk vector uploads**](https://qdrant.tech/doc In some cases, the payload schema is inconsistent across data pipelines, so some fields have mismatched types or are missing altogether. -❓ **Use Case:** A healthcare firm discovered that some pipelines inserted strings where others inserted integers. Filters broke silently or returned inconsistent results, signalling that [**a unified payload schema**](https://qdrant.tech/documentation/concepts/indexing/#payload-index) was not in place. +❓ **Use Case:** A healthcare firm discovered that some pipelines inserted strings where others inserted integers. Filters broke silently or returned inconsistent results, signalling that [**a unified payload schema**](https://qdrant.tech/documentation/manage-data/indexing/#payload-index) was not in place. > When payload fields are typed inconsistently across your ingestion pipelines, filters can break in unpredictable ways. @@ -194,13 +194,13 @@ For example, **some services might write a "status" field as a string ("active") || |-| -|**Read More:** [**Payload Documentation**](https://qdrant.tech/documentation/concepts/payload/)| +|**Read More:** [**Payload Documentation**](https://qdrant.tech/documentation/manage-data/payload/)| ### Decide How to Set Up a Multitenant Collection ❓ **Use Case:** When implementing vector databases, healthcare organizations need to ensure isolation between users' data. Our customer needed to make sure that when they filtered queries to only show a particular patient's documents, and no other patient's documents appeared in the query results. -✅ [**You should almost always consolidate tenants to a single collection**](https://qdrant.tech/documentation/guides/multiple-partitions/) if possible, tagging by tenant. +✅ [**You should almost always consolidate tenants to a single collection**](https://qdrant.tech/documentation/manage-data/multitenancy/) if possible, tagging by tenant. ```text PUT /collections/{collection_name}/index @@ -221,7 +221,7 @@ Figure: For many-tenant setups, spinning up a new collection per tenant can ball || |-| -|**Read More:** [**Multitenancy Documentation**](https://qdrant.tech/documentation/concepts/multitenancy/)| +|**Read More:** [**Multitenancy Documentation**](https://qdrant.tech/documentation/manage-data/multitenancy/)| ## 3. What's the Best Way to Scale the Database and Optimize Resources? ![vector-search-production](/articles_data/vector-search-production/vector-search-production-4.jpg) @@ -230,7 +230,7 @@ Figure: For many-tenant setups, spinning up a new collection per tenant can ball |:-:| |**"How many nodes, CPUs, RAM and storage do I need for my Qdrant Cluster?"**| -It depends. If you're just starting out - we have prepared a tool on our website to help you figure this out. For more information, [**check out the Capacity Planning document as well.**](https://qdrant.tech/documentation/guides/capacity-planning/) +It depends. If you're just starting out - we have prepared a tool on our website to help you figure this out. For more information, [**check out the Capacity Planning document as well.**](https://qdrant.tech/documentation/operations/capacity-planning/) ✅ [**Use the sizing calculator**](https://cloud.qdrant.io/calculator) or performance testing to ensure node specs (RAM/CPU) match your workload. @@ -242,7 +242,7 @@ It depends. If you're just starting out - we have prepared a tool on our website A three-node setup provides a baseline for fault tolerance: if one node goes offline, the remaining two can continue serving queries and maintain a quorum for data consistency. This guards against hardware failures, rolling updates, and network disruptions. Fewer than three nodes leaves you vulnerable to single-point failures that can knock your entire cluster offline. -> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/guides/distributed_deployment/#raft), so check out the docs and learn why this is important. +> [**We follow the Raft Protocol**](https://qdrant.tech/documentation/operations/distributed_deployment/#raft), so check out the docs and learn why this is important. ✅ **Set a replication factor of at least 2** to tolerate node failure without losing availability. @@ -274,7 +274,7 @@ Development and staging environments often run experimental builds, tests, or si > It's quite possible that the user has multiple shards on one node, which end up handling most traffic while other nodes remain underutilized. -In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/guides/distributed_deployment/#sharding) based on your node count and expected RPS. +In this case, you should [**choose the right number of shards**](https://qdrant.tech/documentation/operations/distributed_deployment/#sharding) based on your node count and expected RPS. You need to implement a shard strategy that aligns with real usage patterns. First, distribute your shards across all available nodes. This will help balance the load more effectively. After redistributing the shards, run performance tests to see how it affects your system. Then add replicas and test again to see how that changes performance. @@ -285,14 +285,14 @@ Proper sharding considers data distribution and query patterns. By default, shar || |-| -|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/concepts/sharding/)| +|**Read More:** [**Sharding Documentation**](https://qdrant.tech/documentation/operations/distributed_deployment/#sharding)| ### Manage Your Costs by Scaling Up or Down ![vector-search-production](/articles_data/vector-search-production/vector-search-production-5.jpg) Some teams scale up for daytime surges, then scale down overnight to save resources. If you do this, ensure data is sharded and replicated appropriately, so that scaling up and down won't result in service degradation. -If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/guides/distributed_deployment/#replication-factor), though it may be considered a bit of a hack. +If using Qdrant Cloud you could also do this using the [**Replication Factor**](https://qdrant.tech/documentation/operations/distributed_deployment/#replication-factor), though it may be considered a bit of a hack. > If you have 3 nodes with just 1 shard, and replication factor 6. It will create 3 replicas (one on each node) of that shard, because it can't host more. If you add 3 more nodes at peak times, it'll automatically replicate that shard 3 more times in an attempt to match the factor of 6. @@ -308,7 +308,7 @@ If new nodes remain empty after joining, you waste resources. If departing nodes || |-| -|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/guides/distributed_deployment/)| +|**Read More:** [**Distributed Deployment Documentation**](https://qdrant.tech/documentation/operations/distributed_deployment/)| |**Read More:** [**Resharding**](https://qdrant.tech/documentation/cloud/cluster-scaling/#resharding)| ### How to Predict and Test Cluster Performance @@ -331,7 +331,7 @@ Remember, cold-starts and query behaviour are dataset dependent, which is why yo || |-| -|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/guides/distributed_deployment/) +|**Read More:** [Distributed Deployment Documentation](https://qdrant.tech/documentation/operations/distributed_deployment/) ### How to Design Your Systems to Protect Against Failure @@ -375,7 +375,7 @@ By following these comprehensive load testing practices, you'll be able to ident || |-| -|**Read More:** [**Telemetry and Monitoring Documentation**](https://qdrant.tech/documentation/guides/monitoring/)| +|**Read More:** [**Telemetry and Monitoring Documentation**](https://qdrant.tech/documentation/operations/monitoring/)| |**Read More:** [**Cloud Monitoring Documentation**](https://qdrant.tech/documentation/hybrid-cloud/networking-logging-monitoring/) ## 4. Ensuring Disaster Recovery With Database Backups and Snapshots @@ -413,7 +413,7 @@ If you host tens of billions of vectors, store backups off-node in a different d || |-| -|**Read More:** [**Snapshot Documentation**](https://qdrant.tech/documentation/concepts/snapshots/)| +|**Read More:** [**Snapshot Documentation**](https://qdrant.tech/documentation/operations/snapshots/)| |**Read More:** [**Managed Cloud Backup Documentation**](https://qdrant.tech/documentation/cloud/backups/)| |**Read More:** [**Private Cloud Backup Documentation**](https://qdrant.tech/documentation/private-cloud/backups/)| @@ -438,7 +438,7 @@ Investigations showed they hadn't adjusted the default configuration or reserved || |-| -|**Read More:** [**Qdrant Configuration Documentation**](https://qdrant.tech/documentation/guides/configuration/)| +|**Read More:** [**Qdrant Configuration Documentation**](https://qdrant.tech/documentation/operations/configuration/)| ### Security & Governance @@ -450,7 +450,7 @@ Enabling TLS/HTTPS is essential for meeting compliance requirements in regulated > You need to protect data in transit. To enable TLS/HTTPS for encrypted traffic in production, you need to configure secure communication between clients and your Qdrant database, as well as individual cluster nodes. This involves implementing Transport Layer Security (TLS) certificates to encrypt all traffic, preventing unauthorized access and data interception. -If self-hosting, you can set up encryption yourself by [**incorporating TLS directly from the configuration**](https://qdrant.tech/documentation/guides/security/#tls) +If self-hosting, you can set up encryption yourself by [**incorporating TLS directly from the configuration**](https://qdrant.tech/documentation/operations/security/#tls) ```text service: @@ -469,7 +469,7 @@ tls: || |-| -|**Read More:** [**Security Documentation**](https://qdrant.tech/documentation/guides/security/)| +|**Read More:** [**Security Documentation**](https://qdrant.tech/documentation/operations/security/)| ### Setting up Access Controls in Production diff --git a/qdrant-landing/content/articles/vector-search-resource-optimization.md b/qdrant-landing/content/articles/vector-search-resource-optimization.md index 8dbd9d89c..d6bc25a7c 100644 --- a/qdrant-landing/content/articles/vector-search-resource-optimization.md +++ b/qdrant-landing/content/articles/vector-search-resource-optimization.md @@ -32,12 +32,12 @@ Let's take a look at some common goals and optimization strategies: | Intended Result | Optimization Strategy | |--------------------------------|------------------------------| -| [**High Search Precision + Low Memory Expenditure**](/documentation/guides/optimize/#1-high-speed-search-with-low-memory-usage) | [**On-Disk Indexing**](/documentation/guides/optimize/#1-high-speed-search-with-low-memory-usage) | -| [**Low Memory Expenditure + Fast Search Speed**](/documentation/guides/quantization/) | [**Quantization**](/documentation/guides/quantization/) | -| [**High Search Precision + Fast Search Speed**](/documentation/guides/optimize/#3-high-precision-with-high-speed-search) | [**RAM Storage + Quantization**](/documentation/guides/optimize/#3-high-precision-with-high-speed-search) | -| [**Balance Latency vs Throughput**](/documentation/guides/optimize/#balancing-latency-and-throughput) | [**Segment Configuration**](/documentation/guides/optimize/#balancing-latency-and-throughput) | +| [**High Search Precision + Low Memory Expenditure**](/documentation/operations/optimize/#1-high-speed-search-with-low-memory-usage) | [**On-Disk Indexing**](/documentation/operations/optimize/#1-high-speed-search-with-low-memory-usage) | +| [**Low Memory Expenditure + Fast Search Speed**](/documentation/manage-data/quantization/) | [**Quantization**](/documentation/manage-data/quantization/) | +| [**High Search Precision + Fast Search Speed**](/documentation/operations/optimize/#3-high-precision-with-high-speed-search) | [**RAM Storage + Quantization**](/documentation/operations/optimize/#3-high-precision-with-high-speed-search) | +| [**Balance Latency vs Throughput**](/documentation/operations/optimize/#balancing-latency-and-throughput) | [**Segment Configuration**](/documentation/operations/optimize/#balancing-latency-and-throughput) | -After this article, check out the code samples in our docs on [**Qdrant’s Optimization Methods**](/documentation/guides/optimize/). +After this article, check out the code samples in our docs on [**Qdrant’s Optimization Methods**](/documentation/operations/optimize/). --- @@ -47,7 +47,7 @@ After this article, check out the code samples in our docs on [**Qdrant’s Opti A vector index is the central location where Qdrant calculates vector similarity. It is the backbone of your search process, retrieving relevant results from vast amounts of data. -Qdrant uses the [**HNSW (Hierarchical Navigable Small World Graph) algorithm**](/documentation/concepts/indexing/#vector-index) as its dense vector index, which is both powerful and scalable. +Qdrant uses the [**HNSW (Hierarchical Navigable Small World Graph) algorithm**](/documentation/manage-data/indexing/#vector-index) as its dense vector index, which is both powerful and scalable. **Figure 2:** A sample HNSW vector index with three layers. Follow the blue arrow on the top layer to see how a query travels throughout the database index. The closest result is on the bottom level, nearest to the gray query point. @@ -57,7 +57,7 @@ Qdrant uses the [**HNSW (Hierarchical Navigable Small World Graph) algorithm**]( Working with massive datasets that contain billions of vectors demands significant resources—and those resources come with a price. While Qdrant provides reasonable defaults, tailoring them to your specific use case can unlock optimal performance. Here’s what you need to know. -The following parameters give you the flexibility to fine-tune Qdrant’s performance for your specific workload. You can modify them directly in Qdrant's [**configuration**](https://qdrant.tech/documentation/guides/configuration/) files or at the collection and named vector levels for more granular control. +The following parameters give you the flexibility to fine-tune Qdrant’s performance for your specific workload. You can modify them directly in Qdrant's [**configuration**](https://qdrant.tech/documentation/operations/configuration/) files or at the collection and named vector levels for more granular control. **Figure 3:** A description of three key HNSW parameters. @@ -101,7 +101,7 @@ client.query_points( ) ``` --- -These are just the basics of HNSW. Learn More about [**Indexing**](/documentation/concepts/indexing/). +These are just the basics of HNSW. Learn More about [**Indexing**](/documentation/manage-data/indexing/). --- @@ -110,7 +110,7 @@ These are just the basics of HNSW. Learn More about [**Indexing**](/documentatio Efficient data compression is a cornerstone of resource optimization in vector databases. By reducing memory usage, you can achieve faster query performance without sacrificing too much accuracy. -One powerful technique is [**quantization**](/documentation/guides/quantization/), which transforms high-dimensional vectors into compact representations while preserving relative similarity. Let’s explore the quantization options available in Qdrant. +One powerful technique is [**quantization**](/documentation/manage-data/quantization/), which transforms high-dimensional vectors into compact representations while preserving relative similarity. Let’s explore the quantization options available in Qdrant. #### Scalar Quantization @@ -156,7 +156,7 @@ When working with Qdrant, you can fine-tune the quantization configuration to op Adjust these settings to strike the right balance between precision and efficiency for your specific workload. --- -Learn More about [**Scalar Quantization**](/documentation/guides/quantization/) +Learn More about [**Scalar Quantization**](/documentation/manage-data/quantization/) --- @@ -194,7 +194,7 @@ client.create_collection( > By default, quantized vectors load like original vectors unless you set `always_ram` to `True` for instant access and faster queries. --- -Learn more about [**Binary Quantization**](/documentation/guides/quantization/) +Learn more about [**Binary Quantization**](/documentation/manage-data/quantization/) --- @@ -259,7 +259,7 @@ client.upsert( To ensure proper data isolation in a multitenant environment, you can assign a unique identifier, such as a **group_id**, to each vector. This approach ensures that each user's data remains segregated, allowing users to access only their own data. You can further enhance this setup by applying filters during queries to restrict access to the relevant data. --- -Learn More about [**Multitenancy**](/documentation/guides/multiple-partitions/) +Learn More about [**Multitenancy**](/documentation/manage-data/multitenancy/) --- @@ -325,7 +325,7 @@ Here’s how to choose the shard_number: | **Plan for Scalability** | Start with at least **2 shards per node** to allow room for future growth. | | **Future-Proofing** | Starting with around **12 shards** is a good rule of thumb. This setup allows your system to scale seamlessly from 1 to 12 nodes without requiring re-sharding. | -Learn more about [**Sharding in Distributed Deployment**](/documentation/guides/distributed_deployment/) +Learn more about [**Sharding in Distributed Deployment**](/documentation/operations/distributed_deployment/) --- @@ -358,10 +358,10 @@ results = client.search( ![filterable-vector-index](/articles_data/vector-search-resource-optimization/filterable-vector-index.png) -[**Filterable vector index**](/documentation/concepts/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points. +[**Filterable vector index**](/documentation/manage-data/indexing/): This technique builds additional links **(orange)** between leftover data points. The filtered points which stay behind are now traversible once again. Qdrant uses special category-based methods to connect these data points. --- -Read more about [**Filtering Docs**](/documentation/concepts/filtering/) and check out the [**Complete Filtering Guide**](/articles/vector-search-filtering/). +Read more about [**Filtering Docs**](/documentation/search/filtering/) and check out the [**Complete Filtering Guide**](/articles/vector-search-filtering/). --- #### Batch Processing @@ -417,7 +417,7 @@ ___ #### Hybrid Search -Hybrid search combines **keyword filtering** with **vector similarity search**, enabling faster and more precise results. Keywords help narrow down the dataset quickly, while vector similarity ensures semantic accuracy. This search method combines [**dense and sparse vectors**](/documentation/concepts/vectors/). +Hybrid search combines **keyword filtering** with **vector similarity search**, enabling faster and more precise results. Keywords help narrow down the dataset quickly, while vector similarity ensures semantic accuracy. This search method combines [**dense and sparse vectors**](/documentation/manage-data/vectors/). Hybrid search in Qdrant uses both fusion and reranking. The former is about combining the results from different search methods, based solely on the scores returned by each method. That usually involves some normalization, as the scores returned by different methods might be in different ranges. @@ -428,7 +428,7 @@ Hybrid search in Qdrant uses both fusion and reranking. The former is about comb After that, there is a formula that takes the relevancy measures and calculates the final score that we use later on to reorder the documents. Qdrant has built-in support for the Reciprocal Rank Fusion method, which is the de facto standard in the field. --- -Learn more about [**Hybrid Search**](/articles/hybrid-search/) and read out [**Hybrid Queries docs**](/documentation/concepts/hybrid-queries/). +Learn more about [**Hybrid Search**](/articles/hybrid-search/) and read out [**Hybrid Queries docs**](/documentation/search/hybrid-queries/). --- @@ -494,7 +494,7 @@ client.query_points( ) ``` ___ -Learn more about [**Reranking**](/documentation/search-precision/reranking-hybrid-search/#rerank). +Learn more about [**Reranking**](/documentation/tutorials-search-engineering/reranking-hybrid-search/#rerank). --- @@ -584,7 +584,7 @@ Here are some important metrics to monitor: | grpc_responses_avg_duration_seconds | | Average response duration in gRPC API | | rest_responses_fail_total | | Total number of failed responses (REST) | -Read more about [**Qdrant Open Source Monitoring**](/documentation/guides/monitoring/) and [**Qdrant Cloud Monitoring**](/documentation/cloud/cluster-monitoring/) for managed clusters. +Read more about [**Qdrant Open Source Monitoring**](/documentation/operations/monitoring/) and [**Qdrant Cloud Monitoring**](/documentation/cloud/cluster-monitoring/) for managed clusters. _________________________________________________________________________ ## Recap: When Should You Optimize? diff --git a/qdrant-landing/content/articles/vector-similarity-beyond-search.md b/qdrant-landing/content/articles/vector-similarity-beyond-search.md index 969285d20..621e83f34 100644 --- a/qdrant-landing/content/articles/vector-similarity-beyond-search.md +++ b/qdrant-landing/content/articles/vector-similarity-beyond-search.md @@ -169,7 +169,7 @@ In this loss, the model is trained by fitting the information of relative simila Using the same mechanics, we can look at the training process from the other side. Given a trained model, the user can provide positive and negative examples, and the goal of the discovery process is then to find suitable anchors across the stored collection of vectors. - + {{< figure width=60% src=/articles_data/vector-similarity-beyond-search/discovery.png caption="Reversed triplet loss" >}} Multiple positive-negative pairs can be provided to make the discovery process more accurate. diff --git a/qdrant-landing/content/articles/what-are-embeddings.md b/qdrant-landing/content/articles/what-are-embeddings.md index 4b045e925..d9ae612ae 100644 --- a/qdrant-landing/content/articles/what-are-embeddings.md +++ b/qdrant-landing/content/articles/what-are-embeddings.md @@ -140,7 +140,7 @@ We plan to go deeper into selecting the best model based on performance, cost, i ## Create a neural search service with Fastmbed -Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](/documentation/tutorials/neural-search/)? +Now that you’re familiar with the core concepts around vector embeddings, how about start building your own [Neural Search Service](/documentation/tutorials-search-engineering/neural-search/)? Tutorial guides you through a practical application of how to use Qdrant for document management based on descriptions of companies from [startups-list.com](https://www.startups-list.com/). From embedding data, integrating it with Qdrant's vector database, constructing a search API, and finally deploying your solution with FastAPI. diff --git a/qdrant-landing/content/articles/what-is-a-vector-database.md b/qdrant-landing/content/articles/what-is-a-vector-database.md index c80b3e366..10886ade0 100644 --- a/qdrant-landing/content/articles/what-is-a-vector-database.md +++ b/qdrant-landing/content/articles/what-is-a-vector-database.md @@ -121,7 +121,7 @@ A vector database is made of multiple different entities and relations. Let's un ### Collections -A [collection](https://qdrant.tech/documentation/concepts/collections/) is essentially a group of **vectors** (or “[points](https://qdrant.tech/documentation/concepts/points/)”) that are logically grouped together **based on similarity or a specific task**. Every vector within a collection shares the same dimensionality and can be compared using a single metric. Avoid creating multiple collections unless necessary; instead, consider techniques like **sharding** for scaling across nodes or **multitenancy** for handling different use cases within the same infrastructure. +A [collection](https://qdrant.tech/documentation/manage-data/collections/) is essentially a group of **vectors** (or “[points](https://qdrant.tech/documentation/manage-data/points/)”) that are logically grouped together **based on similarity or a specific task**. Every vector within a collection shares the same dimensionality and can be compared using a single metric. Avoid creating multiple collections unless necessary; instead, consider techniques like **sharding** for scaling across nodes or **multitenancy** for handling different use cases within the same infrastructure. ### Distance Metrics @@ -154,7 +154,7 @@ client.create_collection( ) ``` -For other configurations like `hnsw_config.on_disk` or `memmap_threshold`, see the Qdrant documentation for [Storage.](https://qdrant.tech/documentation/concepts/storage/) +For other configurations like `hnsw_config.on_disk` or `memmap_threshold`, see the Qdrant documentation for [Storage.](https://qdrant.tech/documentation/manage-data/storage/) ### SDKs @@ -186,7 +186,7 @@ In Qdrant, indexing is modular. You can configure indexes for **both vectors and You need to build the payload index for **each field** you'd like to search. The magic here is in the combination: HNSW finds similar vectors, and the payload index makes sure only the ones that fit your criteria come through. Learn more about Qdrant's [Filterable HNSW](https://qdrant.tech/articles/filterable-hnsw/) and why it was built like this. -> Combining [full-text search](https://qdrant.tech/documentation/concepts/indexing/#full-text-index) with vector-based search gives you even more versatility. You can simultaneously search for conceptually similar documents while ensuring specific keywords are present, all within the same query. +> Combining [full-text search](https://qdrant.tech/documentation/manage-data/indexing/#full-text-index) with vector-based search gives you even more versatility. You can simultaneously search for conceptually similar documents while ensuring specific keywords are present, all within the same query. ### 2. Searching: Approximate Nearest Neighbors (ANN) Search @@ -334,7 +334,7 @@ It works by converting high-dimensional vectors, which typically use `4 bytes` p Quantization reduces data precision, and yes, this does lead to some loss of accuracy. However, for binary quantization, **OpenAI embeddings** achieves this performance improvement at a cost of only 5% of accuracy. If you apply techniques like **oversampling** and **rescoring**, this loss can be brought down even further. -However, binary quantization isn’t the only available option. Techniques like [**Scalar Quantization**](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization) and [**Product Quantization**](https://qdrant.tech/documentation/guides/quantization/#product-quantization) are also popular alternatives when optimizing vector compression. +However, binary quantization isn’t the only available option. Techniques like [**Scalar Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization) and [**Product Quantization**](https://qdrant.tech/documentation/manage-data/quantization/#product-quantization) are also popular alternatives when optimizing vector compression. You can set up your chosen quantization method using the `quantization_config` parameter when creating a new collection: @@ -414,7 +414,7 @@ client.create_collection( We recommend using sharding and replication together so that your data is both split across nodes and replicated for availability. -For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/guides/distributed_deployment/) +For more details on features like **user-defined sharding, node failure recovery**, and **consistency guarantees**, see our guide on [Distributed Deployment.](https://qdrant.tech/documentation/operations/distributed_deployment/) ## Multitenancy: Data Isolation for Multi-Tenant Architectures @@ -477,7 +477,7 @@ You can easily setup your access tokens and secure access to sensitive data thro Qdrant Web UI for generating a new access token. -By default, Qdrant instances are **unsecured**, so it's important to configure security measures before moving to production. To learn more about how to configure security for your Qdrant instance and other advanced options, please check out the [official Qdrant documentation on security.](https://qdrant.tech/documentation/guides/security/) +By default, Qdrant instances are **unsecured**, so it's important to configure security measures before moving to production. To learn more about how to configure security for your Qdrant instance and other advanced options, please check out the [official Qdrant documentation on security.](https://qdrant.tech/documentation/operations/security/) ## Time to Experiment diff --git a/qdrant-landing/content/articles/what-is-quantization.md b/qdrant-landing/content/articles/what-is-quantization.md index 34e5ba7c4..baf2de1fc 100644 --- a/qdrant-landing/content/articles/what-is-quantization.md +++ b/qdrant-landing/content/articles/what-is-quantization.md @@ -28,7 +28,7 @@ When working with high-dimensional vectors, such as embeddings from providers li With 1 million vectors needing around 6 GB of memory, as your dataset grows to multiple **millions of vectors**, the memory and processing demands increase significantly. -To understand why this process is so computationally demanding, let's take a look at the nature of the [HNSW index](https://qdrant.tech/documentation/concepts/indexing/#vector-index). +To understand why this process is so computationally demanding, let's take a look at the nature of the [HNSW index](https://qdrant.tech/documentation/manage-data/indexing/#vector-index). The **HNSW (Hierarchical Navigable Small World) index** organizes vectors in a layered graph, connecting each vector to its nearest neighbors. At each layer, the algorithm narrows down the search area until it reaches the lower layers, where it efficiently finds the closest matches to the query. @@ -54,7 +54,7 @@ There are several methods to achieve this, and here we will focus on three main ![](/articles_data/what-is-vector-quantization/astronaut-mars.jpg) -In Qdrant, each dimension is represented by a `float32` value, which uses **4 bytes** of memory. When using [Scalar Quantization](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization), we map our vectors to a range that the smaller `int8` type can represent. An `int8` is only **1 byte** and can represent 256 values (from -128 to 127, or 0 to 255). This results in a **75% reduction** in memory size. +In Qdrant, each dimension is represented by a `float32` value, which uses **4 bytes** of memory. When using [Scalar Quantization](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization), we map our vectors to a range that the smaller `int8` type can represent. An `int8` is only **1 byte** and can represent 256 values (from -128 to 127, or 0 to 255). This results in a **75% reduction** in memory size. For example, if our data lies in the range of -1.0 to 1.0, Scalar Quantization will transform these values to a range that `int8` can represent, i.e., within -128 to 127. The system **maps** the `float32` values into this range. @@ -107,7 +107,7 @@ While the performance gains of Scalar Quantization may not match those achieved ![Astronaut in surreal white environment](/articles_data/what-is-vector-quantization/astronaut-white-surreal.jpg) -[Binary Quantization](https://qdrant.tech/documentation/guides/quantization/#binary-quantization) is an excellent option if you're looking to **reduce memory** usage while also achieving a significant **boost in speed**. It works by converting high-dimensional vectors into simple binary (0 or 1) representations. +[Binary Quantization](https://qdrant.tech/documentation/manage-data/quantization/#binary-quantization) is an excellent option if you're looking to **reduce memory** usage while also achieving a significant **boost in speed**. It works by converting high-dimensional vectors into simple binary (0 or 1) representations. - Values greater than zero are converted to 1. - Values less than or equal to zero are converted to 0. @@ -176,7 +176,7 @@ If you're interested in exploring Binary Quantization in more detail—including ![](/articles_data/what-is-vector-quantization/astronaut-centroids.jpg) -[Product Quantization](https://qdrant.tech/documentation/guides/quantization/#product-quantization) is a method used to compress high-dimensional vectors by representing them with a smaller set of representative points. +[Product Quantization](https://qdrant.tech/documentation/manage-data/quantization/#product-quantization) is a method used to compress high-dimensional vectors by representing them with a smaller set of representative points. The process begins by splitting the original high-dimensional vectors into smaller **sub-vectors.** Each sub-vector represents a segment of the original vector, capturing different characteristics of the data. @@ -268,7 +268,7 @@ Product Quantization can significantly reduce memory usage, potentially offering If your application requires high precision or real-time performance, Product Quantization may not be the best choice. However, if **memory savings** are critical and some accuracy loss is acceptable, it could still be an ideal solution. -Here’s a comparison of speed, accuracy, and compression for all three methods, adapted from [Qdrant's documentation](https://qdrant.tech/documentation/guides/quantization/#how-to-choose-the-right-quantization-method): +Here’s a comparison of speed, accuracy, and compression for all three methods, adapted from [Qdrant's documentation](https://qdrant.tech/documentation/manage-data/quantization/#how-to-choose-the-right-quantization-method): | Quantization method | Accuracy | Speed | Compression | |---------------------|----------|------------|-------------| @@ -519,7 +519,7 @@ Here are some final thoughts to help you choose the right quantization method fo ### Learn More -If you want to learn more about improving accuracy, memory efficiency, and speed when using quantization in Qdrant, we have a dedicated [Quantization tips](https://qdrant.tech/documentation/guides/quantization/#quantization-tips) section in our docs that explains all the quantization tips you can use to enhance your results. +If you want to learn more about improving accuracy, memory efficiency, and speed when using quantization in Qdrant, we have a dedicated [Quantization tips](https://qdrant.tech/documentation/manage-data/quantization/#quantization-tips) section in our docs that explains all the quantization tips you can use to enhance your results. Learn more about optimizing real-time precision with oversampling in Binary Quantization by watching this interview with Qdrant’s CTO, Andrey Vasnetsov: diff --git a/qdrant-landing/content/articles/what-is-rag-in-ai.md b/qdrant-landing/content/articles/what-is-rag-in-ai.md index 45a35d1da..808f9b587 100644 --- a/qdrant-landing/content/articles/what-is-rag-in-ai.md +++ b/qdrant-landing/content/articles/what-is-rag-in-ai.md @@ -133,7 +133,7 @@ The LLM is typically a model like GPT, BART or T5, trained on massive datasets t ![How a Generator works](/articles_data/what-is-rag-in-ai/how-generation-works.png) -The retriever and generator don't operate in isolation. The image bellow shows how the output of the retrieval feeds the generator to produce the final generated response. +The retriever and generator don't operate in isolation. The image below shows how the output of the retrieval feeds the generator to produce the final generated response. ![The entire architecture of a RAG system](/articles_data/what-is-rag-in-ai/rag-system.jpg) diff --git a/qdrant-landing/content/benchmarks/filtered-search-benchmark.md b/qdrant-landing/content/benchmarks/filtered-search-benchmark.md index 6f8756b42..380353ab8 100644 --- a/qdrant-landing/content/benchmarks/filtered-search-benchmark.md +++ b/qdrant-landing/content/benchmarks/filtered-search-benchmark.md @@ -18,7 +18,7 @@ As you can see from the charts, there are three main patterns: - **Speed downturn** - some engines struggle to keep high RPS, it might be related to the requirement of building a filtering mask for the dataset, as described above. -- **Accuracy collapse** - some engines are loosing accuracy dramatically under some filters. It is related to the fact that the HNSW graph becomes disconnected, and the search becomes unreliable. +- **Accuracy collapse** - some engines are losing accuracy dramatically under some filters. It is related to the fact that the HNSW graph becomes disconnected, and the search becomes unreliable. Qdrant avoids all these problems and also benefits from the speed boost, as it implements an advanced [query planning strategy](/documentation/search/#query-planning). diff --git a/qdrant-landing/content/benchmarks/single-node-speed-benchmark.md b/qdrant-landing/content/benchmarks/single-node-speed-benchmark.md index 98a7de5e4..993bcb74b 100644 --- a/qdrant-landing/content/benchmarks/single-node-speed-benchmark.md +++ b/qdrant-landing/content/benchmarks/single-node-speed-benchmark.md @@ -17,7 +17,7 @@ Unlisted: false Most of the engines have improved since [our last run](/benchmarks/single-node-speed-benchmark-2022/). Both life and software have trade-offs but some clearly do better: -* **`Qdrant` achives highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets. +* **`Qdrant` achieves highest RPS and lowest latencies in almost all the scenarios, no matter the precision threshold and the metric we choose.** It has also shown 4x RPS gains on one of the datasets. * `Elasticsearch` has become considerably fast for many cases but it's very slow in terms of indexing time. It can be 10x slower when storing 10M+ vectors of 96 dimensions! (32mins vs 5.5 hrs) * `Milvus` is the fastest when it comes to indexing time and maintains good precision. However, it's not on-par with others when it comes to RPS or latency when you have higher dimension embeddings or more number of vectors. * `Redis` is able to achieve good RPS but mostly for lower precision. It also achieved low latency with single thread, however its latency goes up quickly with more parallel requests. Part of this speed gain comes from their custom protocol. diff --git a/qdrant-landing/content/blog/2025-recap.md b/qdrant-landing/content/blog/2025-recap.md index 912c160a7..c21c0660a 100644 --- a/qdrant-landing/content/blog/2025-recap.md +++ b/qdrant-landing/content/blog/2025-recap.md @@ -46,9 +46,9 @@ In response, our 2025 roadmap centered on four tightly connected capability area In 2025, we focused on giving teams explicit control over retrieval quality as applications moved beyond basic semantic search. Our new capabilities make relevance more explainable, tunable, and aligned with real user intent, especially in agentic and hybrid search workflows. **Related enhancements:** -• [Score-Boosting Reranking](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting) allowing the blending of vector similarity with business signals -• [Full-Text Filtering](https://qdrant.tech/documentation/concepts/filtering/) which brought native multilingual tokenization, stemming, and phrase matching -• [ACORN algorithm](https://qdrant.tech/documentation/concepts/search/#acorn-search-algorithm) for higher-quality filtered HNSW queries +• [Score-Boosting Reranking](https://qdrant.tech/documentation/search/search-relevance/#score-boosting) allowing the blending of vector similarity with business signals +• [Full-Text Filtering](https://qdrant.tech/documentation/search/filtering/) which brought native multilingual tokenization, stemming, and phrase matching +• [ACORN algorithm](https://qdrant.tech/documentation/search/search/#acorn-search-algorithm) for higher-quality filtered HNSW queries • [Maximal Marginal Relevance (MMR)](https://qdrant.tech/blog/mmr-diversity-aware-reranking/) to balance relevance and diversity • ASCII folding for improved multilingual recall @@ -57,19 +57,19 @@ In 2025, we focused on giving teams explicit control over retrieval quality as a To support large, cost-sensitive workloads, we targeted the biggest performance bottlenecks in production systems. New improvements help teams scale indexing and querying without over-provisioning memory or compute. **Related enhancements:** -• [GPU-Accelerated HNSW Indexing](https://qdrant.tech/documentation/guides/running-with-gpu/) unlocks up to an order-of-magnitude faster ingestion -• [Inline Storage](https://qdrant.tech/documentation/guides/optimize/#inline-storage-in-hnsw-index) embedded quantized vectors directly into the graph to dramatically improve disk-based search performance +• [GPU-Accelerated HNSW Indexing](https://qdrant.tech/documentation/operations/running-with-gpu/) unlocks up to an order-of-magnitude faster ingestion +• [Inline Storage](https://qdrant.tech/documentation/operations/optimize/#inline-storage-in-hnsw-index) embedded quantized vectors directly into the graph to dramatically improve disk-based search performance • [Custom storage engine](https://qdrant.tech/articles/gridstore-key-value-storage/) optimized for predictable low-latency access • [Incremental HNSW indexing](https://qdrant.tech/documentation/database-tutorials/bulk-upload/?q=incremental+hnsw#choose-an-indexing-strategy) for upsert-heavy workloads • HNSW graph compression to reduce memory footprint -• Expanded [Quantization](https://qdrant.tech/documentation/guides/quantization/#15-bit-and-2-bit-quantization]) options, including 1.5-bit, 2-bit, and asymmetric quantization +• Expanded [Quantization](https://qdrant.tech/documentation/manage-data/quantization/#15-bit-and-2-bit-quantization]) options, including 1.5-bit, 2-bit, and asymmetric quantization ### Enterprise Scaling & Isolation As Qdrant became shared infrastructure inside larger organizations, we focused on multitenancy, governance, and enterprise needs. **Related enhancements:** -• [Tiered Multitenancy](https://qdrant.tech/documentation/guides/multitenancy/#tiered-multitenancy) enables efficient support for both small and large tenants within a single system +• [Tiered Multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/#tiered-multitenancy) enables efficient support for both small and large tenants within a single system • [Single Sign-On (SSO) and role-based access control (RBAC)](https://qdrant.tech/enterprise-solutions/) • Granular database API keys • [Terraform-enabled Cloud API](https://qdrant.tech/enterprise-solutions/) for automation and governance diff --git a/qdrant-landing/content/blog/azure-marketplace.md b/qdrant-landing/content/blog/azure-marketplace.md index da6c9f50d..290d7ff1f 100644 --- a/qdrant-landing/content/blog/azure-marketplace.md +++ b/qdrant-landing/content/blog/azure-marketplace.md @@ -41,8 +41,8 @@ Ready to experience the benefits of Qdrant on Azure Marketplace? Getting started 1. **Visit the Azure Marketplace**: Navigate to [Qdrant's Marketplace listing](https://azuremarketplace.microsoft.com/en-en/marketplace/apps/qdrantsolutionsgmbh1698769709989.qdrant-db). 2. **Deploy Qdrant**: Follow the simple deployment instructions to set up your instance. -3. **Start Using Qdrant**: Once deployed, start exploring the [features and capabilities of Qdrant](/documentation/concepts/) on Azure. -4. **Read Documentation**: Read Qdrant's [Documentation](/documentation/) and build demo apps using [Tutorials](/documentation/tutorials/). +3. **Start Using Qdrant**: Once deployed, start exploring the [features and capabilities of Qdrant](/documentation/overview/) on Azure. +4. **Read Documentation**: Read Qdrant's [Documentation](/documentation/) and build demo apps using [Tutorials](/documentation/tutorials-lp-overview/). ## Join Us on this Exciting Journey: diff --git a/qdrant-landing/content/blog/beta-database-migration-tool.md b/qdrant-landing/content/blog/beta-database-migration-tool.md index eb9b5d63e..191f97a83 100644 --- a/qdrant-landing/content/blog/beta-database-migration-tool.md +++ b/qdrant-landing/content/blog/beta-database-migration-tool.md @@ -19,7 +19,7 @@ We’ve launched the **beta** of our Qdrant **Vector Data Migration Tool**, desi This powerful tool streams all vectors from a source collection to a target Qdrant instance in live batches. It supports migrations from one Qdrant deployment to another, including from open source to Qdrant Cloud or between cloud regions. But that's not all. You can also migrate your data from other vector databases directly into Qdrant. All with a single command. -Unlike Qdrant’s included [snapshot migration method](https://qdrant.tech/documentation/concepts/snapshots/), which requires consistent node-specific snapshots, our migration tool enables you to easily migrate data between different Qdrant database clusters in streaming batches. The only requirement is that the vector size and distance function must match. +Unlike Qdrant’s included [snapshot migration method](https://qdrant.tech/documentation/operations/snapshots/), which requires consistent node-specific snapshots, our migration tool enables you to easily migrate data between different Qdrant database clusters in streaming batches. The only requirement is that the vector size and distance function must match. This is especially useful if you want to change the collection configuration on the target, for example by choosing a different replication factor or quantization method. diff --git a/qdrant-landing/content/blog/case-study-and-ai.md b/qdrant-landing/content/blog/case-study-and-ai.md index af1c24d17..a7bd8428f 100644 --- a/qdrant-landing/content/blog/case-study-and-ai.md +++ b/qdrant-landing/content/blog/case-study-and-ai.md @@ -71,7 +71,7 @@ By using [Qdrant Cloud](https://qdrant.tech/cloud/), \&AI avoided the need to ma "Patent litigation has huge stakes, one result could influence a billion-dollar case," said Turner. "Accuracy is the top priority, and Qdrant let us optimize for that without compromising on cost or performance." -Qdrant’s support for [payload filters](https://qdrant.tech/documentation/concepts/filtering/), [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/), and quantization let \&AI optimize deeply. Their AI patent agent, Andy, uses natural language to guide attorneys through patent analysis tasks, drastically cutting time-to-result. +Qdrant’s support for [payload filters](https://qdrant.tech/documentation/search/filtering/), [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/), and quantization let \&AI optimize deeply. Their AI patent agent, Andy, uses natural language to guide attorneys through patent analysis tasks, drastically cutting time-to-result. *"With Qdrant, we scaled to a billion vectors and still respond in sub-second latency. That lets us power workflows that used to take hours in just a few minutes."* diff --git a/qdrant-landing/content/blog/case-study-anima-health.md b/qdrant-landing/content/blog/case-study-anima-health.md index ac5e427c7..3c32dc21a 100644 --- a/qdrant-landing/content/blog/case-study-anima-health.md +++ b/qdrant-landing/content/blog/case-study-anima-health.md @@ -56,14 +56,14 @@ Beyond coding, Anima uses Qdrant to understand documents at scale. By working wi Several factors made Qdrant a strong fit for healthcare workloads. -[Deployment flexibility](https://qdrant.tech/documentation/guides/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance. +[Deployment flexibility](https://qdrant.tech/documentation/operations/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance. Cost predictability also played a critical role. With a fixed infrastructure cost for vector search, Anima could use retrieval across multiple passes in their pipelines. This unlocked higher-quality results without eroding margins. >“Knowing that retrieval is reliable and low cost changed how we build. We do not think twice about using vector search as part of our pipelines.” -Colin Cooke, Lead AI Engineer, Anima Health -Finally, Qdrant’s vector-native capabilities mattered. [Payload-based filtering](https://qdrant.tech/documentation/concepts/payload/) allows Anima to scope searches precisely across different electronic health record systems. [Multivector support](https://qdrant.tech/documentation/concepts/payload/) enables experimentation with multiple embedding strategies and providers, reducing long-term lock-in and easing future transitions. +Finally, Qdrant’s vector-native capabilities mattered. [Payload-based filtering](https://qdrant.tech/documentation/manage-data/payload/) allows Anima to scope searches precisely across different electronic health record systems. [Multivector support](https://qdrant.tech/documentation/manage-data/payload/) enables experimentation with multiple embedding strategies and providers, reducing long-term lock-in and easing future transitions. ### Results: Scalable, privacy-first AI in production diff --git a/qdrant-landing/content/blog/case-study-bazaarvoice.md b/qdrant-landing/content/blog/case-study-bazaarvoice.md index dd3199b58..83ee0481b 100644 --- a/qdrant-landing/content/blog/case-study-bazaarvoice.md +++ b/qdrant-landing/content/blog/case-study-bazaarvoice.md @@ -55,8 +55,8 @@ With over a billion vectors, post-filtering meant searching far more data than n Qdrant stood out for a few key reasons: -* [Multitenancy](https://qdrant.tech/documentation/guides/multitenancy/) with payload-based partitioning, allowing searches to be scoped by client, product, or category at query time -* [Quantization](https://qdrant.tech/documentation/guides/quantization/), enabling dramatic reductions in storage and RAM requirements which translates directly to cost +* [Multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) with payload-based partitioning, allowing searches to be scoped by client, product, or category at query time +* [Quantization](https://qdrant.tech/documentation/manage-data/quantization/), enabling dramatic reductions in storage and RAM requirements which translates directly to cost * [Hybrid cloud deployment](https://qdrant.tech/hybrid-cloud/), running inside Bazaarvoice’s VPC on Kubernetes * Operational simplicity, eliminating manual partition management entirely diff --git a/qdrant-landing/content/blog/case-study-dailymotion.md b/qdrant-landing/content/blog/case-study-dailymotion.md index a4b422bad..ebefd72b7 100644 --- a/qdrant-landing/content/blog/case-study-dailymotion.md +++ b/qdrant-landing/content/blog/case-study-dailymotion.md @@ -167,7 +167,7 @@ They aim to work on Perspective feed next and say ![perspective-feed-with-qdrant](/case-studies/dailymotion/perspective-feed-qdrant.jpg) -The team is also interested in leveraging advanced features like [Qdrant’s Discovery API](/documentation/concepts/explore/#recommendation-api) to promote exploration of content to enable finding not only similar but dissimilar content too by using positive and negative vectors in the queries and making it work with the existing collaborative recommendation model. +The team is also interested in leveraging advanced features like [Qdrant’s Discovery API](/documentation/search/explore/#recommendation-api) to promote exploration of content to enable finding not only similar but dissimilar content too by using positive and negative vectors in the queries and making it work with the existing collaborative recommendation model. ### References diff --git a/qdrant-landing/content/blog/case-study-deutsche-telekom.md b/qdrant-landing/content/blog/case-study-deutsche-telekom.md index b00bd2f3f..67cd821db 100644 --- a/qdrant-landing/content/blog/case-study-deutsche-telekom.md +++ b/qdrant-landing/content/blog/case-study-deutsche-telekom.md @@ -76,7 +76,7 @@ When Deutsche Telekom began searching for a scalable, high-performance vector da The team structured its evaluation around two key metrics: 1. **Qualitative metrics**: developer experience, ease of use, memory efficiency features. -2. **Operational simplicity**: how well it fit into their PaaS-first approach and [multitenancy requirements](https://qdrant.tech/documentation/guides/multiple-partitions/). +2. **Operational simplicity**: how well it fit into their PaaS-first approach and [multitenancy requirements](https://qdrant.tech/documentation/manage-data/multitenancy/). Deutsche Telekom's engineers also cited several standout features that made Qdrant the right fit: diff --git a/qdrant-landing/content/blog/case-study-dragonfruit.md b/qdrant-landing/content/blog/case-study-dragonfruit.md index 1f52a7792..8c8ffed29 100644 --- a/qdrant-landing/content/blog/case-study-dragonfruit.md +++ b/qdrant-landing/content/blog/case-study-dragonfruit.md @@ -41,13 +41,13 @@ Retail and warehouse environments are bandwidth-constrained and heterogeneous. A * Operate a vector store at enterprise scale: thousands of locations → thousands of cameras, accumulating into tens to hundreds of billions of vectors and multi-terabyte storage. ### Why Qdrant: Performance headroom and operational control -Dragonfruit chose the open-source version of Qdrant as its vector search engine to meet the twin pressures of real-time reads and high-velocity writes. In head-to-head experiments, Qdrant delivered the QPS targets they needed while giving the team granular, [per-collection](https://qdrant.tech/documentation/concepts/collections/) tuning to match workload diversity. +Dragonfruit chose the open-source version of Qdrant as its vector search engine to meet the twin pressures of real-time reads and high-velocity writes. In head-to-head experiments, Qdrant delivered the QPS targets they needed while giving the team granular, [per-collection](https://qdrant.tech/documentation/manage-data/collections/) tuning to match workload diversity. Key reasons the team highlighted: * **Per-collection configurability.** Collections with heavy reads and low writes use different settings than write-heavy pipelines. Tuning shard counts and HNSW parameters by collection helped hit latency Service Level Objectives (SLOs) without overprovisioning. -* **Efficient numeric formats.** For most vision workloads, [float16](https://qdrant.tech/documentation/concepts/vectors/) vectors were sufficient, improving memory efficiency and cache behavior with no material loss in retrieval accuracy for their use cases. +* **Efficient numeric formats.** For most vision workloads, [float16](https://qdrant.tech/documentation/manage-data/vectors/) vectors were sufficient, improving memory efficiency and cache behavior with no material loss in retrieval accuracy for their use cases. * **Open source and ecosystem fit.** Qdrant’s OSS model aligned with Dragonfruit’s platform strategy and let them co-evolve the deployment with their edge and cloud stack. diff --git a/qdrant-landing/content/blog/case-study-dust.md b/qdrant-landing/content/blog/case-study-dust.md index 62be7af77..16ec5f142 100644 --- a/qdrant-landing/content/blog/case-study-dust.md +++ b/qdrant-landing/content/blog/case-study-dust.md @@ -83,8 +83,8 @@ billing and increase security by having the instance live within the same VPC. 2. **Scale and optimize:** As the load grew, Dust started to take advantage of Qdrant’s features to tune the setup for optimization and scale. They started to look into how they map and cache data, as well as applying some of Qdrant’s [built-in -compression features](/documentation/guides/quantization/). In particular, Dust leveraged the control of the [MMAP -payload threshold](/documentation/concepts/storage/#configuring-memmap-storage) as well as [Scalar Quantization](/articles/scalar-quantization/), which enabled Dust to manage +compression features](/documentation/manage-data/quantization/). In particular, Dust leveraged the control of the [MMAP +payload threshold](/documentation/manage-data/storage/#configuring-memmap-storage) as well as [Scalar Quantization](/articles/scalar-quantization/), which enabled Dust to manage the balance between storing vectors on disk and keeping quantized vectors in RAM, more effectively. “This allowed us to scale smoothly from there,” Polu says. diff --git a/qdrant-landing/content/blog/case-study-fieldy.md b/qdrant-landing/content/blog/case-study-fieldy.md index 1aaa2fdc8..2a9d30e50 100644 --- a/qdrant-landing/content/blog/case-study-fieldy.md +++ b/qdrant-landing/content/blog/case-study-fieldy.md @@ -47,7 +47,7 @@ For the engineering team, these failures had two serious implications. First, mi After evaluating alternatives, Fieldy selected [Qdrant](http://qdrant.tech) for its stability, straightforward configuration, and suitability for self-hosted deployment. They opted to run Qdrant in the same environment as their backend services, ensuring low-latency access and avoiding the cross-region connectivity issues that had contributed to failures in the previous architecture. -The migration process took less than a week. The team reused their existing vector schema to avoid redesigning their indexing logic during the transition. Embeddings, generated using Cohere’s multilingual v3 model, were batched locally before ingestion, replacing the integrated embedding calls previously handled within Weaviate. Following [Qdrant best practices](https://qdrant.tech/documentation/guides/optimize/), they disabled indexing during bulk import to maximize throughput, re-enabling it only after the migration was complete. The backend API calls were updated to use Qdrant’s gRPC interface, and [hybrid search](https://qdrant.tech/articles/hybrid-search/) with BM25 was combined with dense vector retrieval via [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/#hybrid-search) for relevance scoring. +The migration process took less than a week. The team reused their existing vector schema to avoid redesigning their indexing logic during the transition. Embeddings, generated using Cohere’s multilingual v3 model, were batched locally before ingestion, replacing the integrated embedding calls previously handled within Weaviate. Following [Qdrant best practices](https://qdrant.tech/documentation/operations/optimize/), they disabled indexing during bulk import to maximize throughput, re-enabling it only after the migration was complete. The backend API calls were updated to use Qdrant’s gRPC interface, and [hybrid search](https://qdrant.tech/articles/hybrid-search/) with BM25 was combined with dense vector retrieval via [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/search/hybrid-queries/#hybrid-search) for relevance scoring. ### Architecture after migration @@ -63,4 +63,4 @@ Fieldy also achieved a two-thirds reduction in infrastructure costs after moving ### Next steps for retrieval quality -With reliability and cost efficiency achieved, Fieldy’s engineering focus is shifting toward retrieval quality. Planned improvements include adding [location filtering](https://qdrant.tech/documentation/concepts/filtering/#geo) and [datetime filtering](https://qdrant.tech/documentation/concepts/filtering/#datetime-range) within Qdrant to refine result sets, experimenting with late chunking strategies, and testing parallel hybrid searches to increase recall on complex multi-faceted queries. The team is also exploring embedding summaries alongside raw transcript segments to improve retrieval performance on high-level or thematic searches. \ No newline at end of file +With reliability and cost efficiency achieved, Fieldy’s engineering focus is shifting toward retrieval quality. Planned improvements include adding [location filtering](https://qdrant.tech/documentation/search/filtering/#geo) and [datetime filtering](https://qdrant.tech/documentation/search/filtering/#datetime-range) within Qdrant to refine result sets, experimenting with late chunking strategies, and testing parallel hybrid searches to increase recall on complex multi-faceted queries. The team is also exploring embedding summaries alongside raw transcript segments to improve retrieval performance on high-level or thematic searches. \ No newline at end of file diff --git a/qdrant-landing/content/blog/case-study-flipkart.md b/qdrant-landing/content/blog/case-study-flipkart.md index b4482ecc6..c477e23ff 100644 --- a/qdrant-landing/content/blog/case-study-flipkart.md +++ b/qdrant-landing/content/blog/case-study-flipkart.md @@ -4,7 +4,6 @@ title: "Building real-time multimodal similarity search in Flipkart Trust & Safe short_description: "Tackling fraud and abuse with scalable similarity search." description: "Tackling fraud and abuse with scalable similarity search." preview_image: /blog/case-study-flipkart/social_preview_partnership-flipkart.png -social_preview_image: /blog/case-study/social_preview_partnership-flipkart.png date: 2026-01-09 author: "Daniel Azoulai" featured: true diff --git a/qdrant-landing/content/blog/case-study-glassdollar.md b/qdrant-landing/content/blog/case-study-glassdollar.md index 2bcde37fa..ff5508aee 100644 --- a/qdrant-landing/content/blog/case-study-glassdollar.md +++ b/qdrant-landing/content/blog/case-study-glassdollar.md @@ -48,7 +48,7 @@ Instead of optimizing around a single response time target, GlassDollar focused ## Accuracy Meant End-to-end Results, not Just Retriever Scores -For GlassDollar, accuracy was measured at the workflow level: did the system surface the right companies for a given enterprise need, and did users act on the results. That meant optimizing the full architecture, including [query expansion](https://qdrant.tech/documentation/concepts/hybrid-queries/), retrieval, ranking, and contextual ranking, rather than chasing marginal gains in any single component. Faster retrieval mattered most because it enabled more queries to improve query expansion, which raised recall and improved the final shortlist quality. +For GlassDollar, accuracy was measured at the workflow level: did the system surface the right companies for a given enterprise need, and did users act on the results. That meant optimizing the full architecture, including [query expansion](https://qdrant.tech/documentation/search/hybrid-queries/), retrieval, ranking, and contextual ranking, rather than chasing marginal gains in any single component. Faster retrieval mattered most because it enabled more queries to improve query expansion, which raised recall and improved the final shortlist quality. ## Contextual Embeddings Improved Matching Between User Intent and Company Descriptions diff --git a/qdrant-landing/content/blog/case-study-kakao.md b/qdrant-landing/content/blog/case-study-kakao.md index 422508ede..80e6fc741 100644 --- a/qdrant-landing/content/blog/case-study-kakao.md +++ b/qdrant-landing/content/blog/case-study-kakao.md @@ -47,9 +47,9 @@ After evaluating multiple vector databases, the Connectivity Platform team selec The decision came down to a combination of search quality, performance, and operational fit. -Qdrant’s hybrid search capabilities were a key factor. By supporting both dense vectors for semantic search and sparse vectors for keyword-based retrieval—combined using [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/#reciprocal-rank-fusion-rrf), the team could address both conceptual questions and exact-match queries in a single system. Named Vectors made it possible to manage multiple vector types within the same collection. +Qdrant’s hybrid search capabilities were a key factor. By supporting both dense vectors for semantic search and sparse vectors for keyword-based retrieval—combined using [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/search/hybrid-queries/#reciprocal-rank-fusion-rrf), the team could address both conceptual questions and exact-match queries in a single system. Named Vectors made it possible to manage multiple vector types within the same collection. -Performance was another major consideration. Qdrant’s [Rust-based architecture](https://qdrant.tech/articles/why-rust/), efficient [HNSW implementation](https://qdrant.tech/course/essentials/day-2/what-is-hnsw/), and support for [scalar quantization (INT8)](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization) provided low-latency search while optimizing memory usage. This was crucial for an internal service expected to scale over time. +Performance was another major consideration. Qdrant’s [Rust-based architecture](https://qdrant.tech/articles/why-rust/), efficient [HNSW implementation](https://qdrant.tech/course/essentials/day-2/what-is-hnsw/), and support for [scalar quantization (INT8)](https://qdrant.tech/documentation/manage-data/quantization/#scalar-quantization) provided low-latency search while optimizing memory usage. This was crucial for an internal service expected to scale over time. From an operational standpoint, Qdrant fit naturally into Kakao’s environment. Its single-binary design simplified deployment, it ran reliably on Kubernetes, and it allowed Kakao to retain full control over data by self-hosting within internal infrastructure. @@ -63,7 +63,7 @@ The team integrated Qdrant using the asynchronous Python client (`AsyncQdrantCli Collections were designed around data sources, with separate collections for internal technical documentation, historical inquiry data, and a semantic cache used to speed up repeated queries. Metadata filtering allows the system to narrow search scope by service or time period, while maintaining fast response times. -Each collection stores both dense and sparse vectors using [Named Vectors](https://qdrant.tech/documentation/concepts/vectors/#named-vectors). Hybrid search results are merged using RRF to produce more accurate answers across different query types. +Each collection stores both dense and sparse vectors using [Named Vectors](https://qdrant.tech/documentation/manage-data/vectors/#named-vectors). Hybrid search results are merged using RRF to produce more accurate answers across different query types. An automated indexing pipeline handles document ingestion end-to-end. This ranges from cleansing and chunking, to embedding generation, to batch upserts into Qdrant. diff --git a/qdrant-landing/content/blog/case-study-lettria-v2.md b/qdrant-landing/content/blog/case-study-lettria-v2.md index dd791b489..dd4ccfc8c 100644 --- a/qdrant-landing/content/blog/case-study-lettria-v2.md +++ b/qdrant-landing/content/blog/case-study-lettria-v2.md @@ -38,7 +38,7 @@ Enterprises in regulated sectors deal with extensive, complex documentation feat One component of the build was the vector database. Lettria evaluated Weaviate, Milvus, and Qdrant based on their hybrid search capability, deployment simplicity (Docker, Kubernetes), and search performance (latency, RAM usage). -Ultimately, Lettria chose Qdrant. First, it had a simple Kubernetes deployment, superior latency and lower memory footprint in competitive benchmarks. Additionally, there were unique features, such as the grouping API and detailed [payload indexing](https://qdrant.tech/documentation/concepts/payload/), that made Qdrant stand out. +Ultimately, Lettria chose Qdrant. First, it had a simple Kubernetes deployment, superior latency and lower memory footprint in competitive benchmarks. Additionally, there were unique features, such as the grouping API and detailed [payload indexing](https://qdrant.tech/documentation/manage-data/payload/), that made Qdrant stand out. ## Building the document understanding and extraction pipeline @@ -288,7 +288,7 @@ Note that they duplicate the english tags for string properties as they are cons #### Filtering -based on a filter definition. Lettria flattens properties on Neo4J so that they can use similar filters. The nested structure ([more on that here](https://qdrant.tech/documentation/concepts/filtering/#nested-key)) {"foo": { "bar": "qux" }} is kept in Qdrant and dot separated in NeoJ: foo.bar=qux so that they can perform match queries with similar keys from Qdrant. This introduces some complexity as they need to be careful in the handling of url and other 'dot rich' values in properties. +based on a filter definition. Lettria flattens properties on Neo4J so that they can use similar filters. The nested structure ([more on that here](https://qdrant.tech/documentation/search/filtering/#nested-key)) {"foo": { "bar": "qux" }} is kept in Qdrant and dot separated in NeoJ: foo.bar=qux so that they can perform match queries with similar keys from Qdrant. This introduces some complexity as they need to be careful in the handling of url and other 'dot rich' values in properties. If they want to filter based on the onto:surface value, the same keys are used in Qdrant and Neo4J: diff --git a/qdrant-landing/content/blog/case-study-my-askai.md b/qdrant-landing/content/blog/case-study-my-askai.md index f7765411d..b41c35c74 100644 --- a/qdrant-landing/content/blog/case-study-my-askai.md +++ b/qdrant-landing/content/blog/case-study-my-askai.md @@ -48,7 +48,7 @@ Everything changed when OpenAI released its embedding model. Instead of hoping u >"That was a transformational moment for us. Now we could have hundreds of help articles ingested in the system, and a user can ask a question, and we can answer that really specifically and cheaply and quickly." -Over time, the team also learned that semantic search was strong but not universally sufficient, especially when tickets contained product names, error codes, or specific identifiers that benefit from lexical matching. That realization led My AskAI toward experimentation with [hybrid search](https://qdrant.tech/documentation/concepts/hybrid-queries/) as a way to blend semantic similarity with keyword signals, while keeping the operational footprint small. +Over time, the team also learned that semantic search was strong but not universally sufficient, especially when tickets contained product names, error codes, or specific identifiers that benefit from lexical matching. That realization led My AskAI toward experimentation with [hybrid search](https://qdrant.tech/documentation/search/hybrid-queries/) as a way to blend semantic similarity with keyword signals, while keeping the operational footprint small. ## Why My AskAI Chose Qdrant: Scalability, Integrations, and Developer Experience @@ -74,7 +74,7 @@ Once on Qdrant Cloud, My AskAI leaned into a workflow where scaling and day-to-d The ideal infrastructure, as Alex puts it, is the kind you don't have to think about. "I didn't want to have to think about it." -My AskAI also began running customer-specific proofs of concept for [hybrid search](https://qdrant.tech/documentation/concepts/hybrid-queries/), aiming to find the right blend that improved retrieval in the edge cases where semantic-only results were not enough. Before Qdrant, managing hybrid search had required spinning up separate infrastructure on AWS and handling reranking externally. With Qdrant, the team could enable hybrid search per collection and iterate without managing additional systems. +My AskAI also began running customer-specific proofs of concept for [hybrid search](https://qdrant.tech/documentation/search/hybrid-queries/), aiming to find the right blend that improved retrieval in the edge cases where semantic-only results were not enough. Before Qdrant, managing hybrid search had required spinning up separate infrastructure on AWS and handling reranking externally. With Qdrant, the team could enable hybrid search per collection and iterate without managing additional systems. >"Just being able to turn on hybrid search is super useful. It removes that headache and pushes management of hybrid search down to the vendor." diff --git a/qdrant-landing/content/blog/case-study-nyris.md b/qdrant-landing/content/blog/case-study-nyris.md index fc459b29d..977a76ce7 100644 --- a/qdrant-landing/content/blog/case-study-nyris.md +++ b/qdrant-landing/content/blog/case-study-nyris.md @@ -50,14 +50,14 @@ As part of their selection process, Nyris evaluated several critical factors to - **Insert Speed**: Nyris assessed how quickly data could be inserted into the database, including the performance during simultaneous data ingests and query requests. Qdrant excelled in this area, providing the necessary efficiency for their operations. - **Total Cost of Ownership**: Nyris analyzed the infrastructure costs and licensing fees associated with each solution. Qdrant offered a competitive total cost of ownership, making it an economically viable option. - **Data Sovereignty**: The ability to deploy Qdrant in their own clusters was a key aspect for Nyris, ensuring they maintained control over their data and complied with relevant data sovereignty requirements. -- **Dedicated Vector Search Engine:** One of the key advantages of Qdrant, as Lukasson highlights, is its specialization as a dedicated, native vector search engine. "Qdrant, being purpose-built for vector search, can introduce relevant features much faster, like [quantization](https://qdrant.tech/documentation/guides/quantization/), integer8 support, and float32 rescoring. These advancements make searches more precise and cost-effective without sacrificing accuracy—exactly what Nyris needs," said Lukasson. "When optimizing for search accuracy and speed, compromises aren't an option. Just as you wouldn't use a truck to race in Formula 1, we needed a solution designed specifically for vector search, not just a general database with vector search tacked on. With every Qdrant release, we gain new, tailored features that directly enhance our use case.” +- **Dedicated Vector Search Engine:** One of the key advantages of Qdrant, as Lukasson highlights, is its specialization as a dedicated, native vector search engine. "Qdrant, being purpose-built for vector search, can introduce relevant features much faster, like [quantization](https://qdrant.tech/documentation/manage-data/quantization/), integer8 support, and float32 rescoring. These advancements make searches more precise and cost-effective without sacrificing accuracy—exactly what Nyris needs," said Lukasson. "When optimizing for search accuracy and speed, compromises aren't an option. Just as you wouldn't use a truck to race in Formula 1, we needed a solution designed specifically for vector search, not just a general database with vector search tacked on. With every Qdrant release, we gain new, tailored features that directly enhance our use case.” ## Key Benefits of Qdrant in Production Nyris has found several aspects of Qdrant particularly beneficial in their production environment: -- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/guides/security/#granular-access-control-with-jwt) provide enhanced security and performance, critical for safeguarding their data. -- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/guides/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows. +- **Enhanced Security with JWT**: [JSON Web Tokens](https://qdrant.tech/documentation/operations/security/#granular-access-control-with-jwt) provide enhanced security and performance, critical for safeguarding their data. +- **Seamless Scalability**: Qdrant's ability to [scale effortlessly across nodes](https://qdrant.tech/documentation/operations/distributed_deployment/) ensures consistent high performance, even as Nyris's data volume grows. - **Flexible Search Options**: The availability of both graph-based and brute-force search methods offers Nyris the flexibility to tailor the search approach to specific use case requirements. - **Versatile Data Handling**: Qdrant imposes almost no restrictions on data types and vector sizes, allowing Nyris to manage diverse and complex datasets effectively. - **Built with Rust**: The use of [Rust](https://qdrant.tech/articles/why-rust/) ensures superior performance and future-proofing, while its open-source nature allows Nyris to inspect and customize the code as necessary. diff --git a/qdrant-landing/content/blog/case-study-qatech.md b/qdrant-landing/content/blog/case-study-qatech.md index e6789835e..3046f2242 100644 --- a/qdrant-landing/content/blog/case-study-qatech.md +++ b/qdrant-landing/content/blog/case-study-qatech.md @@ -43,7 +43,7 @@ Qdrant’s fast, scalable [vector search](/advanced-search/) enables QA.tech to ## Why QA.tech chose Qdrant for its AI Agent platform -QA.tech’s AI Agents handle high-velocity web actions, requiring efficient real-time operations and scalable infrastructure. The team faced challenges with managing network overhead, CPU load, and the need to store [multiple embeddings](/documentation/concepts/vectors/#multivectors) for different use cases. Qdrant provided the solution to address these issues. +QA.tech’s AI Agents handle high-velocity web actions, requiring efficient real-time operations and scalable infrastructure. The team faced challenges with managing network overhead, CPU load, and the need to store [multiple embeddings](/documentation/manage-data/vectors/#multivectors) for different use cases. Qdrant provided the solution to address these issues. **Reducing Network Overhead with Batch Operations** diff --git a/qdrant-landing/content/blog/case-study-qovery.md b/qdrant-landing/content/blog/case-study-qovery.md index 8dca1fa27..ae6719aea 100644 --- a/qdrant-landing/content/blog/case-study-qovery.md +++ b/qdrant-landing/content/blog/case-study-qovery.md @@ -35,7 +35,7 @@ Qovery’s ambitious vision for the DevOps Copilot ([read more here](https://www ### Seamless Integration of Scalable and Efficient Vector Search -Qovery chose [Qdrant Cloud](https://qdrant.tech/cloud/) after carefully evaluating several options. Romaric Philogène, CEO and co-founder of Qovery, highlighted the importance of open-source credibility, performance, ease of use, and scalability. Qdrant’s native support for [real-time indexing](https://qdrant.tech/documentation/concepts/indexing/) and low-latency queries made it ideal for handling Qovery’s significant data volume and frequency of updates. +Qovery chose [Qdrant Cloud](https://qdrant.tech/cloud/) after carefully evaluating several options. Romaric Philogène, CEO and co-founder of Qovery, highlighted the importance of open-source credibility, performance, ease of use, and scalability. Qdrant’s native support for [real-time indexing](https://qdrant.tech/documentation/manage-data/indexing/) and low-latency queries made it ideal for handling Qovery’s significant data volume and frequency of updates. The integration process was straightforward, with minimal operational overhead, enabling the Qovery team to focus their resources on enhancing the Copilot's capabilities rather than maintaining complex database infrastructure. With its Rust-based architecture, Qdrant delivered the speed, accuracy, and low resource utilization Qovery required. diff --git a/qdrant-landing/content/blog/case-study-sprinklr.md b/qdrant-landing/content/blog/case-study-sprinklr.md index 9043513d7..d0d6635d8 100644 --- a/qdrant-landing/content/blog/case-study-sprinklr.md +++ b/qdrant-landing/content/blog/case-study-sprinklr.md @@ -46,8 +46,8 @@ After evaluating several options of vector DBs, including Pinecone, Weaviate, an - **High Customizability:** Qdrant provided Sprinklr with essential flexibility through high-level abstractions that allowed for extensive customizations. The diverse teams at Sprinklr, working on various GenAI applications, needed a solution that could adapt to different workloads. “The ability to fine-tune configurations at the collection level was crucial for our varied AI applications,” says Sonavane. Qdrant met this need by offering: - **Configuration for high-speed search** that fine-tunes settings for optimal performance. - - [**Quantized vectors**](https://qdrant.tech/documentation/guides/quantization/) for high-dimensional data workloads - - [**Memory map**](https://qdrant.tech/documentation/concepts/storage/#configuring-memmap-storage) for efficient search optimizing memory usage. + - [**Quantized vectors**](https://qdrant.tech/documentation/manage-data/quantization/) for high-dimensional data workloads + - [**Memory map**](https://qdrant.tech/documentation/manage-data/storage/#configuring-memmap-storage) for efficient search optimizing memory usage. - **Speed and Cost Efficiency:** Qdrant provided the best combination of speed and cost, making it the most viable solution for Sprinklr’s needs. “We needed a solution that wouldn’t just meet our performance requirements but also keep costs in check, and Qdrant delivered on both fronts,” says Sonavane. - **Enhanced Monitoring:** Qdrant’s monitoring tools further boosted system efficiency, allowing Sprinklr to maintain high performance across their platforms. @@ -55,7 +55,7 @@ After evaluating several options of vector DBs, including Pinecone, Weaviate, an Sprinklr’s transition to Qdrant was carefully managed, starting with 10% of their workloads before gradually scaling up. The transition was seamless, thanks in part to Qdrant’s configurable [Web UI](https://qdrant.tech/documentation/interfaces/web-ui/), which allowed Sprinklr to fully utilize its capabilities within the existing infrastructure. -“Qdrant’s ability to index [multiple vectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors) simultaneously and retrieve and re-rank with precision brought significant improvements to our workflow,” Sonavane remarks. This feature reduced the need for repeated retrieval processes, significantly improving efficiency. Additionally, Qdrant’s [quantization](https://qdrant.tech/documentation/guides/quantization/) and [memory mapping](https://qdrant.tech/documentation/concepts/storage/#configuring-memmap-storage) features enabled Sprinklr to reduce RAM usage, leading to substantial cost savings. +“Qdrant’s ability to index [multiple vectors](https://qdrant.tech/documentation/manage-data/vectors/#multivectors) simultaneously and retrieve and re-rank with precision brought significant improvements to our workflow,” Sonavane remarks. This feature reduced the need for repeated retrieval processes, significantly improving efficiency. Additionally, Qdrant’s [quantization](https://qdrant.tech/documentation/manage-data/quantization/) and [memory mapping](https://qdrant.tech/documentation/manage-data/storage/#configuring-memmap-storage) features enabled Sprinklr to reduce RAM usage, leading to substantial cost savings. Qdrant now plays a key supportive role in enhancing Sprinklr’s vector search capabilities within its AI-driven applications, which is designed to be cloud- and LLM-agnostic. The platform supports various AI-driven tasks, from retrieval and re-ranking to serving advanced customer experiences. “Retrieval is the foundation of all our AI tasks, and Qdrant’s resilience and speed have made it an integral part of our system,” Sonavane emphasizes. Sprinklr operates [Qdrant as a managed service on AWS](https://qdrant.tech/cloud/), ensuring scalability, reliability, and ease of use. diff --git a/qdrant-landing/content/blog/case-study-visua.md b/qdrant-landing/content/blog/case-study-visua.md index d0a8503da..4dea6e287 100644 --- a/qdrant-landing/content/blog/case-study-visua.md +++ b/qdrant-landing/content/blog/case-study-visua.md @@ -81,6 +81,6 @@ Integrating Qdrant into VISUA's quality control operations has delivered measura #### Expanding Qdrant’s Use Beyond Anomaly Detection -While the primary application of Qdrant is focused on quality control, VISUA's team is actively exploring additional use cases with Qdrant. VISUA's use of Qdrant has inspired new opportunities, notably in content moderation. "The moment we started to experiment with Qdrant, opened up a lot of ideas within the team for new applications,” said Prest on the potential unlocked by Qdrant. For example, this has led them to actively explore the Qdrant [Discovery API](/documentation/concepts/explore/?q=discovery#discovery-api), with an eye on enhancing content moderation processes. +While the primary application of Qdrant is focused on quality control, VISUA's team is actively exploring additional use cases with Qdrant. VISUA's use of Qdrant has inspired new opportunities, notably in content moderation. "The moment we started to experiment with Qdrant, opened up a lot of ideas within the team for new applications,” said Prest on the potential unlocked by Qdrant. For example, this has led them to actively explore the Qdrant [Discovery API](/documentation/search/explore/?q=discovery#discovery-api), with an eye on enhancing content moderation processes. Beyond content moderation, VISUA is set for significant growth by broadening its copyright infringement detection services. As the demand for detecting a wider range of infringements, like unauthorized use of popular characters on merchandise, increases, VISUA plans to expand its technology capabilities. Qdrant will be pivotal in this expansion, enabling VISUA to meet the complex and growing challenges of moderating copyrighted content effectively and ensuring comprehensive protection for brands and creators. \ No newline at end of file diff --git a/qdrant-landing/content/blog/case-study-voiceflow.md b/qdrant-landing/content/blog/case-study-voiceflow.md index f7d7913ed..af9bc572d 100644 --- a/qdrant-landing/content/blog/case-study-voiceflow.md +++ b/qdrant-landing/content/blog/case-study-voiceflow.md @@ -27,7 +27,7 @@ tags: As part of this development, the Voiceflow engineering team was looking for a [vector database](/qdrant-vector-database/) solution to power their RAG setup. They evaluated various vector databases based on several key factors: -- **Performance**: The ability to [handle the scale](/documentation/guides/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently. +- **Performance**: The ability to [handle the scale](/documentation/operations/distributed_deployment/) required by Voiceflow, supporting hundreds of thousands of projects efficiently. - **Metadata**: The capability to tag data and chunks and retrieve based on those values, essential for organizing and accessing specific information swiftly. - **Managed Solution**: The availability of a [managed service](/documentation/cloud/) with automated maintenance, scaling, and security, freeing the team from infrastructure concerns. @@ -82,7 +82,7 @@ Voiceflow achieved significant improvements and efficiencies by leveraging Qdran - **Optimized Performance**: Resolved concerns about retrieval times with a high number of tags by optimizing indexing strategies, achieving efficient performance. - **Minimal Operational Overhead**: Experienced minimal overhead, streamlining their operational processes. - **Future-Ready**: Anticipates further innovation in hybrid search with multi-token attention. -- **Multitenancy Support**: Utilized Qdrant's efficient and [isolated data management](/documentation/guides/multiple-partitions/) to support diverse user needs. +- **Multitenancy Support**: Utilized Qdrant's efficient and [isolated data management](/documentation/manage-data/multitenancy/) to support diverse user needs. Overall, Qdrant's features and infrastructure provided Voiceflow with a stable, scalable, and efficient solution for their data processing and retrieval needs. diff --git a/qdrant-landing/content/blog/case-study-xaver.md b/qdrant-landing/content/blog/case-study-xaver.md index 3151cb39f..0c6521cae 100644 --- a/qdrant-landing/content/blog/case-study-xaver.md +++ b/qdrant-landing/content/blog/case-study-xaver.md @@ -49,7 +49,7 @@ To make this work, Xaver needed a system that could manage knowledge retrieval a ### The solution: Semantic caching, or a two-layer knowledge engine [Qdrant](https://qdrant.tech/documentation/overview/) was selected after extensive evaluation for several reasons: -Xaver’s AI platform includes a “knowledge engine,” an [indexing](https://qdrant.tech/documentation/concepts/indexing/) and [retrieval](https://qdrant.tech/documentation/beginner-tutorials/retrieval-quality/) layer that feeds contextually relevant insights to both automated and human-assisted consultations. It powers two key functions: +Xaver’s AI platform includes a “knowledge engine,” an [indexing](https://qdrant.tech/documentation/manage-data/indexing/) and [retrieval](https://qdrant.tech/documentation/beginner-tutorials/retrieval-quality/) layer that feeds contextually relevant insights to both automated and human-assisted consultations. It powers two key functions: 1. Automated consultation through AI-led sessions via phone, video avatar, messengers, or web chat. diff --git a/qdrant-landing/content/blog/comparing-qdrant-vs-pinecone-vector-databases.md b/qdrant-landing/content/blog/comparing-qdrant-vs-pinecone-vector-databases.md index d01b9faa4..e1efde5e4 100644 --- a/qdrant-landing/content/blog/comparing-qdrant-vs-pinecone-vector-databases.md +++ b/qdrant-landing/content/blog/comparing-qdrant-vs-pinecone-vector-databases.md @@ -27,7 +27,7 @@ Traditional databases, while effective at handling structured data, fall short w - **Indexing Limitations**: Database indexing methods like B-Trees or hash indexes, typically used in relational databases, are inefficient for high-dimensional data and show poor query performance. - **Curse of Dimensionality**: As dimensions increase, data points become sparse, and distance metrics like Euclidean distance lose their effectiveness, leading to poor search query performance. - **Lack of Specialized Algorithms**: Traditional databases do not incorporate advanced algorithms designed to handle high-dimensional data, resulting in slow query processing times. -- **Scalability Challenges**: Managing and querying high-dimensional [vectors](https://qdrant.tech/documentation/concepts/vectors/) require optimized data structures, which traditional databases are not built to handle. +- **Scalability Challenges**: Managing and querying high-dimensional [vectors](https://qdrant.tech/documentation/manage-data/vectors/) require optimized data structures, which traditional databases are not built to handle. - **Storage Inefficiency**: Traditional databases are not optimized for efficiently storing large volumes of high-dimensional data, facing significant challenges in managing space complexity and [retrieval efficiency](https://qdrant.tech/documentation/tutorials/retrieval-quality/). Vector databases address these challenges by efficiently storing and querying high-dimensional vectors. They offer features such as high-dimensional vector storage and retrieval, efficient similarity search, sophisticated indexing algorithms, advanced compression techniques, and integration with various machine learning frameworks. @@ -44,14 +44,14 @@ Qdrant is highly scalable and performant: it can handle billions of vectors effi ### Key Features of Qdrant Vector Database -- **Advanced Similarity Search:** Qdrant supports various similarity [search](https://qdrant.tech/documentation/concepts/search/) metrics like dot product, cosine similarity, Euclidean distance, and Manhattan distance. You can store additional information along with vectors, known as [payload](https://qdrant.tech/documentation/concepts/payload/) in Qdrant terminology. A payload is any JSON formatted data. +- **Advanced Similarity Search:** Qdrant supports various similarity [search](https://qdrant.tech/documentation/search/search/) metrics like dot product, cosine similarity, Euclidean distance, and Manhattan distance. You can store additional information along with vectors, known as [payload](https://qdrant.tech/documentation/manage-data/payload/) in Qdrant terminology. A payload is any JSON formatted data. - **Built Using Rust:** Qdrant is built with Rust, and leverages its performance and efficiency. Rust is famed for its [memory safety](https://arxiv.org/abs/2206.05503) without the overhead of a garbage collector, and rivals C and C++ in speed. -- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/guides/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/guides/multiple-partitions/) where developers can create single collections and partition them using payload. -- **Payload Indexing and Filtering:** Just as Qdrant allows attaching any JSON payload to vectors, it also supports payload indexing and [filtering](https://qdrant.tech/documentation/concepts/filtering/) with a wide range of data types and query conditions, including keyword matching, full-text filtering, numerical ranges, nested object filters, and [geo](https://qdrant.tech/documentation/concepts/filtering/#geo)filtering. +- **Scaling and Multitenancy**: Qdrant supports both vertical and horizontal scaling and uses the Raft consensus protocol for [distributed deployments](https://qdrant.tech/documentation/operations/distributed_deployment/). Developers can run Qdrant clusters with replicas and shards, and seamlessly scale to handle large datasets. Qdrant also supports [multitenancy](https://qdrant.tech/documentation/manage-data/multitenancy/) where developers can create single collections and partition them using payload. +- **Payload Indexing and Filtering:** Just as Qdrant allows attaching any JSON payload to vectors, it also supports payload indexing and [filtering](https://qdrant.tech/documentation/search/filtering/) with a wide range of data types and query conditions, including keyword matching, full-text filtering, numerical ranges, nested object filters, and [geo](https://qdrant.tech/documentation/search/filtering/#geo)filtering. - **Hybrid Search with Sparse Vectors:** Qdrant supports both dense and [sparse vectors](https://qdrant.tech/articles/sparse-vectors/), thereby enabling hybrid search capabilities. Sparse vectors are numerical representations of data where most of the elements are zero. Developers can combine search results from dense and sparse vectors, where sparse vectors ensure that results containing the specific keywords are returned and dense vectors identify semantically similar results. -- **Built-In Vector Quantization:** Qdrant offers three different [quantization](https://qdrant.tech/documentation/guides/quantization/) options to developers to optimize resource usage. Scalar quantization balances accuracy, speed, and compression by converting 32-bit floats to 8-bit integers. Binary quantization, the fastest method, significantly reduces memory usage. Product quantization offers the highest compression, and is perfect for memory-constrained scenarios. -- **Flexible Deployment Options:** Qdrant offers a range of deployment options. Developers can easily set up Qdrant (or Qdrant cluster) [locally](https://qdrant.tech/documentation/quick-start/#download-and-run) using Docker for free. [Qdrant Cloud](https://qdrant.tech/cloud/), on the other hand, is a scalable, managed solution that provides easy access with flexible pricing. Additionally, Qdrant offers [Hybrid Cloud](https://qdrant.tech/hybrid-cloud/) which integrates Kubernetes clusters from cloud, on-premises, or edge, into an enterprise-grade managed service. -- **Security through API Keys, JWT and RBAC:** Qdrant offers developers various ways to [secure](https://qdrant.tech/documentation/guides/security/) their instances. For simple authentication, developers can use API keys (including Read Only API keys). For more granular access control, it offers JSON Web Tokens (JWT) and the ability to build Role-Based Access Control (RBAC). TLS can be enabled to secure connections. Qdrant is also [SOC 2 Type II](https://qdrant.tech/blog/qdrant-soc2-type2-audit/) certified. +- **Built-In Vector Quantization:** Qdrant offers three different [quantization](https://qdrant.tech/documentation/manage-data/quantization/) options to developers to optimize resource usage. Scalar quantization balances accuracy, speed, and compression by converting 32-bit floats to 8-bit integers. Binary quantization, the fastest method, significantly reduces memory usage. Product quantization offers the highest compression, and is perfect for memory-constrained scenarios. +- **Flexible Deployment Options:** Qdrant offers a range of deployment options. Developers can easily set up Qdrant (or Qdrant cluster) [locally](https://qdrant.tech/documentation/quickstart/#download-and-run) using Docker for free. [Qdrant Cloud](https://qdrant.tech/cloud/), on the other hand, is a scalable, managed solution that provides easy access with flexible pricing. Additionally, Qdrant offers [Hybrid Cloud](https://qdrant.tech/hybrid-cloud/) which integrates Kubernetes clusters from cloud, on-premises, or edge, into an enterprise-grade managed service. +- **Security through API Keys, JWT and RBAC:** Qdrant offers developers various ways to [secure](https://qdrant.tech/documentation/operations/security/) their instances. For simple authentication, developers can use API keys (including Read Only API keys). For more granular access control, it offers JSON Web Tokens (JWT) and the ability to build Role-Based Access Control (RBAC). TLS can be enabled to secure connections. Qdrant is also [SOC 2 Type II](https://qdrant.tech/blog/qdrant-soc2-type2-audit/) certified. Additionally, Qdrant integrates seamlessly with popular machine learning frameworks such as [LangChain](https://qdrant.tech/blog/using-qdrant-and-langchain/), LlamaIndex, and Haystack; and Qdrant Hybrid Cloud integrates seamlessly with AWS, DigitalOcean, Google Cloud, Linode, Oracle Cloud, OpenShift, and Azure, among others. @@ -167,7 +167,7 @@ References: - [Pinecone Documentation](https://docs.pinecone.io/) - [Qdrant Documentation](https://qdrant.tech/documentation/) - - If you aren't ready yet, [try out Qdrant locally](/documentation/quick-start/) or sign up for [Qdrant Cloud](https://cloud.qdrant.io/signup). + - If you aren't ready yet, [try out Qdrant locally](/documentation/quickstart/) or sign up for [Qdrant Cloud](https://cloud.qdrant.io/signup). - For more basic information on Qdrant read our [Overview](/documentation/overview/) section or learn more about Qdrant Cloud's [Free Tier](/documentation/cloud/). diff --git a/qdrant-landing/content/blog/cve-2024-2221-response.md b/qdrant-landing/content/blog/cve-2024-2221-response.md index 1079e298f..6a5c83872 100644 --- a/qdrant-landing/content/blog/cve-2024-2221-response.md +++ b/qdrant-landing/content/blog/cve-2024-2221-response.md @@ -45,13 +45,13 @@ guide](/documentation/cloud/authentication/#test-cluster-access). If your Qdrant deployment is local, you do not need an API key. Your next step depends on how you installed Qdrant. For details, read the -[Qdrant Installation](/documentation/guides/installation/) +[Qdrant Installation](/documentation/operations/installation/) guide. #### If you use the Qdrant container or binary Upgrade your deployment. Run the commands in the applicable section of the -[Qdrant Installation](/documentation/guides/installation/) +[Qdrant Installation](/documentation/operations/installation/) guide. The default commands automatically pull the latest version of Qdrant. #### If you use the Qdrant helm chart diff --git a/qdrant-landing/content/blog/cve-2024-3829-response.md b/qdrant-landing/content/blog/cve-2024-3829-response.md index fd66ee12d..0291a3975 100644 --- a/qdrant-landing/content/blog/cve-2024-3829-response.md +++ b/qdrant-landing/content/blog/cve-2024-3829-response.md @@ -45,13 +45,13 @@ guide](https://qdrant.tech/documentation/cloud/quickstart-cloud/#step-2-test-clu If your Qdrant deployment is local, you do not need an API key. Your next step depends on how you installed Qdrant. For details, read the -[Qdrant Installation](https://qdrant.tech/documentation/guides/installation/) +[Qdrant Installation](https://qdrant.tech/documentation/operations/installation/) guide. #### If you use the Qdrant container or binary Upgrade your deployment. Run the commands in the applicable section of the -[Qdrant Installation](https://qdrant.tech/documentation/guides/installation/) +[Qdrant Installation](https://qdrant.tech/documentation/operations/installation/) guide. The default commands automatically pull the latest version of Qdrant. #### If you use the Qdrant helm chart diff --git a/qdrant-landing/content/blog/datatalk-club-podcast-plug.md b/qdrant-landing/content/blog/datatalk-club-podcast-plug.md index ac0297562..08e0d44a6 100644 --- a/qdrant-landing/content/blog/datatalk-club-podcast-plug.md +++ b/qdrant-landing/content/blog/datatalk-club-podcast-plug.md @@ -53,7 +53,7 @@ In the podcast, we addressed the following: - **Model evaluation(LLM)** - Understanding the model at the domain-level for the given use case, supporting required context length and terminology/concept understanding. - **Ingestion pipeline evaluation** - Evaluating factors related to data ingestion and processing such as chunk strategies, chunk size, chunk overlap, and more. - **Retrieval evaluation** - Understanding factors such as average precision, [Distributed cumulative gain](https://en.wikipedia.org/wiki/Discounted_cumulative_gain) (DCG), as well as normalized DCG. -- **Generation evaluation(E2E)** - Establishing guardrails. Evaulating prompts. Evaluating the number of chunks needed to set up the context for generation. +- **Generation evaluation(E2E)** - Establishing guardrails. Evaluating prompts. Evaluating the number of chunks needed to set up the context for generation. ### The recording diff --git a/qdrant-landing/content/blog/decay-functions.md b/qdrant-landing/content/blog/decay-functions.md index 51ae9cfe9..064e2bff1 100644 --- a/qdrant-landing/content/blog/decay-functions.md +++ b/qdrant-landing/content/blog/decay-functions.md @@ -5,7 +5,6 @@ slug: decay-functions # Change this slug to your page slug if needed short_description: Why's and how's of decay functions in Qdrant's relevance score boosting. # Change this description: Understanding decay functions for relevance score boosting. # Change this preview_image: /blog/decay-functions/preview/preview.jpg # Change this -social_preview_image: /blog/decay-functions/preview/social_preview.jpg # Optional image used for link previews title_preview_image: /blog/decay-functions/preview/title.jpg # Optional image used for blog post title date: 2025-09-01T14:55:45+02:00 author: Evgeniya Sukhodolskaya @@ -17,7 +16,7 @@ tags: --- -A problem we've noticed while monitoring the [Qdrant Discord Community](https://discord.gg/d4MPnX3s) is that due to the extensive list of expressions that the [score boosting](https://qdrant.tech/documentation/concepts/search-relevance/#score-boosting) functionality provides, there's room for confusion on how it's supposed to be applied. And that might block you from moving the business logic behind relevance scoring into the Qdrant search engine. We don't want that! +A problem we've noticed while monitoring the [Qdrant Discord Community](https://discord.gg/d4MPnX3s) is that due to the extensive list of expressions that the [score boosting](https://qdrant.tech/documentation/search/search-relevance/#score-boosting) functionality provides, there's room for confusion on how it's supposed to be applied. And that might block you from moving the business logic behind relevance scoring into the Qdrant search engine. We don't want that! In this blog, we'd like to de-spooky-fy the **decay functions** part of the score boosting, or, more precisely: `LinDecayExpression`, `ExpDecayExpression`, and `GaussDecayExpression` -- frequent guests on the Discord *#ask-for-help* channel. @@ -144,13 +143,13 @@ Anything longer than 9 minutes or shorter than 1 minute quickly becomes less rel **Explanation:** Out of all promo codes for different products/events, users will strongly prefer ones uploaded *just now*, as they’re most likely to work. But that relevance drops quickly over time: within a week, it reaches a midpoint of 0.1. After that, if a promo code is still active, it’s a gamble anyway: might work, might not. So old-but-not-expired codes are roughly equally irrelevant. -**Note #5.** For Qdrant [datetime](https://qdrant.tech/documentation/concepts/payload/#datetime) payloads, `scale` should always be provided in seconds! +**Note #5.** For Qdrant [datetime](https://qdrant.tech/documentation/manage-data/payload/#datetime) payloads, `scale` should always be provided in seconds! ### I Don't Know All the Parameters in Advance As you can see, using decay functions in Qdrant's score boosting means you'll have to know the parameters in advance. -What we've seen in our Discord Community quite a few times is that people try to apply decay functions to normalize similarity scores from [prefetches](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries), usually as a way to fuse results from different types of similarity searches. +What we've seen in our Discord Community quite a few times is that people try to apply decay functions to normalize similarity scores from [prefetches](https://qdrant.tech/documentation/search/hybrid-queries/#multi-stage-queries), usually as a way to fuse results from different types of similarity searches. The common question is: @@ -170,10 +169,10 @@ But here's the problem: That 36 might not be a "high" score at all. Maybe your d Now let's see how using decay functions looks in Qdrant. -We'll provide HTTP request examples, but you can use decay functions [analogously in the Python, TypeScript, Rust, Java, C#, and Go clients](/documentation/concepts/search-relevance/#time-based-score-boosting). +We'll provide HTTP request examples, but you can use decay functions [analogously in the Python, TypeScript, Rust, Java, C#, and Go clients](/documentation/search/search-relevance/#time-based-score-boosting). **Note #6.** -Payload variables used within the formula benefit from having [payload indexes](https://qdrant.tech/documentation/concepts/indexing/#payload-index). So, we require you to set up a payload index for any variable used in a formula. +Payload variables used within the formula benefit from having [payload indexes](https://qdrant.tech/documentation/manage-data/indexing/#payload-index). So, we require you to set up a payload index for any variable used in a formula. Let's take our "educational videos in the German language" example and see how it takes shape in Qdrant: @@ -248,7 +247,7 @@ We truly hope this write-up helped untangle things a bit. Now the only thing lef Use the snippets in the article as a starting point and experiment with the relevance score boosting in [Qdrant Cloud](https://qdrant.tech/). We offer a free-forever 1GB cluster: enough to test, tweak, and see how the decay functions behave on your data. -And if you feel like diving deeper into decay functions or score boosting in general, check out our [documentation](/documentation/concepts/search-relevance/#score-boosting), which includes a decay-on-distance example and plenty more to learn from. +And if you feel like diving deeper into decay functions or score boosting in general, check out our [documentation](/documentation/search/search-relevance/#score-boosting), which includes a decay-on-distance example and plenty more to learn from. ### Tell Us What You're Building diff --git a/qdrant-landing/content/blog/facial-recognition.md b/qdrant-landing/content/blog/facial-recognition.md index 343a4fdd1..356db7975 100644 --- a/qdrant-landing/content/blog/facial-recognition.md +++ b/qdrant-landing/content/blog/facial-recognition.md @@ -44,7 +44,7 @@ ___ ## Architecture -**Search Engine & DB:** [**Qdrant**](https://qdrant.tech) stands out as a high-performance [**vector database**](/qdrant-vector-database/) built in Rust, known for its reliability and speed. Its advanced features, such as [**vector visualization**](/documentation/web-ui/) and efficient [**querying**](/documentation/concepts/search/), make it a go-to choice for developers working on embedding-based projects. +**Search Engine & DB:** [**Qdrant**](https://qdrant.tech) stands out as a high-performance [**vector database**](/qdrant-vector-database/) built in Rust, known for its reliability and speed. Its advanced features, such as [**vector visualization**](/documentation/web-ui/) and efficient [**querying**](/documentation/search/search/), make it a go-to choice for developers working on embedding-based projects. ![architecture](/blog/facial-recognition/architecture.png) @@ -121,7 +121,7 @@ If your data is properly embedded, then the visualization tool will appropriatel ## Lessons and Takeaways -Scalability poses challenges when working with large datasets, such as 20,000+ images. Consider optimizations like [**quantization**](/documentation/guides/quantization/) to reduce memory usage or precomputing average embeddings for clusters can significantly minimize storage and computational costs. These strategies ensure the system remains performant as the dataset grows. +Scalability poses challenges when working with large datasets, such as 20,000+ images. Consider optimizations like [**quantization**](/documentation/manage-data/quantization/) to reduce memory usage or precomputing average embeddings for clusters can significantly minimize storage and computational costs. These strategies ensure the system remains performant as the dataset grows. The potential real-world applications of this technology extend far beyond entertainment. Similar systems can be used in security applications for embedding-based facial recognition to secure access to buildings or devices. diff --git a/qdrant-landing/content/blog/hybrid-cloud-airbyte.md b/qdrant-landing/content/blog/hybrid-cloud-airbyte.md index 82a249b67..b99887c78 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-airbyte.md +++ b/qdrant-landing/content/blog/hybrid-cloud-airbyte.md @@ -43,7 +43,7 @@ We put together an end-to-end tutorial to show you how to build a GenAI applicat Learn how to set up a private AI service that addresses customer support issues with high accuracy and effectiveness. By leveraging Airbyte’s data pipelines with Qdrant Hybrid Cloud, you will create a customer support system that is always synchronized with up-to-date knowledge. -[Try the Tutorial](/documentation/tutorials/rag-customer-support-cohere-airbyte-aws/) +[Try the Tutorial](/documentation/examples/rag-customer-support-cohere-airbyte-aws/) #### Documentation: Deploy Qdrant in a Few Clicks diff --git a/qdrant-landing/content/blog/hybrid-cloud-cohere.md b/qdrant-landing/content/blog/hybrid-cloud-cohere.md index 8d825d38e..241390188 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-cohere.md +++ b/qdrant-landing/content/blog/hybrid-cloud-cohere.md @@ -39,7 +39,7 @@ We put together an end-to-end tutorial to show you how to build a GenAI applicat Learn how to set up a private AI service that addresses customer support issues with high accuracy and effectiveness. By leveraging Cohere’s models with Qdrant Hybrid Cloud, you will create a fully private customer support system. -[Try the Tutorial](/documentation/tutorials/rag-customer-support-cohere-airbyte-aws/) +[Try the Tutorial](/documentation/examples/rag-customer-support-cohere-airbyte-aws/) #### Documentation: Deploy Qdrant in a Few Clicks diff --git a/qdrant-landing/content/blog/hybrid-cloud-digitalocean.md b/qdrant-landing/content/blog/hybrid-cloud-digitalocean.md index a607c2079..90a329b4b 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-digitalocean.md +++ b/qdrant-landing/content/blog/hybrid-cloud-digitalocean.md @@ -47,7 +47,7 @@ To get Qdrant Hybrid Cloud setup on DigitalOcean, just follow these steps: We created a tutorial that guides you through setting up and leveraging Qdrant Hybrid Cloud on DigitalOcean for a RAG application. It highlights practical steps to integrate vector search with Jina AI's LLMs, optimizing the generation of high-quality, relevant AI content, while ensuring data sovereignty is maintained throughout. This specific system is tied together via the LlamaIndex framework. -[Try the Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/) +[Try the Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/) For a comprehensive guide, our documentation provides detailed instructions on setting up Qdrant on DigitalOcean. diff --git a/qdrant-landing/content/blog/hybrid-cloud-haystack.md b/qdrant-landing/content/blog/hybrid-cloud-haystack.md index 1a5797a39..d5ac2bdea 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-haystack.md +++ b/qdrant-landing/content/blog/hybrid-cloud-haystack.md @@ -41,7 +41,7 @@ To get you started, we created a comprehensive tutorial that shows how to build Learn how to develop a tutor chatbot from online course materials. You will create a Retrieval Augmented Generation (RAG) pipeline with Haystack for enhanced generative AI capabilities and Qdrant Hybrid Cloud for vector search. By deploying every tool on RedHat OpenShift, you will ensure complete privacy and data sovereignty, whereby no course content leaves your cloud. -[Try the Tutorial](/documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/) +[Try the Tutorial](/documentation/examples/rag-chatbot-red-hat-openshift-haystack/) #### Documentation: Deploy Qdrant in a Few Clicks diff --git a/qdrant-landing/content/blog/hybrid-cloud-jinaai.md b/qdrant-landing/content/blog/hybrid-cloud-jinaai.md index 54a5fe71a..4c3ca5e60 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-jinaai.md +++ b/qdrant-landing/content/blog/hybrid-cloud-jinaai.md @@ -41,7 +41,7 @@ To get you started, we created a comprehensive tutorial that shows how to build Learn how to build an app that retrieves information from PDF user manuals to enhance user experience for companies that sell household appliances. The system will leverage Jina AI embeddings and Qdrant Hybrid Cloud for enhanced generative AI capabilities, while the RAG pipeline will be tied together using the LlamaIndex framework. This example demonstrates how complex tables in PDF documentation can be processed as high quality embeddings with no extra configuration. By introducing Hybrid Search from Qdrant, the RAG functionality is highly accurate. -[Try the Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/) +[Try the Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/) #### Documentation: Deploy Qdrant in a Few Clicks diff --git a/qdrant-landing/content/blog/hybrid-cloud-langchain.md b/qdrant-landing/content/blog/hybrid-cloud-langchain.md index 6a550e1b3..da587d8d4 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-langchain.md +++ b/qdrant-landing/content/blog/hybrid-cloud-langchain.md @@ -41,7 +41,7 @@ To get you started, we’ve put together a tutorial that shows how to create nex We created a comprehensive tutorial to show how you can build a RAG-based system with Qdrant Hybrid Cloud, LangChain and Cohere’s embeddings. This use case is focused on building a question-answering system for internal corporate employee onboarding. -[Try the Tutorial](/documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/) +[Try the Tutorial](/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/) #### Documentation: Deploy Qdrant in a Few Clicks diff --git a/qdrant-landing/content/blog/hybrid-cloud-launch-partners.md b/qdrant-landing/content/blog/hybrid-cloud-launch-partners.md index 5a862e767..6c9b9fd8e 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-launch-partners.md +++ b/qdrant-landing/content/blog/hybrid-cloud-launch-partners.md @@ -34,49 +34,49 @@ Together with our launch partners, we created in-depth tutorials and use cases f > This tutorial shows how to build a private AI customer support system using Cohere's AI models on AWS, Airbyte, and Qdrant Hybrid Cloud for efficient and secure query automation. -[View Tutorial](/documentation/tutorials/rag-customer-support-cohere-airbyte-aws/) +[View Tutorial](/documentation/examples/rag-customer-support-cohere-airbyte-aws/) **RAG System for Employee Onboarding** with Qdrant Hybrid Cloud, Oracle Cloud Infrastructure (OCI), Cohere, and LangChain > This tutorial demonstrates how to use Oracle Cloud Infrastructure (OCI) for a secure setup that integrates Cohere's language models with Qdrant Hybrid Cloud, using LangChain to orchestrate natural language search for corporate documents, enhancing resource discovery and onboarding. -[View Tutorial](/documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/) +[View Tutorial](/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/) **Hybrid Search for Product PDF Manuals** with Qdrant Hybrid Cloud, LlamaIndex, and JinaAI > Create a RAG-based chatbot that enhances customer support by parsing product PDF manuals using Qdrant Hybrid Cloud, LlamaIndex, and JinaAI, with DigitalOcean as the cloud host. This tutorial will guide you through the setup and integration process, enabling your system to deliver precise, context-aware responses for household appliance inquiries. -[View Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/) +[View Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/) **Region-Specific RAG System for Contract Management** with Qdrant Hybrid Cloud, Aleph Alpha, and STACKIT > Learn how to streamline contract management with a RAG-based system in this tutorial, which utilizes Aleph Alpha’s embeddings and a region-specific cloud setup. Hosted on STACKIT with Qdrant Hybrid Cloud, this solution ensures secure, GDPR-compliant storage and processing of data, ideal for businesses with intensive contractual needs. -[View Tutorial](/documentation/tutorials/rag-contract-management-stackit-aleph-alpha/) +[View Tutorial](/documentation/examples/rag-contract-management-stackit-aleph-alpha/) **Movie Recommendation System** with Qdrant Hybrid Cloud and OVHcloud > Discover how to build a recommendation system with our guide on collaborative filtering, using sparse vectors and the Movielens dataset. -[View Tutorial](/documentation/tutorials/recommendation-system-ovhcloud/) +[View Tutorial](/documentation/examples/recommendation-system-ovhcloud/) **Private RAG Information Extraction Engine** with Qdrant Hybrid Cloud and Vultr using DSPy and Ollama > This tutorial teaches you how to handle and structure private documents with large unstructured data. Learn to use DSPy for information extraction, run your LLM with Ollama on Vultr, and manage data with Qdrant Hybrid Cloud on Vultr, perfect for regulated environments needing data privacy. -[View Tutorial](/documentation/tutorials/rag-chatbot-vultr-dspy-ollama/) +[View Tutorial](/documentation/examples/rag-chatbot-vultr-dspy-ollama/) **RAG System That Chats with Blog Contents** with Qdrant Hybrid Cloud and Scaleway using LangChain. > Build a RAG system that combines blog scanning with the capabilities of semantic search. RAG enhances the generation of answers by retrieving relevant documents to aid the question-answering process. This setup showcases the integration of advanced search and AI language processing to improve information retrieval and generation tasks. -[View Tutorial](/documentation/tutorials/rag-chatbot-scaleway/) +[View Tutorial](/documentation/examples/rag-chatbot-scaleway/) **Private Chatbot for Interactive Learning** with Qdrant Hybrid Cloud and Red Hat OpenShift using Haystack. > In this tutorial, you will build a chatbot without public internet access. The goal is to keep sensitive data secure and isolated. Your RAG system will be built with Qdrant Hybrid Cloud on Red Hat OpenShift, leveraging Haystack for enhanced generative AI capabilities. This tutorial especially explores how this setup ensures that not a single data point leaves the environment. -[View Tutorial](/documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/) +[View Tutorial](/documentation/examples/rag-chatbot-red-hat-openshift-haystack/) #### Supporting Documentation diff --git a/qdrant-landing/content/blog/hybrid-cloud-llamaindex.md b/qdrant-landing/content/blog/hybrid-cloud-llamaindex.md index 757b4aa49..d1735fc5b 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-llamaindex.md +++ b/qdrant-landing/content/blog/hybrid-cloud-llamaindex.md @@ -41,7 +41,7 @@ To get you started, we created a comprehensive tutorial that shows how to build Use this end-to-end tutorial to create a system that retrieves information from complex user manuals in PDF format to enhance user experience for companies that sell household appliances. You will build a RAG pipeline with LlamaIndex leveraging Qdrant Hybrid Cloud for enhanced generative AI capabilities. The LlamaIndex integration shows how complex tables inside of items’ PDF documents can be processed via hybrid vector search with no additional configuration. -[Try the Tutorial](/documentation/tutorials/hybrid-search-llamaindex-jinaai/) +[Try the Tutorial](/documentation/examples/hybrid-search-llamaindex-jinaai/) #### Documentation: Deploy Qdrant in a Few Clicks diff --git a/qdrant-landing/content/blog/hybrid-cloud-oracle-cloud-infrastructure.md b/qdrant-landing/content/blog/hybrid-cloud-oracle-cloud-infrastructure.md index 67760e651..d9eecc4d1 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-oracle-cloud-infrastructure.md +++ b/qdrant-landing/content/blog/hybrid-cloud-oracle-cloud-infrastructure.md @@ -37,7 +37,7 @@ Deploying Qdrant Hybrid Cloud on OCI facilitates vector search in production env We created a comprehensive tutorial to show how to leverage the benefits of Qdrant Hybrid Cloud on OCI and build AI applications with a focus on data sovereignty. This use case is focused on building a RAG system for FAQ, leveraging the strengths of Qdrant Hybrid Cloud for vector search, Oracle Cloud Infrastructure (OCI) as a managed Kubernetes provider, Cohere models for embedding, and LangChain as a framework. -[Try the Tutorial](/documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/) +[Try the Tutorial](/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/) Deploying Qdrant Hybrid Cloud on Oracle Cloud Infrastructure only takes a few minutes due to the seamless Kubernetes-native integration. You can get started by following these three steps: diff --git a/qdrant-landing/content/blog/hybrid-cloud-ovhcloud.md b/qdrant-landing/content/blog/hybrid-cloud-ovhcloud.md index 458c08af0..dfc357a9f 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-ovhcloud.md +++ b/qdrant-landing/content/blog/hybrid-cloud-ovhcloud.md @@ -37,7 +37,7 @@ Through the seamless integration between Qdrant Hybrid Cloud and OVHcloud, devel To show how Qdrant Hybrid Cloud deployed on OVHcloud allows developers to leverage the benefits of an AI use case that is completely run within the existing infrastructure, we put together a comprehensive use case tutorial. This tutorial guides you through creating a recommendation system using collaborative filtering and sparse vectors with Qdrant Hybrid Cloud on OVHcloud. It employs the Movielens dataset for practical application, providing insights into building efficient, scalable recommendation engines suitable for developers and data scientists looking to leverage advanced vector search technologies within a secure, GDPR-compliant European cloud infrastructure. -[Try the Tutorial](/documentation/tutorials/recommendation-system-ovhcloud/) +[Try the Tutorial](/documentation/examples/recommendation-system-ovhcloud/) #### Get Started Today and Leverage the Benefits of Qdrant Hybrid Cloud diff --git a/qdrant-landing/content/blog/hybrid-cloud-red-hat-openshift.md b/qdrant-landing/content/blog/hybrid-cloud-red-hat-openshift.md index 4b72b0d24..a75942da9 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-red-hat-openshift.md +++ b/qdrant-landing/content/blog/hybrid-cloud-red-hat-openshift.md @@ -52,7 +52,7 @@ To get started, we created a comprehensive tutorial that shows how to build next In this tutorial, you will build a chatbot without public internet access. The goal is to keep sensitive data secure and isolated. Your RAG system will be built with Qdrant Hybrid Cloud on Red Hat OpenShift, leveraging Haystack for enhanced generative AI capabilities. This tutorial especially explores how this setup ensures that not a single data point leaves the environment. -[Try the Tutorial](/documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/) +[Try the Tutorial](/documentation/examples/rag-chatbot-red-hat-openshift-haystack/) #### Documentation: Deploy Qdrant in a Few Clicks diff --git a/qdrant-landing/content/blog/hybrid-cloud-scaleway.md b/qdrant-landing/content/blog/hybrid-cloud-scaleway.md index bd9c8d56b..90c529fb1 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-scaleway.md +++ b/qdrant-landing/content/blog/hybrid-cloud-scaleway.md @@ -29,7 +29,7 @@ RAG applications often rely on sensitive or proprietary internal data, emphasizi We created a tutorial that guides you through setting up and leveraging Qdrant Hybrid Cloud on Scaleway for a RAG application, providing insights into efficiently managing data within a secure, sovereign framework. It highlights practical steps to integrate vector search with LLMs, optimizing the generation of high-quality, relevant AI content, while ensuring data sovereignty is maintained throughout. -[Try the Tutorial](/documentation/tutorials/rag-chatbot-scaleway/) +[Try the Tutorial](/documentation/examples/rag-chatbot-scaleway/) #### The Benefits of Running Qdrant Hybrid Cloud on Scaleway diff --git a/qdrant-landing/content/blog/hybrid-cloud-stackit.md b/qdrant-landing/content/blog/hybrid-cloud-stackit.md index b2f45e117..0f9272001 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-stackit.md +++ b/qdrant-landing/content/blog/hybrid-cloud-stackit.md @@ -33,7 +33,7 @@ Qdrant Hybrid Cloud is the first managed vector database that can be deployed in To demonstrate the power of Qdrant Hybrid Cloud on STACKIT, we’ve developed a comprehensive tutorial showcasing how to build secure, AI-driven applications focusing on data sovereignty. This tutorial specifically shows how to build a contract management platform that enables users to upload documents (PDF or DOCx), which are then segmented for searchable access. Designed with multitenancy, users can only access their team or organization's documents. It also features custom sharding for location-specific document storage. Beyond search, the application offers rephrasing of document excerpts for clarity to those without context. -[Try the Tutorial](/documentation/tutorials/rag-contract-management-stackit-aleph-alpha/) +[Try the Tutorial](/documentation/examples/rag-contract-management-stackit-aleph-alpha/) #### Start Using Qdrant with STACKIT diff --git a/qdrant-landing/content/blog/hybrid-cloud-vultr.md b/qdrant-landing/content/blog/hybrid-cloud-vultr.md index 8f67e69ea..a1058a904 100644 --- a/qdrant-landing/content/blog/hybrid-cloud-vultr.md +++ b/qdrant-landing/content/blog/hybrid-cloud-vultr.md @@ -45,7 +45,7 @@ We've compiled an in-depth guide for leveraging Qdrant Hybrid Cloud on Vultr to This tutorial outlines creating a personalized AI assistant using Qdrant Hybrid Cloud on Vultr, incorporating advanced vector search to power dynamic, interactive experiences. We will develop a RAG pipeline powered by DSPy and detail how to maintain data privacy within your Vultr environment. -[Try the Tutorial](/documentation/tutorials/rag-chatbot-vultr-dspy-ollama/) +[Try the Tutorial](/documentation/examples/rag-chatbot-vultr-dspy-ollama/) #### Documentation: Effortless Deployment with Qdrant diff --git a/qdrant-landing/content/blog/legal-tech-builders-guide.md b/qdrant-landing/content/blog/legal-tech-builders-guide.md index b7dec940f..cd253a6ea 100644 --- a/qdrant-landing/content/blog/legal-tech-builders-guide.md +++ b/qdrant-landing/content/blog/legal-tech-builders-guide.md @@ -85,7 +85,7 @@ Leveraging Late-Interaction Models for Rich Documents Traditional OCR pipelines can add complexity and create accuracy challenges. But late-interaction models simplify the ingestion pipeline by running at the reranking stage. -Models like ([ColPali](https://qdrant.tech/blog/qdrant-colpali/) and ColQwen) bypass traditional OCR pipelines, directly processing images of complex documents. They enhance accuracy by maintaining original layouts and contextual integrity, simplifying your retrieval pipelines. The tradeoff is a heavier application, but these challenges can be addressed with further [optimization](https://qdrant.tech/documentation/guides/optimize/)*.* +Models like ([ColPali](https://qdrant.tech/blog/qdrant-colpali/) and ColQwen) bypass traditional OCR pipelines, directly processing images of complex documents. They enhance accuracy by maintaining original layouts and contextual integrity, simplifying your retrieval pipelines. The tradeoff is a heavier application, but these challenges can be addressed with further [optimization](https://qdrant.tech/documentation/operations/optimize/)*.* #### Enabling highly granular accuracy for complex legal searches @@ -125,7 +125,7 @@ final_results = reranked[:5] Not every clause is created equal. Legal professionals often care more about specific provisions, jurisdictions, or case types, for example. -Qdrant's [Score Boosting Reranker](/documentation/concepts/search-relevance/#score-boosting) lets you integrate domain-specific logic (e.g., jurisdiction or recent cases) directly into search rankings, ensuring results align precisely with legal business rules. +Qdrant's [Score Boosting Reranker](/documentation/search/search-relevance/#score-boosting) lets you integrate domain-specific logic (e.g., jurisdiction or recent cases) directly into search rankings, ensuring results align precisely with legal business rules. ```json POST /collections/legal-docs/points/query @@ -163,13 +163,13 @@ Legal datasets are growing, and so are the compute bills. From GPU acceleration * [GPU indexing](https://qdrant.tech/blog/qdrant-1.13.x/) accelerates indexing by up to 10x compared to CPU methods, offering vendor-agnostic compatibility with modern GPUs via Vulkan API. -* [Vector quantization](https://qdrant.tech/documentation/guides/quantization/) compresses embeddings, significantly reducing memory and operational costs. It results in lower accuracy, so carefully consider this option. For example, [LawMe](http://qdrant.tech/blog/case-study-lawme), a Qdrant user, uses Binary Quantization to cost-effectively add more data for its AI Legal Assistants. +* [Vector quantization](https://qdrant.tech/documentation/manage-data/quantization/) compresses embeddings, significantly reducing memory and operational costs. It results in lower accuracy, so carefully consider this option. For example, [LawMe](http://qdrant.tech/blog/case-study-lawme), a Qdrant user, uses Binary Quantization to cost-effectively add more data for its AI Legal Assistants. ### Getting Started: Choosing Your Search Infrastructure #### Deploy in private, cloud, or hybrid environments without sacrificing control -No matter your stage—prototype or production—your stack will have to meet both engineering and compliance needs. Qdrant supports flexible deployment strategies, including [managed cloud](https://qdrant.tech/cloud/) and [hybrid cloud](https://qdrant.tech/hybrid-cloud/), along with open-source solutions via [Docker](https://qdrant.tech/documentation/quick-start/), enabling easy scaling and secure management of legal data. +No matter your stage—prototype or production—your stack will have to meet both engineering and compliance needs. Qdrant supports flexible deployment strategies, including [managed cloud](https://qdrant.tech/cloud/) and [hybrid cloud](https://qdrant.tech/hybrid-cloud/), along with open-source solutions via [Docker](https://qdrant.tech/documentation/quickstart/), enabling easy scaling and secure management of legal data. #### Build and iterate quickly with responsive support and built-in tooling diff --git a/qdrant-landing/content/blog/multi-vector-course-release.md b/qdrant-landing/content/blog/multi-vector-course-release.md new file mode 100644 index 000000000..94ece7c56 --- /dev/null +++ b/qdrant-landing/content/blog/multi-vector-course-release.md @@ -0,0 +1,66 @@ +--- +title: "Master Multi-Vector Search With Qdrant" +draft: false +slug: multi-vector-search-course +short_description: "Go beyond single-vector embeddings. Our new advanced course covers ColBERT, ColPali, MaxSim, and production-grade multi-vector pipelines in Qdrant." +description: "Go beyond single-vector embeddings. Our new advanced course covers ColBERT, ColPali, MaxSim, and production-grade multi-vector pipelines in Qdrant." +preview_image: /blog/multi-vector-course-release/hero.png +social_preview_image: /blog/multi-vector-course-release/hero.png +date: 2026-03-24 +author: Neil Kanungo +featured: true +tags: + - qdrant-course + - multi-vector-search + - colbert + - colpali + - certification +--- + +Most vector search tutorials stop at single-vector embeddings: one document, one vector, one similarity score. That works for demos. It falls apart when your retrieval pipeline needs to capture fine-grained token-level interactions across text, images, and PDFs at production scale. + +Until now, engineers who wanted to go deeper had to piece together scattered papers, blog posts, and half-documented repos. There was no structured, hands-on resource that connected the theory of late interaction models to real implementation in a production search engine. + +We built one. + +## Introducing the Multi-Vector Search Course + +[Qdrant's Multi-Vector Search Course](https://qdrant.tech/course/multi-vector-search/) is a free, advanced course created by [Kacper Łukawski](https://www.linkedin.com/in/kacperlukawski/). Kacper designed this course to fill a real gap in the developer community: practical, production-focused education on multi-vector retrieval that goes well beyond "here's how embeddings work." + +This is an advanced course! It's built for ML engineers, backend engineers, and search engineers who already understand vector search fundamentals and want to master what comes next. + +## What You'll Learn + +The course is organized into four modules, each taking roughly one to two hours: + +**Module 0: Setup.** Configure a Qdrant Cloud or local instance and install Python dependencies. + +**Module 1: Text Multi-Vectors.** Understand the late interaction paradigm, learn the MaxSim distance metric, explore real use cases and challenges, and implement ColBERT-based search with Qdrant. + +**Module 2: Multi-Modal Search.** Apply multi-vector representations to images and PDFs using ColPali. Explore model variants in the ColPali family and use visual interpretability for debugging retrieval results. + +**Module 3: Optimization and Evaluation.** Master vector quantization, pooling techniques, and MUVERA indexing for memory-efficient search at billion scale. Build multi-stage retrieval pipelines with Qdrant's Universal Query API and evaluate with industry-standard metrics (Recall@k, NDCG, MRR). + +The course wraps up with a final project: build your own production-ready multi-modal search system from scratch. + +## How It Works + +Every module combines video lessons from the Qdrant team with hands-on Google Colab notebooks. You watch, you build, you evaluate. The progressive structure means each module builds directly on the previous one, so you finish with a complete, working pipeline rather than a collection of disconnected concepts. + +## Earn a Qdrant Certification + +Complete the course and pass the certification exam to earn a shareable Qdrant Multi-Vector Search certificate. It's a concrete way to demonstrate that you can design, implement, and optimize multi-vector retrieval pipelines in production. + +Certifications are available through [Qdrant Academy](https://qdrant.tech/course/multi-vector-search/certification/). + +## Free Swag for the First 20 Certified + +Dive in now: the **first 20 people** who complete their Multi-Vector Search certification and post on LinkedIn with the hashtag **#QdrantCertified** will receive free Qdrant swag. Share your certificate, tag us, and we'll reach out. + +## Start Now + +The course is free, self-paced, and available today. If you've been looking for a structured path from "I understand embeddings" to "I can build and evaluate multi-vector retrieval at scale," this is it. + +[Start the Multi-Vector Search Course](https://qdrant.tech/course/multi-vector-search/) + +Questions? Join our [Discord community](https://discord.gg/qdrant) where Qdrant experts collaborate. diff --git a/qdrant-landing/content/blog/pgvector-tradeoffs.md b/qdrant-landing/content/blog/pgvector-tradeoffs.md index 38be7400a..87c343ba5 100644 --- a/qdrant-landing/content/blog/pgvector-tradeoffs.md +++ b/qdrant-landing/content/blog/pgvector-tradeoffs.md @@ -40,19 +40,19 @@ Here are the six conditions: **1. Your vector dataset is under ~1M vectors.** The community's empirical ceiling is around 10M, but the comfortable range is much lower. Above 1M you'll start hitting index-build times, memory pressure, and recall degradation under load. -This threshold is also easier to hit than you'd expect — especially if you're working with **multivectors**. Techniques like ColBERT-style late interaction generate one embedding *per token* rather than one per document, so a corpus of 100K documents can easily balloon into tens of millions of vectors overnight. Qdrant, by contrast, has [native multivector support](/documentation/concepts/vectors/#multivectors) with dedicated documentation and query APIs built around it. +This threshold is also easier to hit than you'd expect — especially if you're working with **multivectors**. Techniques like ColBERT-style late interaction generate one embedding *per token* rather than one per document, so a corpus of 100K documents can easily balloon into tens of millions of vectors overnight. Qdrant, by contrast, has [native multivector support](/documentation/manage-data/vectors/#multivectors) with dedicated documentation and query APIs built around it. **2. You don't need accurate metadata filtering.** If every search is against the full collection, post-filtering won't limit you. But the moment you need to scope searches to a user, tenant, category, or any selective predicate, pgvector generates unnecessary search overhead. > *"I think the most relevant weakness for pgvector is the lack of 'proper' prefiltering on metadata while leveraging the vector index."* -Qdrant takes an entirely different approach to filtering. Specifically, Qdrant utilizes a [filterable HNSW](/documentation/concepts/indexing/#filtrable-index) which lets you traverse the nearest-neighbor graph while maintaining metadata filters. +Qdrant takes an entirely different approach to filtering. Specifically, Qdrant utilizes a [filterable HNSW](/documentation/manage-data/indexing/#filtrable-index) which lets you traverse the nearest-neighbor graph while maintaining metadata filters. **3. Your embeddings are tightly coupled to relational data.** If vectors are just an attribute of a row (e.g., a product description embedding alongside the product), colocation helps. If vectors are first-class entities, the argument for co-location weakens. **4. You don't need hybrid search.** While pgvector supports dense vector similarity search via HNSW, the Postgres extension ecosystem still lacks a high-quality BM25 implementation — a critical component for hybrid search. -Postgres *does* have full-text search via `tsvector`/`tsquery`, and it's excellent for what it does. But that's lexical search — exact term matches, stemming, and stop words. BM25 is a probabilistic model that considers term frequency, inverse document frequency, and document length. Qdrant supports [native BM25 via sparse vectors](/documentation/concepts/vectors/#sparse-vectors). They're not the same thing. +Postgres *does* have full-text search via `tsvector`/`tsquery`, and it's excellent for what it does. But that's lexical search — exact term matches, stemming, and stop words. BM25 is a probabilistic model that considers term frequency, inverse document frequency, and document length. Qdrant supports [native BM25 via sparse vectors](/documentation/manage-data/vectors/#sparse-vectors). They're not the same thing. **5. Postgres is already doing the heavy lifting for your business logic.** You have existing transactions, schemas, and ACID guarantees that matter. Adding a second data store splits that concern. If Postgres is truly central, the operational argument for colocation is real. @@ -75,7 +75,7 @@ That's three conditions gone before you've even thought about scale. People are When the conditions above don't all hold, dedicated vector stores offer concrete advantages: - **Efficient metadata filtering** — pre-filter on metadata fields before computing similarity, avoiding wasted work on irrelevant vectors -- **Native hybrid search** — combine dense similarity and BM25 keyword matching in a single query with [reciprocal rank fusion](/documentation/concepts/hybrid-queries/) +- **Native hybrid search** — combine dense similarity and BM25 keyword matching in a single query with [reciprocal rank fusion](/documentation/search/hybrid-queries/) - **Scale beyond 10M vectors** — purpose-built sharding, distributed indexing, and memory management - **Decoupled architecture** — scale, optimize, and evolve your search layer independently of your relational database diff --git a/qdrant-landing/content/blog/product-ui-changes.md b/qdrant-landing/content/blog/product-ui-changes.md index 1c6ebee4e..a4c5d999a 100644 --- a/qdrant-landing/content/blog/product-ui-changes.md +++ b/qdrant-landing/content/blog/product-ui-changes.md @@ -59,7 +59,7 @@ When looking at the overview of your cluster, we’ve added new tabs with an imp * **Logs**: get a real-time window into what’s happening inside cluste for transparency, diagnostics, and control (especially important during debugging, performance tuning, or infrastructure troubleshooting\!) * **Backups:** View snapshots of your vector data and metadata that can be used to restore your collections in case of data loss, migration, or rollbacks (not available on free clusters) * **Configuration**: Check your collection defaults and add advanced optimizations (after reading Docs of course) - * For example, we advise against setting up a ton of different collections. Instead segment with [payloads](https://qdrant.tech/documentation/concepts/payload/). + * For example, we advise against setting up a ton of different collections. Instead segment with [payloads](https://qdrant.tech/documentation/manage-data/payload/). When viewing the details of your clusters, you can now view the Cluster UI Dashboard regardless of where you are, and also have easier access to tutorials and resources. @@ -74,7 +74,7 @@ Next we’ve done a major overhaul to the “Get Started” page. Our goal is to ![Image of Get Started webpage](/blog/product-ui-changes/get-started-overview.jpg) **Explore Your Data or Start with Samples** -You’ll see immediately pertinent information to help you get the most out of Qdrant quickly, including the [Cloud Quickstart guide](https://qdrant.tech/documentation/quickstart-cloud/), and resources to help you get your data into Qdrant, or use sample data. +You’ll see immediately pertinent information to help you get the most out of Qdrant quickly, including the [Cloud Quickstart guide](https://qdrant.tech/documentation/cloud-quickstart/), and resources to help you get your data into Qdrant, or use sample data. Learn about the different ways to connect to your cluster, use the Qdrant API, try out sample data, and our personal favorite, use the Qdrant Cluster UI to view your collection data and access tutorials. diff --git a/qdrant-landing/content/blog/qdrant-1.10.x.md b/qdrant-landing/content/blog/qdrant-1.10.x.md index 76d45a73c..0fd9e7e4c 100644 --- a/qdrant-landing/content/blog/qdrant-1.10.x.md +++ b/qdrant-landing/content/blog/qdrant-1.10.x.md @@ -31,14 +31,14 @@ You can now configure the Query API request with the following parameters: |Parameter|Description| |-|-| |no parameter|Returns points by `id`| -|`nearest`|Queries nearest neighbors ([Search](/documentation/concepts/search/))| -|`fusion`|Fuses sparse/dense prefetch queries ([Hybrid Search](/documentation/concepts/hybrid-queries/#hybrid-search))| -|`discover`|Queries `target` with added `context` ([Discovery](/documentation/concepts/explore/#discovery-api))| -|`context` |No target with `context` only ([Context](/documentation/concepts/explore/#context-search))| -|`recommend`|Queries against `positive`/`negative` examples. ([Recommendation](/documentation/concepts/explore/#recommendation-api))| -|`order_by`|Orders results by [payload field](/documentation/concepts/hybrid-queries/#re-ranking-with-payload-values)| +|`nearest`|Queries nearest neighbors ([Search](/documentation/search/search/))| +|`fusion`|Fuses sparse/dense prefetch queries ([Hybrid Search](/documentation/search/hybrid-queries/#hybrid-search))| +|`discover`|Queries `target` with added `context` ([Discovery](/documentation/search/explore/#discovery-api))| +|`context` |No target with `context` only ([Context](/documentation/search/explore/#context-search))| +|`recommend`|Queries against `positive`/`negative` examples. ([Recommendation](/documentation/search/explore/#recommendation-api))| +|`order_by`|Orders results by [payload field](/documentation/search/hybrid-queries/#re-ranking-with-payload-values)| -For example, you can configure Query API to run [Discovery search](/documentation/concepts/explore/#discovery-api). Let's see how that looks: +For example, you can configure Query API to run [Discovery search](/documentation/search/explore/#discovery-api). Let's see how that looks: ```http POST collections/{collection_name}/points/query @@ -57,7 +57,7 @@ POST collections/{collection_name}/points/query } ``` -We will be publishing code samples in [docs](/documentation/concepts/hybrid-queries/) and our new [API specification](http://api.qdrant.tech).
*If you need additional support with this new method, our [Discord](https://qdrant.to/discord) on-call engineers can help you.* +We will be publishing code samples in [docs](/documentation/search/hybrid-queries/) and our new [API specification](http://api.qdrant.tech).
*If you need additional support with this new method, our [Discord](https://qdrant.to/discord) on-call engineers can help you.* ### Native Hybrid Search Support @@ -221,7 +221,7 @@ await client.QueryAsync( Query API can now pre-fetch vectors for requests, which means you can run queries sequentially within the same API call. There are a lot of options here, so you will need to define a strategy to merge these requests using new parameters. For example, you can now include **rescoring within Hybrid Search**, which can open the door to strategies like iterative refinement via matryoshka embeddings. -*To learn more about this, read the [Query API documentation](/documentation/concepts/search/#query-api).* +*To learn more about this, read the [Query API documentation](/documentation/search/search/#query-api).* ## Inverse Document Frequency [IDF] @@ -535,11 +535,11 @@ await client.QueryAsync( ``` **Note:** *The multivector feature is not only useful for ColBERT; it can also be used in other ways.*
-For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/concepts/search/#grouping-api) method. +For instance, in e-commerce, you can use multi-vector to store multiple images of the same item. This serves as an alternative to the [group-by](/documentation/search/search/#grouping-api) method. ## Sparse Vectors Compression -In version 1.9, we introduced the `uint8` [vector datatype](/documentation/concepts/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere. +In version 1.9, we introduced the `uint8` [vector datatype](/documentation/manage-data/vectors/#datatypes) for sparse vectors, in order to support pre-quantized embeddings from companies like JinaAI and Cohere. This time, we are introducing a new datatype **for both sparse and dense vectors**, as well as a different way of **storing** these vectors. **Datatype:** Sparse and dense vectors were previously represented in larger `float32` values, but now they can be turned to the `float16`. `float16` vectors have a lower precision compared to `float32`, which means that there is less numerical accuracy in the vector values - but this is negligible for practical use cases. @@ -669,7 +669,7 @@ documentation, making it easier to navigate and find the information you need. ## S3 Snapshot Storage -Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/concepts/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss. +Qdrant **Collections**, **Shards** and **Storage** can be backed up with [Snapshots](/documentation/operations/snapshots/) and saved in case of data loss or other data transfer purposes. These snapshots can be quite large and the resources required to maintain them can result in higher costs. AWS S3 and other S3-compatible implementations like [min.io](https://min.io/) is a great low-cost alternative that can hold snapshots without incurring high costs. It is globally reliable, scalable and resistant to data loss. You can configure S3 storage settings in the [config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml), specifically with `snapshots_storage`. @@ -697,7 +697,7 @@ storage: secret_key: your_secret_key_here ``` -*Read more about [S3 snapshot storage](/documentation/concepts/snapshots/#s3) and [configuration](/documentation/guides/configuration/).* +*Read more about [S3 snapshot storage](/documentation/operations/snapshots/#s3) and [configuration](/documentation/operations/configuration/).* This integration allows for a more convenient distribution of snapshots. Users of **any S3-compatible object storage** can now benefit from other platform services, such as automated workflows and disaster recovery options. S3's encryption and access control ensure secure storage and regulatory compliance. Additionally, S3 supports performance optimization through various storage classes and efficient data transfer methods, enabling quick and effective snapshot retrieval and management. diff --git a/qdrant-landing/content/blog/qdrant-1.11.x.md b/qdrant-landing/content/blog/qdrant-1.11.x.md index 1a6dedfa5..7c0ed02ad 100644 --- a/qdrant-landing/content/blog/qdrant-1.11.x.md +++ b/qdrant-landing/content/blog/qdrant-1.11.x.md @@ -36,7 +36,7 @@ New Web UI Tools:
Before we dive into the specifics of our optimizations, let's first go over Multitenancy. This is one of our most significant features, [best used for scaling and data isolation](https://qdrant.tech/articles/multitenancy/). -If you’re using Qdrant to manage data for multiple users, regions, or workspaces (tenants), we suggest setting up a [multitenant environment](/documentation/guides/multiple-partitions/). This approach keeps all tenant data in a single global collection, with points separated and isolated by their payload. +If you’re using Qdrant to manage data for multiple users, regions, or workspaces (tenants), we suggest setting up a [multitenant environment](/documentation/manage-data/multitenancy/). This approach keeps all tenant data in a single global collection, with points separated and isolated by their payload. To avoid slow and unnecessary indexing, it’s better to create an index for each relevant payload rather than indexing the entire collection globally. Since some data is indexed more frequently, you can focus on building indexes for specific regions, workspaces, or users. @@ -156,7 +156,7 @@ await client.CreatePayloadIndexAsync( As a result, the storage structure will be organized in a way to co-locate vectors of the same tenant together at the next optimization. -*To learn more about defragmentation, read the [Multitenancy documentation](/documentation/guides/multiple-partitions/).* +*To learn more about defragmentation, read the [Multitenancy documentation](/documentation/manage-data/multitenancy/).* ### On-Disk Support for the Payload Index @@ -281,7 +281,7 @@ await client.CreatePayloadIndexAsync( By moving the index to disk, Qdrant can handle larger datasets that exceed the capacity of RAM, making the system more scalable and capable of storing more data without being constrained by memory limitations. -*To learn more about this, read the [Indexing documentation](/documentation/concepts/indexing/).* +*To learn more about this, read the [Indexing documentation](/documentation/manage-data/indexing/).* ### UUID Datatype for the Payload Index @@ -311,7 +311,7 @@ PUT /collections/{collection_name}/points > For organizations that have numerous users and UUIDs, this simple fix can significantly reduce the cluster size and improve efficiency. -*To learn more about this, read the [Payload documentation](/documentation/concepts/payload/).* +*To learn more about this, read the [Payload documentation](/documentation/manage-data/payload/).* ### Query API: Groups Endpoint @@ -411,7 +411,7 @@ await client.QueryGroupsAsync( This endpoint will retrieve the best N points for each document, assuming that the payload of the points contains the document ID. Sometimes, the best N points cannot be fulfilled due to lack of points or a big distance with respect to the query. In every case, the `group_size` is a best-effort parameter, similar to the limit parameter. -*For more information on grouping capabilities refer to our [Hybrid Queries documentation](/documentation/concepts/hybrid-queries/).* +*For more information on grouping capabilities refer to our [Hybrid Queries documentation](/documentation/search/hybrid-queries/).* ### Query API: Random Sampling @@ -490,7 +490,7 @@ await client.QueryAsync( ); ``` -*To learn more, check out the [Query API documentation](/documentation/concepts/hybrid-queries/).* +*To learn more, check out the [Query API documentation](/documentation/search/hybrid-queries/).* ### Query API: Distribution-Based Score Fusion @@ -658,7 +658,7 @@ await client.QueryAsync( Note that `dbsf` is stateless and calculates the normalization limits only based on the results of each query, not on all the scores that it has seen. -*To learn more, check out the [Hybrid Queries documentation](/documentation/concepts/hybrid-queries/).* +*To learn more, check out the [Hybrid Queries documentation](/documentation/search/hybrid-queries/).* ## Web UI: Search Quality Tool diff --git a/qdrant-landing/content/blog/qdrant-1.12.x.md b/qdrant-landing/content/blog/qdrant-1.12.x.md index e6d7ac0a8..ccd3e5448 100644 --- a/qdrant-landing/content/blog/qdrant-1.12.x.md +++ b/qdrant-landing/content/blog/qdrant-1.12.x.md @@ -113,7 +113,7 @@ Two arrays, `offsets_row` and `offsets_col`, represent the positions of non-zero } } ``` -*To learn more about the distance matrix, read [**The Distance Matrix documentation**](/documentation/concepts/explore/#distance-matrix).* +*To learn more about the distance matrix, read [**The Distance Matrix documentation**](/documentation/search/explore/#distance-matrix).* ## Distance Matrix API in the Graph UI @@ -134,7 +134,7 @@ The new graphing method is cleaner and reveals **relationships and outliers:** ![distance-matrix](/blog/qdrant-1.12.x/distance-matrix.png) -*To learn more about the Web UI Dashboard, read the [**Interfaces documentation**](/documentation/interfaces/web-ui/).* +*To learn more about the Web UI Dashboard, read the [**Interfaces documentation**](/documentation/web-ui/).* ## Facet API for Metadata Cardinality @@ -142,7 +142,7 @@ The new graphing method is cleaner and reveals **relationships and outliers:** In modern applications like e-commerce, users often rely on [**filters**](/articles/vector-search-filtering/), such as **brand** or **color**, to refine search results. The **Facet API** is designed to help users understand the distribution of values in a dataset. -The `facet` endpoint can efficiently count and aggregate values for a specific [**payload field**](/documentation/concepts/payload/) in your dataset. +The `facet` endpoint can efficiently count and aggregate values for a specific [**payload field**](/documentation/manage-data/payload/) in your dataset. You can use it to retrieve unique values for a field, along with the number of points that contain each value. This functionality is similar to `GROUP BY` with `COUNT(*)` in SQL databases. @@ -195,12 +195,12 @@ POST /collections/{collection_name}/facet ``` This feature provides flexibility between performance and precision, depending on the needs of your application. -*To learn more about faceting, read the [**Facet API documentation**](/documentation/concepts/payload/#facet-counts).* +*To learn more about faceting, read the [**Facet API documentation**](/documentation/manage-data/payload/#facet-counts).* ## Text Index on Disk Support ![text-index-disk](/blog/qdrant-1.12.x/text-index-disk.png) -[**Qdrant text indexing**](/documentation/concepts/indexing/#full-text-index) tokenizes text into smaller units (tokens) based on chosen settings (e.g., tokenizer type, token length). These tokens are stored in an inverted index for fast text searches. +[**Qdrant text indexing**](/documentation/manage-data/indexing/#full-text-index) tokenizes text into smaller units (tokens) based on chosen settings (e.g., tokenizer type, token length). These tokens are stored in an inverted index for fast text searches. > With `on_disk` text indexing, the inverted index is stored on disk, reducing memory usage. @@ -222,11 +222,11 @@ PUT /collections/{collection_name}/index } ``` -*To learn more about indexes, read the [**Indexing documentation**](/documentation/concepts/indexing/).* +*To learn more about indexes, read the [**Indexing documentation**](/documentation/manage-data/indexing/).* ## Geo Index on Disk Support -For [**large-scale geographic datasets**](/documentation/concepts/payload/#geo) where storing all indexes in memory is impractical, **geo indexing** allows efficient filtering of points based on geographic coordinates. +For [**large-scale geographic datasets**](/documentation/manage-data/payload/#geo) where storing all indexes in memory is impractical, **geo indexing** allows efficient filtering of points based on geographic coordinates. With `on_disk` geo indexing, the index is written to disk instead of residing in memory, making it possible to handle large datasets without exhausting system memory. @@ -255,11 +255,11 @@ PUT /collections/{collection_name}/index ![geo-index-disk](/blog/qdrant-1.12.x/geo-index-disk.png) -> To learn how to get the best performance from Qdrant, read the [**Optimization Guide**](/documentation/guides/optimize/). +> To learn how to get the best performance from Qdrant, read the [**Optimization Guide**](/documentation/operations/optimize/). ## Just the Beginning -The easiest way to reach that **Hello World** moment is to [**try vector search in a live cluster**](/documentation/quickstart-cloud/). Our **interactive tutorial** will show you how to create a cluster, add data and try some filtering clauses. +The easiest way to reach that **Hello World** moment is to [**try vector search in a live cluster**](/documentation/cloud-quickstart/). Our **interactive tutorial** will show you how to create a cluster, add data and try some filtering clauses. **All of the new features from version 1.12 can be tested in the Web UI:** diff --git a/qdrant-landing/content/blog/qdrant-1.13.x.md b/qdrant-landing/content/blog/qdrant-1.13.x.md index ff3194aca..6540aadcc 100644 --- a/qdrant-landing/content/blog/qdrant-1.13.x.md +++ b/qdrant-landing/content/blog/qdrant-1.13.x.md @@ -62,13 +62,13 @@ This experiment didn't require any changes to the codebase, and everything worke - **Full Feature Support:** GPU indexing supports **all quantization options and datatypes** implemented in Qdrant. - **Large-Scale Benefits:** Fast indexing unlocks larger size of segments, which leads to **higher RPS on the same hardware**. -### [Instructions & Documentation](/documentation/guides/running-with-gpu/) +### [Instructions & Documentation](/documentation/operations/running-with-gpu/) The setup is simple, with pre-configured Docker images [**(check Docker Registry)**](https://hub.docker.com/r/qdrant/qdrant/tags) for GPU environments like NVIDIA and AMD. We've made it so you can enable GPU indexing with minimal configuration changes. > Note: Logs will clearly indicate GPU detection and usage for transparency. -*Read more about this feature in the [**GPU Indexing Documentation**](/documentation/guides/running-with-gpu/)* +*Read more about this feature in the [**GPU Indexing Documentation**](/documentation/operations/running-with-gpu/)* #### Interview With the Creator of GPU Indexing @@ -212,7 +212,7 @@ client.CreateCollection(context.Background(), &qdrant.CreateCollection{ ``` > You may also use the `PATCH` request to enable Strict Mode on an existing collection. -*Read more about Strict Mode in the [**Database Administration Guide**](/documentation/guides/administration/#strict-mode)* +*Read more about Strict Mode in the [**Database Administration Guide**](/documentation/operations/administration/#strict-mode)* ## HNSW Graph Compression @@ -228,7 +228,7 @@ In contrast with traditional compression algorithms, like gzip or lz4, **Delta E > Our experiments didn't observe any measurable performance degradation. However, the memory footprint of the HNSW graph was **reduced by up to 30%**. -*For more general info, read about [**Indexing and Data Structures in Qdrant**](/documentation/concepts/indexing/)* +*For more general info, read about [**Indexing and Data Structures in Qdrant**](/documentation/manage-data/indexing/)* ## Filter by Named Vectors @@ -236,13 +236,13 @@ In contrast with traditional compression algorithms, like gzip or lz4, **Delta E In Qdrant, you can store multiple vectors of different sizes and types in a single data point. This is useful when you have to representing data with multiple embeddings, such as image, text, or video features. -> We previously introduced this feature as [**Named Vectors**](/documentation/concepts/vectors/#named-vectors). Now, you can filter points by checking if a specific named vector exists. +> We previously introduced this feature as [**Named Vectors**](/documentation/manage-data/vectors/#named-vectors). Now, you can filter points by checking if a specific named vector exists. This makes it easy to search for points based on the presence of specific vectors. For example, *if your collection includes image and text vectors, you can filter for points that only have the image vector defined*. ### Create a Collection with Named Vectors -Upon collection [creation](/documentation/concepts/collections/#collection-with-multiple-vectors), you define named vector types, such as `image` or `text`: +Upon collection [creation](/documentation/manage-data/collections/#collection-with-multiple-vectors), you define named vector types, such as `image` or `text`: ```http PUT /collections/{collection_name} @@ -370,7 +370,7 @@ client.Scroll(context.Background(), &qdrant.ScrollPoints{ ``` This feature makes it easier to manage and query collections with heterogeneous data. It will give you more flexibility and control over your vector search workflows. -*To dive deeper into filtering by named vectors, check out the [**Filtering Documentation**](/documentation/concepts/filtering/#has-vector)* +*To dive deeper into filtering by named vectors, check out the [**Filtering Documentation**](/documentation/search/filtering/#has-vector)* ## Custom Storage Engine @@ -401,7 +401,7 @@ There are four elements: the **Data Layer**, **Mask Layer**, **the Region** and ## Get Started with Qdrant ![get-started](/blog/qdrant-1.13.x/image_1.png) -The easiest way to reach that **Hello World** moment is to [**try vector search in a live cluster**](/documentation/quickstart-cloud/). Our **interactive tutorial** will show you how to create a cluster, add data and try some filtering clauses. +The easiest way to reach that **Hello World** moment is to [**try vector search in a live cluster**](/documentation/cloud-quickstart/). Our **interactive tutorial** will show you how to create a cluster, add data and try some filtering clauses. **New features, like named vector filtering, can be tested in the Qdrant Dashboard:** diff --git a/qdrant-landing/content/blog/qdrant-1.14.x.md b/qdrant-landing/content/blog/qdrant-1.14.x.md index 7201bef2d..26acbc78f 100644 --- a/qdrant-landing/content/blog/qdrant-1.14.x.md +++ b/qdrant-landing/content/blog/qdrant-1.14.x.md @@ -175,23 +175,23 @@ POST /collections/{collection_name}/points/query You can tweak parameters like `target`, `scale`, and `midpoint` to shape how quickly the score decays over distance. This is extremely useful for local search scenarios, where location is a major factor but not the only factor. -> This is a very powerful feature that allows for extensive customization. Read more about this feature in the [**Hybrid Queries Documentation**](/documentation/concepts/hybrid-queries/) +> This is a very powerful feature that allows for extensive customization. Read more about this feature in the [**Hybrid Queries Documentation**](/documentation/search/hybrid-queries/) ## Incremental HNSW Indexing ![optimizations](/blog/qdrant-1.14.x/optimizations.jpg) -Rebuilding an entire [**HNSW graph**](/documentation/concepts/indexing/#vector-index) every time new data is added can be computationally expensive. With this release, Qdrant now supports incremental HNSW indexing—an approach that extends existing HNSW graphs rather than recreating them from scratch. +Rebuilding an entire [**HNSW graph**](/documentation/manage-data/indexing/#vector-index) every time new data is added can be computationally expensive. With this release, Qdrant now supports incremental HNSW indexing—an approach that extends existing HNSW graphs rather than recreating them from scratch. > This feature is designed to make indexing faster and more efficient when you’re only adding new points. It reuses the existing structure of the HNSW graph and appends the new data directly onto it. That means much less time spent building and more time searching. Although this initial implementation currently only support upserts, it lays the groundwork for a more dynamic and performance-friendly indexing process. Especially for collections with frequent updates to payload values or growing datasets, incremental HNSW is a big step forward. -> Note that deletes and updates will still trigger a full rebuild. Check out the [**indexing documentation**](/documentation/concepts/indexing/) to learn more. +> Note that deletes and updates will still trigger a full rebuild. Check out the [**indexing documentation**](/documentation/manage-data/indexing/) to learn more. ## Faster Batch Queries ![reranking](/blog/qdrant-1.14.x/gridstore.jpg) -In this release, Qdrant introduces a major performance boost for [**batch query operations**](/documentation/concepts/search/#batch-search-api). Until now, the query batch API used a single thread per segment, which worked well—unless you had just one segment and a large batch of queries in a single request. In such cases, everything was processed on a single thread, significantly slowing things down. This scenario was especially common when using our [**Python client**](https://github.com/qdrant/qdrant-client), which is single-threaded by default. +In this release, Qdrant introduces a major performance boost for [**batch query operations**](/documentation/search/search/#batch-search-api). Until now, the query batch API used a single thread per segment, which worked well—unless you had just one segment and a large batch of queries in a single request. In such cases, everything was processed on a single thread, significantly slowing things down. This scenario was especially common when using our [**Python client**](https://github.com/qdrant/qdrant-client), which is single-threaded by default. The new optimization changes that. Large query batches are now split into chunks, and each chunk is processed on a separate thread. @@ -220,7 +220,7 @@ We ran the same test of large queries for the following configurations: As you can see, the improvement is **most significant (57%) in single-segment configurations** where parallelization was previously limited. Even in already-optimized multi-shard setups, we still see good gains of 12-32%. -> For more on batch queries, check out the [**Search documentation**](/documentation/concepts/search/#batch-search-api). +> For more on batch queries, check out the [**Search documentation**](/documentation/search/search/#batch-search-api). ## Improved Resource Use During Segment Optimization ![segment-optimization](/blog/qdrant-1.14.x/segment-optimization.jpg) @@ -237,7 +237,7 @@ It also gives you **predictable performance**, as there are fewer sudden spikes In our experiment, **we indexed 400 million 512-dimensional vectors**. The previous version of Qdrant took around 40 hours on an 8-core machine, while the new version with this change completed the task in just 28 hours. -> **Tutorial:** If you want to work with a large number of vectors, we can show you how. [**Learn how to upload and search large collections efficiently.**](/documentation/database-tutorials/large-scale-search/) +> **Tutorial:** If you want to work with a large number of vectors, we can show you how. [**Learn how to upload and search large collections efficiently.**](/documentation/tutorials-operations/large-scale-search/) ## Optimized Memory Usage in Immutable Segments diff --git a/qdrant-landing/content/blog/qdrant-1.15.x.md b/qdrant-landing/content/blog/qdrant-1.15.x.md index 3bb972154..ac74180be 100644 --- a/qdrant-landing/content/blog/qdrant-1.15.x.md +++ b/qdrant-landing/content/blog/qdrant-1.15.x.md @@ -66,7 +66,7 @@ This approach maintains storage size and RAM usage similar to binary quantizatio When performing nearest vector search, the query vector is compared against quantized vectors stored in the database. If the query itself remains unquantized and a scoring method exists to evaluate it directly against the compressed vectors, this allows for more accurate results without increasing memory usage. -> Quantization enables efficient storage and search of high-dimensional vectors. Learn more about this from our [**quantization**](/documentation/guides/quantization/) docs. +> Quantization enables efficient storage and search of high-dimensional vectors. Learn more about this from our [**quantization**](/documentation/manage-data/quantization/) docs.
@@ -89,7 +89,7 @@ Dataset: Laion 1 million 512d vectors ![Section 2](/blog/qdrant-1.15.x/section-2.png) Full-text filtering in Qdrant in an efficient way to combine Vector-based scoring with exact keyword match. -And in v1.15 full-text index recieved a number of upgrades which make vector similarity evem more useful. +And in v1.15 full-text index received a number of upgrades which make vector similarity even more useful. ### Multilingual Tokenization @@ -138,7 +138,7 @@ PUT /collections/{collection_name}/index } ``` -For more information about stopwords, see the [documentation](https://qdrant.tech/documentation/concepts/indexing/#stopwords). +For more information about stopwords, see the [documentation](https://qdrant.tech/documentation/manage-data/indexing/#stopwords). ### Stemming @@ -169,10 +169,10 @@ PUT /collections/{collection_name}/index ### Phrase Matching -With [phrase matching](/documentation/concepts/filtering/#phrase-match), you can now perform exact phrase search. +With [phrase matching](/documentation/search/filtering/#phrase-match), you can now perform exact phrase search. It allows you to search for a specific phrase, words in exact order, within a text field. -For efficient phrase seach Qdrant requires to build an additional data structure, +For efficient phrase search Qdrant requires to build an additional data structure, so it needs to be configured during creation of the full-text index: ```http @@ -213,7 +213,7 @@ The above will match: ## MMR Reranking -We introduce [Maximal Marginal Relevance (MMR)](/documentation/concepts/search-relevance/#maximal-marginal-relevance-mmr) reranking to balance relevance and diversity. +We introduce [Maximal Marginal Relevance (MMR)](/documentation/search/search-relevance/#maximal-marginal-relevance-mmr) reranking to balance relevance and diversity. MMR works by selecting the results iteratively, by picking the item with the best combination of similarity to the query and dissimilarity to the already selected items. It prevents your top-k results from being redundant and helps surface varied but relevant answers, particularly in dense datasets with overlapping entries. @@ -225,7 +225,7 @@ It prevents your top-k results from being redundant and helps surface varied but Let’s say you’re building a knowledge assistant or semantic document explorer in which a single query can return multiple highly similar queries. For instance, searching “climate change” in a scientific paper database might return several similar paragraphs. -You can diversify the results with [Maximal Marginal Relevance (MMR)](/documentation/concepts/search-relevance/#maximal-marginal-relevance-mmr). +You can diversify the results with [Maximal Marginal Relevance (MMR)](/documentation/search/search-relevance/#maximal-marginal-relevance-mmr). Instead of returning the top-k results based on pure similarity, MMR helps select a diverse subset of high-quality results. This gives more coverage and avoids redundant results, which is helpful in dense content domains such as academic papers, product catalogs, or search assistants. @@ -273,7 +273,7 @@ As usual, new Qdrant release brings more performance optimization for faster and Qdrant 1.15 introduces HNSW healing. Instead of completely re-building HNSW index during optimization, Qdrant now tries to re-use information from the existing vector index to speed-up construction of the new one. -When points are removed from an existing [HNSW graph](https://qdrant.tech/documentation/concepts/indexing/#vector-index), new links are added to prevent isolation in the graph, and avoid decreasing search quality. +When points are removed from an existing [HNSW graph](https://qdrant.tech/documentation/manage-data/indexing/#vector-index), new links are added to prevent isolation in the graph, and avoid decreasing search quality. {{
}} @@ -283,7 +283,7 @@ This modification, in combinations with [incremental HNSW indexing](/blog/qdrant ### HNSW Graph connectivity estimation -Qdrant builds [addtitional HNSW links](/articles/filterable-hnsw/) to ensure that filtered searches are performed fast and accurate. +Qdrant builds [additional HNSW links](/articles/filterable-hnsw/) to ensure that filtered searches are performed fast and accurate. It does, however, introduce an overhead for indexing complexity, especially when the number of payload indexes is large. With v1.15, Qdrant introduces an optimization, which quickly estimates graph connectivity before creating additional links. diff --git a/qdrant-landing/content/blog/qdrant-1.16.x.md b/qdrant-landing/content/blog/qdrant-1.16.x.md index d14baaa7c..0c1f4e92f 100644 --- a/qdrant-landing/content/blog/qdrant-1.16.x.md +++ b/qdrant-landing/content/blog/qdrant-1.16.x.md @@ -31,32 +31,32 @@ Additionally, version 1.16 introduces a new conditional update API, facilitating Multitenancy is a common requirement for SaaS applications, where multiple customers (tenants) share the same database instance. In Qdrant, when an instance is shared between multiple users, you may need to partition vectors by user. This is done so that each user can only access their own vectors and can’t see the vectors of other users. To implement multitenancy in Qdrant, there are two main approaches: -- [Payload-based multitenancy](/documentation/guides/multiple-partitions/), which works well when you have a large number of small tenants. This causes practically no overhead. Quite the opposite: a query with a tenant payload filter can be faster than a full search. -- [Shard-based multitenancy](/documentation/guides/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead. +- [Payload-based multitenancy](/documentation/manage-data/multitenancy/), which works well when you have a large number of small tenants. This causes practically no overhead. Quite the opposite: a query with a tenant payload filter can be faster than a full search. +- [Shard-based multitenancy](/documentation/operations/distributed_deployment/#user-defined-sharding), designed for when you have a smaller number of larger tenants. This works well when each tenant requires isolation and dedicated resources. Separating tenants by shard prevents a classic noisy neighbor problem where a single high-volume tenant can force the cluster to scale for everyone, increasing costs and reducing performance for smaller tenants. However, shard-based multitenancy is not a good solution when you have a large number of small tenants, as each shard incurs some overhead. Real-world usage patterns often fall between these two use cases. It's common to have a small number of large tenants and a huge tail of smaller ones. You may even have tenants that grow over time, starting small and eventually becoming large enough to require dedicated resources. -In version 1.16, Qdrant can now efficiently combine the two multitenancy approaches with a new feature called [Tiered Multitenancy](/documentation/guides/multitenancy/#tiered-multitenancy). +In version 1.16, Qdrant can now efficiently combine the two multitenancy approaches with a new feature called [Tiered Multitenancy](/documentation/manage-data/multitenancy/#tiered-multitenancy). The main principles behind Tiered Multitenancy are: -- [User-defined Sharding](/documentation/guides/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants. +- [User-defined Sharding](/documentation/operations/distributed_deployment/#user-defined-sharding) allows you to create named shards within a collection. It enables you to isolate large tenants into their own shards. A multitenant collection can consist of a shared "fallback" shard for small tenants and multiple dedicated shards for large tenants. - **Fallback shards** - a special routing mechanism that allows Qdrant to route a request to either a dedicated shard (if it exists) or to a shared fallback shard. This keeps requests unified, without the need to know whether a tenant is dedicated or shared. -- [Tenant promotion](/documentation/guides/multitenancy/#promote-tenant-to-dedicated-shard) - a mechanism that makes it possible to "promote" tenants from the shared Fallback Shard to their own dedicated shard when they grow large enough. This process is based on Qdrant’s internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process. +- [Tenant promotion](/documentation/manage-data/multitenancy/#promote-tenant-to-dedicated-shard) - a mechanism that makes it possible to "promote" tenants from the shared Fallback Shard to their own dedicated shard when they grow large enough. This process is based on Qdrant’s internal shard transfer mechanism, which makes promotion completely transparent for the application. Both read and write requests are supported during the promotion process. {{< figure src="/docs/tenant-promotion.png" alt="Tiered multitenancy with tenant promotion" caption="Tiered multitenancy with tenant promotion" width="90%" >}} -To use Tiered Multitenancy, after [setting up a collection with a shared fallback shard and dedicated shards](/documentation/guides/multitenancy/#configure-tiered-multitenancy), when inserting or querying points, [provide a shard key selector](/documentation/guides/multitenancy/#query-tiered-multitenant-collection). +To use Tiered Multitenancy, after [setting up a collection with a shared fallback shard and dedicated shards](/documentation/manage-data/multitenancy/#configure-tiered-multitenancy), when inserting or querying points, [provide a shard key selector](/documentation/manage-data/multitenancy/#query-tiered-multitenant-collection). ## ACORN - Filtered Vector Search Improvements ![Section 2](/blog/qdrant-1.16.x/section-2.png) -To enhance the scalability and speed of vector search, Qdrant employs a graph-based index structure known as [HNSW (Hierarchical Navigable Small World)](/documentation/concepts/indexing/#vector-index). While traditional HNSW is primarily designed for unfiltered searches, Qdrant has addressed this limitation by implementing a [filterable HSNW index](/articles/filterable-hnsw/). This innovative approach extends the HNSW graph with additional edges that correspond to indexed payload values. This enables Qdrant to maintain search quality even with high filtering selectivity, without introducing any runtime overhead during the search process. +To enhance the scalability and speed of vector search, Qdrant employs a graph-based index structure known as [HNSW (Hierarchical Navigable Small World)](/documentation/manage-data/indexing/#vector-index). While traditional HNSW is primarily designed for unfiltered searches, Qdrant has addressed this limitation by implementing a [filterable HSNW index](/articles/filterable-hnsw/). This innovative approach extends the HNSW graph with additional edges that correspond to indexed payload values. This enables Qdrant to maintain search quality even with high filtering selectivity, without introducing any runtime overhead during the search process. -Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen when you use a combination of high cardinality filters, leading to the HNSW graph becoming [disconnected](/documentation/concepts/indexing/#filterable-index). It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where filtering criteria are not known in advance. +Even with filterable HNSW graphs, there are instances where the quality of search results can deteriorate significantly. This can happen when you use a combination of high cardinality filters, leading to the HNSW graph becoming [disconnected](/documentation/manage-data/indexing/#filterable-index). It is impractical to build additional links for every possible combination of filters in advance due to the potentially vast number of combinations. Another case where filterable HNSW may break down is in situations where filtering criteria are not known in advance. -To address these limitations, in version 1.16 we are introducing support for [ACORN](/documentation/concepts/search/#acorn-search-algorithm), based on the ACORN-1 algorithm described in the paper [ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data](https://arxiv.org/abs/2403.04871). With ACORN enabled, Qdrant not only traverses direct neighbors (the first hop) in the HNSW graph but also examines neighbors of neighbors (the second hop) if the direct neighbors have been filtered out. This enhancement improves search accuracy at the expense of performance, especially when multiple low-selectivity filters are applied. +To address these limitations, in version 1.16 we are introducing support for [ACORN](/documentation/search/search/#acorn-search-algorithm), based on the ACORN-1 algorithm described in the paper [ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data](https://arxiv.org/abs/2403.04871). With ACORN enabled, Qdrant not only traverses direct neighbors (the first hop) in the HNSW graph but also examines neighbors of neighbors (the second hop) if the direct neighbors have been filtered out. This enhancement improves search accuracy at the expense of performance, especially when multiple low-selectivity filters are applied.
@@ -65,7 +65,7 @@ To address these limitations, in version 1.16 we are introducing support for [AC
-You can enable ACORN on a per-query basis, via the optional [query-time `acorn` parameter](/documentation/concepts/search/#acorn-search-algorithm). This doesn't require any changes at index time. +You can enable ACORN on a per-query basis, via the optional [query-time `acorn` parameter](/documentation/search/search/#acorn-search-algorithm). This doesn't require any changes at index time. ### Benchmarks @@ -126,9 +126,9 @@ For instance, querying 1 million vectors with the HNSW parameters `m` set to 16 However, disk-based storage has a property we can exploit to reduce the number of random access reads: paged reading. Disk devices typically read a full page (4KB or more) of data at once. Traditional tree-based data structures, such as B-trees, have used this property effectively. However, in graph-based structures like the HNSW index, grouping connected nodes into pages is not straightforward due to each node potentially having an arbitrary number of connections to other nodes. -With Qdrant version 1.16, you can make use of paged reading through a new feature called [inline storage](/documentation/guides/optimize/#inline-storage-in-hnsw-index). Inline storage allows for storing quantized vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space. +With Qdrant version 1.16, you can make use of paged reading through a new feature called [inline storage](/documentation/operations/optimize/#inline-storage-in-hnsw-index). Inline storage allows for storing quantized vector data directly inside the HNSW nodes. This offers faster read access, at the cost of additional storage space. -Inline storage can be enabled by [setting a collection's HNSW configuration `inline_storage` option to `true`](/documentation/guides/optimize/#inline-storage-in-hnsw-index). It requires quantization to be enabled. +Inline storage can be enabled by [setting a collection's HNSW configuration `inline_storage` option to `true`](/documentation/operations/optimize/#inline-storage-in-hnsw-index). It requires quantization to be enabled.
@@ -218,7 +218,7 @@ However, there was no convenient way to search to match *at least one* of the pr } ``` -In version 1.16, we have added a new [`text_any` condition](/documentation/concepts/filtering/#full-text-any) that simplifies this use case. Now, instead of building complex boolean conditions, Qdrant can handle the tokenization and matching internally. +In version 1.16, we have added a new [`text_any` condition](/documentation/search/filtering/#full-text-any) that simplifies this use case. Now, instead of building complex boolean conditions, Qdrant can handle the tokenization and matching internally. The `text_any` condition matches text fields that contain any of the query terms. In other words, even if a text field contains just one of the query terms, it is considered a match. @@ -270,9 +270,9 @@ Many Latin languages use diacritical marks (accents) to indicate different pronu A solution to this problem is to normalize characters with diacritics to their base ASCII equivalents, a process known as ASCII folding. For example, "café" becomes "cafe" and "naïve" becomes "naive." This normalization allows for more flexible and inclusive search results, improving search recall for multilingual texts. -[An open source contribution by community member eltu](https://github.com/qdrant/qdrant/pull/7408) has added [ASCII folding support](/documentation/concepts/indexing/#ascii-folding) to Qdrant's full-text search capabilities in version 1.16. When enabled, Qdrant automatically normalizes text fields and search terms, for instance by removing diacritical marks. +[An open source contribution by community member eltu](https://github.com/qdrant/qdrant/pull/7408) has added [ASCII folding support](/documentation/manage-data/indexing/#ascii-folding) to Qdrant's full-text search capabilities in version 1.16. When enabled, Qdrant automatically normalizes text fields and search terms, for instance by removing diacritical marks. -To enable ASCII folding, [set the `ascii_folding` option to `true` when creating a full-text payload index](/documentation/concepts/indexing/#ascii-folding). +To enable ASCII folding, [set the `ascii_folding` option to `true` when creating a full-text payload index](/documentation/manage-data/indexing/#ascii-folding). ## Conditional Updates @@ -285,7 +285,7 @@ Point updates in Qdrant are idempotent, meaning that applying the same update mu 3. Client A modifies point P and writes it back to Qdrant. 4. Client B modifies point P (based on stale data) and writes it back to Qdrant, unintentionally overwriting changes made by Client A. -To address this issue, Qdrant 1.16 introduces support for [conditional updates](/documentation/concepts/points/#conditional-updates). With conditional updates, you can specify a condition, in the form of an update filter, that must be met for the update to be applied. If the condition is not met, Qdrant rejects the update, preventing unintended overwrites. +To address this issue, Qdrant 1.16 introduces support for [conditional updates](/documentation/manage-data/points/#conditional-updates). With conditional updates, you can specify a condition, in the form of an update filter, that must be met for the update to be applied. If the condition is not met, Qdrant rejects the update, preventing unintended overwrites. For example, you can add a `version` field to your points to track changes. When updating a point, you can specify a condition that the `version` field must match the expected value. If another client has modified the point in the meantime and incremented the `version`, the update is rejected: @@ -315,10 +315,10 @@ In version 1.16, we have revamped the Web UI with a fresh new look and improved ![Section 7](/blog/qdrant-1.16.x/section-7.png) -- The constant `k` that determines how Reciprocal Rank Fusion (RRF) fuses result sets [is now configurable](/documentation/concepts/hybrid-queries/#parametrized-rrf). -- The Metrics API now exposes [additional metrics](/documentation/guides/monitoring/#metrics) that help monitor your deployment's health. -- In strict mode, it is now possible to [configure the maximum number of payload indices](/documentation/guides/administration/#maximum-number-of-payload-index-count). -- It's now possible to [attach custom metadata to collections](/documentation/concepts/collections/#collection-metadata). +- The constant `k` that determines how Reciprocal Rank Fusion (RRF) fuses result sets [is now configurable](/documentation/search/hybrid-queries/#parametrized-rrf). +- The Metrics API now exposes [additional metrics](/documentation/operations/monitoring/#metrics) that help monitor your deployment's health. +- In strict mode, it is now possible to [configure the maximum number of payload indices](/documentation/operations/administration/#maximum-number-of-payload-index-count). +- It's now possible to [attach custom metadata to collections](/documentation/manage-data/collections/#collection-metadata). For a full list of all changes in version 1.16, please refer to the [change log](https://github.com/qdrant/qdrant/releases/tag/v1.16.0). diff --git a/qdrant-landing/content/blog/qdrant-1.17.x.md b/qdrant-landing/content/blog/qdrant-1.17.x.md index b2d9bab8d..bfa6c9c44 100644 --- a/qdrant-landing/content/blog/qdrant-1.17.x.md +++ b/qdrant-landing/content/blog/qdrant-1.17.x.md @@ -30,9 +30,9 @@ tags: Crafting queries is hard: users often struggle to precisely formulate search queries. At the same time, judging the relevance of a given search result is often much easier. Retrieval systems can leverage this [relevance feedback](/articles/search-feedback-loop/) to iteratively refine results toward user intent. -This release introduces a new [Relevance Feedback Query](/documentation/concepts/search-relevance/#relevance-feedback) as a scalable, vector‑native approach to incorporating relevance feedback. The Relevance Feedback Query uses a small amount of model‑generated feedback to guide the retriever through the entire vector space, effectively nudging search toward “more relevant” results without requiring expensive loops, expensive retrievers, or human labeling. This enables the engine to traverse billions of vectors with improved recall without having to retrain models. +This release introduces a new [Relevance Feedback Query](/documentation/search/search-relevance/#relevance-feedback) as a scalable, vector‑native approach to incorporating relevance feedback. The Relevance Feedback Query uses a small amount of model‑generated feedback to guide the retriever through the entire vector space, effectively nudging search toward “more relevant” results without requiring expensive loops, expensive retrievers, or human labeling. This enables the engine to traverse billions of vectors with improved recall without having to retrain models. -This method works by collecting lightweight feedback on just a few top results, creating “context pairs” of more‑ and less‑relevant examples. These pairs define a signal that adjusts the scoring function during the next retrieval pass. Instead of rewriting queries or rescoring large batches of documents, Qdrant modifies how similarity is computed. Experiments demonstrate substantial gains, especially when pairing expressive retrievers with strong feedback models. For the methodology and experiments behind this feature, see our article [Relevance Feedback in Qdrant](/articles/relevance-feedback). To get started, refer to the [documentation](/documentation/concepts/search-relevance/#relevance-feedback). +This method works by collecting lightweight feedback on just a few top results, creating “context pairs” of more‑ and less‑relevant examples. These pairs define a signal that adjusts the scoring function during the next retrieval pass. Instead of rewriting queries or rescoring large batches of documents, Qdrant modifies how similarity is computed. Experiments demonstrate substantial gains, especially when pairing expressive retrievers with strong feedback models. For the methodology and experiments behind this feature, see our article [Relevance Feedback in Qdrant](/articles/relevance-feedback/). To get started, refer to the [documentation](/documentation/search/search-relevance/#relevance-feedback).
@@ -56,9 +56,9 @@ A common pattern with vector search engines like Qdrant involves bulk uploads. F This release addresses these issues by changing how data is ingested. Shards still process data through the familiar stages: WAL persistence, queued updates, application to unoptimized segments, and eventual full indexing, but two new features reshape how systems behave under heavy write load. -A new [update queue](/documentation/guides/low-latency-search/#query-indexed-data-only) tracks up to one million pending changes. When the queue fills, back pressure slows incoming writes, preventing runaway load and helping clusters stay stable even during large batch operations or recovery after downtime. +A new [update queue](/documentation/search/low-latency-search/#query-indexed-data-only) tracks up to one million pending changes. When the queue fills, back pressure slows incoming writes, preventing runaway load and helping clusters stay stable even during large batch operations or recovery after downtime. -For applications that demand consistently low-latency search, indexed‑only mode ensures queries touch only fully indexed segments. A side-effect of using indexed-only queries was that they could temporarily hide the newest updates, before they were indexed. A new [`prevent_unoptimized` optimizer setting](/documentation/guides/low-latency-search/#query-indexed-data-only) solves this by throttling updates to match the indexing rate, reducing the creation of large unoptimized segments. +For applications that demand consistently low-latency search, indexed‑only mode ensures queries touch only fully indexed segments. A side-effect of using indexed-only queries was that they could temporarily hide the newest updates, before they were indexed. A new [`prevent_unoptimized` optimizer setting](/documentation/search/low-latency-search/#query-indexed-data-only) solves this by throttling updates to match the indexing rate, reducing the creation of large unoptimized segments. Together, these features give developers tighter control over write throughput, indexing behavior, and search performance, especially in high‑volume environments. @@ -66,7 +66,7 @@ Together, these features give developers tighter control over write throughput, By default, a search operation queries a single replica of each shard within a collection. If one of these replicas responds slowly due to load or network issues, this negatively impacts the overall search latency. This phenomenon, where a single slow replica increases the 95th or 99th percentile latency of the entire system, is known as “tail latency.” High tail latency can noticeably degrade the user experience. -To mitigate tail latency for read operations, this release introduces a new [delayed fan-out](/documentation/guides/low-latency-search/#use-delayed-fan-outs) feature. With delayed fan-outs, if the initial request to a replica exceeds a configurable latency threshold, an additional read request is sent to another replica, and Qdrant will use the first available response. Delayed fan-outs help your application provide a consistent, low latency experience to end-users. +To mitigate tail latency for read operations, this release introduces a new [delayed fan-out](/documentation/search/low-latency-search/#use-delayed-fan-outs) feature. With delayed fan-outs, if the initial request to a replica exceeds a configurable latency threshold, an additional read request is sent to another replica, and Qdrant will use the first available response. Delayed fan-outs help your application provide a consistent, low latency experience to end-users. ## Greater Operational Observability @@ -78,11 +78,11 @@ We are continuously working to enhance the operational observability of Qdrant c Qdrant’s API exposes a `/telemetry` endpoint which provides information about the current state of a peer in a cluster, including the number of vectors, shards, and other useful information. However, obtaining a complete view of the entire cluster using this endpoint is not straightforward, requiring querying each peer and piecing together a complete view yourself. -In version 1.17, we’re introducing a new [`/cluster/telemetry` endpoint](/documentation/guides/monitoring/#cluster-wide-telemetry). This API provides information about all peers in a cluster, offering insights into cluster-wide operations such as leader elections, resharding, and shard transfers. +In version 1.17, we’re introducing a new [`/cluster/telemetry` endpoint](/documentation/operations/monitoring/#cluster-wide-telemetry). This API provides information about all peers in a cluster, offering insights into cluster-wide operations such as leader elections, resharding, and shard transfers. ### Segment Optimization Monitoring -Optimization is a background process where Qdrant removes data marked for deletion, merges segments, and creates indexes. To improve visibility into this process, this release introduces [segment optimization monitoring capabilities](/documentation/concepts/optimizer/#optimization-monitoring). +Optimization is a background process where Qdrant removes data marked for deletion, merges segments, and creates indexes. To improve visibility into this process, this release introduces [segment optimization monitoring capabilities](/documentation/operations/optimizer/#optimization-monitoring). A new `/collections/{collection_name}/optimizations` API endpoint provides cluster-wide information about the current optimization status, as well as detailed information for current and past optimization operations. Because the output of the API can be verbose, we’ve added a new Optimizations tab to the Collections interface in the Web UI that makes it easier to analyze the data. Here, you can find an overview of the current optimization status, a timeline of current and past optimization operations, and a breakdown of the tasks in a specific cycle and their durations. @@ -114,17 +114,17 @@ Many people have been asking about point filtering in web UI. And now it's back, As an open source project, we welcome contributions from the Qdrant community. This release features two contributions from community members: -- Not all payload field indexes are used in combination with dense vector queries. With this release, you can [specify whether individual payload field indexes should be reflected in the HNSW index](/documentation/concepts/indexing/#disable-the-creation-of-extra-edges-for-payload-fields). -- A new API endpoint is available to [list all user-defined shard keys](/documentation/guides/distributed_deployment/#user-defined-sharding). +- Not all payload field indexes are used in combination with dense vector queries. With this release, you can [specify whether individual payload field indexes should be reflected in the HNSW index](/documentation/manage-data/indexing/#disable-the-creation-of-extra-edges-for-payload-fields). +- A new API endpoint is available to [list all user-defined shard keys](/documentation/operations/distributed_deployment/#user-defined-sharding). Additionally, this release adds the following features: -- Upserts now support an [update mode](/documentation/concepts/points/#update-mode) for insert-only or update-only operations. +- Upserts now support an [update mode](/documentation/manage-data/points/#update-mode) for insert-only or update-only operations. - To speed up the recovery of the replicas after they’ve been down, shards will [increase the size of their write-ahead log](https://github.com/qdrant/qdrant/pull/7834) when they detect that one of their remote replicas is unavailable. -- Reciprocal Rank Fusion (RRF) combines multiple query results into one list, but its default equal weighting can let weaker rankers dilute stronger ones. [Weighted RRF](/documentation/concepts/hybrid-queries/#reciprocal-rank-fusion-rrf) in Qdrant 1.17 addresses this by letting you assign weights to individual queries. +- Reciprocal Rank Fusion (RRF) combines multiple query results into one list, but its default equal weighting can let weaker rankers dilute stronger ones. [Weighted RRF](/documentation/search/hybrid-queries/#reciprocal-rank-fusion-rrf) in Qdrant 1.17 addresses this by letting you assign weights to individual queries. - A new [user interface in the Web UI enables resharding collections](https://github.com/qdrant/qdrant-web-ui/pull/341) on Qdrant Cloud. -- Qdrant now supports [audit logging](/documentation/guides/security/#audit-logging) to track all API operations that require authentication or authorization. -- [External provider API keys for inference requests](/documentation/concepts/inference/#external-embedding-model-providers) can now be provided in the request header. +- Qdrant now supports [audit logging](/documentation/operations/security/#audit-logging) to track all API operations that require authentication or authorization. +- [External provider API keys for inference requests](/documentation/inference/#external-embedding-model-providers) can now be provided in the request header. For a full list of all changes in version 1.17, please refer to the [change log](https://github.com/qdrant/qdrant/releases/tag/v1.17.0). diff --git a/qdrant-landing/content/blog/qdrant-1.9.x.md b/qdrant-landing/content/blog/qdrant-1.9.x.md index c76f64c05..b2650f284 100644 --- a/qdrant-landing/content/blog/qdrant-1.9.x.md +++ b/qdrant-landing/content/blog/qdrant-1.9.x.md @@ -28,7 +28,7 @@ tags: Historically, our API key supported basic read and write operations. However, recognizing the evolving needs of our user base, especially large organizations, we've implemented additional options for finer control over data access within internal environments. -Qdrant now supports [granular access control using JSON Web Tokens (JWT)](/documentation/guides/security/#granular-access-control-with-jwt). JWT will let you easily limit a user's access to the specific data they are permitted to view. Specifically, JWT-based authentication leverages tokens with restricted access to designated data segments, laying the foundation for implementing role-based access control (RBAC) on top of it. **You will be able to define permissions for users and restrict access to sensitive endpoints.** +Qdrant now supports [granular access control using JSON Web Tokens (JWT)](/documentation/operations/security/#granular-access-control-with-jwt). JWT will let you easily limit a user's access to the specific data they are permitted to view. Specifically, JWT-based authentication leverages tokens with restricted access to designated data segments, laying the foundation for implementing role-based access control (RBAC) on top of it. **You will be able to define permissions for users and restrict access to sensitive endpoints.** **Dashboard users:** For your convenience, we have added a JWT generation tool the Qdrant Web UI under the 🔑 tab. If you're using the default url, you will find it at `http://localhost:6333/dashboard#/jwt`. @@ -36,11 +36,11 @@ Qdrant now supports [granular access control using JSON Web Tokens (JWT)](/docum We highly recommend this feature to enterprises using [Qdrant Hybrid Cloud](/hybrid-cloud/), as it is tailored to those who need additional control over company data and user access. RBAC empowers administrators to define roles and assign specific privileges to users based on their roles within the organization. In combination with [Hybrid Cloud's data sovereign architecture](/documentation/hybrid-cloud/), this feature reinforces internal security and efficient collaboration by granting access only to relevant resources. -> **Documentation:** [Read the access level breakdown](/documentation/guides/security/#table-of-access) to see which actions are allowed or denied. +> **Documentation:** [Read the access level breakdown](/documentation/operations/security/#table-of-access) to see which actions are allowed or denied. ## Faster shard transfers on node recovery -We now offer a streamlined approach to [data synchronization between shards](/documentation/guides/distributed_deployment/#shard-transfer-method) during node upgrades or recovery processes. Traditional methods used to transfer the entire dataset, but our new `wal_delta` method focuses solely on transmitting the difference between two existing shards. By leveraging the Write-Ahead Log (WAL) of both shards, this method selectively transmits missed operations to the target shard, ensuring data consistency. +We now offer a streamlined approach to [data synchronization between shards](/documentation/operations/distributed_deployment/#shard-transfer-method) during node upgrades or recovery processes. Traditional methods used to transfer the entire dataset, but our new `wal_delta` method focuses solely on transmitting the difference between two existing shards. By leveraging the Write-Ahead Log (WAL) of both shards, this method selectively transmits missed operations to the target shard, ensuring data consistency. In some cases, where transfers can take hours, this update **reduces transfers down to a few minutes.** @@ -48,7 +48,7 @@ The advantages of this approach are twofold: 1. **It is faster** since only the differential data is transmitted, avoiding the transfer of redundant information. 2. It upholds robust **ordering guarantees**, crucial for applications reliant on strict sequencing. -For more details on how this works, check out the [shard transfer documentation](/documentation/guides/distributed_deployment/#shard-transfer-method). +For more details on how this works, check out the [shard transfer documentation](/documentation/operations/distributed_deployment/#shard-transfer-method). > **Note:** There are limitations to consider. First, this method only works with existing shards. Second, while the WALs typically retain recent operations, their capacity is finite, potentially impeding the transfer process if exceeded. Nevertheless, for scenarios like rapid node restarts or upgrades, where the WAL content remains manageable, WAL delta transfer is an efficient solution. @@ -56,7 +56,7 @@ Overall, this is a great optional optimization measure and serves as the **auto- ## Native support for uint8 embeddings -Our latest version introduces [support for uint8 embeddings within Qdrant collections](/documentation/concepts/collections/#vector-datatypes). This feature supports embeddings provided by companies in a pre-quantized format. Unlike previous iterations where indirect support was available via [quantization methods](/documentation/guides/quantization/), this update empowers users with direct integration capabilities. +Our latest version introduces [support for uint8 embeddings within Qdrant collections](/documentation/manage-data/collections/#vector-datatypes). This feature supports embeddings provided by companies in a pre-quantized format. Unlike previous iterations where indirect support was available via [quantization methods](/documentation/manage-data/quantization/), this update empowers users with direct integration capabilities. In the case of `uint8`, elements within the vector are represented as unsigned 8-bit integers, encompassing values ranging from 0 to 255. Using these embeddings gives you a **4x memory saving and about a 30% speed-up in search**, while keeping 99.99% of the response quality. As opposed to the original quantization method, with this feature you can spare disk usage if you directly implement pre-quantized embeddings. @@ -73,7 +73,7 @@ PUT /collections/{collection_name} } ``` -> **Note:** When using Quantization to optimize vector search, you can use this feature to `rescore` binary vectors against new byte vectors. With double the speedup, you will be able to achieve a better result than if you rescored with float vectors. With each byte vector quantized at the binary level, the result will deliver unparalleled efficiency and savings. To learn more about this optimization method, read our [Quantization docs](/documentation/guides/quantization/). +> **Note:** When using Quantization to optimize vector search, you can use this feature to `rescore` binary vectors against new byte vectors. With double the speedup, you will be able to achieve a better result than if you rescored with float vectors. With each byte vector quantized at the binary level, the result will deliver unparalleled efficiency and savings. To learn more about this optimization method, read our [Quantization docs](/documentation/manage-data/quantization/). ## Minor improvements and new features diff --git a/qdrant-landing/content/blog/qdrant-cpu-intel-benchmark.md b/qdrant-landing/content/blog/qdrant-cpu-intel-benchmark.md index 01fc6e1e7..ba60a4dae 100644 --- a/qdrant-landing/content/blog/qdrant-cpu-intel-benchmark.md +++ b/qdrant-landing/content/blog/qdrant-cpu-intel-benchmark.md @@ -69,4 +69,4 @@ As large companies continue to integrate sophisticated AI and machine learning t Qdrant is open source and offers a complete SaaS solution, hosted on AWS, GCP, and Azure. -Getting started is easy, either spin up a [container image](https://hub.docker.com/r/qdrant/qdrant) or start a [free Cloud instance](https://cloud.qdrant.io/login). The documentation covers [adding the data](/documentation/tutorials/bulk-upload/) to your Qdrant instance as well as [creating your indices](/documentation/tutorials/optimize/). We would love to hear about what you are building and please connect with our engineering team on [Github](https://github.com/qdrant/qdrant), [Discord](https://discord.com/invite/tdtYvXjC4h), or [LinkedIn](https://www.linkedin.com/company/qdrant). \ No newline at end of file +Getting started is easy, either spin up a [container image](https://hub.docker.com/r/qdrant/qdrant) or start a [free Cloud instance](https://cloud.qdrant.io/login). The documentation covers [adding the data](/documentation/tutorials-develop/bulk-upload/) to your Qdrant instance as well as [creating your indices](/documentation/operations/optimize/). We would love to hear about what you are building and please connect with our engineering team on [Github](https://github.com/qdrant/qdrant), [Discord](https://discord.com/invite/tdtYvXjC4h), or [LinkedIn](https://www.linkedin.com/company/qdrant). \ No newline at end of file diff --git a/qdrant-landing/content/blog/qdrant-n8n.md b/qdrant-landing/content/blog/qdrant-n8n.md index 7862a1c86..bf8a7b4f1 100644 --- a/qdrant-landing/content/blog/qdrant-n8n.md +++ b/qdrant-landing/content/blog/qdrant-n8n.md @@ -20,7 +20,7 @@ Let's go through the process of building a workflow. We'll build a chat with a c ## Prerequisites -- A running Qdrant instance. If you need one, use our [Quick start guide](/documentation/quick-start/) to set it up. +- A running Qdrant instance. If you need one, use our [Quick start guide](/documentation/quickstart/) to set it up. - An OpenAI API Key. Retrieve your key from the [OpenAI API page](https://platform.openai.com/account/api-keys) for your account. - A GitHub access token. If you need to generate one, start at the [GitHub Personal access tokens page](https://github.com/settings/tokens/). diff --git a/qdrant-landing/content/blog/qdrant-relari.md b/qdrant-landing/content/blog/qdrant-relari.md index 08af40320..4638e1c56 100644 --- a/qdrant-landing/content/blog/qdrant-relari.md +++ b/qdrant-landing/content/blog/qdrant-relari.md @@ -192,7 +192,7 @@ def log_retriever_results(retriever, dataset): return log ``` -This is the power of combining Qdrant and Relari. Instead of having to build multiple applications, slowly [upsert](/documentation/concepts/points/#upload-points), and [retrieve](/documentation/concepts/search/) data, you can use both to quickly test different parameters and instantly get results. This evaluation system is built for fast, useful iteration. +This is the power of combining Qdrant and Relari. Instead of having to build multiple applications, slowly [upsert](/documentation/manage-data/points/#upload-points), and [retrieve](/documentation/search/search/) data, you can use both to quickly test different parameters and instantly get results. This evaluation system is built for fast, useful iteration. ### Evaluate results @@ -249,7 +249,7 @@ We can even look at individual cases in the UI to get more insight. Relari and Qdrant can also be integrated to evaluate [hybrid search systems](/articles/hybrid-search/), which combine both sparse (traditional keyword-based) and dense (vector-based) search methods. This combination allows you to leverage the strengths of both approaches, potentially improving the relevance and accuracy of search results. -By using Relari’s evaluation framework alongside Qdrant’s [vector search](/advanced-search/) capabilities, you can experiment with different configurations for hybrid search. For example, you might test varying the ratio of [sparse-to-dense search results](/documentation/concepts/hybrid-queries/#hybrid-search) or adjust how each component contributes to the overall retrieval score. +By using Relari’s evaluation framework alongside Qdrant’s [vector search](/advanced-search/) capabilities, you can experiment with different configurations for hybrid search. For example, you might test varying the ratio of [sparse-to-dense search results](/documentation/search/hybrid-queries/#hybrid-search) or adjust how each component contributes to the overall retrieval score. ## Auto Prompt Optimization diff --git a/qdrant-landing/content/blog/qdrant-stars-announcement copy.md b/qdrant-landing/content/blog/qdrant-stars-announcement copy.md index 6294a2d41..e7daf58f3 100644 --- a/qdrant-landing/content/blog/qdrant-stars-announcement copy.md +++ b/qdrant-landing/content/blog/qdrant-stars-announcement copy.md @@ -75,7 +75,7 @@ Our inaugural Qdrant Stars are a diverse and talented lineup who have shown exce
Owen Colegrove
-

Owen Colegrove is the Co-Founder of SciPhi, making it easy build, deploy, and scale RAG systems using Qdrant vector search tecnology. He has Ph.D. in Physics and was previously a Quantitative Strategist at Citadel and a Researcher at CERN.

+

Owen Colegrove is the Co-Founder of SciPhi, making it easy build, deploy, and scale RAG systems using Qdrant vector search technology. He has Ph.D. in Physics and was previously a Quantitative Strategist at Citadel and a Researcher at CERN.

@@ -172,7 +172,7 @@ Share your journey with vector search technologies and how you plan to contribut #### Nominate a Qdrant Star -Do you know someone who could be our next Qdrant Star? Please submit your nomination through our [nomination form](hhttps://forms.gle/jsEJ9zjdaxqk7F5b9), explaining why they're a great fit. Your recommendation could help us find the next standout ambassador. +Do you know someone who could be our next Qdrant Star? Please submit your nomination through our [nomination form](https://forms.gle/jsEJ9zjdaxqk7F5b9), explaining why they're a great fit. Your recommendation could help us find the next standout ambassador. #### Learn More diff --git a/qdrant-landing/content/blog/qdrant-unstructured.md b/qdrant-landing/content/blog/qdrant-unstructured.md index 0cc9473f7..2d8147004 100644 --- a/qdrant-landing/content/blog/qdrant-unstructured.md +++ b/qdrant-landing/content/blog/qdrant-unstructured.md @@ -18,7 +18,7 @@ In this blog post, we'll demonstrate how to load data into Qdrant from the chann ### Prerequisites -- A running Qdrant instance. Refer to our [Quickstart guide](/documentation/quick-start/) to set up an instance. +- A running Qdrant instance. Refer to our [Quickstart guide](/documentation/quickstart/) to set up an instance. - A Discord bot token. Generate one [here](https://discord.com/developers/applications) after adding the bot to your server. - Unstructured CLI with the required extras. For more information, see the Discord [Getting Started guide](https://discord.com/developers/docs/getting-started). Install it with the following command: @@ -50,7 +50,7 @@ unstructured-ingest discord --help ### Ingesting into Qdrant -Before loading the data, set up a collection with the information you need for the following REST call. In this example we use a local Huggingface model generating 384-dimensional embeddings. You can create a Qdrant [API key](/documentation/cloud/authentication/#create-api-keys) and set names for your Qdrant [collections](/documentation/concepts/collections/). +Before loading the data, set up a collection with the information you need for the following REST call. In this example we use a local Huggingface model generating 384-dimensional embeddings. You can create a Qdrant [API key](/documentation/cloud/authentication/#create-api-keys) and set names for your Qdrant [collections](/documentation/manage-data/collections/). We set up the collection with the following command: diff --git a/qdrant-landing/content/blog/qdrant-x-dust-how-vector-search-helps-make-work-work-better-stan-polu-vector-space-talk-010.md b/qdrant-landing/content/blog/qdrant-x-dust-how-vector-search-helps-make-work-work-better-stan-polu-vector-space-talk-010.md index 1f6feaa3c..1d7ae4c99 100644 --- a/qdrant-landing/content/blog/qdrant-x-dust-how-vector-search-helps-make-work-work-better-stan-polu-vector-space-talk-010.md +++ b/qdrant-landing/content/blog/qdrant-x-dust-how-vector-search-helps-make-work-work-better-stan-polu-vector-space-talk-010.md @@ -243,7 +243,7 @@ So it's not as much as performance, but obviously performance matters, and that' Stanislas Polu: What I mentioned is that it's interesting because today the retrieval is noisy, because the embedders are not perfect, which is an interesting point. -Sorry, I'm double clicking, but I'll come back. The embedded are really not perfect. Are really not perfect. So that's interesting. When Qdrant release kind of optimization for [storage of vectors](https://qdrant.tech/documentation/concepts/storage/), they come with obviously warnings that you may have a loss. +Sorry, I'm double clicking, but I'll come back. The embedded are really not perfect. Are really not perfect. So that's interesting. When Qdrant release kind of optimization for [storage of vectors](https://qdrant.tech/documentation/manage-data/storage/), they come with obviously warnings that you may have a loss. Of precision because of the compression, et cetera, et cetera. And that's funny, like in all kind of retrieval and mental generation world, it really doesn't matter. We take all the performance we can because the loss of precision coming from compression of those vectors at the vector DB level are completely negligible compared to. The holon fuckness of the embedders in. diff --git a/qdrant-landing/content/blog/rag-evaluation-guide.md b/qdrant-landing/content/blog/rag-evaluation-guide.md index 4606c2bee..abc574498 100644 --- a/qdrant-landing/content/blog/rag-evaluation-guide.md +++ b/qdrant-landing/content/blog/rag-evaluation-guide.md @@ -87,7 +87,7 @@ Quotient AI is another platform designed to streamline the evaluation of RAG sys Improper data ingestion can cause the loss of important contextual information, which is critical for generating accurate and coherent responses. Also, inconsistent data ingestion can cause the system to produce unreliable and inconsistent responses, undermining user trust and satisfaction. -Vector databases support different [indexing](https://qdrant.tech/documentation/concepts/indexing/) techniques. In order to know if you are ingesting data properly, you should always check how changes in variables related to indexing techniques affect data ingestion. +Vector databases support different [indexing](https://qdrant.tech/documentation/manage-data/indexing/) techniques. In order to know if you are ingesting data properly, you should always check how changes in variables related to indexing techniques affect data ingestion. #### Solution: Pay attention to how your data is chunked @@ -129,7 +129,7 @@ By evaluating the retrieval quality using these metrics, you can assess the effe Each new LLM with a larger context window claims to render RAG obsolete. However, studies like "[Lost in the Middle](https://arxiv.org/abs/2307.03172)" demonstrate that feeding entire documents to LLMs can diminish their ability to answer questions effectively. Therefore, the retrieval algorithm is crucial for fetching the most relevant data in the RAG system. -**Configure dense vector retrieval:** You need to choose the right [similarity metric](https://qdrant.tech/documentation/concepts/search/) to get the best retrieval quality. Metrics used in dense vector retrieval include Cosine Similarity, Dot Product, Euclidean Distance, and Manhattan Distance. +**Configure dense vector retrieval:** You need to choose the right [similarity metric](https://qdrant.tech/documentation/search/search/) to get the best retrieval quality. Metrics used in dense vector retrieval include Cosine Similarity, Dot Product, Euclidean Distance, and Manhattan Distance. **Use sparse vectors & hybrid search where needed**: For sparse vectors, the algorithm choice of BM-25, SPLADE, or BM-42 will affect retrieval quality. Hybrid Search combines dense vector retrieval with sparse vector-based search. diff --git a/qdrant-landing/content/blog/series-A-funding-round.md b/qdrant-landing/content/blog/series-A-funding-round.md index 356a48ff8..98c845078 100644 --- a/qdrant-landing/content/blog/series-A-funding-round.md +++ b/qdrant-landing/content/blog/series-A-funding-round.md @@ -27,9 +27,9 @@ The rise of generative AI in the last few years has shone a spotlight on vector ## What sets Qdrant apart? -To meet the needs of the next generation of AI applications, Qdrant has always been built with four keys in mind: efficiency, scalability, performance, and flexibility. Our goal is to give our users unmatched speed and reliability, even when they are building massive-scale AI applications requiring the handling of billions of vectors. We did so by building Qdrant on Rust for performance, memory safety, and scale. Additionally, [our custom HNSW search algorithm](/articles/filterable-hnsw/) and unique [filtering](/documentation/concepts/filtering/) capabilities consistently lead to [highest RPS](/benchmarks/), minimal latency, and high control with accuracy when running large-scale, high-dimensional operations. +To meet the needs of the next generation of AI applications, Qdrant has always been built with four keys in mind: efficiency, scalability, performance, and flexibility. Our goal is to give our users unmatched speed and reliability, even when they are building massive-scale AI applications requiring the handling of billions of vectors. We did so by building Qdrant on Rust for performance, memory safety, and scale. Additionally, [our custom HNSW search algorithm](/articles/filterable-hnsw/) and unique [filtering](/documentation/search/filtering/) capabilities consistently lead to [highest RPS](/benchmarks/), minimal latency, and high control with accuracy when running large-scale, high-dimensional operations. -Beyond performance, we provide our users with the most flexibility in cost savings and deployment options. A combination of cutting-edge efficiency features, like [built-in compression options](/documentation/guides/quantization/), [multitenancy](/documentation/guides/multiple-partitions/) and the ability to [offload data to disk](/documentation/concepts/storage/), dramatically reduce memory consumption. Committed to privacy and security, crucial for modern AI applications, Qdrant now also offers on-premise and hybrid SaaS solutions, meeting diverse enterprise needs in a data-sensitive world. This approach, coupled with our open-source foundation, builds trust and reliability with engineers and developers, making Qdrant a game-changer in the vector database domain. +Beyond performance, we provide our users with the most flexibility in cost savings and deployment options. A combination of cutting-edge efficiency features, like [built-in compression options](/documentation/manage-data/quantization/), [multitenancy](/documentation/manage-data/multitenancy/) and the ability to [offload data to disk](/documentation/manage-data/storage/), dramatically reduce memory consumption. Committed to privacy and security, crucial for modern AI applications, Qdrant now also offers on-premise and hybrid SaaS solutions, meeting diverse enterprise needs in a data-sensitive world. This approach, coupled with our open-source foundation, builds trust and reliability with engineers and developers, making Qdrant a game-changer in the vector database domain. ## What's next? diff --git a/qdrant-landing/content/blog/superlinked-multimodal-search.md b/qdrant-landing/content/blog/superlinked-multimodal-search.md index 72d666944..e88e602b2 100644 --- a/qdrant-landing/content/blog/superlinked-multimodal-search.md +++ b/qdrant-landing/content/blog/superlinked-multimodal-search.md @@ -57,7 +57,7 @@ This flexibility with weights allows users to rapidly iterate, experiment, and i **SuperLinked Framework Setup:** Once you [**setup the Superlinked server**](https://github.com/superlinked/superlinked-recipes/tree/main/projects/hotel-search), most of the prototype work is done right out of the [**sample notebook**](https://github.com/superlinked/superlinked-recipes/blob/main/projects/hotel-search/notebooks/superlinked-queries.ipynb). Once ready, you can host from a GitHub repository and deploy via Actions. -**Qdrant Vector Database:** The easiest way to store vectors is to [**create a free Qdrant Cloud cluster**](https://cloud.qdrant.io/login). We have simple docs that show you how to [**grab the API key**](/documentation/quickstart-cloud/) and upsert your new vectors and run some basic searches. For this demo, we have deployed a live Qdrant Cloud cluster. +**Qdrant Vector Database:** The easiest way to store vectors is to [**create a free Qdrant Cloud cluster**](https://cloud.qdrant.io/login). We have simple docs that show you how to [**grab the API key**](/documentation/cloud-quickstart/) and upsert your new vectors and run some basic searches. For this demo, we have deployed a live Qdrant Cloud cluster. **OpenAI API Key:** For natural language queries and generating the weights you will need an OpenAI API key diff --git a/qdrant-landing/content/blog/using-qdrant-and-langchain.md b/qdrant-landing/content/blog/using-qdrant-and-langchain.md index ce805855b..3d27800b1 100644 --- a/qdrant-landing/content/blog/using-qdrant-and-langchain.md +++ b/qdrant-landing/content/blog/using-qdrant-and-langchain.md @@ -74,7 +74,7 @@ Here is what this basic tutorial will teach you: **3. Implement vector similarity search algorithms:** Second, you will create and test a chatbot that only uses the LLM. Then, you will enable the memory component offered by Qdrant. This will allow your chatbot to be modified and updated, giving it long-term memory. -**4. Optimize the chatbot's performance:** In the last step, you will query the chatbot in two ways. First query will retrieve parametric data from the LLM, while the second one will get contexual data via Qdrant. +**4. Optimize the chatbot's performance:** In the last step, you will query the chatbot in two ways. First query will retrieve parametric data from the LLM, while the second one will get contextual data via Qdrant. The goal of this exercise is to show that RAG is simple to implement via LangChain and yields much better results than using LLMs by itself. @@ -94,13 +94,13 @@ Whether you are building a bank fraud-detection system, RAG for e-commerce, or s Now that you know how Qdrant and LangChain can elevate your setup - it's time to try us out. -- Qdrant is open source and you can [quickstart locally](/documentation/quick-start/), [install it via Docker](/documentation/quick-start/), [or to Kubernetes](https://github.com/qdrant/qdrant-helm/). +- Qdrant is open source and you can [quickstart locally](/documentation/quickstart/), [install it via Docker](/documentation/quickstart/), [or to Kubernetes](https://github.com/qdrant/qdrant-helm/). - We also offer [a free-tier of Qdrant Cloud](https://cloud.qdrant.io/) for prototyping and testing. - For best integration with LangChain, read the [official LangChain documentation](https://python.langchain.com/docs/integrations/vectorstores/qdrant/). -- For all other cases, [Qdrant documentation](/documentation/integrations/langchain/) is the best place to get there. +- For all other cases, [Qdrant documentation](/documentation/frameworks/langchain/) is the best place to get there. > We offer additional support tailored to your business needs. [Contact us](https://qdrant.to/contact-us) to learn more about implementation strategies and integrations that suit your company. diff --git a/qdrant-landing/content/blog/vector-image-search-rag-vector-space-talk-008.md b/qdrant-landing/content/blog/vector-image-search-rag-vector-space-talk-008.md index bfeff817b..9d2677f0f 100644 --- a/qdrant-landing/content/blog/vector-image-search-rag-vector-space-talk-008.md +++ b/qdrant-landing/content/blog/vector-image-search-rag-vector-space-talk-008.md @@ -126,7 +126,7 @@ Noe Acache: So the training task was quite simple. And at the end, the embeddings was not learning any very complex features, so it was not really improving it. So jumping onto the areas of improvement, knowing all of that, the first thing I would do if I had to do it again will be to use the managed milboss for a better fine tuning, it would be to labyd hard examples, hard pairs. So, for instance, you know that when you have a matching pair where the similarity score is not too high or not too low, you know, it's where the model kind of struggles and you will find some good matching and also some mistakes. So it's where it kind of is interesting to level to then be able to fine tune your model and make it learn more complex things according to your tasks. Another possibility for fine tuning will be some sort of multilabel classification. So for instance, if you consider tab close, you could say, all right, those disclose contain buttons. It have a color, it have stripes. Noe Acache: -And for all of these categories, you'll get a score between zero and one. And concatenating all these scores together, you can get an embedding which you can put in a vector database for your vector search. It's kind of hard to scale because you need to do a specific model and labeling for each type of object. And I really wonder how Google lens does because their algorithm work very well. So are they working more like with this kind of functioning or this kind of functioning? So if anyone had any thought on that or any idea, again, I'd be happy to talk about it afterwards. And finally, I feel like we made a lot of advancements in multimodal training, trying to combine text inputs with image. We've made input to build some kind of complex embeddings. And how great would it be to have an image embeding you could guide with text. +And for all of these categories, you'll get a score between zero and one. And concatenating all these scores together, you can get an embedding which you can put in a vector database for your vector search. It's kind of hard to scale because you need to do a specific model and labeling for each type of object. And I really wonder how Google lens does because their algorithm work very well. So are they working more like with this kind of functioning or this kind of functioning? So if anyone had any thought on that or any idea, again, I'd be happy to talk about it afterwards. And finally, I feel like we made a lot of advancements in multimodal training, trying to combine text inputs with image. We've made input to build some kind of complex embeddings. And how great would it be to have an image embedding you could guide with text. Noe Acache: So you could just like when creating an embedding of your image, just say, all right, here, I don't care about the movements, I only care about the features on the object, for instance. And then it will learn an embedding according to your task without any fine tuning. I really feel like with the current state of the arts we are able to do this. I mean, we need to do it, but the technology is ready. diff --git a/qdrant-landing/content/blog/vector-search-for-content-based-video-recommendation-gladys-and-sam-vector-space-talk-012.md b/qdrant-landing/content/blog/vector-search-for-content-based-video-recommendation-gladys-and-sam-vector-space-talk-012.md index 82399b8c3..ee99e2b08 100644 --- a/qdrant-landing/content/blog/vector-search-for-content-based-video-recommendation-gladys-and-sam-vector-space-talk-012.md +++ b/qdrant-landing/content/blog/vector-search-for-content-based-video-recommendation-gladys-and-sam-vector-space-talk-012.md @@ -158,7 +158,7 @@ Demetrios: And so you kind of touched on this earlier, but can you say it again? Because I don't know if I fully grasped it. Where are all the places in the system that you are evaluating? Because it's not just the output. Right. And how do you look at evaluation as a system rather than just evaluating the output every once in a while? Sourabh Agrawal: -Yeah, so I mean, what we do is we plug with every part. So even if you start with retrieval, so we have a high level check where we look at the quality of retrieved context. And then we also have evaluations for every part of this retrieval pipeline. So if you're doing query rewrite, if you're doing re ranking, if you're doing sub question, we have evaluations for all of them. In fact, we have worked closely with the llama index team to kind of integrate with all of their modular pipelines. Secondly, once we cross the retrieval step, we have around five to six matrices on this retrieval part. Then we look at the response generation. We have their evaluations for different criterias. +Yeah, so I mean, what we do is we plug with every part. So even if you start with retrieval, so we have a high level check where we look at the quality of retrieved context. And then we also have evaluations for every part of this retrieval pipeline. So if you're doing query rewrite, if you're doing re ranking, if you're doing sub question, we have evaluations for all of them. In fact, we have worked closely with the llama index team to kind of integrate with all of their modular pipelines. Secondly, once we cross the retrieval step, we have around five to six matrices on this retrieval part. Then we look at the response generation. We have their evaluations for different criteria. Sourabh Agrawal: So conciseness, completeness, safety, jailbreaks, prompt injections, as well as you can define your custom guidelines. So you can say that, okay, if the user is asking anything and related to code, the output should also give an example code snippet so you can just in plain English, define this guideline. And we check for that. And then finally, like zooming out, we also have checks. We look at conversations as a whole, how the user is satisfied, how many turns it requires for them to, for the chatbot or the LLM to answer the user. Yeah, that's how we look at the whole evaluations as a whole. diff --git a/qdrant-landing/content/blog/vsd25-post-event.md b/qdrant-landing/content/blog/vsd25-post-event.md index 9dd0e3d40..94695ac9a 100644 --- a/qdrant-landing/content/blog/vsd25-post-event.md +++ b/qdrant-landing/content/blog/vsd25-post-event.md @@ -54,7 +54,7 @@ André highlighted the underlying forces driving this shift: We are convinced that if AI is going to evolve beyond static assistants, it needs a **retrieval layer built for unstructured data and agent workflows**. -Next on stage, our Co-Founder and CTO [**Andrey Vasnetsov**](https://www.linkedin.com/in/andrey-vasnetsov-75268897/) emphazised our belief that ‘vector database’ is actually the wrong term to describe what we are building at Qdrant. **Qdrant is not “a vector database”** because vectors themselves are not data, but representations. +Next on stage, our Co-Founder and CTO [**Andrey Vasnetsov**](https://www.linkedin.com/in/andrey-vasnetsov-75268897/) emphasized our belief that ‘vector database’ is actually the wrong term to describe what we are building at Qdrant. **Qdrant is not “a vector database”** because vectors themselves are not data, but representations. ![Not a Vector DB](/blog/vsd25-post-event/keynote-notvdb.png) diff --git a/qdrant-landing/content/blog/what-is-vector-similarity.md b/qdrant-landing/content/blog/what-is-vector-similarity.md index 9f359db4a..fd5789780 100644 --- a/qdrant-landing/content/blog/what-is-vector-similarity.md +++ b/qdrant-landing/content/blog/what-is-vector-similarity.md @@ -145,15 +145,15 @@ The vector index in Qdrant employs the Hierarchical Navigable Small World (HNSW) ### Scalability -For massive datasets and demanding workloads, Qdrant supports [distributed deployment](/documentation/guides/distributed_deployment/) from v0.8.0. In this mode, you can set up a Qdrant cluster and distribute data across multiple nodes, enabling you to maintain high performance and availability even under increased workloads. Clusters support sharding and replication, and harness the Raft consensus algorithm to manage node coordination. +For massive datasets and demanding workloads, Qdrant supports [distributed deployment](/documentation/operations/distributed_deployment/) from v0.8.0. In this mode, you can set up a Qdrant cluster and distribute data across multiple nodes, enabling you to maintain high performance and availability even under increased workloads. Clusters support sharding and replication, and harness the Raft consensus algorithm to manage node coordination. -Qdrant also supports vector [quantization](/documentation/guides/quantization/) to reduce memory footprint and speed up vector similarity searches, making it very effective for large-scale applications where efficient resource management is critical. +Qdrant also supports vector [quantization](/documentation/manage-data/quantization/) to reduce memory footprint and speed up vector similarity searches, making it very effective for large-scale applications where efficient resource management is critical. There are three quantization strategies you can choose from - scalar quantization, binary quantization and product quantization - which will help you control the trade-off between storage efficiency, search accuracy and speed. ### Security -Qdrant offers several [security features](/documentation/guides/security/) to help protect data and access to the vector store: +Qdrant offers several [security features](/documentation/operations/security/) to help protect data and access to the vector store: - API Key Authentication: This helps secure API access to Qdrant Cloud with static or read-only API keys. - JWT-Based Access Control: You can also enable more granular access control through JSON Web Tokens (JWT), and opt for restricted access to specific parts of the stored data while building Role-Based Access Control (RBAC). @@ -167,7 +167,7 @@ In order to achieve top performance in vector similarity searches, Qdrant employ **Support for Dense and Sparse Vectors**: Qdrant supports both dense and sparse vector representations. While dense vectors are most common, you may encounter situations where the dataset contains a range of specialized domain-specific keywords. [Sparse vectors](/articles/sparse-vectors/) shine in such scenarios. Sparse vectors are vector representations of data where most elements are zero. -**Multitenancy**: Qdrant supports [multitenancy](/documentation/guides/multiple-partitions/) by allowing vectors to be partitioned by payload within a single collection. Using this you can isolate each user's data, and avoid creating separate collections for each user. In order to ensure indexing performance, Qdrant also offers ways to bypass the construction of a global vector index, so that you can index vectors for each user independently. +**Multitenancy**: Qdrant supports [multitenancy](/documentation/manage-data/multitenancy/) by allowing vectors to be partitioned by payload within a single collection. Using this you can isolate each user's data, and avoid creating separate collections for each user. In order to ensure indexing performance, Qdrant also offers ways to bypass the construction of a global vector index, so that you can index vectors for each user independently. **IO Optimizations**: If your data doesn’t fit into the memory, it may require storing on disk. To [optimize disk IO performance](/articles/io_uring/), Qdrant offers io_uring based *async uring* storage backend on Linux-based systems. Benchmarks show that it drastically helps reduce operating system overhead from disk IO. @@ -211,7 +211,7 @@ We have just about witnessed the tip of the iceberg in terms of what vector simi Ready to implement vector similarity in your AI applications? Explore Qdrant's vector database to enhance your data retrieval and AI capabilities. For additional resources and documentation, visit: -- [Quick Start Guide](/documentation/quick-start/) +- [Quick Start Guide](/documentation/quickstart/) - [Documentation](/documentation/) We are always available on our [Discord channel](https://qdrant.to/discord) to answer any questions you might have. You can also sign up for our [newsletter](/subscribe/) to stay ahead of the curve. diff --git a/qdrant-landing/content/course/_index.md b/qdrant-landing/content/course/_index.md index b98f63ab8..d3ed30a61 100644 --- a/qdrant-landing/content/course/_index.md +++ b/qdrant-landing/content/course/_index.md @@ -1,5 +1,5 @@ --- -title: "Welcome to Qdrant Academy" +title: "Qdrant Academy" description: Master vector search and AI-powered applications with Qdrant Academy. Free, self-paced courses guide you from beginner to expert with hands-on projects, code notebooks, and certification. weight: 50 --- @@ -12,13 +12,11 @@ Qdrant Academy is your step-by-step learning hub for mastering vector search, hy Whether you’re new to Qdrant or building production-grade systems, our guided courses help you go from beginner to expert, one module at a time. -Qdrant Academy currently offers one comprehensive course, but more are on the way! Register your interest for upcoming courses below, or take the available course and [get certified](https://train.qdrant.dev)! - ## Available Now -{{< course-card +{{< course-card title="Qdrant Essentials Course" - image="/icons/outline/rocket-white-light.svg" + image="/icons/outline/rocket-white-light.svg" link="/course/essentials/" >}} **What you’ll gain:** @@ -30,7 +28,24 @@ Qdrant Academy currently offers one comprehensive course, but more are on the wa - Ecosystem Integrations (Bonus)

Time to Complete: 9-12 hours
-Includes: videos, code notebooks, projects, walkthroughs +Includes: videos, code notebooks, projects, certification +{{< /course-card >}} + +{{< course-card + title="Multi-Vector Search Course" + image="/icons/outline/similarity-blue.svg" + link="/course/multi-vector-search/" +>}} +**What you’ll gain:** +- Late Interaction Models and MaxSim Scoring +- ColBERT for Text Search +- ColPali for Visual Document Search +- Multi-Stage Retrieval Pipelines +- Quantization and Pooling Techniques +- MUVERA Indexing for Large-Scale Search +

+Time to Complete: 4-6 hours
+Includes: videos, code notebooks, projects, certification {{< /course-card >}} ## Upcoming Courses diff --git a/qdrant-landing/content/course/essentials/_index.md b/qdrant-landing/content/course/essentials/_index.md index 71604f51d..0190cbc98 100644 --- a/qdrant-landing/content/course/essentials/_index.md +++ b/qdrant-landing/content/course/essentials/_index.md @@ -175,7 +175,7 @@ Build the vector search skills that matter: hybrid retrieval, multivector rerank - ML Platforms & Analytics (Tensorlake, Vectorize.io, Superlinked, Quotient)

-

→ Start day 7

+

→ Start Day 7

{{< /accordion >}} diff --git a/qdrant-landing/content/course/essentials/certification/_index.md b/qdrant-landing/content/course/essentials/certification/_index.md index 2c282e524..73ce12ec7 100644 --- a/qdrant-landing/content/course/essentials/certification/_index.md +++ b/qdrant-landing/content/course/essentials/certification/_index.md @@ -1,6 +1,7 @@ --- title: "Qdrant Essentials Certification" description: Get officially certified by Qdrant today! +isLesson: true weight: 100 --- diff --git a/qdrant-landing/content/course/essentials/day-0/building-simple-vector-search.md b/qdrant-landing/content/course/essentials/day-0/building-simple-vector-search.md index 2385813c6..4c02d3e7d 100644 --- a/qdrant-landing/content/course/essentials/day-0/building-simple-vector-search.md +++ b/qdrant-landing/content/course/essentials/day-0/building-simple-vector-search.md @@ -2,6 +2,7 @@ title: "Implementing a Basic Vector Search" description: Learn how to build a basic vector search in Qdrant. Create collections, insert vectors, and run your first similarity search step-by-step with Python. weight: 3 +isLesson: true --- {{< date >}} Day 0 {{< /date >}} @@ -54,7 +55,7 @@ client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API ## Step 4: Create a Collection -A [collection](/documentation/concepts/collections/) in Qdrant is like a table in relational databases - a container for storing vectors and their metadata. When creating a collection, specify: +A [collection](/documentation/manage-data/collections/) in Qdrant is like a table in relational databases - a container for storing vectors and their metadata. When creating a collection, specify: - **Name**: A unique identifier for the collection - **Vector Configuration**: @@ -77,7 +78,7 @@ client.create_collection( Expected output: `True` (indicating successful creation) -**Distance metrics explained** ([learn more](/documentation/concepts/collections/#distance-metrics)): +**Distance metrics explained** ([learn more](/documentation/manage-data/collections/#distance-metrics)): - **Euclidean**: Measures straight-line distance between points in space - **Cosine**: Measures the angle between vectors, focusing on orientation rather than magnitude - **Dot**: Measures the dot product of vectors, capturing both magnitude and direction @@ -96,7 +97,7 @@ The `get_collections()` method returns all collections in your Qdrant instance, ## Step 6: Insert Points into the Collection -[Points](/documentation/concepts/points/) are the core data entities in Qdrant. Each point contains: +[Points](/documentation/manage-data/points/) are the core data entities in Qdrant. Each point contains: - **ID**: A unique identifier - **Vector Data**: An array of numerical values representing the data point in vector space diff --git a/qdrant-landing/content/course/essentials/day-0/pitstop-project.md b/qdrant-landing/content/course/essentials/day-0/pitstop-project.md index 15bce0c75..c1b76f962 100644 --- a/qdrant-landing/content/course/essentials/day-0/pitstop-project.md +++ b/qdrant-landing/content/course/essentials/day-0/pitstop-project.md @@ -2,6 +2,7 @@ title: "Project: Building Your First Vector Search System" description: Apply your Qdrant skills to build a complete vector search system. Create collections, insert data, run similarity and filtered searches, and share your results. weight: 4 +isLesson: true --- {{< date >}} Day 0 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-0/qdrant-cloud.md b/qdrant-landing/content/course/essentials/day-0/qdrant-cloud.md index be143d6d3..f8cea980b 100644 --- a/qdrant-landing/content/course/essentials/day-0/qdrant-cloud.md +++ b/qdrant-landing/content/course/essentials/day-0/qdrant-cloud.md @@ -2,6 +2,7 @@ title: "Qdrant Setup" description: Set up your Qdrant Cloud cluster in minutes. Learn to create collections, manage data, access the Web UI, and connect securely from Python. weight: 2 +isLesson: true --- {{< date >}} Day 0 {{< /date >}} @@ -84,7 +85,7 @@ you’ll get a detailed view with these tabs: * **Search Quality Tab**: Evaluate and benchmark retrieval precision against ground truth. Tune parameters and measure the impact on accuracy. -* **Snapshots Tab**: Manage backups for this collection. Create a [snapshot](/documentation/concepts/snapshots/), restore it later, or migrate it to another cluster. +* **Snapshots Tab**: Manage backups for this collection. Create a [snapshot](/documentation/operations/snapshots/), restore it later, or migrate it to another cluster. * **Visualize Tab**: Explore your vector space with an interactive 2D projection. See clusters, spot outliers, and build intuition about your embeddings. diff --git a/qdrant-landing/content/course/essentials/day-1/chunking-strategies.md b/qdrant-landing/content/course/essentials/day-1/chunking-strategies.md index 229bc4ae1..fc79b646d 100644 --- a/qdrant-landing/content/course/essentials/day-1/chunking-strategies.md +++ b/qdrant-landing/content/course/essentials/day-1/chunking-strategies.md @@ -2,6 +2,7 @@ title: "Text Chunking Strategies" description: Learn how to split text into meaningful chunks for vector search. Compare six chunking strategies and discover how metadata improves retrieval precision in Qdrant. weight: 4 +isLesson: true --- {{< date >}} Day 1 {{< /date >}} @@ -54,7 +55,7 @@ This is where chunking comes in. The goal is to have chunks By breaking a document into focused chunks, each chunk gets its own vector that accurately represents a specific idea. This allows the search to be far more precise. -**Example:** Consider a multi-page Document like the [Qdrant Collection Configuration Guide of Day 7](/course/essentials/day-7/collection-configuration-guide/) covering everything from HNSW to sharding and quantization. +**Example:** Consider a multi-page Document like the [Qdrant Collection Configuration Guide of Day 7](/course/essentials/day-7/) covering everything from HNSW to sharding and quantization. If a user asks: *"What does the m parameter do?"* @@ -419,7 +420,7 @@ The trade-off is computational cost. You're embedding the full document upfront | **Recursive** | Flexible, handles messy input | Heuristic, sometimes brittle | Scraped web content, mixed sources | | **Semantic** | High-quality, meaning-aware | Slower, resource-intensive | Legal, research, critical QA | -**Note**: Sometimes, it's necessary to keep the document intact. If chunking is too complicated, or the document is visually rich (diagrams, graphs etc.), you can use [VLMs](/documentation/advanced-tutorials/pdf-retrieval-at-scale/) to embed the whole page. +**Note**: Sometimes, it's necessary to keep the document intact. If chunking is too complicated, or the document is visually rich (diagrams, graphs etc.), you can use [VLMs](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) to embed the whole page. ## Adding Meaning with Metadata @@ -438,7 +439,7 @@ In Qdrant, this metadata lives in the **payload** - a JSON object attached to ea "section_title": "What Is a Vector", "chunk_index": 7, "chunk_count": 15, - "url": "https://qdrant.tech/documentation/concepts/collections/", + "url": "https://qdrant.tech/documentation/manage-data/collections/", "tags": ["qdrant", "vector search", "point", "vector", "payload"], "source_type": "documentation", "created_at": "2025-01-15T10:00:00Z", @@ -450,7 +451,7 @@ In Qdrant, this metadata lives in the **payload** - a JSON object attached to ea ### What Metadata Enables -**Disclaimer**: For performance reasons, filterable fields must be indexed using the [Payload Index](/documentation/concepts/indexing/#payload-index). +**Disclaimer**: For performance reasons, filterable fields must be indexed using the [Payload Index](/documentation/manage-data/indexing/#payload-index). **1. Filtered Search (Exact Match)** You can filter results based on exact metadata values, which is perfect for categorical data. @@ -468,7 +469,7 @@ filter = models.Filter( ``` **2. Hybrid Search with Text Filtering (Full-Text Search)** -For more powerful text-based filtering, you can combine vector search with traditional keyword search. This requires setting up a [full-text index](/documentation/concepts/indexing/#full-text-index) on a payload field. +For more powerful text-based filtering, you can combine vector search with traditional keyword search. This requires setting up a [full-text index](/documentation/manage-data/indexing/#full-text-index) on a payload field. ```python # Find vectors that also contain the keyword "HNSW" in their content filter = models.Filter( @@ -486,13 +487,13 @@ filter = models.Filter( # Top result per document - get the most relevant chunk from each source group_by = "document_id" ``` -You can read more about grouping [here](/documentation/concepts/hybrid-queries/?q=grouping#grouping). +You can read more about grouping [here](/documentation/search/hybrid-queries/?q=grouping#grouping). **4. Rich Result Display** - Original content with source attribution - Section context for better understanding - Direct links to full documents -- Creation timestamps for [freshness](/documentation/concepts/search-relevance/#time-based-score-boosting) +- Creation timestamps for [freshness](/documentation/search/search-relevance/#time-based-score-boosting) **5. Permission Control** ```python diff --git a/qdrant-landing/content/course/essentials/day-1/distance-metrics.md b/qdrant-landing/content/course/essentials/day-1/distance-metrics.md index 56e7615f4..e8ff7847c 100644 --- a/qdrant-landing/content/course/essentials/day-1/distance-metrics.md +++ b/qdrant-landing/content/course/essentials/day-1/distance-metrics.md @@ -2,13 +2,14 @@ title: "Distance Metrics" description: Learn how distance metrics like cosine, Euclidean, Manhattan, and dot product shape vector similarity in Qdrant. Discover which metric fits your data and use case. weight: 3 +isLesson: true --- {{< date >}} Day 1 {{< /date >}} # Distance Metrics -After vectors are stored, we can use their spatial properties to perform [nearest neighbor searches](/documentation/concepts/search/) that retrieve semantically similar items based on how close they are in this space. +After vectors are stored, we can use their spatial properties to perform [nearest neighbor searches](/documentation/search/search/) that retrieve semantically similar items based on how close they are in this space. The position of a vector in embedding space only reflects meaning as far as the embedding model has learned to encode it. The model and its training objective tell you what "close" means. @@ -165,4 +166,4 @@ If you are training your own model or designing custom features, use these guide * **Dot product** accounts for magnitude and direction. 4. **Experiment:** Qdrant allows you to set distance metrics per named vector, making it easy to A/B test different metrics on your specific data. -Reference: [Distance Metrics in Qdrant Documentation](/documentation/concepts/search/#metrics) \ No newline at end of file +Reference: [Distance Metrics in Qdrant Documentation](/documentation/search/search/#metrics) \ No newline at end of file diff --git a/qdrant-landing/content/course/essentials/day-1/embedding-models.md b/qdrant-landing/content/course/essentials/day-1/embedding-models.md index 078ba2a21..ba7fbcb5a 100644 --- a/qdrant-landing/content/course/essentials/day-1/embedding-models.md +++ b/qdrant-landing/content/course/essentials/day-1/embedding-models.md @@ -2,6 +2,7 @@ title: "Points, Vectors and Payloads" description: Learn Qdrant’s core data model with points, vectors, payloads, and named vectors. Compare dense, sparse, and multivectors, understand dimensionality trade-offs, and master filtering with payload indexes for precise retrieval. weight: 2 +isLesson: true --- {{< date >}} Day 1 {{< /date >}} @@ -81,7 +82,7 @@ The `indices` and `values` arrays must be the same size, and all the `indices` m There is no need to sort the sparse representation by indices, as Qdrant will perform this internally while maintaining the correct link between each index and its value. -We will cover more about sparse vectors on day 3. If you would like to read up on the subject in advance, you can find more documentation [here](/documentation/concepts/vectors/#sparse-vectors). +We will cover more about sparse vectors on day 3. If you would like to read up on the subject in advance, you can find more documentation [here](/documentation/manage-data/vectors/#sparse-vectors). ### Multivectors @@ -233,7 +234,7 @@ While vectors capture the essence of data, payloads hold structured metadata for Payloads can store textual data (descriptions, tags, categories), numerical values (dates, prices, ratings), and complex structures (nested objects, arrays). When searching for dog images, for example, the vector finds visually similar images while payload filters narrow results to images taken within the last year, tagged with "vacation," or meeting specific rating criteria. -Learn more: [Payload Documentation](/documentation/concepts/payload/) +Learn more: [Payload Documentation](/documentation/manage-data/payload/) ### Payload Types @@ -303,7 +304,7 @@ Here are some of the most common condition types: +For the complete, most up-to-date list of all available filtering conditions, please refer to the **[official Filtering documentation](/documentation/search/filtering/#filtering-conditions)**. ### Filtering Capabilities Reference @@ -381,7 +382,7 @@ client.create_payload_index( When filters are highly selective, Qdrant's query planner may bypass vector indexing entirely and use payload indexes for faster results. -For comprehensive filtering examples and advanced usage patterns, see the [Filtering Documentation](/documentation/concepts/filtering/) and [Complete Guide to Filtering in Vector Search](/articles/vector-search-filtering/). +For comprehensive filtering examples and advanced usage patterns, see the [Filtering Documentation](/documentation/search/filtering/) and [Complete Guide to Filtering in Vector Search](/articles/vector-search-filtering/). ## Key Takeaways diff --git a/qdrant-landing/content/course/essentials/day-1/movie-search-system.md b/qdrant-landing/content/course/essentials/day-1/movie-search-system.md index b092dfbe3..79bdc825c 100644 --- a/qdrant-landing/content/course/essentials/day-1/movie-search-system.md +++ b/qdrant-landing/content/course/essentials/day-1/movie-search-system.md @@ -2,6 +2,7 @@ title: "Demo: Semantic Movie Search" description: Build a semantic movie search with Qdrant. Compare chunking strategies, embed descriptions, and combine cosine similarity with metadata filters and grouping for accurate, theme-aware recommendations. weight: 5 +isLesson: true --- {{< date >}} Day 1 {{< /date >}} @@ -234,7 +235,7 @@ Query: 'alien invasion' ## Step 6: Advanced Features -Note: If you are already familiar Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/concepts/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [Day 2](/content/course/essentials/day-2/_index.md) of this course. +Note: If you are already familiar Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/manage-data/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [Day 2](/course/essentials/day-2/) of this course. ### Filtering by Metadata diff --git a/qdrant-landing/content/course/essentials/day-1/pitstop-project.md b/qdrant-landing/content/course/essentials/day-1/pitstop-project.md index 51c230f01..a9a9d6e60 100644 --- a/qdrant-landing/content/course/essentials/day-1/pitstop-project.md +++ b/qdrant-landing/content/course/essentials/day-1/pitstop-project.md @@ -2,6 +2,7 @@ title: "Project: Building a Semantic Search Engine" description: Build a semantic search engine with Qdrant. Compare chunking strategies, index embeddings, and query by meaning to discover what works best for your domain. weight: 6 +isLesson: true --- {{< date >}} Day 1 {{< /date >}} @@ -134,7 +135,7 @@ def paragraph_chunks(text): ### Step 4: Create Collections and Process Data -Note: If you are already familiar with Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/concepts/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [day 2](/content/course/essentials/day-2/_index.md) of this course. +Note: If you are already familiar with Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/manage-data/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [day 2](/course/essentials/day-2/) of this course. ```python collection_name = "day1_semantic_search" diff --git a/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md b/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md index 488c80103..314fea953 100644 --- a/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md +++ b/qdrant-landing/content/course/essentials/day-2/collection-tuning-demo.md @@ -2,6 +2,7 @@ title: "Demo: HNSW Performance Tuning" description: Tune Qdrant’s HNSW index for speed and precision. Optimize bulk uploads, test filters, and benchmark performance on a real 100K OpenAI embedding dataset. weight: 4 +isLesson: true --- {{< date >}} Day 2 {{< /date >}} @@ -408,7 +409,7 @@ else: ## Step 10: Create Payload Indexes -Create a [full‑text index](/documentation/concepts/indexing/#full-text-index) for faster filtering. +Create a [full‑text index](/documentation/manage-data/indexing/#full-text-index) for faster filtering. ```python # Create a payload index for 'text' so filters use an index, not a scan. @@ -525,6 +526,6 @@ print("=" * 60) - [Qdrant Documentation](/documentation/) - Complete technical reference - [HNSW Paper](https://arxiv.org/abs/1603.09320) - Original algorithm research - [Qdrant Cloud](https://cloud.qdrant.io/) - Managed vector search service -- [Performance Tuning Guide](/documentation/guides/optimize/) - Advanced optimization techniques +- [Performance Tuning Guide](/documentation/operations/optimize/) - Advanced optimization techniques **Ready for the pitstop project?** Now it's your turn to optimize performance with your own dataset and use case. You'll apply these same techniques to your domain-specific data and measure the real-world impact of different HNSW parameters and indexing strategies. \ No newline at end of file diff --git a/qdrant-landing/content/course/essentials/day-2/filterable-hnsw.md b/qdrant-landing/content/course/essentials/day-2/filterable-hnsw.md index de12db27c..decae7327 100644 --- a/qdrant-landing/content/course/essentials/day-2/filterable-hnsw.md +++ b/qdrant-landing/content/course/essentials/day-2/filterable-hnsw.md @@ -2,13 +2,14 @@ title: "Combining Vector Search and Filtering" description: Learn how Qdrant combines HNSW vector search with payload filtering. Understand Filterable HNSW, query planning, and payload indexing for accurate, high-performance retrieval. weight: 3 +isLesson: true --- {{< date >}} Day 2 {{< /date >}} # Combining Vector Search and Filtering -We've talked about how Qdrant uses the [HNSW](/documentation/concepts/indexing/#filterable-index) graph to efficiently search dense vectors. But in real-world applications, you'll often want to constrain your search using filters. This creates unique challenges for graph traversal that Qdrant solves elegantly. +We've talked about how Qdrant uses the [HNSW](/documentation/manage-data/indexing/#filterable-index) graph to efficiently search dense vectors. But in real-world applications, you'll often want to constrain your search using filters. This creates unique challenges for graph traversal that Qdrant solves elegantly.
-Production vector search engines face an inevitable scaling challenge: memory requirements grow with dataset size, while search latency demands vectors remain in fast storage. [Quantization](/documentation/guides/quantization/) provides the solution by compressing vector representations while maintaining retrieval quality - but the method you choose fundamentally determines your system's performance characteristics. +Production vector search engines face an inevitable scaling challenge: memory requirements grow with dataset size, while search latency demands vectors remain in fast storage. [Quantization](/documentation/manage-data/quantization/) provides the solution by compressing vector representations while maintaining retrieval quality - but the method you choose fundamentally determines your system's performance characteristics. ## The Memory Economics @@ -64,7 +65,7 @@ client.create_collection( ) ``` -> [Check out](/documentation/guides/quantization/#setting-up-scalar-quantization) how to set up scalar quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients. +> [Check out](/documentation/manage-data/quantization/#setting-up-scalar-quantization) how to set up scalar quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients. ## Binary Quantization @@ -74,7 +75,7 @@ The computational advantages are substantial. Bitwise operations enable distance However, binary quantization demands specific model characteristics for optimal performance. The technique works best with high-dimensional vectors (≥1024 dimensions) that exhibit centered value distributions around zero. Models like OpenAI's text-embedding-ada-002 and Cohere's embed-english-v2.0 have been validated for binary compatibility, but other models may experience significant accuracy degradation. -> **Update:** Starting from Qdrant **v1.15.0**, [two additional quantization types](/documentation/guides/quantization/#15-bit-and-2-bit-quantization) were introduced: **1.5-bit** and **2-bit binary quantization**. +> **Update:** Starting from Qdrant **v1.15.0**, [two additional quantization types](/documentation/manage-data/quantization/#15-bit-and-2-bit-quantization) were introduced: **1.5-bit** and **2-bit binary quantization**. > These methods provide a useful middle ground: they are more aggressive than scalar quantization but offer better precision than standard binary quantization. They also address one of binary quantization’s main weaknesses: **handling values close to zero**. > > **Additionally:** [Asymmetric quantization](#asymmetric-quantization) was added. This method allows combining different quantization strategies for queries and documents, helping balance **accuracy** and **compression efficiency**. @@ -97,7 +98,7 @@ client.create_collection( ) ``` -> [Check out](/documentation/guides/quantization/#setting-up-binary-quantization) how to set up binary quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients. +> [Check out](/documentation/manage-data/quantization/#setting-up-binary-quantization) how to set up binary quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients. ## Product Quantization @@ -125,7 +126,7 @@ client.create_collection( ) ``` -> [Check out](/documentation/guides/quantization/#setting-up-product-quantization) how to set up product quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients. +> [Check out](/documentation/manage-data/quantization/#setting-up-product-quantization) how to set up product quantization in **TypeScript**, **Rust**, **Java**, **C#**, and **Go** clients. ## Quantization Comparison @@ -136,14 +137,14 @@ client.create_collection( | Product | 0.7 | 0.5x | up to 64x |å *For compatible models -> [Check out](/documentation/guides/quantization/#how-to-choose-the-right-quantization-method) how the new **1.5-bit** and **2-bit binary quantization** methods compare to classical binary quantization. They offer a balanced middle ground between **binary** and **scalar** approaches. +> [Check out](/documentation/manage-data/quantization/#how-to-choose-the-right-quantization-method) how the new **1.5-bit** and **2-bit binary quantization** methods compare to classical binary quantization. They offer a balanced middle ground between **binary** and **scalar** approaches. ## Dual Storage Architecture Qdrant's quantization implementation maintains both compressed and original vectors, enabling flexible deployment strategies and safe experimentation. This dual storage approach allows you to switch quantization methods, adjust parameters, or disable quantization entirely without data re-ingestion - a critical advantage for production systems where data pipeline complexity must be minimized. -> Check out our **[quantization tips](/documentation/guides/quantization/#quantization-tips)** +> Check out our **[quantization tips](/documentation/manage-data/quantization/#quantization-tips)** The default configuration stores both representations in RAM, providing fast search with quantized vectors and exact scoring with originals when needed. However, this negates memory savings. The optimal production pattern places original vectors on disk (`on_disk=True`) while keeping quantized vectors in RAM (`always_ram=True`). This configuration delivers the best of both worlds: rapid quantized search with the ability to perform exact rescoring by reading only the small candidate set from disk storage. diff --git a/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md b/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md index 735b7b975..acf56d373 100644 --- a/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md +++ b/qdrant-landing/content/course/essentials/day-5/colbert-multivectors.md @@ -2,6 +2,7 @@ title: "Multivectors for Late Interaction Models" description: Learn how Qdrant supports late interaction models like ColBERT and ColPali using multivectors for token-level precision, enabling fine-grained, context-aware text and visual document retrieval. weight: 2 +isLesson: true --- {{< date >}} Day 5 {{< /date >}} @@ -26,9 +27,9 @@ Late-interaction models such as ColBERT retain per-token document vectors. At se ## Late Interaction: Token-Level Precision -Qdrant implements this powerful technique through [multivector representations](/documentation/concepts/vectors/#multivectors). A multivector field holds an ordered list of subvectors, each of which captures a different token of the document. +Qdrant implements this powerful technique through [multivector representations](/documentation/manage-data/vectors/#multivectors). A multivector field holds an ordered list of subvectors, each of which captures a different token of the document. -At query time, Qdrant performs the late interaction scoring. It compares every query token embedding $ q_i $ with each document token embedding $ d_j $. Only the highest score per query token is retained and these top scores are then summed. This mechanism, called MaxSim, delivers fine-grained relevance that respects the structure of your content. +At query time, Qdrant performs late interaction scoring. It compares every query token embedding $ q_i $ with each document token embedding $ d_j $. Only the highest score per query token is retained, and those top scores are then summed. This mechanism, called MaxSim, delivers fine-grained relevance that respects the structure of your content. $$ MaxSim_{\text{norm}}(Q, D) = \frac{1}{|Q|} \sum_{i=1}^{|Q|} \max_{j=1}^{|D|} \text{sim}(q_i, d_j) @@ -48,7 +49,7 @@ doc_multivectors = list(encoder.embed(["A long document about AI in medicine."]) # Returns [[token_vec1, token_vec2, ...]] ``` -The model `colbert-ir/colbertv2.0` outputs 128-dimensional vectors and is available through FastEmbed's optimized ONNX runtime. Use `.embed` for documents and `.query_embed` for queries. +The model `colbert-ir/colbertv2.0` outputs 128-dimensional vectors and is available through FastEmbed's optimized ONNX Runtime. Use `.embed` for documents and `.query_embed` for queries. ## Collection Configuration: Multivector Setup @@ -80,7 +81,7 @@ client.create_collection( ) ``` -By specifying `MAX_SIM`, you tell Qdrant to apply the late interaction scoring at query time. We explicitly disable HNSW indexing with `m=0` because the graph typically won't be used for multivectors (except in rare edge cases), so disabling it saves RAM. Without HNSW, queries use brute-force MaxSim scoring across all points, which provides maximum precision but may be slower on large collections. For better performance on larger datasets, you'll learn about retrieval-reranking patterns in the next lesson. +By specifying `MAX_SIM`, you tell Qdrant to apply late interaction scoring at query time. We explicitly disable HNSW indexing with `m=0` because the graph typically won't be used for multivectors except in rare edge cases, so disabling it saves RAM. Without HNSW, queries use brute-force MaxSim scoring across all points, which provides maximum precision but may be slower on large collections. For better performance on larger datasets, you'll learn about retrieval and reranking patterns in the next lesson. ## Querying with ColBERT @@ -102,13 +103,13 @@ hits = client.query_points( ) ``` -Qdrant performs brute-force MaxSim scoring between your query tokens and the document tokens for all points in the collection. This delivers highly precise results based on fine-grained token-level matching. Keep in mind that without HNSW indexing, this approach may be slower on large collections - in the next lesson you'll see how to combine fast approximate retrieval with ColBERT reranking for better performance. +Qdrant performs brute-force MaxSim scoring between your query tokens and the document tokens for all points in the collection. This delivers highly precise results based on fine-grained token-level matching. Keep in mind that without HNSW indexing, this approach may be slower on large collections. In the next lesson, you'll see how to combine fast approximate retrieval with ColBERT reranking for better performance. ## ColPali for Visual Documents -For documents with rich layouts, PDFs, invoices, slide decks, ColPali (Contextualized Late Interaction over PaliGemma) extends the same idea to vision. ColPali divides each page into a 32×32 grid (1,024 patches), encodes each patch with a vision-language model into 128-dimensional vectors, and treats those patch embeddings as subvectors. You use the identical multivector configuration, and Qdrant applies MaxSim on token embeddings, no matter how they were created. +For documents with rich layouts such as PDFs, invoices, and slide decks, ColPali (Contextualized Late Interaction over PaliGemma) extends the same idea to vision. ColPali divides each page into a 32x32 grid (1,024 patches), encodes each patch with a vision-language model into 128-dimensional vectors, and treats those patch embeddings as subvectors. You use the same multivector configuration, and Qdrant applies MaxSim to those embeddings regardless of how they were created. The visual approach eliminates traditional OCR and layout detection steps, processing document images directly to capture both textual content and visual structure in a single pass. This makes ColPali particularly effective for complex documents where layout and visual elements are crucial for understanding. ## Next -With multivectors in your toolkit, you can achieve high-precision retrieval for both text and visual documents. In the next lesson, we'll explore the Universal Query API, where you'll learn how to combine multiple retrieval strategies and use ColBERT for reranking - a more common production pattern that balances speed and precision when working with large collections. \ No newline at end of file +With multivectors in your toolkit, you can achieve high-precision retrieval for both text and visual documents. In the next lesson, we'll explore the Universal Query API, where you'll learn how to combine multiple retrieval strategies and use ColBERT for reranking, a more common production pattern that balances speed and precision when working with large collections. diff --git a/qdrant-landing/content/course/essentials/day-5/pitstop-project.md b/qdrant-landing/content/course/essentials/day-5/pitstop-project.md index 01e5513e6..9cca79281 100644 --- a/qdrant-landing/content/course/essentials/day-5/pitstop-project.md +++ b/qdrant-landing/content/course/essentials/day-5/pitstop-project.md @@ -2,6 +2,7 @@ title: "Project: Building a Recommendation System" description: Build a hybrid AI recommendation system with Qdrant’s Universal Query API—combining dense, sparse, and multivector retrieval, ColBERT reranking, and RRF fusion in one atomic query. weight: 5 +isLesson: true --- {{< date >}} Day 5 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-5/universal-query-api.md b/qdrant-landing/content/course/essentials/day-5/universal-query-api.md index 2f17e0662..fa7ff2c9e 100644 --- a/qdrant-landing/content/course/essentials/day-5/universal-query-api.md +++ b/qdrant-landing/content/course/essentials/day-5/universal-query-api.md @@ -2,17 +2,18 @@ title: "The Universal Query API" description: Learn how to run dense, sparse, and ColBERT multivector retrieval with Qdrant’s Universal Query API—fusing, filtering, and reranking results in a single atomic request. weight: 3 +isLesson: true --- {{< date >}} Day 5 {{< /date >}} # The Universal Query API -Picture this: a customer types "leather jackets" into your store's search bar. You want to show items that match the style semantically - so a bomber jacket surfaces even if it doesn't mention "leather jackets" verbatim - but you also need to enforce your business rules. Only products under $200, only items in stock, only jackets released within the past year. Traditionally, you'd fire off a search, gather results, then apply filters and glue code. With Qdrant's [Universal Query API](/documentation/concepts/hybrid-queries/), all of that happens in one declarative request. +Picture this: a customer types "leather jackets" into your store's search bar. You want to show items that match the style semantically - so a bomber jacket surfaces even if it doesn't mention "leather jackets" verbatim - but you also need to enforce your business rules. Only products under $200, only items in stock, only jackets released within the past year. Traditionally, you'd fire off a search, gather results, then apply filters and glue code. With Qdrant's [Universal Query API](/documentation/search/hybrid-queries/), all of that happens in one declarative request. ## Run dense + sparse retrieval in parallel with RRF -First, you retrieve candidates from multiple sources in parallel and fuse their ranks. Below, we blend dense semantics from a BGE model with sparse keyword matching from SPLADE by using [Reciprocal Rank Fusion](/documentation/concepts/hybrid-queries/#hybrid-search) to merge the two lists: +First, you retrieve candidates from multiple sources in parallel and fuse their ranks. Below, we blend dense semantics from a BGE model with sparse keyword matching from SPLADE by using [Reciprocal Rank Fusion](/documentation/search/hybrid-queries/#hybrid-search) to merge the two lists: ```python from qdrant_client import QdrantClient, models diff --git a/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md b/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md index 23619a520..258c62728 100644 --- a/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md +++ b/qdrant-landing/content/course/essentials/day-5/universal-query-demo.md @@ -2,6 +2,7 @@ title: "Demo: Universal Query for Hybrid Retrieval" description: Build a hybrid research discovery system using Qdrant’s Universal Query API—combine dense, sparse, and ColBERT vectors for semantic, keyword, and reranked retrieval in one query. weight: 4 +isLesson: true --- {{< date >}} Day 5 {{< /date >}} @@ -108,7 +109,7 @@ client.create_payload_index( ## Prepare and Ingest Research Paper Data -Now that our collection is configured with vectors and payload indexes, let's take some sample research papers: +Now that our collection is configured with vectors and payload indexes, let's define a few sample research papers: ```python sample_data = [ @@ -133,13 +134,13 @@ sample_data = [ "open_access": True, }, { - "title": "Zero-Shot Retrieval for Scalable Visual Search in a Two-Sided Marketplace", - "authors": ["Andre Rusli", "Shoma Ishimoto", "Sho Akiyama", "Aman Kumar Singh"], - "abstract": "Visual search offers an intuitive way for customers to explore diverse product catalogs, particularly in consumer-to-consumer (C2C) marketplaces where listings are often unstructured and visually driven. This paper presents a scalable visual search system deployed in Mercari's C2C marketplace...", - "research_area": "computer_vision", - "published_date": "2025-07-31", - "impact_score": 0.78, - "citation_count": 12, + "title": "MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings", + "authors": ["Jason Lee", "Vahab Mirrokni", "Rajesh Jayaram"], + "abstract": "We present MUVERA, a retrieval approach that compresses multi-vector representations into fixed-dimensional encodings for efficient first-stage retrieval while preserving the quality benefits of late interaction models. The method reduces serving costs and latency without giving up strong reranking performance...", + "research_area": "machine_learning", + "published_date": "2024-05-29", + "impact_score": 0.84, + "citation_count": 27, "open_access": True, }, ] @@ -188,7 +189,7 @@ client.upload_points( ) ``` -## Step 3: The Universal Query in Action +## Step 2: The Universal Query in Action Let's build a sophisticated research discovery query step by step. We'll orchestrate dense search, sparse search, RRF fusion, and ColBERT reranking - all in a single API call. @@ -327,11 +328,11 @@ for i, hit in enumerate(response.points or [], 1): print(f" Score: {hit.score:.4f}\n") ``` -And there you have it - a sophisticated multi-stage research discovery system in a single declarative query! +And there you have it: a sophisticated multi-stage research discovery system in a single declarative query. ## Real ArXiv Dataset Integration -Here's how you could populate the collection with real data (if the endpoint wasn't broken): +Here's how you could populate the collection with real arXiv data: ```python # ! pip install arxiv @@ -391,10 +392,10 @@ print(f"Uploaded {len(points)} research papers to collection") - **Single Request**: Complex multi-stage research discovery in one API call - **Parallel Execution**: Dense and sparse searches run concurrently - **Smart Filtering**: Apply research quality filters at optimal stages -- **Real Data**: Works with actual arXiv dataset and research metadata +- **Real Data**: Works with actual arXiv data and research metadata - **Production Ready**: Scales to millions of papers with sub-second latency -The Universal Query API eliminates the complexity of building multi-turn retrieval systems. What used to require coordination between semantic search engines, keyword systems, and reranking models now happens in a single, optimized request - perfect for academic search, literature reviews, and research recommendation systems. +The Universal Query API eliminates the complexity of building multi-turn retrieval systems. What used to require coordination between semantic search engines, keyword systems, and reranking models now happens in a single, optimized request, which makes it a good fit for academic search, literature reviews, and research recommendation systems. ## Next diff --git a/qdrant-landing/content/course/essentials/day-6/congratulations.md b/qdrant-landing/content/course/essentials/day-6/congratulations.md index 3dcecedc4..af05ff59f 100644 --- a/qdrant-landing/content/course/essentials/day-6/congratulations.md +++ b/qdrant-landing/content/course/essentials/day-6/congratulations.md @@ -2,6 +2,7 @@ title: "Course Completion and Next Steps" description: Complete your Qdrant course by earning certification and mastering hybrid, multivector, and production-ready vector search techniques—skills to design, evaluate, and deploy real-world AI search systems. weight: 3 +isLesson: true --- {{< date >}} Day 6 {{< /date >}} @@ -47,7 +48,7 @@ Get recognized for completing Day 0–6 and the final project. Add it to your Li ## What's Next? -**Explore Advanced Integrations**: Check out [Day 7 Partner Integrations](../../day-7/) to see how Qdrant works with leading AI frameworks and data platforms. +**Explore Advanced Integrations**: Check out [Day 7 Partner Integrations](/course/essentials/day-7/) to see how Qdrant works with leading AI frameworks and data platforms. **Join the Community**: Share your final project results and connect with other practitioners building vector search systems. The Qdrant community is always excited to see what people build. diff --git a/qdrant-landing/content/course/essentials/day-6/final-project.md b/qdrant-landing/content/course/essentials/day-6/final-project.md index f1f1d58e2..d7ed7f95b 100644 --- a/qdrant-landing/content/course/essentials/day-6/final-project.md +++ b/qdrant-landing/content/course/essentials/day-6/final-project.md @@ -2,6 +2,7 @@ title: "Final Project: Production-Ready Documentation Search Engine" description: Create a complete documentation search system with Qdrant, featuring hybrid retrieval, multivector reranking, and performance evaluation for a portfolio-ready, production-quality vector search application. weight: 2 +isLesson: true --- {{< date >}} Day 6 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/_index.md b/qdrant-landing/content/course/essentials/day-7/_index.md index c4ce13c96..293edf5fc 100644 --- a/qdrant-landing/content/course/essentials/day-7/_index.md +++ b/qdrant-landing/content/course/essentials/day-7/_index.md @@ -27,40 +27,40 @@ Learn about the Qdrant ecosystem and integration strategies. - icon: /courses/course-integrations/haystack.png title: Haystack content: Build end-to-end agentic pipelines with Qdrant - link: haystack/ + link: /course/essentials/day-7/haystack/ - icon: /courses/course-integrations/tensorlake.svg title: Tensorlake content: Build scalable data lakes with vector capabilities - link: tensorlake/ + link: /course/essentials/day-7/tensorlake/ - icon: /courses/course-integrations/llamaindex.svg title: LlamaIndex content: Build agentic workflows for complex enterprise documents - link: llamaindex/ + link: /course/essentials/day-7/llamaindex/ - icon: /courses/course-integrations/unstructured.svg title: Unstructured.io content: Process and vectorize documents from any format - link: unstructured/ + link: /course/essentials/day-7/unstructured/ - icon: /courses/course-integrations/quotient.svg title: Quotient content: Advanced analytics with vector data - link: quotient/ + link: /course/essentials/day-7/quotient/ - icon: /courses/course-integrations/superlinked.svg title: Superlinked content: Advanced feature engineering for vectors - link: superlinked/ + link: /course/essentials/day-7/superlinked/ - icon: /courses/course-integrations/camel-ai.svg title: Camel AI content: Agentic RAG with multi-agent systems - link: camel/ + link: /course/essentials/day-7/camel/ - icon: /courses/course-integrations/jina.svg title: Jina AI content: Advanced multimodal embeddings with Qdrant - link: jina/ + link: /course/essentials/day-7/jina/ {{< /cards-list >}} diff --git a/qdrant-landing/content/course/essentials/day-7/camel.md b/qdrant-landing/content/course/essentials/day-7/camel.md index 80fc9fc3f..92fa0a9f7 100644 --- a/qdrant-landing/content/course/essentials/day-7/camel.md +++ b/qdrant-landing/content/course/essentials/day-7/camel.md @@ -2,6 +2,7 @@ title: "Integrating with Camel AI" description: Learn how Camel AI and Qdrant enable automated RAG pipelines with multi-agent communication, vector-based memory, and seamless integration into live environments like Discord bots. weight: 8 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/haystack.md b/qdrant-landing/content/course/essentials/day-7/haystack.md index b484beb7c..a5c201b30 100644 --- a/qdrant-landing/content/course/essentials/day-7/haystack.md +++ b/qdrant-landing/content/course/essentials/day-7/haystack.md @@ -2,6 +2,7 @@ title: "Integrating with Haystack" description: Learn how Qdrant and Haystack combine to deliver end-to-end search and recommendation systems with hybrid retrieval, semantic filtering, and agentic AI orchestration. weight: 2 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/jina.md b/qdrant-landing/content/course/essentials/day-7/jina.md index 046f3062d..f0ac57fa0 100644 --- a/qdrant-landing/content/course/essentials/day-7/jina.md +++ b/qdrant-landing/content/course/essentials/day-7/jina.md @@ -2,6 +2,7 @@ title: "Integrating with Jina AI" description: Learn how Jina AI’s Embeddings v4 and Qdrant enable advanced multimodal retrieval, supporting text-to-image, image-to-text, and hybrid search with high-performance vector storage. weight: 9 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/llamaindex.md b/qdrant-landing/content/course/essentials/day-7/llamaindex.md index 62d445145..6bacd0521 100644 --- a/qdrant-landing/content/course/essentials/day-7/llamaindex.md +++ b/qdrant-landing/content/course/essentials/day-7/llamaindex.md @@ -2,6 +2,7 @@ title: "Integrating with LlamaIndex" description: Learn how LlamaIndex and Qdrant power intelligent RAG pipelines, function-calling agents, and cloud-synced vector search systems with structured workflows and dynamic query handling. weight: 6 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/quotient.md b/qdrant-landing/content/course/essentials/day-7/quotient.md index de4083e4b..adda538a1 100644 --- a/qdrant-landing/content/course/essentials/day-7/quotient.md +++ b/qdrant-landing/content/course/essentials/day-7/quotient.md @@ -2,6 +2,7 @@ title: "Integrating with Quotient" description: Learn how Quotient and Qdrant combine to deliver end-to-end AI monitoring, hallucination detection, and performance analytics for reliable, high-quality retrieval-augmented generation systems. weight: 7 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/superlinked.md b/qdrant-landing/content/course/essentials/day-7/superlinked.md index 32c23a060..56e569fbd 100644 --- a/qdrant-landing/content/course/essentials/day-7/superlinked.md +++ b/qdrant-landing/content/course/essentials/day-7/superlinked.md @@ -2,6 +2,7 @@ title: "Integrating with Superlinked" description: Learn how Superlinked’s Mixture of Encoders and Qdrant enable rich, multi-modal embeddings that fuse semantic, numerical, and temporal data for optimized vector retrieval. weight: 5 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/tensorlake.md b/qdrant-landing/content/course/essentials/day-7/tensorlake.md index 280a6479a..392d51b3b 100644 --- a/qdrant-landing/content/course/essentials/day-7/tensorlake.md +++ b/qdrant-landing/content/course/essentials/day-7/tensorlake.md @@ -2,6 +2,7 @@ title: "Integrating with Tensorlake" description: Learn how TensorLake and Qdrant combine document parsing, knowledge graphs, and vector search to build scalable, structured data lakes for advanced RAG and research discovery applications. weight: 4 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} diff --git a/qdrant-landing/content/course/essentials/day-7/unstructured.md b/qdrant-landing/content/course/essentials/day-7/unstructured.md index e3beadd4d..47e2edbd0 100644 --- a/qdrant-landing/content/course/essentials/day-7/unstructured.md +++ b/qdrant-landing/content/course/essentials/day-7/unstructured.md @@ -2,6 +2,7 @@ title: "Integrating with Unstructured.io" description: Learn how Unstructured.io and Qdrant transform unstructured enterprise data into structured embeddings through VLM document understanding, smart chunking, and secure, production-ready ETL pipelines. weight: 3 +isLesson: true --- {{< date >}} Day 7 {{< /date >}} @@ -93,7 +94,7 @@ This architecture enables various enterprise use cases: - [Unstructured Qdrant Destination](https://docs.unstructured.io/ui/destinations/qdrant): Official Unstructured documentation for sending processed data to Qdrant. Learn about Qdrant Cloud integration, collection setup, and workflow configuration. -- [Qdrant & Unstructured Integration Guide](/documentation/frameworks/unstructured/): +- [Qdrant & Unstructured Integration Guide](/documentation/data-management/unstructured/): Official Qdrant documentation for Unstructured.io integration, covering setup and best practices for document processing pipelines. ⭐ **Show your support!** Give Unstructured a star on their GitHub repository: [github.com/Unstructured-IO/unstructured](https://github.com/Unstructured-IO/unstructured) diff --git a/qdrant-landing/content/course/essentials/faq/_index.md b/qdrant-landing/content/course/essentials/faq/_index.md index a0a2ba90b..51b77b7ad 100644 --- a/qdrant-landing/content/course/essentials/faq/_index.md +++ b/qdrant-landing/content/course/essentials/faq/_index.md @@ -45,7 +45,7 @@ Check each Day’s page of content. {{< course-card title="Why Start Today" image="/icons/outline/rocket-white-light.svg" -link="/course/day-0/">}} +link="/course/essentials/day-0/">}} - Seeing practical examples (e.g., hybrid search, sparse+dense vectors) - Learning key deployment tactics (multi-node clusters, on-disk indexing, RBAC) diff --git a/qdrant-landing/content/course/multi-vector-search/_index.md b/qdrant-landing/content/course/multi-vector-search/_index.md new file mode 100644 index 000000000..408803c38 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/_index.md @@ -0,0 +1,166 @@ +--- +title: "Multi-Vector Search Course" +page_title: "Qdrant Multi-Vector Search Course" +description: Master late interaction models, ColPali, and production optimization. Build scalable multi-vector search pipelines. +content: + sidebarTitle: "Multi-Vector Search Course" + menuTitle: + text: Course Overview + url: /course/multi-vector-search/ + nextButton: Continue to Next Step + nextDay: Complete + title: "Multi-Vector Search" + description: Master late interaction models, ColPali, and production optimization. Build scalable multi-vector search pipelines. +partition: course +isLesson: true +--- + +# Multi-Vector Search + +**Build production-ready multi-vector search pipelines** + +Go beyond single-vector embeddings with late interaction models like ColBERT and ColPali. Learn the MaxSim distance metric, optimize for billion-scale search, and evaluate your retrieval pipelines with industry-standard metrics. + +
+ +
+ +
+ +{{< cards-list >}} +- icon: /icons/outline/play-white.svg + title: 4 modules + content: Focused lessons building from fundamentals to production +- icon: /icons/outline/cloud-check-blue.svg + title: Shareable certificate + content: Earn a digital certificate upon completion +- icon: /icons/outline/time-blue.svg + title: Flexible schedule + content: Learn at your own pace (1–2 hours/module) +- icon: /icons/outline/plan.svg + title: Advanced level + content: Assumes familiarity with vector search basics + +{{< /cards-list >}} + +
+ +## What you'll learn +{{< course-card + title="Skills you'll gain:" + image="/icons/outline/training-white.svg" + type="wide-list">}} + +- Late interaction paradigm and MaxSim distance metric +- ColBERT for text and ColPali for visual documents +- Multi-stage retrieval with prefetch and reranking +- Quantization and pooling techniques for memory optimization +- MUVERA indexing for billion-scale search +- Evaluation metrics: Recall@k, NDCG, MRR + +{{< /course-card >}} + +### The Path + +**Module 0**: Setup. Configure Qdrant Cloud or local instance and install Python dependencies. + +**Module 1**: Text multi-vectors. Understand the late interaction paradigm, learn the MaxSim distance metric, explore use cases and challenges, and implement ColBERT with Qdrant. + +**Module 2**: Multi-modal search. Apply multi-vector representations to images and PDFs with ColPali. Explore model variants and leverage visual interpretability for debugging. + +**Module 3**: Optimization and evaluation. Master quantization, pooling, and MUVERA for memory-efficient search. Build multi-stage retrieval pipelines and evaluate with standard metrics. + +## How the course works + +{{< cards-list >}} + +- icon: /icons/outline/training-purple.svg + title: Video-first lessons + content: Clear, concise modules by the Qdrant team +- icon: /icons/outline/hacker-purple.svg + title: Final project + content: Build a production-ready multi-modal search system +- icon: /icons/outline/similarity-blue.svg + title: Hands-on notebooks + content: Practice each concept with Colab notebooks +- icon: /icons/outline/copy.svg + title: Progressive learning + content: Build from fundamentals to advanced optimization + {{< /cards-list >}} + +
+ +## Syllabus + +{{< accordion >}} +- title: "Module 0: Setting Up Dependencies" + content: | + - Qdrant Setup + - Installing Dependencies +
+
+

→ Start Module 0

+ +- title: "Module 1: Multi-Vector Representations for Textual Data" + content: | + - Late Interaction Basics + - MaxSim Distance Metric + - Use Cases for Multi-Vector Search + - Problems of Multi-Vector Search + - Multi-Vector Embeddings in Qdrant +
+
+

→ Start Module 1

+ +- title: "Module 2: Multi-Vector Representations for Multi-Modal Data" + content: | + - How ColPali Models Work + - ColPali Family Overview + - Visual Interpretability of ColPali +
+
+

→ Start Module 2

+ +- title: "Module 3: Scalability and Optimization" + content: | + - Multi-Stage Retrieval with Universal Query API + - Vector Quantization Techniques + - Pooling Techniques + - MUVERA Indexing + - Evaluating Search Pipelines + - Final Project +
+
+

→ Start Module 3

+{{< /accordion >}} + + +## Who it's for + +ML, backend, and search engineers who want to go beyond single-vector embeddings. Requires intermediate Python, basic familiarity with vector search concepts (embeddings, similarity metrics), and comfort with APIs. + +## Time commitment + +- Duration: 4 modules at 1 hour/module +- Video learning: 1.5 hours +- Hands-on notebooks: 1.5 hours +- Final project: 1-3 hours +- Total: 4-6 hours + + +{{< course-card + title="Ready to master multi-vector search?" + image="/icons/outline/rocket-white-light.svg" + link="/course/multi-vector-search/module-0/">}} +**What you'll get** +- Build production-ready multi-vector pipelines +- Practice with real Colab notebooks +- Learn optimization techniques for scale +- Portfolio project and community support +{{< /course-card >}} diff --git a/qdrant-landing/content/course/multi-vector-search/certification.md b/qdrant-landing/content/course/multi-vector-search/certification.md new file mode 100644 index 000000000..15ba4eb3e --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/certification.md @@ -0,0 +1,26 @@ +--- +title: "Qdrant Multi-Vector Certification" +description: "Get officially certified in multi-vector search by Qdrant." +url: /course/multi-vector-search/certification/ +isLesson: true +weight: 50 +--- + +# Qdrant Multi-Vector Search Certification + +Congratulations! You’ve completed the **Multi-Vector Search course**. You didn’t just learn how to store vectors; you learned how to build high-performance retrieval systems using late interaction models and multi-vector representations. + +You’ve moved past single-vector embeddings and dove deep into ColBERT, ColPali, MaxSim scoring, MUVERA, and production-grade multi-vector pipelines. That effort deserves more than just a “finished” status. It deserves professional recognition! + +## Get #QdrantCertified + +Your expertise is now production-ready. It’s time to validate those skills with our official certification. + +Head over to [train.qdrant.dev](https://train.qdrant.dev) to take the exam. + +Passing this exam proves you aren’t just a user; you are a Search Engineer capable of: + +- Designing multi-vector retrieval pipelines with late interaction models. +- Applying ColPali and its variants for visual document search. +- Optimizing multi-vector search for both memory and latency. +- Mastering MaxSim scoring and the nuances of multi-vector architectures in Qdrant. \ No newline at end of file diff --git a/qdrant-landing/content/course/multi-vector-search/module-0/_index.md b/qdrant-landing/content/course/multi-vector-search/module-0/_index.md new file mode 100644 index 000000000..f1d84afd0 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-0/_index.md @@ -0,0 +1,21 @@ +--- +title: "Module 0: Setting Up Dependencies" +description: "Set up your development environment for multi-vector search. Install required dependencies and prepare your workspace for the course." +isLesson: true +weight: 10 +--- + +{{< date >}} Module 0 {{< /date >}} + +# Setting Up Dependencies + +Get your environment ready for exploring multi-vector search with Qdrant. + +--- + +## Today's path + +1. Qdrant Setup +2. Installing Dependencies + +By the end, you'll have a working development environment ready for multi-vector search experiments. diff --git a/qdrant-landing/content/course/multi-vector-search/module-0/installing-dependencies.md b/qdrant-landing/content/course/multi-vector-search/module-0/installing-dependencies.md new file mode 100644 index 000000000..9263c0a0e --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-0/installing-dependencies.md @@ -0,0 +1,132 @@ +--- +title: "Installing Dependencies" +description: Install Python dependencies including FastEmbed and Qdrant client. +weight: 2 +isLesson: true +--- + +{{< date >}} Module 0 {{< /date >}} + +# Installing Dependencies + +To work with multi-vector search in Qdrant, you'll need several Python libraries: Qdrant client for search and FastEmbed for multi-vector embeddings. + +We'll set up a clean Python environment and install everything you need to start experimenting with multi-vector representations. + +## Python Environment Setup + +### Using uv (Recommended) + +For this course, we recommend using [uv](https://docs.astral.sh/uv/), a modern Python package manager that's significantly faster and more reliable than traditional pip. It handles virtual environments and dependencies with better performance and dependency resolution. + +**Install uv:** + +On macOS and Linux: +```bash +curl -LsSf https://astral.sh/uv/install.sh | sh +``` + +On Windows: +```bash +powershell -c "irm https://astral.sh/uv/install.ps1 | iex" +``` + +**Create a new virtual environment:** + +```bash +uv venv +source .venv/bin/activate # On macOS/Linux +# or +.venv\Scripts\activate # On Windows +``` + +### Alternative: Using Poetry + +If you prefer Poetry for dependency management, it offers robust project management with automatic virtual environment handling and dependency lock files. + +**Install Poetry:** + +On macOS and Linux: +```bash +curl -sSL https://install.python-poetry.org | python3 - +``` + +On Windows (PowerShell): +```bash +(Invoke-WebRequest -Uri https://install.python-poetry.org -UseBasicParsing).Content | py - +``` + +**Create a new project or add dependencies:** + +```bash +# Initialize a new Poetry project +poetry init + +# Activate the virtual environment +poetry shell +``` + +**Python Version Requirements:** +You'll need Python 3.10 or higher, as required by the qdrant-client library. + +## Installing Dependencies + +With your virtual environment activated, install the required libraries: + +**Using uv (recommended):** + +```bash +uv pip install "qdrant-client>=1.16.2" "fastembed>=0.8.0" +``` + +**Using Poetry:** + +```bash +poetry add "qdrant-client>=1.16.2" "fastembed>=0.8.0" +``` + +**Using pip:** + +```bash +pip install "qdrant-client>=1.16.2" "fastembed>=0.8.0" +``` + +### What These Libraries Do + +- **qdrant-client**: The official Python client for Qdrant, providing both synchronous and asynchronous APIs for vector search operations. This library contains full type definitions and supports all Qdrant features. + +- **fastembed**: A fast, lightweight library for generating embeddings, maintained by the Qdrant team. It includes support for multi-vector embeddings which we'll use extensively in this course. **(Note: fastembed=0.7.5 or above required for this course)** + +## Verification Steps + +Let's verify that everything is installed correctly. + +**Test your imports:** + +```python +from qdrant_client import QdrantClient +from fastembed import TextEmbedding + +print("All dependencies installed successfully!") +``` + +**Quick connection test:** + +If you set up Qdrant in the previous lesson, verify you can connect: + +```python +# For Qdrant Cloud +client = QdrantClient( + url="https://your-cluster-url.cloud.qdrant.io", + api_key="your-api-key" +) + +# For local Qdrant +# client = QdrantClient(url="http://localhost:6333") + +print(f"Connected to Qdrant: {client.get_collections()}") +``` + +## Next Steps + +With your Python environment configured and dependencies installed, you're ready to dive into Module 1, where we'll explore the fundamentals of multi-vector search and understand how it differs from traditional single-vector approaches. diff --git a/qdrant-landing/content/course/multi-vector-search/module-0/qdrant-setup.md b/qdrant-landing/content/course/multi-vector-search/module-0/qdrant-setup.md new file mode 100644 index 000000000..f21ddda67 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-0/qdrant-setup.md @@ -0,0 +1,95 @@ +--- +title: "Qdrant Setup" +description: Set up Qdrant for multi-vector search. Learn how to create a collection and configure it for multi-vector embeddings. +weight: 1 +isLesson: true +--- + +{{< date >}} Module 0 {{< /date >}} + +# Qdrant Setup + +Before diving into multi-vector search, you need a running Qdrant instance. Whether you choose Qdrant Cloud for a managed solution or a local deployment, this lesson will get you up and running. + +Multi-vector search requires specific collection configurations that differ from traditional single-vector setups. We'll cover the essentials to prepare your environment. + +--- + +## Qdrant Cloud Setup (Recommended) + +Qdrant Cloud is the fastest way to get started with multi-vector search. It provides a fully managed, production-ready vector database with automatic backups, high availability, and secure TLS connections. Both Qdrant Cloud and the open-source version provide the same feature set - Cloud simply handles the infrastructure for you. + +### Create Your Cluster + +1. Sign up at [cloud.qdrant.io](https://cloud.qdrant.io/signup) using your email, Google, or GitHub account. + +2. Navigate to **Clusters** -> **Create a Free Cluster**. The Free Tier provides sufficient resources for this course. + + ![Create cluster](/docs/gettingstarted/gui-quickstart/create-cluster.png) + +3. Select a region closest to your location or application. + +4. Once your cluster is ready, copy the API key from the cluster dashboard and store it securely. You can generate additional keys later from the **API Keys** section. + + ![Get API key](/docs/gettingstarted/gui-quickstart/api-key.png) + +### Access the Web UI + +Click **Cluster UI** in the top-right corner of your cluster page to open the dashboard. + +![Access dashboard](/docs/gettingstarted/gui-quickstart/access-dashboard.png) + +The Web UI provides several useful tools: +- **Console**: Test REST API calls directly in your browser +- **Collections**: Manage all your collections and their configurations +- **Tutorial**: Interactive walkthrough with sample data + +### Save Your Credentials + +Store your cluster URL and API key for use in upcoming lessons. Create an `.env` file in your working directory: + +```env +QDRANT_URL=https://YOUR-CLUSTER.cloud.qdrant.io:6333 +QDRANT_API_KEY=YOUR_API_KEY +``` + +Replace `YOUR-CLUSTER` with your actual cluster URL from the dashboard, and `YOUR_API_KEY` with the API key you copied earlier. + +You'll use these credentials in the next lesson when we install and configure the Python client. + +## Local Qdrant Installation + +Qdrant's open-source version provides the same features as Qdrant Cloud but requires you to manage the infrastructure yourself. This option works well for development, testing, or when you need full control over your deployment. + +### Docker Installation (Recommended) + +The fastest way to run Qdrant locally is with Docker: + +```bash +docker run -p 6333:6333 -p 6334:6334 \ + -v $(pwd)/qdrant_storage:/qdrant/storage:z \ + qdrant/qdrant +``` + +This command: +- Exposes port `6333` for the REST API +- Exposes port `6334` for the gRPC API +- Mounts a local directory for persistent storage + +Once running, you can access the Web UI at `http://localhost:6333/dashboard` to verify the installation. + +### Alternative Installation Methods + +For production deployments or other installation methods, see the [Qdrant Installation Guide](/documentation/operations/installation/). + +## Verifying Your Setup + +Open the Qdrant Web UI: +- **Cloud users**: Click **Cluster UI** in the top-right corner of your cluster dashboard +- **Local users**: Navigate to `http://localhost:6333/dashboard` + +If the Web UI loads and you can see the **Collections** tab, your setup is complete. In the next lesson, you'll install the Python dependencies to connect programmatically. + +--- + +Next, you'll install the Python dependencies needed to work with multi-vector embeddings. diff --git a/qdrant-landing/content/course/multi-vector-search/module-1/_index.md b/qdrant-landing/content/course/multi-vector-search/module-1/_index.md new file mode 100644 index 000000000..8c697edf6 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-1/_index.md @@ -0,0 +1,25 @@ +--- +title: "Module 1: Multi-Vector Representations for Textual Data" +description: "Learn about multi-vector representations for text with ColBERT. Understand how they differ from single vector embeddings and when to use them." +isLesson: true +weight: 20 +--- + +{{< date >}} Module 1 {{< /date >}} + +# Multi-Vector Representations for Textual Data + +Dive into multi-vector text representations and discover how ColBERT changes the vector search landscape. + +--- + +## Today's path + +1. Late Interaction Basics +2. MaxSim Distance Metric +3. Use Cases for Multi-Vector Search +4. Problems of Multi-Vector Search +5. Multi-Vector Embeddings in Qdrant + +You'll understand when multi-vector representations outperform traditional single-vector embeddings, and what kind of +problems to expect when you start working with multi-vector search at scale. diff --git a/qdrant-landing/content/course/multi-vector-search/module-1/late-interaction-basics.md b/qdrant-landing/content/course/multi-vector-search/module-1/late-interaction-basics.md new file mode 100644 index 000000000..0f6a7177a --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-1/late-interaction-basics.md @@ -0,0 +1,222 @@ +--- +title: "Late Interaction Basics" +description: Understand the late interaction paradigm and how it differs from traditional dense embeddings for text search. +weight: 1 +isLesson: true +--- + +{{< date >}} Module 1 {{< /date >}} + +# Late Interaction Basics + +When building a search system, one fundamental question emerges: **when should a query and document interact?** The answer to this question may affect both the quality of search results and the system's scalability. + +This lesson introduces the late interaction paradigm - the foundation of multi-vector search - and explores how it compares to other approaches. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## Understanding the Alternatives + +Before diving into late interaction, let's establish what we mean by "interaction." + +In search systems, **interaction** refers to when and how the query and document representations influence each other. Do they interact during encoding, or only during comparison? This timing fundamentally shapes the system's architecture. + +We can categorize approaches based on when this interaction occurs: + +- **No interaction:** Query and document are encoded independently into fixed representations, then compared. They never "see" each other during encoding. +- **Early interaction:** Query and document are encoded together, with each word attending to the other during the encoding process. Maximum interaction, but no pre-computation. +- **Late interaction:** Query and document are encoded independently (like no interaction), but we preserve fine-grained representations that interact during scoring (late in the process). + +![Side-by-side comparison showing no interaction vs early interaction vs late interaction](/courses/multi-vector-search/module-1/interaction-comparison.png) + +Let's examine each paradigm to understand the trade-offs. A simple example will illustrate the differences between all the methods. + +```python +# Example documents and query we'll use throughout this lesson +documents = [ + "Qdrant is an AI-native vector database and a semantic search engine", + "Relational databases are not well-suited for search", +] +query = "What is Qdrant?" +``` + +### Single-Vector Embeddings (No Interaction) + +The most common approach encodes each document and query into a single dense vector, then compares them using similarity (most often cosine similarity). + +**The strength:** This method is simple, fast, and scales well. Document vectors can be pre-computed and stored, making search efficient even across billions of documents. + +**The limitation:** Compressing an entire document into a single vector means losing fine-grained details. Think of it like summarizing a book in one sentence - you capture the general theme but miss the nuances that might be relevant to a specific query. + +Let's load a dense embedding model and generate vector representations for our documents. + +```python +from fastembed import TextEmbedding + +# Load the BAAI/bge-small-en-v1.5 model +dense_model = TextEmbedding("BAAI/bge-small-en-v1.5") +# Pass the documents through the model. The .passage_embed +# method returns a generator we can iterate over and is +# supposed to be used for the documents only. +dense_generator = dense_model.passage_embed(documents) +# Running next on the generator yields one vector at +# the time, representing a single document. +dense_vector = next(dense_generator) +``` + +We also need to generate a vector for the query using the same model. + +```python +# Generate a dense vector for the query as well, using +# the .query_embed method this time. +dense_query_vector = next(dense_model.query_embed(query)) +``` + +Now we can calculate the similarity between the query and each document using the dot product. + +```python +import numpy as np + +# Calculate the dot product between the query +# and the first document vector +np.dot(dense_query_vector, dense_vector) +``` + +Let's calculate the similarity with the second document as well. + +```python +# Calculate the dot product between the same query +# and the second document vectors +np.dot(dense_query_vector, next(dense_generator)) +``` + +Notice how each document and query produces exactly **one vector** of 384 dimensions. To achieve this compression, the model internally generates embeddings for each token, then uses **pooling** (typically mean pooling or a special [CLS] token) to aggregate them into a single fixed-size representation. + +![Single vector compresses all details, multi-vector maintains token-level granularity](/courses/multi-vector-search/module-1/single-vector-process.png) + +### Cross-Encoders (Early Interaction) + +At the other extreme, cross-encoders process the query and document together through a neural network, producing a relevance score. This is "early interaction" because the query and document interact during the encoding phase itself. + +**The strength:** This approach achieves deep contextual understanding. Every query word can "attend to" every document word during encoding, enabling precise relevance judgments. + +**The limitation:** You must process every query-document pair from scratch. For a collection of a million documents, that means a million forward passes through a neural network for each query - prohibitively expensive for initial retrieval. + +Cross-encoders excel at re-ranking a small candidate set but don't scale for searching large collections. + +Let's see how cross-encoders work differently by processing the query and documents together. + +```python +from fastembed.rerank.cross_encoder import TextCrossEncoder + +# Load the Xenova/ms-marco-MiniLM-L-6-v2 cross encoder model +cross_encoder = TextCrossEncoder("Xenova/ms-marco-MiniLM-L-6-v2") +# Run .rerank method on the query and all the documents. +# It does not create any vector representations, but gives +# the score indicating the relevance of the document for +# the provided query. +cross_encoder.rerank(query, documents) +``` + +The key difference: cross-encoders **cannot pre-compute** document representations. Each query requires fresh computation for every candidate, making them impractical for initial retrieval over large collections. + +### The Gap + +We need an approach that combines the best of both worlds: the scalability of pre-computed single-vector representations and the fine-grained matching capability of cross-encoders. + +## Late Interaction: The Core Paradigm + +Late interaction solves this challenge through a simple but powerful idea: **encode documents and queries into multiple token-level vectors, then defer the comparison until search time.** + +### How It Works + +Instead of compressing a document into a single vector, late interaction represents it as a collection of contextualized token embeddings + +1. **Encode:** Pass the document through an encoder (like ColBERT) to generate one embedding vector per token +2. **Store multi-vector representations:** Keep all token vectors instead of aggregating them into a single vector +3. **Defer comparison:** At search time, compare query token vectors against document token vectors +4. **Late interaction:** The actual "interaction" between query and document happens late - only when computing relevance scores + +![Indexing and search pipelines](/courses/multi-vector-search/module-1/indexing-and-search.png) + +This isn't that much fundamentally different from both single-vector and cross-encoder approaches, yet there are some differences: +- Unlike single-vector: We preserve fine-grained, token-level information +- Unlike cross-encoders: We encode documents independently, enabling pre-computation + +### Key Benefits + +**Pre-computation:** Document embeddings can be computed once and stored, just like single-vector approaches. You don't need to re-encode documents for every query. + +**Fine-grained matching:** Different query terms can match different parts of the document. A query about "apple computer" can distinguish contextual meaning from "apple fruit" based on which document tokens match strongly. + +**Contextual understanding:** Token embeddings are contextualized by the surrounding text. The word "bank" has different embeddings in "river bank" versus "financial bank." + +![River vs financial bank](/courses/multi-vector-search/module-1/river-vs-financial-bank.png) + +**Scalability:** While requiring more storage than single vectors, the deferred comparison enables searching large collections efficiently. + +### The ColBERT Approach + +The canonical implementation of late interaction is **ColBERT** (Contextualized Late Interaction over BERT), developed at Stanford. ColBERT popularized this paradigm and demonstrated that you can achieve cross-encoder-level effectiveness with single-vector-level efficiency. + +The core innovation: maintaining bags of contextualized embeddings and delaying the interaction computation until the final stage. + +Let's load a ColBERT model and generate multi-vector representations for our documents. + +```python +from fastembed import LateInteractionTextEmbedding + +# Load the colbert-ir/colbertv2.0 model +colbert_model = LateInteractionTextEmbedding("colbert-ir/colbertv2.0") +# Run .passage_embed on all the documents and create +# a generator of the multi-vector representations +colbert_generator = colbert_model.passage_embed(documents) +colbert_vector = next(colbert_generator) +``` + +Similarly, we create a multi-vector representation for the query. + +```python +# Create multi-vector representation for the query +colbert_query_vector = next(colbert_model.query_embed(query)) +``` + +**Key observation:** Unlike single-vector search, each document is represented by **multiple vectors**. At search time, we compare each query token against all document tokens to compute a relevance score. + +## Why This Matters for Multi-Vector Search + +Late interaction isn't just a technical optimization - it represents a fundamental shift in how we think about semantic search. + +**Captures semantic nuance:** Because we maintain multiple vectors per document, the system can capture complex, multi-faceted content. A document about "Python programming for data science" can match queries about programming languages, data analysis, and scientific computing - each matching different token sets. + +**Enables scale:** Pre-computed multi-vector representations mean you can build practical search systems over large document collections. The computational cost grows with collection size, not quadratically with query-document pairs. + +**Foundation for this course:** Everything we'll explore in subsequent lessons builds on this paradigm - from the distance metrics that enable multi-vector comparison to multi-modal extensions like ColPali to optimization techniques for production deployment. + +**Beyond text:** The late interaction paradigm extends naturally to other modalities. Module 2 explores how ColPali applies these same principles to visual documents, enabling semantic search over images and PDFs. + +## What's Next + +Understanding the conceptual foundation of late interaction is the first step. But how exactly do we compare sets of query vectors against sets of document vectors? + +In the next lesson, you'll learn about **MaxSim** - the distance metric that powers late interaction search. MaxSim defines the specific mathematical operation for computing similarity between multi-vector representations. + +From there, we'll explore use cases where multi-vector search excels, challenges you'll face in production, and how to implement these techniques in Qdrant. diff --git a/qdrant-landing/content/course/multi-vector-search/module-1/maxsim-distance.md b/qdrant-landing/content/course/multi-vector-search/module-1/maxsim-distance.md new file mode 100644 index 000000000..9d19db9ea --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-1/maxsim-distance.md @@ -0,0 +1,194 @@ +--- +title: "MaxSim Distance Metric" +description: Learn about the MaxSim distance metric used in multi-vector search and how it computes similarity between multi-vector representations. +weight: 2 +isLesson: true +--- + +{{< date >}} Module 1 {{< /date >}} + +# MaxSim Distance Metric + +MaxSim (Maximum Similarity) is the core distance metric for late interaction models. Unlike traditional vector similarity metrics that operate on pairs of single vectors, MaxSim computes similarity between sequences of vectors. + +Understanding MaxSim is important for working with multi-vector search effectively and understanding its performance characteristics. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## The MaxSim Formula + +In late interaction, we represent documents and queries as sequences of token vectors. But how do we measure similarity between two sets of vectors? + +The answer is **MaxSim** (Maximum Similarity), defined mathematically as: + +$$ +\text{MaxSim}(Q, D) = \sum_{i=1}^{|Q|} \max_{j=1}^{|D|} \text{sim}(q_i, d_j) +$$ + +Where: +- $Q$ represents the query token vectors +- $D$ represents the document token vectors +- $\text{sim}(q_i, d_j)$ is a base similarity function that measures the distance between two vectors + +The $\text{sim}(q_i, d_j)$ function can be any distance metric - dot product, cosine similarity, Euclidean distance, or others. Let's break down what this formula actually computes. + +## Understanding MaxSim Step-by-Step + +### The Computation Process + +Let's make this concrete with an example. Consider: + +- **Query**: "apple computer" +- **Document**: "Apple makes the MacBook laptop" + +MaxSim formula will choose the strongest connection from the query tokens to document tokens and sum the strengths of each connection for all the query tokens. + +![MaxSim(Q, D)](/courses/multi-vector-search/module-1/maxsim-q-d.png) + +The key insight here: **each query token finds its best match in the document**. This enables **fine-grained semantic matching** - instead of comparing holistic document representations, we're allowing each part of the query to independently find its most relevant counterpart. + +Let's walk through this with a concrete example: + +```python +query = "apple computer" +document = "Apple makes the MacBook laptop" +``` + +Next, we load the ColBERT model to generate multi-vector representations for both the query and document: + +```python +from fastembed import LateInteractionTextEmbedding + +# Load the colbert-ir/colbertv2.0 model +colbert_model = LateInteractionTextEmbedding("colbert-ir/colbertv2.0") + +# Create multi-vector representations of the query and document +query_vector = next(colbert_model.query_embed(query)) +document_vector = next(colbert_model.passage_embed(document)) +``` + +### Understanding Tokenization + +Before we compute MaxSim, let's understand how ColBERT tokenizes our query and document. This tokenization is crucial because each token gets its own embedding vector. + +```python +query_tokenization = colbert_model.model.tokenize([query])[0] +query_tokenization.tokens +``` + +This shows how the query tokenizes into individual tokens including special tokens like `[CLS]` and `[SEP]`. + +```python +document_tokenization = colbert_model.model.tokenize([document])[0] +document_tokenization.tokens +``` + +Similarly, the document is tokenized. Notice that words like 'MacBook' may be split into subword tokens using WordPiece tokenization (e.g., 'mac' and '##book'). Each token will get its own embedding vector. + +Now let's compute MaxSim step-by-step. For each query token, we'll find its maximum similarity score across all document tokens, then sum these maxima: + +```python +import numpy as np + +similarity = 0.0 +for qt, qt_vector in zip(query_tokenization.tokens, + query_vector): + max_idx, max_sim = 0, np.dot(qt_vector, document_vector[0]) + for i, dt_vector in enumerate(document_vector[1:], start=1): + distance = np.dot(qt_vector, dt_vector) + if distance > max_sim: + max_idx, max_sim = i, distance + + print(qt, max_idx, max_sim) + similarity += max_sim +``` + +The code iterates through each query token, computes its dot product similarity with every document token, and keeps track of the maximum. The `print` statement shows which document position each query token matched best with and the similarity score. Each query token independently seeks its strongest counterpart in the document. + +```python +print("MaxSim(Q, D) =", similarity) +``` + +The final MaxSim(Q, D) score is the sum of all these maximum similarities. + +### The Intuition Behind MaxSim + +MaxSim implements **token-level relevance matching**. Each query term actively seeks its most relevant counterpart in the document. Unlike single-vector search where you compare one query vector to one document vector (one-to-one), MaxSim performs many-to-many matching. This enables **contextual precision**. When you search for "apple computer", the word "apple" in your query will match strongly with "Apple" (the company name) in documents about technology, not "apple" (the fruit) in documents about nutrition, as contextualized token embeddings should capture these semantic distinctions. + +But why maximum and not average? Consider a query about "Python programming". A document might discuss Python extensively in one section but also mention cooking recipes in another. MaxSim focuses on the **strong matches** (Python-related tokens) rather than diluting the score with irrelevant tokens. This design choice reflects a fundamental insight: relevance is often concentrated, not uniform. + +## The HNSW Challenge + +There's a significant challenge related to MaxSim when it comes to indexing these representations efficiently. HNSW (Hierarchical Navigable Small World) graphs enable fast approximate nearest neighbor search by building static proximity graphs. They rely on the assumption that distance functions must be symmetric and query-independent. This allows HNSW to construct fixed neighbor relationships that makes the effective graph traversal possible. + +**MaxSim breaks this assumption by design.** Looking back at the formula, notice that $Q$ and $D$ play fundamentally different roles: + +- We **iterate over query tokens** - each query token contributes to the final sum +- Documents are **what we search within** - we take the maximum similarity from document tokens for each query token + +This non-symmetrical structure means `MaxSim(Q, D) ≠ MaxSim(D, Q)`. When you swap the parameters, you change which tokens contribute to the sum. + +![MaxSim(D, Q)](/courses/multi-vector-search/module-1/maxsim-d-q.png) + +To illustrate this asymmetry concretely, let's compute MaxSim in reverse - iterating over document tokens instead of query tokens: + +```python +similarity = 0.0 +for dt, dt_vector in zip(document_tokenization.tokens, + document_vector): + max_idx, max_sim = 0, np.dot(dt_vector, query_vector[0]) + for i, qt_vector in enumerate(query_vector[1:], start=1): + distance = np.dot(dt_vector, qt_vector) + if distance > max_sim: + max_idx, max_sim = i, distance + + print(dt, max_idx, max_sim) + similarity += max_sim +``` + +Now the iteration is different: we loop over document tokens instead of query tokens. Since the document has more tokens than the query, we're summing over more terms, which fundamentally changes the computation. + +```python +print("MaxSim(D, Q) =", similarity) +``` + +The `MaxSim(D, Q)` score will be different from `MaxSim(Q, D)` because we iterated over a different number of tokens. This asymmetry occurs because: +- We summed over document tokens instead of query tokens (different number of terms) +- The iteration direction changed the fundamental computation +- Query and document play fundamentally different roles in the formula + +This non-symmetry is why HNSW indexing becomes problematic for MaxSim - nearest neighbor relationships change depending on which direction you query. + +A document's nearest neighbors are query-dependent and change based on which query you're processing, making it impossible to build the static proximity graph that HNSW requires. The practical reality is you must compute MaxSim against every document at query time with brute force comparison. For large collections, this becomes slow, or even impossible. + +The common solution is **two-stage retrieval**, and we'll explore that pattern in later lessons and see how Qdrant optimizes multi-vector search. + +## Connecting the Concepts + +In the previous lesson, you learned about the late interaction paradigm - encoding queries and documents into multiple token vectors and deferring comparison until search time. MaxSim is the mathematical operation that makes this deferred comparison work. It's the bridge between the conceptual model (multi-vector representations) and practical implementation. Without MaxSim or a similar aggregation function, we'd have no way to score documents against queries when both are represented as sequences of vectors. + +## What's Next + +MaxSim gives us a way to compare multi-vector representations, capturing fine-grained semantic matching that single-vector search cannot achieve. But as we've seen, it also introduces significant computational challenges - particularly the asymmetry that breaks traditional indexing approaches like HNSW. + +In the next lesson, we'll explore **use cases where multi-vector search excels** despite these challenges. You'll see scenarios where the improved relevance and semantic precision justify the additional computational cost. + +From there, we'll learn how to implement efficient multi-vector search in Qdrant, including the hybrid retrieval patterns that make late interaction practical at scale. diff --git a/qdrant-landing/content/course/multi-vector-search/module-1/multi-vector-in-qdrant.md b/qdrant-landing/content/course/multi-vector-search/module-1/multi-vector-in-qdrant.md new file mode 100644 index 000000000..09b843e55 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-1/multi-vector-in-qdrant.md @@ -0,0 +1,188 @@ +--- +title: "Multi-Vector Embeddings in Qdrant" +description: Configure Qdrant collections for multi-vector embeddings and learn how to index and query multi-vector data. +weight: 5 +isLesson: true +--- + +{{< date >}} Module 1 {{< /date >}} + +# Multi-Vector Embeddings in Qdrant + +You've learned how MaxSim enables fine-grained token-level matching and explored both the benefits and challenges of multi-vector search. Now it's time to put that knowledge into practice. + +Qdrant provides first-class support for multi-vector embeddings, making it straightforward to build search systems that leverage late interaction. In this lesson, you'll learn how to configure Qdrant collections for multi-vector search, index documents with token-level embeddings, and execute queries using MaxSim distance. + +By the end, you'll understand the key configuration parameters, know when to use storage optimization strategies, and be ready to build your own multi-vector search applications. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## Creating a Multi-Vector Collection + +Setting up a collection for multi-vector search requires specific configuration to enable late interaction and MaxSim distance calculation. Here's how to create a collection configured for ColBERT embeddings: + +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient("http://localhost:6333") + +client.create_collection( + collection_name="colbert-search", + vectors_config={ + "colbert": models.VectorParams( + # Size of an individual token vector + size=128, + # Distance function for token similarity + distance=models.Distance.DOT, + # Enable multi-vector mode with MaxSim + multivector_config=models.MultiVectorConfig( + comparator=models.MultiVectorComparator.MAX_SIM, + ), + # Disable HNSW indexing + hnsw_config=models.HnswConfigDiff(m=0), + ), + }, +) +``` + +Let's break down each parameter: + +- **`size`**: The dimensionality of each individual token embedding. +- **`distance`**: The similarity metric used to compare individual token vectors. This is the per-token comparison, and MaxSim will aggregate these scores. +- **`multivector_config`**: Enables multi-vector mode for this collection. The `comparator=MultiVectorComparator.MAX_SIM` parameter tells Qdrant to use MaxSim distance when comparing multi-vector documents to queries. +- **`hnsw_config`**: Setting `m=0` disables HNSW indexing. HNSW graphs don't work with MaxSim either way, so it should be disabled to not create an index we won't use. + +You now have a collection ready for multi-vector search. The configuration handles MaxSim automatically - you just need to provide the embeddings. + +## Indexing Multi-Vector Documents + +With your collection configured, you can now index documents. Each document needs multiple token embeddings - one for each token in the text: + +```python +import uuid + +documents = [ + # Document A: Highly relevant - addresses all query aspects + "When async tasks fail to return database connections, the pool " + "becomes exhausted and requests start failing. Ensuring " + "connections are closed after awaits prevents this.", + + # Document B: Partially relevant - mentions some concepts + "Database resource exhaustion can occur due to limited pool sizes.", + + # Document C: Keyword-stuffed - contains related terms without substance + "Understanding concurrency, async IO, and database performance in " + "Python web applications.", + + # Document D: Completely irrelevant + "Handling training for pythons should be done gradually, starting " + "with short sessions and increasing duration as the snake becomes " + "more comfortable.", +] + +client.upsert( + collection_name="colbert-search", + points=[ + models.PointStruct( + id=uuid.uuid4().hex, + vector={ + "colbert": models.Document( + text=doc, + model="colbert-ir/colbertv2.0", + ) + }, + payload={ + "text": doc, + } + ) + for doc in documents + ] +) +``` + +This example uses `models.Document` for convenience - Qdrant's **local inference** feature powered by FastEmbed integration. You provide the text and model name, and Qdrant handles tokenization and encoding automatically. You can also generate embeddings externally and provide them directly as lists of vectors. + +The `payload` field stores the original text and metadata, returned with search results. + +## Querying with MaxSim + +Querying works the same way - provide your query text and let Qdrant handle MaxSim computation: + +```python +query = "How can I prevent Python database connection pool exhaustion?" + +results = client.query_points( + collection_name="colbert-search", + query=models.Document( + text=query, + model="colbert-ir/colbertv2.0", + ), + using="colbert", + limit=2, +) +``` + +The results show which documents have the highest token-level semantic overlap with your query. Documents about async connection pool exhaustion rank higher than generic database documents because more query tokens find strong matches. + +## Storage Optimization: Offloading to Disk + +As discussed in the previous lesson, multi-vector search requires significantly more memory than traditional single-vector search. A document with 500 tokens stores 500 separate embeddings, creating substantial memory pressure for large collections. + +When you prioritize maximum precision over low latency, Qdrant offers a solution: **offload vectors to disk**. This strategy keeps embeddings on disk rather than loading them into RAM, dramatically reducing memory requirements at the cost of slower query performance: + +```python +client.create_collection( + collection_name="colbert-search-on-disk", + vectors_config={ + "colbert": models.VectorParams( + size=128, + distance=models.Distance.DOT, + multivector_config=models.MultiVectorConfig( + comparator=models.MultiVectorComparator.MAX_SIM, + ), + hnsw_config=models.HnswConfigDiff(m=0), + # Offload vectors to disk + on_disk=True, + ), + }, +) +``` + +The single `on_disk=True` parameter changes the storage strategy. Here's the trade-off: + +**Benefits**: Dramatically reduced memory footprint. Instead of storing thousands of token embeddings in RAM, Qdrant reads them from disk during search. This enables multi-vector search on collections that would otherwise exceed available memory. + +**Costs**: Higher query latency due to disk I/O. Every MaxSim calculation requires reading document embeddings from disk, which is orders of magnitude slower than RAM access. Query times increase from milliseconds to potentially seconds, depending on collection size and hardware. + +**When to use it**: Disk offloading works best when precision is critical but latency constraints are relaxed. Research applications, offline batch processing, or scenarios where you're willing to wait a few seconds for the most accurate results all benefit from this approach. For production systems requiring sub-second response times, you'll need other optimization strategies. + +This is just one optimization technique. Module 3 covers additional approaches including quantization, pooling, and multi-stage retrieval that balance precision, memory, and latency more effectively for production deployments. + +## What's Next + +You've learned how to configure Qdrant for multi-vector search, from creating collections with MaxSim comparators to indexing documents and executing queries. The key takeaways: + +- **Multi-vector collections** require specific configuration: `multivector_config` with `MAX_SIM` comparator +- **HNSW indexing is disabled** because MaxSim doesn't work with static proximity graphs +- **Disk offloading** (`on_disk=True`) reduces memory usage when precision matters more than latency +- **Local inference** with FastEmbed integration lets you provide text directly instead of pre-computed embeddings + +The examples in this lesson focused on text search using ColBERT. But late interaction and MaxSim aren't limited to text. In Module 2, you'll discover how multi-vector embeddings extend to multi-modal data - searching PDFs and images using visual token embeddings with ColPali. diff --git a/qdrant-landing/content/course/multi-vector-search/module-1/problems-multi-vector.md b/qdrant-landing/content/course/multi-vector-search/module-1/problems-multi-vector.md new file mode 100644 index 000000000..ed2fbbf09 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-1/problems-multi-vector.md @@ -0,0 +1,64 @@ +--- +title: "Problems of Multi-Vector Search" +description: Understand the challenges and limitations of multi-vector search at scale, including memory and performance considerations. +weight: 4 +isLesson: true +--- + +{{< date >}} Module 1 {{< /date >}} + +# Problems of Multi-Vector Search + +Multi-vector search delivers impressive retrieval quality, but it comes with significant challenges. Before deploying multi-vector search in production, you need to understand these limitations and plan accordingly. + +The good news: Module 3 covers optimization techniques that address many of these challenges. + +--- + +
+ +
+ +--- + +## The Indexing Challenge: Why HNSW Doesn't Work + +One of the fundamental challenges with multi-vector search stems from **HNSW indexing incompatibility**. As you learned in the MaxSim lesson, traditional vector search relies on HNSW (Hierarchical Navigable Small World) graphs to enable fast approximate nearest neighbor search. HNSW works by building static proximity graphs that connect similar documents, allowing efficient traversal during queries. + +![HNSW search](/courses/multi-vector-search/module-1/hnsw-search.png) + +However, HNSW requires distance functions to be **symmetric and query-independent**. MaxSim breaks both assumptions by design. Remember that `MaxSim(Q, D) ≠ MaxSim(D, Q)` - when you swap the parameters, you iterate over different token sets, fundamentally changing the computation. A query with 3 tokens produces a different score than iterating over a document's 200 tokens. + +This asymmetry creates **query-dependent nearest neighbor relationships**. A document's closest neighbors change based on which query you're processing, making it impossible to build the static proximity graph that HNSW requires. The practical consequence: **you must compute MaxSim against every document at query time using brute force comparison**. For collections with millions of documents, this becomes prohibitively slow without optimization strategies. + +![Brute-force scan](/courses/multi-vector-search/module-1/brute-force-scan.png) + +## The Resource Challenge: Memory and Computational Overhead + +Beyond indexing challenges, multi-vector search demands significantly more resources than traditional single-vector approaches. You've seen the precision benefits in the previous lesson on use cases - now let's understand the costs. + +The overhead manifests in three ways: **storage, memory, and computation**. While single-vector search represents each document with one embedding (typically at least 384 dimensions), late interaction models like ColBERT use smaller per-token dimensions - around **128 dimensions per token** - but maintain separate embeddings for every token. + +Even with these smaller dimensions, the total storage explodes. Consider a technical article with 500 tokens: you're storing 500 × 128 = 64,000 floats versus just 384 floats for a single dense embedding - roughly **167x more storage** for that document. The lower dimensionality per token doesn't compensate for the sheer number of vectors. + +![Memory requirements](/courses/multi-vector-search/module-1/memory-requirements.png) + +Each query-document pair requires computing multiple dot products (query tokens × document tokens), creating substantial computational overhead. A query with 10 tokens against a document with 200 tokens means 2,000 dot product operations instead of a single comparison. + +This **memory overhead** combined with brute-force scanning creates fundamental scalability limits. While multi-vector search may work acceptably for collections with thousands or even millions of documents, **brute-force MaxSim doesn't scale indefinitely**. At billion-scale collections, computing MaxSim against every document becomes prohibitively expensive - both in terms of memory required to keep all vectors and latency from scanning the entire dataset. Whether you can deploy multi-vector search in production depends on your collection size, infrastructure budget, and latency requirements. For some applications with manageable collection sizes, the precision gains justify the resource investment. For billion-scale systems, the optimization techniques covered in Module 3 become essential rather than optional. + +## These Challenges Are Solvable + +While these limitations are real, **solutions exist**. The precision benefits of multi-vector search sometimes justify the costs for applications that demand fine-grained semantic matching. Understanding these trade-offs helps you make informed decisions about when and how to deploy multi-vector search in production. + +## What's Next + +Understanding these fundamental limitations prepares you for practical implementation. In the next lesson, you'll learn how to use **Qdrant multi-vector search**, including collection configuration and storage optimization strategies. + +From theory to practice - let's see how Qdrant makes multi-vector search work. diff --git a/qdrant-landing/content/course/multi-vector-search/module-1/use-cases-multi-vector.md b/qdrant-landing/content/course/multi-vector-search/module-1/use-cases-multi-vector.md new file mode 100644 index 000000000..965541c66 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-1/use-cases-multi-vector.md @@ -0,0 +1,200 @@ +--- +title: "Use Cases for Multi-Vector Search" +description: Discover scenarios where multi-vector search outperforms single-vector embeddings and provides better retrieval quality. +weight: 3 +isLesson: true +--- + +{{< date >}} Module 1 {{< /date >}} + +# Use Cases for Multi-Vector Search + +**When is the added complexity of multi-vector search actually worth it?** Multi-vector representations require more storage, more computation, and more careful implementation than simple single-vector embeddings. So why bother? + +The answer comes down to one core capability: **fine-grained matching**. In the previous lessons, you learned how late interaction preserves token-level representations and how MaxSim computes similarity through independent token matching. Now you'll see when this precision actually matters - and when it doesn't. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## The Power of Fine-Grained Matching + +Single-vector embeddings compress entire documents and queries into single points in embedding space. This compression is a form of **lossy averaging** - all the nuanced details, specific requirements, and fine-grained semantics get blended into one representative vector. For many tasks, this works fine. But it fundamentally loses information. + +**Multi-vector search preserves token-level information.** Instead of averaging "Python async database connection pooling" into one vector that represents the general topic, it maintains separate representations for each concept. As you learned in the MaxSim lesson, each query token independently finds its best match in the document, and these matches are summed, not averaged, to produce the final score. + +This distinction is crucial. With single-vector search, a document that mentions Python and databases will likely get a moderate similarity score, even if it never discusses the specific combination you're looking for. With multi-vector search, **every query token must find a strong match** for the overall score to be high. This is token-level verification, not just topical matching. + +## A Concrete Demonstration + +Let's move from theory to practice with a real-world example that shows exactly when multi-vector search makes a difference. + +### The Scenario: A Technical Support Query + +Consider a developer searching for a specific technical solution: + +```python +query = "How can I prevent Python database connection " \ + "pool exhaustion in async web applications?" +``` + +This query has multiple specific requirements: Python, database connections, connection pooling, exhaustion problems, and async web applications. A truly relevant document should address all of these aspects together. + +Now consider four documents with varying relevance: + +```python +documents = [ + # Document A: Highly relevant - addresses all query aspects + "When async tasks fail to return database connections, the pool " + "becomes exhausted and requests start failing. Ensuring " + "connections are closed after awaits prevents this.", + + # Document B: Partially relevant - mentions some concepts + "Database resource exhaustion can occur due to limited pool sizes.", + + # Document C: Keyword-stuffed - contains related terms without substance + "Understanding concurrency, async IO, and database performance in " + "Python web applications.", + + # Document D: Completely irrelevant + "Handling training for pythons should be done gradually, starting " + "with short sessions and increasing duration as the snake becomes " + "more comfortable.", +] +``` + +**What should happen?** Document A should rank highest - it's the only one that directly addresses the problem. Document B is partially relevant. Document C sounds relevant (it mentions Python, async, database, web applications) but provides no substantive answer. Document D is obviously irrelevant (despite containing "python"). + +Let's see how single-vector and multi-vector approaches handle this. + +### Single-Vector Embeddings: Missing the Details + +First, let's try the traditional approach using a single dense vector per document: + +```python +from fastembed import TextEmbedding + +# Load the BAAI/bge-small-en-v1.5 model +dense_model = TextEmbedding("BAAI/bge-small-en-v1.5") +``` + +Encode the query into a single 384-dimensional vector: + +```python +dense_query_vector = next(dense_model.query_embed(query)) +dense_query_vector.shape +``` + +Encode all documents into single vectors: + +```python +import numpy as np + +dense_vectors = np.array(list(dense_model.passage_embed(documents))) +dense_vectors.shape +``` + +Compute similarity scores using dot product: + +```python +np.dot(dense_query_vector, dense_vectors.T) +``` + +**What happened?** The single-vector approach assigns very similar scores to Document A (highly relevant) and Document C (keyword-stuffed). Document C ranks nearly as high as Document A, even though it provides no actual solution! + +The problem: by compressing all information into a single vector, the model captures **topical similarity** but misses whether the document actually addresses the specific requirements. Document C mentions the right topics (Python, async, database, web applications) but doesn't connect them meaningfully. + +### Multi-Vector with ColBERT: Token-Level Verification + +Now let's use ColBERT's multi-vector approach: + +```python +from fastembed import LateInteractionTextEmbedding + +# Load the colbert-ir/colbertv2.0 model +colbert_model = LateInteractionTextEmbedding("colbert-ir/colbertv2.0") +``` + +Encode the query into multiple token-level vectors: + +```python +colbert_query_vector = next(colbert_model.query_embed(query)) +colbert_query_vector.shape +``` + +Encode documents - note that each has a different number of token vectors: + +```python +colbert_vectors = list(colbert_model.passage_embed(documents)) +[cv.shape for cv in colbert_vectors] +``` + +Compute MaxSim scores for each document: + +```python +for colbert_doc_vector in colbert_vectors: + # For each document, compute similarity between all query-doc token pairs + dot_product = np.dot(colbert_query_vector, colbert_doc_vector.T) + # For each query token, take the maximum similarity with any doc token + max_scores = dot_product.max(axis=1) + # Sum these maximum similarities to get the final MaxSim score + print(max_scores.sum()) +``` + +**What happened?** ColBERT produces a clear ranking: +1. **Document A** (highly relevant) - highest score +2. **Document B** (partially relevant) - moderate score +3. **Document C** (keyword-stuffed) - lower score +4. **Document D** (irrelevant) - lowest score + +Document C now correctly ranks **lower** than Documents A and B. The token-level verification catches that while Document C mentions related keywords, it lacks the specific technical details the query requires. + +### Why the Difference? + +**Single-vector behavior**: Document C gets inflated scores because it contains many topically-related terms. The averaging process captures "this document is about Python web development and databases" but can't verify whether it actually addresses connection pool exhaustion. + +**ColBERT behavior**: Each query token must find strong matches in the document: +- "prevent" needs a match -> Document C has no solution-oriented content +- "connection pool exhaustion" needs specific matches -> Document C mentions these words separately but not in context +- "async" + "database" + "Python" must all connect -> Document C has them but not in the right relationship + +**The key insight**: MaxSim's requirement that **every query token finds a strong match** prevents keyword-stuffing from inflating scores. It's not enough to mention the right topics - the document must contain those concepts with the right semantic relationships. + +## When NOT to Use Multi-Vector Search + +Multi-vector search isn't always necessary. Two scenarios where single-vector embeddings are preferable: + +**Simple, broad queries**: When you're searching for general topical relevance rather than verifying specific requirements, single-vector embeddings work well. Queries like "Python tutorials" or "machine learning basics" seek documents about a topic, not documents containing multiple specific concepts. For these queries, the added precision of token-level matching doesn't provide significant value - topical similarity is exactly what you need. + +**Resource-constrained environments**: Multi-vector search comes with substantial overhead: +- **Storage**: Multiple vectors per document instead of one (even thousands of them) +- **Memory**: All token vectors must be accessible during search +- **Computation**: MaxSim is incompatible with HNSW and requires computing similarities across all query-document token pairs + +When deploying to systems with strict latency requirements, these costs may be prohibitive. In these cases, you're making a trade-off: accepting lower precision on complex queries to meet resource constraints. The precision benefits of multi-vector search apply regardless of collection size. Collection size affects whether you can afford the overhead, not whether you need the precision. + +These resource considerations are real challenges you'll face in production. The next lesson examines them in detail. + +## Conclusion + +Multi-vector search excels at **fine-grained matching** - scenarios where you need to verify that ALL aspects of a query are present in a document, not just that they're topically related. By preserving token-level information rather than averaging it away, ColBERT and other late interaction models enable precision that single-vector embeddings cannot achieve. + +When you have multi-requirement queries, need contextual precision, or must distinguish between partial and complete information, MaxSim's token-level verification provides clear advantages. Each query token independently finds its best match, and the aggregation ensures all requirements are strongly present. + +But this power comes at a cost. In the next lesson, we'll examine the challenges: storing hundreds of vectors per document, computing MaxSim efficiently, and the indexing limitations you learned about in the MaxSim lesson. diff --git a/qdrant-landing/content/course/multi-vector-search/module-2/_index.md b/qdrant-landing/content/course/multi-vector-search/module-2/_index.md new file mode 100644 index 000000000..b0c5dfd0d --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-2/_index.md @@ -0,0 +1,22 @@ +--- +title: "Module 2: Multi-Vector Representations for Multi-Modal Data" +description: "Explore multi-modal multi-vector search with ColPali. Learn how to search across images and text, and configure Qdrant for multi-vector embeddings." +isLesson: true +weight: 30 +--- + +{{< date >}} Module 2 {{< /date >}} + +# Multi-Vector Representations for Multi-Modal Data + +Extend multi-vector representations beyond text to unlock powerful multi-modal search capabilities. + +--- + +## Today's path + +1. How ColPali Models Work +2. ColPali Family Overview +3. Visual Interpretability of ColPali + +You'll learn to build multi-modal search systems that understand both images and text. diff --git a/qdrant-landing/content/course/multi-vector-search/module-2/colpali-family.md b/qdrant-landing/content/course/multi-vector-search/module-2/colpali-family.md new file mode 100644 index 000000000..9c30129e4 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-2/colpali-family.md @@ -0,0 +1,130 @@ +--- +title: "ColPali Family Overview" +description: Explore the ColPali model family and their capabilities for multi-modal document understanding and retrieval. +weight: 2 +isLesson: true +--- + +{{< date >}} Module 2 {{< /date >}} + +# ColPali Family Overview + +The ColPali is not only the name of a model. Still, it is also often used to refer to an entire family of models that convert images and text into multi-vector representations, based on Vision Language Models. + +Let's explore what the options are and which model to choose depending on the data you work with. + +--- + +
+ +
+ +--- + +The ColPali family includes several model variants. When selecting a model for your application, you'll need to consider factors like model size, supported languages, computational requirements, and licensing constraints - each variant offers different trade-offs along these dimensions. + +## Model Size + +While the original ColPali model delivers good performance, its multi-billion parameter size can be challenging for resource-constrained environments, demos, or CPU-only deployments. Fortunately, smaller alternatives maintain competitive performance while dramatically reducing computational requirements. + +### ColSmol: Efficient Small-Scale Models + +The **[ColSmol](https://huggingface.co/vidore/colSmol-256M)** family offers compact variants built on SmolVLM, available in [256M](https://huggingface.co/vidore/colSmol-256M) and [500M](https://huggingface.co/vidore/colSmol-500M) parameter sizes. These Apache 2.0 licensed models generate ColBERT-style multi-vector representations while being small enough for resource-constrained environments, like browser-based applications, or edge computing. + +### ColFlor: Ultra-Compact Retrieval + +**[ColFlor](https://huggingface.co/ahmed-masry/ColFlor)** pushes efficiency even further with only 174 million parameters, achieving performance 17× smaller and up to 9.8× faster than ColPali with only a 1.8% drop in accuracy on text-rich English documents. Built on Florence-2's architecture, ColFlor is particularly attractive for demo environments, educational purposes, and scenarios where computational efficiency outweighs marginal performance differences. + +## Supported Languages + +Original ColPali model primarily focuses on English documents, but several multilingual alternatives have emerged that extend multi-vector capabilities across different languages. + +### NVIDIA Multilingual Models + +The **[NVIDIA Llama-NeMoRetriever-ColEmbed-3B-v1](https://huggingface.co/nvidia/llama-nemoretriever-colembed-3b-v1)** was among the leading multilingual visual document retrieval models when it was released. Built on top of Google's SigLIP-2 vision encoder and Meta's Llama 3.2-3B language model, this late interaction embedding model demonstrated strong performance on multilingual retrieval benchmarks including ViDoRe and MIRACL-VISION. + +However, **potential users should be aware of licensing restrictions**. The model is available **for non-commercial and research use only** due to multiple overlapping licenses: NVIDIA's Non-Commercial License, Apache 2.0 for the SigLIP-2 component, and Meta's Llama 3.2 Community License Agreement. Organizations requiring commercial deployment should carefully review these license terms or consider alternatives. + +### Open-Source Multilingual Alternative + +For commercial applications, **[Nomic AI's ColNomic-Embed-Multimodal-7B](https://huggingface.co/nomic-ai/colnomic-embed-multimodal-7b)** offers a compelling fully open-source alternative. Released in early 2025, this model demonstrated competitive performance on multilingual retrieval benchmarks. The model is available under an open-source license that permits commercial use, making it suitable for production deployments without licensing concerns. Nomic AI released a complete suite including both multi-vector (ColNomic) and single-vector variants in 3B and 7B parameter sizes, giving developers flexibility in choosing the right trade-off between performance and resource requirements. + +## Benchmarking Visual Document Retrieval + +The **ViDoRe (Visual Document Retrieval) Benchmark** has emerged as the leading evaluation framework for visual retrieval models on document understanding tasks. As of January 2026, it stands as the largest and most comprehensive benchmark in the field, evaluating models across multiple domains, 6 languages (including English and French), and realistic retrieval scenarios including cross-document and long-form queries. You can explore the latest model performance and compare different approaches on the [ViDoRe leaderboard](https://huggingface.co/spaces/vidore/vidore-leaderboard). + +![ViDoRe V3](/courses/multi-vector-search/module-2/vidore-v3.png) + +***Source:** https://huggingface.co/vidore* + +The current version, **ViDoRe V3**, represents the benchmark's scale and ambition with 26,000+ pages across 3,099 queries in 6 languages, spanning 10 datasets (8 public, 2 private). The benchmark uses **nDCG** (Normalized Discounted Cumulative Gain) as its primary metric and includes challenging datasets spanning diverse domains - including HR, finance, industrial, pharmaceuticals, physics, computer science, and energy sectors. What sets ViDoRe V3 apart is its focus on real-world complexity: models are tested on truly challenging retrieval tasks with human-verified annotations that mirror actual user behavior in enterprise document retrieval scenarios. + +## The Impact of Bidirectional Attention + +Bidirectional attention has emerged as a promising approach for multi-vector representations of multi-modal data. + +The choice between **unidirectional** and **bidirectional** attention mechanisms significantly impacts model performance for embedding tasks. Understanding this distinction helps explain why certain architectures excel at retrieval while others are optimized for generation. Let's first examine unidirectional attention and its limitations, then see how bidirectional attention addresses these constraints. + +### Unidirectional Attention: Designed for Generation + +Most large language models (LLMs) and vision-language models (VLMs) like GPT, Llama, and the base models of [ColPali](https://huggingface.co/vidore/colpali) use **unidirectional (causal) attention**. In this approach, each token can only attend to tokens that came before it in the sequence - the model looks backward but never forward. + +This design makes perfect sense for generative tasks: when predicting the next token, the model should only use past context, not future information it hasn't generated yet. However, this constraint creates limitations for embedding tasks. Token representations encode only information from previous context, missing crucial contextual information from subsequent tokens in the sequence. + +![Unidirectional Attention: Only Past Context](/courses/multi-vector-search/module-2/unidirectional-attention.png) + +The default ColPali model uses the unidirectional attention, derived from the underlying VLM. As a reminder, here is how we load the ColPali v.1.3 with FastEmbed. + +```python +from fastembed import LateInteractionMultimodalEmbedding + +# Load the Qdrant/colpali-v1.3-fp16 model from HF hub +colpali_model = LateInteractionMultimodalEmbedding( + model_name="Qdrant/colpali-v1.3-fp16" +) +``` + +### Bidirectional Attention: Optimized for Embeddings + +**Bidirectional attention**, as used in encoder models like BERT, allows each token to attend to the entire input sequence - both past and future tokens. This creates richer, more contextually informed representations since each token's embedding incorporates information from the complete surrounding context. + +![Bidirectional Attention: Full Context](/courses/multi-vector-search/module-2/bidirectional-attention.png) + +Research has shown that [bidirectional encoder models are often the best option when training visual retrievers](https://arxiv.org/html/2510.01149). The [**ColModernVBERT**](https://huggingface.co/ModernVBERT/colmodernvbert) model exemplifies this approach: with only ~250 million parameters - over 10 times fewer than ColPali - it achieves performance only slightly lower on the ViDoRe benchmark. + +![Benchmark results on ViDoRe](/courses/multi-vector-search/module-2/benchmark-results.png) + +***Source:** Teiletche, P., Macé, Q., Conti, M., Loison, A., Viaud, G., Colombo, P., & Faysse, M. (2025). *ModernVBERT: Towards Smaller Visual Document Retrievers*. arXiv preprint arXiv:2510.01149. https://arxiv.org/abs/2510.01149* + +The chart above compares performance on the ViDoRe V2 benchmark. While ColPali and ColQwen achieve strong results with unidirectional attention, **ColModernVBERT stands out as the only bidirectional model** - proving that full-context attention enables competitive performance with dramatically fewer parameters. By allowing context to flow in both directions, these compact models generate embeddings that capture the nuanced semantic relationships critical for accurate multi-modal retrieval where visual and textual information must be jointly understood. + +This efficiency makes ModernVBERT particularly attractive for resource-constrained environments or CPU-only deployments, where the slight performance trade-off is worthwhile for the substantial gains in speed and reduced computational requirements. + +ColModernVBERT is available in FastEmbed and might be used like any other late interaction model for multi-modal data. + +```python +from fastembed import LateInteractionMultimodalEmbedding + +# Load the Qdrant/colmodernvbert model from HF hub +colpali_model = LateInteractionMultimodalEmbedding( + model_name="Qdrant/colmodernvbert" +) +``` + +## Future of Multi-Vector Representations + +The multi-vector approach extends naturally to new modalities beyond images. Models like **[TomoroAI/tomoro-colqwen3-embed-8b](https://huggingface.co/TomoroAI/tomoro-colqwen3-embed-8b)**, based on ColQwen3, demonstrate this extensibility by adding support for short video retrieval. Based on preliminary findings, ColQwen3 generalizes to videos while learning from image-text retrieval tasks - it samples video clips, encodes frames, then pools frame embeddings with per-dimension max before MaxSim scoring. + +While not yet fine-tuned on large-scale video retrieval datasets, this lightweight approach highlights a key strength of the multi-vector paradigm, where **new modalities can be added** while preserving the core late interaction architecture. + +## What's next + +You've explored the ColPali family ecosystem - from ultra-compact models like ColFlor (174M parameters) to multilingual alternatives from NVIDIA and Nomic AI. You've learned how bidirectional attention architectures can achieve competitive results with far fewer parameters, and seen the benchmark results that guide model selection based on your performance and resource requirements. + +Now that you know which model to use, let's learn how to interpret what ColPali "sees" in your documents. diff --git a/qdrant-landing/content/course/multi-vector-search/module-2/how-colpali-works.md b/qdrant-landing/content/course/multi-vector-search/module-2/how-colpali-works.md new file mode 100644 index 000000000..01e1a1ede --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-2/how-colpali-works.md @@ -0,0 +1,305 @@ +--- +title: "How ColPali Models Work" +description: Understand the inner workings of ColPali models and how they generate multi-vector representations for images and documents. +weight: 1 +isLesson: true +--- + +{{< date >}} Module 2 {{< /date >}} + +# How ColPali Models Work + +ColPali extends the late interaction paradigm from text to visual documents. It can process PDFs, images, and scanned documents, generating multi-vector representations that capture both textual and visual information. + +Understanding ColPali's architecture helps you leverage its full potential for multi-modal document retrieval. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## From Text to Visual Documents + +**What about documents that aren't just text?** PDFs often contain diagrams, tables, charts, equations, and complex layouts where the visual presentation carries as much meaning as the text itself. + +Traditional approaches to searching visual documents typically involve two steps: first, extract text using OCR (Optical Character Recognition), then search the extracted text. This pipeline has significant limitations: + +- **Lost layout information**: OCR converts documents to plain text, discarding spatial relationships and visual structure +- **Diagram blindness**: Charts, graphs, and diagrams are either ignored or poorly represented +- **Format fragility**: Tables, mathematical notation, and multi-column layouts often get mangled during text extraction + +**ColPali takes a fundamentally different approach**: it treats the document image itself as the primary representation. ColPali "tokenizes" images into spatial patches. Each patch becomes a visual token with its own embedding, enabling token-level matching without ever extracting text. + +This means ColPali can match queries directly to visual regions of document pages - no OCR required. + +## The Vision Language Model Foundation + +Before diving into how ColPali processes images, let's understand its architectural foundation. **ColPali isn't built from scratch** - it's based on sophisticated Vision Language Models (VLMs) that already understand the relationship between visual and textual information. + + + +### Building on PaliGemma + +ColPali v1.3 is built on **PaliGemma-3B**, a Vision Language Model that combines two powerful components: + +- **SigLIP-So400m** (Vision Encoder): Processes images into visual features +- **Gemma-2B** (Language Model): Contextualizes those features using transformer layers + +Why use a VLM instead of training a vision model from scratch? VLMs are pre-trained on massive datasets of image-text pairs, so they already "understand" how visual content relates to language. This makes them ideal for document retrieval: they can naturally connect text queries to corresponding visual content. + +PaliGemma was specifically designed for tasks requiring both visual understanding and language processing - perfect for searching documents that blend text, diagrams, tables, and equations. + +**Fine-tuning for Document Retrieval**: The base PaliGemma model is adapted for document retrieval using LoRA (Low-Rank Adaptation), a parameter-efficient technique that specializes the model without retraining everything. The vision encoder stays frozen while attention layers are fine-tuned to optimize for document search. + +### The Two-Component Architecture + +Think of ColPali's architecture as having "eyes" and a "brain": + +**Vision Encoder (SigLIP) - The Eyes:** +- Takes the document image and divides it into patches (the 32×32 grid we'll explore next) +- Processes each patch through a vision transformer +- Outputs visual features that capture what's in each patch: text characters, diagram elements, table cells, etc. + +**Language Model (Gemma-2B) - The Brain:** +- Receives the visual features from SigLIP +- Runs them through transformer layers to add context and semantic understanding +- Each patch's features get enriched by understanding neighboring patches +- Outputs contextualized representations that understand relationships across the document + +**The Flow:** Image patches are first processed through the SigLIP encoder to generate visual features. These visual features then pass through the Gemma-2B transformer, which produces contextualized representations. Finally, the contextualized representations flow through a projection layer to produce 128-dimensional embeddings. This pipeline ensures that each 128-dim embedding doesn't just capture what's in one patch - it understands that patch in the context of the entire document. + +![ColPali architecture diagram](/courses/multi-vector-search/module-2/colpali-architecture.png) + +For text inputs, there isn't any additional preprocessing step, but it is tokenized and passed through the language model. This architecture can use the same "brain" to process different modalities. + +## Patch-Based Image Processing + +Vision transformers, the foundation of ColPali, don't process entire images at once. Instead, they divide images into a grid of fixed-size patches - think of it like a checkerboard overlaid on your document. + +The ColPali pipeline starts with a document. It is usually a screenshot of a single PDF page, or anything you want to encode. While you might convert a PDF page to a high-resolution screenshot to preserve visual details during conversion, **the ColPali preprocessing itself resizes images to a fixed resolution** of **448×448 pixels** before processing. + +This fixed-size input is then divided into patches: +- **Patch size**: 14×14 pixels per patch +- **Patch grid**: 32×32 patches (448 / 14 = 32 patches per side) +- **Total patches**: 1024 patches per image (32 × 32 = 1024) + +![ColPali image processing](/courses/multi-vector-search/module-2/colpali-image-processing.png) + +**Each patch becomes a visual "token"** representing a spatial region of the document. A patch might contain part of a diagram, a few words of text, a piece of an equation, or even just whitespace - each gets its own embedding. This consistent structure ensures predictable embedding sizes and simplifies the search pipeline. + +## Token Structure and Embeddings + +Now that you understand ColPali's VLM foundation - SigLIP processing patches, Gemma-2B contextualizing features, and the projection layer generating embeddings - let's see how this architecture processes images into the final multi-vector representation. + +When ColPali processes an image, the PaliGemma model doesn't just generate patch embeddings directly. It creates a sequence of tokens that flows through the entire architecture we just described. This sequence includes both the image patches and additional instruction tokens that help guide the model's behavior. + +For a single document image, the token sequence contains: +- **1024 special `` tokens** - one for each patch in the 32×32 grid +- **Instruction tokens** - additional tokens that help the model understand its task + +`...Describe the image.\n` + +The final embedding output contains **1030 vectors** of 128 dimensions each: 1024 for the image patches plus 6 additional instruction tokens. For search purposes, all vectors participate in the MaxSim calculation - each query token finds its best match among all 1030 vectors. + +With this multi-vector representation, each query word finds its best match among all 1030 visual vectors using **late interaction**. Query terms match to the most semantically similar patches - whether those patches contain diagrams, text, tables, or equations - exactly the same MaxSim scoring we learned about in Module 1, but extended to visual documents. + +## Implementing Multi-Modal Search with FastEmbed and Qdrant + +Now that you understand ColPali's architecture - from its VLM foundation to patch-based processing to multi-vector embeddings - let's build a practical multi-modal search system. We'll again use **FastEmbed**, as it also provides a unified interface for multi-modal late interaction models like ColPali. + +### Loading ColPali with FastEmbed + +Let's start by loading a ColPali model. FastEmbed supports several ColPali variants, but in this lesson we'll use ColPali v1.3: + +```python +from fastembed import LateInteractionMultimodalEmbedding + +# Load the Qdrant/colpali-v1.3-fp16 model from HF hub +colpali_model = LateInteractionMultimodalEmbedding( + model_name="Qdrant/colpali-v1.3-fp16" +) +``` + +This single initialization loads the entire ColPali architecture we discussed: SigLIP vision encoder, Gemma-2B language model, and the projection layer. The `LateInteractionMultimodalEmbedding` class handles both image and text encoding through a unified interface, encapsulating the image and text processing logic so you can just pass image paths or queries. + +### Processing Images and Understanding Embeddings + +Now let's process a document image through ColPali to see the multi-vector embeddings in action: + +```python +from PIL import Image + +image = Image.open("images/einstein-newspaper.jpg") +image +``` + +Now let's generate the embeddings: + +```python +# Create the representation of the image +colpali_generator = colpali_model.embed_image(["images/einstein-newspaper.jpg"]) +document_vectors = next(colpali_generator) +document_vectors.shape +``` + +The output shape shows **1030 vectors of 128 dimensions each**. These 1030 vectors include the 1024 image patch embeddings (from the 32×32 grid) plus 6 additional instruction tokens used by the model. + +**Processing query text:** + +ColPali also encodes text queries into multi-vector representations: + +```python +# Create the representation of the query +query = "When did dr. Einstein die?" +query_vectors = next(colpali_model.embed_text(query)) +query_vectors.shape +``` + +Let's define a helper function to compute the MaxSim score: + +```python +import numpy as np + +def maxsim(Q, D): + sims = np.dot(Q, D.T) + max_sims = sims.max(axis=1) + return max_sims.sum() +``` + +Now finally compute the score: + +```python +maxsim(query_vectors, document_vectors) +``` + +During search, each query token finds its best match among all 1030 document vectors using **MaxSim** - the same late interaction scoring we learned in Module 1. A higher score indicates better semantic overlap between the query and the visual document. + +### Setting Up Qdrant for Multi-Vector Search + +To store and search these multi-vector embeddings, we need a Qdrant collection configured for late interaction: + +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient("http://localhost:6333") + +client.create_collection( + collection_name="colpali", + vectors_config={ + "colpali-v1.3": models.VectorParams( + size=128, + distance=models.Distance.DOT, + multivector_config=models.MultiVectorConfig( + comparator=models.MultiVectorComparator.MAX_SIM, + ), + hnsw_config=models.HnswConfigDiff(m=0), + ), + }, +) +``` + +This configuration mirrors what you learned in [Module 1](/course/multi-vector-search/module-1/multi-vector-in-qdrant/) - multi-vector config with MAX_SIM comparator, dot product distance, and HNSW disabled. + +### Indexing Visual Documents + +Now let's index visual documents using Qdrant's **local inference** feature. This approach uses `models.Image()` to let Qdrant handle embedding generation automatically via FastEmbed: + +```python +import uuid + +documents = [ + "images/einstein-newspaper.jpg", + "images/titanic-newspaper.jpg", + "images/men-walk-on-moon-newspaper.jpg", +] +client.upsert( + collection_name="colpali", + points=[ + models.PointStruct( + id=uuid.uuid4().hex, + vector={ + "colpali-v1.3": models.Image( + image=doc, + model="Qdrant/colpali-v1.3-fp16", + ) + }, + payload={ + "image": doc, + } + ) + for doc in documents + ] +) +``` + +This uses **local inference** - similar to `models.Document()` for text that you saw in Module 1. Qdrant's FastEmbed integration processes images locally (not on a remote server), generating the multi-vector embeddings automatically. This is much simpler than manually generating embeddings with FastEmbed and then upserting them. + +Once indexed, each document image is represented by its 1030 vectors (1024 patches + 6 instruction tokens), ready for late interaction search. When a query comes in, Qdrant will compute MaxSim scores between the query tokens and these vectors. + +### Querying with Late Interaction + +The power of ColPali becomes clear when searching. Using local inference, you can query with text and match against visual documents: + +```python +client.query_points( + collection_name="colpali", + query=models.Document( + text="Who was the first man on the moon?", + model="Qdrant/colpali-v1.3-fp16", + ), + using="colpali-v1.3", + limit=2, +) +``` + +This uses the same local inference pattern you saw in Module 1 with `models.Document()`. Qdrant handles query tokenization, embedding generation, and MaxSim computation automatically. + +**Let's try different queries:** + +```python +client.query_points( + collection_name="colpali", + query=models.Document( + text="Why did Titanic sink?", + model="Qdrant/colpali-v1.3-fp16", + ), + using="colpali-v1.3", + limit=2, +) +``` + +**How late interaction works for visual search:** + +1. Your query is tokenized into query embeddings (e.g., 6-8 tokens for "Who was the first man on the moon?") +2. For each query token, Qdrant finds the maximum similarity across all 1030 visual vectors of each document image +3. The maximum similarities are summed (MaxSim) to score each image +4. Images with content matching the query score highest + +This is **text-to-image search without OCR**. You're matching the semantic meaning of text queries directly to visual content - headlines, photos, captions, and text in its original layout. The late interaction paradigm enables fine-grained matching at the token level, just like ColBERT for text, but extended to visual documents. + +## What's Next + +Now that you understand **how ColPali works** and can build basic multi-modal search systems, several questions naturally arise: + +- **Which ColPali model should you use?** The ColPali family includes variants optimized for different hardware constraints and accuracy requirements. +- **How can you interpret what the model is finding?** Unlike black-box embeddings, ColPali offers powerful visualization capabilities to see exactly which image regions match your query terms. +- **How do you optimize for production?** Multi-vector search can be resource-intensive at scale - you'll need techniques like quantization, pooling, and multi-stage retrieval. + +In the next lesson, we'll explore the **ColPali family** and learn when to use each model variant. Then, we'll dive into visual interpretability to understand exactly what ColPali sees in your documents. diff --git a/qdrant-landing/content/course/multi-vector-search/module-2/visual-interpretability.md b/qdrant-landing/content/course/multi-vector-search/module-2/visual-interpretability.md new file mode 100644 index 000000000..2ee719709 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-2/visual-interpretability.md @@ -0,0 +1,399 @@ +--- +title: "Visual Interpretability of ColPali" +description: Learn how to visualize and interpret ColPali embeddings to understand what the model focuses on in images. +weight: 3 +isLesson: true +--- + +{{< date >}} Module 2 {{< /date >}} + +# Visual Interpretability of ColPali + +**Why did this document match my query?** Unlike traditional black-box embedding models that produce a single opaque vector, ColPali's multi-vector architecture offers something remarkable: you can see exactly where the model "looks" when matching a query to a document. + +This visual interpretability is invaluable for building trust in multi-modal search systems, debugging unexpected results, and understanding model behavior and limitations. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## The Key Insight: Spatial Correspondence + +In the [previous lesson on ColPali's architecture](/course/multi-vector-search/module-2/how-colpali-works/), you learned that ColPali divides images into a 32×32 grid of patches: + +- **Input image**: 448×448 pixels +- **Patch size**: 14×14 pixels +- **Patch grid**: 32×32 patches +- **Total patch embeddings**: 1024 vectors + +The crucial insight for interpretability is that **each embedding maintains a known spatial location**. Patch index `i` (where `i` ranges from 0 to 1023) maps directly to a position in the grid: + +$$ +\text{row} = \left\lfloor \frac{i}{32} \right\rfloor +$$ + +$$ +\text{col} = i \mod 32 +$$ + +This position corresponds to a specific pixel region in the original image: + +$$ +\text{pixel\_region} = \text{image}\left[\text{row} \cdot 14 : (\text{row}+1) \cdot 14, \; \text{col} \cdot 14 : (\text{col}+1) \cdot 14\right] +$$ + + + +This spatial correspondence is what makes visual interpretability possible. When a query token has high similarity with a particular patch embedding, we know exactly where in the document that match occurred. + +## Computing Token-Patch Similarities + +To visualize what ColPali focuses on, we compute the similarity between each query token and all document patches. For a given query token embedding, we can calculate its similarity with each of the 1024 patch embeddings: + +```python +import numpy as np + +def compute_similarity_map(query_token_vec, doc_vectors): + """Compute similarity map for a single query token.""" + # Take only the 1024 patch embeddings (excluding instruction tokens) + patch_vectors = doc_vectors[:1024] + + # Compute dot product similarity with all patches + similarities = np.dot(patch_vectors, query_token_vec) + + # Reshape to 32×32 spatial grid + return similarities.reshape(32, 32) +``` + +The result is a 32×32 similarity map - essentially a heatmap showing where in the document this particular query token has the strongest matches. High values indicate regions where the model finds semantic relevance to that token. + +## Manual Implementation with FastEmbed + +Let's build a complete interpretability system from scratch using only FastEmbed and standard Python libraries. This approach works with any late interaction model and gives you full control over the visualization process. + +### Step 1: Generate Embeddings + +First, let's load a model and generate embeddings for both a document image and a query: + +```python +from fastembed import LateInteractionMultimodalEmbedding + +# Load ColPali model +model = LateInteractionMultimodalEmbedding( + model_name="Qdrant/colpali-v1.3-fp16" +) + +# Load and embed a document image +image_path = "images/einstein-newspaper.jpg" +doc_vectors = next(model.embed_image([image_path])) + +# Embed a query +query = "When did Einstein die?" +query_vectors = next(model.embed_text(query)) + +print(f"Document embeddings shape: {doc_vectors.shape}") # (1030, 128) +print(f"Query embeddings shape: {query_vectors.shape}") # (N, 128) where N = number of tokens +``` + +### Step 2: Compute Similarity Maps for All Query Tokens + +Now we compute a similarity map for each query token: + +```python +def compute_all_similarity_maps(query_vectors, doc_vectors): + """Compute similarity maps for all query tokens.""" + return np.array([ + compute_similarity_map(query_token_vec, doc_vectors) + for query_token_vec in query_vectors + ]) + +# Compute similarity maps +similarity_maps = compute_all_similarity_maps(query_vectors, doc_vectors) +print(f"Similarity maps shape: {similarity_maps.shape}") # (N, 32, 32) +``` + +## Creating Heatmap Visualizations + +To visualize the similarity maps, we need to: +1. Upsample the 32×32 map to match the image dimensions +2. Overlay it as a semi-transparent heatmap on the original image + +### Step 3: Upsample and Overlay + +```python +from scipy.ndimage import zoom +import matplotlib.pyplot as plt +import matplotlib.cm as cm + +def create_heatmap_overlay(image, similarity_map, alpha=0.5): + """Create a heatmap overlay on the original image.""" + # Ensure image is in RGB and resized to 448×448 (ColPali's input size) + if isinstance(image, str): + image = Image.open(image) + image = image.convert("RGB").resize((448, 448)) + image_array = np.array(image) + + # Upsample similarity map from 32×32 to 448×448 + # zoom factor = 448/32 = 14 + # Convert to float64 as scipy.ndimage.zoom doesn't support all dtypes (e.g., float16) + upsampled_map = zoom(similarity_map.astype(np.float64), 14, order=1) + + # Normalize to [0, 1] range + min_val = upsampled_map.min() + max_val = upsampled_map.max() + if max_val > min_val: + normalized_map = (upsampled_map - min_val) / (max_val - min_val) + else: + normalized_map = np.zeros_like(upsampled_map) + + # Apply colormap (using 'jet' for red=high, blue=low) + heatmap = cm.jet(normalized_map)[:, :, :3] # Remove alpha channel + heatmap = (heatmap * 255).astype(np.uint8) + + # Blend with original image + blended = (alpha * heatmap + (1 - alpha) * image_array).astype(np.uint8) + + return Image.fromarray(blended), normalized_map +``` + +### Step 4: Visualize Multiple Query Tokens + +Let's create a side-by-side visualization of what each query token focuses on: + +```python +def visualize_query_tokens(image_path, query, model, num_tokens_to_show=5): + """Visualize similarity maps for each query token.""" + # Generate embeddings + doc_vectors = next(model.embed_image([image_path])) + query_vectors = next(model.embed_text(query)) + + # Compute similarity maps + similarity_maps = compute_all_similarity_maps(query_vectors, doc_vectors) + + # Load original image + original_image = Image.open(image_path).convert("RGB").resize((448, 448)) + + # Limit number of tokens to display + n_tokens = min(len(similarity_maps), num_tokens_to_show) + + # Create figure + fig, axes = plt.subplots(1, n_tokens + 1, figsize=(4 * (n_tokens + 1), 4)) + + # Show original image + axes[0].imshow(original_image) + axes[0].set_title("Original") + axes[0].axis("off") + + # Show heatmap for each token + for i in range(n_tokens): + overlay, _ = create_heatmap_overlay(original_image, similarity_maps[i]) + axes[i + 1].imshow(overlay) + axes[i + 1].set_title(f"Token {i}") + axes[i + 1].axis("off") + + plt.suptitle(f'Query: "{query}"', fontsize=14) + plt.tight_layout() + plt.show() + +# Visualize what each token focuses on +visualize_query_tokens( + "images/einstein-newspaper.jpg", + "When did Einstein die?", + model, + num_tokens_to_show=6 +) +``` + +![Multi-token visualization](/courses/multi-vector-search/module-2/multi-token-visualization.png) + +## Practical Example: Debugging a Search + +Visual interpretability becomes powerful when debugging search results. Let's walk through a complete example to understand why certain documents match (or don't match) specific queries. + +### Scenario: Investigating an Unexpected Match + +Imagine you're searching for "bar chart showing revenue" and get an unexpected result. Let's visualize what's happening: + +```python +from transformers import AutoTokenizer + +# Load tokenizer for ColPali (based on PaliGemma) +tokenizer = AutoTokenizer.from_pretrained("google/paligemma-3b-pt-224") + +def debug_search_result(image_path, query, model, tokenizer): + """Debug why a document matched a query.""" + # Generate embeddings + doc_vectors = next(model.embed_image([image_path])) + query_vectors = next(model.embed_text(query)) + + # Tokenize query to get actual token strings + # ColPali uses "Query: " prefix internally + query_with_prefix = f"Query: {query}" + tokens = tokenizer.tokenize(query_with_prefix) + + # Compute MaxSim score + similarities = np.dot(query_vectors, doc_vectors.T) + max_sims = similarities.max(axis=1) + total_score = max_sims.sum() + + print(f"Query: {query}") + print(f"Total MaxSim Score: {total_score:.2f}") + print(f"\nPer-token contributions:") + + # Show contribution of each token + for i, (max_sim, token_sims) in enumerate(zip(max_sims, similarities)): + # Find which patch this token matched best with + best_patch_idx = token_sims[:1024].argmax() + row, col = best_patch_idx // 32, best_patch_idx % 32 + # Display actual token text (fall back to index if out of range) + token_str = tokens[i] if i < len(tokens) else f"[pad_{i}]" + print(f" '{token_str}': score={max_sim:.3f}, best match at patch ({row}, {col})") + + return total_score, max_sims + +# Debug the search result +score, token_scores = debug_search_result( + "images/financial-report.png", + "bar chart showing revenue", + model, + tokenizer +) +``` + +This analysis shows you: +- The total relevance score +- How much each query token contributes +- Where in the document each token found its best match + +If a token like "revenue" is matching in an unexpected location, the visualization reveals whether the model is correctly identifying revenue-related content or making an error. + +### Interpreting the Results + +When analyzing heatmaps: + +- **Concentrated heat**: The token is focusing on a specific region - good for precise matches +- **Diffuse heat**: The token finds multiple relevant regions or isn't strongly matched anywhere +- **Unexpected locations**: May indicate the model is matching based on visual similarity rather than semantic meaning + +## Aggregated MaxSim Visualization + +Sometimes you want to see which patches contribute most to the overall score, regardless of which query token they matched. This **aggregated MaxSim view** shows the document-level relevance: + +```python +def compute_maxsim_contribution(query_vectors, doc_vectors): + """Compute how much each patch contributes to the MaxSim score.""" + # Use only patch embeddings + patch_vectors = doc_vectors[:1024] + + # Compute all pairwise similarities + similarities = np.dot(query_vectors, patch_vectors.T) # (n_query, 1024) + + # For each patch, take the maximum contribution across all query tokens + # This shows which patches are most "useful" for any query token + max_contribution = similarities.max(axis=0) # (1024,) + + # Reshape to spatial grid + contribution_map = max_contribution.reshape(32, 32) + + return contribution_map + +def visualize_maxsim_contribution(image_path, query, model): + """Visualize which patches contribute most to the MaxSim score.""" + # Generate embeddings + doc_vectors = next(model.embed_image([image_path])) + query_vectors = next(model.embed_text(query)) + + # Compute contribution map + contribution_map = compute_maxsim_contribution(query_vectors, doc_vectors) + + # Create visualization + original_image = Image.open(image_path).convert("RGB").resize((448, 448)) + overlay, _ = create_heatmap_overlay(original_image, contribution_map) + + fig, axes = plt.subplots(1, 2, figsize=(10, 5)) + + axes[0].imshow(original_image) + axes[0].set_title("Original Document") + axes[0].axis("off") + + axes[1].imshow(overlay) + axes[1].set_title(f"MaxSim Contribution\nQuery: \"{query}\"") + axes[1].axis("off") + + plt.tight_layout() + plt.show() + +# Visualize overall contribution +visualize_maxsim_contribution( + "images/einstein-newspaper.jpg", + "When did Einstein die?", + model +) +``` + +![MaxSim contribution visualization](/courses/multi-vector-search/module-2/maxsim-contribution.png) + +This aggregated view is particularly useful for: +- **Understanding document-level relevance**: See which regions make this document match the query +- **Identifying key content**: Highlights the most semantically important patches +- **Quality assessment**: Check if the model focuses on relevant content (text, diagrams) rather than noise + +## A Note on Newer Architectures + +The interpretability techniques we've covered work directly with ColPali because of its simple spatial mapping: 448×448 pixels -> 32×32 patches -> 1024 embeddings. Each patch index maps directly to a spatial location. + +However, **newer architectures use more complex image processing** that makes precise visualization more challenging. + +### Split-Image Processing + +Models like ColModernVBERT and ColIdefics3 use a **split-image approach**: + +1. **Resize**: The image is resized to fit a maximum edge constraint (e.g., 1344 pixels) +2. **Split into sub-patches**: The resized image is divided into fixed-size sub-patches (typically 512×512 pixels) +3. **Token grids per sub-patch**: Each sub-patch becomes a grid of tokens (e.g., 8×8 = 64 tokens) +4. **Global patch**: A downscaled view of the entire image is appended as a final set of tokens + +This means tokens arrive in **sub-patch-sequential order** rather than row-major spatial order. To reconstruct spatial correspondence for visualization, you need to: + +1. Exclude the global patch tokens (they lack spatial correspondence to specific regions) +2. Rearrange tokens from sub-patch order back to a 2D spatial grid +3. Account for varying image dimensions (different images produce different numbers of sub-patches) + +The core insight remains: **multi-vector representations enable interpretability** because each embedding has semantic meaning. The mapping from embedding to image location just becomes more involved with advanced architectures. + +## What's Next + +You've now learned one of ColPali's most powerful features: the ability to see exactly where the model focuses when matching queries to documents. This transparency helps you: + +- Debug unexpected search results +- Build trust in your retrieval system +- Understand model behavior and limitations +- Validate that the model focuses on relevant content + +With a solid understanding of how ColPali works, the model variants available, and how to interpret what the model sees, you're ready to tackle the next challenge: **making these systems production-ready**. + +In Module 3, we'll explore the scalability and optimization techniques you need for real-world deployments - from memory optimization and quantization to multi-stage retrieval pipelines that can handle millions of documents efficiently. diff --git a/qdrant-landing/content/course/multi-vector-search/module-3/_index.md b/qdrant-landing/content/course/multi-vector-search/module-3/_index.md new file mode 100644 index 000000000..95a889bd0 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-3/_index.md @@ -0,0 +1,25 @@ +--- +title: "Module 3: Scalability and Optimization" +description: "Address scalability challenges in multi-vector search. Learn optimization techniques including quantization, pooling, MUVERA, and multi-stage retrieval." +isLesson: true +weight: 40 +--- + +{{< date >}} Module 3 {{< /date >}} + +# Scalability and Optimization + +Tackle the memory and performance challenges of production-scale multi-vector search. + +--- + +## Today's path + +1. Multi-Stage Retrieval with Universal Query API +2. Vector Quantization Techniques +3. Pooling Techniques +4. MUVERA +5. Evaluating Search Pipelines +6. Final Project + +You'll master the optimization strategies needed to deploy multi-vector search at scale. The module concludes with a final project where you'll apply all the learned skills to build a real-world multi-vector search system for multi-modal data. diff --git a/qdrant-landing/content/course/multi-vector-search/module-3/evaluating-pipelines.md b/qdrant-landing/content/course/multi-vector-search/module-3/evaluating-pipelines.md new file mode 100644 index 000000000..51fd642c9 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-3/evaluating-pipelines.md @@ -0,0 +1,589 @@ +--- +title: "Evaluating Search Pipelines" +description: Learn how to evaluate different search configurations in terms of cost, latency, and retrieval quality using ground truth datasets and standardized metrics. +weight: 5 +isLesson: true +--- + +{{< date >}} Module 3 {{< /date >}} + +# Evaluating Search Pipelines + +Throughout this module, you've learned many optimization techniques: quantization to reduce memory, pooling to compress representations, MUVERA for efficient indexing, and multi-stage retrieval to balance speed with accuracy. But how do you know which combination is right for *your* data? + +The answer lies in systematic evaluation across three dimensions: **cost** (memory and compute resources), **latency** (query response time), and **quality** (retrieval accuracy). Cost and latency are straightforward to measure - you can observe memory usage and time queries directly. Quality, however, requires a more principled approach: you need to measure whether your system returns the *right* documents. + +> **Note:** This lesson demonstrates evaluation methodology on a **small, comprehensible dataset** that you can manually inspect and understand. We use 4 document images with 8 queries where you can verify relevance judgments yourself. Real production benchmarks would use larger datasets like ViDoRe, but the methodology remains the same. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## Quick Reference: Evaluation Metrics + +Before diving in, here's the essential guide to choosing metrics. For multi-stage pipelines, focus on **Recall@k for the prefetch stage** (did we capture candidates?) and **NDCG@k for the final ranking** (did we order them correctly?). + +| Metric | When to Use | Key Insight | +|-----------------|----------------------------|----------------------------------------------| +| **Recall@k** | Prefetch stage (k=50, 100) | Did we capture relevant candidates? | +| **NDCG@k** | Final ranking (k=5, 10) | Are most relevant results ranked highest? | +| **MRR** | Single-answer retrieval | How quickly do we find THE right answer? | +| **Precision@k** | Result page quality | What fraction of shown results are relevant? | + +## Ground Truth: Defining "Correct" Retrieval + +To measure retrieval quality, you need **relevance judgments** (called **qrels**) - data that defines which documents are relevant for which queries. + +### What Are Qrels? + +Each qrel is a triplet: a query, a document, and a relevance score indicating how relevant that document is to the query. + +```python +# Ground truth: which documents are relevant to which queries? +qrels_dict = { + "company quarterly financial results and revenue": { + "images/financial-report.png": 3, # Highly relevant + }, + "historic ship disaster at sea": { + "images/titanic-newspaper.jpg": 3, # Highly relevant + }, + "news headline from early 1900s": { + "images/titanic-newspaper.jpg": 3, # Highly relevant + "images/einstein-newspaper.jpg": 2, # Somewhat relevant + }, +} +``` + +Relevance can be **binary** (0 = not relevant, 1 = relevant) or **graded** (0-3 scale where higher means more relevant). Graded relevance provides more nuance for ranking evaluation. + +### Building Ground Truth + +Three practical approaches: + +1. **Manual Annotation** - Have domain experts review query-document pairs. Highest quality but time-intensive. Even 50-100 queries provides valuable signal. + +2. **Synthetic Generation with LLMs** - Use language models to generate relevant queries for documents. Scales well but may not capture real user query patterns. + +3. **Existing Benchmarks** - For visual document retrieval, **ViDoRe** provides standardized evaluation. For text retrieval, **BEIR** offers diverse domains. + +For this lesson, we'll use a small manually-annotated dataset where you can inspect the relevance judgments yourself. + +## The Degrees of Freedom Problem + +With so many optimization techniques, the configuration space explodes quickly: + +| Dimension | Options | +|------------------|---------------------------------------------------------------------| +| **Quantization** | None, Scalar (int8), Binary (1-bit, 1.5-bit, 2-bit), Product | +| **Pooling** | None, Hierarchical (k=16, 32, 64, ...) | +| **MUVERA** | Disabled, or Enabled (differenr k_sim, dim_proj, r_reps variations) | +| **Multi-stage** | Single-stage, Two-stage (prefetch: 50, 100, 500, ...) | + +Just considering these options yields hundreds of possible configurations. The key insight: **you don't need to test them all**. Instead, you need a systematic approach to test *representative* configurations efficiently. + +## Unified Collection Architecture + +The traditional approach creates a separate collection for each configuration - tedious and storage-intensive. A better approach: **store all vector representations in a single collection** using named vectors. + +### Four Named Vectors + +We'll store four different representations of each document: + +```python +from qdrant_client import QdrantClient, models +from qdrant_client.models import ( + VectorParams, Distance, MultiVectorConfig, MultiVectorComparator, + ScalarQuantization, ScalarQuantizationConfig, ScalarType, +) + +client = QdrantClient(url="http://localhost:6333") + +COLLECTION_NAME = "eval-multi-vector" + +client.create_collection( + collection_name=COLLECTION_NAME, + vectors_config={ + # Full ColModernVBERT multi-vector (no quantization) + "colmodernvbert": VectorParams( + size=128, + distance=Distance.DOT, + multivector_config=MultiVectorConfig( + comparator=MultiVectorComparator.MAX_SIM + ), + hnsw_config=models.HnswConfigDiff(m=0), # Disable HNSW for multi-vector + ), + # ColModernVBERT with scalar quantization enabled + "colmodernvbert_sq": VectorParams( + size=128, + distance=Distance.DOT, + multivector_config=MultiVectorConfig( + comparator=MultiVectorComparator.MAX_SIM + ), + hnsw_config=models.HnswConfigDiff(m=0), + quantization_config=ScalarQuantization( + scalar=ScalarQuantizationConfig( + type=ScalarType.INT8, + quantile=0.99, + always_ram=True, + ) + ), + ), + # MUVERA single-vector approximation for fast HNSW search + "muvera": VectorParams( + size=40960, # muvera.embedding_size from k_sim=6, dim_proj=32, r_reps=20 + distance=Distance.COSINE, + ), + # Hierarchical pooled multi-vector (k=32 clusters) + "hierarchical": VectorParams( + size=128, + distance=Distance.DOT, + multivector_config=MultiVectorConfig( + comparator=MultiVectorComparator.MAX_SIM + ), + hnsw_config=models.HnswConfigDiff(m=0), + ), + }, +) +``` + +**Key insight:** The `colmodernvbert_sq` vector stores the same embeddings as `colmodernvbert` but with scalar quantization enabled. This allows direct comparison of quantized vs. non-quantized search without re-indexing. + +### Embedding and Uploading Documents + +A single function generates all four representations for each document: + +```python +from fastembed import LateInteractionMultimodalEmbedding +from scipy.cluster.vq import kmeans2 +import numpy as np +from fastembed.postprocess import Muvera + +# Load the embedding model +model = LateInteractionMultimodalEmbedding( + model_name="Qdrant/colmodernvbert" +) + +# Initialize MUVERA with the same configuration as the collection +muvera = Muvera.from_multivector_model(model=model, k_sim=6, dim_proj=32, r_reps=20) + + +def hierarchical_pool(embeddings: np.ndarray, k: int = 32) -> np.ndarray: + """Pool multi-vector to k centroids using k-means clustering.""" + if len(embeddings) <= k: + return embeddings # No pooling needed + centroids, labels = kmeans2(embeddings.astype(np.float64), k, minit="++") + # Return mean of embeddings in each cluster + pooled = np.array([ + embeddings[labels == i].mean(axis=0) + for i in range(k) + if (labels == i).any() + ]) + return pooled.astype(np.float32) + + +def embed_and_upload_document(doc_path: str, doc_id: int) -> None: + """Embed a document and upload all four vector representations.""" + # Generate full multi-vector embeddings + full_multivec = np.array(list(model.embed_image([doc_path]))[0]) + + # Generate MUVERA approximation + muvera_vec = muvera.process_document(full_multivec) + + # Generate hierarchical pooled version (k=32) + hierarchical_vec = hierarchical_pool(full_multivec, k=32) + + # Upload all representations in one point + client.upsert( + collection_name=COLLECTION_NAME, + points=[ + models.PointStruct( + id=doc_id, + payload={"filename": doc_path}, + vector={ + "colmodernvbert": full_multivec.tolist(), + "colmodernvbert_sq": full_multivec.tolist(), # Same data, quantized config + "muvera": muvera_vec.tolist(), + "hierarchical": hierarchical_vec.tolist(), + }, + ) + ], + ) +``` + +### Sample Dataset + +For this lesson, we use a small dataset you can manually inspect: + +```python +# Small dataset of document images you can manually inspect +DOC_PATHS = [ + "images/financial-report.png", + "images/titanic-newspaper.jpg", + "images/men-walk-on-moon-newspaper.jpg", + "images/einstein-newspaper.jpg", +] + +# Upload all documents with all 4 vector representations +for doc_id, doc_path in enumerate(DOC_PATHS): + embed_and_upload_document(doc_path, doc_id) +``` + +## Building Evaluation Pipelines + +With all vectors stored in a single collection, we can build different pipelines **without re-indexing** - we simply query different named vectors. + +### The Search Function + +```python +def search_pipeline( + query_embedding: np.ndarray, + using: str, + prefetch_using: str | None = None, + prefetch_limit: int = 50, + limit: int = 10, +) -> list[tuple[str, float]]: + """ + Execute a search pipeline with optional prefetch stage. + + Args: + query_embedding: The query's multi-vector embedding + using: Named vector for final ranking + prefetch_using: Named vector for prefetch (None = single-stage) + prefetch_limit: How many candidates to retrieve in prefetch + limit: Final number of results + + Returns: + List of (filename, score) tuples + """ + if prefetch_using is None: + # Single-stage search + response = client.query_points( + collection_name=COLLECTION_NAME, + query=query_embedding.tolist(), + using=using, + limit=limit, + ) + else: + # Two-stage search: prefetch with one vector, rerank with another + # For MUVERA prefetch, we need the MUVERA query embedding + if prefetch_using == "muvera": + prefetch_query = muvera.process_query(query_embedding).tolist() + else: + prefetch_query = query_embedding.tolist() + + response = client.query_points( + collection_name=COLLECTION_NAME, + prefetch=[ + models.Prefetch( + query=prefetch_query, + using=prefetch_using, + limit=prefetch_limit, + ) + ], + query=query_embedding.tolist(), + using=using, + limit=limit, + ) + + return [ + (point.payload["filename"], point.score) + for point in response.points + ] +``` + +### Representative Pipeline Configurations + +Rather than testing all combinations, we select **6 representative pipelines** that cover the key trade-offs: + +```python +PIPELINES = { + # Baseline: full quality, no optimization + "baseline": { + "using": "colmodernvbert", + "prefetch_using": None, + }, + + # Scalar quantized: reduced memory, minimal quality loss + "scalar_quantized": { + "using": "colmodernvbert_sq", + "prefetch_using": None, + }, + + # Hierarchical pooling: fewer vectors per document + "hierarchical": { + "using": "hierarchical", + "prefetch_using": None, + }, + + # Two-stage: fast MUVERA prefetch + full quality rerank + "muvera_rerank": { + "using": "colmodernvbert", + "prefetch_using": "muvera", + "prefetch_limit": 50, + }, + + # Two-stage with quantized rerank + "muvera_quantized": { + "using": "colmodernvbert_sq", + "prefetch_using": "muvera", + "prefetch_limit": 50, + }, + + # Maximum compression: MUVERA prefetch + pooled rerank + "muvera_hierarchical": { + "using": "hierarchical", + "prefetch_using": "muvera", + "prefetch_limit": 50, + }, +} +``` + +**Key insight:** All six pipelines query the **same indexed data** - we just configure which named vectors to use for prefetch and final ranking. + +### Running All Pipelines + +```python +def evaluate_all_pipelines( + queries: dict[str, np.ndarray], + qrels: dict, +) -> dict[str, dict]: + """Run all pipeline configurations and collect results.""" + from ranx import Qrels, Run, evaluate + import time + + ranx_qrels = Qrels(qrels) + results = {} + + for pipeline_name, config in PIPELINES.items(): + # Collect search results for all queries + pipeline_results = {} + latencies = [] + + for query_text, query_embedding in queries.items(): + start = time.perf_counter() + search_results = search_pipeline(query_embedding, **config) + latencies.append((time.perf_counter() - start) * 1000) + + # Convert to ranx format: {doc_id: score} + pipeline_results[query_text] = { + filename: score for filename, score in search_results + } + + # Create ranx Run and evaluate + run = Run(pipeline_results, name=pipeline_name) + metrics = evaluate(ranx_qrels, run, ["ndcg@10", "recall@10", "mrr"]) + + results[pipeline_name] = { + "metrics": metrics, + "avg_latency_ms": np.mean(latencies), + } + + return results +``` + +## Evaluation with ranx + +The **ranx** library provides battle-tested implementations of IR metrics. Here's how to compare all our pipelines: + +```python +from ranx import Qrels, Run, compare +import time + +# Define ground truth - which documents are relevant to which queries +qrels_dict = { + "company quarterly financial results and revenue": { + "images/financial-report.png": 3, # Highly relevant + }, + "historic ship disaster at sea": { + "images/titanic-newspaper.jpg": 3, + }, + "space exploration and astronauts": { + "images/men-walk-on-moon-newspaper.jpg": 3, + }, + "physics theory and scientist": { + "images/einstein-newspaper.jpg": 3, + }, + "news headline from early 1900s": { + "images/titanic-newspaper.jpg": 3, + "images/einstein-newspaper.jpg": 2, # Somewhat relevant + }, + "business earnings report": { + "images/financial-report.png": 3, + }, + "NASA moon landing mission": { + "images/men-walk-on-moon-newspaper.jpg": 3, + }, + "ocean liner sinking": { + "images/titanic-newspaper.jpg": 3, + }, +} + +qrels = Qrels(qrels_dict) + +# Generate query embeddings +query_embeddings = { + query: np.array(list(model.embed_text([query]))[0]) + for query in qrels_dict.keys() +} + +# Collect runs from each pipeline +runs = [] +latency_results = {} + +for pipeline_name, config in PIPELINES.items(): + pipeline_results = {} + latencies = [] + + for query_text, query_embedding in query_embeddings.items(): + start = time.perf_counter() + search_results = search_pipeline(query_embedding, **config, limit=10) + latencies.append((time.perf_counter() - start) * 1000) + + # Convert to ranx format: {doc_id: score} + pipeline_results[query_text] = { + filename: score for filename, score in search_results + } + + runs.append(Run(pipeline_results, name=pipeline_name)) + latency_results[pipeline_name] = np.mean(latencies) + +# Compare all pipelines +report = compare( + qrels=qrels, + runs=runs, + metrics=["ndcg@10", "recall@10", "mrr"], + max_p=0.05, # Statistical significance threshold +) +print(report) +``` + +This produces a formatted comparison table showing how each pipeline performs across all metrics, with statistical significance indicators. + +## Choosing a Pipeline: Trade-off Analysis + +With multiple pipelines evaluated across quality and latency, how do you choose? The concept of **Pareto optimality** helps frame the decision. A pipeline is Pareto-optimal if no other pipeline beats it in *all* dimensions simultaneously. For example, `muvera_rerank` might offer 94% of baseline quality at 4x the speed - you can't get both faster *and* higher quality without trade-offs. + +When analyzing your results, plot quality (NDCG@10) against latency or memory usage. Pipelines that fall below the "frontier" of best options are **dominated** - there's always a better choice available. Focus your attention on configurations along this frontier, then choose based on your constraints: quality-first applications should stay closer to baseline, while latency-critical systems can move toward faster approximations like `muvera_hierarchical`. + +In practice, multi-stage pipelines with MUVERA prefetch often provide the best balance - they leverage fast HNSW search for candidate retrieval while preserving full multi-vector quality for final ranking. Start with `muvera_rerank` as a strong default, then adjust the prefetch limit based on your recall requirements. + +## Putting It All Together + +Here's the complete evaluation workflow: + +```python +from fastembed import LateInteractionMultimodalEmbedding +from fastembed.postprocess import Muvera +from ranx import Qrels, Run, compare +import numpy as np + +# 1. Load model and initialize MUVERA +model = LateInteractionMultimodalEmbedding( + model_name="Qdrant/colmodernvbert" +) +muvera = Muvera.from_multivector_model(model=model, k_sim=6, dim_proj=32, r_reps=20) + +# 2. Create unified collection with all named vectors +# (See collection creation code above) + +# 3. Embed and upload documents (generates all 4 representations) +DOC_PATHS = [ + "images/financial-report.png", + "images/titanic-newspaper.jpg", + "images/men-walk-on-moon-newspaper.jpg", + "images/einstein-newspaper.jpg", +] + +for doc_id, doc_path in enumerate(DOC_PATHS): + embed_and_upload_document(doc_path, doc_id) + +# 4. Define ground truth (qrels) +qrels_dict = { + "company quarterly financial results and revenue": { + "images/financial-report.png": 3, + }, + "historic ship disaster at sea": { + "images/titanic-newspaper.jpg": 3, + }, + # ... more queries with relevance judgments +} +qrels = Qrels(qrels_dict) + +# 5. Embed queries +query_embeddings = { + query: np.array(list(model.embed_text([query]))[0]) + for query in qrels_dict.keys() +} + +# 6. Evaluate all pipeline configurations and collect runs +runs = [] +for pipeline_name, config in PIPELINES.items(): + pipeline_results = {} + for query_text, query_embedding in query_embeddings.items(): + search_results = search_pipeline(query_embedding, **config, limit=10) + pipeline_results[query_text] = { + filename: score for filename, score in search_results + } + runs.append(Run(pipeline_results, name=pipeline_name)) + +# 7. Compare all pipelines +report = compare( + qrels=qrels, + runs=runs, + metrics=["ndcg@10", "recall@10", "mrr"], +) +print(report) + +# 8. Select pipeline based on requirements: +# - Quality-first? Use baseline or muvera_rerank +# - Latency-critical? Use muvera_hierarchical +# - Balanced? Use muvera_quantized +``` + +## Summary + +Evaluating multi-vector search pipelines requires: + +1. **Ground truth data** (qrels) that defines what "correct" retrieval means +2. **A unified collection architecture** that stores all vector representations together +3. **Representative pipeline configurations** that cover the key trade-offs +4. **Appropriate metrics** chosen for your use case (NDCG for ranking, Recall for prefetch) +5. **Trade-off analysis** using Pareto frontiers to identify optimal configurations + +The key insight from this lesson: **you don't need separate collections for each configuration**. By storing multiple named vectors (full multi-vector, quantized, MUVERA, hierarchical pooled), you can evaluate many pipeline configurations efficiently without re-indexing. + +--- + +## Course Conclusion + +Congratulations on completing the Multi-Vector Search course! + +You've journeyed from the foundations of late interaction and MaxSim distance, through multi-modal applications with ColPali, to production optimization techniques. You now have the knowledge to: + +- **Understand** why multi-vector representations capture richer semantic relationships +- **Implement** multi-vector search with Qdrant's native support +- **Apply** these techniques to visual document retrieval using ColPali models +- **Optimize** for production with quantization, pooling, MUVERA, and multi-stage retrieval +- **Evaluate** different configurations to find the right trade-offs for your use case + +Multi-vector search is a powerful paradigm that's particularly well-suited for complex retrieval tasks where single-vector representations fall short. As you apply these techniques to your own projects, remember that the best configuration depends on your specific data, queries, and constraints. Use the evaluation framework from this lesson to make data-driven decisions. + +We encourage you to experiment with your own datasets, try different model variants, and share what you learn with the community. The field of multi-vector search is evolving rapidly, and practical insights from real-world applications are invaluable. + +Happy searching! diff --git a/qdrant-landing/content/course/multi-vector-search/module-3/final-project.md b/qdrant-landing/content/course/multi-vector-search/module-3/final-project.md new file mode 100644 index 000000000..4471a1949 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-3/final-project.md @@ -0,0 +1,146 @@ +--- +title: "Final Project: Build Your Own Multi-Vector Search System" +description: Apply everything you've learned to build a multi-vector search system that solves a real problem of your choosing. +weight: 7 +isLesson: true +--- + +{{< date >}} Module 3 {{< /date >}} + +# Final Project: Build Your Own Multi-Vector Search System + +--- + +## Your Mission + +It's time to bring together everything you've learned about multi-vector search, late interaction models, and production optimization. You'll build a sophisticated document retrieval system that leverages late interaction's token-level matching for superior search quality. + +Your search engine will understand the nuanced relationships between query terms and document content. When someone searches for "machine learning applications in healthcare," your system will find documents that discuss relevant concepts even when they use different terminology, thanks to late interaction's fine-grained matching. + +This mirrors real-world challenges in enterprise search, research libraries, and knowledge management. You'll implement the complete pipeline: multi-vector embedding using **ColModernVBERT** or **ColPali**, memory-efficient storage with quantization or pooling, and optimized retrieval with multi-stage search. + +--- + +## What You'll Build + +A working multi-vector search system that: + +- Indexes real documents using late interaction embeddings +- Applies optimization techniques to manage memory and latency +- Retrieves relevant results with measurable quality +- Documents your design decisions and trade-offs + +Both ColModernVBERT and ColPali work well for visual document understanding - choose based on your preference or experiment with both. + +The specific dataset, optimization strategy, and retrieval configuration are up to you. + +--- + +## Choose Your Challenge + +This is your project. Pick a problem that matters to you. + +### Use Your Own Data + +The most valuable learning comes from working with documents you actually care about: + +- **Technical documentation** you reference frequently +- **Research papers** in your field of interest +- **Internal documents** (reports, manuals, wikis) from your work +- **Personal collection** of PDFs, articles, or notes + +Aim for **at least 50 documents** to have enough variety for meaningful evaluation. More is better for seeing how your system scales. + +### Public Datasets (If You Prefer) + +If you don't have a suitable personal dataset: + +- **ArXiv papers** on a topic you're curious about +- **Wikipedia articles** from a category you'd like to explore +- **Open-source documentation** from projects you use +- **News articles** from a domain you follow + +The key is picking something where you can judge search quality intuitively. You'll need to create evaluation queries, and that's easier when you understand the content. + +--- + +## Project Requirements + +Your project should demonstrate: + +### 1. Working Multi-Vector Search + +**Use ColModernVBERT** to index your documents and retrieve results using MaxSim scoring. The system should return relevant documents for natural language queries. Alternatively, consider **ColPali**. + +### 2. At Least One Optimization Technique + +Apply something you learned in this module: + +- Binary or scalar quantization +- Token pooling (clustering, attention-based, or hierarchical) +- MUVERA indexing +- Multi-stage retrieval pipeline + +Measure the impact of your chosen technique on memory, latency, or search quality. + +### 3. Evaluation with Ground Truth + +Create a test set of queries with known relevant documents. This doesn't need to be exhaustive - 10-20 queries with 3-5 relevant documents each is enough to see meaningful patterns. Measure at least one retrieval metric (precision@k, recall@k, or MRR). + +### 4. Brief Write-Up + +Document your decisions: + +- Why you chose your dataset +- Whether you used ColModernVBERT or ColPali, and why +- What optimization technique(s) you applied and why +- What worked well and what surprised you +- Key metrics from your evaluation + +--- + +## Suggested Approach + +These hints are optional. Feel free to chart your own path. + +**Start simple.** Get a basic multi-vector search working before adding optimizations. It's easier to measure the impact of changes when you have a baseline. + +**Create ground truth early.** Before optimizing, write your evaluation queries and identify relevant documents. This lets you measure whether changes improve or hurt quality. + +**Compare configurations.** Try at least two different setups (e.g., with and without quantization, or different pooling strategies). The comparison will teach you more than a single configuration. + +**Keep notes as you go.** Document what you try and what happens. Your future self will thank you, and it makes the write-up easier. + +--- + +## Share Your Results + +We'd love to see what you build. Share your project on the [Qdrant Discord](https://discord.gg/qdrant) in the [#course-submissions](https://discord.com/channels/907569970500743200/1429673887590776832) channel. + +Tell us about: + +- **Your dataset** and why you chose it +- **Your model choice** (ColModernVBERT or ColPali, and why) +- **Your collection configuration** (quantization, pooling, indexing) +- **Your retrieval metrics** (precision@10, recall, latency) +- **What you learned** and what surprised you + +Seeing how others approached the same challenge is one of the best ways to deepen your understanding. + +--- + +## What You've Accomplished + +By completing this project, you've built a production-ready multi-vector search system that: + +* **Leverages late interaction** for superior search quality through token-level matching +* **Scales efficiently** using quantization, pooling, or optimized indexing +* **Delivers fast results** through multi-stage retrieval and HNSW optimization +* **Measures quality** with comprehensive evaluation metrics +* **Documents trade-offs** between accuracy, speed, and memory + +You've mastered the full pipeline from multi-vector embeddings to production optimization - skills directly applicable to enterprise search, document management, research platforms, and AI-powered knowledge systems. + +**Congratulations on completing the Multi-Vector Search course.** + +Continue your learning journey by exploring advanced topics like fine-tuning late interaction models for domain-specific documents, or building Retrieval Augmented Generation pipelines to derive insights from your retrieved documents. diff --git a/qdrant-landing/content/course/multi-vector-search/module-3/multi-stage-retrieval.md b/qdrant-landing/content/course/multi-vector-search/module-3/multi-stage-retrieval.md new file mode 100644 index 000000000..9a4f272ae --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-3/multi-stage-retrieval.md @@ -0,0 +1,334 @@ +--- +title: "Multi-Stage Retrieval with Universal Query API" +description: Combine multiple optimization techniques in multi-stage retrieval pipelines using Qdrant's Universal Query API. +weight: 1 +isLesson: true +--- + +{{< date >}} Module 3 {{< /date >}} + +# Multi-Stage Retrieval with Universal Query API + +The most effective production deployments combine multiple optimization techniques in multi-stage pipelines. Fast approximate methods retrieve candidates, which are then reranked with higher-quality methods. + +Qdrant's Universal Query API makes it easy to build sophisticated multi-stage retrieval systems. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## Why Multi-Stage Retrieval? + +You've learned that multi-vector representations like ColBERT provide superior search quality compared to single-vector embeddings. But there's a challenge: **computing MaxSim for every document in a large collection is expensive**. + +Here's the dilemma: single-vector models are fast but less accurate, while multi-vector models are accurate but computationally intensive. What if you could combine the strengths of both? + +Multi-stage retrieval offers an elegant solution: use a fast method to narrow down candidates, then apply a high-quality method to rerank only those candidates. + +## The Multi-Stage Retrieval Pattern + +The key insight is that you don't need to use your most expensive model on every document in your collection. Instead, you can split the search into stages: + +1. **Stage 1 (Prefetch)**: Use a fast single-vector embedding model to retrieve a large set of candidates +2. **Stage 2 (Rerank)**: Use ColBERT's multi-vector representations to rerank only those candidates + +This pattern dramatically reduces computational cost while maintaining high search quality. + +![Multi-Stage Retrieval](/courses/multi-vector-search/module-3/multi-stage-retrieval.png) + +## The Critical Role of Oversampling + +Here's the most important concept in multi-stage retrieval: **you must retrieve more candidates in the prefetch stage than you want in your final results**. + +This is called **oversampling**, and it's essential for maintaining search quality. + +### Why Oversample? + +Consider what happens if you retrieve exactly 10 results in both stages: + +- **Stage 1**: Single-vector model retrieves its "top 10" documents +- **Stage 2**: ColBERT reranks those same 10 documents + +The problem? You're limited to ColBERT reranking only the 10 documents the single-vector model selected. If the truly best document is ranked 11th by the single-vector model, ColBERT will never see it. + +By oversampling - retrieving 100, 500, or even 1000 candidates in the prefetch stage - you give ColBERT a much larger pool to work with. This dramatically increases the chance that the truly best documents are in the candidate set. + +### Choosing the Oversampling Factor + +The oversampling factor (how many candidates to retrieve vs. how many final results you want) is a trade-off: + +- **Higher oversampling** (e.g., retrieve 1000 to return 10): + - Better final search quality + - Higher computational cost in Stage 2 + +- **Lower oversampling** (e.g., retrieve 50 to return 10): + - Faster overall query time + - Lower computational cost + - Risk of missing relevant documents + +![Oversampling impact](/courses/multi-vector-search/module-3/oversampling-impact.png) + +## Implementing Multi-Stage Retrieval in Qdrant + +Qdrant's Universal Query API makes multi-stage retrieval straightforward using the `prefetch` parameter. The pattern is simple: whenever a query has at least one prefetch, Qdrant: + +1. Performs the prefetch query (or queries) +2. Applies the main query over the results of the prefetch + +### Basic Example: Single-Vector to ColBERT + +Let's say you want to search for documents about "quantum computing applications" and return the top 10 results. + +First, create a collection with both single-vector and multi-vector representations. We'll also add sparse vectors for BM25 to enable hybrid search later: + +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient("http://localhost:6333") + +# Create collection with both single-vector and multi-vector representations +client.create_collection( + collection_name="hybrid-search", + vectors_config={ + # Fast single-vector for prefetch stage + "bge-small-en-v1.5": models.VectorParams( + size=384, + distance=models.Distance.COSINE, + ), + # High-quality multi-vector for reranking stage + "colbert": models.VectorParams( + size=128, + distance=models.Distance.DOT, + multivector_config=models.MultiVectorConfig( + comparator=models.MultiVectorComparator.MAX_SIM, + ), + hnsw_config=models.HnswConfigDiff(m=0), + ), + }, + sparse_vectors_config={ + "bm25": models.SparseVectorParams( + modifier=models.Modifier.IDF, + ), + }, +) +``` + +Next, ingest your documents. Using `models.Document`, Qdrant handles the embedding automatically on the server side: + +```python +documents = [ + ("research", "Quantum computing applications are emerging in cryptography..."), + ("research", "Researchers are exploring quantum computing applications in drug discovery..."), + ("finance", "Quantum computing applications in finance include portfolio optimization..."), + # ... more documents +] + +# Ingest data to the collection - Qdrant embeds automatically +client.upsert( + collection_name="hybrid-search", + points=[ + models.PointStruct( + id=i, + vector={ + "bge-small-en-v1.5": models.Document( + text=doc, + model="BAAI/bge-small-en-v1.5", + ), + "colbert": models.Document( + text=doc, + model="colbert-ir/colbertv2.0", + ), + "bm25": models.Document( + text=doc, + model="Qdrant/bm25", + ), + }, + payload={ + "text": doc, + "category": category, + }, + ) + for i, (category, doc) in enumerate(documents) + ] +) +``` + +Now, here's how you structure the multi-stage query. Notice that we use `models.Document` for the query as well - Qdrant embeds it server-side using the appropriate model for each stage: + +```python +query = "quantum computing applications" + +# Multi-stage query: prefetch with single-vector, rerank with ColBERT +results = client.query_points( + collection_name="hybrid-search", + prefetch=[ + models.Prefetch( + query=models.Document( + text=query, + model="BAAI/bge-small-en-v1.5", + ), + using="bge-small-en-v1.5", + limit=500, # Retrieve 500 candidates for reranking + ), + ], + query=models.Document( + text=query, + model="colbert-ir/colbertv2.0", + ), + using="colbert", + limit=10, # Return top 10 after reranking +) +``` + +### Understanding the Query Structure + +The key parts of a multi-stage query: + +- **`prefetch`**: Defines the first stage search + - Uses `models.Document` to specify the text and embedding model + - The `using` parameter specifies which named vector to search + - Has its own `limit` parameter (this is your oversampling size) + - Quickly narrows down the candidate set + +- **Main query**: Defines the reranking stage + - Uses `models.Document` with the ColBERT model for multi-vector embedding + - The `using` parameter specifies the multi-vector named vector (e.g., `"colbert"`) + - Its `limit` parameter determines final result count + - Only runs on the candidates from prefetch + +Using `models.Document` lets Qdrant handle all embedding with FastEmbed, simplifying your client code and ensuring consistency between indexing and querying. + +## Advanced Multi-Stage Patterns + +Multi-stage retrieval isn't limited to just two stages. You can chain multiple prefetch operations for three or more stages by nesting prefetch operations. + +### Combining Multiple Weak Retrievers + +You can also use multiple retrieval methods in the prefetch stage. For example, you might combine both dense and sparse vectors using query fusion to create a stronger initial candidate set, then rerank with ColBERT. + +This hybrid approach in the prefetch stage can improve recall - ensuring that the candidate pool contains relevant documents that might be missed by either dense or sparse search alone. Since we configured BM25 sparse vectors when creating the collection, we can use them directly in the prefetch: + +```python +# Multi-stage with hybrid prefetch: combine dense and sparse retrieval +results = client.query_points( + collection_name="hybrid-search", + prefetch=[ + # Dense retrieval using single-vector embeddings + models.Prefetch( + query=models.Document( + text=query, + model="BAAI/bge-small-en-v1.5", + ), + using="bge-small-en-v1.5", + limit=500, + ), + # Sparse retrieval using BM25 + models.Prefetch( + query=models.Document( + text=query, + model="Qdrant/bm25", + ), + using="bm25", + limit=500, + ), + ], + # Results from both prefetch queries are combined, then reranked + query=models.Document( + text=query, + model="colbert-ir/colbertv2.0", + ), + using="colbert", + limit=10, +) +``` + +Multi-stage retrieval works seamlessly with filters. An important behavior to understand: **filters in the main query are automatically propagated to all prefetch stages**. This means when you add a filter to your main query, it applies to the entire multi-stage pipeline. This is efficient because it narrows the candidate pool early in the prefetch stage, reducing computational cost throughout. + +```python +# Filters in the main query automatically propagate to prefetch stages +results = client.query_points( + collection_name="hybrid-search", + prefetch=[ + models.Prefetch( + query=models.Document( + text=query, + model="BAAI/bge-small-en-v1.5", + ), + using="bge-small-en-v1.5", + limit=500, + ), + ], + query=models.Document( + text=query, + model="colbert-ir/colbertv2.0", + ), + using="colbert", + limit=10, + # This filter applies to BOTH prefetch and reranking stages + query_filter=models.Filter( + must=[ + models.FieldCondition( + key="category", + match=models.MatchValue(value="research"), + ), + ], + ), +) +``` + +## When to Use Multi-Stage Retrieval + +Multi-stage retrieval is valuable for **large collections** (100K+ documents) where you need **fast queries** while maintaining **multi-vector quality**. It's particularly effective when combining different embedding models or optimizing compute costs. + +Skip multi-stage retrieval for small collections (< 10K documents), scenarios where single-vector embeddings suffice, or real-time indexing where maintaining dual embeddings adds excessive overhead. + +## Performance Characteristics + +Multi-stage retrieval's key advantage is **reducing the number of documents the multi-vector model scans**. Since late interaction performs full scans, limiting candidates dramatically improves performance. + +**Direct late interaction**: 1M documents = 1M MaxSim calculations +**Multi-stage**: 1M documents -> prefetch 500 candidates = 500 MaxSim calculations + +Speedup is roughly **size of the collection / number of candidates**: +- 1M documents, prefetch 1000 -> ~1000x fewer calculations +- 100K documents, prefetch 500 -> ~200x fewer calculations + +Higher oversampling improves quality but increases Stage 2 cost + +## Bringing It All Together + +Multi-stage retrieval is a powerful optimization technique, but it's just one tool in your arsenal. Real-world production systems often combine multiple techniques: + +- **Multi-stage retrieval** for computational efficiency +- **Score boosting** to adjust relevance based on metadata +- **Diversification** to reduce redundancy in results +- **Filtering** to narrow results by business rules +- **Query fusion** to combine multiple search strategies + +The beauty of Qdrant's Universal Query API is that all these techniques can be composed together in a single query. You might use multi-stage retrieval with oversampling, apply filters at each stage, boost scores based on recency, and diversify final results - all in one request. + +However, **there's no one-size-fits-all solution**. The optimal search pipeline depends on your specific data characteristics, quality requirements, latency constraints, and computational budget. This is why **evaluation is critical**. You need to systematically measure and compare different pipeline configurations to find what works best for your use case. + +In the final lesson of this module, you'll learn exactly how to evaluate and compare different search pipelines - giving you the tools to make informed, data-driven decisions about your search architecture. + +## What's Next + +You've learned how to build sophisticated multi-stage retrieval pipelines that combine different techniques for optimal results. But multi-vector representations can be memory-intensive, especially at scale. + +In the next lesson, you'll discover **vector quantization techniques** - powerful compression methods that can reduce memory usage by 4-32x while maintaining search quality. diff --git a/qdrant-landing/content/course/multi-vector-search/module-3/muvera.md b/qdrant-landing/content/course/multi-vector-search/module-3/muvera.md new file mode 100644 index 000000000..b33ee70c1 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-3/muvera.md @@ -0,0 +1,199 @@ +--- +title: "MUVERA" +description: Understand MUVERA and how it enables HNSW indexing for multi-vector search despite MaxSim asymmetry. +weight: 4 +isLesson: true +--- + +{{< date >}} Module 3 {{< /date >}} + +# MUVERA + +MUVERA (Multi-Vector Retrieval with Approximation) solves a fundamental problem: MaxSim's asymmetry makes traditional indexing methods like HNSW ineffective. MUVERA enables fast approximate search for multi-vector representations. + +Understanding MUVERA is key to scaling multi-vector search to millions of documents. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## The HNSW Incompatibility Problem + +Traditional vector indexes like HNSW are designed for single-vector search with symmetric distance metrics. Multi-vector representations break this assumption: **MaxSim is inherently asymmetric and non-metric**. + +When comparing query tokens to document tokens, the direction matters - searching documents with a query gives different results than searching queries with a document. This asymmetry makes HNSW's graph-based navigation ineffective, forcing us back to full scans across millions of documents. + +## MUVERA: Making Multi-Vector Search Fast + +MUVERA (Multi-Vector Retrieval Algorithm) solves this incompatibility by creating a **single approximation vector** for each document that HNSW can efficiently index. The algorithm works in three stages: + +1. **SimHash Clustering**: Groups token vectors into spatial regions using random hyperplanes +2. **Fixed Dimensional Encoding (FDE)**: Aggregates clustered vectors into a single representative vector per document +3. **Dimensionality Reduction**: Applies random projection to create compact, robust representations + +![MUVERA high-level](/courses/multi-vector-search/module-3/muvera-high-level.png) + +The resulting single vector approximates the multi-vector representation well enough for fast retrieval, achieving massive speedup. A reranking step with the original multi-vector data then aims to recover full accuracy. + +For a deep dive into MUVERA's architecture and mathematical foundations, see our detailed article: [MUVERA: Making Multivectors More Performant](/articles/muvera-embeddings/). + +## Applying MUVERA Multi-Stage Retrieval + +FastEmbed provides built-in support for MUVERA postprocessing, making it straightforward to implement this optimization. Let's see how to apply MUVERA to multi-vector embeddings for fast retrieval with multi-stage reranking. + +### Setting Up MUVERA with ColModernVBERT + +FastEmbed 0.7.2+ includes MUVERA as a postprocessing technique that transforms variable-length multi-vector sequences into fixed-dimensional single vectors. You can apply MUVERA to any late interaction model, including **ColModernVBERT**. + +The workflow involves loading a multi-vector embedding model and wrapping it with a MUVERA processor: + +```python +from fastembed import LateInteractionMultimodalEmbedding +from fastembed.postprocess import Muvera + +# Load ColModernVBERT model +model = LateInteractionMultimodalEmbedding(model_name="Qdrant/colmodernvbert") + +# Wrap with MUVERA processor +muvera = Muvera.from_multivector_model( + model=model, + k_sim=6, # 2^6 = 64 similarity buckets + dim_proj=32, # Projection dimensionality + r_reps=20 # Random projection repetitions +) +``` + +The MUVERA processor accepts several key parameters that control the speed-accuracy tradeoff: + +- **`k_sim`**: Number of similarity buckets for SimHash clustering (more buckets = higher precision) +- **`dim_proj`**: Target dimensionality for the projection (lower = faster search, higher = better accuracy) +- **`r_reps`**: Number of random projections for robust encoding (higher = more stable representations) + +### Creating the Collection + +For multi-stage retrieval to work, you need to index **both** the MUVERA approximation vectors and the original multi-vector representations in Qdrant. The MUVERA vectors enable fast HNSW retrieval, while the multi-vector representations provide accurate reranking. + +First, create a collection with both vector types: + +```python +from qdrant_client import QdrantClient, models + +# Create collection with both vector types +client = QdrantClient("http://localhost:6333") +client.create_collection( + collection_name="documents-muvera", + vectors_config={ + "muvera": models.VectorParams( + size=muvera.embedding_size, + distance=models.Distance.COSINE + ), + "colmodernvbert": models.VectorParams( + size=model.embedding_size, + distance=models.Distance.COSINE, + multivector_config=models.MultiVectorConfig( + comparator=models.MultiVectorComparator.MAX_SIM + ) + ) + } +) +``` + +The dual-vector approach stores: +1. **MUVERA embeddings**: Single vector per document, fast HNSW retrieval +2. **Multi-vector representation**: Full token sequences per document, precise MaxSim scoring + +### Embedding and Indexing Documents + +Next, embed your documents and upload both representations. Generate multi-vector embeddings with ColModernVBERT, then process them through MUVERA: + +```python +# Embed documents +image_path = "images/financial-report.png" +doc_multivec = list(model.embed_image([image_path]))[0] +doc_muvera = muvera.process_document(doc_multivec) +``` + +Upload both the MUVERA approximation and the original multi-vector representation: + +```python +# Upload both representations +client.upsert( + collection_name="documents-muvera", + points=[ + models.PointStruct( + id=0, + payload={"source": image_path}, + vector={"muvera": doc_muvera, "colmodernvbert": doc_multivec} + ) + ] +) +``` + +![MUVERA search](/courses/multi-vector-search/module-3/muvera-search.png) + +### Querying with Multi-Stage Retrieval + +At query time, you first retrieve candidates using the fast MUVERA vectors, then rescore those candidates using the original multi-vector representations. This hybrid approach maintains search quality while dramatically reducing compute requirements. + +First, encode the query in both formats: + +```python +# Encode query in both formats +query = "quarterly revenue growth" +query_multivec = list(model.embed_text(query))[0] +query_muvera = muvera.process_query(query_multivec) +``` + +Then perform two-stage retrieval using Qdrant's prefetch mechanism: + +```python +# Two-stage retrieval with Qdrant's prefetch +results = client.query_points( + collection_name="documents-muvera", + prefetch=models.Prefetch( + query=query_muvera, + using="muvera", + limit=100, # Stage 1: Fast MUVERA retrieval + ), + query=query_multivec, + using="colmodernvbert", # Stage 2: Precise MaxSim reranking + limit=10, + with_payload=True +) +``` + +This two-stage process achieves near-identical accuracy to full multi-vector search across all documents, but only computes expensive MaxSim operations for a small candidate set. + +For complete implementation details and parameter tuning guidance, see the [FastEmbed Postprocessing documentation](/documentation/fastembed/fastembed-postprocessing/). + +### Trade-offs and Considerations + +**Storage Requirements**: MUVERA requires storing both representations - the single approximation vector and the full multi-vector sequence. This doubles storage compared to single-vector search, but remains practical for production systems. Offloading original vectors to disk might be a solution, if you can afford using that much memory. + +**Speed Gains**: The performance improvement is substantial. Instead of computing MaxSim against millions of documents, you only compute it for your candidate set (typically 100-1000 documents). The MUVERA-powered HNSW retrieval scales logarithmically, making multi-vector search viable at scale. + +**Accuracy Preservation**: Research shows MUVERA maintains nearly the same accuracy as full multi-vector search when properly configured. The key is choosing appropriate parameter values based on your dataset and quality requirements. + +## What's Next + +You've learned why MaxSim's asymmetry prevents HNSW from efficiently indexing multi-vector representations, and how MUVERA solves this with single-vector approximations. By combining MUVERA's fast retrieval with multi-vector reranking, you can achieve significant speedups while maintaining search quality. + +The optimization techniques covered in this module - quantization, pooling, and MUVERA - are **complementary techniques** that can be combined for maximum efficiency. For example, you can test quantization on both your MUVERA vectors and your stored multi-vector representations, reducing memory footprint while maintaining the speed benefits of HNSW indexing. Or keep MUVERA vectors unchanged and play with quantization and pooling for the original vectors. These optimizations stack together, allowing you to build highly efficient production pipelines. + +Now that you have multiple optimization tools in your toolkit, how do you evaluate which combinations work best for your use case? Let's learn how to measure and compare different search pipeline configurations. \ No newline at end of file diff --git a/qdrant-landing/content/course/multi-vector-search/module-3/pooling-techniques.md b/qdrant-landing/content/course/multi-vector-search/module-3/pooling-techniques.md new file mode 100644 index 000000000..1388f1340 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-3/pooling-techniques.md @@ -0,0 +1,209 @@ +--- +title: "Pooling Techniques" +description: Reduce the number of vectors per document using row/column pooling and hierarchical token pooling strategies. +weight: 3 +isLesson: true +--- + +{{< date >}} Module 3 {{< /date >}} + +# Pooling Techniques + +While quantization reduces the size of each vector, pooling reduces the number of vectors per document. By intelligently combining token embeddings, you can achieve significant memory savings while preserving retrieval quality. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## Pooling in Embedding Models + +Pooling isn't new to vector search - it's fundamental to how most embedding models work. When you encode text with models like Sentence Transformers, the model first generates embeddings for each token in your input. But to create a single vector representing the entire text, the model must **pool** these token embeddings together. + +Common pooling strategies in dense embedding models include: + +- **Mean pooling**: Average all token embeddings into a single vector +- **CLS token pooling**: Use the special `[CLS]` token's embedding as the document representation +- **Max pooling**: Take the maximum value for each dimension across all tokens +- **Weighted pooling**: Assign different importance to different tokens (e.g., using attention weights) + +These techniques compress variable-length sequences of token embeddings into fixed-size vectors, making them compatible with traditional vector search systems. + +With multi-vector representations, we face a similar but more nuanced challenge. Instead of reducing tokens to a single vector upfront, we maintain multiple vectors per document to preserve richer semantic information. However, as you learned in the previous lessons, this creates memory and performance challenges. **Pooling techniques for multi-vector search** let you strategically reduce the number of vectors while retaining the benefits of late interaction. + +**Important:** Pooling is typically applied only to **document embeddings**, not queries. Why? Queries are usually short (a few tokens), so there's little memory to save. More importantly, we want to preserve full query resolution - every query token should have the opportunity to find its best match among document tokens. The memory savings come from compressing the large document collection, not the ephemeral query vectors. + +## Pooling for Multi-Vector Representations + +### Image-Specific Methods + +For visual document representations like ColPali, spatial relationships in the patch grid enable effective pooling strategies. As you learned in Module 2's visual interpretability lesson, patches in the same row or column often capture semantically related content - a row might contain a line of text, while a column might capture a vertical element like a table border or sidebar. + +**Row pooling** groups patches by their horizontal position: + +1. Organize the 1024 patch embeddings into a 32×32 grid +2. Apply mean pooling across each row (combining 32 patches) +3. Result: 32 vectors instead of 1024 + +Mathematically: + +

$$\text{RowPool}_i = \text{Mean}(\lbrace p_{i,j} : j \in [0, 31] \rbrace)$$

+ +Where $p_{i,j}$ is the patch embedding at row $i$, column $j$. + +**Column pooling** works similarly but along the vertical axis: + +

$$\text{ColPool}_j = \text{Mean}(\lbrace p_{i,j} : i \in [0, 31] \rbrace)$$

+ +This also produces 32 vectors, but captures vertical content relationships instead. + +![Row/column pooling](/courses/multi-vector-search/module-3/row-column-pooling.png) + +**Memory savings** are substantial. FastEmbed returns embeddings in float16 format by default, which already halves the memory compared to float32: + +| Representation | Vectors | Memory (float16) | Memory (float32) | +|----------------|---------|------------------|------------------| +| Full patches | 1024 | 256 KB | 512 KB | +| Row pooling | 32 | 8 KB | 16 KB | +| Column pooling | 32 | 8 KB | 16 KB | + +That's a **32× reduction** in vector count and memory footprint. + +**Trade-offs to consider:** + +- **Loss of fine-grained resolution**: Small details that span partial rows may blend together +- **Row pooling** may work better for horizontally-oriented content, like text +- **Column pooling** may better capture vertical structures like tables, sidebars, or vertically-oriented text +- You can combine both (64 vectors) for a balanced approach + +```python +import numpy as np +from fastembed import LateInteractionMultimodalEmbedding + +# Load ColPali model +model = LateInteractionMultimodalEmbedding(model_name="Qdrant/colpali-v1.3-fp16") + +# Embed a document image (returns ~1030 vectors × 128 dimensions) +image_path = "images/financial-report.png" # Your document image +embeddings = list(model.embed_image([image_path]))[0] +print(f"Original shape: {embeddings.shape}") # (1030, 128) + +# Reshape to spatial grid: (rows, columns, embedding_dim) +# Get only the first 1024 embeddings, as instruction tokens do +# not represent images +grid = embeddings[:1024].reshape(32, 32, 128) + +# Row pooling: average across columns (axis=1) +row_pooled = grid.mean(axis=1) # Shape: (32, 128) + +# Column pooling: average across rows (axis=0) +col_pooled = grid.mean(axis=0) # Shape: (32, 128) + +# Combined approach (optional): concatenate row and column pooled +combined = np.vstack([row_pooled, col_pooled]) # Shape: (64, 128) + +# Memory comparison (FastEmbed uses float16 by default) +original_memory = embeddings.nbytes # 1030 × 128 × 2 = 263,680 bytes +pooled_memory = row_pooled.nbytes # 32 × 128 × 2 = 8,192 bytes + +print(f"Original: {original_memory:,} bytes ({original_memory // 1024} KB)") +print(f"Row pooled: {pooled_memory:,} bytes ({pooled_memory // 1024} KB)") +print(f"Reduction: {original_memory // pooled_memory}×") +``` + +### Generic Methods + +While row/column pooling exploits the spatial structure of image embeddings, **hierarchical token pooling** works for any multi-vector representation - text, images, or hybrid documents. The core idea: instead of grouping by fixed spatial positions, cluster tokens by **semantic similarity**. + +**How hierarchical pooling works:** + +1. Apply k-means clustering to group similar token embeddings +2. Pool within each cluster using mean pooling +3. Output: $k$ vectors instead of $n$ original tokens + +This approach adapts to the content itself. For a document with dense text and sparse images, clustering naturally allocates more representative vectors to the text regions where semantic variation is higher. + +**Key parameters:** + +- **Number of clusters ($k$)**: Controls the compression ratio. $k=32$ gives similar compression to row pooling, while $k=64$ preserves more detail +- **Clustering algorithm**: k-means is fast and effective. Although hierarchical clustering can capture nested semantic structures but adds overhead + +**Comparison: Row/Column vs. Hierarchical Pooling** + +| Aspect | Row/Column Pooling | Hierarchical Pooling | +|-----------------------|-------------------------------------|---------------------------------| +| **Works with** | Images only (requires spatial grid) | Any multi-vector representation | +| **Grouping strategy** | Fixed spatial positions | Semantic similarity | +| **Compression ratio** | Fixed (32×) | Configurable via $k$ | +| **Indexing overhead** | None | Clustering computation | +| **Preserves** | Spatial structure | Semantic diversity | + +**Trade-offs:** + +- **Higher indexing cost**: Clustering adds computational overhead during document encoding +- **Content-adaptive**: Allocates representation capacity where semantic variation is highest +- **Loses spatial interpretability**: Unlike row pooling, you can't easily map pooled vectors back to document regions +- **Hyperparameter sensitivity**: The choice of $k$ affects retrieval quality and must be tuned + +```python +from scipy.cluster.vq import kmeans2 + +# Embed a document image +image_path = "images/financial-report.png" +embeddings = list(model.embed_image([image_path]))[0] + +def hierarchical_pool(embeddings: np.ndarray, k: int) -> np.ndarray: + """Pool embeddings using k-means clustering.""" + # Cluster embeddings into k groups + centroids, labels = kmeans2(embeddings, k, minit='++') + + # Pool within each cluster using mean + pooled = np.array([ + embeddings[labels == i].mean(axis=0) + for i in range(k) + ]) + return pooled + +# Compare different compression levels +for k in [16, 32, 64, 128]: + pooled = hierarchical_pool(embeddings, k) + reduction = len(embeddings) / k + print(f"k={k:3d}: {len(embeddings)} → {k} vectors ({reduction:.0f}× reduction)") +``` + +## What's Next + +This lesson covered two complementary strategies for reducing the number of vectors per document: + +- **Row/column pooling**: Exploits spatial structure in image embeddings for a fixed reduction (32x for ColPali) +- **Hierarchical pooling**: Content-adaptive clustering that works for any multi-vector representation + +Combined with quantization from the previous lesson, you can achieve dramatic memory savings: + +| Technique | Memory per Document | +|-----------------------------------|---------------------| +| Baseline (1024 vectors × float32) | 512 KB | +| Row pooling only | 16 KB | +| Row pooling + scalar quantization | 4 KB | +| Row pooling + binary quantization | 512 bytes | + +That's a **1000× reduction** from baseline to the most aggressive combination - making multi-vector search practical even for large document collections. + +However, there's still one challenge we haven't addressed: **indexing**. Even with pooled representations, we're still performing brute-force MaxSim comparisons. For millions of documents, this becomes a bottleneck. + +In the next lesson, you'll learn about **MUVERA** - a technique that enables HNSW indexing for multi-vector representations, unlocking fast approximate search at scale. diff --git a/qdrant-landing/content/course/multi-vector-search/module-3/quantization-techniques.md b/qdrant-landing/content/course/multi-vector-search/module-3/quantization-techniques.md new file mode 100644 index 000000000..aa9abc505 --- /dev/null +++ b/qdrant-landing/content/course/multi-vector-search/module-3/quantization-techniques.md @@ -0,0 +1,336 @@ +--- +title: "Vector Quantization Techniques" +description: Learn how to reduce memory usage with scalar quantization, binary quantization, and other compression methods. +weight: 2 +isLesson: true +--- + +{{< date >}} Module 3 {{< /date >}} + +# Vector Quantization Techniques + +Vector quantization compresses vectors by reducing the precision of each component. Qdrant supports several quantization methods that can reduce memory usage by 4-64x, sometimes with minimal quality loss. + +Choosing the right quantization method depends on your quality requirements and memory constraints. + +--- + +
+ +
+ +--- + +**Follow along in Colab:** + Open In Colab + + +--- + +## The Memory Challenge with Multi-Vector Models + +By default, embedding models produce vectors with **float32 precision** - each component uses 32 bits (4 bytes) of memory. For single-vector embeddings, this is manageable. But multi-vector models like **ColModernVBERT** change the equation dramatically. + +Consider a typical ColPali scenario using **ColModernVBERT**: +- **~1024 vectors per document** (one per visual patch) +- **128 dimensions per vector** (model embedding size) +- **float32 precision** (4 bytes per component) + +Let's calculate the memory for a single document: + +$$ +\text{Memory per document} = 1024 \text{ vectors} \times 128 \text{ dims} \times 4 \text{ bytes} = 524{,}288 \text{ bytes} = 512 \text{ KB} +$$ + +For a collection of **1 million documents**: + +$$ +\text{Total memory} = 1{,}000{,}000 \times 512 \text{ KB} = 512 \text{ GB} +$$ + +Compare this to a traditional single-vector model (e.g., 768-dimensional): + +$$ +\text{Single-vector memory} = 768 \text{ dims} \times 4 \text{ bytes} = 3{,}072 \text{ bytes} = 3 \text{ KB per document} +$$ + +**Multi-vector representations use ~170x more memory** than single-vector models for the same number of documents. This is where quantization becomes essential. + +## Quantization: Compressing Without Losing Quality + +**Vector quantization** reduces memory by representing vectors with fewer bits while preserving the relative distances between them. Qdrant supports several quantization methods optimized for different scenarios. + +**Important: What Quantization Does (and Doesn't) Do** + +Quantization is a **memory optimization technique**, not an indexing solution: +- ✅ **Reduces memory footprint** by 4-32x for storing multi-vector representations +- ✅ **Reduces infrastructure costs** by requiring less RAM +- ✅ **Provides speed improvements** through SIMD operations and smaller data transfers +- ❌ **Does NOT enable HNSW indexing** for multi-vector search - brute force scan is still required + +Multi-vector search with MaxSim fundamentally requires comparing query tokens against all document tokens. HNSW and other graph-based indexes cannot efficiently navigate this token-level comparison space. Quantization makes the brute force search faster and cheaper, but the search strategy remains exhaustive. + +The key insight: **you don't need perfect precision to find the right matches**. If document A is closer to a query than document B in full precision, it usually remains closer after quantization. + +### Scalar Quantization: The Reliable Default + +**Scalar quantization** converts float32 values to 8-bit integers (uint8), reducing memory by **4x**. + +**How it works:** +1. Find the min and max values across all vector components +2. Map the float range to [0, 255] +3. Store the scaling parameters for reconstruction + +For our ColModernVBERT example: + +$$ +\text{Quantized memory} = \frac{512 \text{ GB}}{4} = 128 \text{ GB} +$$ + +**Benefits:** +- **4x memory reduction** with <1% accuracy loss +- **Up to 2x faster brute force search** via SIMD optimization (still exhaustive, but more efficient) +- **Lower infrastructure costs** by reducing RAM requirements +- Works universally across all vector types and dimensions + +**Configuration parameter:** +- `quantile`: Excludes outliers (e.g., 0.99 excludes 1% of extreme values for better scaling) + +### Binary Quantization: Maximum Compression + +**Binary quantization** represents each component as a single bit (positive/negative), achieving **32x compression**. Qdrant also supports **1.5-bit** and **2-bit** variants for better accuracy with moderate compression. + +For ColModernVBERT vectors: + +**1-bit binary quantization:** + +$$ +\text{Memory per document} = 1024 \text{ vectors} \times 128 \text{ dims} \times \frac{1}{8} \text{ bytes} = 16{,}384 \text{ bytes} = 16 \text{ KB} +$$ + +$$ +\text{Total for 1M docs} = \frac{512 \text{ GB}}{32} = 16 \text{ GB} +$$ + +**1.5-bit binary quantization:** + +$$ +\text{Total for 1M docs} = \frac{512 \text{ GB}}{24} \approx 21.3 \text{ GB} +$$ + +**2-bit binary quantization:** + +$$ +\text{Total for 1M docs} = \frac{512 \text{ GB}}{16} = 32 \text{ GB} +$$ + +**Typical dimension ranges for binary quantization:** +- **1-bit**: Often used with high-dimensional vectors (1536+ dimensions) +- **1.5-bit**: Commonly applied to 1024-1536 dimensions +- **2-bit**: Frequently used with 768-1024 dimensions + +For ColModernVBERT's **128 dimensions**, binary quantization presents unique challenges. With such low dimensionality, each bit of precision has a larger impact on the representation. The choice between scalar and binary quantization - and which binary variant to use - depends on your specific use case and quality requirements. We'll explore how to evaluate these trade-offs systematically in the final lesson of this module. + +### Real-World Impact for ColPali Collections + +Let's compare all options for a **1 million document** ColModernVBERT collection: + +| Method | Memory | Compression | Speed Boost | +|----------------------|---------|---------------|-------------| +| **No quantization** | 512 GB | 1x (baseline) | 1x | +| **Scalar (int8)** | 128 GB | 4x | ~2x | +| **Binary (2-bit)** | 32 GB | 16x | ~20x | +| **Binary (1.5-bit)** | 21.3 GB | 24x | ~30x | +| **Binary (1-bit)** | 16 GB | 32x | ~40x | + + + +The table above shows the theoretical memory savings and search performance characteristics for each quantization method. The actual impact on retrieval quality is **not included** because it varies significantly based on your specific documents, queries, and quality requirements. + +**The choice between these methods requires systematic evaluation** of your complete search pipeline. We'll cover evaluation methodologies in detail in the final lesson of this module, where you'll learn how to measure the impact of quantization on your specific use case. + +## Enabling Quantization in Qdrant + +One of Qdrant's powerful features: **you can enable quantization on an existing collection** without changing your inference or ingestion pipelines. The quantization happens transparently during indexing. + + +First, let's load the ColPali model and prepare some sample documents: + +```python +from fastembed import LateInteractionMultimodalEmbedding + +# Load ColPali model for generating multi-vector embeddings +model = LateInteractionMultimodalEmbedding( + model_name="Qdrant/colpali-v1.3-fp16" +) + +# Sample document images and metadata +image_paths = [ + "images/financial-report.png", + "images/titanic-newspaper.jpg", + "images/moon-landing.jpg", + "images/einstein-newspaper.jpg", +] + +documents = [ + {"title": "Financial Report", "type": "report", "topic": "finance"}, + {"title": "Titanic Sinking", "type": "newspaper", "topic": "history"}, + {"title": "Moon Landing", "type": "newspaper", "topic": "space"}, + {"title": "Einstein Theory", "type": "newspaper", "topic": "science"}, +] + +# Generate embeddings for all images +image_embeddings = list(model.embed_image(image_paths)) +``` + +Now create a collection with scalar quantization: + +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient("http://localhost:6333") + +client.create_collection( + collection_name="colpali-scalar", + vectors_config={ + "colpali": models.VectorParams( + size=128, + distance=models.Distance.DOT, + multivector_config=models.MultiVectorConfig( + comparator=models.MultiVectorComparator.MAX_SIM, + ), + hnsw_config=models.HnswConfigDiff(m=0), # Disable HNSW for multi-vector + ), + }, + quantization_config=models.ScalarQuantization( + scalar=models.ScalarQuantizationConfig( + type=models.ScalarType.INT8, + quantile=0.99, # Exclude 1% outliers for better scaling + always_ram=True, + ), + ), +) +``` + +Ingest the documents into the scalar-quantized collection: + +```python +client.upsert( + collection_name="colpali-scalar", + points=[ + models.PointStruct( + id=i, + vector={"colpali": embedding.tolist()}, + payload=documents[i], + ) + for i, embedding in enumerate(image_embeddings) + ], +) +``` + +Enabling a different type of quantization requires setting a different quantization configuration. + +```python +client.create_collection( + collection_name="colpali-binary", + vectors_config={ + "colpali": models.VectorParams( + size=128, + distance=models.Distance.DOT, + multivector_config=models.MultiVectorConfig( + comparator=models.MultiVectorComparator.MAX_SIM, + ), + hnsw_config=models.HnswConfigDiff(m=0), + ), + }, + quantization_config=models.BinaryQuantization( + binary=models.BinaryQuantizationConfig( + always_ram=True, + ), + ), +) + +# Ingest the same data into the binary-quantized collection +client.upsert( + collection_name="colpali-binary", + points=[ + models.PointStruct( + id=i, + vector={"colpali": embedding.tolist()}, + payload=documents[i], + ) + for i, embedding in enumerate(image_embeddings) + ], +) +``` + +## Search-Time Control with Rescoring + +Qdrant provides **automatic rescoring**: the quantized index quickly finds candidates, then re-ranks them using the original float32 vectors for accuracy. + +**Key search parameters:** +- `rescore`: Re-evaluate top candidates with original vectors (default: true) +- `oversampling`: Fetch more candidates before rescoring (e.g., 2.0 = fetch 2x results) + +For ColPali searches: + +```python +# Generate query embeddings from a text query +query = "financial quarterly results revenue" +query_embeddings = list(model.embed_text([query]))[0] + +# Search with rescoring enabled +results = client.query_points( + collection_name="colpali-scalar", + query=query_embeddings.tolist(), + using="colpali", + limit=10, + search_params=models.SearchParams( + quantization=models.QuantizationSearchParams( + ignore=False, # Use quantized vectors for initial search + rescore=True, # Re-rank with original float32 vectors + oversampling=2.0, # Fetch 2x candidates before rescoring + ), + ), +) +``` + +The rescoring step is **critical for multi-vector search** because MaxSim aggregates many token-level similarities - small quantization errors can compound. Rescoring with original vectors ensures your final results maintain high quality. + +## Quantization Impact on Search Quality + +The compression ratios and speed improvements shown above are only part of the story. **The real question is how quantization affects your retrieval quality** - and that answer depends entirely on your specific use case. + +Several factors influence quantization's impact: + +1. **Model dimensionality** - ColModernVBERT's 128-dimensional vectors behave differently under quantization than higher-dimensional models (768+ dims) +2. **Document characteristics** - Text-heavy documents, image-heavy pages, and structured forms each respond differently to compression +3. **Query patterns** - Keyword-like queries vs. semantic questions may show different sensitivity to quantization errors +4. **Quality thresholds** - Your application's tolerance for retrieval quality changes + +Published benchmarks provide general guidance, but **your specific documents and queries will behave differently**. A quantization method that works well for one dataset might perform poorly on another. + +This is why the final lesson in this module focuses entirely on **evaluating multi-vector search pipelines**. You'll learn systematic approaches to measure quantization's impact on your specific use case, helping you make informed trade-offs between memory savings and retrieval quality. + +## What's Next + +In this lesson, you learned how quantization dramatically reduces memory usage and improves brute force search performance for multi-vector representations. Key takeaways: +- ColModernVBERT's memory footprint: **512 KB per document** without quantization +- Scalar quantization reduces this to **128 KB** (4x compression) with ~2x faster brute force search +- Binary quantization can achieve **16-32 GB** total memory for 1 million documents (vs 512 GB uncompressed) with up to 40x faster brute force search +- **Quantization is a memory and speed optimization**, not an indexing solution - HNSW remains incompatible with multi-vector search +- The search strategy remains exhaustive (brute force), but becomes significantly cheaper and faster + +Crucially, **quantization can be enabled on existing collections** without modifying your ingestion or inference code, making it straightforward to experiment with different approaches. + +However, choosing the right quantization method requires measuring its impact on your specific retrieval quality. We'll cover systematic evaluation approaches in the final lesson of this module. + +Next, we'll explore **pooling techniques** that reduce the number of vectors per document - a complementary approach to reducing memory that works alongside quantization. diff --git a/qdrant-landing/content/documentation/_index.md b/qdrant-landing/content/documentation/_index.md index 8758ca8c4..cdf38cd15 100644 --- a/qdrant-landing/content/documentation/_index.md +++ b/qdrant-landing/content/documentation/_index.md @@ -10,7 +10,7 @@ content: linkDescription: Clone this repo now and build a search engine in five minutes. cloudButton: text: Cloud Quickstart - url: /documentation/quickstart-cloud/ + url: /documentation/cloud-quickstart/ localButton: text: Local Quickstart url: /documentation/quickstart/ @@ -31,7 +31,7 @@ content: description: Boost search speed, reduce latency, and improve the accuracy and memory usage of your Qdrant deployment. button: text: Learn More - url: /documentation/guides/optimize/ + url: /documentation/operations/optimize/ cardsPartial: documentation/cards/docs-cards cards: - id: 1 @@ -42,7 +42,7 @@ content: title: Distributed Deployment description: Scale Qdrant beyond a single node and optimize for high availability, fault tolerance, and billion-scale performance. link: - url: /documentation/guides/distributed_deployment/ + url: /documentation/operations/distributed_deployment/ text: Read More - id: 2 tag: Documents @@ -52,7 +52,7 @@ content: title: Multitenancy description: Build vector search apps that serve millions of users. Learn about data isolation, security, and performance tuning. link: - url: /documentation/guides/multiple-partitions/ + url: /documentation/manage-data/multitenancy/ text: Read More - id: 3 tag: Blog @@ -76,7 +76,7 @@ Qdrant is an AI-native vector search and a semantic search engine. You can use i ||| |-:|:-| -|[Cloud Quickstart](/documentation/quickstart-cloud/)|[Local Quickstart](/documentation/quick-start/)| +|[Cloud Quickstart](/documentation/cloud-quickstart/)|[Local Quickstart](/documentation/quickstart/)| ## Ready to start developing? @@ -88,9 +88,9 @@ Qdrant is an AI-native vector search and a semantic search engine. You can use i ## Qdrant's most popular features: |||| |:-|:-|:-| -|[Filterable HNSW](/documentation/filtering/)
Single-stage payload filtering | [Recommendations & Context Search](/documentation/concepts/explore/#explore-the-data)
Exploratory advanced search| [Pure-Vector Hybrid Search](/documentation/hybrid-queries/)
Full text and semantic search in one| -|[Multitenancy](/documentation/guides/multiple-partitions/)
Payload-based partitioning|[Custom Sharding](/documentation/guides/distributed_deployment/#sharding)
For data isolation and distribution|[Role Based Access Control](/documentation/guides/security/?q=jwt#granular-access-control-with-jwt)
Secure JWT-based access | -|[Quantization](/documentation/guides/quantization/)
Compress data for drastic speedups|[Multivector Support](/documentation/concepts/vectors/?q=multivect#multivectors)
For ColBERT late interaction |[Built-in IDF](/documentation/concepts/indexing/?q=inverse+docu#idf-modifier)
Advanced similarity calculation| +|[Filterable HNSW](/documentation/search/filtering/)
Single-stage payload filtering | [Recommendations & Context Search](/documentation/search/explore/#explore-the-data)
Exploratory advanced search| [Pure-Vector Hybrid Search](/documentation/search/hybrid-queries/)
Full text and semantic search in one| +|[Multitenancy](/documentation/manage-data/multitenancy/)
Payload-based partitioning|[Custom Sharding](/documentation/operations/distributed_deployment/#sharding)
For data isolation and distribution|[Role Based Access Control](/documentation/operations/security/?q=jwt#granular-access-control-with-jwt)
Secure JWT-based access | +|[Quantization](/documentation/manage-data/quantization/)
Compress data for drastic speedups|[Multivector Support](/documentation/manage-data/vectors/?q=multivect#multivectors)
For ColBERT late interaction |[Built-in IDF](/documentation/manage-data/indexing/?q=inverse+docu#idf-modifier)
Advanced similarity calculation| ## Developer guidebooks: diff --git a/qdrant-landing/content/documentation/cloud-account-setup.md b/qdrant-landing/content/documentation/cloud-account-setup.md index 115ecf4b9..4a541d0fa 100644 --- a/qdrant-landing/content/documentation/cloud-account-setup.md +++ b/qdrant-landing/content/documentation/cloud-account-setup.md @@ -69,7 +69,7 @@ If you use multiple accounts for different purposes, it is a good idea to give t ### Changing the Account Owner -Every account has one owner. The owner is granted full admin permissions for the account as well as futher unique permissions allowing them to either delete the account or transfer account ownership. +Every account has one owner. The owner is granted full admin permissions for the account as well as further unique permissions allowing them to either delete the account or transfer account ownership. To transfer ownership of an account, as the owner, visit the *Access Management* page. In the actions menu of the user you wish to transfer to, you will find the option 'Make Account Owner' which begins the transfer. @@ -94,6 +94,6 @@ Qdrant Cloud supports Enterprise Single-Sign-On for Premium Tier customers. The * SAML * Azure Active Directory -Enterprise Sign-On is available as an add-on for [Premium Tier](/documentation/cloud/premium/) customers. If you are interested in using SSO, please [contact us](/contact-us/). +Enterprise Sign-On is available as an add-on for [Premium Tier](/documentation/cloud-premium/) customers. If you are interested in using SSO, please [contact us](/contact-us/). diff --git a/qdrant-landing/content/documentation/cloud-api.md b/qdrant-landing/content/documentation/cloud-api.md index 99dae5546..293e42d2f 100644 --- a/qdrant-landing/content/documentation/cloud-api.md +++ b/qdrant-landing/content/documentation/cloud-api.md @@ -19,7 +19,7 @@ To cater to diverse integration needs, the Qdrant Cloud API offers two primary i * **REST/JSON API**: A conventional HTTP/1.1 (and HTTP/2) interface with JSON payloads. This API is provided via a gRPC Gateway, translating RESTful calls into gRPC messages, offering ease of use for web clients, scripts, and broader tool compatibility. You can find the API definitions and generated client libraries in our Qdrant Cloud Public API [GitHub repository](https://github.com/qdrant/qdrant-cloud-public-api). -**Note:** The API is splitted into multiple services to make it easier to use. +**Note:** The API is split into multiple services to make it easier to use. ### Qdrant Cloud API Endpoints @@ -44,4 +44,4 @@ For samples on how to use the API, with a tool like grpcurl, curl or any of the ## Terraform Provider -Qdrant Cloud also provides a Terraform provider to manage your Qdrant Cloud resources. [Learn more](/documentation/infrastructure/terraform/). +Qdrant Cloud also provides a Terraform provider to manage your Qdrant Cloud resources. [Learn more](/documentation/cloud-tools/terraform/). diff --git a/qdrant-landing/content/documentation/cloud-getting-started.md b/qdrant-landing/content/documentation/cloud-getting-started.md index 224c75168..47040499b 100644 --- a/qdrant-landing/content/documentation/cloud-getting-started.md +++ b/qdrant-landing/content/documentation/cloud-getting-started.md @@ -12,15 +12,15 @@ Welcome to Qdrant Managed Cloud! This document contains all the information you ## Prerequisites -Before creating a cluster, make sure you have a Qdrant Cloud account. Detailed instructions for signing up can be found in the [Qdrant Cloud Setup](/documentation/cloud/qdrant-cloud-setup/) guide. Qdrant Cloud supports granular [role-based access control](/documentation/cloud-rbac/). +Before creating a cluster, make sure you have a Qdrant Cloud account. Detailed instructions for signing up can be found in the [Qdrant Cloud Setup](/documentation/cloud-account-setup/) guide. Qdrant Cloud supports granular [role-based access control](/documentation/cloud-rbac/). -You also need to provide [payment details](/documentation/cloud/pricing-payments/). If you have a custom payment agreement, first create your account, then [contact our Support Team](https://support.qdrant.io/) to finalize the setup. +You also need to provide [payment details](/documentation/cloud-pricing-payments/). If you have a custom payment agreement, first create your account, then [contact our Support Team](https://support.qdrant.io/) to finalize the setup. Premium Plan subscribers can enable single sign-on (SSO) for their organizations. To activate SSO, please reach out to the Support Team at [https://support.qdrant.io/](https://support.qdrant.io/) for guidance. ## Cluster Sizing -Before deploying any cluster, consider the resources needed for your specific workload. Our [Capacity Planning guide](/documentation/guides/capacity-planning/) describes how to assess the required CPU, memory, and storage. Additionally, the [Pricing Calculator](https://cloud.qdrant.io/calculator) helps you estimate associated costs based on your projected usage. +Before deploying any cluster, consider the resources needed for your specific workload. Our [Capacity Planning guide](/documentation/operations/capacity-planning/) describes how to assess the required CPU, memory, and storage. Additionally, the [Pricing Calculator](https://cloud.qdrant.io/calculator) helps you estimate associated costs based on your projected usage. ## Creating and Managing Clusters @@ -28,9 +28,9 @@ After setting up your account, you can create a Qdrant Cluster by following the ## Preparing for Production -For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. More details are available in the [Distributed Deployment](/documentation/guides/distributed_deployment/) guide. For more information on how to create a production-ready cluster, see our [Vector Search in Production](/articles/vector-search-production/) article. +For a production-ready environment, consider deploying a multi-node Qdrant cluster (at least three nodes) with replication enabled. More details are available in the [Distributed Deployment](/documentation/operations/distributed_deployment/) guide. For more information on how to create a production-ready cluster, see our [Vector Search in Production](/articles/vector-search-production/) article. -If you are looking to optimize costs, you can reduce memory usage through [Quantization](/documentation/guides/quantization/) or by [offloading vectors to disk](/documentation/concepts/storage/#configuring-memmap-storage). +If you are looking to optimize costs, you can reduce memory usage through [Quantization](/documentation/manage-data/quantization/) or by [offloading vectors to disk](/documentation/manage-data/storage/#configuring-memmap-storage). ## Infrastructure as Code Automation diff --git a/qdrant-landing/content/documentation/cloud-pricing-payments.md b/qdrant-landing/content/documentation/cloud-pricing-payments.md index f62ea2d42..0c71f63a4 100644 --- a/qdrant-landing/content/documentation/cloud-pricing-payments.md +++ b/qdrant-landing/content/documentation/cloud-pricing-payments.md @@ -58,7 +58,7 @@ To subscribe: 2. Select **GCP Marketplace** as the payment method. You will be redirected to the GCP Marketplace listing for Qdrant. 3. Select **Subscribe**. (If you have already subscribed, select **Manage on Provider**.) 4. On the next screen, choose options as required, and select **Subscribe**. -5. On the pop-up window that appers, select **Sign up with Qdrant**. +5. On the pop-up window that appears, select **Sign up with Qdrant**. You will be redirected to the Billing Details screen in the [Qdrant Cloud Console](https://cloud.qdrant.io/). From there you can start to create Qdrant database clusters. diff --git a/qdrant-landing/content/documentation/cloud-quickstart.md b/qdrant-landing/content/documentation/cloud-quickstart.md index 43def9ae5..231135e3d 100644 --- a/qdrant-landing/content/documentation/cloud-quickstart.md +++ b/qdrant-landing/content/documentation/cloud-quickstart.md @@ -7,14 +7,14 @@ aliases: - cloud-quick-start - cloud-quickstart - cloud/quickstart-cloud/ - - /documentation/quickstart-cloud/ + - /documentation/cloud-quickstart/ --- # Quick Start with Qdrant Cloud

-Learn how to set up Qdrant Cloud and perform your first semantic search in just a few minutes. We'll use a sample dataset of menu items embedded with the `sentence-transformers/all-MiniLM-L6-v2` model via [Cloud Inference](/documentation/concepts/inference/). This is one of the free embedding models available on Qdrant Cloud. For a list of the available free and paid models, refer to the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. +Learn how to set up Qdrant Cloud and perform your first semantic search in just a few minutes. We'll use a sample dataset of menu items embedded with the `sentence-transformers/all-MiniLM-L6-v2` model via [Cloud Inference](/documentation/inference/). This is one of the free embedding models available on Qdrant Cloud. For a list of the available free and paid models, refer to the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. ## 1. Create a Cloud Cluster @@ -1036,6 +1036,6 @@ You've just performed semantic search on real menu item data. The query "vegetar ## What's Next? -- Explore [filtering](/documentation/concepts/filtering/) to combine semantic search with structured queries -- Learn about [collections](/documentation/concepts/collections/) and advanced configuration options -- Check out more [examples and tutorials](/documentation/tutorials-overview/) \ No newline at end of file +- Explore [filtering](/documentation/search/filtering/) to combine semantic search with structured queries +- Learn about [collections](/documentation/manage-data/collections/) and advanced configuration options +- Check out more [examples and tutorials](/documentation/tutorials-lp-overview/) \ No newline at end of file diff --git a/qdrant-landing/content/documentation/cloud-rbac/permission-reference.md b/qdrant-landing/content/documentation/cloud-rbac/permission-reference.md index 48001d4eb..24009c82c 100644 --- a/qdrant-landing/content/documentation/cloud-rbac/permission-reference.md +++ b/qdrant-landing/content/documentation/cloud-rbac/permission-reference.md @@ -57,8 +57,8 @@ Permissions for API Keys, backups, clusters, and backup schedules. ### **Cluster Data** | Permission | Description | |------------|------------| -| `read:cluster_data` | View cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-only` JWT access for the cluster.](/documentation/guides/security/) | -| `write:cluster_data` | View and modify cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-write` JWT access for the cluster.](/documentation/guides/security/) | +| `read:cluster_data` | View cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-only` JWT access for the cluster.](/documentation/operations/security/) | +| `write:cluster_data` | View and modify cluster data, used for the Cluster UI button on Cluster Details. [Maps to global `read-write` JWT access for the cluster.](/documentation/operations/security/) | ### **Backup Schedules** | Permission | Description | diff --git a/qdrant-landing/content/documentation/cloud-security.md b/qdrant-landing/content/documentation/cloud-security.md index ed16268ab..81002ab78 100644 --- a/qdrant-landing/content/documentation/cloud-security.md +++ b/qdrant-landing/content/documentation/cloud-security.md @@ -20,13 +20,13 @@ All Qdrant clusters running in Qdrant Managed Cloud are isolated from each other All storage volumes are encrypted at rest. [Premium customers](/documentation/cloud-premium/) can also bring their own encryption keys for storage volumes. -Data in transit is protected with Transport-Layer-Security (TLS). It is possible to restrict the [IP ranges](/cloud/configure-cluster/#client-ip-restrictions) that are allowed to access a cluster. +Data in transit is protected with Transport-Layer-Security (TLS). It is possible to restrict the [IP ranges](/documentation/cloud/configure-cluster/#client-ip-restrictions) that are allowed to access a cluster. Infrastructure access is restricted and audited according to the policies described in our [Trust Center](https://qdrant.to/trust-center). All infrastructure components are regularly patched and updated to ensure up to date security. Access to Qdrant Cloud accounts can be configured with granular [Role-Based Access Control](/documentation/cloud-rbac/). [Premium customers](/documentation/cloud-premium/) can also enable [Single Sign-On (SSO)](/documentation/cloud-account-setup/#enterprise-single-sign-on-sso) with their identity provider. -[API keys](/cloud/authentication/) can be configured with granular access controls and rotated at any time. API Keys are never stored in plaintext. +[API keys](/documentation/cloud-api/#authentication) can be configured with granular access controls and rotated at any time. API Keys are never stored in plaintext. We recommend configuring an expiration date and rotating your API keys regularly as a security best practice. @@ -42,7 +42,7 @@ Qdrant clusters in Hybrid Cloud also run in hardened, unprivileged containers wi ### Private Cloud -In Qdrant Private Cloud, Qdrant clusters run completely isolated and air-gapped within your infrastucture without any connection to the Qdrant Cloud Console. +In Qdrant Private Cloud, Qdrant clusters run completely isolated and air-gapped within your infrastructure without any connection to the Qdrant Cloud Console. Since there is no connection or communication with Qdrant, you are fully responsible for the security of the entire Qdrant Private Cloud installation. This also means that you do not benefit from the integrated management and observability features of Qdrant Managed Cloud and Hybrid Cloud. diff --git a/qdrant-landing/content/documentation/cloud-tools/pulumi.md b/qdrant-landing/content/documentation/cloud-tools/pulumi.md index 78c6c7917..9756329a8 100644 --- a/qdrant-landing/content/documentation/cloud-tools/pulumi.md +++ b/qdrant-landing/content/documentation/cloud-tools/pulumi.md @@ -13,7 +13,7 @@ A Qdrant SDK in any of Pulumi's supported languages can be generated based on th ## Pre-requisites 1. A [Pulumi Installation](https://www.pulumi.com/docs/install/). -2. An [API key](/documentation/qdrant-cloud-api/#authentication-connecting-to-cloud-api) to access the Qdrant cloud API. +2. An [API key](/documentation/cloud-api/#authentication-connecting-to-cloud-api) to access the Qdrant cloud API. ## Setup diff --git a/qdrant-landing/content/documentation/cloud-tools/terraform.md b/qdrant-landing/content/documentation/cloud-tools/terraform.md index a6c8333ad..7ddcd3fff 100644 --- a/qdrant-landing/content/documentation/cloud-tools/terraform.md +++ b/qdrant-landing/content/documentation/cloud-tools/terraform.md @@ -15,7 +15,7 @@ With the [Qdrant Terraform Provider](https://registry.terraform.io/providers/qdr To use the Qdrant Terraform Provider, you'll need: 1. A [Terraform installation](https://developer.hashicorp.com/terraform/install). -2. An [API key](/documentation/qdrant-cloud-api/#authentication-connecting-to-cloud-api) to access the Qdrant cloud API. +2. An [API key](/documentation/cloud-api/#authentication) to access the Qdrant cloud API. ## Example Usage diff --git a/qdrant-landing/content/documentation/cloud/authentication.md b/qdrant-landing/content/documentation/cloud/authentication.md index 85c1ac5ce..fe8577f22 100644 --- a/qdrant-landing/content/documentation/cloud/authentication.md +++ b/qdrant-landing/content/documentation/cloud/authentication.md @@ -7,7 +7,7 @@ weight: 30 This page describes what Database API keys are and shows you how to use the Qdrant Cloud Console to create a Database API key for a cluster. You will learn how to connect to your cluster using the new API key. -Database API keys can be configured with granular access control. Database API keys with granular access control can be recognized by starting with `eyJhb`. Please refer to the [Table of access](/documentation/guides/security/#table-of-access) to understand what permissions you can configure. +Database API keys can be configured with granular access control. Database API keys with granular access control can be recognized by starting with `eyJhb`. Please refer to the [Table of access](/documentation/operations/security/#table-of-access) to understand what permissions you can configure. Database API keys with granular access control are available for clusters using version **v1.11.0** and above. diff --git a/qdrant-landing/content/documentation/cloud/backups.md b/qdrant-landing/content/documentation/cloud/backups.md index 35be1f81d..5377d94a6 100644 --- a/qdrant-landing/content/documentation/cloud/backups.md +++ b/qdrant-landing/content/documentation/cloud/backups.md @@ -25,7 +25,7 @@ set up your cluster, as described in the following sections: - [Create a cluster](/documentation/cloud/create-cluster/) - Set up [Authentication](/documentation/cloud/authentication/) -- Configure one or more [Collections](/documentation/concepts/collections/) +- Configure one or more [Collections](/documentation/manage-data/collections/) ## Automatic Backups @@ -71,7 +71,7 @@ Or you can restore the backup into a new cluster. Qdrant also offers a snapshot API which allows you to create a snapshot of a specific collection or your entire cluster. For more information, see our -[snapshot documentation](/documentation/concepts/snapshots/). +[snapshot documentation](/documentation/operations/snapshots/). Here is how you can take a snapshot and recover a collection: @@ -79,11 +79,11 @@ Here is how you can take a snapshot and recover a collection: - For a single node cluster, call the snapshot endpoint on the exposed URL. - For a multi node cluster call a snapshot on each node of the collection. Specifically, prepend `node-{num}-` to your cluster URL. - Then call the [snapshot endpoint](/documentation/concepts/snapshots/#create-snapshot) on the individual hosts. Start with node 0. + Then call the [snapshot endpoint](/documentation/operations/snapshots/#create-snapshot) on the individual hosts. Start with node 0. - In the response, you'll see the name of the snapshot. 2. Delete and recreate the collection. 3. Recover the snapshot: - - Call the [recover endpoint](/documentation/concepts/snapshots/#recover-in-cluster-deployment). Set a location which points to the snapshot file (`file:///qdrant/snapshots/{collection_name}/{snapshot_file_name}`) for each host. + - Call the [recover endpoint](/documentation/operations/snapshots/#recover-in-cluster-deployment). Set a location which points to the snapshot file (`file:///qdrant/snapshots/{collection_name}/{snapshot_file_name}`) for each host. ## Backup Considerations diff --git a/qdrant-landing/content/documentation/cloud/cluster-access.md b/qdrant-landing/content/documentation/cloud/cluster-access.md index 3329babd6..d17778644 100644 --- a/qdrant-landing/content/documentation/cloud/cluster-access.md +++ b/qdrant-landing/content/documentation/cloud/cluster-access.md @@ -27,9 +27,9 @@ Have a look at the [API reference](/documentation/interfaces/#api-reference) and ## Node Specific Endpoints -Next to the cluster endpoint which loadbalances requests across all healthy Qdrant nodes, each node in the cluster has its own endpoint as well. This is mainly usefull for monitoring or manual shard management purpuses. +Next to the cluster endpoint which loadbalances requests across all healthy Qdrant nodes, each node in the cluster has its own endpoint as well. This is mainly useful for monitoring or manual shard management purposes. -You can finde the node specific endpoints on the cluster detail page in the Qdrant Cloud Console. +You can find the node specific endpoints on the cluster detail page in the Qdrant Cloud Console. ![Cluster node endpoints](/documentation/cloud/cloud-node-endpoints.png) diff --git a/qdrant-landing/content/documentation/cloud/cluster-monitoring.md b/qdrant-landing/content/documentation/cloud/cluster-monitoring.md index 78498b60f..a24503690 100644 --- a/qdrant-landing/content/documentation/cloud/cluster-monitoring.md +++ b/qdrant-landing/content/documentation/cloud/cluster-monitoring.md @@ -120,7 +120,7 @@ The account owner will receive automatic alerts via email if your cluster has an **Where can I learn more about this alert?** - You can learn more about disk capacity [here](/documentation/guides/capacity-planning/#scaling-disk-space-in-qdrant-cloud). + You can learn more about disk capacity [here](/documentation/operations/capacity-planning/#scaling-disk-space-in-qdrant-cloud). You can learn about vertical scaling [here](/documentation/cloud/cluster-scaling/#vertical-scaling). @@ -142,11 +142,11 @@ The account owner will receive automatic alerts via email if your cluster has an A single collection with payload index partitioning is usually optimal compared to many small individual tenant collections.* - It is also possible to split collections across clusters, the [Qdrant Migration CLI](/documentation/database-tutorials/migration/) can help you with this. + It is also possible to split collections across clusters, the [Qdrant Migration CLI](/documentation/tutorials-operations/migration/) can help you with this. **Where can I learn more about this alert?** - You can learn more about how to set up multi-tenancy with a Qdrant collection [here](/documentation/guides/multiple-partitions/). + You can learn more about how to set up multi-tenancy with a Qdrant collection [here](/documentation/manage-data/multitenancy/). - title: Cluster is Unhealthy content: | @@ -202,7 +202,7 @@ The account owner will receive automatic alerts via email if your cluster has an Learn about the SDKs [here](/documentation/interfaces/). - Learn more about JWT Keys and permissions [here](/documentation/guides/security/?q=jwt#granular-access-control-with-jwt). + Learn more about JWT Keys and permissions [here](/documentation/operations/security/?q=jwt#granular-access-control-with-jwt). - title: A Node is CPU Throttled content: | @@ -232,9 +232,9 @@ The account owner will receive automatic alerts via email if your cluster has an Learn how to optimize Qdrant for performance and configure indexing [here](/documentation/concepts/indexing/) and [here](/documentation/guides/optimize/). - Learn about optimizers [here](/documentation/concepts/optimizer/). + Learn about optimizers [here](/documentation/operations/optimizer/). - Learn more about hybrid search [here](/documentation/concepts/hybrid-queries/). + Learn more about hybrid search [here](/documentation/search/hybrid-queries/). - title: Node CPU Usage is Not Distributed Equally content: | @@ -280,7 +280,7 @@ The account owner will receive automatic alerts via email if your cluster has an **Where can I learn more about this alert?** - Learn more about distributed deployments and resharding [here](/documentation/guides/distributed_deployment/#resharding). + Learn more about distributed deployments and resharding [here](/documentation/operations/distributed_deployment/#resharding). Learn more about cloud rebalancing [here](/documentation/cloud/configure-cluster/#shard-rebalancing). @@ -298,7 +298,7 @@ Metrics in a Prometheus-compatible format are available at the `/metrics` endpoi You can also access the `/telemetry` [endpoint](https://api.qdrant.tech/api-reference/service/telemetry) of your database. This endpoint is available on the cluster endpoint and provides information about the current state of the database, including the number of vectors, shards, and other useful information. -For more information, see [Qdrant monitoring](/documentation/guides/monitoring/). +For more information, see [Qdrant monitoring](/documentation/operations/monitoring/). ### Cluster System Metrics diff --git a/qdrant-landing/content/documentation/cloud/cluster-scaling.md b/qdrant-landing/content/documentation/cloud/cluster-scaling.md index b5af278f1..efb2dc589 100644 --- a/qdrant-landing/content/documentation/cloud/cluster-scaling.md +++ b/qdrant-landing/content/documentation/cloud/cluster-scaling.md @@ -27,7 +27,7 @@ Vertical scaling can be an effective way to improve the performance of a cluster In such cases, horizontal scaling may be a more effective solution. -Horizontal scaling is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/guides/distributed_deployment/#sharding) section for details. +Horizontal scaling is the process of increasing the capacity of a cluster by adding more nodes and distributing the load and data among them. The horizontal scaling at Qdrant starts on the collection level. You have to choose the number of shards you want to distribute your collection around while creating the collection. Please refer to the [sharding documentation](/documentation/operations/distributed_deployment/#sharding) section for details. When scaling up horizontally, the cloud platform will automatically rebalance all available shards across nodes to ensure that the data is evenly distributed. See [Configuring Clusters](/documentation/cloud/configure-cluster/#shard-rebalancing) for more details. diff --git a/qdrant-landing/content/documentation/cloud/cluster-upgrades.md b/qdrant-landing/content/documentation/cloud/cluster-upgrades.md index 772bd9e51..22e13b913 100644 --- a/qdrant-landing/content/documentation/cloud/cluster-upgrades.md +++ b/qdrant-landing/content/documentation/cloud/cluster-upgrades.md @@ -9,7 +9,9 @@ As soon as a new Qdrant version is available. Qdrant Cloud will show you an upda To update to a new version, go to the Cluster Details page, choose the new version from the version dropdown and click **Update**. -If you are several versions behind, multiple updates might be required to reach the latest version. In this case, Qdrant Cloud will automatically perform the required intermediate updates to ensure a supported update path. You should still ensure that your client applications and used SKDs are compatible with the target version. +If you are several versions behind, multiple updates might be required to reach the latest version. In this case, Qdrant Cloud will automatically perform the required intermediate updates to ensure a supported update path. You need to ensure that your client applications and used SDKs are compatible with the target version. + +We recommend first updating the client SDKs, and after that update the cluster to ensure a smooth update process. All client SDKs are tested to be backwards compatible with the latest 3 minor versions of Qdrant. ![Cluster Updates](/documentation/cloud/cluster-upgrades.png) diff --git a/qdrant-landing/content/documentation/cloud/configure-cluster.md b/qdrant-landing/content/documentation/cloud/configure-cluster.md index a0a0fd5b3..71a48fc34 100644 --- a/qdrant-landing/content/documentation/cloud/configure-cluster.md +++ b/qdrant-landing/content/documentation/cloud/configure-cluster.md @@ -7,12 +7,12 @@ weight: 55 Qdrant Cloud offers several advanced configuration options to optimize clusters for your specific needs. You can access these options from the Cluster Details page in the Qdrant Cloud console. -The cloud platform does not expose all [configuration options](/documentation/guides/configuration/) available in Qdrant. We have selected the relevant options that are explained in detail below. +The cloud platform does not expose all [configuration options](/documentation/operations/configuration/) available in Qdrant. We have selected the relevant options that are explained in detail below. -In adition the cloud platform automatically configures the following settings for your cluster to ensure optimal performance and reliability: +In addition the cloud platform automatically configures the following settings for your cluster to ensure optimal performance and reliability: -* The maximum number of collections in a cluster is set to 1000. Larger numbers of collections lead to performance degradation. For more information see [Multitenancy](/documentation/guides/multiple-partitions/). -* Strict mode is activated by default for new collections enforcing that all filters being used in retrieve and udpate queries are indexed. This improves performance and reliability. You can disable this individually for each collection. For more information see [Strict Mode](/documentation/guides/administration/#strict-mode). +* The maximum number of collections in a cluster is set to 1000. Larger numbers of collections lead to performance degradation. For more information see [Multitenancy](/documentation/manage-data/multitenancy/). +* Strict mode is activated by default for new collections enforcing that all filters being used in retrieve and update queries are indexed. This improves performance and reliability. You can disable this individually for each collection. For more information see [Strict Mode](/documentation/operations/administration/#strict-mode). * The cluster mode is automatically enabled to allow distributed deployments and horizontal scaling. * The maximum amount of payload indexes per collection is set to 100. Larger numbers of payload indexes lead to performance degradation (starting with Qdrant v1.16.0). @@ -24,7 +24,7 @@ You can set default values for the configuration of new collections in your clus You can configure the default *Replication Factor*, the default *Write Consistency Factor*, and if vectors should be stored on disk only, instead of being cached in RAM. -Refer to [Qdrant Configuration](/documentation/guides/configuration/#configuration-options) for more details. +Refer to [Qdrant Configuration](/documentation/operations/configuration/#configuration-options) for more details. ## Advanced Optimizations @@ -40,7 +40,7 @@ Configures how many CPUs (threads) to allocate for optimization and indexing job *Async Scorer* -Enables async scorer which uses io_uring when rescoring. See [Qdrant under the hood: io_uring](/articles/io_uring/#and-what-about-qdrant) and [Large Scale Search](/documentation/database-tutorials/large-scale-search/) for more details. +Enables async scorer which uses io_uring when rescoring. See [Qdrant under the hood: io_uring](/articles/io_uring/#and-what-about-qdrant) and [Large Scale Search](/documentation/tutorials-operations/large-scale-search/) for more details. ## Client IP Restrictions diff --git a/qdrant-landing/content/documentation/cloud/create-cluster.md b/qdrant-landing/content/documentation/cloud/create-cluster.md index cf6c4d62b..75346e715 100644 --- a/qdrant-landing/content/documentation/cloud/create-cluster.md +++ b/qdrant-landing/content/documentation/cloud/create-cluster.md @@ -20,7 +20,7 @@ A free tier cluster only includes 1 single node with the following resources: | Disk space | 4 GB | | Nodes | 1 | -This configuration supports serving about 1 M vectors of 768 dimensions. To calculate your needs, refer to our documentation on [Capacity Planning](/documentation/guides/capacity-planning/). +This configuration supports serving about 1 M vectors of 768 dimensions. To calculate your needs, refer to our documentation on [Capacity Planning](/documentation/operations/capacity-planning/). The choice of cloud providers and regions is limited. @@ -52,7 +52,7 @@ On top of the Free cluster features, Standard clusters offer: You have a broad choice of regions on AWS, Azure and Google Cloud. -For payment information see [**Pricing and Payments**](/documentation/cloud/pricing-payments/). +For payment information see [**Pricing and Payments**](/documentation/cloud-pricing-payments/). ## Create a Cluster @@ -75,7 +75,7 @@ This page shows you how to use the Qdrant Cloud Console to create a custom Qdran 1. Choose your data center region or Hybrid Cloud environment. 1. Configure RAM for each node. - > For more information, see our [Capacity Planning](/documentation/guides/capacity-planning/) guidance. + > For more information, see our [Capacity Planning](/documentation/operations/capacity-planning/) guidance. 1. Choose the number of vCPUs per node. If you add more RAM, the menu provides different options for vCPUs. 1. Select the number of nodes you want the cluster to be deployed on. @@ -111,7 +111,7 @@ You should create a backup schedule for your cluster. This ensures that you can **Collection Sharding** -To allow your cluster to easily scale horizontally, you should configure at least twice as many shards per collection than the number of nodes in your cluster. You can configure the number of shards when creating a collection. See [**Sharding**](/documentation/guides/distributed_deployment/#sharding) for more information. +To allow your cluster to easily scale horizontally, you should configure at least twice as many shards per collection than the number of nodes in your cluster. You can configure the number of shards when creating a collection. See [**Sharding**](/documentation/operations/distributed_deployment/#sharding) for more information. If you did not configure enough shards in a collection, you can use the [**Resharding**](/documentation/cloud/cluster-scaling/#resharding) feature to change the number of shards in an existing collection. diff --git a/qdrant-landing/content/documentation/cloud/inference.md b/qdrant-landing/content/documentation/cloud/inference.md index 87d716874..94cd90c78 100644 --- a/qdrant-landing/content/documentation/cloud/inference.md +++ b/qdrant-landing/content/documentation/cloud/inference.md @@ -5,7 +5,7 @@ weight: 81 # Inference in Qdrant Managed Cloud -[Inference](/documentation/concepts/inference/) is the process of creating vector embeddings from text, images, or other data types using a machine learning model. +[Inference](/documentation/inference/) is the process of creating vector embeddings from text, images, or other data types using a machine learning model. Qdrant Managed Cloud allows you to use inference directly in the cloud, without the need to set up and maintain your own inference infrastructure. You can use [embedding models hosted on Qdrant Cloud](#cloud-inference), or use [externally hosted models](#use-external-models). @@ -21,7 +21,7 @@ Inference is enabled by default for all new clusters, created after July, 7th 20 ## Using Inference -Inference can be easily used through the Qdrant SDKs and the REST or GRPC APIs when upserting points and when querying the database. Refer to the [Inference documentation](/documentation/concepts/inference/) for details. +Inference can be easily used through the Qdrant SDKs and the REST or GRPC APIs when upserting points and when querying the database. Refer to the [Inference documentation](/documentation/inference/) for details. ## Cloud Inference diff --git a/qdrant-landing/content/documentation/data-management/_index.md b/qdrant-landing/content/documentation/data-management/_index.md index 8435d53c4..13b352837 100644 --- a/qdrant-landing/content/documentation/data-management/_index.md +++ b/qdrant-landing/content/documentation/data-management/_index.md @@ -1,7 +1,7 @@ --- title: Data Management -weight: 11 -partition: build +weight: 600 +partition: ecosystem --- ## Data Management Integrations diff --git a/qdrant-landing/content/documentation/data-management/airbyte.md b/qdrant-landing/content/documentation/data-management/airbyte.md index 19839644e..c5cf0fa55 100644 --- a/qdrant-landing/content/documentation/data-management/airbyte.md +++ b/qdrant-landing/content/documentation/data-management/airbyte.md @@ -26,7 +26,7 @@ Before you start, make sure you have the following: 1. Airbyte instance, either [Open Source](https://airbyte.com/solutions/airbyte-open-source), [Self-Managed](https://airbyte.com/solutions/airbyte-enterprise), or [Cloud](https://airbyte.com/solutions/airbyte-cloud). 2. Running instance of Qdrant. It has to be accessible by URL from the machine where Airbyte is running. - You can follow the [installation guide](/documentation/guides/installation/) to set up Qdrant. + You can follow the [installation guide](/documentation/operations/installation/) to set up Qdrant. ## Setting up Qdrant as a destination diff --git a/qdrant-landing/content/documentation/data-management/airflow.md b/qdrant-landing/content/documentation/data-management/airflow.md index aa9191ab8..45e196ce9 100644 --- a/qdrant-landing/content/documentation/data-management/airflow.md +++ b/qdrant-landing/content/documentation/data-management/airflow.md @@ -13,7 +13,7 @@ Qdrant is available as a [provider](https://airflow.apache.org/docs/apache-airfl Before configuring Airflow, you need: -1. A Qdrant instance to connect to. You can set one up in our [installation guide](/documentation/guides/installation/). +1. A Qdrant instance to connect to. You can set one up in our [installation guide](/documentation/operations/installation/). 2. A running Airflow instance. You can use their [Quick Start Guide](https://airflow.apache.org/docs/apache-airflow/stable/start.html). diff --git a/qdrant-landing/content/documentation/data-management/confluent.md b/qdrant-landing/content/documentation/data-management/confluent.md index 6d58202ac..bd6295414 100644 --- a/qdrant-landing/content/documentation/data-management/confluent.md +++ b/qdrant-landing/content/documentation/data-management/confluent.md @@ -53,7 +53,7 @@ _Click each to expand._
Unnamed/Default vector -Reference: [Creating a collection with a default vector](https://qdrant.tech/documentation/concepts/collections/#create-a-collection). +Reference: [Creating a collection with a default vector](/documentation/manage-data/collections/#create-a-collection). ```json { @@ -82,7 +82,7 @@ Reference: [Creating a collection with a default vector](https://qdrant.tech/doc
Named multiple vectors -Reference: [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors). +Reference: [Creating a collection with multiple vectors](/documentation/manage-data/collections/#collection-with-multiple-vectors). ```json { @@ -123,7 +123,7 @@ Reference: [Creating a collection with multiple vectors](https://qdrant.tech/doc
Sparse vectors -Reference: [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors). +Reference: [Creating a collection with sparse vectors](/documentation/manage-data/collections/#collection-with-sparse-vectors). ```json { @@ -172,7 +172,7 @@ Reference: [Creating a collection with sparse vectors](https://qdrant.tech/docum Reference: -- [Multi-vectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors) +- [Multi-vectors](/documentation/manage-data/vectors/#multivectors) ```json { @@ -221,9 +221,9 @@ Reference: Reference: -- [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors). +- [Creating a collection with multiple vectors](/documentation/manage-data/collections/#collection-with-multiple-vectors). -- [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors). +- [Creating a collection with sparse vectors](/documentation/manage-data/collections/#collection-with-sparse-vectors). ```json { diff --git a/qdrant-landing/content/documentation/data-management/fluvio.md b/qdrant-landing/content/documentation/data-management/fluvio.md index 09d90f84a..fb368a416 100644 --- a/qdrant-landing/content/documentation/data-management/fluvio.md +++ b/qdrant-landing/content/documentation/data-management/fluvio.md @@ -71,7 +71,7 @@ _Click each to expand._
Unnamed/Default vector -Reference: [Creating a collection with a default vector](https://qdrant.tech/documentation/concepts/collections/#create-a-collection). +Reference: [Creating a collection with a default vector](/documentation/manage-data/collections/#create-a-collection). ```json { @@ -100,7 +100,7 @@ Reference: [Creating a collection with a default vector](https://qdrant.tech/doc
Named multiple vectors -Reference: [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors). +Reference: [Creating a collection with multiple vectors](/documentation/manage-data/collections/#collection-with-multiple-vectors). ```json { @@ -141,7 +141,7 @@ Reference: [Creating a collection with multiple vectors](https://qdrant.tech/doc
Sparse vectors -Reference: [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors). +Reference: [Creating a collection with sparse vectors](/documentation/manage-data/collections/#collection-with-sparse-vectors). ```json { @@ -235,9 +235,9 @@ Reference: [Creating a collection with sparse vectors](https://qdrant.tech/docum Reference: -- [Creating a collection with multiple vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-multiple-vectors). +- [Creating a collection with multiple vectors](/documentation/manage-data/collections/#collection-with-multiple-vectors). -- [Creating a collection with sparse vectors](https://qdrant.tech/documentation/concepts/collections/#collection-with-sparse-vectors). +- [Creating a collection with sparse vectors](/documentation/manage-data/collections/#collection-with-sparse-vectors). ```json { diff --git a/qdrant-landing/content/documentation/data-management/spark.md b/qdrant-landing/content/documentation/data-management/spark.md index cbb41eaec..184bb8b14 100644 --- a/qdrant-landing/content/documentation/data-management/spark.md +++ b/qdrant-landing/content/documentation/data-management/spark.md @@ -71,7 +71,7 @@ public class QdrantSparkJavaExample { ### Loading data -Before loading the data using this connector, a collection has to be [created](https://qdrant.tech/documentation/concepts/collections/#create-a-collection) in advance with the appropriate vector dimensions and configurations. +Before loading the data using this connector, a collection has to be [created](/documentation/manage-data/collections/#create-a-collection) in advance with the appropriate vector dimensions and configurations. The connector supports ingesting multiple named/unnamed, dense/sparse vectors. diff --git a/qdrant-landing/content/documentation/data-synchronization/_index.md b/qdrant-landing/content/documentation/data-synchronization/_index.md new file mode 100644 index 000000000..eb1577b0b --- /dev/null +++ b/qdrant-landing/content/documentation/data-synchronization/_index.md @@ -0,0 +1,21 @@ +--- +title: Data Synchronization +weight: 400 +partition: ecosystem +--- + +# Keeping Your Data in Sync with Qdrant + +After migrating to Qdrant, maintaining consistency between Qdrant and your other data stores can be an ongoing operational concern. As records are created, updated, or deleted in your other sources, those changes often need to be reflected in Qdrant to keep search results accurate and fresh. + +The guides in this section cover database-specific synchronization patterns for keeping Qdrant aligned with your existing infrastructure. + +## Synchronization Strategies + +How you sync depends on your source system and write volume: + +- **Event-driven sync** — listen for change events (CDC, triggers, or message queues) and upsert or delete points in Qdrant in real time. +- **Batch sync** — periodically re-index records that have changed since the last run, using a timestamp or version field to identify updates. +- **Dual-write** — write to both systems simultaneously in application code, with a background reconciliation job to catch divergence. + +Each approach involves trade-offs between latency, complexity, and consistency guarantees. The guides below cover patterns for specific databases. diff --git a/qdrant-landing/content/documentation/data-synchronization.md b/qdrant-landing/content/documentation/data-synchronization/with-postgres.md similarity index 99% rename from qdrant-landing/content/documentation/data-synchronization.md rename to qdrant-landing/content/documentation/data-synchronization/with-postgres.md index e73cf34c8..e841ac24b 100644 --- a/qdrant-landing/content/documentation/data-synchronization.md +++ b/qdrant-landing/content/documentation/data-synchronization/with-postgres.md @@ -1,7 +1,7 @@ --- -title: Keeping Postgres in Sync -weight: 33 -partition: build +title: With Postgres +weight: 5 +partition: ecosystem --- # Keeping Postgres and Qdrant in Sync diff --git a/qdrant-landing/content/documentation/datasets.md b/qdrant-landing/content/documentation/datasets.md index c34afe6b4..f80c71a5d 100644 --- a/qdrant-landing/content/documentation/datasets.md +++ b/qdrant-landing/content/documentation/datasets.md @@ -1,7 +1,7 @@ --- title: Practice Datasets -weight: 28 -partition: build +weight: 1500 +partition: ecosystem --- # Common Datasets in Snapshot Format @@ -22,7 +22,7 @@ on a dataset name to see its detailed description. | [Arxiv.org abstracts](#arxivorg-abstracts) | [InstructorXL](https://huggingface.co/hkunlp/instructor-xl) | 768 | 2.3M | 8.4 GB | [Download](https://snapshots.qdrant.io/arxiv_abstracts-3083016565637815127-2023-06-02-07-26-29.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/arxiv-abstracts-instructorxl-embeddings) | | [Wolt food](#wolt-food) | [clip-ViT-B-32](https://huggingface.co/sentence-transformers/clip-ViT-B-32) | 512 | 1.7M | 7.9 GB | [Download](https://snapshots.qdrant.io/wolt-clip-ViT-B-32-2446808438011867-2023-12-14-15-55-26.snapshot) | [Open](https://huggingface.co/datasets/Qdrant/wolt-food-clip-ViT-B-32-embeddings) | -Once you download a snapshot, you need to [restore it](/documentation/concepts/snapshots/#restore-snapshot) +Once you download a snapshot, you need to [restore it](/documentation/operations/snapshots/#restore-snapshot) using the Qdrant CLI upon startup or through the API. ## Qdrant on Hugging Face @@ -40,7 +40,7 @@ and build your applications based on semantic search. **Please let us know if yo a specific dataset!** If you are not familiar with [Hugging Face datasets](https://huggingface.co/docs/datasets/index), -or would like to know how to combine it with Qdrant, please refer to the [tutorial](/documentation/tutorials/huggingface-datasets/). +or would like to know how to combine it with Qdrant, please refer to the [tutorial](/documentation/tutorials-basics/huggingface-datasets/). ## Arxiv.org diff --git a/qdrant-landing/content/documentation/dl-integration-examples.md b/qdrant-landing/content/documentation/dl-integration-examples.md index 5f44e6b22..b7151683c 100644 --- a/qdrant-landing/content/documentation/dl-integration-examples.md +++ b/qdrant-landing/content/documentation/dl-integration-examples.md @@ -1,9 +1,9 @@ --- #Delimiter files are used to separate the list of documentation pages into sections. -title: "Integration Guides" +title: "Ecosystem Guides" type: delimiter -weight: 20 # Change this weight to change order of sections -partition: build +weight: 1100 # Change this weight to change order of sections +partition: ecosystem sitemapExclude: True _build: publishResources: false diff --git a/qdrant-landing/content/documentation/dl-integrations.md b/qdrant-landing/content/documentation/dl-integrations.md index 7d0ef8fb5..ec706d073 100644 --- a/qdrant-landing/content/documentation/dl-integrations.md +++ b/qdrant-landing/content/documentation/dl-integrations.md @@ -2,10 +2,10 @@ #Delimiter files are used to separate the list of documentation pages into sections. title: "Integrations" type: delimiter -weight: 10 # Change this weight to change order of sections +weight: 500 # Change this weight to change order of sections sitemapExclude: True _build: publishResources: false render: never -partition: build +partition: ecosystem --- \ No newline at end of file diff --git a/qdrant-landing/content/documentation/dl-migrate.md b/qdrant-landing/content/documentation/dl-migrate.md index a6a2fffb7..3ebf67282 100644 --- a/qdrant-landing/content/documentation/dl-migrate.md +++ b/qdrant-landing/content/documentation/dl-migrate.md @@ -2,10 +2,10 @@ #Delimiter files are used to separate the list of documentation pages into sections. title: "Migrate to Qdrant" type: delimiter -weight: 30 # Change this weight to change order of sections +weight: 100 # Change this weight to change order of sections sitemapExclude: True _build: publishResources: false render: never -partition: build +partition: ecosystem --- diff --git a/qdrant-landing/content/documentation/dl-ecosystem.md b/qdrant-landing/content/documentation/dl-tools.md similarity index 92% rename from qdrant-landing/content/documentation/dl-ecosystem.md rename to qdrant-landing/content/documentation/dl-tools.md index fb6b9e74f..a9bf11237 100644 --- a/qdrant-landing/content/documentation/dl-ecosystem.md +++ b/qdrant-landing/content/documentation/dl-tools.md @@ -1,6 +1,6 @@ --- #Delimiter files are used to separate the list of documentation pages into sections. -title: "Ecosystem" +title: "Qdrant Tools" type: delimiter weight: 300 # Change this weight to change order of sections sitemapExclude: True diff --git a/qdrant-landing/content/documentation/build-tab.md b/qdrant-landing/content/documentation/ecosystem-tab.md similarity index 85% rename from qdrant-landing/content/documentation/build-tab.md rename to qdrant-landing/content/documentation/ecosystem-tab.md index 38b294a70..b463391d0 100644 --- a/qdrant-landing/content/documentation/build-tab.md +++ b/qdrant-landing/content/documentation/ecosystem-tab.md @@ -1,11 +1,13 @@ --- -title: Build World-Class Applications -slug: build +title: Explore the Qdrant Ecosystem +slug: ecosystem +aliases: + - /documentation/build/ breadcrumb: false content: - partial: documentation/banners/banner-c - title: Build World-Class Applications - description: Dev-portal Build + title: Explore the Qdrant Ecosystem + description: Dev-portal Ecosystem image: src: /img/dev-portal-build/spanish-ai-app-hero.png alt: Spanish AI app.png @@ -24,7 +26,7 @@ content: title: Search description: Build a simple neural search service with Qdrant and FastEmbed. Learn how to upload data, create indexes, and run search queries. link: - url: /documentation/beginner-tutorials/hybrid-search-fastembed/ + url: /documentation/tutorials-search-engineering/hybrid-search-fastembed/ text: Read More - id: 2 image: @@ -44,6 +46,6 @@ content: link: url: /documentation/send-data/ text: Read More -partition: build +partition: ecosystem hideInSidebar: true --- diff --git a/qdrant-landing/content/documentation/edge/_index.md b/qdrant-landing/content/documentation/edge/_index.md index 497123540..89fb5a891 100644 --- a/qdrant-landing/content/documentation/edge/_index.md +++ b/qdrant-landing/content/documentation/edge/_index.md @@ -1,6 +1,6 @@ --- title: "Qdrant Edge" -weight: 315 +weight: 225 partition: qdrant --- diff --git a/qdrant-landing/content/documentation/edge/edge-data-synchronization-patterns.md b/qdrant-landing/content/documentation/edge/edge-data-synchronization-patterns.md index 1ad8643e1..0bead20a8 100644 --- a/qdrant-landing/content/documentation/edge/edge-data-synchronization-patterns.md +++ b/qdrant-landing/content/documentation/edge/edge-data-synchronization-patterns.md @@ -1,6 +1,7 @@ --- title: "Data Synchronization Patterns" weight: 20 +partition: qdrant --- # Data Synchronization Patterns @@ -13,7 +14,7 @@ Instead of starting with an empty Edge Shard, you may want to initialize it with ![Qdrant Edge Shards can be initialized from snapshots of server-side shards](/documentation/edge/qdrant-edge-restore-snapshot.png) -When creating a snapshot for synchronization, specify the applicable server-side shard ID in the snapshot URL. This allows for a single collection to serve multiple independent users or devices, each with its own Edge Shard. Read more about Qdrant's sharding strategy in the [Tiered Multitenancy Documentation](/documentation/guides/multitenancy/#tiered-multitenancy). +When creating a snapshot for synchronization, specify the applicable server-side shard ID in the snapshot URL. This allows for a single collection to serve multiple independent users or devices, each with its own Edge Shard. Read more about Qdrant's sharding strategy in the [Tiered Multitenancy Documentation](/documentation/manage-data/multitenancy/#tiered-multitenancy). First, craft a snapshot URL: diff --git a/qdrant-landing/content/documentation/edge/edge-fastembed-embeddings.md b/qdrant-landing/content/documentation/edge/edge-fastembed-embeddings.md index b83aac7b9..159f92343 100644 --- a/qdrant-landing/content/documentation/edge/edge-fastembed-embeddings.md +++ b/qdrant-landing/content/documentation/edge/edge-fastembed-embeddings.md @@ -1,6 +1,7 @@ --- title: "On-Device Embeddings" weight: 15 +partition: qdrant --- # On-Device Embeddings with Qdrant Edge and FastEmbed diff --git a/qdrant-landing/content/documentation/edge/edge-quickstart.md b/qdrant-landing/content/documentation/edge/edge-quickstart.md index f9e00d59c..8cd6f5404 100644 --- a/qdrant-landing/content/documentation/edge/edge-quickstart.md +++ b/qdrant-landing/content/documentation/edge/edge-quickstart.md @@ -1,6 +1,7 @@ --- title: "Quickstart" weight: 10 +partition: qdrant --- # Qdrant Edge Quickstart diff --git a/qdrant-landing/content/documentation/edge/edge-synchronization-guide.md b/qdrant-landing/content/documentation/edge/edge-synchronization-guide.md index ef839d1be..a5069329b 100644 --- a/qdrant-landing/content/documentation/edge/edge-synchronization-guide.md +++ b/qdrant-landing/content/documentation/edge/edge-synchronization-guide.md @@ -1,6 +1,7 @@ --- title: "Synchronize with a Server" weight: 30 +partition: qdrant --- # Synchronize Qdrant Edge with a Server diff --git a/qdrant-landing/content/documentation/embeddings/_index.md b/qdrant-landing/content/documentation/embeddings/_index.md index fade68ec7..325c61347 100644 --- a/qdrant-landing/content/documentation/embeddings/_index.md +++ b/qdrant-landing/content/documentation/embeddings/_index.md @@ -1,7 +1,7 @@ --- title: Embeddings -weight: 12 -partition: build +weight: 700 +partition: ecosystem --- # Supported Embedding Providers & Models diff --git a/qdrant-landing/content/documentation/examples/Qdrant-DSPy-medicalbot.md b/qdrant-landing/content/documentation/examples/Qdrant-DSPy-medicalbot.md index 25a8e5fb3..bed138f38 100644 --- a/qdrant-landing/content/documentation/examples/Qdrant-DSPy-medicalbot.md +++ b/qdrant-landing/content/documentation/examples/Qdrant-DSPy-medicalbot.md @@ -1,6 +1,6 @@ --- title: Building a Chain-of-Thought Medical Chatbot with Qdrant and DSPy -weight: 18 +weight: 20 --- # Building a Chain-of-Thought Medical Chatbot with Qdrant and DSPy diff --git a/qdrant-landing/content/documentation/examples/_index.md b/qdrant-landing/content/documentation/examples/_index.md index 75b054a53..c1d047e08 100644 --- a/qdrant-landing/content/documentation/examples/_index.md +++ b/qdrant-landing/content/documentation/examples/_index.md @@ -1,7 +1,7 @@ --- title: Build Prototypes -weight: 25 -partition: build +weight: 1300 +partition: ecosystem --- # Build Prototypes @@ -22,13 +22,13 @@ The following guided samples help you get started with real-world projects using | [Blog-Reading RAG Chatbot](/documentation/examples/rag-chatbot-scaleway/) | Develop a RAG-based Chatbot on Scaleway and with LangChain | Qdrant, LangChain, GPT-4o | [Movie Recommendation System](/documentation/examples/recommendation-system-ovhcloud/) | Build a Movie Recommendation System with LlamaIndex and With JinaAI | Qdrant | | [GraphRAG Agent](/documentation/examples/graphrag-qdrant-neo4j/) | Build a GraphRAG Agent with Neo4J and Qdrant | Qdrant, Neo4j | -| [Building a Chain-of-Thought Medical Chatbot with Qdrant and DSPy](/documentation/examples/Qdrant-DSPy-medicalbot/) | How to build a medical chatbot grounded in medical literature with Qdrant and DSPy. | Qdrant, DSPy | +| [Building a Chain-of-Thought Medical Chatbot with Qdrant and DSPy](/documentation/examples/qdrant-dspy-medicalbot/) | How to build a medical chatbot grounded in medical literature with Qdrant and DSPy. | Qdrant, DSPy | ## Example Notebooks -Our Notebooks offer complex instructions that are supported with a throrough explanation. Follow along by trying out the code and get the most out of each example. +Our Notebooks offer complex instructions that are supported with a thorough explanation. Follow along by trying out the code and get the most out of each example. | Example | Description | Stack | |---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------|----------------------------| diff --git a/qdrant-landing/content/documentation/examples/aleph-alpha-search.md b/qdrant-landing/content/documentation/examples/aleph-alpha-search.md index 9921bd958..4ac7713f1 100644 --- a/qdrant-landing/content/documentation/examples/aleph-alpha-search.md +++ b/qdrant-landing/content/documentation/examples/aleph-alpha-search.md @@ -1,6 +1,6 @@ --- title: Aleph Alpha Search -weight: 16 +weight: 10 draft: true --- diff --git a/qdrant-landing/content/documentation/examples/cohere-rag-connector.md b/qdrant-landing/content/documentation/examples/cohere-rag-connector.md index 9608d7148..ce295adfd 100644 --- a/qdrant-landing/content/documentation/examples/cohere-rag-connector.md +++ b/qdrant-landing/content/documentation/examples/cohere-rag-connector.md @@ -1,6 +1,6 @@ --- title: Implement Cohere RAG connector -weight: 24 +weight: 35 aliases: - /documentation/tutorials/cohere-rag-connector/ --- @@ -258,7 +258,7 @@ The output should look like following: Our web service is implemented, yet running only on our local machine. It has to be exposed to the public before Command-R can interact with it. For a quick experiment, it might be enough to set up tunneling using services such as [ngrok](https://ngrok.com/). We won't cover all the details in the tutorial, but their -[Quickstart](https://ngrok.com/docs/guides/getting-started/) is a great resource describing the process step-by-step. +[Quickstart](https://ngrok.com/docs/getting-started/) is a great resource describing the process step-by-step. Alternatively, you can also deploy the service with a public URL. Once it's done, we can create the connector first, and then tell the model to use it, while interacting through the chat diff --git a/qdrant-landing/content/documentation/examples/hybrid-search-llamaindex-jinaai.md b/qdrant-landing/content/documentation/examples/hybrid-search-llamaindex-jinaai.md index 19f233510..c73f388e6 100644 --- a/qdrant-landing/content/documentation/examples/hybrid-search-llamaindex-jinaai.md +++ b/qdrant-landing/content/documentation/examples/hybrid-search-llamaindex-jinaai.md @@ -1,6 +1,6 @@ --- title: Chat With Product PDF Manuals Using Hybrid Search -weight: 27 +weight: 45 social_preview_image: /blog/hybrid-cloud-llamaindex/hybrid-cloud-llamaindex-tutorial.png aliases: - /documentation/tutorials/hybrid-search-llamaindex-jinaai/ diff --git a/qdrant-landing/content/documentation/examples/llama-index-multitenancy.md b/qdrant-landing/content/documentation/examples/llama-index-multitenancy.md index 31954d2ae..3e057f16b 100644 --- a/qdrant-landing/content/documentation/examples/llama-index-multitenancy.md +++ b/qdrant-landing/content/documentation/examples/llama-index-multitenancy.md @@ -1,6 +1,6 @@ --- title: Multitenancy with LlamaIndex -weight: 18 +weight: 25 aliases: - /documentation/tutorials/llama-index-multitenancy/ --- @@ -9,8 +9,8 @@ aliases: If you are building a service that serves vectors for many independent users, and you want to isolate their data, the best practice is to use a single collection with payload-based partitioning. This approach is -called **multitenancy**. Our guide on the [Separate Partitions](/documentation/guides/multiple-partitions/) describes -how to set it up in general, but if you use [LlamaIndex](/documentation/integrations/llama-index/) as a +called **multitenancy**. Our guide on the [Separate Partitions](/documentation/manage-data/multitenancy/) describes +how to set it up in general, but if you use [LlamaIndex](/documentation/frameworks/llama-index/) as a backend, you may prefer reading a more specific instruction. So here it is! ## Prerequisites @@ -147,7 +147,7 @@ for document in documents: Our documents have been split into nodes, encoded using the embedding model, and stored in the vector store. However, we don't want to allow our users to search for all the documents in the collection, but only for the documents that belong to a library they are interested in. For that reason, we need -to set up the Qdrant [payload index](/documentation/concepts/indexing/#payload-index), so the search +to set up the Qdrant [payload index](/documentation/manage-data/indexing/#payload-index), so the search is more efficient. ```python @@ -162,7 +162,7 @@ client.create_payload_index( The payload index is not the only thing we want to change. Since none of the search queries will be executed on the whole collection, we can also change its configuration, so the HNSW -graph is not built globally. This is also done due to [performance reasons](/documentation/guides/multiple-partitions/#calibrate-performance). +graph is not built globally. This is also done due to [performance reasons](/documentation/manage-data/multitenancy/#calibrate-performance). **You should not be changing these parameters, if you know there will be some global search operations done on the collection.** diff --git a/qdrant-landing/content/documentation/examples/mighty.md b/qdrant-landing/content/documentation/examples/mighty.md index bd8552a9f..40be97f5e 100644 --- a/qdrant-landing/content/documentation/examples/mighty.md +++ b/qdrant-landing/content/documentation/examples/mighty.md @@ -2,7 +2,7 @@ title: "Inference with Mighty" short_description: "Mighty offers a speedy scalable embedding, a perfect fit for the speedy scalable Qdrant search. Let's combine them!" description: "We combine Mighty and Qdrant to create a semantic search service in Rust with just a few lines of code." -weight: 17 +weight: 15 author: Andre Bogus author_link: https://llogiq.github.io date: 2023-06-01T11:24:20+01:00 diff --git a/qdrant-landing/content/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain.md b/qdrant-landing/content/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain.md index dd6f0c4ea..1e829f2be 100644 --- a/qdrant-landing/content/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain.md +++ b/qdrant-landing/content/documentation/examples/natural-language-search-oracle-cloud-infrastructure-cohere-langchain.md @@ -1,6 +1,6 @@ --- title: RAG System for Employee Onboarding -weight: 30 +weight: 55 social_preview_image: /blog/hybrid-cloud-oracle-cloud-infrastructure/hybrid-cloud-oracle-cloud-infrastructure-tutorial.png aliases: - /documentation/tutorials/natural-language-search-oracle-cloud-infrastructure-cohere-langchain/ diff --git a/qdrant-landing/content/documentation/examples/rag-chatbot-red-hat-openshift-haystack.md b/qdrant-landing/content/documentation/examples/rag-chatbot-red-hat-openshift-haystack.md index bb5294b61..a06e852b1 100644 --- a/qdrant-landing/content/documentation/examples/rag-chatbot-red-hat-openshift-haystack.md +++ b/qdrant-landing/content/documentation/examples/rag-chatbot-red-hat-openshift-haystack.md @@ -1,6 +1,6 @@ --- title: Private Chatbot for Interactive Learning -weight: 23 +weight: 30 social_preview_image: /blog/hybrid-cloud-red-hat-openshift/hybrid-cloud-red-hat-openshift-tutorial.png aliases: - /documentation/tutorials/rag-chatbot-red-hat-openshift-haystack/ @@ -237,7 +237,7 @@ search_pipeline = Pipeline() ``` Our second process takes user input, converts it into embeddings and then searches for the most relevant documents -using the query embedding. This might look familiar, but we arent working with `Document` instances +using the query embedding. This might look familiar, but we aren't working with `Document` instances anymore, since the query only accepts raw text. Thus, some of the components will be different, especially the embedder, as it has to accept a single string as an input and produce a single embedding as an output: @@ -458,4 +458,4 @@ The response should be similar to the one we got in the Python before: - [Haystack's documentation](https://docs.haystack.deepset.ai/docs/kubernetes) describes [how to deploy the Hayhooks service in a Kubernetes environment](https://docs.haystack.deepset.ai/docs/kubernetes), so you can easily move it to your own OpenShift infrastructure. -- If you are just getting started and need more guidance on Qdrant, read the [quickstart](/documentation/quick-start/) or try out our [beginner tutorial](/documentation/tutorials/neural-search/). \ No newline at end of file +- If you are just getting started and need more guidance on Qdrant, read the [quickstart](/documentation/quickstart/) or try out our [beginner tutorial](/documentation/tutorials-search-engineering/neural-search/). \ No newline at end of file diff --git a/qdrant-landing/content/documentation/examples/rag-chatbot-scaleway.md b/qdrant-landing/content/documentation/examples/rag-chatbot-scaleway.md index e969615ab..7c533a649 100644 --- a/qdrant-landing/content/documentation/examples/rag-chatbot-scaleway.md +++ b/qdrant-landing/content/documentation/examples/rag-chatbot-scaleway.md @@ -1,6 +1,6 @@ --- title: Blog-Reading Chatbot with GPT-4o -weight: 35 +weight: 70 social_preview_image: /blog/hybrid-cloud-scaleway/hybrid-cloud-scaleway-tutorial.png aliases: - /documentation/tutorials/rag-chatbot-scaleway/ diff --git a/qdrant-landing/content/documentation/examples/rag-chatbot-vultr-dspy-ollama.md b/qdrant-landing/content/documentation/examples/rag-chatbot-vultr-dspy-ollama.md index 6dc92538b..dc939a7b7 100644 --- a/qdrant-landing/content/documentation/examples/rag-chatbot-vultr-dspy-ollama.md +++ b/qdrant-landing/content/documentation/examples/rag-chatbot-vultr-dspy-ollama.md @@ -1,6 +1,6 @@ --- title: Private RAG Information Extraction Engine -weight: 32 +weight: 60 social_preview_image: /blog/hybrid-cloud-vultr/hybrid-cloud-vultr-tutorial.png aliases: - /documentation/tutorials/rag-chatbot-vultr-dspy-ollama/ diff --git a/qdrant-landing/content/documentation/examples/rag-contract-management-stackit-aleph-alpha.md b/qdrant-landing/content/documentation/examples/rag-contract-management-stackit-aleph-alpha.md index 35b22def8..fc9fcc873 100644 --- a/qdrant-landing/content/documentation/examples/rag-contract-management-stackit-aleph-alpha.md +++ b/qdrant-landing/content/documentation/examples/rag-contract-management-stackit-aleph-alpha.md @@ -1,6 +1,6 @@ --- title: Region-Specific Contract Management System -weight: 28 +weight: 50 social_preview_image: /blog/hybrid-cloud-aleph-alpha/hybrid-cloud-aleph-alpha-tutorial.png aliases: - /documentation/tutorials/rag-contract-management-stackit-aleph-alpha/ @@ -93,7 +93,7 @@ developing business logic. Aleph Alpha embeddings are high dimensional vectors by default, with a dimensionality of `5120`. However, a pretty unique feature of that model is that they might be compressed to a size of `128`, with a small drop in accuracy performance (4-6%, according to the docs). Qdrant can store even the original vectors easily, and this sounds like a -good idea to enable [Binary Quantization](/documentation/guides/quantization/#binary-quantization) to save space and +good idea to enable [Binary Quantization](/documentation/manage-data/quantization/#binary-quantization) to save space and make the retrieval faster. Let's create a collection with such settings: ```python @@ -234,7 +234,7 @@ llm = AlephAlpha( Then, we can glue the components together and build the search process. `RetrievalQA` is a class that takes implements the Question Retrieval process, with a specified retriever and Large Language Model. The instance of `Qdrant` might be converted into a retriever, with additional filter that will be passed to the `similarity_search` method. The filter -is created as [in a regular Qdrant query](/documentation/concepts/filtering/), with the `roles` field set to the +is created as [in a regular Qdrant query](/documentation/search/filtering/), with the `roles` field set to the user's roles. ```python diff --git a/qdrant-landing/content/documentation/examples/rag-customer-support-cohere-airbyte-aws.md b/qdrant-landing/content/documentation/examples/rag-customer-support-cohere-airbyte-aws.md index 53bf6c2b5..d815b72bf 100644 --- a/qdrant-landing/content/documentation/examples/rag-customer-support-cohere-airbyte-aws.md +++ b/qdrant-landing/content/documentation/examples/rag-customer-support-cohere-airbyte-aws.md @@ -1,6 +1,6 @@ --- title: Question-Answering System for AI Customer Support -weight: 26 +weight: 40 social_preview_image: /blog/hybrid-cloud-airbyte/hybrid-cloud-airbyte-tutorial.png aliases: - /documentation/tutorials/rag-customer-support-cohere-airbyte-aws/ diff --git a/qdrant-landing/content/documentation/examples/recommendation-system-ovhcloud.md b/qdrant-landing/content/documentation/examples/recommendation-system-ovhcloud.md index 60d2de8d6..f8d8ec516 100644 --- a/qdrant-landing/content/documentation/examples/recommendation-system-ovhcloud.md +++ b/qdrant-landing/content/documentation/examples/recommendation-system-ovhcloud.md @@ -1,6 +1,6 @@ --- title: Movie Recommendation System -weight: 34 +weight: 65 social_preview_image: /blog/hybrid-cloud-ovhcloud/hybrid-cloud-ovhcloud-tutorial.png aliases: - /documentation/tutorials/recommendation-system-ovhcloud/ diff --git a/qdrant-landing/content/documentation/faq/database-optimization.md b/qdrant-landing/content/documentation/faq/database-optimization.md index 2d900d674..7d20b49bd 100644 --- a/qdrant-landing/content/documentation/faq/database-optimization.md +++ b/qdrant-landing/content/documentation/faq/database-optimization.md @@ -9,11 +9,11 @@ weight: 2 The primary source of memory usage is vector data. There are several ways to address that: -- Configure [Quantization](/documentation/guides/quantization/) to reduce the memory usage of vectors. +- Configure [Quantization](/documentation/manage-data/quantization/) to reduce the memory usage of vectors. - Configure on-disk vector storage The choice of the approach depends on your requirements. -Read more about [configuring the optimal](/documentation/tutorials/optimize/) use of Qdrant. +Read more about [configuring the optimal](/documentation/operations/optimize/) use of Qdrant. ### How do you choose the machine configuration? @@ -38,6 +38,6 @@ If you want to limit the memory usage of the service, we recommend using [limits There are several possible reasons for that: -- **Using filters without payload index** -- If you're performing a search with a filter but you don't have a payload index, Qdrant will have to load whole payload data from disk to check the filtering condition. Ensure you have adequately configured [payload indexes](/documentation/concepts/indexing/#payload-index). +- **Using filters without payload index** -- If you're performing a search with a filter but you don't have a payload index, Qdrant will have to load whole payload data from disk to check the filtering condition. Ensure you have adequately configured [payload indexes](/documentation/manage-data/indexing/#payload-index). - **Usage of on-disk vector storage with slow disks** -- If you're using on-disk vector storage, ensure you have fast enough disks. We recommend using local SSDs with at least 50k IOPS. Read more about the influence of the disk speed on the search latency in the article about [Memory Consumption](/articles/memory-consumption/). - **Large limit or non-optimal query parameters** -- A large limit or offset might lead to significant performance degradation. Please pay close attention to the query/collection parameters that significantly diverge from the defaults. They might be the reason for the performance issues. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md index 6b7e1f6ba..5d853b31f 100644 --- a/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md +++ b/qdrant-landing/content/documentation/faq/qdrant-fundamentals.md @@ -18,7 +18,7 @@ In dense vectors, Qdrant supports up to 65,535 dimensions. ### What is the maximum size of vector metadata that can be stored? -There is no inherent limitation on metadata size, but it should be [optimized for performance and resource usage](/documentation/guides/optimize/). Users can set upper limits in the configuration. +There is no inherent limitation on metadata size, but it should be [optimized for performance and resource usage](/documentation/operations/optimize/). Users can set upper limits in the configuration. ### Can the same similarity search query yield different results on different machines? @@ -30,7 +30,7 @@ This depends on the nature of your data and the specific application. Consider f ### How does Qdrant handle different vector embeddings from various providers in the same collection? -Qdrant natively [supports multiple vectors per data point](/documentation/concepts/vectors/#multivectors), allowing different embeddings from various providers to coexist within the same collection. +Qdrant natively [supports multiple vectors per data point](/documentation/manage-data/vectors/#multivectors), allowing different embeddings from various providers to coexist within the same collection. ### Can I migrate my embeddings from another vector store to Qdrant? @@ -46,14 +46,14 @@ Make sure to check that the collection status is `green` and that the number of ### Why collection info shows inaccurate number of points? Collection info API in Qdrant returns an approximate number of points in the collection. -If you need an exact number, you can use the [count](/documentation/concepts/points/#counting-points) API. +If you need an exact number, you can use the [count](/documentation/manage-data/points/#counting-points) API. ### Vectors in the collection don't match what I uploaded. There are two possible reasons for this: -- You used the `Cosine` distance metric in the [collection settings](/concepts/collections/#collections). In this case, Qdrant pre-normalizes your vectors for faster distance computation. If you strictly need the original vectors to be preserved, consider using the `Dot` distance metric instead. -- You used the `uint8` [datatype](/documentation/concepts/vectors/#datatypes) to store vectors. `uint8` requires a special format for input values, which might not be compatible with the typical output of embedding models. +- You used the `Cosine` distance metric in the [collection settings](/documentation/manage-data/collections/#collections). In this case, Qdrant pre-normalizes your vectors for faster distance computation. If you strictly need the original vectors to be preserved, consider using the `Dot` distance metric instead. +- You used the `uint8` [datatype](/documentation/manage-data/vectors/#datatypes) to store vectors. `uint8` requires a special format for input values, which might not be compatible with the typical output of embedding models. ## Search @@ -71,7 +71,7 @@ If you're still seeing `"vector": null` in your results, it might be that the ve ### How can I search without a vector? -You are likely looking for the [scroll](/documentation/concepts/points/#scroll-points) method. It allows you to retrieve the records based on filters or even iterate over all the records in the collection. +You are likely looking for the [scroll](/documentation/manage-data/points/#scroll-points) method. It allows you to retrieve the records based on filters or even iterate over all the records in the collection. ### Does Qdrant support a full-text search or a hybrid search? @@ -84,8 +84,8 @@ What Qdrant can do: - Apply full-text filters to the vector search (i.e., perform vector search among the records with specific words or phrases) - Do prefix search and semantic [search-as-you-type](/articles/search-as-you-type/) - Sparse vectors, as used in [SPLADE](https://github.com/naver/splade) or similar models -- [Multi-vectors](/documentation/concepts/vectors/#multivectors), for example ColBERT and other late-interaction models -- Combination of the [multiple searches](/documentation/concepts/hybrid-queries/) +- [Multi-vectors](/documentation/manage-data/vectors/#multivectors), for example ColBERT and other late-interaction models +- Combination of the [multiple searches](/documentation/search/hybrid-queries/) What Qdrant doesn't plan to support: @@ -105,11 +105,11 @@ It is _highly_ recommended not to create many small collections, as it will lead We consider creating a collection for each user/dialog/document as an antipattern. -Please read more about collections, isolation, and multiple users in our [Multitenancy](/documentation/tutorials/multiple-partitions/) tutorial. +Please read more about collections, isolation, and multiple users in our [Multitenancy](/documentation/manage-data/collections/#multitenancy) tutorial. ### How do I upload a large number of vectors into a Qdrant collection? -Read about our recommendations in the [bulk upload](/documentation/tutorials/bulk-upload/) tutorial. +Read about our recommendations in the [bulk upload](/documentation/tutorials-develop/bulk-upload/) tutorial. ### Can I only store quantized vectors and discard full precision vectors? @@ -144,7 +144,7 @@ You should always index first if you know your filters upfront. If you need to i ## Should I create one Qdrant collection per user? No. Creating one collection per user is more resource intensive. -Instead of creating separate collections for each user, we recommend creating a [single collection](https://qdrant.tech/documentation/guides/multiple-partitions/) and separate access using payloads. Each Qdrant point can have a payload as metadata. For multitenancy, you can include a `user_id` or `tenant_id` for each point. To optimize storage further, you can enable [tenant indexing](https://qdrant.tech/documentation/concepts/indexing/#tenant-index) for payload fields. +Instead of creating separate collections for each user, we recommend creating a [single collection](/documentation/manage-data/multitenancy/) and separate access using payloads. Each Qdrant point can have a payload as metadata. For multitenancy, you can include a `user_id` or `tenant_id` for each point. To optimize storage further, you can enable [tenant indexing](/documentation/manage-data/indexing/#tenant-index) for payload fields. ## Cloud diff --git a/qdrant-landing/content/documentation/fastembed/fastembed-colbert.md b/qdrant-landing/content/documentation/fastembed/fastembed-colbert.md index 08f8a1560..f6bddcb04 100644 --- a/qdrant-landing/content/documentation/fastembed/fastembed-colbert.md +++ b/qdrant-landing/content/documentation/fastembed/fastembed-colbert.md @@ -30,10 +30,10 @@ All interactions between these parts are expected to be done "later" outside the ## Using ColBERT in Qdrant -Qdrant supports [multivector representations](https://qdrant.tech/documentation/concepts/vectors/#multivectors) out of the box so that you can use any late interaction model as `ColBERT` or `ColPali` in Qdrant without any additional pre/post-processing. +Qdrant supports [multivector representations](/documentation/manage-data/vectors/#multivectors) out of the box so that you can use any late interaction model as `ColBERT` or `ColPali` in Qdrant without any additional pre/post-processing. This tutorial uses ColBERT as a first-stage retriever on a toy dataset. -You can see how to use ColBERT as a reranker in our [multi-stage queries documentation](https://qdrant.tech/documentation/concepts/hybrid-queries/#multi-stage-queries). +You can see how to use ColBERT as a reranker in our [multi-stage queries documentation](/documentation/search/hybrid-queries/#multi-stage-queries). ## Setup Install `fastembed`. @@ -144,8 +144,8 @@ from qdrant_client import QdrantClient, models qdrant_client = QdrantClient(":memory:") # Qdrant is running from RAM. ``` -Now, let's create a small [collection](https://qdrant.tech/documentation/concepts/collections/) with our movie data. -For that, we will use the [multivectors](https://qdrant.tech/documentation/concepts/vectors/#multivectors) functionality supported in Qdrant. +Now, let's create a small [collection](/documentation/manage-data/collections/) with our movie data. +For that, we will use the [multivectors](/documentation/manage-data/vectors/#multivectors) functionality supported in Qdrant. To configure multivector collection, we need to specify: - similarity metric between vectors; - the size of each vector (for ColBERT, it's **128**); diff --git a/qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md b/qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md index 4d7cb8815..b2cdd4ef1 100644 --- a/qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md +++ b/qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md @@ -13,7 +13,7 @@ $$ $$ A detailed breakdown of the idea behind miniCOIL can be found in the -["miniCOIL: on the road to Usable Sparse Neural Retreival" article](https://qdrant.tech/articles/minicoil/) or, in a [recorded talk "miniCOIL: Sparse Neural Retrieval Done Right"](https://youtu.be/f1sBJMSgBXA?si=G3C5--UVRKAW5WJ0). +["miniCOIL: on the road to Usable Sparse Neural Retrieval" article](https://qdrant.tech/articles/minicoil/) or, in a [recorded talk "miniCOIL: Sparse Neural Retrieval Done Right"](https://youtu.be/f1sBJMSgBXA?si=G3C5--UVRKAW5WJ0). This tutorial will demonstrate how miniCOIL-based sparse neural retrieval performs compared to BM25-based lexical retrieval. @@ -77,7 +77,7 @@ documents = [ ## Create Collection Let's create a collection to store and index titles. -As miniCOIL was designed with Qdrant's ability to calculate the keywords Inverse Document Frequency (IDF) in mind, we need to configure miniCOIL sparse vectors with [IDF modifier](https://qdrant.tech/documentation/concepts/indexing/#idf-modifier). +As miniCOIL was designed with Qdrant's ability to calculate the keywords Inverse Document Frequency (IDF) in mind, we need to configure miniCOIL sparse vectors with [IDF modifier](/documentation/manage-data/indexing/#idf-modifier).