From 3d318f07653591019372f39ae96bbf179d938224 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?=D0=95=D0=B2=D0=B3=D0=B5=D0=BD=D0=B8=D1=8F=20=D0=A1=D1=83?= =?UTF-8?q?=D1=85=D0=BE=D0=B4=D0=BE=D0=BB=D1=8C=D1=81=D0=BA=D0=B0=D1=8F?= Date: Fri, 4 Jul 2025 11:44:17 +0200 Subject: [PATCH] minicoil tutorial --- .../content/documentation/fastembed/_index.md | 15 +- .../fastembed/fastembed-minicoil.md | 160 ++++++++++++++++++ .../bm25/create-collection/_description.md | 1 + .../bm25/create-collection/python.md | 10 ++ .../bm25/query-points/_description.md | 1 + .../fastembed/bm25/query-points/python.md | 13 ++ .../bm25/upsert-points/_description.md | 1 + .../fastembed/bm25/upsert-points/python.md | 27 +++ .../create-collection/_description.md | 1 + .../minicoil/create-collection/python.md | 10 ++ .../minicoil/query-points/_description.md | 1 + .../fastembed/minicoil/query-points/python.md | 13 ++ .../minicoil/upsert-points/_description.md | 1 + .../minicoil/upsert-points/python.md | 27 +++ 14 files changed, 275 insertions(+), 6 deletions(-) create mode 100644 qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/_description.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/python.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/_description.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/python.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/_description.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/python.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/_description.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/python.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/_description.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/python.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/_description.md create mode 100644 qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/python.md diff --git a/qdrant-landing/content/documentation/fastembed/_index.md b/qdrant-landing/content/documentation/fastembed/_index.md index 837dc0ac2..f50374b81 100644 --- a/qdrant-landing/content/documentation/fastembed/_index.md +++ b/qdrant-landing/content/documentation/fastembed/_index.md @@ -11,11 +11,16 @@ By using FastEmbed, you can ensure that your embedding generation process is not FastEmbed easily integrates with Qdrant for a variety of multimodal search purposes. -## How to get started with FastEmbed +## Using FastEmbed -|Beginner|Advanced| -|:-:|:-:| -|[Generate Text Embedings with FastEmbed](/documentation/fastembed/fastembed-quickstart/)|[Combine FastEmbed with Qdrant for Vector Search](/documentation/fastembed/fastembed-semantic-search/)| +| Type | Guide | What you'll learn | +|---|-------|--------------------| +| **Beginner** | [Generate Text Embeddings](/documentation/fastembed/fastembed-quickstart/) | Install FastEmbed and generate dense text embeddings | +| | [Dense Embeddings + Qdrant](/documentation/fastembed/fastembed-semantic-search/) | Generate and index dense embeddings for semantic similarity search | +| **Advanced** | [miniCOIL Sparse Embeddings + Qdrant](/documentation/fastembed/fastembed-minicoil/) | Use Qdrant's sparse neural retriever for exact text search | +| | [SPLADE Sparse Embeddings + Qdrant](/documentation/fastembed/fastembed-splade/) | Generate sparse neural embeddings for exact text search | +| | [ColBERT Multivector Embeddings + Qdrant](/documentation/fastembed/fastembed-colbert/) | Generate and index multi-vector representations; **ideal for rescoring, or small-scale retrieval** | +| | [Reranking with FastEmbed](/documentation/fastembed/fastembed-rerankers/) | Re-rank top-K results using FastEmbed cross-encoders | ## Why is FastEmbed useful? @@ -23,5 +28,3 @@ FastEmbed easily integrates with Qdrant for a variety of multimodal search purpo - Fast: By using ONNX, FastEmbed ensures high-performance inference across various hardware platforms. - Accurate: FastEmbed aims for better accuracy and recall than models like OpenAI's `Ada-002`. It always uses model which demonstrate strong results on the MTEB leaderboard. - Support: FastEmbed supports a wide range of models, including multilingual ones, to meet diverse use case needs. - - diff --git a/qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md b/qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md new file mode 100644 index 000000000..f4f643c54 --- /dev/null +++ b/qdrant-landing/content/documentation/fastembed/fastembed-minicoil.md @@ -0,0 +1,160 @@ +--- +title: Working with miniCOIL +weight: 4 +--- + +# How to use miniCOIL, Qdrant's Sparse Neural Retriever + +**miniCOIL** is an open-sourced sparse neural retrieval model that acts as if a BM25-based retriever understood the contextual meaning of keywords and ranked results accordingly. + +**miniCOIL** scoring is based on the BM25 formula scaled by the semantic similarity between matched keywords in a query and a document. +$$ +\text{miniCOIL}(D,Q) = \sum_{i=1}^{N} \text{IDF}(q_i) \cdot \text{Importance}^{q_i}_{D} \cdot {\color{YellowGreen}\text{Meaning}^{q_i \times d_j}} \text{, where keyword } d_j \in D \text{ equals } q_i +$$ + +A detailed breakdown of the idea behind miniCOIL can be found in the +["miniCOIL: on the road to Usable Sparse Neural Retreival" article](https://qdrant.tech/articles/minicoil/) or, in a [recorded talk "miniCOIL: Sparse Neural Retrieval Done Right"](https://youtu.be/f1sBJMSgBXA?si=G3C5--UVRKAW5WJ0). + +This tutorial will demonstrate how miniCOIL-based sparse neural retrieval performs compared to BM25-based lexical retrieval. + +## When to use miniCOIL + +When exact keyword matches in the retrieved results are a requirement, and all matches should be ranked based on the contextual meaning of keywords. + +If results should be similar by meaning but are expressed differently, with no overlapping keywords, you should use dense embeddings or combine them with miniCOIL in a hybrid search setting. + +## Setup + +Install `qdrant-client` integration with `fastembed`. + +```python +pip install "qdrant-client[fastembed]" +``` + +Then, initialize the Qdrant client. You could use for experiments [a free cluster](https://qdrant.tech/documentation/cloud-quickstart/#authenticate-via-sdks) in Qdrant Cloud or run a [local Qdrant instance via Docker](https://qdrant.tech/documentation/quickstart/#initialize-the-client). + +We'll run our search on a list of book and article titles containing the keywords "*vector*" and "*search*" used in different contexts, to demonstrate how miniCOIL captures the meaning of these keywords as opposed to BM25. + +
+ A dataset + +```python +documents = [ + "Vector Graphics in Modern Web Design", + "The Art of Search and Self-Discovery", + "Efficient Vector Search Algorithms for Large Datasets", + "Searching the Soul: A Journey Through Mindfulness", + "Vector-Based Animations for User Interface Design", + "Search Engines: A Technical and Social Overview", + "The Rise of Vector Databases in AI Systems", + "Search Patterns in Human Behavior", + "Vector Illustrations: A Guide for Creatives", + "Search and Rescue: Technologies in Emergency Response", + "Vectors in Physics: From Arrows to Equations", + "Searching for Lost Time in the Digital Age", + "Vector Spaces and Linear Transformations", + "The Endless Search for Truth in Philosophy", + "3D Modeling with Vectors in Blender", + "Search Optimization Strategies for E-commerce", + "Vector Drawing Techniques with Open-Source Tools", + "In Search of Meaning: A Psychological Perspective", + "Advanced Vector Calculus for Engineers", + "Search Interfaces: UX Principles and Case Studies", + "The Use of Vector Fields in Meteorology", + "Search and Destroy: Cybersecurity in the 21st Century", + "From Bitmap to Vector: A Designer’s Guide", + "Search Engines and the Democratization of Knowledge", + "Vector Geometry in Game Development", + "The Human Search for Connection in a Digital World", + "AI-Powered Vector Search in Recommendation Systems", + "Searchable Archives: The History of Digital Retrieval", + "Vector Control Strategies in Public Health", + "The Search for Extraterrestrial Intelligence" +] +``` +
+ +## Create Collection +Let's create a collection to store and index titles. + +As miniCOIL was designed with Qdrant's ability to calculate the keywords Inverse Document Frequency (IDF) in mind, we need to configure miniCOIL sparse vectors with [IDF modifier](https://qdrant.tech/documentation/concepts/indexing/#idf-modifier). + + + +{{< code-snippet path="/documentation/headless/snippets/fastembed/minicoil/create-collection" >}} + +
+ Analogously, we configure a collection with BM25-based sparse vectors + +{{< code-snippet path="/documentation/headless/snippets/fastembed/bm25/create-collection" >}} + +
+ +## Convert to Sparse Vectors & Upload to Qdrant + +Next, we need to convert titles to miniCOIL sparse representations and upsert them into the configured collection. + +Qdrant and FastEmbed integration allows for hiding the inference process under the hood. + +That means: + +- FastEmbed downloads the selected model from Hugging Face; +- FastEmbed runs local inference under the hood; +- Inferenced sparse representations are uploaded to Qdrant. + +{{< code-snippet path="/documentation/headless/snippets/fastembed/minicoil/upsert-points" >}} + +
+ Analogously, we convert & upsert BM25-based sparse vectors + +{{< code-snippet path="/documentation/headless/snippets/fastembed/bm25/upsert-points" >}} + +
+ +## Retrieve with miniCOIL +Using query *"Vectors in Medicine"*, we'll demo the difference between miniCOIL and BM25-based retrieval. + +None of the indexed titles contain the keyword *"medicine"*, so it won't contribute to the similarity score. +At the same time, the word *"vector"* appears once in many titles, and its role is roughly equal in all of them from the perspective of the BM25-based retriever. +miniCOIL, however, can capture the meaning of the keyword *"vector"* in the context of *"medicine"* and match a document where *"vector"* is used in a medicine-related context. + +For BM25-based retrieval: + +{{< code-snippet path="/documentation/headless/snippets/fastembed/bm25/query-points" >}} + +Result will be: + +```bash +QueryResponse( + points=[ + ScoredPoint( + id=18, version=1, score=0.8405092, + payload={ + 'title': 'Advanced Vector Calculus for Engineers' + }, + vector=None, shard_key=None, order_value=None) + ] + ) +``` + +While for miniCOIL-based retrieval: + +{{< code-snippet path="/documentation/headless/snippets/fastembed/minicoil/query-points" >}} + +We will get: + +```bash +QueryResponse( + points=[ + ScoredPoint( + id=28, version=1, score=0.7005557, + payload={ + 'title': 'Vector Control Strategies in Public Health' + }, + vector=None, shard_key=None, order_value=None) + ] + ) +``` + diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/_description.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/_description.md new file mode 100644 index 000000000..4b18cbf45 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/_description.md @@ -0,0 +1 @@ +This code snippet creates a collection configured for BM25-based lexical retrieval. It defines a sparse named vector with the IDF modifier enabled,ensuring that the Inverse Document Frequency, a core component of the BM25 scoring formula, is calculated on Qdrant’s side. This setup allows Qdrant to perform BM25-style retrieval based on keyword frequency and rarity. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/python.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/python.md new file mode 100644 index 000000000..b7d3997c3 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/create-collection/python.md @@ -0,0 +1,10 @@ +```python +client.create_collection( + collection_name="{bm25_collection_name}", + sparse_vectors_config={ + "bm25": models.SparseVectorParams( + modifier=models.Modifier.IDF + ) + } +) +``` \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/_description.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/_description.md new file mode 100644 index 000000000..8a60fcea4 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/_description.md @@ -0,0 +1 @@ +This code snippet performs a search in a collection configured for BM25 sparse vectors using the Qdrant and FastEmbed integration. It infers a sparse BM25 vector for the query and retrieves the most relevant document (limit=1) based on the BM25 scoring. The response includes the top-matching document’s ID, score, and payload. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/python.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/python.md new file mode 100644 index 000000000..e157e8eed --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/query-points/python.md @@ -0,0 +1,13 @@ +```python +query = "Vectors in Medicine" + +client.query_points( + collection_name="{bm25_collection_name}", + query=models.Document( + text=query, + model="Qdrant/bm25" + ), + using="bm25", + limit=1, +) +``` \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/_description.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/_description.md new file mode 100644 index 000000000..fbde049c7 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/_description.md @@ -0,0 +1 @@ +This code snippet uses the Qdrant and FastEmbed integration to infer and upsert documents into a collection configured for BM25 sparse vectors. Each document is converted into a sparse BM25 vector, with the conversion incorporating the avg_len parameter of the BM25 scoring formula — the average document length in the corpus — which must be provided by the user. The resulting vector and the document’s text, stored as payload, are then upserted to the collection. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/python.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/python.md new file mode 100644 index 000000000..59b60aac4 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/bm25/upsert-points/python.md @@ -0,0 +1,27 @@ +```python +#Estimating the average length of the documents in the corpus +avg_documents_length = sum(len(document.split()) for document in documents) / len(documents) + +client.upsert( + collection_name="{bm25_collection_name}", + points=[ + models.PointStruct( + id=i, + payload={ + "text": documents[i] + }, + vector={ + # Sparse vector from BM25 + "bm25": models.Document( + text=documents[i], + model="Qdrant/bm25", + options={"avg_len": avg_documents_length} + #Average length of documents in the corpus + # (a part of the BM25 formula) + ) + }, + ) + for i in range(len(documents)) + ], +) +``` \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/_description.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/_description.md new file mode 100644 index 000000000..3520579af --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/_description.md @@ -0,0 +1 @@ +This code snippet creates a collection configured to store and index documents using miniCOIL sparse neural retrieval. It defines a sparse named vector with the IDF modifier enabled, so that the Inverse Document Frequency is calculated and applied to keywords as required by the miniCOIL scoring formula. This setup allows Qdrant to use miniCOIL’s BM25-inspired retrieval with contextual semantic weighting of matched keywords. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/python.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/python.md new file mode 100644 index 000000000..11c5d8e96 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/create-collection/python.md @@ -0,0 +1,10 @@ +```python +client.create_collection( + collection_name="{minicoil_collection_name}", + sparse_vectors_config={ + "minicoil": models.SparseVectorParams( + modifier=models.Modifier.IDF #Inverse Document Frequency + ) + } +) +``` \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/_description.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/_description.md new file mode 100644 index 000000000..0c12cbfc8 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/_description.md @@ -0,0 +1 @@ +This code snippet performs a search in a collection configured for miniCOIL sparse vectors using the Qdrant and FastEmbed integration. It infers a sparse miniCOIL vector for the query and retrieves the most relevant document (limit=1) based on the miniCOIL scoring. The response includes the top-matching document’s ID, score, and payload. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/python.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/python.md new file mode 100644 index 000000000..0638f81c9 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/query-points/python.md @@ -0,0 +1,13 @@ +```python +query = "Vectors in Medicine" + +client.query_points( + collection_name="{minicoil_collection_name}", + query=models.Document( + text=query, + model="Qdrant/minicoil-v1" + ), + using="minicoil", + limit=1 +) +``` \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/_description.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/_description.md new file mode 100644 index 000000000..a2393cc8c --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/_description.md @@ -0,0 +1 @@ +This code snippet uses the Qdrant and FastEmbed integration to infer and upsert documents into a collection configured for miniCOIL sparse vectors. Each document is converted into a sparse miniCOIL vector, with the conversion incorporating the avg_len parameter of the BM25-based miniCOIL scoring formula — the average document length in the corpus — which must be provided by the user. The resulting vector and the document’s text, stored as payload, are then upserted to the collection. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/python.md b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/python.md new file mode 100644 index 000000000..1735f3555 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/fastembed/minicoil/upsert-points/python.md @@ -0,0 +1,27 @@ +```python +#Estimating the average length of the documents in the corpus +avg_documents_length = sum(len(document.split()) for document in documents) / len(documents) + +client.upsert( + collection_name="{minicoil_collection_name}", + points=[ + models.PointStruct( + id=i, + payload={ + "text": documents[i] + }, + vector={ + # Sparse miniCOIL vectors + "minicoil": models.Document( + text=documents[i], + model="Qdrant/minicoil-v1", + options={"avg_len": avg_documents_length} + #Average length of documents in the corpus + # (a part of the BM25 formula on which miniCOIL is built) + ) + }, + ) + for i in range(len(documents)) + ], +) +``` \ No newline at end of file