mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-06 19:38:30 +02:00
minicoil tutorial
This commit is contained in:
+1
@@ -0,0 +1 @@
|
||||
This code snippet creates a collection configured for BM25-based lexical retrieval. It defines a sparse named vector with the IDF modifier enabled,ensuring that the Inverse Document Frequency, a core component of the BM25 scoring formula, is calculated on Qdrant’s side. This setup allows Qdrant to perform BM25-style retrieval based on keyword frequency and rarity.
|
||||
+10
@@ -0,0 +1,10 @@
|
||||
```python
|
||||
client.create_collection(
|
||||
collection_name="{bm25_collection_name}",
|
||||
sparse_vectors_config={
|
||||
"bm25": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF
|
||||
)
|
||||
}
|
||||
)
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet performs a search in a collection configured for BM25 sparse vectors using the Qdrant and FastEmbed integration. It infers a sparse BM25 vector for the query and retrieves the most relevant document (limit=1) based on the BM25 scoring. The response includes the top-matching document’s ID, score, and payload.
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
```python
|
||||
query = "Vectors in Medicine"
|
||||
|
||||
client.query_points(
|
||||
collection_name="{bm25_collection_name}",
|
||||
query=models.Document(
|
||||
text=query,
|
||||
model="Qdrant/bm25"
|
||||
),
|
||||
using="bm25",
|
||||
limit=1,
|
||||
)
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet uses the Qdrant and FastEmbed integration to infer and upsert documents into a collection configured for BM25 sparse vectors. Each document is converted into a sparse BM25 vector, with the conversion incorporating the avg_len parameter of the BM25 scoring formula — the average document length in the corpus — which must be provided by the user. The resulting vector and the document’s text, stored as payload, are then upserted to the collection.
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
```python
|
||||
#Estimating the average length of the documents in the corpus
|
||||
avg_documents_length = sum(len(document.split()) for document in documents) / len(documents)
|
||||
|
||||
client.upsert(
|
||||
collection_name="{bm25_collection_name}",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=i,
|
||||
payload={
|
||||
"text": documents[i]
|
||||
},
|
||||
vector={
|
||||
# Sparse vector from BM25
|
||||
"bm25": models.Document(
|
||||
text=documents[i],
|
||||
model="Qdrant/bm25",
|
||||
options={"avg_len": avg_documents_length}
|
||||
#Average length of documents in the corpus
|
||||
# (a part of the BM25 formula)
|
||||
)
|
||||
},
|
||||
)
|
||||
for i in range(len(documents))
|
||||
],
|
||||
)
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet creates a collection configured to store and index documents using miniCOIL sparse neural retrieval. It defines a sparse named vector with the IDF modifier enabled, so that the Inverse Document Frequency is calculated and applied to keywords as required by the miniCOIL scoring formula. This setup allows Qdrant to use miniCOIL’s BM25-inspired retrieval with contextual semantic weighting of matched keywords.
|
||||
+10
@@ -0,0 +1,10 @@
|
||||
```python
|
||||
client.create_collection(
|
||||
collection_name="{minicoil_collection_name}",
|
||||
sparse_vectors_config={
|
||||
"minicoil": models.SparseVectorParams(
|
||||
modifier=models.Modifier.IDF #Inverse Document Frequency
|
||||
)
|
||||
}
|
||||
)
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet performs a search in a collection configured for miniCOIL sparse vectors using the Qdrant and FastEmbed integration. It infers a sparse miniCOIL vector for the query and retrieves the most relevant document (limit=1) based on the miniCOIL scoring. The response includes the top-matching document’s ID, score, and payload.
|
||||
+13
@@ -0,0 +1,13 @@
|
||||
```python
|
||||
query = "Vectors in Medicine"
|
||||
|
||||
client.query_points(
|
||||
collection_name="{minicoil_collection_name}",
|
||||
query=models.Document(
|
||||
text=query,
|
||||
model="Qdrant/minicoil-v1"
|
||||
),
|
||||
using="minicoil",
|
||||
limit=1
|
||||
)
|
||||
```
|
||||
+1
@@ -0,0 +1 @@
|
||||
This code snippet uses the Qdrant and FastEmbed integration to infer and upsert documents into a collection configured for miniCOIL sparse vectors. Each document is converted into a sparse miniCOIL vector, with the conversion incorporating the avg_len parameter of the BM25-based miniCOIL scoring formula — the average document length in the corpus — which must be provided by the user. The resulting vector and the document’s text, stored as payload, are then upserted to the collection.
|
||||
+27
@@ -0,0 +1,27 @@
|
||||
```python
|
||||
#Estimating the average length of the documents in the corpus
|
||||
avg_documents_length = sum(len(document.split()) for document in documents) / len(documents)
|
||||
|
||||
client.upsert(
|
||||
collection_name="{minicoil_collection_name}",
|
||||
points=[
|
||||
models.PointStruct(
|
||||
id=i,
|
||||
payload={
|
||||
"text": documents[i]
|
||||
},
|
||||
vector={
|
||||
# Sparse miniCOIL vectors
|
||||
"minicoil": models.Document(
|
||||
text=documents[i],
|
||||
model="Qdrant/minicoil-v1",
|
||||
options={"avg_len": avg_documents_length}
|
||||
#Average length of documents in the corpus
|
||||
# (a part of the BM25 formula on which miniCOIL is built)
|
||||
)
|
||||
},
|
||||
)
|
||||
for i in range(len(documents))
|
||||
],
|
||||
)
|
||||
```
|
||||
Reference in New Issue
Block a user