minicoil tutorial

This commit is contained in:
Евгения Суходольская
2025-07-04 11:44:17 +02:00
parent 309d7151da
commit 3d318f0765
14 changed files with 275 additions and 6 deletions
@@ -0,0 +1 @@
This code snippet creates a collection configured for BM25-based lexical retrieval. It defines a sparse named vector with the IDF modifier enabled,ensuring that the Inverse Document Frequency, a core component of the BM25 scoring formula, is calculated on Qdrant’s side. This setup allows Qdrant to perform BM25-style retrieval based on keyword frequency and rarity.
@@ -0,0 +1,10 @@
```python
client.create_collection(
collection_name="{bm25_collection_name}",
sparse_vectors_config={
"bm25": models.SparseVectorParams(
modifier=models.Modifier.IDF
)
}
)
```
@@ -0,0 +1 @@
This code snippet performs a search in a collection configured for BM25 sparse vectors using the Qdrant and FastEmbed integration. It infers a sparse BM25 vector for the query and retrieves the most relevant document (limit=1) based on the BM25 scoring. The response includes the top-matching document’s ID, score, and payload.
@@ -0,0 +1,13 @@
```python
query = "Vectors in Medicine"
client.query_points(
collection_name="{bm25_collection_name}",
query=models.Document(
text=query,
model="Qdrant/bm25"
),
using="bm25",
limit=1,
)
```
@@ -0,0 +1 @@
This code snippet uses the Qdrant and FastEmbed integration to infer and upsert documents into a collection configured for BM25 sparse vectors. Each document is converted into a sparse BM25 vector, with the conversion incorporating the avg_len parameter of the BM25 scoring formula — the average document length in the corpus — which must be provided by the user. The resulting vector and the document’s text, stored as payload, are then upserted to the collection.
@@ -0,0 +1,27 @@
```python
#Estimating the average length of the documents in the corpus
avg_documents_length = sum(len(document.split()) for document in documents) / len(documents)
client.upsert(
collection_name="{bm25_collection_name}",
points=[
models.PointStruct(
id=i,
payload={
"text": documents[i]
},
vector={
# Sparse vector from BM25
"bm25": models.Document(
text=documents[i],
model="Qdrant/bm25",
options={"avg_len": avg_documents_length}
#Average length of documents in the corpus
# (a part of the BM25 formula)
)
},
)
for i in range(len(documents))
],
)
```