new: replace .add and .query with local inference in fastembed semant… (#1575)

* new: replace .add and .query with local inference in fastembed semantic search

* do not hardcode model dim, format example to avoid scrollbars

* avoid scroll bars

* joint installation of qdrant client and fastembed
This commit is contained in:
George
2025-04-26 13:30:39 +03:00
committed by GitHub
parent d15bf50fd1
commit 5388a2cb76
@@ -5,21 +5,15 @@ weight: 3
# Using FastEmbed with Qdrant for Vector Search # Using FastEmbed with Qdrant for Vector Search
## Install Qdrant Client ## Install Qdrant Client and FastEmbed
```python ```python
pip install qdrant-client pip install "qdrant-client[fastembed]>=1.14.2"
```
## Install FastEmbed
Installing FastEmbed will let you quickly turn data to vectors, so that Qdrant can search over them.
```python
pip install fastembed
``` ```
## Initialize the client ## Initialize the client
Qdrant Client has a simple in-memory mode that lets you try semantic search locally. Qdrant Client has a simple in-memory mode that lets you try semantic search locally.
```python ```python
from qdrant_client import QdrantClient from qdrant_client import QdrantClient, models
client = QdrantClient(":memory:") # Qdrant is running from RAM. client = QdrantClient(":memory:") # Qdrant is running from RAM.
``` ```
@@ -28,21 +22,49 @@ client = QdrantClient(":memory:") # Qdrant is running from RAM.
Now you can add two sample documents, their associated metadata, and a point `id` for each. Now you can add two sample documents, their associated metadata, and a point `id` for each.
```python ```python
docs = ["Qdrant has a LangChain integration for chatbots.", "Qdrant has a LlamaIndex integration for agents."] docs = [
"Qdrant has a LangChain integration for chatbots.",
"Qdrant has a LlamaIndex integration for agents.",
]
metadata = [ metadata = [
{"source": "langchain-docs"}, {"source": "langchain-docs"},
{"source": "llamaindex-docs"}, {"source": "llamaindex-docs"},
] ]
ids = [42, 2] ids = [42, 2]
``` ```
## Load data to a collection ## Create a collection
Create a test collection and upsert your two documents to it.
Qdrant stores vectors and associated metadata in collections.
Collection requires vector parameters to be set during creation.
In this tutorial, we'll be using `BAAI/bge-small-en` to compute embeddings.
```python ```python
client.add( model_name = "BAAI/bge-small-en"
client.create_collection(
collection_name="test_collection", collection_name="test_collection",
documents=docs, vectors_config=models.VectorParams(
metadata=metadata, size=client.get_embedding_size(model_name),
ids=ids distance=models.Distance.COSINE
), # size and distance are model dependent
)
```
## Upsert documents to the collection
Qdrant client can do inference implicitly within its methods via FastEmbed integration.
It requires wrapping your data in models, like `models.Document` (or `models.Image` if you're working with images)
```python
metadata_with_docs = [
{"document": doc, "source": meta["source"]} for doc, meta in zip(docs, metadata)
]
client.upload_collection(
collection_name="test_collection",
vectors=[models.Document(text=doc, model=model_name) for doc in docs],
payload=metadata_with_docs,
ids=ids,
) )
``` ```
## Run vector search ## Run vector search
@@ -50,21 +72,36 @@ client.add(
Here, you will ask a dummy question that will allow you to retrieve a semantically relevant result. Here, you will ask a dummy question that will allow you to retrieve a semantically relevant result.
```python ```python
search_result = client.query( search_result = client.query_points(
collection_name="test_collection", collection_name="test_collection",
query_text="Which integration is best for agents?" query=models.Document(
) text="Which integration is best for agents?",
model=model_name
)
).points
print(search_result) print(search_result)
``` ```
The semantic search engine will retrieve the most similar result in order of relevance. In this case, the second statement about LlamaIndex is more relevant. The semantic search engine will retrieve the most similar result in order of relevance. In this case, the second statement about LlamaIndex is more relevant.
```bash ```python
[QueryResponse(id=2, embedding=None, sparse_embedding=None, [
metadata={'document': 'Qdrant has a LlamaIndex integration for agents', ScoredPoint(
'source': 'llamaindex-docs'}, document='Qdrant has a LlamaIndex integration for agents.', id=2,
score=0.8749180370667156), score=0.87491801319731,
QueryResponse(id=42, embedding=None, sparse_embedding=None, payload={
metadata={'document': 'Qdrant has a LangChain integration for chatbots.', "document": "Qdrant has a LlamaIndex integration for agents.",
'source': 'langchain-docs'}, document='Qdrant has a LangChain integration for chatbots.', "source": "llamaindex-docs",
score=0.8351846822959111)] },
...
),
ScoredPoint(
id=42,
score=0.8351846627714035,
payload={
"document": "Qdrant has a LangChain integration for chatbots.",
"source": "langchain-docs",
},
...
),
]
``` ```