Add guide about optimizing FastEmbed throughput (#2195)

* Add guide about optimizing FastEmbed throughput

* Update weights

* Review feedback
This commit is contained in:
Abdon Pijpelink
2026-03-23 16:53:06 +01:00
committed by GitHub
parent ea0de4c12e
commit e9a691a15a
19 changed files with 252 additions and 7 deletions
@@ -1,6 +1,6 @@
---
title: Working with ColBERT
weight: 6
weight: 60
---
# How to Generate ColBERT Multivectors with FastEmbed
@@ -1,6 +1,6 @@
---
title: Working with miniCOIL
weight: 4
weight: 40
---
# How to use miniCOIL, Qdrant's Sparse Neural Retriever
@@ -0,0 +1,58 @@
---
title: "Optimize Throughput"
weight: 30
---
# Optimize FastEmbed Throughput
By default, FastEmbed processes documents sequentially in the main processing thread. To optimize throughput, FastEmbed supports processing documents in parallel.
When parallel processing is enabled, FastEmbed splits a dataset across multiple workers, each running an independent copy of the embedding model. Internally, documents are split into batches and put on a shared input queue. Each batch is then processed by one of the workers, put on a shared output queue, and then collected and reordered to match the original input order.
To configure FastEmbed for parallel processing, use the following parameters:
- `parallel`: the number of workers.
- When set to `None` (default), the embedding model runs in the main process.
- When set to `0`, FastEmbed detects the number of available CPU cores and parallelizes across that many workers.
- When set to `1` or higher, FastEmbed uses the specified number of workers.
- Batch size: the number of documents that each worker processes in each batch. Adjusting this to balance memory usage and processing speed. Lower it if you're running out of memory during local inference. Raise it to improve throughput if you have plenty of memory and are processing large document sets.
To configure the batch size, use:
- `batch_size`: stand-alone FastEmbed parameter for batch processing. Defaults to `256` for text and `16` for images.
- `local_inference_batch_size`: Qdrant Client parameter for batch processing. Defaults to `8`.
- `lazy_load`: set to `True` to avoid loading the embedding model until it's needed for inference. Enabling lazy loading prevents loading the model in the main process when using multiple workers, which saves memory and reduces startup time.
## Parallelize FastEmbed with the Qdrant Client
When using FastEmbed with Qdrant Client, specify the `local_inference_batch_size` parameter when initializing the client to configure the batch size. For example:
{{< code-snippet path="/documentation/headless/snippets/fastembed/optimize/qdrant-client/" block="client-connection" >}}
Next, when creating points, set `lazy_load` to `True` in the inference object to avoid loading the embedding model in the main process:
{{< code-snippet path="/documentation/headless/snippets/fastembed/optimize/qdrant-client/" block="lazy-load" >}}
When using [`fastembed-gpu`](https://qdrant.github.io/fastembed/examples/FastEmbed_GPU/), also set `cuda` to `True` to enable GPU acceleration:
{{< code-snippet path="/documentation/headless/snippets/fastembed/optimize/qdrant-client/" block="lazy-load-gpu" >}}
Finally, when uploading points, set the `parallel` parameter to the desired number of workers:
{{< code-snippet path="/documentation/headless/snippets/fastembed/optimize/qdrant-client/" block="upload-data" >}}
## Parallelize Standalone FastEmbed
When using FastEmbed as a standalone library, first enable lazy loading of the embedding model:
{{< code-snippet path="/documentation/headless/snippets/fastembed/optimize/standalone/" block="lazy-load" >}}
FastEmbed supports [distributing the workload across multiple GPU devices](https://qdrant.github.io/fastembed/examples/FastEmbed_GPU/). To enable this:
- Install `fastembed-gpu`.
- Set `cuda` to `True` to enable GPU acceleration.
- Configure `device_ids` with a list of GPU device IDs to assign workers to. For example, `device_ids=[0, 1]` assigns workers to GPUs 0 and 1. If not specified, FastEmbed will assign all workers to the default GPU device.
{{< code-snippet path="/documentation/headless/snippets/fastembed/optimize/standalone/" block="lazy-load-gpu" >}}
When generating embeddings, set the batch size with `batch_size` and the number of workers with `parallel`:
{{< code-snippet path="/documentation/headless/snippets/fastembed/optimize/standalone/" block="embed" >}}
@@ -1,6 +1,6 @@
---
title: Multi-Vector Postprocessing
weight: 9
weight: 90
---
# Multi-Vector Postprocessing
@@ -1,6 +1,6 @@
---
title: "Quickstart"
weight: 2
weight: 10
---
# How to Generate Text Embedings with FastEmbed
@@ -1,6 +1,6 @@
---
title: Reranking with FastEmbed
weight: 8
weight: 80
---
# How to use rerankers with FastEmbed
@@ -1,6 +1,6 @@
---
title: "FastEmbed & Qdrant"
weight: 3
weight: 20
---
# Using FastEmbed with Qdrant for Vector Search
@@ -1,6 +1,6 @@
---
title: Working with SPLADE
weight: 5
weight: 50
---
# How to Generate Sparse Vectors with SPLADE
@@ -0,0 +1,7 @@
```python
client = QdrantClient(
url=QDRANT_URL,
api_key=QDRANT_API_KEY,
local_inference_batch_size=256, # FastEmbed batch size
)
```
@@ -0,0 +1,13 @@
```python
point = models.PointStruct(
id=1,
vector=models.Document(
text="The text to embed",
model="BAAI/bge-small-en-v1.5",
options={
"lazy_load": True,
"cuda": True,
},
)
)
```
@@ -0,0 +1,12 @@
```python
point = models.PointStruct(
id=1,
vector=models.Document(
text="The text to embed",
model="BAAI/bge-small-en-v1.5",
options={
"lazy_load": True,
},
)
)
```
@@ -0,0 +1,38 @@
```python
from qdrant_client import QdrantClient, models
client = QdrantClient(
url=QDRANT_URL,
api_key=QDRANT_API_KEY,
local_inference_batch_size=256, # FastEmbed batch size
)
point = models.PointStruct(
id=1,
vector=models.Document(
text="The text to embed",
model="BAAI/bge-small-en-v1.5",
options={
"lazy_load": True,
},
)
)
point = models.PointStruct(
id=1,
vector=models.Document(
text="The text to embed",
model="BAAI/bge-small-en-v1.5",
options={
"lazy_load": True,
"cuda": True,
},
)
)
client.upload_points(
collection_name=COLLECTION_NAME,
points=points,
parallel=4 # use 4 workers to process documents in parallel
)
```
@@ -0,0 +1,7 @@
```python
client.upload_points(
collection_name=COLLECTION_NAME,
points=points,
parallel=4 # use 4 workers to process documents in parallel
)
```
@@ -0,0 +1,51 @@
from qdrant_client import QdrantClient, models
# @hide-start
QDRANT_URL=""
QDRANT_API_KEY=""
points: list[models.PointStruct] = []
COLLECTION_NAME=""
# @hide-end
# @block-start client-connection
client = QdrantClient(
url=QDRANT_URL,
api_key=QDRANT_API_KEY,
local_inference_batch_size=256, # FastEmbed batch size
)
# @block-end client-connection
# @block-start lazy-load
point = models.PointStruct(
id=1,
vector=models.Document(
text="The text to embed",
model="BAAI/bge-small-en-v1.5",
options={
"lazy_load": True,
},
)
)
# @block-end lazy-load
# @block-start lazy-load-gpu
point = models.PointStruct(
id=1,
vector=models.Document(
text="The text to embed",
model="BAAI/bge-small-en-v1.5",
options={
"lazy_load": True,
"cuda": True,
},
)
)
# @block-end lazy-load-gpu
# @block-start upload-data
client.upload_points(
collection_name=COLLECTION_NAME,
points=points,
parallel=4 # use 4 workers to process documents in parallel
)
# @block-end upload-data
@@ -0,0 +1,3 @@
```python
embeddings = list(model.embed(docs, batch_size=256, parallel=4))
```
@@ -0,0 +1,8 @@
```python
model = TextEmbedding(
model_name="BAAI/bge-small-en-v1.5",
lazy_load=True, # don't load the model until first embed call
cuda=True, # enable GPU acceleration
device_ids=[0, 1], # spread workers across GPUs 0 and 1
)
```
@@ -0,0 +1,6 @@
```python
model = TextEmbedding(
model_name="BAAI/bge-small-en-v1.5",
lazy_load=True, # don't load the model until first embed call
)
```
@@ -0,0 +1,17 @@
```python
from fastembed import TextEmbedding
model = TextEmbedding(
model_name="BAAI/bge-small-en-v1.5",
lazy_load=True, # don't load the model until first embed call
)
model = TextEmbedding(
model_name="BAAI/bge-small-en-v1.5",
lazy_load=True, # don't load the model until first embed call
cuda=True, # enable GPU acceleration
device_ids=[0, 1], # spread workers across GPUs 0 and 1
)
embeddings = list(model.embed(docs, batch_size=256, parallel=4))
```
@@ -0,0 +1,25 @@
from fastembed import TextEmbedding
# @block-start lazy-load
model = TextEmbedding(
model_name="BAAI/bge-small-en-v1.5",
lazy_load=True, # don't load the model until first embed call
)
# @block-end lazy-load
# @block-start lazy-load-gpu
model = TextEmbedding(
model_name="BAAI/bge-small-en-v1.5",
lazy_load=True, # don't load the model until first embed call
cuda=True, # enable GPU acceleration
device_ids=[0, 1], # spread workers across GPUs 0 and 1
)
# @block-end lazy-load-gpu
# @hide-start
docs = ["", "", ""]
# @hide-end
# @block-start embed
embeddings = list(model.embed(docs, batch_size=256, parallel=4))
# @block-end embed