* Break Inference page into several pages * Make all inference code snippets testable and clean up * Make more snippets testable * Edits * Document automatic query and passage prefix injection in Cloud Inference Qdrant Cloud Inference silently applies model-specific prefixes (e.g. "query: "/"passage: " for E5, BGE-style instruction prefix for BGE/mxbai/ Snowflake arctic-embed) so users don't need to manage them manually. Add a section explaining this behavior, the idempotency guarantee, and the scope (Qdrant-hosted models only; external providers handle their own). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * Document short query optimization in Cloud Inference Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * Update links * Expand on external provider API key usage * Add section about external provider API keys * Default to header for external API keys --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
1.3 KiB
title, weight
| title | weight |
|---|---|
| BM25 | 20 |
Server-side Inference: BM25
BM25 (Best Matching 25) is a ranking function for text search. BM25 uses sparse vectors that represent documents, where each dimension corresponds to a word. Qdrant can generate these sparse embeddings from input text directly on the server.
While upserting points, provide the text and the qdrant/bm25 embedding model:
{{< code-snippet path="/documentation/headless/snippets/inference/ingest/" >}}
Qdrant uses the model to generate the embeddings and stores the point with the resulting vector. Retrieving the point shows the embeddings that were generated:
....
"my-bm25-vector": {
"indices": [
112174620,
177304315,
662344706,
771857363,
1617337648
],
"values": [
1.6697302,
1.6697302,
1.6697302,
1.6697302,
1.6697302
]
}
....
]
Similarly, use the BM25 model at query time by providing the query string and the qdrant/bm25 embedding model:
{{< code-snippet path="/documentation/headless/snippets/inference/query/" >}}
Read more about full-text search with BM25 in the text search guide.