Restructure inference docs (#2225)

* Break Inference page into several pages

* Make all inference code snippets testable and clean up

* Make more snippets testable

* Edits

* Document automatic query and passage prefix injection in Cloud Inference

Qdrant Cloud Inference silently applies model-specific prefixes (e.g.
"query: "/"passage: " for E5, BGE-style instruction prefix for BGE/mxbai/
Snowflake arctic-embed) so users don't need to manage them manually.
Add a section explaining this behavior, the idempotency guarantee, and
the scope (Qdrant-hosted models only; external providers handle their own).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Document short query optimization in Cloud Inference

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* Update links

* Expand on external provider API key usage

* Add section about external provider API keys

* Default to header for external API keys

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Abdon Pijpelink
2026-06-24 08:11:03 +02:00
committed by GitHub
co-authored by Claude Sonnet 4.6
parent 90d072eb03
commit 478b96554f
250 changed files with 3259 additions and 2187 deletions
@@ -1,22 +1,20 @@
use qdrant_client::{
Qdrant,
qdrant::{Document, Query, QueryPointsBuilder, Value},
qdrant::{Document, Query, QueryPointsBuilder},
};
use std::collections::HashMap;
pub async fn main() -> anyhow::Result<()> {
let client = Qdrant::from_url("<your-qdrant-url>").build().unwrap();
let mut options = HashMap::<String, Value>::new();
options.insert("openrouter-api-key".to_string(), "<YOUR_OPENROUTER_API_KEY>".into());
let client = Qdrant::from_url("<your-qdrant-url>").build().unwrap(); // @hide
client
.with_header("openrouter-api-key", "<YOUR_OPENROUTER_API_KEY>")
.query(
QueryPointsBuilder::new("{collection_name}")
.query(Query::new_nearest(Document {
text: "How to bake cookies?".into(),
model: "openrouter/mistralai/mistral-embed-2312".into(),
options,
options: HashMap::new(),
}))
.build(),
)