add image examples and update docs

This commit is contained in:
generall
2025-07-15 02:10:35 +02:00
parent 305e6969d7
commit 6ae287b12b
10 changed files with 446 additions and 1 deletions
@@ -29,7 +29,81 @@ Inference is billed based on the number of tokens processed by the model. The co
## Using Inference
Inference can be easily used through the Qdrant SDKs and the REST or GRPC APIs. Inference is available when upserting points as well as when querying the database.
Inference can be easily used through the Qdrant SDKs and the REST or GRPC APIs.
Inference is available when upserting points as well as when querying the database.
It is can be done with special *Interface Objects*, defined in Qdrant API.
There are
* **`Document`** object, used for text inference. Example:
```js
// Document
{
// Model input
text: "Your text",
// Name of the model, to do inference with
model: "<the-model-to-use>",
// Extra parameters for the model, Optional
options: {}
}
```
* **`Image`** object, used for image inference. Example:
```js
// Image
{
// Image input
image: "<url>", // Or base64 of the image
// Name of the model, to do inference with
model: "<the-model-to-use>",
// Extra parameters for the model, Optional
options: {}
}
```
* **`Object`** object, reserved for all other types of input, which might be implemented in future.
Qdrant API supports usage of Inference Objects in all places, where regular vectors can be used.
For example:
```http
POST /collections/<your-collection>/points/query
{
"query": {
"nearest": [0.12, 0.34, 0.56, 0.78, ...]
}
}
```
Can be replaced with
```http
POST /collections/<your-collection>/points/query
{
"query": {
"nearest":{
"text": "My Query Text",
"model": "<the-model-to-use>"
}
}
}
```
In this case, the Qdrant server will call the inference server, automatically replace the Inference Object, and perform the search query.
The obtained embedding will only be transferred within the low-latency network and will never be transmitted between the client and Qdrant Cloud.
The input used for inference will not be saved anywhere. If you need to persist it in Qdrant, make sure to explicitly include it in the payload.
### Text Inference
Let's consider a simple example of using Cloud Inference with text model.
In this in this example we create one point and use simple search query with `Document` Inference Object.
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/simple/" >}}
@@ -38,3 +112,34 @@ Usage examples, specific to each cluster and model, can also be found in the Inf
Note that each model has a context window, which is the maximum number of tokens that can be processed by the model in a single request. If the input text exceeds the context window, it will be truncated to fit within the limit. The context window size is displayed in the Inference tab of the Cluster Detail page.
For dense vector models, you also have to ensure that the vector size configured in the collection matches the output size of the model. If the vector size does not match, the upsert will fail with an error.
### Image Inference
Here is another simple example of using Cloud Inference with an image model.
This time, we will use the `CLIP` model to encode an image and then use a text query to search for it.
Since the `CLIP` model is multimodal, we can use both image and text inputs on the same vector field.
{{< code-snippet path="/documentation/headless/snippets/cloud-inference/image/" >}}
The Qdrant Inference server will download images using the provided link.
Note that each model has limitations on the file size and extensions it can work with.
Please refer to the model card for details.
### Local Inference Compatibility
The Python SDK offers a unique capability: it supports both [local](/documentation/fastembed/fastembed-semantic-search/) and cloud inference through an identical interface.
You can easily switch between local and cloud inference by setting the cloud_inference flag when initializing the QdrantClient. For example:
```python
client = QdrantClient(
url="https://your-cluster.qdrant.io",
api_key="<your-api-key>",
cloud_inference=True, # Set to False to use local inference
)
```
This flexibility allows you to develop and test your applications locally or in continuous integration (CI) environments without requiring access to cloud inference resources.
When `cloud_inference` is set to `False`, inference is performed locally usign `fastembed`.
When set to `True`, inference requests are handled by Qdrant Cloud.
@@ -0,0 +1,7 @@
This code snippet shows how to use Qdrant Cloud's cloud-side inference to automatically generate vector embeddings from images and text during upsert and search operations.
In the example, a new point is inserted with an image URL and a specified model. The vector embedding is generated on the Qdrant Cloud side using the provided model. The `image` and `model` fields specify the image to embed and the model to use.
After the point is inserted, it becomes searchable. The snippet also demonstrates how to perform a search query using cloud-side inference: a `text` and `model` are provided, and Qdrant Cloud generates the query vector automatically. This allows you to search your collection using natural language queries without manually generating embeddings.
Example demonstrates multimodal search with `CLIP` model and cloud-side inference.
@@ -0,0 +1,31 @@
```bash
# Create a new vector
curl -X PUT "https://xyz-example.qdrant.io:6333/collections/<your-collection>/points?wait=true" \
-H "Content-Type: application/json" \
-H "api-key: <paste-your-api-key-here>" \
-d '{
"points": [
{
"id": 1,
"vector": {
"text": "https://qdrant.tech/example.png",
"model": "qdrant/clip-vit-b-32-vision"
},
"payload": {
"title": "Example Image"
}
}
]
}'
# Perform a search query
curl -X POST "https://xyz-example.qdrant.io:6333/collections/<your-collection>/points/query" \
-H "Content-Type: application/json" \
-H "api-key: <paste-your-api-key-here>" \
-d '{
"query": {
"text": "Mission to Mars",
"model": "qdrant/clip-vit-b-32-text"
}
}'
```
@@ -0,0 +1,40 @@
```csharp
using Qdrant.Client;
using Qdrant.Client.Grpc;
using Value = Qdrant.Client.Grpc.Value;
var client = new QdrantClient(
host: "xyz-example.qdrant.io",
port: 6334,
https: true,
apiKey: "<paste-your-api-key-here>"
);
await client.UpsertAsync(
collectionName: "<your-collection>",
points: new List <PointStruct> {
new() {
Id = 1,
Vectors = new Image() {
Image = "https://qdrant.tech/example.png",
Model = "qdrant/clip-vit-b-32-vision",
},
Payload = {
["title"] = "Example Image"
},
},
}
);
var points = await client.QueryAsync(
collectionName: "<your-collection>",
query: new Document() {
Text = "Mission to Mars",
Model = "qdrant/clip-vit-b-32-text"
}
);
foreach(var point in points) {
Console.WriteLine(point);
}
```
@@ -0,0 +1,57 @@
```go
package main
import (
"context"
"log"
"time"
"github.com/qdrant/go-client/qdrant"
)
func main() {
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
client, err := qdrant.NewClient(&qdrant.Config{
Host: "xyz-example.qdrant.io",
Port: 6334,
APIKey: "<paste-your-api-key-here>",
UseTLS: true,
})
if err != nil {
log.Fatalf("did not connect: %v", err)
}
defer client.Close()
_, err = client.GetPointsClient().Upsert(ctx, &qdrant.UpsertPoints{
CollectionName: "<your-collection>",
Points: []*qdrant.PointStruct{
{
Id: qdrant.NewIDNum(uint64(1)),
Vectors: qdrant.NewVectorsImage(&qdrant.Image{
Image: "https://qdrant.tech/example.png",
Model: "qdrant/clip-vit-b-32-vision",
}),
Payload: qdrant.NewValueMap(map[string]any{
"title": "Example image",
}),
},
},
})
if err != nil {
log.Fatalf("error creating point: %v", err)
}
points, err := client.Query(ctx, &qdrant.QueryPoints{
CollectionName: "<your-collection>",
Query: qdrant.NewQueryNearest(
qdrant.NewVectorInputDocument(&qdrant.Document{
Text: "Mission to Mars",
Model: "qdrant/clip-vit-b-32-text",
}),
),
})
log.Printf("List of points: %s", points)
}
```
@@ -0,0 +1,27 @@
```http
# Insert new points with cloud-side inference
PUT /collections/<your-collection>/points?wait=true
{
"points": [
{
"id": 1,
"vector": {
"image": "https://qdrant.tech/example.png",
"model": "qdrant/clip-vit-b-32-vision"
},
"payload": {
"title": "Example Image"
}
}
]
}
# Search in the collection using cloud-side inference
POST /collections/<your-collection>/points/query
{
"query": {
"text": "Mission to Mars",
"model": "qdrant/clip-vit-b-32-text"
}
}
```
@@ -0,0 +1,59 @@
```java
package org.example;
import static io.qdrant.client.PointIdFactory.id;
import static io.qdrant.client.QueryFactory.nearest;
import static io.qdrant.client.ValueFactory.value;
import static io.qdrant.client.VectorsFactory.vectors;
import io.qdrant.client.grpc.Points;
import io.qdrant.client.grpc.Points.Document;
import io.qdrant.client.grpc.Points.Image;
import io.qdrant.client.grpc.Points.PointStruct;
import java.util.List;
import java.util.Map;
import java.util.concurrent.ExecutionException;
public class Main {
public static void main(String[] args)
throws ExecutionException, InterruptedException {
QdrantClient client =
new QdrantClient(
QdrantGrpcClient.newBuilder("xyz-example.qdrant.io", 6334, true)
.withApiKey("<paste-your-api-key-here>")
.build());
client
.upsertAsync(
"<your-collection>",
List.of(
PointStruct.newBuilder()
.setId(id(1))
.setVectors(
vectors(
Image.newBuilder()
.setImage("https://qdrant.tech/example.png")
.setModel("qdrant/clip-vit-b-32-vision")
.build()))
.putAllPayload(Map.of("title", value("Example Image")))
.build()))
.get();
List <Points.ScoredPoint> points =
client
.queryAsync(
Points.QueryPoints.newBuilder()
.setCollectionName("<your-collection>")
.setQuery(
nearest(
Document.newBuilder()
.setText("Mission to Mars")
.setModel("qdrant/clip-vit-b-32-text")
.build()))
.build())
.get();
System.out.printf(points.toString());
}
}
```
@@ -0,0 +1,37 @@
```python
from qdrant_client import QdrantClient
from qdrant_client.models import PointStruct, Image, Document
client = QdrantClient(
url="https://xyz-example.qdrant.io:6333",
api_key="<paste-your-api-key-here>",
# IMPORTANT
# If not enabled, inference will be performed locally
cloud_inference=True,
)
points = [
PointStruct(
id=1,
vector=Image(
image="https://qdrant.tech/example.png",
model="qdrant/clip-vit-b-32-vision"
),
payload={
"title": "Example Image"
}
)
]
client.upsert(collection_name="<your-collection>", points=points)
result = client.query_points(
collection_name="<your-collection>",
query=Document(
text="Mission to Mars",
model="qdrant/clip-vit-b-32-text"
)
)
print(result)
```
@@ -0,0 +1,47 @@
```rust
use qdrant_client::qdrant::Query;
use qdrant_client::qdrant::QueryPointsBuilder;
use qdrant_client::Payload;
use qdrant_client::Qdrant;
use qdrant_client::qdrant::{Document, Image};
use qdrant_client::qdrant::{PointStruct, UpsertPointsBuilder};
#[tokio::main]
async fn main() {
let client = Qdrant::from_url("https://xyz-example.qdrant.io:6334")
.api_key("<paste-your-api-key-here>")
.build()
.unwrap();
let points = vec![
PointStruct::new(
1,
Image::new_from_url(
"https://qdrant.tech/example.png",
"qdrant/clip-vit-b-32-vision"
),
Payload::try_from(serde_json::json!({
"title": "Example Image"
})).unwrap(),
)
];
let upsert_request = UpsertPointsBuilder::new(
"<your-collection>",
points
).wait(true);
let _ = client.upsert_points(upsert_request).await;
let query_document = Document::new(
"Mission to Mars",
"qdrant/clip-vit-b-32-text"
);
let query_request = QueryPointsBuilder::new("<your-collection>")
.query(Query::new_nearest(query_document));
let result = client.query(query_request).await.unwrap();
println!("Result: {:?}", result);
}
```
@@ -0,0 +1,35 @@
```typescript
import {QdrantClient} from "@qdrant/js-client-rest";
const client = new QdrantClient({
url: 'https://xyz-example.qdrant.io:6333',
apiKey: '<paste-your-api-key-here>',
});
const points = [
{
id: 1,
vector: {
image: "https://qdrant.tech/example.png",
model: "qdrant/clip-vit-b-32-vision"
},
payload: {
title: "Example Image"
}
}
];
await client.upsert("<your-collection>", { wait: true, points });
const result = await client.query(
"<your-collection>",
{
query: {
text: "Mission to Mars",
model: "qdrant/clip-vit-b-32-text"
},
}
)
console.log(result);
```