mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-28 23:48:31 +02:00
docs: Added code snippets
This commit is contained in:
@@ -42,7 +42,7 @@ To avoid slow and unnecessary indexing, it’s better to create an index for eac
|
||||
|
||||
*For more details on scaling best practices, read [How to Implement Multitenancy and Custom Sharding](https://qdrant.tech/articles/multitenancy/).*
|
||||
|
||||
### Defragmentation of Tenant Storage
|
||||
### Defragmentation of Tenant Storage
|
||||
|
||||
With version 1.11, Qdrant changes how vectors from the same tenant are stored on disk, placing them **closer together** for faster bulk reading and reduced scaling costs. This approach optimizes storage and retrieval operations for different tenants, leading to more efficient system performance and better resource utilization.
|
||||
|
||||
@@ -70,7 +70,7 @@ client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
field_name="workspace_2",
|
||||
field_schema=models.KeywordIndexParams(
|
||||
type="keywprd",
|
||||
type="keyword",
|
||||
is_tenant=True,
|
||||
),
|
||||
)
|
||||
@@ -137,30 +137,32 @@ client
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreatePayloadIndexAsync(
|
||||
collectionName: "{collection_name}",
|
||||
fieldName: "workspace_2",
|
||||
schemaType: PayloadSchemaType.Keyword,
|
||||
indexParams: new PayloadIndexParams
|
||||
{
|
||||
KeywordIndexParams = new KeywordIndexParams
|
||||
{
|
||||
IsTenant = true
|
||||
}
|
||||
}
|
||||
collectionName: "{collection_name}",
|
||||
fieldName: "workspace_2",
|
||||
schemaType: PayloadSchemaType.Keyword,
|
||||
indexParams: new PayloadIndexParams
|
||||
{
|
||||
KeywordIndexParams = new KeywordIndexParams
|
||||
{
|
||||
IsTenant = true
|
||||
}
|
||||
}
|
||||
);
|
||||
|
||||
```
|
||||
As a result, the storage structure will be organized in a way to co-locate vectors of the same tenant together.
|
||||
|
||||
As a result, the storage structure will be organized in a way to co-locate vectors of the same tenant together.
|
||||
|
||||
*To learn more about defragmentation, read the [Multitenancy documentation](/documentation/guides/multiple-partitions/).*
|
||||
|
||||
### On-Disk Support for the Payload Index
|
||||
|
||||
When managing billions of records across millions of tenants, keeping all data in RAM is inefficient, especially when only a small subset is frequently accessed. As of 1.11, you can offload "cold" data to disk and cache the “hot” data in RAM.
|
||||
When managing billions of records across millions of tenants, keeping all data in RAM is inefficient, especially when only a small subset is frequently accessed. As of 1.11, you can offload "cold" data to disk and cache the “hot” data in RAM.
|
||||
|
||||
*This feature can help you manage a high number of different payload indexes, which is beneficial if you are working with large varied datasets.*
|
||||
|
||||
@@ -168,9 +170,9 @@ When managing billions of records across millions of tenants, keeping all data i
|
||||
|
||||

|
||||
|
||||
**Example:** As you create an index for Workspace 2, set the `on_disk` parameter.
|
||||
**Example:** As you create an index for Workspace 2, set the `on_disk` parameter.
|
||||
|
||||
```html
|
||||
```http
|
||||
PUT /collections/{collection_name}/index
|
||||
{
|
||||
"field_name": "workspace_2",
|
||||
@@ -182,19 +184,116 @@ PUT /collections/{collection_name}/index
|
||||
}
|
||||
```
|
||||
|
||||
```python
|
||||
client.create_payload_index(
|
||||
collection_name="{collection_name}",
|
||||
field_name="workspace_2",
|
||||
field_schema=models.KeywordIndexParams(
|
||||
type="keyword",
|
||||
is_tenant=True,
|
||||
on_disk=True,
|
||||
),
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
client.createPayloadIndex("{collection_name}", {
|
||||
field_name: "workspace_2",
|
||||
field_schema: {
|
||||
type: "keyword",
|
||||
is_tenant: true,
|
||||
on_disk: true
|
||||
},
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::qdrant::{
|
||||
CreateFieldIndexCollectionBuilder,
|
||||
KeywordIndexParamsBuilder,
|
||||
FieldType
|
||||
};
|
||||
use qdrant_client::{Qdrant, QdrantError};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
client.create_field_index(
|
||||
CreateFieldIndexCollectionBuilder::new(
|
||||
"{collection_name}",
|
||||
"workspace_2",
|
||||
FieldType::Keyword,
|
||||
)
|
||||
.field_index_params(
|
||||
KeywordIndexParamsBuilder::default()
|
||||
.is_tenant(true)
|
||||
.on_disk(true),
|
||||
),
|
||||
);
|
||||
```
|
||||
|
||||
```java
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Collections.PayloadIndexParams;
|
||||
import io.qdrant.client.grpc.Collections.PayloadSchemaType;
|
||||
import io.qdrant.client.grpc.Collections.KeywordIndexParams;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.createPayloadIndexAsync(
|
||||
"{collection_name}",
|
||||
"workspace_2",
|
||||
PayloadSchemaType.Keyword,
|
||||
PayloadIndexParams.newBuilder()
|
||||
.setKeywordIndexParams(
|
||||
KeywordIndexParams.newBuilder()
|
||||
.setIsTenant(true)
|
||||
.setOnDisk(true)
|
||||
.build())
|
||||
.build(),
|
||||
null,
|
||||
null,
|
||||
null)
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.CreatePayloadIndexAsync(
|
||||
collectionName: "{collection_name}",
|
||||
fieldName: "workspace_2",
|
||||
schemaType: PayloadSchemaType.Keyword,
|
||||
indexParams: new PayloadIndexParams
|
||||
{
|
||||
KeywordIndexParams = new KeywordIndexParams
|
||||
{
|
||||
IsTenant = true,
|
||||
OnDisk = true
|
||||
}
|
||||
}
|
||||
);
|
||||
|
||||
```
|
||||
|
||||
By moving the index to disk, Qdrant can handle larger datasets that exceed the capacity of RAM, making the system more scalable and capable of storing more data without being constrained by memory limitations.
|
||||
|
||||
*To learn more about this, read the [Indexing documentation](/documentation/concepts/indexing/).*
|
||||
|
||||
### UUID Datatype for the Payload Index
|
||||
|
||||
Many Qdrant users rely on UUIDs in their payloads, but storing these as strings comes with a substantial memory overhead—approximately 60 bytes per UUID. In reality, UUIDs only require 16 bytes of storage when stored as raw bytes.
|
||||
Many Qdrant users rely on UUIDs in their payloads, but storing these as strings comes with a substantial memory overhead—approximately 60 bytes per UUID. In reality, UUIDs only require 16 bytes of storage when stored as raw bytes.
|
||||
|
||||
To address this inefficiency, we’ve developed a new index type tailored specifically for UUIDs that stores them internally as bytes, **reducing memory usage by up to 3.75x.**
|
||||
|
||||
**Example:** When adding two separate points, indicate their UUID in the payload. In this example, both data points belong to the same user (with the same UUID).
|
||||
|
||||
```html
|
||||
```http
|
||||
PUT /collections/{collection_name}/points
|
||||
{
|
||||
"points": [
|
||||
@@ -216,7 +315,7 @@ PUT /collections/{collection_name}/points
|
||||
|
||||
*To learn more about this, read the [Payload documentation](/documentation/concepts/payload/).*
|
||||
|
||||
### Query API: Groups Endpoint
|
||||
### Query API: Groups Endpoint
|
||||
|
||||
When searching over data, you can group results by specific payload field, which is useful when you have multiple data points for the same item and you want to avoid redundant entries in the results.
|
||||
|
||||
@@ -316,14 +415,13 @@ This endpoint will retrieve the best N points for each document, assuming that t
|
||||
|
||||
*For more information on grouping capabilities refer to our [Hybrid Queries documentation](/documentation/concepts/hybrid-queries/).*
|
||||
|
||||
|
||||
### Query API: Random Sampling
|
||||
### Query API: Random Sampling
|
||||
|
||||
Our [Food Discovery Demo](https://food-discovery.qdrant.tech) always shows a random sample of foods from the larger dataset. Now you can do the same and set the randomization from a basic Query API endpoint.
|
||||
|
||||
When calling the Query API, you will be able to select a subset of data points from a larger dataset in a random manner.
|
||||
When calling the Query API, you will be able to select a subset of data points from a larger dataset in a random manner.
|
||||
|
||||
*This technique is often used to reduce the computational load, improve query response times, or provide a representative sample of the data for various analytical purposes.*
|
||||
*This technique is often used to reduce the computational load, improve query response times, or provide a representative sample of the data for various analytical purposes.*
|
||||
|
||||
**Example:** When querying the collection, you can configure it to retrieve a random sample of data.
|
||||
|
||||
@@ -338,17 +436,73 @@ sampled = client.query_points(
|
||||
query=models.SampleQuery(sample=models.Sample.Random)
|
||||
)
|
||||
```
|
||||
|
||||
```typescript
|
||||
import { QdrantClient } from "@qdrant/js-client-rest";
|
||||
|
||||
const client = new QdrantClient({ host: "localhost", port: 6333 });
|
||||
|
||||
let sampled = client.query("{collection_name}", {
|
||||
query: { sample: "random" },
|
||||
});
|
||||
```
|
||||
|
||||
```rust
|
||||
use qdrant_client::Qdrant;
|
||||
use qdrant_client::qdrant::{Query, QueryPointsBuilder, Sample};
|
||||
|
||||
let client = Qdrant::from_url("http://localhost:6334").build()?;
|
||||
|
||||
let sampled = client
|
||||
.query(
|
||||
QueryPointsBuilder::new("{collection_name}").query(Query::new_sample(Sample::Random)),
|
||||
)
|
||||
.await?;
|
||||
```
|
||||
|
||||
```java
|
||||
import static io.qdrant.client.QueryFactory.sample;
|
||||
|
||||
import io.qdrant.client.QdrantClient;
|
||||
import io.qdrant.client.QdrantGrpcClient;
|
||||
import io.qdrant.client.grpc.Points.Sample;
|
||||
import io.qdrant.client.grpc.Points.QueryPoints;
|
||||
|
||||
QdrantClient client =
|
||||
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
|
||||
|
||||
client
|
||||
.queryAsync(
|
||||
QueryPoints.newBuilder()
|
||||
.setCollectionName("{collection_name}")
|
||||
.setQuery(sample(Sample.Random))
|
||||
.build())
|
||||
.get();
|
||||
```
|
||||
|
||||
```csharp
|
||||
using Qdrant.Client;
|
||||
using Qdrant.Client.Grpc;
|
||||
|
||||
var client = new QdrantClient("localhost", 6334);
|
||||
|
||||
await client.QueryAsync(
|
||||
collectionName: "{collection_name}",
|
||||
query: Sample.Random
|
||||
);
|
||||
```
|
||||
|
||||
*To learn more, check out the [Query API documentation](/documentation/concepts/hybrid-queries/).*
|
||||
|
||||
### Query API: Distribution-Based Score Fusion
|
||||
|
||||
In version 1.10, we added Reciprocal Rank Fusion (RRF) as a way of fusing results from Hybrid Queries. Now we are adding Distribution-Based Score Fusion (DBSF). Michelangiolo Mazzeschi talks more about this fusion method in his latest [Medium article](https://medium.com/plain-simple-software/distribution-based-score-fusion-dbsf-a-new-approach-to-vector-search-ranking-f87c37488b18).
|
||||
|
||||
*DBSF normalizes the scores of the points in each query, using the mean +/- the 3rd standard deviation as limits, and then sums the scores of the same point across different queries.*
|
||||
|
||||
*DBSF normalizes the scores of the points in each query, using the mean +/- the 3rd standard deviation as limits, and then sums the scores of the same point across different queries.*
|
||||
|
||||
**Example:** To fuse `prefetch` results from sparse and dense queries, set `"fusion": "dbsf"`
|
||||
```html
|
||||
|
||||
```http
|
||||
POST /collections/{collection_name}/points/query
|
||||
{
|
||||
"prefetch": [
|
||||
@@ -508,22 +662,21 @@ Note that `dbsf` is stateless and calculates the normalization limits only based
|
||||
|
||||
*To learn more, check out the [Hybrid Queries documentation](/documentation/concepts/hybrid-queries/).*
|
||||
|
||||
## Web UI: Search Quality Tool
|
||||
## Web UI: Search Quality Tool
|
||||
|
||||
We have updated the Qdrant Web UI with additional testing functionality. Now you can check the quality of your search requests in real time and measure it against exact search.
|
||||
We have updated the Qdrant Web UI with additional testing functionality. Now you can check the quality of your search requests in real time and measure it against exact search.
|
||||
|
||||
**Try it:** In the Dashboard, go to collection settings and test the **Precision** from the Search Quality menu tab.
|
||||
**Try it:** In the Dashboard, go to collection settings and test the **Precision** from the Search Quality menu tab.
|
||||
|
||||
> The feature will conduct semantic search for each point and produce a report below.
|
||||
|
||||
<iframe width="560" height="315" src="https://www.youtube.com/embed/PJHzeVay_nQ?si=u-6lqCVECd-A319M" title="YouTube video player" frameborder="0" allow="accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen></iframe>
|
||||
|
||||
|
||||
## Web UI: Graph Exploration Tool
|
||||
|
||||
Deeper exploration is highly dependant on expanding context. This is something we previously covered in the [Discovery Needs Context](/articles/discovery-search/) article earlier this year. Now, we have developed a UI feature to help you visualize how semantic search can be used for exploratory and recommendation purposes.
|
||||
Deeper exploration is highly dependant on expanding context. This is something we previously covered in the [Discovery Needs Context](/articles/discovery-search/) article earlier this year. Now, we have developed a UI feature to help you visualize how semantic search can be used for exploratory and recommendation purposes.
|
||||
|
||||
**Try it:** Using the feature is pretty self-explanatory. Each collection's dataset can be explored from the **Graph** tab. As you see the images change, you can steer your search in the direction of specific characteristics that interest you.
|
||||
**Try it:** Using the feature is pretty self-explanatory. Each collection's dataset can be explored from the **Graph** tab. As you see the images change, you can steer your search in the direction of specific characteristics that interest you.
|
||||
|
||||
> Search results will become more "distilled" and tailored to your preferences.
|
||||
|
||||
@@ -533,4 +686,4 @@ Deeper exploration is highly dependant on expanding context. This is something w
|
||||
|
||||
If you’re new to Qdrant, now is the perfect time to start. Check out our [documentation](/documentation/) guides and see why Qdrant is the go-to solution for vector search.
|
||||
|
||||
We’re very happy to bring you this latest version of Qdrant, and we can’t wait to see what you build with it. As always, your feedback is invaluable—feel free to reach out with any questions or comments on our [community forum](https://qdrant.to/discord).
|
||||
We’re very happy to bring you this latest version of Qdrant, and we can’t wait to see what you build with it. As always, your feedback is invaluable—feel free to reach out with any questions or comments on our [community forum](https://qdrant.to/discord).
|
||||
|
||||
Reference in New Issue
Block a user