mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-29 07:58:31 +02:00
v0.8.x docs (#54)
* v0.8.x docs * docs auto-sync * docs auto-sync * update doc sync script * docs auto-sync * docs auto-sync * docs auto-sync Co-authored-by: qdrant <qdrant@users.noreply.github.com>
This commit is contained in:
co-authored by
qdrant
parent
0bf25c9913
commit
cd0cb0f528
@@ -20,4 +20,5 @@ jobs:
|
|||||||
git config --global user.email 'qdrant@users.noreply.github.com'
|
git config --global user.email 'qdrant@users.noreply.github.com'
|
||||||
git remote set-url origin https://x-access-token:${{ secrets.GITHUB_TOKEN }}@github.com/$GITHUB_REPOSITORY
|
git remote set-url origin https://x-access-token:${{ secrets.GITHUB_TOKEN }}@github.com/$GITHUB_REPOSITORY
|
||||||
git checkout $GITHUB_HEAD_REF
|
git checkout $GITHUB_HEAD_REF
|
||||||
|
git add qdrant-landing/content/documentation/*.md
|
||||||
git commit -am "docs auto-sync" && git push || true
|
git commit -am "docs auto-sync" && git push || true
|
||||||
|
|||||||
@@ -48,9 +48,9 @@ keywords = "search engine, neural network, matching, filter, SaaS, approximate n
|
|||||||
|
|
||||||
gdpr = "We use cookies to learn more about you. At any time you can delete or block cookies through your browser settings."
|
gdpr = "We use cookies to learn more about you. At any time you can delete or block cookies through your browser settings."
|
||||||
|
|
||||||
githubDocPrefix = "https://github.com/qdrant/docs/tree/master/qdrant/v0.7.x/"
|
githubDocPrefix = "https://github.com/qdrant/docs/tree/master/qdrant/v0.8.x/"
|
||||||
|
|
||||||
docVersion = "v0.7.x"
|
docVersion = "v0.8.x"
|
||||||
|
|
||||||
googleTagManager = "GTM-KRLCXD5"
|
googleTagManager = "GTM-KRLCXD5"
|
||||||
|
|
||||||
|
|||||||
@@ -24,6 +24,6 @@ These can be:
|
|||||||
|
|
||||||
In addition to this documentation, you may be interested in looking at examples of projects made with Qdrant:
|
In addition to this documentation, you may be interested in looking at examples of projects made with Qdrant:
|
||||||
|
|
||||||
* [Semantic Search for startups](https://qdrant.to/semantic-search-demo) + [Source Code](https://github.com/qdrant/qdrant_demo)
|
* [Semantic Search for startups](https://demo.qdrant.tech/) + [Source Code](https://github.com/qdrant/qdrant_demo)
|
||||||
* [Visual Food Discovery](https://qdrant.to/food-discovery)
|
* [Visual Food Discovery](https://food-discovery.qdrant.tech/)
|
||||||
* [Step-by-Step tutorial on building neural search](/articles/neural-search-tutorial/)
|
* [Step-by-Step tutorial on building neural search](http://localhost:1313/articles/neural-search-tutorial/)
|
||||||
@@ -22,10 +22,8 @@ These settings can be changed at any time by a corresponding request.
|
|||||||
|
|
||||||
### Create collection
|
### Create collection
|
||||||
|
|
||||||
With REST API
|
```http
|
||||||
|
PUT /collections/{collection_name}
|
||||||
```
|
|
||||||
PUT /collections/example_collection
|
|
||||||
|
|
||||||
{
|
{
|
||||||
"name": "example_collection",
|
"name": "example_collection",
|
||||||
@@ -34,13 +32,29 @@ PUT /collections/example_collection
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
|
||||||
|
client = QdrantClient(host="localhost", port=6333)
|
||||||
|
|
||||||
|
client.recreate_collection(
|
||||||
|
name="{collection_name}",
|
||||||
|
distance="Cosine",
|
||||||
|
vector_size=300,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
In addition to the required options, you can also specify custom values for the following collection options:
|
In addition to the required options, you can also specify custom values for the following collection options:
|
||||||
|
|
||||||
- `hnsw_config`
|
- `hnsw_config` - see [indexing](../indexing/#vector-index) for details.
|
||||||
- `wal_config`
|
- `wal_config` - Write-Ahead-Log related configuration. See more details about [WAL](../storage/#versioning)
|
||||||
- `optimizers_config`
|
- `optimizers_config` - see [optimizer](../optimizer) for details.
|
||||||
|
- `shard_number` - which defines how many shards the collection should have. See [distributed deployment](../distributed_deployment#sharding) section for details.
|
||||||
|
- `on_disk_payload` - defines where to store payload data. If `true` - payload will be stored on disk only. Might be useful for limiting the RAM usage in case of large payload.
|
||||||
|
|
||||||
See [schema definitions](https://qdrant.github.io/qdrant/redoc/index.html#operation/create_collection) and a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml) for more information about collection parameters.
|
Default parameters for the optional collection parameters are defined in [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml).
|
||||||
|
|
||||||
|
See [schema definitions](https://qdrant.github.io/qdrant/redoc/index.html#operation/create_collection) and a [configuration file](https://github.com/qdrant/qdrant/blob/master/config/config.yaml) for more information about collection parameters.
|
||||||
|
|
||||||
|
|
||||||
<!--
|
<!--
|
||||||
@@ -52,18 +66,13 @@ See [schema definitions](https://qdrant.github.io/qdrant/redoc/index.html#operat
|
|||||||
|
|
||||||
### Delete collection
|
### Delete collection
|
||||||
|
|
||||||
With REST API
|
```http
|
||||||
|
DELETE /collections/{collection_name}
|
||||||
```
|
```
|
||||||
DELETE /collections/example_collection
|
|
||||||
```
|
|
||||||
|
|
||||||
<!--
|
|
||||||
#### Python
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.delete_collection(collection_name="{collection_name}")
|
||||||
```
|
```
|
||||||
-->
|
|
||||||
|
|
||||||
|
|
||||||
### Update collection parameters
|
### Update collection parameters
|
||||||
@@ -72,8 +81,8 @@ Dynamic parameter updates may be helpful, for example, for more efficient initia
|
|||||||
With these settings, you can disable indexing during the upload process. And enable it immediately after the upload is finished.
|
With these settings, you can disable indexing during the upload process. And enable it immediately after the upload is finished.
|
||||||
As a result, you will not waste extra computation resources on rebuilding the index.
|
As a result, you will not waste extra computation resources on rebuilding the index.
|
||||||
|
|
||||||
```
|
```http
|
||||||
PATCH /collections/example_collection
|
PATCH /collections/{collection_name}
|
||||||
|
|
||||||
{
|
{
|
||||||
"optimizers_config": {
|
"optimizers_config": {
|
||||||
@@ -108,7 +117,7 @@ Since all changes of aliases happen atomically, no concurrent requests will be a
|
|||||||
|
|
||||||
### Create alias
|
### Create alias
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/aliases
|
POST /collections/aliases
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -133,7 +142,7 @@ POST /collections/aliases
|
|||||||
|
|
||||||
### Remove alias
|
### Remove alias
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/aliases
|
POST /collections/aliases
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -160,7 +169,7 @@ Multiple alias actions are performed atomically.
|
|||||||
For example, you can switch underlying collection with the following command:
|
For example, you can switch underlying collection with the following command:
|
||||||
|
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/aliases
|
POST /collections/aliases
|
||||||
|
|
||||||
{
|
{
|
||||||
|
|||||||
@@ -0,0 +1,149 @@
|
|||||||
|
---
|
||||||
|
title: Configuration
|
||||||
|
weight: 45
|
||||||
|
---
|
||||||
|
|
||||||
|
To change or correct Qdrant's behavior, default collection settings, and network interface parameters, you can use the configuration file.
|
||||||
|
|
||||||
|
Default configuration file is located in [config/config.yaml](https://github.com/qdrant/qdrant/blob/master/config/config.yaml).
|
||||||
|
|
||||||
|
In the production environment, you can override any value of this file by providing new values in `/qdrant/config/production.yaml` inside the docker.
|
||||||
|
|
||||||
|
Here is an example of how you can pass custom configuration inside the docker container:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run -p 6333:6333 \
|
||||||
|
-v $(pwd)/path/to/custom_config.yaml:/qdrant/config/production.yaml \
|
||||||
|
qdrant/qdrant
|
||||||
|
```
|
||||||
|
|
||||||
|
Example of the configuration file:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
debug: false
|
||||||
|
log_level: INFO
|
||||||
|
|
||||||
|
|
||||||
|
storage:
|
||||||
|
# Where to store all the data
|
||||||
|
storage_path: ./storage
|
||||||
|
|
||||||
|
# If true - point's payload will not be stored in memory.
|
||||||
|
# It will be read from the disk every time it is requested.
|
||||||
|
# This setting saves RAM by (slightly) increasing the response time.
|
||||||
|
# Note: those payload values that are involved in filtering and are indexed - remain in RAM.
|
||||||
|
on_disk_payload: false
|
||||||
|
|
||||||
|
# Write-ahead-log related configuration
|
||||||
|
wal:
|
||||||
|
# Size of a single WAL segment
|
||||||
|
wal_capacity_mb: 32
|
||||||
|
|
||||||
|
# Number of WAL segments to create ahead of actual data requirement
|
||||||
|
wal_segments_ahead: 0
|
||||||
|
|
||||||
|
|
||||||
|
performance:
|
||||||
|
# Number of parallel threads used for search operations. If 0 - auto selection.
|
||||||
|
max_search_threads: 0
|
||||||
|
|
||||||
|
optimizers:
|
||||||
|
# The minimal fraction of deleted vectors in a segment, required to perform segment optimization
|
||||||
|
deleted_threshold: 0.2
|
||||||
|
|
||||||
|
# The minimal number of vectors in a segment, required to perform segment optimization
|
||||||
|
vacuum_min_vector_number: 1000
|
||||||
|
|
||||||
|
# Target amount of segments optimizer will try to keep.
|
||||||
|
# Real amount of segments may vary depending on multiple parameters:
|
||||||
|
# - Amount of stored points
|
||||||
|
# - Current write RPS
|
||||||
|
#
|
||||||
|
# It is recommended to select default number of segments as a factor of the number of search threads,
|
||||||
|
# so that each segment would be handled evenly by one of the threads.
|
||||||
|
# If `default_segment_number = 0`, will be automatically selected by the number of available CPUs
|
||||||
|
default_segment_number: 0
|
||||||
|
|
||||||
|
# Do not create segments larger this size (in KiloBytes).
|
||||||
|
# Large segments might require disproportionately long indexation times,
|
||||||
|
# therefore it makes sense to limit the size of segments.
|
||||||
|
#
|
||||||
|
# If indexation speed have more priority for your - make this parameter lower.
|
||||||
|
# If search speed is more important - make this parameter higher.
|
||||||
|
# Note: 1Kb = 1 vector of size 256
|
||||||
|
max_segment_size_kb: 200000
|
||||||
|
|
||||||
|
# Maximum size (in KiloBytes) of vectors to store in-memory per segment.
|
||||||
|
# Segments larger than this threshold will be stored as read-only memmaped file.
|
||||||
|
# To enable memmap storage, lower the threshold
|
||||||
|
# Note: 1Kb = 1 vector of size 256
|
||||||
|
memmap_threshold_kb: 200000
|
||||||
|
|
||||||
|
# Maximum size (in KiloBytes) of vectors allowed for plain index.
|
||||||
|
# Default value based on https://github.com/google-research/google-research/blob/master/scann/docs/algorithms.md
|
||||||
|
# Note: 1Kb = 1 vector of size 256
|
||||||
|
indexing_threshold_kb: 20000
|
||||||
|
|
||||||
|
# Interval between forced flushes.
|
||||||
|
flush_interval_sec: 1
|
||||||
|
|
||||||
|
# Max number of threads, which can be used for optimization. If 0 - `NUM_CPU - 1` will be used
|
||||||
|
max_optimization_threads: 0
|
||||||
|
|
||||||
|
# Default parameters of HNSW Index. Could be override for each collection individually
|
||||||
|
hnsw_index:
|
||||||
|
# Number of edges per node in the index graph. Larger the value - more accurate the search, more space required.
|
||||||
|
m: 16
|
||||||
|
# Number of neighbours to consider during the index building. Larger the value - more accurate the search, more time required to build index.
|
||||||
|
ef_construct: 100
|
||||||
|
# Minimal size (in KiloBytes) of vectors for additional payload-based indexing.
|
||||||
|
# If payload chunk is smaller than `full_scan_threshold_kb` additional indexing won't be used -
|
||||||
|
# in this case full-scan search should be preferred by query planner and additional indexing is not required.
|
||||||
|
# Note: 1Kb = 1 vector of size 256
|
||||||
|
full_scan_threshold_kb: 10000
|
||||||
|
|
||||||
|
service:
|
||||||
|
|
||||||
|
# Maximum size of POST data in a single request in megabytes
|
||||||
|
max_request_size_mb: 32
|
||||||
|
|
||||||
|
# Number of parallel workers used for serving the api. If 0 - equal to the number of available cores.
|
||||||
|
# If missing - Same as storage.max_search_threads
|
||||||
|
max_workers: 0
|
||||||
|
|
||||||
|
# Host to bind the service on
|
||||||
|
host: 0.0.0.0
|
||||||
|
|
||||||
|
# HTTP port to bind the service on
|
||||||
|
http_port: 6333
|
||||||
|
|
||||||
|
# gRPC port to bind the service on.
|
||||||
|
# If `null` - gRPC is disabled. Default: null
|
||||||
|
grpc_port: null
|
||||||
|
# Uncomment to enable gRPC:
|
||||||
|
# grpc_port: 6334
|
||||||
|
|
||||||
|
# Enable CORS headers in REST API.
|
||||||
|
# If enabled, browsers would be allowed to query REST endpoints regardless of query origin.
|
||||||
|
# More info: https://developer.mozilla.org/en-US/docs/Web/HTTP/CORS
|
||||||
|
# Default: true
|
||||||
|
enable_cors: true
|
||||||
|
|
||||||
|
cluster:
|
||||||
|
# Use `enabled: true` to run Qdrant in distributed deployment mode
|
||||||
|
enabled: true
|
||||||
|
|
||||||
|
# Configuration of the inter-cluster communication
|
||||||
|
p2p:
|
||||||
|
# Port for internal communication between peers
|
||||||
|
port: 6335
|
||||||
|
|
||||||
|
# Configuration related to distributed consensus algorithm
|
||||||
|
consensus:
|
||||||
|
# How frequently peers should ping each other.
|
||||||
|
# Setting this parameter to lower value will allow consensus
|
||||||
|
# to detect disconnected nodes earlier, but too frequent
|
||||||
|
# tick period may create significant network and CPU overhead.
|
||||||
|
# We encourage you NOT to change this parameter unless you know what you are doing.
|
||||||
|
tick_period_ms: 100
|
||||||
|
```
|
||||||
@@ -3,18 +3,166 @@ title: Distributed deployment
|
|||||||
weight: 50
|
weight: 50
|
||||||
---
|
---
|
||||||
|
|
||||||
# Distributed Deployment
|
Since version v0.8.0 Qdrant supports an experimental mode of distributed deployment.
|
||||||
|
In this mode, multiple Qdrant services communicate with each other to distribute the data across the peers to extend the storage capabilities and increase stability.
|
||||||
|
|
||||||
Currently work-in-progress, see [Roadmap](https://github.com/qdrant/qdrant/blob/master/docs/roadmap/README.md)
|
To enable distributed deployment - enable the cluster mode in the [configuration](../configuration) or using the ENV variable: `QDRANT__CLUSTER__ENABLED=true`.
|
||||||
|
|
||||||
## Replication
|
```yaml
|
||||||
|
cluster:
|
||||||
|
# Use `enabled: true` to run Qdrant in distributed deployment mode
|
||||||
|
enabled: true
|
||||||
|
# Configuration of the inter-cluster communication
|
||||||
|
p2p:
|
||||||
|
# Port for internal communication between peers
|
||||||
|
port: 6335
|
||||||
|
|
||||||
Currently work-in-progress, see [Roadmap](https://github.com/qdrant/qdrant/blob/master/docs/roadmap/README.md)
|
# Configuration related to distributed consensus algorithm
|
||||||
|
consensus:
|
||||||
|
# How frequently peers should ping each other.
|
||||||
|
# Setting this parameter to lower value will allow consensus
|
||||||
|
# to detect disconnected node earlier, but too frequent
|
||||||
|
# tick period may create significant network and CPU overhead.
|
||||||
|
# We encourage you NOT to change this parameter unless you know what you are doing.
|
||||||
|
tick_period_ms: 100
|
||||||
|
```
|
||||||
|
|
||||||
|
With default configuration, Qdrant will use port `6335` for its internal communication.
|
||||||
|
All peers should be accessible on this port from within the cluster, but make sure to isolate this port from outside access, as it might be used to perform write operations.
|
||||||
|
|
||||||
|
Additionally, the first peer of the cluster should be provided with its URL, so it could tell other nodes how it should be reached.
|
||||||
|
Use the `uri` CLI argument to provide the URL to the peer:
|
||||||
|
|
||||||
|
```
|
||||||
|
./qdrant --uri 'http://qdrant_node_1:6335'
|
||||||
|
```
|
||||||
|
|
||||||
|
Subsequent peers in a cluster must know at least one node of the existing cluster to synchronize through it with the rest of the cluster.
|
||||||
|
|
||||||
|
To do this, they need to be provided with a bootstrap URL:
|
||||||
|
|
||||||
|
```
|
||||||
|
./qdrant --bootstrap 'http://qdrant_node_1:6335'
|
||||||
|
```
|
||||||
|
|
||||||
|
The URL of the new peers themselves will be calculated automatically from the IP address of their request.
|
||||||
|
But it is also possible to provide them individually using `--uri` argument.
|
||||||
|
|
||||||
|
```text
|
||||||
|
USAGE:
|
||||||
|
qdrant [OPTIONS]
|
||||||
|
|
||||||
|
OPTIONS:
|
||||||
|
--bootstrap <URI>
|
||||||
|
Uri of the peer to bootstrap from in case of multi-peer deployment. If not specified -
|
||||||
|
this peer will be considered as a first in a new deployment
|
||||||
|
|
||||||
|
--uri <URI>
|
||||||
|
Uri of this peer. Other peers should be able to reach it by this uri.
|
||||||
|
|
||||||
|
This value has to be supplied if this is the first peer in a new deployment.
|
||||||
|
|
||||||
|
In case this is not the first peer and it bootstraps the value is optional. If not
|
||||||
|
supplied then qdrant will take internal grpc port from config and derive the IP address
|
||||||
|
of this peer on bootstrap peer (receiving side)
|
||||||
|
|
||||||
|
```
|
||||||
|
|
||||||
|
After a successful synchronization you can observe the state of the cluster through the [REST API](https://qdrant.github.io/qdrant/redoc/index.html?v=master#tag/cluster):
|
||||||
|
|
||||||
|
```
|
||||||
|
GET /cluster
|
||||||
|
```
|
||||||
|
|
||||||
|
Example result:
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"result": {
|
||||||
|
"status": "enabled",
|
||||||
|
"peer_id": 11532566549086892000,
|
||||||
|
"peers": {
|
||||||
|
"9834046559507417430": {
|
||||||
|
"uri": "http://172.18.0.3:6335/"
|
||||||
|
},
|
||||||
|
"11532566549086892528": {
|
||||||
|
"uri": "http://qdrant_node_1:6335/"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"raft_info": {
|
||||||
|
"term": 1,
|
||||||
|
"commit": 4,
|
||||||
|
"pending_operations": 1,
|
||||||
|
"leader": 11532566549086892000,
|
||||||
|
"role": "Leader"
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"status": "ok",
|
||||||
|
"time": 5.731e-06
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
## Raft
|
||||||
|
|
||||||
|
Qdrant is using the [Raft](https://raft.github.io/) consensus protocol to maintain consistency regarding the cluster topology and the collections structure.
|
||||||
|
|
||||||
|
Operation with points, on the other hand, are not going through the consensus infrastructure.
|
||||||
|
Qdrant is not intended to have strong transaction guarantees, which allows it to perform point operations with low overhead.
|
||||||
|
In practice, it means that Qdrant does not guarantee atomic distributed updates but allows you to wait until the [operation is complete](../points/#awaiting-result) to see the results of your writes.
|
||||||
|
|
||||||
|
Collection operations, on the contrary, are part of the consensus.
|
||||||
|
It means that all nodes should agree on what operations should be applied before the service will perform them.
|
||||||
|
|
||||||
|
Practically, it means that if the cluster is in a transition state - either electing a new leader after a failure or starting up, the collection update operations will be denied.
|
||||||
|
|
||||||
|
You may use the cluster [REST API](https://qdrant.github.io/qdrant/redoc/index.html?v=master#tag/cluster) to check the state of the consensus.
|
||||||
|
|
||||||
## Sharding
|
## Sharding
|
||||||
|
|
||||||
Currently work-in-progress, see [Roadmap](https://github.com/qdrant/qdrant/blob/master/docs/roadmap/README.md)
|
A Collection in Qdrant is made of one or several shards.
|
||||||
|
Each shard is an independent storage of points which is able to perform all operations provided by collections.
|
||||||
|
Points are distributed among shards according to the [consistent hashing](https://en.wikipedia.org/wiki/Consistent_hashing) algorithm, so that shards are managing non-intersecting subsets of points.
|
||||||
|
|
||||||
## RAFT
|
During the creation of the collection, shards are evenly distributed across all existing nodes.
|
||||||
|
Each node knows where all parts of the collection are stored through the [consensus protocol](./#raft), so if it is time to search - each node could query all other nodes to obtain the full search result.
|
||||||
|
|
||||||
Currently work-in-progress, see [Roadmap](https://github.com/qdrant/qdrant/blob/master/docs/roadmap/README.md)
|
You can define number of shards in your create-collection request:
|
||||||
|
|
||||||
|
```http
|
||||||
|
PUT /collections/{collection_name}
|
||||||
|
|
||||||
|
{
|
||||||
|
"name": "example_collection",
|
||||||
|
"distance": "Cosine",
|
||||||
|
"vector_size": 300,
|
||||||
|
"shard_number": 6
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
|
||||||
|
client = QdrantClient(host="localhost", port=6333)
|
||||||
|
|
||||||
|
client.recreate_collection(
|
||||||
|
name="{collection_name}",
|
||||||
|
distance="Cosine",
|
||||||
|
vector_size=300,
|
||||||
|
shard_number=6
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
We recommend selecting the number of shards as a factor of the number of nodes you are currently running in your cluster.
|
||||||
|
For example, if you have 3 nodes, 6 shards could be a good option.
|
||||||
|
|
||||||
|
### Shard re-balancing
|
||||||
|
|
||||||
|
In case you want to extend your cluster with new nodes or some nodes become slower than the others, it might be helpful to re-balance shard alignment in the cluster.
|
||||||
|
|
||||||
|
Shard re-balancing operation will move shards from over-loaded nodes to less loaded, creating more even data distribution.
|
||||||
|
|
||||||
|
Shard re-balancing is currently work-in-progress, see [Roadmap](https://qdrant.to/roadmap)
|
||||||
|
|
||||||
|
## Replication
|
||||||
|
|
||||||
|
Currently work-in-progress, see [Roadmap](https://qdrant.to/roadmap)
|
||||||
|
|||||||
@@ -19,7 +19,7 @@ Let's take a look at the clauses implemented in Qdrant.
|
|||||||
|
|
||||||
Suppose we have a set of points with the following payload:
|
Suppose we have a set of points with the following payload:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 1, "city": "London", "color": "green"},
|
{"id": 1, "city": "London", "color": "green"},
|
||||||
{"id": 2, "city": "London", "color": "red"},
|
{"id": 2, "city": "London", "color": "red"},
|
||||||
@@ -34,7 +34,9 @@ Suppose we have a set of points with the following payload:
|
|||||||
|
|
||||||
Example:
|
Example:
|
||||||
|
|
||||||
```
|
```http
|
||||||
|
POST /collections/{collection_name}/points/scroll
|
||||||
|
|
||||||
{
|
{
|
||||||
"filter": {
|
"filter": {
|
||||||
"must": [
|
"must": [
|
||||||
@@ -46,9 +48,32 @@ Example:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
from qdrant_client.http import models
|
||||||
|
|
||||||
|
client = QdrantClient(host="localhost", port=6333)
|
||||||
|
|
||||||
|
client.scroll(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
scroll_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.Match(value="London"),
|
||||||
|
),
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.Match(value="red"),
|
||||||
|
),
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Filtered points would be:
|
Filtered points would be:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 2, "city": "London", "color": "red"}
|
{"id": 2, "city": "London", "color": "red"}
|
||||||
]
|
]
|
||||||
@@ -62,7 +87,9 @@ In this sense, `must` is equivalent to the operator `AND`.
|
|||||||
|
|
||||||
Example:
|
Example:
|
||||||
|
|
||||||
```
|
```http
|
||||||
|
POST /collections/{collection_name}/points/scroll
|
||||||
|
|
||||||
{
|
{
|
||||||
"filter": {
|
"filter": {
|
||||||
"should": [
|
"should": [
|
||||||
@@ -74,9 +101,27 @@ Example:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.scroll(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
scroll_filter=models.Filter(
|
||||||
|
should=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.Match(value="London"),
|
||||||
|
),
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.Match(value="red"),
|
||||||
|
),
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Filtered points would be:
|
Filtered points would be:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 1, "city": "London", "color": "green"},
|
{"id": 1, "city": "London", "color": "green"},
|
||||||
{"id": 2, "city": "London", "color": "red"},
|
{"id": 2, "city": "London", "color": "red"},
|
||||||
@@ -93,7 +138,9 @@ In this sense, `should` is equivalent to the operator `OR`.
|
|||||||
|
|
||||||
Example:
|
Example:
|
||||||
|
|
||||||
```
|
```http
|
||||||
|
POST /collections/{collection_name}/points/scroll
|
||||||
|
|
||||||
{
|
{
|
||||||
"filter": {
|
"filter": {
|
||||||
"must_not": [
|
"must_not": [
|
||||||
@@ -105,9 +152,27 @@ Example:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.scroll(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
scroll_filter=models.Filter(
|
||||||
|
must_not=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.Match(value="London")
|
||||||
|
),
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.Match(value="red")
|
||||||
|
),
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Filtered points would be:
|
Filtered points would be:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 5, "city": "Moscow", "color": "green"},
|
{"id": 5, "city": "Moscow", "color": "green"},
|
||||||
{"id": 6, "city": "Moscow", "color": "blue"}
|
{"id": 6, "city": "Moscow", "color": "blue"}
|
||||||
@@ -122,7 +187,9 @@ In this sense, `must_not` is equivalent to the expression `(NOT A) AND (NOT B) A
|
|||||||
|
|
||||||
It is also possible to use several clauses simultaneously:
|
It is also possible to use several clauses simultaneously:
|
||||||
|
|
||||||
```
|
```http
|
||||||
|
POST /collections/{collection_name}/points/scroll
|
||||||
|
|
||||||
{
|
{
|
||||||
"filter": {
|
"filter": {
|
||||||
"must": [
|
"must": [
|
||||||
@@ -136,9 +203,29 @@ It is also possible to use several clauses simultaneously:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.scroll(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
scroll_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.Match(value="London")
|
||||||
|
),
|
||||||
|
],
|
||||||
|
must_not=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.Match(value="red")
|
||||||
|
),
|
||||||
|
],
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Filtered points would be:
|
Filtered points would be:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 1, "city": "London", "color": "green"},
|
{"id": 1, "city": "London", "color": "green"},
|
||||||
{"id": 3, "city": "London", "color": "blue"},
|
{"id": 3, "city": "London", "color": "blue"},
|
||||||
@@ -147,9 +234,11 @@ Filtered points would be:
|
|||||||
|
|
||||||
In this case, the conditions are combined by `AND`.
|
In this case, the conditions are combined by `AND`.
|
||||||
|
|
||||||
Also the conditions could be recursively nested. Example:
|
Also, the conditions could be recursively nested. Example:
|
||||||
|
|
||||||
|
```http
|
||||||
|
POST /collections/{collection_name}/points/scroll
|
||||||
|
|
||||||
```
|
|
||||||
{
|
{
|
||||||
"filter": {
|
"filter": {
|
||||||
"must_not": [
|
"must_not": [
|
||||||
@@ -165,9 +254,31 @@ Also the conditions could be recursively nested. Example:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.scroll(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
scroll_filter=models.Filter(
|
||||||
|
must_not=[
|
||||||
|
models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.Match(value="London")
|
||||||
|
),
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.Match(value="red")
|
||||||
|
),
|
||||||
|
],
|
||||||
|
),
|
||||||
|
],
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Filtered points would be:
|
Filtered points would be:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 1, "city": "London", "color": "green"},
|
{"id": 1, "city": "London", "color": "green"},
|
||||||
{"id": 3, "city": "London", "color": "blue"},
|
{"id": 3, "city": "London", "color": "blue"},
|
||||||
@@ -184,7 +295,7 @@ Let's look at the existing condition variants and what types of data they apply
|
|||||||
|
|
||||||
### Match
|
### Match
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"key": "color",
|
"key": "color",
|
||||||
"match": {
|
"match": {
|
||||||
@@ -193,7 +304,16 @@ Let's look at the existing condition variants and what types of data they apply
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.MatchValue(value="red"),
|
||||||
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
|
For the other types, the match condition will look exactly the same, except for the type used:
|
||||||
|
|
||||||
|
```json
|
||||||
{
|
{
|
||||||
"key": "count",
|
"key": "count",
|
||||||
"match": {
|
"match": {
|
||||||
@@ -202,6 +322,13 @@ Let's look at the existing condition variants and what types of data they apply
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
models.FieldCondition(
|
||||||
|
key="count",
|
||||||
|
match=models.MatchValue(value=0),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
The simplest kind of condition is one that checks if the stored value equals the given one.
|
The simplest kind of condition is one that checks if the stored value equals the given one.
|
||||||
If several values are stored, at least one of them should match the condition.
|
If several values are stored, at least one of them should match the condition.
|
||||||
@@ -209,7 +336,7 @@ You can apply it to [keyword](../payload/#keyword), [integer](../payload/#intege
|
|||||||
|
|
||||||
### Range
|
### Range
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"key": "price",
|
"key": "price",
|
||||||
"range": {
|
"range": {
|
||||||
@@ -221,6 +348,18 @@ You can apply it to [keyword](../payload/#keyword), [integer](../payload/#intege
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
models.FieldCondition(
|
||||||
|
key="price",
|
||||||
|
range=models.Range(
|
||||||
|
gt=None,
|
||||||
|
gte=100.0,
|
||||||
|
lt=None,
|
||||||
|
lte=450.0,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
The `range` condition sets the range of possible values for stored payload values.
|
The `range` condition sets the range of possible values for stored payload values.
|
||||||
If several values are stored, at least one of them should match the condition.
|
If several values are stored, at least one of them should match the condition.
|
||||||
|
|
||||||
@@ -237,7 +376,7 @@ Can be applied to [float](../payload/#float) and [integer](../payload/#integer)
|
|||||||
|
|
||||||
#### Geo Bounding Box
|
#### Geo Bounding Box
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"key": "location",
|
"key": "location",
|
||||||
"geo_bounding_box": {
|
"geo_bounding_box": {
|
||||||
@@ -253,13 +392,28 @@ Can be applied to [float](../payload/#float) and [integer](../payload/#integer)
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
models.FieldCondition(
|
||||||
|
key="location",
|
||||||
|
geo_bounding_box=models.GeoBoundingBox(
|
||||||
|
bottom_right=models.GeoPoint(
|
||||||
|
lat=52.495862,
|
||||||
|
lon=13.455868,
|
||||||
|
),
|
||||||
|
top_left=models.GeoPoint(
|
||||||
|
lat=52.520711,
|
||||||
|
lon=13.403683,
|
||||||
|
),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
It matches with `location`s inside a rectangle with the coordinates of the upper left corner in `bottom_right` and the coordinates of the lower right corner in `top_left`.
|
It matches with `location`s inside a rectangle with the coordinates of the upper left corner in `bottom_right` and the coordinates of the lower right corner in `top_left`.
|
||||||
|
|
||||||
|
|
||||||
#### Geo Radius
|
#### Geo Radius
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"key": "location",
|
"key": "location",
|
||||||
"geo_radius": {
|
"geo_radius": {
|
||||||
@@ -272,6 +426,20 @@ It matches with `location`s inside a rectangle with the coordinates of the upper
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
models.FieldCondition(
|
||||||
|
key="location",
|
||||||
|
geo_radius=models.GeoRadius(
|
||||||
|
center=models.GeoPoint(
|
||||||
|
lat=52.520711,
|
||||||
|
lon=13.403683,
|
||||||
|
),
|
||||||
|
radius=1000.0,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
It matches with `location`s inside a circle with the `center` at the center and a radius of `radius` meters.
|
It matches with `location`s inside a circle with the `center` at the center and a radius of `radius` meters.
|
||||||
|
|
||||||
If several values are stored, at least one of them should match the condition.
|
If several values are stored, at least one of them should match the condition.
|
||||||
@@ -283,7 +451,7 @@ In addition to the direct value comparison, it is also possible to filter by the
|
|||||||
|
|
||||||
For example, given the data:
|
For example, given the data:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 1, "name": "product A", "comments": ["Very good!", "Excellent"]},
|
{"id": 1, "name": "product A", "comments": ["Very good!", "Excellent"]},
|
||||||
{"id": 2, "name": "product B", "comments": ["meh", "expected more", "ok"]},
|
{"id": 2, "name": "product B", "comments": ["meh", "expected more", "ok"]},
|
||||||
@@ -292,7 +460,7 @@ For example, given the data:
|
|||||||
|
|
||||||
We can perform the search only among the items with more than two comments:
|
We can perform the search only among the items with more than two comments:
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"key": "comments",
|
"key": "comments",
|
||||||
"values_count": {
|
"values_count": {
|
||||||
@@ -301,9 +469,16 @@ We can perform the search only among the items with more than two comments:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
models.FieldCondition(
|
||||||
|
key="comments",
|
||||||
|
values_count=models.ValuesCount(gt=2),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
The result would be:
|
The result would be:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 2, "name": "product B", "comments": ["meh", "expected more", "ok"]},
|
{"id": 2, "name": "product B", "comments": ["meh", "expected more", "ok"]},
|
||||||
]
|
]
|
||||||
@@ -316,7 +491,7 @@ If stored value is not an array - it is assumed that the amount of values is equ
|
|||||||
Sometimes it is also useful to filter out records that are missing some value.
|
Sometimes it is also useful to filter out records that are missing some value.
|
||||||
The `IsEmpty` condition may help you with that:
|
The `IsEmpty` condition may help you with that:
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"is_empty": {
|
"is_empty": {
|
||||||
"key": "reports"
|
"key": "reports"
|
||||||
@@ -324,7 +499,13 @@ The `IsEmpty` condition may help you with that:
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
This condition will match all records where the field `reports` either does not exists, or have `NULL` or `[]` value.
|
```python
|
||||||
|
models.IsEmptyCondition(
|
||||||
|
is_empty=models.PayloadField(key="reports"),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
This condition will match all records where the field `reports` either does not exist, or have `NULL` or `[]` value.
|
||||||
|
|
||||||
<aside role="status">The <b>IsEmpty</b> is often useful together with the logical negation <b>must_not</b>. In this case all non-empty values will be selected.</aside>
|
<aside role="status">The <b>IsEmpty</b> is often useful together with the logical negation <b>must_not</b>. In this case all non-empty values will be selected.</aside>
|
||||||
|
|
||||||
@@ -333,7 +514,9 @@ This condition will match all records where the field `reports` either does not
|
|||||||
This type of query is not related to payload, but can be very useful in some situations.
|
This type of query is not related to payload, but can be very useful in some situations.
|
||||||
For example, the user could mark some specific search results as irrelevant, or we want to search only among the specified points.
|
For example, the user could mark some specific search results as irrelevant, or we want to search only among the specified points.
|
||||||
|
|
||||||
```
|
```http
|
||||||
|
POST /collections/{collection_name}/points/scroll
|
||||||
|
|
||||||
{
|
{
|
||||||
"filter": {
|
"filter": {
|
||||||
"must": [
|
"must": [
|
||||||
@@ -344,10 +527,21 @@ For example, the user could mark some specific search results as irrelevant, or
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.scroll(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
scroll_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.HasIdCondition(has_id=[1, 3, 5, 7, 9, 11]),
|
||||||
|
],
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
|
||||||
Filtered points would be:
|
Filtered points would be:
|
||||||
|
|
||||||
```
|
```json
|
||||||
[
|
[
|
||||||
{"id": 1, "city": "London", "color": "green"},
|
{"id": 1, "city": "London", "color": "green"},
|
||||||
{"id": 3, "city": "London", "color": "blue"},
|
{"id": 3, "city": "London", "color": "blue"},
|
||||||
|
|||||||
@@ -22,9 +22,7 @@ Creating an index requires additional computational resources and memory, so cho
|
|||||||
|
|
||||||
To mark a field as indexable, you can use the following:
|
To mark a field as indexable, you can use the following:
|
||||||
|
|
||||||
REST API
|
```http
|
||||||
|
|
||||||
```
|
|
||||||
PUT /collections/{collection_name}/index
|
PUT /collections/{collection_name}/index
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -33,12 +31,15 @@ PUT /collections/{collection_name}/index
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
With Python client
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
|
||||||
|
client = QdrantClient(host="localhost", port=6333)
|
||||||
|
|
||||||
|
client.create_payload_index(collection_name="{collection_name}",
|
||||||
|
field_name="name_of_the_field_to_index",
|
||||||
|
field_type="keyword")
|
||||||
```
|
```
|
||||||
-->
|
|
||||||
|
|
||||||
Available field types are:
|
Available field types are:
|
||||||
|
|
||||||
@@ -75,10 +76,10 @@ storage:
|
|||||||
# Number of neighbours to consider during the index building.
|
# Number of neighbours to consider during the index building.
|
||||||
# Larger the value - more accurate the search, more time required to build index.
|
# Larger the value - more accurate the search, more time required to build index.
|
||||||
ef_construct: 100
|
ef_construct: 100
|
||||||
# Minimal amount of points for additional payload-based indexing.
|
# Minimal size (in KiloBytes) of vectors for additional payload-based indexing.
|
||||||
# If payload chunk is smaller than `full_scan_threshold` additional indexing won't be used -
|
# If payload chunk is smaller than `full_scan_threshold_kb` additional indexing won't be used -
|
||||||
# in this case full-scan search should be preferred by query planner
|
# in this case full-scan search should be preferred by query planner and additional indexing is not required.
|
||||||
# and additional indexing is not required.
|
# Note: 1Kb = 1 vector of size 256
|
||||||
full_scan_threshold: 10000
|
full_scan_threshold: 10000
|
||||||
|
|
||||||
```
|
```
|
||||||
|
|||||||
@@ -71,7 +71,14 @@ After a successful build, the binary is available at `./target/release/qdrant`.
|
|||||||
|
|
||||||
## With Kubernetes
|
## With Kubernetes
|
||||||
|
|
||||||
ToDo
|
You can use a ready-made [Helm Chart](https://helm.sh/docs/) to run Qdrant in your Kubeternetes cluster.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
helm repo add qdrant https://qdrant.to/helm
|
||||||
|
helm install qdrant-release qdrant/qdrant
|
||||||
|
```
|
||||||
|
|
||||||
|
Read further instructions in [qdrant-helm](https://github.com/qdrant/qdrant-helm) repository.
|
||||||
|
|
||||||
## Configuration
|
## Configuration
|
||||||
|
|
||||||
|
|||||||
@@ -75,16 +75,16 @@ Here is an example of parameter values:
|
|||||||
```yaml
|
```yaml
|
||||||
storage:
|
storage:
|
||||||
optimizers:
|
optimizers:
|
||||||
# Maximum number of vectors to store in-memory per segment.
|
# Maximum size (in KiloBytes) of vectors to store in-memory per segment.
|
||||||
# Segments larger than this threshold will be stored as read-only memmaped file.
|
# Segments larger than this threshold will be stored as read-only memmaped file.
|
||||||
memmap_threshold: 50000
|
# To enable memmap storage, lower the threshold
|
||||||
# Maximum number of vectors allowed for plain index.
|
# Note: 1Kb = 1 vector of size 256
|
||||||
# Default value based on
|
memmap_threshold_kb: 200000
|
||||||
# https://github.com/google-research/google-research/blob/master/scann/docs/algorithms.md
|
|
||||||
indexing_threshold: 20000
|
# Maximum size (in KiloBytes) of vectors allowed for plain index.
|
||||||
# Starting from this amount of vectors per-segment
|
# Default value based on https://github.com/google-research/google-research/blob/master/scann/docs/algorithms.md
|
||||||
# the engine will start building index for payload.
|
# Note: 1Kb = 1 vector of size 256
|
||||||
payload_indexing_threshold: 10000
|
indexing_threshold_kb: 20000
|
||||||
```
|
```
|
||||||
|
|
||||||
In addition to the configuration file, you can also set optimizer parameters separately for each [collection](../collections).
|
In addition to the configuration file, you can also set optimizer parameters separately for each [collection](../collections).
|
||||||
|
|||||||
@@ -138,39 +138,66 @@ Coordinate should be described as an object containing two fields: `lon` - for l
|
|||||||
|
|
||||||
## Create point with payload
|
## Create point with payload
|
||||||
|
|
||||||
With REST API
|
```http
|
||||||
|
PUT /collections/{collection_name}/points
|
||||||
```
|
|
||||||
PUT http://localhost:6333/collections/{collection_name}/points
|
|
||||||
|
|
||||||
{
|
{
|
||||||
"points": [
|
"points": [
|
||||||
{
|
{
|
||||||
"id": 1,
|
"id": 1,
|
||||||
"vector": [0.05, 0.61, 0.76, 0.74],
|
"vector": [0.05, 0.61, 0.76, 0.74],
|
||||||
"payload": {"city": "Berlin", price: 1.99}
|
"payload": {"city": "Berlin", "price": 1.99}
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"id": 2,
|
"id": 2,
|
||||||
"vector": [0.19, 0.81, 0.75, 0.11],
|
"vector": [0.19, 0.81, 0.75, 0.11],
|
||||||
"payload": {"city": ["Berlin", "London"], price: 1.99}
|
"payload": {"city": ["Berlin", "London"], "price": 1.99}
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"id": 3,
|
"id": 3,
|
||||||
"vector": [0.36, 0.55, 0.47, 0.94],
|
"vector": [0.36, 0.55, 0.47, 0.94],
|
||||||
"payload": {"city": ["Berlin", "Moscow"], price: [1.99, 2.99]}
|
"payload": {"city": ["Berlin", "Moscow"], "price": [1.99, 2.99]}
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
With python
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
```
|
from qdrant_client import QdrantClient
|
||||||
-->
|
from qdrant_client.http import models
|
||||||
|
|
||||||
|
client = QdrantClient(host="localhost", port=6333)
|
||||||
|
|
||||||
|
client.upsert(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=1,
|
||||||
|
vector=[0.05, 0.61, 0.76, 0.74],
|
||||||
|
payload={
|
||||||
|
"city": "Berlin",
|
||||||
|
"price": 1.99,
|
||||||
|
},
|
||||||
|
),
|
||||||
|
models.PointStruct(
|
||||||
|
id=2,
|
||||||
|
vector=[0.19, 0.81, 0.75, 0.11],
|
||||||
|
payload={
|
||||||
|
"city": ["Berlin", "London"],
|
||||||
|
"price": 1.99,
|
||||||
|
},
|
||||||
|
),
|
||||||
|
models.PointStruct(
|
||||||
|
id=3,
|
||||||
|
vector=[0.36, 0.55, 0.47, 0.94],
|
||||||
|
payload={
|
||||||
|
"city": ["Berlin", "Moscow"],
|
||||||
|
"price": [1.99, 2.99],
|
||||||
|
},
|
||||||
|
),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
## Update payload
|
## Update payload
|
||||||
|
|
||||||
@@ -179,7 +206,7 @@ PUT http://localhost:6333/collections/{collection_name}/points
|
|||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/set_payload)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/set_payload)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/payload
|
POST /collections/{collection_name}/points/payload
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -193,14 +220,16 @@ POST /collections/{collection_name}/points/payload
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
```
|
client.set_payload(
|
||||||
|
collection_name="{collection_name}",
|
||||||
-->
|
payload={
|
||||||
|
"property1": "string",
|
||||||
|
"property2": "string",
|
||||||
|
},
|
||||||
|
points=[0, 3, 10],
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
### Delete payload
|
### Delete payload
|
||||||
|
|
||||||
@@ -209,7 +238,7 @@ This method removes specified payload keys from specified points
|
|||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/delete_payload)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/delete_payload)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/payload/delete
|
POST /collections/{collection_name}/points/payload/delete
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -218,22 +247,21 @@ POST /collections/{collection_name}/points/payload/delete
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.delete_payload(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
keys=["color", "price"],
|
||||||
|
points=[0, 3, 100],
|
||||||
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
-->
|
|
||||||
|
|
||||||
### Clear payload
|
### Clear payload
|
||||||
|
|
||||||
This method removes all payload keys from specified points
|
This method removes all payload keys from specified points
|
||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/clear_payload)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/clear_payload)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/payload/clear
|
POST /collections/{collection_name}/points/payload/clear
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -241,13 +269,16 @@ POST /collections/{collection_name}/points/payload/clear
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
With Python client
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.clear_payload(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points_selector=models.PointIdsList(
|
||||||
|
points=[0, 3, 100],
|
||||||
|
)
|
||||||
|
)
|
||||||
```
|
```
|
||||||
-->
|
|
||||||
|
|
||||||
|
<aside role="status">You can also use `models.FilterSelector` to remove the points matching given filter criteria, instead of providing the ids.</aside>
|
||||||
|
|
||||||
## Payload indexing
|
## Payload indexing
|
||||||
|
|
||||||
@@ -265,7 +296,7 @@ To create index for the field, you can use the following:
|
|||||||
|
|
||||||
REST API
|
REST API
|
||||||
|
|
||||||
```
|
```http
|
||||||
PUT /collections/{collection_name}/index
|
PUT /collections/{collection_name}/index
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -274,12 +305,13 @@ PUT /collections/{collection_name}/index
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
Python client
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.create_payload_index(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
field_name="name_of_the_field_to_index",
|
||||||
|
field_type="keyword",
|
||||||
|
)
|
||||||
```
|
```
|
||||||
-->
|
|
||||||
|
|
||||||
The index usage flag is displayed in the payload schema with the [collection info API](https://qdrant.github.io/qdrant/redoc/index.html#operation/get_collection).
|
The index usage flag is displayed in the payload schema with the [collection info API](https://qdrant.github.io/qdrant/redoc/index.html#operation/get_collection).
|
||||||
|
|
||||||
|
|||||||
@@ -16,6 +16,8 @@ At the first stage, the operation is written to the Write-ahead-log.
|
|||||||
|
|
||||||
After this moment, the service will not lose the data, even if the machine loses power supply.
|
After this moment, the service will not lose the data, even if the machine loses power supply.
|
||||||
|
|
||||||
|
## Awaiting result
|
||||||
|
|
||||||
If the API is called with the `&wait=false` parameter, or if it is not explicitly specified, the client will receive an acknowledgment of receiving data:
|
If the API is called with the `&wait=false` parameter, or if it is not explicitly specified, the client will receive an acknowledgment of receiving data:
|
||||||
|
|
||||||
```json
|
```json
|
||||||
@@ -59,7 +61,7 @@ Examples of UUID string representations:
|
|||||||
That means that in every request UUID string could be used instead of numerical id.
|
That means that in every request UUID string could be used instead of numerical id.
|
||||||
Example:
|
Example:
|
||||||
|
|
||||||
```
|
```http
|
||||||
PUT /collections/{collection_name}/points
|
PUT /collections/{collection_name}/points
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -73,18 +75,29 @@ PUT /collections/{collection_name}/points
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
```
|
from qdrant_client import QdrantClient
|
||||||
|
from qdrant_client.http import models
|
||||||
|
|
||||||
-->
|
client = QdrantClient(host="localhost", port=6333)
|
||||||
|
|
||||||
|
client.upsert(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id="5c56c793-69f3-4fbf-87e6-c4bf54c28c26",
|
||||||
|
payload={
|
||||||
|
"color": "red",
|
||||||
|
},
|
||||||
|
vector=[0.9, 0.1, 0.1],
|
||||||
|
),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
and
|
and
|
||||||
|
|
||||||
```
|
```http
|
||||||
PUT /collections/{collection_name}/points
|
PUT /collections/{collection_name}/points
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -98,16 +111,22 @@ PUT /collections/{collection_name}/points
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.upsert(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=1,
|
||||||
|
payload={
|
||||||
|
"color": "red",
|
||||||
|
},
|
||||||
|
vector=[0.9, 0.1, 0.1],
|
||||||
|
),
|
||||||
|
]
|
||||||
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
-->
|
are both possible.
|
||||||
|
|
||||||
both are possible.
|
|
||||||
|
|
||||||
## Upload points
|
## Upload points
|
||||||
|
|
||||||
@@ -119,7 +138,7 @@ Internally, these options do not differ and are made only for the convenience of
|
|||||||
|
|
||||||
Create points with REST API :
|
Create points with REST API :
|
||||||
|
|
||||||
```
|
```http
|
||||||
PUT /collections/{collection_name}/points
|
PUT /collections/{collection_name}/points
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -133,15 +152,34 @@ PUT /collections/{collection_name}/points
|
|||||||
"vectors": [
|
"vectors": [
|
||||||
[0.9, 0.1, 0.1],
|
[0.9, 0.1, 0.1],
|
||||||
[0.1, 0.9, 0.1],
|
[0.1, 0.9, 0.1],
|
||||||
[0.1, 0.1, 0.9],
|
[0.1, 0.1, 0.9]
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.upsert(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points=models.Batch(
|
||||||
|
ids=[1, 2, 3],
|
||||||
|
payloads=[
|
||||||
|
{"color": "red"},
|
||||||
|
{"color": "green"},
|
||||||
|
{"color": "blue"},
|
||||||
|
],
|
||||||
|
vectors=[
|
||||||
|
[0.9, 0.1, 0.1],
|
||||||
|
[0.1, 0.9, 0.1],
|
||||||
|
[0.1, 0.1, 0.9],
|
||||||
|
]
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
or record-oriented equivalent:
|
or record-oriented equivalent:
|
||||||
|
|
||||||
```
|
```http
|
||||||
PUT /collections/{collection_name}/points
|
PUT /collections/{collection_name}/points
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -165,6 +203,35 @@ PUT /collections/{collection_name}/points
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.upsert(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points=[
|
||||||
|
models.PointStruct(
|
||||||
|
id=1,
|
||||||
|
payload={
|
||||||
|
"color": "red",
|
||||||
|
},
|
||||||
|
vector=[0.9, 0.1, 0.1],
|
||||||
|
),
|
||||||
|
models.PointStruct(
|
||||||
|
id=2,
|
||||||
|
payload={
|
||||||
|
"color": "green",
|
||||||
|
},
|
||||||
|
vector=[0.1, 0.9, 0.1],
|
||||||
|
),
|
||||||
|
models.PointStruct(
|
||||||
|
id=3,
|
||||||
|
payload={
|
||||||
|
"color": "blue",
|
||||||
|
},
|
||||||
|
vector=[0.1, 0.1, 0.9],
|
||||||
|
),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
<!--
|
<!--
|
||||||
|
|
||||||
The Python client has additional features for loading points.
|
The Python client has additional features for loading points.
|
||||||
@@ -195,7 +262,7 @@ The second is to modify the payload, for which there are several methods.
|
|||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/set_payload)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/set_payload)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/payload
|
POST /collections/{collection_name}/points/payload
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -209,20 +276,22 @@ POST /collections/{collection_name}/points/payload
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
```
|
client.set_payload(
|
||||||
|
collection_name="{collection_name}",
|
||||||
-->
|
payload={
|
||||||
|
"property1": "string",
|
||||||
|
"property2": "string",
|
||||||
|
},
|
||||||
|
points=[0, 3, 10],
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
#### Delete payload keys
|
#### Delete payload keys
|
||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/delete_payload)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/delete_payload)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/payload/delete
|
POST /collections/{collection_name}/points/payload/delete
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -231,22 +300,21 @@ POST /collections/{collection_name}/points/payload/delete
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.delete_payload(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
keys=["color", "price"],
|
||||||
|
points=[0, 3, 100],
|
||||||
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
-->
|
|
||||||
|
|
||||||
#### Clear payload
|
#### Clear payload
|
||||||
|
|
||||||
This method removes all payload keys from specified points
|
This method removes all payload keys from specified points
|
||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/clear_payload)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/clear_payload)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/payload/clear
|
POST /collections/{collection_name}/points/payload/clear
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -254,21 +322,21 @@ POST /collections/{collection_name}/points/payload/clear
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
```
|
client.clear_payload(
|
||||||
|
collection_name="{collection_name}",
|
||||||
-->
|
points_selector=models.PointIdsList(
|
||||||
|
points=[0, 3, 100],
|
||||||
|
)
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
## Delete points
|
## Delete points
|
||||||
|
|
||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/delete_points)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/delete_points)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/delete
|
POST /collections/{collection_name}/points/delete
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -276,18 +344,18 @@ POST /collections/{collection_name}/points/delete
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.delete(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points_selector=models.PointIdsList(
|
||||||
|
points=[0, 3, 100],
|
||||||
|
),
|
||||||
|
)
|
||||||
```
|
```
|
||||||
|
|
||||||
-->
|
|
||||||
|
|
||||||
Alternative way to specify which points to remove is to use filter.
|
Alternative way to specify which points to remove is to use filter.
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/delete
|
POST /collections/{collection_name}/points/delete
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -304,6 +372,22 @@ POST /collections/{collection_name}/points/delete
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.delete(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
points_selector=models.FilterSelector(
|
||||||
|
filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.MatchValue(value="red"),
|
||||||
|
),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
This example removes all points with `{ "color": "red" }` from the collection.
|
This example removes all points with `{ "color": "red" }` from the collection.
|
||||||
|
|
||||||
|
|
||||||
@@ -314,7 +398,7 @@ There is a method for retrieving points by their ids.
|
|||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/get_points)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/get_points)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points
|
POST /collections/{collection_name}/points
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -322,13 +406,12 @@ POST /collections/{collection_name}/points
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
<!--
|
|
||||||
|
|
||||||
Python client:
|
|
||||||
|
|
||||||
```python
|
```python
|
||||||
|
client.retrieve(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
ids=[0, 3, 10],
|
||||||
|
)
|
||||||
```
|
```
|
||||||
-->
|
|
||||||
|
|
||||||
This method has additional parameters `with_vector` and `with_payload`.
|
This method has additional parameters `with_vector` and `with_payload`.
|
||||||
Using these parameters, you can select parts of the point you want as a result.
|
Using these parameters, you can select parts of the point you want as a result.
|
||||||
@@ -339,7 +422,7 @@ The single point can also be retrieved via the API:
|
|||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/get_point)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/get_point)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
GET /collections/{collection_name}/points/{point_id}
|
GET /collections/{collection_name}/points/{point_id}
|
||||||
```
|
```
|
||||||
|
|
||||||
@@ -357,7 +440,7 @@ Sometimes it might be necessary to get all stored points without knowing ids, or
|
|||||||
|
|
||||||
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/scroll_points)):
|
REST API ([Schema](https://qdrant.github.io/qdrant/redoc/index.html#operation/scroll_points)):
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/scroll
|
POST /collections/{collection_name}/points/scroll
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -377,6 +460,23 @@ POST /collections/{collection_name}/points/scroll
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.scroll(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
scroll_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="color",
|
||||||
|
match=models.Match(value="red")
|
||||||
|
),
|
||||||
|
]
|
||||||
|
),
|
||||||
|
limit=1,
|
||||||
|
with_payload=True,
|
||||||
|
with_vector=False,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Returns all point with `color` = `red`.
|
Returns all point with `color` = `red`.
|
||||||
|
|
||||||
```json
|
```json
|
||||||
|
|||||||
@@ -79,7 +79,7 @@ curl 'http://localhost:6333/collections/test_collection'
|
|||||||
|
|
||||||
Expected response:
|
Expected response:
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"result": {
|
"result": {
|
||||||
"status": "green",
|
"status": "green",
|
||||||
|
|||||||
@@ -60,8 +60,9 @@ Let's look at an example of a search query.
|
|||||||
|
|
||||||
REST API - API Schema definition is available [here](https://qdrant.github.io/qdrant/redoc/index.html#operation/search_points)
|
REST API - API Schema definition is available [here](https://qdrant.github.io/qdrant/redoc/index.html#operation/search_points)
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/search
|
POST /collections/{collection_name}/points/search
|
||||||
|
|
||||||
{
|
{
|
||||||
"filter": {
|
"filter": {
|
||||||
"must": [
|
"must": [
|
||||||
@@ -81,11 +82,31 @@ POST /collections/{collection_name}/points/search
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
||||||
<!--
|
|
||||||
```python
|
```python
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
from qdrant_client.http import models
|
||||||
|
|
||||||
|
client = QdrantClient(host="localhost", port=6333)
|
||||||
|
|
||||||
|
client.search(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
query_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.MatchValue(
|
||||||
|
value="London",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
]
|
||||||
|
),
|
||||||
|
search_params=models.SearchParams(
|
||||||
|
hnsw_ef=128
|
||||||
|
),
|
||||||
|
query_vector=[0.2, 0.1, 0.9, 0.7],
|
||||||
|
top=3,
|
||||||
|
)
|
||||||
```
|
```
|
||||||
-->
|
|
||||||
|
|
||||||
In this example, we are looking for vectors similar to vector `[0.2, 0.1, 0.9, 0.7]`.
|
In this example, we are looking for vectors similar to vector `[0.2, 0.1, 0.9, 0.7]`.
|
||||||
Parameter `top` specifies the amount of most similar results we would like to retrieve.
|
Parameter `top` specifies the amount of most similar results we would like to retrieve.
|
||||||
@@ -101,7 +122,7 @@ See details of possible filters and their work in the [filtering](../filtering)
|
|||||||
|
|
||||||
Example result of this API would be
|
Example result of this API would be
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"result": [
|
"result": [
|
||||||
{ "id": 10, "score": 0.81 },
|
{ "id": 10, "score": 0.81 },
|
||||||
@@ -115,6 +136,15 @@ Example result of this API would be
|
|||||||
|
|
||||||
The `result` contains ordered by `score` list of found point ids.
|
The `result` contains ordered by `score` list of found point ids.
|
||||||
|
|
||||||
|
### Filtering results by score
|
||||||
|
|
||||||
|
In addition to payload filtering, it might be useful to filter out results with a low similarity score.
|
||||||
|
For example, if you know the minimal acceptance score for your model and do not want any results which are less similar than the threshold.
|
||||||
|
In this case, you can use `score_threshold` parameter of the search query.
|
||||||
|
It will exclude all results with a score worse than the given.
|
||||||
|
|
||||||
|
<aside role="status">This parameter may exclude lower or higher scores depending on the used metric. For example, higher scores of Euclidean metric are considered more distant and, therefore, will be excluded.</aside>
|
||||||
|
|
||||||
### Payload in vector in the result
|
### Payload in vector in the result
|
||||||
|
|
||||||
By default, retrieval methods do not return any stored information.
|
By default, retrieval methods do not return any stored information.
|
||||||
@@ -122,8 +152,9 @@ Additional parameters `with_vector` and `with_payload` could alter this behavior
|
|||||||
|
|
||||||
Example:
|
Example:
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/search
|
POST /collections/{collection_name}/points/search
|
||||||
|
|
||||||
{
|
{
|
||||||
"vector": [0.2, 0.1, 0.9, 0.7],
|
"vector": [0.2, 0.1, 0.9, 0.7],
|
||||||
"with_vector": true,
|
"with_vector": true,
|
||||||
@@ -131,10 +162,20 @@ POST /collections/{collection_name}/points/search
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.search(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
query_vector=[0.2, 0.1, 0.9, 0.7],
|
||||||
|
with_vector=True,
|
||||||
|
with_payload=True,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Parameter `with_payload` might also be used to include or exclude specific fields only:
|
Parameter `with_payload` might also be used to include or exclude specific fields only:
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/search
|
POST /collections/{collection_name}/points/search
|
||||||
|
|
||||||
{
|
{
|
||||||
"vector": [0.2, 0.1, 0.9, 0.7],
|
"vector": [0.2, 0.1, 0.9, 0.7],
|
||||||
"with_payload": {
|
"with_payload": {
|
||||||
@@ -143,6 +184,16 @@ POST /collections/{collection_name}/points/search
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.search(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
query_vector=[0.2, 0.1, 0.9, 0.7],
|
||||||
|
with_payload=models.PayloadSelectorExclude(
|
||||||
|
exclude=["city"],
|
||||||
|
),
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
## Recommendation API
|
## Recommendation API
|
||||||
|
|
||||||
<aside role="alert">Negative vectors is an experimental functionality that is not guaranteed to work with all kind of embeddings.</aside>
|
<aside role="alert">Negative vectors is an experimental functionality that is not guaranteed to work with all kind of embeddings.</aside>
|
||||||
@@ -154,13 +205,15 @@ The recommendation API allows specifying several positive and negative vector ID
|
|||||||
|
|
||||||
` average_vector = avg(positive_vectors) + ( avg(positive_vectors) - avg(negative_vectors) )`
|
` average_vector = avg(positive_vectors) + ( avg(positive_vectors) - avg(negative_vectors) )`
|
||||||
|
|
||||||
|
If there is only one positive ID provided - this request is equivalent to the regular search with vector of that point.
|
||||||
|
|
||||||
Vector components that have a greater value in a negative vector are penalized, and those that have a greater value in a positive vector, on the contrary, are amplified.
|
Vector components that have a greater value in a negative vector are penalized, and those that have a greater value in a positive vector, on the contrary, are amplified.
|
||||||
This average vector will be used to find the most similar vectors in the collection.
|
This average vector will be used to find the most similar vectors in the collection.
|
||||||
|
|
||||||
|
|
||||||
REST API - API Schema definition is available [here](https://qdrant.github.io/qdrant/redoc/index.html#operation/recommend_points)
|
REST API - API Schema definition is available [here](https://qdrant.github.io/qdrant/redoc/index.html#operation/recommend_points)
|
||||||
|
|
||||||
```
|
```http
|
||||||
POST /collections/{collection_name}/points/recommend
|
POST /collections/{collection_name}/points/recommend
|
||||||
|
|
||||||
{
|
{
|
||||||
@@ -180,9 +233,28 @@ POST /collections/{collection_name}/points/recommend
|
|||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
client.recommend(
|
||||||
|
collection_name="{collection_name}",
|
||||||
|
query_filter=models.Filter(
|
||||||
|
must=[
|
||||||
|
models.FieldCondition(
|
||||||
|
key="city",
|
||||||
|
match=models.MatchValue(
|
||||||
|
value="London",
|
||||||
|
),
|
||||||
|
)
|
||||||
|
]
|
||||||
|
),
|
||||||
|
negative=[718],
|
||||||
|
positive=[100, 231],
|
||||||
|
top=10,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
Example result of this API would be
|
Example result of this API would be
|
||||||
|
|
||||||
```
|
```json
|
||||||
{
|
{
|
||||||
"result": [
|
"result": [
|
||||||
{ "id": 10, "score": 0.81 },
|
{ "id": 10, "score": 0.81 },
|
||||||
@@ -193,9 +265,3 @@ Example result of this API would be
|
|||||||
"time": 0.001
|
"time": 0.001
|
||||||
}
|
}
|
||||||
```
|
```
|
||||||
|
|
||||||
|
|
||||||
<!--
|
|
||||||
```python
|
|
||||||
```
|
|
||||||
-->
|
|
||||||
|
|||||||
@@ -33,8 +33,21 @@ Thus, segments using mmap storage are `non-appendable` and can only be construed
|
|||||||
|
|
||||||
## Payload storage
|
## Payload storage
|
||||||
|
|
||||||
In the current version of Qdrant, payload storage is organized in the same way as in-memory vectors.
|
Qdrant supports two types of payload storages: InMemory and OnDisk.
|
||||||
Payload is loaded into RAM at service startup while disk and [RocksDB](https://rocksdb.org/) are used for persistence.
|
|
||||||
|
InMemory payload storage is organized in the same way as in-memory vectors.
|
||||||
|
Payload is loaded into RAM at service startup while disk and [RocksDB](https://rocksdb.org/) are used for persistence only.
|
||||||
|
This type of storage works quite fast, but it may require a lot of space to keep all the data in RAM, especially if the payload has large values attached - abstracts of text or even images.
|
||||||
|
|
||||||
|
In the case of large payload values, it might be better to use OnDisk payload storage.
|
||||||
|
This type of storage will read and write payload directly to RocksDB, so it won't require any significant amount of RAM to store.
|
||||||
|
The downside, however, is the access latency.
|
||||||
|
If you need to query vectors with some payload-based conditions - checking values stored on disk might take too much time.
|
||||||
|
In this scenario, we recommend creating a payload index for each field used in filtering conditions to avoid disk access.
|
||||||
|
Once you create the field index, Qdrant will preserve all values of the indexed field in RAM regardless of the payload storage type.
|
||||||
|
|
||||||
|
You can specify the desired type of payload storage with [configuration file](../configuration/) or with collection parameter `on_disk_payload` during [creation](../collections/#create-collection) of the collection.
|
||||||
|
|
||||||
|
|
||||||
## Versioning
|
## Versioning
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user