minor touches

This commit is contained in:
Arnaud Gourlay
2023-12-08 06:45:01 +01:00
parent 9094d94d9c
commit 235d173e0a
3 changed files with 13 additions and 9 deletions
@@ -370,9 +370,11 @@ client
.await?;
```
There are no required configuration parameters for named sparse vectors.
Outside of a unique name, there are no required configuration parameters for sparse vectors.
However, there are optional parameters to tune the underlying [sparse index](../indexing/#sparse-vector-index).
The distance function for sparse vectors is always `Dot` and does not need to be specified.
However, there are optional parameters to tune the underlying [sparse vector index](../indexing/#sparse-vector-index).
### Delete collection
@@ -234,20 +234,20 @@ performance.
Qdrant supports sparse vectors, which are vectors with a large number of zeroes.
We can take advantage of this property to index vector in a specialized way, which allows to save space and speed up search.
We can take advantage of this property to index the vectors in a specialized way, which allows to save space and speed up search.
The underlying index is an inverted index, which stores the list of vectors for each non-zero dimension.
The underlying structure is an inverted index, which stores the list of vectors for each non-zero dimension.
Upon search, the index is used to find the list of vectors that have non-zero values in the query dimensions.
Then, the vectors are scored using the dot product.
There are optimizations in place to reduce the number of vectors to score for dimensions with a large number of vectors.
The sparse vector index supports filtering by payload fields, which allows to use it in combination with the payload index.
Similar to dense vectors, the sparse vector index supports filtering by payload fields, which allows to use it in combination with indexed payload fields.
Similar to the dense vector, it is possible configure `full_scan_threshold` to control when to drive the search from the payload index to decrease the number of vectors to score.
It is possible configure `full_scan_threshold` to control when to drive the search from the payload index to decrease the number of vectors to score.
In the case of sparse vectors, the threshold is specified in the number of vectors, not in the size of the payload.
In the case of sparse vectors, the threshold is specified in the number of matching vectors found by the query planner.
The index always resides in memory for appendable segments providing fast search and update operations by default.
@@ -258,8 +258,10 @@ You can still use payload filtering and other features of the search API with sp
There are however important differences between dense and sparse vector search:
- only `Dot` metric is supported for sparse vectors (no need to specify it in the request)
- the search is not approximate, it is always returning the exact match
- it returns only the vectors which have non-zero values in the same indices as the query vector, for this reason, you can can receive less than `limit` results.
- the sparse search is not approximate, it is always returning the exact match
- the spearse search returns only the vectors which have non-zero values in the same indices as the query vector, for this reason, you can can receive less than `limit` results.
In general, the speed of the search is proportional to the number of non-zero values in the query vector.
```http
POST /collections/{collection_name}/points/search