mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-05 19:08:32 +02:00
Review feedback
This commit is contained in:
@@ -11,7 +11,7 @@ Qdrant is a vector search engine, making it a great tool for [semantic search](#
|
||||
|
||||
### Semantic Search
|
||||
|
||||
Semantic search is a search technique that focuses on the meaning of the text rather than just matching on keywords. This is achieved by converting text into [dense vectors](/documentation/concepts/vectors/#dense-vectors) (embeddings) using machine learning models. These vectors capture the semantic meaning of the text, enabling you to find similar text even if it doesn't share exact keywords.
|
||||
Semantic search is a search technique that focuses on the meaning of the text rather than just matching on keywords. This is achieved by converting text into [vectors](/documentation/concepts/vectors/) (embeddings) using machine learning models. These vectors capture the semantic meaning of the text, enabling you to find similar text even if it doesn't share exact keywords.
|
||||
|
||||
For example, to search through a collection of books, you could use a model like the `all-MiniLM-L6-v2` sentence transformer model. First, create a collection and configure a dense vector for the book descriptions:
|
||||
|
||||
@@ -82,11 +82,11 @@ When it comes to lexical search in Qdrant, it's important to distinguish between
|
||||
|
||||
## Filtering
|
||||
|
||||
Qdrant supports [filtering](/documentation/concepts/filtering) on a wide range of datatypes: numbers, dates, booleans, geolocations, and strings. In Qdrant, a filter is always combined with a vector query. The vector query is used to score and rank the results, while the filter is used to narrow down the results based on specific criteria.
|
||||
Qdrant supports [filtering](/documentation/concepts/filtering) on a wide range of datatypes: numbers, dates, booleans, geolocations, and strings. In Qdrant, a filter is typically combined with a vector query. The vector query is used to score and rank the results, while the filter is used to narrow down the results based on specific criteria.
|
||||
|
||||
### Text and Keyword Strings
|
||||
|
||||
When it comes to filtering on strings, it is important to understand the difference between the two types of strings in Qdrant: text and keyword. These two string types are designed for different use cases: filtering on exact string values or filtering on individual search terms. To filter on exact string values, Qdrant uses **keyword** strings. Keyword strings are ideal for filtering on strings like IDs, categories, or tags. To filter on individual terms within a larger body of text, Qdrant uses **text** strings.
|
||||
When it comes to filtering on strings, it is important to understand the difference between the two types of strings in Qdrant: text and keyword. These two string types are designed for different use cases: filtering on exact string values or filtering on individual search terms. To filter on exact string values, Qdrant uses **keyword** strings. Keyword strings are ideal for filtering on strings like IDs, categories, or tags. To filter on individual terms or phrases within a larger body of text, Qdrant uses **text** strings.
|
||||
|
||||
For example, take a string like "United States". If you want to filter on all points with this exact string in the payload, use a keyword filter. On the other hand, if you want to filter on all points that contain the word "united" (matching "United States" as well as "United Kingdom"), use a text filter.
|
||||
|
||||
@@ -399,7 +399,7 @@ The response contains three separate result sets. You can return the first non-e
|
||||
|
||||
Full-text search is similar to full-text filtering, with the key difference being that full-text queries are used for ranking. For each document that matches the search terms, Qdrant calculates a relevance score based on how well the document matches the search terms. That score is used to rank the results. Qdrant supports several full-text search scoring algorithms.
|
||||
|
||||
Full-text search in Qdrant is powered by [sparse vectors](/articles/sparse-vectors/). Why sparse vectors? Because they are a flexible way to represent data for search purposes, from classic BM25-based search, to semantic search, and [collaborative filtering](/documentation/advanced-tutorials/collaborative-filtering/). Each dimension in a sparse vector corresponds to a specific term in the vocabulary, and the value in that dimension represents the weight of that term in the document. Weights can be calculated using document statistics for use with the [BM25](#bm25) ranking algorithm, or you can use transformer-based models that can capture semantic meaning, like [SPLADE++](#splade), and [miniCOIL](#minicoil).
|
||||
Full-text search in Qdrant is powered by [sparse vectors](/articles/sparse-vectors/). Why sparse vectors? Because they are a flexible way to represent data for search purposes, from classic BM25-based search, to semantic search, and [collaborative filtering](/documentation/advanced-tutorials/collaborative-filtering/). Each term in the vocabulary corresponds to one or more dimension of the sparse vector, and the values in those dimensions represent the weight of that term in the document. Weights can be calculated using document statistics for use with the [BM25](#bm25) ranking algorithm, or you can use transformer-based models that can capture semantic meaning, like [SPLADE++](#splade), and [miniCOIL](#minicoil).
|
||||
|
||||
### BM25
|
||||
|
||||
@@ -447,7 +447,7 @@ POST /collections/books/points/query
|
||||
|
||||
The SPLADE (Sparse Lexical and Dense) family of models are transformer-based models that generate sparse vectors out of text. These models combine the benefits of traditional lexical search with the power of transformer-based models by generating homonyms and synonyms.
|
||||
|
||||
The advantage of using SPLADE models is that they [perform better](/articles/sparse-vectors/#splade) than traditional BM25. The downside is that they use a fixed vocabulary, which means that you can't use them to find terms that are not in the vocabulary, such as product IDs.
|
||||
The advantage of using SPLADE models is that they [perform better](/articles/sparse-vectors/#splade) than traditional BM25. They also have several downsides though. First, because they use a fixed vocabulary, you can't use SPLADE models to find terms that are not in the vocabulary, such as product IDs and out-of-domain language (words not seen in training). Secondly, because they are transformer-based models, SPLADE models are slower and require more computational resources than the traditional BM25 model.
|
||||
|
||||
On [Qdrant Cloud](/documentation/concepts/inference/#qdrant-cloud-inference), you can use the SPLADE++ model with inference. Alternatively, you can generate vectors on the client side using the [FastEmbed](/documentation/fastembed/) library.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user