diff --git a/qdrant-landing/content/documentation/concepts/filtering.md b/qdrant-landing/content/documentation/concepts/filtering.md index 228e96444..30f2b887a 100644 --- a/qdrant-landing/content/documentation/concepts/filtering.md +++ b/qdrant-landing/content/documentation/concepts/filtering.md @@ -319,6 +319,23 @@ If there is no full-text index for the field, the condition will work as exact s If the query has several words, then the condition will be satisfied only if all of them are present in the text. +### Phrase Match + +*Available as of v1.15.0* + +A match `phrase` condition also leverages [full-text index](/documentation/concepts/indexing/#full-text-index), to perform exact phrase comparisons. +It allows you to search for a specific token phrase within the text field. + +For example, the text `"quick brown fox"` will be matched by the query `"brown fox"`, but not by `"fox brown"`. + + + +If there is no full-text index for the field, the condition will work as exact substring match. + +{{< code-snippet path="/documentation/headless/snippets/filter-condition/phrase-match/" >}} + ### Range {{< code-snippet path="/documentation/headless/snippets/filter-condition/range/" >}} @@ -337,7 +354,7 @@ Can be applied to [float](/documentation/concepts/payload/#float) and [integer]( ### Datetime Range -The datetime range is a unique range condition, used for [datetime](/documentation/concepts/payload/#datetime) payloads, which supports RFC 3339 formats. +The datetime range is a unique range condition, used for [datetime](/documentation/concepts/payload/#datetime) payloads, which supports RFC 3339 formats. You do not need to convert dates to UNIX timestaps. During comparison, timestamps are parsed and converted to UTC. _Available as of v1.8.0_ @@ -371,9 +388,9 @@ If several values are stored, at least one of them should match the condition. These conditions can only be applied to payloads that match the [geo-data format](/documentation/concepts/payload/#geo). #### Geo Polygon -Geo Polygons search is useful for when you want to find points inside an irregularly shaped area, for example a country boundary or a forest boundary. A polygon always has an exterior ring and may optionally include interior rings. A lake with an island would be an example of an interior ring. If you wanted to find points in the water but not on the island, you would make an interior ring for the island. +Geo Polygons search is useful for when you want to find points inside an irregularly shaped area, for example a country boundary or a forest boundary. A polygon always has an exterior ring and may optionally include interior rings. A lake with an island would be an example of an interior ring. If you wanted to find points in the water but not on the island, you would make an interior ring for the island. -When defining a ring, you must pick either a clockwise or counterclockwise ordering for your points. The first and last point of the polygon must be the same. +When defining a ring, you must pick either a clockwise or counterclockwise ordering for your points. The first and last point of the polygon must be the same. Currently, we only support unprojected global coordinates (decimal degrees longitude and latitude) and we are datum agnostic. @@ -381,7 +398,7 @@ Currently, we only support unprojected global coordinates (decimal degrees longi A match is considered any point location inside or on the boundaries of the given polygon's exterior but not inside any interiors. -If several location values are stored for a point, then any of them matching will include that point as a candidate in the resultset. +If several location values are stored for a point, then any of them matching will include that point as a candidate in the resultset. These conditions can only be applied to payloads that match the [geo-data format](/documentation/concepts/payload/#geo). ### Values count @@ -422,7 +439,7 @@ This condition will match all records where the field `reports` either does not ### Is Null -It is not possible to test for `NULL` values with the match condition. +It is not possible to test for `NULL` values with the match condition. We have to use `IsNull` condition instead: {{< code-snippet path="/documentation/headless/snippets/filter-condition/is-null/" >}} diff --git a/qdrant-landing/content/documentation/concepts/indexing.md b/qdrant-landing/content/documentation/concepts/indexing.md index bb8385c61..c6efad38d 100644 --- a/qdrant-landing/content/documentation/concepts/indexing.md +++ b/qdrant-landing/content/documentation/concepts/indexing.md @@ -7,7 +7,7 @@ aliases: # Indexing -A key feature of Qdrant is the effective combination of vector and traditional indexes. It is essential to have this because for vector search to work effectively with filters, having vector index only is not enough. In simpler terms, a vector index speeds up vector search, and payload indexes speed up filtering. +A key feature of Qdrant is the effective combination of vector and traditional indexes. It is essential to have this because for vector search to work effectively with filters, having a vector index only is not enough. In simpler terms, a vector index speeds up vector search, and payload indexes speed up filtering. The indexes in the segments exist independently, but the parameters of the indexes themselves are configured for the whole collection. @@ -37,37 +37,14 @@ Available field types are: * `bool` - for [bool](/documentation/concepts/payload/#bool) payload, affects [Match](/documentation/concepts/filtering/#match) filtering conditions (available as of v1.4.0). * `geo` - for [geo](/documentation/concepts/payload/#geo) payload, affects [Geo Bounding Box](/documentation/concepts/filtering/#geo-bounding-box) and [Geo Radius](/documentation/concepts/filtering/#geo-radius) filtering conditions. * `datetime` - for [datetime](/documentation/concepts/payload/#datetime) payload, affects [Range](/documentation/concepts/filtering/#range) filtering conditions (available as of v1.8.0). -* `text` - a special kind of index, available for [keyword](/documentation/concepts/payload/#keyword) / string payloads, affects [Full Text search](/documentation/concepts/filtering/#full-text-match) filtering conditions. +* `text` - a special kind of index, available for [keyword](/documentation/concepts/payload/#keyword) / string payloads, affects [Full Text search](/documentation/concepts/filtering/#full-text-match) filtering conditions. Read more about [text index configuration](#full-text-index) * `uuid` - a special type of index, similar to `keyword`, but optimized for [UUID values](/documentation/concepts/payload/#uuid). Affects [Match](/documentation/concepts/filtering/#match) filtering conditions. (available as of v1.11.0) -Payload index may occupy some additional memory, so it is recommended to only use index for those fields that are used in filtering conditions. -If you need to filter by many fields and the memory limits does not allow to index all of them, it is recommended to choose the field that limits the search result the most. +Payload index may occupy some additional memory, so it is recommended to only use the index for those fields that are used in filtering conditions. +If you need to filter by many fields and the memory limits do not allow for indexing all of them, it is recommended to choose the field that limits the search result the most. As a rule, the more different values a payload value has, the more efficiently the index will be used. -### Full-text index - -*Available as of v0.10.0* - -Qdrant supports full-text search for string payload. -Full-text index allows you to filter points by the presence of a word or a phrase in the payload field. - -Full-text index configuration is a bit more complex than other indexes, as you can specify the tokenization parameters. -Tokenization is the process of splitting a string into tokens, which are then indexed in the inverted index. - -To create a full-text index, you can use the following: - -{{< code-snippet path="/documentation/headless/snippets/create-payload-index/simple-full-text/" >}} - -Available tokenizers are: - -* `word` - splits the string into words, separated by spaces, punctuation marks, and special characters. -* `whitespace` - splits the string into words, separated by spaces. -* `prefix` - splits the string into words, separated by spaces, punctuation marks, and special characters, and then creates a prefix index for each word. For example: `hello` will be indexed as `h`, `he`, `hel`, `hell`, `hello`. -* `multilingual` - special type of tokenizer based on [charabia](https://github.com/meilisearch/charabia) package. It allows proper tokenization and lemmatization for multiple languages, including those with non-latin alphabets and non-space delimiters. See [charabia documentation](https://github.com/meilisearch/charabia) for full list of supported languages supported normalization options. In the default build configuration, qdrant does not include support for all languages, due to the increasing size of the resulting binary. Chinese, Japanese and Korean languages are not enabled by default, but can be enabled by building qdrant from source with `--features multiling-chinese,multiling-japanese,multiling-korean` flags. - -See [Full Text match](/documentation/concepts/filtering/#full-text-match) for examples of querying with full-text index. - ### Parameterized index *Available as of v1.8.0* @@ -78,9 +55,9 @@ you to fine-tune indexing and search performance. Both the regular and parameterized `integer` indexes use the following flags: - `lookup`: enables support for direct lookup using - [Match](/documentation/concepts/filtering/#match) filters. + [Match](/documentation/concepts/filtering/#match) filters. - `range`: enables support for - [Range](/documentation/concepts/filtering/#range) filters. + [Range](/documentation/concepts/filtering/#range) filters. The regular `integer` index assumes both `lookup` and `range` are `true`. In contrast, to configure a parameterized index, you would set only one of these @@ -88,10 +65,10 @@ filters to `true`: | `lookup` | `range` | Result | |----------|---------|-----------------------------| -| `true` | `true` | Regular integer index | -| `true` | `false` | Parameterized integer index | -| `false` | `true` | Parameterized integer index | -| `false` | `false` | No integer index | +| `true` | `true` | Regular integer index | +| `true` | `false` | Parameterized integer index | +| `false` | `true` | Parameterized integer index | +| `false` | `false` | No integer index | The parameterized index can enhance performance in collections with millions of points. We encourage you to try it out. If it does not enhance performance @@ -115,14 +92,14 @@ As latency in this case is critical, it is recommended to keep hot payload index There are, however, cases when payload indexes are too large or rarely used. In those cases, it is possible to store payload indexes on disk. To configure on-disk payload index, you can use the following index parameters: {{< code-snippet path="/documentation/headless/snippets/create-payload-index/keyword-on-disk/" >}} -Payload index on-disk is supported for following types: +Payload index on-disk is supported for the following types: * `keyword` * `integer` @@ -138,12 +115,12 @@ The list will be extended in future versions. *Available as of v1.11.0* -Many vector search use-cases require multitenancy. In a multi-tenant scenario the collection is expected to contain multiple subsets of data, where each subset belongs to a different tenant. +Many vector search use-cases require multitenancy. In a multi-tenant scenario the collection is expected to contain multiple subsets of data, where each subset belongs to a different tenant. Qdrant supports efficient multi-tenant search by enabling [special configuration](/documentation/guides/multiple-partitions/) vector index, which disables global search and only builds sub-indexes for each tenant. However, knowing that the collection contains multiple tenants unlocks more opportunities for optimization. @@ -178,6 +155,73 @@ Principal optimization is supported for following types: * `datetime` +## Full-text index + +Qdrant supports full-text search for string payload. +Full-text index allows you to filter points by the presence of a word or a phrase in the payload field. + +Full-text index configuration is a bit more complex than other indexes, as you can specify the tokenization parameters. +Tokenization is the process of splitting a string into tokens, which are then indexed in the inverted index. + +See [Full Text match](/documentation/concepts/filtering/#full-text-match) for examples of querying with a full-text index. + +To create a full-text index, you can use the following: + +{{< code-snippet path="/documentation/headless/snippets/create-payload-index/simple-full-text/" >}} + +### Tokenizers + +Tokenizers are algorithms used to split text into smaller units called tokens, which are then indexed and searched in a full-text index. +In the context of Qdrant, tokenizers determine how string payloads are broken down for efficient searching and filtering. +The choice of tokenizer affects how queries match the indexed text, supporting different languages, word boundaries, and search behaviours such as prefix or phrase matching. + +Available tokenizers are: + +* `word` - splits the string into words, separated by spaces, punctuation marks, and special characters. +* `whitespace` - splits the string into words, separated by spaces. +* `prefix` - splits the string into words, separated by spaces, punctuation marks, and special characters, and then creates a prefix index for each word. For example: `hello` will be indexed as `h`, `he`, `hel`, `hell`, `hello`. +* `multilingual` - a special type of tokenizer based on multiple packages like [charabia](https://github.com/meilisearch/charabia) and [vaporetto](https://github.com/daac-tools/vaporetto) to deliver fast and accurate tokenization for a large variety of languages. It allows proper tokenization and lemmatization for multiple languages, including those with non-Latin alphabets and non-space delimiters. See the [charabia documentation](https://github.com/meilisearch/charabia) for a full list of supported languages and normalization options. Note: For the Japanese language, Qdrant relies on the `vaporetto` project, which has much less overhead compared to `charabia`, while maintaining comparable performance. + +### Stemmer + +A **stemmer** is an algorithm used in text processing to reduce words to their root or base form, known as the "stem." For example, the words "running", "runner and "runs" can all be reduced to the stem "run." +When configuring a full-text index in Qdrant, you can specify a stemmer to be used for a particular language. This enables the index to recognize and match different inflections or derivations of a word. + +Qdrant provides an implementation of [Snowball stemmer](https://snowballstem.org/), a widely used and performant variant for some of the most popular languages. +For the list of supported languages, please visit the [rust-stemmers repository](https://github.com/qdrant/rust-stemmers). + +Here is an example of full-text Index configuration with Snowball stemmer: + +{{< code-snippet path="/documentation/headless/snippets/create-payload-index/stemmer-full-text/" >}} + +### Stopwords + +Stopwords are common words (such as "the", "is", "at", "which", and "on") that are often filtered out during text processing because they carry little meaningful information for search and retrieval tasks. + +In Qdrant, you can specify a list of stopwords to be ignored during full-text indexing and search. This helps simplify search queries and improves relevance. + +You can configure stopwords based on predefined languages, as well as extend existing stopword lists with custom words. + +Here is an example of configuring a full-text index with custom stopwords: + + +{{< code-snippet path="/documentation/headless/snippets/create-payload-index/stopwords-full-text/" >}} + +### Phrase Search + +Phrase search in Qdrant allows you to find documents or points where a specific sequence of words appears together, in the same order, within a text payload field. +This is useful when you want to match exact phrases rather than individual words scattered throughout the text. + +When using a full-text index with phrase search enabled, you can perform phrase search by enclosing the desired phrase in double quotes in your filter query. +For example, searching for `"machine learning"` will only return results where the words "machine" and "learning" appear together as a phrase, not just anywhere in the text. + +For efficient phrase search, Qdrant requires building an additional data structure, so it needs to be configured during the creation of the full-text index: + +{{< code-snippet path="/documentation/headless/snippets/create-payload-index/phrase-full-text/" >}} + +See [Phrase Match](/documentation/concepts/filtering/#phrase-match) for examples of querying phrases with a full-text index. + + ## Vector Index A vector index is a data structure built on vectors through a specific mathematical model. @@ -187,7 +231,7 @@ Qdrant currently only uses HNSW as a dense vector index. [HNSW](https://arxiv.org/abs/1603.09320) (Hierarchical Navigable Small World Graph) is a graph-based indexing algorithm. It builds a multi-layer navigation structure for an image according to certain rules. In this structure, the upper layers are more sparse and the distances between nodes are farther. The lower layers are denser and the distances between nodes are closer. The search starts from the uppermost layer, finds the node closest to the target in this layer, and then enters the next layer to begin another search. After multiple iterations, it can quickly approach the target position. -In order to improve performance, HNSW limits the maximum degree of nodes on each layer of the graph to `m`. In addition, you can use `ef_construct` (when building index) or `ef` (when searching targets) to specify a search range. +In order to improve performance, HNSW limits the maximum degree of nodes on each layer of the graph to `m`. In addition, you can use `ef_construct` (when building an index) or `ef` (when searching targets) to specify a search range. The corresponding parameters could be configured in the configuration file: @@ -240,10 +284,10 @@ To configure a sparse vector index, create a collection with the following param The following parameters may affect performance: -- `on_disk: true` - The index is stored on disk, which lets you save memory. This may slow down search performance. +- `on_disk: true` - The index is stored on disk, which lets you save memory. This may slow down search performance. - `on_disk: false` - The index is still persisted on disk, but it is also loaded into memory for faster search. -Unlike a dense vector index, a sparse vector index does not require a pre-defined vector size. It automatically adjusts to the size of the vectors added to the collection. +Unlike a dense vector index, a sparse vector index does not require a predefined vector size. It automatically adjusts to the size of the vectors added to the collection. **Note:** A sparse vector index only supports dot-product similarity searches. It does not support other distance metrics. @@ -252,7 +296,7 @@ Unlike a dense vector index, a sparse vector index does not require a pre-define *Available as of v1.10.0* For many search algorithms, it is important to consider how often an item occurs in a collection. -Intuitively speaking, the less frequently an item appears in a collection, the more important it is in a search. +Intuitively speaking, the less frequently an item appears in a collection, the more important it is in a search. This is also known as the Inverse Document Frequency (IDF). It is used in text search engines to rank search results based on the rarity of a word in a collection. diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/_description.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/_description.md new file mode 100644 index 000000000..001783154 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/_description.md @@ -0,0 +1 @@ +This code snippet demonstrates how to create a full-text index with phrase match support for a specified field in a collection. The index configuration includes details such as the field name, type (text), tokenizer (word), and whether to convert tokens to lowercase. This setup enables filtering points based on the presence of specific words or phrases in the field, allowing for efficient full-text search functionality within the payload. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/csharp.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/csharp.md new file mode 100644 index 000000000..1458d40fa --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/csharp.md @@ -0,0 +1,21 @@ +```csharp +using Qdrant.Client; +using Qdrant.Client.Grpc; + +var client = new QdrantClient("localhost", 6334); + +await client.CreatePayloadIndexAsync( + collectionName: "{collection_name}", + fieldName: "name_of_the_field_to_index", + schemaType: PayloadSchemaType.Text, + indexParams: new PayloadIndexParams + { + TextIndexParams = new TextIndexParams + { + Tokenizer = TokenizerType.Word, + Lowercase = true, + PhraseMatching = true + } + } +); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/go.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/go.md new file mode 100644 index 000000000..2d67a6952 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/go.md @@ -0,0 +1,24 @@ +```go +import ( + "context" + + "github.com/qdrant/go-client/qdrant" +) + +client, err := qdrant.NewClient(&qdrant.Config{ + Host: "localhost", + Port: 6334, +}) + +client.CreateFieldIndex(context.Background(), &qdrant.CreateFieldIndexCollection{ + CollectionName: "{collection_name}", + FieldName: "name_of_the_field_to_index", + FieldType: qdrant.FieldType_FieldTypeText.Enum(), + FieldIndexParams: qdrant.NewPayloadIndexParamsText( + &qdrant.TextIndexParams{ + Tokenizer: qdrant.TokenizerType_Whitespace, + Lowercase: qdrant.PtrOf(true), + PhraseMatching: qdrant.PtrOf(true), + }), +}) +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/http.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/http.md new file mode 100644 index 000000000..304dcb237 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/http.md @@ -0,0 +1,12 @@ +```http +PUT /collections/{collection_name}/index +{ + "field_name": "name_of_the_field_to_index", + "field_schema": { + "type": "text", + "tokenizer": "word", + "lowercase": true, + "phrase_matching": true + } +} +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/java.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/java.md new file mode 100644 index 000000000..3433a177d --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/java.md @@ -0,0 +1,29 @@ +```java +import io.qdrant.client.QdrantClient; +import io.qdrant.client.QdrantGrpcClient; +import io.qdrant.client.grpc.Collections.PayloadIndexParams; +import io.qdrant.client.grpc.Collections.PayloadSchemaType; +import io.qdrant.client.grpc.Collections.TextIndexParams; +import io.qdrant.client.grpc.Collections.TokenizerType; + +QdrantClient client = + new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build()); + +client + .createPayloadIndexAsync( + "{collection_name}", + "name_of_the_field_to_index", + PayloadSchemaType.Text, + PayloadIndexParams.newBuilder() + .setTextIndexParams( + TextIndexParams.newBuilder() + .setTokenizer(TokenizerType.Word) + .setLowercase(true) + .setPhraseMatching(true) + .build()) + .build(), + null, + null, + null) + .get(); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/python.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/python.md new file mode 100644 index 000000000..32b80c447 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/python.md @@ -0,0 +1,16 @@ +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient(url="http://localhost:6333") + +client.create_payload_index( + collection_name="{collection_name}", + field_name="name_of_the_field_to_index", + field_schema=models.TextIndexParams( + type="text", + tokenizer=models.TokenizerType.WORD, + lowercase=True, + phrase_matching=True, + ), +) +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/rust.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/rust.md new file mode 100644 index 000000000..5925eb5aa --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/rust.md @@ -0,0 +1,25 @@ +```rust +use qdrant_client::qdrant::{ + CreateFieldIndexCollectionBuilder, + TextIndexParamsBuilder, + FieldType, + TokenizerType, +}; +use qdrant_client::Qdrant; + +let client = Qdrant::from_url("http://localhost:6334").build()?; + +let text_index_params = TextIndexParamsBuilder::new(TokenizerType::Word) + .phrase_matching(true) + .lowercase(true); + +client + .create_field_index( + CreateFieldIndexCollectionBuilder::new( + "{collection_name}", + "name_of_the_field_to_index", + FieldType::Text, + ).field_index_params(text_index_params.build()), + ) + .await?; +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/typescript.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/typescript.md new file mode 100644 index 000000000..1d8be6a9a --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/phrase-full-text/typescript.md @@ -0,0 +1,15 @@ +```typescript +import { QdrantClient } from "@qdrant/js-client-rest"; + +const client = new QdrantClient({ host: "localhost", port: 6333 }); + +client.createPayloadIndex("{collection_name}", { + field_name: "name_of_the_field_to_index", + field_schema: { + type: "text", + tokenizer: "word", + lowercase: true, + phrase_matching: true, + }, +}); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/_description.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/_description.md index 9a7ba6df5..45a981e5d 100644 --- a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/_description.md +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/_description.md @@ -1 +1 @@ -This code snippet demonstrates how to create a full-text index for a specified field in a collection. The index configuration includes details such as the field name, type (text), tokenizer (word), minimum and maximum token length, and whether to convert tokens to lowercase. This setup enables filtering points based on the presence of specific words or phrases in the field, allowing for efficient full-text search functionality within the payload. \ No newline at end of file +This code snippet demonstrates how to create a full-text index for a specified field in a collection. The index configuration includes details such as the field name, type (text), tokenizer (word), minimum and maximum token length, and whether to convert tokens to lowercase. This setup enables filtering points based on the presence of specific words in the field, allowing for efficient full-text search functionality within the payload. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/http.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/http.md index 3af0ed3d9..bc1baf0ab 100644 --- a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/http.md +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/http.md @@ -6,7 +6,7 @@ PUT /collections/{collection_name}/index "type": "text", "tokenizer": "word", "min_token_len": 2, - "max_token_len": 20, + "max_token_len": 10, "lowercase": true } } diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/python.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/python.md index 360fe6a9d..08060af34 100644 --- a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/python.md +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/python.md @@ -10,7 +10,7 @@ client.create_payload_index( type="text", tokenizer=models.TokenizerType.WORD, min_token_len=2, - max_token_len=15, + max_token_len=10, lowercase=True, ), ) diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/rust.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/rust.md index a1a5244dd..bc48e706e 100644 --- a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/rust.md +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/rust.md @@ -1,27 +1,26 @@ ```rust use qdrant_client::qdrant::{ - payload_index_params::IndexParams, CreateFieldIndexCollectionBuilder, FieldType, - PayloadIndexParams, TextIndexParams, TokenizerType, + CreateFieldIndexCollectionBuilder, + TextIndexParamsBuilder, + FieldType, + TokenizerType, }; use qdrant_client::Qdrant; let client = Qdrant::from_url("http://localhost:6334").build()?; +let text_index_params = TextIndexParamsBuilder::new(TokenizerType::Word) + .min_token_len(2) + .max_token_len(10) + .lowercase(true); + client .create_field_index( CreateFieldIndexCollectionBuilder::new( "{collection_name}", "name_of_the_field_to_index", FieldType::Text, - ) - .field_index_params(PayloadIndexParams { - index_params: Some(IndexParams::TextIndexParams(TextIndexParams { - tokenizer: TokenizerType::Word as i32, - min_token_len: Some(2), - max_token_len: Some(10), - lowercase: Some(true), - })), - }), + ).field_index_params(text_index_params.build()), ) .await?; ``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/typescript.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/typescript.md index 87a69d43d..bee47999e 100644 --- a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/typescript.md +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/typescript.md @@ -9,7 +9,7 @@ client.createPayloadIndex("{collection_name}", { type: "text", tokenizer: "word", min_token_len: 2, - max_token_len: 15, + max_token_len: 10, lowercase: true, }, }); diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/_description.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/_description.md new file mode 100644 index 000000000..17c93287c --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/_description.md @@ -0,0 +1,2 @@ +This code snippet demonstrates how to create a full-text index for a specified field in a collection with stemmer configuration. +Stemmer configuration allows to apply stemming to input words, stemmer configuration contains 2 options - type of the stemmer (snowball), and desired language (english). \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/csharp.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/csharp.md new file mode 100644 index 000000000..4a8170e24 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/csharp.md @@ -0,0 +1,26 @@ +```csharp +using Qdrant.Client; +using Qdrant.Client.Grpc; + +var client = new QdrantClient("localhost", 6334); + +await client.CreatePayloadIndexAsync( + collectionName: "{collection_name}", + fieldName: "name_of_the_field_to_index", + schemaType: PayloadSchemaType.Text, + indexParams: new PayloadIndexParams + { + TextIndexParams = new TextIndexParams + { + Tokenizer = TokenizerType.Word, + Stemmer = new StemmingAlgorithm + { + Snowball = new SnowballParams + { + Language = "english" + } + } + } + } +); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/go.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/go.md new file mode 100644 index 000000000..b1fb62c4d --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/go.md @@ -0,0 +1,25 @@ +```go +import ( + "context" + + "github.com/qdrant/go-client/qdrant" +) + +client, err := qdrant.NewClient(&qdrant.Config{ + Host: "localhost", + Port: 6334, +}) + +client.CreateFieldIndex(context.Background(), &qdrant.CreateFieldIndexCollection{ + CollectionName: "{collection_name}", + FieldName: "name_of_the_field_to_index", + FieldType: qdrant.FieldType_FieldTypeText.Enum(), + FieldIndexParams: qdrant.NewPayloadIndexParamsText( + &qdrant.TextIndexParams{ + Tokenizer: qdrant.TokenizerType_Word, + Stemmer: qdrant.NewStemmingAlgorithmSnowball(&qdrant.SnowballParams{ + Language: "english", + }), + }), +}) +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/http.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/http.md new file mode 100644 index 000000000..acffb68b5 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/http.md @@ -0,0 +1,14 @@ +```http +PUT /collections/{collection_name}/index +{ + "field_name": "name_of_the_field_to_index", + "field_schema": { + "type": "text", + "tokenizer": "word", + "stemmer": { + "type": "snowball", + "language": "english" + } + } +} +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/java.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/java.md new file mode 100644 index 000000000..915b65e9a --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/java.md @@ -0,0 +1,34 @@ +```java +import io.qdrant.client.QdrantClient; +import io.qdrant.client.QdrantGrpcClient; +import io.qdrant.client.grpc.Collections.PayloadIndexParams; +import io.qdrant.client.grpc.Collections.PayloadSchemaType; +import io.qdrant.client.grpc.Collections.SnowballParams; +import io.qdrant.client.grpc.Collections.StemmingAlgorithm; +import io.qdrant.client.grpc.Collections.TextIndexParams; +import io.qdrant.client.grpc.Collections.TokenizerType; + +QdrantClient client = + new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build()); + +client + .createPayloadIndexAsync( + "{collection_name}", + "name_of_the_field_to_index", + PayloadSchemaType.Text, + PayloadIndexParams.newBuilder() + .setTextIndexParams( + TextIndexParams.newBuilder() + .setTokenizer(TokenizerType.Word) + .setStemmer( + StemmingAlgorithm.newBuilder() + .setSnowball( + SnowballParams.newBuilder().setLanguage("english").build()) + .build()) + .build()) + .build(), + true, + null, + null) + .get(); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/python.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/python.md new file mode 100644 index 000000000..868427854 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/python.md @@ -0,0 +1,18 @@ +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient(url="http://localhost:6333") + +client.create_payload_index( + collection_name="{collection_name}", + field_name="name_of_the_field_to_index", + field_schema=models.TextIndexParams( + type="text", + tokenizer=models.TokenizerType.WORD, + stemmer=models.SnowballParams( + type=models.Snowball.SNOWBALL, + language=models.SnowballLanguage.ENGLISH + ) + ), +) +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/rust.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/rust.md new file mode 100644 index 000000000..d0ebf1617 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/rust.md @@ -0,0 +1,24 @@ +```rust +use qdrant_client::qdrant::{ + CreateFieldIndexCollectionBuilder, + TextIndexParamsBuilder, + FieldType, + TokenizerType, +}; +use qdrant_client::Qdrant; + +let client = Qdrant::from_url("http://localhost:6334").build()?; + +let text_index_params = TextIndexParamsBuilder::new(TokenizerType::Word) + .snowball_stemmer("english".to_string()); + +client + .create_field_index( + CreateFieldIndexCollectionBuilder::new( + "{collection_name}", + "{field_name}", + FieldType::Text, + ).field_index_params(text_index_params.build()), + ) + .await?; +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/typescript.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/typescript.md new file mode 100644 index 000000000..be5ef9a67 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stemmer-full-text/typescript.md @@ -0,0 +1,17 @@ +```typescript +import { QdrantClient } from "@qdrant/js-client-rest"; + +const client = new QdrantClient({ host: "localhost", port: 6333 }); + +client.createPayloadIndex("{collection_name}", { + field_name: "name_of_the_field_to_index", + field_schema: { + type: "text", + tokenizer: "word", + stemmer: { + type: "snowball", + language: "english" + } + } +}); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/_description.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/_description.md new file mode 100644 index 000000000..ae3287330 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/_description.md @@ -0,0 +1 @@ +This code snippet demonstrates how to create a full-text index for a specified field in a collection with stopwords configuration. There are 2 examples of Stopwords configuration, simple and explicit. Simple configuration specifies only a signle language, only pre-defined stopwords from this language will be used. Explicit configurations configures combination of stopwords from 2 languages and one custom stopword. \ No newline at end of file diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/csharp.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/csharp.md new file mode 100644 index 000000000..f03e98aee --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/csharp.md @@ -0,0 +1,24 @@ +```csharp +using Qdrant.Client; +using Qdrant.Client.Grpc; + +var client = new QdrantClient("localhost", 6334); + +await client.CreatePayloadIndexAsync( + collectionName: "{collection_name}", + fieldName: "name_of_the_field_to_index", + schemaType: PayloadSchemaType.Text, + indexParams: new PayloadIndexParams + { + TextIndexParams = new TextIndexParams + { + Tokenizer = TokenizerType.Word, + Stopwords = new StopwordsSet + { + Languages = { "english", "spanish" }, + Custom = { "example" } + } + } + } +); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/go.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/go.md new file mode 100644 index 000000000..30c9eac9a --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/go.md @@ -0,0 +1,26 @@ +```go +import ( + "context" + + "github.com/qdrant/go-client/qdrant" +) + +client, err := qdrant.NewClient(&qdrant.Config{ + Host: "localhost", + Port: 6334, +}) + +client.CreateFieldIndex(context.Background(), &qdrant.CreateFieldIndexCollection{ + CollectionName: "{collection_name}", + FieldName: "name_of_the_field_to_index", + FieldType: qdrant.FieldType_FieldTypeText.Enum(), + FieldIndexParams: qdrant.NewPayloadIndexParamsText( + &qdrant.TextIndexParams{ + Tokenizer: qdrant.TokenizerType_Word, + Stopwords: &qdrant.StopwordsSet{ + Languages: []string{"english", "spanish"}, + Custom: []string{"example"}, + }, + }), +}) +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/http.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/http.md new file mode 100644 index 000000000..9b4af648c --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/http.md @@ -0,0 +1,31 @@ +```http +// Simple +PUT collections/{collection_name}/index +{ + "field_name": "name_of_the_field_to_index", + "field_schema": { + "type": "text", + "tokenizer": "word", + "stopwords": "english" + } +} + +// Explicit +PUT collections/{collection_name}/index +{ + "field_name": "name_of_the_field_to_index", + "field_schema": { + "type": "text", + "tokenizer": "word", + "stopwords": { + "languages": [ + "english", + "spanish" + ], + "custom": [ + "example" + ] + } + } +} +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/java.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/java.md new file mode 100644 index 000000000..26a8f46f8 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/java.md @@ -0,0 +1,35 @@ +```java +import java.util.List; + +import io.qdrant.client.QdrantClient; +import io.qdrant.client.QdrantGrpcClient; +import io.qdrant.client.grpc.Collections.PayloadIndexParams; +import io.qdrant.client.grpc.Collections.PayloadSchemaType; +import io.qdrant.client.grpc.Collections.StopwordsSet; +import io.qdrant.client.grpc.Collections.TextIndexParams; +import io.qdrant.client.grpc.Collections.TokenizerType; + +QdrantClient client = + new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build()); + +client + .createPayloadIndexAsync( + "{collection_name}", + "name_of_the_field_to_index", + PayloadSchemaType.Text, + PayloadIndexParams.newBuilder() + .setTextIndexParams( + TextIndexParams.newBuilder() + .setTokenizer(TokenizerType.Word) + .setStopwords( + StopwordsSet.newBuilder() + .addAllLanguages(List.of("english", "spanish")) + .addAllCustom(List.of("example")) + .build()) + .build()) + .build(), + true, + null, + null) + .get(); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/python.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/python.md new file mode 100644 index 000000000..44de65fac --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/python.md @@ -0,0 +1,35 @@ +```python +from qdrant_client import QdrantClient, models + +client = QdrantClient(url="http://localhost:6333") + +# Simple +client.create_payload_index( + collection_name="{collection_name}", + field_name="name_of_the_field_to_index", + field_schema=models.TextIndexParams( + type="text", + tokenizer=models.TokenizerType.WORD, + stopwords=models.Language.ENGLISH, + ), +) + +# Explicit +client.create_payload_index( + collection_name="{collection_name}", + field_name="name_of_the_field_to_index", + field_schema=models.TextIndexParams( + type="text", + tokenizer=models.TokenizerType.WORD, + stopwords=models.StopwordsSet( + languages=[ + models.Language.ENGLISH, + models.Language.SPANISH, + ], + custom=[ + "example" + ] + ), + ), +) +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/rust.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/rust.md new file mode 100644 index 000000000..48d84a984 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/rust.md @@ -0,0 +1,46 @@ +```rust +use qdrant_client::qdrant::{ + CreateFieldIndexCollectionBuilder, + TextIndexParamsBuilder, + FieldType, + TokenizerType, + StopwordsSet, +}; +use qdrant_client::Qdrant; + +let client = Qdrant::from_url("http://localhost:6334").build()?; + +// Simple +let text_index_params = TextIndexParamsBuilder::new(TokenizerType::Word) + .stopwords_language("english".to_string()); + +client + .create_field_index( + CreateFieldIndexCollectionBuilder::new( + "{collection_name}", + "name_of_the_field_to_index", + FieldType::Text, + ).field_index_params(text_index_params.build()), + ) + .await?; + +// Explicit +let text_index_params = TextIndexParamsBuilder::new(TokenizerType::Word) + .stopwords(StopwordsSet { + languages: vec![ + "english".to_string(), + "spanish".to_string(), + ], + custom: vec!["example".to_string()], + }); + +client + .create_field_index( + CreateFieldIndexCollectionBuilder::new( + "{collection_name}", + "{field_name}", + FieldType::Text, + ).field_index_params(text_index_params.build()), + ) + .await?; +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/typescript.md b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/typescript.md new file mode 100644 index 000000000..7913cfda4 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/create-payload-index/stopwords-full-text/typescript.md @@ -0,0 +1,34 @@ +```typescript +import { QdrantClient } from "@qdrant/js-client-rest"; + +const client = new QdrantClient({ host: "localhost", port: 6333 }); + + +// Simple +client.createPayloadIndex("{collection_name}", { + field_name: "name_of_the_field_to_index", + field_schema: { + type: "text", + tokenizer: "word", + stopwords: "english" + }, +}); + +// Explicit +client.createPayloadIndex("{collection_name}", { + field_name: "name_of_the_field_to_index", + field_schema: { + type: "text", + tokenizer: "word", + stopwords: { + languages: [ + "english", + "spanish" + ], + custom: [ + "example" + ] + } + }, +}); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/_description.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/_description.md new file mode 100644 index 000000000..6274ea632 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/_description.md @@ -0,0 +1 @@ +This code snippet sets up a field condition to search for an exact (token) phrase within a text field. In this case, a special `phrase` match is defined with the target text being "brown fox". The behavior may vary depending on the configuration of the full-text index for the field. If there is no full-text index configured for the field, the condition will work as an exact substring match. diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/csharp.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/csharp.md new file mode 100644 index 000000000..1e5fb408a --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/csharp.md @@ -0,0 +1,5 @@ +```csharp +using static Qdrant.Client.Grpc.Conditions; + +MatchPhrase("description", "brown fox"); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/go.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/go.md new file mode 100644 index 000000000..bcefa854b --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/go.md @@ -0,0 +1,5 @@ +```go +import "github.com/qdrant/go-client/qdrant" + +qdrant.NewMatchPhrase("description", "brown fox") +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/java.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/java.md new file mode 100644 index 000000000..8fc659850 --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/java.md @@ -0,0 +1,5 @@ +```java +import static io.qdrant.client.ConditionFactory.matchPhrase; + +matchPhrase("description", "brown fox"); +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/json.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/json.md new file mode 100644 index 000000000..6e5c9cd5a --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/json.md @@ -0,0 +1,8 @@ +```json +{ + "key": "description", + "match": { + "phrase": "brown fox" + } +} +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/python.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/python.md new file mode 100644 index 000000000..d7387507b --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/python.md @@ -0,0 +1,6 @@ +```python +models.FieldCondition( + key="description", + match=models.MatchPhrase(phrase="brown fox"), +) +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/rust.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/rust.md new file mode 100644 index 000000000..76dfb0eee --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/rust.md @@ -0,0 +1,5 @@ +```rust +use qdrant_client::qdrant::Condition; + +Condition::matches_phrase("description", "brown fox") +``` diff --git a/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/typescript.md b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/typescript.md new file mode 100644 index 000000000..6f441984f --- /dev/null +++ b/qdrant-landing/content/documentation/headless/snippets/filter-condition/phrase-match/typescript.md @@ -0,0 +1,6 @@ +```typescript +{ + key: 'description', + match: {phrase: 'brown fox'} +} +```