Add ASCII Folding docs for 1.16

This commit is contained in:
Abdon Pijpelink
2025-11-11 17:14:15 +01:00
parent 388c8fd929
commit 40f808568d
7 changed files with 72 additions and 0 deletions
@@ -182,6 +182,26 @@ Available tokenizers are:
* `prefix` - splits the string into words, separated by spaces, punctuation marks, and special characters, and then creates a prefix index for each word. For example: `hello` will be indexed as `h`, `he`, `hel`, `hell`, `hello`.
* `multilingual` - a special type of tokenizer based on multiple packages like [charabia](https://github.com/meilisearch/charabia) and [vaporetto](https://github.com/daac-tools/vaporetto) to deliver fast and accurate tokenization for a large variety of languages. It allows proper tokenization and lemmatization for multiple languages, including those with non-Latin alphabets and non-space delimiters. See the [charabia documentation](https://github.com/meilisearch/charabia) for a full list of supported languages and normalization options. Note: For the Japanese language, Qdrant relies on the `vaporetto` project, which has much less overhead compared to `charabia`, while maintaining comparable performance.
### Lowercasing
By default, full-text search in Qdrant is case-insensitive. For example, users can search for the lowercase term `tv` and find text fields containing the uppercase word `TV`. Case-insensitivity is achieved by converting both the words in the index and the query terms to lowercase.
Lowercasing is enabled by default. To enable case-sensitive full-text search, configure a full-text index with `lowercase` set to `false`:
{{< code-snippet path="/documentation/headless/snippets/create-payload-index/lowercase-full-text/" >}}
### ASCII Folding
*Available as of v1.16.0*
When enabled, ASCII folding converts Unicode characters into their corresponding ASCII equivalents, for example, by removing diacritics. For instance, the character `ã` is changed into `a`, `ç` becomes `c`, and `é` is converted to `e`.
Because ASCII folding is applied to both the words in the index and the query terms, it increases recall. For example, users can search for `cafe` and also find text fields containing the word `café`.
ASCII folding is not enabled by default. To enable it, configure a full-text index with `ascii_folding` set to `true`:
{{< code-snippet path="/documentation/headless/snippets/create-payload-index/asciifolding-full-text/" >}}
### Stemmer
A **stemmer** is an algorithm used in text processing to reduce words to their root or base form, known as the "stem." For example, the words "running", "runner and "runs" can all be reduced to the stem "run."
@@ -0,0 +1 @@
This code snippet demonstrates how to create a full-text index with ASCII folding support for a specified field in a collection. The index configuration includes details such as the field name, type (text), tokenizer (word), and whether to ASCII-fold tokens. This setup enables filtering points based on the presence of specific words while ignoring diacritics in the field, allowing for efficient full-text search functionality within the payload.
@@ -0,0 +1,11 @@
```http
PUT /collections/{collection_name}/index
{
"field_name": "name_of_the_field_to_index",
"field_schema": {
"type": "text",
"tokenizer": "word",
"ascii_folding": true
}
}
```
@@ -0,0 +1,14 @@
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({ host: "localhost", port: 6333 });
client.createPayloadIndex("{collection_name}", {
field_name: "name_of_the_field_to_index",
field_schema: {
type: "text",
tokenizer: "word",
ascii_folding: true,
},
});
```
@@ -0,0 +1 @@
This code snippet demonstrates how to configure a full-text index for case-sensitive match support for a specified field in a collection. The index configuration includes details such as the field name, type (text), tokenizer (word), and whether to convert tokens to lowercase. Lowercasing is enabled by default, enabling case-sensitive matching. This setup disables lowercasing, enabling the filtering og points based on the presence of exact words in the field.
@@ -0,0 +1,11 @@
```http
PUT /collections/{collection_name}/index
{
"field_name": "name_of_the_field_to_index",
"field_schema": {
"type": "text",
"tokenizer": "word",
"lowercase": false
}
}
```
@@ -0,0 +1,14 @@
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({ host: "localhost", port: 6333 });
client.createPayloadIndex("{collection_name}", {
field_name: "name_of_the_field_to_index",
field_schema: {
type: "text",
tokenizer: "word",
lowercase: false,
},
});
```