[v1.15.0] Text index documentation (#1698)

* phrase matching documentation

* mention incompatibility with prefix tokenizer

* nvm, it works well

* fix rust snippet

* fix python snippet

* Update qdrant-landing/content/documentation/headless/snippets/create-payload-index/simple-full-text/csharp.md

* [1.15] Text Index Documentation (#1796)

* wip

* add snippets for simple text index creation and phrase index separatelly

* add some stopwords snippets + stopwords description

* add stemmer snippets

* docs: Stemmer, stop words Go snippets

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* docs: Stemmer, stop words Java snippets

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* docs: Stemmer, stop words C# snippets

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* fix: http, TS stemmer snippets

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* docs: Nit fixes indexing.md

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* lowercase stemmer is intended for consistency

---------

Signed-off-by: Anush008 <anushshetty90@gmail.com>
Co-authored-by: Anush008 <anushshetty90@gmail.com>

---------

Signed-off-by: Anush008 <anushshetty90@gmail.com>
Co-authored-by: Anush <anushshetty90@gmail.com>
Co-authored-by: Andrey Vasnetsov <andrey@vasnetsov.com>
This commit is contained in:
Luis Cossío
2025-07-18 13:24:00 +02:00
committed by GitHub
co-authored by Anush008 Andrey Vasnetsov
parent fe7a45a4ad
commit ef31a4afe0
39 changed files with 697 additions and 61 deletions
@@ -0,0 +1,2 @@
This code snippet demonstrates how to create a full-text index for a specified field in a collection with stemmer configuration.
Stemmer configuration allows to apply stemming to input words, stemmer configuration contains 2 options - type of the stemmer (snowball), and desired language (english).
@@ -0,0 +1,26 @@
```csharp
using Qdrant.Client;
using Qdrant.Client.Grpc;
var client = new QdrantClient("localhost", 6334);
await client.CreatePayloadIndexAsync(
collectionName: "{collection_name}",
fieldName: "name_of_the_field_to_index",
schemaType: PayloadSchemaType.Text,
indexParams: new PayloadIndexParams
{
TextIndexParams = new TextIndexParams
{
Tokenizer = TokenizerType.Word,
Stemmer = new StemmingAlgorithm
{
Snowball = new SnowballParams
{
Language = "english"
}
}
}
}
);
```
@@ -0,0 +1,25 @@
```go
import (
"context"
"github.com/qdrant/go-client/qdrant"
)
client, err := qdrant.NewClient(&qdrant.Config{
Host: "localhost",
Port: 6334,
})
client.CreateFieldIndex(context.Background(), &qdrant.CreateFieldIndexCollection{
CollectionName: "{collection_name}",
FieldName: "name_of_the_field_to_index",
FieldType: qdrant.FieldType_FieldTypeText.Enum(),
FieldIndexParams: qdrant.NewPayloadIndexParamsText(
&qdrant.TextIndexParams{
Tokenizer: qdrant.TokenizerType_Word,
Stemmer: qdrant.NewStemmingAlgorithmSnowball(&qdrant.SnowballParams{
Language: "english",
}),
}),
})
```
@@ -0,0 +1,14 @@
```http
PUT /collections/{collection_name}/index
{
"field_name": "name_of_the_field_to_index",
"field_schema": {
"type": "text",
"tokenizer": "word",
"stemmer": {
"type": "snowball",
"language": "english"
}
}
}
```
@@ -0,0 +1,34 @@
```java
import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
import io.qdrant.client.grpc.Collections.PayloadIndexParams;
import io.qdrant.client.grpc.Collections.PayloadSchemaType;
import io.qdrant.client.grpc.Collections.SnowballParams;
import io.qdrant.client.grpc.Collections.StemmingAlgorithm;
import io.qdrant.client.grpc.Collections.TextIndexParams;
import io.qdrant.client.grpc.Collections.TokenizerType;
QdrantClient client =
new QdrantClient(QdrantGrpcClient.newBuilder("localhost", 6334, false).build());
client
.createPayloadIndexAsync(
"{collection_name}",
"name_of_the_field_to_index",
PayloadSchemaType.Text,
PayloadIndexParams.newBuilder()
.setTextIndexParams(
TextIndexParams.newBuilder()
.setTokenizer(TokenizerType.Word)
.setStemmer(
StemmingAlgorithm.newBuilder()
.setSnowball(
SnowballParams.newBuilder().setLanguage("english").build())
.build())
.build())
.build(),
true,
null,
null)
.get();
```
@@ -0,0 +1,18 @@
```python
from qdrant_client import QdrantClient, models
client = QdrantClient(url="http://localhost:6333")
client.create_payload_index(
collection_name="{collection_name}",
field_name="name_of_the_field_to_index",
field_schema=models.TextIndexParams(
type="text",
tokenizer=models.TokenizerType.WORD,
stemmer=models.SnowballParams(
type=models.Snowball.SNOWBALL,
language=models.SnowballLanguage.ENGLISH
)
),
)
```
@@ -0,0 +1,24 @@
```rust
use qdrant_client::qdrant::{
CreateFieldIndexCollectionBuilder,
TextIndexParamsBuilder,
FieldType,
TokenizerType,
};
use qdrant_client::Qdrant;
let client = Qdrant::from_url("http://localhost:6334").build()?;
let text_index_params = TextIndexParamsBuilder::new(TokenizerType::Word)
.snowball_stemmer("english".to_string());
client
.create_field_index(
CreateFieldIndexCollectionBuilder::new(
"{collection_name}",
"{field_name}",
FieldType::Text,
).field_index_params(text_index_params.build()),
)
.await?;
```
@@ -0,0 +1,17 @@
```typescript
import { QdrantClient } from "@qdrant/js-client-rest";
const client = new QdrantClient({ host: "localhost", port: 6333 });
client.createPayloadIndex("{collection_name}", {
field_name: "name_of_the_field_to_index",
field_schema: {
type: "text",
tokenizer: "word",
stemmer: {
type: "snowball",
language: "english"
}
}
});
```