From b13106e7f021f7ea17017c90c7e20261eeef3c76 Mon Sep 17 00:00:00 2001 From: Bastian Hofmann Date: Wed, 25 Jun 2025 16:59:27 +0200 Subject: [PATCH] Add notes about context window and vector size --- qdrant-landing/content/documentation/cloud/inference.md | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/qdrant-landing/content/documentation/cloud/inference.md b/qdrant-landing/content/documentation/cloud/inference.md index fb2fb651e..f237c0180 100644 --- a/qdrant-landing/content/documentation/cloud/inference.md +++ b/qdrant-landing/content/documentation/cloud/inference.md @@ -338,4 +338,8 @@ func main() { }) ``` -Usage examples, specific to each cluster and model can also be found in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. \ No newline at end of file +Usage examples, specific to each cluster and model can also be found in the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. + +Note that, each model has a context window, which is the maximum number of tokens that can be processed by the model in a single request. If the input text exceeds the context window, it will be truncated to fit within the limit. The context window size is displayed in the Inference tab of the Cluster Detail page. + +For dense vector models, you also have to ensure that the vector size configured in the collection matches the output size of the model. If the vector size does not match, the upsert will fail with an error.