update blog with benchmark results, new examples, title, date

This commit is contained in:
Thierry Damiba
2026-03-30 20:11:03 -07:00
parent 9fb095c6f0
commit ee543a7002
@@ -1,12 +1,12 @@
---
title: "[Working Title]"
title: "Qdrant Skills for AI Agents"
draft: false
slug: qdrant-skills-release
short_description: "Agent skills for Qdrant: solutions architect knowledge encoded for AI agents. Diagnose, tune, and scale vector search without guessing."
description: "Agent skills for Qdrant: solutions architect knowledge encoded for AI agents. Diagnose, tune, and scale vector search without guessing."
preview_image: /blog/qdrant-skills-release/hero.png
social_preview_image: /blog/qdrant-skills-release/hero.png
date: 2026-03-26
date: 2026-03-31
author: Thierry Damiba
featured: true
tags:
@@ -36,7 +36,7 @@ Each one of these is a lever. Pull one and it changes the others. Enable binary
These primitives compound. The right combination depends on your data distribution, query patterns, and production constraints.
[Cosmos](https://qdrant.tech/blog/case-study-cosmos/), a visual search platform for creative professionals, is a canonical example of composable vector search in production. By combining multiple Qdrant features, they created a refined, intuitive search experience that understands the users taste. In a single collection using named vectors, they store CLIP vectors (text to image search), CNNs (for style), pHash (to identify duplicates), and color embeddings (for palette search). Color searches leverage five different vectors held in memory for fast distance calculations; text search relies on CLIP vectors. Hybrid queries blend the two.
[Cosmos](https://qdrant.tech/blog/case-study-cosmos/), a visual search platform for creative professionals, is a canonical example of composable vector search in production. By combining multiple Qdrant features, they created a refined, intuitive search experience that understands the user's taste. In a single collection using named vectors, they store CLIP vectors (text to image search), CNNs (for style), pHash (to identify duplicates), and color embeddings (for palette search). Color searches leverage five different vectors held in memory for fast distance calculations; text search relies on CLIP vectors. Hybrid queries blend the two.
They started with built-in reciprocal rank fusion, then moved to application-side fusion because they needed custom scoring that balanced relevance, engagement, and aesthetics. Every one of those choices was a tradeoff: memory for color precision, latency for ranking quality, simplicity for control.
@@ -137,63 +137,69 @@ The result is a hierarchical navigation system over our documentation, organized
## Real examples: same agent, same question, with and without skills
We ran each question 5 times on Claude Opus 4.6 (1M context), with and without Qdrant skills loaded. Same model, same system prompt, no cherry-picking. With skills: 150/150 assertions passed (100%). Without skills: 78/150 (52%). The full prompts, outputs, and grading are [reproducible in the repo](https://github.com/qdrant/skills/tree/main/skills-eval-workspace).
We ran each question 5 times on Claude Opus 4.6 (1M context), with and without Qdrant skills loaded. Same model, same system prompt, no cherry-picking.
**With skills: 144/150 assertions passed (96%). Without skills: 97/150 (65%).**
| Metric | With Skill | Without Skill | Delta |
|--------|-----------|---------------|-------|
| **Pass Rate** | **96% ± 8%** | **65% ± 24%** | **+31%** |
| **Response Time** | **24.8s ± 6.1s** | **45.5s ± 23.8s** | **1.8x faster** |
| **Output Size** | **~759 tokens** | **~875 tokens** | **13% fewer** |
Skills produce more accurate, faster, and more focused responses. The full prompts, outputs, and grading are [reproducible in the eval repo](https://github.com/qdrant-labs/skills-eval).
---
### "We index social media posts and only care about the last 3 months. Our cron job deletes old posts but performance keeps degrading."
### "We have a social media search index with a 3-month retention window. We run a nightly cron that deletes old posts but after 6 months in production our searches went from 15ms to 200ms and disk usage keeps climbing even though the dataset size is stable."
*Baseline pass rate: 28% across 5 runs. With skill: 100%. The largest delta in our evaluation.*
*Baseline pass rate: 32% across 5 runs. With skill: 100%. The largest delta in our evaluation.*
**Without skills:**
> The core issue: Qdrant uses soft deletes internally. When you delete points via a cron job, the data is marked as deleted but not immediately reclaimed. The HNSW index still contains the deleted nodes, queries still traverse them, and segment compaction has to work harder. Over time, you accumulate tombstones and both search latency and memory usage climb.
> Qdrant uses soft deletes. When you call `client.delete(filter=...)`, points are marked with tombstones but not physically removed from the HNSW graph. The optimizer merges segments and cleans up tombstones, but with a nightly bulk delete pattern, you're creating a steady buildup of dead points that the default optimizer config can't keep up with.
>
> Instead of deleting individual points, rotate entire collections on a schedule. Create monthly collections (`social_posts_2026_01`, `social_posts_2026_02`, ...), point an alias at the active window, search across all active collections in parallel, and merge results client-side. Each month, create the new collection and drop the oldest one.
> Switch from filter-delete to collection aliasing with rolling collections. Instead of deleting from one collection, use time-bucketed collections behind an alias. Create monthly collections, point the alias at the active set, and drop the expired month entirely. Search across multiple collections by querying each and merging results client-side.
**With skills:**
> Your performance degradation is expected. Filter-and-delete leaves tombstoned points that degrade search until the optimizer compacts segments. Disk is not reclaimed instantly either. For a 3-month rolling window, this is the wrong strategy.
> You're hitting the known degradation pattern of filter-and-delete at scale. `client.delete(filter=...)` doesn't physically remove data. It marks points as tombstoned. These tombstoned points still exist on disk and still get visited during search, which is why latency climbed from 15ms to 200ms. The optimizer creates new compacted segments while your next delete batch creates new tombstones.
>
> Switch to user-defined sharding with one shard key per month. Create a collection with `sharding_method: "custom"`, create shard keys `2026-01`, `2026-02`, `2026-03`. Ingest into the current month's shard key. When a new month starts, create the next shard key, redirect writes, and delete the oldest shard key. Deleting a shard key reclaims all resources instantly. No tombstones, no compaction, no degradation. Pre-create the next month's shard key before rotation to avoid write failures during the transition.
> Switch to shard rotation with user-defined sharding. Create a collection with `sharding_method: "custom"`, create shard keys per month (`2026-01`, `2026-02`, `2026-03`). Ingest into the current month's shard key. On rotation, pre-create the next month's shard key, then drop the oldest. Dropping a shard key is instant: no tombstones, no fragmentation, no optimizer overhead, disk reclaimed immediately. Queries without `shard_key_selector` automatically search all shards, so no application changes needed for reads.
**The difference:** Both agents correctly diagnose tombstone degradation. But without the skill, Claude defaults to collection rotation with alias swapping, which requires application-level fan-out search across multiple collections and client-side result merging. The skill teaches shard rotation: a single collection with user-defined shard keys, where queries automatically span all time periods without fan-out and dropping a shard key reclaims resources instantly. The unskilled agent doesn't know this pattern exists.
---
### "We're running Qdrant 1.14 and want to upgrade to 1.17. Our cluster has replication factor 1 and 2 nodes."
### "We're maxing out at 150 QPS but need 500. Single node, 32GB RAM, 8 cores, 20M vectors at 384 dims. We tried increasing hnsw_ef and segment count to 16 to make it faster but QPS actually got worse."
*Baseline pass rate: 48% across 5 runs. With skill: 100%.*
*Baseline pass rate: 60% across 5 runs. With skill: 100%.*
**Without skills:**
> With replication factor 1, you have zero redundancy. Each shard exists on exactly one node. Increase replication factor to 2 before upgrading. Upgrade one minor version at a time: 1.14 to 1.15, then 1.15 to 1.16, then 1.16 to 1.17. Do not skip versions. Take snapshots before each hop. Perform rolling upgrades: stop one node, upgrade it, wait for it to rejoin and sync, then do the other.
> The problem is too many segments. 16 segments means every query touches all 16 in parallel, starving other queries of CPU. Set `default_segment_number` to 2. Also lower `hnsw_ef` back to 64, it controls accuracy not throughput. Enable scalar quantization to compress vectors from ~29GB to ~7.5GB so everything fits in RAM.
**With skills:**
> Two critical issues. First: your replication factor of 1 means rolling upgrades cannot guarantee zero downtime. Each shard only exists on one node, so when that node goes down for upgrade, those shards are unavailable. Increase to RF=2 before starting.
>
> Second: Qdrant only guarantees storage compatibility between consecutive minor versions. You must upgrade step by step: 1.14 to 1.15, 1.15 to 1.16, 1.16 to 1.17. For each hop, upgrade the SDK to the target minor version first, then upgrade the server. The server is backward-compatible with the next SDK version for available features. Qdrant Cloud automates this process and can jump directly from 1.14 to 1.17, but self-hosted clusters must do each step manually. Use the [qcloud CLI](https://github.com/qdrant/qcloud-cli) for Cloud upgrades.
> You're tuning in the wrong direction. Both changes hurt throughput: increasing `hnsw_ef` makes each query do more work (it's a latency/accuracy knob, not throughput), and 16 segments means more overhead per query. For throughput, you want fewer, larger segments (`default_segment_number: 2`). Enable scalar quantization with `always_ram: true` to cut vector memory by 4x. Use the batch search API to amortize per-request overhead. Set `optimizer_cpu_budget: 2` so indexing only uses 2 of your 8 cores, reserving 6 for queries. If still not enough after all tuning, add read replicas for horizontal scaling.
**The difference:** The unskilled agent gets the big picture right (step-by-step, increase RF). But it misses three things the skill provides: the explicit storage compatibility guarantee (why you can't skip), the SDK-before-server upgrade ordering, and that Qdrant Cloud can automate multi-version jumps. These are the details that prevent a failed migration at 2am.
**The difference:** The unskilled agent catches the segment count issue and quantization but misses two things the skill provides: the batch search API for amortizing overhead, and limiting optimizer CPU budget to reserve cores for queries. It also doesn't explain the latency vs throughput distinction clearly enough to prevent the user from making the same mistake again.
---
### "We have 5,000 tenants with 10k-100k vectors each. We've been creating one collection per tenant but hitting limits."
### "We have a multi-tenant SaaS with 5,000 tenants, each with 10k-100k vectors. We started with one collection per tenant but at around 800 collections Qdrant started using 90%+ memory just for overhead. We tried grouping small tenants into shared collections but search got really slow because the HNSW graph is global."
*Baseline pass rate: 48% across 5 runs. With skill: 100%.*
*Baseline pass rate: 68% across 5 runs. With skill: 100%.*
**Without skills:**
> Creating 5,000 separate collections is not sustainable. Each collection carries its own HNSW index, WAL, optimizer threads, and segment management. The correct approach is to consolidate into a single shared collection with payload-based partitioning. Add a `tenant_id` field to every point, create a keyword payload index on it, and filter by `tenant_id` on every query. This drops resource usage dramatically and scales to 100,000+ tenants.
> Your two attempts hit the two classic anti-patterns: collection-per-tenant has unsustainable overhead, and shared collection without tenant-aware indexing means HNSW builds one global graph where most traversed nodes get rejected by the filter. Put all tenants in one collection, add a `tenant_id` payload field, create a keyword index on it, and filter every query by `tenant_id`. Qdrant will use the payload index to restrict HNSW traversal to only that tenant's vectors.
**With skills:**
> One collection per tenant does not scale past a few hundred. One company hit the 1,000 collection limit after a year of collection-per-repo and had to migrate to payload partitioning. Use a shared collection with a `tenant_id` field and create a keyword index with `is_tenant: true`. This is critical: without `is_tenant: true`, Qdrant builds a single global HNSW graph and filters after traversal, which kills performance. With it, Qdrant builds per-tenant HNSW segments so each tenant's search stays fast. Do not build a global HNSW for multi-tenant collections; use `payload_m` instead.
>
> Your tenants range 10x in size (10k to 100k vectors). If some grow significantly larger, use tiered multitenancy to promote large tenants to dedicated shards while keeping small tenants on shared shards. Custom sharding by tenant ID hash becomes relevant at 100k+ tenants, not at your current 5,000.
> Your shared collection approach failed because you're building a global HNSW graph. Create one collection with a `tenant_id` payload field and a keyword index with `is_tenant: true`. This flag is critical: without it, Qdrant builds one global HNSW graph and scans sequentially per tenant. With it, Qdrant builds per-tenant HNSW segments so each tenant search only touches that tenant's vectors. Use `payload_m` instead of the default `m` parameter and set `m: 0` to disable the global HNSW entirely since you'll never search without a tenant filter. For your 10x size range across tenants, use tiered multitenancy to promote large tenants to dedicated shards while keeping small tenants on shared shards.
**The difference:** Both agents give the correct high-level answer: consolidate into one collection with payload filtering. But the unskilled agent misses the specific configuration that makes it work well: `is_tenant: true` for per-tenant HNSW segments, `payload_m` over global HNSW, and tiered multitenancy for uneven tenant sizes. It also lacks the real-world incident (the 1,000 collection limit) that makes the recommendation concrete. The skilled agent gives you the architecture; the unskilled agent gives you the idea.
**The difference:** Both agents give the correct high-level answer: consolidate into one collection with payload filtering. But the unskilled agent misses the specific configuration that makes it work: `is_tenant: true` for per-tenant HNSW segments, `payload_m` over global HNSW, and tiered multitenancy for uneven tenant sizes. The skilled agent gives you the architecture; the unskilled agent gives you the idea.
## Try it