Merge remote-tracking branch 'upstream/master' into docs/go-ecommerce

This commit is contained in:
Nathan LeRoy
2026-01-30 11:44:10 -05:00
30 changed files with 624 additions and 32 deletions
@@ -0,0 +1,183 @@
---
title: "Two Approaches to Helping AI Agents Use Your API (And Why You Need Both)"
short_description: "Mintlify's skill.md and Armin Ronacher's REPL-first MCP solve different failure modes. Together, they define how agents should interact with developer tools."
description: "Two emerging patterns for agent-assisted development: static knowledge files and dynamic tool access. How they complement each other using Qdrant as a case study."
preview_dir: /articles_data/skill-md-meets-repl/preview
social_preview_image: /articles_data/skill-md-meets-repl/preview/social_preview.jpg
author: Thierry Damiba
draft: false
date: 2026-01-28T00:00:00-08:00
weight: -200
category: rag-and-genai
---
AI coding agents fail in predictable ways when working with APIs. Two recent approaches from Mintlify and Armin Ronacher attack different failure modes. Understanding both reveals something useful about how agents should interact with developer tools.
## Two Failure Modes
When an agent writes code against your API, it can fail because:
1. **It doesn't know what it doesn't know.** The agent uses a deprecated method, misconfigures a parameter, or violates a constraint that isn't obvious from type signatures. This is the "known unknowns" problem: things the API maintainer knows but the agent doesn't.
2. **It can't discover what exists.** The agent doesn't know what collections exist, what the payload schema looks like, or what data is actually in the system. This is the "unknown unknowns" problem: things specific to the user's environment that no amount of documentation covers.
Most agent failures trace back to one of these. Mintlify's skill.md addresses the first. Armin Ronacher's REPL-first MCP addresses the second.
## What skill.md Gives You
[Michael Ryaboy](https://www.linkedin.com/in/michael-ryaboy-software-engineer) and the team at [Mintlify](https://mintlify.com) introduced [skill.md](https://mintlify.com/blog/skill-md): a static file that ships knowledge to the agent before it writes code. It's not documentation. It's a briefing. Decision tables, not tutorials. Gotchas, not explanations.
For Qdrant, a skill.md might include:
<div style="max-width: 640px; margin: 2rem auto; border-radius: 12px; overflow: hidden; font-family: 'JetBrains Mono', 'Fira Code', monospace; font-size: 14px; box-shadow: 0 4px 24px rgba(0,0,0,0.12);">
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-bottom: 2px solid #dc3545;">Decision Table</div>
<div style="background: #16213e; padding: 0;">
<table style="width: 100%; border-collapse: collapse; color: #e0e0e0; font-size: 14px;">
<thead>
<tr style="border-bottom: 1px solid #2a2a4e;">
<th style="text-align: left; padding: 12px 20px; color: #8890a8; font-weight: 400;">Want to...</th>
<th style="text-align: left; padding: 12px 20px; color: #8890a8; font-weight: 400;">Do</th>
</tr>
</thead>
<tbody>
<tr style="border-bottom: 1px solid #2a2a4e;">
<td style="padding: 10px 20px; color: #e0e0f0;">Search</td>
<td style="padding: 10px 20px;"><code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">client.query_points(col, query=vec, limit=10)</code></td>
</tr>
<tr style="border-bottom: 1px solid #2a2a4e;">
<td style="padding: 10px 20px; color: #e0e0f0;">Filter</td>
<td style="padding: 10px 20px;">Add <code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">query_filter=Filter(must=[...])</code></td>
</tr>
<tr style="border-bottom: 1px solid #2a2a4e;">
<td style="padding: 10px 20px; color: #e0e0f0;">Hybrid</td>
<td style="padding: 10px 20px;"><code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">prefetch</code> dense+sparse, <code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">FusionQuery(Fusion.RRF)</code></td>
</tr>
<tr>
<td style="padding: 10px 20px; color: #e0e0f0;">Multi-tenant</td>
<td style="padding: 10px 20px;"><code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">create_payload_index(field, is_tenant=True)</code></td>
</tr>
</tbody>
</table>
</div>
<div style="background: #1a1a2e; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-top: 1px solid #2a2a4e; border-bottom: 2px solid #dc3545; color: #e0e0e0;">Gotchas</div>
<div style="background: #16213e; padding: 16px 20px; line-height: 1.8; color: #c0c0d8;">
<span style="color: #dc3545;">&#x2716;</span> <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">query_points</code> not <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">search</code> - all search variants deprecated<br>
<span style="color: #dc3545;">&#x2716;</span> Never one collection per user - use <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">is_tenant=True</code> index<br>
<span style="color: #dc3545;">&#x2716;</span> BM25 requires <code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">Modifier.IDF</code> - sparse search needs it
</div>
</div>
This prevents the agent from using `client.search()` (deprecated), creating a collection per user (anti-pattern), or misconfiguring sparse vectors (common mistake). These are known failure modes that documentation alone doesn't prevent. Agents don't read docs the way humans do.
## What REPL-First MCP Gives You
[Armin Ronacher](https://lucumr.pocoo.org/), creator of Flask and now building [Earendil](https://earendil.dev/), proposed a [different approach](https://lucumr.pocoo.org/2025/1/22/what-i-want-for-ai-tools/). Instead of 30 narrow MCP tools, give the agent a Python shell with the SDK pre-configured:
```python
# Agent can just run this
collections = client.get_collections()
print([c.name for c in collections.collections])
# Then inspect the actual schema
info = client.get_collection("products")
print(info.config.params.vectors)
```
The agent discovers what exists by asking the system directly. No tool for "list collections." No tool for "get schema." Just Python. The REPL handles the unknown unknowns: what's actually in your Qdrant instance right now.
## Why Neither Alone Works
**skill.md without REPL:** The agent knows *how* to use `query_points` but not *what* to query. It guesses collection names. It assumes payload fields. It writes syntactically correct code that fails at runtime.
**REPL without skill.md:** The agent can discover what exists but still uses deprecated methods. It creates collections with wrong configurations. It makes the same mistakes it would have made without the REPL, just with more information about the data.
Together, the agent workflow looks like this:
<div style="max-width: 640px; margin: 2rem auto; border-radius: 12px; overflow: hidden; font-family: 'JetBrains Mono', 'Fira Code', monospace; font-size: 14px; box-shadow: 0 4px 24px rgba(0,0,0,0.12);">
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-bottom: 2px solid #dc3545;">Agent Workflow</div>
<div style="background: #16213e; padding: 24px 28px;">
<div style="margin-bottom: 20px;">
<div style="color: #dc3545; font-weight: 700; margin-bottom: 6px;">1. Read skill.md</div>
<div style="padding-left: 20px; line-height: 1.7;">
<code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">"Use query_points, not search"</code><br>
<code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">"Never one collection per user"</code><br>
<code class="repl" style="color: #7fdbca !important; background: rgba(127,219,202,0.1) !important;">"BM25 needs Modifier.IDF"</code>
</div>
</div>
<div style="margin-bottom: 20px;">
<div style="color: #dc3545; font-weight: 700; margin-bottom: 6px;">2. Use REPL</div>
<div style="padding-left: 20px; line-height: 1.7;">
<code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">client.get_collections()</code><br>
<code class="repl" style="color: #addb67 !important; background: rgba(173,219,103,0.1) !important;">client.scroll("products", limit=1)</code><br>
<span style="color: #8890a8;"># Now knows: collection exists, payload has "category" field</span>
</div>
</div>
<div>
<div style="color: #dc3545; font-weight: 700; margin-bottom: 6px;">3. Write code</div>
<div style="padding-left: 20px; line-height: 1.7;">
<span style="color: #e0e0f0;">Correct method + correct collection + correct filter fields</span>
</div>
</div>
</div>
</div>
The skill.md prevents known mistakes. The REPL handles environment-specific discovery. Both failure modes addressed.
## Implementation
The skill.md is just a file you drop into your project. The REPL is an MCP tool. A minimal implementation:
```python
@server.call_tool()
async def handle_tool_call(name: str, arguments: dict):
if name == "qdrant-repl":
code = arguments["code"]
# client, models pre-configured in repl_globals
try:
result = eval(compile(code, "<repl>", "eval"), repl_globals)
return repr(result) if result else "(ok)"
except SyntaxError:
exec(compile(code, "<repl>", "exec"), repl_globals)
return "(executed)"
```
State persists between calls. The agent builds up context incrementally.
## The Broader Pattern
This isn't specific to Qdrant. Any API with:
- Deprecated methods or migration paths → needs skill.md
- User-specific state (databases, collections, schemas) → needs REPL
The two approaches complement because they address orthogonal problems. Static knowledge for static mistakes. Dynamic access for dynamic discovery.
Most developer tools need both.
## This Doesn't Make It Easy
Adding skill.md and a REPL doesn't mean agents suddenly work flawlessly. They still hallucinate. They still misunderstand requirements. They still write code that technically runs but doesn't do what you wanted.
What these tools do is eliminate *unnecessary* failures. The agent won't fail because it used a deprecated method. That's a solved problem with skill.md. It won't fail because it guessed a collection name. The REPL lets it check. But it can still fail because it misunderstood what you meant by "similar products" or because the embedding model you're using doesn't capture the semantics you care about.
You might ask: why not just pull context from docs automatically? Tools that generate context from documentation solve a real problem, but they solve a different one. Auto-generated context gives you API surface area. A skill.md gives you judgment. "Don't create one collection per user" isn't in any API reference. "BM25 is broken without Modifier.IDF" isn't in any tutorial. Those lessons come from production pain, not documentation. The two approaches aren't in competition. Use both.
The goal isn't perfect agents. The goal is agents that fail for interesting reasons instead of boring ones. Deprecated API calls are boring failures. Wrong collection names are boring failures. Skill.md and REPL handle the boring stuff so you can focus on the hard problems.
## Try It
The [Qdrant skill.md](https://github.com/thierrypdamiba/skill-md) is 94 lines. It's a minimal example you can use, update, and configure for your own setup.
Drop it into your project:
```bash
mkdir -p .claude/skills/qdrant
curl -o .claude/skills/qdrant/skill.md \
https://raw.githubusercontent.com/thierrypdamiba/skill-md/main/README.md
```
For the REPL side, configure the [Qdrant MCP server with REPL](https://github.com/thierrypdamiba/mcp-server-qdrant-repl). It's a fork of the official MCP server that adds a `qdrant-repl` tool giving agents a stateful Python shell with a pre-configured `QdrantClient`. The agent gets both the briefing and the toolkit.
Your agents will stop using deprecated methods. They'll stop guessing collection names. They'll write correct queries on the first attempt. Not because they got smarter, but because they got the right information at the right time.
We'd love to hear how you're using Qdrant with AI agents. What's working, what's breaking, what patterns you've found. If you want to stay ahead of the curve on this stuff, [sign up for our newsletter](https://qdrant.tech/subscribe/). And thanks to you, our readers, for pushing us to make the developer experience better for everyone.
@@ -98,7 +98,7 @@ For that reason, when comparing two similar sentences, their embeddings will tur
<img src="/articles_data/what-is-a-vector-database/two-similar-vectors.png" alt="Comparison of the embeddings of 2 similar sentences" width="500">
That’s the beauty of embeddings. Tthe complexity of the data is distilled into something that can be compared across a multi-dimensional space.
That’s the beauty of embeddings. The complexity of the data is distilled into something that can be compared across a multi-dimensional space.
### 3. The Payload: Adding Context with Metadata
@@ -112,7 +112,7 @@ For example, if you’re searching for a picture of a dog, the vector helps the
<img src="/articles_data/what-is-a-vector-database/filtering-example.png" alt="Filtering Example" width="500">
The payload can help you narrow down those results by ignoring vectors that doesn't match your query vector filtering criteria. If you want the full picture of how filtering works in Qdrant, check out our [Complete Guide to Filtering.](https://qdrant.tech/articles/vector-search-filtering/)
The payload can help you narrow down those results by ignoring vectors that don't match your query vector filtering criteria. If you want the full picture of how filtering works in Qdrant, check out our [Complete Guide to Filtering.](https://qdrant.tech/articles/vector-search-filtering/)
## The Architecture of a Vector Database
@@ -290,7 +290,7 @@ Sparse vectors are ideal for tasks like **keyword search** or **metadata filteri
Sometimes context alone isn’t enough. Sometimes you need precision, too. Dense vectors are fantastic when you need to retrieve results based on the context or meaning behind the data. Sparse vectors are useful when you also need **keyword or specific attribute matching**.
> With hybrid search you don’t have to choose one over the othe and use both to get searches that are more **relevant** and **filtered**.
> With hybrid search you don’t have to choose one over the other and use both to get searches that are more **relevant** and **filtered**.
To achieve this balance, Qdrant uses **normalization** and **fusion** techniques to blend results from multiple search methods. One common approach is **Reciprocal Rank Fusion (RRF)**, where results from different methods are merged, giving higher importance to items ranked highly by both methods. This ensures that the best candidates, whether identified through dense or sparse vectors, appear at the top of the results.
@@ -324,9 +324,9 @@ This is just a simple example and there's so much more you can do with it. See o
![vector-database-architecture](/articles_data/what-is-a-vector-database/vector-database-2.jpeg)
As your vector dataset grow larger, so do the computational demands of searching through it.
As your vector dataset grows larger, so do the computational demands of searching through it.
Quantized vectors are much smaller and easier to compare. With methods like [**Binary Quantization**](https://qdrant.tech/articles/binary-quantization/), you can see **search speeds improve by up to 40x while memory usage decreases by 32x**. Improvements that can be decicive when dealing with large datasets or needing low-latency results.
Quantized vectors are much smaller and easier to compare. With methods like [**Binary Quantization**](https://qdrant.tech/articles/binary-quantization/), you can see **search speeds improve by up to 40x while memory usage decreases by 32x**. Improvements that can be decisive when dealing with large datasets or needing low-latency results.
It works by converting high-dimensional vectors, which typically use `4 bytes` per dimension, into binary representations, using just `1 bit` per dimension. Values above zero become "1", and everything else becomes "0".
@@ -473,7 +473,7 @@ In more advanced setups, Qdrant uses **JWT (JSON Web Tokens)** to enforce **Role
RBAC defines roles and assigns permissions, while JWT securely encodes these roles into tokens. Each request is validated against the user's JWT, ensuring they can only access or modify data based on their assigned permissions.
You can easily setup you access tokens and secure access to sensitive data through the **Qdrant Web UI:**
You can easily setup your access tokens and secure access to sensitive data through the **Qdrant Web UI:**
<img src="/articles_data/what-is-a-vector-database/jwt-web-ui.png" alt="Qdrant Web UI for generating a new access token." width="1000">
@@ -0,0 +1,95 @@
---
draft: false
title: "How Anima Health scaled clinical document intelligence with Qdrant"
short_description: "Anima Health scaled privacy-first clinical intelligence with Qdrant."
description: "Discover how Anima Health used Qdrant to power vector search for clinical document coding, privacy-first retrieval, and agentic workflows in UK primary care."
preview_image: /blog/case-study-anima-health/social_preview_partnership-anima-health.png
social_preview_image: /blog/case-study-anima-health/social_preview_partnership-anima-health.png
date: 2026-01-28
author: "Daniel Azoulai"
featured: true
tags:
- Anima Health
- vector search
- healthcare
- clinical document intelligence
- retrieval-augmented generation
- agentic ai
- privacy
- case study
---
![Anima Health scaled privacy-first clinical intelligence with Qdrant](/blog/case-study-anima-health/anima-bento.png)
Primary care systems across the UK are under intense strain. General practitioners (GPs) balance their time with patient demand, understaffing, administrative burden vs. delivering care. <a href="https://animahealth.com/" target="_blank">Anima Health</a> set out to address this challenge by building a clinical operating system designed to make primary care more efficient, more informed, and more humane for both clinicians and patients.
At the heart of Anima’s platform is the ability to process large volumes of unstructured clinical data, including documents, test results, referral letters, and notes, while maintaining strict privacy guarantees. To achieve this at scale, Anima relies on Qdrant as a core infrastructure component for vector search, similarity analysis, and agentic AI workflows.
### The challenge: Under-capacity clinics and unstructured data overload
Anima focuses on GP practices in the UK, where under-capacity is the defining operational constraint. Clinics must triage and treat more patients than ever, with too few clinicians and limited administrative resources.
A major bottleneck lies in unstructured clinical documents. These can include PDFs, handwritten notes, blood test results, and referral letters. Important information often arrives late, is poorly indexed, or requires manual review by clinicians.
>“The main bottleneck is GPs. We are completely understaffed across the UK, and there is no end in sight. Optimizing GP time is essential.”*
-Colin Cooke, Lead AI Engineer, Anima Health
This overload creates two compounding problems: First, clinicians lack timely access to information that could inform better decisions, and second, highly trained medical staff spend a disproportionate amount of time on administrative work rather than patient care.
### The solution: Vector-powered clinical intelligence with Qdrant
From the earliest stages of the product, Anima identified vector search as foundational infrastructure. Qdrant became the backbone of several key workflows, most notably clinical document coding.
Clinical documents are analyzed using large language models (LLMs) to extract meaning from unstructured text. However, LLMs alone cannot reliably handle medical ontologies such as SNOMED (Systematized Nomenclature of Medicine) codes, which are numeric identifiers with precise clinical meaning and downstream implications. These codes must be exact.
Anima represents SNOMED codes inside Qdrant as vector embeddings, enriched with metadata. During document processing, Qdrant is used as a retrieval layer inside an agentic pipeline. It narrows the search space, surfaces candidate codes, and enables high-confidence recommendations that clinicians can review and approve.
>“LLMs are great at understanding unstructured data, but they cannot free recall SNOMED codes. Those numeric IDs matter, and getting them wrong has real consequences.”*
-Colin Cooke, Lead AI Engineer, Anima Health
Beyond coding, Anima uses Qdrant to understand documents at scale. By working with embedded representations of documents rather than raw text, the system can identify patterns across documents while preserving patient privacy. These signals influence downstream workflows without exposing sensitive content to models or operators.
>“Embeddings let us do a lot with clinical data without ever coming close to violating patient privacy. That is something we lean on heavily.”*
-Colin Cooke, Lead AI Engineer, Anima Health
### Why Qdrant: Deployment control, cost predictability, and flexibility
Several factors made Qdrant a strong fit for healthcare workloads.
[Deployment flexibility](https://qdrant.tech/documentation/guides/installation/) was non-negotiable. Anima requires data to remain at rest in the UK, allowing them to make strong guarantees to customers about compliance and residency. Qdrant’s self-hosted and region-controlled deployment options enabled this without compromising performance.
Cost predictability also played a critical role. With a fixed infrastructure cost for vector search, Anima could use retrieval across multiple passes in their pipelines. This unlocked higher-quality results without eroding margins.
>“Knowing that retrieval is reliable and low cost changed how we build. We do not think twice about using vector search as part of our pipelines.”*
-Colin Cooke, Lead AI Engineer, Anima Health
Finally, Qdrant’s vector-native capabilities mattered. [Payload-based filtering](https://qdrant.tech/documentation/concepts/payload/) allows Anima to scope searches precisely across different electronic health record systems. [Multivector support](https://qdrant.tech/documentation/concepts/payload/) enables experimentation with multiple embedding strategies and providers, reducing long-term lock-in and easing future transitions.
### Results: Scalable, privacy-first AI in production
Since moving into production, Qdrant has scaled quietly alongside Anima’s growth. Despite rapid increases in workload, the vector layer required minimal operational attention, even as it became deeply embedded across multiple pipelines.
>“We experienced significant growth, and Qdrant was never something I had to think about. It just continued to work.”*
-Colin Cooke, Lead AI Engineer, Anima Health
Clinicians benefit from faster document processing, better prioritization, and reduced administrative overhead. Patients benefit from more timely and informed care decisions. Internally, Anima gained the confidence to build increasingly agentic workflows on top of a reliable retrieval foundation.
## What’s next for Anima Health
As Anima Health continues to scale, the team is looking beyond individual workflows toward more fully agentic clinical systems. These systems are designed to operate with greater autonomy while remaining tightly governed by clinical and regulatory constraints.
A key focus area is expanding how medical knowledge and patient context are represented and retrieved. Today, vector search already plays a central role in document understanding and classification. Going forward, Anima sees embeddings as a way to model longer-term patient state and clinical context that can evolve over time.
>“As we move toward more agentic systems, having reliable tools between the agent and our medical knowledge is essential. Qdrant is becoming one of those core tools.”*
-Colin Cooke, Lead AI Engineer, Anima Health
Another area of exploration is temporal relevance. Clinical information can become stale at different rates depending on its nature. Anima is interested in systems that can reason about how recent or uncertain a piece of information is, and adjust retrieval and decision-making accordingly.
The team is also investing in flexibility across their AI stack. This includes experimenting with multiple embedding models in parallel and transitioning between providers without disrupting production systems. Qdrant’s capabilities make this kind of controlled experimentation possible, allowing Anima to adopt better models as they emerge.
>“I do not want to be locked into a single embedding provider. Multivector support gives us the freedom to test and transition as models improve.”*
-Colin Cooke, Lead AI Engineer, Anima Health
Ultimately, Anima’s roadmap points toward AI systems that can support clinicians continuously rather than reactively. By grounding these systems in robust retrieval, strict privacy boundaries, and predictable infrastructure, Anima aims to scale clinical intelligence without sacrificing trust.
As healthcare organizations look to adopt more advanced AI workflows, Anima’s approach shows how vector search can serve not just as a technical optimization, but as a foundation for safe, future-proof clinical AI.
@@ -0,0 +1,97 @@
---
draft: false
title: "How Kakao Built an AI-Powered Internal Service Desk with Qdrant"
short_description: "Kakao built an AI-powered internal service desk with Qdrant."
description: "Discover how Kakao’s Connectivity Platform team built an AI-powered internal Service Desk using Qdrant to enable hybrid search, scale securely on Kubernetes, and improve employee productivity."
preview_image: /blog/case-study-kakao/social_preview_partnership-kakao.png
social_preview_image: /blog/case-study-kakao/social_preview_partnership-kakao.png
date: 2026-01-27
author: "David Koh - Kakao Connectivity Platform"
featured: true
tags:
- Kakao
- vector search
- hybrid search
- retrieval-augmented generation
- internal knowledge search
- case study
---
<a href="https://www.kakaocorp.com/" target="_blank">Kakao</a> is one of South Korea's leading technology companies, best known for KakaoTalk, the country's dominant messaging platform with over 48 million monthly active users. Beyond messaging, Kakao operates a broad ecosystem of services including maps, mobility, fintech, and enterprise solutions.
## Helping employees find answers faster without sacrificing precision or control
Kakao’s Connectivity Platform team set out to solve a familiar internal problem: employees across the organization needed a faster, more reliable way to get answers about internal systems, APIs, and operational procedures. The result was **Service Desk Agent**, an AI-powered internal service desk designed to answer questions in natural language using Kakao’s internal documentation and historical inquiry data.
Built as a Retrieval-Augmented Generation (RAG) system on top of LangGraph, Service Desk Agent acts as a conversational interface to Kakao’s internal knowledge. This helps employees resolve issues quickly, while reducing repetitive work for their support staff.
## The challenge: searching complex internal knowledge at scale
From the beginning, the team faced a search problem that couldn’t be solved with a single retrieval approach.
Service Desk Agent needed to work across two very different types of data:
1. Long-form technical documentation, including project guides and API specifications, where understanding context matters. But, exact system names and proper nouns still need to be matched.
2. Historical Q\&A and incident data, which often includes precise error messages, commands, and configuration details.
Pure keyword search struggled with semantic questions. Pure vector search struggled with exact terms and proper nouns. Neither approach alone was sufficient.
At the same time, Kakao had strict infrastructure and operational requirements. All data needed to remain within internal infrastructure, the system had to be deployable on Kubernetes, and the solution needed a permissive open-source license with strong official documentation for operations like upgrades, backup, and recovery.
This was a greenfield project; Kakao wasn’t replacing an existing vector database. Instead, the team evaluated several options, including Milvus, Weaviate, Qdrant, and Elasticsearch, to determine the best foundation for RAG-based internal AI services.
## Why Kakao chose Qdrant
After evaluating multiple vector databases, the Connectivity Platform team selected Qdrant as their first vector search solution.
The decision came down to a combination of search quality, performance, and operational fit.
Qdrant’s hybrid search capabilities were a key factor. By supporting both dense vectors for semantic search and sparse vectors for keyword-based retrieval—combined using [Reciprocal Rank Fusion (RRF)](https://qdrant.tech/documentation/concepts/hybrid-queries/#reciprocal-rank-fusion-rrf), the team could address both conceptual questions and exact-match queries in a single system. Named Vectors made it possible to manage multiple vector types within the same collection.
Performance was another major consideration. Qdrant’s [Rust-based architecture](https://qdrant.tech/articles/why-rust/), efficient [HNSW implementation](https://qdrant.tech/course/essentials/day-2/what-is-hnsw/), and support for [scalar quantization (INT8)](https://qdrant.tech/documentation/guides/quantization/#scalar-quantization) provided low-latency search while optimizing memory usage. This was crucial for an internal service expected to scale over time.
From an operational standpoint, Qdrant fit naturally into Kakao’s environment. Its single-binary design simplified deployment, it ran reliably on Kubernetes, and it allowed Kakao to retain full control over data by self-hosting within internal infrastructure.
Finally, the team highlighted the developer experience: a well-designed Python SDK with async support, an intuitive Web UI for debugging and exploration, and detailed official documentation covering both usage and performance tuning.
## How Service Desk Agent uses Qdrant
Qdrant sits at the core of Service Desk Agent’s RAG architecture, acting as the system’s primary vector store.
The team integrated Qdrant using the asynchronous Python client (`AsyncQdrantClient`) to handle high query concurrency, with secure HTTPS communication for all data access.
Collections were designed around data sources, with separate collections for internal technical documentation, historical inquiry data, and a semantic cache used to speed up repeated queries. Metadata filtering allows the system to narrow search scope by service or time period, while maintaining fast response times.
Each collection stores both dense and sparse vectors using [Named Vectors](https://qdrant.tech/documentation/concepts/vectors/#named-vectors). Hybrid search results are merged using RRF to produce more accurate answers across different query types.
An automated indexing pipeline handles document ingestion end-to-end. This ranges from cleansing and chunking, to embedding generation, to batch upserts into Qdrant.
The entire system runs on Kakao’s internal Kubernetes cluster, using a replication-based Qdrant setup for high availability and rolling updates for zero-downtime deployments. The system integrates with distributed locking and state management solutions within the broader application architecture.
![indexing-pipeline](/blog/case-study-kakao/indexing-pipeline.png)
![search-pipeline-diagram](/blog/case-study-kakao/search-pipeline-diagram.png)
## The result: faster answers, lower support load
Today, Service Desk Agent supports approximately 1 million vectors, with the dataset continuing to grow as more internal knowledge is indexed. The system operates across multiple collections and uses 3072-dimensional embeddings from OpenAI’s model.
By combining hybrid search, metadata filtering, and semantic caching, the team was able to improve both search quality and response times. Scalar quantization reduced memory usage, while tuned HNSW parameters helped maintain low latency under load.
From a business perspective, the impact has been clear:
* Reduced workload for support staff, as more inquiries are resolved automatically before tickets are created.
* Improved employee satisfaction, with faster, self-service access to internal knowledge.
* Higher development productivity, enabled by Qdrant’s SDK, documentation, and ease of integration during rapid prototyping.
Before Service Desk Agent, employees searched across multiple internal knowledge bases manually; a process that could take several minutes depending on query complexity. Now, end-to-end response time averages under 30 seconds, from query submission to complete answer delivery.
## What’s next
Kakao plans to continue expanding Service Desk Agent by indexing more of its internal knowledge base. The team is also exploring deeper GraphRAG integration, potential multimodal search capabilities, and sharding strategies to support future scale and performance needs.
Within the Connectivity Platform team, Qdrant has become the foundation for RAG-based internal AI services. It provides the hybrid search, flexibility, and operational stability required to support AI-driven knowledge access at scale.
*This post authored by David Koh — Kakao Connectivity Platform*
@@ -4,6 +4,8 @@ draft: false
slug: qdrant-1.15.x
short_description: "Smarter Quantization, Healing Indexes, and Multilingual Text Filtering"
description: "Qdrant v1.15 release presents new Quantization Features, advanced Full-Text filtering and a bunch of performance optimizations"
preview_image: /blog/qdrant-1.15.x/social_preview.jpg
social_preview_image: /blog/qdrant-1.15.x/social_preview.jpg
date: 2025-07-18T00:00:00-08:00
author: Derrick Mwiti
featured: true
@@ -4,6 +4,8 @@ draft: false
slug: qdrant-1.16.x
short_description: "v1.16 of Qdrant focuses on tiered multitenancy with tenant promotion and disk-efficient vector search."
description: "v1.16 of Qdrant focuses on tiered multitenancy with tenant promotion, disk-efficient vector search with inline storage, and improved filtered vector search with ACORN."
preview_image: /blog/qdrant-1.16.x/social_preview.jpg
social_preview_image: /blog/qdrant-1.16.x/social_preview.jpg
date: 2025-11-19T00:00:00-08:00
author: Abdon Pijpelink
featured: true
@@ -0,0 +1,68 @@
---
title: "Qdrant Academy Expands with Official Certification"
draft: false
slug: qdrant-certification-launch
short_description: "Qdrant Academy launched with first course, Qdrant Essentials"
description: "Master the art of production-grade retrieval with Qdrant Academy’s new certification. Earn credentials, score exclusive swag, and level up your engineering skills."
preview_image: /blog/qdrant-certification-launch/hero-graphic.png
social_preview_image: /blog/qdrant-certification-launch/hero-graphic.png
date: 2026-01-28
author: Neil Kanungo
featured: true
tags:
- Community
- Academy
---
Since we first announced **[Qdrant Academy](https://qdrant.tech/course/)**, our mission has been to provide developers with more than just documentation. We wanted to build a structured path to mastering vector search. As the AI search landscape matures, the distinction between a simple storage layer and a high-performance vector search engine has become the defining factor in production-grade RAG and recommendation systems.
Today, we are thrilled to take the next step in that mission. It’s time to move from learning to proving your expertise with the launch of our first official certification.
### Introducing the "Qdrant Essentials" Certification
The [Qdrant Essentials course](https://qdrant.tech/course/essentials/) has already helped thousands of developers understand the "why" behind high-dimensional search. Now, you can officially validate that knowledge.
By completing the course and passing the final exam at **[train.qdrant.dev](https://train.qdrant.dev)**, you’ll earn a digital credential that proves you can architect search systems that are as efficient as they are accurate.
#### What the Essentials Track Covers:
* **Engine Architecture:** Deep dives into HNSW, distance metrics, and collection structures.
* **Precision Filtering:** Mastering payload-based filtering without sacrificing search speed.
* **Hybrid Search:** Implementing a mix of dense and sparse vectors for superior retrieval.
* **Production Optimization:** Utilizing quantization and rescoring to scale your engine efficiently.
### Why Get Certified?
In a field as fast-moving as AI, "knowing a bit of Python" isn't enough. Moving from a prototype to a production-ready system requires specialized engineering judgment. Becoming **#QdrantCertified** can be a game-changer for your career:
* **Verified Expertise:** It proves you understand the critical trade-offs—like balancing latency vs. accuracy—that separate a hobbyist project from enterprise infrastructure.
* **Career Differentiation:** As companies hunt for RAG and Agentic AI experts, this badge signals that you can handle high-scale vector search, reducing your onboarding time and making you an immediate asset.
* **Standardized Knowledge:** You aren't just learning from assorted tutorials; you’re learning the industry standard for high-performance retrieval directly from the creators of Qdrant.
* **Engineering Authority:** Gain the confidence to lead internal AI workshops or architect your company's next-gen search platform using verified best practices.
### Get Certified. Get Swag.
We want to see those certificates! To celebrate the launch of our certification platform, we’re sending out some exclusive gear to our early achievers.
> **The first 30 people** to post their Qdrant Essentials certification to LinkedIn with the hashtag **#QdrantCertified** will receive a free Qdrant swag pack.
It’s simple: Learn, pass the exam at [train.qdrant.dev](https://train.qdrant.dev), and share your success with the community to claim your prize.
### More Courses Launching Soon
The "Essentials" course is just the foundation. Qdrant Academy is expanding rapidly to support developers at every stage of their journey:
#### The 2-Hour Beginner Launchpad
Coming soon, we are launching a **2-hour Basic Course**. This is designed for those who need a high-impact, low-time-commitment introduction to the world of vector search. You’ll go from "What is an embedding?" to "I have a running search engine" in a quick yet comprehensive Qdrant intro.
#### Advanced Retrieval Topics
For the power users, our upcoming **Multivectors Course** will tackle the cutting edge of retrieval. It will focus on Late Interaction models (like ColBERT), and will cover sophisticated retrieval with MUVERA. You’ll learn how to handle token-level embeddings to achieve incredible retrieval precision for complex datasets.
## Ready to Level Up?
Come grow with Qdrant, and prove your knowledge with Qdrant Certifications:
1. **Learn:** Head over to the [Qdrant Essentials course](https://qdrant.tech/course/essentials/).
2. **Certify:** Take the exam and claim your badge at **[train.qdrant.dev](https://train.qdrant.dev)**.
3. **Win:** Post it on LinkedIn with **#QdrantCertified** and grab your swag.
As always, happy coding!
+1 -1
View File
@@ -12,7 +12,7 @@ Qdrant Academy is your step-by-step learning hub for mastering vector search, hy
Whether you’re new to Qdrant or building production-grade systems, our guided courses help you go from beginner to expert, one module at a time.
Qdrant Academy has just launched in Fall of 2025 and currently offers one comprehensive course, but more are on the way! Register your interest in each upcoming below.
Qdrant Academy currently offers one comprehensive course, but more are on the way! Register your interest for upcoming courses below, or take the available course and [get certified](https://train.qdrant.dev)!
## Available Now
@@ -6,4 +6,17 @@ weight: 100
# Qdrant Essentials Certification
Coming soon! [Click here](https://forms.gle/QPSfdMjs3QpUCtGT9) to be notified when certifications become available.
Congratulations! You’ve officially navigated the complexities of the **Qdrant Essentials** course. You didn’t just learn how to store data; you learned how to architect a high-performance **Vector Search Engine**.
You’ve moved past simple "Hello World" tutorials and dove deep into HNSW indexing, hybrid search, and production-grade optimization. That effort deserves more than just a "finished" status. It deserves professional recognition.
## 🏆 Get #QdrantCertified
Your expertise is now production-ready. It’s time to validate those skills with our official certification.
**Head over to [train.qdrant.dev](https://train.qdrant.dev) to take the exam.**
Passing this exam proves you aren't just a user; you are a **Search Engineer** capable of:
* **Designing** scalable retrieval systems.
* **Optimizing** for both memory and latency.
* **Mastering** the nuances of the Qdrant architecture.
@@ -0,0 +1,127 @@
---
title: "On-Device Embeddings"
weight: 15
---
# On-Device Embeddings with Qdrant Edge and FastEmbed
To generate embeddings for use with Qdrant Edge directly on a device, you can use the [FastEmbed](/documentation/fastembed/) library. FastEmbed provides multimodal models that run efficiently on edge devices to generate vector embeddings from text and images.
# Provision the Device
Assuming the devices on which you will run Qdrant Edge have intermittent or no internet connectivity, you need to provision them with the necessary dependencies and model files ahead of time. First, install FastEmbed and the Qdrant Edge Python bindings:
```python
pip install fastembed qdrant-edge-py
```
Next, download the embedding models and save them locally on the device. Instantiate instances of `ImageEmbedding` and `TextEmbedding`, setting the `cache_dir` parameter to a local directory:
```python
from fastembed import ImageEmbedding, TextEmbedding
TEXT_MODEL_NAME='Qdrant/clip-ViT-B-32-text'
VISION_MODEL_NAME='Qdrant/clip-ViT-B-32-vision'
MODELS_DIR="./qdrant-edge-directory/models"
ImageEmbedding(
model_name=VISION_MODEL_NAME,
cache_dir=MODELS_DIR
)
TextEmbedding(
model_name=TEXT_MODEL_NAME,
cache_dir=MODELS_DIR
)
```
The models will be downloaded and cached in the specified `MODELS_DIR` directory, from where you can use them to generate embeddings.
# Generate Image Embeddings
First, initialize an Edge Shard as described in the [Qdrant Edge Quickstart Guide](/documentation/edge/edge-quickstart/).
<details>
<summary>Details</summary>
```python
from pathlib import Path
from qdrant_edge import (
Distance,
EdgeConfig,
EdgeShard,
VectorDataConfig,
)
SHARD_DIRECTORY = "./qdrant-edge-directory"
VECTOR_DIMENSION = 512
VECTOR_NAME="my-vector"
Path(SHARD_DIRECTORY).mkdir(parents=True, exist_ok=True)
config = EdgeConfig(
vector_data={
VECTOR_NAME: VectorDataConfig(
size=VECTOR_DIMENSION,
distance=Distance.Cosine,
)
}
)
edge_shard = EdgeShard(SHARD_DIRECTORY, config)
```
</details>
Assuming you have an image file `temp.jpg`, you can generate an embedding for it using FastEmbed's `ImageEmbedding` class and then store it in the Edge Shard:
```python
from pathlib import Path
from qdrant_edge import Point, UpdateOperation
import uuid
IMAGES_DIR = "images"
model = ImageEmbedding(
model_name=VISION_MODEL_NAME,
cache_dir=MODELS_DIR,
local_files_only=True
)
embeddings = list(model.embed([Path(IMAGES_DIR) / "temp.jpg"]))[0]
point = Point(
id=str(uuid.uuid4()),
vector={VECTOR_NAME: embeddings.tolist()}
)
edge_shard.update(UpdateOperation.upsert_points([point]))
```
Note the use of `cache_dir=MODELS_DIR` and `local_files_only=True` to load the image embedding model from the local directory where it was previously downloaded.
# Generate Text Embeddings
At query time, you can generate text embeddings using FastEmbed's `TextEmbedding` class. For example, to query the Edge Shard:
```python
from qdrant_edge import Query, QueryRequest
model = TextEmbedding(
model_name=TEXT_MODEL_NAME,
cache_dir=MODELS_DIR,
local_files_only=True
)
embeddings = list(model.embed(["<search terms>"]))[0]
results = edge_shard.query(
QueryRequest(
query=Query.Nearest(embeddings.tolist(),using=VECTOR_NAME),
limit=10,
with_vector=False,
with_payload=True
)
)
```
Again, using `cache_dir=MODELS_DIR` and `local_files_only=True` ensures the text embedding model is loaded from the local directory.
@@ -75,7 +75,7 @@ class QdrantQueryTool(Tool):
We can now set up `CodeAgent` to use our `QdrantQueryTool`.
```python
from smolagents import CodeAgent, HfApiModel
from smolagents import CodeAgent, InferenceClientModel, LogLevel
import os
# HuggingFace Access Token
@@ -83,7 +83,7 @@ import os
os.environ["HF_TOKEN"] = "----------"
agent = CodeAgent(
tools=[QdrantQueryTool()], model=HfApiModel(), max_iterations=4, verbose=True
tools=[QdrantQueryTool()], model=InferenceClientModel(), max_steps=4, verbosity_level=LogLevel.DEBUG
)
```
+2 -2
View File
@@ -1,6 +1,6 @@
---
title: Edge beta
description: Edge beta
title: Qdrant Edge | Embedded and Edge AI Systems
description: Qdrant Edge | Embedded and Edge AI Systems
url: edge
aliases: [edge-beta, edge-beta/]
build:
@@ -6,12 +6,13 @@ questions:
answer: Teams building AI systems that need fast, local vector search on embedded or resource-constrained devices, such as robots, mobile apps, or IoT hardware.
- id: 1
question: Is this available to all Qdrant users?
answer: Not yet. Qdrant Edge is in private beta. We're selecting a limited number of partners based on technical fit and active edge deployment scenarios.
answer: Yes. Read the <a href="https://qdrant.tech/documentation/edge/edge-quickstart/">Quick Start guide</a>,
and view the <a href="https://github.com/qdrant/qdrant-edge-demo">demo</a> on GitHub.
- id: 2
question: What are the minimum requirements to join the beta?
answer: You should have a clear use case for on-device or offline vector search. Preference is given to companies working with embedded hardware or deploying agents at the edge.
- id: 3
question: How do I get access?
answer: Qdrant Edge is currently in private beta. If you're building edge-native or embedded AI systems and want early access, apply to join the beta.
answer: If you're building edge-native or embedded AI systems, apply to join the beta. Or, read the <a href="https://qdrant.tech/documentation/edge/edge-quickstart/">Quick Start guide</a>, and view the <a href="https://github.com/qdrant/qdrant-edge-demo">demo</a> on GitHub
sitemapExclude: true
---
@@ -1,6 +1,6 @@
---
title: Apply to Join the Beta
description: Private beta available to selected teams building embedded or edge-native AI systems.
description: Beta available to selected teams building embedded or edge-native AI systems.
form:
id: edge-beta-form
hubspotFormOptions: '{
@@ -8,7 +8,7 @@ label:
subtitle: Run Vector Search Inside Embedded and Edge AI Systems
description: Qdrant Edge is a lightweight, in-process vector search engine designed for embedded devices, autonomous systems, and mobile agents. It enables on-device retrieval with minimal memory footprint, no background services, and optional synchronization with Qdrant Cloud.
startFree:
text: Apply to Join the Beta
text: Join the Beta
url: "#form"
image:
src: /img/qdrant-edge-scheme.svg
+20 -16
View File
@@ -23,6 +23,17 @@ mainCards:
name: Robert Caulk
position: Founder of Emergent Methods
description: Robert is working with a team on AskNews.app to adaptively enrich, index, and report on over 1 million news articles per day
- id: 2
headTitle: Distinguished Ambassador
headIcon:
src: /icons/fill/sparks-purple.svg
alt: Sparks
image:
src: /img/stars/tarun-jain.jpg
alt: Tarun Jain Photo
name: Tarun Jain
position: Founding Engineer at Stealth and YouTube @AiwithTarun
description: Google Developer Expert in AI and two-time Google Summer of Code contributor (Red Hen Lab and caMicroscope). Creates educational AI content on YouTube and regularly speaks at tech events.
cards:
- id: 0
image:
@@ -137,76 +148,69 @@ cards:
position: Global Solutions Lead @ LTIMindtree
description: Tech leader with 14+ years of experience in AI and GenAI, helping global clients drive innovation through strategy, architecture, and data.
- id: 16
image:
src: /img/stars/tarun-jain.jpg
alt: Tarun Jain Photo
name: Tarun Jain
position: Super Agent @ AI Planet
description: Google Developer Expert in AI and two-time Google Summer of Code contributor (Red Hen Lab and caMicroscope). Creates educational AI content on YouTube and regularly speaks at tech events.
- id: 17
image:
src: /img/stars/athos-g.jpg
alt: Athos Georgiou Photo
name: Athos Georgiou
position: R&D Technical AI Lead @ CGI
description: R&D Technical AI Lead with 10+ years across AI, finance, academia, and energy. Open-source contributor and a track record in AI product development, training.
- id: 18
- id: 17
image:
src: /img/stars/deepak-chawla.jpg
alt: Deepak Chawla Photo
name: Deepak Chawla
position: Founder @ HiDevs
description: Founder of HiDevs with 10+ years of experience in AI, ML, and GenAI. Has mentored over 3,000 students, led 50+ webinars and workshops, teaching vector search and RAG using Qdrant.
- id: 19
- id: 18
image:
src: /img/stars/saurabh-rai.jpg
alt: Saurabh Rai Photo
name: Saurabh Rai
position: Lead Engineer @ Resume Matcher
description: Creator of Resume Matcher, an open-source tool trusted by 30K+ users worldwide to optimize resumes for ATS and real-world roles, powered by Qdrant.
- id: 20
- id: 19
image:
src: /img/stars/samir-akarioh.jpg
alt: Samir Akarioh Photo
name: Samir Akarioh
position: Developer Advocate @ Gatling
description: DevOps and AI educator, speaking globally. Former Qdrant support engineer, now creating educational content that often features real-world use of vector search.
- id: 21
- id: 20
image:
src: /img/stars/roan-weigert.jpg
alt: Roan Weigert Photo
name: Roan Weigert
position: DevRel AI Engineer @ APARAVI
description: AI creator spotlighting San Francisco’s tech scene through interviews with leading founders, investors, and experts. Shares videos on AI innovation AI innovation with Qdrant as the vector DB.
- id: 22
- id: 21
image:
src: /img/stars/faraz-photo.jpg
alt: Faraz Gurramkonda Photo
name: Faraz Gurramkonda
position: AI Software Engineer @ Great Lakes Civil Services
description: A software engineer focused on scalable AI, automation, and developer tools. Founder of Xautomation and creator of VRTestSniffer (ASE 2025).
- id: 23
- id: 22
image:
src: /img/stars/pavan-photo.jpg
alt: Pavan Vemuri Photo
name: Pavan Vemuri
position: Director of Product Engineering @ SDVerse
description: AI strategist and MLOps expert with 15+ years of experience in software-defined vehicles and scalable systems. Specialized in Gen AI, digital transformation, and vector database solutions.
- id: 24
- id: 23
image:
src: /img/stars/mihir-photo.jpg
alt: Mihir Inamdar Photo
name: Mihir Inamdar
position: AI @Sutherland
description: Mihir Inamdar is an Open Source Developer Advocate focused on scalable AI, generative models, and vector database–powered RAG systems.
- id: 25
- id: 24
image:
src: /img/stars/mohammed-photo.jpg
alt: Mohammed Arbi
name: Mohammed Arbi
position: Machine Learning Engineer
description: Building intelligent AI systems with Generative AI, RAG, MLOps, and Qdrant. Active GDG mentor, sharing knowledge and supporting the developer community.
- id: 26
- id: 25
image:
src: /img/stars/vaibhava-photo.jpg
alt: Vaibhava Laxmi