mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 10:28:29 +02:00
new migration section initial iteration
This commit is contained in:
@@ -0,0 +1,42 @@
|
|||||||
|
---
|
||||||
|
title: Data Synchronization
|
||||||
|
weight: 25
|
||||||
|
is_empty: false
|
||||||
|
partition: qdrant
|
||||||
|
---
|
||||||
|
|
||||||
|
# Keeping Postgres and Qdrant in Sync
|
||||||
|
|
||||||
|
If you've migrated your vectors to Qdrant but still use Postgres as your source of truth, the next challenge is keeping both systems in sync as data changes.
|
||||||
|
|
||||||
|
This section covers three progressively robust sync architectures — from simple application-level dual-writes to production-grade Change Data Capture — with working code, failure mode analysis, and clear guidance on when to use each.
|
||||||
|
|
||||||
|
Not sure if you need a dedicated vector store alongside Postgres? Read [pgvector Tradeoffs](/documentation/data-synchronization/pgvector-tradeoffs/) to understand the six conditions under which pgvector is sufficient — and when you'll outgrow it.
|
||||||
|
|
||||||
|
## Three Tiers of Sync
|
||||||
|
|
||||||
|
| Tier | Pattern | Best For | Sync Lag |
|
||||||
|
| :--- | :--- | :--- | :--- |
|
||||||
|
| [Tier 1: Dual-Write](/documentation/data-synchronization/dual-writes/) | Write to both in request handler | Prototypes, < 10K records | None (synchronous) |
|
||||||
|
| [Tier 2: Transactional Outbox](/documentation/data-synchronization/transactional-outbox/) | Outbox table + background worker | Most production apps | Seconds |
|
||||||
|
| [Tier 3: Change Data Capture](/documentation/data-synchronization/change-data-capture/) | Debezium + Redpanda/Kafka | High-throughput, multi-consumer | Seconds |
|
||||||
|
|
||||||
|
## Choosing Your Tier
|
||||||
|
|
||||||
|
```
|
||||||
|
Do you have < 10K records and low write volume?
|
||||||
|
└── Yes → Tier 1 (dual-write) is fine to start
|
||||||
|
|
||||||
|
Does Qdrant downtime need to be invisible to your write path?
|
||||||
|
└── Yes → Go to Tier 2
|
||||||
|
|
||||||
|
Do you already run Kafka/Redpanda infrastructure?
|
||||||
|
└── Yes → Tier 3 is a natural fit
|
||||||
|
|
||||||
|
Do multiple services (not just Qdrant) need to react to data changes?
|
||||||
|
└── Yes → Tier 3
|
||||||
|
|
||||||
|
Otherwise → Tier 2
|
||||||
|
```
|
||||||
|
|
||||||
|
These tiers aren't permanent decisions. Start with Tier 1. When you hit its limits — Qdrant outages generating too much drift, write latency becoming noticeable — move to Tier 2. Only when Tier 2 becomes a bottleneck or you need replay capability should you invest in Tier 3.
|
||||||
@@ -0,0 +1,144 @@
|
|||||||
|
---
|
||||||
|
title: "Tier 3: Change Data Capture"
|
||||||
|
weight: 40
|
||||||
|
---
|
||||||
|
|
||||||
|
# Tier 3: Change Data Capture with Debezium + Redpanda
|
||||||
|
|
||||||
|
> "Let Postgres tell Qdrant what changed"
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
CDC is architecturally different from [dual-writes](/documentation/data-synchronization/dual-writes/) and the [transactional outbox](/documentation/data-synchronization/transactional-outbox/) in a fundamental way: **the application code has no awareness of Qdrant**. The FastAPI routes are pure Postgres CRUD — they don't import the Qdrant client, they don't write to an outbox. Sync is handled entirely in the infrastructure layer.
|
||||||
|
|
||||||
|
## How It Works
|
||||||
|
|
||||||
|
Postgres's Write-Ahead Log (WAL) is a sequential log of every change to the database — it exists for crash recovery and replication. With `wal_level = logical`, external consumers can read this log in a structured format.
|
||||||
|
|
||||||
|
Debezium connects to Postgres via a logical replication slot and captures every INSERT, UPDATE, and DELETE on the `products` table as a JSON event. These events are published to a Redpanda topic. A Python consumer service reads from that topic and calls Qdrant.
|
||||||
|
|
||||||
|
## Postgres Setup
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- Enable in postgresql.conf (or docker-compose command args):
|
||||||
|
-- wal_level = logical
|
||||||
|
|
||||||
|
-- Create a publication for the products table
|
||||||
|
CREATE PUBLICATION products_publication FOR TABLE products;
|
||||||
|
```
|
||||||
|
|
||||||
|
## Debezium Connector Config
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"name": "products-connector",
|
||||||
|
"config": {
|
||||||
|
"connector.class": "io.debezium.connector.postgresql.PostgresConnector",
|
||||||
|
"database.hostname": "postgres",
|
||||||
|
"database.port": "5432",
|
||||||
|
"database.user": "debezium",
|
||||||
|
"database.dbname": "fashiondb",
|
||||||
|
"database.server.name": "pgserver",
|
||||||
|
"table.include.list": "public.products",
|
||||||
|
"plugin.name": "pgoutput",
|
||||||
|
"publication.name": "products_publication",
|
||||||
|
"slot.name": "qdrant_sync",
|
||||||
|
"topic.prefix": "pgserver",
|
||||||
|
"transforms": "unwrap",
|
||||||
|
"transforms.unwrap.type": "io.debezium.transforms.ExtractNewRecordState",
|
||||||
|
"transforms.unwrap.delete.handling.mode": "rewrite"
|
||||||
|
}
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
The `ExtractNewRecordState` transform unwraps Debezium's envelope format into a flat payload, and adds an `__op` field (`c` = create, `u` = update, `d` = delete) so the consumer doesn't have to parse the raw Debezium schema.
|
||||||
|
|
||||||
|
## The Consumer Service
|
||||||
|
|
||||||
|
```py
|
||||||
|
from confluent_kafka import Consumer
|
||||||
|
|
||||||
|
def consume_and_sync():
|
||||||
|
consumer = Consumer({
|
||||||
|
"bootstrap.servers": "redpanda:9092",
|
||||||
|
"group.id": "qdrant-sync",
|
||||||
|
"auto.offset.reset": "earliest",
|
||||||
|
"enable.auto.commit": "false", # manual commit after successful processing
|
||||||
|
})
|
||||||
|
consumer.subscribe(["pgserver.public.products"])
|
||||||
|
|
||||||
|
while True:
|
||||||
|
msg = consumer.poll(1.0)
|
||||||
|
if msg is None:
|
||||||
|
continue
|
||||||
|
|
||||||
|
event = json.loads(msg.value())
|
||||||
|
op = event.get("__op") # 'c', 'u', or 'd'
|
||||||
|
|
||||||
|
if op in ("c", "u"):
|
||||||
|
upsert_to_qdrant(event)
|
||||||
|
elif op == "d":
|
||||||
|
delete_from_qdrant(event["article_id"])
|
||||||
|
|
||||||
|
consumer.commit() # commit only after successful processing
|
||||||
|
```
|
||||||
|
|
||||||
|
Manual offset commits (committing only after successfully processing each message) give you at-least-once semantics: if the consumer crashes mid-message, it reprocesses from the last committed offset on restart.
|
||||||
|
|
||||||
|
## Redpanda vs. Kafka
|
||||||
|
|
||||||
|
The example uses Redpanda as the event bus. Redpanda is Kafka-wire-protocol-compatible but ships as a single binary with no ZooKeeper dependency — simpler to run. The consumer code is identical whether you point it at Redpanda or Apache Kafka; just change `bootstrap.servers`.
|
||||||
|
|
||||||
|
## The Killer Feature: Replaying History
|
||||||
|
|
||||||
|
CDC's most underappreciated advantage is replayability. If Qdrant needs to be rebuilt from scratch — new index configuration, migration to a new cluster, disaster recovery — you can do it by replaying the Redpanda topic from the beginning. Every change that ever happened to the products table is recorded in the event stream.
|
||||||
|
|
||||||
|
With [dual-write](/documentation/data-synchronization/dual-writes/) or the [outbox](/documentation/data-synchronization/transactional-outbox/), a full rebuild requires running a bulk export from Postgres. With CDC, the event stream *is* the rebuild mechanism.
|
||||||
|
|
||||||
|
## Failure Modes
|
||||||
|
|
||||||
|
| Failure | Consequence | Mitigation |
|
||||||
|
| :---- | :---- | :---- |
|
||||||
|
| Qdrant is down | Consumer pauses; Redpanda retains events | Consumer retries; resumes from offset on recovery |
|
||||||
|
| Redpanda is down | Debezium buffers; WAL grows | Monitor WAL size; Redpanda HA with replication in production |
|
||||||
|
| Debezium crashes | WAL retains changes since last checkpoint | Debezium resumes from replication slot on restart |
|
||||||
|
| Schema change in Postgres | Connector may need restart | Monitor connector status; test schema migrations in staging |
|
||||||
|
|
||||||
|
## The WAL Disk Bloat Problem
|
||||||
|
|
||||||
|
One operational hazard specific to CDC: Postgres holds WAL segments until the replication slot consumer acknowledges them. If your consumer is down for a long time, Postgres can accumulate significant disk usage. Monitor replication slot lag:
|
||||||
|
|
||||||
|
```sql
|
||||||
|
SELECT
|
||||||
|
slot_name,
|
||||||
|
pg_size_pretty(
|
||||||
|
pg_wal_lsn_diff(pg_current_wal_lsn(), confirmed_flush_lsn)
|
||||||
|
) AS lag
|
||||||
|
FROM pg_replication_slots;
|
||||||
|
```
|
||||||
|
|
||||||
|
Set up alerting if this exceeds a few GB. In extreme cases (consumer permanently down), you may need to drop the replication slot — which means re-seeding Qdrant from Postgres rather than from the event stream.
|
||||||
|
|
||||||
|
## When to Use This
|
||||||
|
|
||||||
|
- High write throughput systems where the outbox table would become a bottleneck
|
||||||
|
- When multiple downstream consumers need to react to changes (not just Qdrant)
|
||||||
|
- When application code must be fully decoupled from sync concerns
|
||||||
|
- Teams with existing Redpanda/Kafka infrastructure
|
||||||
|
- When replay-from-scratch capability is a requirement
|
||||||
|
|
||||||
|
Tier 3 is genuinely powerful, but it comes with real operational costs: Redpanda and Debezium to deploy, configure, monitor, and upgrade. If you're not already running streaming infrastructure, think hard before adding it for a single sync use case. [Tier 2](/documentation/data-synchronization/transactional-outbox/) handles most production scenarios with far less complexity.
|
||||||
|
|
||||||
|
## Comparison Matrix
|
||||||
|
|
||||||
|
| Dimension | Tier 1: Dual-Write | Tier 2: Outbox | Tier 3: CDC |
|
||||||
|
| :---- | :---- | :---- | :---- |
|
||||||
|
| **Sync complexity** | Low | Medium | High |
|
||||||
|
| **Extra infrastructure** | None | Outbox table + worker | Redpanda + Debezium + consumer |
|
||||||
|
| **Consistency model** | Best-effort | At-least-once, eventual | At-least-once, eventual |
|
||||||
|
| **Write latency impact** | Adds Qdrant round-trip | None (async) | None (async) |
|
||||||
|
| **Qdrant downtime impact** | Generates drift | Events queue in Postgres | Events queue in Redpanda |
|
||||||
|
| **Captures direct SQL** | No | No | Yes (all WAL changes) |
|
||||||
|
| **Replay capability** | No | Limited (outbox retention) | Yes (Redpanda retention) |
|
||||||
|
| **Operational overhead** | Minimal | Low-Medium | High |
|
||||||
|
| **Best for** | Prototypes, internal tools | Most production apps | High-throughput, multi-consumer |
|
||||||
@@ -0,0 +1,99 @@
|
|||||||
|
---
|
||||||
|
title: "Tier 1: Dual-Writes"
|
||||||
|
weight: 20
|
||||||
|
---
|
||||||
|
|
||||||
|
# Tier 1: Application-Level Dual-Write
|
||||||
|
|
||||||
|
> "Just do it in your app code"
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
Every CRUD endpoint writes to Postgres first, then to Qdrant, in the same request handler. If the Qdrant write fails, the error is logged but the request succeeds — Postgres is the source of truth, and a reconciliation job can fix drift later.
|
||||||
|
|
||||||
|
## The Code
|
||||||
|
|
||||||
|
The route handler is exactly what you'd expect: one write after the other, with error handling around the Qdrant call:
|
||||||
|
|
||||||
|
```py
|
||||||
|
@router.post("/products", response_model=ProductResponse, status_code=201)
|
||||||
|
async def create_product(product: ProductCreate):
|
||||||
|
# 1. Write to Postgres first — it is the source of truth
|
||||||
|
row = await insert_product(product.model_dump())
|
||||||
|
|
||||||
|
# 2. Write to Qdrant — non-blocking on failure; reconcile catches drift
|
||||||
|
try:
|
||||||
|
await upsert_product(row)
|
||||||
|
except Exception as exc:
|
||||||
|
logger.error("Qdrant upsert failed for %s: %s", product.article_id, exc)
|
||||||
|
|
||||||
|
return row
|
||||||
|
```
|
||||||
|
|
||||||
|
The same pattern applies to every mutating operation: Postgres first, Qdrant second, exceptions caught and logged but not re-raised.
|
||||||
|
|
||||||
|
## Failure Modes
|
||||||
|
|
||||||
|
| Failure | Consequence | Mitigation |
|
||||||
|
| :---- | :---- | :---- |
|
||||||
|
| Qdrant is down | Postgres write succeeds; Qdrant write silently skipped | Logged; `reconcile` fixes drift |
|
||||||
|
| Qdrant is slow | Request latency spikes (blocks on Qdrant call) | Client has configurable timeout |
|
||||||
|
| Postgres fails after Qdrant write | Orphaned point in Qdrant | Write Postgres first; `reconcile --fix` cleans up |
|
||||||
|
| Network partition | Partial writes | Reconciliation script |
|
||||||
|
|
||||||
|
## What This Approach Gets Right
|
||||||
|
|
||||||
|
The strongest argument for dual-write isn't correctness — it's cognitive simplicity. A new engineer can read this code and immediately understand the entire sync story. There are no background workers, no queues, no separate processes. The request handler is the sync mechanism.
|
||||||
|
|
||||||
|
This simplicity has real value for prototypes, internal tools, and early-stage products where iteration speed matters more than operational rigor.
|
||||||
|
|
||||||
|
## Where It Falls Apart
|
||||||
|
|
||||||
|
The write path is coupled to Qdrant availability. If Qdrant has a hiccup — even a brief one — you're generating drift that has to be cleaned up later. There's no guarantee that every write will reach Qdrant; you're relying on reconciliation to eventually make things right.
|
||||||
|
|
||||||
|
More subtly: the request latency includes the Qdrant round-trip. For write-heavy workloads, this becomes a bottleneck.
|
||||||
|
|
||||||
|
## When to Use This
|
||||||
|
|
||||||
|
- Prototypes and MVPs
|
||||||
|
- Internal tools where occasional inconsistency is tolerable
|
||||||
|
- Low write throughput (< 10K products, < a few hundred writes/day)
|
||||||
|
- Teams that want to ship fast and revisit operational concerns later
|
||||||
|
|
||||||
|
## The Universal Safety Net: Reconciliation
|
||||||
|
|
||||||
|
Every sync architecture drifts eventually. The reconciliation script is what catches the residue:
|
||||||
|
|
||||||
|
```py
|
||||||
|
async def reconcile(fix: bool = False) -> ReconcileResult:
|
||||||
|
pg_ids = set(await get_all_article_ids_from_postgres())
|
||||||
|
qdrant_ids = set(await get_all_point_ids_from_qdrant())
|
||||||
|
|
||||||
|
missing_in_qdrant = pg_ids - qdrant_ids # need to sync
|
||||||
|
orphaned_in_qdrant = qdrant_ids - pg_ids # need to delete
|
||||||
|
|
||||||
|
if fix:
|
||||||
|
for article_id in missing_in_qdrant:
|
||||||
|
product = await get_product(article_id)
|
||||||
|
await upsert_product(product)
|
||||||
|
|
||||||
|
if orphaned_in_qdrant:
|
||||||
|
await qdrant_client.delete(
|
||||||
|
collection_name="products",
|
||||||
|
points_selector=orphaned_in_qdrant,
|
||||||
|
)
|
||||||
|
|
||||||
|
return ReconcileResult(
|
||||||
|
postgres_count=len(pg_ids),
|
||||||
|
qdrant_count=len(qdrant_ids),
|
||||||
|
missing_in_qdrant=len(missing_in_qdrant),
|
||||||
|
orphaned_in_qdrant=len(orphaned_in_qdrant),
|
||||||
|
in_sync=len(missing_in_qdrant) == 0 and len(orphaned_in_qdrant) == 0,
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
Postgres is the source of truth; Qdrant is a derived read store. When they diverge, Postgres wins. Run this on a schedule — nightly is usually sufficient — and on-demand when you suspect drift.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Next:** [Tier 2: Transactional Outbox](/documentation/data-synchronization/transactional-outbox/) — decouple Qdrant from your write path.
|
||||||
@@ -0,0 +1,85 @@
|
|||||||
|
---
|
||||||
|
title: pgvector Tradeoffs
|
||||||
|
weight: 10
|
||||||
|
---
|
||||||
|
|
||||||
|
# "Start with pgvector": Why You Might Outgrow It Faster Than You Think
|
||||||
|
|
||||||
|
The most common advice in every vector database thread online is some version of "start with pgvector, graduate later." We analyzed 110+ community threads from Hacker News and Reddit to see if the data supports this heuristic. The short answer is that it's more nuanced than it sounds, and most applications will hit its limits sooner than expected.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Appeal of "Just Use pgvector"
|
||||||
|
|
||||||
|
This advice is attractive for obvious reasons. If you're already running Postgres — and most teams are — pgvector gives you vector search without new infrastructure, new ops burden, or new sync headaches. One system, one deployment.
|
||||||
|
|
||||||
|
> *"My decision tree looks like this: Use pgvector until I have a very specific reason not to."*
|
||||||
|
|
||||||
|
> *"Default to pgvector, avoid premature optimization."*
|
||||||
|
|
||||||
|
The people giving this advice are usually running Postgres for transactional data and found pgvector sufficient for a few thousand vectors. For that use case it works.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Six Conditions That Must All Hold
|
||||||
|
|
||||||
|
pgvector is a reasonable default, but only when six specific conditions hold *simultaneously*.
|
||||||
|
|
||||||
|
**1. Your vector dataset is under ~1M vectors.** The community's empirical ceiling is around 10M, but the comfortable range is much lower. Above 1M you'll start hitting index-build times, memory pressure, and recall degradation under load.
|
||||||
|
|
||||||
|
This threshold is also easier to hit than you'd expect — especially if you're working with **multivectors**. Techniques like ColBERT-style late interaction generate one embedding *per token* rather than one per document, so a corpus of 100K documents can easily balloon into tens of millions of vectors overnight. Qdrant, by contrast, has [native multivector support](/documentation/concepts/vectors/#multivectors) with dedicated documentation and query APIs built around it.
|
||||||
|
|
||||||
|
**2. You don't need accurate metadata filtering.** If every search is against the full collection, post-filtering won't limit you. But the moment you need to scope searches to a user, tenant, category, or any selective predicate, pgvector generates unnecessary search overhead.
|
||||||
|
|
||||||
|
> *"I think the most relevant weakness for pgvector is the lack of 'proper' prefiltering on metadata while leveraging the vector index."*
|
||||||
|
|
||||||
|
Qdrant takes an entirely different approach to filtering. Specifically, Qdrant utilizes a [filterable HNSW](/documentation/concepts/indexing/#filtrable-index) which lets you traverse the nearest-neighbor graph while maintaining metadata filters.
|
||||||
|
|
||||||
|
**3. Your embeddings are tightly coupled to relational data.** If vectors are just an attribute of a row (e.g., a product description embedding alongside the product), colocation helps. If vectors are first-class entities, the argument for co-location weakens.
|
||||||
|
|
||||||
|
**4. You don't need hybrid search.** While pgvector supports dense vector similarity search via HNSW, the Postgres extension ecosystem still lacks a high-quality BM25 implementation — a critical component for hybrid search.
|
||||||
|
|
||||||
|
Postgres *does* have full-text search via `tsvector`/`tsquery`, and it's excellent for what it does. But that's lexical search — exact term matches, stemming, and stop words. BM25 is a probabilistic model that considers term frequency, inverse document frequency, and document length. Qdrant supports [native BM25 via sparse vectors](/documentation/concepts/vectors/#sparse-vectors). They're not the same thing.
|
||||||
|
|
||||||
|
**5. Postgres is already doing the heavy lifting for your business logic.** You have existing transactions, schemas, and ACID guarantees that matter. Adding a second data store splits that concern. If Postgres is truly central, the operational argument for colocation is real.
|
||||||
|
|
||||||
|
**6. Your team is small and search logic in SQL is manageable.** Embedding search logic in the database is typically an anti-pattern, but for small teams with simple search needs, pgvector's colocation reduces cognitive overhead.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Most Applications Fail at Least Two of These
|
||||||
|
|
||||||
|
Individually, each of these six criteria are not hard to satisfy. However, you need *all six* to hold in order to comfortably stay within pgvector, and hitting two or three disqualifiers happens fast.
|
||||||
|
|
||||||
|
Consider a typical B2B SaaS product that adds search. You'll almost certainly need tenant-scoped filtering (condition 2 fails). If you're searching structured content — product names, SKUs, technical specs — you'll want hybrid search (condition 4 fails). And your dataset will cross 1M faster than you expect once you're embedding documents, document chunks, and metadata (condition 1 is under pressure).
|
||||||
|
|
||||||
|
That's three conditions gone before you've even thought about scale. People are quick to point out that you might not immediately outgrow pgvector in terms of *scale*, but you almost certainly very quickly outgrow it in terms of *features*.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Where Dedicated Stores Win
|
||||||
|
|
||||||
|
When the conditions above don't all hold, dedicated vector stores offer concrete advantages:
|
||||||
|
|
||||||
|
- **Efficient metadata filtering** — pre-filter on metadata fields before computing similarity, avoiding wasted work on irrelevant vectors
|
||||||
|
- **Native hybrid search** — combine dense similarity and BM25 keyword matching in a single query with [reciprocal rank fusion](/documentation/concepts/hybrid-queries/)
|
||||||
|
- **Scale beyond 10M vectors** — purpose-built sharding, distributed indexing, and memory management
|
||||||
|
- **Decoupled architecture** — scale, optimize, and evolve your search layer independently of your relational database
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## The Sync Problem Is Real — But Solvable
|
||||||
|
|
||||||
|
There's a reason the "start with pgvector" advice persists despite these limitations. The #1 pain point developers report when running a dedicated vector store alongside Postgres is keeping them in sync:
|
||||||
|
|
||||||
|
> *"It was a PITA keeping data synced between pgsql <> qdrant."*
|
||||||
|
|
||||||
|
> *"One of the biggest nightmares with Pinecone was keeping the data in sync... With pg_vector integrated directly into my main database, this synchronization problem has completely disappeared."*
|
||||||
|
|
||||||
|
This is a legitimate concern, and we don't want to dismiss it. However, it's also a solved problem with well-known patterns, ranging from simple dual-writes for prototypes to transactional outbox patterns for production, to full CDC pipelines for high-throughput systems.
|
||||||
|
|
||||||
|
If you've decided you need a dedicated vector store, don't let sync anxiety push you back to pgvector. This guide walks through three progressively robust sync architectures — each with working code, failure mode analysis, and clear guidance on when to use which:
|
||||||
|
|
||||||
|
1. **[Dual-Writes](/documentation/data-synchronization/dual-writes/)** — simple application-level sync for prototypes
|
||||||
|
2. **[Transactional Outbox](/documentation/data-synchronization/transactional-outbox/)** — production-grade at-least-once delivery
|
||||||
|
3. **[Change Data Capture](/documentation/data-synchronization/change-data-capture/)** — infrastructure-level sync for high-throughput systems
|
||||||
@@ -0,0 +1,150 @@
|
|||||||
|
---
|
||||||
|
title: "Tier 2: Transactional Outbox"
|
||||||
|
weight: 30
|
||||||
|
---
|
||||||
|
|
||||||
|
# Tier 2: Transactional Outbox Pattern
|
||||||
|
|
||||||
|
> "Never lose a sync event"
|
||||||
|
|
||||||
|
## Architecture
|
||||||
|
|
||||||
|
Instead of writing to Qdrant directly from the request handler, we write an *event* into a `sync_outbox` table in the **same Postgres transaction** as the product write. A background worker picks up these events and syncs them to Qdrant asynchronously.
|
||||||
|
|
||||||
|
The outbox event exists if and only if the product write succeeded. There's no window between the two — they commit atomically.
|
||||||
|
|
||||||
|
## The Outbox Table
|
||||||
|
|
||||||
|
```sql
|
||||||
|
CREATE TABLE sync_outbox (
|
||||||
|
id BIGSERIAL PRIMARY KEY,
|
||||||
|
entity_id VARCHAR(20) NOT NULL, -- article_id
|
||||||
|
operation VARCHAR(10) NOT NULL, -- 'upsert' | 'delete'
|
||||||
|
payload JSONB, -- product snapshot at write time
|
||||||
|
status VARCHAR(20) DEFAULT 'pending',
|
||||||
|
attempts INT DEFAULT 0,
|
||||||
|
max_attempts INT DEFAULT 5,
|
||||||
|
last_error TEXT,
|
||||||
|
created_at TIMESTAMPTZ DEFAULT NOW(),
|
||||||
|
processed_at TIMESTAMPTZ
|
||||||
|
);
|
||||||
|
```
|
||||||
|
|
||||||
|
Note the `payload` column: it stores the full product data at write time. The worker doesn't re-query the products table — by the time it processes an event, the product may have been updated again. The payload is the snapshot that was valid when the event was created.
|
||||||
|
|
||||||
|
## The Write Path
|
||||||
|
|
||||||
|
Every CRUD endpoint opens a single Postgres transaction that writes both the product and the outbox event:
|
||||||
|
|
||||||
|
```py
|
||||||
|
@router.post("/products", response_model=ProductResponse, status_code=201)
|
||||||
|
async def create_product(product: ProductCreate):
|
||||||
|
pool = await get_pool()
|
||||||
|
async with pool.acquire() as conn:
|
||||||
|
async with conn.transaction():
|
||||||
|
row = await conn.fetchrow(INSERT_SQL, *product_values())
|
||||||
|
await enqueue_upsert(conn, product.article_id, dict(row))
|
||||||
|
return dict(row)
|
||||||
|
```
|
||||||
|
|
||||||
|
`enqueue_upsert` inserts a row into `sync_outbox` on the same connection, within the same transaction. If anything fails, both writes roll back together.
|
||||||
|
|
||||||
|
The request returns as soon as the Postgres transaction commits. Qdrant isn't touched during request handling.
|
||||||
|
|
||||||
|
## The Worker
|
||||||
|
|
||||||
|
A background async task processes outbox events in batches. The critical concurrency primitive is `FOR UPDATE SKIP LOCKED` — it lets multiple workers claim events safely without stepping on each other:
|
||||||
|
|
||||||
|
```py
|
||||||
|
async def process_batch() -> int:
|
||||||
|
pool = await get_pool()
|
||||||
|
async with pool.acquire() as conn:
|
||||||
|
rows = await conn.fetch(
|
||||||
|
"""
|
||||||
|
UPDATE sync_outbox
|
||||||
|
SET status = 'processing', attempts = attempts + 1
|
||||||
|
WHERE id IN (
|
||||||
|
SELECT id FROM sync_outbox
|
||||||
|
WHERE status IN ('pending', 'failed')
|
||||||
|
AND attempts < max_attempts
|
||||||
|
ORDER BY created_at ASC
|
||||||
|
LIMIT $1
|
||||||
|
FOR UPDATE SKIP LOCKED
|
||||||
|
)
|
||||||
|
RETURNING *
|
||||||
|
""",
|
||||||
|
BATCH_SIZE,
|
||||||
|
)
|
||||||
|
|
||||||
|
for row in rows:
|
||||||
|
await _process_event(dict(row))
|
||||||
|
|
||||||
|
return len(rows)
|
||||||
|
```
|
||||||
|
|
||||||
|
Events that fail get their `attempts` counter incremented and their `status` set back to `'pending'` (or `'failed'` if they've exceeded `max_attempts`). Qdrant upserts are naturally idempotent by point ID, so duplicate processing is harmless.
|
||||||
|
|
||||||
|
## Two Worker Modes
|
||||||
|
|
||||||
|
The implementation supports two delivery strategies:
|
||||||
|
|
||||||
|
**Polling** (default): the worker wakes up every N seconds, checks for pending events, and processes them. Simple, robust, and doesn't require a persistent database connection.
|
||||||
|
|
||||||
|
**LISTEN/NOTIFY**: a Postgres trigger fires `NOTIFY sync_outbox_insert` when a new row is inserted. The worker wakes up immediately, giving near-real-time sync without busy-polling:
|
||||||
|
|
||||||
|
```py
|
||||||
|
async def run_listen_worker() -> None:
|
||||||
|
conn = await asyncpg.connect(dsn)
|
||||||
|
await conn.execute("LISTEN sync_outbox_insert")
|
||||||
|
|
||||||
|
while True:
|
||||||
|
try:
|
||||||
|
await asyncio.wait_for(conn.wait_for_notify(), timeout=30.0)
|
||||||
|
except asyncio.TimeoutError:
|
||||||
|
pass # fall back to sweep anyway
|
||||||
|
|
||||||
|
await process_batch()
|
||||||
|
```
|
||||||
|
|
||||||
|
The 30-second sweep fallback is important: if a notification is missed (e.g., the worker was down briefly), the polling fallback ensures nothing stays stuck indefinitely.
|
||||||
|
|
||||||
|
## Failure Modes
|
||||||
|
|
||||||
|
| Failure | Consequence | Mitigation |
|
||||||
|
| :---- | :---- | :---- |
|
||||||
|
| Qdrant is down | Events queue in `sync_outbox` | Worker retries with backoff on recovery |
|
||||||
|
| Worker crashes | Events remain `pending` | Worker picks up on restart; `SKIP LOCKED` prevents double-processing |
|
||||||
|
| Duplicate processing | Same event processed twice | Qdrant upserts are idempotent by point ID |
|
||||||
|
| Outbox table grows large | Storage/performance impact | Prune `completed` events older than N days |
|
||||||
|
|
||||||
|
## What This Approach Gets Right
|
||||||
|
|
||||||
|
The write path is completely decoupled from Qdrant. If Qdrant is down for an hour, writes succeed normally and events queue up. When Qdrant comes back, the worker drains the queue. The application never blocks on Qdrant.
|
||||||
|
|
||||||
|
The delivery semantics are clear and enforceable: at-least-once, durable in Postgres, retried until success or `max_attempts`. Failed events are visible in the database and inspectable with SQL.
|
||||||
|
|
||||||
|
You also get a natural sync status endpoint:
|
||||||
|
|
||||||
|
```py
|
||||||
|
@router.get("/sync/status")
|
||||||
|
async def sync_status():
|
||||||
|
return await get_sync_status()
|
||||||
|
# → {"pending": 0, "failed": 0, "avg_lag_seconds": 1.2}
|
||||||
|
```
|
||||||
|
|
||||||
|
## The Trade-off
|
||||||
|
|
||||||
|
You have eventual consistency. There's a window — typically milliseconds to seconds — between when a product is written to Postgres and when it appears in Qdrant search results. For most applications this is perfectly acceptable. For applications where writes must be immediately searchable, this requires a different approach.
|
||||||
|
|
||||||
|
You also have a new table to manage: the outbox table grows over time and needs periodic cleanup of completed events.
|
||||||
|
|
||||||
|
## When to Use This
|
||||||
|
|
||||||
|
- **Most production applications** with moderate write throughput
|
||||||
|
- When Qdrant availability shouldn't impact your write path
|
||||||
|
- When you can tolerate seconds of sync lag
|
||||||
|
- **The recommended default for the majority of production Postgres + Qdrant deployments**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Next:** [Tier 3: Change Data Capture](/documentation/data-synchronization/change-data-capture/) — infrastructure-level sync for high-throughput systems.
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
---
|
||||||
|
#Delimiter files are used to separate the list of documentation pages into sections.
|
||||||
|
title: "Migrate to Qdrant"
|
||||||
|
type: delimiter
|
||||||
|
weight: 22 # Change this weight to change order of sections
|
||||||
|
sitemapExclude: True
|
||||||
|
_build:
|
||||||
|
publishResources: false
|
||||||
|
render: never
|
||||||
|
partition: qdrant
|
||||||
|
---
|
||||||
@@ -0,0 +1,11 @@
|
|||||||
|
| Tutorial | Objective | Stack | Time | Level |
|
||||||
|
| :--- | :--- | :--- | :--- | :--- |
|
||||||
|
| [Migration Tool Overview](/documentation/migrate-to-qdrant/) | Migrate vectors from any supported source. | <span class="pill">CLI</span> | Varies | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [From Pinecone](/documentation/migrate-to-qdrant/from-pinecone/) | Migrate from Pinecone serverless indexes. | <span class="pill">CLI</span> | 15m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [From Weaviate](/documentation/migrate-to-qdrant/from-weaviate/) | Migrate from Weaviate (pre-create collection). | <span class="pill">CLI</span> | 20m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [From Milvus](/documentation/migrate-to-qdrant/from-milvus/) | Migrate from Milvus/Zilliz with partitions. | <span class="pill">CLI</span> | 15m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [From Elasticsearch](/documentation/migrate-to-qdrant/from-elasticsearch/) | Migrate dense vectors from Elasticsearch. | <span class="pill">CLI</span> | 15m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [From pgvector](/documentation/migrate-to-qdrant/from-pgvector/) | Migrate from PostgreSQL pgvector tables. | <span class="pill">CLI</span> | 15m | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [Migration Verification](/documentation/migration-verification/) | Verify data integrity and search quality. | <span class="pill">Python</span> | 1h+ | <span class="text-yellow">Intermediate</span> |
|
||||||
|
| [pgvector Tradeoffs](/documentation/data-synchronization/pgvector-tradeoffs/) | When to outgrow pgvector for Qdrant. | <span class="pill">None</span> | 15m | <span class="text-green">Beginner</span> |
|
||||||
|
| [Postgres-Qdrant Sync](/documentation/data-synchronization/) | Keep Postgres and Qdrant in sync. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
|
||||||
@@ -0,0 +1,56 @@
|
|||||||
|
---
|
||||||
|
title: Migration Tool
|
||||||
|
weight: 23
|
||||||
|
is_empty: false
|
||||||
|
partition: qdrant
|
||||||
|
---
|
||||||
|
|
||||||
|
# Migrate to Qdrant
|
||||||
|
|
||||||
|
The [Qdrant Migration Tool](https://github.com/qdrant/migration) is a CLI that moves your vectors, metadata, and sparse embeddings from other vector databases into Qdrant. It runs as a Docker container, streams data in batches, and can resume interrupted migrations.
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker pull registry.cloud.qdrant.io/library/qdrant-migration
|
||||||
|
```
|
||||||
|
|
||||||
|
## Supported Sources
|
||||||
|
|
||||||
|
| Source | CLI Subcommand | Auto-Creates Collection? |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| [Pinecone](/documentation/migrate-to-qdrant/from-pinecone/) | `pinecone` | Yes |
|
||||||
|
| [Weaviate](/documentation/migrate-to-qdrant/from-weaviate/) | `weaviate` | No (must pre-create) |
|
||||||
|
| [Milvus](/documentation/migrate-to-qdrant/from-milvus/) | `milvus` | Yes |
|
||||||
|
| [Elasticsearch](/documentation/migrate-to-qdrant/from-elasticsearch/) | `elasticsearch` | Yes |
|
||||||
|
| [pgvector](/documentation/migrate-to-qdrant/from-pgvector/) | `pg` | Yes |
|
||||||
|
|
||||||
|
The tool also supports Chroma, Redis, MongoDB, OpenSearch, S3 Vectors, FAISS, Apache Solr, and [Qdrant-to-Qdrant](/documentation/tutorials-operations/migration/) migrations.
|
||||||
|
|
||||||
|
## General Advice
|
||||||
|
|
||||||
|
1. **Run the tool close to your databases.** Direct connectivity between source and target is not required — the tool streams through the machine it runs on. For best performance, use a machine with low latency to both.
|
||||||
|
|
||||||
|
2. **Use `--net=host` for local instances.** If either database runs on the host machine, the container needs host networking to reach `localhost`.
|
||||||
|
|
||||||
|
3. **The tool resumes by default.** Migration progress is tracked in a `_migration_offsets` collection in Qdrant. If a migration is interrupted, re-running the same command picks up where it left off. Use `--migration.restart` to force a fresh start.
|
||||||
|
|
||||||
|
4. **Batch size is tunable.** The default batch size is 50. For large migrations, increase it with `--migration.batch-size` (e.g., 256 or 512) to improve throughput.
|
||||||
|
|
||||||
|
## Universal CLI Options
|
||||||
|
|
||||||
|
These flags apply to all source types:
|
||||||
|
|
||||||
|
| Flag | Default | Description |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `--migration.batch-size` | 50 | Points per upsert batch |
|
||||||
|
| `--migration.restart` | false | Ignore saved progress, start fresh |
|
||||||
|
| `--migration.create-collection` | true | Auto-create target collection |
|
||||||
|
| `--migration.batch-delay` | 0 | Milliseconds between batches |
|
||||||
|
| `--migration.num-workers` | CPU cores | Parallel workers |
|
||||||
|
| `--debug` / `--trace` | — | Verbose logging |
|
||||||
|
|
||||||
|
## After Migration
|
||||||
|
|
||||||
|
Once your data is in Qdrant, verify that everything arrived correctly:
|
||||||
|
|
||||||
|
- **[Migration Verification Guide](/documentation/migration-verification/)** — a structured framework covering data integrity checks and search quality validation.
|
||||||
|
- **[Data Synchronization](/documentation/data-synchronization/)** — if you're running Postgres alongside Qdrant, learn how to keep them in sync.
|
||||||
@@ -0,0 +1,76 @@
|
|||||||
|
---
|
||||||
|
title: From Elasticsearch
|
||||||
|
weight: 40
|
||||||
|
---
|
||||||
|
|
||||||
|
# Migrate from Elasticsearch to Qdrant
|
||||||
|
|
||||||
|
## What You Need from Elasticsearch
|
||||||
|
|
||||||
|
- **Elasticsearch URL** — the HTTP endpoint
|
||||||
|
- **Index name** — the index containing your vectors
|
||||||
|
- **Credentials** — username/password or API key
|
||||||
|
|
||||||
|
## Concept Mapping
|
||||||
|
|
||||||
|
| Elasticsearch | Qdrant | Notes |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| Index | Collection | One-to-one mapping |
|
||||||
|
| Document | Point | Each document becomes a point |
|
||||||
|
| `dense_vector` field | Vector | Mapped automatically |
|
||||||
|
| Document fields | Payload | Non-vector fields become payload |
|
||||||
|
| `cosine` | `Cosine` | ES returns `1 - cosine_distance`; Qdrant returns cosine similarity directly |
|
||||||
|
| `l2_norm` | `Euclid` | Direct mapping |
|
||||||
|
| `dot_product` | `Dot` | Direct mapping |
|
||||||
|
|
||||||
|
## Run the Migration
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration elasticsearch \
|
||||||
|
--elasticsearch.url 'https://your-es-host:9200' \
|
||||||
|
--elasticsearch.index 'your-index' \
|
||||||
|
--elasticsearch.username 'elastic' \
|
||||||
|
--elasticsearch.password 'your-password' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### Using API Key Authentication
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration elasticsearch \
|
||||||
|
--elasticsearch.url 'https://your-es-host:9200' \
|
||||||
|
--elasticsearch.index 'your-index' \
|
||||||
|
--elasticsearch.api-key 'your-es-api-key' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### All Elasticsearch-Specific Flags
|
||||||
|
|
||||||
|
| Flag | Required | Description |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `--elasticsearch.url` | Yes | Elasticsearch HTTP endpoint |
|
||||||
|
| `--elasticsearch.index` | Yes | Index to migrate |
|
||||||
|
| `--elasticsearch.username` | No | Username for basic auth |
|
||||||
|
| `--elasticsearch.password` | No | Password for basic auth |
|
||||||
|
| `--elasticsearch.api-key` | No | API key for authentication |
|
||||||
|
| `--elasticsearch.insecure-skip-verify` | No | Skip TLS certificate verification |
|
||||||
|
|
||||||
|
## Hybrid Search Considerations
|
||||||
|
|
||||||
|
If your Elasticsearch setup uses hybrid BM25 + kNN scoring, you'll need to reconstruct this in Qdrant using [sparse vectors](/documentation/concepts/vectors/#sparse-vectors) (for BM25-like behavior) alongside dense vectors. The migration tool transfers the dense vectors; you'll need to generate sparse vectors separately if you want hybrid search in Qdrant.
|
||||||
|
|
||||||
|
Qdrant supports native hybrid search with [Reciprocal Rank Fusion (RRF)](/documentation/concepts/hybrid-queries/) to combine dense and sparse results.
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
- **Nested documents:** Elasticsearch nested documents need to be flattened or restructured for Qdrant's payload model.
|
||||||
|
- **Score normalization:** Elasticsearch `_score` values are not comparable to Qdrant scores. Use rank-based metrics (recall@k, Spearman correlation) rather than raw score comparison when [verifying your migration](/documentation/migration-verification/).
|
||||||
|
- **BM25 is not migrated:** The migration tool transfers vectors and document fields. If you relied on Elasticsearch's BM25 scoring, you'll need to set up sparse vectors in Qdrant separately.
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
After migration, verify your data arrived correctly with the [Migration Verification Guide](/documentation/migration-verification/).
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
---
|
||||||
|
title: From Milvus
|
||||||
|
weight: 30
|
||||||
|
---
|
||||||
|
|
||||||
|
# Migrate from Milvus to Qdrant
|
||||||
|
|
||||||
|
## What You Need from Milvus
|
||||||
|
|
||||||
|
- **Milvus URL** — the gRPC endpoint of your Milvus instance
|
||||||
|
- **Collection name** — the collection to migrate
|
||||||
|
- **API key** — if using Zilliz Cloud or authenticated Milvus
|
||||||
|
|
||||||
|
## Concept Mapping
|
||||||
|
|
||||||
|
| Milvus | Qdrant | Notes |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| Collection | Collection | One-to-one mapping |
|
||||||
|
| Partition | Payload field or separate collection | Use `--milvus.partitions` to specify which partitions to migrate |
|
||||||
|
| Schema fields | Payload | Non-vector fields become payload |
|
||||||
|
| `COSINE` | `Cosine` | Direct mapping |
|
||||||
|
| `L2` | `Euclid` | Direct mapping |
|
||||||
|
| `IP` (inner product) | `Dot` | Direct mapping |
|
||||||
|
| Dynamic fields | Payload | JSON-typed dynamic fields are preserved |
|
||||||
|
|
||||||
|
## Run the Migration
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration milvus \
|
||||||
|
--milvus.url 'your-milvus-host:19530' \
|
||||||
|
--milvus.collection 'your-collection' \
|
||||||
|
--milvus.api-key 'your-milvus-api-key' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### Migrating Specific Partitions
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration milvus \
|
||||||
|
--milvus.url 'your-milvus-host:19530' \
|
||||||
|
--milvus.collection 'your-collection' \
|
||||||
|
--milvus.partitions 'partition_a,partition_b' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### All Milvus-Specific Flags
|
||||||
|
|
||||||
|
| Flag | Required | Description |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `--milvus.url` | Yes | Milvus gRPC endpoint |
|
||||||
|
| `--milvus.collection` | Yes | Collection name to migrate |
|
||||||
|
| `--milvus.api-key` | No | API key (for Zilliz Cloud) |
|
||||||
|
| `--milvus.username` | No | Username for authentication |
|
||||||
|
| `--milvus.password` | No | Password for authentication |
|
||||||
|
| `--milvus.db-name` | No | Database name |
|
||||||
|
| `--milvus.partitions` | No | Comma-separated partition names |
|
||||||
|
| `--milvus.server-version` | No | Override detected server version |
|
||||||
|
| `--milvus.enable-tls-auth` | No | Enable TLS authentication |
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
- **Partition handling:** Milvus partitions can map to Qdrant collections or payload filters. If you merge partitions into a single collection, add a partition name as a payload field for filtering.
|
||||||
|
- **Schema strictness:** Milvus enforces schema on write; Qdrant is schema-flexible. Verify that the schema-less flexibility didn't cause payload fields to drift during migration.
|
||||||
|
- **Dynamic fields:** Milvus dynamic fields (introduced in 2.3) may serialize differently. Check that JSON-typed dynamic fields survived the migration with correct structure.
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
After migration, verify your data arrived correctly with the [Migration Verification Guide](/documentation/migration-verification/).
|
||||||
@@ -0,0 +1,72 @@
|
|||||||
|
---
|
||||||
|
title: From pgvector
|
||||||
|
weight: 50
|
||||||
|
---
|
||||||
|
|
||||||
|
# Migrate from pgvector to Qdrant
|
||||||
|
|
||||||
|
## What You Need from Postgres
|
||||||
|
|
||||||
|
- **Connection URL** — a standard Postgres connection string
|
||||||
|
- **Table name** — the table containing your vector data
|
||||||
|
|
||||||
|
## Concept Mapping
|
||||||
|
|
||||||
|
| pgvector | Qdrant | Notes |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| Table | Collection | One-to-one mapping |
|
||||||
|
| Row | Point | Each row becomes a point |
|
||||||
|
| `vector` column | Vector | Mapped automatically |
|
||||||
|
| Other columns | Payload | All non-vector columns become payload fields |
|
||||||
|
| `vector_cosine_ops` | `Cosine` | pgvector returns distance (1 - similarity); Qdrant returns similarity |
|
||||||
|
| `vector_l2_ops` | `Euclid` | Direct mapping |
|
||||||
|
| `vector_ip_ops` | `Dot` | pgvector uses negative inner product for ordering; scores will be inverted |
|
||||||
|
|
||||||
|
## Run the Migration
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration pg \
|
||||||
|
--pg.url 'postgres://user:password@host:5432/dbname' \
|
||||||
|
--pg.table 'your_embeddings_table' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### Selecting Specific Columns
|
||||||
|
|
||||||
|
By default, all columns are migrated. Use `--pg.columns` to select specific ones:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration pg \
|
||||||
|
--pg.url 'postgres://user:password@host:5432/dbname' \
|
||||||
|
--pg.table 'your_embeddings_table' \
|
||||||
|
--pg.columns 'id,embedding,title,category' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### All pgvector-Specific Flags
|
||||||
|
|
||||||
|
| Flag | Required | Description |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `--pg.url` | Yes | Postgres connection string |
|
||||||
|
| `--pg.table` | Yes | Table name to migrate |
|
||||||
|
| `--pg.key-column` | No | Column to use as point ID |
|
||||||
|
| `--pg.columns` | No | Comma-separated columns to migrate (default: all) |
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
- **Partition structure:** If you had manual partitions in pgvector (common at scale), verify that all partitions were migrated, not just the primary table.
|
||||||
|
- **NULL handling:** PostgreSQL NULLs may be dropped during export. Check that optional fields are represented correctly in Qdrant payloads.
|
||||||
|
- **Index type and recall:** pgvector supports IVFFlat and HNSW. If your baseline was captured with IVFFlat (lower recall), Qdrant's HNSW may return better results. This looks like a "mismatch" but is an improvement.
|
||||||
|
- **Row count approximation:** Postgres's `n_live_tup` is an estimate, not an exact count. Use `SELECT COUNT(*) FROM your_table` for accurate comparison during [migration verification](/documentation/migration-verification/).
|
||||||
|
|
||||||
|
## After Migration: Keeping Postgres and Qdrant in Sync
|
||||||
|
|
||||||
|
If you continue using Postgres as your source of truth alongside Qdrant, you'll need a sync strategy. The [Data Synchronization Guide](/documentation/data-synchronization/) covers three approaches from simple dual-writes to production-grade Change Data Capture.
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
After migration, verify your data arrived correctly with the [Migration Verification Guide](/documentation/migration-verification/).
|
||||||
@@ -0,0 +1,71 @@
|
|||||||
|
---
|
||||||
|
title: From Pinecone
|
||||||
|
weight: 10
|
||||||
|
---
|
||||||
|
|
||||||
|
# Migrate from Pinecone to Qdrant
|
||||||
|
|
||||||
|
## What You Need from Pinecone
|
||||||
|
|
||||||
|
- **API key** — from the [Pinecone console](https://app.pinecone.io/)
|
||||||
|
- **Index name** — the name of the index to migrate
|
||||||
|
- **Index host URL** — the host endpoint shown in your index dashboard
|
||||||
|
|
||||||
|
<aside role="status">Only Pinecone <strong>serverless</strong> indexes support listing all vectors for migration. Legacy pod-based indexes may require additional steps.</aside>
|
||||||
|
|
||||||
|
## Concept Mapping
|
||||||
|
|
||||||
|
| Pinecone | Qdrant | Notes |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| Index | Collection | One-to-one mapping |
|
||||||
|
| Namespace | Payload field or separate collection | No direct equivalent — the tool migrates all namespaces. Use `--pinecone.namespace` to migrate a specific one |
|
||||||
|
| Metadata | Payload | Direct mapping |
|
||||||
|
| Sparse values | Sparse vectors | Mapped to `sparse_vector` named vector by default |
|
||||||
|
| `cosine` | `Cosine` | Direct mapping |
|
||||||
|
| `dotproduct` | `Dot` | Pinecone requires unit-normalized vectors for dotproduct |
|
||||||
|
| `euclidean` | `Euclid` | Direct mapping |
|
||||||
|
|
||||||
|
## Run the Migration
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration pinecone \
|
||||||
|
--pinecone.index-host 'https://your-index-host.pinecone.io' \
|
||||||
|
--pinecone.index-name 'your-index' \
|
||||||
|
--pinecone.api-key 'pcsk_...' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### Migrating a Specific Namespace
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration pinecone \
|
||||||
|
--pinecone.index-host 'https://your-index-host.pinecone.io' \
|
||||||
|
--pinecone.index-name 'your-index' \
|
||||||
|
--pinecone.api-key 'pcsk_...' \
|
||||||
|
--pinecone.namespace 'my-namespace' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection'
|
||||||
|
```
|
||||||
|
|
||||||
|
### All Pinecone-Specific Flags
|
||||||
|
|
||||||
|
| Flag | Required | Description |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `--pinecone.index-name` | Yes | Name of the Pinecone index |
|
||||||
|
| `--pinecone.index-host` | Yes | Host URL of the Pinecone index |
|
||||||
|
| `--pinecone.api-key` | Yes | Pinecone API key |
|
||||||
|
| `--pinecone.namespace` | No | Specific namespace to migrate |
|
||||||
|
| `--pinecone.service-host` | No | Custom Pinecone service host |
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
- **Score scaling:** Pinecone cosine similarity returns values in [0, 1] (rescaled). Qdrant returns [-1, 1]. Rankings are identical, but raw scores won't match.
|
||||||
|
- **Metadata size limits:** Pinecone limits metadata to 40KB per vector. Qdrant has no per-payload size limit, so data is preserved as-is.
|
||||||
|
- **Namespace strategy:** If you have multiple namespaces, decide upfront whether to merge them into a single Qdrant collection (using a `namespace` payload field for filtering) or create separate collections.
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
After migration, verify your data arrived correctly with the [Migration Verification Guide](/documentation/migration-verification/).
|
||||||
@@ -0,0 +1,83 @@
|
|||||||
|
---
|
||||||
|
title: From Weaviate
|
||||||
|
weight: 20
|
||||||
|
---
|
||||||
|
|
||||||
|
# Migrate from Weaviate to Qdrant
|
||||||
|
|
||||||
|
## What You Need from Weaviate
|
||||||
|
|
||||||
|
- **Host URL** — the Weaviate instance address
|
||||||
|
- **Class name** — the class to migrate
|
||||||
|
- **Authentication** — API key, username/password, or bearer token depending on your setup
|
||||||
|
- **Vector dimensions** — Weaviate does not expose vector dimensions through its API, so you must know this value
|
||||||
|
|
||||||
|
<aside role="alert"><strong>Important:</strong> Because Weaviate does not expose vector dimensions, the migration tool cannot auto-create the Qdrant collection. You must create the collection manually before running the migration.</aside>
|
||||||
|
|
||||||
|
## Pre-Create Your Qdrant Collection
|
||||||
|
|
||||||
|
```bash
|
||||||
|
curl -X PUT 'https://your-instance.cloud.qdrant.io:6333/collections/your-collection' \
|
||||||
|
-H 'api-key: your-qdrant-api-key' \
|
||||||
|
-H 'Content-Type: application/json' \
|
||||||
|
-d '{
|
||||||
|
"vectors": {
|
||||||
|
"size": 384,
|
||||||
|
"distance": "Cosine"
|
||||||
|
}
|
||||||
|
}'
|
||||||
|
```
|
||||||
|
|
||||||
|
Replace `384` with your actual vector dimensions. Set the distance metric to match your Weaviate configuration.
|
||||||
|
|
||||||
|
## Concept Mapping
|
||||||
|
|
||||||
|
| Weaviate | Qdrant | Notes |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| Class | Collection | One-to-one mapping |
|
||||||
|
| Properties | Payload | Direct mapping |
|
||||||
|
| `cosine` | `Cosine` | Direct mapping |
|
||||||
|
| `l2-squared` | `Euclid` | Qdrant uses L2, not L2-squared; scores differ in magnitude but ranking is identical |
|
||||||
|
| `dot` | `Dot` | Direct mapping |
|
||||||
|
| Cross-references | Payload fields | Store referenced IDs as payload fields and rebuild linking in your application |
|
||||||
|
| Tenants | Payload field or separate collections | Use `--weaviate.tenant` to migrate a specific tenant |
|
||||||
|
|
||||||
|
## Run the Migration
|
||||||
|
|
||||||
|
```bash
|
||||||
|
docker run --net=host --rm -it registry.cloud.qdrant.io/library/qdrant-migration weaviate \
|
||||||
|
--weaviate.host 'your-weaviate-host.example.com' \
|
||||||
|
--weaviate.scheme https \
|
||||||
|
--weaviate.class-name 'YourClass' \
|
||||||
|
--weaviate.auth-type apiKey \
|
||||||
|
--weaviate.api-key 'your-weaviate-api-key' \
|
||||||
|
--qdrant.url 'https://your-instance.cloud.qdrant.io:6334' \
|
||||||
|
--qdrant.api-key 'your-qdrant-api-key' \
|
||||||
|
--qdrant.collection 'your-collection' \
|
||||||
|
--migration.create-collection false
|
||||||
|
```
|
||||||
|
|
||||||
|
<aside role="status">Note the <code>--migration.create-collection false</code> flag — since you pre-created the collection, the tool should skip auto-creation.</aside>
|
||||||
|
|
||||||
|
### All Weaviate-Specific Flags
|
||||||
|
|
||||||
|
| Flag | Required | Description |
|
||||||
|
| :--- | :--- | :--- |
|
||||||
|
| `--weaviate.host` | Yes | Weaviate host address |
|
||||||
|
| `--weaviate.scheme` | No | `http` or `https` (default: `http`) |
|
||||||
|
| `--weaviate.class-name` | Yes | Weaviate class to migrate |
|
||||||
|
| `--weaviate.auth-type` | No | `none`, `apiKey`, `password`, `client`, or `bearer` |
|
||||||
|
| `--weaviate.api-key` | No | API key (when auth-type is `apiKey`) |
|
||||||
|
| `--weaviate.username` | No | Username (when auth-type is `password`) |
|
||||||
|
| `--weaviate.password` | No | Password (when auth-type is `password`) |
|
||||||
|
| `--weaviate.tenant` | No | Specific tenant to migrate |
|
||||||
|
|
||||||
|
## Gotchas
|
||||||
|
|
||||||
|
- **Vector dimensions not exposed:** Always pre-create the Qdrant collection. If you don't know your dimensions, check a sample vector from Weaviate or your embedding model's documentation.
|
||||||
|
- **Cross-references:** Weaviate cross-references don't have a direct equivalent in Qdrant. Store referenced IDs as payload fields and rebuild the linking in your application layer.
|
||||||
|
- **Module dependencies:** If you used Weaviate vectorizer modules (e.g., `text2vec-openai`), ensure you exported the actual vectors, not just the source text.
|
||||||
|
|
||||||
|
## Next Steps
|
||||||
|
|
||||||
|
After migration, verify your data arrived correctly with the [Migration Verification Guide](/documentation/migration-verification/).
|
||||||
@@ -0,0 +1,66 @@
|
|||||||
|
---
|
||||||
|
title: Migration Verification
|
||||||
|
weight: 24
|
||||||
|
is_empty: false
|
||||||
|
partition: qdrant
|
||||||
|
---
|
||||||
|
|
||||||
|
# Migration Verification Guide
|
||||||
|
|
||||||
|
Switching databases is often necessary to improve the performance and costs of large software systems. As the size of data continues to grow, navigating this migration can be tricky. While a migration might appear to have been successful, several silent errors may be occurring. Some examples include: data loss, metadata drift, and search quality regressions. These problems can hide behind a successful import. To help mitigate these issues, this guide gives you a structured framework to verify that your migration worked correctly.
|
||||||
|
|
||||||
|
## Who This Guide Is For
|
||||||
|
|
||||||
|
Engineers and teams migrating to Qdrant from another vector search system (Pinecone, Weaviate, Milvus/Zilliz, Elasticsearch, pgvector, or any custom FAISS/ScaNN deployment). The verification steps are system-agnostic on the source side. That is, you capture baselines from your current system, then validate against Qdrant.
|
||||||
|
|
||||||
|
## Prerequisites
|
||||||
|
|
||||||
|
Before starting verification, you need:
|
||||||
|
|
||||||
|
- Access to your source system (to capture baselines)
|
||||||
|
- A Qdrant instance with your migrated data loaded
|
||||||
|
- Python 3.8+ with the `qdrant-client` library installed (code examples use Python, but the concepts apply to any client)
|
||||||
|
- The [Qdrant Migration Tool](/documentation/migrate-to-qdrant/) already run
|
||||||
|
|
||||||
|
## The Four Stages of Verification
|
||||||
|
|
||||||
|
This guide proposes four distinct stages of verification, each building on the next to ensure a successful migration:
|
||||||
|
|
||||||
|
1. **[Pre-Migration Baseline](/documentation/migration-verification/pre-migration-baseline/):** Before starting the migration, capture what "correct" looks like in your source system. This includes vector counts, metadata samples, collection configuration, and baseline search results. Without this, post-migration comparison is impossible. *Budget 15 to 30 minutes*.
|
||||||
|
2. **[Data Integrity](/documentation/migration-verification/data-integrity/):** After migration, verify that the data arrived intact. This layer catches missing vectors, dropped metadata fields, type coercion errors, and misconfigured collections. These checks take only minutes to run and catch the most common migration failures.
|
||||||
|
3. **[Search Quality](/documentation/migration-verification/search-quality/):** This layer catches result ranking changes, recall degradation, and relevance shifts caused by differences in indexing, quantization, or scoring between systems. *Depending on the tier you choose, this takes anywhere from 15 minutes to several hours.*
|
||||||
|
4. **[Diagnosing Discrepancies](/documentation/migration-verification/diagnosing-discrepancies/):** When any of the previous layers catches a problem, this section provides a decision tree for root-causing it: is the issue in the data or the configuration? Which vendor-specific gotcha is responsible? What's the fastest path to resolution?
|
||||||
|
|
||||||
|
## Quick Reference: Verification Checklist
|
||||||
|
|
||||||
|
Use this as a summary after reading the full guide.
|
||||||
|
|
||||||
|
```
|
||||||
|
PRE-MIGRATION
|
||||||
|
[ ] Capture vector count per collection/index
|
||||||
|
[ ] Export sample metadata records (at least 1,000 or 1% of data, whichever is larger)
|
||||||
|
[ ] Record collection configuration (distance metric, dimensions, index params)
|
||||||
|
[ ] Run and record baseline queries (10-50 representative queries with top-k results)
|
||||||
|
[ ] Note your source system's version, quantization settings, and index configuration
|
||||||
|
|
||||||
|
DATA INTEGRITY (post-migration)
|
||||||
|
[ ] Vector count matches source (exact or within expected tolerance)
|
||||||
|
[ ] Sample metadata spot-check passes (field names, types, values)
|
||||||
|
[ ] Collection configuration matches intent (distance metric, vector dimensions)
|
||||||
|
[ ] No orphaned or duplicate point IDs
|
||||||
|
|
||||||
|
SEARCH QUALITY
|
||||||
|
[ ] Tier 1: Spot-check 5-10 queries, results look reasonable
|
||||||
|
[ ] Tier 2: Recall@k on 50+ sampled queries meets threshold (e.g., ≥0.9)
|
||||||
|
[ ] Tier 3: (If applicable) NDCG/MRR on labeled evaluation set meets target
|
||||||
|
|
||||||
|
DISCREPANCY DIAGNOSIS (if checks fail)
|
||||||
|
[ ] Identify whether issue is data-level or configuration-level
|
||||||
|
[ ] Check distance metric alignment
|
||||||
|
[ ] Check quantization and indexing parameter differences
|
||||||
|
[ ] Check metadata type mapping
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Next:** [Pre-Migration Baseline](/documentation/migration-verification/pre-migration-baseline/)
|
||||||
@@ -0,0 +1,279 @@
|
|||||||
|
---
|
||||||
|
title: Data Integrity
|
||||||
|
weight: 20
|
||||||
|
---
|
||||||
|
|
||||||
|
# Data Integrity Verification
|
||||||
|
|
||||||
|
Once you've established a [baseline](/documentation/migration-verification/pre-migration-baseline/), you first need to check data integrity. Data integrity answers the question: "Did all my data arrive, and did it arrive correctly?" These are the fastest checks to run and catch the most common migration failures.
|
||||||
|
|
||||||
|
## 1. Vector Count Verification
|
||||||
|
|
||||||
|
The simplest check: does the number of vectors in Qdrant match your source system?
|
||||||
|
|
||||||
|
```py
|
||||||
|
from qdrant_client import QdrantClient
|
||||||
|
|
||||||
|
client = QdrantClient("localhost", port=6333)
|
||||||
|
|
||||||
|
# Get collection info
|
||||||
|
collection_info = client.get_collection("your_collection")
|
||||||
|
qdrant_count = collection_info.points_count
|
||||||
|
|
||||||
|
# Compare against baseline
|
||||||
|
source_count = baseline["total_vector_count"] # From pre-migration capture
|
||||||
|
|
||||||
|
if qdrant_count == source_count:
|
||||||
|
print(f"✓ Vector count matches: {qdrant_count}")
|
||||||
|
else:
|
||||||
|
diff = source_count - qdrant_count
|
||||||
|
pct = (diff / source_count) * 100
|
||||||
|
print(f"✗ Count mismatch: source={source_count}, qdrant={qdrant_count}, "
|
||||||
|
f"missing={diff} ({pct:.2f}%)")
|
||||||
|
```
|
||||||
|
|
||||||
|
**Common causes of count mismatches:**
|
||||||
|
|
||||||
|
| Symptom | Likely Cause |
|
||||||
|
| ----- | ----- |
|
||||||
|
| Qdrant count is lower | Migration script failed partway through; duplicate IDs in source were deduplicated; source count included soft-deleted records |
|
||||||
|
| Qdrant count is higher | Duplicate inserts from a retried migration; source count didn't include all namespaces/partitions |
|
||||||
|
| Counts match but data is wrong | ID collision: different vectors mapped to the same point ID |
|
||||||
|
|
||||||
|
**When exact match isn't expected:** Some source systems count differently. Pinecone's `describe_index_stats` counts across all namespaces; if you migrated only a subset, the counts won't match. pgvector's `n_live_tup` is an estimate. Document these expected discrepancies before concluding the migration failed.
|
||||||
|
|
||||||
|
## 2. Vector Dimension Verification
|
||||||
|
|
||||||
|
Confirm that vector dimensions match your source configuration:
|
||||||
|
|
||||||
|
```py
|
||||||
|
collection_info = client.get_collection("your_collection")
|
||||||
|
qdrant_dim = collection_info.config.params.vectors.size
|
||||||
|
# For named vectors:
|
||||||
|
# qdrant_dim = collection_info.config.params.vectors["dense"].size
|
||||||
|
|
||||||
|
source_dim = baseline["dimension"]
|
||||||
|
|
||||||
|
assert qdrant_dim == source_dim, (
|
||||||
|
f"Dimension mismatch: source={source_dim}, qdrant={qdrant_dim}"
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
**If dimensions don't match:** This almost always indicates a migration script error (e.g., truncated vectors, wrong embedding model used for re-embedding). Do not proceed with further verification until this is resolved.
|
||||||
|
|
||||||
|
## 3. Distance Metric Verification
|
||||||
|
|
||||||
|
Verify the distance metric matches your source system's configuration:
|
||||||
|
|
||||||
|
```py
|
||||||
|
qdrant_metric = collection_info.config.params.vectors.distance
|
||||||
|
# Returns: "Cosine", "Euclid", or "Dot"
|
||||||
|
|
||||||
|
# Map source system metrics to Qdrant equivalents
|
||||||
|
METRIC_MAP = {
|
||||||
|
# Pinecone
|
||||||
|
"cosine": "Cosine",
|
||||||
|
"euclidean": "Euclid",
|
||||||
|
"dotproduct": "Dot",
|
||||||
|
# Weaviate
|
||||||
|
"l2-squared": "Euclid",
|
||||||
|
# Milvus
|
||||||
|
"COSINE": "Cosine",
|
||||||
|
"L2": "Euclid",
|
||||||
|
"IP": "Dot",
|
||||||
|
}
|
||||||
|
|
||||||
|
expected_metric = METRIC_MAP.get(baseline["metric"])
|
||||||
|
assert qdrant_metric == expected_metric, (
|
||||||
|
f"Distance metric mismatch: source={baseline['metric']} "
|
||||||
|
f"(expected {expected_metric}), qdrant={qdrant_metric}"
|
||||||
|
)
|
||||||
|
```
|
||||||
|
|
||||||
|
A distance metric mismatch is a silent error. When migrating, the vectors still load, and queries still return results. For example, cosine similarity and dot product produce identical rankings only when vectors are unit-normalized. If your vectors aren't normalized and you switch between cosine and dot product, every search result changes.
|
||||||
|
|
||||||
|
## 4. Metadata (Payload) Verification
|
||||||
|
|
||||||
|
Metadata verification checks three things: field presence, field types, and field values.
|
||||||
|
|
||||||
|
### 4a. Field Presence
|
||||||
|
|
||||||
|
Check that all expected metadata fields exist in Qdrant:
|
||||||
|
|
||||||
|
```py
|
||||||
|
import random
|
||||||
|
|
||||||
|
# Sample points from Qdrant using scroll
|
||||||
|
records, _next = client.scroll(
|
||||||
|
collection_name="your_collection",
|
||||||
|
limit=1000,
|
||||||
|
with_payload=True,
|
||||||
|
with_vectors=False, # Skip vectors to speed up the check
|
||||||
|
)
|
||||||
|
|
||||||
|
# Collect all field names across sampled records
|
||||||
|
qdrant_fields = set()
|
||||||
|
for record in records:
|
||||||
|
if record.payload:
|
||||||
|
qdrant_fields.update(record.payload.keys())
|
||||||
|
|
||||||
|
source_fields = set(baseline["metadata_fields"])
|
||||||
|
missing = source_fields - qdrant_fields
|
||||||
|
extra = qdrant_fields - source_fields
|
||||||
|
|
||||||
|
if missing:
|
||||||
|
print(f"✗ Fields missing in Qdrant: {missing}")
|
||||||
|
if extra:
|
||||||
|
print(f"⚠ Extra fields in Qdrant (may be expected): {extra}")
|
||||||
|
if not missing and not extra:
|
||||||
|
print(f"✓ All {len(source_fields)} metadata fields present")
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4b. Field Type Consistency
|
||||||
|
|
||||||
|
Check that field types survived the migration:
|
||||||
|
|
||||||
|
```py
|
||||||
|
def check_field_types(source_record, qdrant_record):
|
||||||
|
"""Compare field types between source and Qdrant records."""
|
||||||
|
issues = []
|
||||||
|
for field, source_value in source_record.items():
|
||||||
|
if field not in qdrant_record:
|
||||||
|
issues.append(f" {field}: missing in Qdrant")
|
||||||
|
continue
|
||||||
|
qdrant_value = qdrant_record[field]
|
||||||
|
if type(source_value) != type(qdrant_value):
|
||||||
|
issues.append(
|
||||||
|
f" {field}: type changed from "
|
||||||
|
f"{type(source_value).__name__} to {type(qdrant_value).__name__} "
|
||||||
|
f"(source={source_value!r}, qdrant={qdrant_value!r})"
|
||||||
|
)
|
||||||
|
return issues
|
||||||
|
```
|
||||||
|
|
||||||
|
**Common type coercion issues:**
|
||||||
|
|
||||||
|
| Source Type | Qdrant Arrival | Impact |
|
||||||
|
| ----- | ----- | ----- |
|
||||||
|
| Integer → Float | `42` → `42.0` | Filter `= 42` may fail; use range filter instead |
|
||||||
|
| Boolean → String | `true` → `"true"` | Filter `= true` returns no results |
|
||||||
|
| Nested object → Flattened | `{"a": {"b": 1}}` → `{"a.b": 1}` | Nested filter syntax won't match |
|
||||||
|
| Array → Single value | `["tag1", "tag2"]` → `"tag1"` | Array containment filters break |
|
||||||
|
| Null → Missing field | `null` → (field absent) | `is_null` filter won't find it |
|
||||||
|
|
||||||
|
### 4c. Field Value Spot-Check
|
||||||
|
|
||||||
|
For your sampled records, compare actual values:
|
||||||
|
|
||||||
|
```py
|
||||||
|
def spot_check_values(source_sample, qdrant_collection, client):
|
||||||
|
"""Compare metadata values for sampled records."""
|
||||||
|
mismatches = []
|
||||||
|
|
||||||
|
for source_record in source_sample:
|
||||||
|
point_id = source_record["id"]
|
||||||
|
qdrant_points = client.retrieve(
|
||||||
|
collection_name=qdrant_collection,
|
||||||
|
ids=[point_id],
|
||||||
|
with_payload=True,
|
||||||
|
)
|
||||||
|
if not qdrant_points:
|
||||||
|
mismatches.append({"id": point_id, "issue": "Point not found in Qdrant"})
|
||||||
|
continue
|
||||||
|
|
||||||
|
qdrant_payload = qdrant_points[0].payload
|
||||||
|
for field, source_value in source_record["metadata"].items():
|
||||||
|
qdrant_value = qdrant_payload.get(field)
|
||||||
|
if source_value != qdrant_value:
|
||||||
|
mismatches.append({
|
||||||
|
"id": point_id,
|
||||||
|
"field": field,
|
||||||
|
"source": source_value,
|
||||||
|
"qdrant": qdrant_value,
|
||||||
|
})
|
||||||
|
|
||||||
|
return mismatches
|
||||||
|
```
|
||||||
|
|
||||||
|
## 5. Point ID Verification
|
||||||
|
|
||||||
|
Check for duplicate or orphaned point IDs:
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Scroll through all points and collect IDs
|
||||||
|
all_ids = []
|
||||||
|
next_offset = None
|
||||||
|
while True:
|
||||||
|
records, next_offset = client.scroll(
|
||||||
|
collection_name="your_collection",
|
||||||
|
limit=1000,
|
||||||
|
offset=next_offset,
|
||||||
|
with_payload=False,
|
||||||
|
with_vectors=False,
|
||||||
|
)
|
||||||
|
all_ids.extend([r.id for r in records])
|
||||||
|
if next_offset is None:
|
||||||
|
break
|
||||||
|
|
||||||
|
# Check for duplicates
|
||||||
|
if len(all_ids) != len(set(all_ids)):
|
||||||
|
duplicates = [id for id in all_ids if all_ids.count(id) > 1]
|
||||||
|
print(f"✗ Found {len(duplicates)} duplicate point IDs")
|
||||||
|
else:
|
||||||
|
print(f"✓ No duplicate point IDs ({len(all_ids)} unique)")
|
||||||
|
```
|
||||||
|
|
||||||
|
**Note on ID mapping:** If your source system uses string IDs and you mapped them to integer IDs (or vice versa) during migration, maintain a mapping file and verify it's consistent.
|
||||||
|
|
||||||
|
## 6. Vector Value Spot-Check
|
||||||
|
|
||||||
|
For a small sample, verify that the actual vector values match:
|
||||||
|
|
||||||
|
```py
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
def verify_vectors(source_vectors, qdrant_collection, client, tolerance=1e-6):
|
||||||
|
"""Spot-check that vector values match between source and Qdrant."""
|
||||||
|
mismatches = []
|
||||||
|
|
||||||
|
for source in source_vectors:
|
||||||
|
qdrant_points = client.retrieve(
|
||||||
|
collection_name=qdrant_collection,
|
||||||
|
ids=[source["id"]],
|
||||||
|
with_vectors=True,
|
||||||
|
)
|
||||||
|
if not qdrant_points:
|
||||||
|
mismatches.append({"id": source["id"], "issue": "not found"})
|
||||||
|
continue
|
||||||
|
|
||||||
|
source_vec = np.array(source["vector"])
|
||||||
|
qdrant_vec = np.array(qdrant_points[0].vector)
|
||||||
|
|
||||||
|
if not np.allclose(source_vec, qdrant_vec, atol=tolerance):
|
||||||
|
max_diff = np.max(np.abs(source_vec - qdrant_vec))
|
||||||
|
mismatches.append({
|
||||||
|
"id": source["id"],
|
||||||
|
"max_difference": float(max_diff),
|
||||||
|
})
|
||||||
|
|
||||||
|
return mismatches
|
||||||
|
```
|
||||||
|
|
||||||
|
**Expected tolerance:** Exact float equality (tolerance=0) is too strict if quantization is applied on either side. If you're using scalar quantization in Qdrant, expect small differences. If neither system uses quantization, values should match exactly.
|
||||||
|
|
||||||
|
## Passing Criteria
|
||||||
|
|
||||||
|
| Check | Pass | Investigate |
|
||||||
|
| ----- | ----- | ----- |
|
||||||
|
| Vector count | Exact match (or within documented tolerance) | Any unexplained difference |
|
||||||
|
| Dimensions | Exact match | Any mismatch (stop here) |
|
||||||
|
| Distance metric | Maps correctly to Qdrant equivalent | Any mismatch (stop here) |
|
||||||
|
| Metadata fields | All source fields present | Missing fields |
|
||||||
|
| Metadata types | Types preserved or intentionally converted | Unexpected type changes |
|
||||||
|
| Metadata values | Spot-check sample matches | >1% mismatch rate |
|
||||||
|
| Point IDs | No duplicates, all source IDs present | Missing or duplicate IDs |
|
||||||
|
| Vector values | Within tolerance (1e-6 without quantization) | Differences exceeding tolerance |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Next:** [Search Quality Verification](/documentation/migration-verification/search-quality/)
|
||||||
+288
@@ -0,0 +1,288 @@
|
|||||||
|
---
|
||||||
|
title: Diagnosing Discrepancies
|
||||||
|
weight: 40
|
||||||
|
---
|
||||||
|
|
||||||
|
# Diagnosing Discrepancies
|
||||||
|
|
||||||
|
When verification catches a problem, you need to determine whether it's a data issue (something went wrong during migration) or a configuration issue (the data is correct but the systems behave differently). This page provides a diagnostic decision tree and vendor-specific gotchas.
|
||||||
|
|
||||||
|
## Decision Tree
|
||||||
|
|
||||||
|
Start here when any verification check fails:
|
||||||
|
|
||||||
|
```
|
||||||
|
Is the vector count wrong?
|
||||||
|
├─ Yes → Data-level issue
|
||||||
|
│ ├─ Count lower than expected → Check migration script logs for errors,
|
||||||
|
│ │ timeouts, or partial failures. Re-run for missing segments.
|
||||||
|
│ ├─ Count higher than expected → Check for duplicate inserts (retried batches)
|
||||||
|
│ │ or source count excluding namespaces/partitions.
|
||||||
|
│ └─ Count matches but IDs differ → ID mapping error during migration.
|
||||||
|
│
|
||||||
|
└─ No (count matches) → Continue
|
||||||
|
│
|
||||||
|
Are metadata fields missing or wrong type?
|
||||||
|
├─ Yes → Payload mapping issue
|
||||||
|
│ ├─ Fields missing → Source system may omit null fields on export.
|
||||||
|
│ │ Check migration script's null handling.
|
||||||
|
│ ├─ Types changed → See "Type Coercion" section below.
|
||||||
|
│ └─ Values differ → Encoding issue (UTF-8, special characters, unicode normalization).
|
||||||
|
│
|
||||||
|
└─ No (metadata looks correct) → Continue
|
||||||
|
│
|
||||||
|
Are search results completely different?
|
||||||
|
├─ Yes → Configuration-level issue
|
||||||
|
│ ├─ Check distance metric (most common cause)
|
||||||
|
│ ├─ Check if index is built (HNSW may not be built yet on fresh data)
|
||||||
|
│ └─ Check if vectors are normalized (affects cosine vs. dot product)
|
||||||
|
│
|
||||||
|
└─ No (results overlap but differ at the margins) → Expected behavior
|
||||||
|
│
|
||||||
|
Is recall@10 below 0.85?
|
||||||
|
├─ Yes → Indexing parameter mismatch
|
||||||
|
│ ├─ Compare HNSW ef_construction and M values
|
||||||
|
│ ├─ Compare ef (search-time) parameters
|
||||||
|
│ └─ Check quantization settings
|
||||||
|
│
|
||||||
|
└─ No → Migration is working correctly.
|
||||||
|
Results differ on borderline cases due to
|
||||||
|
ANN approximation. This is normal.
|
||||||
|
```
|
||||||
|
|
||||||
|
## Configuration-Level Issues
|
||||||
|
|
||||||
|
### Distance Metric Mismatch
|
||||||
|
|
||||||
|
The most impactful configuration error. Here's how metrics map across systems:
|
||||||
|
|
||||||
|
| Source System | Source Metric | Qdrant Equivalent | Notes |
|
||||||
|
| ----- | ----- | ----- | ----- |
|
||||||
|
| Pinecone | `cosine` | `Cosine` | Direct mapping |
|
||||||
|
| Pinecone | `dotproduct` | `Dot` | Pinecone requires unit-normalized vectors for dotproduct |
|
||||||
|
| Pinecone | `euclidean` | `Euclid` | Direct mapping |
|
||||||
|
| Weaviate | `cosine` | `Cosine` | Direct mapping |
|
||||||
|
| Weaviate | `l2-squared` | `Euclid` | Qdrant uses L2, not L2-squared; scores will differ in magnitude but ranking is identical |
|
||||||
|
| Weaviate | `dot` | `Dot` | Direct mapping |
|
||||||
|
| Milvus | `COSINE` | `Cosine` | Direct mapping |
|
||||||
|
| Milvus | `L2` | `Euclid` | Direct mapping |
|
||||||
|
| Milvus | `IP` (inner product) | `Dot` | Direct mapping |
|
||||||
|
| Elasticsearch | `cosine` | `Cosine` | ES returns `1 - cosine_distance`; Qdrant returns cosine similarity directly |
|
||||||
|
| pgvector | `vector_cosine_ops` | `Cosine` | pgvector returns distance (1 - similarity); Qdrant returns similarity |
|
||||||
|
| pgvector | `vector_l2_ops` | `Euclid` | Direct mapping |
|
||||||
|
| pgvector | `vector_ip_ops` | `Dot` | pgvector uses negative inner product for ordering; scores will be inverted |
|
||||||
|
|
||||||
|
**Diagnostic test:** Take a single query vector, compute its distance to a known target vector manually (using numpy), and compare against both systems:
|
||||||
|
|
||||||
|
```py
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
query = np.array([...]) # Your query vector
|
||||||
|
target = np.array([...]) # A known result vector
|
||||||
|
|
||||||
|
# Manual distance calculations
|
||||||
|
cosine_sim = np.dot(query, target) / (np.linalg.norm(query) * np.linalg.norm(target))
|
||||||
|
dot_product = np.dot(query, target)
|
||||||
|
euclidean = np.linalg.norm(query - target)
|
||||||
|
|
||||||
|
print(f"Cosine similarity: {cosine_sim:.6f}")
|
||||||
|
print(f"Dot product: {dot_product:.6f}")
|
||||||
|
print(f"Euclidean distance: {euclidean:.6f}")
|
||||||
|
|
||||||
|
# Compare against Qdrant's reported score
|
||||||
|
qdrant_result = client.query_points(
|
||||||
|
collection_name="your_collection",
|
||||||
|
query=query.tolist(),
|
||||||
|
limit=1,
|
||||||
|
)
|
||||||
|
print(f"Qdrant score: {qdrant_result.points[0].score:.6f}")
|
||||||
|
|
||||||
|
# The Qdrant score should match one of the manual calculations.
|
||||||
|
# If it doesn't match the expected metric, the collection is misconfigured.
|
||||||
|
```
|
||||||
|
|
||||||
|
### HNSW Index Not Built
|
||||||
|
|
||||||
|
On a freshly migrated collection, the HNSW index may still be building. During this period, Qdrant falls back to brute-force search, which returns exact results (recall = 1.0). Once the index finishes building, results shift to approximate.
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Check index status
|
||||||
|
collection_info = client.get_collection("your_collection")
|
||||||
|
print(f"Indexed vectors: {collection_info.indexed_vectors_count}")
|
||||||
|
print(f"Total vectors: {collection_info.points_count}")
|
||||||
|
|
||||||
|
if collection_info.indexed_vectors_count < collection_info.points_count:
|
||||||
|
print("⚠ Index is still building. Wait for completion before running search quality checks.")
|
||||||
|
```
|
||||||
|
|
||||||
|
**Gotcha:** If you run Tier 2 verification while the index is building, you'll get artificially high recall (brute-force is exact). Re-run after indexing completes to get the real numbers.
|
||||||
|
|
||||||
|
### Vector Normalization
|
||||||
|
|
||||||
|
Cosine similarity and dot product produce identical rankings when vectors are unit-normalized (L2 norm = 1.0). If your source system assumed normalized vectors and you switch to dot product (or vice versa) during migration, results will differ.
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Check if vectors are normalized
|
||||||
|
sample_points = client.scroll(
|
||||||
|
collection_name="your_collection",
|
||||||
|
limit=100,
|
||||||
|
with_vectors=True,
|
||||||
|
)[0]
|
||||||
|
|
||||||
|
norms = [np.linalg.norm(p.vector) for p in sample_points]
|
||||||
|
print(f"Vector norms: min={min(norms):.4f}, max={max(norms):.4f}, mean={np.mean(norms):.4f}")
|
||||||
|
|
||||||
|
if all(abs(n - 1.0) < 0.001 for n in norms):
|
||||||
|
print("Vectors are unit-normalized. Cosine and Dot produce equivalent rankings.")
|
||||||
|
else:
|
||||||
|
print("Vectors are NOT normalized. Cosine and Dot will produce different rankings.")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Quantization Differences
|
||||||
|
|
||||||
|
If your source system uses one quantization scheme and Qdrant uses another (or none), scores will differ. This is expected and doesn't indicate data corruption.
|
||||||
|
|
||||||
|
| Source Quantization | Qdrant Quantization | Expected Impact |
|
||||||
|
| ----- | ----- | ----- |
|
||||||
|
| None | None | Scores should match closely |
|
||||||
|
| None | Scalar (int8) | Small score differences, recall may change by 1-2% |
|
||||||
|
| None | Product Quantization | Larger score differences, recall may drop 2-5% (tune `rescore` to compensate) |
|
||||||
|
| PQ | None | Qdrant results will be more accurate than source |
|
||||||
|
| PQ | PQ | Scores will differ (different codebooks), but recall should be comparable |
|
||||||
|
|
||||||
|
## Data-Level Issues
|
||||||
|
|
||||||
|
### Partial Migration Failures
|
||||||
|
|
||||||
|
The most common data-level issue: a batch upload timed out or errored, and the migration script didn't retry.
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Find missing IDs by comparing source and Qdrant
|
||||||
|
all_ids = set()
|
||||||
|
offset = None
|
||||||
|
while True:
|
||||||
|
records, offset = client.scroll(
|
||||||
|
collection_name="your_collection",
|
||||||
|
limit=1000,
|
||||||
|
offset=offset,
|
||||||
|
with_payload=False,
|
||||||
|
with_vectors=False,
|
||||||
|
)
|
||||||
|
all_ids.update(r.id for r in records)
|
||||||
|
if offset is None:
|
||||||
|
break
|
||||||
|
|
||||||
|
# Compare against source IDs
|
||||||
|
source_ids = set(baseline["all_ids"]) # Or load from your mapping file
|
||||||
|
missing = source_ids - all_ids
|
||||||
|
if missing:
|
||||||
|
print(f"Missing {len(missing)} IDs. First 10: {list(missing)[:10]}")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Type Coercion Problems
|
||||||
|
|
||||||
|
When metadata types change during migration, filtered search breaks silently. The filter executes without error but matches zero documents.
|
||||||
|
|
||||||
|
**Debugging approach:**
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Verify what types Qdrant stored
|
||||||
|
sample = client.scroll(
|
||||||
|
collection_name="your_collection",
|
||||||
|
limit=1,
|
||||||
|
with_payload=True,
|
||||||
|
)[0][0]
|
||||||
|
|
||||||
|
for field, value in sample.payload.items():
|
||||||
|
print(f" {field}: {type(value).__name__} = {value!r}")
|
||||||
|
```
|
||||||
|
|
||||||
|
**Common fixes:**
|
||||||
|
|
||||||
|
| Problem | Fix |
|
||||||
|
| ----- | ----- |
|
||||||
|
| Integer stored as float | Use range filter (`gte`/`lte`) instead of exact match, or re-upload with explicit int casting |
|
||||||
|
| Boolean stored as string | Re-upload the affected payload field with `client.set_payload()` |
|
||||||
|
| Array flattened to single value | Re-upload; check your migration script's array handling |
|
||||||
|
| Nested object lost structure | Re-upload with correct nesting; Qdrant supports nested payloads |
|
||||||
|
|
||||||
|
### Encoding and Unicode Issues
|
||||||
|
|
||||||
|
Metadata strings with non-ASCII characters, emoji, or special Unicode can be mangled during migration if encoding isn't handled consistently.
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Spot-check strings with non-ASCII content
|
||||||
|
import unicodedata
|
||||||
|
|
||||||
|
for record in sample_records:
|
||||||
|
for field, value in record.payload.items():
|
||||||
|
if isinstance(value, str) and not value.isascii():
|
||||||
|
# Check for common encoding issues
|
||||||
|
try:
|
||||||
|
value.encode("utf-8").decode("utf-8")
|
||||||
|
except UnicodeError:
|
||||||
|
print(f" Encoding issue: {field} in record {record.id}")
|
||||||
|
```
|
||||||
|
|
||||||
|
## Vendor-Specific Gotchas
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>From Pinecone</b></summary>
|
||||||
|
|
||||||
|
* **Namespace handling:** Pinecone namespaces don't have a direct Qdrant equivalent. Common approach: migrate each namespace as a separate collection, or merge into one collection with a `namespace` payload field. Verify your approach preserved the separation correctly.
|
||||||
|
* **Metadata size limits:** Pinecone limits metadata to 40KB per vector. Qdrant has no per-payload size limit, so this shouldn't cause issues. But if your migration script truncated metadata to fit Pinecone's limit, the truncated version is what you're migrating.
|
||||||
|
* **Score scaling:** Pinecone cosine similarity returns values in [0, 1] (rescaled). Qdrant returns [-1, 1]. Rankings are identical, but raw scores won't match.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>From Weaviate</b></summary>
|
||||||
|
|
||||||
|
* **GraphQL to REST:** Weaviate's GraphQL query model is structurally different from Qdrant's REST/gRPC API. Filter translation is the most error-prone step. Verify each filter type (string match, numeric range, boolean, array containment) individually.
|
||||||
|
* **Cross-references:** Weaviate cross-references don't have a direct equivalent. Store referenced IDs as payload fields and rebuild the linking in your application layer.
|
||||||
|
* **Module dependencies:** If you used Weaviate modules (e.g., `text2vec-openai`), the vectorization happened server-side. Ensure you exported the actual vectors, not the source text alone.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>From Milvus / Zilliz</b></summary>
|
||||||
|
|
||||||
|
* **Schema strictness:** Milvus enforces schema on write; Qdrant is schema-flexible. Verify that schema-less flexibility didn't cause payload fields to drift during migration.
|
||||||
|
* **Partition mapping:** Milvus partitions can map to Qdrant collections or payload filters. Verify the mapping preserved query isolation.
|
||||||
|
* **Dynamic fields:** Milvus dynamic fields (introduced in 2.3) may serialize differently. Check that JSON-typed dynamic fields survived the migration with correct structure.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>From Elasticsearch</b></summary>
|
||||||
|
|
||||||
|
* **BM25 + vector hybrid:** If your ES setup used hybrid BM25 + kNN scoring, you'll need to reconstruct this in Qdrant using sparse vectors (for BM25-like behavior) alongside dense vectors. The scores won't match 1:1 because the ranking models are different.
|
||||||
|
* **Nested documents:** ES nested documents need to be flattened or restructured for Qdrant's payload model.
|
||||||
|
* **Score normalization:** ES `_score` values are not comparable to Qdrant scores. Don't use raw score comparison; use rank-based metrics (recall@k, Spearman correlation).
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>From pgvector</b></summary>
|
||||||
|
|
||||||
|
* **Partition structure:** If you had manual partitions in pgvector (common at scale), verify that all partitions were migrated, not just the primary table.
|
||||||
|
* **NULL handling:** PostgreSQL NULLs may be dropped during export. Check that optional fields are represented correctly in Qdrant payloads.
|
||||||
|
* **Index type:** pgvector supports IVFFlat and HNSW. The index type affects which results you captured in your baseline. If your baseline was captured with IVFFlat (lower recall), Qdrant's HNSW may return better results. This looks like a "mismatch" but is an improvement.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
## When to Re-Migrate vs. Adjust Configuration
|
||||||
|
|
||||||
|
| Diagnosis | Action |
|
||||||
|
| ----- | ----- |
|
||||||
|
| Distance metric wrong | Re-create collection with correct metric; re-upload vectors |
|
||||||
|
| HNSW parameters suboptimal | Adjust parameters and wait for re-indexing (no re-upload needed) |
|
||||||
|
| Missing vectors | Re-run migration for missing batches only (use upsert) |
|
||||||
|
| Metadata types wrong | Use `set_payload` to fix affected fields (no vector re-upload needed) |
|
||||||
|
| Payload fields missing | Use `set_payload` to add missing fields from source export |
|
||||||
|
| Quantization causing recall drop | Adjust quantization settings or enable rescoring |
|
||||||
|
| Everything checks out but "feels wrong" | Build Tier 3 evaluation data. "Feels wrong" without metrics isn't actionable. |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Previous:** [Search Quality Verification](/documentation/migration-verification/search-quality/) | **Start:** [Migration Verification Overview](/documentation/migration-verification/)
|
||||||
+238
@@ -0,0 +1,238 @@
|
|||||||
|
---
|
||||||
|
title: Pre-Migration Baseline
|
||||||
|
weight: 10
|
||||||
|
---
|
||||||
|
|
||||||
|
# Pre-Migration Baseline
|
||||||
|
|
||||||
|
Establishing a baseline is paramount for migration verification. If you don't capture what "correct" looks like before you migrate, you have nothing to compare against afterward. This page covers what to record from your source system before starting the migration.
|
||||||
|
|
||||||
|
## What to Capture
|
||||||
|
|
||||||
|
There are four pieces of information that need to be accounted for when establishing a baseline: collection/index inventory, metadata samples, baseline search results, and system configuration snapshots.
|
||||||
|
|
||||||
|
### 1. Collection/Index Inventory
|
||||||
|
|
||||||
|
For every index/collection you plan to migrate, record the following information:
|
||||||
|
|
||||||
|
```
|
||||||
|
For each collection:
|
||||||
|
- Name / identifier
|
||||||
|
- Vector count
|
||||||
|
- Vector dimensions
|
||||||
|
- Distance metric (cosine, dot product, euclidean)
|
||||||
|
- Index type and parameters (e.g., HNSW ef_construction, M)
|
||||||
|
- Quantization settings (if any)
|
||||||
|
- Replication factor (if applicable)
|
||||||
|
```
|
||||||
|
|
||||||
|
Pay close attention to the distance metric. Distance metric mismatches are the single most common cause of search quality regressions after migration. Cosine similarity vs. dot product vs. Euclidean distance will produce different rankings from the same vectors. If your source system uses cosine and you accidentally configure Qdrant for dot product, every search result changes.
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>Pinecone</b></summary>
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Pinecone baseline capture
|
||||||
|
import pinecone
|
||||||
|
|
||||||
|
# Record index stats
|
||||||
|
index = pinecone.Index("your-index")
|
||||||
|
stats = index.describe_index_stats()
|
||||||
|
|
||||||
|
baseline = {
|
||||||
|
"total_vector_count": stats.total_vector_count,
|
||||||
|
"dimension": stats.dimension,
|
||||||
|
"namespaces": {
|
||||||
|
ns: {"vector_count": ns_stats.vector_count}
|
||||||
|
for ns, ns_stats in stats.namespaces.items()
|
||||||
|
},
|
||||||
|
# Pinecone doesn't expose distance metric via API;
|
||||||
|
# check your index creation code or dashboard
|
||||||
|
"metric": "cosine", # VERIFY THIS MANUALLY
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>Weaviate</b></summary>
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Weaviate baseline capture
|
||||||
|
import weaviate
|
||||||
|
|
||||||
|
client = weaviate.Client("http://localhost:8080")
|
||||||
|
|
||||||
|
schema = client.schema.get()
|
||||||
|
for cls in schema["classes"]:
|
||||||
|
baseline = {
|
||||||
|
"class_name": cls["class"],
|
||||||
|
"vector_count": client.query.aggregate(cls["class"]).with_meta_count().do(),
|
||||||
|
"distance_metric": cls.get("vectorIndexConfig", {}).get("distance", "cosine"),
|
||||||
|
"ef_construction": cls.get("vectorIndexConfig", {}).get("efConstruction"),
|
||||||
|
"vector_dimensions": None, # Weaviate infers from data; check a sample vector
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>Milvus / Zilliz</b></summary>
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Milvus baseline capture
|
||||||
|
from pymilvus import connections, Collection
|
||||||
|
|
||||||
|
connections.connect("default", host="localhost", port="19530")
|
||||||
|
collection = Collection("your_collection")
|
||||||
|
collection.load()
|
||||||
|
|
||||||
|
baseline = {
|
||||||
|
"collection_name": collection.name,
|
||||||
|
"num_entities": collection.num_entities,
|
||||||
|
"schema_fields": [
|
||||||
|
{"name": f.name, "dtype": str(f.dtype), "dim": getattr(f, "dim", None)}
|
||||||
|
for f in collection.schema.fields
|
||||||
|
],
|
||||||
|
"index_params": collection.indexes, # Capture index type + params
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>Elasticsearch</b></summary>
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Elasticsearch baseline capture
|
||||||
|
from elasticsearch import Elasticsearch
|
||||||
|
|
||||||
|
es = Elasticsearch("http://localhost:9200")
|
||||||
|
|
||||||
|
# Get mapping to find vector field config
|
||||||
|
mapping = es.indices.get_mapping(index="your_index")
|
||||||
|
stats = es.count(index="your_index")
|
||||||
|
|
||||||
|
baseline = {
|
||||||
|
"index_name": "your_index",
|
||||||
|
"document_count": stats["count"],
|
||||||
|
"mapping": mapping, # Contains vector field type, dims, similarity metric
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>pgvector</b></summary>
|
||||||
|
|
||||||
|
```sql
|
||||||
|
-- pgvector baseline capture
|
||||||
|
SELECT
|
||||||
|
relname AS table_name,
|
||||||
|
n_live_tup AS approximate_row_count
|
||||||
|
FROM pg_stat_user_tables
|
||||||
|
WHERE relname = 'your_embeddings_table';
|
||||||
|
|
||||||
|
-- Vector dimensions (check first row)
|
||||||
|
SELECT vector_dims(embedding) FROM your_embeddings_table LIMIT 1;
|
||||||
|
|
||||||
|
-- Index configuration
|
||||||
|
SELECT indexname, indexdef
|
||||||
|
FROM pg_indexes
|
||||||
|
WHERE tablename = 'your_embeddings_table';
|
||||||
|
|
||||||
|
-- Distance metric: check your index definition
|
||||||
|
-- ivfflat with vector_cosine_ops = cosine
|
||||||
|
-- ivfflat with vector_l2_ops = euclidean
|
||||||
|
-- ivfflat with vector_ip_ops = inner product (dot)
|
||||||
|
```
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
### 2. Metadata Sample
|
||||||
|
|
||||||
|
Export a representative sample of metadata (or payloads) from your source system. You'll use this for field-by-field comparison after migration.
|
||||||
|
|
||||||
|
**How much to sample:** At least 1,000 records or 1% of your data, *whichever is larger*. For datasets under 100K vectors, consider exporting all metadata.
|
||||||
|
|
||||||
|
**What to record for each sample:**
|
||||||
|
|
||||||
|
```
|
||||||
|
- Point/document ID
|
||||||
|
- All metadata fields with their values
|
||||||
|
- Metadata field types (string, integer, float, boolean, array, nested object)
|
||||||
|
- Any null/missing fields (important: some systems drop nulls on export)
|
||||||
|
```
|
||||||
|
|
||||||
|
Metadata type coercion is a subtle migration failure. A field stored as an integer in Pinecone might arrive as a float in Qdrant. A boolean stored as `"true"` (string) in Elasticsearch will need explicit type handling. These mismatches don't cause errors during import, but they break filtered search queries.
|
||||||
|
|
||||||
|
### 3. Baseline Search Queries
|
||||||
|
|
||||||
|
The most valuable baseline you can capture is search quality. Select 10 to 50 queries that represent your actual search workload:
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Structure for recording baseline queries
|
||||||
|
baseline_queries = [
|
||||||
|
{
|
||||||
|
"query_id": "q001",
|
||||||
|
"description": "Product search: running shoes",
|
||||||
|
"query_vector": [...], # The actual query vector
|
||||||
|
"filters": {"category": "footwear", "in_stock": True}, # If applicable
|
||||||
|
"top_k": 10,
|
||||||
|
"source_results": [
|
||||||
|
{"id": "doc_123", "score": 0.95, "rank": 1},
|
||||||
|
{"id": "doc_456", "score": 0.91, "rank": 2},
|
||||||
|
# ... full top-k
|
||||||
|
],
|
||||||
|
"timestamp": "2026-03-10T14:30:00Z",
|
||||||
|
"source_system": "pinecone",
|
||||||
|
"source_index": "products-v2",
|
||||||
|
},
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
**How to choose representative queries:**
|
||||||
|
|
||||||
|
* Include your most frequent production queries (check logs)
|
||||||
|
* Include edge cases: queries with highly selective filters, queries that return few results, queries across multiple data types
|
||||||
|
* Include queries from different parts of the vector space (test across multiple clusters, not a single region of similar queries)
|
||||||
|
* If you use hybrid search (dense + sparse), capture both components
|
||||||
|
|
||||||
|
**What to record for each query:**
|
||||||
|
|
||||||
|
* The query vector itself (exact floats, not re-embedded)
|
||||||
|
* Any metadata filters applied
|
||||||
|
* The top-k value used
|
||||||
|
* The full ranked result list with scores
|
||||||
|
* Whether re-ranking was applied
|
||||||
|
|
||||||
|
### 4. System Configuration Snapshot
|
||||||
|
|
||||||
|
Record the configuration of your source system that affects search behavior:
|
||||||
|
|
||||||
|
```
|
||||||
|
- Software version (e.g., Pinecone API version, Weaviate 1.24, Milvus 2.3)
|
||||||
|
- Index/collection creation parameters
|
||||||
|
- Quantization settings (PQ, SQ, none)
|
||||||
|
- HNSW parameters (ef_construction, M, ef_search) if applicable
|
||||||
|
- Segment/shard configuration
|
||||||
|
- Any custom scoring, re-ranking, or post-processing logic
|
||||||
|
- Client library version
|
||||||
|
```
|
||||||
|
|
||||||
|
When search results differ post-migration, you need to determine whether the difference comes from the data or the configuration. Without a configuration snapshot, you can't distinguish between "the vectors migrated incorrectly" and "the indexing parameters produce different recall characteristics."
|
||||||
|
|
||||||
|
## Output
|
||||||
|
|
||||||
|
After completing this step, you should have four artifacts:
|
||||||
|
|
||||||
|
1. **Collection inventory** (JSON or YAML): names, counts, dimensions, metrics, index params
|
||||||
|
2. **Metadata sample** (JSONL): representative records with all fields and types
|
||||||
|
3. **Baseline queries** (JSON): query vectors, filters, and source system results
|
||||||
|
4. **Configuration snapshot** (text): source system settings that affect search behavior
|
||||||
|
|
||||||
|
Store these alongside your migration scripts. You'll reference them in every subsequent verification step.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Next:** [Data Integrity Verification](/documentation/migration-verification/data-integrity/)
|
||||||
@@ -0,0 +1,387 @@
|
|||||||
|
---
|
||||||
|
title: Search Quality
|
||||||
|
weight: 30
|
||||||
|
---
|
||||||
|
|
||||||
|
# Search Quality Verification
|
||||||
|
|
||||||
|
Two systems can hold identical vectors and produce different search results because of differences in indexing, quantization, scoring, and filtering implementation.
|
||||||
|
|
||||||
|
This is perhaps the hardest part of migration verification. The guide breaks it into **three tiers** so you can pick the level of rigor that matches your resources and risk tolerance.
|
||||||
|
|
||||||
|
## Three-Tiered Search Quality Checks
|
||||||
|
|
||||||
|
| Tier | Effort | What It Catches | When to Use |
|
||||||
|
| ----- | ----- | ----- | ----- |
|
||||||
|
| **Tier 1: Spot-Check** | 15 min | Gross failures: wrong metric, broken filters, obviously wrong results | Every migration |
|
||||||
|
| **Tier 2: Statistical Sampling** | 1-2 hours | Systematic recall degradation, filter interaction bugs, score distribution shifts | Production workloads, >100K vectors |
|
||||||
|
| **Tier 3: Gold-Standard Evaluation** | Half day to days | Measurable relevance changes with confidence intervals | High-stakes search (revenue, safety), regulated industries |
|
||||||
|
|
||||||
|
**Our recommendation:** Every migration should run Tier 1 and Tier 2. Tier 3 is for teams that have (or can build) labeled evaluation data. If you don't have labeled data today, Tier 2 gives you a strong quantitative baseline and this guide shows you how to build toward Tier 3 over time.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tier 1: Spot-Check (Every Migration)
|
||||||
|
|
||||||
|
Run your [baseline queries](/documentation/migration-verification/pre-migration-baseline/) against Qdrant and eyeball the results. This catches configuration-level errors that would affect every query: wrong distance metric, missing index, broken filter logic.
|
||||||
|
|
||||||
|
```py
|
||||||
|
from qdrant_client import QdrantClient, models
|
||||||
|
|
||||||
|
client = QdrantClient("localhost", port=6333)
|
||||||
|
|
||||||
|
def run_baseline_queries(baseline_queries, collection_name):
|
||||||
|
"""Run pre-recorded baseline queries against Qdrant."""
|
||||||
|
results = []
|
||||||
|
for bq in baseline_queries:
|
||||||
|
qdrant_results = client.query_points(
|
||||||
|
collection_name=collection_name,
|
||||||
|
query=bq["query_vector"],
|
||||||
|
limit=bq["top_k"],
|
||||||
|
query_filter=build_qdrant_filter(bq["filters"]) if bq.get("filters") else None,
|
||||||
|
)
|
||||||
|
|
||||||
|
results.append({
|
||||||
|
"query_id": bq["query_id"],
|
||||||
|
"description": bq["description"],
|
||||||
|
"source_results": bq["source_results"],
|
||||||
|
"qdrant_results": [
|
||||||
|
{"id": hit.id, "score": hit.score, "rank": i + 1}
|
||||||
|
for i, hit in enumerate(qdrant_results.points)
|
||||||
|
],
|
||||||
|
})
|
||||||
|
return results
|
||||||
|
```
|
||||||
|
|
||||||
|
### What to Look For
|
||||||
|
|
||||||
|
For each query, compare the source results against Qdrant results:
|
||||||
|
|
||||||
|
```py
|
||||||
|
def tier1_report(comparison_results):
|
||||||
|
"""Generate a human-readable spot-check report."""
|
||||||
|
for result in comparison_results:
|
||||||
|
source_ids = [r["id"] for r in result["source_results"]]
|
||||||
|
qdrant_ids = [r["id"] for r in result["qdrant_results"]]
|
||||||
|
|
||||||
|
overlap = set(source_ids) & set(qdrant_ids)
|
||||||
|
overlap_pct = len(overlap) / len(source_ids) * 100
|
||||||
|
|
||||||
|
print(f"\nQuery: {result['description']} ({result['query_id']})")
|
||||||
|
print(f" Result overlap: {len(overlap)}/{len(source_ids)} ({overlap_pct:.0f}%)")
|
||||||
|
|
||||||
|
# Check if top result matches
|
||||||
|
if source_ids and qdrant_ids:
|
||||||
|
if source_ids[0] == qdrant_ids[0]:
|
||||||
|
print(f" Top result: ✓ matches")
|
||||||
|
else:
|
||||||
|
print(f" Top result: ✗ differs "
|
||||||
|
f"(source={source_ids[0]}, qdrant={qdrant_ids[0]})")
|
||||||
|
|
||||||
|
# Check score distribution
|
||||||
|
if result["qdrant_results"]:
|
||||||
|
scores = [r["score"] for r in result["qdrant_results"]]
|
||||||
|
print(f" Score range: {min(scores):.4f} to {max(scores):.4f}")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Tier 1 Pass/Fail Criteria
|
||||||
|
|
||||||
|
* **Top-1 match rate ≥80%:** 8 out of 10 queries return the same top result
|
||||||
|
* **Top-10 overlap ≥70%:** At least 7 of the same documents appear in the top 10 (order may differ)
|
||||||
|
* **No empty results:** If a query returned results on the source, it should return results on Qdrant
|
||||||
|
* **Score range is reasonable:** Cosine similarity scores should be between -1 and 1; dot product scores vary by vector magnitude
|
||||||
|
|
||||||
|
**If Tier 1 fails:** Stop. The issue is almost certainly a configuration problem (distance metric, missing index, filter translation error). Go to [Diagnosing Discrepancies](/documentation/migration-verification/diagnosing-discrepancies/) before running further checks.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tier 2: Statistical Sampling (Recommended)
|
||||||
|
|
||||||
|
Tier 2 quantifies search quality using recall@k: the fraction of source system results that appear in Qdrant's results. Instead of eyeballing 10 queries, you measure recall across 50 or more and compute statistics.
|
||||||
|
|
||||||
|
### Recall@k
|
||||||
|
|
||||||
|
Recall@k measures: "Of the top-k results from the source system, what fraction also appears in Qdrant's top-k?"
|
||||||
|
|
||||||
|
```py
|
||||||
|
def recall_at_k(source_results, qdrant_results, k):
|
||||||
|
"""Compute recall@k: fraction of source top-k present in Qdrant top-k."""
|
||||||
|
source_ids = set(r["id"] for r in source_results[:k])
|
||||||
|
qdrant_ids = set(r["id"] for r in qdrant_results[:k])
|
||||||
|
if not source_ids:
|
||||||
|
return 1.0 # No source results = vacuously correct
|
||||||
|
return len(source_ids & qdrant_ids) / len(source_ids)
|
||||||
|
```
|
||||||
|
|
||||||
|
### Running the Evaluation
|
||||||
|
|
||||||
|
```py
|
||||||
|
import numpy as np
|
||||||
|
import json
|
||||||
|
|
||||||
|
def tier2_evaluation(baseline_queries, collection_name, client, k=10):
|
||||||
|
"""Run Tier 2 recall evaluation across all baseline queries."""
|
||||||
|
recalls = []
|
||||||
|
|
||||||
|
for bq in baseline_queries:
|
||||||
|
qdrant_results = client.query_points(
|
||||||
|
collection_name=collection_name,
|
||||||
|
query=bq["query_vector"],
|
||||||
|
limit=k,
|
||||||
|
query_filter=build_qdrant_filter(bq["filters"]) if bq.get("filters") else None,
|
||||||
|
)
|
||||||
|
|
||||||
|
qdrant_ranked = [
|
||||||
|
{"id": hit.id, "score": hit.score}
|
||||||
|
for hit in qdrant_results.points
|
||||||
|
]
|
||||||
|
|
||||||
|
r_at_k = recall_at_k(bq["source_results"], qdrant_ranked, k)
|
||||||
|
recalls.append({
|
||||||
|
"query_id": bq["query_id"],
|
||||||
|
"recall_at_k": r_at_k,
|
||||||
|
})
|
||||||
|
|
||||||
|
# Compute aggregate statistics
|
||||||
|
recall_values = [r["recall_at_k"] for r in recalls]
|
||||||
|
stats = {
|
||||||
|
"num_queries": len(recalls),
|
||||||
|
"k": k,
|
||||||
|
"mean_recall": float(np.mean(recall_values)),
|
||||||
|
"median_recall": float(np.median(recall_values)),
|
||||||
|
"min_recall": float(np.min(recall_values)),
|
||||||
|
"p5_recall": float(np.percentile(recall_values, 5)),
|
||||||
|
"p25_recall": float(np.percentile(recall_values, 25)),
|
||||||
|
"std_recall": float(np.std(recall_values)),
|
||||||
|
}
|
||||||
|
|
||||||
|
return recalls, stats
|
||||||
|
```
|
||||||
|
|
||||||
|
### Interpreting Tier 2 Results
|
||||||
|
|
||||||
|
```py
|
||||||
|
def tier2_report(recalls, stats):
|
||||||
|
"""Print Tier 2 evaluation summary."""
|
||||||
|
print(f"Recall@{stats['k']} across {stats['num_queries']} queries:")
|
||||||
|
print(f" Mean: {stats['mean_recall']:.3f}")
|
||||||
|
print(f" Median: {stats['median_recall']:.3f}")
|
||||||
|
print(f" Min: {stats['min_recall']:.3f}")
|
||||||
|
print(f" P5: {stats['p5_recall']:.3f}")
|
||||||
|
print(f" Std: {stats['std_recall']:.3f}")
|
||||||
|
|
||||||
|
# Flag low-recall queries for investigation
|
||||||
|
low_recall = [r for r in recalls if r["recall_at_k"] < 0.7]
|
||||||
|
if low_recall:
|
||||||
|
print(f"\n⚠ {len(low_recall)} queries with recall < 0.7:")
|
||||||
|
for r in low_recall:
|
||||||
|
print(f" {r['query_id']}: {r['recall_at_k']:.3f}")
|
||||||
|
```
|
||||||
|
|
||||||
|
### Why Recall Won't Be 1.0 (And That's OK)
|
||||||
|
|
||||||
|
Even a correct migration will often show recall@10 between 0.85 and 0.95 rather than 1.0. This isn't a bug. Here's why:
|
||||||
|
|
||||||
|
* **HNSW is approximate:** Both systems use approximate nearest neighbor algorithms. Different HNSW parameters (`ef_construction`, `M`, `ef` at search time) produce slightly different traversal paths and retrieve slightly different neighbors.
|
||||||
|
* **Index build order matters:** HNSW graph structure depends on insertion order. The same data inserted in a different order produces a different graph with different (but statistically equivalent) recall.
|
||||||
|
* **Quantization introduces noise:** If either system uses quantization, distance calculations have reduced precision. Two systems with different quantization schemes will disagree on borderline results.
|
||||||
|
* **Score ties:** When multiple vectors have nearly identical distances to the query, tie-breaking is arbitrary. The 10th and 11th results may swap between systems.
|
||||||
|
|
||||||
|
**What matters:** The distribution of recall, not individual query recall. If mean recall@10 ≥0.85 and no queries have recall <0.5, the migration is working correctly. The systems are disagreeing on borderline results, not on clear matches.
|
||||||
|
|
||||||
|
### Tier 2 Pass/Fail Criteria
|
||||||
|
|
||||||
|
| Metric | Pass | Investigate | Fail |
|
||||||
|
| ----- | ----- | ----- | ----- |
|
||||||
|
| Mean recall@10 | ≥0.85 | 0.70 to 0.85 | <0.70 |
|
||||||
|
| Median recall@10 | ≥0.90 | 0.75 to 0.90 | <0.75 |
|
||||||
|
| Min recall@10 | ≥0.50 | 0.30 to 0.50 | <0.30 |
|
||||||
|
| P5 recall@10 | ≥0.60 | 0.40 to 0.60 | <0.40 |
|
||||||
|
|
||||||
|
**If Tier 2 passes but a few queries have low recall:** This is normal. Check whether the low-recall queries involve highly selective filters or edge cases. See [Diagnosing Discrepancies](/documentation/migration-verification/diagnosing-discrepancies/).
|
||||||
|
|
||||||
|
### Extending Tier 2: Score Correlation
|
||||||
|
|
||||||
|
Beyond recall, check whether the relative ordering of scores is consistent:
|
||||||
|
|
||||||
|
```py
|
||||||
|
from scipy import stats as scipy_stats
|
||||||
|
|
||||||
|
def score_correlation(source_results, qdrant_results):
|
||||||
|
"""Compute rank correlation between source and Qdrant scores for overlapping results."""
|
||||||
|
# Find overlapping IDs
|
||||||
|
source_map = {r["id"]: r["score"] for r in source_results}
|
||||||
|
qdrant_map = {r["id"]: r["score"] for r in qdrant_results}
|
||||||
|
common_ids = set(source_map.keys()) & set(qdrant_map.keys())
|
||||||
|
|
||||||
|
if len(common_ids) < 3:
|
||||||
|
return None # Not enough overlap to compute correlation
|
||||||
|
|
||||||
|
source_scores = [source_map[id] for id in common_ids]
|
||||||
|
qdrant_scores = [qdrant_map[id] for id in common_ids]
|
||||||
|
|
||||||
|
# Spearman rank correlation (order matters more than magnitude)
|
||||||
|
correlation, p_value = scipy_stats.spearmanr(source_scores, qdrant_scores)
|
||||||
|
return {"correlation": correlation, "p_value": p_value, "n_common": len(common_ids)}
|
||||||
|
```
|
||||||
|
|
||||||
|
A Spearman correlation >0.8 across overlapping results means the ranking is preserved even if the exact scores differ (which they will, since different systems scale scores differently).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Tier 3: Gold-Standard Evaluation (When You Have Labeled Data)
|
||||||
|
|
||||||
|
Tier 3 measures whether search results are *relevant*, not merely consistent with the source system. This requires labeled relevance judgments: for a set of queries, human-verified labels indicating which documents are relevant.
|
||||||
|
|
||||||
|
### Why Tier 3 Matters
|
||||||
|
|
||||||
|
Tier 2 assumes the source system's results are the ground truth. But if you're migrating because your source system's search quality was insufficient, matching its results perfectly is the wrong goal. Tier 3 measures against what the results *should* be, not what they were.
|
||||||
|
|
||||||
|
### Building an Evaluation Set
|
||||||
|
|
||||||
|
If you don't have labeled data, here are three practical approaches to build one:
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>Approach A: Click/Conversion Logs (Lowest Effort)</b></summary>
|
||||||
|
|
||||||
|
If your application logs user interactions (clicks, purchases, bookmarks, shares), these are implicit relevance signals:
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Structure for click-based evaluation data
|
||||||
|
eval_queries = [
|
||||||
|
{
|
||||||
|
"query_id": "eval_001",
|
||||||
|
"query_vector": [...],
|
||||||
|
"filters": {...},
|
||||||
|
"relevant_docs": [
|
||||||
|
# Documents users engaged with after this query
|
||||||
|
{"id": "doc_789", "relevance": 2}, # Purchased/converted
|
||||||
|
{"id": "doc_012", "relevance": 1}, # Clicked but didn't convert
|
||||||
|
],
|
||||||
|
},
|
||||||
|
]
|
||||||
|
```
|
||||||
|
|
||||||
|
**Caveat:** Click data has position bias (users click top results more) and only captures what users saw, not what they would have found relevant. It's a reasonable starting point, not a perfect label set.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>Approach B: Expert Labeling (Moderate Effort)</b></summary>
|
||||||
|
|
||||||
|
Have domain experts label 50 to 100 queries with 20 to 50 candidate documents each:
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Labeling guidelines (customize for your domain)
|
||||||
|
RELEVANCE_SCALE = {
|
||||||
|
3: "Exactly what the user is looking for",
|
||||||
|
2: "Useful and related",
|
||||||
|
1: "Tangentially related",
|
||||||
|
0: "Not relevant",
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
**Practical tip:** Don't label from scratch. Run your baseline queries against both systems, pool the top 20 results from each, deduplicate, and have experts label the combined set.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
<details>
|
||||||
|
<summary><b>Approach C: Synthetic Evaluation (Lowest Barrier)</b></summary>
|
||||||
|
|
||||||
|
If you have structured metadata, construct queries where you know the correct answer:
|
||||||
|
|
||||||
|
```py
|
||||||
|
# Example: for a product catalog with known categories
|
||||||
|
synthetic_queries = []
|
||||||
|
for product in sample_products:
|
||||||
|
synthetic_queries.append({
|
||||||
|
"query_vector": product["embedding"],
|
||||||
|
"known_relevant": [product["id"]], # The product itself should be top-1
|
||||||
|
"known_category": product["category"],
|
||||||
|
# Top results should share the category
|
||||||
|
})
|
||||||
|
```
|
||||||
|
|
||||||
|
This doesn't measure relevance in the human sense, but it catches systematic retrieval failures.
|
||||||
|
|
||||||
|
</details>
|
||||||
|
|
||||||
|
### Computing Tier 3 Metrics
|
||||||
|
|
||||||
|
With labeled data, compute standard information retrieval metrics:
|
||||||
|
|
||||||
|
```py
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
def ndcg_at_k(retrieved_ids, relevance_map, k):
|
||||||
|
"""Normalized Discounted Cumulative Gain at k."""
|
||||||
|
dcg = 0.0
|
||||||
|
for i, doc_id in enumerate(retrieved_ids[:k]):
|
||||||
|
rel = relevance_map.get(doc_id, 0)
|
||||||
|
dcg += (2**rel - 1) / np.log2(i + 2) # i+2 because log2(1) = 0
|
||||||
|
|
||||||
|
# Ideal DCG: sort all relevance scores descending
|
||||||
|
ideal_rels = sorted(relevance_map.values(), reverse=True)[:k]
|
||||||
|
idcg = sum((2**rel - 1) / np.log2(i + 2) for i, rel in enumerate(ideal_rels))
|
||||||
|
|
||||||
|
return dcg / idcg if idcg > 0 else 0.0
|
||||||
|
|
||||||
|
def mrr(retrieved_ids, relevant_ids):
|
||||||
|
"""Mean Reciprocal Rank: how high is the first relevant result?"""
|
||||||
|
for i, doc_id in enumerate(retrieved_ids):
|
||||||
|
if doc_id in relevant_ids:
|
||||||
|
return 1.0 / (i + 1)
|
||||||
|
return 0.0
|
||||||
|
|
||||||
|
def tier3_evaluation(eval_queries, collection_name, client, k=10):
|
||||||
|
"""Full Tier 3 evaluation with NDCG and MRR."""
|
||||||
|
ndcgs = []
|
||||||
|
mrrs = []
|
||||||
|
|
||||||
|
for eq in eval_queries:
|
||||||
|
qdrant_results = client.query_points(
|
||||||
|
collection_name=collection_name,
|
||||||
|
query=eq["query_vector"],
|
||||||
|
limit=k,
|
||||||
|
query_filter=build_qdrant_filter(eq["filters"]) if eq.get("filters") else None,
|
||||||
|
)
|
||||||
|
retrieved_ids = [hit.id for hit in qdrant_results.points]
|
||||||
|
|
||||||
|
relevance_map = {d["id"]: d["relevance"] for d in eq["relevant_docs"]}
|
||||||
|
relevant_ids = set(relevance_map.keys())
|
||||||
|
|
||||||
|
ndcgs.append(ndcg_at_k(retrieved_ids, relevance_map, k))
|
||||||
|
mrrs.append(mrr(retrieved_ids, relevant_ids))
|
||||||
|
|
||||||
|
return {
|
||||||
|
"mean_ndcg": float(np.mean(ndcgs)),
|
||||||
|
"mean_mrr": float(np.mean(mrrs)),
|
||||||
|
"median_ndcg": float(np.median(ndcgs)),
|
||||||
|
"num_queries": len(eval_queries),
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Tier 3 Pass/Fail Criteria
|
||||||
|
|
||||||
|
Tier 3 targets are domain-specific. Here are starting points:
|
||||||
|
|
||||||
|
| Metric | Good | Acceptable | Investigate |
|
||||||
|
| ----- | ----- | ----- | ----- |
|
||||||
|
| NDCG@10 | ≥0.70 | 0.50 to 0.70 | <0.50 |
|
||||||
|
| MRR | ≥0.60 | 0.40 to 0.60 | <0.40 |
|
||||||
|
|
||||||
|
**The real test:** Compare Tier 3 metrics between your source system and Qdrant. If Qdrant's NDCG is equal to or higher than the source, the migration improved search quality. If it's lower, investigate whether the difference is caused by configuration (fixable) or a genuine capability gap.
|
||||||
|
|
||||||
|
### Building Toward Tier 3 Incrementally
|
||||||
|
|
||||||
|
Most teams don't have labeled evaluation data on migration day. That's fine. Here's a practical path:
|
||||||
|
|
||||||
|
1. **Day 0 (migration):** Run Tier 1 and Tier 2. You now have a quantitative baseline.
|
||||||
|
2. **Week 1:** Start logging search queries and user interactions in production.
|
||||||
|
3. **Month 1:** Use click logs to build Approach A evaluation data. Run Tier 3 against it.
|
||||||
|
4. **Quarter 1:** Have domain experts label the hardest queries (the ones where Tier 2 recall was lowest). This becomes your gold-standard set.
|
||||||
|
5. **Ongoing:** Re-run Tier 3 whenever you change embeddings, indexing parameters, or quantization settings.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Next:** [Diagnosing Discrepancies](/documentation/migration-verification/diagnosing-discrepancies/)
|
||||||
@@ -36,6 +36,13 @@ partition: qdrant
|
|||||||
|
|
||||||
{{% include "content/documentation/headless/content/tutorials/develop.md" %}}
|
{{% include "content/documentation/headless/content/tutorials/develop.md" %}}
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### Migrate to Qdrant
|
||||||
|
*Move your vectors from other databases and keep them in sync.*
|
||||||
|
|
||||||
|
{{% include "content/documentation/headless/content/tutorials/migrate.md" %}}
|
||||||
|
|
||||||
<!-- KEEP BELOW FOR REFERENCE -->
|
<!-- KEEP BELOW FOR REFERENCE -->
|
||||||
<!--
|
<!--
|
||||||
|
|
||||||
|
|||||||
@@ -89,7 +89,9 @@ When the migration is complete, you will see the new collection on Qdrant with a
|
|||||||
|
|
||||||
## Conclusion
|
## Conclusion
|
||||||
|
|
||||||
The **Qdrant Migration Tool** makes data transfer across vector database instances effortless. Whether you're moving between cloud regions, upgrading from self-hosted to Qdrant Cloud, or switching from other databases such as Pinecone, this tool saves you hours of manual effort. [Try it today](https://github.com/qdrant/migration).
|
The **Qdrant Migration Tool** makes data transfer across vector database instances effortless. Whether you're moving between cloud regions, upgrading from self-hosted to Qdrant Cloud, or switching from other databases such as Pinecone, this tool saves you hours of manual effort. [Try it today](https://github.com/qdrant/migration).
|
||||||
|
|
||||||
|
For detailed per-provider migration guides (Pinecone, Weaviate, Milvus, Elasticsearch, pgvector), see the [Migrate to Qdrant](/documentation/migrate-to-qdrant/) section. After migrating, use the [Migration Verification Guide](/documentation/migration-verification/) to confirm data integrity and search quality.
|
||||||
|
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user