use x as letter

This commit is contained in:
Ivan Pleshkov
2026-05-12 17:16:38 +02:00
parent 962ac60a8f
commit 4e5361d5b3
@@ -18,7 +18,7 @@ category: qdrant-internals
weight: -200
---
If you run production vector workloads, you already know the compression ladder in Qdrant: float32 is the baseline, **Scalar Quantization (SQ)** compresses vectors by 4× with almost no recall hit, and **Binary Quantization (BQ)** packs vectors at 16× or 32×.
If you run production vector workloads, you already know the compression ladder in Qdrant: float32 is the baseline, **Scalar Quantization (SQ)** compresses vectors by 4x with almost no recall hit, and **Binary Quantization (BQ)** packs vectors at 16x or 32x.
Qdrant 1.18 ships **TurboQuant** — a new rotation-based vector quantization method from [Google Research](https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/), with extensions that make it work on real production embeddings. Summarizing the results of benchmarks across public embedding datasets:
@@ -33,10 +33,10 @@ This article walks through what TurboQuant is, what we added on top to make it p
Before TurboQuant, Qdrant offered two primary production-grade quantization paths:
* **[Scalar Quantization (SQ)](https://qdrant.tech/articles/scalar-quantization/)** — int8 per coordinate. 4× compression. Recall is essentially indistinguishable from float32 on most embeddings. The default first step when memory matters.
* **[Binary Quantization (BQ)](https://qdrant.tech/articles/binary-quantization/)** — 1- or 2-bit storage (32× or 16× compression). Recall depends heavily on the embedding model — it works beautifully on isotropic, well-trained models.
* **[Scalar Quantization (SQ)](https://qdrant.tech/articles/scalar-quantization/)** — int8 per coordinate. 4x compression. Recall is essentially indistinguishable from float32 on most embeddings. The default first step when memory matters.
* **[Binary Quantization (BQ)](https://qdrant.tech/articles/binary-quantization/)** — 1- or 2-bit storage (32x or 16x compression). Recall depends heavily on the embedding model — it works beautifully on isotropic, well-trained models.
TurboQuant adds a new path with four operating points: 8× (4 bits/dim), 16× (2 bits/dim), ~21× (1.5 bits/dim), and 32× (1 bit/dim).
TurboQuant adds a new path with four operating points: 8x (4 bits/dim), 16x (2 bits/dim), ~21x (1.5 bits/dim), and 32x (1 bit/dim).
### Enabling TurboQuant
@@ -54,15 +54,15 @@ Recall, HNSW (`m=16`, `ef_construct=128`), on four representative datasets — [
**1. TQ 4-bit is competitive with SQ at half the storage.** On `arxiv-instructorxl` and `dbpedia-gemini` it is about 1 pp below SQ; on `dbpedia-openai-ada` and `wiki-cohere-v3` it actually *beats* SQ by up to 4.6 pp.
{{< figure src="/articles_data/turboquant/at-a-glance-4bit.svg" alt="Recall comparison: float32 baseline vs SQ (4×) vs TurboQuant 4-bit (8×) across four datasets" caption="float32 baseline, SQ (4× compression), and TurboQuant 4-bit (8× compression)." width="100%" >}}
{{< figure src="/articles_data/turboquant/at-a-glance-4bit.svg" alt="Recall comparison: float32 baseline vs SQ (4x) vs TurboQuant 4-bit (8x) across four datasets" caption="float32 baseline, SQ (4x compression), and TurboQuant 4-bit (8x compression)." width="100%" >}}
**2. TQ 2-bit beats BQ 2-bit by 11–15 pp** on these four datasets (and 9–24 pp across all ten datasets), at the same 16× storage.
**2. TQ 2-bit beats BQ 2-bit by 11–15 pp** on these four datasets (and 9–24 pp across all ten datasets), at the same 16x storage.
{{< figure src="/articles_data/turboquant/at-a-glance-2bit.svg" alt="Recall comparison at 16× compression: TurboQuant 2-bit vs Binary Quantization 2-bit" caption="At 16× compression — TurboQuant 2-bit vs Binary Quantization 2-bit." width="100%" >}}
{{< figure src="/articles_data/turboquant/at-a-glance-2bit.svg" alt="Recall comparison at 16x compression: TurboQuant 2-bit vs Binary Quantization 2-bit" caption="At 16x compression — TurboQuant 2-bit vs Binary Quantization 2-bit." width="100%" >}}
**3. TQ 1-bit beats vanilla BQ 1-bit by 9–21 pp** on these four datasets (and 9–21 pp across all ten datasets), at the same 32× storage. `BQ 1-bit` here is the vanilla 1-bit configuration (1-bit storage, 1-bit query); the asymmetric variant (8-bit query) is in the [detailed table](#detailed-benchmarks).
**3. TQ 1-bit beats vanilla BQ 1-bit by 9–21 pp** on these four datasets (and 9–21 pp across all ten datasets), at the same 32x storage. `BQ 1-bit` here is the vanilla 1-bit configuration (1-bit storage, 1-bit query); the asymmetric variant (8-bit query) is in the [detailed table](#detailed-benchmarks).
{{< figure src="/articles_data/turboquant/at-a-glance-1bit.svg" alt="Recall comparison at 32× compression: TurboQuant 1-bit vs vanilla Binary Quantization 1-bit" caption="At 32× compression — TurboQuant 1-bit vs vanilla Binary Quantization 1-bit." width="100%" >}}
{{< figure src="/articles_data/turboquant/at-a-glance-1bit.svg" alt="Recall comparison at 32x compression: TurboQuant 1-bit vs vanilla Binary Quantization 1-bit" caption="At 32x compression — TurboQuant 1-bit vs vanilla Binary Quantization 1-bit." width="100%" >}}
## What Is TurboQuant?
@@ -108,15 +108,15 @@ The rotation step gives every coordinate a roughly N(0, 1) distribution **on iso
Because Qdrant stores data in segments, we can fix this per segment. For each segment we do a single **pre-pass** before quantization: estimate a `(shift, scale)` pair per coordinate after rotation, then apply `x → (x + shift) · scale` to pull the empirical per-coordinate distribution back onto the codebook's grid. The same `(shift, scale)` is baked into the segment's metadata and reused for every query that hits the segment.
**This is free at search time** thanks to the asymmetric scoring scheme. The stored code is `x⁺ = (x + shift) · scale`, so the original vector is `x = x⁺ / scale − shift`. Plugging that into the dot product gives
* **This is free at search time** thanks to the asymmetric scoring scheme. The stored code is `x⁺ = (x + shift) · scale`, so the original vector is `x = x⁺ / scale − shift`. Plugging that into the dot product gives
`⟨q, x⟩ = ⟨q / scale, x⁺⟩ − ⟨q, shift⟩`
— the per-coordinate `1/scale` collapses into the query, and the `⟨q, shift⟩` term is a single scalar that depends only on the query. Both are computed **once per query**. The hot path still scores the raw `b·D`-bit code against a precomputed query, with one scalar added at the end; the scoring kernel does not change shape, and storage stays at exactly `b·D` bits per vector. All of the per-coordinate precision lives on the query side, where we have full float room to spend.
**Why not just mean + stddev?** Mean-and-stddev rescaling assumes the post-rotation coordinates are Gaussian — exactly the assumption that breaks on anisotropic data, which is the case where we need calibration in the first place. We anchor calibration to the codebook itself instead: the `(shift, scale)` pair is fit so the empirical quantiles at the probability levels of the **outermost codebook centroid** land at that centroid. The quantiles themselves are estimated with the [P-Square algorithm](https://www.cse.wustl.edu/~jain/papers/ftp/psqr.pdf) (Jain & Chlamtac, 1985) — streaming, no parametric fit, constant memory per coordinate.
* **Why not just mean + stddev?** Mean-and-stddev rescaling assumes the post-rotation coordinates are Gaussian — exactly the assumption that breaks on anisotropic data, which is the case where we need calibration in the first place. We anchor calibration to the codebook itself instead: the `(shift, scale)` pair is fit so the empirical quantiles at the probability levels of the **outermost codebook centroid** land at that centroid. The quantiles themselves are estimated with the [P-Square algorithm](https://www.cse.wustl.edu/~jain/papers/ftp/psqr.pdf) (Jain & Chlamtac, 1985) — streaming, no parametric fit, constant memory per coordinate.
**Sampling, not full scan.** Running P-Square over every vector in the segment increases index-build time. We instead sample a random subset of segment vectors using [Vitter's Algorithm R](https://en.wikipedia.org/wiki/Reservoir_sampling#Algorithm_R) (classical reservoir sampling), then run P-Square on the reservoir.
* **Sampling, not full scan.** Running P-Square over every vector in the segment increases index-build time. We instead sample a random subset of segment vectors using [Vitter's Algorithm R](https://en.wikipedia.org/wiki/Reservoir_sampling#Algorithm_R) (classical reservoir sampling), then run P-Square on the reservoir.
Truly isotropic data matches the theoretical Gaussian quantiles, the formula collapses to `(shift=0, scale=1)`, and the encoded vector is bit-identical to vanilla TurboQuant. So the `(shift, scale)` correction never degrades isotropic data.
@@ -169,36 +169,36 @@ Setup: HNSW index (`m=16`, `ef_construct=128`). Rows are ordered by storage clas
| Variant (compression) | arxiv-384 | arxiv-iXL | dbp-gem | dbp-3s | dbp-3l | dbp-oai | cohere | h&m | laion | ads-1M |
| --------------------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- | ---------- |
| f32 (1×) | 0.9855 | 0.9419 | 0.9167 | 0.9384 | 0.9348 | 0.9625 | 0.9446 | 0.9967 | 0.9897 | 0.9298 |
| SQ (4×) | 0.9674 | 0.9285 | 0.9134 | 0.9362 | 0.9339 | 0.8839 | 0.9014 | 0.9789 | 0.9276 | 0.9187 |
| **TQ 4-bit (8×)** | **0.9442** | **0.9193** | **0.9020** | **0.9313** | **0.9271** | **0.9299** | **0.9271** | **0.9739** | **0.9438** | **0.9169** |
| TQ 2-bit (16×) | 0.8477 | 0.8227 | 0.8170 | 0.8838 | 0.8806 | 0.8480 | 0.8303 | 0.9195 | 0.8349 | 0.8706 |
| BQ 2-bit (16×) | 0.6948 | 0.6756 | 0.6689 | 0.7630 | 0.7513 | 0.7332 | 0.6880 | 0.7018 | 0.5953 | 0.7808 |
| TQ 1.5-bit (~21×) | 0.7567 | 0.7143 | 0.7391 | 0.8278 | 0.8197 | 0.7690 | 0.6460 | 0.8756 | 0.7213 | 0.7941 |
| TQ 1-bit (32×) | 0.7127 | 0.6763 | 0.6990 | 0.7997 | 0.7924 | 0.7356 | 0.6300 | 0.8540 | 0.6807 | 0.7717 |
| BQ asymmetric (32×) | 0.7070 | 0.5919 | 0.6112 | 0.7910 | 0.7824 | 0.7072 | 0.6287 | 0.7802 | 0.5800 | 0.7570 |
| BQ 1-bit (32×) | 0.6028 | 0.4683 | 0.4945 | 0.7041 | 0.6921 | 0.6098 | 0.5409 | 0.6989 | 0.4762 | 0.6760 |
| f32 (1x) | 0.9855 | 0.9419 | 0.9167 | 0.9384 | 0.9348 | 0.9625 | 0.9446 | 0.9967 | 0.9897 | 0.9298 |
| SQ (4x) | 0.9674 | 0.9285 | 0.9134 | 0.9362 | 0.9339 | 0.8839 | 0.9014 | 0.9789 | 0.9276 | 0.9187 |
| **TQ 4-bit (8x)** | **0.9442** | **0.9193** | **0.9020** | **0.9313** | **0.9271** | **0.9299** | **0.9271** | **0.9739** | **0.9438** | **0.9169** |
| TQ 2-bit (16x) | 0.8477 | 0.8227 | 0.8170 | 0.8838 | 0.8806 | 0.8480 | 0.8303 | 0.9195 | 0.8349 | 0.8706 |
| BQ 2-bit (16x) | 0.6948 | 0.6756 | 0.6689 | 0.7630 | 0.7513 | 0.7332 | 0.6880 | 0.7018 | 0.5953 | 0.7808 |
| TQ 1.5-bit (~21x) | 0.7567 | 0.7143 | 0.7391 | 0.8278 | 0.8197 | 0.7690 | 0.6460 | 0.8756 | 0.7213 | 0.7941 |
| TQ 1-bit (32x) | 0.7127 | 0.6763 | 0.6990 | 0.7997 | 0.7924 | 0.7356 | 0.6300 | 0.8540 | 0.6807 | 0.7717 |
| BQ asymmetric (32x) | 0.7070 | 0.5919 | 0.6112 | 0.7910 | 0.7824 | 0.7072 | 0.6287 | 0.7802 | 0.5800 | 0.7570 |
| BQ 1-bit (32x) | 0.6028 | 0.4683 | 0.4945 | 0.7041 | 0.6921 | 0.6098 | 0.5409 | 0.6989 | 0.4762 | 0.6760 |
The pattern repeats across all ten datasets:
* **TQ 4-bit is competitive with SQ at half the storage.** On 9 of 10 datasets the gap to SQ is within 2 pp in either direction; on 3 of those (`dbp-oai`, `cohere`, `laion`) TQ 4-bit *beats* SQ — by up to 4.6 pp on `dbp-oai`. The single exception is `arxiv-384`, where TQ 4-bit trails SQ by 2.3 pp. The pattern is consistent: when SQ's int8-per-coordinate grid is mismatched with the embedding distribution, an adaptive 4-bit quantizer with anisotropy compensation does better, despite using half the bits.
* **TQ 2-bit beats BQ 2-bit by 9–24 pp** on every dataset, at the same 16× storage class. The largest margins are on `laion` (+24.0 pp) and `h&m` (+21.8 pp); the smallest is `ads-1M` (+9.0 pp).
* **TQ 1-bit beats vanilla BQ 1-bit by 9–21 pp** on every dataset, at the same 32× storage class. Against the stronger asymmetric BQ configuration (1-bit storage, 8-bit query), TQ 1-bit is still ahead on every dataset, though the margin narrows — between 0.1 pp (`cohere`, essentially tied) and 10 pp (`laion`).
* **TQ 1.5-bit (~21×)** sits between the 2-bit and 1-bit operating points and is the right pick when 32× is too aggressive but 16× leaves storage on the table.
* **TQ 2-bit beats BQ 2-bit by 9–24 pp** on every dataset, at the same 16x storage class. The largest margins are on `laion` (+24.0 pp) and `h&m` (+21.8 pp); the smallest is `ads-1M` (+9.0 pp).
* **TQ 1-bit beats vanilla BQ 1-bit by 9–21 pp** on every dataset, at the same 32x storage class. Against the stronger asymmetric BQ configuration (1-bit storage, 8-bit query), TQ 1-bit is still ahead on every dataset, though the margin narrows — between 0.1 pp (`cohere`, essentially tied) and 10 pp (`laion`).
* **TQ 1.5-bit (~21x)** sits between the 2-bit and 1-bit operating points and is the right pick when 32x is too aggressive but 16x leaves storage on the table.
## When to Use TurboQuant
A practical guide:
* **You currently run SQ** → try TQ 4-bit. Comparable recall (often within 1–2 pp; sometimes higher) at half the memory. The easiest upgrade call on the ladder.
* **You currently run BQ at any bit depth** → try TurboQuant at the same storage budget (BQ 2-bit → TQ 2-bit, BQ 1.5-bit → TQ 1.5-bit, BQ 1-bit → TQ 1-bit). On the benchmarks described here, it consistently delivers higher recall — typically 10–20 pp at both the 16× and 32× storage classes. Stay on BQ if you observe a noticeable drop in throughput on your workload, or if the recall improvement is too small to matter for your use case.
* **You currently run BQ at any bit depth** → try TurboQuant at the same storage budget (BQ 2-bit → TQ 2-bit, BQ 1.5-bit → TQ 1.5-bit, BQ 1-bit → TQ 1-bit). On the benchmarks described here, it consistently delivers higher recall — typically 10–20 pp at both the 16x and 32x storage classes. Stay on BQ if you observe a noticeable drop in throughput on your workload, or if the recall improvement is too small to matter for your use case.
* **You need cosine, dot, or L2** → all three are first-class in TurboQuant. **L1** → stay on SQ.
A word on indexing: TurboQuant has a small one-time pre-pass per segment (the calibration scan) that runs in a few seconds at production segment sizes. Once a segment is calibrated, the calibration is reused across queries and segment merges; it is paid once per segment, never per query.
## Conclusion
TurboQuant gives Qdrant a new path on the compression ladder: 8× compression at SQ-level recall, and at 16× / 32× a consistent 10–20 percentage points of recall above BQ on every embedding model we have benchmarked. What makes that work is a hybrid — TurboQuant's MSE codebook and integer-arithmetic SIMD kernels, RaBitQ's per-vector length rescaling and bit-plane scoring at 1-bit, and the anisotropy-compensation pre-pass we developed on top to make all of it land on real production embeddings. The whole stack is shipping in **Qdrant 1.18**, on Cloud and in the standard Docker image. Migration from SQ or BQ is a config change and a re-index; the rest of the application stays identical.
TurboQuant gives Qdrant a new path on the compression ladder: 8x compression at SQ-level recall, and at 16x / 32x a consistent 10–20 percentage points of recall above BQ on every embedding model we have benchmarked. What makes that work is a hybrid — TurboQuant's MSE codebook and integer-arithmetic SIMD kernels, RaBitQ's per-vector length rescaling and bit-plane scoring at 1-bit, and the anisotropy-compensation pre-pass we developed on top to make all of it land on real production embeddings. The whole stack is shipping in **Qdrant 1.18**, on Cloud and in the standard Docker image. Migration from SQ or BQ is a config change and a re-index; the rest of the application stays identical.
## Further Reading