better calibration and simd parts

This commit is contained in:
Ivan Pleshkov
2026-05-11 15:14:25 +02:00
parent bb835f71ec
commit eb7e866208
@@ -123,25 +123,23 @@ TurboQuant's PROD variant spends an entire QJL random projection plus extra bits
### Per-Coordinate Calibration (Anisotropy Compensation)
After rotation, we run a single pre-pass over the segment that estimates a `(shift, scale)` pair per coordinate. Each coordinate is then mapped via `x → (x + shift) · scale` before quantization, pulling its empirical distribution onto the Lloyd-Max codebook's grid no matter how skewed it was originally. This is the extension that turns vanilla MSE from "elegant but only competitive on isotropic data" into something that beats BQ across every embedding model we have benchmarked.
The rotation step gives every coordinate a roughly N(0, 1) distribution **on isotropic data** — the proof goes through the central limit theorem and does not extend to anisotropic embeddings, where a few directions concentrate most of the variance. After rotation those high-variance directions get spread across coordinates, but the per-coordinate distributions are not all identical Gaussians — they have different scales, different shapes, sometimes heavy tails. The Lloyd-Max codebook is fitted once for N(0, 1) and stays fixed, so coordinates that drift off the codebook grid waste centroids and lose recall.
The non-obvious part is **what we calibrate to**. Standard mean-and-stddev rescaling implicitly assumes the rotated coordinates are Gaussian, which is exactly the assumption that breaks on anisotropic embeddings — that is why vanilla TurboQuant struggles on them in the first place. Instead, we anchor calibration directly to the codebook itself: the `(shift, scale)` pair is fit so that the empirical quantiles at the probability levels of the **outermost codebook centroid** land at that centroid. The calibration is robust to whatever distributional weirdness the post-rotation coordinates actually have — heavy tails, bimodality, anisotropy, all of it — without making any parametric assumption.
Because Qdrant stores data in segments, we can fix this per segment. For each segment we do a single **pre-pass** before quantization: estimate a `(shift, scale)` pair per coordinate after rotation, then apply `x → (x + shift) · scale` to pull the empirical per-coordinate distribution back onto the codebook's grid. The same `(shift, scale)` is baked into the segment's metadata and reused for every query that hits the segment.
The quantile estimation itself is done with the [P-Square algorithm](https://www.cse.wustl.edu/~jain/papers/ftp/psqr.pdf) (Jain & Chlamtac, 1985) on a reservoir-sampled prefix of the segment: streaming, no parametric fit, constant memory per coordinate. The reservoir size is picked per bit-depth — the 4-bit codebook anchors at a deeper tail quantile than 1-bit, so it gets a larger reservoir to keep the quantile estimator's variance flat. The full P-Square + reservoir setup is detailed in [Ivan's deep-dive](#further-reading).
**This is free at search time** thanks to the asymmetric scoring scheme. Applying `(shift, scale)` on the data side would mean an extra multiply per coordinate per scored vector — unaffordable in the hot path. Instead we move that work onto the query: at query-encoding time we precompute the rescaled query and a single scalar score-bias term, once per query. The scoring kernel does not change shape, and storage stays at exactly `b·D` bits per vector.
Three properties make this extension safe to ship by default:
**Why not just mean + stddev?** Mean-and-stddev rescaling assumes the post-rotation coordinates are Gaussian — exactly the assumption that breaks on anisotropic data, which is the case where we need calibration in the first place. We anchor calibration to the codebook itself instead: the `(shift, scale)` pair is fit so the empirical quantiles at the probability levels of the **outermost codebook centroid** land at that centroid. The quantiles themselves are estimated with the [P-Square algorithm](https://www.cse.wustl.edu/~jain/papers/ftp/psqr.pdf) (Jain & Chlamtac, 1985) — streaming, no parametric fit, constant memory per coordinate.
* **On truly isotropic data, calibration vanishes.** The empirical quantiles match the theoretical Gaussian quantiles, the formula collapses to `(shift=0, scale=1)`, and the encoded vector is bit-identical to vanilla TurboQuant MSE. So the extension never makes already-good data worse — that is a mathematical property of the formula, not an empirical observation across the benchmark suite.
* **The compensation is paid by the query, not by storage.** The shift and scale fold into a one-time precomputed scaling of the rotated query plus a single scalar correction term. Storage stays at exactly `b·D` bits per vector — no per-vector overhead beyond the codebook indices themselves.
* **Symmetric scoring stays well-defined.** The same `(shift, scale)` pair is baked into the segment's metadata and applied consistently whether one or both sides of a comparison are quantized. HNSW build and segment merges work directly on the encoded vectors; the calibration parameters compose into the score formula symmetrically.
**Sampling, not full scan.** Running P-Square over every vector in the segment would dominate index-build time. We instead sample a random subset of segment vectors using [Vitter's Algorithm R](https://en.wikipedia.org/wiki/Reservoir_sampling#Algorithm_R) (classical reservoir sampling), then run P-Square on the reservoir. The sample size is picked per bit-depth — the 4-bit codebook anchors at a deeper tail quantile than 1-bit, so it needs more samples to keep the estimator's variance flat. Full setup is in [Ivan's deep-dive](#further-reading).
Ablation in the deep-dive quantifies the impact: at 4-bit on highly anisotropic embeddings (e.g. arxiv-instructorxl), this extension alone is worth roughly 14 percentage points of recall vs. vanilla MSE; at 1-bit it is worth ~8 pp. On already-isotropic embeddings (e.g. text-embedding-3-large) the contribution is below 1 pp at every bit depth — exactly as the "vanishes on isotropic data" property predicts.
On truly isotropic data the empirical quantiles match the theoretical Gaussian quantiles, the formula collapses to `(shift=0, scale=1)`, and the encoded vector is bit-identical to vanilla TurboQuant MSE. The extension never makes already-good data worse.
### L2 and Unnormalized Dot
Vanilla TurboQuant assumes all inputs live on the unit sphere — that is, cosine distance only. We extend the algorithm beyond the sphere by **storing the original L2 norm** and restore L2 and Unnormalized dot from normalized one, so dot and L2 cost the same as cosine in the hot path.
L2 distances are reconstructed via the identity `‖q − v‖² = ‖q‖² + ‖v‖² − 2⟨q, v⟩ = ‖q‖² + ‖v‖² − 2 ‖v‖² ‖q‖²⟨q_normalized, v_normalized⟩`, where all components on the right-hand side are already available.
L2 distances are reconstructed via the identity `‖q − v‖² = ‖q‖² + ‖v‖² − 2⟨q, v⟩ = ‖q‖² + ‖v‖² − 2 ‖v‖ ‖q‖ ⟨q_normalized, v_normalized⟩`, where all components on the right-hand side are already available.
Net result: **cosine, dot, and L2 are all first-class** in Qdrant's TurboQuant — same storage layout, same kernels, no precision tax for the non-cosine metrics.