Add links to TurboQuant article (#2354)

* Add links to TQ article

* Add link to articles for other quantization methods

* Update TQ article preview images
This commit is contained in:
Abdon Pijpelink
2026-05-15 16:26:18 +02:00
committed by GitHub
parent 33410a4795
commit 1a0f06231c
7 changed files with 5 additions and 5 deletions
+1 -1
View File
@@ -66,7 +66,7 @@ The following table shows recall for 1-bit TurboQuant (TQ1) compared to uncompre
Compared to 1-bit binary quantization, 1-bit TurboQuant offers better recall at equivalent storage budgets, albeit at a lower speed. Similar trends are observed for 1.5-bit and 2-bit configurations.
We will soon publish an article with detailed numbers, including throughput and indexing times.
For more details about Qdrant's TurboQuant implementation, refer to our article [TurboQuant in Qdrant](/articles/turboquant-quantization/).
### Get Started with TurboQuant
@@ -54,7 +54,7 @@ Depending on your requirements for recall, compression, and distance metrics, co
TurboQuant is [a quantization method developed by Google](https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/). It operates by applying a fast random rotation to vectors before compression, which evenly redistributes data across coordinates. This allows applying a single pre-computed, globally optimized quantization mapping across the dataset, enabling TurboQuant to work effectively with any vector distribution and overcoming a key limitation found in binary quantization.
Qdrant's implementation of TurboQuant extends the original algorithm to close the gap between the algorithm's theoretical assumptions and real-world embeddings.
[Qdrant's implementation of TurboQuant](/articles/turboquant-quantization/) extends the original algorithm to close the gap between the algorithm's theoretical assumptions and real-world embeddings.
TurboQuant uses asymmetric quantization automatically: only stored vectors are compressed, while queries are scored in full precision. This improves accuracy and requires no additional configuration.
@@ -85,7 +85,7 @@ Manhattan (L1) distance is supported but requires full vector reconstruction per
*Available as of v1.1.0*
Scalar quantization, in the context of vector search engines, is a compression technique that compresses vectors by reducing the number of bits used to represent each vector component.
[Scalar quantization](/articles/scalar-quantization/), in the context of vector search engines, is a compression technique that compresses vectors by reducing the number of bits used to represent each vector component.
For instance, Qdrant uses 32-bit floating numbers to represent the original vector components. Scalar quantization allows you to reduce the number of bits used to 8.
In other words, Qdrant performs `float32 -> uint8` conversion for each vector component.
@@ -106,7 +106,7 @@ Please refer to the [Quantization Tips](#quantization-tips) section for more inf
*Available as of v1.5.0*
Binary quantization is an extreme case of scalar quantization.
[Binary quantization](/articles/binary-quantization/) is an extreme case of scalar quantization.
This feature lets you represent each vector component as a single bit, effectively reducing the memory footprint by a factor of 32. This is the fastest quantization method, since it lets you perform a vector comparison with a few CPU instructions. Binary quantization can achieve up to a 40x speedup compared to the original vectors.
However, binary quantization is only efficient for high-dimensional vectors and require a centered distribution of vector components.
@@ -191,7 +191,7 @@ See how to set up Asymmetric Quantization quantization in the [following section
*Available as of v1.2.0*
Product quantization is a method of compressing vectors to minimize their memory usage by dividing them into
[Product quantization](/articles/product-quantization/) is a method of compressing vectors to minimize their memory usage by dividing them into
chunks and quantizing each segment individually.
Each chunk is approximated by a centroid index that represents the original vector component.
The positions of the centroids are determined through the utilization of a clustering algorithm such as k-means.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 120 KiB

After

Width:  |  Height:  |  Size: 53 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 46 KiB

After

Width:  |  Height:  |  Size: 16 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.1 MiB

After

Width:  |  Height:  |  Size: 1.3 MiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 349 KiB

After

Width:  |  Height:  |  Size: 219 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 130 KiB

After

Width:  |  Height:  |  Size: 68 KiB