Add links to TurboQuant article (#2354)
* Add links to TQ article * Add link to articles for other quantization methods * Update TQ article preview images
@@ -66,7 +66,7 @@ The following table shows recall for 1-bit TurboQuant (TQ1) compared to uncompre
|
||||
|
||||
Compared to 1-bit binary quantization, 1-bit TurboQuant offers better recall at equivalent storage budgets, albeit at a lower speed. Similar trends are observed for 1.5-bit and 2-bit configurations.
|
||||
|
||||
We will soon publish an article with detailed numbers, including throughput and indexing times.
|
||||
For more details about Qdrant's TurboQuant implementation, refer to our article [TurboQuant in Qdrant](/articles/turboquant-quantization/).
|
||||
|
||||
### Get Started with TurboQuant
|
||||
|
||||
|
||||
@@ -54,7 +54,7 @@ Depending on your requirements for recall, compression, and distance metrics, co
|
||||
|
||||
TurboQuant is [a quantization method developed by Google](https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/). It operates by applying a fast random rotation to vectors before compression, which evenly redistributes data across coordinates. This allows applying a single pre-computed, globally optimized quantization mapping across the dataset, enabling TurboQuant to work effectively with any vector distribution and overcoming a key limitation found in binary quantization.
|
||||
|
||||
Qdrant's implementation of TurboQuant extends the original algorithm to close the gap between the algorithm's theoretical assumptions and real-world embeddings.
|
||||
[Qdrant's implementation of TurboQuant](/articles/turboquant-quantization/) extends the original algorithm to close the gap between the algorithm's theoretical assumptions and real-world embeddings.
|
||||
|
||||
TurboQuant uses asymmetric quantization automatically: only stored vectors are compressed, while queries are scored in full precision. This improves accuracy and requires no additional configuration.
|
||||
|
||||
@@ -85,7 +85,7 @@ Manhattan (L1) distance is supported but requires full vector reconstruction per
|
||||
|
||||
*Available as of v1.1.0*
|
||||
|
||||
Scalar quantization, in the context of vector search engines, is a compression technique that compresses vectors by reducing the number of bits used to represent each vector component.
|
||||
[Scalar quantization](/articles/scalar-quantization/), in the context of vector search engines, is a compression technique that compresses vectors by reducing the number of bits used to represent each vector component.
|
||||
|
||||
For instance, Qdrant uses 32-bit floating numbers to represent the original vector components. Scalar quantization allows you to reduce the number of bits used to 8.
|
||||
In other words, Qdrant performs `float32 -> uint8` conversion for each vector component.
|
||||
@@ -106,7 +106,7 @@ Please refer to the [Quantization Tips](#quantization-tips) section for more inf
|
||||
|
||||
*Available as of v1.5.0*
|
||||
|
||||
Binary quantization is an extreme case of scalar quantization.
|
||||
[Binary quantization](/articles/binary-quantization/) is an extreme case of scalar quantization.
|
||||
This feature lets you represent each vector component as a single bit, effectively reducing the memory footprint by a factor of 32. This is the fastest quantization method, since it lets you perform a vector comparison with a few CPU instructions. Binary quantization can achieve up to a 40x speedup compared to the original vectors.
|
||||
|
||||
However, binary quantization is only efficient for high-dimensional vectors and require a centered distribution of vector components.
|
||||
@@ -191,7 +191,7 @@ See how to set up Asymmetric Quantization quantization in the [following section
|
||||
|
||||
*Available as of v1.2.0*
|
||||
|
||||
Product quantization is a method of compressing vectors to minimize their memory usage by dividing them into
|
||||
[Product quantization](/articles/product-quantization/) is a method of compressing vectors to minimize their memory usage by dividing them into
|
||||
chunks and quantizing each segment individually.
|
||||
Each chunk is approximated by a centroid index that represents the original vector component.
|
||||
The positions of the centroids are determined through the utilization of a clustering algorithm such as k-means.
|
||||
|
||||
|
Before Width: | Height: | Size: 120 KiB After Width: | Height: | Size: 53 KiB |
|
Before Width: | Height: | Size: 46 KiB After Width: | Height: | Size: 16 KiB |
|
Before Width: | Height: | Size: 1.1 MiB After Width: | Height: | Size: 1.3 MiB |
|
Before Width: | Height: | Size: 349 KiB After Width: | Height: | Size: 219 KiB |
|
Before Width: | Height: | Size: 130 KiB After Width: | Height: | Size: 68 KiB |