Add Binary Q (#462)

* * docs(gemini.md): add section on using Gemini Embedding Models with Binary Quantization

* * docs(gemini.md): provide comparison results of search with Binary Quantization and original model

* * docs(gemini.md): update Gemini Embedding Models performance table with improved recall at 100 oversampling limit

* * docs(gemini.md): fix typo in search results table caption
This commit is contained in:
Nirant
2023-12-13 17:06:17 +05:30
committed by GitHub
parent 49106b9f09
commit 6720253abd
@@ -91,4 +91,21 @@ qdrant_client.search(
)
```
That's it! You can now use Gemini Embedding Models with Qdrant.
## Using Gemini Embedding Models with Binary Quantization
You can use Gemini Embedding Models with [Binary Quantization](../../articles/binary-quantization.md) - a technique that allows you to reduce the size of the embeddings by 32 times without losing the quality of the search results too much.
In this table, you can see the results of the search with the `models/embedding-001` model with Binary Quantization in comparison with the original model:
At an oversampling of 3 and a limit of 100, we've a 95% recall against the exact nearest neighbors with rescore enabled.
| oversampling | | 1 | 1 | 2 | 2 | 3 | 3 |
|--------------|---------|----------|----------|----------|----------|----------|----------|
| limit | | | | | | | |
| | rescore | False | True | False | True | False | True |
| 10 | | 0.523333 | 0.831111 | 0.523333 | 0.915556 | 0.523333 | 0.950000 |
| 20 | | 0.510000 | 0.836667 | 0.510000 | 0.912222 | 0.510000 | 0.937778 |
| 50 | | 0.489111 | 0.841556 | 0.489111 | 0.913333 | 0.488444 | 0.947111 |
| 100 | | 0.485778 | 0.846556 | 0.485556 | 0.929000 | 0.486000 | **0.956333** |
That's it! You can now use Gemini Embedding Models with Qdrant!