Add Tantivy back to the BM42 article (#1014)

This commit is contained in:
Kacper Łukawski
2024-07-09 12:45:50 +02:00
committed by GitHub
parent e855ec4fb9
commit ad59119bf6
+10 -7
View File
@@ -335,19 +335,22 @@ After encoding with BM42, the average vector size is only **5.6 elements per doc
With `datatype: uint8` available in Qdrant, the total size of the sparse vector index is about **13Mb** for ~530k documents.
As a referene point, we use the BM25 implementation with the same preprocessing pipeline: tokenization, stop-words removal, and lemmatization. There are other implementations of BM25 which might use different components and produce different results.
As a reference point, we use:
Those implementations are out of the scope for this article.
- BM25 with tantivy
- the [sparse vector BM25 implementation](https://huggingface.co/Qdrant/bm25) with the same preprocessing pipeline like for BM42: tokenization, stop-words removal, and lemmatization
| | BM25 (Sparse) | BM42 |
|-------------------|------|----------|
|~~Precision @ 10~~ * | ~~0.45~~ | ~~0.49~~ |
| Recall @ 10 | 0.83 | **0.85** |
| | BM25 (tantivy) | BM25 (Sparse) | BM42 |
|----------------------|-------------------|---------------|----------|
| ~~Precision @ 10~~ * | ~~0.45~~ | ~~0.45~~ | ~~0.49~~ |
| Recall @ 10 | ~~0.71~~ **0.89** | 0.83 | 0.85 |
\* - values were corrected after the publication due to a mistake in the evaluation script.
<aside role="status">
When used properly, BM25 with tantivy achieves the best results. Our initial implementation performed the character escaping that led to understating the value of <code>recall@10</code> for tantivy.
</aside>
To make our benchmarks transparent, we have published scripts we used for the evaluation: see [github repo](https://github.com/qdrant/bm42_eval).