mirror of
https://github.com/qdrant/landing_page.git
synced 2026-09-28 23:48:31 +02:00
Add Tantivy back to the BM42 article (#1014)
This commit is contained in:
@@ -335,19 +335,22 @@ After encoding with BM42, the average vector size is only **5.6 elements per doc
|
||||
|
||||
With `datatype: uint8` available in Qdrant, the total size of the sparse vector index is about **13Mb** for ~530k documents.
|
||||
|
||||
As a referene point, we use the BM25 implementation with the same preprocessing pipeline: tokenization, stop-words removal, and lemmatization. There are other implementations of BM25 which might use different components and produce different results.
|
||||
As a reference point, we use:
|
||||
|
||||
Those implementations are out of the scope for this article.
|
||||
- BM25 with tantivy
|
||||
- the [sparse vector BM25 implementation](https://huggingface.co/Qdrant/bm25) with the same preprocessing pipeline like for BM42: tokenization, stop-words removal, and lemmatization
|
||||
|
||||
|
||||
| | BM25 (Sparse) | BM42 |
|
||||
|-------------------|------|----------|
|
||||
|~~Precision @ 10~~ * | ~~0.45~~ | ~~0.49~~ |
|
||||
| Recall @ 10 | 0.83 | **0.85** |
|
||||
| | BM25 (tantivy) | BM25 (Sparse) | BM42 |
|
||||
|----------------------|-------------------|---------------|----------|
|
||||
| ~~Precision @ 10~~ * | ~~0.45~~ | ~~0.45~~ | ~~0.49~~ |
|
||||
| Recall @ 10 | ~~0.71~~ **0.89** | 0.83 | 0.85 |
|
||||
|
||||
|
||||
\* - values were corrected after the publication due to a mistake in the evaluation script.
|
||||
|
||||
<aside role="status">
|
||||
When used properly, BM25 with tantivy achieves the best results. Our initial implementation performed the character escaping that led to understating the value of <code>recall@10</code> for tantivy.
|
||||
</aside>
|
||||
|
||||
To make our benchmarks transparent, we have published scripts we used for the evaluation: see [github repo](https://github.com/qdrant/bm42_eval).
|
||||
|
||||
|
||||
Reference in New Issue
Block a user