diff --git a/qdrant-landing/content/articles/bm42.md b/qdrant-landing/content/articles/bm42.md index ee80d8b88..0e675ba3e 100644 --- a/qdrant-landing/content/articles/bm42.md +++ b/qdrant-landing/content/articles/bm42.md @@ -335,19 +335,22 @@ After encoding with BM42, the average vector size is only **5.6 elements per doc With `datatype: uint8` available in Qdrant, the total size of the sparse vector index is about **13Mb** for ~530k documents. -As a referene point, we use the BM25 implementation with the same preprocessing pipeline: tokenization, stop-words removal, and lemmatization. There are other implementations of BM25 which might use different components and produce different results. +As a reference point, we use: -Those implementations are out of the scope for this article. +- BM25 with tantivy +- the [sparse vector BM25 implementation](https://huggingface.co/Qdrant/bm25) with the same preprocessing pipeline like for BM42: tokenization, stop-words removal, and lemmatization - -| | BM25 (Sparse) | BM42 | -|-------------------|------|----------| -|~~Precision @ 10~~ * | ~~0.45~~ | ~~0.49~~ | -| Recall @ 10 | 0.83 | **0.85** | +| | BM25 (tantivy) | BM25 (Sparse) | BM42 | +|----------------------|-------------------|---------------|----------| +| ~~Precision @ 10~~ * | ~~0.45~~ | ~~0.45~~ | ~~0.49~~ | +| Recall @ 10 | ~~0.71~~ **0.89** | 0.83 | 0.85 | \* - values were corrected after the publication due to a mistake in the evaluation script. + To make our benchmarks transparent, we have published scripts we used for the evaluation: see [github repo](https://github.com/qdrant/bm42_eval).