This commit is contained in:
generall
2024-07-09 13:03:43 +02:00
parent ad59119bf6
commit c00615b50e
+1 -1
View File
@@ -338,7 +338,7 @@ With `datatype: uint8` available in Qdrant, the total size of the sparse vector
As a reference point, we use:
- BM25 with tantivy
- the [sparse vector BM25 implementation](https://huggingface.co/Qdrant/bm25) with the same preprocessing pipeline like for BM42: tokenization, stop-words removal, and lemmatization
- the [sparse vector BM25 implementation](https://github.com/qdrant/bm42_eval/blob/master/index_bm25_qdrant.py) with the same preprocessing pipeline like for BM42: tokenization, stop-words removal, and lemmatization
| | BM25 (tantivy) | BM25 (Sparse) | BM42 |
|----------------------|-------------------|---------------|----------|