imgs fixed + disclaimer

This commit is contained in:
Евгения Суходольская
2025-05-05 12:52:53 +02:00
committed by Евгения Суходольская
parent a0fbdc092c
commit 8fd03af204
3 changed files with 6 additions and 3 deletions
+6 -3
View File
@@ -144,7 +144,7 @@ Unless... We do a simple trick!
Imagine a bag-of-words sparse vector. Every word from the vocabulary takes up one cell. If the word is present in the encoded text — we assign some weight; if it isn't — it equals zero.
If we have a mini COIL vector describing a word's meaning, for example, in 4D semantic space, we could just dedicate 4 consecutive cells for TBD word in the sparse vector, one cell per "meaning" dimension. If we don't, we could fall back to a classic one-cell description with a pure BM25 score.
If we have a mini COIL vector describing a word's meaning, for example, in 4D semantic space, we could just dedicate 4 consecutive cells for word in the sparse vector, one cell per "meaning" dimension. If we don't, we could fall back to a classic one-cell description with a pure BM25 score.
**Such representations can be used in any standard inverted index.**
@@ -216,8 +216,7 @@ Then we could train all the words we're interested in and simply combine (stack)
### Implementation Details
The code of the training approach sketched above is open-sourced [in this repository](https://github.com/qdrant/miniCOIL).
TBD: NEEDS UPDATED README CC ANDREY
The code of the training approach sketched above is open-sourced [in this repository](https://github.com/qdrant/miniCOIL).
Here are the specific characteristics of the miniCOIL model we trained based on this approach:
@@ -253,6 +252,10 @@ We're testing our 4D miniCOIL model versus [our BM25 implementation](https://hug
- `k = 1.2`, `b = 0.75` default values recommended to use with BM25 scoring;
- `avg_len` estimated on 50_000 documents from a respective dataset.
<aside role="status">
BM25 results depend on implementation details, such as the choice of stemmer, tokenizer, stop word list, etc. So, to make miniCOIL comparable to BM25, we use our own BM25 implementation and reuse all its implementation choices for miniCOIL.
</aside>
We compare models based on the `NDCG@10` metric, as we're interested in the ranking performance of miniCOIL compared to BM25. Both retrieve the same subset of indexed documents based on exact matches, but miniCOIL should ideally rank this subset better based on its semantics understanding.
The result on several domains we tested is the following:
Binary file not shown.

Before

Width:  |  Height:  |  Size: 12 KiB

After

Width:  |  Height:  |  Size: 30 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 39 KiB

After

Width:  |  Height:  |  Size: 28 KiB