results based on a 2nd review

This commit is contained in:
Евгения Суходольская
2025-05-05 12:52:53 +02:00
committed by Евгения Суходольская
parent a8e51435e5
commit 5cb74da9db
3 changed files with 109 additions and 106 deletions
+109 -106
View File
@@ -95,9 +95,9 @@ Experiments showed that SPLADE without its term expansion tells the same old sto
## Eyes on the Prize: Useable Sparse Neural Retrieval
Striving for perfection on specific benchmarks, the sparse neural retrieval field either produced models performing out-of-domain worse than BM25 (ironically, [trained with BM25-based hard negatives](https://arxiv.org/pdf/2307.10488)) or ones based on heavy document expansion, disrupting sparsity.
Striving for perfection on specific benchmarks, the sparse neural retrieval field either produced models performing out-of-domain worse than BM25 (ironically, [trained with BM25-based hard negatives](https://arxiv.org/pdf/2307.10488)) or ones based on heavy document expansion, lowering sparsity.
To be usable in production, the minimal criteria a sparse neural retriever should meet are:
So, to be usable in production, the minimal criteria a sparse neural retriever should meet are:
- **Producing lightweight sparse representations (it's in the name!).** Inheriting the perks of term-based retrieval, it should be lightweight and simple. For broader semantic search, there are dense retrievers, and they work.
- **Being better than BM25 at ranking in different domains.** The goal is a term-based retriever capable of distinguishing word meanings — what BM25 can't do — preserving BM25's out-of-domain, time-proven performance.
@@ -108,15 +108,15 @@ To be usable in production, the minimal criteria a sparse neural retriever shoul
One of the attempts in the field of Sparse Neural Retrieval — [Contextualized Inverted Lists (COIL)](https://qdrant.tech/articles/modern-sparse-neural-retrieval/#sparse-neural-retriever-which-understood-homonyms) — stands out with its approach to term weights encoding.
Instead of squishing high-dimensional token representations (usually 768-dimensional BERT embeddings) into a single number, COIL authors project them to smaller vectors of 32 dimensions. They propose storing these vectors in **inverted lists* of an **inverted index** (used in term-based retrieval) **as is** and comparing vector representations through dot product.
Instead of squishing high-dimensional token representations (usually 768-dimensional BERT embeddings) into a single number, COIL authors project them to smaller vectors of 32 dimensions. They propose storing these vectors in **inverted lists** of an **inverted index** (used in term-based retrieval) as is and comparing vector representations through dot product.
This approach captures deeper semantics — one number can't perfectly convey all shades of meaning that one word can have. Yet it didn't catch on, and COIL didn't become popular, presumably due to the following reasons:
- Inverted indexes are usually not designed to store vectors and perform vector operations.
- Trained end-to-end with a relevance objective on [MS MARCO dataset](https://microsoft.github.io/msmarco/), COIL's performance is heavily domain-bound.
- Additionally, COIL operates on tokens, reusing BERT's tokenizer. Yet, working at a word level is far better for term-based retrieval. Say we want to search for a "retriever" in our documentation. COIL will break it down into `re`, `#trie`, and `#ver` 32-dimensional vectors and match all three parts separately -- not so convenient.
- Additionally, COIL operates on tokens, reusing BERT's tokenizer. Yet, working at a word level is far better for term-based retrieval. Say we want to search for a *"retriever"* in our documentation. COIL will break it down into `re`, `#trie`, and `#ver` 32-dimensional vectors and match all three parts separately -- not so convenient.
However, COIL representations allow distinguishing **homographs**, a skill BM25 lacks. The best ideas don't start from zero. We could try to **build on top of COIL, keeping in mind what needs fixing**:
However, COIL representations allow distinguishing homographs, a skill BM25 lacks. The best ideas don't start from zero. We could try to **build on top of COIL, keeping in mind what needs fixing**:
1. To get a model performant on out-of-domain data, we should abandon end-to-end training on a relevance objective — there is not enough data to train a model able to generalize.
2. We should keep representations sparse and reusable in a classic inverted index.
@@ -126,80 +126,90 @@ However, COIL representations allow distinguishing **homographs**, a skill BM25
BM25 has been a decent baseline across various domains for many years -- and for a good reason. So why discard a time-proven formula?
Instead of learning to assign importance scores to words based on text relevance (relevance objective), let's add a semantic COIL-inspired component to BM25 scoring. Then, if we manage to capture a word's meaning, our solution alone could work like BM25 combined with a semantically aware reranker -- or, in other words:
Instead of learning to assign words' importance scores based on query-document relevance, let's add a semantic COIL-inspired component to BM25 scoring. Then, if we manage to capture a word's meaning, our solution alone could work like BM25 combined with a semantically aware reranker -- or, in other words:
- It could see the difference between a *"fruit **bat**"* and a *"baseball **bat**"*
- When used with word stems, it could distinguish parts of speech. For example, *"inform"*, *"informant"*, *"informational"*, *"informed"*, and *"informally"* won't get lost in one *"inform"* stem.
- When used with word stems, it could distinguish parts of speech. For example, *"inform"*, *"informant"*, *"informational"*, *"informed"*, and *"informally"* won't get lost in one *"**inform**"* stem.
$$
\text{score}(D,Q) = \sum_{i=1}^{N} \text{IDF}(q_i) \cdot \text{Importance}^{q_i}_{D} \cdot \text{Meaning}^{q_i \times D}_{\text{semantic match}}
$$
We won't need labelled retrieval-related datasets to acquire this semantic component since word meaning is hidden in the surrounding context. And if we stumble upon a word we don't know, we can just use the BM25 formula for it as is!
We won't need labelled retrieval-related datasets to acquire this semantic component as word meaning is hidden in the surrounding context, aka in any texts including this word.
And if our model stumbles upon a word it hasn't "seen" during training, we can just fall back to the original BM25 formula!
### Bag-of-words in 4D
COIL uses 32 values per one term. Do we need these much? How many examples of a word with 32 non-intersecting meanings could you give without additional research?
COIL uses 32 values to describe a term. Do we need this many? How many words with 32 separate meanings could you name without additional research? Eight or even four seem already sufficient.
Yet even if we use less values for COIL representations, the initial problem of inadaptability of dense vectors to inverted index persists.
Yet, even if we use fewer values in COIL representations, the initial problem of dense vectors not fitting into a classical inverted index persists.
Unless... We do a simple trick!
Imagine a classic bag-of-words sparse vector. Every word from vocabulary takes up one cell in this vector, if it's present in encoded text -- we assign it some score, if it doesn't -- it's equal to zero.
TBD IMAGE 1D -> 4D BAG-OF-WORD
Now, imagine we have a very small dense vector, for example, of a size four, describing word's meanings within a 4D semantical space.
We could simply dedicate fours consequitive cells for each word in our sparse vector, one per "meaning" dimension. Then vectors would be only 4 times less sparse than a classical bag-of-words version, one word fitting in four bytes if needed.
Imagine a bag-of-words sparse vector. Every word from the vocabulary takes up one cell. If the word is present in the encoded text — we assign some weight; if it isn't — it equals zero.
In this setting, classical term-based retrieval approach -- multiply score from matched cells (words) and sum over -- fully resembles dot product of word vectors, as proposed in COIL.
If we have a mini dense vector describing a word's meaning, for example, in 4D semantic space, we could just dedicate 4 consecutive cells for this word in the sparse vector, one cell per "meaning" dimension. If we don't, we could fall back to a classic one-cell description with a pure BM25 score.
Such representation of **mini COIL vectors** allows for flexibility:
- It can be used in any classical inverted index
- If word is unknown (so, consequently, it's meanings), we could dedicate to it only one cell, as in classical bag-of-words, and assing it pure BM25 score.
**Such representation of mini COIL vectors can be used in any standard inverted index.**
TBD PIC EXPLAINING BAG-OF-WORDS TO 4D BAG-OF-WORDS (miniCOIL)
## Training miniCOIL
## How to Train miniCOIL
Now we're coming to the part where we need to somehow get this low dimensional encapsulation of a word meaning in context.
Now we're coming to the part where we need to understand how to get this low dimensional encapsulation of a word meaning.
We want to work smarter, not harder, and rely as much as possible on time proven solutions. Dense encoders are good at encoding word's meaning in its context, it would be convenient to reuse their output. Moreover, we could kill two birds with one stone, if we wanted to add miniCOIL to hybrid search -- where dense encoders inference is done regardless.
NO RELEVANCE OBJECTIVE
WE DO NOT WANT TO TRAIN (why DENSE) -- WE WANT TO REUSE + HYBRID SEARCH
Yet dense encoders outputs are high-dimensional, so need we need to perform **meaning-in-context preserving dimensionality reduction**. The goal is to:
To capture contextual meaning of words in input data, we get a contextuazlied word embedding from a dense encoder. Yet it's high-dimensional, and need we need to perform a meaning preserving dimensionality reduction, so the word's low dimensional representation is comparable within an encoded corpora.
- Avoid relevance objective and dependance on labelled datasets;
- Find reflecting words' meanings spatial relations target, reusable for different dense encoders;
- Use the simplest architecture possible.
We don't want to depend on corpora and, moreover, we would want find a way to easily adapt to different dense encoders, as they're suitable for different domains.
### Spatial Relations of Meaning
What could we possibly choose as a training target?
What does it mean, meaning-in-context preserving dimensionality reduction? We want miniCOIL vectors to be comparable in their low dimensional vector space, *fruit **bat*** and *vampire **bat*** closer to each other than to *baseball **bat***, while correctly preserving input's meaning.
### Everything new is well-forgotten old
We need something to calibrate words meaning-dependent spatial relations on, when reducing the dimensionality of input context vectors. This target should reflect spatial relations though similar metric as input dense encoder (usually, COSINE similarity), be precise in capturing word meanings and reusable.
We need a target which reflects word's context in the best way. We want to be able to re-use this target for different dense encoders and we want to efficiently recycle this target for training different words.
Well, why to reinvent the wheel, when linguistics taught us that word's meaning is in its context, and there are dense encoders (usually, on the bigger side), which are able to reflect these context relationships pretty well.
Well, why to reinvent the wheel, when linguistics and based on this concept dense retrieval, teaches us that word's meaning is hidden in its context.
Let's assume that sentences containing the same words with different meanings (defined by sentence context), should cluster in vector space based on these meanings.
Let's assume that sentences sharing one word should cluster in vector space in a way, that each cluster contains sentences with this word in one specific meaning. If it's true, we could encode a humongous amount of various sentences with a sophisticated dense encoder, and form a reusable spatial target:
If it's true, we could encode humongous amount of various sentences with a sophisticated smart model, perfectly capturing contexts, and reuse this pool for TBD training.
- It's unlabeled data, so we can get truly huge amounts of it, which will help with generalization. For example, we could use CommonCrawl or OpenWebText.
- It will have to be inferenced with a heavy model once and could be used for training different encoder-based dimensionality reductions TBD.
So, will it work?
- This approach doesn't require labelled data, so we can get truly huge amounts of it, which will help with generalization. For example, we could use data from web, as [OpenWebText dataset](https://paperswithcode.com/dataset/openwebtext)
- We could use a sophisticated model once, embedding all of these sentences. Once inferenced, these embeddings can guide dimensionality reduction for various input dense encoders.
- Using sentences allows to reuse a spatial relations target for different words within one sentence. We could use word contextualized embeddings directly, however, then data preparation should have been much more complicated and required far bigger storage.
### It's Going to Work, I Bat
Let’s take a look at the word “bat”.
We took several thousands of sentences containing the word “bat”, sampled from OpenWebText.
Let’s test our assumption and take a look at the word “bat”.
We took several thousands of sentences containing the word “bat”, which we sampled from OpenWebText and encoded with a `mxbai-embed-large-v1` encoder, selected by its decent performace on MTEB benhcmark among English language models.
We encoded these sentences with a smart big transformer, which certainly knows how to capture context (`mxbai-embed-large-v1`).
Now let’s TBD them to 2D with UMAP, to see if we can visually distinguish any clusters, containing sentences where “bat” has the same meaning.
Let's project the result to 2D, to see if we can visually distinguish any clusters, containing sentences where “bat” has the same meaning.
![bat-umap](/articles_data/minicoil/bat.png)
CAPTION: Looks like a bat
Ok, so we found a working target (sentence embeddings of a powerful transformer), which will help us to guide meaning-in-a-context-preserving TBD for a transformer-of-choice input.
The result has to two big clusters related to *"bat"* as an animal and *"bat"* as a sports equipment, and two smaller ones at their intersection, related to fluttering motion and *"bat"* as verb in sports. As a bonus, 2D projection also resembles a bat.
Works for one word - right? SO let's work like this
Or course, meanings blend into each other, yet points close describe the same "type" of bats. Then it seems like we found a working target, which will help us to guide meaning-in-a-context-preserving projection for a dense encoder input of our choice.
Let's keep it simple and flexible, from the inference, training and explainability perspective. Let's learn to meaningfully compress one word. Then we can scale to all words in vocabulary that we're interested and simply combine (stack) all word encoders in one miniCOIL.
#### Training Objective
Let's continue dealing with *"bats"*. We have pool of sentences containing word *"bat"* in different meanings, from which we get an input -- *"bat"* contextualized embeddings from a dense encoder of choice -- and spatial relations target -- embedded with `mxbai-embed-large-v1` sentences.
We want to align spatial relations of compressed from input representations based on our target. For that, as a training objective, we can select minimization of the [triplet loss](https://qdrant.tech/articles/triplet-loss/). We rely on the confidence (size of margin) of a `mxbai-embed-large-v1` to guide our projection model.
![minicoil-training](/articles_data/minicoil/minicoil-training.png)
Since we're dealing with one word, it's enough to train one projection layer (Input Transformer DIM x miniCOIL vector DIM) with Tahn activation on top -- this choice of activation function is due to using COSINE similarity as measure of spatial relations.
<aside role="status">
Since miniCOIL vectors are trained to reflect spatial relationships based on a COSINE metric, they should be normalized before inserting them in bag-og-words sparse vectors, which are compared thought dot product.
</aside>
### Eating Elephant One Bite at a Time
We have an idea how to train a dimensionality reduction layer for one word. Let's keep it simple and flexible, from the inference, training and explainability perspective -- keep models on per-word level.
This will come with:
@@ -209,92 +219,85 @@ This will come with:
4. Flexibility to discover and tune underperforming words.
5. Flexibility to extend and shrink vocabulary depending on desired domain.
### A simple Encoder and no Decoders
NO NEED TO EXPLAIN WHY DON'T WE HAVE A DECODER
Yet training encoder (for dimensionality reduction) and decoder (to throw away later) from 4 dimensions to 1024 dimensions of mxbai-embed-large-v1 is too expensive -- for each word (say, we want to have 30k words in miniCOIl vocab) for each input transformer of choice.
Then, instead of learning encoder through decoder, let's directly learn to align spatial relations (calculated through Cosine Distance -- as dense encoders, as, for example, mxbai, are taught to operate on it) of compressed representations based on [triplet loss](https://qdrant.tech/articles/triplet-loss/) guided by our target.
We want to train a simple dimensionality reduction model, one per word. For that it's enough to have one layer with Tahn actication (so we can comfortably use COSINE similarity to evaluate and compare spatial relations).
TBD ENCODER PIC 512 x 4 + Tahn
Then, we use the confidence of a big model (margin between positive and negative examples) & it’s knowledge (what’s positive and what’s negative compared to anchor) to tune our downprojection layer.
![minicoil-training](/articles_data/minicoil/minicoil-training.png)
Then we can scale to all words in vocabulary that we're interested in and simply combine (stack) all word models in one miniCOIL model.
### Realization Details
TBD : make it probably into a table with specs
The code of the training approach sketched above is open sourced [in this repository](https://github.com/qdrant/miniCOIL)
**Input encoder**: jina-small (512x dim)
Specific characteristics of the existing model are the following:
**miniCOIL vocab**
| Component | Description |
|:---|:---|
| **Input Dense Encoder** | `jina-embeddings-v2-small-en` (512 dim) |
| **miniCOIL Vectors Size** | 4 dimensions. |
| **Dimensionality Reduction Layer** | 512x4 + Tanh() |
| **miniCOIL Vocabulary** | List of 30,000 most common English words, cleaned out of stop words and words of size smaller than 3 letters and stemmed, [taken from here](https://github.com/arstgit/high-frequency-vocabulary/tree/master). |
| **Target Data** | 40 million sentences encoded with `mxbai-embed-large-v1` -- a random subset of [OpenWebText dataset](https://paperswithcode.com/dataset/openwebtext). To sample training data, we uploaded sentences and their embeddings to Qdrant, and built a [full text index](https://qdrant.tech/documentation/concepts/indexing/#full-text-index) on sentences with a tokenizer `word`. |
| **Data per Word** | We sample 8000 sentences per word, calculate a cosine distance matrix between them, and use it to get triplets with a margin of at least 0.1.<br>Additionally we apply augmentation – we take a sentence and cut a target word + 1-3 (randomly chosen number) words around it, forming a new sentence. We use the same similarity score between original and augmented sentences for simplicity. This augmentation allows us to train miniCOIL to grasp context better (in big sentences it gets blended out). |
| **Training Parameters** | - Epochs: 60<br>- Optimizer: Adam with a learning rate of 1e-4<br>- Validation: 20% |
List of 30k commonly used words (after cleaning it out of stop words and 1-2 letter words + stemming)
Fun Fact: Firstly he tried to generate them with ChatGPT, aka “give me 30000 most frequently used english words”. ChatGPT generated a file with:
“Word_1
Word_2
…
Word_30000”
**Target data**
https://paperswithcode.com/dataset/openwebtext a subset of OpenWebText dataset, split into sentences. In the end we had 40 mln sentences, all uploaded to Qdrant with their mxbai-large embeddings and full text word index on sentences, so it’s easy to sample training data per word.
How many sentences do we sample per training a word stem? 8000, random sampled. We calculate a cosine distance matrix between them, and sample from it triplets for a triplet loss (anchor, positive, negative) with a margin of at least 0.1
Additionally we apply augmentation – we take a sentence and cut a target word + 1-3 (randomly chosen number) words around it, forming a new sentence. We use the same similarity score between original and augmented sentences for simplicity. This augmentation allows us to train miniCOIL to grasp context better (in big sentences it gets blended out).
**Parameters**
Epoches - 60
Trained on 1 CPU (!!! – so super cheap to train!) – takes 50 seconds per word;
Optimizer - Adam (1e-4)
Validation - 20%
**Each word was trained on 1 CPU, and it took approximately fifty seconds per word to train.**
## Results
#### Validation Loss
### Validation Loss
Theoretical difference between the small transformer (jina-small) and the “role model” transformer (mixbread-large) is 83% (in 17% jina will make a mistake in distinguishing positive and negative examples in relation to anchor)
Direct difference between the input transformer `jina-embeddings-v2-small-en` and the “role model” transformer `mxbai-embed-large-v1` is 83% (in 17% `jina-embeddings-v2-small-en` will make a mistake in distinguishing positive and negative examples in relation to anchor if compared to `mxbai-embed-large-v1`)
Our validation loss gets very close to this number – cc https://github.com/qdrant/miniCOIL/pull/7#issuecomment-2585474300
Our validation loss gets very close to this theoretical limit. Depending on output size (???? WHICH) we get from 60(76%) to 38(85%) failed triplets per batch (256).
#### Demo
![validation-loss](/articles_data/minicoil/validation_loss.png)
### Benchmarking
We’re running miniCOIL versus BM25 (our implementation, suitable for vector storage) on the BEIR benchmark.
Code of benhmark can be seen here [FOR THAT WE NEED A PR APPROVED AND MERGED].
k = 1.2, b = 0.75 (bm25 defaults), avg_len estimated on 50k documents of a dataset.
As a metric we use NDCG@10, as we're interested in ranking performance of miniCOIL compared to BM25.
It’s important to note that all sparse neural retrieval models are usually trained end-to-end on msmarco, and miniCOIL is not, as our goal was to make it as least domain- and dataset-dependent as possible, so measuring on msmarco in this case means checking out-of-domain performance
| Dataset | BM25 (NDCG@10) | MiniCOIL (NDCG@10) |
|:-----------|:--------------|:------------------|
| MS MARCO | 0.237 | 0.244 |
| NQ | - | - |
| Quora | - | - |
| FiQA-2018 | - | - |
miniCOIL performs better than BM25 in different domains, without being trained specifically on them. It shows that we’re moving in the right direction of making sparse neural retrieval usable.
ANYTHING ELSE TO SAY?
SMTH ABOUT TIME OF INFERENCE?
WE ARE GOING TO EXTEND IT? WE ARE HAPPY TO MEASURE ON ANYTHING ELSE, AREN'T WE?
### Demo
https://minicoil.qdrant.tech/ here is the demo. It uses miniCOIL vectors, projecting onto 2D first 2 [0-1] and second two [2-3] coordinates of them.
#### Benchmarks
We’re running miniCOIL versus BM25 (our implementation) on the BEIR benchmark.
k = 1.2, b = 0.75 (bm25 defaults), avg_len estimated on 50k documents of a dataset.
DO WE NEED IT HERE? DO WE NEED TO WRITE MORE? PROVIDE EXAMPLES? TELL WHICH MODEL IT USES?
The main goal was to check on msmarco (it’s the biggest one). It’s important to note that all sparse neural retrieval models are usually trained end-to-end on msmarco, and miniCOIL is not, our goal was to make it as least domain- and dataset-dependent as possible.
## Key Takeaways
BM25 NDCG@10 MsMarco 0.237
MiniCOIL NDCG@10 MsMarco 0.244
This article portays miniCOIL as an attempt to make sparse neural retrieval useable.
So, when it's the right tool for the right job?
miniCOIL performs a little better, which shows that we’re moving in the right direction with our research, dealing successfully with cons of sparse neural retrieval without being domain-dependent
If you're in a need of precise exact matching, and BM25 is not satisfying you -- you know that you corpora contains right answers and yet BM25 ranking is off, returning top results of a wrong meaning, then miniCOIL is a way to go. It works as BM25 with word meaning understanding. If you struggle to match results exactly, as they're expressed in different words, add dense encoders to retrieval.
## Conclusion
miniCOIL is beneficial to use as a part of a hybrid search system, as it enhances it without any noticeable increase in cost of usage, reusing output of a dense encoder.
How preson should get that miniCOIL is needed
### Why We Can Call it Useable
Search on exact matches and there's a part where we see troubles -- non-relevant docs at the top, ranking problem with BM25, miniCOIL is useful then. If it's an empy result -- miniCOIL won't help, you need dense.
This approach to training sparse neural encoder:
### miniCOIL strengths
1. Allow people to use transformers of their choice, which they will (regardless) use for inference in hybrid search, to cover the dense part of it. miniCOIL should be trained per transformer.
2. Simple architecture, where 1 word – 1 uber small trainable model, leads to EXTREMELY fast inference (and also super fast training) & it doesn’t weigh too much. And EXTREMELY fast and cheap training - 1 CPU 50 seconds per word.
3. No dependency on relevance objective (if document is relevant or irrelevant to the query), we don’t need labelled data for training – so (1) we can find a lot of training data (2) avoid risk of overfitting to a domain, as most sparse retrievers do with MsMarco dataset
4. This 1 word = 1 model allows for flexibility: you want to extend vocabulary of words, usable by miniCOIL, for your particular dataset – just train it for unknown words, don’t have to redo all the training.
**A useful part of a hybrid search system.** Instead of trying to replace hybrid search and dense encoders within it, miniCOIL could enhance hybrid retrieval -- and basically without significantly increasing its cost. Re-use and recycle. Dense encoders are capable of generating contextualized token representations. If we're working in a hybrid search scenario, we have to use dense encoder regardless.
![minicoil-inference](/articles_data/minicoil/minicoil-inference-hybrid.png)
1. Allows fully reusing dense encoder output, and, therefore, easily adaptable to different dense encoders. This also makes it's a free enchancement of a hybrid search solution.
2. Has a simple architecture, 1 word – 1-layer model, which leads to extremely fast inference and traing and low memory footprint.
3. Is not dependent on relevance objective, and doesn't require labelled data for training. Since it can be trained in a self-supervised manner, it's scalable and genelizeable.
4. Is fleaxible due to training on a word-level. If you want to extend miniCOIL's vocabulary for your particular use case, you could simply extend training to unknown words.
### What's next?
link to repo - commentt, like, subscribe
See it in FastEmbed, Inference (?)
More transformers
DO WE PROMISE ANYTHING?
And most importantly, you using it in practice, as we created it useable.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 77 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 75 KiB