fixes after the first review

This commit is contained in:
Евгения Суходольская
2025-05-05 12:52:53 +02:00
committed by Евгения Суходольская
parent 044ed1bad7
commit d5a31b0485
+93 -96
View File
@@ -6,7 +6,7 @@ social_preview_image: /articles_data/minicoil/social-preview.jpg
preview_dir: /articles_data/minicoil/preview
weight: -190
author: Evgeniya Sukhodolskaya
date: 2025-04-13T12:00:00+03:00
date: 2025-04-23T12:00:00+03:00
draft: false
keywords:
- hybrid search
@@ -16,22 +16,13 @@ keywords:
category: machine-learning
---
How would you describe an ideal retriever?
Have you ever heard of sparse neural retrieval? If so, have you used it in production?
It provides explainable results so we understand why we get what we get.
It is precise -- the topmost results are the most relevant.
Simultaneously, we want high recall so as not to miss relevant information.
The representation of documents in a retrieval system has a minimal memory footprint.
And, of course, the retrieval process is fast (and furious).
It's a field with excellent potential -- who would not want to use an approach combining the strengths of dense and term-based text retrieval? Yet it's not so popular. Is it due to the common curse of *"What looks good on paper is not guaranteed to work in practice?"*
An ideal retriever is the Roman Empire meme of the information retrieval field -- we all think about it daily.
Let's consider text retrieval, the oldest and the most researched area in the field. Could you name an ideal text retriever?
This article describes our step toward sparse neural retrieval *as it should be in practice* -- a field of lightweight term-based retrievers capable of distinguishing word meanings.
It seems impossible to make a universal solution even just for text -- by trying to improve one characteristic, one usually has to compromise on another. This is greatly illustrated by sparse neural retrieval -- an approach developed out of attempts to combine the strengths of dense and term-based text retrieval while getting rid of their weaknesses.
Sparse neural retrieval has great potential -- if we remove from it the burden of being a pitch-perfect for every scenario.
This article describes our step toward sparse neural retrieval *as it should be* -- a field of lightweight term-based retrievers capable of distinguishing word meanings. Learning from the mistakes of previous attempts, we created **miniCOIL**, a new sparse neural candidate to take BM25’s place in hybrid searches. We're happy to share it with you and awaiting your feedback.
Learning from the mistakes of previous attempts, we created **miniCOIL**, a new sparse neural candidate to take BM25's place in hybrid searches. We're happy to share it with you and awaiting your feedback.
## The Good, the Bad and the Ugly
@@ -46,13 +37,14 @@ Term-based retrieval usually works with a text as a bag-of-words. These words pl
Famous **BM25** estimates words' contribution based on their
1. Importance in a particular text (term frequency (TF)-based).
2. Significance within the whole corpus (Inverse Document Frequency (IDF)-based).
It's also based on several parameters reflecting typical text length in the corpus, the exact meaning of which you can check in our [detailed breakdown of the BM25 formula](https://qdrant.tech/articles/bm42/).
It also has several parameters reflecting typical text length in the corpus, the exact meaning of which you can check in [our detailed breakdown of the BM25 formula](https://qdrant.tech/articles/bm42/).
Precisely defining word importance within a text is untrivial.
BM25 is built on the idea that the term importance can be defined statistically.
It isn’t far from the truth in long texts, where frequent repetition of a certain word signals that the text is related to this concept.
Yet words have different meanings, which drastically affect the text's relevance. Think of *"fruit bat"* and *"baseball bat"*—same frequency, different meaning.
It isn't far from the truth in long texts, where frequent repetition of a certain word signals that the text is related to this concept. In very short texts (say, chunks for retrieval augmented generation (RAG)), it's less applicable, with TF of 0 or 1. We approached fixing it in our [BM42 modification of BM25 algorithm](https://qdrant.tech/articles/bm42/)
Yet there is one component of a word's importance for retrieval, which is not considered in BM25 at all -- word meaning. The same words have different meanings in different contexts, and it drastically affects the text's relevance. Think of *"fruit **bat**"* and *"baseball **bat**"*—the same importance in the text, different meanings.
### Dense Retrieval
@@ -60,134 +52,133 @@ How to capture the meaning? Bag-of-words models like BM25 assume that words are
> You shall know a word by the company it keeps
This idea, together with a motivation to numerically express word relationships, powered the development of the second branch of retrieval -- dense vectors-based.
This idea, together with a motivation to numerically express word relationships, powered the development of the second branch of retrieval -- dense vectors-based. Transformer models with attention mechanisms solved distinguishing a word's meaning within the text's context, making it a part of relevance matching in retrieval.
Perhaps you remember the famous formula of **word2vec**. *king - man + woman = queen*, quantifiable through vectors encoding these words. The meaning of different words is captured, yet the problem of distinguishing "bats" persists-**word2vec** learns only **static embeddings**. The "bat" vector representation mixes in itself everything from bloodsucking animals to sports matches.
Yet dense retrieval didn't (and can't) become a replacement for a term-based one. Dense retrievers are capable of broad semantic similarity searches, yet they lack precision when we need results including a specific keyword.
A major breakthrough in machine learning—**transformer models with attention mechanisms** (for example, **BERT**)—solved distinguishing a word's meaning within the text’s context, making it a part of relevance matching in retrieval.
What's not ideal here? Attention-based encoders combine (by pooling) word vector representations into a representation of a text chunk. This makes the similarity of texts low-interpretable and less precise. What exactly is the average between a "fruit" and a "bat"?
We could compare texts based on the similarity of each word's dense vector representation, as **late interaction models** do, but then we get as far as possible from fast retrieval at scale.
Imagine: each text chunk has at least a few dozen words, each word is encoded as a 128-dimensional vector (as in **ColBERT**), and you have a billion documents in vector storage...
It's a fool's errand -- trying to make dense retrievers do exact matching, as they're built in a paradigm that every word matches every other word semantically to some extent, and this semantic similarity depends on a training data of a particular model.
### Sparse Neural Retrieval
So, on one side, we have weak explainability, sometimes leading to frustrating retrieval results, and on the other—lightweight, explainable and fast term-based retrievers like BM25, absolutely incapable of capturing semantics.
So, on one side, we have a weak control over matching, sometimes leading to too broad retrieval results, and on the other—lightweight, explainable and fast term-based retrievers like BM25, absolutely incapable of capturing semantics.
Of course, we want the best of both worlds. A pretty standard answer here is **[hybrid search](https://qdrant.tech/articles/hybrid-search/)**, which fuses the results of dense and term-based retrievers of choice. What could even possibly be not ideal here? Well, it pulls the double duty when it comes to memory and speed.
Of course, we want the best of both worlds. A pretty standard answer here is **[hybrid search](https://qdrant.tech/articles/hybrid-search/)**, which fuses the results of dense and term-based retrievers of choice. Yet it pulls double duty when it comes to memory and speed and requires experimentation with balancing the impact of both retrievers.
Sparse neural retrieval was pushed by the idea that this drawback could also be avoided (as all others).
Sparse neural retrieval was pushed by the idea that this drawback could also be avoided.
- Why **sparse**? Term-based retrieval can operate on sparse vectors, where each word in a text is assigned a non-zero value (its importance for a text).
- Why **neural**? Instead of deriving an importance score for a word based on its statistics in a text, let's use machine learning models capable of encoding words' meaning.
Well, most oftenly perfect is the enemy of the good:)
- Why **sparse**? Term-based retrieval can operate on sparse vectors, where each word in a text is assigned a non-zero value (its importance in this text).
- Why **neural**? Instead of deriving an importance score for a word based on its statistics, let's use machine learning models capable of encoding words' meaning.
#### So Why is it Not Widely Used?
![sparse-neural-retrieval-problems](/articles_data/minicoil/models_problems.png)
[The detailed history of sparse neural retrieval makes a whole other article](https://qdrant.tech/articles/modern-sparse-neural-retrieval/). Summing a big part of it up, there were many attempts to map a word representation produced by a dense encoder to a single-valued importance score, and most of them never saw the real world outside of research papers (**DeepImpact**, **TILDEv2**, **uniCOIL**).
Squishing high-dimensional word representations (usually 768-dimensional BERT embeddings) into a single number inevitably leads to a blended word's meaning. On the other hand, stepping in a late-interaction-like direction with vector-per-word representations, as done by the authors of **Contextualized Inverted Lists**, goes against one of the main benefits of using sparse representations -- retrieval speed and a small memory footprint.
In general, trained end-to-end on a relevance objective, most of the **sparse encoders** estimated word importance well only for a particular domain. [Their out-of-domain accuracy (on datasets they hadn't "seen" during training) was worse than BM25.](https://arxiv.org/pdf/2307.10488).
Trained end-to-end on a relevance objective, most of the **sparse encoders** estimated word importance well only for a particular domain. [Their out-of-domain accuracy (on datasets they hadn't "seen" during training) was worse than BM25.](https://arxiv.org/pdf/2307.10488).
The SOTA of sparse neural retrieval is (Sparse Lexical and Expansion Model) -- **SPLADE**. This one surely made its way into retrieval systems -- you could [use SPLADE++ in Qdrant with FastEmbed](https://qdrant.tech/documentation/fastembed/fastembed-splade/).
Yet there's a catch. The "expansion" part of SPLADE's name refers to a technique against another weakness of term-based retrieval -- **vocabulary mismatch**. Where dense encoders succeed in matching a *"fruit bat"* and *"flying fox"*, term-based retrieval is powerless.
SPLADE solves this problem by expanding documents and queries with additional fitting terms. However, it leads to SPLADE representations becoming not-so-sparse (so, consequently, not lightweight) and far less explainable as expansion choices are made by machine learning models.
SPLADE solves this problem by **expanding documents and queries with additional fitting terms**. However, it leads to SPLADE inference becoming heavy, produced representations becoming not-so-sparse (so, consequently, not lightweight) and far less explainable as expansion choices are made by machine learning models.
> Big man in a suit of armor. Take that off, what are you?
Experiments showed that SPLADE without its term expansion tells the same old story of sparse encoders — [it performs worse than BM25](https://arxiv.org/pdf/2307.10488).
## Eyes on the Prize: Useable Sparse Neural Retrieval
#### Eyes on the Prize: Useable Sparse Neural Retrieval
Striving for absolute perfection, the sparse neural retrieval field either produced models performing worse than BM25 (funnily enough, [trained with BM25-based hard negatives](https://arxiv.org/pdf/2307.10488)) — so not-so-"neural", or representations which are not-so-sparse.
Striving for perfection on specific benchmarks, sparse neural retrieval field either produced models performing out-of-domain worse than BM25 (funnily enough, [trained with BM25-based hard negatives](https://arxiv.org/pdf/2307.10488)) or ones based on heavy document expansion, disrupting sparsity.
What if instead we lift this burden of being ideal from sparse neural retrievers and select and fulfil the minimal amount of criteria which a usable sparse neural retriever should fit:
To be useable in production, the minimal criteria which a sparse neural retriever should fit are:
- **It should produce sparse representations (it's in the name!).** Inheriting the perks of a term-based retrieval, it should be lightweight and simple. For a broader semantic search, there are dense retrievers, and they work.
- **It should be better than BM25 at ranking.** BM25 is a decent baseline, and the goal is to make a term-based retriever capable of distinguishing word meanings -- what BM25 can't do. Better in not just one domain, the result should be reusable, so able to generalize.
**Then it could become a useful part of a hybrid search system.** Instead of trying to replace hybrid search and dense encoders within it, sparse neural model could enhance hybrid retrieval -- and ideally, without significantly increasing its cost.
Our "what if" of this kind resulted in a **miniCOIl** sparse neural retriever.
- **Producing lightweight sparse representations (it's in the name!).** Inheriting the perks of a term-based retrieval, it should be lightweight and simple. For a broader semantic search, there are dense retrievers, and they work.
- **Being better than BM25 at ranking in different domains.** The goal is a term-based retriever capable of distinguishing word meanings -- what BM25 can't do -- and yet keep BM25 out-of-domain time-proven performance.
## miniCOIL
![minicoil](/articles_data/minicoil/minicoil.png)
STARTING FROM HERE VERY DRAFT-Y
### The Idea Behind It
To get a model performant on out-of-domain data, we should abandon the end-to-end training on relevance objective -- there is not enough data to train a model able to generalize.
#### Inspired by COIL
### Standing on the Shoulders of BM25
One of the attempts in the field of Sparse Neural Retrieval -- [Contextualized Inverted Lists (COIL)]((https://qdrant.tech/articles/modern-sparse-neural-retrieval/#sparse-neural-retriever-which-understood-homonyms)) stood out with its approach to word's importance score encoding.
BM25 is a decent baseline in various domains for many years for a reason, so why to discard it, when we can build on top of it, using the proved formula as-is.
Instead of squishing high-dimensional word representations (usually 768-dimensional BERT embeddings) into a single number, authors downprojected them to a smaller vectors (32 dimensions). These vectors are supposed to be to stored in inverted index (used in term-based retrieval) as-is, and during matching compared through dot product operation.
This way of defining the importance score captures deeper semantics -- one single number can't perfectly convey all shades of meaning one word can have. Yet this approach didn't become popular due to following problems:
- Inverted indexes are not designed for storing vectors and performing vector operations.
- When trained end-to-end with a relevance objective on [MS MARCO](https://microsoft.github.io/msmarco/), COIL's performance is heavily bound to one domain.
- Additionally, COIL works with tokens, directly reusing transformers tokenizer. Yet working directly with words is far better for term-based retrieval. Say, we want to search for a specific meaning of a word “retriever” in our documentation. COIL will break it down into `re`, `#trie` and `#ver` 32-dimensional vectors and match all three parts separately.
The best ideas don't start from zero -- let's build on top of COIL, keeping in mind what needs fixing:
1. To get a model performant on out-of-domain data, we should abandon the end-to-end training on relevance objective -- there is not enough data to train a model able to generalize.
2. We should keep representation sparse and reusable in a classic inverted index.
3. We should fix tokenization. This one is easy, as it was already done in several sparse neural retrievers, and [we also learned to do it in our BM42](https://qdrant.tech/articles/bm42/#wordpiece-retokenization).
#### Standing on the Shoulders of BM25
BM25 is a decent baseline in various domains for many years for a reason, so why to discard the time-proved formula.
The only thing it lacks is ability to distinguish word meanings.
- It doesn't see a difference between a *"fruit bat"* and a *"baseball bat"*
- When used on with word stemms, it can't distinguish parts of speech. For example, *"inform"*, *"informant"*, *"informational"*, *"informed"* and *"informally"* will be all mixed in a one *"inform"* stem.
So, instead of learning to assign importance score of words based on texts relevance, let's just add a semantic component to BM25, to make it rank better, understanding word meanings.
So, instead of learning to assign importance score of words based on texts relevance, let's just add a semantic component to BM25, to make it rank better, understanding word meanings. To learn word meanings we don't need labeled retrieval-based datasets, as word meaning can be leanred from its context -- dense encoders are a proof of this.
### Keep it Sparse
#### Keeping it Sparse
How many meanings one word can have?
TBD PICTURE OF MAKING SPARSE VECTOR OUT OF DENSE
One number is not enough to express several meanings. Using 128 (as ColBERT) and 32 (as COIL) numbers goes against sparsity criteria.
The trick is simple - a small enough dense vector can be a part of a big sparse one. Instead of doing dot product separately for each word vector, we can do it once for all of them -- on text sparse representation.
What if we use four dimensions to describe word's meaning? Then vectors would be only 4 times less sparse than a classical bag-of-words version, 1 token fitting in 4 bytes if needed.
Yet 32 values per word is a lot -- our vectors will become 32 times less sparse. How many separate meanings one word can have? Can you give many examples of a word with 32 non-intersecting meanings?
How many words have much more than 4 meanings? If they do, though, we can choose a bigger number, for example, 8.
Feels like we can use much less numbers, for example, four. Then vectors would be only 4 times less sparse than a classical bag-of-words version, 1 token fitting in 4 bytes if needed.
### Searching for Meaning
### How to Make miniCOIL
Now we're coming to the part where we need to understand how to get this 4 dimensional encapsulation of a word meaning.
Now we're coming to the part where we need to understand how to get this low dimensional encapsulation of a word meaning.
Firstly, we need to somehow get any representation of a word meaning within input, as it will vary each time based on the context Well, re-use and recycle. Dense encoders are capable of generating contextualized token representations. If we're working in a hybrid search scenario, we have to use dense encoder regardless, so why not re-use token high-dimensional representations.
To capture contextual meaning of words in input data, we get a contextuazlied word embedding from a dense encoder. Yet it's high-dimensional, and need we need to perform a meaning preserving dimensionality reduction, so the word's low dimensional representation is comparable within an encoded corpora.
Yet we need a condenced one. Then we can reformulate our task to training a model capable of a meaning preserving dimensionality reduction.
We don't want to depend on corpora and, moreover, we would want find a way to easily adapt to different dense encoders, as they're suitable for different domains.
![minicoil-inference](/articles_data/minicoil/minicoil-inference-hybrid.png)
What could we possibly choose as a training target?
### Words not Tokens!
#### Everything new is well-forgotten old
Transformers provide contextualized token embeddings, not word ones.
We need a target which reflects word's context in the best way. We want to be able to re-use this target for different dense encoders and we want to efficiently recycle this target for training different words.
Yet working directly on words is far better for retrieval, say, we want to use a word “retriever” for semantic search in our documentation, since we write about different types of retrievers and this word has several meanings within the documentation.
Well, why to reinvent the wheel, when linguistics and based on this concept dense retrieval, teaches us that word's meaning is hidden in its context.
Let's assume that sentences containing the same words with different meanings (defined by sentence context), should cluster in vector space based on these meanings.
Transformer tokenizer will break it down into re, #trie and #ver, and we wouldn’t want to work with sparse vectors having separate cells for each of 3 pieces.
If it's true, we could encode humongous amount of various sentences with a sophisticated smart model, perfectly capturing contexts, and reuse this pool for downprojection training.
- It's unlabeled data, so we can get truly huge amounts of it, which will help with generalization. For example, we could use CommonCrawl or OpenWebText.
- It will have to be inferenced with a heavy model once and could be used for training different encoder-based downprojections.
Our goal is to work on words, and we will resolve tokens into a word representation in a same way we did for BM42.
So, will it work?
### Everything new is well-forgotten old
Our goal is to learn downprojecting contextualized embeddings of one word, so the resulting four dimensional vector reflects word meaning in general, and we could match and compare it with word meanings in our corpora.
We need a target which reflects this context in the best way. We want to re-use this target for different dense encoders used in a hybrid search and we want to re-use this target for training different words.
Well, do you remember word2vec skipgramms? Word's meaning is engraved in its context.
Then let's assume that sentences containing the same words with different meanings (defined by sentence context), should cluster in vector space based on these meanings.
### It's Going to Work, I Bat
#### It's Going to Work, I Bat
Let’s take a look at the word “bat”.
We took several thousands of sentences containing the word “bat”, sampled from OpenWebText.
Let’s encode these sentences with a smart big transformer, which certainly knows how to capture context (mxbai-embed-large-v1).
We encoded these sentences with a smart big transformer, which certainly knows how to capture context (mxbai-embed-large-v1).
Now let’s downproject them to 2D with UMAP, to see if we can visually distinguish any clusters, containing sentences where “bat” has the same meaning.
![bat-umap](/articles_data/minicoil/bat.png)
Ok, so we found a working target (sentence embeddings of a powerful transformer), which will help us to guide meaning-in-a-context-preserving downprojection for a transformer-of-choice input.
### How do you eat an elephant? One Bite at a Time
#### How do you eat an elephant? One Bite at a Time
So, we see that for 1 word it seems to work out. Then let's keep it simple and flexible, from the inference, training and explainability perspective. Let's learn to meaningfully compress one word. Then we can scale to all words in vocabulary that we're interested and simply combine (stack) all word encoders in one miniCOIL.
Let's keep it simple and flexible, from the inference, training and explainability perspective. Let's learn to meaningfully compress one word. Then we can scale to all words in vocabulary that we're interested and simply combine (stack) all word encoders in one miniCOIL.
This will come with:
@@ -197,52 +188,52 @@ This will come with:
4. Flexibility to discover and tune underperforming words.
5. Flexibility to extend and shrink vocabulary depending on desired domain.
### A simple encoder and no heavy Decoders
#### A simple encoder and no heavy Decoders
Usually dimensionality reduction in this case would ask for autoencoder architecture.
Yet training encoder + decoder from 4 dimensions to 1024 dimensions of mxbai-embed-large-v1 which was chosen as a target context embedder is too expensive for each word (say, we want to have 30k words in miniCOIl vocab) for each input transformer of choice (we don’t want to stop on jina-small, in an ideal scenario).
Yet training encoder (for dimensionality reduction) and decoder (to throw away later) from 4 dimensions to 1024 dimensions of mxbai-embed-large-v1 is too expensive -- for each word (say, we want to have 30k words in miniCOIl vocab) for each input transformer of choice.
So we embed target sentences as-is with mxbai-embed-large-v1 only once (and store them in Qdrant, obviously).
Then, instead of learning encoder through decoder, let's directly learn to align spatial relations (calculated through Cosine Distance -- as dense encoders, as, for example, mxbai, are taught to operate on it) of compressed representations based on [triplet loss](https://qdrant.tech/articles/triplet-loss/) guided by our target.
And instead of learning encoder through decoder, we directly learn to align spatial relations (based on a Cosine Distance) of compressed representations based on [triplet loss](https://qdrant.tech/articles/triplet-loss/) guided by our target.
We want to train a simple dimensionality reduction model, one per word. For that it's enough to have one layer with Tahn actication (so we can comfortably use COSINE similarity to evaluate and compare spatial relations).
Basically, we use the confidence of a big model (margin between positive and negative examples) & it’s knowledge (what’s positive and what’s negative compared to anchor) to tune our downprojection.
TBD ENCODER PIC 512 x 4 + Tahn
Then, we use the confidence of a big model (margin between positive and negative examples) & it’s knowledge (what’s positive and what’s negative compared to anchor) to tune our downprojection layer.
![minicoil-training](/articles_data/minicoil/minicoil-training.png)
### Architecture & Training details
### Realization Details
For ones interested, maybe hide
TBD : make it probably into a table with specs
#### Encoder
**Input encoder**: jina-small (512x dim)
512 x 4 + Tahn (TBD Picture)
**miniCOIL vocab**
#### miniCOIL vocab
List of 30k commonly used words (after cleaning it out of stop words and 1-2 letter words + stemming)
Fun Fact: Firstly he tried to generate them with ChatGPT, aka “give me 30000 most frequently used english words”. ChatGPT generated a file with:
“Word_1
Word_2
…
Word_30000”
#### Target data
**Target data**
https://paperswithcode.com/dataset/openwebtext a subset of OpenWebText dataset, split into sentences. In the end we had 40 mln sentences, all uploaded to Qdrant with their mxbai-large embeddings and full text word index on sentences, so it’s easy to sample training data per word.
How many sentences do we sample per training a word stem? 8000, random sampled. We calculate a cosine distance matrix between them, and sample from it triplets for a triplet loss (anchor, positive, negative) with a margin of at least 0.1
Additionally we apply augmentation – we take a sentence and cut a target word + 1-3 (randomly chosen number) words around it, forming a new sentence. We use the same similarity score between original and augmented sentences for simplicity. This augmentation allows us to train miniCOIL to grasp context better (in big sentences it gets blended out).
#### Input Data
The same sentences, we filter them based on a word from the train vocabulary and generate this word’s contextualized embedding (in a sentence) as input.
#### Parameters
**Parameters**
Epoches - 60
Trained on 1 CPU (!!! – so super cheap to train!) – takes 50 seconds per word;
Optimizer - Adam (1e-4)
Validation - 20%
## Results
### Results
#### Validation Loss
@@ -250,10 +241,10 @@ Theoretical difference between the small transformer (jina-small) and the “rol
Our validation loss gets very close to this number – cc https://github.com/qdrant/miniCOIL/pull/7#issuecomment-2585474300
### Demo
#### Demo
https://minicoil.qdrant.tech/ here is the demo. It uses miniCOIL vectors, projecting onto 2D first 2 [0-1] and second two [2-3] coordinates of them.
### Benchmarks
#### Benchmarks
We’re running miniCOIL versus BM25 (our implementation) on the BEIR benchmark.
k = 1.2, b = 0.75 (bm25 defaults), avg_len estimated on 50k documents of a dataset.
@@ -264,6 +255,7 @@ MiniCOIL NDCG@10 MsMarco 0.244
miniCOIL performs a little better, which shows that we’re moving in the right direction with our research, dealing successfully with cons of sparse neural retrieval without being domain-dependent
## Conclusion
### miniCOIL strengths
1. Allow people to use transformers of their choice, which they will (regardless) use for inference in hybrid search, to cover the dense part of it. miniCOIL should be trained per transformer.
@@ -271,7 +263,12 @@ miniCOIL performs a little better, which shows that we’re moving in the right
3. No dependency on relevance objective (if document is relevant or irrelevant to the query), we don’t need labelled data for training – so (1) we can find a lot of training data (2) avoid risk of overfitting to a domain, as most sparse retrievers do with MsMarco dataset
4. This 1 word = 1 model allows for flexibility: you want to extend vocabulary of words, usable by miniCOIL, for your particular dataset – just train it for unknown words, don’t have to redo all the training.
**A useful part of a hybrid search system.** Instead of trying to replace hybrid search and dense encoders within it, miniCOIL could enhance hybrid retrieval -- and basically without significantly increasing its cost. Re-use and recycle. Dense encoders are capable of generating contextualized token representations. If we're working in a hybrid search scenario, we have to use dense encoder regardless.
![minicoil-inference](/articles_data/minicoil/minicoil-inference-hybrid.png)
### What's next?
link to repo
See it in inference (?)
More transformers