mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-04 18:38:30 +02:00
Update wording miniCOIL.md
Made changes to the wording in some of the sentences but kept the content the same.
This commit is contained in:
@@ -18,32 +18,32 @@ category: machine-learning
|
||||
|
||||
Have you ever heard of sparse neural retrieval? If so, have you used it in production?
|
||||
|
||||
It's a field with excellent potential -- who would not want to use an approach combining the strengths of dense and term-based text retrieval? Yet it's not so popular. Is it due to the common curse of *"What looks good on paper is not going work in practice"*?
|
||||
It's a field with excellent potential -- who wouldn't want to use an approach that combines the strengths of dense and term-based text retrieval? Yet it's not so popular. Is it due to the common curse of *“What looks good on paper is not going to work in practice”?*?
|
||||
|
||||
This article describes our step towards sparse neural retrieval *as it should be* -- lightweight term-based retrievers capable of distinguishing word meanings.
|
||||
This article describes our path towards sparse neural retrieval *as it should be* -- lightweight term-based retrievers capable of distinguishing word meanings.
|
||||
|
||||
Learning from the mistakes of previous attempts, we created **miniCOIL**, a new sparse neural candidate to take BM25's place in hybrid searches. We're happy to share it with you and awaiting your feedback.
|
||||
Learning from the mistakes of previous attempts, we created **miniCOIL**, a new sparse neural candidate to take BM25's place in hybrid searches. We're happy to share it with you and are awaiting your feedback.
|
||||
|
||||
## The Good, the Bad and the Ugly
|
||||
|
||||
Sparse neural retrieval is not so well known, as opposed to methods it's based on -- term-based and dense retrieval. Their weaknesses motivated this field development, guiding it's evolution. Let's follow its path.
|
||||
Sparse neural retrieval is not so well known, as opposed to methods it's based on -- term-based and dense retrieval. Their weaknesses motivated this field's development, guiding its evolution. Let's follow its path.
|
||||
|
||||
{{< figure src="/articles_data/minicoil/models_evolution.png" alt="Retrievers evolation" caption="Retrievers evolation" width="100%" >}}
|
||||
|
||||
### Term-based Retrieval
|
||||
|
||||
Term-based retrieval usually works with a text as a bag-of-words. These words play roles of different importance, contributing to the overall relevance score between a document and a query.
|
||||
Term-based retrieval usually treats text as a bag of words. These words play roles of different importance, contributing to the overall relevance score between a document and a query.
|
||||
|
||||
Famous **BM25** estimates words' contribution based on their
|
||||
Famous **BM25** estimates words' contribution based on their:
|
||||
1. Importance in a particular text -- Term Frequency (TF) based.
|
||||
2. Significance within the whole corpus -- Inverse Document Frequency (IDF) based.
|
||||
|
||||
It also has several parameters reflecting typical text length in the corpus, the exact meaning of which you can check in [our detailed breakdown of the BM25 formula](https://qdrant.tech/articles/bm42/).
|
||||
|
||||
Precisely defining word importance within a text is untrivial.
|
||||
Precisely defining word importance within a text is nontrivial.
|
||||
|
||||
BM25 is built on the idea that the term importance can be defined statistically.
|
||||
It isn't far from the truth in long texts, where frequent repetition of a certain word signals that the text is related to this concept. In very short texts -- say, chunks for Retrieval Augmented Generation (RAG) -- it's less applicable, with TF of 0 or 1. We approached fixing it in our [BM42 modification of BM25 algorithm](https://qdrant.tech/articles/bm42/).
|
||||
BM25 is built on the idea that term importance can be defined statistically.
|
||||
This isn't far from the truth in long texts, where frequent repetition of a certain word signals that the text is related to this concept. In very short texts -- say, chunks for Retrieval Augmented Generation (RAG) -- it's less applicable, with TF of 0 or 1. We approached fixing it in our [BM42 modification of BM25 algorithm.](https://qdrant.tech/articles/bm42/)
|
||||
|
||||
Yet there is one component of a word's importance for retrieval, which is not considered in BM25 at all -- word meaning. The same words have different meanings in different contexts, and it affects the text's relevance. Think of *"fruit **bat**"* and *"baseball **bat**"*—the same importance in the text, different meanings.
|
||||
|
||||
@@ -51,47 +51,47 @@ Yet there is one component of a word's importance for retrieval, which is not co
|
||||
|
||||
How to capture the meaning? Bag-of-words models like BM25 assume that words are placed in a text independently, while linguists say:
|
||||
|
||||
> "You shall know a word by the company it keeps" John Rupert Firth
|
||||
> "You shall know a word by the company it keeps" - John Rupert Firth
|
||||
|
||||
This idea, together with a motivation to numerically express word relationships, powered the development of the second branch of retrieval -- dense vectors-based. Transformer models with attention mechanisms solved distinguishing a word's meaning within the text's context, making it a part of relevance matching in retrieval.
|
||||
This idea, together with the motivation to numerically express word relationships, powered the development of the second branch of retrieval -- dense vectors. Transformer models with attention mechanisms solved the challenge of distinguishing word meanings within text context, making it a part of relevance matching in retrieval.
|
||||
|
||||
Yet dense retrieval didn't (and can't) become a complete replacement for a term-based one. Dense retrievers are capable of broad semantic similarity searches, yet they lack precision when we need results including a specific keyword.
|
||||
Yet dense retrieval didn't (and can't) become a complete replacement for term-based retreival. Dense retrievers are capable of broad semantic similarity searches, yet they lack precision when we need results including a specific keyword.
|
||||
|
||||
It's a fool's errand -- trying to make dense retrievers do exact matching, as they're built in a paradigm that every word matches every other word semantically to some extent, and this semantic similarity depends on a training data of a particular model.
|
||||
It's a fool's errand -- trying to make dense retrievers do exact matching, as they're built in a paradigm where every word matches every other word semantically to some extent, and this semantic similarity depends on the training data of a particular model.
|
||||
|
||||
### Sparse Neural Retrieval
|
||||
|
||||
So, on one side, we have a weak control over matching, sometimes leading to too broad retrieval results, and on the other—lightweight, explainable and fast term-based retrievers like BM25, incapable of capturing semantics.
|
||||
So, on one side, we have weak control over matching, sometimes leading to too broad retrieval results, and on the other—lightweight, explainable and fast term-based retrievers like BM25, incapable of capturing semantics.
|
||||
|
||||
Of course, we want the best of both worlds, fuzed in one model, no drawbacks included. Sparse neural retrieval evolution was pushed by this desire.
|
||||
Of course, we want the best of both worlds, fused in one model, no drawbacks included. Sparse neural retrieval evolution was pushed by this desire.
|
||||
|
||||
- Why **sparse**? Term-based retrieval can operate on sparse vectors, where each word in a text is assigned a non-zero value (its importance in this text).
|
||||
- Why **sparse**? Term-based retrieval can operate on sparse vectors, where each word in the text is assigned a non-zero value (its importance in this text).
|
||||
- Why **neural**? Instead of deriving an importance score for a word based on its statistics, let's use machine learning models capable of encoding words' meaning.
|
||||
|
||||
**So why is it not widely used?**
|
||||
{{< figure src="/articles_data/minicoil/models_problems.png" alt="Problems of modern sparse neural retrievers" caption="Problems of modern sparse neural retrievers" width="100%" >}}
|
||||
|
||||
The detailed history of sparse neural retrieval makes [a whole other article](https://qdrant.tech/articles/modern-sparse-neural-retrieval/). Summing a big part of it up, there were many attempts to map a word representation produced by a dense encoder to a single-valued importance score, and most of them never saw the real world outside of research papers (**DeepImpact**, **TILDEv2**, **uniCOIL**).
|
||||
The detailed history of sparse neural retrieval makes for [a whole other article](https://qdrant.tech/articles/modern-sparse-neural-retrieval/). Summing a big part of it up, there were many attempts to map a word representation produced by a dense encoder to a single-valued importance score, and most of them never saw the real world outside of research papers (**DeepImpact**, **TILDEv2**, **uniCOIL**).
|
||||
|
||||
Trained end-to-end on a relevance objective, most of the **sparse encoders** estimated word importance well only for a particular domain. Their out-of-domain accuracy, on datasets they hadn't "seen" during training, [was worse than BM25.](https://arxiv.org/pdf/2307.10488).
|
||||
Trained end-to-end on a relevance objective, most of the **sparse encoders** estimated word importance well only for a particular domain. Their out-of-domain accuracy, on datasets they hadn't "seen" during training, [was worse than BM25.](https://arxiv.org/pdf/2307.10488)
|
||||
|
||||
The SOTA of sparse neural retrieval is (Sparse Lexical and Expansion Model) -- **SPLADE**. This one surely made its way into retrieval systems -- you could [use SPLADE++ in Qdrant with FastEmbed](https://qdrant.tech/documentation/fastembed/fastembed-splade/).
|
||||
The SOTA of sparse neural retrieval is **SPLADE** -- (Sparse Lexical and Expansion Model). This model has made its way into retrieval systems - you can [use SPLADE++ in Qdrant with FastEmbed](https://qdrant.tech/documentation/fastembed/fastembed-splade/).
|
||||
|
||||
Yet there's a catch. The "expansion" part of SPLADE's name refers to a technique against another weakness of term-based retrieval -- **vocabulary mismatch**. Where dense encoders succeed in matching a *"fruit bat"* and *"flying fox"*, term-based retrieval is powerless.
|
||||
Yet there's a catch. The "expansion" part of SPLADE's name refers to a technique that combats against another weakness of term-based retrieval -- **vocabulary mismatch**. While dense encoders can successfully connect related terms like "fruit bat" and "flying fox", term-based retrieval fails at this task.
|
||||
|
||||
SPLADE solves this problem by **expanding documents and queries with additional fitting terms**. However, it leads to SPLADE inference becoming heavy. Additionally, produced representations become not-so-sparse (so, consequently, not lightweight) and far less explainable as expansion choices are made by machine learning models.
|
||||
|
||||
> "Big man in a suit of armor. Take that off, what are you?"
|
||||
|
||||
Experiments showed that SPLADE without its term expansion tells the same old story of sparse encoders — [it performs worse than BM25](https://arxiv.org/pdf/2307.10488).
|
||||
Experiments showed that SPLADE without its term expansion tells the same old story of sparse encoders — [it performs worse than BM25.](https://arxiv.org/pdf/2307.10488)
|
||||
|
||||
## Eyes on the Prize: Usable Sparse Neural Retrieval
|
||||
|
||||
Striving for perfection on specific benchmarks, the sparse neural retrieval field either produced models performing out-of-domain worse than BM25 (ironically, [trained with BM25-based hard negatives](https://arxiv.org/pdf/2307.10488)) or ones based on heavy document expansion, lowering sparsity.
|
||||
Striving for perfection on specific benchmarks, the sparse neural retrieval field either produced models performing worse than BM25 out-of-domain(ironically, [trained with BM25-based hard negatives](https://arxiv.org/pdf/2307.10488)) or models based on heavy document expansion, lowering sparsity.
|
||||
|
||||
So, to be usable in production, the minimal criteria a sparse neural retriever should meet are:
|
||||
To be usable in production, the minimal criteria a sparse neural retriever should meet are:
|
||||
|
||||
- **Producing lightweight sparse representations (it's in the name!).** Inheriting the perks of term-based retrieval, it should be lightweight and simple. For broader semantic search, there are dense retrievers, and they work.
|
||||
- **Producing lightweight sparse representations (it's in the name!).** Inheriting the perks of term-based retrieval, it should be lightweight and simple. For broader semantic search, there are dense retrievers.
|
||||
- **Being better than BM25 at ranking in different domains.** The goal is a term-based retriever capable of distinguishing word meanings — what BM25 can't do — preserving BM25's out-of-domain, time-proven performance.
|
||||
|
||||
{{< figure src="/articles_data/minicoil/minicoil.png" alt="The idea behind miniCOIL" caption="The idea behind miniCOIL" width="100%" >}}
|
||||
@@ -102,15 +102,17 @@ One of the attempts in the field of Sparse Neural Retrieval — [Contextualized
|
||||
|
||||
Instead of squishing high-dimensional token representations (usually 768-dimensional BERT embeddings) into a single number, COIL authors project them to smaller vectors of 32 dimensions. They propose storing these vectors in **inverted lists** of an **inverted index** (used in term-based retrieval) as is and comparing vector representations through dot product.
|
||||
|
||||
This approach captures deeper semantics — one number can't perfectly convey all shades of meaning that one word can have. Yet it didn't catch on, and COIL didn't become popular, presumably due to the following reasons:
|
||||
This approach captures deeper semantics, a single number simply cannot convey all the nuanced meanings a word can have.
|
||||
|
||||
Despite this advantage, COIL failed to gain widespread adoption for several key reasons:
|
||||
|
||||
- Inverted indexes are usually not designed to store vectors and perform vector operations.
|
||||
- Trained end-to-end with a relevance objective on [MS MARCO dataset](https://microsoft.github.io/msmarco/), COIL's performance is heavily domain-bound.
|
||||
- Additionally, COIL operates on tokens, reusing BERT's tokenizer. Yet, working at a word level is far better for term-based retrieval. Say we want to search for a *"retriever"* in our documentation. COIL will break it down into `re`, `#trie`, and `#ver` 32-dimensional vectors and match all three parts separately -- not so convenient.
|
||||
- Additionally, COIL operates on tokens, reusing BERT's tokenizer. However, working at a word level is far better for term-based retrieval. Imagine we want to search for a *"retriever"* in our documentation. COIL will break it down into `re`, `#trie`, and `#ver` 32-dimensional vectors and match all three parts separately -- not so convenient.
|
||||
|
||||
However, COIL representations allow distinguishing homographs, a skill BM25 lacks. The best ideas don't start from zero. We could try to **build on top of COIL, keeping in mind what needs fixing**:
|
||||
However, COIL representations allow distinguishing homographs, a skill BM25 lacks. The best ideas don't start from zero. We propose an approach **built on top of COIL, keeping in mind what needs fixing**:
|
||||
|
||||
1. To get a model performant on out-of-domain data, we should **abandon end-to-end training on a relevance objective** — there is not enough data to train a model able to generalize.
|
||||
1. We should **abandon end-to-end training on a relevance objective** to get a model performant on out-of-domain data. There is not enough data to train a model able to generalize.
|
||||
2. We should **keep representations sparse and reusable in a classic inverted index**.
|
||||
3. We should **fix tokenization**. This problem is the easiest one to solve, as it was already done in several sparse neural retrievers, and [we also learned to do it in our BM42](https://qdrant.tech/articles/bm42/#wordpiece-retokenization).
|
||||
|
||||
@@ -126,7 +128,7 @@ $$
|
||||
|
||||
Then, if we manage to capture a word's meaning, our solution alone could work like BM25 combined with a semantically aware reranker -- or, in other words:
|
||||
|
||||
- It could see the difference between homophrapgs;
|
||||
- It could see the difference between homographs;
|
||||
- When used with word stems, it could distinguish parts of speech.
|
||||
|
||||
{{< figure src="/articles_data/minicoil/examples.png" alt="Meaning component" caption="Meaning component" width="100%" >}}
|
||||
@@ -138,7 +140,7 @@ And if our model stumbles upon a word it hasn't "seen" during training, we can j
|
||||
COIL uses 32 values to describe one term. Do we need this many? How many words with 32 separate meanings could we name without additional research?
|
||||
|
||||
Yet, even if we use fewer values in COIL representations, the initial problem of dense vectors not fitting into a classical inverted index persists.
|
||||
Unless... We do a simple trick!
|
||||
Unless... We perform a simple trick!
|
||||
|
||||
{{< figure src="/articles_data/minicoil/bow_4D.png" alt="miniCOIL vectors to sparse representation" caption="miniCOIL vectors to sparse representation" width="80%" >}}
|
||||
|
||||
@@ -156,7 +158,7 @@ We want to work smarter, not harder, and rely as much as possible on time-proven
|
||||
|
||||
### Reducing Dimensions
|
||||
|
||||
Yet dense encoder outputs are high-dimensional, so we need to perform **dimensionality reduction, which should preserve the word's meaning in context**. The goal is to:
|
||||
Dense encoder outputs are high-dimensional, so we need to perform **dimensionality reduction, which should preserve the word's meaning in context**. The goal is to:
|
||||
|
||||
- Avoid relevance objective and dependence on labelled datasets;
|
||||
- Find reflecting meaning spatial relations target;
|
||||
@@ -178,7 +180,7 @@ We took several thousand sentences with this word, which we sampled from [OpenWe
|
||||
|
||||
{{< figure src="/articles_data/minicoil/bat.png" alt="Sentences with \"bat\" in 2D" caption="Sentences with \"bat\" in 2D. <br>A very important observation: *Looks like a bat*:)" width="80%" >}}
|
||||
|
||||
The result had two big clusters related to *"bat"* as an animal and *"bat"* as a sports equipment, and two smaller ones related to fluttering motion and verb used in sports. Seems like it could work!
|
||||
The result had two big clusters related to *"bat"* as an animal and *"bat"* as a sports equipment, and two smaller ones related to fluttering motion and the verb used in sports. Seems like it could work!
|
||||
|
||||
### Architecture and Training Objective
|
||||
|
||||
@@ -200,7 +202,7 @@ Since miniCOIL vectors are trained to reflect spatial relationships based on cos
|
||||
|
||||
#### Eating Elephant One Bite at a Time
|
||||
|
||||
Now, we have a full idea of how to train miniCOIL for one word. How do we scale to a whole vocabulary?
|
||||
Now, we have the full idea of how to train miniCOIL for one word. How do we scale to a whole vocabulary?
|
||||
|
||||
What if we keep it simple and continue training a model per word? It has certain benefits:
|
||||
|
||||
@@ -224,7 +226,7 @@ Here are the specific characteristics of the miniCOIL model we trained based on
|
||||
|:---|:---|
|
||||
| **Input Dense Encoder** | [`jina-embeddings-v2-small-en`](https://huggingface.co/jinaai/jina-embeddings-v2-small-en) (512 dimensions) |
|
||||
| **miniCOIL Vectors Size** | 4 dimensions |
|
||||
| **miniCOIL Vocabulary** | List of 30,000 most common English words, cleaned of stop words and words shorter than 3 letters, [taken from here](https://github.com/arstgit/high-frequency-vocabulary/tree/master). Words are stemmed to align miniCOIL with our BM25 implementation. |
|
||||
| **miniCOIL Vocabulary** | List of 30,000 of the most common English words, cleaned of stop words and words shorter than 3 letters, [taken from here](https://github.com/arstgit/high-frequency-vocabulary/tree/master). Words are stemmed to align miniCOIL with our BM25 implementation. |
|
||||
| **Training Data** | 40 million sentences — a random subset of the [OpenWebText dataset](https://paperswithcode.com/dataset/openwebtext). To make sampling convenient, we uploaded sentences and their [`mxbai-embed-large-v1`](https://huggingface.co/mixedbread-ai/mxbai-embed-large-v1) embeddings to Qdrant and built a [full-text payload index](https://qdrant.tech/documentation/concepts/indexing/#full-text-index) on sentences with a tokenizer of type `word`. |
|
||||
| **Training Data per Word** | We sample 8000 sentences per word and form triplets with a margin of at least **0.1**.<br>Additionally, we apply **augmentation** — take a sentence and cut out the target word plus its 1–3 neighbours. We reuse the same similarity score between original and augmented sentences for simplicity. |
|
||||
| **Training Parameters** | **Epochs**: 60<br>**Optimizer**: Adam with a learning rate of 1e-4<br>**Validation set**: 20% |
|
||||
@@ -250,10 +252,10 @@ To check our 4D miniCOIL version performance in different domains, we, ironicall
|
||||
|
||||
We're testing our 4D miniCOIL model versus [our BM25 implementation](https://huggingface.co/Qdrant/bm25). BEIR datasets are indexed to Qdrant using the following parameters for both methods:
|
||||
- `k = 1.2`, `b = 0.75` default values recommended to use with BM25 scoring;
|
||||
- `avg_len` estimated on 50_000 documents from a respective dataset.
|
||||
- `avg_len` estimated on 50,000 documents from a respective dataset.
|
||||
|
||||
<aside role="status">
|
||||
BM25 results depend on implementation details, such as the choice of stemmer, tokenizer, stop word list, etc. So, to make miniCOIL comparable to BM25, we use our own BM25 implementation and reuse all its implementation choices for miniCOIL.
|
||||
BM25 results depend on implementation details, such as the choice of stemmer, tokenizer, stop word list, etc. To make miniCOIL comparable to BM25, we use our own BM25 implementation and reuse all its implementation choices for miniCOIL.
|
||||
</aside>
|
||||
|
||||
We compare models based on the `NDCG@10` metric, as we're interested in the ranking performance of miniCOIL compared to BM25. Both retrieve the same subset of indexed documents based on exact matches, but miniCOIL should ideally rank this subset better based on its semantics understanding.
|
||||
@@ -278,7 +280,7 @@ To use any model for your specific use case, always benchmark it yourself!<br> P
|
||||
|
||||
This article describes our attempt to make a lightweight sparse neural retriever that is able to generalize to out-of-domain data. Sparse neural retrieval has a lot of potential, and we hope to see it gain more traction.
|
||||
|
||||
### Why this Approach can be Called Usable?
|
||||
### Why is this Approach Useful?
|
||||
|
||||
This approach to training sparse neural retrievers:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user