diff --git a/qdrant-landing/content/articles/bm42.md b/qdrant-landing/content/articles/bm42.md index 81c074dd5..e60107f28 100644 --- a/qdrant-landing/content/articles/bm42.md +++ b/qdrant-landing/content/articles/bm42.md @@ -67,11 +67,11 @@ The last component of the formula can be intuitively interpreted as **term impor This might look a bit complicated, so let's break it down. $$ -\text{Term importance in document }(q_i) = \color{red}\frac{f(q_i, D)\color{black} \cdot \color{blue}(k_1 + 1) \color{black} }{\color{red}f(q_i, D)\color{black} + \color{blue}k_1\color{black} \cdot \left(1 - \color{blue}b\color{black} + \color{blue}b\color{black} \cdot \frac{|D|}{\text{avgdl}}\right)} +\text{Term importance in document }(q_i) = \color{red}\frac{f(q_i, D)\color{gray} \cdot \color{blue}(k_1 + 1) \color{gray} }{\color{red}f(q_i, D)\color{gray} + \color{blue}k_1\color{gray} \cdot \left(1 - \color{blue}b\color{gray} + \color{blue}b\color{gray} \cdot \frac{|D|}{\text{avgdl}}\right)} $$ -- The $\color{red}f(q_i, D)\color{black}$ - is the frequency of the term $q_i$ in the document $D$. Or in other words, the number of times the term $q_i$ appears in the document $D$. -- The $\color{blue}k_1\color{black}$ and $\color{blue}b\color{black}$ are the hyperparameters of the BM25 formula. In most implementations, they are constants set to $k_1=1.5$ and $b=0.75$. Those constants define relative implications of the term frequency and the document length in the formula. +- The $\color{red}f(q_i, D)\color{gray}$ - is the frequency of the term $q_i$ in the document $D$. Or in other words, the number of times the term $q_i$ appears in the document $D$. +- The $\color{blue}k_1\color{gray}$ and $\color{blue}b\color{gray}$ are the hyperparameters of the BM25 formula. In most implementations, they are constants set to $k_1=1.5$ and $b=0.75$. Those constants define relative implications of the term frequency and the document length in the formula. - The $\frac{|D|}{\text{avgdl}}$ - is the relative length of the document $D$ compared to the average document length in the corpora. The intuition befind this part is following: if the token is found in the smaller document, it is more likely that this token is important for this document. #### Will BM25 term importance in the document work for RAG? @@ -168,7 +168,7 @@ weights = torch.mean(attentions[-1][0,:,0], axis=0) # │ │ │ └─── [CLS] token is the first one # │ │ └─────── First item of the batch # │ └────────── Last transformer layer -# └────────────────────────── Averate all 6 attention heads +# └────────────────────────── Average all 6 attention heads for weight, token in zip(weights, tokens): print(f"{token}: {weight}") diff --git a/qdrant-landing/static/articles_data/modern-sparse-neural-retrieval/SPLADE++.png b/qdrant-landing/static/articles_data/modern-sparse-neural-retrieval/SPLADE++.png index 2e7ff28f5..ff7761833 100644 Binary files a/qdrant-landing/static/articles_data/modern-sparse-neural-retrieval/SPLADE++.png and b/qdrant-landing/static/articles_data/modern-sparse-neural-retrieval/SPLADE++.png differ