fix ML articles

This commit is contained in:
Евгения Суходольская
2025-04-10 12:02:16 +02:00
parent 407478c630
commit 55491a538d
2 changed files with 4 additions and 4 deletions
+4 -4
View File
@@ -67,11 +67,11 @@ The last component of the formula can be intuitively interpreted as **term impor
This might look a bit complicated, so let's break it down. This might look a bit complicated, so let's break it down.
$$ $$
\text{Term importance in document }(q_i) = \color{red}\frac{f(q_i, D)\color{black} \cdot \color{blue}(k_1 + 1) \color{black} }{\color{red}f(q_i, D)\color{black} + \color{blue}k_1\color{black} \cdot \left(1 - \color{blue}b\color{black} + \color{blue}b\color{black} \cdot \frac{|D|}{\text{avgdl}}\right)} \text{Term importance in document }(q_i) = \color{red}\frac{f(q_i, D)\color{gray} \cdot \color{blue}(k_1 + 1) \color{gray} }{\color{red}f(q_i, D)\color{gray} + \color{blue}k_1\color{gray} \cdot \left(1 - \color{blue}b\color{gray} + \color{blue}b\color{gray} \cdot \frac{|D|}{\text{avgdl}}\right)}
$$ $$
- The $\color{red}f(q_i, D)\color{black}$ - is the frequency of the term $q_i$ in the document $D$. Or in other words, the number of times the term $q_i$ appears in the document $D$. - The $\color{red}f(q_i, D)\color{gray}$ - is the frequency of the term $q_i$ in the document $D$. Or in other words, the number of times the term $q_i$ appears in the document $D$.
- The $\color{blue}k_1\color{black}$ and $\color{blue}b\color{black}$ are the hyperparameters of the BM25 formula. In most implementations, they are constants set to $k_1=1.5$ and $b=0.75$. Those constants define relative implications of the term frequency and the document length in the formula. - The $\color{blue}k_1\color{gray}$ and $\color{blue}b\color{gray}$ are the hyperparameters of the BM25 formula. In most implementations, they are constants set to $k_1=1.5$ and $b=0.75$. Those constants define relative implications of the term frequency and the document length in the formula.
- The $\frac{|D|}{\text{avgdl}}$ - is the relative length of the document $D$ compared to the average document length in the corpora. The intuition befind this part is following: if the token is found in the smaller document, it is more likely that this token is important for this document. - The $\frac{|D|}{\text{avgdl}}$ - is the relative length of the document $D$ compared to the average document length in the corpora. The intuition befind this part is following: if the token is found in the smaller document, it is more likely that this token is important for this document.
#### Will BM25 term importance in the document work for RAG? #### Will BM25 term importance in the document work for RAG?
@@ -168,7 +168,7 @@ weights = torch.mean(attentions[-1][0,:,0], axis=0)
# │ │ │ └─── [CLS] token is the first one # │ │ │ └─── [CLS] token is the first one
# │ │ └─────── First item of the batch # │ │ └─────── First item of the batch
# │ └────────── Last transformer layer # │ └────────── Last transformer layer
# └────────────────────────── Averate all 6 attention heads # └────────────────────────── Average all 6 attention heads
for weight, token in zip(weights, tokens): for weight, token in zip(weights, tokens):
print(f"{token}: {weight}") print(f"{token}: {weight}")
Binary file not shown.

Before

Width:  |  Height:  |  Size: 1.2 MiB

After

Width:  |  Height:  |  Size: 1.1 MiB