Improve grammar in document clustering section

This commit is contained in:
Kacper Łukawski
2025-09-01 09:05:03 +02:00
parent 3e600478dc
commit aba56e0f2f
@@ -74,9 +74,9 @@ with the same number of dimensions, but the way we compute these vectors differs
#### Document clustering
Once we assigned all the document token vectors to clusters, we compute the cluster centroids by averaging all vectors
in each cluster. This gives us a representative vector for each cluster. If a cluster ends up being empty (i.e., no
vectors were assigned to it), we fill it using vector from the nearest non-empty cluster. The distance between
Once we have assigned all the document token vectors to clusters, we compute the cluster centroids by averaging all
vectors in each cluster. This gives us a representative vector for each cluster. If a cluster ends up being empty (i.e.,
no vectors were assigned to it), we fill it using vector from the nearest non-empty cluster. The distance between
clusters is determined by the Hamming distance between their cluster IDs.
![FDE document processing](/articles_data/muvera-embeddings/fde-document-processing.png)