add fastembed docs

This commit is contained in:
davidmyriel
2024-08-05 18:54:29 -07:00
parent 6d1f7e9aec
commit 3ed1bde499
5 changed files with 596 additions and 0 deletions
@@ -0,0 +1,77 @@
---
title: "FastEmbed"
weight: 6
---
# What is FastEmbed?
FastEmbed is a lightweight Python library built for embedding generation. It supports popular embedding models and offers a user-friendly experience for embedding data into vector space.
By using FastEmbed, you can ensure that your embedding generation process is not only fast and efficient but also highly accurate, meeting the needs of various machine learning and natural language processing applications.
FastEmbed easily integrates with Qdrant for a variety of multimodal search purposes.
## How to get started with FastEmbed
|Beginner|Advanced|
|:-:|:-:|
|[Generate Text Embedings with FastEmbed](fastembed-quickstart/)|[Combine FastEmbed with Qdrant for Vector Search](fastembed-semantic-search/)|
## Why is FastEmbed popular?
- Light & Fast: FastEmbed is designed to be efficient and quick, utilizing quantized model weights to reduce the computational load.
- ONNX Runtime for Inference: The library leverages ONNX Runtime, which allows for fast inference on various hardware platforms, ensuring high performance.
- Accuracy and Recall: FastEmbed aims to provide better accuracy and recall compared to other models like OpenAI’s Ada-002. The default embedding model is `BAAI/bge-small-en-v1.5`, which has shown strong results on the MTEB leaderboard.
- Supported Models: FastEmbed supports many models, including multilingual ones, to cater to a wide range of use cases.
## Supported Models
We are continuously updating the list with better and useful models. Please note that this list changes often.
### Text Embedding Models
| model | dim | description | size_in_GB |
|--------------------------------------------|-----|---------------------------------------------------------------------------------|------------|
| BAAI/bge-small-en-v1.5 | 384 | Fast and Default English model | 0.067 |
| BAAI/bge-small-zh-v1.5 | 512 | Fast and recommended Chinese model | 0.090 |
| snowflake/snowflake-arctic-embed-xs | 384 | Based on all-MiniLM-L6-v2 model with only 22m ... | 0.090 |
| sentence-transformers/all-MiniLM-L6-v2 | 384 | Sentence Transformer model, MiniLM-L6-v2 | 0.090 |
| jinaai/jina-embeddings-v2-small-en | 512 | English embedding model supporting 8192 sequen... | 0.120 |
| BAAI/bge-small-en | 384 | Fast English model | 0.130 |
| snowflake/snowflake-arctic-embed-s | 384 | Based on infloat/e5-small-unsupervised, does n... | 0.130 |
| nomic-ai/nomic-embed-text-v1.5-Q | 768 | Quantized 8192 context length english model | 0.130 |
| BAAI/bge-base-en-v1.5 | 768 | Base English model, v1.5 | 0.210 |
| sentence-transformers/paraphrase-multilingual-... | 384 | Sentence Transformer model, paraphrase-multili... | 0.220 |
| Qdrant/clip-ViT-B-32-text | 512 | CLIP text encoder | 0.250 |
| jinaai/jina-embeddings-v2-base-de | 768 | German embedding model supporting 8192 sequenc... | 0.320 |
| BAAI/bge-base-en | 768 | Base English model | 0.420 |
| snowflake/snowflake-arctic-embed-m | 768 | Based on intfloat/e5-base-unsupervised model, ... | 0.430 |
| nomic-ai/nomic-embed-text-v1.5 | 768 | 8192 context length english model | 0.520 |
| jinaai/jina-embeddings-v2-base-en | 768 | English embedding model supporting 8192 sequen... | 0.520 |
| nomic-ai/nomic-embed-text-v1 | 768 | 8192 context length english model | 0.520 |
| snowflake/snowflake-arctic-embed-m-long | 768 | Based on nomic-ai/nomic-embed-text-v1-unsuperv... | 0.540 |
| mixedbread-ai/mxbai-embed-large-v1 | 1024 | MixedBread Base sentence embedding model, does... | 0.640 |
| jinaai/jina-embeddings-v2-base-code | 768 | Source code embedding model supporting 8192 se... | 0.640 |
| sentence-transformers/paraphrase-multilingual-... | 768 | Sentence-transformers model for tasks like clu... | 1.000 |
| snowflake/snowflake-arctic-embed-l | 1024 | Based on intfloat/e5-large-unsupervised, large... | 1.020 |
| thenlper/gte-large | 1024 | Large general text embeddings model | 1.200 |
| BAAI/bge-large-en-v1.5 | 1024 | Large English model, v1.5 | 1.200 |
| intfloat/multilingual-e5-large | 1024 | Multilingual model, e5-large. Recommend using ... | 2.240 |
### Image Embedding Models
| model | dim | description | size_in_GB |
|-----------------------------|------|------------------------------------------------------|------------|
| Qdrant/resnet50-onnx | 2048 | ResNet-50 from `Deep Residual Learning for Ima... | 0.10 |
| Qdrant/clip-ViT-B-32-vision | 512 | CLIP vision encoder based on ViT-B/32 | 0.34 |
| Qdrant/Unicom-ViT-B-32 | 512 | Unicom Unicom-ViT-B-32 from open-metric-learning | 0.48 |
| Qdrant/Unicom-ViT-B-16 | 768 | Unicom Unicom-ViT-B-16 from open-metric-learning | 0.82 |
### Sparse Embedding Models
| model | vocab_size | description | size_in_GB | requires_idf |
|---------------------------------------------|------------|-----------------------------------------------------------------------------------|------------|--------------|
| Qdrant/bm25 | NaN | BM25 as sparse embeddings meant to be used with... | 0.010 | True |
| Qdrant/bm42-all-minilm-l6-v2-attentions | 30522.0 | Light sparse embedding model, which assigns an... | 0.090 | True |
| prithvida/Splade_PP_en_v1 | 30522.0 | Misspelled version of the model. Retained for... | 0.532 | NaN |
| prithivida/Splade_PP_en_v1 | 30522.0 | Independent Implementation of SPLADE++ Model for dense and sparse retrieval tasks | 0.532 | NaN |
### Late Interaction Embedding Models
| model | dim | description | size_in_GB |
|--------------------------|-----|-------------------------|------------|
| colbert-ir/colbertv2.0 | 128 | Late interaction model | 0.44 |
@@ -0,0 +1,146 @@
---
title: Working with ColBERT
weight: 6
---
# How to Generate ColBERT Multivectors with FastEmbed
With FastEmbed, you can use ColBERT to generate multivector embeddings. ColBERT is a powerful retrieval model that combines the strength of BERT embeddings with efficient late interaction techniques. FastEmbed will provide you with an optimized pipeline to utilize these embeddings in your search tasks.
Please note that ColBERT requires more resources than other no-interaction models. We recommend you use ColBERT as a re-ranker instead of a first-stage retriever.
The first-stage retriever can retrieve 100-500 examples. This task would be done by a simpler model. Then, you can rank the leftover results with ColBERT.
## Setup
This command imports all late interaction models for text embedding.
```python
from fastembed import LateInteractionTextEmbedding
```
You can list which models are supported in your version of FastEmbed.
```python
LateInteractionTextEmbedding.list_supported_models()
```
This command displays the available models. The output shows details about the ColBERT model, including its dimensions, description, size, sources, and model file.
```python
[{'model': 'colbert-ir/colbertv2.0',
'dim': 128,
'description': 'Late interaction model',
'size_in_GB': 0.44,
'sources': {'hf': 'colbert-ir/colbertv2.0'},
'model_file': 'model.onnx'}]
```
Now, load the model.
```python
embedding_model = LateInteractionTextEmbedding("colbert-ir/colbertv2.0")
```
The model files will be fetched and downloaded, with progress showing.
## Embed data
First, you need to define both documents and queries.
```python
documents = [
"ColBERT is a late interaction text embedding model, however, there are also other models such as TwinBERT.",
"On the contrary to the late interaction models, the early interaction models contains interaction steps at embedding generation process",
]
queries = [
"Are there any other late interaction text embedding models except ColBERT?",
"What is the difference between late interaction and early interaction text embedding models?",
]
```
**Note:** ColBERT computes document and query embeddings differently. Make sure to use the corresponding methods.
Now, create embeddings from both documents and queries.
```python
document_embeddings = list(
embedding_model.embed(documents)
) # embed and qury_embed return generators,
# which we need to evaluate by writing them to a list
query_embeddings = list(embedding_model.query_embed(queries))
```
Display the shapes of document and query embeddings.
```python
document_embeddings[0].shape, query_embeddings[0].shape
```
You should get something like this:
```python
((26, 128), (32, 128))
```
Don't worry about query embeddings having the bigger shape in this case. ColBERT authors recommend to pad queries with [MASK] tokens to 32 tokens. They also recommends to truncate queries to 32 tokens, however we don't do that in FastEmbed, so you can put some straight into the queries.
## Compute similarity
This function calculates the relevance scores using the MaxSim operator, sorts the documents based on these scores, and returns the indices of the top-k documents.
```python
import numpy as np
def compute_relevance_scores(query_embedding: np.array, document_embeddings: np.array, k: int):
"""
Compute relevance scores for top-k documents given a query.
:param query_embedding: Numpy array representing the query embedding, shape: [num_query_terms, embedding_dim]
:param document_embeddings: Numpy array representing embeddings for documents, shape: [num_documents, max_doc_length, embedding_dim]
:param k: Number of top documents to return
:return: Indices of the top-k documents based on their relevance scores
"""
# Compute batch dot-product of query_embedding and document_embeddings
# Resulting shape: [num_documents, num_query_terms, max_doc_length]
scores = np.matmul(query_embedding, document_embeddings.transpose(0, 2, 1))
# Apply max-pooling across document terms (axis=2) to find the max similarity per query term
# Shape after max-pool: [num_documents, num_query_terms]
max_scores_per_query_term = np.max(scores, axis=2)
# Sum the scores across query terms to get the total score for each document
# Shape after sum: [num_documents]
total_scores = np.sum(max_scores_per_query_term, axis=1)
# Sort the documents based on their total scores and get the indices of the top-k documents
sorted_indices = np.argsort(total_scores)[::-1][:k]
return sorted_indices
```
Calculate sorted indices.
```python
sorted_indices = compute_relevance_scores(
np.array(query_embeddings[0]), np.array(document_embeddings), k=3
)
print("Sorted document indices:", sorted_indices)
```
The output shows the sorted document indices based on the relevance to the query.
```python
Sorted document indices: [0 1]
```
## Show results
```python
print(f"Query: {queries[0]}")
for index in sorted_indices:
print(f"Document: {documents[index]}")
```
The query and corresponding sorted documents are displayed, showing the relevance of each document to the query.
```bash
Query: Are there any other late interaction text embedding models except ColBERT?
Document: ColBERT is a late interaction text embedding model, however, there are also other models such as TwinBERT.
Document: On the contrary to the late interaction models, the early interaction models contains interaction steps at embedding generation process
```
@@ -0,0 +1,107 @@
---
title: "Quickstart"
weight: 2
---
# How to Generate Text Embedings with FastEmbed
## Install FastEmbed
```python
pip install fastembed
```
Just for demo purposes, you will use Lists and NumPy to work with sample data.
```python
from typing import List
import numpy as np
```
## Load default model
In this example, you will use the default text embedding model, `BAAI/bge-small-en-v1.5`.
```python
from fastembed import TextEmbedding
```
## Add sample data
Now, add two sample documents. Your documents must be in a list, and each document must be a string
```python
documents: List[str] = [
"FastEmbed is lighter than Transformers & Sentence-Transformers.",
"FastEmbed is supported by and maintained by Qdrant.",
]
```
Download and initialize the model. Print a message to verify the process.
```python
embedding_model = TextEmbedding()
print("The model BAAI/bge-small-en-v1.5 is ready to use.")
```
## Embed data
Generate embeddings for both documents.
```python
embeddings_generator = embedding_model.embed(documents)
embeddings_list = list(embeddings_generator)
len(embeddings_list[0])
```
Here is the sample document list. The default model creates vectors with 384 dimensions.
```bash
Document: This is built to be faster and lighter than other embedding libraries e.g. Transformers, Sentence-Transformers, etc.
Vector of type: <class 'numpy.ndarray'> with shape: (384,)
Document: fastembed is supported by and maintained by Qdrant.
Vector of type: <class 'numpy.ndarray'> with shape: (384,)
```
## Verify embedding shape
Next, take the embeddings and convert them into a NumPy array.
Then, check the shape of this array to confirm the dimensions of the generated embeddings.
```python
embeddings_list = np.array(list(embedding_model.embed(documents)))
embeddings_list.shape
```
The shape will be a tuple indicating the number of documents and the dimensions of the embeddings. For example, if there are 2 documents and each embedding has 384 dimensions, the shape will be (2, 384).
```python
(2, 384)
```
## See the embeddings
```python
print("Embeddings:\n", embeddings_list)
```
The embeddings don't look too interesting, but here is a visual.
```bash
Embeddings:
[[-0.11154681 0.00976555 0.00524559 0.01951888 -0.01934952 0.02943449
-0.10519084 -0.00890122 0.01831438 0.01486796 -0.05642502 0.02561352
-0.00120165 0.00637456 0.02633459 0.0089221 0.05313658 0.03955453
-0.04400245 -0.02929407 0.04691846 -0.02515868 0.00778646 -0.05410657
-0.04362096 0.01275103 -0.02304643 -0.02250824 0.01992303 -0.19920594
0.01895355 -0.02651559 0.08252289 -0.02281932 -0.05781299 -0.01367694
-0.0357074 0.05386209 -0.10155275 0.02097272 0.00652704 0.03738071
-0.03508432 -0.00091192 0.03425014 -0.02445892 -0.00739993 -0.03084338
-0.03605132 -0.02814267 0.01547157 -0.0215851 0.0254423 0.02438808
-0.02049713 -0.02665791 0.03727325 0.08809511 0.02471923 -0.0085119
-0.0201654 0.05734601 -0.05127536 0.02370393 -0.02993909 -0.02091447
-0.03927833 -0.02315613 0.00927736 0.07043561 0.02359607 -0.01002697
0.04758466 0.00795123 0.00199829 0.05515845 0.01843713 -0.00297784
-0.01339653 -0.00348842 0.07951678 -0.06366211 -0.01295009 -0.00558423
-0.02937807 -0.03242922 0.03339408 0.01734223 0.01248069 0.06825151
-0.08455556 -0.04574722 -0.0095825 0.05233807 -0.00915513 -0.06187493
-0.01228709 0.033805 0.03197376 0.35104704 -0.06548341 -0.00566747
0.05857922 -0.05616153 0.01007555 -0.03780003 0.04616216 -0.01843753
-0.09301556 -0.02079745 -0.04339451 -0.02648945 -0.02158683 0.02083867
0.02295375 0.01272577 -0.03758859 0.00559671 -0.06286892 0.04843555
0.00935006 0.01959503 0.00112758 -0.01437079 0.01551349 0.03026596
0.04408653 0.0547558 0.0532619 0.07697894 -0.01663396 0.01228552
-0.01277487 0.03190343 0.02996432 0.04491728 -0.02979672 -0.0327898
0.01730017 -0.00223152 -0.04579118 0.11361726 0.01470178 -0.04264669
...
-0.00243012 -0.01820582 0.02938612 0.02108984 -0.02178085 0.02971899
-0.00790564 0.03561783 0.0652488 -0.04371546 -0.05550042 0.02651665
-0.01116153 -0.01682246 -0.05976734 -0.03143916 0.06522726 0.01801389
-0.02611006 0.01627177 -0.0368538 0.03968835 0.027597 0.03305927]]
```
@@ -0,0 +1,70 @@
---
title: "FastEmbed & Qdrant"
weight: 3
---
# Using FastEmbed with Qdrant for Vector Search
## Install Qdrant Client
The Qdrant Python Client has a simple in-memory mode that lets you try semantic search locally.
```python
pip install qdrant-client
```
## Install FastEmbed
Installing FastEmbed will let you quickly turn data to vectors, so that Qdrant can search over them.
```python
pip install fastembed
```
## Initialize the client
```python
from qdrant_client import QdrantClient
client = QdrantClient(":memory:") # Qdrant is running from RAM.
```
## Add data
Now you can add two sample documents, their associated metadata, and a point `id` for each.
```python
docs = ["Qdrant has a LangChain integration for chatbots.", "Qdrant has a LlamaIndex integration for agents."]
metadata = [
{"source": "langchain-docs"},
{"source": "llamaindex-docs"},
]
ids = [42, 2]
```
## Load data to a collection
Create a test collection and upsert your two documents to it.
```python
client.add(
collection_name="test_collection",
documents=docs,
metadata=metadata,
ids=ids
)
```
## Run vector search
Here, you will ask a dummy question that will allow you to retrieve a semantically relevant result.
```python
search_result = client.query(
collection_name="test_collection",
query_text="Which integration is best for agents?"
)
print(search_result)
```
The semantic search engine will retrieve the most similar result in order of releance. In this case, the second statement about LlamaIndex is more relevant.
```bash
[QueryResponse(id=2, embedding=None, sparse_embedding=None,
metadata={'document': 'Qdrant has a LlamaIndex integration for agents',
'source': 'llamaindex-docs'}, document='Qdrant has a LlamaIndex integration for agents.',
score=0.8749180370667156),
QueryResponse(id=42, embedding=None, sparse_embedding=None,
metadata={'document': 'Qdrant has a LangChain integration for chatbots.',
'source': 'langchain-docs'}, document='Qdrant has a LangChain integration for chatbots.',
score=0.8351846822959111)]
```
@@ -0,0 +1,196 @@
---
title: Working with SPLADE
weight: 5
---
# How to Generate Sparse Vectors with SPLADE
SPLADE is a novel method for learning sparse text representation vectors, outperforming BM25 in tasks like information retrieval and document classification. Its main advantage is generating efficient and interpretable sparse vectors, making it effective for large-scale text data.
## Setup
First, install FastEmbed.
```python
pip install -q fastembed
```
Next, import the required modules for sparse embeddings and Python’s typing module.
```python
from fastembed import SparseTextEmbedding, SparseEmbedding
from typing import List
```
You may always check the list of all supported sparse embedding models.
```python
SparseTextEmbedding.list_supported_models()
```
This will return a list of models, each with its details such as model name, vocabulary size, description, and sources.
```python
[{'model': 'prithivida/Splade_PP_en_v1',
'vocab_size': 30522,
'description': 'Independent Implementation of SPLADE++ Model for English',
'size_in_GB': 0.532,
'sources': {'hf': 'Qdrant/SPLADE_PP_en_v1'}}]
```
Now, load the model.
```python
model_name = "prithvida/Splade_PP_en_v1"
# This triggers the model download
model = SparseTextEmbedding(model_name=model_name)
```
## Embed data
You need to define a list of documents to be embedded.
```python
documents: List[str] = [
"Chandrayaan-3 is India's third lunar mission",
"It aimed to land a rover on the Moon's surface - joining the US, China and Russia",
"The mission is a follow-up to Chandrayaan-2, which had partial success",
"Chandrayaan-3 will be launched by the Indian Space Research Organisation (ISRO)",
"The estimated cost of the mission is around $35 million",
"It will carry instruments to study the lunar surface and atmosphere",
"Chandrayaan-3 landed on the Moon's surface on 23rd August 2023",
"It consists of a lander named Vikram and a rover named Pragyan similar to Chandrayaan-2. Its propulsion module would act like an orbiter.",
"The propulsion module carries the lander and rover configuration until the spacecraft is in a 100-kilometre (62 mi) lunar orbit",
"The mission used GSLV Mk III rocket for its launch",
"Chandrayaan-3 was launched from the Satish Dhawan Space Centre in Sriharikota",
"Chandrayaan-3 was launched earlier in the year 2023",
]
```
Then, generate sparse embeddings for each document.
Here,`batch_size` is optional and helps to process documents in batches.
```python
sparse_embeddings_list: List[SparseEmbedding] = list(
model.embed(documents, batch_size=6)
)
```
## Retrieve embeddings
`Sparse_embeddings_list` contains sparse embeddings for the documents provided earlier. Each element in this list is a `SparseEmbedding` object that contains the sparse vector representation of a document. Set the variable `index` to 0. This will retrieve data from the first document in the list.
```python
index = 0
sparse_embeddings_list[index]
```
This output is a `SparseEmbedding` object for the first document in our list. It contains two arrays: `values` and `indices`. - The `values` array represents the weights of the features (tokens) in the document. - The `indices` array represents the indices of these features in the model's vocabulary.
Each pair of corresponding `values` and `indices` represents a token and its weight in the document.
```python
SparseEmbedding(values=array([0.05297208, 0.01963477, 0.36459631, 1.38508618, 0.71776593,
0.12667948, 0.46230844, 0.446771 , 0.26897505, 1.01519883,
1.5655334 , 0.29412213, 1.53102326, 0.59785569, 1.1001817 ,
0.02079751, 0.09955651, 0.44249091, 0.09747757, 1.53519952,
1.36765671, 0.15740395, 0.49882549, 0.38629025, 0.76612782,
1.25805044, 0.39058095, 0.27236196, 0.45152301, 0.48262018,
0.26085234, 1.35912788, 0.70710695, 1.71639752]), indices=array([ 1010, 1011, 1016, 1017, 2001, 2018, 2034, 2093, 2117,
2319, 2353, 2509, 2634, 2686, 2796, 2817, 2922, 2959,
3003, 3148, 3260, 3390, 3462, 3523, 3822, 4231, 4316,
4774, 5590, 5871, 6416, 11926, 12076, 16469]))
```
## Examine weights
Now, print the first 5 features and their weights for better understanding.
```python
for i in range(5):
print(f"Token at index {sparse_embeddings_list[0].indices[i]} has weight {sparse_embeddings_list[0].values[i]}")
```
The output will display the token indices and their corresponding weights for the first document.
```python
Token at index 1010 has weight 0.05297207832336426
Token at index 1011 has weight 0.01963476650416851
Token at index 1016 has weight 0.36459630727767944
Token at index 1017 has weight 1.385086178779602
Token at index 2001 has weight 0.7177659273147583
```
## Analyze results
Let's use the tokenizer vocab to make sense of these indices.
```python
import json
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(SparseTextEmbedding.list_supported_models()[0]["sources"]["hf"])
```
The `get_tokens_and_weights` function takes a `SparseEmbedding` object and a `tokenizer` as input. It will construct a dictionary where the keys are the decoded tokens, and the values are their corresponding weights.
```python
def get_tokens_and_weights(sparse_embedding, tokenizer):
token_weight_dict = {}
for i in range(len(sparse_embedding.indices)):
token = tokenizer.decode([sparse_embedding.indices[i]])
weight = sparse_embedding.values[i]
token_weight_dict[token] = weight
# Sort the dictionary by weights
token_weight_dict = dict(sorted(token_weight_dict.items(), key=lambda item: item[1], reverse=True))
return token_weight_dict
# Test the function with the first SparseEmbedding
print(json.dumps(get_tokens_and_weights(sparse_embeddings_list[index], tokenizer), indent=4))
```
## Dictionary output
The dictionary is then sorted by weights in descending order.
```python
{
"chandra": 1.7163975238800049,
"third": 1.5655333995819092,
"##ya": 1.535199522972107,
"india": 1.5310232639312744,
"3": 1.385086178779602,
"mission": 1.3676567077636719,
"lunar": 1.3591278791427612,
"moon": 1.2580504417419434,
"indian": 1.1001816987991333,
"##an": 1.015198826789856,
"3rd": 0.7661278247833252,
"was": 0.7177659273147583,
"spacecraft": 0.7071069478988647,
"space": 0.5978556871414185,
"flight": 0.4988254904747009,
"satellite": 0.4826201796531677,
"first": 0.46230843663215637,
"expedition": 0.4515230059623718,
"three": 0.4467709958553314,
"fourth": 0.44249090552330017,
"vehicle": 0.390580952167511,
"iii": 0.3862902522087097,
"2": 0.36459630727767944,
"##3": 0.2941221296787262,
"planet": 0.27236196398735046,
"second": 0.26897504925727844,
"missions": 0.2608523368835449,
"launched": 0.15740394592285156,
"had": 0.12667948007583618,
"largest": 0.09955651313066483,
"leader": 0.09747757017612457,
",": 0.05297207832336426,
"study": 0.02079751156270504,
"-": 0.01963476650416851
}
```
## Observations
- The relative order of importance is quite useful. The most important tokens in the sentence have the highest weights.
- **Term Expansion:** The model can expand the terms in the document. This means that the model can generate weights for tokens that are not present in the document but are related to the tokens in the document. This is a powerful feature that allows the model to capture the context of the document. Here, you'll see that the model has added the tokens '3' from 'third' and 'moon' from 'lunar' to the sparse vector.
## Design choices
- The weights are not normalized. This means that the sum of the weights is not 1 or 100. This is a common practice in sparse embeddings, as it allows the model to capture the importance of each token in the document.
- Tokens are included in the sparse vector only if they are present in the model's vocabulary. This means that the model will not generate a weight for tokens that it has not seen during training.
- Tokens do not map to words directly -- allowing you to gracefully handle typo errors and out-of-vocabulary tokens.