diff --git a/qdrant-landing/content/articles/fastembed.md b/qdrant-landing/content/articles/fastembed.md index ea77d10cb..fdb38f3c6 100644 --- a/qdrant-landing/content/articles/fastembed.md +++ b/qdrant-landing/content/articles/fastembed.md @@ -1,3 +1,21 @@ +--- +title: "FastEmbed: 2x faster Embeddings" +short_description: "FastEmbed is a Python library engineered for speed, efficiency, and above all, usability." +description: "FastEmbed is a Python library engineered for speed, efficiency, and accuracy. It's more accurate than OpenAI and 1.5x faster than the PyTorch implementation with fewer dependencies" +social_preview_image: /articles_data/fastembed/social_preview.png +preview_dir: /articles_data/fastembed/preview +weight: -40 +author: Nirant Kasliwal +author_link: https://nirantk.com/about/ +date: 2023-10-18T13:00:00+03:00 +draft: false +keywords: + - vector search + - embedding models + - Flag Embedding + - OpenAI Ada + - quantized embedding model +--- # FastEmbed In the ever-changing landscape of Data Science and Machine Learning, practitioners often find themselves navigating through a labyrinth of models, libraries, and frameworks. Among the plethora of choices, the need for a specialized, efficient, and easy-to-implement solution for embedding generation is increasingly evident. This is where FastEmbed (docs: [https://qdrant.github.io/fastembed/](https://qdrant.github.io/fastembed/)) comes into play—a Python library engineered for speed, efficiency, and above all, usability. @@ -69,29 +87,22 @@ Suggested Illustration: Graphical for computational efficiency and accuracy metr ## Key Features - ### Computational Efficiency -ONNX Runtime: Examine how FastEmbed leverages ONNX Runtime for inference - -Resource Utilization: Analyze the resource footprint and computational benefits arising from the lightweight nature of FastEmbed. - - - -

>>>>> gd2md-html alert: inline image link here (to images/image1.png). Store image on your image server and adjust path/filename/extension if necessary.
(Back to top)(Next alert)
>>>>>

+FastEmbed is fast because of a lot of small things we've taken care of for you: +1. **Quantized Models**: We quantize the models for CPU (and Mac Metal) – giving you the best buck for your compute model. Our models are so small, you can run this in AWS Lambda if you'd like! +2. **1.5x Throughput**: This is the fastest CPU model which beats OpenAI Embedding model as well. And we do so while being 1.5x faster than the Open Source implementation. ![alt_text](images/image1.png "image_tooltip") - - ### Retaining Accuracy and Recall We support quantized models for State of the Art Embedding models e.g. those from [MTEB](https://huggingface.co/spaces/mteb/leaderboard). For FastEmbed's DefaultEmbedding model, we give this throughput improvement without sacrificing Accuracy or Recall. How do we measure this? The cosine similarity between the Transformers/PyTorch implementation and our quantized model is 0.999999. -We strongly recommend that you pin the FastEmbed version in your usage to a specific version, since the DefaultEmbedding will also be continuously updated to give a strong speed vs accuracy balance. +**No decision fatigue: The DefaultEmbedding model will always be the best Open Source model for English. And if this changes, we'll make a new minor version release e.g. 0.0.6 to 0.1. We strongly recommend that you pin the FastEmbed version in your usage to a specific version. ### Comparison Against OpenAI @@ -100,8 +111,6 @@ For retrieval, FastEmbed does almost 3% better than OpenAI. We're also faster be On every metric that you care about: speed, accuracy and ease of use – we do better and intend to continue to do so! - -

>>>>> gd2md-html alert: inline image link here (to images/image2.png). Store image on your image server and adjust path/filename/extension if necessary.
(Back to top)(Next alert)
>>>>>

diff --git a/qdrant-landing/static/articles_data/fastembed/image1.png b/qdrant-landing/static/articles_data/fastembed/image1.png new file mode 100644 index 000000000..e89a34398 Binary files /dev/null and b/qdrant-landing/static/articles_data/fastembed/image1.png differ diff --git a/qdrant-landing/static/articles_data/fastembed/image2.png b/qdrant-landing/static/articles_data/fastembed/image2.png new file mode 100644 index 000000000..f11aa5e72 Binary files /dev/null and b/qdrant-landing/static/articles_data/fastembed/image2.png differ diff --git a/qdrant-landing/static/articles_data/fastembed/image3.png b/qdrant-landing/static/articles_data/fastembed/image3.png new file mode 100644 index 000000000..a58347394 Binary files /dev/null and b/qdrant-landing/static/articles_data/fastembed/image3.png differ diff --git a/qdrant-landing/static/articles_data/fastembed/image4.png b/qdrant-landing/static/articles_data/fastembed/image4.png new file mode 100644 index 000000000..43d920d83 Binary files /dev/null and b/qdrant-landing/static/articles_data/fastembed/image4.png differ