mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-03 01:48:32 +02:00
added sep images
This commit is contained in:
@@ -8,7 +8,7 @@ preview_image: /blog/hitchhikers-guide/preview/preview.jpg
|
|||||||
social_preview_image: /blog/hitchhikers-guide/preview/social_preview.jpg # Optional image used for link previews
|
social_preview_image: /blog/hitchhikers-guide/preview/social_preview.jpg # Optional image used for link previews
|
||||||
title_preview_image: /blog/hitchhikers-guide/preview/title.jpg # Optional image used for blog post title
|
title_preview_image: /blog/hitchhikers-guide/preview/title.jpg # Optional image used for blog post title
|
||||||
|
|
||||||
date: 2025-07-08T13:50:57+02:00
|
date: 2025-07-09T00:00:00+02:00
|
||||||
author: Clelia Astra Bertelli & Evgeniya Sukhodolskaya
|
author: Clelia Astra Bertelli & Evgeniya Sukhodolskaya
|
||||||
featured: false
|
featured: false
|
||||||
tags:
|
tags:
|
||||||
@@ -17,7 +17,8 @@ tags:
|
|||||||
- blog
|
- blog
|
||||||
---
|
---
|
||||||
|
|
||||||
> From lecture halls to production pipelines, [Qdrant Stars](https://qdrant.tech/stars/) -- founders, mentors and open-source contributors -- share how they’re building with vectors in the wild. In this post, Clelia distils tips from her talk at the [“Bavaria, Advancements in SEarch Development” meetup](https://lu.ma/based_meetup), where she covered hard-won lessons from her extensive open-source building.
|
> From lecture halls to production pipelines, [Qdrant Stars](https://qdrant.tech/stars/) -- founders, mentors and open-source contributors -- share how they’re building with vectors in the wild.
|
||||||
|
> In this post, Clelia distils tips from her talk at the [“Bavaria, Advancements in SEarch Development” meetup](https://lu.ma/based_meetup), where she covered hard-won lessons from her extensive open-source building.
|
||||||
|
|
||||||
*Hey there, vector space astronauts!*
|
*Hey there, vector space astronauts!*
|
||||||
|
|
||||||
@@ -43,6 +44,8 @@ When the user asks a question, context will be *retrieved* from the database and
|
|||||||
|
|
||||||
## Text Extraction: Your Best Friend and Worst Enemy
|
## Text Extraction: Your Best Friend and Worst Enemy
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
Text extraction is a crucial step: having clean, well-structured raw text can be game-changing for all the downstream steps of your RAG, especially to make the retrieved context easily “understandable” for the LLM.
|
Text extraction is a crucial step: having clean, well-structured raw text can be game-changing for all the downstream steps of your RAG, especially to make the retrieved context easily “understandable” for the LLM.
|
||||||
|
|
||||||
You can perform text extraction in various ways, for example:
|
You can perform text extraction in various ways, for example:
|
||||||
@@ -60,6 +63,8 @@ It allows you to use LlamaParse or simple parsing with PyPDF to extract text fro
|
|||||||
|
|
||||||
## Chunking Is All You Need
|
## Chunking Is All You Need
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
Chunking might really make the difference between a successful and a failing RAG pipeline.
|
Chunking might really make the difference between a successful and a failing RAG pipeline.
|
||||||
|
|
||||||
Chunking means breaking large text down into pieces, which will be given to the LLM to perform augmented generation. It is then crucial to break the text down into chunks that are meaningful.
|
Chunking means breaking large text down into pieces, which will be given to the LLM to perform augmented generation. It is then crucial to break the text down into chunks that are meaningful.
|
||||||
@@ -80,6 +85,8 @@ You can also check out how **abstract syntax trees can be used to parse and chun
|
|||||||
|
|
||||||
## Embeddings: Catch ‘em All!
|
## Embeddings: Catch ‘em All!
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
Embedding text equals generating a numerical representation of it. For example, as a target representation, you could choose *dense* or *sparse* vectors. The difference is in what they capture: dense embeddings are the best at broadly catching the semantic nuances of the text, while sparse embeddings precisely pick up its keywords.
|
Embedding text equals generating a numerical representation of it. For example, as a target representation, you could choose *dense* or *sparse* vectors. The difference is in what they capture: dense embeddings are the best at broadly catching the semantic nuances of the text, while sparse embeddings precisely pick up its keywords.
|
||||||
|
|
||||||
The good news is you don't have to pick one with a [hybrid search](https://qdrant.tech/articles/hybrid-search/). Hybrid search combines results from both a dense (semantic) search and a sparse (keyword) search.
|
The good news is you don't have to pick one with a [hybrid search](https://qdrant.tech/articles/hybrid-search/). Hybrid search combines results from both a dense (semantic) search and a sparse (keyword) search.
|
||||||
@@ -97,6 +104,8 @@ If you want a quick start on a hybrid search, check out [Pokemon-Bot](https://gi
|
|||||||
|
|
||||||
## Search Boosting 101
|
## Search Boosting 101
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
Search boosting is something everybody wants: less compute, reduced latency and, overall, faster and more efficient pipelines that can make the UX way smoother. I’ll mention two of them.
|
Search boosting is something everybody wants: less compute, reduced latency and, overall, faster and more efficient pipelines that can make the UX way smoother. I’ll mention two of them.
|
||||||
|
|
||||||
### Semantic caching
|
### Semantic caching
|
||||||
@@ -111,7 +120,7 @@ Then, before running the whole RAG pipeline, you perform a quick search within y
|
|||||||
|
|
||||||
Binary quantization is also something that can help you, especially if you have tons of documents (we’re talking millions). A large dataset is a performance challenge and a memory problem: embeddings from providers like OpenAI can have 1536 dimensions, meaning almost 6 kB per full-precision embedding!
|
Binary quantization is also something that can help you, especially if you have tons of documents (we’re talking millions). A large dataset is a performance challenge and a memory problem: embeddings from providers like OpenAI can have 1536 dimensions, meaning almost 6 kB per full-precision embedding!
|
||||||
|
|
||||||
*And here it comes:* taking the vector as a list of float32 numbers, binary quantization converts it to a list of 0s and 1s based on mathematical rounding. Such a representation significantly reduces the memory footprint and makes it easier for your search algorithm to compare vector representations.
|
*And here it comes:* taking the vector as a list of floating point numbers, binary quantization converts it to a list of 0s and 1s based on mathematical rounding. Such a representation significantly reduces the memory footprint and makes it easier for your search algorithm to compare vector representations.
|
||||||
|
|
||||||
Using binary quantization comes with the natural question: *“Are the search results as good as if I were using the non-quantized vectors?”* Generally, they aren’t *as* good, which is why they should be combined with the other techniques discussed above, such as rescoring.
|
Using binary quantization comes with the natural question: *“Are the search results as good as if I were using the non-quantized vectors?”* Generally, they aren’t *as* good, which is why they should be combined with the other techniques discussed above, such as rescoring.
|
||||||
|
|
||||||
@@ -121,6 +130,8 @@ Do you want to build a production-ready system with semantic caching and quantiz
|
|||||||
|
|
||||||
## Querying Makes the Difference
|
## Querying Makes the Difference
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
A common error in RAG pipelines is that you curate every detail, but you do not take into account one key aspect: **queries**.
|
A common error in RAG pipelines is that you curate every detail, but you do not take into account one key aspect: **queries**.
|
||||||
|
|
||||||
Most of the time, queries are *too generic* or *too specific* to pick up the knowledge embedded in your vector database. There are, nevertheless, some magic tricks you can apply to optimize querying for your use case.
|
Most of the time, queries are *too generic* or *too specific* to pick up the knowledge embedded in your vector database. There are, nevertheless, some magic tricks you can apply to optimize querying for your use case.
|
||||||
@@ -141,6 +152,8 @@ Well, you can build an agentic system for automated choice of query transformati
|
|||||||
|
|
||||||
## Don’t Drown in Evals
|
## Don’t Drown in Evals
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
Ideas turned into implementations are cool, yet only eval metrics can tell whether your project delivers real value and has a go-to-market potential.
|
Ideas turned into implementations are cool, yet only eval metrics can tell whether your project delivers real value and has a go-to-market potential.
|
||||||
|
|
||||||
It is very easy, though, to drown in all of the evaluation frameworks, strategies and metrics out there. It is also common to get medium-to-good results on some metrics when evaluating the first implementation of your product and happily stopping the development while actually there is a huge room for improvement.
|
It is very easy, though, to drown in all of the evaluation frameworks, strategies and metrics out there. It is also common to get medium-to-good results on some metrics when evaluating the first implementation of your product and happily stopping the development while actually there is a huge room for improvement.
|
||||||
|
|||||||
Binary file not shown.
|
After Width: | Height: | Size: 548 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 591 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 666 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 580 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 606 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 530 KiB |
Reference in New Issue
Block a user