added sep images

This commit is contained in:
Evgeniya Sukhodolskaya
2025-07-09 12:04:55 +02:00
parent a754d61007
commit 2c62efc6a6
7 changed files with 16 additions and 3 deletions
@@ -8,7 +8,7 @@ preview_image: /blog/hitchhikers-guide/preview/preview.jpg
social_preview_image: /blog/hitchhikers-guide/preview/social_preview.jpg # Optional image used for link previews social_preview_image: /blog/hitchhikers-guide/preview/social_preview.jpg # Optional image used for link previews
title_preview_image: /blog/hitchhikers-guide/preview/title.jpg # Optional image used for blog post title title_preview_image: /blog/hitchhikers-guide/preview/title.jpg # Optional image used for blog post title
date: 2025-07-08T13:50:57+02:00 date: 2025-07-09T00:00:00+02:00
author: Clelia Astra Bertelli & Evgeniya Sukhodolskaya author: Clelia Astra Bertelli & Evgeniya Sukhodolskaya
featured: false featured: false
tags: tags:
@@ -17,7 +17,8 @@ tags:
- blog - blog
--- ---
> From lecture halls to production pipelines, [Qdrant Stars](https://qdrant.tech/stars/) -- founders, mentors and open-source contributors -- share how they’re building with vectors in the wild. In this post, Clelia distils tips from her talk at the [“Bavaria, Advancements in SEarch Development” meetup](https://lu.ma/based_meetup), where she covered hard-won lessons from her extensive open-source building. > From lecture halls to production pipelines, [Qdrant Stars](https://qdrant.tech/stars/) -- founders, mentors and open-source contributors -- share how they’re building with vectors in the wild.
> In this post, Clelia distils tips from her talk at the [“Bavaria, Advancements in SEarch Development” meetup](https://lu.ma/based_meetup), where she covered hard-won lessons from her extensive open-source building.
*Hey there, vector space astronauts!* *Hey there, vector space astronauts!*
@@ -43,6 +44,8 @@ When the user asks a question, context will be *retrieved* from the database and
## Text Extraction: Your Best Friend and Worst Enemy ## Text Extraction: Your Best Friend and Worst Enemy
![text-extraction](/blog/hitchhikers-guide/sep_1.png)
Text extraction is a crucial step: having clean, well-structured raw text can be game-changing for all the downstream steps of your RAG, especially to make the retrieved context easily “understandable” for the LLM. Text extraction is a crucial step: having clean, well-structured raw text can be game-changing for all the downstream steps of your RAG, especially to make the retrieved context easily “understandable” for the LLM.
You can perform text extraction in various ways, for example: You can perform text extraction in various ways, for example:
@@ -60,6 +63,8 @@ It allows you to use LlamaParse or simple parsing with PyPDF to extract text fro
## Chunking Is All You Need ## Chunking Is All You Need
![chunking](/blog/hitchhikers-guide/sep_2.png)
Chunking might really make the difference between a successful and a failing RAG pipeline. Chunking might really make the difference between a successful and a failing RAG pipeline.
Chunking means breaking large text down into pieces, which will be given to the LLM to perform augmented generation. It is then crucial to break the text down into chunks that are meaningful. Chunking means breaking large text down into pieces, which will be given to the LLM to perform augmented generation. It is then crucial to break the text down into chunks that are meaningful.
@@ -80,6 +85,8 @@ You can also check out how **abstract syntax trees can be used to parse and chun
## Embeddings: Catch ‘em All! ## Embeddings: Catch ‘em All!
![embeddings](/blog/hitchhikers-guide/sep_3.png)
Embedding text equals generating a numerical representation of it. For example, as a target representation, you could choose *dense* or *sparse* vectors. The difference is in what they capture: dense embeddings are the best at broadly catching the semantic nuances of the text, while sparse embeddings precisely pick up its keywords. Embedding text equals generating a numerical representation of it. For example, as a target representation, you could choose *dense* or *sparse* vectors. The difference is in what they capture: dense embeddings are the best at broadly catching the semantic nuances of the text, while sparse embeddings precisely pick up its keywords.
The good news is you don't have to pick one with a [hybrid search](https://qdrant.tech/articles/hybrid-search/). Hybrid search combines results from both a dense (semantic) search and a sparse (keyword) search. The good news is you don't have to pick one with a [hybrid search](https://qdrant.tech/articles/hybrid-search/). Hybrid search combines results from both a dense (semantic) search and a sparse (keyword) search.
@@ -97,6 +104,8 @@ If you want a quick start on a hybrid search, check out [Pokemon-Bot](https://gi
## Search Boosting 101 ## Search Boosting 101
![search-boosting](/blog/hitchhikers-guide/sep_4.png)
Search boosting is something everybody wants: less compute, reduced latency and, overall, faster and more efficient pipelines that can make the UX way smoother. I’ll mention two of them. Search boosting is something everybody wants: less compute, reduced latency and, overall, faster and more efficient pipelines that can make the UX way smoother. I’ll mention two of them.
### Semantic caching ### Semantic caching
@@ -111,7 +120,7 @@ Then, before running the whole RAG pipeline, you perform a quick search within y
Binary quantization is also something that can help you, especially if you have tons of documents (we’re talking millions). A large dataset is a performance challenge and a memory problem: embeddings from providers like OpenAI can have 1536 dimensions, meaning almost 6 kB per full-precision embedding! Binary quantization is also something that can help you, especially if you have tons of documents (we’re talking millions). A large dataset is a performance challenge and a memory problem: embeddings from providers like OpenAI can have 1536 dimensions, meaning almost 6 kB per full-precision embedding!
*And here it comes:* taking the vector as a list of float32 numbers, binary quantization converts it to a list of 0s and 1s based on mathematical rounding. Such a representation significantly reduces the memory footprint and makes it easier for your search algorithm to compare vector representations. *And here it comes:* taking the vector as a list of floating point numbers, binary quantization converts it to a list of 0s and 1s based on mathematical rounding. Such a representation significantly reduces the memory footprint and makes it easier for your search algorithm to compare vector representations.
Using binary quantization comes with the natural question: *“Are the search results as good as if I were using the non-quantized vectors?”* Generally, they aren’t *as* good, which is why they should be combined with the other techniques discussed above, such as rescoring. Using binary quantization comes with the natural question: *“Are the search results as good as if I were using the non-quantized vectors?”* Generally, they aren’t *as* good, which is why they should be combined with the other techniques discussed above, such as rescoring.
@@ -121,6 +130,8 @@ Do you want to build a production-ready system with semantic caching and quantiz
## Querying Makes the Difference ## Querying Makes the Difference
![querying](/blog/hitchhikers-guide/sep_5.png)
A common error in RAG pipelines is that you curate every detail, but you do not take into account one key aspect: **queries**. A common error in RAG pipelines is that you curate every detail, but you do not take into account one key aspect: **queries**.
Most of the time, queries are *too generic* or *too specific* to pick up the knowledge embedded in your vector database. There are, nevertheless, some magic tricks you can apply to optimize querying for your use case. Most of the time, queries are *too generic* or *too specific* to pick up the knowledge embedded in your vector database. There are, nevertheless, some magic tricks you can apply to optimize querying for your use case.
@@ -141,6 +152,8 @@ Well, you can build an agentic system for automated choice of query transformati
## Don’t Drown in Evals ## Don’t Drown in Evals
![evalustion](/blog/hitchhikers-guide/sep_6.png)
Ideas turned into implementations are cool, yet only eval metrics can tell whether your project delivers real value and has a go-to-market potential. Ideas turned into implementations are cool, yet only eval metrics can tell whether your project delivers real value and has a go-to-market potential.
It is very easy, though, to drown in all of the evaluation frameworks, strategies and metrics out there. It is also common to get medium-to-good results on some metrics when evaluating the first implementation of your product and happily stopping the development while actually there is a huge room for improvement. It is very easy, though, to drown in all of the evaluation frameworks, strategies and metrics out there. It is also common to get medium-to-good results on some metrics when evaluating the first implementation of your product and happily stopping the development while actually there is a huge room for improvement.
Binary file not shown.

After

Width:  |  Height:  |  Size: 548 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 591 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 666 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 580 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 606 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 530 KiB