fix: linkchecker include filter port mismatch (#2242)

* initial commit; fixed anchor links on internal docs pages

* add back in absolute paths for links in code comments

* fix: update linkchecker include filter to match server port 1314

PR #1629 changed the Hugo server to port 1314 but forgot to update
the --include filter, which still matched port 1313. This caused all
links to be excluded, making the checker a no-op (0 checked, 82277 excluded).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Fix url rewrite regex so images are not impacted

* Fix links from non-documentation pages

* Fix broken links

* more broken links

* more broken links

* broken link

* Add srcset width descriptor to .lycheeignore

* Ignore URLs that contain a % character

* Anchor regex so it matches the entire URL

---------

Co-authored-by: kanungle <neil.kanungo@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
This commit is contained in:
Andrey Vasnetsov
2026-03-30 17:21:32 +02:00
committed by GitHub
co-authored by Claude Opus 4.6 kanungle Abdon Pijpelink
parent dc0080fffa
commit 1a40961d62
272 changed files with 872 additions and 879 deletions
@@ -202,7 +202,7 @@ pip install "qdrant-client[fastembed]"
```
Of course, we need a running Qdrant server for vector search. If you need one,
you can [use a local Docker container](/documentation/quick-start/)
you can [use a local Docker container](/documentation/quickstart/)
or deploy it using the [Qdrant Cloud](https://cloud.qdrant.io/).
You can use either to follow this tutorial. Configure the connection parameters:
@@ -255,7 +255,7 @@ Now that all the preparations are complete, let's start building a neural search
In order to process incoming requests, the hybrid search class will need 3 things: 1) models to convert the query into a vector, 2) the Qdrant client to perform search queries, 3) fusion function to re-rank dense and sparse search results.
Qdrant supports 2 fusion functions for combining the results: [reciprocal rank fusion](https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf) and [distribution based score fusion](https://qdrant.tech/documentation/concepts/hybrid-queries/?q=distribution+based+sc#:~:text=Distribution%2DBased%20Score%20Fusion)
Qdrant supports 2 fusion functions for combining the results: [reciprocal rank fusion](https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf) and [distribution based score fusion](/documentation/search/hybrid-queries/?q=distribution+based+sc#:~:text=Distribution%2DBased%20Score%20Fusion)
1. Create a file named `hybrid_searcher.py` and specify the following.
@@ -17,7 +17,7 @@ A neural search service uses artificial neural networks to improve the accuracy
<aside role="status">
There is a version of this tutorial that uses <a href="https://github.com/qdrant/fastembed">Fastembed</a> model inference engine instead of Sentence Transformers.
Check it out <a href="/documentation/beginner-tutorials/hybrid-search-fastembed/">here</a>.
Check it out <a href="/documentation/tutorials-search-engineering/hybrid-search-fastembed/">here</a>.
</aside>
@@ -27,7 +27,7 @@ Recent advancements in **Vision Large Language Models (VLLMs)**, such as [**ColP
VLLMs like **ColPali** and **ColQwen** generate **multivector representations** for each PDF page; the representations are stored and indexed in a vector database. During the retrieval process, models dynamically create multivector representations for (textual) user queries, and precise retrieval -- matching between PDF pages and queries -- is achieved through [late-interaction mechanism](/blog/qdrant-colpali/#how-colpali-works-under-the-hood).
<aside role="status"> Qdrant supports <a href="/documentation/concepts/vectors/#multivectors">multivector representations</a>, making it well-suited for using embedding models such as ColPali, ColQwen, or <a href="/documentation/fastembed/fastembed-colbert/">ColBERT</a></aside>
<aside role="status"> Qdrant supports <a href="/documentation/manage-data/vectors/#multivectors">multivector representations</a>, making it well-suited for using embedding models such as ColPali, ColQwen, or <a href="/documentation/fastembed/fastembed-colbert/">ColBERT</a></aside>
## Challenges of Scaling VLLMs
@@ -40,7 +40,7 @@ The heavy multivector representations produced by VLLMs make PDF retrieval at sc
To understand the impact, consider the construction of an [**HNSW index**](/articles/what-is-a-vector-database/#1-indexing-hnsw-index-and-sending-data-to-qdrant), a common indexing algorithm for vector databases. Let's roughly estimate the number of comparisons needed to insert a new PDF page into the index.
- **Vectors per page:** ~700 (ColQwen) or ~1,000 (ColPali)
- **[ef_construct](/documentation/concepts/indexing/#vector-index):** 100 (default)
- **[ef_construct](/documentation/manage-data/indexing/#vector-index):** 100 (default)
The lower bound estimation for the number of vector comparisons would be:
@@ -58,7 +58,7 @@ the approximate search results. Then, we will call the exact search endpoint to
in terms of precision.
Before we start, let's create a collection, fill it with some data and then start our evaluation. We will use the same dataset as in the
[Loading a dataset from Hugging Face hub](/documentation/tutorials/huggingface-datasets/) tutorial, `Qdrant/arxiv-titles-instructorxl-embeddings`
[Loading a dataset from Hugging Face hub](/documentation/tutorials-basics/huggingface-datasets/) tutorial, `Qdrant/arxiv-titles-instructorxl-embeddings`
from the [Hugging Face hub](https://huggingface.co/datasets/Qdrant/arxiv-titles-instructorxl-embeddings). Let's download it in a streaming
mode, as we are only going to use part of it.
@@ -135,7 +135,7 @@ ranking quality of search results, with higher scores indicating better performa
Binary Quantization definitely speeds up the retrieval, and make it cheaper, but also seems not to affect the quality of
the retrieval much in some cases. **However, that's something you should carefully verify on your own data**. If you are
a Qdrant user, then you can just enable quantization on an existing collection and [measure the impact on the retrieval
quality](/documentation/beginner-tutorials/retrieval-quality/).
quality](/documentation/tutorials-search-engineering/retrieval-quality/).
All the tests we did were performed using [`beir-qdrant`](https://github.com/kacperlukawski/beir-qdrant), and might be
reproduced by running [the script available on the project
@@ -27,7 +27,7 @@ As you will see later in the tutorial, Qdrant supports multivectors and thus lat
With token-level vectors, models like ColBERT can match specific query tokens to the most relevant parts of a document, enabling high-accuracy retrieval through Late Interaction.
In late interaction, each document is converted into multiple token-level vectors instead of a single vector. The query is also tokenized and embedded into various vectors. Then, the query and document vectors are matched using a similarity function: MaxSim. You can see how it is calculated [here](https://qdrant.tech/documentation/concepts/vectors/#multivectors).
In late interaction, each document is converted into multiple token-level vectors instead of a single vector. The query is also tokenized and embedded into various vectors. Then, the query and document vectors are matched using a similarity function: MaxSim. You can see how it is calculated [here](/documentation/manage-data/vectors/#multivectors).
In traditional retrieval, the query and document are converted into single embeddings, after which similarity is computed. This is an early interaction because the information is compressed before retrieval.
@@ -7,7 +7,7 @@ weight: 2
| Time: 30 min | Level: Intermediate | Output: [GitHub](https://github.com/qdrant/examples/blob/master/using-relevance-feedback/Customizing_Relevance_Feedback.ipynb) | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://githubtocolab.com/qdrant/examples/blob/master/using-relevance-feedback/Customizing_Relevance_Feedback.ipynb) |
| --- | ----------- | ----------- | ----------- |
In Qdrant 1.17 we introduced a new [Relevance Feedback Query](/documentation/concepts/search-relevance/#relevance-feedback), our scalable, first ever vector index-native approach to [incorporating relevance feedback](/articles/search-feedback-loop/) in retrieval.
In Qdrant 1.17 we introduced a new [Relevance Feedback Query](/documentation/search/search-relevance/#relevance-feedback), our scalable, first ever vector index-native approach to [incorporating relevance feedback](/articles/search-feedback-loop/) in retrieval.
In this tutorial, you'll see how to:
1. Customize Relevance Feedback Query for your Qdrant collection, retriever and feedback model.
@@ -25,7 +25,7 @@ A detailed description of how it works can be found in the article [Relevance Fe
### Strategy
To use the feedback for guiding a retriever in the vector space, there are several possible strategies. For now, only the **naive strategy** is available -- [a simple 3-parameter formula](https://qdrant.tech/documentation/concepts/search-relevance/#naive-strategy) which adjusts similarity scoring based on the feedback.
To use the feedback for guiding a retriever in the vector space, there are several possible strategies. For now, only the **naive strategy** is available -- [a simple 3-parameter formula](/documentation/search/search-relevance/#naive-strategy) which adjusts similarity scoring based on the feedback.
For the strategy to work well, the parameters of this naive formula should be customized for your data, retriever and feedback model.
For convenience, we provide you with a [`qdrant-relevance-feedback` Python package](https://pypi.org/project/qdrant-relevance-feedback/) that gives you the corresponding parameters for your use case.
@@ -94,7 +94,7 @@ Check what a point in this collection looks like.
![Point in the documentation collection](/documentation/tutorials/using-relevance-feedback/point.png)
Our documentation collection has only one vector per point -- `Default vector`. However, in Qdrant, one can have several [named vectors](https://qdrant.tech/documentation/concepts/vectors/#named-vectors) per point.
Our documentation collection has only one vector per point -- `Default vector`. However, in Qdrant, one can have several [named vectors](/documentation/manage-data/vectors/#named-vectors) per point.
We need to provide the name of the vector associated with the retriever that we're planning to optimize with feedback.
@@ -102,7 +102,7 @@ We need to provide the name of the vector associated with the retriever that we'
RETRIEVER_VECTOR_NAME = None # None if it's a default vector or your named vector handle in Qdrant's collection
```
We also need to point our framework to the raw data which is vectorized with our retriever model. Here, the raw data is the `text` field in the point's [payload](https://qdrant.tech/documentation/concepts/payload/#payload). Simply put, we search on `text` snippets -- in their vectorized form.
We also need to point our framework to the raw data which is vectorized with our retriever model. Here, the raw data is the `text` field in the point's [payload](/documentation/manage-data/payload/#payload). Simply put, we search on `text` snippets -- in their vectorized form.
```python
PAYLOAD_KEY = "text"