mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-08 12:28:32 +02:00
fix: linkchecker include filter port mismatch (#2242)
* initial commit; fixed anchor links on internal docs pages * add back in absolute paths for links in code comments * fix: update linkchecker include filter to match server port 1314 PR #1629 changed the Hugo server to port 1314 but forgot to update the --include filter, which still matched port 1313. This caused all links to be excluded, making the checker a no-op (0 checked, 82277 excluded). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Fix url rewrite regex so images are not impacted * Fix links from non-documentation pages * Fix broken links * more broken links * more broken links * broken link * Add srcset width descriptor to .lycheeignore * Ignore URLs that contain a % character * Anchor regex so it matches the entire URL --------- Co-authored-by: kanungle <neil.kanungo@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Abdon Pijpelink <abdon.pijpelink@qdrant.com>
This commit is contained in:
co-authored by
Claude Opus 4.6
kanungle
Abdon Pijpelink
parent
dc0080fffa
commit
1a40961d62
@@ -55,7 +55,7 @@ This is where chunking comes in. The goal is to have chunks
|
||||
|
||||
By breaking a document into focused chunks, each chunk gets its own vector that accurately represents a specific idea. This allows the search to be far more precise.
|
||||
|
||||
**Example:** Consider a multi-page Document like the [Qdrant Collection Configuration Guide of Day 7](/course/essentials/day-7/collection-configuration-guide/) covering everything from HNSW to sharding and quantization.
|
||||
**Example:** Consider a multi-page Document like the [Qdrant Collection Configuration Guide of Day 7](/course/essentials/day-7/) covering everything from HNSW to sharding and quantization.
|
||||
|
||||
If a user asks: *"What does the m parameter do?"*
|
||||
|
||||
@@ -420,7 +420,7 @@ The trade-off is computational cost. You're embedding the full document upfront
|
||||
| **Recursive** | Flexible, handles messy input | Heuristic, sometimes brittle | Scraped web content, mixed sources |
|
||||
| **Semantic** | High-quality, meaning-aware | Slower, resource-intensive | Legal, research, critical QA |
|
||||
|
||||
**Note**: Sometimes, it's necessary to keep the document intact. If chunking is too complicated, or the document is visually rich (diagrams, graphs etc.), you can use [VLMs](/documentation/advanced-tutorials/pdf-retrieval-at-scale/) to embed the whole page.
|
||||
**Note**: Sometimes, it's necessary to keep the document intact. If chunking is too complicated, or the document is visually rich (diagrams, graphs etc.), you can use [VLMs](/documentation/tutorials-search-engineering/pdf-retrieval-at-scale/) to embed the whole page.
|
||||
|
||||
## Adding Meaning with Metadata
|
||||
|
||||
@@ -439,7 +439,7 @@ In Qdrant, this metadata lives in the **payload** - a JSON object attached to ea
|
||||
"section_title": "What Is a Vector",
|
||||
"chunk_index": 7,
|
||||
"chunk_count": 15,
|
||||
"url": "https://qdrant.tech/documentation/concepts/collections/",
|
||||
"url": "https://qdrant.tech/documentation/manage-data/collections/",
|
||||
"tags": ["qdrant", "vector search", "point", "vector", "payload"],
|
||||
"source_type": "documentation",
|
||||
"created_at": "2025-01-15T10:00:00Z",
|
||||
@@ -451,7 +451,7 @@ In Qdrant, this metadata lives in the **payload** - a JSON object attached to ea
|
||||
|
||||
### What Metadata Enables
|
||||
|
||||
**Disclaimer**: For performance reasons, filterable fields must be indexed using the [Payload Index](/documentation/concepts/indexing/#payload-index).
|
||||
**Disclaimer**: For performance reasons, filterable fields must be indexed using the [Payload Index](/documentation/manage-data/indexing/#payload-index).
|
||||
|
||||
**1. Filtered Search (Exact Match)**
|
||||
You can filter results based on exact metadata values, which is perfect for categorical data.
|
||||
@@ -469,7 +469,7 @@ filter = models.Filter(
|
||||
```
|
||||
|
||||
**2. Hybrid Search with Text Filtering (Full-Text Search)**
|
||||
For more powerful text-based filtering, you can combine vector search with traditional keyword search. This requires setting up a [full-text index](/documentation/concepts/indexing/#full-text-index) on a payload field.
|
||||
For more powerful text-based filtering, you can combine vector search with traditional keyword search. This requires setting up a [full-text index](/documentation/manage-data/indexing/#full-text-index) on a payload field.
|
||||
```python
|
||||
# Find vectors that also contain the keyword "HNSW" in their content
|
||||
filter = models.Filter(
|
||||
@@ -487,13 +487,13 @@ filter = models.Filter(
|
||||
# Top result per document - get the most relevant chunk from each source
|
||||
group_by = "document_id"
|
||||
```
|
||||
You can read more about grouping [here](/documentation/concepts/hybrid-queries/?q=grouping#grouping).
|
||||
You can read more about grouping [here](/documentation/search/hybrid-queries/?q=grouping#grouping).
|
||||
|
||||
**4. Rich Result Display**
|
||||
- Original content with source attribution
|
||||
- Section context for better understanding
|
||||
- Direct links to full documents
|
||||
- Creation timestamps for [freshness](/documentation/concepts/search-relevance/#time-based-score-boosting)
|
||||
- Creation timestamps for [freshness](/documentation/search/search-relevance/#time-based-score-boosting)
|
||||
|
||||
**5. Permission Control**
|
||||
```python
|
||||
|
||||
@@ -9,7 +9,7 @@ isLesson: true
|
||||
|
||||
# Distance Metrics
|
||||
|
||||
After vectors are stored, we can use their spatial properties to perform [nearest neighbor searches](/documentation/concepts/search/) that retrieve semantically similar items based on how close they are in this space.
|
||||
After vectors are stored, we can use their spatial properties to perform [nearest neighbor searches](/documentation/search/search/) that retrieve semantically similar items based on how close they are in this space.
|
||||
|
||||
The position of a vector in embedding space only reflects meaning as far as the embedding model has learned to encode it. The model and its training objective tell you what "close" means.
|
||||
|
||||
@@ -166,4 +166,4 @@ If you are training your own model or designing custom features, use these guide
|
||||
* **Dot product** accounts for magnitude and direction.
|
||||
4. **Experiment:** Qdrant allows you to set distance metrics per named vector, making it easy to A/B test different metrics on your specific data.
|
||||
|
||||
Reference: [Distance Metrics in Qdrant Documentation](/documentation/concepts/search/#metrics)
|
||||
Reference: [Distance Metrics in Qdrant Documentation](/documentation/search/search/#metrics)
|
||||
@@ -82,7 +82,7 @@ The `indices` and `values` arrays must be the same size, and all the `indices` m
|
||||
|
||||
There is no need to sort the sparse representation by indices, as Qdrant will perform this internally while maintaining the correct link between each index and its value.
|
||||
|
||||
We will cover more about sparse vectors on day 3. If you would like to read up on the subject in advance, you can find more documentation [here](/documentation/concepts/vectors/#sparse-vectors).
|
||||
We will cover more about sparse vectors on day 3. If you would like to read up on the subject in advance, you can find more documentation [here](/documentation/manage-data/vectors/#sparse-vectors).
|
||||
|
||||
|
||||
### Multivectors
|
||||
@@ -234,7 +234,7 @@ While vectors capture the essence of data, payloads hold structured metadata for
|
||||
|
||||
Payloads can store textual data (descriptions, tags, categories), numerical values (dates, prices, ratings), and complex structures (nested objects, arrays). When searching for dog images, for example, the vector finds visually similar images while payload filters narrow results to images taken within the last year, tagged with "vacation," or meeting specific rating criteria.
|
||||
|
||||
Learn more: [Payload Documentation](/documentation/concepts/payload/)
|
||||
Learn more: [Payload Documentation](/documentation/manage-data/payload/)
|
||||
|
||||
|
||||
### Payload Types
|
||||
@@ -304,7 +304,7 @@ Here are some of the most common condition types:
|
||||
|
||||
<aside role="alert"> This list covers the most common conditions available at the time of this course. Qdrant is constantly evolving, and new filtering capabilities may have been added.
|
||||
|
||||
For the complete, most up-to-date list of all available filtering conditions, please refer to the **[official Filtering documentation](/documentation/concepts/filtering/#filtering-conditions)**.</aside>
|
||||
For the complete, most up-to-date list of all available filtering conditions, please refer to the **[official Filtering documentation](/documentation/search/filtering/#filtering-conditions)**.</aside>
|
||||
|
||||
|
||||
### Filtering Capabilities Reference
|
||||
@@ -382,7 +382,7 @@ client.create_payload_index(
|
||||
|
||||
When filters are highly selective, Qdrant's query planner may bypass vector indexing entirely and use payload indexes for faster results.
|
||||
|
||||
For comprehensive filtering examples and advanced usage patterns, see the [Filtering Documentation](/documentation/concepts/filtering/) and [Complete Guide to Filtering in Vector Search](/articles/vector-search-filtering/).
|
||||
For comprehensive filtering examples and advanced usage patterns, see the [Filtering Documentation](/documentation/search/filtering/) and [Complete Guide to Filtering in Vector Search](/articles/vector-search-filtering/).
|
||||
|
||||
## Key Takeaways
|
||||
|
||||
|
||||
@@ -235,7 +235,7 @@ Query: 'alien invasion'
|
||||
|
||||
## Step 6: Advanced Features
|
||||
|
||||
Note: If you are already familiar Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/concepts/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [Day 2](/content/course/essentials/day-2/_index.md) of this course.
|
||||
Note: If you are already familiar Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/manage-data/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [Day 2](/course/essentials/day-2/) of this course.
|
||||
|
||||
|
||||
### Filtering by Metadata
|
||||
|
||||
@@ -135,7 +135,7 @@ def paragraph_chunks(text):
|
||||
|
||||
### Step 4: Create Collections and Process Data
|
||||
|
||||
Note: If you are already familiar with Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/concepts/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [day 2](/content/course/essentials/day-2/_index.md) of this course.
|
||||
Note: If you are already familiar with Qdrant's filterable HNSW, you will know that effective filtering and grouping often relies on creating a [payload index](/documentation/manage-data/indexing/#payload-index) before building HNSW indexes. To keep things simple in this tutorial, we will do a basic search with filters without payload indexes and talk about proper usage of payload indexes on [day 2](/course/essentials/day-2/) of this course.
|
||||
|
||||
```python
|
||||
collection_name = "day1_semantic_search"
|
||||
|
||||
Reference in New Issue
Block a user