new design

This commit is contained in:
Dylan Couzon
2026-05-28 00:19:34 -04:00
parent 7b20d2ec9e
commit 26ea463cf1
5 changed files with 242 additions and 162 deletions
@@ -7,4 +7,4 @@
| [Multivectors and Late Interaction](/documentation/tutorials-search-engineering/using-multivector-representations/) | Effective use of multivector representations. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) | Fuse title, summary, chunk, and tag vectors with named vectors and the Query API. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Static Embeddings](/documentation/tutorials-search-engineering/static-embeddings/) | Evaluate the utility of static embeddings. | <span class="pill">Python</span> | 20m | <span class="text-yellow">Intermediate</span> |
| [Branch-Aware Search](/documentation/tutorials-search-engineering/branch-aware-search/) | Scope semantic search to a branch's HEAD in a versioned corpus. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Branch-Aware Search](/documentation/tutorials-search-engineering/branch-aware-search/) | Scope search to a branch's live view in a versioned corpus, inherited from its ancestors. | <span class="pill">Python</span> | 25m | <span class="text-yellow">Intermediate</span> |
@@ -76,7 +76,7 @@ partition: develop
| [Semantic Search for Code](/documentation/tutorials-search-engineering/code-search/) | Navigate codebases using vector similarity. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Multi-Representation Search](/documentation/tutorials-search-engineering/multi-representation-search/) | Fuse title, summary, chunk, and tag vectors with named vectors and the Query API. | <span class="pill">Python</span> | 45m | <span class="text-yellow">Intermediate</span> |
| [Static Embeddings](/documentation/tutorials-search-engineering/static-embeddings/) | Evaluate the renaissance of static embeddings. | <span class="pill">Python</span> | 20m | <span class="text-yellow">Intermediate</span> |
| [Branch-Aware Search](/documentation/tutorials-search-engineering/branch-aware-search/) | Scope semantic search to a branch's HEAD in a versioned corpus. | <span class="pill">Python</span> | 30m | <span class="text-yellow">Intermediate</span> |
| [Branch-Aware Search](/documentation/tutorials-search-engineering/branch-aware-search/) | Scope search to a branch's live view in a versioned corpus, inherited from its ancestors. | <span class="pill">Python</span> | 25m | <span class="text-yellow">Intermediate</span> |
---
@@ -1,7 +1,7 @@
---
title: Branch-Aware Search
short_description: "Build branch-aware semantic search on Qdrant: query a corpus version, get back what's live on that branch."
description: "Index a versioned corpus in Qdrant and scope queries to a branch's HEAD using a materialized per-branch live-set of point IDs."
short_description: "Scope search to a version-controlled branch in Qdrant so a query returns that branch's live view, inherited from its ancestors."
description: "Index a versioned corpus in Qdrant and scope queries to a branch's live view with an ancestor-walk filter over per-version branch and seq payload fields."
weight: 12
aliases:
- /documentation/tutorials/branch-aware-search/
@@ -9,18 +9,16 @@ aliases:
# Branch-Aware Search Over Versioned Documents
| Time: 30 min | Level: Intermediate | Output: [GitHub](https://github.com/qdrant/examples/tree/master/branch-aware-search) | [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://githubtocolab.com/qdrant/examples/blob/master/branch-aware-search/branch_aware_search.ipynb) |
| --- | ----------- | ----------- | ----------- |
| Time: 25 min | Level: Intermediate |
| --- | ----------- |
Searching a versioned corpus needs to scope to "what's live on this branch right now." Without that, a query on a feature branch returns content from `main` the branch has overridden, or misses content the branch added. This tutorial builds the pattern: materialize each branch's HEAD as a set of point IDs and filter queries against that set. Creating a new branch copies that set. No new embeddings, no new Qdrant writes.
A search on a specific branch should return exactly the content that is live on that branch, and nothing from another. Without that scoping, a query returns content from other branches, or a version this branch already replaced.
The pipeline applies anywhere a corpus has a branching history: IDE assistants searching a developer's feature branch, CMS systems searching draft content separately from published, policy or contract repositories with regional or jurisdictional variants. We use a synthetic mini-repo here so the mechanic is the focus.
This tutorial assumes you're comfortable with [hybrid search](/documentation/search/text-search/#combining-semantic-and-lexical-search-with-hybrid-search), [filtering](/documentation/search/filtering/), and the [Query API](/documentation/search/hybrid-queries/).
Each version of a file is one point in the Qdrant collection, tagged with the `branch` that wrote it and a `seq` (the commit number on that branch). <br>When a later commit overwrites or deletes a file, the previous version's point records which branch did it, and when. A search on a branch is then a single filter: it reads from that branch and the branches it forked from, and skips anything they later replaced.
## Setup
Install the Python packages used throughout the tutorial:
Install the Python client:
```bash
pip install "qdrant-client>=1.18"
@@ -32,50 +30,92 @@ This tutorial uses <a href="/documentation/inference/#qdrant-cloud-inference">Qd
## The Synthetic Corpus
The corpus is a small in-memory documentation site: pricing pages, policies, getting-started guides, and API reference. Three branches: `main` (the live published version), plus two draft branches forked from main, `pricing-refresh` (marketing's draft) and `compliance-update` (legal's draft).
The corpus is a small documentation site: pricing pages, policies, getting-started guides, and API reference. There are three branches: `main` is the published version, with `pricing-refresh` (marketing's) and `compliance-update` (legal's) forked from it.
```python
main_docs = {
"pricing/pro-tier.md": "The Pro tier costs $29 per month and includes 100 GB of storage, unlimited API calls, and team collaboration for up to 10 seats.",
"pricing/enterprise.md": "Enterprise pricing is customized based on volume. Contact sales for a quote, custom SLAs, and an SSO demo.",
"pricing/free-tier.md": "The Free tier is permanently free with 1 GB of storage and 100 API calls per day. No credit card required.",
"policies/refunds.md": "Refunds are processed within 30 days of the original purchase. Pro-rated charges apply for partial periods.",
"policies/data-retention.md": "Customer data is retained for the duration of the contract and 90 days after termination, then permanently deleted.",
"policies/acceptable-use.md": "Our acceptable use policy prohibits scraping, automated abuse, and any content that violates applicable laws.",
"getting-started/install.md": "Install the CLI with pip install our-tool. The default config writes to ~/.our-tool/config.json.",
"getting-started/auth.md": "Authenticate with an API key from your dashboard. Set it via OUR_TOOL_API_KEY or in the config file.",
"getting-started/first-query.md":"Run your first query with our-tool query 'your search here'. Results return as JSON or YAML.",
"guides/python-sdk.md": "The Python SDK wraps every API endpoint. Install with pip install our-tool-sdk and import from our_tool.",
"guides/typescript-sdk.md": "The TypeScript SDK ships type definitions for all API responses. Install with npm install @our-tool/sdk.",
"guides/migrations.md": "Schema migrations run automatically on deployment. Roll back with our-tool migrate --rollback.",
"guides/monitoring.md": "Metrics export in Prometheus format on port 9090. Dashboard templates ship for Grafana.",
"guides/troubleshooting.md": "Common errors and their resolutions. Check the status page first if multiple users are affected.",
"guides/best-practices.md": "Production-grade recommendations for security, performance, and reliability.",
"api/overview.md": "The REST API uses standard HTTPS with bearer token auth. All responses are JSON.",
"api/rate-limits.md": "Free plan is 100 requests per day. Pro plan is unlimited within fair use. Enterprise has custom limits.",
"api/errors.md": "All errors return standard HTTP status codes with a JSON body describing the cause.",
"README.md": "Overview of our tool for new users. Start with the Getting Started guide.",
"CHANGELOG.md": "Release notes by version, newest first. Major releases follow semantic versioning.",
"pricing/pro-tier.md": "The Pro tier costs $29 per month and includes 100 GB of storage, unlimited API calls, and team collaboration for up to 10 seats.",
"pricing/enterprise.md": "Enterprise pricing is customized based on volume. Contact sales for a quote, custom SLAs, and an SSO demo.",
"pricing/free-tier.md": "The Free tier is permanently free with 1 GB of storage and 100 API calls per day. No credit card required.",
"policies/refunds.md": "Refunds are processed within 30 days of the original purchase. Pro-rated charges apply for partial periods.",
"policies/data-retention.md": "Customer data is retained for the duration of the contract and 90 days after termination, then permanently deleted.",
"policies/acceptable-use.md": "Our acceptable use policy prohibits scraping, automated abuse, and any content that violates applicable laws.",
"getting-started/install.md": "Install the CLI with pip install our-tool. The default config writes to ~/.our-tool/config.json.",
"getting-started/auth.md": "Authenticate with an API key from your dashboard. Set it via OUR_TOOL_API_KEY or in the config file.",
"getting-started/first-query.md": "Run your first query with our-tool query 'your search here'. Results return as JSON or YAML.",
"guides/python-sdk.md": "The Python SDK wraps every API endpoint. Install with pip install our-tool-sdk and import from our_tool.",
"guides/typescript-sdk.md": "The TypeScript SDK ships type definitions for all API responses. Install with npm install @our-tool/sdk.",
"guides/migrations.md": "Schema migrations run automatically on deployment. Roll back with our-tool migrate --rollback.",
"guides/monitoring.md": "Metrics export in Prometheus format on port 9090. Dashboard templates ship for Grafana.",
"guides/troubleshooting.md": "Common errors and their resolutions. Check the status page first if multiple users are affected.",
"guides/best-practices.md": "Production-grade recommendations for security, performance, and reliability.",
"api/overview.md": "The REST API uses standard HTTPS with bearer token auth. All responses are JSON.",
"api/rate-limits.md": "Free plan is 100 requests per day. Pro plan is unlimited within fair use. Enterprise has custom limits.",
"api/errors.md": "All errors return standard HTTP status codes with a JSON body describing the cause.",
"README.md": "Overview of our tool for new users. Start with the Getting Started guide.",
"CHANGELOG.md": "Release notes by version, newest first. Major releases follow semantic versioning.",
}
# Marketing draft: bump Pro tier pricing and refresh enterprise messaging.
pricing_refresh_overrides = {
"pricing/pro-tier.md": "The Pro tier costs $39 per month and includes 200 GB of storage, unlimited API calls, team collaboration for unlimited seats, and SSO support.",
"pricing/enterprise.md": "Enterprise pricing starts at $499 per month with custom volume discounts. SSO, audit logs, and dedicated support included.",
# Marketing branch: raise the Pro price and refresh enterprise messaging.
pricing_refresh = {
"pricing/pro-tier.md": "The Pro tier costs $39 per month and includes 200 GB of storage, unlimited API calls, team collaboration for unlimited seats, and SSO support.",
"pricing/enterprise.md": "Enterprise pricing starts at $499 per month with custom volume discounts. SSO, audit logs, and dedicated support included.",
}
# Compliance draft: add GDPR language to the Pro tier page and tighten the refund window.
compliance_update_overrides = {
"pricing/pro-tier.md": "The Pro tier costs $29 per month and includes 100 GB of storage, unlimited API calls, and team collaboration for up to 10 seats. EU customers: data is processed under GDPR with EU-region storage.",
"policies/refunds.md": "Refunds are processed within 14 days for Pro customers and 30 days for Free tier. No questions asked within the first 7 days.",
# Compliance branch: add GDPR language to the Pro tier page and tighten the refund window.
compliance_update = {
"pricing/pro-tier.md": "The Pro tier costs $29 per month and includes 100 GB of storage, unlimited API calls, and team collaboration for up to 10 seats. EU customers: data is processed under GDPR with EU-region storage.",
"policies/refunds.md": "Refunds are processed within 14 days for Pro customers and 30 days for Free tier. No questions asked within the first 7 days.",
}
# main keeps committing after the two branches fork off: it edits a shared file and adds a new one.
main_later_edit = {
"api/rate-limits.md": "Free plan is 60 requests per minute. Pro and Enterprise are unlimited within fair use, with burst credits.",
"policies/sla.md": "Service level agreement: 99.9% uptime for Pro and Enterprise, measured monthly with service credits for breaches.",
}
```
Both drafts edit `pricing/pro-tier.md` with different angles: `pricing-refresh` changes the price and adds SSO, `compliance-update` adds GDPR copy at the original price. Each also touches one additional file unique to it. Everything else is shared across all three branches. Same path, three branches, three distinct versions of `pricing/pro-tier.md`: that's the picture we want to verify.
Two cases this design has to get right:
## Collection Schema and Branch State
1. **Per-branch versions of the same file.** Both forks edit `pricing/pro-tier.md`: `pricing-refresh` raises the price and adds SSO, while `compliance-update` adds GDPR copy at the original price. The same path should resolve to three different versions across the three branches.
2. **Timing.** After both branches fork off `main`, `main` keeps moving: it edits the shared `api/rate-limits.md` and adds `policies/sla.md`. Each fork should still see `main`'s files as they were at the fork point, not these later changes.
The content lives in a Qdrant collection. For simplicity, the branch graph and per-branch live-set live in plain Python dictionaries.
## How a Branch Sees Content
A branch records its fork point as `(parent, parent's seq at fork)`. From there, what the branch sees is its own commits plus everything inherited from its ancestors up to the fork point, minus anything superseded along the way:
```text
main ──● seq0 ───────────────────● seq1
│ (20 base files) (edit api/rate-limits.md, add policies/sla.md)
├── pricing-refresh ● seq0
└── compliance-update ● seq0
```
Both branches forked at `main seq0`, so each sees `main` as it stood at that fork.
Three pieces of runtime state encode this in plain Python:
```python
lineage: dict[str, tuple[str, int] | None] = {} # branch -> (parent, fork_seq)
last_seq: dict[str, int] = {} # branch -> its latest commit seq (-1 before first)
head: dict[str, dict[str, str]] = {} # branch -> {path -> current point id}
```
`lineage` is the durable per-branch record. <br>`head` and `last_seq` are replay bookkeeping, rebuilt from version control on startup.
<aside role="status">
In the tutorial these dicts live in process memory. In production, persist the per-branch <code>lineage</code> records (the only durable state) somewhere small and reachable: a JSON manifest checked into the repo, a metadata table, or a separate Qdrant collection with one point per branch. The <code>head</code> map stays a replay cache, rebuild it from version control alongside the collection.
</aside>
## Collection Schema
Each point in the collection has three payload fields:
- `branch`: the branch that wrote a version.
- `seq`: a per-branch counter we assign on each commit (`0`, `1`, `2`, ...), not the git commit hash. It orders each branch's commits, so the filter can include only the commits at or before each ancestor's fork point.
- `overwritten_in`: a list of `{by, seq}` records, each meaning "replaced in branch `by` at that branch's commit `seq`." Overwrites and deletes both append one.
Every field the visibility filter touches needs a payload index:
```python
from qdrant_client import QdrantClient, models
@@ -89,162 +129,193 @@ client = QdrantClient(
client.create_collection(
collection_name="content",
vectors_config={
"dense": models.VectorParams(size=384, distance=models.Distance.COSINE),
},
sparse_vectors_config={
"bm25": models.SparseVectorParams(modifier=models.Modifier.IDF),
},
)
client.create_payload_index(
collection_name="content",
field_name="branch",
field_schema=models.PayloadSchemaType.KEYWORD,
)
client.create_payload_index(
collection_name="content",
field_name="seq",
field_schema=models.PayloadSchemaType.INTEGER,
)
client.create_payload_index(
collection_name="content",
field_name="overwritten_in[].by",
field_schema=models.PayloadSchemaType.KEYWORD,
)
client.create_payload_index(
collection_name="content",
field_name="overwritten_in[].seq",
field_schema=models.PayloadSchemaType.INTEGER,
)
```
Branch membership lives outside the point payload, so no payload indexes are needed. In this design, where content-derived point IDs are shared across branches, storing live membership on each point (scalar tag, array of branches, or similar) would conflict with that sharing: each branch's update would either overwrite another branch's data or require writes across every shared point.
An alternative design that includes the branch in the point ID itself (e.g., `uuid5(NS, "branch|path|hash")`) would let branch info live on the point safely, at the cost of re-ingesting all of a parent's content on every fork. This tutorial trades that for cheap branching.
Branch state in Python:
```python
# Per-branch live-set: which point IDs are currently visible on this branch's HEAD
live_sets: dict[str, set[str]] = {}
# Per-branch (path -> set of point IDs). The set keeps the structure correct
# for the chunked-file case even though this tutorial uses one chunk per file.
path_to_points: dict[str, dict[str, set[str]]] = {}
```
<aside role="status">
These dictionaries live in memory, so a restart loses them. Production deployments would need to back them with a database, or rebuild them from the version-control system.
</aside>
## Ingest
Point IDs are derived from `(path, content_hash)`, so the same content at the same path always maps to the same point. Branches that share an unchanged file share the point too.
A commit writes a new point for each file it adds or changes, tagged with the branch and its current `seq`. If the file already had a version on this branch, that prior point gets a `{by, seq}` appended to the `overwritten_in` field, marking it superseded.
<br>The point ID is derived from `(branch, seq, path)`, so replaying the same history produces the same IDs and the rebuild path can `upsert` without duplicating points.
<aside role="status">
This tutorial uses one point per file. For chunked corpora the branch logic is unchanged, the point ID gains a stable chunk anchor, becoming <code>(branch, seq, path, anchor)</code>. See <a href="/course/essentials/day-1/chunking-strategies/">chunking strategies</a> for how to split each content type.
</aside>
```python
import hashlib
import uuid
CONTENT_NS = uuid.UUID("00000000-0000-0000-0000-000000000001")
NS = uuid.UUID("00000000-0000-0000-0000-000000000042")
def point_id(path: str, content: str) -> str:
"""Deterministic UUID from (path, content): same inputs always produce the same ID."""
content_hash = hashlib.sha256(content.encode()).hexdigest()
return str(uuid.uuid5(CONTENT_NS, f"{path}|{content_hash}"))
# BM25's key parameter for short fields: the average word count.
word_counts = [len(text.split()) for text in main_docs.values()]
AVG_LEN = round(sum(word_counts) / len(word_counts), 1) # ~15.3 here
def point_id(branch: str, seq: int, path: str) -> str:
return str(uuid.uuid5(NS, f"{branch}|{seq}|{path}"))
def create_branch(name: str, parent: str | None):
"""Create a branch. A fork copies the parent's live-set, so the child starts identical and diverges as it commits."""
"""Record a branch. A fork inherits the parent's view, no points written."""
last_seq[name] = -1
if parent is None:
live_sets[name] = set()
path_to_points[name] = {}
lineage[name] = None
head[name] = {}
else:
live_sets[name] = set(live_sets[parent])
path_to_points[name] = {p: set(ids) for p, ids in path_to_points[parent].items()}
# fork point: parent + parent's latest seq
lineage[name] = (parent, last_seq[parent])
head[name] = dict(head[parent]) # start from parent's view
def ingest_commit(branch: str, files: dict[str, str]):
"""Upsert a commit's files, then update the branch's live-set: evict each path's old point IDs and add the new ones."""
def supersede(branch: str, path: str, at_seq: int):
"""Mark the version of `path` that `branch` currently sees, at `at_seq`."""
prev = head[branch].get(path)
if prev is None:
return
point = client.retrieve("content", ids=[prev])[0]
marks = point.payload["overwritten_in"]
marks.append({"by": branch, "seq": at_seq})
client.set_payload(
"content",
payload={"overwritten_in": marks},
points=[prev],
)
def commit(branch: str, writes: dict[str, str] | None = None, deletes: list[str] | None = None):
"""One commit at one seq: write or overwrite files, and/or delete files."""
writes, deletes = writes or {}, deletes or []
seq = last_seq[branch] + 1
last_seq[branch] = seq
points = []
# New point IDs grouped by path (one per file here; a set so the multi-chunk case works too).
new_ids_by_path: dict[str, set[str]] = {}
for path, content in files.items():
pid = point_id(path, content)
new_ids_by_path.setdefault(path, set()).add(pid)
for path, content in writes.items():
supersede(branch, path, seq) # mark the version this commit replaces
pid = point_id(branch, seq, path)
points.append(models.PointStruct(
id=pid,
vector={
"dense": models.Document(text=content, model="sentence-transformers/all-minilm-l6-v2"),
"bm25": models.Document(text=content, model="qdrant/bm25", options={"avg_len": 12}),
"bm25": models.Document(
text=content,
model="qdrant/bm25",
options={"avg_len": AVG_LEN},
),
},
payload={
"path": path,
"content": content,
"branch": branch,
"seq": seq,
"overwritten_in": [],
},
payload={"path": path, "content": content},
))
# Write to Qdrant first; only mutate branch state after the upsert is durably
# acknowledged. If the upsert fails the exception propagates and the in-memory
# branch HEAD is unchanged.
client.upsert(collection_name="content", points=points, wait=True)
for path, new_ids in new_ids_by_path.items():
prev_ids = path_to_points[branch].get(path, set())
live_sets[branch] -= prev_ids
live_sets[branch] |= new_ids
path_to_points[branch][path] = new_ids
def delete_files(branch: str, paths: list[str]):
"""Delete files from a branch by dropping their point IDs from the live-set."""
for path in paths:
ids = path_to_points[branch].pop(path, set())
live_sets[branch] -= ids
head[branch][path] = pid
for path in deletes:
supersede(branch, path, seq) # same record, no replacement point
head[branch].pop(path, None)
if points:
client.upsert(collection_name="content", points=points, wait=True)
```
Removing a file from a branch drops its point IDs from that branch's live-set, so queries on that branch stop returning it. The point itself stays in Qdrant: the tutorial never deletes points, so it remains both for other branches that reference it and after the last branch drops it (reclaiming those is the storage-growth concern in Scaling to Production). A filter-based design that walks ancestors would need a tombstone here, because the ancestor clause keeps matching the parent's copy. The live-set sidesteps that.
A delete writes the same record as an overwrite, with no replacement point behind it, and it is a real commit, so it carries a `seq` like any other. That `seq` is what lets a branch that forked before the delete keep the file, while the branch that deleted it (and anything forked after) does not.
Set up the fixture:
Build the fixture:
```python
create_branch("main", parent=None)
ingest_commit("main", main_docs)
commit("main", writes=main_docs)
create_branch("pricing-refresh", parent="main") # marketing draft forks from main's HEAD
ingest_commit("pricing-refresh", pricing_refresh_overrides)
create_branch("pricing-refresh", parent="main") # forks at main seq0
commit("pricing-refresh", writes=pricing_refresh)
create_branch("compliance-update", parent="main") # compliance draft forks from main's HEAD as a sibling
ingest_commit("compliance-update", compliance_update_overrides)
create_branch("compliance-update", parent="main") # also at main seq0
commit(
"compliance-update",
writes=compliance_update,
deletes=["policies/acceptable-use.md"],
)
commit("main", writes=main_later_edit) # main seq1, after both forks
```
Forking copied main's set, and each draft swapped in new point IDs only for the files it changed. Everything else stays shared:
![Three branch live-sets (main, pricing-refresh, and compliance-update), each holding 20 point IDs. The two drafts forked from main and copied its set, then swapped in new point IDs only for the files they changed; all other files keep the same IDs as main. One Qdrant collection stores 24 physical points: seven versions of the three edited files plus 17 shared files.](/documentation/tutorials/branch-aware-search/branch-live-sets.png)
After ingest:
- Each live-set has 20 point IDs (one per document at that branch's HEAD).
- `main`'s 20 IDs include the original `pricing/pro-tier.md`, `pricing/enterprise.md`, and `policies/refunds.md`.
- `pricing-refresh`'s 20 IDs swap in the new pricing copy for `pricing/pro-tier.md` and the new enterprise copy for `pricing/enterprise.md`. The other 18 are shared with main.
- `compliance-update`'s 20 IDs swap in the GDPR-extended `pricing/pro-tier.md` and the tightened `policies/refunds.md`. The other 18 are shared with main.
- The Qdrant collection physically stores 24 points: 20 from main, plus four new versions written by the two drafts. Old versions stay in the collection but drop out of the branches that overrode them.
Each fork wrote only the files it changed, and `main`'s later commit wrote two points (the edited `api/rate-limits.md` and the new `policies/sla.md`). The collection holds 26 physical points: the 20 base files, two versions from each fork, and the two from `main`'s later commit.
## Query
Branch-aware search is a single hybrid query restricted to the points in the branch's live-set, using Qdrant's [Has id](/documentation/search/filtering/#has-id) filter:
A query on a branch is one filter built from a lineage walk. <br>For each branch on the path there's a **cutoff**, the highest `seq` to include from that branch: for the querying branch itself, that's its latest commit; for an ancestor, the `seq` at which this line forked from it. <br>Both filter clauses key off those cutoffs. The `should` clause (the candidates to consider) gathers each branch's versions up to its cutoff. The `must_not` clause (the ones to drop) excludes anything that branch superseded at or before its cutoff, using a [nested filter](/documentation/search/filtering/#nested-object-filter) so `by` and `seq` match on the same record:
```python
def candidate_clause(b: str, cut: int) -> models.Filter:
"""Versions written on branch b with seq up to cut."""
return models.Filter(must=[
models.FieldCondition(key="branch", match=models.MatchValue(value=b)),
models.FieldCondition(key="seq", range=models.Range(lte=cut)),
])
def exclusion_clause(b: str, cut: int) -> models.NestedCondition:
"""Match if overwritten_in has {by: b, seq <= cut} on the same record."""
return models.NestedCondition(nested=models.Nested(
key="overwritten_in",
filter=models.Filter(must=[
models.FieldCondition(key="by", match=models.MatchValue(value=b)),
models.FieldCondition(key="seq", range=models.Range(lte=cut)),
]),
))
def visibility_filter(branch: str) -> models.Filter:
# Walk lineage: (branch, cutoff) for this branch then each ancestor.
cutoffs = [(branch, last_seq[branch])]
node = lineage[branch]
while node is not None:
parent, fork = node
cutoffs.append((parent, fork))
node = lineage[parent]
should, must_not = [], []
for b, cut in cutoffs:
should.append(candidate_clause(b, cut))
must_not.append(exclusion_clause(b, cut))
return models.Filter(should=should, must_not=must_not)
def search(branch: str, query: str, limit: int = 5):
branch_filter = models.Filter(
must=[models.HasIdCondition(has_id=list(live_sets[branch]))]
)
return client.query_points(
collection_name="content",
prefetch=[
models.Prefetch(
query=models.Document(text=query, model="sentence-transformers/all-minilm-l6-v2"),
using="dense",
filter=branch_filter,
limit=50,
),
models.Prefetch(
query=models.Document(text=query, model="qdrant/bm25", options={"avg_len": 12}),
using="bm25",
filter=branch_filter,
limit=50,
),
],
query=models.FusionQuery(fusion=models.Fusion.RRF),
query=models.Document(
text=query,
model="qdrant/bm25",
options={"avg_len": AVG_LEN},
),
using="bm25",
query_filter=visibility_filter(branch),
limit=limit,
with_payload=True,
).points
```
<aside role="status">
The branch filter works with any search type. This tutorial uses hybrid (dense + BM25) as it is a solid default for mixed text. Use dense alone when relevance is mostly semantic, sparse/BM25 when exact terms matter most (identifiers, error strings, or codes), and hybrid when you need both. See <a href="/documentation/search/text-search/">Text Search</a> and <a href="/documentation/search/hybrid-queries/">Hybrid Queries</a> for details.
</aside>
Run the same query against all three branches and print each one's top result:
Run a query that lands on `pricing/pro-tier.md`, against all three branches:
```python
query = "Pro tier monthly cost"
print("main:", search("main", query, limit=1)[0].payload["content"])
print("pricing-refresh:", search("pricing-refresh", query, limit=1)[0].payload["content"])
print("compliance-update:", search("compliance-update", query, limit=1)[0].payload["content"])
for branch in ["main", "pricing-refresh", "compliance-update"]:
print(f"{branch}:", search(branch, "Pro tier storage and seats", limit=1)[0].payload["content"])
# Output:
# main: The Pro tier costs $29 per month and includes 100 GB of storage, unlimited API calls, and team collaboration for up to 10 seats.
@@ -252,24 +323,33 @@ print("compliance-update:", search("compliance-update", query, limit=1)[0].paylo
# compliance-update: The Pro tier costs $29 per month and includes 100 GB of storage, unlimited API calls, and team collaboration for up to 10 seats. EU customers: data is processed under GDPR with EU-region storage.
```
Same query, three branches, three answers. The drafts stay isolated: `pricing-refresh`'s result never mentions GDPR, and `compliance-update`'s never mentions SSO.
Same query, three branches, three different answers. `pricing-refresh` returns its $39 page with SSO. `compliance-update` returns its GDPR variant. `main` returns the original.
## Adapting to Other Content Types
Now the timing case. `main` edited `api/rate-limits.md` at `seq 1`, after both branches forked. `main` now shows the new text; the forks still see the version they inherited:
This tutorial stores one chunk per document because each document is short. Real corpora are usually chunked, and the mechanic carries over: the live-set still holds point IDs, the query is still a Has id filter, and the point ID gains a chunk anchor, becoming `(path, anchor, content_hash)`, so chunks within a file get distinct IDs. Editing one chunk changes only that chunk's point ID:
```python
for branch in ["main", "pricing-refresh"]:
print(f"{branch}:", search(branch, "rate limits requests free plan", limit=1)[0].payload["content"])
![A code file split by function. The point ID for one function is built as uuid5 of three parts: the path, the anchor (its AST symbol path), and the content_hash. Editing that function changes only its content_hash, so only its point ID changes; the unchanged functions keep their IDs and stay shared across branches.](/documentation/tutorials/branch-aware-search/chunk-id-mapping.png)
# Output:
# main: Free plan is 60 requests per minute. Pro and Enterprise are unlimited within fair use, with burst credits.
# pricing-refresh: Free plan is 100 requests per day. Pro plan is unlimited within fair use. Enterprise has custom limits.
```
The anchor has to be stable, like a heading path, an AST symbol path, or a clause number, so unchanged content keeps its ID across edits. Unstable anchors like line numbers or byte offsets shift on every insertion and force the whole file to re-embed. For how to split each content type, see the [chunking strategies course](/course/essentials/day-1/chunking-strategies/).
On update, re-chunk the whole file and pass all of its chunks, not only the changed ones: eviction is path-scoped, so a partial update would drop the unchanged chunks from the live-set. Passing every chunk means re-submitting the unchanged ones, which still carry the same IDs. Set the upsert's [update mode](/documentation/manage-data/points/#update-mode) to `models.UpdateMode.INSERT_ONLY` so those existing IDs are skipped and only the new or changed chunks get written.
`pricing-refresh` forked at `main`'s `seq 0`, before the `seq 1` edit. Its cutoff for `main` is `0`, so the record stamped at `seq 1` falls past the cutoff, the exclusion doesn't apply, and the pre-edit version stays visible. The same cutoff hides `policies/sla.md` (which `main` added at `seq 1`) from both forks.
## Scaling to Production
This pattern uses the simplest Qdrant primitives on purpose. The inline Has id filter carries the branch's whole ID set on every query, which stays comfortable into the thousands of points per branch, enough for most corpora. A few spots still need hardening before production:
- **The filter grows with lineage depth, not corpus size.** The query walks the lineage once, adding one `should` clause and one `must_not` clause per ancestor. The filter's size tracks how deep a branch sits, not how many files or versions the collection holds.
- **Storage growth.** Every edit adds a new point and keeps the old one, so long-lived branches accumulate versions. Reclaiming superseded points is a production concern, much like log compaction.
- **Qdrant is a derived index.** Branch lineage and history live in version control. Build the collection by replaying that history, and if the index ever drifts, rebuild it. Nothing here needs the index to be the source of truth.
- **Durable branch state.** The live-sets here live in Python dictionaries for clarity. In production, back them with a database, or rebuild them from the version-control system, which is the real source of truth.
- **Storage grows with edits.** Every overwrite keeps the old version as a point, so long-lived branches accumulate versions. Reclaiming superseded versions is a production concern.
- **Out of scope.** Three-way merges, force-push, and concurrent writers on a branch are each their own design decision and aren't needed to show the core mechanic. Treat Qdrant as the search index and git as the source of truth: if the index drifts, rebuild it.
- **Out of scope.** Three-way merges, force-push, and concurrent writers on a branch are each their own design decision and are not needed to show the core mechanic.
## Related Reading
- [Filtering](/documentation/search/filtering/) for the full set of filter clauses.
- [Text Search](/documentation/search/text-search/) for BM25 and the `avg_len` parameter.
- [Payload Indexing](/documentation/manage-data/indexing/#payload-index) for the index types the visibility fields use.
Binary file not shown.

Before

Width:  |  Height:  |  Size: 48 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 75 KiB