Apply adversarial-review fixes to hybrid-queries and snippets

Address findings from a code-bot adversarial review of the hybrid-search
materials:

- `hybrid-formula-decay/` (all 7 language sources + http.md +
  `_description.md`): wrap the decay term in `MultExpression(mult=[0.1, ...])`
  in every language. Previously the snippets summed `$score` with
  `ExpDecayExpression` directly, modeling the failure mode the docs
  explicitly warn against (un-weighted decay crowds out small RRF scores).
  Also document the `defaults` requirement and the recommended datetime
  payload index in `_description.md`. Build validated across all 6 SDKs.

- `hybrid-rrf/go.go`: add `Limit: qdrant.PtrOf(uint64(20))` to both
  prefetches so the Go snippet matches the other language tabs.

- `hybrid-queries.md`:
  - Reframe the weighted-RRF intro to drop "semantic search model
    understands meaning better than a simple keyword matcher". On
    SciFact (the corpus in the companion notebook) BM25 actually beats
    dense, so the universal claim was contradicted by our own data.
  - Clarify that the notebook provides a tuning helper to adapt to a
    train/val split, not that it demonstrates the split itself.
  - Add a one-line note that Qdrant uses zero-based rank positions so
    readers can verify the RRF formula against actual scores.
  - Apply brand-voice fixes: Title Case on "Multi-Stage Queries" and
    "Re-Scoring Examples", replace "all the above techniques" with
    "all of these techniques".

`generated/*.md` regenerated via `./docker.sh ./generate-md.py`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
Dylan Couzon
2026-05-15 21:54:47 -04:00
co-authored by Claude Opus 4.7
parent fc299a0e9c
commit c0128a53bc
17 changed files with 169 additions and 94 deletions
@@ -1 +1,3 @@
This code snippet shows the canonical pattern for combining hybrid search fusion with business-logic ranking. The inner prefetch fuses sparse and dense results with RRF, then the outer `FormulaQuery` applies exponential decay on a `published_at` payload field to boost recent documents. Use this pattern any time you want fusion plus recency, popularity, geo decay, or category-conditional multipliers, rather than trying to encode those signals as fusion weights.
This code snippet shows the canonical pattern for combining hybrid search fusion with business-logic ranking. The inner prefetch fuses sparse and dense results with RRF, then the outer `FormulaQuery` applies exponential decay on a `published_at` payload field to boost recent documents. The decay term is wrapped in `mult` with a `0.1` coefficient so it nudges the ranking rather than crowding out the small RRF scores. Use this pattern any time you want fusion plus recency, popularity, geo decay, or category-conditional multipliers, rather than trying to encode those signals as fusion weights.
For production: `published_at` must exist on every point, or supply `defaults` in the `FormulaQuery` to fill missing values. A datetime payload index on `published_at` keeps query latency low.
@@ -35,15 +35,22 @@ public class Snippet
Sum =
{
"$score", // the fused score from the RRF prefetch
Expression.FromExpDecay(
new()
new MultExpression
{
Mult =
{
X = Expression.FromDateTimeKey("published_at"),
Target = Expression.FromDateTime("YYYY-MM-DDT00:00:00Z"),
Scale = 86400 * 180, // 180 days in seconds
Midpoint = 0.5f
0.1f, // caps decay contribution
Expression.FromExpDecay(
new()
{
X = Expression.FromDateTimeKey("published_at"),
Target = Expression.FromDateTime("YYYY-MM-DDT00:00:00Z"),
Scale = 86400 * 180, // 180 days in seconds
Midpoint = 0.5f
}
)
}
)
}
}
}
},
@@ -32,15 +32,22 @@ await client.QueryAsync(
Sum =
{
"$score", // the fused score from the RRF prefetch
Expression.FromExpDecay(
new()
new MultExpression
{
Mult =
{
X = Expression.FromDateTimeKey("published_at"),
Target = Expression.FromDateTime("YYYY-MM-DDT00:00:00Z"),
Scale = 86400 * 180, // 180 days in seconds
Midpoint = 0.5f
0.1f, // caps decay contribution
Expression.FromExpDecay(
new()
{
X = Expression.FromDateTimeKey("published_at"),
Target = Expression.FromDateTime("YYYY-MM-DDT00:00:00Z"),
Scale = 86400 * 180, // 180 days in seconds
Midpoint = 0.5f
}
)
}
)
}
}
}
},
@@ -34,11 +34,16 @@ client.Query(context.Background(), &qdrant.QueryPoints{
Expression: qdrant.NewExpressionSum(&qdrant.SumExpression{
Sum: []*qdrant.Expression{
qdrant.NewExpressionVariable("$score"), // the fused score from the RRF prefetch
qdrant.NewExpressionExpDecay(&qdrant.DecayParamsExpression{
X: qdrant.NewExpressionDatetimeKey("published_at"),
Target: qdrant.NewExpressionDatetime("YYYY-MM-DDT00:00:00Z"),
Scale: qdrant.PtrOf(float32(86400 * 180)), // 180 days in seconds
Midpoint: qdrant.PtrOf(float32(0.5)),
qdrant.NewExpressionMult(&qdrant.MultExpression{
Mult: []*qdrant.Expression{
qdrant.NewExpressionConstant(0.1), // caps decay contribution
qdrant.NewExpressionExpDecay(&qdrant.DecayParamsExpression{
X: qdrant.NewExpressionDatetimeKey("published_at"),
Target: qdrant.NewExpressionDatetime("YYYY-MM-DDT00:00:00Z"),
Scale: qdrant.PtrOf(float32(86400 * 180)), // 180 days in seconds
Midpoint: qdrant.PtrOf(float32(0.5)),
}),
},
}),
},
}),
@@ -1,7 +1,9 @@
```java
import static io.qdrant.client.ExpressionFactory.constant;
import static io.qdrant.client.ExpressionFactory.datetime;
import static io.qdrant.client.ExpressionFactory.datetimeKey;
import static io.qdrant.client.ExpressionFactory.expDecay;
import static io.qdrant.client.ExpressionFactory.mult;
import static io.qdrant.client.ExpressionFactory.sum;
import static io.qdrant.client.ExpressionFactory.variable;
import static io.qdrant.client.QueryFactory.formula;
@@ -12,6 +14,7 @@ import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
import io.qdrant.client.grpc.Points.DecayParamsExpression;
import io.qdrant.client.grpc.Points.Formula;
import io.qdrant.client.grpc.Points.MultExpression;
import io.qdrant.client.grpc.Points.PrefetchQuery;
import io.qdrant.client.grpc.Points.QueryPoints;
import io.qdrant.client.grpc.Points.Rrf;
@@ -49,13 +52,18 @@ client.queryAsync(
SumExpression.newBuilder()
.addSum(variable("$score"))
.addSum(
expDecay(
DecayParamsExpression.newBuilder()
.setX(datetimeKey("published_at"))
.setTarget(
datetime("YYYY-MM-DDT00:00:00Z"))
.setScale(86400 * 180)
.setMidpoint(0.5f)
mult(
MultExpression.newBuilder()
.addMult(constant(0.1f))
.addMult(
expDecay(
DecayParamsExpression.newBuilder()
.setX(datetimeKey("published_at"))
.setTarget(
datetime("YYYY-MM-DDT00:00:00Z"))
.setScale(86400 * 180)
.setMidpoint(0.5f)
.build()))
.build()))
.build()))
.build()))
@@ -25,14 +25,17 @@ client.query_points(
formula=models.SumExpression(
sum=[
"$score", # the fused score from the RRF prefetch
models.ExpDecayExpression(
exp_decay=models.DecayParamsExpression(
x=models.DatetimeKeyExpression(datetime_key="published_at"),
target=models.DatetimeExpression(datetime="YYYY-MM-DDT00:00:00Z"),
scale=86400 * 180, # 180 days in seconds
midpoint=0.5,
)
),
models.MultExpression(mult=[
0.1, # caps decay contribution; un-weighted decay [0, 1] would otherwise crowd out small RRF scores
models.ExpDecayExpression(
exp_decay=models.DecayParamsExpression(
x=models.DatetimeKeyExpression(datetime_key="published_at"),
target=models.DatetimeExpression(datetime="YYYY-MM-DDT00:00:00Z"),
scale=86400 * 180, # 180 days in seconds
midpoint=0.5,
)
),
]),
]
)
),
@@ -29,12 +29,15 @@ client.query(
.query(
FormulaBuilder::new(Expression::sum_with([
Expression::score(),
Expression::exp_decay(
DecayParamsExpressionBuilder::new(Expression::datetime_key("published_at"))
.target(Expression::datetime("YYYY-MM-DDT00:00:00Z"))
.scale(86400.0 * 180.0)
.midpoint(0.5),
),
Expression::mult_with([
Expression::constant(0.1),
Expression::exp_decay(
DecayParamsExpressionBuilder::new(Expression::datetime_key("published_at"))
.target(Expression::datetime("YYYY-MM-DDT00:00:00Z"))
.scale(86400.0 * 180.0)
.midpoint(0.5),
),
]),
])),
)
.limit(10u64),
@@ -28,12 +28,17 @@ await client.query("{collection_name}", {
sum: [
"$score", // the fused score from the RRF prefetch
{
exp_decay: {
x: { datetime_key: "published_at" },
target: { datetime: "YYYY-MM-DDT00:00:00Z" },
scale: 86400 * 180, // 180 days in seconds
midpoint: 0.5,
},
mult: [
0.1, // caps decay contribution; un-weighted decay [0, 1] would otherwise crowd out small RRF scores
{
exp_decay: {
x: { datetime_key: "published_at" },
target: { datetime: "YYYY-MM-DDT00:00:00Z" },
scale: 86400 * 180, // 180 days in seconds
midpoint: 0.5,
},
},
],
},
],
},
@@ -38,11 +38,16 @@ func Main() {
Expression: qdrant.NewExpressionSum(&qdrant.SumExpression{
Sum: []*qdrant.Expression{
qdrant.NewExpressionVariable("$score"), // the fused score from the RRF prefetch
qdrant.NewExpressionExpDecay(&qdrant.DecayParamsExpression{
X: qdrant.NewExpressionDatetimeKey("published_at"),
Target: qdrant.NewExpressionDatetime("YYYY-MM-DDT00:00:00Z"),
Scale: qdrant.PtrOf(float32(86400 * 180)), // 180 days in seconds
Midpoint: qdrant.PtrOf(float32(0.5)),
qdrant.NewExpressionMult(&qdrant.MultExpression{
Mult: []*qdrant.Expression{
qdrant.NewExpressionConstant(0.1), // caps decay contribution
qdrant.NewExpressionExpDecay(&qdrant.DecayParamsExpression{
X: qdrant.NewExpressionDatetimeKey("published_at"),
Target: qdrant.NewExpressionDatetime("YYYY-MM-DDT00:00:00Z"),
Scale: qdrant.PtrOf(float32(86400 * 180)), // 180 days in seconds
Midpoint: qdrant.PtrOf(float32(0.5)),
}),
},
}),
},
}),
@@ -25,16 +25,21 @@ POST /collections/{collection_name}/points/query
"sum": [
"$score", // the fused score from the RRF prefetch
{
"exp_decay": {
"x": {
"datetime_key": "published_at"
},
"target": {
"datetime": "YYYY-MM-DDT00:00:00Z"
},
"scale": 15552000, // 180 days in seconds
"midpoint": 0.5
}
"mult": [
0.1, // caps decay contribution
{
"exp_decay": {
"x": {
"datetime_key": "published_at"
},
"target": {
"datetime": "YYYY-MM-DDT00:00:00Z"
},
"scale": 15552000, // 180 days in seconds
"midpoint": 0.5
}
}
]
}
]
}
@@ -1,8 +1,10 @@
package com.example.snippets_amalgamation;
import static io.qdrant.client.ExpressionFactory.constant;
import static io.qdrant.client.ExpressionFactory.datetime;
import static io.qdrant.client.ExpressionFactory.datetimeKey;
import static io.qdrant.client.ExpressionFactory.expDecay;
import static io.qdrant.client.ExpressionFactory.mult;
import static io.qdrant.client.ExpressionFactory.sum;
import static io.qdrant.client.ExpressionFactory.variable;
import static io.qdrant.client.QueryFactory.formula;
@@ -13,6 +15,7 @@ import io.qdrant.client.QdrantClient;
import io.qdrant.client.QdrantGrpcClient;
import io.qdrant.client.grpc.Points.DecayParamsExpression;
import io.qdrant.client.grpc.Points.Formula;
import io.qdrant.client.grpc.Points.MultExpression;
import io.qdrant.client.grpc.Points.PrefetchQuery;
import io.qdrant.client.grpc.Points.QueryPoints;
import io.qdrant.client.grpc.Points.Rrf;
@@ -52,13 +55,18 @@ public class Snippet {
SumExpression.newBuilder()
.addSum(variable("$score"))
.addSum(
expDecay(
DecayParamsExpression.newBuilder()
.setX(datetimeKey("published_at"))
.setTarget(
datetime("YYYY-MM-DDT00:00:00Z"))
.setScale(86400 * 180)
.setMidpoint(0.5f)
mult(
MultExpression.newBuilder()
.addMult(constant(0.1f))
.addMult(
expDecay(
DecayParamsExpression.newBuilder()
.setX(datetimeKey("published_at"))
.setTarget(
datetime("YYYY-MM-DDT00:00:00Z"))
.setScale(86400 * 180)
.setMidpoint(0.5f)
.build()))
.build()))
.build()))
.build()))
@@ -24,14 +24,17 @@ client.query_points(
formula=models.SumExpression(
sum=[
"$score", # the fused score from the RRF prefetch
models.ExpDecayExpression(
exp_decay=models.DecayParamsExpression(
x=models.DatetimeKeyExpression(datetime_key="published_at"),
target=models.DatetimeExpression(datetime="YYYY-MM-DDT00:00:00Z"),
scale=86400 * 180, # 180 days in seconds
midpoint=0.5,
)
),
models.MultExpression(mult=[
0.1, # caps decay contribution; un-weighted decay [0, 1] would otherwise crowd out small RRF scores
models.ExpDecayExpression(
exp_decay=models.DecayParamsExpression(
x=models.DatetimeKeyExpression(datetime_key="published_at"),
target=models.DatetimeExpression(datetime="YYYY-MM-DDT00:00:00Z"),
scale=86400 * 180, # 180 days in seconds
midpoint=0.5,
)
),
]),
]
)
),
@@ -29,12 +29,15 @@ pub async fn main() -> anyhow::Result<()> {
.query(
FormulaBuilder::new(Expression::sum_with([
Expression::score(),
Expression::exp_decay(
DecayParamsExpressionBuilder::new(Expression::datetime_key("published_at"))
.target(Expression::datetime("YYYY-MM-DDT00:00:00Z"))
.scale(86400.0 * 180.0)
.midpoint(0.5),
),
Expression::mult_with([
Expression::constant(0.1),
Expression::exp_decay(
DecayParamsExpressionBuilder::new(Expression::datetime_key("published_at"))
.target(Expression::datetime("YYYY-MM-DDT00:00:00Z"))
.scale(86400.0 * 180.0)
.midpoint(0.5),
),
]),
])),
)
.limit(10u64),
@@ -27,12 +27,17 @@ await client.query("{collection_name}", {
sum: [
"$score", // the fused score from the RRF prefetch
{
exp_decay: {
x: { datetime_key: "published_at" },
target: { datetime: "YYYY-MM-DDT00:00:00Z" },
scale: 86400 * 180, // 180 days in seconds
midpoint: 0.5,
},
mult: [
0.1, // caps decay contribution; un-weighted decay [0, 1] would otherwise crowd out small RRF scores
{
exp_decay: {
x: { datetime_key: "published_at" },
target: { datetime: "YYYY-MM-DDT00:00:00Z" },
scale: 86400 * 180, // 180 days in seconds
midpoint: 0.5,
},
},
],
},
],
},
@@ -16,10 +16,12 @@ client.Query(context.Background(), &qdrant.QueryPoints{
{
Query: qdrant.NewQuerySparse([]uint32{1, 42}, []float32{0.22, 0.8}),
Using: qdrant.PtrOf("sparse"),
Limit: qdrant.PtrOf(uint64(20)),
},
{
Query: qdrant.NewQueryDense([]float32{0.01, 0.45, 0.67}),
Using: qdrant.PtrOf("dense"),
Limit: qdrant.PtrOf(uint64(20)),
},
},
Query: qdrant.NewQueryRRF(&qdrant.Rrf{}),
@@ -20,10 +20,12 @@ func Main() {
{
Query: qdrant.NewQuerySparse([]uint32{1, 42}, []float32{0.22, 0.8}),
Using: qdrant.PtrOf("sparse"),
Limit: qdrant.PtrOf(uint64(20)),
},
{
Query: qdrant.NewQueryDense([]float32{0.01, 0.45, 0.67}),
Using: qdrant.PtrOf("dense"),
Limit: qdrant.PtrOf(uint64(20)),
},
},
Query: qdrant.NewQueryRRF(&qdrant.Rrf{}),
@@ -53,6 +53,8 @@ Where:
- $r_d$ is the rank of document $d$ in ranking $r$
- $w_r$ is the weight of ranking $r$ (set to 1 by default)
_Qdrant uses zero-based rank positions; the top result has $r_d = 0$._
Because $w_r$ defaults to 1, without setting explicit weights, the formula can be simplified to the original RRF function:
$$ score(d\in D) = \sum_{r_d\in R(d)} \frac{1}{k + r_d} $$
@@ -71,7 +73,7 @@ To change the value of constant $k$ in the formula, use the dedicated `rrf` quer
#### Weighted RRF
_Available as of v1.17.0_
By default, each query is assigned an equal weight. In reality, some queries are stronger, more discriminative, or more domain-specific than others. For example, a semantic search model understands meaning better than a simple keyword matcher. Assigning equal weight to both can cause the weaker model to negatively influence results, leading to a suboptimal search experience. To address this, you can assign greater weight to rankers that perform well.
By default, each query is assigned an equal weight. In reality, one retriever is often stronger than the other for a given workload. For example, a dense retriever may dominate on natural-language queries, while BM25 may win on identifier-heavy ones. Assigning equal weight to both can let the weaker retriever drag down results. To address this, you can assign greater weight to rankers that perform well on your evaluation set.
The `rrf` query allows you to configure relative weights for each of the prefetches. For example, if you have two prefetches and assign a weight of 3.0 to the first and 1.0 to the second, a document ranked third in the first query scores the same as a document ranked first in the second query. In the case of non-overlapping result sets, these weights return three results from the first set for every one result from the second set.
@@ -81,7 +83,7 @@ Weights should be provided as an array of numbers, where each weight is applied
Weights are a hyperparameter, not a free knob. A held-out eval is the most defensible way to set them.
- **With an eval set (queries paired with known-relevant docs):** split your eval queries in two. Search the weight space on the first half, then report metrics on the second half (held out from the search). Reporting on the same set you tuned on inflates the result. The [Choosing a Fusion Method notebook](https://githubtocolab.com/qdrant/examples/blob/master/fusion-methods/Choosing_a_Fusion_Method.ipynb) demonstrates this with a reusable `tune_rrf_weights` grid-search helper. Random search and Bayesian optimization (Optuna, hyperopt) work equally well for two-retriever fusion.
- **With an eval set (queries paired with known-relevant docs):** split your eval queries in two. Search the weight space on the first half, then report metrics on the second half (held out from the search). Reporting on the same set you tuned on inflates the result. The [Choosing a Fusion Method notebook](https://githubtocolab.com/qdrant/examples/blob/master/fusion-methods/Choosing_a_Fusion_Method.ipynb) provides a reusable `tune_rrf_weights` grid-search helper you can adapt to a train/val split. Random search and Bayesian optimization (Optuna, hyperopt) work equally well for two-retriever fusion.
- **Without an eval set:** leave weights at the default `(1.0, 1.0)`. Hand-tuned weights without measurement are unlikely to beat the default reliably.
Retune when your retrievers change (new embedding model, new chunking), when your corpus drifts substantially, or on a fixed cadence with a fresh eval sample.
@@ -130,7 +132,7 @@ For a deeper breakdown of when to prefer each, see the [FAQ on RRF vs. DBSF](/do
<aside role="status">A common request is "alpha-weighted linear combination of dense and sparse scores." This is unreliable without first normalizing the scores: dense (cosine, bounded) and sparse (BM25, unbounded) scores live on different scales that also shift per query, so a fixed alpha over raw scores tends to be dominated by whichever retriever has larger raw magnitudes on a given query. RRF sidesteps this by using ranks. DBSF sidesteps it by normalizing distributions. <code>FormulaQuery</code> can do it explicitly if you write the normalization yourself.</aside>
## Multi-stage queries
## Multi-Stage Queries
In general, larger vector representations give more accurate search results, but makes them more expensive to compute.
@@ -150,7 +152,7 @@ such that the coarse results are fetched first, and then they are refined later
<aside role="status">Disable the HNSW index for vectors used only for rescoring by setting <code>m=0</code> in the vector's HNSW configuration. Rescoring does not use the HNSW index, so disabling it will free up memory.</aside>
### Re-scoring examples
### Re-Scoring Examples
Fetch 1000 results using a shorter MRL byte vector, then re-score them using the full vector and get the top 10.
@@ -160,7 +162,7 @@ Fetch 100 results using the default vector, then re-score them using a multi-vec
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rescoring-multivector/" >}}
It is possible to combine all the above techniques in a single query:
You can combine all of these techniques in a single query:
{{< code-snippet path="/documentation/headless/snippets/query-points/hybrid-rescoring-multistage/" >}}