Fix SIMD index links in SPLADE series parts 1 & 3

Point to the actual posting_list.rs implementation instead of the
sparse-vectors article.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Evgeniya Sukhodolskaya
2026-06-04 11:43:21 +02:00
co-authored by Claude Sonnet 4.6
parent 9ab2af0542
commit 2d8e995a8f
2 changed files with 2 additions and 2 deletions
@@ -132,7 +132,7 @@ client.query_points(
) )
``` ```
**Production-ready scaling.** Rust + [SIMD-optimized inverted index](https://qdrant.tech/articles/sparse-vectors/) with an on-disk option keeps RAM low even with 200+ active terms per doc across millions of products. **Production-ready scaling.** Rust + [SIMD-optimized inverted index](https://github.com/qdrant/qdrant/blob/master/lib/sparse/src/index/posting_list.rs) with an on-disk option keeps RAM low even with 200+ active terms per doc across millions of products.
**No ANN approximation.** Sparse retrieval uses an [inverted index](https://qdrant.tech/articles/sparse-vectors/), the same data structure powering BM25. Results are exact - no recall tradeoffs from approximate nearest neighbor search. **No ANN approximation.** Sparse retrieval uses an [inverted index](https://qdrant.tech/articles/sparse-vectors/), the same data structure powering BM25. Results are exact - no recall tradeoffs from approximate nearest neighbor search.
@@ -238,7 +238,7 @@ A common concern: isn't running a transformer on every query slow?
| Sparse retrieval (Qdrant) | <1ms | Negligible | | Sparse retrieval (Qdrant) | <1ms | Negligible |
| **Total** | **10-20ms** | Real-time | | **Total** | **10-20ms** | Real-time |
The retrieval itself is negligible. Qdrant's Rust + [SIMD-optimized inverted index](https://qdrant.tech/articles/sparse-vectors/) scans millions of posting lists in sub-millisecond time. All the latency is in the encoder, which runs once per query regardless of catalog size. The retrieval itself is negligible. Qdrant's Rust + [SIMD-optimized inverted index](https://github.com/qdrant/qdrant/blob/master/lib/sparse/src/index/posting_list.rs) scans millions of posting lists in sub-millisecond time. All the latency is in the encoder, which runs once per query regardless of catalog size.
Optimization strategies if 15ms isn't fast enough: Optimization strategies if 15ms isn't fast enough:
- **Batch queries**: Encode multiple queries together (autocomplete, related searches) - **Batch queries**: Encode multiple queries together (autocomplete, related searches)