diff --git a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md index d7fd1fdae..bc9acdc11 100644 --- a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md +++ b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md @@ -132,7 +132,7 @@ client.query_points( ) ``` -**Production-ready scaling.** Rust + [SIMD-optimized inverted index](https://qdrant.tech/articles/sparse-vectors/) with an on-disk option keeps RAM low even with 200+ active terms per doc across millions of products. +**Production-ready scaling.** Rust + [SIMD-optimized inverted index](https://github.com/qdrant/qdrant/blob/master/lib/sparse/src/index/posting_list.rs) with an on-disk option keeps RAM low even with 200+ active terms per doc across millions of products. **No ANN approximation.** Sparse retrieval uses an [inverted index](https://qdrant.tech/articles/sparse-vectors/), the same data structure powering BM25. Results are exact - no recall tradeoffs from approximate nearest neighbor search. diff --git a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md index de0ef2875..3289c0fe5 100644 --- a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md +++ b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md @@ -238,7 +238,7 @@ A common concern: isn't running a transformer on every query slow? | Sparse retrieval (Qdrant) | <1ms | Negligible | | **Total** | **10-20ms** | Real-time | -The retrieval itself is negligible. Qdrant's Rust + [SIMD-optimized inverted index](https://qdrant.tech/articles/sparse-vectors/) scans millions of posting lists in sub-millisecond time. All the latency is in the encoder, which runs once per query regardless of catalog size. +The retrieval itself is negligible. Qdrant's Rust + [SIMD-optimized inverted index](https://github.com/qdrant/qdrant/blob/master/lib/sparse/src/index/posting_list.rs) scans millions of posting lists in sub-millisecond time. All the latency is in the encoder, which runs once per query regardless of catalog size. Optimization strategies if 15ms isn't fast enough: - **Batch queries**: Encode multiple queries together (autocomplete, related searches)