mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-06 19:38:30 +02:00
Add Parts 2-4 of sparse embeddings e-commerce series, simplify styling
This commit is contained in:
@@ -1,5 +1,5 @@
|
||||
---
|
||||
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search, Part 1: Why Sparse Embeddings Beat BM25"
|
||||
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 1: Why Sparse Embeddings Beat BM25"
|
||||
short_description: "Dense embeddings blur exact matches. Sparse embeddings keep the details that matter in e-commerce search."
|
||||
description: "Part 1 of a 4-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Learn why sparse embeddings outperform BM25 and dense models for product search, how SPLADE works, and why Qdrant's native sparse vector support matters."
|
||||
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-1/preview
|
||||
@@ -13,13 +13,19 @@ category: practicle-examples
|
||||
|
||||
*This is Part 1 of a 4-part series on fine-tuning sparse embeddings for e-commerce search. We'll go from "why bother?" to a production system that beats BM25 by 29%.*
|
||||
|
||||
**Series:**
|
||||
- Part 1: Why Sparse Embeddings Beat BM25 (here)
|
||||
- [Part 2: Training on Modal](/articles/sparse-embeddings-ecommerce-part-2/)
|
||||
- [Part 3: Evaluation & Hard Negatives](/articles/sparse-embeddings-ecommerce-part-3/)
|
||||
- [Part 4: Specialization vs Generalization](/articles/sparse-embeddings-ecommerce-part-4/)
|
||||
|
||||
---
|
||||
|
||||
Search "iPhone 15 Pro Max 256GB" on a dense embedding system and it happily returns the 128GB model. The semantic similarity is high - it's the same phone! But the customer specified 256GB for a reason. In e-commerce, the details aren't noise. They're the whole point.
|
||||
|
||||
This is the gap that sparse embeddings fill. And with fine-tuning, they fill it dramatically well - we achieved a **29% improvement over BM25** on Amazon's ESCI dataset, one of the largest public e-commerce search benchmarks.
|
||||
|
||||
In this series, we'll build the entire system: data loading, GPU training on Modal, evaluation with Qdrant, and hard negative mining. But first, let's understand why sparse embeddings are the right tool for this job.
|
||||
In this series, we'll build the entire system: data loading, GPU training on Modal, evaluation with Qdrant, and hard negative mining. The [full code is on GitHub](https://github.com/thierrypdamiba/finetune-ecommerce-search) and the [fine-tuned models are on HuggingFace](https://huggingface.co/thierrydamiba/splade-ecommerce-esci). But first, let's understand why sparse embeddings are the right tool for this job.
|
||||
|
||||
## The Problem with Dense Embeddings in E-Commerce
|
||||
|
||||
@@ -37,13 +43,15 @@ But this strength becomes a weakness in e-commerce:
|
||||
|
||||
Sparse embeddings take a fundamentally different approach. Instead of compressing text into a small, dense vector, they project it onto a large vocabulary space - typically 30,000+ dimensions (one per token in the vocabulary). But only 100-300 of those dimensions are non-zero.
|
||||
|
||||
| Dimension | Dense Embeddings | Sparse Embeddings |
|
||||
|-----------|-----------------|-------------------|
|
||||
| **Vector size** | 384-1024 dimensions | ~30,000 dimensions (vocabulary size) |
|
||||
| **Non-zero values** | All dimensions active | Only 100-300 terms active |
|
||||
| **Index type** | Approximate Nearest Neighbor (HNSW) | Inverted index |
|
||||
| | Dense | Sparse |
|
||||
|---|---|---|
|
||||
| **Vector size** | 384-1024 dims | ~30,000 dims |
|
||||
| **Non-zero values** | All active | 100-300 terms |
|
||||
| **Index type** | ANN (HNSW) | Inverted index |
|
||||
| **Exact matching** | Weak | Strong |
|
||||
| **Interpretability** | Black box | Transparent (see which terms matched) |
|
||||
| **Interpretability** | Black box | Transparent |
|
||||
|
||||
Both approaches encode text into vectors, but sparse embeddings preserve individual term signals that dense models compress away.
|
||||
|
||||
The key difference: each dimension in a sparse vector corresponds to an actual word in the vocabulary. You can inspect the vector and see exactly which terms the model considers important and how much weight it gives each one.
|
||||
|
||||
@@ -51,38 +59,21 @@ The key difference: each dimension in a sparse vector corresponds to an actual w
|
||||
|
||||
SPLADE (Sparse Lexical and Expansion) is the model architecture that makes this work. It passes text through a transformer with a masked language model (MLM) head, then applies max pooling and log saturation to produce sparse weights:
|
||||
|
||||
<div style="max-width: 640px; margin: 2rem auto; border-radius: 12px; overflow: hidden; font-family: 'JetBrains Mono', 'Fira Code', monospace; font-size: 14px; box-shadow: 0 4px 24px rgba(0,0,0,0.12);">
|
||||
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-bottom: 2px solid #dc3545;">SPLADE Encoding Pipeline</div>
|
||||
<div style="background: #16213e; padding: 20px; color: #c0c0d8; line-height: 2; text-align: center;">
|
||||
<div style="color: #e0e0f0;">Input: <code style="color: #addb67; background: rgba(173,219,103,0.1);">"noise canceling headphones"</code></div>
|
||||
<div style="color: #8890a8;">▼</div>
|
||||
<div style="color: #e0e0f0;">DistilBERT + MLM Head</div>
|
||||
<div style="color: #8890a8;">▼</div>
|
||||
<div style="color: #e0e0f0;">Max Pooling over tokens</div>
|
||||
<div style="color: #8890a8;">▼</div>
|
||||
<div style="color: #e0e0f0;">ReLU + Log Saturation: <code style="color: #7fdbca; background: rgba(127,219,202,0.1);">log(1 + ReLU(x))</code></div>
|
||||
<div style="color: #8890a8; font-size: 12px; font-style: italic;">A learned version of BM25's saturation curve</div>
|
||||
</div>
|
||||
<div style="background: #1a1a2e; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-top: 1px solid #2a2a4e; border-bottom: 2px solid #dc3545; color: #e0e0e0;">Output (~200 non-zero terms out of 30,522)</div>
|
||||
<div style="background: #16213e; padding: 0;">
|
||||
<table style="width: 100%; border-collapse: collapse; color: #e0e0e0; font-size: 14px;">
|
||||
<thead>
|
||||
<tr style="border-bottom: 1px solid #2a2a4e;">
|
||||
<th style="text-align: left; padding: 12px 20px; color: #8890a8; font-weight: 400;">Token</th>
|
||||
<th style="text-align: left; padding: 12px 20px; color: #8890a8; font-weight: 400;">Weight</th>
|
||||
</tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr style="border-bottom: 1px solid #2a2a4e;"><td style="padding: 10px 20px;">headphones</td><td style="padding: 10px 20px; color: #addb67;">2.3</td></tr>
|
||||
<tr style="border-bottom: 1px solid #2a2a4e;"><td style="padding: 10px 20px;">noise</td><td style="padding: 10px 20px; color: #addb67;">1.9</td></tr>
|
||||
<tr style="border-bottom: 1px solid #2a2a4e;"><td style="padding: 10px 20px;">canceling</td><td style="padding: 10px 20px; color: #addb67;">1.7</td></tr>
|
||||
<tr style="border-bottom: 1px solid #2a2a4e;"><td style="padding: 10px 20px;">audio</td><td style="padding: 10px 20px; color: #addb67;">1.2</td></tr>
|
||||
<tr style="border-bottom: 1px solid #2a2a4e;"><td style="padding: 10px 20px;">wireless</td><td style="padding: 10px 20px; color: #addb67;">0.8</td></tr>
|
||||
<tr><td style="padding: 10px 20px;">sound</td><td style="padding: 10px 20px; color: #addb67;">0.6</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
</div>
|
||||
For an input like `"noise canceling headphones"`, SPLADE encodes it in four steps:
|
||||
|
||||
1. **Tokenize and encode** the input through DistilBERT with a masked language model (MLM) head
|
||||
2. **Max pool** across all token positions to get a single score per vocabulary term
|
||||
3. **Apply log saturation** — `log(1 + ReLU(x))` — a learned version of BM25's saturation curve that prevents any single term from dominating
|
||||
4. **Output a sparse vector** with ~200 non-zero values out of 30,522 vocabulary dimensions
|
||||
|
||||
| Token | Weight |
|
||||
|---|---|
|
||||
| headphones | 2.3 |
|
||||
| noise | 1.9 |
|
||||
| canceling | 1.7 |
|
||||
| audio | 1.2 |
|
||||
| wireless | 0.8 |
|
||||
| sound | 0.6 |
|
||||
|
||||
The **log saturation** step is important. Without it, a single high-confidence term could dominate the score. The log compression keeps results balanced - "headphones" matters more than "audio", but not 10x more.
|
||||
|
||||
@@ -96,20 +87,11 @@ The model learns three things simultaneously:
|
||||
|
||||
This expansion is what separates SPLADE from traditional keyword search. BM25 can only match terms that literally appear in both the query and the document. SPLADE adds related terms that the model learned from training data:
|
||||
|
||||
<div style="max-width: 640px; margin: 2rem auto; font-family: -apple-system, sans-serif; text-align: center;">
|
||||
<div style="color: #888; font-size: 13px; margin-bottom: 12px; text-transform: uppercase; letter-spacing: 1px;">Query: "summer dress"</div>
|
||||
<div style="display: flex; flex-wrap: wrap; gap: 8px; margin-bottom: 16px; justify-content: center;">
|
||||
<span style="background: #1a4d2e; color: #4ade80; padding: 6px 14px; border-radius: 20px; font-size: 15px; font-weight: 500;">dress <span style="opacity: 0.6; font-size: 12px;">2.5</span></span>
|
||||
<span style="background: #1a4d2e; color: #4ade80; padding: 6px 14px; border-radius: 20px; font-size: 15px; font-weight: 500;">summer <span style="opacity: 0.6; font-size: 12px;">2.1</span></span>
|
||||
</div>
|
||||
<div style="color: #7fdbca; font-size: 11px; text-transform: uppercase; letter-spacing: 1px; margin-bottom: 12px;">+ expanded by SPLADE</div>
|
||||
<div style="display: flex; flex-wrap: wrap; gap: 8px; justify-content: center;">
|
||||
<span style="background: #1a2e4d; color: #7fdbca; padding: 6px 14px; border-radius: 20px; font-size: 14px;">sundress <span style="opacity: 0.6; font-size: 12px;">1.8</span></span>
|
||||
<span style="background: #1a2e4d; color: #7fdbca; padding: 6px 14px; border-radius: 20px; font-size: 14px;">floral <span style="opacity: 0.6; font-size: 12px;">0.9</span></span>
|
||||
<span style="background: #1a2e4d; color: #7fdbca; padding: 6px 14px; border-radius: 20px; font-size: 14px;">lightweight <span style="opacity: 0.6; font-size: 12px;">0.7</span></span>
|
||||
<span style="background: #1a2e4d; color: #7fdbca; padding: 6px 14px; border-radius: 20px; font-size: 14px;">cotton <span style="opacity: 0.6; font-size: 12px;">0.6</span></span>
|
||||
</div>
|
||||
</div>
|
||||
**Query: "summer dress"**
|
||||
|
||||
Original terms: `dress` (2.5), `summer` (2.1)
|
||||
|
||||
Expanded by SPLADE: `sundress` (1.8), `floral` (0.9), `lightweight` (0.7), `cotton` (0.6)
|
||||
|
||||
The model adds "sundress", "floral", and "cotton" - terms that appear in product titles even when "summer" doesn't. This matches products like *"Floral Sundress for Women - Lightweight Cotton"* that BM25 would miss entirely.
|
||||
|
||||
@@ -151,26 +133,20 @@ client.query_points(
|
||||
|
||||
Our training pipeline combines three components:
|
||||
|
||||
<div style="max-width: 640px; margin: 2rem auto; border-radius: 12px; overflow: hidden; font-family: 'JetBrains Mono', 'Fira Code', monospace; font-size: 14px; box-shadow: 0 4px 24px rgba(0,0,0,0.12);">
|
||||
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-bottom: 2px solid #dc3545;">Modal (GPU Training)</div>
|
||||
<div style="background: #16213e; padding: 16px 20px; line-height: 1.8; color: #c0c0d8;">
|
||||
<span style="color: #addb67;">•</span> A100 GPUs on demand<br>
|
||||
<span style="color: #addb67;">•</span> Persistent volumes for checkpoints<br>
|
||||
<span style="color: #addb67;">•</span> Detached runs for long training
|
||||
</div>
|
||||
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-top: 1px solid #2a2a4e; border-bottom: 2px solid #dc3545;">Sentence Transformers v5</div>
|
||||
<div style="background: #16213e; padding: 16px 20px; line-height: 1.8; color: #c0c0d8;">
|
||||
<span style="color: #addb67;">•</span> SparseEncoder architecture<br>
|
||||
<span style="color: #addb67;">•</span> SpladeLoss with regularization<br>
|
||||
<span style="color: #addb67;">•</span> Built-in training utilities
|
||||
</div>
|
||||
<div style="background: #1a1a2e; color: #e0e0e0; padding: 12px 20px; text-align: center; font-size: 13px; letter-spacing: 1px; text-transform: uppercase; border-top: 1px solid #2a2a4e; border-bottom: 2px solid #dc3545;">Qdrant (Sparse Vector Store)</div>
|
||||
<div style="background: #16213e; padding: 16px 20px; line-height: 1.8; color: #c0c0d8;">
|
||||
<span style="color: #addb67;">•</span> Native sparse vector support<br>
|
||||
<span style="color: #addb67;">•</span> Inverted index<br>
|
||||
<span style="color: #addb67;">•</span> Hybrid search ready
|
||||
</div>
|
||||
</div>
|
||||
**[Modal](https://modal.com/)** (GPU Training)
|
||||
- A100 GPUs on demand
|
||||
- Persistent volumes for checkpoints
|
||||
- Detached runs for long training
|
||||
|
||||
**[Sentence Transformers v5](https://www.sbert.net/)** (Training Framework)
|
||||
- SparseEncoder architecture
|
||||
- SpladeLoss with regularization
|
||||
- Built-in training utilities
|
||||
|
||||
**[Qdrant](https://qdrant.tech/)** (Sparse Vector Store)
|
||||
- Native sparse vector support
|
||||
- Inverted index
|
||||
- Hybrid search ready
|
||||
|
||||
Modal gives us serverless A100 GPUs - no idle hardware, no queue management. Sentence Transformers v5 introduced the `SparseEncoder` class that makes SPLADE training straightforward. And Qdrant handles storage, indexing, and retrieval with native sparse vector support.
|
||||
|
||||
@@ -178,12 +154,11 @@ Modal gives us serverless A100 GPUs - no idle hardware, no queue management. Sen
|
||||
|
||||
Over the next three articles, we'll walk through the full pipeline:
|
||||
|
||||
- <span style="color: #888;">**Part 2: Training on Modal** (coming soon)</span> - Loading the Amazon ESCI dataset, creating the SPLADE model, configuring loss functions with sparsity regularization, and running GPU training with persistent checkpoints.
|
||||
- [**Part 2: Training on Modal**](/articles/sparse-embeddings-ecommerce-part-2/) - Loading the Amazon ESCI dataset, creating the SPLADE model, configuring loss functions with sparsity regularization, and running GPU training with persistent checkpoints.
|
||||
|
||||
- <span style="color: #888;">**Part 3: Evaluation and Hard Negative Mining** (coming soon)</span> - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE hard negative mining loops, and analyzing what fine-tuning actually changes in the model.
|
||||
- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE hard negative mining loops, and analyzing what fine-tuning actually changes in the model.
|
||||
|
||||
- <span style="color: #888;">**Part 4: Specialization vs Generalization** (coming soon)</span> - Cross-domain evaluation on Wayfair and Home Depot data, multi-domain training, when to specialize vs generalize, and production deployment guidance.
|
||||
- [**Part 4: Specialization vs Generalization**](/articles/sparse-embeddings-ecommerce-part-4/) - Cross-domain evaluation on Wayfair and Home Depot data, multi-domain training, when to specialize vs generalize, and production deployment guidance.
|
||||
|
||||
The end result: a fine-tuned SPLADE model that achieves **nDCG@10 of 0.388** on Amazon ESCI, compared to **0.301** for BM25 and **0.324** for off-the-shelf SPLADE. That 29% improvement over BM25 translates to meaningfully better search results for real e-commerce queries.
|
||||
The end result: a fine-tuned SPLADE model that achieves **nDCG@10 of 0.388** on Amazon ESCI, compared to **0.301** for BM25 and **0.324** for off-the-shelf SPLADE. That 29% improvement over BM25 translates to meaningfully better search results for real e-commerce queries. You can try the models directly from HuggingFace: [splade-ecommerce-esci](https://huggingface.co/thierrydamiba/splade-ecommerce-esci) (best in-domain) and [splade-ecommerce-multidomain](https://huggingface.co/thierrydamiba/splade-ecommerce-multidomain) (better generalization).
|
||||
|
||||
Stay tuned for Part 2, where we'll dive into the training pipeline on Modal.
|
||||
|
||||
Reference in New Issue
Block a user