mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-03 01:48:32 +02:00
Fix ANCE references across SPLADE series
Replace "ANCE" with "ANCE-inspired" throughout since the implementation uses Qdrant's exact inverted index, not ANN. Reference the original ANCE paper (Xiong et al., arxiv 2007.00808) in Part 3. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 4.6
parent
f4d989b1cb
commit
9ab2af0542
@@ -163,7 +163,7 @@ Over the next four articles, we'll walk through the full pipeline:
|
||||
|
||||
- [**Part 2: Training on Modal**](/articles/sparse-embeddings-ecommerce-part-2/) - Loading the Amazon ESCI dataset, creating the SPLADE model, configuring loss functions with sparsity regularization, and running GPU training with persistent checkpoints.
|
||||
|
||||
- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE hard negative mining loops, and analyzing what fine-tuning actually changes in the model.
|
||||
- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE-inspired hard negative mining loops, and analyzing what fine-tuning actually changes in the model.
|
||||
|
||||
- [**Part 4: Specialization vs Generalization**](/articles/sparse-embeddings-ecommerce-part-4/) - Cross-domain evaluation on Wayfair and Home Depot data, multi-domain training, when to specialize vs generalize, and production deployment guidance.
|
||||
|
||||
|
||||
@@ -1,7 +1,7 @@
|
||||
---
|
||||
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 3: Evaluation and Hard Negatives"
|
||||
short_description: "Evaluate fine-tuned SPLADE with Qdrant and boost results with hard negative mining."
|
||||
description: "Part 3 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Index products in Qdrant, run retrieval benchmarks, and implement ANCE hard negative mining for a 28% improvement over BM25."
|
||||
description: "Part 3 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Index products in Qdrant, run retrieval benchmarks, and implement ANCE-inspired hard negative mining for a 28% improvement over BM25."
|
||||
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-3/preview
|
||||
social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-3/preview/social_preview.jpg
|
||||
weight: -198
|
||||
@@ -157,13 +157,13 @@ With the **fine-tuned SPLADE**, hybrid actually hurts: SPLADE-only scored 0.413
|
||||
|
||||
This is a useful finding. Hybrid search isn't always better. It depends on the relative strength of your signals. If your sparse model is domain-tuned and your dense model is generic, the dense component can actively harm results.
|
||||
|
||||
## Hard Negative Mining with ANCE
|
||||
## ANCE-inspired Hard Negative Mining
|
||||
|
||||

|
||||
|
||||
The training in Part 2 used in-batch negatives: other products in the same batch serve as negatives for a given query. This works but has a limitation: random products are easy negatives. The model doesn't learn to distinguish between genuinely confusable products.
|
||||
|
||||
[ANCE](https://www.sbert.net/examples/training/quora_duplicate_questions/README.html) (Approximate Nearest Neighbor Negative Contrastive Estimation) fixes this by mining hard negatives from the current model's own retrieval results:
|
||||
Inspired by [ANCE](https://arxiv.org/abs/2007.00808), this approach mines hard negatives from the current model's own retrieval results:
|
||||
|
||||
1. **Index** products into Qdrant with the current model
|
||||
2. **Retrieve** top-K products for each query
|
||||
@@ -199,9 +199,9 @@ hard_neg_examples = miner.mine_for_training(
|
||||
|
||||
Sparse retrieval keeps mining cheap, with sub-millisecond per query in Qdrant. For 100K queries, the mining step takes seconds, not minutes. Payload filters exclude known positives so you don't accidentally treat a relevant product as a negative.
|
||||
|
||||
### When to Use ANCE
|
||||
### When to Use ANCE-inspired Mining
|
||||
|
||||
ANCE adds complexity. You need to:
|
||||
This approach adds complexity. You need to:
|
||||
1. Index products with the current model
|
||||
2. Run retrieval for all training queries
|
||||
3. Filter and format the results
|
||||
@@ -252,7 +252,7 @@ For most e-commerce applications, 15ms is fine, especially when it delivers 28%
|
||||
|
||||
- **Fine-tuned SPLADE beats BM25 by 28% and off-the-shelf SPLADE by 19%.** Domain-specific training matters, even for sparse models.
|
||||
- **Hybrid search isn't always better.** A strong domain-tuned sparse model can outperform sparse+dense fusion when the dense component is generic.
|
||||
- **Hard negative mining (ANCE) adds 5-10%** on top of basic training. Qdrant's sparse retrieval makes the mining step cheap.
|
||||
- **Hard negative mining (ANCE-inspired) adds 5-10%** on top of basic training. Qdrant's sparse retrieval makes the mining step cheap.
|
||||
- **Production latency is 10-20ms total.** Transformer encoding is the bottleneck, not retrieval.
|
||||
- **The model learns domain-specific patterns**: query expansion, term weighting, and e-commerce vocabulary all improve with fine-tuning.
|
||||
|
||||
|
||||
@@ -22,11 +22,11 @@ category: practicle-examples
|
||||
|
||||
---
|
||||
|
||||
In Parts 1 through 4, we built a SPLADE fine-tuning pipeline piece by piece: data loading, Modal GPU training, Qdrant evaluation, ANCE hard negative mining, cross-domain experiments. The code worked. The results were strong: 28% over BM25 on Amazon ESCI.
|
||||
In Parts 1 through 4, we built a SPLADE fine-tuning pipeline piece by piece: data loading, Modal GPU training, Qdrant evaluation, ANCE-inspired hard negative mining, cross-domain experiments. The code worked. The results were strong: 28% over BM25 on Amazon ESCI.
|
||||
|
||||
Using it required reading four articles, cloning a repo, understanding the training loop internals, wiring up Modal volumes, and configuring Qdrant connections manually. That's fine for a series walkthrough. It's not fine for someone who has a product catalog and wants a better search model by end of day.
|
||||
|
||||
So we packaged everything into [`qdrant-sparse-finetune`](https://github.com/qdrant/sparse-finetune): an open-source CLI and web dashboard that runs the entire pipeline (synthetic query generation, SPLADE training with ANCE, evaluation, and HuggingFace publishing) with a single command.
|
||||
So we packaged everything into [`qdrant-sparse-finetune`](https://github.com/qdrant/sparse-finetune): an open-source CLI and web dashboard that runs the entire pipeline (synthetic query generation, SPLADE training with ANCE-inspired hard negative mining, evaluation, and HuggingFace publishing) with a single command.
|
||||
|
||||
## The Problem We're Solving
|
||||
|
||||
@@ -38,7 +38,7 @@ The [series repo](https://github.com/thierrypdamiba/finetune-ecommerce-search) i
|
||||
2. Either provide labeled queries or set up an LLM API for synthetic generation
|
||||
3. Configure Modal volumes and GPU settings
|
||||
4. Wire up Qdrant credentials for indexing and mining
|
||||
5. Run training, manually trigger ANCE iterations
|
||||
5. Run training, manually trigger hard negative mining iterations
|
||||
6. Evaluate, interpret metrics
|
||||
7. Publish to HuggingFace if you want to share the model
|
||||
|
||||
@@ -96,7 +96,7 @@ qdrant-finetune studio
|
||||
|
||||
This launches a web dashboard with tabs for each stage of the pipeline:
|
||||
|
||||
**Train.** Configure the base model, ANCE iterations, batch size, and GPU backend. Submit a job and watch live logs with a loss chart that updates as training progresses.
|
||||
**Train.** Configure the base model, hard negative mining iterations, batch size, and GPU backend. Submit a job and watch live logs with a loss chart that updates as training progresses.
|
||||
|
||||
**Evaluate.** Point at a trained model and test queries. Get metric cards for nDCG@10, MRR@10, Recall, and Precision.
|
||||
|
||||
@@ -151,7 +151,7 @@ That single line:
|
||||
1. Loads your product data
|
||||
2. Generates synthetic queries via LLM
|
||||
3. Creates a SPLADE encoder
|
||||
4. Runs 3 rounds of ANCE hard negative mining against Qdrant
|
||||
4. Runs 3 rounds of ANCE-inspired hard negative mining against Qdrant
|
||||
5. Saves the fine-tuned model
|
||||
|
||||
For more control:
|
||||
@@ -174,7 +174,7 @@ metrics = trainer.evaluate(queries="test_queries.csv")
|
||||
print(metrics) # {'ndcg@10': 0.389, 'mrr@10': 0.387, ...}
|
||||
```
|
||||
|
||||
Same ANCE loop from Part 3, same evaluation metrics. The `Trainer` class wraps the training loop, Qdrant indexing, hard negative mining, and evaluation into a coherent API.
|
||||
Same ANCE-inspired loop from Part 3, same evaluation metrics. The `Trainer` class wraps the training loop, Qdrant indexing, hard negative mining, and evaluation into a coherent API.
|
||||
|
||||
## Getting Started
|
||||
|
||||
@@ -217,7 +217,7 @@ The 28% improvement over BM25 from Part 3 isn't locked behind a research repo an
|
||||
|
||||
- **[Part 1: Why sparse embeddings for e-commerce](/articles/sparse-embeddings-ecommerce-part-1/)**: SPLADE combines keyword precision with learned expansion
|
||||
- **[Part 2: Training pipeline on Modal](/articles/sparse-embeddings-ecommerce-part-2/)**: A100 training with persistent checkpoints
|
||||
- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)**: +28% vs BM25, ANCE mining with Qdrant
|
||||
- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)**: +28% vs BM25, ANCE-inspired mining with Qdrant
|
||||
- **[Part 4: Specialization vs generalization](/articles/sparse-embeddings-ecommerce-part-4/)**: Domain-specific vs multi-domain tradeoffs
|
||||
- **[Part 5: From research to product](/articles/sparse-embeddings-ecommerce-part-5/)**: CLI + dashboard that runs the full pipeline
|
||||
|
||||
|
||||
Reference in New Issue
Block a user