Fix ANCE references across SPLADE series

Replace "ANCE" with "ANCE-inspired" throughout since the implementation
uses Qdrant's exact inverted index, not ANN. Reference the original
ANCE paper (Xiong et al., arxiv 2007.00808) in Part 3.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Evgeniya Sukhodolskaya
2026-06-04 11:42:36 +02:00
co-authored by Claude Sonnet 4.6
parent f4d989b1cb
commit 9ab2af0542
3 changed files with 14 additions and 14 deletions
@@ -163,7 +163,7 @@ Over the next four articles, we'll walk through the full pipeline:
- [**Part 2: Training on Modal**](/articles/sparse-embeddings-ecommerce-part-2/) - Loading the Amazon ESCI dataset, creating the SPLADE model, configuring loss functions with sparsity regularization, and running GPU training with persistent checkpoints.
- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE hard negative mining loops, and analyzing what fine-tuning actually changes in the model.
- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE-inspired hard negative mining loops, and analyzing what fine-tuning actually changes in the model.
- [**Part 4: Specialization vs Generalization**](/articles/sparse-embeddings-ecommerce-part-4/) - Cross-domain evaluation on Wayfair and Home Depot data, multi-domain training, when to specialize vs generalize, and production deployment guidance.
@@ -1,7 +1,7 @@
---
title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 3: Evaluation and Hard Negatives"
short_description: "Evaluate fine-tuned SPLADE with Qdrant and boost results with hard negative mining."
description: "Part 3 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Index products in Qdrant, run retrieval benchmarks, and implement ANCE hard negative mining for a 28% improvement over BM25."
description: "Part 3 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Index products in Qdrant, run retrieval benchmarks, and implement ANCE-inspired hard negative mining for a 28% improvement over BM25."
preview_dir: /articles_data/sparse-embeddings-ecommerce-part-3/preview
social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-3/preview/social_preview.jpg
weight: -198
@@ -157,13 +157,13 @@ With the **fine-tuned SPLADE**, hybrid actually hurts: SPLADE-only scored 0.413
This is a useful finding. Hybrid search isn't always better. It depends on the relative strength of your signals. If your sparse model is domain-tuned and your dense model is generic, the dense component can actively harm results.
## Hard Negative Mining with ANCE
## ANCE-inspired Hard Negative Mining
![The ANCE hard negative mining loop](/articles_data/sparse-embeddings-ecommerce-part-3/ance-loop.png)
The training in Part 2 used in-batch negatives: other products in the same batch serve as negatives for a given query. This works but has a limitation: random products are easy negatives. The model doesn't learn to distinguish between genuinely confusable products.
[ANCE](https://www.sbert.net/examples/training/quora_duplicate_questions/README.html) (Approximate Nearest Neighbor Negative Contrastive Estimation) fixes this by mining hard negatives from the current model's own retrieval results:
Inspired by [ANCE](https://arxiv.org/abs/2007.00808), this approach mines hard negatives from the current model's own retrieval results:
1. **Index** products into Qdrant with the current model
2. **Retrieve** top-K products for each query
@@ -199,9 +199,9 @@ hard_neg_examples = miner.mine_for_training(
Sparse retrieval keeps mining cheap, with sub-millisecond per query in Qdrant. For 100K queries, the mining step takes seconds, not minutes. Payload filters exclude known positives so you don't accidentally treat a relevant product as a negative.
### When to Use ANCE
### When to Use ANCE-inspired Mining
ANCE adds complexity. You need to:
This approach adds complexity. You need to:
1. Index products with the current model
2. Run retrieval for all training queries
3. Filter and format the results
@@ -252,7 +252,7 @@ For most e-commerce applications, 15ms is fine, especially when it delivers 28%
- **Fine-tuned SPLADE beats BM25 by 28% and off-the-shelf SPLADE by 19%.** Domain-specific training matters, even for sparse models.
- **Hybrid search isn't always better.** A strong domain-tuned sparse model can outperform sparse+dense fusion when the dense component is generic.
- **Hard negative mining (ANCE) adds 5-10%** on top of basic training. Qdrant's sparse retrieval makes the mining step cheap.
- **Hard negative mining (ANCE-inspired) adds 5-10%** on top of basic training. Qdrant's sparse retrieval makes the mining step cheap.
- **Production latency is 10-20ms total.** Transformer encoding is the bottleneck, not retrieval.
- **The model learns domain-specific patterns**: query expansion, term weighting, and e-commerce vocabulary all improve with fine-tuning.
@@ -22,11 +22,11 @@ category: practicle-examples
---
In Parts 1 through 4, we built a SPLADE fine-tuning pipeline piece by piece: data loading, Modal GPU training, Qdrant evaluation, ANCE hard negative mining, cross-domain experiments. The code worked. The results were strong: 28% over BM25 on Amazon ESCI.
In Parts 1 through 4, we built a SPLADE fine-tuning pipeline piece by piece: data loading, Modal GPU training, Qdrant evaluation, ANCE-inspired hard negative mining, cross-domain experiments. The code worked. The results were strong: 28% over BM25 on Amazon ESCI.
Using it required reading four articles, cloning a repo, understanding the training loop internals, wiring up Modal volumes, and configuring Qdrant connections manually. That's fine for a series walkthrough. It's not fine for someone who has a product catalog and wants a better search model by end of day.
So we packaged everything into [`qdrant-sparse-finetune`](https://github.com/qdrant/sparse-finetune): an open-source CLI and web dashboard that runs the entire pipeline (synthetic query generation, SPLADE training with ANCE, evaluation, and HuggingFace publishing) with a single command.
So we packaged everything into [`qdrant-sparse-finetune`](https://github.com/qdrant/sparse-finetune): an open-source CLI and web dashboard that runs the entire pipeline (synthetic query generation, SPLADE training with ANCE-inspired hard negative mining, evaluation, and HuggingFace publishing) with a single command.
## The Problem We're Solving
@@ -38,7 +38,7 @@ The [series repo](https://github.com/thierrypdamiba/finetune-ecommerce-search) i
2. Either provide labeled queries or set up an LLM API for synthetic generation
3. Configure Modal volumes and GPU settings
4. Wire up Qdrant credentials for indexing and mining
5. Run training, manually trigger ANCE iterations
5. Run training, manually trigger hard negative mining iterations
6. Evaluate, interpret metrics
7. Publish to HuggingFace if you want to share the model
@@ -96,7 +96,7 @@ qdrant-finetune studio
This launches a web dashboard with tabs for each stage of the pipeline:
**Train.** Configure the base model, ANCE iterations, batch size, and GPU backend. Submit a job and watch live logs with a loss chart that updates as training progresses.
**Train.** Configure the base model, hard negative mining iterations, batch size, and GPU backend. Submit a job and watch live logs with a loss chart that updates as training progresses.
**Evaluate.** Point at a trained model and test queries. Get metric cards for nDCG@10, MRR@10, Recall, and Precision.
@@ -151,7 +151,7 @@ That single line:
1. Loads your product data
2. Generates synthetic queries via LLM
3. Creates a SPLADE encoder
4. Runs 3 rounds of ANCE hard negative mining against Qdrant
4. Runs 3 rounds of ANCE-inspired hard negative mining against Qdrant
5. Saves the fine-tuned model
For more control:
@@ -174,7 +174,7 @@ metrics = trainer.evaluate(queries="test_queries.csv")
print(metrics) # {'ndcg@10': 0.389, 'mrr@10': 0.387, ...}
```
Same ANCE loop from Part 3, same evaluation metrics. The `Trainer` class wraps the training loop, Qdrant indexing, hard negative mining, and evaluation into a coherent API.
Same ANCE-inspired loop from Part 3, same evaluation metrics. The `Trainer` class wraps the training loop, Qdrant indexing, hard negative mining, and evaluation into a coherent API.
## Getting Started
@@ -217,7 +217,7 @@ The 28% improvement over BM25 from Part 3 isn't locked behind a research repo an
- **[Part 1: Why sparse embeddings for e-commerce](/articles/sparse-embeddings-ecommerce-part-1/)**: SPLADE combines keyword precision with learned expansion
- **[Part 2: Training pipeline on Modal](/articles/sparse-embeddings-ecommerce-part-2/)**: A100 training with persistent checkpoints
- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)**: +28% vs BM25, ANCE mining with Qdrant
- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)**: +28% vs BM25, ANCE-inspired mining with Qdrant
- **[Part 4: Specialization vs generalization](/articles/sparse-embeddings-ecommerce-part-4/)**: Domain-specific vs multi-domain tradeoffs
- **[Part 5: From research to product](/articles/sparse-embeddings-ecommerce-part-5/)**: CLI + dashboard that runs the full pipeline