diff --git a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md index b0b2c5768..d7fd1fdae 100644 --- a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md +++ b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-1.md @@ -163,7 +163,7 @@ Over the next four articles, we'll walk through the full pipeline: - [**Part 2: Training on Modal**](/articles/sparse-embeddings-ecommerce-part-2/) - Loading the Amazon ESCI dataset, creating the SPLADE model, configuring loss functions with sparsity regularization, and running GPU training with persistent checkpoints. -- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE hard negative mining loops, and analyzing what fine-tuning actually changes in the model. +- [**Part 3: Evaluation and Hard Negative Mining**](/articles/sparse-embeddings-ecommerce-part-3/) - Indexing products in Qdrant, running retrieval benchmarks (nDCG, MRR, Recall), implementing ANCE-inspired hard negative mining loops, and analyzing what fine-tuning actually changes in the model. - [**Part 4: Specialization vs Generalization**](/articles/sparse-embeddings-ecommerce-part-4/) - Cross-domain evaluation on Wayfair and Home Depot data, multi-domain training, when to specialize vs generalize, and production deployment guidance. diff --git a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md index e506ed2f2..de0ef2875 100644 --- a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md +++ b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-3.md @@ -1,7 +1,7 @@ --- title: "Fine-Tuning Sparse Embeddings for E-Commerce Search | Part 3: Evaluation and Hard Negatives" short_description: "Evaluate fine-tuned SPLADE with Qdrant and boost results with hard negative mining." -description: "Part 3 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Index products in Qdrant, run retrieval benchmarks, and implement ANCE hard negative mining for a 28% improvement over BM25." +description: "Part 3 of a 5-part series on fine-tuning SPLADE sparse embeddings for e-commerce search. Index products in Qdrant, run retrieval benchmarks, and implement ANCE-inspired hard negative mining for a 28% improvement over BM25." preview_dir: /articles_data/sparse-embeddings-ecommerce-part-3/preview social_preview_image: /articles_data/sparse-embeddings-ecommerce-part-3/preview/social_preview.jpg weight: -198 @@ -157,13 +157,13 @@ With the **fine-tuned SPLADE**, hybrid actually hurts: SPLADE-only scored 0.413 This is a useful finding. Hybrid search isn't always better. It depends on the relative strength of your signals. If your sparse model is domain-tuned and your dense model is generic, the dense component can actively harm results. -## Hard Negative Mining with ANCE +## ANCE-inspired Hard Negative Mining ![The ANCE hard negative mining loop](/articles_data/sparse-embeddings-ecommerce-part-3/ance-loop.png) The training in Part 2 used in-batch negatives: other products in the same batch serve as negatives for a given query. This works but has a limitation: random products are easy negatives. The model doesn't learn to distinguish between genuinely confusable products. -[ANCE](https://www.sbert.net/examples/training/quora_duplicate_questions/README.html) (Approximate Nearest Neighbor Negative Contrastive Estimation) fixes this by mining hard negatives from the current model's own retrieval results: +Inspired by [ANCE](https://arxiv.org/abs/2007.00808), this approach mines hard negatives from the current model's own retrieval results: 1. **Index** products into Qdrant with the current model 2. **Retrieve** top-K products for each query @@ -199,9 +199,9 @@ hard_neg_examples = miner.mine_for_training( Sparse retrieval keeps mining cheap, with sub-millisecond per query in Qdrant. For 100K queries, the mining step takes seconds, not minutes. Payload filters exclude known positives so you don't accidentally treat a relevant product as a negative. -### When to Use ANCE +### When to Use ANCE-inspired Mining -ANCE adds complexity. You need to: +This approach adds complexity. You need to: 1. Index products with the current model 2. Run retrieval for all training queries 3. Filter and format the results @@ -252,7 +252,7 @@ For most e-commerce applications, 15ms is fine, especially when it delivers 28% - **Fine-tuned SPLADE beats BM25 by 28% and off-the-shelf SPLADE by 19%.** Domain-specific training matters, even for sparse models. - **Hybrid search isn't always better.** A strong domain-tuned sparse model can outperform sparse+dense fusion when the dense component is generic. -- **Hard negative mining (ANCE) adds 5-10%** on top of basic training. Qdrant's sparse retrieval makes the mining step cheap. +- **Hard negative mining (ANCE-inspired) adds 5-10%** on top of basic training. Qdrant's sparse retrieval makes the mining step cheap. - **Production latency is 10-20ms total.** Transformer encoding is the bottleneck, not retrieval. - **The model learns domain-specific patterns**: query expansion, term weighting, and e-commerce vocabulary all improve with fine-tuning. diff --git a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md index 234f38a44..d728ab87a 100644 --- a/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md +++ b/qdrant-landing/content/articles/sparse-embeddings-ecommerce-part-5.md @@ -22,11 +22,11 @@ category: practicle-examples --- -In Parts 1 through 4, we built a SPLADE fine-tuning pipeline piece by piece: data loading, Modal GPU training, Qdrant evaluation, ANCE hard negative mining, cross-domain experiments. The code worked. The results were strong: 28% over BM25 on Amazon ESCI. +In Parts 1 through 4, we built a SPLADE fine-tuning pipeline piece by piece: data loading, Modal GPU training, Qdrant evaluation, ANCE-inspired hard negative mining, cross-domain experiments. The code worked. The results were strong: 28% over BM25 on Amazon ESCI. Using it required reading four articles, cloning a repo, understanding the training loop internals, wiring up Modal volumes, and configuring Qdrant connections manually. That's fine for a series walkthrough. It's not fine for someone who has a product catalog and wants a better search model by end of day. -So we packaged everything into [`qdrant-sparse-finetune`](https://github.com/qdrant/sparse-finetune): an open-source CLI and web dashboard that runs the entire pipeline (synthetic query generation, SPLADE training with ANCE, evaluation, and HuggingFace publishing) with a single command. +So we packaged everything into [`qdrant-sparse-finetune`](https://github.com/qdrant/sparse-finetune): an open-source CLI and web dashboard that runs the entire pipeline (synthetic query generation, SPLADE training with ANCE-inspired hard negative mining, evaluation, and HuggingFace publishing) with a single command. ## The Problem We're Solving @@ -38,7 +38,7 @@ The [series repo](https://github.com/thierrypdamiba/finetune-ecommerce-search) i 2. Either provide labeled queries or set up an LLM API for synthetic generation 3. Configure Modal volumes and GPU settings 4. Wire up Qdrant credentials for indexing and mining -5. Run training, manually trigger ANCE iterations +5. Run training, manually trigger hard negative mining iterations 6. Evaluate, interpret metrics 7. Publish to HuggingFace if you want to share the model @@ -96,7 +96,7 @@ qdrant-finetune studio This launches a web dashboard with tabs for each stage of the pipeline: -**Train.** Configure the base model, ANCE iterations, batch size, and GPU backend. Submit a job and watch live logs with a loss chart that updates as training progresses. +**Train.** Configure the base model, hard negative mining iterations, batch size, and GPU backend. Submit a job and watch live logs with a loss chart that updates as training progresses. **Evaluate.** Point at a trained model and test queries. Get metric cards for nDCG@10, MRR@10, Recall, and Precision. @@ -151,7 +151,7 @@ That single line: 1. Loads your product data 2. Generates synthetic queries via LLM 3. Creates a SPLADE encoder -4. Runs 3 rounds of ANCE hard negative mining against Qdrant +4. Runs 3 rounds of ANCE-inspired hard negative mining against Qdrant 5. Saves the fine-tuned model For more control: @@ -174,7 +174,7 @@ metrics = trainer.evaluate(queries="test_queries.csv") print(metrics) # {'ndcg@10': 0.389, 'mrr@10': 0.387, ...} ``` -Same ANCE loop from Part 3, same evaluation metrics. The `Trainer` class wraps the training loop, Qdrant indexing, hard negative mining, and evaluation into a coherent API. +Same ANCE-inspired loop from Part 3, same evaluation metrics. The `Trainer` class wraps the training loop, Qdrant indexing, hard negative mining, and evaluation into a coherent API. ## Getting Started @@ -217,7 +217,7 @@ The 28% improvement over BM25 from Part 3 isn't locked behind a research repo an - **[Part 1: Why sparse embeddings for e-commerce](/articles/sparse-embeddings-ecommerce-part-1/)**: SPLADE combines keyword precision with learned expansion - **[Part 2: Training pipeline on Modal](/articles/sparse-embeddings-ecommerce-part-2/)**: A100 training with persistent checkpoints -- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)**: +28% vs BM25, ANCE mining with Qdrant +- **[Part 3: Evaluation and hard negatives](/articles/sparse-embeddings-ecommerce-part-3/)**: +28% vs BM25, ANCE-inspired mining with Qdrant - **[Part 4: Specialization vs generalization](/articles/sparse-embeddings-ecommerce-part-4/)**: Domain-specific vs multi-domain tradeoffs - **[Part 5: From research to product](/articles/sparse-embeddings-ecommerce-part-5/)**: CLI + dashboard that runs the full pipeline