Add metric comparability notes to SPLADE series parts 1-4

Metrics were measured on a subsample (100k products, 10k queries) with
all relevant documents included, so they are not comparable to official
Amazon ESCI benchmarks. Add inline notes at each metric table/result.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Evgeniya Sukhodolskaya
2026-06-04 12:21:09 +02:00
co-authored by Claude Sonnet 4.6
parent 2d8e995a8f
commit 02b56a52c5
4 changed files with 12 additions and 0 deletions
@@ -171,3 +171,5 @@ Over the next four articles, we'll walk through the full pipeline:
The end result: a fine-tuned SPLADE model that achieves **nDCG@10 of 0.388** on Amazon ESCI, compared to **0.301** for BM25 and **0.324** for off-the-shelf SPLADE. That 29% improvement over BM25 translates to meaningfully better search results for real e-commerce queries. You can try the models directly from HuggingFace: [splade-ecommerce-esci](https://huggingface.co/thierrydamiba/splade-ecommerce-esci) (best in-domain) and [splade-ecommerce-multidomain](https://huggingface.co/thierrydamiba/splade-ecommerce-multidomain) (better generalization).
> **Note:** These metrics were measured on a subsample of 100k products and 10k queries where all relevant documents are included. They are not directly comparable to official Amazon ESCI benchmarks and should be treated as a comparative signal only.
@@ -336,6 +336,8 @@ The results were disastrous:
| Standard SPLADE (contextual) | **0.389** |
| Inference-Free (static) | 0.065 |
> **Note:** These metrics were measured on a subsample of 100k products and 10k queries where all relevant documents are included. They are not directly comparable to official Amazon ESCI benchmarks and should be treated as a comparative signal only.
That's 6x worse without contextual encoding.
The static embedding completely failed because e-commerce queries are highly contextual. "Apple" means different things in "apple iphone" vs "apple fruit". The static embedding can't disambiguate. It looks up "apple" and returns the same vector regardless of context.
@@ -131,6 +131,8 @@ Here's what we found, evaluated on 2,000 test queries:
| SPLADE (off-the-shelf) | 0.326 | 0.339 | +7.2% |
| **SPLADE (fine-tuned)** | **0.389** | **0.387** | **+27.5%** |
> **Note:** These metrics were measured on a subsample of 100k products and 10k queries where all relevant documents are included. They are not directly comparable to official Amazon ESCI benchmarks and should be treated as a comparative signal only.
The fine-tuned model beats BM25 by nearly 28%. More telling: it beats the off-the-shelf SPLADE by 19%. The off-the-shelf model was trained on MS MARCO (web search queries), not e-commerce. That 19% gap is the value of domain-specific training.
### What About Hybrid Search?
@@ -155,6 +157,8 @@ With the **off-the-shelf SPLADE**, hybrid helps: +1.3% over sparse alone. Both s
With the **fine-tuned SPLADE**, hybrid actually hurts: SPLADE-only scored 0.413 vs hybrid at 0.405. The fine-tuned sparse model is strong enough that adding a generic dense signal dilutes the ranking. The dense model retrieves semantically similar but irrelevant products that drag down nDCG.
> **Note:** These metrics were measured on a subsample of 100k products and 10k queries where all relevant documents are included. They are not directly comparable to official Amazon ESCI benchmarks and should be treated as a comparative signal only.
This is a useful finding. Hybrid search isn't always better. It depends on the relative strength of your signals. If your sparse model is domain-tuned and your dense model is generic, the dense component can actively harm results.
## ANCE-inspired Hard Negative Mining
@@ -43,6 +43,8 @@ We took our Amazon ESCI-trained model and tested it on three additional datasets
| Home Depot | 0.349 | **0.391** | 0.384* | +10.0% |
| MS MARCO (web) | 0.915 | 0.982 | 0.751 | -17.9% |
> **Note:** These metrics were measured on a subsample of 100k products and 10k queries where all relevant documents are included. They are not directly comparable to official Amazon ESCI benchmarks and should be treated as a comparative signal only.
*On Home Depot, the off-the-shelf model edges out the fine-tuned one (0.391 vs 0.384).
Three patterns emerge:
@@ -82,6 +84,8 @@ The hypothesis: exposure to diverse e-commerce catalogs should improve cross-dom
| Home Depot | 0.384 | **0.410** | +6.8% |
| MS MARCO | 0.751 | **0.829** | +10.4% |
> **Note:** These metrics were measured on a subsample of 100k products and 10k queries where all relevant documents are included. They are not directly comparable to official Amazon ESCI benchmarks and should be treated as a comparative signal only.
Multi-domain training does exactly what you'd expect:
- **ESCI drops 4%**: Less specialization means less Amazon-specific optimization. The model can't memorize Amazon's vocabulary as deeply when it's also learning Wayfair and Home Depot patterns.