Add metric comparability notes to SPLADE series parts 1-4

Metrics were measured on a subsample (100k products, 10k queries) with
all relevant documents included, so they are not comparable to official
Amazon ESCI benchmarks. Add inline notes at each metric table/result.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Evgeniya Sukhodolskaya
2026-06-04 12:21:09 +02:00
co-authored by Claude Sonnet 4.6
parent 2d8e995a8f
commit 02b56a52c5
4 changed files with 12 additions and 0 deletions
@@ -336,6 +336,8 @@ The results were disastrous:
| Standard SPLADE (contextual) | **0.389** |
| Inference-Free (static) | 0.065 |
> **Note:** These metrics were measured on a subsample of 100k products and 10k queries where all relevant documents are included. They are not directly comparable to official Amazon ESCI benchmarks and should be treated as a comparative signal only.
That's 6x worse without contextual encoding.
The static embedding completely failed because e-commerce queries are highly contextual. "Apple" means different things in "apple iphone" vs "apple fruit". The static embedding can't disambiguate. It looks up "apple" and returns the same vector regardless of context.