mirror of
https://github.com/qdrant/landing_page.git
synced 2026-10-03 01:48:32 +02:00
Merge pull request #1876 from qdrant/case-study-alhena
Update case-study-alhena.md
This commit is contained in:
@@ -60,9 +60,9 @@ With Qdrant handling retrieval, Alhena no longer needed to customize infrastruct
|
||||
|
||||
## Hitting production-grade performance targets
|
||||
|
||||
Latency was a critical metric for Alhena. With FAISS, vector search on catalogs with 100,000+ items often took three seconds or more. That delayed the start of agent response streaming, making the AI feel sluggish and hurting the user experience. Pinecone helped on large indexes, but introduced latency on small ones, and couldn’t handle hybrid filtering needs.
|
||||
Latency was a critical metric for Alhena. With FAISS, vector search on catalogs with 100,000+ items went far above their latency budget. That delayed the start of agent response streaming, making the AI feel sluggish and hurting the user experience. Pinecone helped on large indexes, but introduced latency on small ones, and couldn’t handle hybrid filtering needs.
|
||||
|
||||
Qdrant reduced retrieval latency on the same datasets to approximately 300 milliseconds. That enabled Alhena to meet its internal P95 SLA of 3.5 seconds from query to first token, even after accounting for hallucination detection, policy enforcement, and contextual rewriting.
|
||||
Qdrant reduced retrieval latency by up to 90% on the same datasets. That enabled Alhena to meet its internal P95 SLA from query to first token, even after accounting for hallucination detection, policy enforcement, and contextual rewriting.
|
||||
|
||||
*“We track every millisecond. Qdrant helped us cut vector retrieval time by 90 percent at scale. That’s what made it possible to stay under our latency SLA.”*
|
||||
— Kang-Chi Ho
|
||||
|
||||
Reference in New Issue
Block a user