From 28f46d73653a2840aa6f0e463760901a300394eb Mon Sep 17 00:00:00 2001 From: daniel-azoulai Date: Wed, 27 Aug 2025 09:23:50 -0700 Subject: [PATCH] Update case-study-alhena.md --- qdrant-landing/content/blog/case-study-alhena.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/qdrant-landing/content/blog/case-study-alhena.md b/qdrant-landing/content/blog/case-study-alhena.md index 67182f05d..b2bea3d9d 100644 --- a/qdrant-landing/content/blog/case-study-alhena.md +++ b/qdrant-landing/content/blog/case-study-alhena.md @@ -60,9 +60,9 @@ With Qdrant handling retrieval, Alhena no longer needed to customize infrastruct ## Hitting production-grade performance targets -Latency was a critical metric for Alhena. With FAISS, vector search on catalogs with 100,000+ items often took three seconds or more. That delayed the start of agent response streaming, making the AI feel sluggish and hurting the user experience. Pinecone helped on large indexes, but introduced latency on small ones, and couldn’t handle hybrid filtering needs. +Latency was a critical metric for Alhena. With FAISS, vector search on catalogs with 100,000+ items went far above their latency budget. That delayed the start of agent response streaming, making the AI feel sluggish and hurting the user experience. Pinecone helped on large indexes, but introduced latency on small ones, and couldn’t handle hybrid filtering needs. -Qdrant reduced retrieval latency on the same datasets to approximately 300 milliseconds. That enabled Alhena to meet its internal P95 SLA of 3.5 seconds from query to first token, even after accounting for hallucination detection, policy enforcement, and contextual rewriting. +Qdrant reduced retrieval latency by up to 90% on the same datasets. That enabled Alhena to meet its internal P95 SLA from query to first token, even after accounting for hallucination detection, policy enforcement, and contextual rewriting. *“We track every millisecond. Qdrant helped us cut vector retrieval time by 90 percent at scale. That’s what made it possible to stay under our latency SLA.”* — Kang-Chi Ho