New search-as-you-type article (#231)

* New search-as-you-type article

* fix image links

* add benchmark before

* fix some code and verbiage, add final benchmark table

* fix link and table

* stray period removed

* fix typos

* grammar and capitalization

* one more code fix

* one more code fix

* explain recommend more, add benchmark results bar chart

* extend the article with info about batching and concurrent requests

* sum up

* add code link

* re-add deduplication

* aside on benchmark interpretation, thanks to @generall

* fix missing " in json

* review changes for sayt

* remove stray quotation marks

* upd images + minor renaming

* Update qdrant-landing/content/articles/search-as-you-type.md

---------

Co-authored-by: generall <andrey@vasnetsov.com>
This commit is contained in:
llogiq
2023-08-14 05:15:01 +02:00
committed by GitHub
co-authored by generall
parent ec873ec73b
commit d82f65a6c2
14 changed files with 247 additions and 0 deletions
@@ -0,0 +1,136 @@
---
title: Semantic Search As You Type
short_description: "A demo that does what the title says"
description: To show off Qdrant's performance, we show how to do a quick search-as-you-type that will come back within a few milliseconds.
social_preview_image: /articles_data/search-as-you-type/preview/social_preview.jpg
small_preview_image: /articles_data/search-as-you-type/icon.svg
preview_dir: /articles_data/search-as-you-type/preview
weight: -2
author: Andre Bogus
author_link: https://llogiq.github.io
date: 2023-08-14T00:00:00+01:00
draft: false
keywords: search, semantic, vector, llm, integration, benchmark, recommend, performance, rust
---
Qdrant is one of the fastest vector search engines out there, so while looking for a demo to show off, we came upon the idea to do a search-as-you-type box with a fully semantic search backend. Now we already have a semantic/keyword hybrid search on our website. But that one is written in Python, which incurs some overhead for the interpreter. Naturally, I wanted to see how fast I could go using Rust.
Since Qdrant doesn't embed by itself, I had to decide on an embedding model. The prior version used the [SentenceTransformers](https://www.sbert.net/) package, which in turn employs Bert-based [All-MiniLM-L6-V2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/tree/main) model. This model is battle-tested and delivers fair results at speed, so not experimenting on this front I took an [ONNX version](https://huggingface.co/optimum/all-MiniLM-L6-v2/tree/main) and ran that within the service.
The workflow looks like this:
![Search Qdrant by Embedding](/articles_data/search-as-you-type/Qdrant_Search_by_Embedding.png)
This will, after tokenizing and embedding send a `/collections/site/points/search` POST request to Qdrant, sending the following JSON:
```json
POST collections/site/points/search
{
"vector": [-0.06716014,-0.056464013, ...(382 values omitted)],
"limit": 5,
"with_payload": true,
}
```
Even with avoiding a network round-trip, the embedding still takes some time. As always in optimization, if you cannot do the work faster, a good solution is to avoid work altogether (please don't tell my employer). This can be done by pre-computing common prefixes and calculating embeddings for them, then storing them in a `prefix_cache` collection. Now the [`recommend`](https://docs.rs/qdrant-client/latest/qdrant_client/client/struct.QdrantClient.html#method.recommend) API method can find the best matches without doing any embedding. For now, I use short (up to and including 5 letters) prefixes, but I can also parse the logs to get the most common search terms and add them to the cache later.
![Qdrant Recommendation](/articles_data/search-as-you-type/Qdrant_Recommendation.png)
Making that work requires setting up the `prefix_cache` collection with points that have the prefix as their `point_id` and the embedding as their `vector`, which lets us do the lookup with no search or index. The `prefix_to_id` function currently uses the `u64` variant of `PointId`, which can hold eight bytes, enough for this use. If the need arises, one could instead encode the names as UUID, hashing the input. Since I know all our prefixes are within 8 bytes, I decided against this for now.
The `recommend` endpoint works roughly the same as `search_points`, but instead of searching for a vector, Qdrant searches for one or more points (you can also give negative example points the search engine will try to avoid in the results). It was built to help drive recommendation engines, saving the round-trip of sending the current point's vector back to Qdrant to find more similar ones. However Qdrant goes a bit further by allowing us to select a different collection to lookup the points, which allows us to keep our `prefix_cache` collection separate from the site data. So in our case, Qdrant first looks up the point from the `prefix_cache`, takes its vector and searches for that in the `site` collection, using the precomputed embeddings from the cache. The API endpoint expects a POST of the following JSON to `/collections/site/points/recommend`:
```json
POST collections/site/points/recommend
{
"positive": [1936024932],
"limit": 5,
"with_payload": true,
"lookup_from": {
"collection": "prefix_cache"
}
}
```
Now I have, in the best Rust tradition, a blazingly fast semantic search.
To demo it, I used our [Qdrant documentation website](https://qdrant.tech/documentation)'s page search, replacing our previous Python implementation. So in order to not just spew empty words, here is a benchmark, showing different queries that exercise different code paths.
Since the operations themselves are far faster than the network whose fickle nature would have swamped most measurable differences, I benchmarked both the Python and Rust services locally. I'm measuring both versions on the same AMD Ryzen 9 5900HX with 16GB RAM running Linux. The table shows the average time and error bound in milliseconds. I only measured up to a thousand concurrent requests. None of the services showed any slowdown with more requests in that range. I do not expect our service to become DDOS'd, so I didn't benchmark with more load.
Without further ado, here are the results:
| query length | Short | Long |
|---------------|-----------|------------|
| Python 🐍 | 16 ± 4 ms | 16 ± 4 ms |
| Rust 🦀 | 1½ ± ½ ms | 5 ± 1 ms |
The Rust version consistently outperforms the Python version and offers a semantic search even on few-character queries. If the prefix cache is hit (as in the short query length), the semantic search can even get more than ten times faster than the Python version. The general speed-up is due to both the relatively lower overhead of Rust + Actix Web compared to Python + FastAPI (even if that already performs admirably), as well as using ONNX Runtime instead of SentenceTransformers for the embedding. The prefix cache gives the Rust version a real boost by doing a semantic search without doing any embedding work.
As an aside, while the millisecond differences shown here may mean relatively little for our users, whose latency will be dominated by the network in between, when typing, every millisecond more or less can make a difference in user perception. Also search-as-you-type generates between three and five times as much load as a plain search, so the service will experience more traffic. Less time per request means being able to handle more of them.
Mission accomplished! But wait, there's more!
### Prioritizing Exact Matches and Headings
To improve on the quality of the results, Qdrant can do multiple searches in parallel, and then the service puts the results in sequence, taking the first best matches. The extended code searches:
1. Text matches in titles
2. Text matches in body (paragraphs or lists)
3. Semantic matches in titles
4. Any Semantic matches
Those are put together by taking them in the above order, deduplicating as necessary.
![merge workflow](/articles_data/search-as-you-type/sayt_merge.png)
Instead of sending a `search` or `recommend` request, one can also send a `search/batch` or `recommend/batch` request, respectively. Each of those contain a `"searches"` property with any number of search/recommend JSON requests:
```json
POST collections/site/points/search/batch
{
"searches": [
{
"vector": [-0.06716014,-0.056464013, ...],
"filter": {
"must": [
{ "key": "text", "match": { "text": <query> }},
{ "key": "tag", "match": { "any": ["h1", "h2", "h3"] }},
]
}
...,
},
{
"vector": [-0.06716014,-0.056464013, ...],
"filter": {
"must": [ { "key": "body", "match": { "text": <query> }} ]
}
...,
},
{
"vector": [-0.06716014,-0.056464013, ...],
"filter": {
"must": [ { "key": "tag", "match": { "any": ["h1", "h2", "h3"] }} ]
}
...,
},
{
"vector": [-0.06716014,-0.056464013, ...],
...,
},
]
}
```
As the queries are done in a batch request, there isn't any additional network overhead and only very modest computation overhead, yet the results will be better in many cases.
The only additional complexity is to flatten the result lists and take the first 5 results, deduplicating by point ID. Now there is one final problem: The query may be short enough to take the recommend code path, but still not be in the prefix cache. In that case, doing the search *sequentially* would mean two round-trips between the service and the Qdrant instance. The solution is to *concurrently* start both requests and take the first successful non-empty result.
![sequential vs. concurrent flow](/articles_data/search-as-you-type/sayt_concurrency.png)
While this means more load for the Qdrant vector search engine, this is not the limiting factor. The relevant data is already in cache in many cases, so the overhead stays within acceptable bounds, and the maximum latency in case of prefix cache misses is measurably reduced.
The code is available on the [Qdrant github](https://github.com/qdrant/page-search)
To sum up: Rust is fast, recommend lets us use precomputed embeddings, batch requests are awesome and one can do a semantic search in mere milliseconds.
Binary file not shown.

After

Width:  |  Height:  |  Size: 22 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 36 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 145 KiB

@@ -0,0 +1,111 @@
<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<svg
version="1.0"
width="512.000000pt"
height="512.000000pt"
viewBox="0 0 512.000000 512.000000"
preserveAspectRatio="xMidYMid meet"
id="svg38"
sodipodi:docname="icon.svg"
inkscape:version="1.1.2 (0a00cf5339, 2022-02-04)"
xmlns:inkscape="http://www.inkscape.org/namespaces/inkscape"
xmlns:sodipodi="http://sodipodi.sourceforge.net/DTD/sodipodi-0.dtd"
xmlns="http://www.w3.org/2000/svg"
xmlns:svg="http://www.w3.org/2000/svg">
<defs
id="defs42" />
<sodipodi:namedview
id="namedview40"
pagecolor="#ffffff"
bordercolor="#999999"
borderopacity="1"
inkscape:pageshadow="0"
inkscape:pageopacity="0"
inkscape:pagecheckerboard="true"
inkscape:document-units="pt"
showgrid="false"
inkscape:zoom="0.7130816"
inkscape:cx="241.20662"
inkscape:cy="395.46666"
inkscape:window-width="2560"
inkscape:window-height="1011"
inkscape:window-x="0"
inkscape:window-y="32"
inkscape:window-maximized="1"
inkscape:current-layer="svg38" />
<g
transform="translate(0.000000,512.000000) scale(0.100000,-0.100000)"
fill="#000000"
stroke="none"
id="g36"
style="fill:#ffffff">
<path
d="M2504 5089 c-18 -19 -19 -42 -22 -301 -2 -154 1 -316 6 -359 17 -142 89 -244 210 -302 57 -26 76 -30 200 -36 119 -6 143 -10 179 -31 84 -47 112 -119 113 -282 0 -145 -13 -194 -69 -249 -57 -57 -94 -69 -227 -69 -151 0 -235 -31 -319 -117 -21 -21 -50 -63 -64 -93 -22 -47 -26 -71 -29 -167 l-4 -113 -64 0 c-75 0 -130 -23 -160 -67 -15 -22 -20 -52 -24 -148 l-5 -120 -1015 -5 c-965 -5 -1017 -6 -1050 -24 -59 -31 -119 -96 -140 -154 -20 -51 -20 -79 -20 -1132 0 -1059 0 -1080 20 -1133 26 -68 77 -122 149 -156 l56 -26 2335 0 2335 0 56 26 c72 34 123 88 149 156 20 53 20 74 20 1133 0 1053 0 1081 -20 1132 -23 63 -85 128 -150 157 -43 20 -64 21 -481 21 -402 0 -439 -1 -458 -18 -30 -24 -31 -76 -2 -105 21 -22 25 -22 448 -25 411 -3 427 -3 455 -23 61 -44 58 22 58 -1139 0 -1156 3 -1092 -56 -1139 l-27 -21 -2327 0 -2327 0 -27 21 c-59 47 -56 -17 -56 1139 0 1162 -3 1095 59 1139 l29 21 1727 2 1727 3 19 24 c25 30 24 76 -1 101 -19 19 -33 20 -419 20 l-399 0 -4 119 c-3 102 -6 124 -24 147 -41 55 -70 69 -155 73 l-79 3 0 90 c0 96 11 130 59 181 48 53 75 61 227 67 129 6 146 9 205 37 155 73 219 197 219 428 0 211 -49 327 -170 405 -67 43 -137 60 -250 60 -115 0 -160 11 -206 52 -69 61 -68 56 -74 434 -5 318 -6 343 -24 363 -26 29 -86 29 -112 0z m236 -2364 l0 -95 -180 0 -180 0 0 95 0 95 180 0 180 0 0 -95z"
id="path2"
style="fill:#ffffff" />
<path
d="M393 2264 c-64 -32 -68 -48 -68 -256 0 -173 2 -189 21 -215 43 -58 69 -64 279 -61 274 4 275 6 275 275 0 267 -7 273 -290 273 -148 0 -191 -4 -217 -16z m357 -259 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path4"
style="fill:#ffffff" />
<path
d="M1043 2264 c-64 -32 -68 -48 -68 -254 0 -165 2 -188 19 -217 33 -54 50 -58 266 -58 288 0 284 -4 288 260 4 280 -1 285 -288 285 -148 0 -191 -4 -217 -16z m357 -259 l0 -126 -137 3 -138 3 -3 123 -3 122 141 0 140 0 0 -125z"
id="path6"
style="fill:#ffffff" />
<path
d="M1695 2266 c-17 -7 -40 -26 -50 -42 -18 -26 -20 -47 -20 -219 0 -275 -5 -270 285 -270 290 0 285 -5 285 270 0 276 1 275 -287 275 -129 -1 -192 -5 -213 -14z m355 -261 l0 -126 -137 3 -138 3 -3 123 -3 122 141 0 140 0 0 -125z"
id="path8"
style="fill:#ffffff" />
<path
d="M2345 2266 c-17 -7 -40 -26 -50 -42 -18 -26 -20 -47 -20 -219 0 -275 -5 -270 285 -270 290 0 285 -5 285 270 0 276 1 275 -287 275 -129 -1 -192 -5 -213 -14z m355 -261 l0 -125 -140 0 -140 0 0 125 0 125 140 0 140 0 0 -125z"
id="path10"
style="fill:#ffffff" />
<path
d="M2993 2265 c-63 -27 -68 -47 -68 -260 0 -275 -5 -270 285 -270 290 0 285 -5 285 270 0 276 1 275 -287 275 -133 -1 -192 -5 -215 -15z m355 -257 l-3 -123 -137 -3 -138 -3 0 126 0 125 140 0 141 0 -3 -122z"
id="path12"
style="fill:#ffffff" />
<path
d="M3640 2262 c-63 -31 -71 -62 -68 -267 4 -264 0 -260 288 -260 216 0 233 4 266 58 17 29 19 52 19 217 0 269 -1 270 -287 270 -154 0 -188 -3 -218 -18z m358 -254 l-3 -123 -137 -3 -138 -3 0 126 0 125 140 0 141 0 -3 -122z"
id="path14"
style="fill:#ffffff" />
<path
d="M4290 2262 c-61 -30 -70 -63 -70 -255 0 -269 1 -271 275 -275 210 -3 236 3 279 61 19 26 21 42 21 215 0 271 -1 272 -287 272 -154 0 -188 -3 -218 -18z m350 -257 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path16"
style="fill:#ffffff" />
<path
d="M392 1543 c-60 -29 -67 -56 -67 -258 0 -158 2 -184 19 -211 36 -59 32 -59 386 -62 214 -2 338 0 362 8 20 5 48 24 62 41 26 30 26 31 26 221 0 212 -6 234 -66 263 -49 23 -674 22 -722 -2z m638 -258 l0 -125 -275 0 -275 0 0 125 0 125 275 0 275 0 0 -125z"
id="path18"
style="fill:#ffffff" />
<path
d="M1340 1542 c-61 -30 -70 -63 -70 -255 0 -185 8 -220 56 -252 26 -18 47 -20 237 -20 l209 0 34 37 34 38 0 190 c0 282 2 280 -282 280 -154 0 -188 -3 -218 -18z m350 -257 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path20"
style="fill:#ffffff" />
<path
d="M2015 1546 c-67 -30 -70 -38 -73 -250 -3 -187 -3 -192 20 -226 35 -51 69 -59 266 -59 189 -1 222 6 260 54 21 26 22 38 22 217 0 279 1 278 -282 278 -129 -1 -192 -5 -213 -14z m345 -261 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path22"
style="fill:#ffffff" />
<path
d="M2674 1544 c-58 -29 -64 -51 -64 -262 0 -179 1 -191 22 -217 38 -48 71 -55 260 -54 197 0 231 8 266 59 23 34 23 39 20 226 -3 213 -6 221 -75 250 -48 21 -387 19 -429 -2z m356 -259 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path24"
style="fill:#ffffff" />
<path
d="M3342 1543 c-57 -28 -62 -50 -62 -263 l0 -190 34 -38 34 -37 209 0 c294 0 293 -1 293 272 0 267 -7 273 -290 273 -150 0 -191 -4 -218 -17z m358 -258 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path26"
style="fill:#ffffff" />
<path
d="M4004 1544 c-58 -29 -64 -51 -64 -262 0 -190 0 -191 26 -221 14 -17 42 -36 62 -41 24 -8 148 -10 362 -8 354 3 350 3 386 62 17 27 19 53 19 211 0 288 20 275 -432 275 -272 0 -333 -3 -359 -16z m636 -259 l0 -125 -275 0 -275 0 0 125 0 125 275 0 275 0 0 -125z"
id="path28"
style="fill:#ffffff" />
<path
d="M415 833 c-34 -9 -63 -33 -76 -66 -10 -23 -14 -82 -14 -207 0 -155 2 -179 19 -207 33 -55 49 -58 270 -58 287 0 286 -1 286 272 0 266 -7 274 -285 272 -99 -1 -189 -4 -200 -6z m335 -268 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path30"
style="fill:#ffffff" />
<path
d="M1103 824 c-64 -32 -68 -48 -68 -259 0 -208 4 -225 62 -256 39 -21 2887 -21 2926 0 58 31 62 48 62 256 0 167 -2 193 -18 217 -42 61 23 58 -1201 58 -807 0 -1122 -3 -1140 -11 -37 -17 -51 -60 -32 -97 8 -16 23 -32 31 -36 9 -3 511 -6 1116 -6 l1099 0 0 -125 0 -125 -1380 0 -1380 0 0 125 0 125 103 0 c95 0 105 2 125 23 31 33 29 80 -4 106 -24 19 -40 21 -148 21 -90 0 -129 -5 -153 -16z"
id="path32"
style="fill:#ffffff" />
<path
d="M4290 823 c-61 -31 -70 -64 -70 -256 0 -273 -1 -272 286 -272 221 0 237 3 270 58 17 28 19 52 19 207 0 276 3 274 -275 278 -171 2 -199 1 -230 -15z m350 -258 l0 -125 -135 0 -135 0 0 125 0 125 135 0 135 0 0 -125z"
id="path34"
style="fill:#ffffff" />
</g>
</svg>

After

Width:  |  Height:  |  Size: 7.2 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 13 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 29 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 28 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 497 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 90 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 86 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 34 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 18 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 26 KiB