diff --git a/qdrant-landing/content/articles/what-qdrant-optimizers-cost-your-queries.md b/qdrant-landing/content/articles/what-qdrant-optimizers-cost-your-queries.md index f28fe0c48..173af088f 100644 --- a/qdrant-landing/content/articles/what-qdrant-optimizers-cost-your-queries.md +++ b/qdrant-landing/content/articles/what-qdrant-optimizers-cost-your-queries.md @@ -124,18 +124,27 @@ In steady state, however, the single-segment configuration delivered the best se ## Optimizer Threads: Smoother Queries or a Shorter Wait, Pick One -We tested `max_optimization_threads` and `max_indexing_threads` on two settings: both set to 1 (serialized) and Qdrant's auto-select default. + + +Optimizers run on the same threads as your Qdrant instance: controlling the access that optimizers have to the CPU, by limiting or increasing the number of threads they can run on, is a way of controlling how fast they clear your collection's backlog and how much space is left for search workloads. + +By setting both `max_optimizers_threads` and `max_indexing_threads` to 1, we saw the draining window stretch for 3,244.1s, 6.8 times longer than it took when leaving those settings at Qdrant's default values. This, of course, comes with decreased search latency during the draining phase, since optimizers have a limited CPU budget and do not compete as much with search operations: p95, in this case, was capeped at 373.7 ms, less than half of what the default configuration got to (820.4ms). + +This quantifies the read/write contention tradeoff that the docs explain: serializing optimizers work by reducing their CPU budget to a single thread means a smoother, more predictable per-query latency while the backlog clears, at the direct cost of how long the backlog takes to clear. + +{{ Chart on draining duration and draining latency for the two runs }} + ## Vacuum: The Same Deletion, Two Opposite Outcomes -This comparison depends entirely on a threshold most people never touch. We loaded two identical collections, deleted the same 25% of points from each, and compared what happened next. The difference between them was `deleted_threshold`, the fraction of a segment's points that has to be marked deleted before Qdrant will vacuum it. + + +An area that requires optimization, and that most users rarely touch, is deletion: as many other databases do, Qdrant applies only a "soft delete" operation to its data when a `DELETE` request is sent. Simply put, deleted points are marked as such and simply skipped at query time: this avoids expensive disk writes to record deletions, and keeps the overall latency of the operation down. + +When a sufficiently high number of deleted points is reached (determined bz the `deleted_threshold` config), Qdrant starts vacuuming the points, i.e. removing them from the persistent storage layer. This job is handled by the `vacuum` optimizer, and its activity slows down reads, for the same read/write contention principle we saw above. + +We tried setting the `deleted_threshold` to 20% and deleting ~25% of the points, triggering the optimizer, and then doing the same but with a 50% threshold. This single configuration change caused a big divide between the two runs: search latency jumped from a 5.0ms in steady state to 22.7ms during vacuuming in the first case, while it improved in the second, going from 5.4ms to 4.3ms (improvement due to the minor number of searchable points). + +Setting a higher `deleted_threshold` could then seem the perfect move: no vacuuming, no latency spikes, improved performance over a drop in the points count. As for all decisions, though, it comes with a tradeoff, in this case higher storage space required to fit the collection. While this can be fine for smaller datasets, disk space can become a constraint in larger collections, thus the `deleted_threshold` becomes less of free win and more of a choice to carefully ponderate. + +{{ chart on steady state vs vacuuming latency }} ## Turning Indexing on Later Reopens the Whole Backlog at Once + -## What We Would Do With This +One last strategy for keeping read-write contention under control is to turn optimizations off during upload, and resume them once the upload is finished. + +We put that to the test by running the benchmark with indexing turned off, then reconfiguring the collection to activate indexing by lowering the indexing threshold, and finally we measured query latency on 5 rounds of 1000 queries each. We tried with both `prevent_unoptimized` set to `true` and `false`. + +What we could see is that, without continuous indxing, collections do not benefit from the incremental buildout of the index during upload, which means longer indexing times and high latency during this phase. There are, however, big differences between how the two runs handled the optimization progress: in the first case, with `prevent_unoptimized` set to `false`, we did not record optimizers going idle and search latency climbed up to an overall median of 2.7s, with the tail reaching 12.1s. This is due to the fact that search queries come in continuously, constantly hitting partially optimized or fully unoptimized segments, and they compete for the same I/O and CPU resources the optimizers are also trying to access. + +On the other hand, preventing unoptimized segments from being touched by queries allowed the segments to progress in their optimizations much faster, being able to complete them by the end of the first round of qeuries: latency benefits from it, with a p95 of 621.7ms, and a steady-state p95 of 5.7ms. As said before, this improvement is only applicable if a temporary recall and results loss is acceptable in favor of better query latency. + +## Takeaways - Turn on `prevent_unoptimized` before a bulk load if incomplete results during that window are acceptable, and switch writes to `wait=false` first if your client defaults to `wait=true`. - Cap segment size for large loads if slower steady-state queries are an acceptable trade.