* Modernize Day 4 Large-Scale Data Ingestion lesson The upload example on the published page cannot run: upload_collection has no `points` parameter, so it raises TypeError before sending anything. Its `show_progress` argument is also swallowed by **kwargs and does nothing. The example numbered point IDs from zero for every chunk. Because writes are idempotent, each chunk overwrote the one before it. Verified against Qdrant Cloud 1.19.0: three chunks of 100 points left 100 points in the collection, with no error raised. Other corrections: - recreate_collection is deprecated and destroys the collection on re-run. Replaced with collection_exists + create_collection. - on_disk, on_disk_payload and always_ram were replaced by the per-structure `memory` parameter in v1.19. The docs list the old flags under Legacy Settings, and `pinned` has no equivalent among them. - Batch-size guidance recommended up to 10,000 points per request. Measured on real data, requests of 4,000 and above fail outright with a connection reset rather than degrading. - Two anchors did not resolve: #hnsw-config appears nowhere else in this repo, and #on-disk-storage does not exist on the storage page. - #batch-update pointed at batching mixed operation types, not batched uploads. Changed to #upload-points. - max_segment_size and indexing_threshold are measured in kilobytes, which the lesson did not say. Adds a diagram covering the choice between upsert, upload_points and upload_collection, and links idempotence and inference, which the lesson did not previously reference. Structure, headings and section order follow the published page. Both code blocks were executed against Qdrant Cloud 1.19.0 with qdrant-client 1.19.0 and the stored collection config was read back and checked against the prose. * Address review: restore client setup, rework diagram Restore the imports and QdrantClient construction ahead of the create_collection call, matching the opening of the code blocks in what-is-quantization.md, rescoring-oversampling-indexing.md and pitstop-project.md. Diagram: drop the in-image title, raise type from 11-13px to 15-17px, and move the description out of the SVG into body text below the image. The explanation that previously sat above the image now sits below it, so it is not duplicated. * Address review: drop prefer_grpc, match callout format Remove prefer_grpc from the client constructor so it matches the other course modules. It stays in the parameter list as a suggestion. Collapse the two closing callouts to the single-line link format used across Days 3, 4 and 5, rather than multi-line boxes with their own headings. * Remove callout containers and red Note labels Convert the six blockquote callouts to plain body text and drop the red "Note" and "Best Practice" labels. Wording and links are unchanged. * Restore Note callouts, drop the LAION link Put back the three Note callouts higher up the page, which match the style used in the other course modules. The three containers flagged in review stay as plain text, and the LAION benchmark link is removed. * Restore the LAION link as plain text The review asked for the containers to be removed, not the content, so put the link back without its box. * Remove all callout boxes and red labels Review asked for these throughout the page, not only the three flagged in the screenshot. Wording and links are unchanged. * Point the lesson at the re-recorded video
7.3 KiB
title, short_description, description, weight, isLesson
| title | short_description | description | weight | isLesson |
|---|---|---|---|---|
| Large-Scale Data Ingestion | Choose the right ingestion strategy for Qdrant: batched upserts, upload_points, and streaming uploads for million- and billion-scale workloads. | Master large-scale vector ingestion in Qdrant. Compare upsert, upload_points, and upload_collection, and learn how to stream a large dataset into a collection without loading it into memory. | 4 | true |
{{< date >}} Day 4 {{< /date >}}
Large-Scale Data Ingestion
In vector search applications inserting a few thousand data points is straightforward but the dynamics change completely when dealing with millions or billions of records. Tiny inefficiencies in the ingestion process compound into significant time losses, increased memory pressure, and degraded search performance.
Every individual upsert call initiates a transaction that consumes memory and disk I/O to build parts of the index. At scale, this naive approach can overwhelm your system, causing upload times to spike and search quality to decrease. Efficiently preparing and loading your data into Qdrant is paramount for building a reliable and scalable AI application.
Choosing Your Ingestion Strategy
The Qdrant client gives you three ways to get points in. The first has you managing the batching; the other two hand that to the client.
-
upsert is the basic write operation, and the one every client library has. Send points one at a time for real-time updates, or in batches for a bulk load, which minimizes the overhead of opening a connection per point. You decide the batch size and you send the requests.
-
upload_points takes an iterable of
models.PointStruct, the record-oriented shape: one object per point, carrying its own id, vector, and payload. -
upload_collection takes
vectors,payload, andidsas separate arguments, the column-oriented shape.
Those last two are the same tool in two shapes. Both add parallelization, retries, and lazy batching on top of upsert, and because both accept iterators, neither needs the whole dataset in memory. Pick whichever matches how your data already sits: the docs note the two formats are equivalent internally and offered for convenience.
You can also skip generating embeddings yourself. With inference, you send the text or image and the model name, and Qdrant produces the vector on upsert.
upload_points and upload_collection are helpers in the client library rather than server endpoints, so what's available depends on your language:
| Client | Bulk helper |
|---|---|
| Python | upload_points, upload_collection |
| Rust | upsert_points_chunked(request, chunk_size) |
| TypeScript, Go, Java, C# | Batched upsert calls |
The bottleneck during upload is usually the client library, not the Qdrant server. If ingestion speed is your priority, the Rust client is the fastest option.
The Collection Configuration
When a collection is too large to hold in memory, each structure takes a memory parameter that says where it lives. pinned stays on the heap, cached is memory-mapped and pre-warmed, and cold is memory-mapped and read on demand.
from qdrant_client import QdrantClient, models
import os
client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))
client.create_collection(
collection_name="my_collection",
vectors_config=models.VectorParams(
size=512,
distance=models.Distance.COSINE,
datatype=models.Datatype.FLOAT16,
memory=models.Memory.COLD,
),
payload=models.PayloadStorageParams(memory=models.Memory.COLD),
quantization_config=models.BinaryQuantization(
binary=models.BinaryQuantizationConfig(memory=models.Memory.PINNED),
),
hnsw_config=models.HnswConfigDiff(memory=models.Memory.PINNED),
)
Two things catch people out. pinned is rejected for dense vectors, which support only cached or cold. And max_segment_size and indexing_threshold are both measured in kilobytes rather than points. See Memory Tiers for which tier suits which structure.
memory arrived in Qdrant v1.19. If you are following older material, it replaces on_disk on the vectors, on_disk_payload on the collection, always_ram in the quantization config, and on_disk in the HNSW config. Those still work but are deprecated, and pinned has no equivalent among them.
The Upload Process
Hand the method an iterable and it takes care of the requests. Because it accepts an iterator, you can feed it a generator that reads from disk as it goes, rather than materializing the whole set first.
import tqdm
client.upload_collection(
collection_name="my_collection",
vectors=embeddings,
payload=payloads,
ids=tqdm.tqdm(ids),
batch_size=256,
parallel=4,
)
A few things worth knowing about these parameters:
idsmust be unique across the whole upload. Writes are idempotent, so a point sent under an id that already exists overwrites it instead of erroring. That is what you want on a retry, and what bites you if two batches reuse the same numbers.batch_sizecontrols how many points go in each request. The Bulk Upload guide covers how to pick it, along with sharding and payload indexes.parallelstarts worker processes. Each one opens its own connection, so if batches begin failing after you raise it, drop back to1.tqdmaround any of the iterables gives you progress. There is noshow_progressparameter.prefer_grpc=Trueon the client skips JSON serialization on every batch.update_mode=models.UpdateMode.INSERT_ONLY(v1.17) makes a resumed upload skip points already in the collection rather than rewriting them. See Update Mode.- On Windows and macOS,
parallelgreater than 1 needs the upload call behind aif __name__ == "__main__":guard in a script, because Python starts workers by re-importing your module. Without it the upload hangs instead of failing. Notebooks are unaffected.
Start small and test. Before attempting to upload your entire dataset, ingest a smaller chunk to validate your configuration and process.
Try it hands-on in the Google Colab notebook.
See it at 400 million points in the LAION-400M benchmark.