Files
landing_page/qdrant-landing/content/course/essentials/day-4/large-scale-ingestion.md
T
John Kupchanko | Qdrant 4c69107c5b Modernize Day 4: Large-Scale Data Ingestion (#2664)
* Modernize Day 4 Large-Scale Data Ingestion lesson

The upload example on the published page cannot run: upload_collection has
no `points` parameter, so it raises TypeError before sending anything. Its
`show_progress` argument is also swallowed by **kwargs and does nothing.

The example numbered point IDs from zero for every chunk. Because writes are
idempotent, each chunk overwrote the one before it. Verified against Qdrant
Cloud 1.19.0: three chunks of 100 points left 100 points in the collection,
with no error raised.

Other corrections:

- recreate_collection is deprecated and destroys the collection on re-run.
  Replaced with collection_exists + create_collection.
- on_disk, on_disk_payload and always_ram were replaced by the per-structure
  `memory` parameter in v1.19. The docs list the old flags under Legacy
  Settings, and `pinned` has no equivalent among them.
- Batch-size guidance recommended up to 10,000 points per request. Measured
  on real data, requests of 4,000 and above fail outright with a connection
  reset rather than degrading.
- Two anchors did not resolve: #hnsw-config appears nowhere else in this
  repo, and #on-disk-storage does not exist on the storage page.
- #batch-update pointed at batching mixed operation types, not batched
  uploads. Changed to #upload-points.
- max_segment_size and indexing_threshold are measured in kilobytes, which
  the lesson did not say.

Adds a diagram covering the choice between upsert, upload_points and
upload_collection, and links idempotence and inference, which the lesson did
not previously reference.

Structure, headings and section order follow the published page. Both code
blocks were executed against Qdrant Cloud 1.19.0 with qdrant-client 1.19.0
and the stored collection config was read back and checked against the prose.

* Address review: restore client setup, rework diagram

Restore the imports and QdrantClient construction ahead of the
create_collection call, matching the opening of the code blocks in
what-is-quantization.md, rescoring-oversampling-indexing.md and
pitstop-project.md.

Diagram: drop the in-image title, raise type from 11-13px to 15-17px,
and move the description out of the SVG into body text below the image.
The explanation that previously sat above the image now sits below it,
so it is not duplicated.

* Address review: drop prefer_grpc, match callout format

Remove prefer_grpc from the client constructor so it matches the other
course modules. It stays in the parameter list as a suggestion.

Collapse the two closing callouts to the single-line link format used
across Days 3, 4 and 5, rather than multi-line boxes with their own
headings.

* Remove callout containers and red Note labels

Convert the six blockquote callouts to plain body text and drop the
red "Note" and "Best Practice" labels. Wording and links are unchanged.

* Restore Note callouts, drop the LAION link

Put back the three Note callouts higher up the page, which match the
style used in the other course modules. The three containers flagged in
review stay as plain text, and the LAION benchmark link is removed.

* Restore the LAION link as plain text

The review asked for the containers to be removed, not the content, so
put the link back without its box.

* Remove all callout boxes and red labels

Review asked for these throughout the page, not only the three flagged
in the screenshot. Wording and links are unchanged.

* Point the lesson at the re-recorded video
2026-09-04 10:42:56 -07:00

7.3 KiB

title, short_description, description, weight, isLesson
title short_description description weight isLesson
Large-Scale Data Ingestion Choose the right ingestion strategy for Qdrant: batched upserts, upload_points, and streaming uploads for million- and billion-scale workloads. Master large-scale vector ingestion in Qdrant. Compare upsert, upload_points, and upload_collection, and learn how to stream a large dataset into a collection without loading it into memory. 4 true

{{< date >}} Day 4 {{< /date >}}

Large-Scale Data Ingestion


In vector search applications inserting a few thousand data points is straightforward but the dynamics change completely when dealing with millions or billions of records. Tiny inefficiencies in the ingestion process compound into significant time losses, increased memory pressure, and degraded search performance.

Every individual upsert call initiates a transaction that consumes memory and disk I/O to build parts of the index. At scale, this naive approach can overwhelm your system, causing upload times to spike and search quality to decrease. Efficiently preparing and loading your data into Qdrant is paramount for building a reliable and scalable AI application.

Choosing Your Ingestion Strategy

The Qdrant client gives you three ways to get points in. The first has you managing the batching; the other two hand that to the client.

  • upsert is the basic write operation, and the one every client library has. Send points one at a time for real-time updates, or in batches for a bulk load, which minimizes the overhead of opening a connection per point. You decide the batch size and you send the requests.

  • upload_points takes an iterable of models.PointStruct, the record-oriented shape: one object per point, carrying its own id, vector, and payload.

  • upload_collection takes vectors, payload, and ids as separate arguments, the column-oriented shape.

upsert takes models.Batch or a list and you batch it yourself. upload_points is record-oriented, an iterable of PointStruct. upload_collection is column-oriented, taking vectors, payload and ids as parallel columns. All three write into the collection.

Those last two are the same tool in two shapes. Both add parallelization, retries, and lazy batching on top of upsert, and because both accept iterators, neither needs the whole dataset in memory. Pick whichever matches how your data already sits: the docs note the two formats are equivalent internally and offered for convenience.

You can also skip generating embeddings yourself. With inference, you send the text or image and the model name, and Qdrant produces the vector on upsert.

upload_points and upload_collection are helpers in the client library rather than server endpoints, so what's available depends on your language:

Client Bulk helper
Python upload_points, upload_collection
Rust upsert_points_chunked(request, chunk_size)
TypeScript, Go, Java, C# Batched upsert calls

The bottleneck during upload is usually the client library, not the Qdrant server. If ingestion speed is your priority, the Rust client is the fastest option.

The Collection Configuration

When a collection is too large to hold in memory, each structure takes a memory parameter that says where it lives. pinned stays on the heap, cached is memory-mapped and pre-warmed, and cold is memory-mapped and read on demand.

from qdrant_client import QdrantClient, models
import os

client = QdrantClient(url=os.getenv("QDRANT_URL"), api_key=os.getenv("QDRANT_API_KEY"))

client.create_collection(
    collection_name="my_collection",
    vectors_config=models.VectorParams(
        size=512,
        distance=models.Distance.COSINE,
        datatype=models.Datatype.FLOAT16,
        memory=models.Memory.COLD,
    ),
    payload=models.PayloadStorageParams(memory=models.Memory.COLD),
    quantization_config=models.BinaryQuantization(
        binary=models.BinaryQuantizationConfig(memory=models.Memory.PINNED),
    ),
    hnsw_config=models.HnswConfigDiff(memory=models.Memory.PINNED),
)

Two things catch people out. pinned is rejected for dense vectors, which support only cached or cold. And max_segment_size and indexing_threshold are both measured in kilobytes rather than points. See Memory Tiers for which tier suits which structure.

memory arrived in Qdrant v1.19. If you are following older material, it replaces on_disk on the vectors, on_disk_payload on the collection, always_ram in the quantization config, and on_disk in the HNSW config. Those still work but are deprecated, and pinned has no equivalent among them.

The Upload Process

Hand the method an iterable and it takes care of the requests. Because it accepts an iterator, you can feed it a generator that reads from disk as it goes, rather than materializing the whole set first.

import tqdm

client.upload_collection(
    collection_name="my_collection",
    vectors=embeddings,
    payload=payloads,
    ids=tqdm.tqdm(ids),
    batch_size=256,
    parallel=4,
)

A few things worth knowing about these parameters:

  • ids must be unique across the whole upload. Writes are idempotent, so a point sent under an id that already exists overwrites it instead of erroring. That is what you want on a retry, and what bites you if two batches reuse the same numbers.
  • batch_size controls how many points go in each request. The Bulk Upload guide covers how to pick it, along with sharding and payload indexes.
  • parallel starts worker processes. Each one opens its own connection, so if batches begin failing after you raise it, drop back to 1.
  • tqdm around any of the iterables gives you progress. There is no show_progress parameter.
  • prefer_grpc=True on the client skips JSON serialization on every batch.
  • update_mode=models.UpdateMode.INSERT_ONLY (v1.17) makes a resumed upload skip points already in the collection rather than rewriting them. See Update Mode.
  • On Windows and macOS, parallel greater than 1 needs the upload call behind a if __name__ == "__main__": guard in a script, because Python starts workers by re-importing your module. Without it the upload hangs instead of failing. Notebooks are unaffected.

Start small and test. Before attempting to upload your entire dataset, ingest a smaller chunk to validate your configuration and process.

Try it hands-on in the Google Colab notebook.

See it at 400 million points in the LAION-400M benchmark.