* * feat(gemini.md): add documentation for integrating Gemini embeddings with Qdrant * * refactor(integrations): move cohere.md to embeddings folder * refactor(integrations): move openai.md to embeddings folder * refactor(integrations): move autogen.md to frameworks folder * refactor(integrations): move langchain.md to frameworks folder * blacken * * feat(embedding, frameworks): reorganise integrations into embedding and frameworks, add _index.md to both * * chore(gemini.md): remove old Gemini integration documentation * * chore(embedding/_index.md): update weight from 24 to 23 and set is_empty to false * chore(frameworks/_index.md): update weight from 24 to 23 and set is_empty to false * Split integrations into embedding and frameworks * Update heading level for embedding a document * Update Gemini embedding documentation * Update titles for embedding and frameworks sections * Try again with nesting * Add documentation for integrated frameworks and embedding options * Delete integrations documentation file * Add Delimiter; unknown weights * Change all weights to 3x * Delimiter reorg * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * * docs(embedding/gemini.md): update Gemini Embedding Model API documentation * * - Add information about the new Gemini Embedding Model and its compatibility with Qdrant * - Clarify the usage of the `task_type` parameter in the API call * - Provide a list of supported task types and * * docs(embedding): update list of embedding integrations * * refactor(fifty-one.md): Rename file from embedding/fifty-one.md to frameworks/fifty-one.md * refactor(txtai.md): Rename file from embedding/txtai.md to frameworks/txtai.md * * chore(embedding): update is_empty value to true in _index.md * chore(embedding): remove Fifty One from embedding/_index.md * embedding -> embeddings --------- Co-authored-by: Atita Arora <atarora@users.noreply.github.com>
3.4 KiB
title, weight
| title | weight |
|---|---|
| DLT | 1300 |
DLT(Data Load Tool)
DLT is an open-source library that you can add to your Python scripts to load data from various and often messy data sources into well-structured, live datasets.
With the DLT-Qdrant integration, you can now select Qdrant as a DLT destination to load data into.
DLT Enables
- Automated maintenance - with schema inference, alerts and short declarative code, maintenance becomes simple.
- Run it where Python runs - on Airflow, serverless functions, notebooks. Scales on micro and large infrastructure alike.
- User-friendly, declarative interface that removes knowledge obstacles for beginners while empowering senior professionals.
Usage
To get started, install dlt with the qdrant extra.
pip install "dlt[qdrant]"
Configure the destination in the DLT secrets file. The file is located at ~/.dlt/secrets.toml by default. Add the following section to the secrets file.
[destination.qdrant.credentials]
location = "https://your-qdrant-url"
api_key = "your-qdrant-api-key"
The location will default to http://localhost:6333 and api_key is not defined - which are the defaults for a local Qdrant instance.
Find more information about DLT configurations here.
Define the source of the data.
import dlt
from dlt.destinations.qdrant import qdrant_adapter
movies = [
{
"title": "Blade Runner",
"year": 1982,
"description": "The film is about a dystopian vision of the future that combines noir elements with sci-fi imagery."
},
{
"title": "Ghost in the Shell",
"year": 1995,
"description": "The film is about a cyborg policewoman and her partner who set out to find the main culprit behind brain hacking, the Puppet Master."
},
{
"title": "The Matrix",
"year": 1999,
"description": "The movie is set in the 22nd century and tells the story of a computer hacker who joins an underground group fighting the powerful computers that rule the earth."
}
]
Define the pipeline.
pipeline = dlt.pipeline(
pipeline_name="movies",
destination="qdrant",
dataset_name="movies_dataset",
)
Run the pipeline.
info = pipeline.run(
qdrant_adapter(
movies,
embed=["title", "description"]
)
)
The data is now loaded into Qdrant.
To use vector search after the data has been loaded, you must specify which fields Qdrant needs to generate embeddings for. You do that by wrapping the data (or DLT resource) with the qdrant_adapter function.
Write disposition
A DLT write disposition defines how the data should be written to the destination. All write dispositions are supported by the Qdrant destination.
DLT Sync
Qdrant destination supports syncing of the DLT state.
Next steps
- The comprehensive Qdrant DLT destination documentation can be found here.