* * feat(gemini.md): add documentation for integrating Gemini embeddings with Qdrant * * refactor(integrations): move cohere.md to embeddings folder * refactor(integrations): move openai.md to embeddings folder * refactor(integrations): move autogen.md to frameworks folder * refactor(integrations): move langchain.md to frameworks folder * blacken * * feat(embedding, frameworks): reorganise integrations into embedding and frameworks, add _index.md to both * * chore(gemini.md): remove old Gemini integration documentation * * chore(embedding/_index.md): update weight from 24 to 23 and set is_empty to false * chore(frameworks/_index.md): update weight from 24 to 23 and set is_empty to false * Split integrations into embedding and frameworks * Update heading level for embedding a document * Update Gemini embedding documentation * Update titles for embedding and frameworks sections * Try again with nesting * Add documentation for integrated frameworks and embedding options * Delete integrations documentation file * Add Delimiter; unknown weights * Change all weights to 3x * Delimiter reorg * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * Update qdrant-landing/content/documentation/embedding/gemini.md Co-authored-by: Atita Arora <atarora@users.noreply.github.com> * * docs(embedding/gemini.md): update Gemini Embedding Model API documentation * * - Add information about the new Gemini Embedding Model and its compatibility with Qdrant * - Clarify the usage of the `task_type` parameter in the API call * - Provide a list of supported task types and * * docs(embedding): update list of embedding integrations * * refactor(fifty-one.md): Rename file from embedding/fifty-one.md to frameworks/fifty-one.md * refactor(txtai.md): Rename file from embedding/txtai.md to frameworks/txtai.md * * chore(embedding): update is_empty value to true in _index.md * chore(embedding): remove Fifty One from embedding/_index.md * embedding -> embeddings --------- Co-authored-by: Atita Arora <atarora@users.noreply.github.com>
6.6 KiB
title, weight
| title | weight |
|---|---|
| Apache Spark | 1400 |
Apache Spark
Spark is a leading distributed computing framework that empowers you to work with massive datasets efficiently. When it comes to leveraging the power of Spark for your data processing needs, the Qdrant-Spark Connector is to be considered. This connector enables Qdrant to serve as a storage destination in Spark, offering a seamless bridge between the two.
Installation
You can set up the Qdrant-Spark Connector in a few different ways, depending on your preferences and requirements.
GitHub Releases
The simplest way to get started is by downloading pre-packaged JAR file releases from the Qdrant-Spark GitHub releases page. These JAR files come with all the necessary dependencies to get you going.
Building from Source
If you prefer to build the JAR from source, you'll need JDK 17 and Maven installed on your system. Once you have the prerequisites in place, navigate to the project's root directory and run the following command:
mvn package -P assembly
This command will compile the source code and generate a fat JAR, which will be stored in the target directory by default.
Maven Central
For Java and Scala projects, you can also obtain the Qdrant-Spark Connector from Maven Central.
<dependency>
<groupId>io.qdrant</groupId>
<artifactId>spark</artifactId>
<version>1.6</version>
</dependency>
Getting Started
After successfully installing the Qdrant-Spark Connector, you can start integrating Qdrant with your Spark applications. Below, we'll walk through the basic steps of creating a Spark session with Qdrant support and loading data into Qdrant.
Creating a single-node Spark session with Qdrant Support
To begin, import the necessary libraries and create a Spark session with Qdrant support. Here's how:
from pyspark.sql import SparkSession
spark = SparkSession.builder.config(
"spark.jars",
"spark-1.0-assembly.jar", # Specify the downloaded JAR file
)
.master("local[*]")
.appName("qdrant")
.getOrCreate()
import org.apache.spark.sql.SparkSession
val spark = SparkSession.builder
.config("spark.jars", "spark-1.0-assembly.jar") // Specify the downloaded JAR file
.master("local[*]")
.appName("qdrant")
.getOrCreate()
import org.apache.spark.sql.SparkSession;
public class QdrantSparkJavaExample {
public static void main(String[] args) {
SparkSession spark = SparkSession.builder()
.config("spark.jars", "spark-1.0-assembly.jar") // Specify the downloaded JAR file
.master("local[*]")
.appName("qdrant")
.getOrCreate();
...
}
}
Loading Data into Qdrant
Here's how you can use the Qdrant-Spark Connector to upsert data:
<YourDataFrame>
.write
.format("io.qdrant.spark.Qdrant")
.option("qdrant_url", <QDRANT_URL>) # REST URL of the Qdrant instance
.option("collection_name", <QDRANT_COLLECTION_NAME>) # Name of the collection to write data into
.option("embedding_field", <EMBEDDING_FIELD_NAME>) # Name of the field holding the embeddings
.option("schema", <YourDataFrame>.schema.json()) # JSON string of the dataframe schema
.mode("append")
.save()
<YourDataFrame>
.write
.format("io.qdrant.spark.Qdrant")
.option("qdrant_url", QDRANT_URL) // REST URL of the Qdrant instance
.option("collection_name", QDRANT_COLLECTION_NAME) // Name of the collection to write data into
.option("embedding_field", EMBEDDING_FIELD_NAME) // Name of the field holding the embeddings
.option("schema", <YourDataFrame>.schema.json()) // JSON string of the dataframe schema
.mode("append")
.save()
<YourDataFrame>
.write()
.format("io.qdrant.spark.Qdrant")
.option("qdrant_url", QDRANT_URL) // REST URL of the Qdrant instance
.option("collection_name", QDRANT_COLLECTION_NAME) // Name of the collection to write data into
.option("embedding_field", EMBEDDING_FIELD_NAME) // Name of the field holding the embeddings
.option("schema", <YourDataFrame>.schema().json()) // JSON string of the dataframe schema
.mode("append")
.save();
Datatype Support
Qdrant supports all the Spark data types, and the appropriate data types are mapped based on the provided schema.
Options and Spark Types
The Qdrant-Spark Connector provides a range of options to fine-tune your data integration process. Here's a quick reference:
| Option | Description | DataType | Required |
|---|---|---|---|
qdrant_url |
REST URL of the Qdrant instance | StringType |
✅ |
collection_name |
Name of the collection to write data into | StringType |
✅ |
embedding_field |
Name of the field holding the embeddings | ArrayType(FloatType) |
✅ |
schema |
JSON string of the dataframe schema | StringType |
✅ |
mode |
Write mode of the dataframe | StringType |
✅ |
id_field |
Name of the field holding the point IDs. Default: A random UUID is generated | StringType |
❌ |
batch_size |
Max size of the upload batch. Default: 100 | IntType |
❌ |
retries |
Number of upload retries. Default: 3 | IntType |
❌ |
api_key |
Qdrant API key for authenticated requests. Default: null | StringType |
❌ |
For more information, be sure to check out the Qdrant-Spark GitHub repository. The Apache Spark guide is available here. Happy data processing!