diff --git a/qdrant-landing/content/documentation/frameworks/spark.md b/qdrant-landing/content/documentation/frameworks/spark.md index bf9a75d67..f3d6984ef 100644 --- a/qdrant-landing/content/documentation/frameworks/spark.md +++ b/qdrant-landing/content/documentation/frameworks/spark.md @@ -6,7 +6,7 @@ aliases: [ ../integrations/spark/ ] # Apache Spark -[Spark](https://spark.apache.org/) is a leading distributed computing framework that empowers you to work with massive datasets efficiently. When it comes to leveraging the power of Spark for your data processing needs, the [Qdrant-Spark Connector](https://github.com/qdrant/qdrant-spark) is to be considered. This connector enables Qdrant to serve as a storage destination in Spark, offering a seamless bridge between the two. +[Apache Spark](https://spark.apache.org/) is a distributed computing framework designed for big data processing and analytics. This connector enables [Qdrant](https://qdrant.tech/) to be a storage destination in Spark. ## Installation @@ -23,19 +23,12 @@ If you prefer to build the JAR from source, you'll need [JDK 8](https://www.azul ```bash mvn package ``` + This command will compile the source code and generate a fat JAR, which will be stored in the `target` directory by default. ### Maven Central -For Java and Scala projects, you can also obtain the Qdrant-Spark Connector from [Maven Central](https://central.sonatype.com/artifact/io.qdrant/spark). - -```xml - - io.qdrant - spark - 2.0.0 - -``` +For use with Java and Scala projects, the package can be found [here](https://central.sonatype.com/artifact/io.qdrant/spark). ## Getting Started @@ -50,7 +43,7 @@ from pyspark.sql import SparkSession spark = SparkSession.builder.config( "spark.jars", - "spark-2.0.jar", # Specify the downloaded JAR file + "spark-VERSION.jar", # Specify the downloaded JAR file ) .master("local[*]") .appName("qdrant") @@ -61,7 +54,7 @@ spark = SparkSession.builder.config( import org.apache.spark.sql.SparkSession val spark = SparkSession.builder - .config("spark.jars", "spark-2.0.jar") // Specify the downloaded JAR file + .config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file .master("local[*]") .appName("qdrant") .getOrCreate() @@ -73,63 +66,159 @@ import org.apache.spark.sql.SparkSession; public class QdrantSparkJavaExample { public static void main(String[] args) { SparkSession spark = SparkSession.builder() - .config("spark.jars", "spark-2.0.jar") // Specify the downloaded JAR file + .config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file .master("local[*]") .appName("qdrant") - .getOrCreate(); - ... + .getOrCreate(); } } ``` -### Loading Data into Qdrant +### Loading data into Qdrant - + -Here's how you can use the Qdrant-Spark Connector to upsert data: +The connector supports ingesting multiple named/unnamed, dense/sparse vectors. + +
+ Unnamed/Default vector ```python - - .write - .format("io.qdrant.spark.Qdrant") - .option("qdrant_url", ) # REST URL of the Qdrant instance - .option("collection_name", ) # Name of the collection to write data into - .option("embedding_field", ) # Name of the field holding the embeddings - .option("schema", .schema.json()) # JSON string of the dataframe schema - .mode("append") - .save() + + .write + .format("io.qdrant.spark.Qdrant") + .option("qdrant_url", ) + .option("collection_name", ) + .option("embedding_field", ) # Expected to be a field of type ArrayType(FloatType) + .option("schema", .schema.json()) + .mode("append") + .save() ``` -```scala - - .write - .format("io.qdrant.spark.Qdrant") - .option("qdrant_url", QDRANT_GRPC_URL) // REST URL of the Qdrant instance - .option("collection_name", QDRANT_COLLECTION_NAME) // Name of the collection to write data into - .option("embedding_field", EMBEDDING_FIELD_NAME) // Name of the field holding the embeddings - .option("schema", .schema.json()) // JSON string of the dataframe schema - .mode("append") - .save() +
+
+ Named vector + +```python + + .write + .format("io.qdrant.spark.Qdrant") + .option("qdrant_url", ) + .option("collection_name", ) + .option("embedding_field", ) # Expected to be a field of type ArrayType(FloatType) + .option("vector_name", ) + .option("schema", .schema.json()) + .mode("append") + .save() ``` -```java - - .write() - .format("io.qdrant.spark.Qdrant") - .option("qdrant_url", QDRANT_GRPC_URL) // REST URL of the Qdrant instance - .option("collection_name", QDRANT_COLLECTION_NAME) // Name of the collection to write data into - .option("embedding_field", EMBEDDING_FIELD_NAME) // Name of the field holding the embeddings - .option("schema", .schema().json()) // JSON string of the dataframe schema - .mode("append") - .save(); +> #### NOTE +> +> The `embedding_field` and `vector_name` options are maintained for backward compatibility. It is recommended to use `vector_fields` and `vector_names` for named vectors as shown below. + +
+ +
+ Multiple named vectors + +```python + + .write + .format("io.qdrant.spark.Qdrant") + .option("qdrant_url", "") + .option("collection_name", "") + .option("vector_fields", ",") + .option("vector_names", ",") + .option("schema", .schema.json()) + .mode("append") + .save() ``` +
+ +
+ Sparse vectors + +```python + + .write + .format("io.qdrant.spark.Qdrant") + .option("qdrant_url", "") + .option("collection_name", "") + .option("sparse_vector_value_fields", "") + .option("sparse_vector_index_fields", "") + .option("sparse_vector_names", "") + .option("schema", .schema.json()) + .mode("append") + .save() +``` + +
+ +
+ Multiple sparse vectors + +```python + + .write + .format("io.qdrant.spark.Qdrant") + .option("qdrant_url", "") + .option("collection_name", "") + .option("sparse_vector_value_fields", ",") + .option("sparse_vector_index_fields", ",") + .option("sparse_vector_names", ",") + .option("schema", .schema.json()) + .mode("append") + .save() +``` + +
+ +
+ Combination of named dense and sparse vectors + +```python + + .write + .format("io.qdrant.spark.Qdrant") + .option("qdrant_url", "") + .option("collection_name", "") + .option("vector_fields", ",") + .option("vector_names", ",") + .option("sparse_vector_value_fields", ",") + .option("sparse_vector_index_fields", ",") + .option("sparse_vector_names", ",") + .option("schema", .schema.json()) + .mode("append") + .save() +``` + +
+ +
+ No vectors - Entire dataframe is stored as payload + +```python + + .write + .format("io.qdrant.spark.Qdrant") + .option("qdrant_url", "") + .option("collection_name", "") + .option("schema", .schema.json()) + .mode("append") + .save() +``` + +
+ ## Databricks + You can use the `qdrant-spark` connector as a library in [Databricks](https://www.databricks.com/) to ingest data into Qdrant. + - Go to the `Libraries` section in your cluster dashboard. - Select `Install New` to open the library installation modal. -- Search for `io.qdrant:spark:2.0.0` in the Maven packages and click `Install`. +- Search for `io.qdrant:spark:VERSION` in the Maven packages and click `Install`. ![Databricks](/documentation/frameworks/spark/databricks.png) @@ -137,21 +226,23 @@ You can use the `qdrant-spark` connector as a library in [Databricks](https://ww Qdrant supports all the Spark data types, and the appropriate data types are mapped based on the provided schema. -## Options and Spark Types - -The Qdrant-Spark Connector provides a range of options to fine-tune your data integration process. Here's a quick reference: - -| Option | Description | DataType | Required | -| :---------------- | :------------------------------------------------------------------------ | :--------------------- | :------- | -| `qdrant_url` | GRPC URL of the Qdrant instance. Eg: | `StringType` | ✅ | -| `collection_name` | Name of the collection to write data into | `StringType` | ✅ | -| `embedding_field` | Name of the field holding the embeddings | `ArrayType(FloatType)` | ✅ | -| `schema` | JSON string of the dataframe schema | `StringType` | ✅ | -| `id_field` | Name of the field holding the point IDs. Default: Generates a random UUId | `StringType` | ❌ | -| `batch_size` | Max size of the upload batch. Default: 100 | `IntType` | ❌ | -| `retries` | Number of upload retries. Default: 3 | `IntType` | ❌ | -| `api_key` | Qdrant API key to be sent in the header. Default: null | `StringType` | ❌ | -| `vector_name` | Name of the vector in the collection. Default: null | `StringType` | ❌ +## Options and Spark types 🛠️ +| Option | Description | Column DataType | Required | +| :--------------------------- | :------------------------------------------------------------------ | :---------------------------- | :------- | +| `qdrant_url` | GRPC URL of the Qdrant instance. Eg: | - | ✅ | +| `collection_name` | Name of the collection to write data into | - | ✅ | +| `schema` | JSON string of the dataframe schema | - | ✅ | +| `embedding_field` | Name of the column holding the embeddings | `ArrayType(FloatType)` | ❌ | +| `id_field` | Name of the column holding the point IDs. Default: Random UUID | `StringType` or `IntegerType` | ❌ | +| `batch_size` | Max size of the upload batch. Default: 64 | - | ❌ | +| `retries` | Number of upload retries. Default: 3 | - | ❌ | +| `api_key` | Qdrant API key for authentication | - | ❌ | +| `vector_name` | Name of the vector in the collection. | - | ❌ | +| `vector_fields` | Comma-separated names of columns holding the vectors. | `ArrayType(FloatType)` | ❌ | +| `vector_names` | Comma-separated names of vectors in the collection. | - | ❌ | +| `sparse_vector_index_fields` | Comma-separated names of columns holding the sparse vector indices. | `ArrayType(IntegerType)` | ❌ | +| `sparse_vector_value_fields` | Comma-separated names of columns holding the sparse vector values. | `ArrayType(FloatType)` | ❌ | +| `sparse_vector_names` | Comma-separated names of the sparse vectors in the collection. | - | ❌ | For more information, be sure to check out the [Qdrant-Spark GitHub repository](https://github.com/qdrant/qdrant-spark). The Apache Spark guide is available [here](https://spark.apache.org/docs/latest/quick-start.html). Happy data processing!