docs: Misc updates to /frameworks/spark.md and /concepts/vectors.md (#1321)

* Update spark.md

* Update vectors.md
This commit is contained in:
Anush
2024-11-27 08:03:23 +05:30
committed by GitHub
parent 1761181d37
commit eef4cd1afb
2 changed files with 11 additions and 15 deletions
@@ -1043,7 +1043,7 @@ client
.setDistance(Distance.Cosine) .setDistance(Distance.Cosine)
.build())))) .build()))))
.setSparseVectorsConfig(SparseVectorConfig.newBuilder().putMap( .setSparseVectorsConfig(SparseVectorConfig.newBuilder().putMap(
"text", SparseVectorParams.getDefaultInstance())) "text-sparse", SparseVectorParams.getDefaultInstance()))
.build()) .build())
.get(); .get();
``` ```
@@ -13,27 +13,25 @@ You can set up the Qdrant-Spark Connector in a few different ways, depending on
### GitHub Releases ### GitHub Releases
The simplest way to get started is by downloading pre-packaged JAR file releases from the [GitHub releases page](https://github.com/qdrant/qdrant-spark/releases). These JAR files come with all the necessary dependencies. You can download the packaged JAR file from the [GitHub releases](https://github.com/qdrant/qdrant-spark/releases). It comes with all the required dependencies.
### Building from Source ### Building from Source
If you prefer to build the JAR from source, you'll need [JDK 8](https://www.azul.com/downloads/#zulu) and [Maven](https://maven.apache.org/) installed on your system. Once you have the prerequisites in place, navigate to the project's root directory and run the following command: To build the JAR from source, you'll need [JDK 8](https://www.azul.com/downloads/#zulu) and [Maven](https://maven.apache.org/) installed on your system. Once you have those in place, navigate to the project's root directory and run the following command:
```bash ```bash
mvn package mvn package -DskipTests
``` ```
This command will compile the source code and generate a fat JAR, which will be stored in the `target` directory by default. This will compile the source code and generate a fat JAR, which will be stored in the `target` directory by default.
### Maven Central ### Maven Central
For use with Java and Scala projects, the package can be found [here](https://central.sonatype.com/artifact/io.qdrant/spark). The package can be found [here](https://central.sonatype.com/artifact/io.qdrant/spark).
## Usage ## Usage
Below, we'll walk through the steps of creating a Spark session with Qdrant support and loading data into Qdrant. Below, we'll walk through the steps of creating a Spark session and ingesting data into Qdrant.
### Creating a single-node Spark session with Qdrant Support
To begin, import the necessary libraries and create a Spark session with Qdrant support: To begin, import the necessary libraries and create a Spark session with Qdrant support:
@@ -42,7 +40,7 @@ from pyspark.sql import SparkSession
spark = SparkSession.builder.config( spark = SparkSession.builder.config(
"spark.jars", "spark.jars",
"spark-VERSION.jar", # Specify the downloaded JAR file "spark-VERSION.jar", # Specify the path to the downloaded JAR file
) )
.master("local[*]") .master("local[*]")
.appName("qdrant") .appName("qdrant")
@@ -53,7 +51,7 @@ spark = SparkSession.builder.config(
import org.apache.spark.sql.SparkSession import org.apache.spark.sql.SparkSession
val spark = SparkSession.builder val spark = SparkSession.builder
.config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file .config("spark.jars", "spark-VERSION.jar") // Specify the path to the downloaded JAR file
.master("local[*]") .master("local[*]")
.appName("qdrant") .appName("qdrant")
.getOrCreate() .getOrCreate()
@@ -65,7 +63,7 @@ import org.apache.spark.sql.SparkSession;
public class QdrantSparkJavaExample { public class QdrantSparkJavaExample {
public static void main(String[] args) { public static void main(String[] args) {
SparkSession spark = SparkSession.builder() SparkSession spark = SparkSession.builder()
.config("spark.jars", "spark-VERSION.jar") // Specify the downloaded JAR file .config("spark.jars", "spark-VERSION.jar") // Specify the path to the downloaded JAR file
.master("local[*]") .master("local[*]")
.appName("qdrant") .appName("qdrant")
.getOrCreate(); .getOrCreate();
@@ -73,8 +71,6 @@ public class QdrantSparkJavaExample {
} }
``` ```
### Loading data into Qdrant
<aside role="status">Before loading the data using this connector, a collection has to be <a href="/documentation/concepts/collections/#create-a-collection">created</a> in advance with the appropriate vector dimensions and configurations.</aside> <aside role="status">Before loading the data using this connector, a collection has to be <a href="/documentation/concepts/collections/#create-a-collection">created</a> in advance with the appropriate vector dimensions and configurations.</aside>
The connector supports ingesting multiple named/unnamed, dense/sparse vectors. The connector supports ingesting multiple named/unnamed, dense/sparse vectors.
@@ -229,7 +225,7 @@ You can use the `qdrant-spark` connector as a library in [Databricks](https://ww
## Datatype Support ## Datatype Support
Qdrant supports all the Spark data types, and the appropriate data types are mapped based on the provided schema. Qdrant supports most Spark data types, and the appropriate data types are mapped based on the provided schema.
## Configuration Options ## Configuration Options