Merge branch 'master' into bq-updates
@@ -17,12 +17,12 @@ jobs:
|
||||
with:
|
||||
hugo-version: "latest"
|
||||
- name: Run hugo
|
||||
run: cd qdrant-landing && hugo -b 'https://qdrant.tech/'
|
||||
run: cd qdrant-landing && hugo -b ''
|
||||
- name: Link Checker
|
||||
id: lychee
|
||||
uses: lycheeverse/lychee-action@v1.8.0
|
||||
with:
|
||||
args: --offline qdrant-landing/public
|
||||
args: --offline --base qdrant-landing/public qdrant-landing/public
|
||||
fail: true
|
||||
env:
|
||||
GITHUB_TOKEN: ${{secrets.GITHUB_TOKEN}}
|
||||
|
||||
@@ -2,6 +2,6 @@ https://qdrant.to/twitter
|
||||
https://fonts.gstatic.com/
|
||||
https://twitter.com/intent.*
|
||||
https://www.linkedin.com/sharing.*
|
||||
file://.*
|
||||
http://localhost.*
|
||||
https://fonts.googleapis.com/
|
||||
admin/emails/*
|
||||
@@ -114,7 +114,7 @@ In the open-source world, you pay for the resources you use, not the number of d
|
||||
Resources depend more on the optimal solution for each use case.
|
||||
As a result, running a dedicated vector search engine can be even cheaper, as it allows optimization specifically for vector search use cases.
|
||||
|
||||
For instance, Qdrant implements a number of [quantization techniques](documentation/guides/quantization/) that can significantly reduce the memory footprint of embeddings.
|
||||
For instance, Qdrant implements a number of [quantization techniques](/documentation/guides/quantization/) that can significantly reduce the memory footprint of embeddings.
|
||||
|
||||
In terms of data transfer costs, on most cloud providers, network use within a region is usually free. As long as you put the original source data and the vector store in the same region, there are no added data transfer costs.
|
||||
|
||||
|
||||
@@ -199,8 +199,8 @@ Here are some terms that are added: "Berlin", and "founder" - despite having no
|
||||
|
||||
If you're interested in using the higher-performance approach, check out the following models:
|
||||
|
||||
1. [naver/efficient-splade-VI-BT-large-doc](huggingface.co/naver/efficient-splade-vi-bt-large-doc)
|
||||
2. [naver/efficient-splade-VI-BT-large-query](huggingface.co/naver/efficient-splade-vi-bt-large-doc)
|
||||
1. [naver/efficient-splade-VI-BT-large-doc](https://huggingface.co/naver/efficient-splade-vi-bt-large-doc)
|
||||
2. [naver/efficient-splade-VI-BT-large-query](https://huggingface.co/naver/efficient-splade-vi-bt-large-doc)
|
||||
|
||||
## Why SPLADE works? Term Expansion
|
||||
|
||||
|
||||
@@ -1,5 +1,4 @@
|
||||
---
|
||||
title: Qdrant Blog
|
||||
subtitle: Check out our latest posts
|
||||
sitemapExclude: True
|
||||
---
|
||||
@@ -38,7 +38,7 @@ more time shipping features and fixing bugs.
|
||||
bloop’s mission is to make software engineers autonomous and semantic code search is the cornerstone
|
||||
of that vision. The project is maintained by a group of Rust and Typescript engineers and ML researchers.
|
||||
It leverages many prominent nascent technologies, such as [Tauri](http://tauri.app), [tantivy](https://docs.rs/tantivy),
|
||||
[Qdrant](http://qdrant.tech) and [Anthropic](https://www.anthropic.com/).
|
||||
[Qdrant](https://qdrant.tech) and [Anthropic](https://www.anthropic.com/).
|
||||
|
||||
## About Qdrant
|
||||
|
||||
|
||||
@@ -0,0 +1,10 @@
|
||||
---
|
||||
draft: false
|
||||
title: Semantic code search
|
||||
short_description: Searching over Qdrant source code using semantic search
|
||||
description: It can be difficult to go through an unknown codebase. This demo shows how to implement a semantic search application for code search tasks, with two neural encoders. These encoders are a general-purpose sentence transformer and a code-specific model. This supports both natural and code-like queries, which covers a broad range of interactions.
|
||||
preview_image: /demo/code-search.png
|
||||
link: https://code-search.qdrant.tech/
|
||||
weight: 4
|
||||
sitemapExclude: True
|
||||
---
|
||||
@@ -39,4 +39,4 @@ Now that you have signed up via AWS Marketplace, please read our instructions to
|
||||
|
||||
2. Learn how to [authenticate and access your cluster](../../cloud/authentication/).
|
||||
|
||||
3. Additional open source [documentation](../../troubleshooting/).
|
||||
3. Additional open source [documentation](/documentation/guides/common-errors/).
|
||||
|
||||
@@ -0,0 +1,98 @@
|
||||
---
|
||||
title: Mistral
|
||||
weight: 700
|
||||
---
|
||||
|
||||
| Time: 10 min | Level: Beginner | [](https://githubtocolab.com/qdrant/examples/blob/mistral-getting-started/mistral-embed-getting-started/mistral_qdrant_getting_started.ipynb) |
|
||||
| --- | ----------- | ----------- |
|
||||
|
||||
# Mistral
|
||||
Qdrant is compatible with the new released Mistral Embed and its official Python SDK that can be installed as any other package:
|
||||
|
||||
## Setup
|
||||
|
||||
### Install the client
|
||||
|
||||
```bash
|
||||
pip install mistralai
|
||||
```
|
||||
|
||||
And then we set this up:
|
||||
|
||||
```python
|
||||
from mistralai.client import MistralClient
|
||||
from qdrant_client import QdrantClient
|
||||
from qdrant_client.http.models import PointStruct, VectorParams, Distance
|
||||
collection_name = "example_collection"
|
||||
|
||||
MISTRAL_API_KEY = "your_mistral_api_key"
|
||||
search_client = QdrantClient(":memory:")
|
||||
mistral_client = MistralClient(api_key=MISTRAL_API_KEY)
|
||||
texts = [
|
||||
"Qdrant is the best vector search engine!",
|
||||
"Loved by Enterprises and everyone building for low latency, high performance, and scale.",
|
||||
]
|
||||
```
|
||||
|
||||
Let's see how to use the Embedding Model API to embed a document for retrieval.
|
||||
|
||||
The following example shows how to embed a document with the `models/embedding-001` with the `retrieval_document` task type:
|
||||
|
||||
## Embedding a document
|
||||
|
||||
```python
|
||||
result = mistral_client.embeddings(
|
||||
model="mistral-embed",
|
||||
input=texts,
|
||||
)
|
||||
```
|
||||
|
||||
The returned result has a data field with a key: `embedding`. The value of this key is a list of floats representing the embedding of the document.
|
||||
|
||||
### Converting this into Qdrant Points
|
||||
|
||||
```python
|
||||
points = [
|
||||
PointStruct(
|
||||
id=idx,
|
||||
vector=response.embedding,
|
||||
payload={"text": text},
|
||||
)
|
||||
for idx, (response, text) in enumerate(zip(result.data, texts))
|
||||
]
|
||||
```
|
||||
|
||||
## Create a collection and Insert the documents
|
||||
|
||||
```python
|
||||
search_client.create_collection(collection_name, vectors_config=
|
||||
VectorParams(
|
||||
size=1024,
|
||||
distance=Distance.COSINE,
|
||||
)
|
||||
)
|
||||
search_client.upsert(collection_name, points)
|
||||
```
|
||||
|
||||
## Searching for documents with Qdrant
|
||||
|
||||
Once the documents are indexed, you can search for the most relevant documents using the same model with the `retrieval_query` task type:
|
||||
|
||||
```python
|
||||
search_client.search(
|
||||
collection_name=collection_name,
|
||||
query_vector=mistral_client.embeddings(
|
||||
model="mistral-embed", input=["What is the best to use for vector search scaling?"]
|
||||
).data[0].embedding,
|
||||
)
|
||||
```
|
||||
|
||||
## Using Mistral Embedding Models with Binary Quantization
|
||||
|
||||
You can use Mistral Embedding Models with [Binary Quantization](/articles/binary-quantization/) - a technique that allows you to reduce the size of the embeddings by 32 times without losing the quality of the search results too much.
|
||||
|
||||
At an oversampling of 3 and a limit of 100, we've a 95% recall against the exact nearest neighbors with rescore enabled.
|
||||
|
||||

|
||||
|
||||
That's it! You can now use Mistral Embedding Models with Qdrant!
|
||||
@@ -0,0 +1,28 @@
|
||||
---
|
||||
title: DocsGPT
|
||||
weight: 2600
|
||||
---
|
||||
|
||||
# DocsGPT
|
||||
|
||||
[DocsGPT](https://docsgpt.arc53.com/) is an open-source documentation assistant that enables you to build conversational user experiences on top of your data.
|
||||
|
||||
Qdrant is supported as a vectorstore in DocsGPT to ingest and semantically retrieve documents.
|
||||
|
||||
## Configuration
|
||||
|
||||
Learn how to setup DocsGPT in their [Quickstart guide](https://docs.docsgpt.co.uk/Deploying/Quickstart).
|
||||
|
||||
You can configure DocsGPT with environment variables in a `.env` file.
|
||||
|
||||
To configure DocsGPT to use Qdrant as the vector store, set `VECTOR_STORE` to `"qdrant"`.
|
||||
|
||||
```bash
|
||||
echo "VECTOR_STORE=qdrant" >> .env
|
||||
```
|
||||
|
||||
DocsGPT includes a list of the Qdrant configuration options that you can set as environment variables [here](https://github.com/arc53/DocsGPT/blob/00dfb07b15602319bddb95089e3dab05fac56240/application/core/settings.py#L46-L59).
|
||||
|
||||
## Further reading
|
||||
|
||||
- [DocsGPT Reference](https://github.com/arc53/DocsGPT)
|
||||
@@ -21,7 +21,7 @@ learn about one of the most popular and fastest growing vector databases in the
|
||||
|
||||
## What is Qdrant?
|
||||
|
||||
[Qdrant](http://qdrant.tech) "is a vector similarity search engine that provides a production-ready
|
||||
[Qdrant](https://qdrant.tech) "is a vector similarity search engine that provides a production-ready
|
||||
service with a convenient API to store, search, and manage points (i.e. vectors) with an additional
|
||||
payload." You can think of the payloads as additional pieces of information that can help you
|
||||
hone in on your search and also receive useful information that you can give to your users.
|
||||
|
||||
@@ -67,7 +67,7 @@ There are also various community-driven projects aimed to provide the support fo
|
||||
maintained, thus not mentioned here. However, it is still possible to interact with both engines through the HTTP REST or gRPC API.
|
||||
That makes it easy to integrate with any technology of your choice.
|
||||
|
||||
If you are a Python user, then both tools are well-integrated with the most popular libraries like [LangChain](../integrations/langchain/), [LlamaIndex](../integrations/llama-index/), [Haystack](../integrations/haystack/), and more.
|
||||
If you are a Python user, then both tools are well-integrated with the most popular libraries like [LangChain](/documentation/frameworks/langchain/), [LlamaIndex](/documentation/frameworks/llama-index/), [Haystack](/documentation/frameworks/haystack/), and more.
|
||||
Using any of those libraries makes it easier to experiment with different vector databases, as the transition should be seamless.
|
||||
|
||||
## Planning to migrate?
|
||||
@@ -92,6 +92,6 @@ Migrating from Pinecone to Qdrant involves a series of well-planned steps to ens
|
||||
|
||||
1. If you aren't ready yet, [try out Qdrant locally](/documentation/quick-start/) or sign up for [Qdrant Cloud](https://cloud.qdrant.io/).
|
||||
|
||||
2. For more basic information on Qdrant read our [Overview](overview/) section or learn more about Qdrant Cloud's [Free Tier](documentation/cloud/).
|
||||
2. For more basic information on Qdrant read our [Overview](/documentation/overview/) section or learn more about Qdrant Cloud's [Free Tier](/documentation/cloud/).
|
||||
|
||||
3. If ready to migrate, please consult our [Comprehensive Guide](https://github.com/NirantK/qdrant_tools) for further details on migration steps.
|
||||
|
||||
@@ -12,18 +12,19 @@ aliases:
|
||||
|
||||
These tutorials demonstrate different ways you can build vector search into your applications.
|
||||
|
||||
| Tutorial | Description | Stack |
|
||||
|------------------------------------------------------------------------|-------------------------------------------------------------------|----------------------------|
|
||||
| [Configure Optimal Use](../tutorials/optimize/) | Configure Qdrant collections for best resource use. | Qdrant |
|
||||
| [Separate Partitions](../tutorials/multiple-partitions/) | Serve vectors for many independent users. | Qdrant |
|
||||
| [Bulk Upload Vectors](../tutorials/bulk-upload/) | Upload a large scale dataset. | Qdrant |
|
||||
| [Create Dataset Snapshots](../tutorials/create-snapshot/) | Turn a dataset into a snapshot by exporting it from a collection. | Qdrant |
|
||||
| [Semantic Search for Beginners](../tutorials/search-beginners/) | Create a simple search engine locally in minutes. | Qdrant |
|
||||
| [Simple Neural Search](../tutorials/neural-search/) | Build and deploy a neural search that browses startup data. | Qdrant, BERT, FastAPI |
|
||||
| [Aleph Alpha Search](../tutorials/aleph-alpha-search/) | Build a multimodal search that combines text and image data. | Qdrant, Aleph Alpha |
|
||||
| [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
||||
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
||||
| [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
||||
| [HuggingFace datasets](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||
| [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
|
||||
| Tutorial | Description | Stack |
|
||||
|----------------------------------------------------------------------------|-------------------------------------------------------------------|---------------------------------------------|
|
||||
| [Configure Optimal Use](../tutorials/optimize/) | Configure Qdrant collections for best resource use. | Qdrant |
|
||||
| [Separate Partitions](../tutorials/multiple-partitions/) | Serve vectors for many independent users. | Qdrant |
|
||||
| [Bulk Upload Vectors](../tutorials/bulk-upload/) | Upload a large scale dataset. | Qdrant |
|
||||
| [Create Dataset Snapshots](../tutorials/create-snapshot/) | Turn a dataset into a snapshot by exporting it from a collection. | Qdrant |
|
||||
| [Semantic Search for Beginners](../tutorials/search-beginners/) | Create a simple search engine locally in minutes. | Qdrant |
|
||||
| [Simple Neural Search](../tutorials/neural-search/) | Build and deploy a neural search that browses startup data. | Qdrant, BERT, FastAPI |
|
||||
| [Aleph Alpha Search](../tutorials/aleph-alpha-search/) | Build a multimodal search that combines text and image data. | Qdrant, Aleph Alpha |
|
||||
| [Mighty Semantic Search](../tutorials/mighty/) | Build a simple semantic search with an on-demand NLP service. | Qdrant, Mighty |
|
||||
| [Asynchronous API](../tutorials/async-api/) | Communicate with Qdrant server asynchronously with Python SDK. | Qdrant, Python |
|
||||
| [Multitenancy with LlamaIndex](../tutorials/llama-index-multitenancy/) | Handle data coming from multiple users in LlamaIndex. | Qdrant, Python, LlamaIndex |
|
||||
| [HuggingFace datasets](../tutorials/huggingface-datasets/) | Load a Hugging Face dataset to Qdrant | Qdrant, Python, datasets |
|
||||
| [Measure retrieval quality](../tutorials/retrieval-quality/) | Measure and fine-tune the retrieval quality | Qdrant, Python, datasets |
|
||||
| [Use semantic search to navigate your codebase](../tutorials/code-search/) | Implement semantic search application for code search task | Qdrant, Python, sentence-transformers, Jina |
|
||||
| [Troubleshooting](../tutorials/common-errors/) | Solutions to common errors and fixes | Qdrant |
|
||||
|
||||
@@ -0,0 +1,444 @@
|
||||
---
|
||||
title: Semantic code search
|
||||
weight: 22
|
||||
---
|
||||
|
||||
# Use semantic search to navigate your codebase
|
||||
|
||||
| Time: 45 min | Level: Intermediate | [](https://colab.research.google.com/github/qdrant/examples/blob/master/code-search/code-search.ipynb) | |
|
||||
|--------------|---------------------|--|----|
|
||||
|
||||
You too can enrich your applications with Qdrant semantic search. In this
|
||||
tutorial, we describe how you can use Qdrant to navigate a codebase, to help
|
||||
you find relevant code snippets. As an example, we will use the [Qdrant](https://github.com/qdrant/qdrant)
|
||||
source code itself, which is mostly written in Rust.
|
||||
|
||||
<aside role="status">This tutorial might not work on code bases that are not disciplined or structured. For good code search, you may need to refactor the project first.</aside>
|
||||
|
||||
## The approach
|
||||
|
||||
We want to search codebases using natural semantic queries, and searching for
|
||||
code based on similar logic. You can set up these tasks with embeddings:
|
||||
|
||||
1. General usage neural encoder for natural-like queries, in our case `all-MiniLM-L6-v2`
|
||||
from the
|
||||
[sentence-transformers](https://www.sbert.net/docs/pretrained_models.html) library.
|
||||
2. Specialized embeddings for code-to-code similarity search. We use the
|
||||
`jina-embeddings-v2-base-code` model.
|
||||
|
||||
To prepare our code for `all-MiniLM-L6-v2`, we preprocess the code to text that
|
||||
more closely resembles natural language. The Jina embeddings model supports a
|
||||
variety of standard programming languages, so there is no need to preprocess the
|
||||
snippets. We can use the code as is.
|
||||
|
||||
## Data preparation
|
||||
|
||||
Chunking the application sources into smaller parts is a non-trivial task. In
|
||||
general, functions, class methods, structs, enums, and all the other language-specific
|
||||
constructs are good candidates for chunks. They are big enough to
|
||||
contain some meaningful information, but small enough to be processed by
|
||||
embedding models with a limited context window. You can also use docstrings,
|
||||
comments, and other metadata can be used to enrich the chunks with additional
|
||||
information.
|
||||
|
||||

|
||||
|
||||
### Parsing the codebase
|
||||
|
||||
While our example uses Rust, you can use our approach with any other language.
|
||||
You can parse code with a [Language Server Protocol](https://microsoft.github.io/language-server-protocol/) (**LSP**)
|
||||
compatible tool. You can use an LSP to build a graph of the codebase, and then extract chunks.
|
||||
We did our work with the [rust-analyzer](https://rust-analyzer.github.io/).
|
||||
We exported the parsed codebase into the [LSIF](https://microsoft.github.io/language-server-protocol/specifications/lsif/0.4.0/specification/)
|
||||
format, a standard for code intelligence data. Next, we used the LSIF data to
|
||||
navigate the codebase and extract the chunks. For details, see our [code search
|
||||
demo](https://github.com/qdrant/demo-code-search).
|
||||
|
||||
<aside role="status">
|
||||
For other languages, you can use the same approach. There are
|
||||
<a href="https://microsoft.github.io/language-server-protocol/implementors/servers/">plenty of implementations available
|
||||
</a>.
|
||||
</aside>
|
||||
|
||||
We then exported the chunks into JSON documents with not only the code itself,
|
||||
but also context with the location of the code in the project. For example, see
|
||||
the description of the `await_ready_for_timeout` function from the `IsReady`
|
||||
struct in the `common` module:
|
||||
|
||||
```json
|
||||
{
|
||||
"name":"await_ready_for_timeout",
|
||||
"signature":"fn await_ready_for_timeout (& self , timeout : Duration) -> bool",
|
||||
"code_type":"Function",
|
||||
"docstring":"= \" Return `true` if ready, `false` if timed out.\"",
|
||||
"line":44,
|
||||
"line_from":43,
|
||||
"line_to":51,
|
||||
"context":{
|
||||
"module":"common",
|
||||
"file_path":"lib/collection/src/common/is_ready.rs",
|
||||
"file_name":"is_ready.rs",
|
||||
"struct_name":"IsReady",
|
||||
"snippet":" /// Return `true` if ready, `false` if timed out.\n pub fn await_ready_for_timeout(&self, timeout: Duration) -> bool {\n let mut is_ready = self.value.lock();\n if !*is_ready {\n !self.condvar.wait_for(&mut is_ready, timeout).timed_out()\n } else {\n true\n }\n }\n"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
You can examine the Qdrant structures, parsed in JSON, in the [`structures.jsonl`
|
||||
file](https://storage.googleapis.com/tutorial-attachments/code-search/structures.jsonl)
|
||||
in our Google Cloud Storage bucket. Download it and use it as a source of data for our code search.
|
||||
|
||||
```shell
|
||||
wget https://storage.googleapis.com/tutorial-attachments/code-search/structures.jsonl
|
||||
```
|
||||
|
||||
Next, load the file and parse the lines into a list of dictionaries:
|
||||
|
||||
```python
|
||||
import json
|
||||
|
||||
structures = []
|
||||
with open("structures.jsonl", "r") as fp:
|
||||
for i, row in enumerate(fp):
|
||||
entry = json.loads(row)
|
||||
structures.append(entry)
|
||||
```
|
||||
|
||||
### Code to *natural language* conversion
|
||||
|
||||
Each programming language has its own syntax which is not a part of the natural
|
||||
language. Thus, a general-purpose model probably does not understand the code
|
||||
as is. We can, however, normalize the data by removing code specifics and
|
||||
including additional context, such as module, class, function, and file name.
|
||||
We took the following steps:
|
||||
|
||||
1. Extract the signature of the function, method, or other code construct.
|
||||
2. Divide camel case and snake case names into separate words.
|
||||
3. Take the docstring, comments, and other important metadata.
|
||||
4. Build a sentence from the extracted data using a predefined template.
|
||||
5. Remove the special characters and replace them with spaces.
|
||||
|
||||
As input, expect dictionaries with the same structure. Define a `textify`
|
||||
function to do the conversion. We'll use an `inflection` library to convert
|
||||
with different naming conventions.
|
||||
|
||||
```shell
|
||||
pip install inflection
|
||||
```
|
||||
|
||||
Once all dependencies are installed, we define the `textify` function:
|
||||
|
||||
```python
|
||||
import inflection
|
||||
import re
|
||||
|
||||
from typing import Dict, Any
|
||||
|
||||
def textify(chunk: Dict[str, Any]) -> str:
|
||||
# Get rid of all the camel case / snake case
|
||||
# - inflection.underscore changes the camel case to snake case
|
||||
# - inflection.humanize converts the snake case to human readable form
|
||||
name = inflection.humanize(inflection.underscore(chunk["name"]))
|
||||
signature = inflection.humanize(inflection.underscore(chunk["signature"]))
|
||||
|
||||
# Check if docstring is provided
|
||||
docstring = ""
|
||||
if chunk["docstring"]:
|
||||
docstring = f"that does {chunk['docstring']} "
|
||||
|
||||
# Extract the location of that snippet of code
|
||||
context = (
|
||||
f"module {chunk['context']['module']} "
|
||||
f"file {chunk['context']['file_name']}"
|
||||
)
|
||||
if chunk["context"]["struct_name"]:
|
||||
struct_name = inflection.humanize(
|
||||
inflection.underscore(chunk["context"]["struct_name"])
|
||||
)
|
||||
context = f"defined in struct {struct_name} {context}"
|
||||
|
||||
# Combine all the bits and pieces together
|
||||
text_representation = (
|
||||
f"{chunk['code_type']} {name} "
|
||||
f"{docstring}"
|
||||
f"defined as {signature} "
|
||||
f"{context}"
|
||||
)
|
||||
|
||||
# Remove any special characters and concatenate the tokens
|
||||
tokens = re.split(r"\W", text_representation)
|
||||
tokens = filter(lambda x: x, tokens)
|
||||
return " ".join(tokens)
|
||||
```
|
||||
|
||||
Now we can use `textify` to convert all chunks into text representations:
|
||||
|
||||
```python
|
||||
text_representations = list(map(textify, structures))
|
||||
```
|
||||
|
||||
This is how the `await_ready_for_timeout` function description appears:
|
||||
|
||||
```text
|
||||
Function Await ready for timeout that does Return true if ready false if timed out defined as Fn await ready for timeout self timeout duration bool defined in struct Is ready module common file is_ready rs
|
||||
```
|
||||
|
||||
## Ingestion pipeline
|
||||
|
||||
Next, we build the code search engine to vectorizing data and set up a semantic
|
||||
search mechanism for both embedding models.
|
||||
|
||||
### Natural language embeddings
|
||||
|
||||
We can encode text representations through the `all-MiniLM-L6-v2` model from
|
||||
`sentence-transformers`. With the following command, we install `sentence-transformers`
|
||||
with dependencies:
|
||||
|
||||
```shell
|
||||
pip install sentence-transformers optimum onnx
|
||||
```
|
||||
|
||||
Then we can use the model to encode the text representations:
|
||||
|
||||
```python
|
||||
from sentence_transformers import SentenceTransformer
|
||||
|
||||
nlp_model = SentenceTransformer("all-MiniLM-L6-v2")
|
||||
nlp_embeddings = nlp_model.encode(
|
||||
text_representations, show_progress_bar=True,
|
||||
)
|
||||
```
|
||||
|
||||
### Code embeddings
|
||||
|
||||
The `jina-embeddings-v2-base-code` model is a good candidate for this task.
|
||||
You can also get it from the `sentence-transformers` library, with conditions.
|
||||
Visit [the model page](https://huggingface.co/jinaai/jina-embeddings-v2-base-code),
|
||||
accept the rules, and generate the access token in your [account settings](https://huggingface.co/settings/tokens).
|
||||
Once you have the token, you can use the model as follows:
|
||||
|
||||
```python
|
||||
HF_TOKEN = "THIS_IS_YOUR_TOKEN"
|
||||
|
||||
# Extract the code snippets from the structures to a separate list
|
||||
code_snippets = [
|
||||
structure["context"]["snippet"] for structure in structures
|
||||
]
|
||||
|
||||
code_model = SentenceTransformer(
|
||||
"jinaai/jina-embeddings-v2-base-code",
|
||||
token=HF_TOKEN,
|
||||
trust_remote_code=True
|
||||
)
|
||||
code_model.max_seq_length = 8192 # increase the context length window
|
||||
code_embeddings = code_model.encode(
|
||||
code_snippets, batch_size=4, show_progress_bar=True,
|
||||
)
|
||||
```
|
||||
|
||||
Remember to set the `trust_remote_code` parameter to `True`. Otherwise, the
|
||||
model does not produce meaningful vectors. Setting this parameter allows the
|
||||
library to download and possibly launch some code on your machine, so be sure
|
||||
to trust the source.
|
||||
|
||||
With both the natural language and code embeddings, we can build and store them
|
||||
in the Qdrant collection.
|
||||
|
||||
### Building Qdrant collection
|
||||
|
||||
We use the `qdrant-client` library to interact with the Qdrant server. Let's
|
||||
install that client:
|
||||
|
||||
```shell
|
||||
pip install qdrant-client
|
||||
```
|
||||
|
||||
Of course, we need a running Qdrant server for vector search. If you need one,
|
||||
you can [use a local Docker container](https://qdrant.tech/documentation/quick-start/)
|
||||
or deploy it using the [Qdrant Cloud](https://cloud.qdrant.io/).
|
||||
You can use either to follow this tutorial. Configure the connection parameters:
|
||||
|
||||
```python
|
||||
QDRANT_URL = "https://my-cluster.cloud.qdrant.io:6333" # http://localhost:6333 for local instance
|
||||
QDRANT_API_KEY = "THIS_IS_YOUR_API_KEY" # None for local instance
|
||||
```
|
||||
|
||||
Then use the library to create a collection:
|
||||
|
||||
```python
|
||||
from qdrant_client import QdrantClient, models
|
||||
|
||||
client = QdrantClient(QDRANT_URL, api_key=QDRANT_API_KEY)
|
||||
client.create_collection(
|
||||
"qdrant-sources",
|
||||
vectors_config={
|
||||
"text": models.VectorParams(
|
||||
size=nlp_embeddings.shape[1],
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
"code": models.VectorParams(
|
||||
size=code_embeddings.shape[1],
|
||||
distance=models.Distance.COSINE,
|
||||
),
|
||||
}
|
||||
)
|
||||
```
|
||||
|
||||
Our newly created collection is ready to accept the data. Let's upload the embeddings:
|
||||
|
||||
```python
|
||||
import uuid
|
||||
|
||||
points = [
|
||||
models.PointStruct(
|
||||
id=uuid.uuid4().hex,
|
||||
vector={
|
||||
"text": text_embedding,
|
||||
"code": code_embedding,
|
||||
},
|
||||
payload=structure,
|
||||
)
|
||||
for text_embedding, code_embedding, structure in zip(nlp_embeddings, code_embeddings, structures)
|
||||
]
|
||||
|
||||
client.upload_points("qdrant-sources", points=points, batch_size=64)
|
||||
```
|
||||
|
||||
The uploaded points are immediately available for search. Next, query the
|
||||
collection to find relevant code snippets.
|
||||
|
||||
## Querying the codebase
|
||||
|
||||
We use one of the models to search the collection. Start with text embeddings.
|
||||
Run the following query "*How do I count points in a collection?*". Review the
|
||||
results.
|
||||
|
||||
```python
|
||||
query = "How do I count points in a collection?"
|
||||
|
||||
hits = client.search(
|
||||
"qdrant-sources",
|
||||
query_vector=(
|
||||
"text", nlp_model.encode(query).tolist()
|
||||
),
|
||||
limit=5,
|
||||
)
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
| module | file_name | score | signature |
|
||||
|--------------------|---------------------|------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| toc | point_ops.rs | 0.59448624 | ` async fn count (& self , collection_name : & str , request : CountRequestInternal , read_consistency : Option < ReadConsistency > , shard_selection : ShardSelectorInternal ,) -> Result < CountResult , StorageError > ` |
|
||||
| operations | types.rs | 0.5493385 | ` # [doc = " Count Request"] # [doc = " Counts the number of points which satisfy the given filter."] # [doc = " If filter is not provided, the count of all points in the collection will be returned."] # [derive (Debug , Deserialize , Serialize , JsonSchema , Validate)] # [serde (rename_all = "snake_case")] pub struct CountRequestInternal { # [doc = " Look only for points which satisfies this conditions"] # [validate] pub filter : Option < Filter > , # [doc = " If true, count exact number of points. If false, count approximate number of points faster."] # [doc = " Approximate count might be unreliable during the indexing process. Default: true"] # [serde (default = "default_exact_count")] pub exact : bool , } ` |
|
||||
| collection_manager | segments_updater.rs | 0.5121002 | ` fn upsert_points < 'a , T > (segments : & SegmentHolder , op_num : SeqNumberType , points : T ,) -> CollectionResult < usize > where T : IntoIterator < Item = & 'a PointStruct > , ` |
|
||||
| collection | point_ops.rs | 0.5063539 | ` async fn count (& self , request : CountRequestInternal , read_consistency : Option < ReadConsistency > , shard_selection : & ShardSelectorInternal ,) -> CollectionResult < CountResult > ` |
|
||||
| map_index | mod.rs | 0.49973983 | ` fn get_points_with_value_count < Q > (& self , value : & Q) -> Option < usize > where Q : ? Sized , N : std :: borrow :: Borrow < Q > , Q : Hash + Eq , ` |
|
||||
|
||||
It seems we were able to find some relevant code structures. Let's try the same with the code embeddings:
|
||||
|
||||
```python
|
||||
hits = client.search(
|
||||
"qdrant-sources",
|
||||
query_vector=(
|
||||
"code", code_model.encode(query).tolist()
|
||||
),
|
||||
limit=5,
|
||||
)
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
| module | file_name | score | signature |
|
||||
|---------------|----------------------------|------------|-----------------------------------------------|
|
||||
| field_index | geo_index.rs | 0.73278356 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| numeric_index | mod.rs | 0.7254976 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| map_index | mod.rs | 0.7124739 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| map_index | mod.rs | 0.7124739 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| fixtures | payload_context_fixture.rs | 0.706204 | ` fn total_point_count (& self) -> usize ` |
|
||||
|
||||
While the scores retrieved by different models are not comparable, but we can
|
||||
see that the results are different. Code and text embeddings can capture
|
||||
different aspects of the codebase. We can use both models to query the collection
|
||||
and then combine the results to get the most relevant code snippets, from a single batch request.
|
||||
|
||||
```python
|
||||
results = client.search_batch(
|
||||
"qdrant-sources",
|
||||
requests=[
|
||||
models.SearchRequest(
|
||||
vector=models.NamedVector(
|
||||
name="text",
|
||||
vector=nlp_model.encode(query).tolist()
|
||||
),
|
||||
with_payload=True,
|
||||
limit=5,
|
||||
),
|
||||
models.SearchRequest(
|
||||
vector=models.NamedVector(
|
||||
name="code",
|
||||
vector=code_model.encode(query).tolist()
|
||||
),
|
||||
with_payload=True,
|
||||
limit=5,
|
||||
),
|
||||
]
|
||||
)
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
| module | file_name | score | signature |
|
||||
|--------------------|----------------------------|------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
|
||||
| toc | point_ops.rs | 0.59448624 | ` async fn count (& self , collection_name : & str , request : CountRequestInternal , read_consistency : Option < ReadConsistency > , shard_selection : ShardSelectorInternal ,) -> Result < CountResult , StorageError > ` |
|
||||
| operations | types.rs | 0.5493385 | ` # [doc = " Count Request"] # [doc = " Counts the number of points which satisfy the given filter."] # [doc = " If filter is not provided, the count of all points in the collection will be returned."] # [derive (Debug , Deserialize , Serialize , JsonSchema , Validate)] # [serde (rename_all = "snake_case")] pub struct CountRequestInternal { # [doc = " Look only for points which satisfies this conditions"] # [validate] pub filter : Option < Filter > , # [doc = " If true, count exact number of points. If false, count approximate number of points faster."] # [doc = " Approximate count might be unreliable during the indexing process. Default: true"] # [serde (default = "default_exact_count")] pub exact : bool , } ` |
|
||||
| collection_manager | segments_updater.rs | 0.5121002 | ` fn upsert_points < 'a , T > (segments : & SegmentHolder , op_num : SeqNumberType , points : T ,) -> CollectionResult < usize > where T : IntoIterator < Item = & 'a PointStruct > , ` |
|
||||
| collection | point_ops.rs | 0.5063539 | ` async fn count (& self , request : CountRequestInternal , read_consistency : Option < ReadConsistency > , shard_selection : & ShardSelectorInternal ,) -> CollectionResult < CountResult > ` |
|
||||
| map_index | mod.rs | 0.49973983 | ` fn get_points_with_value_count < Q > (& self , value : & Q) -> Option < usize > where Q : ? Sized , N : std :: borrow :: Borrow < Q > , Q : Hash + Eq , ` |
|
||||
| field_index | geo_index.rs | 0.73278356 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| numeric_index | mod.rs | 0.7254976 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| map_index | mod.rs | 0.7124739 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| map_index | mod.rs | 0.7124739 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| fixtures | payload_context_fixture.rs | 0.706204 | ` fn total_point_count (& self) -> usize ` |
|
||||
|
||||
|
||||
This is one example of how you can use different models and combine the results.
|
||||
In a real-world scenario, you might run some reranking and deduplication, as
|
||||
well as additional processing of the results.
|
||||
|
||||
### Grouping the results
|
||||
|
||||
You can improve the search results, by grouping them by payload properties.
|
||||
In our case, we can group the results by the module. If we use code embeddings,
|
||||
we can see multiple results from the `map_index` module. Let's group the
|
||||
results and assume a single result per module:
|
||||
|
||||
```python
|
||||
results = client.search_groups(
|
||||
"qdrant-sources",
|
||||
query_vector=(
|
||||
"code", code_model.encode(query).tolist()
|
||||
),
|
||||
group_by="context.module",
|
||||
limit=5,
|
||||
group_size=1,
|
||||
)
|
||||
```
|
||||
|
||||
Output:
|
||||
|
||||
| module | file_name | score | signature |
|
||||
|---------------|----------------------------|------------|-----------------------------------------------|
|
||||
| field_index | geo_index.rs | 0.73278356 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| numeric_index | mod.rs | 0.7254976 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| map_index | mod.rs | 0.7124739 | ` fn count_indexed_points (& self) -> usize ` |
|
||||
| fixtures | payload_context_fixture.rs | 0.706204 | ` fn total_point_count (& self) -> usize ` |
|
||||
| hnsw_index | graph_links.rs | 0.6998417 | ` fn num_points (& self) -> usize ` |
|
||||
|
||||
With the grouping feature, we get more diverse results.
|
||||
|
||||
## Summary
|
||||
|
||||
This tutorial demonstrates how to use Qdrant to navigate a codebase. For an
|
||||
end-to-end implementation, review the [code search
|
||||
notebook](https://colab.research.google.com/github/qdrant/examples/blob/master/code-search/code-search.ipynb).
|
||||
|
After Width: | Height: | Size: 56 KiB |
|
Before Width: | Height: | Size: 2.1 MiB |
|
Before Width: | Height: | Size: 2.1 MiB After Width: | Height: | Size: 44 KiB |
|
After Width: | Height: | Size: 427 KiB |
|
After Width: | Height: | Size: 178 KiB |
|
Before Width: | Height: | Size: 2.1 MiB |
|
Before Width: | Height: | Size: 2.1 MiB After Width: | Height: | Size: 145 KiB |
|
After Width: | Height: | Size: 154 KiB |
|
After Width: | Height: | Size: 30 KiB |
|
After Width: | Height: | Size: 598 KiB |
|
After Width: | Height: | Size: 604 KiB |
|
After Width: | Height: | Size: 596 KiB |
|
After Width: | Height: | Size: 134 KiB |