From 08c296ae0233b2b72d08fac8e1507f618b2ba383 Mon Sep 17 00:00:00 2001 From: mjang Date: Fri, 8 Mar 2024 08:31:26 -0800 Subject: [PATCH 1/4] Add clarifying process to code search demo --- .../documentation/tutorials/code-search.md | 22 +++++++++++++++++-- 1 file changed, 20 insertions(+), 2 deletions(-) diff --git a/qdrant-landing/content/documentation/tutorials/code-search.md b/qdrant-landing/content/documentation/tutorials/code-search.md index cd0628fdb..26f7feea1 100644 --- a/qdrant-landing/content/documentation/tutorials/code-search.md +++ b/qdrant-landing/content/documentation/tutorials/code-search.md @@ -20,8 +20,8 @@ source code itself, which is mostly written in Rust. We want to search codebases using natural semantic queries, and searching for code based on similar logic. You can set up these tasks with embeddings: -1. General usage neural encoder for natural-like queries, in our case `all-MiniLM-L6-v2` - from the +1. General usage neural encoder for Natural Language Processing (NLP), in our case + `all-MiniLM-L6-v2` from the [sentence-transformers](https://www.sbert.net/docs/pretrained_models.html) library. 2. Specialized embeddings for code-to-code similarity search. We use the `jina-embeddings-v2-base-code` model. @@ -31,6 +31,24 @@ more closely resembles natural language. The Jina embeddings model supports a variety of standard programming languages, so there is no need to preprocess the snippets. We can use the code as is. +## The process + +Once configured, our code search demo uses the following process: + +1. The user sends a query. +1. Both models vectorize that query simultaneously. We get two different vectors. +1. Both vectors are used in parallel to find relevant snippets. +1. Once we retrieve results for both vectors, we merge them in one of the + following scenarios: + 1. If both methods return different results, we display those results. + 1. If there is an overlap between the search results, we merge overlapping + snippets. + +NLP-based search is based on function signatures, but code search may return +smaller pieces, such as loops. So, if we receive a particular function signature +from the NLP model and part of its implementation from the code model, we merge +the results and highlight the overlap. + ## Data preparation Chunking the application sources into smaller parts is a non-trivial task. In From f1610ca6f10064c83cde41c9384cffe326422f3c Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Kacper=20=C5=81ukawski?= Date: Thu, 14 Mar 2024 10:55:13 +0100 Subject: [PATCH 2/4] Move the process description to the code search demo section --- .../documentation/tutorials/code-search.md | 35 +++++++++++-------- 1 file changed, 20 insertions(+), 15 deletions(-) diff --git a/qdrant-landing/content/documentation/tutorials/code-search.md b/qdrant-landing/content/documentation/tutorials/code-search.md index 26f7feea1..8b46e90c1 100644 --- a/qdrant-landing/content/documentation/tutorials/code-search.md +++ b/qdrant-landing/content/documentation/tutorials/code-search.md @@ -31,19 +31,6 @@ more closely resembles natural language. The Jina embeddings model supports a variety of standard programming languages, so there is no need to preprocess the snippets. We can use the code as is. -## The process - -Once configured, our code search demo uses the following process: - -1. The user sends a query. -1. Both models vectorize that query simultaneously. We get two different vectors. -1. Both vectors are used in parallel to find relevant snippets. -1. Once we retrieve results for both vectors, we merge them in one of the - following scenarios: - 1. If both methods return different results, we display those results. - 1. If there is an overlap between the search results, we merge overlapping - snippets. - NLP-based search is based on function signatures, but code search may return smaller pieces, such as loops. So, if we receive a particular function signature from the NLP model and part of its implementation from the code model, we merge @@ -429,14 +416,32 @@ This is one example of how you can use different models and combine the results. In a real-world scenario, you might run some reranking and deduplication, as well as additional processing of the results. -Our [Code search demo](https://github.com/qdrant/demo-code-search) uses -both models. In the screenshot, we search for `flush of wal`. The result +## Code search demo + +Our [Code search demo](https://github.com/qdrant/demo-code-search) uses the following process: + +1. The user sends a query. +1. Both models vectorize that query simultaneously. We get two different + vectors. +1. Both vectors are used in parallel to find relevant snippets. We expect + 5 examples from the NLP search and 20 examples from the code search. +1. Once we retrieve results for both vectors, we merge them in one of the + following scenarios: + 1. If both methods return different results, we prefer the results from + the general usage model (NLP). + 1. If there is an overlap between the search results, we merge overlapping + snippets. + +In the screenshot, we search for `flush of wal`. The result shows relevant code, merged from both models. Note the highlighted code in lines 621-629. It's where both models agree. ![Results from both models, with overlap](/documentation/tutorials/code-search/code-search-demo-example.png) Now you see semantic code intelligence, in action. + +Our demo is available online + ### Grouping the results You can improve the search results, by grouping them by payload properties. From d99e34334af6f6d2eae00e42e945056c9b810fa3 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Kacper=20=C5=81ukawski?= Date: Thu, 14 Mar 2024 10:58:37 +0100 Subject: [PATCH 3/4] Change the link to point to the online version --- qdrant-landing/content/documentation/tutorials/code-search.md | 4 +--- 1 file changed, 1 insertion(+), 3 deletions(-) diff --git a/qdrant-landing/content/documentation/tutorials/code-search.md b/qdrant-landing/content/documentation/tutorials/code-search.md index 8b46e90c1..203b47ea3 100644 --- a/qdrant-landing/content/documentation/tutorials/code-search.md +++ b/qdrant-landing/content/documentation/tutorials/code-search.md @@ -418,7 +418,7 @@ well as additional processing of the results. ## Code search demo -Our [Code search demo](https://github.com/qdrant/demo-code-search) uses the following process: +Our [Code search demo](https://code-search.qdrant.tech/) uses the following process: 1. The user sends a query. 1. Both models vectorize that query simultaneously. We get two different @@ -440,8 +440,6 @@ code in lines 621-629. It's where both models agree. Now you see semantic code intelligence, in action. -Our demo is available online - ### Grouping the results You can improve the search results, by grouping them by payload properties. From bab7014e7c68a9d8ec74aa038f4d569e130203dd Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Kacper=20=C5=81ukawski?= Date: Thu, 14 Mar 2024 11:00:16 +0100 Subject: [PATCH 4/4] Adjust the headings --- qdrant-landing/content/documentation/tutorials/code-search.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/qdrant-landing/content/documentation/tutorials/code-search.md b/qdrant-landing/content/documentation/tutorials/code-search.md index 203b47ea3..284f9ef3b 100644 --- a/qdrant-landing/content/documentation/tutorials/code-search.md +++ b/qdrant-landing/content/documentation/tutorials/code-search.md @@ -416,7 +416,7 @@ This is one example of how you can use different models and combine the results. In a real-world scenario, you might run some reranking and deduplication, as well as additional processing of the results. -## Code search demo +### Code search demo Our [Code search demo](https://code-search.qdrant.tech/) uses the following process: