Merge pull request #223 from qdrant/fix-snippets

properly assign language to snippets
This commit is contained in:
David Sertic
2023-07-13 09:18:25 +02:00
committed by GitHub
13 changed files with 32 additions and 27 deletions
@@ -205,7 +205,7 @@ The command will build the image using the Dockerfile and deploy the service at
URL. Once the command is finished, the service should be running on the hostname we got
previously:
```
```text
https://your-application-name.fly.dev
```
@@ -321,20 +321,20 @@ fio --randrepeat=1 \
Initially, we tested on a network-mounted disk, but its performance was too slow, with a read IOPS of 6366 and a bandwidth of 24.9 MiB/s:
```
```text
read: IOPS=6366, BW=24.9MiB/s (26.1MB/s)(8192MiB/329424msec)
```
To improve performance, we switched to a local disk, which showed much faster results, with a read IOPS of 63.2k and a bandwidth of 247 MiB/s:
```
```text
read: IOPS=63.2k, BW=247MiB/s (259MB/s)(8192MiB/33207msec)
```
That gave us a significant speed boost, but we wanted to see if we could improve performance even further.
To do that, we switched to a machine with a local SSD, which showed even better results, with a read IOPS of 183k and a bandwidth of 716 MiB/s:
```
```text
read: IOPS=183k, BW=716MiB/s (751MB/s)(8192MiB/11438msec)
```
@@ -212,7 +212,7 @@ This payload allows additional filtering during the search request.
Qdrant has a pre-built docker image and start working with it is just as simple as running
```
```bash
docker run -p 6333:6333 qdrant/qdrant
```
@@ -118,20 +118,20 @@ To start Qdrant, use the instructions on its [homepage](https://github.com/qdran
Download image from [DockerHub](https://hub.docker.com/r/qdrant/qdrant):
```
```bash
docker pull qdrant/qdrant
```
And run the service inside the docker:
```
```bash
docker run -p 6333:6333 \
-v $(pwd)/qdrant_storage:/qdrant/storage \
qdrant/qdrant
```
You should see output like this
```
```text
...
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting 12 workers
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting "actix-web-service-0.0.0.0:6333" service on 0.0.0.0:6333
@@ -149,7 +149,7 @@ To interact with Qdrant from python, I recommend using an out-of-the-box client
To install it, use the following command
```
```bash
pip install qdrant-client
```
@@ -229,7 +229,7 @@ The full code for this step could be found [here](https://github.com/qdrant/qdra
Now that all the preparations are complete, let's start building a neural search class.
First, install all the requirements:
```
```bash
pip install sentence-transformers numpy
```
@@ -313,7 +313,7 @@ It is super easy to use and requires minimal code writing.
To install it, use the command
```
```bash
pip install fastapi uvicorn
```
@@ -347,7 +347,7 @@ if __name__ == "__main__":
Now, if you run the service with
```
```bash
python service.py
```
@@ -23,7 +23,7 @@ Quantum quantization is a novel approach that leverages the power of quantum com
The conversion of float32 vectors to qbit vectors can be represented by the following formula:
```
```text
qbit_vector = Q( float32_vector )
```
@@ -29,7 +29,7 @@ So a single number needs 4 bytes of the memory and a 512-dimensional vector occu
2 kB. That's only the memory used to store the vector. There is also an overhead of the
HNSW graph, so as a rule of thumb we estimate the memory size with the following formula:
```
```text
memory_size = 1.5 * number_of_vectors * vector_dimension * 4 bytes
```
@@ -49,7 +49,7 @@ And 8 CPUs + 16Gb RAM for client machine. We were trying to make the bottleneck
### Experiment setup
```
```text
┌────────┐ ┌──────────┐
│ ├─────►│ │
│ Client │ │ Engine │
@@ -14,7 +14,7 @@ It depends on a number of factors and options you can choose for your collection
If you need to keep all vectors in memory for maximum performance, there is a very rough formula for estimating the needed memory size looks like this:
```
```text
memory_size = number_of_vectors * vector_dimension * 4 bytes * 1.5
```
@@ -46,6 +46,6 @@ the fast search for the most active and recent users.
In this case you can estimate required memory size as follows:
```
```text
memory_size = number_of_active_vectors * vector_dimension * 4 bytes * 1.5
```
@@ -40,13 +40,13 @@ With python client it is possible to try Qdrant without running docker container
Simply install it
```
```bash
pip install qdrant-client
```
and initialize client like this:
```
```python
from qdrant_client import QdrantClient
client = QdrantClient(":memory:")
@@ -9,7 +9,7 @@ weight: 100
Each collection segment needs some files to be open. At some point you may encounter the following errors in your server log:
```
```text
Error: Too many files open (OS error 24)
```
@@ -22,7 +22,7 @@ Much like Qdrant, the [Mighty](https://max.io/) inference server is written in R
For Mighty, start up a [docker container](https://hub.docker.com/layers/maxdotio/mighty-sentence-transformers/0.9.9/images/sha256-0d92a89fbdc2c211d927f193c2d0d34470ecd963e8179798d8d391a4053f6caf?context=explore) with an open port 5050. Just loading the port in a window shows the following:
```
```json
{
"name": "sentence-transformers/all-MiniLM-L6-v2",
"architectures": [
@@ -109,7 +109,7 @@ docker run -p 6333:6333 \
```
You should see output like this
```
```text
...
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting 12 workers
[2021-02-05T00:08:51Z INFO actix_server::builder] Starting "actix-web-service-0.0.0.0:6333" service on 0.0.0.0:6333
@@ -310,7 +310,7 @@ if __name__ == "__main__":
3. Run the service.
```
```bash
python service.py
```
@@ -19,24 +19,28 @@ Before you begin, you need to have a [recent version of Python](https://www.pyth
## 1. Installation
You need to process your data so that the search engine can work with it. The [Sentence Transformers] framework gives you access to common [Large Language Models] that turn raw data into embeddings.
```python
```bash
pip install -U sentence-transformers
```
Once encoded, this data needs to be kept somewhere. Qdrant lets you store data as embeddings. You can also use Qdrant to run search queries against this data. This means that you can ask the engine to give you relevant answers that go way beyond keyword matching.
```python
```bash
pip install qdrant-client
```
### Import the models
Once the two main frameworks are defined, you need to specify the exact models this engine will use.
```python
from qdrant_client import models, QdrantClient
from sentence_transformers import SentenceTransformer
```
The [Sentence Transformers](https://www.sbert.net/index.html) framework contains many embedding models. However, [all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2) is the fastest encoder for this tutorial.
```python
encoder = SentenceTransformer('all-MiniLM-L6-v2')
```
@@ -127,11 +131,12 @@ for hit in hits:
The search engine shows three of the most likely responses that have to do with the alien invasion. Each of the responses is assigned a score to show how close the response is to the original inquiry.
```python
```text
{'name': 'The War of the Worlds', 'description': 'A Martian invasion of Earth throws humanity into chaos.', 'author': 'H.G. Wells', 'year': 1898} score: 0.570093257022374
{'name': "The Hitchhiker's Guide to the Galaxy", 'description': 'A comedic science fiction series following the misadventures of an unwitting human and his alien friend.', 'author': 'Douglas Adams', 'year': 1979} score: 0.5040468703143637
{'name': 'The Three-Body Problem', 'description': 'Humans encounter an alien civilization that lives in a dying system.', 'author': 'Liu Cixin', 'year': 2008} score: 0.45902943411768216
```
### Narrow down the query
How about the most recent book from the early 2000s?
@@ -160,7 +165,7 @@ for hit in hits:
The query has been narrowed down to one result from 2008.
```python
```text
{'name': 'The Three-Body Problem', 'description': 'Humans encounter an alien civilization that lives in a dying system.', 'author': 'Liu Cixin', 'year': 2008} score: 0.45902943411768216
```