docs: Use local inference modern-sparse-neural-retrieval.md (#1600)

* docs: Update to use local inference

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: Review updates

Signed-off-by: Anush008 <anushshetty90@gmail.com>

* chore: Formatting

Signed-off-by: Anush008 <anushshetty90@gmail.com>

---------

Signed-off-by: Anush008 <anushshetty90@gmail.com>
This commit is contained in:
Anush
2025-05-15 19:37:07 +05:30
committed by GitHub
parent 9577b88aac
commit 57fb83b9bc
5 changed files with 72 additions and 7 deletions
@@ -438,6 +438,7 @@ metadata = [{"movie_name": "The Passion of Joan of Arc", "movie_watch_time_min":
</details> </details>
Upload embedded descriptions with movie metadata into the collection. Upload embedded descriptions with movie metadata into the collection.
```python ```python
qdrant_client.upsert( qdrant_client.upsert(
collection_name="movies", collection_name="movies",
@@ -456,6 +457,34 @@ qdrant_client.upsert(
], ],
) )
``` ```
<aside role="status">
You can also implicitly generate sparse vectors using built-in FastEmbed integration.
</aside>
<details>
<summary>Implicitly generate sparse vectors (Click to expand)</summary>
```python
qdrant_client.upsert(
collection_name="movies",
points=[
models.PointStruct(
id=idx,
payload=metadata[idx],
vector={
"film_description": models.Document(
text=description, model=sparse_model_name
)
},
)
for idx, description in enumerate(descriptions)
],
)
```
</details>
#### Querying #### Querying
Let’s query our collection! Let’s query our collection!
@@ -472,6 +501,24 @@ response = qdrant_client.query_points(
) )
print(response) print(response)
``` ```
<details>
<summary>Implicitly generate sparse vectors (Click to expand)</summary>
```python
response = qdrant_client.query_points(
collection_name="movies",
query=models.Document(text="A movie about music", model=sparse_model_name),
using="film_description",
limit=1,
with_vectors=True,
with_payload=True,
)
print(response)
```
</details>
Output looks like this: Output looks like this:
```bash ```bash
points=[ScoredPoint( points=[ScoredPoint(
@@ -586,6 +633,24 @@ response = qdrant_client.query_points(
print(get_tokens_and_weights(response.points[0].vector['film_description'], tokenizer)) print(get_tokens_and_weights(response.points[0].vector['film_description'], tokenizer))
``` ```
<details>
<summary>Implicitly generate sparse vectors (Click to expand)</summary>
```python
response = qdrant_client.query_points(
collection_name="movies",
query=models.Document(text="A movie about music", model=sparse_model_name),
using="film_description",
limit=1,
with_vectors=True,
with_payload=True,
)
print(get_tokens_and_weights(response.points[0].vector["film_description"], tokenizer))
```
</details>
And that's how SPLADE++ expanded the answer. And that's how SPLADE++ expanded the answer.
```python ```python
@@ -219,11 +219,11 @@ You are not restricted to using one LLM in your program; you can use [multiple](
You can easily set up [Qdrant](/documentation/frameworks/dspy/) vector store to act as the retrieval model. To do so, follow these steps: You can easily set up [Qdrant](/documentation/frameworks/dspy/) vector store to act as the retrieval model. To do so, follow these steps:
```python ```python
# pip install dspy-ai[qdrant] # pip install dspy-ai dspy-qdrant
import dspy import dspy
from dspy.retrieve.qdrant_rm import QdrantRM from dspy_qdrant import QdrantRM
from qdrant_client import QdrantClient from qdrant_client import QdrantClient
@@ -156,7 +156,7 @@ LangGraph works with a state-based system. We define our state like this:
```python ```python
class State(TypedDict): class State(TypedDict):
messages: Annotated[list, add_messages] messages: Annotated[list, add_messages]
``` ```
--- ---
@@ -42,7 +42,7 @@ We are going to need a couple of Python packages to run our application. They mi
`dspy-ai` package and `qdrant` extra: `dspy-ai` package and `qdrant` extra:
```shell ```shell
pip install dspy-ai[qdrant] pip install dspy-ai dspy-qdrant
``` ```
### Qdrant Hybrid Cloud ### Qdrant Hybrid Cloud
@@ -181,7 +181,7 @@ gemma_model = dspy.OllamaLocal(
Similarly, we have to define connection to our Qdrant Hybrid Cloud cluster: Similarly, we have to define connection to our Qdrant Hybrid Cloud cluster:
```python ```python
from dspy.retrieve.qdrant_rm import QdrantRM from dspy_qdrant import QdrantRM
from qdrant_client import QdrantClient, models from qdrant_client import QdrantClient, models
client = QdrantClient( client = QdrantClient(
@@ -17,7 +17,7 @@ Qdrant can be used as a retrieval mechanism in the DSPy flow.
For the Qdrant retrieval integration, include `dspy-ai` with the `qdrant` extra: For the Qdrant retrieval integration, include `dspy-ai` with the `qdrant` extra:
```bash ```bash
pip install dspy-ai[qdrant] pip install dspy-ai dspy-qdrant
``` ```
## Usage ## Usage
@@ -25,7 +25,7 @@ pip install dspy-ai[qdrant]
We can configure `DSPy` settings to use the Qdrant retriever model like so: We can configure `DSPy` settings to use the Qdrant retriever model like so:
```python ```python
import dspy import dspy
from dspy.retrieve.qdrant_rm import QdrantRM from dspy_qdrant import QdrantRM
from qdrant_client import QdrantClient from qdrant_client import QdrantClient