improved langchain integration with updated api and diagrams

This commit is contained in:
Manas Chopra
2026-07-30 16:41:20 +05:30
parent 8e1dd68bfa
commit c287a9d1f5
7 changed files with 186 additions and 52 deletions
@@ -1,7 +1,7 @@
--- ---
title: "Using LangChain for Question Answering with Qdrant" title: "Question Answering with LangChain and Qdrant"
short_description: "Large Language Models might be developed fast with modern tool. Here is how!" short_description: "Build a retrieval-augmented question answering pipeline with just a few lines of code."
description: "We combined LangChain, a pre-trained LLM from OpenAI, SentenceTransformers & Qdrant to create a question answering system with just a few lines of code. Learn more!" description: "We combined LangChain, a modern chat model like Claude or GPT, FastEmbed & Qdrant to create a question answering system with just a few lines of code. Learn more!"
social_preview_image: /articles_data/langchain-integration/preview/social_preview.jpg social_preview_image: /articles_data/langchain-integration/preview/social_preview.jpg
small_preview_image: /articles_data/langchain-integration/chain.svg small_preview_image: /articles_data/langchain-integration/chain.svg
preview_dir: /articles_data/langchain-integration/preview preview_dir: /articles_data/langchain-integration/preview
@@ -17,35 +17,43 @@ keywords:
- large language models - large language models
- question answering - question answering
- openai - openai
- anthropic
- claude
- fastembed
- embeddings - embeddings
category: demos-and-tutorials category: demos-and-tutorials
--- ---
# Streamlining Question Answering: Simplifying Integration with LangChain and Qdrant # Question Answering with LangChain and Qdrant
**Follow along in Colab:** [![Open In Colab](https://colab.research.google.com/assets/colab-badge.svg)](https://colab.research.google.com/github/qdrant/examples/blob/add-langchain-integration/langchain-integration/langchain_integration.ipynb)
Building applications with Large Language Models doesn't have to be complicated. A lot has been going on recently to simplify the development, Building applications with Large Language Models doesn't have to be complicated. A lot has been going on recently to simplify the development,
so you can utilize already pre-trained models and support even complex pipelines with a few lines of code. [LangChain](https://langchain.readthedocs.io) so you can utilize already pre-trained models and support even complex pipelines with a few lines of code. [LangChain](https://docs.langchain.com/oss/python/langchain/overview)
provides unified interfaces to different libraries, so you can avoid writing boilerplate code and focus on the value you want to bring. provides unified interfaces to different libraries, so you can avoid writing boilerplate code and focus on the value you want to bring.
## Why Use Qdrant for Question Answering with LangChain? ## Why Use Qdrant for Question Answering with LangChain?
It has been reported millions of times recently, but let's say that again. ChatGPT-like models struggle with generating factual statements if no context It has been reported millions of times, but let's say it again. Modern LLMs, whether that's Claude, GPT, or any other chat model, still struggle to
is provided. They have some general knowledge but cannot guarantee to produce a valid answer consistently. Thus, it is better to provide some facts we generate factual statements if no context is provided. They have some general knowledge but cannot guarantee to produce a valid answer consistently. Thus,
know are actual, so it can just choose the valid parts and extract them from all the provided contextual data to give a comprehensive answer. [Vector database, it is better to provide some facts we know are actual, so it can just choose the valid parts and extract them from all the provided contextual data to give
such as Qdrant](https://qdrant.tech/), is of great help here, as their ability to perform a [semantic search](https://qdrant.tech/documentation/tutorials/search-beginners/) over a huge knowledge base is crucial to preselect some possibly valid a comprehensive answer. A [vector search engine, such as Qdrant](https://qdrant.tech/), is of great help here, as its ability to perform a
documents, so they can be provided into the LLM. That's also one of the **chains** implemented in [LangChain](https://qdrant.tech/documentation/frameworks/langchain/), which is called `VectorDBQA`. And Qdrant got [semantic search](https://qdrant.tech/documentation/tutorials/search-beginners/) over a huge knowledge base is crucial to preselect some possibly valid
integrated with the library, so it might be used to build it effortlessly. documents, so they can be provided into the LLM. This pattern is commonly known as retrieval-augmented generation, and it is one of the core building
blocks of [LangChain](https://qdrant.tech/documentation/frameworks/langchain/), which got Qdrant integrated as a first-class vector store, so it might be
used to build such pipelines effortlessly.
### The Two-Model Approach ### The Two-Model Approach
Surprisingly enough, there will be two models required to set things up. First of all, we need an embedding model that will convert the set of facts into Surprisingly enough, there will be two models required to set things up. First of all, we need an embedding model that will convert the set of facts into
vectors, and store those into Qdrant. That's an identical process to any other semantic search application. We're going to use one of the vectors, and store those into Qdrant. That's an identical process to any other semantic search application. We're going to use
`SentenceTransformers` models, so it can be hosted locally. The embeddings created by that model will be put into Qdrant and used to retrieve the most [FastEmbed](https://qdrant.tech/articles/fastembed/), Qdrant's own lightweight embedding library, so it can be hosted locally without pulling in a full
similar documents, given the query. PyTorch or TensorFlow stack. The embeddings created by that model will be put into Qdrant and used to retrieve the most similar documents, given the query.
However, when we receive a query, there are two steps involved. First of all, we ask Qdrant to provide the most relevant documents and simply combine all However, when we receive a query, there are two steps involved. First of all, we ask Qdrant to provide the most relevant documents and simply combine all
of them into a single text. Then, we build a prompt to the LLM (in our case [OpenAI](https://openai.com/)), including those documents as a context, of course together with the of them into a single text. Then, we build a prompt to the chat model (in our examples below, either [Anthropic's Claude](https://www.anthropic.com/claude)
question asked. So the input to the LLM looks like the following: or [OpenAI's GPT](https://openai.com/)), including those documents as a context, of course together with the question asked. So the input to the LLM
looks like the following:
```text ```text
Use the following pieces of context to answer the question at the end. If you don't know the answer, just say that you don't know, don't try to make up an answer. Use the following pieces of context to answer the question at the end. If you don't know the answer, just say that you don't know, don't try to make up an answer.
@@ -56,54 +64,158 @@ Question: How much is 2 + 2?
Helpful Answer: Helpful Answer:
``` ```
There might be several context documents combined, and it is solely up to LLM to choose the right piece of content. But our expectation is, the model should There might be several context documents combined, and it is solely up to the LLM to choose the right piece of content. But our expectation is, the model
respond with just `4`. should respond with just `4`.
## Why do we need two different models? ## Why do we need two different models?
Both solve some different tasks. The first model performs feature extraction, by converting the text into vectors, while Both solve some different tasks. The first model performs feature extraction, by converting the text into vectors, while
the second one helps in text generation or summarization. Disclaimer: This is not the only way to solve that task with LangChain. Such a chain is called `stuff` the second one helps in text generation or summarization. Disclaimer: this is not the only way to solve that task with LangChain. Since we simply stuff
in the library nomenclature. all the retrieved documents into a single prompt, this pattern is often called a **stuff** chain.
![](/articles_data/langchain-integration/flow-diagram.png) ![](/articles_data/langchain-integration/flow-diagram.png)
Enough theory! This sounds like a pretty complex application, as it involves several systems. But with LangChain, it might be implemented in just a few lines Enough theory! This sounds like a pretty complex application, as it involves several systems. But with LangChain, it might be implemented in just a few
of code, thanks to the recent integration with [Qdrant](https://qdrant.tech/). We're not even going to work directly with `QdrantClient`, as everything is already done in the background lines of code, thanks to the integration with [Qdrant](https://qdrant.tech/). We're not even going to work directly with `QdrantClient`, as everything is
by LangChain. If you want to get into the source code right away, all the processing is available as a already done in the background by LangChain.
[Google Colab notebook](https://colab.research.google.com/drive/19RxxkZdnq_YqBH5kBV10Rt0Rax-kminD?usp=sharing).
## How to Implement Question Answering with LangChain and Qdrant ## How to Implement Question Answering with LangChain and Qdrant
### Step 1: Configuration ### Step 1: Configuration
A journey of a thousand miles begins with a single step, in our case with the configuration of all the services. We'll be using [Qdrant Cloud](https://cloud.qdrant.io), Before anything else, install the packages this pipeline touches - LangChain's Qdrant integration, FastEmbed, the `datasets` library for loading Natural
so we need an API key. The same is for OpenAI - the API key has to be obtained from their website. Questions, and whichever chat model provider you'd like to call:
![](/articles_data/langchain-integration/code-configuration.png) ```shell
pip install langchain langchain-qdrant fastembed datasets langchain-anthropic langchain-openai
```
A journey of a thousand miles begins with a single step, in our case with the configuration of all the services. We'll be using [Qdrant
Cloud](https://cloud.qdrant.io), so we need a URL and an API key. On the generation side, you can plug in whichever chat model you prefer - an API key
from [Anthropic](https://console.anthropic.com/) or [OpenAI](https://platform.openai.com/) is all you need, since LangChain exposes the same interface for
both.
```python
import os
os.environ["QDRANT_URL"] = "https://xxxxxx-xxxxxx.xxx.aws.cloud.qdrant.io"
os.environ["QDRANT_API_KEY"] = "<your-qdrant-api-key>"
# Pick whichever provider you'd like to use for generating the answers
os.environ["ANTHROPIC_API_KEY"] = "<your-anthropic-api-key>" # for Claude
os.environ["OPENAI_API_KEY"] = "<your-openai-api-key>" # for GPT
```
### Step 2: Building the knowledge base ### Step 2: Building the knowledge base
We also need some facts from which the answers will be generated. There is plenty of public datasets available, and We also need some facts from which the answers will be generated. There is plenty of public datasets available, and
[Natural Questions](https://ai.google.com/research/NaturalQuestions/visualization) is one of them. It consists of the whole HTML content of the websites they were [Natural Questions](https://ai.google.com/research/NaturalQuestions/visualization) is one of them - a collection of real Google search queries paired
scraped from. That means we need some preprocessing to extract plain text content. As a result, we’re going to have two lists of strings - one for questions and with the relevant passage from Wikipedia that answers them. Rather than parsing the raw, HTML-heavy release ourselves, we can pull the
the other one for the answers. already-cleaned `query`/`answer` pairs published on the Hugging Face Hub, which gets us two lists of strings - one for questions and the other one for
the answers - in a single call.
The answers have to be vectorized with the first of our models. The `sentence-transformers/all-mpnet-base-v2` is one of the possibilities, but there are some ```python
other options available. LangChain will handle that part of the process in a single function call. from datasets import load_dataset
![](/articles_data/langchain-integration/code-qdrant.png) dataset = load_dataset("sentence-transformers/natural-questions", split="train")
# 100 pairs is enough to experiment with; drop the .select() call entirely to index all 100k+ rows
dataset = dataset.select(range(100))
### Step 3: Setting up QA with Qdrant in a loop questions = dataset["query"]
answers = dataset["answer"]
```
`VectorDBQA` is a chain that performs the process described above. So it, first of all, loads some facts from Qdrant and then feeds them into OpenAI LLM which The answers have to be vectorized with our embedding model. FastEmbed defaults to
should analyze them to find the answer to a given question. The only last thing to do before using it is to put things together, also with a single function call. [`BAAI/bge-small-en-v1.5`](https://huggingface.co/BAAI/bge-small-en-v1.5), a small, quantized model that runs comfortably on CPU, but there are
[several other models](https://qdrant.github.io/fastembed/examples/Supported_Models/) to pick from. `langchain-community`, which used to ship a
`FastEmbedEmbeddings` wrapper, is [being sunset](https://github.com/langchain-ai/langchain-community/issues/674), so instead we wrap FastEmbed's
`TextEmbedding` directly with LangChain's `Embeddings` interface - it's a handful of lines, and it keeps the pipeline free of a deprecated dependency.
LangChain will still handle vectorizing the documents and creating the Qdrant collection in a single function call.
![](/articles_data/langchain-integration/code-vectordbqa.png) ```python
from typing import List
from fastembed import TextEmbedding
from langchain_core.embeddings import Embeddings
from langchain_qdrant import QdrantVectorStore
class FastEmbedEmbeddings(Embeddings):
def __init__(self, model_name: str = "BAAI/bge-small-en-v1.5"):
self._model = TextEmbedding(model_name=model_name)
def embed_documents(self, texts: List[str]) -> List[List[float]]:
return [vector.tolist() for vector in self._model.embed(texts)]
def embed_query(self, text: str) -> List[float]:
return self.embed_documents([text])[0]
embeddings = FastEmbedEmbeddings()
doc_store = QdrantVectorStore.from_texts(
answers,
embeddings,
url=os.environ["QDRANT_URL"],
api_key=os.environ["QDRANT_API_KEY"],
collection_name="natural-questions",
)
```
### Step 3: Setting up the retrieval chain
With the knowledge base in place, the only thing left is to combine the retriever with a chat model. [`init_chat_model`](https://docs.langchain.com/oss/python/langchain/models)
lets us pass in the name of any supported model - Claude, GPT, or otherwise - without changing the rest of the pipeline. We then compose the retriever, the
prompt, and the model using LangChain's expression language (LCEL), so the whole chain is defined with a single `|`-piped expression.
```python
from langchain.chat_models import init_chat_model
from langchain_core.output_parsers import StrOutputParser
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
retriever = doc_store.as_retriever()
prompt = ChatPromptTemplate.from_template(
"""Use the following pieces of context to answer the question at the end. If you don't know the answer, just
say that you don't know, don't try to make up an answer.
{context}
Question: {question}
Helpful Answer:"""
)
def format_docs(docs):
return "\n\n".join(doc.page_content for doc in docs)
# Swap the model name for any other provider LangChain supports, e.g. "gpt-5.1"
llm = init_chat_model("claude-sonnet-4-5", model_provider="anthropic")
chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
```
## Step 4: Testing out the chain ## Step 4: Testing out the chain
And that's it! We can put some queries, and LangChain will perform all the required processing to find the answer in the provided context. And that's it! We can put in some queries, and LangChain will perform all the required processing to find the answer in the provided context. Since we
already have a `questions` list from the dataset, let's just sample a handful of them and see how the chain responds:
![](/articles_data/langchain-integration/code-answering.png) ```python
import random
random.seed(76)
selected_questions = random.choices(questions, k=5)
for question in selected_questions:
print(">", question)
print(chain.invoke(question), end="\n\n")
```
The exact wording will vary depending on which chat model you plug in, but running the chain against the same knowledge base as the original experiment
produces answers along these lines:
```text ```text
> what kind of music is scott joplin most famous for > what kind of music is scott joplin most famous for
@@ -123,7 +235,4 @@ And that's it! We can put some queries, and LangChain will perform all the requi
``` ```
The great thing about such a setup is that the knowledge base might be easily extended with some new facts and those will be included in the prompts The great thing about such a setup is that the knowledge base might be easily extended with some new facts and those will be included in the prompts
sent to LLM later on. Of course, assuming their similarity to the given question will be in the top results returned by Qdrant. sent to the LLM later on. Of course, assuming their similarity to the given question will be in the top results returned by Qdrant.
If you want to run the chain on your own, the simplest way to reproduce it is to open the
[Google Colab notebook](https://colab.research.google.com/drive/19RxxkZdnq_YqBH5kBV10Rt0Rax-kminD?usp=sharing).
@@ -73,7 +73,7 @@ qdrant = QdrantVectorStore.from_documents(
Local mode, without using the Qdrant server, may also store your vectors on disk so they’re persisted between runs. Local mode, without using the Qdrant server, may also store your vectors on disk so they’re persisted between runs.
```python ```python
qdrant = Qdrant.from_documents( qdrant = QdrantVectorStore.from_documents(
docs, docs,
embeddings, embeddings,
path="/tmp/local_qdrant", path="/tmp/local_qdrant",
@@ -111,7 +111,7 @@ qdrant = QdrantVectorStore.from_documents(
To search with only dense vectors, To search with only dense vectors,
- The `retrieval_mode` parameter should be set to `RetrievalMode.DENSE`(default). - The `retrieval_mode` parameter should be set to `RetrievalMode.DENSE`(default).
- A [dense embeddings](https://python.langchain.com/v0.2/docs/integrations/text_embedding/) value should be provided for the `embedding` parameter. - A [dense embeddings](https://docs.langchain.com/oss/python/integrations/text_embedding) value should be provided for the `embedding` parameter.
```py ```py
from langchain_qdrant import RetrievalMode from langchain_qdrant import RetrievalMode
@@ -128,6 +128,31 @@ query = "What did the president say about Ketanji Brown Jackson"
found_docs = qdrant.similarity_search(query) found_docs = qdrant.similarity_search(query)
``` ```
If you'd rather not depend on an embedding provider's API, [FastEmbed](https://github.com/qdrant/fastembed) also lets you generate dense embeddings
locally. `langchain-community`, which used to ship a `FastEmbedEmbeddings` class, is [being sunset](https://github.com/langchain-ai/langchain-community/issues/674),
so wrap FastEmbed's `TextEmbedding` directly with LangChain's `Embeddings` interface instead:
```py
from typing import List
from fastembed import TextEmbedding
from langchain_core.embeddings import Embeddings
class FastEmbedEmbeddings(Embeddings):
def __init__(self, model_name: str = "BAAI/bge-small-en-v1.5"):
self._model = TextEmbedding(model_name=model_name)
def embed_documents(self, texts: List[str]) -> List[List[float]]:
return [vector.tolist() for vector in self._model.embed(texts)]
def embed_query(self, text: str) -> List[float]:
return self.embed_documents([text])[0]
embeddings = FastEmbedEmbeddings() # defaults to BAAI/bge-small-en-v1.5
```
### Sparse Vector Search ### Sparse Vector Search
To search with only sparse vectors, To search with only sparse vectors,
@@ -142,7 +167,7 @@ To use it, install the [FastEmbed package](https://github.com/qdrant/fastembed#-
```python ```python
from langchain_qdrant import FastEmbedSparse, RetrievalMode from langchain_qdrant import FastEmbedSparse, RetrievalMode
sparse_embeddings = FastEmbedSparse(model_name="Qdrant/BM25") sparse_embeddings = FastEmbedSparse(model_name="Qdrant/bm25")
qdrant = QdrantVectorStore.from_documents( qdrant = QdrantVectorStore.from_documents(
docs, docs,
@@ -161,7 +186,7 @@ found_docs = qdrant.similarity_search(query)
To perform a hybrid search using dense and sparse vectors with score fusion, To perform a hybrid search using dense and sparse vectors with score fusion,
- The `retrieval_mode` parameter should be set to `RetrievalMode.HYBRID`. - The `retrieval_mode` parameter should be set to `RetrievalMode.HYBRID`.
- A [dense embeddings](https://python.langchain.com/v0.2/docs/integrations/text_embedding/) value should be provided for the `embedding` parameter. - A [dense embeddings](https://docs.langchain.com/oss/python/integrations/text_embedding) value should be provided for the `embedding` parameter.
- An implementation of the [SparseEmbeddings interface](https://github.com/langchain-ai/langchain/blob/master/libs/partners/qdrant/langchain_qdrant/sparse_embeddings.py) using any sparse embeddings provider has to be provided as value to the `sparse_embedding` parameter. - An implementation of the [SparseEmbeddings interface](https://github.com/langchain-ai/langchain/blob/master/libs/partners/qdrant/langchain_qdrant/sparse_embeddings.py) using any sparse embeddings provider has to be provided as value to the `sparse_embedding` parameter.
```python ```python
@@ -187,7 +212,7 @@ Note that if you've added documents with HYBRID mode, you can switch to any retr
## Next steps ## Next steps
If you'd like to know more about running Qdrant in a LangChain-based application, please read our article If you'd like to know more about running Qdrant in a LangChain-based application, please read our article
[Question Answering with LangChain and Qdrant without boilerplate](/articles/langchain-integration/). Some more information [Question Answering with LangChain and Qdrant](/articles/langchain-integration/). Some more information
might also be found in the [LangChain documentation](https://python.langchain.com/docs/integrations/vectorstores/qdrant). might also be found in the [LangChain documentation](https://python.langchain.com/docs/integrations/vectorstores/qdrant).
- [Source Code](https://github.com/langchain-ai/langchain/tree/master/libs%2Fpartners%2Fqdrant) - [Source Code](https://github.com/langchain-ai/langchain/tree/master/libs%2Fpartners%2Fqdrant)
Binary file not shown.

Before

Width:  |  Height:  |  Size: 97 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 53 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 116 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 102 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 297 KiB

After

Width:  |  Height:  |  Size: 374 KiB