From f0ba324abe09dde13dfd29950f47804189fb57dd Mon Sep 17 00:00:00 2001 From: Anush008 Date: Mon, 19 May 2025 21:17:24 +0530 Subject: [PATCH] docs: NL Web integration Signed-off-by: Anush008 --- .../documentation/frameworks/_index.md | 3 +- .../content/documentation/frameworks/nlweb.md | 97 +++++++++++++++++++ 2 files changed, 99 insertions(+), 1 deletion(-) create mode 100644 qdrant-landing/content/documentation/frameworks/nlweb.md diff --git a/qdrant-landing/content/documentation/frameworks/_index.md b/qdrant-landing/content/documentation/frameworks/_index.md index bfaf6e374..ef7eeec3d 100644 --- a/qdrant-landing/content/documentation/frameworks/_index.md +++ b/qdrant-landing/content/documentation/frameworks/_index.md @@ -24,7 +24,7 @@ aliases: ["/documentation/frameworks/memgpt/"] | [Fifty-One](/documentation/frameworks/fifty-one/) | Toolkit for building high-quality datasets and computer vision models. | | [Genkit](/documentation/frameworks/genkit/) | Framework to build, deploy, and monitor production-ready AI-powered apps. | | [Haystack](/documentation/frameworks/haystack/) | LLM orchestration framework to build customizable, production-ready LLM applications. | -| [HoneyHive](/documentation/frameworks/honeyhive/) | AI observability and evaluation platform that provides tracing and monitoring tools for GenAI pipelines. | +| [HoneyHive](/documentation/frameworks/honeyhive/) | AI observability and evaluation platform that provides tracing and monitoring tools for GenAI pipelines. | | [Lakechain](/documentation/frameworks/lakechain/) | Python framework for deploying document processing pipelines on AWS using infrastructure-as-code. | | [Langchain](/documentation/frameworks/langchain/) | Python framework for building context-aware, reasoning applications using LLMs. | | [Langchain-Go](/documentation/frameworks/langchain-go/) | Go framework for building context-aware, reasoning applications using LLMs. | @@ -35,6 +35,7 @@ aliases: ["/documentation/frameworks/memgpt/"] | [Mirror Security](/documentation/frameworks/mirror-security/) | Python framework for vector encryption and access control. | | [Mem0](/documentation/frameworks/mem0/) | Self-improving memory layer for LLM applications, enabling personalized AI experiences. | | [Neo4j GraphRAG](/documentation/frameworks/neo4j-graphrag/) | Package to build graph retrieval augmented generation (GraphRAG) applications using Neo4j and Python. | +| [NLWeb](/documentation/frameworks/nlweb/) | A framework to turn websites into chat-ready data using schema.org and associated data formats. | | [OpenAI Agents](/documentation/frameworks/openai-agents/) | Python framework for managing multiple AI agents that can work together. | | [Pandas-AI](/documentation/frameworks/pandas-ai/) | Python library to query/visualize your data (CSV, XLSX, PostgreSQL, etc.) in natural language | | [Ragbits](/documentation/frameworks/ragbits/) | Python package that offers essential "bits" for building powerful Retrieval-Augmented Generation (RAG) applications. | diff --git a/qdrant-landing/content/documentation/frameworks/nlweb.md b/qdrant-landing/content/documentation/frameworks/nlweb.md new file mode 100644 index 000000000..a6349789f --- /dev/null +++ b/qdrant-landing/content/documentation/frameworks/nlweb.md @@ -0,0 +1,97 @@ +--- +title: Microsoft NLWeb +--- + +# NLWeb + +Microsoft's [NLWeb](https://github.com/microsoft/NLWeb) is a proposed framework that enables natural language interfaces for websites, using Schema.org, formats like RSS and the emerging [MCP protocol](https://github.com/microsoft/NLWeb/blob/main/docs/RestAPI.md). + +Qdrant is supported as a vector store backend within NLWeb for embedding storage and context retrieval. + +## Usage + +NLWeb includes Qdrant integration by default. You can install and configure it to use Qdrant as the retrieval engine. + +### Installation + +Clone the repo and set up your environment: + +```bash +git clone https://github.com/microsoft/NLWeb +cd NLWeb +python -m venv .venv +source venv/bin/activate # or `venv\Scripts\activate` on Windows +cd code +pip install -r requirements.txt +``` + +### Configuring Qdrant + +To use **Qdrant**, update your configuration. + +#### 1. Copy and edit the environment variables file + +```bash +cp .env.template .env +``` + +Ensure the following values are set in your `.env` file: + +```text +QDRANT_URL="https://xyz-example.cloud-region.cloud-provider.cloud.qdrant.io:6333" +QDRANT_API_KEY="" +``` + +#### 2. Update config files in `code/config` + +* **`config_retrieval.yaml`** + +```yaml +retrieval_engine: qdrant_url +``` + +Alternatively you can use an in-memory Qdrant instance for experimentation. + +```yaml +retrieval_engine: qdrant_local + +endpoints: + qdrant_local: + # Use local file-based storage with a specific path + database_path: "../data/db" + # Set the collection name to use + index_name: nlweb_collection + # Specify the database type + db_type: qdrant +``` + +### Loading Data + +Once configured, load your content using RSS feeds or other supported formats. + +From the `code` directory: + +```bash +python -m tools.db_load https://feeds.libsyn.com/121695/rss Behind-the-Tech +python -m tools.db_load https://feeds.megaphone.fm/recodedecode Decoder +``` + +This will ingest the content into your local Qdrant instance. + +### Running the Server + +To start your NLWeb server with Qdrant: + +```bash +python -m server.main +``` + +You can now query your content via natural language using either the web UI at or directly through the MCP-compatible [REST API](https://github.com/microsoft/NLWeb/blob/main/docs/RestAPI.md). + +## Further Reading + +* [Source](https://github.com/microsoft/NLWeb) +* [Life of a Chat Query](https://github.com/microsoft/NLWeb/tree/main/docs/LifeOfAChatQuery.md) +* [Modifying behaviour by changing prompts](https://github.com/microsoft/NLWeb/tree/main/docs/Prompts.md) +* [Modifying control flow](https://github.com/microsoft/NLWeb/tree/main/docs/ControlFlow.md) +* [Modifying the user interface](https://github.com/microsoft/NLWeb/tree/main/docs/UserInterface.md)