diff --git a/LlamaIndex_Icon_onTransparentBackground.svg b/LlamaIndex_Icon_onTransparentBackground.svg new file mode 100755 index 000000000..a74e8cd27 --- /dev/null +++ b/LlamaIndex_Icon_onTransparentBackground.svg @@ -0,0 +1,12 @@ + diff --git a/qdrant-landing/content/course/essentials/day-7/_index.md b/qdrant-landing/content/course/essentials/day-7/_index.md new file mode 100644 index 000000000..373876adf --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/_index.md @@ -0,0 +1,65 @@ +--- +title: "Day 7: Partner Ecosystem Integrations (Bonus)" +isLesson: true +weight: 10 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Partner Ecosystem Integrations (Bonus) + +Explore the Qdrant ecosystem and learn how to integrate with leading AI and data platforms. + +--- + +## Partner Integrations Overview + +Learn about the Qdrant ecosystem and integration strategies. + +[**➡️ Partner Integrations**](https://qdrant.tech/partners/) + +--- + +## Choose Your Integration + +{{< cards-list >}} +- icon: /courses/course-integrations/haystack.svg + title: Haystack + content: Build end-to-end NLP pipelines with Qdrant + link: haystack/ + +- icon: /courses/course-integrations/tensorlake.svg + title: Tensorlake + content: Build scalable data lakes with vector capabilities + link: tensorlake/ + +- icon: /courses/course-integrations/llamaindex.svg + title: LlamaIndex + content: Build agentic workflows for complex enterprise documents + link: llamaindex/ + +- icon: /courses/course-integrations/unstructured.svg + title: Unstructured.io + content: Process and vectorize documents from any format + link: unstructured/ + +- icon: /courses/course-integrations/quotient.svg + title: Quotient + content: Advanced analytics with vector data + link: quotient/ + +- icon: /courses/course-integrations/superlinked.svg + title: Superlinked + content: Advanced feature engineering for vectors + link: superlinked/ + +- icon: /courses/course-integrations/camel-ai.svg + title: Camel AI + content: Agentic RAG with multi-agent systems + link: camel/ + +- icon: /courses/course-integrations/jina.svg + title: Jina AI + content: Advanced multimodal embeddings with Qdrant + link: jina/ +{{< /cards-list >}} diff --git a/qdrant-landing/content/course/essentials/day-7/camel.md b/qdrant-landing/content/course/essentials/day-7/camel.md new file mode 100644 index 000000000..84f109102 --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/camel.md @@ -0,0 +1,78 @@ +--- +title: Integrating with Camel AI +weight: 11 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with Camel AI + +Agentic RAG with multi-agent systems using Camel AI and Qdrant. + +{{< youtube "Kz59XG_blY8" >}} + +## What You'll Learn + +- Multi-agent system architectures +- Agentic RAG patterns and best practices +- Agent collaboration and communication +- Building autonomous AI systems with Qdrant +- Auto-Retrieval with CAMEL for automated RAG processes +- Discord bot integration with vector databases + +## CAMEL Auto-Retrieval Architecture + +CAMEL (Communicative Agents for "Mind" Exploration of Large Language Model Society) provides an advanced framework for building multi-agent systems with automated RAG capabilities. The Auto-Retrieval module streamlines the process of expanding agent capabilities by automatically handling context retrieval from vector databases like Qdrant. + +### Core Concept + +Traditional agent systems require manual context management and retrieval setup. CAMEL's Auto-Retrieval approach automates this process by: + +- **Expanding Agent Capability**: Using RAG techniques to provide additional context to agents, enabling them to understand and work with internal data more effectively. +- **Automated Vector Storage**: CAMEL handles the complexity of storing and retrieving embeddings from vector databases like Qdrant. +- **Multi-Model Support**: Supporting various large language models and embedding platforms through a unified interface. +- **Real-time Integration**: Enabling seamless integration with platforms like Discord for interactive agent deployment. + +### Auto-Retrieval Process + +The CAMEL Auto-Retrieval workflow follows these key steps: + +1. **Environment Setup**: Install necessary libraries and configure your chosen large model API (supports various platforms including "gamma" and others). + +2. **Vector Database Configuration**: + - Specify Qdrant as your vector storage backend + - Provide local path for vector storage + - Choose appropriate embedding model for your use case + +3. **Automated RAG Implementation**: + - The `camel.AutoRetrieval` module handles the entire RAG process + - Automatically processes and stores document embeddings + - Manages similarity search and context retrieval + +4. **Agent Integration**: + - Retrieved information is automatically provided as context to your agent + - Works with powerful base models like "game 2.5 flash" for fast, accurate responses + - Enables agents to answer complex questions using your knowledge base + +5. **Platform Integration**: + - Deploy agents as Discord bots for real-time interaction + - Test with queries like "What is Qdrant?" and "Why do we need a vector database?" + - Agents provide accurate, context-aware responses based on your knowledge base + +### Vector Retrieval Demonstration + +When you query "What is Qdrant?" with a Qdrant website link, the system: +- Retrieves relevant content with similarity scores +- Includes metadata for context understanding +- Provides comprehensive answers based on the retrieved information +- Maintains conversation context for follow-up questions + +## Resources + +- [CAMEL Qdrant Integration](https://docs.camel-ai.org/cookbooks/applications/customer_service_Discord_bot_with_agentic_RAG#integrating-qdrant-for-large-files-to-build-a-more-powerful-discord-bot): + Official CAMEL documentation for integrating Qdrant with Discord bots and agentic RAG. Learn about Auto-Retrieval, vector storage, and building powerful customer service bots. + +- [Qdrant & CAMEL Integration Guide](https://qdrant.tech/documentation/frameworks/camel/): + Official Qdrant documentation on integrating with CAMEL-AI. Learn how to use Qdrant as a storage mechanism for ingesting and retrieving semantically similar data in your multi-agent systems. + +⭐ **Show your support!** Give CAMEL a star on their GitHub repository: [github.com/camel-ai/camel](https://github.com/camel-ai/camel) diff --git a/qdrant-landing/content/course/essentials/day-7/haystack.md b/qdrant-landing/content/course/essentials/day-7/haystack.md new file mode 100644 index 000000000..c3cd87918 --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/haystack.md @@ -0,0 +1,90 @@ +--- +title: Integrating with Haystack +weight: 3 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with Haystack + +Build end-to-end NLP pipelines with Haystack and Qdrant. + +{{< youtube "lMinhPZufTc" >}} + +## What You'll Learn + +- Haystack pipeline integration +- Document processing workflows +- Question answering systems +- Search and retrieval optimization +- Sparse vector search and metadata filtering +- LLM-based agent development +- Movie recommendation system architecture + +## Haystack Movie Recommendation Assistant + +Haystack provides a powerful framework for building sophisticated recommendation systems that combine multiple search strategies. The movie recommendation assistant demonstrates how to leverage sparse vector search, metadata filtering, and LLM-based agents to handle complex natural language queries like "find me a highly-rated action movie about car racing" or "recommend five Japanese thrillers." + +### Core Architecture + +The Haystack recommendation system uses a multi-layered approach to deliver accurate and relevant results: + +- **Sparse Vector Search**: Utilizes sparse embeddings to capture keyword-based relevance and semantic meaning +- **Metadata Filtering**: Enables precise filtering by movie attributes like genre, rating, year, and language +- **LLM-Based Agents**: Intelligent agents that can interpret complex queries and dynamically choose between search strategies +- **Qdrant Integration**: Seamless storage and retrieval of both dense and sparse vector representations + +### Implementation Workflow + +The movie recommendation system follows these key steps: + +1. **Data Preparation**: + - Convert movie data into Haystack documents with rich metadata + - Structure information including title, genre, rating, year, language, and plot descriptions + +2. **Sparse Embedding Creation**: + - Generate sparse embeddings that capture both semantic and keyword-based relevance + - Optimize embeddings for movie recommendation use cases + +3. **Qdrant Cloud Integration**: + - Write sparse embeddings and metadata to Qdrant Cloud + - Configure collections for optimal retrieval performance + - Set up proper indexing for fast metadata filtering + +4. **Query Pipeline Development**: + - Build retrieval pipelines that combine semantic search and metadata filtering + - Implement intelligent routing based on query complexity and intent + +5. **Agent Implementation**: + - Create LLM-based agents that can interpret natural language queries + - Enable dynamic strategy selection between semantic search and metadata filtering + - Implement query understanding for complex requests + +### Advanced Query Handling + +The system excels at processing sophisticated queries by: + +- **Natural Language Understanding**: Interpreting queries like "highly-rated action movie about car racing" +- **Multi-Criteria Filtering**: Combining genre, rating, and thematic requirements +- **Dynamic Strategy Selection**: Choosing between semantic search, metadata filtering, or hybrid approaches +- **Contextual Recommendations**: Providing relevant suggestions based on user preferences and movie characteristics + +### Real-World Applications + +This architecture extends beyond movie recommendations to various domains: + +- **E-commerce**: Product recommendations with complex attribute filtering +- **Content Discovery**: Finding relevant articles, videos, or resources +- **Enterprise Search**: Intelligent document retrieval with metadata constraints +- **Personalized Recommendations**: User-specific content suggestions + +## Resources + +- [Haystack Qdrant Integration](https://haystack.deepset.ai/integrations/qdrant-document-store): + Official Haystack documentation for using Qdrant as a document store. Learn about installation, usage, and connecting to Qdrant Cloud clusters. + +- [Qdrant & Haystack Integration Guide](https://qdrant.tech/documentation/frameworks/haystack/): + Official Qdrant documentation on integrating with Haystack. Learn how to build powerful NLP pipelines with vector search capabilities. + +⭐ **Show your support!** Give Haystack a star on their GitHub repository: [github.com/deepset-ai/haystack](https://github.com/deepset-ai/haystack) + diff --git a/qdrant-landing/content/course/essentials/day-7/jina.md b/qdrant-landing/content/course/essentials/day-7/jina.md new file mode 100644 index 000000000..c25baf447 --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/jina.md @@ -0,0 +1,128 @@ +--- +title: Integrating with Jina AI +weight: 12 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with Jina AI + +Advanced multimodal embeddings with Jina AI and Qdrant. + +{{< youtube "lJ7mkvHETfg" >}} + +## What You'll Learn + +- Jina Embeddings v4 model capabilities +- Multimodal text and image embeddings +- Multi-vector embeddings for enhanced performance +- API integration and self-hosting options +- Text-to-image retrieval systems +- Late chunking for long documents +- Performance optimization strategies + +## Jina AI Multimodal Embeddings + +Jina AI provides state-of-the-art deep neural networks for transforming text and images into high-quality vector representations. The Jina Embeddings v4 model represents a breakthrough in multimodal embedding technology, enabling seamless integration of text and image data within a unified vector space for sophisticated search and retrieval applications. + +### Core Architecture + +Jina AI's embedding system offers several key capabilities: + +- **Multimodal Support**: Jina Embeddings v4 supports both text and images on document and query sides +- **Unified Vector Space**: All data types embedded in the same vector space, enabling cross-modal search +- **Flexible Deployment**: API-based service with 10 million free tokens or self-hosted options +- **Multi-Vector Embeddings**: Enhanced performance for visually rich documents with multiple vector representations +- **Late Chunking**: Intelligent processing of long documents with optimized chunking strategies + +### Multimodal Search Capabilities + +The Jina Embeddings v4 model enables sophisticated search scenarios: + +- **Text-to-Text Search**: Traditional semantic search within text databases +- **Image-to-Image Search**: Visual similarity search across image libraries +- **Text-to-Image Search**: Finding images using text descriptions +- **Image-to-Text Search**: Locating relevant text content using image queries +- **Cross-Modal Retrieval**: Seamless search across mixed content types + +### Implementation Workflow + +The complete Jina AI integration with Qdrant follows these steps: + +1. **Model Selection and Configuration**: + - Choose between Jina AI API or self-hosted deployment + - Select appropriate embedding types (`retrieval.query` or `retrieval.passage`) + - Configure API parameters for optimal performance + +2. **Data Storage Process**: + - Send documents to Jina API to generate embeddings + - Process both text and image content through the embedding model + - Store documents and their corresponding embeddings in Qdrant collections + - Preserve metadata for enhanced retrieval capabilities + +3. **Query Processing**: + - Send queries to Jina API to generate query embeddings + - Support both text and image queries + - Use generated embeddings to search Qdrant database + - Retrieve relevant results with similarity scores + +4. **Multi-Vector Implementation**: + - Enable `return_multi_vector` parameter for visually rich documents + - Generate multiple vectors per document for enhanced detail capture + - Implement multi-vector storage in Qdrant collections + - Achieve 5-10% improvement in typical retrieval metrics + +### Advanced Features + +**Multi-Vector Embeddings**: +- **Enhanced Detail Capture**: Multiple vectors per document capture more nuanced information +- **Visual Content Optimization**: Dramatically improved performance for papers, graphs, and tables +- **Performance Gains**: 5-10% increase in retrieval metrics for visually rich content +- **Flexible Implementation**: Easy integration with existing Qdrant workflows + +**Late Chunking Strategy**: +- **Long Document Processing**: Intelligent handling of extended text content +- **Context Preservation**: Maintains semantic coherence across document sections +- **Optimized Chunking**: Automatic optimization for embedding model requirements +- **Scalable Processing**: Efficient handling of large document collections + +**API Customization**: +- **Flexible Configuration**: Customizable API parameters for specific use cases +- **Code Generation**: Export generated code with API parameters +- **Integration Ready**: Seamless integration with existing development workflows +- **Performance Tuning**: Optimized settings for different content types + +### Real-World Applications + +This architecture enables various sophisticated use cases: + +- **Content Discovery**: Multimodal search across text and image libraries +- **E-commerce**: Product search using both text descriptions and visual features +- **Research Platforms**: Academic paper discovery with text and figure search +- **Media Management**: Intelligent organization and retrieval of mixed media content +- **Document Intelligence**: Advanced document analysis with visual element understanding + +### Performance Optimization + +**Retrieval Enhancement**: +- **Multi-Vector Benefits**: Improved accuracy for complex visual documents +- **Cross-Modal Search**: Enhanced user experience with flexible query types +- **Scalable Architecture**: Efficient processing of large-scale multimodal datasets +- **Quality Metrics**: Measurable improvements in retrieval performance + +**Deployment Strategies**: +- **API Integration**: Quick setup with Jina AI's managed service +- **Self-Hosting**: Full control with on-premises model deployment +- **Hybrid Approaches**: Flexible deployment options for different requirements +- **Cost Optimization**: Efficient token usage with intelligent caching strategies + +## Resources + +- [Build a RAG System with Jina Embeddings and Qdrant](https://jina.ai/news/build-a-rag-system-with-jina-embeddings-and-qdrant/): + Official Jina AI guide on building RAG systems with Jina Embeddings v2 and Qdrant. Learn how to create retrieval-augmented generation engines using LlamaIndex and multimodal embeddings. + +- [Jina AI & Qdrant Integration Guide](https://qdrant.tech/documentation/embeddings/jina-embeddings/): + Official Qdrant documentation on integrating Jina AI embeddings with Qdrant. Learn how to implement multimodal search with text and image embeddings. + +⭐ **Show your support!** Give Jina AI a star on their GitHub repository: [github.com/jina-ai/jina](https://github.com/jina-ai/jina) + diff --git a/qdrant-landing/content/course/essentials/day-7/llamaindex.md b/qdrant-landing/content/course/essentials/day-7/llamaindex.md new file mode 100644 index 000000000..418ed083f --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/llamaindex.md @@ -0,0 +1,117 @@ +--- +title: Integrating with LlamaIndex +weight: 9 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with LlamaIndex + +Data framework for building LLM applications with Qdrant. + +{{< youtube "ytWskQWsAA4" >}} + +## What You'll Learn + +- Building data pipelines with LlamaIndex +- Connecting LlamaIndex to Qdrant +- Query engines and retrieval strategies +- Advanced RAG patterns with LlamaIndex +- Agent workflows and function calling +- Custom workflow development with events and steps +- Qdrant Cloud integration with Llama Cloud +- Real-world data ingestion and processing + +## LlamaIndex Agent Development Framework + +LlamaIndex provides a comprehensive framework for building sophisticated LLM applications with Qdrant integration. The platform supports multiple deployment options including local Qdrant instances, Qdrant Cloud, and Llama Cloud with Qdrant Cloud synchronization, enabling flexible and scalable agent development. + +### Core Architecture + +LlamaIndex's agent framework offers several key components: + +- **Multi-Deployment Support**: Local Qdrant, Qdrant Cloud, and Llama Cloud integration options +- **Agent Workflows**: Custom workflow development with events and steps +- **Function Agents**: LLM-powered function calling for dynamic operations +- **Sparse Embeddings**: Support for both dense and sparse vector representations +- **Real-World Data Integration**: Comprehensive data loaders and readers from Llama Hub + +### Implementation Workflow + +The complete LlamaIndex integration follows these key steps: + +1. **Environment Setup**: + - Install required dependencies and libraries + - Configure environment variables for OpenAI and Llama Cloud API keys + - Set up Qdrant client connections + +2. **Basic Agent Development**: + - Create agents capable of querying databases and writing new statements + - Set up Qdrant client and vector store configuration + - Define sparse embedding models for optimal performance + - Create document chunks (nodes) for processing + +3. **Basic RAG Implementation**: + - Perform simple retrieval-augmented generation queries + - Use vector index with document nodes + - Implement basic semantic search capabilities + +4. **Function Agent Development**: + - Build function agents with Python functions: `write_statement` and `query_collection` + - Enable LLM decision-making for function selection based on user queries + - Implement dynamic function calling based on query intent + +5. **Custom Workflow Creation**: + - Develop custom agent workflows using LlamaIndex's events and steps + - Classify incoming queries into "save to docs" or "ask" actions + - Use Pydantic models and structured LLMs for query classification + - Trigger appropriate events based on query classification + +6. **Real-World Data Integration**: + - Utilize data loaders and readers from Llama Hub + - Ingest real-world data (web pages, documents) into Qdrant indexes + - Process and embed various data formats + +7. **Qdrant Cloud with Llama Cloud**: + - Demonstrate Qdrant Cloud integration via Llama Cloud + - Show indexes using Qdrant as vector database with Google Drive as data source + - Leverage Llama Parse for advanced document parsing + +### Advanced Features + +**Agent Workflow Capabilities**: +- **Event-Driven Architecture**: Custom workflows using events and steps +- **Query Classification**: Intelligent routing of queries to appropriate handlers +- **Structured LLM Integration**: Pydantic models for reliable data processing +- **Dynamic Function Selection**: LLM-powered decision making for function calls + +**Multi-Modal Data Support**: +- **Web Content**: Crawling and processing web pages +- **Document Processing**: Advanced parsing with Llama Parse +- **Cloud Integration**: Seamless Google Drive and cloud storage integration +- **Real-Time Processing**: Live data ingestion and embedding + +**Deployment Flexibility**: +- **Local Development**: Qdrant running locally for development and testing +- **Cloud Production**: Qdrant Cloud for scalable production deployments +- **Hybrid Solutions**: Llama Cloud with Qdrant Cloud synchronization + +### Real-World Applications + +This architecture enables various sophisticated use cases: + +- **Knowledge Management**: Building intelligent knowledge bases with dynamic content updates +- **Document Processing**: Advanced document parsing and search systems +- **Customer Support**: AI-powered support agents with real-time knowledge access +- **Research Tools**: Intelligent research assistants with multi-source data integration + +## Resources + +- [LlamaIndex Qdrant Integration](https://docs.llamaindex.ai/en/stable/examples/vector_stores/qdrant_hybrid/): + Official LlamaIndex documentation for using Qdrant as a vector store. Learn about hybrid search, vector storage configuration, and query examples. + +- [Qdrant & LlamaIndex Integration Guide](https://qdrant.tech/documentation/frameworks/llama-index/): + Official Qdrant documentation on integrating with LlamaIndex. Learn how to build sophisticated RAG applications and AI agents with function calling capabilities. + +⭐ **Show your support!** Give LlamaIndex a star on their GitHub repository: [github.com/run-llama/llama_index](https://github.com/run-llama/llama_index) + diff --git a/qdrant-landing/content/course/essentials/day-7/quotient.md b/qdrant-landing/content/course/essentials/day-7/quotient.md new file mode 100644 index 000000000..9111c8aed --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/quotient.md @@ -0,0 +1,124 @@ +--- +title: Integrating with Quotient +weight: 10 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with Quotient + +Advanced analytics with vector data using Quotient platform. + +{{< youtube "QeQuCsh1SHs" >}} + +## What You'll Learn + +- Analytics platform integration +- Vector data analysis techniques +- Business intelligence applications +- Reporting and visualization +- RAG monitoring and quality assurance +- AI application monitoring and debugging +- Hallucination detection and document relevance scoring + +## Quotient AI Monitoring Platform + +Quotient AI provides critical monitoring capabilities for AI applications and agents, automatically detecting quality issues and providing comprehensive insights into system performance. The platform serves as an essential monitoring layer for RAG (Retrieval Augmented Generation) applications, helping maintain reliability and enabling effective debugging of AI systems. + +### Core Architecture + +The Quotient AI monitoring platform offers several key components: + +- **Automatic Quality Monitoring**: Continuously monitors agent performance and response quality +- **Hallucination Detection**: Identifies when AI responses contain fabricated or incorrect information +- **Document Relevance Scoring**: Evaluates how well retrieved documents match user queries +- **Monitoring Dashboards**: Provides visual insights into system performance and quality metrics +- **Root Cause Analysis**: Offers tools to identify and debug issues in AI applications + +### Documentation Q&A Agent with Monitoring + +The demonstration showcases building a comprehensive documentation Q&A agent with built-in monitoring using a modular stack: + +**Core Components**: +- **Tavilli**: Crawling and extracting documentation from company websites +- **Qdrant**: Fast semantic search over document chunks +- **LangChain**: Orchestrating document processing, retrieval, and LLM calls +- **OpenAI**: Generating answers and creating high-quality embeddings +- **Quotient AI**: Critical monitoring layer for quality assurance + +### Implementation Workflow + +The complete system follows these steps: + +1. **Environment Setup**: + - Configure API keys for OpenAI, Tavilli, and Quotient AI + - Install necessary libraries and dependencies + - Set up monitoring configurations + +2. **Core Component Configuration**: + - Qdrant as vector store for document embeddings + - OpenAI for embeddings (`text-embedding-3-small`) and answer generation (`GPT-4o`) + - Quotient for monitoring with hallucination detection and document relevance scoring enabled by default + +3. **Content Extraction and Processing**: + - Tavilli extracts real text content from live documentation websites + - Crawls and collects relevant URLs automatically + - Processes structured content for embedding + +4. **Document Processing Pipeline**: + - LangChain's recursive character text splitter creates overlapping 700-token chunks + - 50-token overlap preserves context across chunks + - Optimized chunking strategy for embedding models + +5. **Embedding and Storage**: + - Chunked documents are embedded using OpenAI's embedding model + - Embeddings and metadata stored in Qdrant collections + - Optimized for fast retrieval and similarity search + +6. **RAG Chain Implementation**: + - Orchestrates question-answering process + - Retrieves relevant chunks based on user queries + - Formats retrieved content as context for answer generation + - Constrains answers to use only relevant documentation + +7. **Monitoring Integration**: + - Every interaction logged to Quotient AI + - Asynchronous detection pipelines identify potential hallucinations + - Document relevance scoring provides retrieval quality insights + - Real-time monitoring dashboards track system performance + +### Key Monitoring Features + +**Hallucination Detection**: +- Automatically identifies fabricated or incorrect information in AI responses +- Provides confidence scores for response accuracy +- Enables proactive quality control + +**Document Relevance Scoring**: +- Evaluates how well retrieved documents match user queries +- Identifies retrieval quality issues +- Helps optimize document chunking and embedding strategies + +**Performance Dashboards**: +- Visual insights into system performance metrics +- Quality trends and anomaly detection +- Root cause analysis tools for debugging + +### Real-World Applications + +This monitoring architecture enables various enterprise use cases: + +- **Enterprise Documentation**: Monitoring Q&A systems for technical documentation +- **Customer Support**: Quality assurance for AI-powered customer service agents +- **Content Management**: Monitoring AI systems that process and respond to content queries +- **Research Applications**: Ensuring accuracy in AI-powered research and analysis tools + +## Resources + +- [Optimizing RAG Through an Evaluation-Based Methodology](https://qdrant.tech/articles/rapid-rag-optimization-with-qdrant-and-quotient/): + Learn how to optimize RAG systems using Qdrant and Quotient through systematic evaluation. Covers experimentation with chunking, retrieval strategies, and model selection. + +- [Building High-Quality RAG Applications with Qdrant and Quotient](https://blog.quotientai.co/building-high-quality-rag-applications-with-qdrant-and-quotient/): + Comprehensive guide on building production-ready RAG applications with comprehensive monitoring using Quotient AI and Qdrant for quality assurance and performance tracking. + +**Note**: Visit [quotientai.co](https://www.quotientai.co/) for more information. \ No newline at end of file diff --git a/qdrant-landing/content/course/essentials/day-7/superlinked.md b/qdrant-landing/content/course/essentials/day-7/superlinked.md new file mode 100644 index 000000000..5733480d5 --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/superlinked.md @@ -0,0 +1,42 @@ +--- +title: Integrating with Superlinked +weight: 8 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with Superlinked + +Advanced feature engineering for vector search applications. + +{{< youtube "55iDpaHwKJo" >}} + +## What You'll Learn + +- Advanced feature engineering techniques +- Vector space optimization +- Multi-modal data handling +- Performance enhancement strategies + +## Mixture of Encoders Architecture in Superlinked + +The Mixture of Encoders architecture is Superlinked's modular system for combining multiple data-specific embedding models into one unified representation. It creates metadata-aware vector embeddings that integrate signals from text, images, popularity, user interaction, numbers, categories, and time, producing richer and more accurate results for search, retrieval, and recommendation tasks. + +### Core Concept + +Traditional embedding systems rely on a single model to handle all types of input. Superlinked's Mixture of Encoders takes a different approach by letting users define multiple spaces, each powered by an encoder specialized for a particular data type. + +- **Text encoders** (for example, sentence transformers or LLMs) capture semantic meaning from descriptions or notes. +- **Numerical encoders** represent quantitative metrics such as ratings, prices, or counts. +- **Categorical encoders** handle tags, IDs, or labels to model discrete entities. +- **Temporal encoders** learn patterns of recency, popularity, frequency, and event timing to capture how relevance changes over time. + +Each encoder produces its own vector representation that reflects its data domain. These embeddings are then merged into a composite embedding, forming a single, context-rich representation that captures multiple dimensions of meaning. + + +## Resources + +- [Superlinked & Qdrant Integration Guide](https://qdrant.tech/documentation/frameworks/superlinked/): + Official Qdrant documentation on integrating Superlinked with Qdrant. Learn how to build advanced vector search applications with multiple encoder types for text, images, numbers, categories, and temporal data. + +⭐ **Show your support!** Give Superlinked a star on their GitHub repository: [github.com/superlinked/superlinked](https://github.com/superlinked/superlinked) \ No newline at end of file diff --git a/qdrant-landing/content/course/essentials/day-7/tensorlake.md b/qdrant-landing/content/course/essentials/day-7/tensorlake.md new file mode 100644 index 000000000..7a4615fae --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/tensorlake.md @@ -0,0 +1,116 @@ +--- +title: Integrating with Tensorlake +weight: 7 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with TensorLake + +Build scalable data lakes with vector search capabilities using TensorLake's advanced document parsing techniques. + +{{< youtube "IVfoVS0KfPM" >}} + +## What You'll Learn + +- Data lake architecture with vectors +- Large-scale data management +- Analytics and vector search integration +- ETL pipeline optimization +- Knowledge graph creation from unstructured documents +- Document parsing and structured data extraction +- LangGraph agent integration for natural language querying + +## TensorLake Knowledge Graph Integration + +TensorLake introduces an innovative approach to enhancing Qdrant collection querying through advanced document parsing and knowledge graph creation. The platform transforms unstructured documents into structured knowledge graphs, providing comprehensive data extraction and intelligent summarization of complex tables and figures, leading to more accurate embeddings and fine-tuned searches in RAG applications. + +### Core Architecture + +TensorLake's document parsing engine provides several key capabilities: + +- **Knowledge Graph Creation**: Transforms unstructured documents into structured knowledge graphs with preserved relationships +- **Document Layout Preservation**: Maintains reading order and groups related content like authors and references +- **Table and Figure Summarization**: Creates intelligent summaries of complex tables and figures for semantic searchability +- **Structured Data Extraction**: Extracts metadata including titles, authors, conferences, keywords, and references + +### Academic Research Paper Processing + +The demonstration showcases TensorLake's capabilities using academic research papers: + +1. **Document Parsing**: + - TensorLake's engine preserves reading order and document structure + - Groups authors and creates complete document layout + - Extracts structured metadata from research papers + +2. **Knowledge Graph Generation**: + - Creates comprehensive knowledge graphs from parsed documents + - Maintains relationships between authors, institutions, and references + - Preserves hierarchical document structure + +3. **Table and Figure Processing**: + - Summarizes complex tables and figures for embedding + - Makes large tables semantically searchable + - Enables fine-grained queries on tabular data + +### Qdrant Integration Workflow + +The complete integration process follows these steps: + +1. **Document Processing with TensorLake**: + - Parse documents to extract structured information + - Generate knowledge graphs with preserved relationships + - Create summaries of tables and figures + +2. **Embedding Creation and Storage**: + - Create embeddings from processed document content + - Generate detailed payloads including title, authors, conference, keywords, and references + - Upsert embeddings and metadata into Qdrant collections + +3. **Index Creation**: + - Create indices for easier filtering by metadata attributes + - Optimize collections for both semantic search and metadata filtering + - Enable efficient querying across different data types + +4. **LangGraph Agent Integration**: + - Implement natural language querying capabilities + - Enable intelligent filtering based on query context + - Provide summaries of relevant document sections + +### Advanced Query Capabilities + +The system supports multiple query types: + +- **Simple Semantic Queries**: Basic vector similarity search without filtering +- **Filtered Queries**: Search by specific authors, conferences, or other metadata +- **Combined Queries**: Semantic search with metadata filtering for precise results +- **Natural Language Queries**: LangGraph agent interprets complex questions and applies appropriate filters + +### Key Benefits + +TensorLake's integration with Qdrant provides several advantages: + +- **Enhanced Accuracy**: More accurate embeddings through structured data extraction +- **Complete Document Understanding**: Preserves document hierarchy and relationships +- **Fine-tuned Search**: Enables precise queries combining semantic and metadata filtering +- **Robust Collections**: More complete and reliable query results +- **Natural Language Interface**: Intuitive querying through LangGraph agents + +### Real-World Applications + +This architecture enables various advanced use cases: + +- **Research Discovery**: Finding relevant academic papers with complex criteria +- **Legal Document Analysis**: Processing contracts and legal documents with structured extraction +- **Technical Documentation**: Creating searchable knowledge bases from technical manuals +- **Enterprise Knowledge Management**: Building comprehensive search systems for large document collections + +## Resources + +- [TensorLake Qdrant Integration](https://docs.tensorlake.ai/integrations/qdrant#qdrant): + Official TensorLake documentation for integrating with Qdrant. Learn about document parsing, knowledge graph creation, and structured data extraction for enhanced RAG applications. + +- [Qdrant & TensorLake Integration Guide](https://www.tensorlake.ai/blog/announcing-qdrant-tensorlake): + Explore how combining TensorLake's document parsing capabilities with Qdrant's vector search enhances RAG applications with structured filters and semantic search. + +**Note**: Visit [tensorlake.ai](https://www.tensorlake.ai/) for more information. diff --git a/qdrant-landing/content/course/essentials/day-7/unstructured.md b/qdrant-landing/content/course/essentials/day-7/unstructured.md new file mode 100644 index 000000000..ceaad1e61 --- /dev/null +++ b/qdrant-landing/content/course/essentials/day-7/unstructured.md @@ -0,0 +1,99 @@ +--- +title: Integrating with Unstructured.io +weight: 6 +--- + +{{< date >}} Day 7 {{< /date >}} + +# Integrating with Unstructured.io + +Process and vectorize documents with Unstructured.io and Qdrant. + +{{< youtube "FRIOwOy6VZk" >}} + +## What You'll Learn + +- Document processing with Unstructured.io +- Multi-format document ingestion +- Structured data extraction +- Automated vectorization pipelines +- Enterprise data transformation workflows +- VLM-powered document understanding +- Production-ready ETL pipelines + +## Unstructured Enterprise Data Processing + +Unstructured.io addresses the critical challenge of processing unstructured enterprise data, which typically accounts for 80% of enterprise information. The platform provides a composable solution to transform PDFs, Word documents, emails, and other unstructured formats into structured outputs optimized for GenAI initiatives, eliminating the complexity of custom scripts and tools. + +### Core Architecture + +Unstructured.io's enterprise-grade platform offers several key components: + +- **VLM Partitioner**: Uses GPT-4o vision for intelligent document understanding, recognizing layout, tables, images, and text hierarchy +- **Smart Chunker**: Employs `chunk_by_title` strategy for semantic coherence, creating optimized chunks for embeddings +- **Vector Embedder**: Uses OpenAI's text embedding 3 small model to create embeddings that understand semantic meaning +- **Enterprise Security**: Production-ready, scalable, and composable pipeline with enterprise security and compliance + +### End-to-End Workflow + +The complete workflow demonstrates seamless integration between AWS S3, Unstructured, and Qdrant: + +1. **Data Ingestion**: + - Raw content stored in AWS S3 buckets + - Unstructured's S3 connector pulls documents into the SaaS platform + - Supports multiple document formats (PDF, Word, emails, etc.) + +2. **Document Processing**: + - Documents are partitioned into structured elements using VLM technology + - Intelligent recognition of layout, tables, images, and text hierarchy + - Smart chunking strategy preserves semantic coherence + +3. **Vector Generation**: + - Chunked documents are embedded using OpenAI's text embedding 3 small model + - Dense vector representations capture semantic meaning + - Rich metadata preserved including document hierarchy and element classification + +4. **Qdrant Integration**: + - Qdrant connector deposits processed data into vector database + - One uploaded file can result in 43 points in Qdrant due to chunking + - Metadata includes document hierarchy, element classification, and sample text + +### Key Features + +**VLM Partitioner Capabilities**: +- Intelligent document layout recognition +- Table and figure extraction with context preservation +- Image understanding and text hierarchy analysis +- Multi-modal content processing + +**Smart Chunking Strategy**: +- `chunk_by_title` approach maintains semantic coherence +- Optimized chunk sizes for embedding models +- Context preservation across document sections +- Hierarchical document structure maintenance + +**Enterprise-Grade Processing**: +- Production-ready scalability +- Enterprise security and compliance features +- Automated ETL pipeline management +- Reduced complexity compared to custom solutions + +### Real-World Applications + +This architecture enables various enterprise use cases: + +- **Document Intelligence**: Automated processing of legal documents, contracts, and reports +- **Knowledge Management**: Enterprise-wide document search and retrieval systems +- **Compliance**: Automated document analysis for regulatory requirements +- **RAG Applications**: Enhanced retrieval for AI applications with structured document understanding + +## Resources + +- [Unstructured Qdrant Destination](https://docs.unstructured.io/ui/destinations/qdrant): + Official Unstructured documentation for sending processed data to Qdrant. Learn about Qdrant Cloud integration, collection setup, and workflow configuration. + +- [Qdrant & Unstructured Integration Guide](https://qdrant.tech/documentation/frameworks/unstructured/): + Official Qdrant documentation for Unstructured.io integration, covering setup and best practices for document processing pipelines. + +⭐ **Show your support!** Give Unstructured a star on their GitHub repository: [github.com/Unstructured-IO/unstructured](https://github.com/Unstructured-IO/unstructured) + diff --git a/qdrant-landing/content/course/essentials/day-9/_index.md b/qdrant-landing/content/course/essentials/day-9/_index.md deleted file mode 100644 index e08cb2a85..000000000 --- a/qdrant-landing/content/course/essentials/day-9/_index.md +++ /dev/null @@ -1,59 +0,0 @@ ---- -title: Day 9 (Bonus) -isLesson: true -weight: 9 ---- - -{{< date >}} Day 9 {{< /date >}} - -# Bonus: Advanced Configurations - -{{< cards-list >}} -- icon: /courses/course-integrations/crew-ai.svg - title: Building Agents with CrewAI and Qdrant - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/haystack.svg - title: Haystack - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/jina.svg - title: Jina - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/n8n.svg - title: n8n - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/camel-ai.svg - title: Camel-AI - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/tensorlake.svg - title: Tensorlake - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/vectorize.svg - title: Vectorize.io - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/unstructured.svg - title: Unstructured.io - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/quotient.svg - title: Quotient - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/superlinked.svg - title: Superlinked - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/twelveLabs.svg - title: TwelveLabs - content: Qdrant is compatible with Cohere co.embed API. - -- icon: /courses/course-integrations/aparavi.svg - title: APARAVI - content: Qdrant is compatible with Cohere co.embed API. -{{< /cards-list >}} diff --git a/qdrant-landing/static/courses/course-integrations/llamaindex.svg b/qdrant-landing/static/courses/course-integrations/llamaindex.svg new file mode 100755 index 000000000..f238b8097 --- /dev/null +++ b/qdrant-landing/static/courses/course-integrations/llamaindex.svg @@ -0,0 +1,12 @@ + \ No newline at end of file diff --git a/qdrant-landing/themes/qdrant-2024/layouts/shortcodes/cards-list.html b/qdrant-landing/themes/qdrant-2024/layouts/shortcodes/cards-list.html index d43bace77..d75f0ee2f 100644 --- a/qdrant-landing/themes/qdrant-2024/layouts/shortcodes/cards-list.html +++ b/qdrant-landing/themes/qdrant-2024/layouts/shortcodes/cards-list.html @@ -1,12 +1,22 @@ {{ $data := .Inner | transform.Unmarshal }}