diff --git a/qdrant-landing/content/documentation/frameworks/_index.md b/qdrant-landing/content/documentation/frameworks/_index.md index 91132561c..195cdce83 100644 --- a/qdrant-landing/content/documentation/frameworks/_index.md +++ b/qdrant-landing/content/documentation/frameworks/_index.md @@ -7,6 +7,7 @@ weight: 33 | ------------------------------------- | ---------------------------------------------------------------------------------------------------- | | [Airbyte](./airbyte/) | Data integration platform specialising in ELT pipelines. | | [Airflow](./airflow/) | Platform designed for developing, scheduling, and monitoring batch-oriented workflows. | +| [Apify](./apify/) | Platform to build web scrapers and automate web browser tasks. | | [AutoGen](./autogen/) | Framework from Microsoft building LLM applications using multiple conversational agents. | | [Bubble](./bubble) | Development platform for application development with a no-code interface | | [Canopy](./canopy/) | Framework from Pinecone for building RAG applications using LLMs and knowledge bases. | diff --git a/qdrant-landing/content/documentation/frameworks/apify.md b/qdrant-landing/content/documentation/frameworks/apify.md new file mode 100644 index 000000000..64cedc2cf --- /dev/null +++ b/qdrant-landing/content/documentation/frameworks/apify.md @@ -0,0 +1,82 @@ +--- +title: Apify +weight: 3600 +--- + +# Apify + +[Apify](https://apify.com/) is a web scraping and browser automation platform featuring an [app store](https://apify.com/store) with over 1,500 pre-built micro-apps known as Actors. These serverless cloud programs, which are essentially dockers under the hood, are designed for various web automation applications, including data collection. + +One such Actor, built especially for AI and RAG applications, is [Website Content Crawler](https://apify.com/apify/website-content-crawler). + +It's ideal for this purpose because it has built-in HTML processing and data-cleaning functions. That means you can easily remove fluff, duplicates, and other things on a web page that aren't relevant, and provide only the necessary data to the language model. + +The Markdown can then be used to feed Qdrant to train AI models or supply them with fresh web content. + +Qdrant is available as an [official integration](https://apify.com/apify/qdrant-integration) to load Apify datasets into a collection. + +You can refer to the [Apify documentation](https://docs.apify.com/platform/integrations/qdrant) to set up the integration via the Apify UI. + +## Programmatic Usage + +Apify also supports programmatic access to integrations via the [Apify Python SDK](https://docs.apify.com/sdk/python/). + +1. Install the Apify Python SDK by running the following command: + + ```sh + pip install apify-client + ``` + +2. Create a Python script and import all the necessary modules: + + ```python + from apify_client import ApifyClient + + APIFY_API_TOKEN = "YOUR-APIFY-TOKEN" + OPENAI_API_KEY = "YOUR-OPENAI-API-KEY" + # COHERE_API_KEY = "YOUR-COHERE-API-KEY" + + QDRANT_URL = "YOUR-QDRANT-URL" + QDRANT_API_KEY = "YOUR-QDRANT-API-KEY" + + client = ApifyClient(APIFY_API_TOKEN) + ``` + +3. Call the [Website Content Crawler](https://apify.com/apify/website-content-crawler) Actor to crawl the Qdrant documentation and extract text content from the web pages: + + ```python + actor_call = client.actor("apify/website-content-crawler").call( + run_input={"startUrls": [{"url": "https://qdrant.tech/documentation/"}]} + ) + ``` + +4. Call the Qdrant integration and store all data in the Qdrant Vector Database: + + ```python + qdrant_integration_inputs = { + "qdrantUrl": QDRANT_URL, + "qdrantApiKey": QDRANT_API_KEY, + "qdrantCollectionName": "apify", + "qdrantAutoCreateCollection": True, + "datasetId": actor_call["defaultDatasetId"], + "datasetFields": ["text"], + "enableDeltaUpdates": True, + "deltaUpdatesPrimaryDatasetFields": ["url"], + "expiredObjectDeletionPeriodDays": 30, + "embeddingsProvider": "OpenAI", # "Cohere" + "embeddingsApiKey": OPENAI_API_KEY, + "performChunking": True, + "chunkSize": 1000, + "chunkOverlap": 0, + } + actor_call = client.actor("apify/qdrant-integration").call(run_input=qdrant_integration_inputs) + + ``` + +Upon running the script, the data from will be scraped, transformed into vector embeddings and stored in the Qdrant collection. + +## Further Reading + +- Apify [Documentation](https://docs.apify.com/) +- Apify [Templates](https://apify.com/templates) +- Integration [Source Code](https://github.com/apify/actor-vector-database-integrations) diff --git a/qdrant-landing/static/documentation/frameworks/apify-social-preview.png b/qdrant-landing/static/documentation/frameworks/apify-social-preview.png new file mode 100644 index 000000000..e1def1189 Binary files /dev/null and b/qdrant-landing/static/documentation/frameworks/apify-social-preview.png differ