ClaudeCodeMod

All shelves / MCP servers

Crawl4AI RAG

coleam00/mcp-crawl4ai-rag · 2.2k stars · Python · MIT

MCP server Web Crawling and RAG Capabilities for AI Agents and AI Coding Assistants

Install

In your shell
docker run --env-file .env -p 8051:8051 mcp/crawl4ai-rag

These repos do not share one command. When an entry shows a command, it was copied as published. Check the repo's README before you run it.

Open the repo

Files

README.md

A powerful implementation of the Model Context Protocol (MCP) integrated with Crawl4AI and Supabase for providing AI agents and AI coding assistants with advanced web crawling and RAG capabilities.

With this MCP server, you can scrape anything and then use that knowledge anywhere for RAG.

The primary goal is to bring this MCP server into Archon as I evolve it to be more of a knowledge engine for AI coding assistants to build AI agents. This first version of the Crawl4AI/RAG MCP server will be improved upon greatly soon, especially making it more configurable so you can use different embedding models and run everything locally with Ollama.

Consider this GitHub repository a testbed, hence why I haven't been super actively address issues and pull requests yet. I certainly will though as I bring this into Archon V2!

Overview

This MCP server provides tools that enable AI agents to crawl websites, store content in a vector database (Supabase), and perform RAG over the crawled content. It follows the best practices for building MCP servers based on the Mem0 MCP server template I provided on my channel previously.

The server includes several advanced RAG strategies that can be enabled to enhance retrieval quality:

  • Contextual Embeddings for enriched semantic understanding
  • Hybrid Search combining vector and keyword search
  • Agentic RAG for specialized code example extraction
  • Reranking for improved result relevance using cross-encoder models
  • Knowledge Graph for AI hallucination detection and repository code analysis

See the Configuration section below for details on how to enable and configure these strategies.

Vision

The Crawl4AI RAG MCP server is just the beginning. Here's where we're headed:

  1. Integration with Archon: Building this system directly into Archon to create a comprehensive knowledge engine for AI coding assistants to build better AI agents.
  1. Multiple Embedding Models: Expanding beyond OpenAI to support a variety of embedding models, including the ability to run everything locally with Ollama for complete control and privacy.
  1. Advanced RAG Strategies: Implementing sophisticated retrieval techniques like contextual retrieval, late chunking, and others to move beyond basic "naive lookups" and significantly enhance the power and precision of the RAG system, especially as it integrates with Archon.
  1. Enhanced Chunking Strategy: Implementing a Context 7-inspired chunking approach that focuses on examples and creates distinct, semantically meaningful sections for each chunk, improving retrieval precision.
  1. Performance Optimization: Increasing crawling and indexing speed to make it more realistic to "quickly" index new documentation to then leverage it within the same prompt in an AI coding assistant.

Features

  • Smart URL Detection: Automatically detects and handles different URL types (regular webpages, sitemaps, text files)
  • Recursive Crawling: Follows internal links to discover content
  • Parallel Processing: Efficiently crawls multiple pages simultaneously
  • Content Chunking: Intelligently splits content by headers and size for better processing
  • Vector Search: Performs RAG over crawled content, optionally filtering by data source for precision
  • Source Retrieval: Retrieve sources available for filtering to guide the RAG process

Tools

The server provides essential web crawling and search tools:

Core Tools (Always Available)

  1. crawl_single_page: Quickly crawl a single web page and store its content in the vector database
  2. smart_crawl_url: Intelligently crawl a full website based on the type of URL provided (sitemap, llms-full.txt, or a regular webpage that needs to be crawled recursively)
  3. get_available_sources: Get a list of all available sources (domains) in the database
  4. perform_rag_query: Search for relevant content using semantic search with optional source filtering

Conditional Tools

  1. search_code_examples (requires USE_AGENTIC_RAG=true): Search specifically for code examples and their summaries from crawled documentation. This tool provides targeted code snippet retrieval for AI coding assistants.

Knowledge Graph Tools (requires USE_KNOWLEDGE_GRAPH=true, see below)

  1. parse_github_repository: Parse a GitHub repository into a Neo4j knowledge graph, extracting classes, methods, functions, and their relationships for hallucination detection
  2. check_ai_script_hallucinations: Analyze Python scripts for AI hallucinations by validating imports, method calls, and class usage against the knowledge graph
  3. query_knowledge_graph: Explore and query the Neo4j knowledge graph with commands like repos, classes, methods, and custom Cypher queries

Prerequisites

Installation

Using Docker (Recommended)

  1. Clone this repository:
   git clone https://github.com/coleam00/mcp-crawl4ai-rag.git
   cd mcp-crawl4ai-rag
  1. Build the Docker image:
   docker build -t mcp/crawl4ai-rag --build-arg PORT=8051 .
  1. Create a .env file based on the configuration section below

Using uv directly (no Docker)

  1. Clone this repository:
   git clone https://github.com/coleam00/mcp-crawl4ai-rag.git
   cd mcp-crawl4ai-rag
  1. Install uv if you don't have it:
   pip install uv
  1. Create and activate a virtual environment:
   uv venv
   .venv\Scripts\activate
   # on Mac/Linux: source .venv/bin/activate
  1. Install dependencies:
   uv pip install -e .
   crawl4ai-setup
  1. Create a .env file based on the configuration section below

Database Setup

Before running the server, you need to set up the database with the pgvector extension:

  1. Go to the SQL Editor in your Supabase dashboard (create a new project first if necessary)
  1. Create a new query and paste the contents of crawled_pages.sql
  1. Run the query to create the necessary tables and functions

Knowledge Graph Setup (Optional)

To enable AI hallucination detection and repository analysis features, you need to set up Neo4j.

Also, the knowledge graph implementation isn't fully compatible with Docker yet, so I would recommend right now running directly through uv if you want to use the hallucination detection within the MCP server!

For installing Neo4j:

Local AI Package (Recommended)

The easiest way to get Neo4j running locally is with the Local AI Package - a curated collection of local AI services including Neo4j:

  1. Clone the Local AI Package:
   git clone https://github.com/coleam00/local-ai-packaged.git
   cd local-ai-packaged
  1. Start Neo4j:

Follow the instructions in the Local AI Package repository to start Neo4j with Docker Compose

  1. Default connection details:
  • URI: bolt://localhost:7687
  • Username: neo4j
  • Password: Check the Local AI Package documentation for the default password

Manual Neo4j Installation

Alternatively, install Neo4j directly:

  1. Install Neo4j Desktop: Download from neo4j.com/download
  1. Create a new database:
  • Open Neo4j Desktop
  • Create a new project and database
  • Set a password for the neo4j user
  • Start the database
  1. Note your connection details:
  • URI: bolt://localhost:7687 (default)
  • Username: neo4j (default)
  • Password: Whatever you set during creation

Configuration

Create a .env file in the project root with the following variables:

# MCP Server Configuration
HOST=0.0.0.0
PORT=8051
TRANSPORT=sse

# OpenAI API Configuration
OPENAI_API_KEY=your_openai_api_key

# LLM for summaries and contextual embeddings
MODEL_CHOICE=gpt-4.1-nano

# RAG Strategies (set to "true" or "false", default to "false")
USE_CONTEXTUAL_EMBEDDINGS=false
USE_HYBRID_SEARCH=false
USE_AGENTIC_RAG=false
USE_RERANKING=false
USE_KNOWLEDGE_GRAPH=false

# Supabase Configuration

Facts

Kind
MCP server
Repo
coleam00/mcp-crawl4ai-rag
Group
Uncategorized
Stars
2.2k
License
MIT
Language
Python
Last push
2025-07-25
Forks
575

More on this shelf

  1. 1Everythingmodelcontextprotocol/serversThis MCP server attempts to exercise all the features of the MCP protocol. It is not intended to be a useful server, but rather a test server for builders of MCP clients. It implements prompts, tools, resources, sampling, and more to showcase MCP capabilities.85.8k
  2. 2Fetchmodelcontextprotocol/serversA Model Context Protocol server that provides web content fetching capabilities. This server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.85.8k
  3. 3Gitmodelcontextprotocol/serversA Model Context Protocol server for Git repository interaction and automation. This server provides tools to read, search, and manipulate Git repositories via Large Language Models.85.8k
  4. 4Memorymodelcontextprotocol/serversA basic implementation of persistent memory using a local knowledge graph. This lets Claude remember information about the user across chats.85.8k
  5. 5Sequential Thinkingmodelcontextprotocol/serversAn MCP server implementation that provides a tool for dynamic and reflective problem-solving through a structured thinking process.85.8k
  6. 6Timemodelcontextprotocol/serversA Model Context Protocol server that provides time and timezone conversion capabilities. This server enables LLMs to get current time information and perform timezone conversions using IANA timezone names, with automatic system timezone detection.85.8k