MCP Documentation Server
andrea9293/mcp-documentation-server · 313 stars · TypeScript · MIT
MCP server MCP Documentation Server - Bridge the AI Knowledge Gap. ✨ Features: Document management • Gemini integration • AI-powered semantic search • File uploads • Smart chunking • Multilingual support • Zero-setup 🎯 Perfect for: New frameworks • API docs • Internal guides
Install
npx skills add https://github.com/andrea9293/mcp-documentation-server --skill documentation-serverThese repos do not share one command. When an entry shows a command, it was copied as published. Check the repo's README before you run it.
Files
MCP Documentation Server
Local-first document management and semantic search for AI coding agents. No external databases, no cloud APIs, no vendor lock-in.
Unlike other MCP servers that are CLI-only, this one ships with a full web dashboard — browse, search, upload, and manage your knowledge base from your browser. Every MCP tool is also exposed as a REST API, giving AI agents a lean, schema-free interface.
- 🏠 Runs fully offline — Orama vector DB with local AI embeddings (Transformers.js)
- 🌐 Built-in Web UI — starts automatically on port 3080 alongside the MCP server
- 🔍 Hybrid search — full-text + vector similarity with parent-child chunking
- 🤖 Optional AI search — Google Gemini for advanced document analysis (bring your own key)
- 📁 Drag & drop uploads —
.txt,.md,.pdfsupport - 📦 Published on the MCP Registry — installable via npx, no clone needed
Quick Start
{
"mcpServers": {
"documentation": {
"command": "npx",
"args": ["-y", "@andrea9293/mcp-documentation-server"]
}
}
}
Open your browser at http://localhost:3080 — the web UI starts automatically.
🤖 Agent Skill (REST API) — recommended for AI agents
Every MCP tool is also accessible via the REST API on http://127.0.0.1:3080/api/. This is the recommended way to interact from AI agents (Claude Code, OpenCode, Gemini CLI, Cursor) because it avoids loading MCP tool schemas into the conversation context — only the response JSON enters.
curl -s http://127.0.0.1:3080/api/config
curl -s http://127.0.0.1:3080/api/documents
curl -s -X POST http://127.0.0.1:3080/api/search-all \
-H "Content-Type: application/json" \
-d '{"query": "your search", "limit": 5}'
A ready-to-use skill is included at skills/documentation-server/SKILL.md — it teaches your agent every endpoint with examples. Install it:
npx skills add https://github.com/andrea9293/mcp-documentation-server --skill documentation-server
Basic workflow
- Add documents using
add_documentor place.txt/.md/.pdffiles in the uploads folder and callprocess_uploads. - Search across everything with
search_all_documents, or within a single document withsearch_documents. - Use
get_context_windowto fetch neighboring chunks and give the LLM broader context.
Web UI
The web interface starts automatically on port 3080 when the MCP server launches. From the web UI you can:
- 📊 Dashboard — overview of all documents and stats
- 📄 Documents — browse, view, and delete documents
- ➕ Add Document — create documents with title, content, and metadata
- 🔍 Search All — semantic search across all documents
- 🎯 Search in Doc — search within a specific document
- 🤖 AI Search — Gemini-powered analysis (if
GEMINI_API_KEYis set) - 📁 Upload Files — drag & drop files and process them into the knowledge base
- 🪟 Context Window — explore chunks around a specific index
Configure an MCP client
#### Minimal
{
"mcpServers": {
"documentation": {
"command": "npx",
"args": ["-y", "@andrea9293/mcp-documentation-server"]
}
}
}
#### With environment variables (all optional)
{
"mcpServers": {
"documentation": {
"command": "npx",
"args": ["-y", "@andrea9293/mcp-documentation-server"],
"env": {
"MCP_BASE_DIR": "/path/to/workspace",
"GEMINI_API_KEY": "your-api-key-here",
"MCP_EMBEDDING_MODEL": "Xenova/all-MiniLM-L6-v2",
"START_WEB_UI": "true",
"WEB_HOST": "127.0.0.1",
"WEB_PORT": "3080"
}
}
}
}
All environment variables are optional. Without GEMINI_API_KEY, only the local embedding-based search tools are available.
MCP Tools
The server registers the following tools (all validated with Zod schemas):
📄 Document Management
📁 File Processing
🔍 Search
Configuration
Configure via environment variables or a .env file in the project root:
Storage layout
~/.mcp-documentation-server/ # Or custom path via MCP_BASE_DIR
├── data/
│ ├── orama-chunks.msp # Orama vector DB (child chunks + embeddings)
│ ├── orama-docs.msp # Orama document DB (full content + metadata)
│ ├── orama-parents.msp # Orama parent chunks DB (context sections)
│ ├── migration-complete.flag # Written after legacy JSON migration
│ └── *.md # Markdown copies of documents
└── uploads/ # Drop .txt, .md, .pdf files here
Embedding Models
Set via MCP_EMBEDDING_MODEL:
Models are downloaded on first use (~80–420 MB). The vector dimension is determined automatically from the provider.
⚠️ Important: Changing the embedding model requires re-adding all documents — embeddings from different models are incompatible. The Orama database is recreated automatically when the dimension changes.
Architecture
Server (FastMCP, stdio)
├─ Web UI (Express, port 3080)
│ └─ REST API → DocumentManager
└─ MCP Tools
└─ DocumentManager
├─ OramaStore — Orama vector DB (chunks DB + docs DB + parents DB), persistence, migration
├─ IntelligentChunker — Parent-child chunking (code, markdown, text, PDF)
├─ EmbeddingProvider — Local embeddings via @xenova/transformers
│ └─ EmbeddingCache — LRU in-memory cache
└─ GeminiSearchService — Optional AI search via Google Gemini
- OramaStore manages three Orama instances: one for document metadata/content, one for child chunks with vector embeddings, and one for parent chunks (context sections). All are persisted to binary files on disk and restored on startup.
- IntelligentChunker implements the Parent-Child Chunking pattern: documents are first split into large parent chunks that preserve full context (sections, paragraphs), then each parent is further split into small child chunks for precise vector search. At query time, results are deduplicated by parent so that the LLM receives both the matched fragment and the broader context.
Facts
- Kind
- MCP server
- Repo
- andrea9293/mcp-documentation-server
- Group
- Uncategorized
- Stars
- 313
- License
- MIT
- Language
- TypeScript
- Last push
- 2026-08-27
- Forks
- 45
- Topics
- documents, gemini, knowledge-base, mcp-server, model-context-protocol
- 1Everythingmodelcontextprotocol/serversThis MCP server attempts to exercise all the features of the MCP protocol. It is not intended to be a useful server, but rather a test server for builders of MCP clients. It implements prompts, tools, resources, sampling, and more to showcase MCP capabilities.85.8k
- 2Fetchmodelcontextprotocol/serversA Model Context Protocol server that provides web content fetching capabilities. This server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.85.8k
- 3Gitmodelcontextprotocol/serversA Model Context Protocol server for Git repository interaction and automation. This server provides tools to read, search, and manipulate Git repositories via Large Language Models.85.8k
- 4Memorymodelcontextprotocol/serversA basic implementation of persistent memory using a local knowledge graph. This lets Claude remember information about the user across chats.85.8k
- 5Sequential Thinkingmodelcontextprotocol/serversAn MCP server implementation that provides a tool for dynamic and reflective problem-solving through a structured thinking process.85.8k
- 6Timemodelcontextprotocol/serversA Model Context Protocol server that provides time and timezone conversion capabilities. This server enables LLMs to get current time information and perform timezone conversions using IANA timezone names, with automatic system timezone detection.85.8k