Paper Search
openags/paper-search-mcp · 1.5k stars · Python · MIT
MCP server MCP, CLI, Skills for searching and downloading academic papers from multiple sources like arXiv, PubMed, bioRxiv, etc.
Install
The repo has no one-line install. Follow its README.
Files
Paper Search MCP
A Model Context Protocol (MCP) server for searching and downloading academic papers from multiple sources. The project follows a free-first strategy: prioritize open and public data sources, support optional API keys when they improve stability or coverage, and keep source-specific connectors extensible for advanced users.
Table of Contents
- Overview
- Project Principles
- MCP Authorization Compatibility
- Features
- Source Strategy
- Sci-Hub Notice
- Installation
- Claude Code (Skill)
- Method 1 — Smithery
- Method 2 — uvx
- Method 3 — uv
- Method 4 — pip
- Method 5 — npx
- Method 6 — Docker
- Method 7 — Clone & run from source
- DeepSeek Harness (DSH)
- Environment Variables
- Contributing
- Demo
- Star History
- License
- TODO
Overview
paper-search-mcp is a Python-based tool for searching and downloading academic papers from various platforms. It provides tools for searching papers, downloading PDFs, and extracting text, making it ideal for researchers and AI-driven workflows. It can be used as an MCP server (for Claude Desktop and other MCP clients) or as a Claude Code skill with a CLI interface.
Project Principles
- Free-First: Public and open sources are the default roadmap. Paid or restricted sources are not the core direction of this project.
- Optional API Keys: API keys are supported only when they improve stability, rate limits, or metadata quality. The MCP should still be usable without them whenever possible.
- LLM-Friendly Retrieval: Search results should be standardized, deduplicated, and as complete as possible for downstream LLM workflows.
- Source Transparency: Different sources have different strengths. The MCP should make those tradeoffs explicit instead of pretending every source supports full-text retrieval.
MCP Authorization Compatibility
The bundled MCP server supports stdio (the default), sse, and streamable-http. Network transports bind to 127.0.0.1 by default. Optional OAuth protected-resource mode adds standard discovery, JWT bearer-token validation and required scopes to both HTTP transports, using an external authorization server.
Enable it with --auth oauth or PAPER_SEARCH_MCP_AUTH=oauth and the explicit issuer/JWKS/resource/audience/scope configuration in OAuth protected-resource setup. Invalid or incomplete HTTP auth settings fail startup. stdio remains independent of HTTP authentication. Local HTTP without OAuth configuration remains open; do not publish that listener directly. A protected remote deployment still needs TLS, rate limits and filesystem isolation. This feature does not host login/accounts or replace the external issuer's OAuth flow; live identity-provider interoperability must be validated for your deployment.
Features
- Two-Layer Architecture:
- Layer 1 (Unified Tooling): High-level
search_papersfor multi-source concurrent search & deduplication, anddownload_with_fallbackrelying on publisher open access links with sequential fallbacks. - Layer 2 (Platform Connectors): Modular connectors for specific academic platforms (arXiv, PubMed, bioRxiv, Semantic Scholar, etc.) equipped with intelligent DOI extraction via regex text analysis or API fields.
- Multi-Source Support: Search and download papers from arXiv, PubMed, bioRxiv, medRxiv, Google Scholar, IACR ePrint Archive, Semantic Scholar, Crossref, OpenAlex, PubMed Central (PMC), CORE, Europe PMC, dblp, OpenAIRE, CiteSeerX, DOAJ, BASE, Zenodo, HAL, SSRN, OpenReview, Unpaywall (DOI lookup), and optional Sci-Hub workflows.
- Opt-in Fast Search: CLI search keeps broad coverage by default. Use
-s fast(OpenAlex, Crossref, arXiv, PubMed, Europe PMC) or-s fastest(OpenAlex and Crossref) when lower latency matters more than coverage. - Standardized Output: Papers are returned in a consistent dictionary format via the
Paperclass. - Free-First Design: Open and public sources are prioritized before any optional commercial or restricted integrations.
- Optional API-Key Enhancement: Sources like Semantic Scholar can work better with a user-provided API key, but are not intended to force paid usage.
- Discovery + Retrieval Workflow: Google Scholar and Crossref can be used for discovery and DOI backfilling, while open repositories and publisher links are used for lawful full-text resolution where available.
- OA-First Fallback Chain:
download_with_fallbacknow follows source-native download → OpenAIRE/CORE/Europe PMC/PMC discovery → Unpaywall DOI resolution → optional Sci-Hub. - Bounded reference lookups: OpenAlex references and citing papers from a DOI or OpenAlex ID, with explicit budgets and truncation metadata.
- MCP Integration: Compatible with MCP clients for LLM context enhancement.
- DeepSeek Harness (DSH) Integration: A dsh profile bundle (clone the repo,
npx @deepseek-ai/dsh@0.2.0-rc.2 plugin --profile web add link:./dsh) exposing every tool asmcp__paper-search__*, with an optional guidance skill. - Extensible Design: Easily add new academic platforms by extending the
academic_platformsmodule.
Source Strategy
The long-term goal is not to depend on a single search engine, but to combine multiple free and public sources with clear roles:
- Open metadata backbone: Crossref, OpenAlex, Semantic Scholar, dblp, CiteSeerX, SSRN, Unpaywall (DOI-centric OA metadata).
- Discipline-specific sources: arXiv, PubMed, PubMed Central, Europe PMC, IACR.
- Open-access full-text sources: arXiv, PMC, CORE, OpenAIRE, DOAJ, BASE, Zenodo, HAL, publisher open-access links.
- Discovery and DOI recovery: Google Scholar can be useful for finding titles, versions, and DOI clues when other public metadata sources are incomplete.
Recommended free-first roadmap:
- Keep current public sources stable.
- Add OpenAlex as a broad free metadata source.
- Add PubMed Central and Europe PMC for stronger biomedical full-text access.
- Add CORE and OpenAIRE for repository-based open-access retrieval.
- Use Google Scholar mainly as a discovery fallback, not as the primary canonical source.
Platform Capability Matrix
This matrix reflects verified live-integration results from functional and end-to-end regression tests in this repository. Columns show the highest capability level observed under normal conditions.
Facts
- Kind
- MCP server
- Repo
- openags/paper-search-mcp
- Group
- Uncategorized
- Stars
- 1.5k
- License
- MIT
- Language
- Python
- Last push
- 2026-10-02
- Forks
- 290
- Topics
- ai-scientist, arxiv-papers, mcp-server, paper-search
- 1Everythingmodelcontextprotocol/serversThis MCP server attempts to exercise all the features of the MCP protocol. It is not intended to be a useful server, but rather a test server for builders of MCP clients. It implements prompts, tools, resources, sampling, and more to showcase MCP capabilities.85.8k
- 2Fetchmodelcontextprotocol/serversA Model Context Protocol server that provides web content fetching capabilities. This server enables LLMs to retrieve and process content from web pages, converting HTML to markdown for easier consumption.85.8k
- 3Gitmodelcontextprotocol/serversA Model Context Protocol server for Git repository interaction and automation. This server provides tools to read, search, and manipulate Git repositories via Large Language Models.85.8k
- 4Memorymodelcontextprotocol/serversA basic implementation of persistent memory using a local knowledge graph. This lets Claude remember information about the user across chats.85.8k
- 5Sequential Thinkingmodelcontextprotocol/serversAn MCP server implementation that provides a tool for dynamic and reflective problem-solving through a structured thinking process.85.8k
- 6Timemodelcontextprotocol/serversA Model Context Protocol server that provides time and timezone conversion capabilities. This server enables LLMs to get current time information and perform timezone conversions using IANA timezone names, with automatic system timezone detection.85.8k