commoncrawl-mcp
PulseGate's liveness check found it on 24 Sep 2026; it has been in the index since 24 Sep 2026. How this is checked
commoncrawl-mcp is an open-source Model Context Protocol server for accessing Common Crawl's public web archive data. Developers can run it locally and connect MCP-compatible clients or agents to search and retrieve archived web content.
Inferred · not functionally tested
Overview
6 featuresPurpose: Accessing and querying Common Crawl public web data through an MCP-compatible server.
Inferred · not functionally tested
Audience: developers building MCP clients and AI agents
Inferred · not functionally tested
Functions: data_extraction
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: indicated (inferred, not tested) · CLI: indicated (inferred, not tested) · Self-hosting: indicated (inferred, not tested)
Recorded constraints: pricing: open_source · license: Open Source · platforms: CLI, WEB · deployment: browser, cli, self_hosted, api_only
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: github.com. These links do not verify the individual claims.
commoncrawl-mcp is a Frameworks & SDKs project. Inferred · not functionally tested: It focuses on accessing and querying Common Crawl public web data through an MCP-compatible server. Inferred · not functionally tested: commoncrawl-mcp is an open-source project aimed at developers building MCP clients and AI agents. Basis unknown · not verified: commoncrawl-mcp is open source under the Open Source license. Basis unknown · not verified: It ships for the web, the command line, and API, and it can be self-hosted.
Behind commoncrawl-mcp is mrfentmen. Inferred · not functionally tested: Key capabilities include MCP server, Common Crawl access, and web archive search. Inferred · not functionally tested: Catalogued interfaces include an MCP server.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- MCP server
- Common Crawl access
- Web archive search
- Archived content retrieval
- Public data queries
- npm installation
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed21 Sep · 20:53 UTCio.github.mrfentmen/commoncrawl-mcp seen via MCP Registry (official)Source: MCP Registry (official) · Open
Frequently asked questions about commoncrawl-mcp
- What does commoncrawl-mcp do?
- Inferred · not functionally tested: Commoncrawl-mcp focuses on accessing and querying Common Crawl public web data through an MCP-compatible server. It is catalogued under Frameworks & SDKs on PulseGate.
- Who should use commoncrawl-mcp?
- Inferred · not functionally tested: commoncrawl-mcp is an open-source project built for developers building MCP clients and AI agents.
- Does commoncrawl-mcp have a free plan?
- Basis unknown · not verified: Yes — commoncrawl-mcp is open source under the Open Source license and free to use.
- What platforms does commoncrawl-mcp run on?
- Basis unknown · not verified: commoncrawl-mcp runs on the web, the command line, and API. It can also be self-hosted.
- Is commoncrawl-mcp still active?
- PulseGate's liveness check found it on 24 Sep 2026.
- What are alternatives to commoncrawl-mcp?
- Similar projects tracked by PulseGate include CrawlForge MCP, ietf-mcp, and wikisource-mcp.CrawlForge MCPietf-mcpwikisource-mcp
- Who develops commoncrawl-mcp?
- commoncrawl-mcp is developed by mrfentmen.
- Is commoncrawl-mcp open source?
- Basis unknown · not verified: Yes — commoncrawl-mcp is open source under the Open Source license.
Similar projects
Closest matches by what these projects do