commoncrawl
No liveness check has reached it yet; it has been in the index since 9 Oct 2026. How this is checked
Commoncrawl is an HTTP JSON-RPC MCP server exposing tools to list Common Crawl collections, search CDX indexes, and retrieve archived page records. It is intended for developers and researchers building applications that need historical web data.
Inferred · not functionally tested
Overview
6 featuresPurpose: Finding and retrieving historical web pages and crawl records from Common Crawl archives.
Inferred · not functionally tested
Audience: developers and researchers
Inferred · not functionally tested
Functions: data_extraction
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: indicated (inferred, not tested) · CLI: unknown · Self-hosting: unknown
Recorded constraints: pricing: free · license: Proprietary · platforms: WEB · deployment: browser, api_only, cloud_managed
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: gateway.pipeworx.io. These links do not verify the individual claims.
In the API design, testing & docs space, commoncrawl takes a focused approach. Inferred · not functionally tested: It focuses on finding and retrieving historical web pages and crawl records from Common Crawl archives. Inferred · not functionally tested: It is built as a B2B product for developers and researchers. Basis unknown · not verified: commoncrawl is free to use. Basis unknown · not verified: It ships for the web and API.
Pipeworx builds and maintains commoncrawl. Inferred · not functionally tested: Among its 6 catalogued features are crawl listing, CDX index search, and archived record retrieval. Inferred · not functionally tested: Catalogued interfaces include an MCP server.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Crawl listing
- CDX index search
- Archived record retrieval
- JSON-RPC endpoint
- MCP server
- Keyless access
Topics: Inferred · not functionally tested
Built with & integrations
- Cloudflare
- cf-ray header
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed9 Oct · 18:01 UTCCommoncrawl seen via MCP Registry (official)Source: MCP Registry (official) · Open
Frequently asked questions about commoncrawl
- What does commoncrawl do?
- Inferred · not functionally tested: Commoncrawl focuses on finding and retrieving historical web pages and crawl records from Common Crawl archives. It is catalogued under API design, testing & docs on PulseGate.
- Who should use commoncrawl?
- Inferred · not functionally tested: commoncrawl is a B2B product built for developers and researchers.
- Does commoncrawl have a free plan?
- Basis unknown · not verified: Yes — commoncrawl is free to use.
- What platforms does commoncrawl run on?
- Basis unknown · not verified: commoncrawl runs on the web and API.
- Is commoncrawl still maintained?
- Unverified. commoncrawl has not been re-checked since it entered the index, so there is no finding either way — and only a positive finding would say otherwise.
- What projects are similar to commoncrawl?
- Similar projects tracked by PulseGate include commoncrawl-mcp, crawlbase, and firecrawl.commoncrawl-mcpcrawlbasefirecrawl
- Who makes commoncrawl?
- commoncrawl is developed by Pipeworx.
- Does commoncrawl have an API or integrations?
- Inferred · not functionally tested: Yes — commoncrawl exposes an MCP server.
Similar projects
Closest matches by what these projects do