Skip to content
Back to the index

commoncrawl

pipeworx.ioInfrastructure

No liveness check has reached it yet; it has been in the index since 9 Oct 2026. How this is checked

Commoncrawl is an HTTP JSON-RPC MCP server exposing tools to list Common Crawl collections, search CDX indexes, and retrieve archived page records. It is intended for developers and researchers building applications that need historical web data.

Inferred · not functionally tested

FreeWebAPICloud-managed
commoncrawl preview
Visit pipeworx.io
1alternative
6features
2026since

Overview

6 features

Purpose: Finding and retrieving historical web pages and crawl records from Common Crawl archives.

Inferred · not functionally tested

Audience: developers and researchers

Inferred · not functionally tested

Functions: data_extraction

Inferred · not functionally tested

Interfaces: API: indicated (inferred, not tested) · MCP: indicated (inferred, not tested) · CLI: unknown · Self-hosting: unknown

Recorded constraints: pricing: free · license: Proprietary · platforms: WEB · deployment: browser, api_only, cloud_managed

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: gateway.pipeworx.io. These links do not verify the individual claims.

In the API design, testing & docs space, commoncrawl takes a focused approach. Inferred · not functionally tested: It focuses on finding and retrieving historical web pages and crawl records from Common Crawl archives. Inferred · not functionally tested: It is built as a B2B product for developers and researchers. Basis unknown · not verified: commoncrawl is free to use. Basis unknown · not verified: It ships for the web and API.

Pipeworx builds and maintains commoncrawl. Inferred · not functionally tested: Among its 6 catalogued features are crawl listing, CDX index search, and archived record retrieval. Inferred · not functionally tested: Catalogued interfaces include an MCP server.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Crawl listing
  • CDX index search
  • Archived record retrieval
  • JSON-RPC endpoint
  • MCP server
  • Keyless access

Topics: Inferred · not functionally tested

Tags
common-crawlweb-archivescdx-indexmcp-server

JSON profile · Text profile · Access guide

Built with & integrations

Hosting
Cloudflare
Connectors
MCP
Runs on
BrowserAPI-onlyCloud-managed
Detected from
Cloudflare
cf-ray header

Trust & compliance

Public signals
HTTPSFree tier

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed9 Oct · 18:01 UTC
    Commoncrawl seen via MCP Registry (official)
    Source: MCP Registry (official) · Open

Frequently asked questions about commoncrawl

What does commoncrawl do?
Inferred · not functionally tested: Commoncrawl focuses on finding and retrieving historical web pages and crawl records from Common Crawl archives. It is catalogued under API design, testing & docs on PulseGate.
Who should use commoncrawl?
Inferred · not functionally tested: commoncrawl is a B2B product built for developers and researchers.
Does commoncrawl have a free plan?
Basis unknown · not verified: Yes — commoncrawl is free to use.
What platforms does commoncrawl run on?
Basis unknown · not verified: commoncrawl runs on the web and API.
Is commoncrawl still maintained?
Unverified. commoncrawl has not been re-checked since it entered the index, so there is no finding either way — and only a positive finding would say otherwise.
What projects are similar to commoncrawl?
Similar projects tracked by PulseGate include commoncrawl-mcp, crawlbase, and firecrawl.commoncrawl-mcpcrawlbasefirecrawl
Who makes commoncrawl?
commoncrawl is developed by Pipeworx.
Does commoncrawl have an API or integrations?
Inferred · not functionally tested: Yes — commoncrawl exposes an MCP server.

Similar projects

Closest matches by what these projects do