Skip to content
Back to the index

Common Crawl

commoncrawl.orgInfrastructure🇺🇸

PulseGate's liveness check found it on 17 Sep 2026; it is registered on GitHub and has been in the index since 17 Sep 2026. How this is checked

Common Crawl maintains a free, open repository of web crawl data collected from the public internet. Researchers and developers can access crawl archives, indexes, statistics, and related datasets for web research, data extraction, and machine learning.

Inferred · not functionally tested

FreeWebAPICloud-managed
Common Crawl preview
Visit commoncrawl.org
31stars
4forks
7features
2018since

Overview

6 features

Purpose: Accessing and analyzing large-scale web crawl data without collecting the web independently.

Inferred · not functionally tested

Audience: researchers, developers, and machine learning teams

Inferred · not functionally tested

Functions: data_extraction

Inferred · not functionally tested

Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: unknown · Self-hosting: unknown

Recorded constraints: pricing: free · license: Proprietary · platforms: WEB · deployment: browser, api_only, cloud_managed

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: commoncrawl.org · github.com. These links do not verify the individual claims.

Common Crawl is a Data integration & ETL project. Inferred · not functionally tested: It focuses on accessing and analyzing large-scale web crawl data without collecting the web independently. Inferred · not functionally tested: It is built as an open-source project for researchers, developers, and machine learning teams. Basis unknown · not verified: Common Crawl is free to use. Basis unknown · not verified: It runs on the web and API.

It is developed by Common Crawl (United States), and it first shipped in 2018. The project is developed in the open on GitHub with 31 stars and 7 commits in the last 90 days. Inferred · not functionally tested: Among its 6 catalogued features are crawl archives, CDXJ index, and URL index.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Crawl archives
  • CDXJ index
  • URL index
  • Web graphs
  • Crawl statistics
  • Graph statistics

Topics: Inferred · not functionally tested

Tags
web-crawl-dataopen-web-corpusresearch-datasetsdata-extraction

JSON profile · Text profile · Access guide

Built with & integrations

Hosting
Cloudflare
Runs on
BrowserAPI-onlyCloud-managed
Detected from
Cloudflare
cf-ray header · cf-cache-status header

Trust & compliance

Public signals
HTTPSPrivacy PolicyTerms of ServiceFree tierGitHub · ★ 31Active maintenance

Indexing history

2

What PulseGate has recorded for this listing

  1. Indexed8 Oct · 18:08 UTC
    Commoncrawl seen via Saashub
    Source: Saashub · Open
  2. Indexed17 Sep · 15:54 UTC
    Common Crawl Data Stored on a Hugging Face Bucket seen via Hacker News firehose (Algolia)
    Source: Hacker News firehose (Algolia) · Open

Frequently asked questions about Common Crawl

What does Common Crawl do?
Inferred · not functionally tested: Common Crawl focuses on accessing and analyzing large-scale web crawl data without collecting the web independently. It is catalogued under Data integration & ETL on PulseGate.
Who should use Common Crawl?
Inferred · not functionally tested: Common Crawl is an open-source project built for researchers, developers, and machine learning teams.
Does Common Crawl have a free plan?
Basis unknown · not verified: Yes — Common Crawl is free to use.
What platforms does Common Crawl run on?
Basis unknown · not verified: Common Crawl runs on the web and API.
Is Common Crawl still active?
PulseGate's liveness check found it on 17 Sep 2026. Its GitHub repository shows 7 commits in the last 90 days.
Who makes Common Crawl?
Common Crawl is developed by Common Crawl, based in the United States.
How long has Common Crawl been around?
Common Crawl first shipped in 2018.
Is Common Crawl open source?
Basis unknown · not verified: Common Crawl has a public GitHub repository.

Similar projects

Closest matches by what these projects do