Common Crawl
PulseGate's liveness check found it on 17 Sep 2026; it is registered on GitHub and has been in the index since 17 Sep 2026. How this is checked
Common Crawl maintains a free, open repository of web crawl data collected from the public internet. Researchers and developers can access crawl archives, indexes, statistics, and related datasets for web research, data extraction, and machine learning.
Inferred · not functionally tested
Overview
6 featuresPurpose: Accessing and analyzing large-scale web crawl data without collecting the web independently.
Inferred · not functionally tested
Audience: researchers, developers, and machine learning teams
Inferred · not functionally tested
Functions: data_extraction
Inferred · not functionally tested
Interfaces: API: indicated (inferred, not tested) · MCP: unknown · CLI: unknown · Self-hosting: unknown
Recorded constraints: pricing: free · license: Proprietary · platforms: WEB · deployment: browser, api_only, cloud_managed
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: commoncrawl.org · github.com. These links do not verify the individual claims.
Common Crawl is a Data integration & ETL project. Inferred · not functionally tested: It focuses on accessing and analyzing large-scale web crawl data without collecting the web independently. Inferred · not functionally tested: It is built as an open-source project for researchers, developers, and machine learning teams. Basis unknown · not verified: Common Crawl is free to use. Basis unknown · not verified: It runs on the web and API.
It is developed by Common Crawl (United States), and it first shipped in 2018. The project is developed in the open on GitHub with 31 stars and 7 commits in the last 90 days. Inferred · not functionally tested: Among its 6 catalogued features are crawl archives, CDXJ index, and URL index.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Crawl archives
- CDXJ index
- URL index
- Web graphs
- Crawl statistics
- Graph statistics
Topics: Inferred · not functionally tested
Built with & integrations
- Cloudflare
- cf-ray header · cf-cache-status header
Trust & compliance
Indexing history
2What PulseGate has recorded for this listing
- Indexed17 Sep · 15:54 UTCCommon Crawl Data Stored on a Hugging Face Bucket seen via Hacker News firehose (Algolia)Source: Hacker News firehose (Algolia) · Open
Frequently asked questions about Common Crawl
- What does Common Crawl do?
- Inferred · not functionally tested: Common Crawl focuses on accessing and analyzing large-scale web crawl data without collecting the web independently. It is catalogued under Data integration & ETL on PulseGate.
- Who should use Common Crawl?
- Inferred · not functionally tested: Common Crawl is an open-source project built for researchers, developers, and machine learning teams.
- Does Common Crawl have a free plan?
- Basis unknown · not verified: Yes — Common Crawl is free to use.
- What platforms does Common Crawl run on?
- Basis unknown · not verified: Common Crawl runs on the web and API.
- Is Common Crawl still active?
- PulseGate's liveness check found it on 17 Sep 2026. Its GitHub repository shows 7 commits in the last 90 days.
- Who makes Common Crawl?
- Common Crawl is developed by Common Crawl, based in the United States.
- How long has Common Crawl been around?
- Common Crawl first shipped in 2018.
- Is Common Crawl open source?
- Basis unknown · not verified: Common Crawl has a public GitHub repository.
Similar projects
Closest matches by what these projects do