Block Common Crawl via robots.txt
PulseGate's liveness check found it on 7 Oct 2026; it is registered on WordPress.org and has been in the index since 6 Sep 2026. How this is checked
Block Common Crawl via robots.txt is a WordPress plugin that adds rules to WordPress’s virtual robots.txt output to discourage the Common Crawl bot from crawling site content. It is intended for site owners who want to limit inclusion in datasets used for AI training.
Inferred · not functionally tested
Overview
4 featuresPurpose: Preventing Common Crawl bots from collecting and indexing content from a WordPress website.
Inferred · not functionally tested
Audience: WordPress site owners
Inferred · not functionally tested
Functions: Unknown
Interfaces: API: unknown · MCP: unknown · CLI: unknown · Self-hosting: unknown
Recorded constraints: pricing: free · license: Proprietary · platforms: WEB · deployment: browser, cloud_managed
Constraint provenance is unknown; confirm requirements with the publisher.
Record sources: wordpress.org · wordpress.org. These links do not verify the individual claims.
In the Website firewall, bot & spam protection space, Block Common Crawl via robots.txt takes a focused approach. Inferred · not functionally tested: It focuses on preventing Common Crawl bots from collecting and indexing content from a WordPress website. Inferred · not functionally tested: Block Common Crawl via robots.txt is a B2B product aimed at wordPress site owners. Basis unknown · not verified: It is available for free. Basis unknown · not verified: It ships for the web.
Behind Block Common Crawl via robots.txt is apasionados, and it first shipped in 2023. Inferred · not functionally tested: Among its 4 catalogued features are virtual robots.txt rules, CCBot blocking, and wordPress integration. The interface is available in English and Gujarati. It is being sunset.
Summary written by a language model from the project’s public pages.
Tasks: Inferred · not functionally tested
- Virtual robots.txt rules
- CCBot blocking
- WordPress integration
- AI crawler control
Topics: Inferred · not functionally tested
Built with & integrations
Trust & compliance
Indexing history
1What PulseGate has recorded for this listing
- Indexed6 Sep · 12:18 UTCBlock Common Crawl via robots.txt seen via WordPress PluginsSource: WordPress Plugins · Open
Frequently asked questions about Block Common Crawl via robots.txt
- What is Block Common Crawl via robots.txt?
- Inferred · not functionally tested: Block Common Crawl via robots.txt focuses on preventing Common Crawl bots from collecting and indexing content from a WordPress website. It is catalogued under Website firewall, bot & spam protection on PulseGate.
- Who is Block Common Crawl via robots.txt for?
- Inferred · not functionally tested: Block Common Crawl via robots.txt is a B2B product built for wordPress site owners.
- Is Block Common Crawl via robots.txt free?
- Basis unknown · not verified: Yes — Block Common Crawl via robots.txt is free to use.
- What platforms does Block Common Crawl via robots.txt run on?
- Basis unknown · not verified: Block Common Crawl via robots.txt runs on the web.
- Is Block Common Crawl via robots.txt still maintained?
- It is being sunset according to the latest indexed signals.
- What projects are similar to Block Common Crawl via robots.txt?
- Similar projects tracked by PulseGate include Block Chat GPT via robots.txt, Block Archive.org via WordPress robots.txt, and Block AI Crawlers.Block Chat GPT via robots.txtBlock Archive.org via WordPress robots.txtBlock AI Crawlers
- Who develops Block Common Crawl via robots.txt?
- Block Common Crawl via robots.txt is developed by apasionados.
- When did Block Common Crawl via robots.txt launch?
- Block Common Crawl via robots.txt first shipped in 2023.
Similar projects
Closest matches by what these projects do