Skip to content
Back to the index

Block Common Crawl via robots.txt

wordpress.orgWebsite firewall, bot & spam protection

PulseGate's liveness check found it on 7 Oct 2026; it is registered on WordPress.org and has been in the index since 6 Sep 2026. How this is checked

Block Common Crawl via robots.txt is a WordPress plugin that adds rules to WordPress’s virtual robots.txt output to discourage the Common Crawl bot from crawling site content. It is intended for site owners who want to limit inclusion in datasets used for AI training.

Inferred · not functionally tested

FreeWebCloud-managed
Block Common Crawl via robots.txt preview
Visit wordpress.org

Overview

4 features

Purpose: Preventing Common Crawl bots from collecting and indexing content from a WordPress website.

Inferred · not functionally tested

Audience: WordPress site owners

Inferred · not functionally tested

Functions: Unknown

Interfaces: API: unknown · MCP: unknown · CLI: unknown · Self-hosting: unknown

Recorded constraints: pricing: free · license: Proprietary · platforms: WEB · deployment: browser, cloud_managed

Constraint provenance is unknown; confirm requirements with the publisher.

Record sources: wordpress.org · wordpress.org. These links do not verify the individual claims.

In the Website firewall, bot & spam protection space, Block Common Crawl via robots.txt takes a focused approach. Inferred · not functionally tested: It focuses on preventing Common Crawl bots from collecting and indexing content from a WordPress website. Inferred · not functionally tested: Block Common Crawl via robots.txt is a B2B product aimed at wordPress site owners. Basis unknown · not verified: It is available for free. Basis unknown · not verified: It ships for the web.

Behind Block Common Crawl via robots.txt is apasionados, and it first shipped in 2023. Inferred · not functionally tested: Among its 4 catalogued features are virtual robots.txt rules, CCBot blocking, and wordPress integration. The interface is available in English and Gujarati. It is being sunset.

Summary written by a language model from the project’s public pages.

Tasks: Inferred · not functionally tested

  • Virtual robots.txt rules
  • CCBot blocking
  • WordPress integration
  • AI crawler control

Topics: Inferred · not functionally tested

Tags
common-crawl-blockingrobots-txtwordpress-securityai-crawler-control

JSON profile · Text profile · Access guide

Built with & integrations

Runs on
BrowserCloud-managed

Trust & compliance

Public signals
HTTPSPrivacy PolicyFree tierMulti-language

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed6 Sep · 12:18 UTC
    Block Common Crawl via robots.txt seen via WordPress Plugins
    Source: WordPress Plugins · Open

Frequently asked questions about Block Common Crawl via robots.txt

What is Block Common Crawl via robots.txt?
Inferred · not functionally tested: Block Common Crawl via robots.txt focuses on preventing Common Crawl bots from collecting and indexing content from a WordPress website. It is catalogued under Website firewall, bot & spam protection on PulseGate.
Who is Block Common Crawl via robots.txt for?
Inferred · not functionally tested: Block Common Crawl via robots.txt is a B2B product built for wordPress site owners.
Is Block Common Crawl via robots.txt free?
Basis unknown · not verified: Yes — Block Common Crawl via robots.txt is free to use.
What platforms does Block Common Crawl via robots.txt run on?
Basis unknown · not verified: Block Common Crawl via robots.txt runs on the web.
Is Block Common Crawl via robots.txt still maintained?
It is being sunset according to the latest indexed signals.
What projects are similar to Block Common Crawl via robots.txt?
Similar projects tracked by PulseGate include Block Chat GPT via robots.txt, Block Archive.org via WordPress robots.txt, and Block AI Crawlers.Block Chat GPT via robots.txtBlock Archive.org via WordPress robots.txtBlock AI Crawlers
Who develops Block Common Crawl via robots.txt?
Block Common Crawl via robots.txt is developed by apasionados.
When did Block Common Crawl via robots.txt launch?
Block Common Crawl via robots.txt first shipped in 2023.

Similar projects

Closest matches by what these projects do