PulseGateCategoriesMethodologyCompanyThe global software index— through the gate this hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets91 categories · 140 niches
PulseGate

The global index of software taking shape now.

Stores show what passed through a store. Launch sites show what launched there. Catalogs show what entered their catalog. Each sees the market through its own gate. PulseGate reads across them.

FollowGitHubX (Twitter)LinkedIn
Platform
IndexIndex by setCategoriesIndustry UpdatesMethodologySupply IndexDead ProjectsData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutTeamDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGateCRcrawler.sh
Visit↗
Skip to content
  1. Index›
  2. RAG, search & retrieval›
  3. crawler.sh
← Back to the index
CR

crawler.sh

crawler.sh·Infrastructure

crawler.sh is a local CLI tool that extracts clean Markdown from any website, rendering JavaScript and respecting robots.txt. It is designed for AI engineers and data scientists building RAG pipelines, fine-tuning corpora, or agent contexts, providing bulk export and SEO analysis features without cloud dependencies or per-page fees.

FreeWebCLIWindowsmacOSLinux
Ccrawler.sh preview
Visit crawler.sh↗

Overview

5 features

In the RAG, search & retrieval space, crawler.sh takes a focused approach. It focuses on extracting clean, structured Markdown content from websites for use in AI training and retrieval-augmented generation pipelines. It is built as a B2B product for AI engineers and data scientists. crawler.sh is free to use. crawler.sh is available on the web, the command line, Windows, macOS, and Linux.

crawler.sh first shipped in 2024. Among its 5 catalogued features are markdown export, javaScript rendering, and bulk site crawling.

Summary written by a language model from the project’s public pages.

  • ✓Markdown export
  • ✓JavaScript rendering
  • ✓Bulk site crawling
  • ✓robots.txt respect
  • ✓SEO analysis
Tags
rag-markdownlocal-crawlerjs-rendering
AI capabilities
Text
Inference: Local

Built with & integrations

Framework
Astro
Hosting
Cloudflare
Runs on
BrowserCLIWindowsmacOSLinux
Detected from
Astro
astro-island in the HTML · /_astro/ in the HTML
Cloudflare
cf-ray header · cf-cache-status header

Trust & compliance

Verified signals
✓HTTPS✓Privacy Policy✓Terms of Service✓Free tier
Legal
Privacy Policy →Terms of Service →

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed17 Jun · 15:09 UTC
    Listing verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about crawler.sh

What is crawler.sh?
Crawler.sh focuses on extracting clean, structured Markdown content from websites for use in AI training and retrieval-augmented generation pipelines. It is catalogued under RAG, search & retrieval on PulseGate.
Who is crawler.sh for?
crawler.sh is a B2B product built for AI engineers and data scientists.
Does crawler.sh have a free plan?
Yes — crawler.sh is free to use.
What platforms does crawler.sh run on?
crawler.sh runs on the web, the command line, Windows, macOS, and Linux.
Is crawler.sh still active?
Unverified. crawler.sh has not been re-checked since it entered the index, so there is no finding either way — and only a positive finding would say otherwise.
How long has crawler.sh been around?
crawler.sh first shipped in 2024.

At a glance

Pricing
Free
Platforms
Cli · Web
Languages
English
Built for
AI engineers and data scientists
Model
B2B
Solves
Extracting clean, structured Markdown content from websites for use in AI training and retrieval-augmented generation pipelines.

Index record

Identity confidence
Medium · 72.8
Indexed
17 Jun 2026
Lifecycle
Alive
Last seen
17 Jun 2026
Identity audit (12)
Slug
crawler-sh-local-markdown-extractor-for-ai-training-and-rag-crawler-sh
Lifecycle last checked
8 Aug 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
17 Jun 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Derived from the project's own page and URL.
Category from
Assigned by a language model.
Summary from
Written by a language model from public pages.
Languages from
Detected by a language model, checked against the page's own declaration.
Canonical URL
https://crawler.sh/

Ship this? Send a correction — no account, and you get a link to follow it.

Similar projects

Closest matches by what these projects do

  • SIsimple-markdown-crawlerpypi.org