PulseGateWireMethodologyCompanyThe global software index— through the gate this hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets91 categories · 140 niches
PulseGate

The global index of software taking shape now.

Stores show what passed through a store. Launch sites show what launched there. Catalogs show what entered their catalog. Each sees the market through its own gate. PulseGate reads across them.

FollowGitHubX (Twitter)LinkedIn
Platform
IndexIndex by setCategoriesWireMethodologySupply IndexDead ProjectsData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutTeamDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGateSIsimple-markdown-crawler
Visit↗
Skip to content
  1. Index›
  2. CLI tools & terminal›
  3. simple-markdown-crawler
← Back to the index
SI

simple-markdown-crawler

PyPI·Infrastructure

simple-markdown-crawler is an open-source command-line web crawler that recursively visits websites and converts each page into a separate Markdown file. It supports multithreaded crawling and is intended for web scraping, RAG pipelines, and LLM workflows.

Open SourceMITCLISelf-hosted
Visit PyPI↗

Overview

5 features

In the CLI tools & terminal space, simple-markdown-crawler takes a focused approach. It focuses on collecting website content as structured Markdown for scraping, RAG, and LLM workflows. simple-markdown-crawler is an open-source project aimed at developers building web scraping, RAG, and LLM data pipelines. The project is open source (MIT). simple-markdown-crawler is available on the command line, and it can be self-hosted.

Behind simple-markdown-crawler is yukiteruamano, and it first shipped in 2023. The project is developed in the open on GitHub with 4 commits in the last 90 days. Key capabilities include recursive crawling, multithreaded crawling, and HTML to Markdown.

Summary written by a language model from the project’s public pages.

  • ✓Recursive crawling
  • ✓Multithreaded crawling
  • ✓HTML to Markdown
  • ✓Per-page files
  • ✓RAG-ready output
Tags
web-crawlerhtml-to-markdownrag-ingestionllm-data-prep

Built with & integrations

Connectors
github
Runs on
CLISelf-hosted

Trust & compliance

License
MIT
Verified signals
✓HTTPS✓Open Source✓Free tier✓GitHub✓Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed8 Aug · 00:57 UTC
    simple-markdown-crawler verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about simple-markdown-crawler

What is simple-markdown-crawler?
Simple-markdown-crawler focuses on collecting website content as structured Markdown for scraping, RAG, and LLM workflows. It is catalogued under CLI tools & terminal on PulseGate.
Who is simple-markdown-crawler for?
simple-markdown-crawler is an open-source project built for developers building web scraping, RAG, and LLM data pipelines.
Is simple-markdown-crawler free?
Yes — simple-markdown-crawler is open source under the MIT license and free to use.
What platforms does simple-markdown-crawler run on?
simple-markdown-crawler runs on the command line. It can also be self-hosted.
Is simple-markdown-crawler still maintained?
The GitHub repository shows 4 commits in the last 90 days.
What projects are similar to simple-markdown-crawler?
Similar projects tracked by PulseGate include crawldown, html_docs_crawler, and crawler.sh.crawldownhtml_docs_crawlercrawler.sh
Who develops simple-markdown-crawler?
simple-markdown-crawler is developed by yukiteruamano.
When did simple-markdown-crawler launch?
simple-markdown-crawler first shipped in 2023.

At a glance

Platforms
Cli
Languages
English
Open source
Yes (GitHub)
License
MIT
Built for
developers building web scraping, RAG, and LLM data pipelines
Model
Open source
Solves
Collecting website content as structured Markdown for scraping, RAG, and LLM workflows.

Registered as

GitHub
yukiteruamano/simple-markdown-crawler
PyPI
simple-markdown-crawler

Developer

yukiteruamano
Small team
↗ GitHub

Open source

View on GitHub →
Stars
0
Forks
0
Open issues
0
Last commit
8 Aug 2026
Commits 90d
4
Contributors
2
Authorship
Small team
Default branch
main

Index record

Identity confidence
Low · 64
Indexed
8 Aug 2026
Lifecycle
Alive
Last seen
8 Aug 2026
Identity audit (13)
Slug
simple-markdown-crawler-pypi-org
Lifecycle last checked
9 Aug 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
8 Aug 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Written by a language model from the project's public pages.
Category from
Assigned by a language model.
Summary from
Written by a language model from public pages.
Languages from
Detected by a language model from page content.
Last updated
8 Aug 2026
Canonical URL
https://pypi.org/project/simple-markdown-crawler

Ship this? Send a correction — no account, and you get a link to follow it.

Similar projects

Closest matches by what these projects do

  • CRcrawldownpypi.org
  • HThtml_docs_crawlergithub.com
  • CRcrawler.shcrawler.sh