PulseGateWireMethodologyCompanyThe global software index— through the gate this hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets91 categories · 140 niches
PulseGate

The global index of software taking shape now.

Stores show what passed through a store. Launch sites show what launched there. Catalogs show what entered their catalog. Each sees the market through its own gate. PulseGate reads across them.

FollowGitHubX (Twitter)LinkedIn
Platform
IndexIndex by setCategoriesWireMethodologySupply IndexDead ProjectsData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutTeamDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGateHThtml_docs_crawler
Visit↗
Skip to content
  1. Index›
  2. CLI tools & terminal›
  3. html_docs_crawler
← Back to the index
HT

html_docs_crawler

github.com·Infrastructure

html_docs_crawler is an open-source command-line tool that crawls HTML documentation sites and converts their content to Markdown format, correcting internal links in the process. It is designed for developers and technical writers who need to migrate or archive documentation efficiently. The tool is distributed via PyPI and is fully open source under the MIT license.

Open SourceCLI
Hhtml_docs_crawler preview
Visit github.com↗

Overview

5 features

In the CLI tools & terminal space, html_docs_crawler takes a focused approach. It focuses on automating the conversion of HTML documentation to Markdown with correct internal links for developers and technical writers. html_docs_crawler is an open-source project aimed at developers. html_docs_crawler is open source under the Open Source license. It runs on the command line.

html_docs_crawler first shipped in 2026. Development happens publicly on GitHub with 26 commits in the last 90 days. Among its 5 catalogued features are HTML to Markdown, internal link correction, and documentation crawling.

Summary written by a language model from the project’s public pages.

  • ✓HTML to Markdown
  • ✓Internal link correction
  • ✓Documentation crawling
  • ✓Web scraping
  • ✓Command-line interface
Tags
html-to-markdowndocumentation-crawlerweb-scraping

Built with & integrations

Runs on
CLI

Trust & compliance

License
Open Source
Verified signals
✓HTTPS✓Open Source✓Free tier✓GitHub✓Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed28 Jun · 17:29 UTC
    Listing verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about html_docs_crawler

What does html_docs_crawler do?
Html_docs_crawler focuses on automating the conversion of HTML documentation to Markdown with correct internal links for developers and technical writers. It is catalogued under CLI tools & terminal on PulseGate.
Who should use html_docs_crawler?
html_docs_crawler is an open-source project built for developers.
Is html_docs_crawler free?
Yes — html_docs_crawler is open source under the Open Source license and free to use.
What platforms does html_docs_crawler run on?
html_docs_crawler runs on the command line.
Is html_docs_crawler still maintained?
The GitHub repository shows 26 commits in the last 90 days.
When did html_docs_crawler launch?
html_docs_crawler first shipped in 2026.
Is html_docs_crawler open source?
Yes — html_docs_crawler is open source under the Open Source license, developed on GitHub.

At a glance

Platforms
Cli
Languages
English
Open source
Yes (GitHub)
License
Open Source
Built for
developers
Model
Open source
Solves
Automating the conversion of HTML documentation to Markdown with correct internal links for developers and technical writers.

Registered as

GitHub
zwidny/doc_crawler
PyPI
html_docs_crawler

Developer

Zwidny
Solo developer
↗ GitHub

Open source

View on GitHub →
Stars
0
Forks
0
Open issues
0
Last commit
3 Jun 2026
Commits 90d
26
Contributors
1
Authorship
Solo
Default branch
master

Index record

Identity confidence
Low · 64
Indexed
28 Jun 2026
Lifecycle
Alive
Last seen
28 Jun 2026
Identity audit (12)
Slug
html-docs-crawler-pypi-org
Lifecycle last checked
8 Aug 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
28 Jun 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Written by a language model from the project's public pages.
Category from
Assigned by a language model.
Summary from
Written by a language model from public pages.
Languages from
Detected by a language model from page content.
Canonical URL
https://github.com/zwidny/doc_crawler

Ship this? Send a correction — no account, and you get a link to follow it.

Similar projects

Closest matches by what these projects do

  • SIsimple-markdown-crawlerpypi.org