PulseGateCategoriesMethodologyCompanyThe global software index— through the gate this hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets90 categories · 139 niches
PulseGate

The global index of software taking shape now.

Stores show what passed through a store. Launch sites show what launched there. Catalogs show what entered their catalog. Each sees the market through its own gate. PulseGate reads across them.

FollowGitHubX (Twitter)LinkedIn
Platform
IndexIndex by setCategoriesIndustry UpdatesMethodologySupply IndexData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutTeamDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGatePApagewise-pdf-extractor
Visit↗
Skip to content
  1. Index›
  2. AI›
  3. pagewise-pdf-extractor
← Back to the index
PA

pagewise-pdf-extractor

github.com·Infrastructure

Page-wise PDF to Markdown extraction with text extraction, OCR, LLM fallback, and progress metadata.

Open SourceApache-2.0WebCLISelf-hosted
Ppagewise-pdf-extractor preview
Visit github.com↗

Overview

5 features

pagewise-pdf-extractor sits in PulseGate's AI category. It focuses on extracting structured text and content from PDFs, including scanned documents, into Markdown format efficiently. pagewise-pdf-extractor is an open-source project aimed at developers and data engineers. pagewise-pdf-extractor is open source under the Apache-2.0 license. It runs on the web and the command line, and it can be self-hosted.

pagewise-pdf-extractor first shipped in 2026. The project is developed in the open on GitHub with 14 commits in the last 90 days. Key capabilities include PDF to Markdown, OCR support, and LLM fallback.

  • ✓PDF to Markdown
  • ✓OCR support
  • ✓LLM fallback
  • ✓Progress metadata
  • ✓Page-wise extraction
Tags
pdf-extractionocr-processingmarkdown-conversionrag-pipelinecli-tool
AI capabilities
TextStructured
Inference: LocalWeights: Open

Built with & integrations

Runs on
BrowserCLISelf-hosted

Trust & compliance

License
Apache-2.0
Verified signals
✓HTTPS✓Open Source✓Free tier✓GitHub✓Active maintenance

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed14 Jun · 21:28 UTC
    Listing verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about pagewise-pdf-extractor

What is pagewise-pdf-extractor?
Pagewise-pdf-extractor focuses on extracting structured text and content from PDFs, including scanned documents, into Markdown format efficiently. It is catalogued under AI on PulseGate.
Who is pagewise-pdf-extractor for?
pagewise-pdf-extractor is an open-source project built for developers and data engineers.
Does pagewise-pdf-extractor have a free plan?
Yes — pagewise-pdf-extractor is open source under the Apache-2.0 license and free to use.
What platforms does pagewise-pdf-extractor run on?
pagewise-pdf-extractor runs on the web and the command line. It can also be self-hosted.
Is pagewise-pdf-extractor still active?
The GitHub repository shows 14 commits in the last 90 days.
When did pagewise-pdf-extractor launch?
pagewise-pdf-extractor first shipped in 2026.
Is pagewise-pdf-extractor open source?
Yes — pagewise-pdf-extractor is open source under the Apache-2.0 license, developed on GitHub.

At a glance

Platforms
Cli
Languages
English
Open source
Yes (GitHub)
License
Apache-2.0
Built for
developers and data engineers
Model
Open source
Solves
Extracting structured text and content from PDFs, including scanned documents, into Markdown format efficiently.

Registered as

GitHub
ebmurha/pagewise-pdf-extractor
PyPI
pagewise-pdf-extractor

Developer

Ebmurha
Solo developer
↗ GitHub

Open source

View on GitHub →
Stars
0
Forks
0
Open issues
0
Last commit
8 Jun 2026
Commits 90d
14
Contributors
1
Authorship
Solo
Default branch
main

Index record

Identity confidence
Medium · 77.6
Indexed
14 Jun 2026
Lifecycle
Alive
Last seen
14 Jun 2026
Identity audit (12)
Slug
pagewise-pdf-extractor-pypi-org
Lifecycle last checked
8 Aug 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
14 Jun 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Written by a language model from the project's public pages.
Category from
Assigned by the 2026 taxonomy migration.
Summary from
Taken from the package registry entry.
Languages from
Detected by a language model from page content.
Canonical URL
https://github.com/ebmurha/pagewise-pdf-extractor

Ship this? Send a correction — no account, and you get a link to follow it.

Similar projects

Closest matches by what these projects do

  • MDmdweavergithub.com