pdfmux provides robust, self-healing PDF extraction optimized for Retrieval-Augmented Generation pipelines. It detects and flags unreadable pages instead of silently failing, offers per-page confidence scores, and generates verifiable signed manifests. It includes loaders for LangChain and LlamaIndex and serves as a strong alternative to tools like LlamaParse. Available as a Python package with an MCP server.
pdfmux is an AI & ML project. Unreliable PDF parsing that silently drops or corrupts content when building RAG systems from documents. pdfmux is an open-source project aimed at AI developers building RAG applications. The project is open source (MIT). It runs on the web, the command line, and API.
Behind pdfmux is NameetP, and it first shipped in 2026. Development happens publicly on GitHub with 77 stars and 46 commits in the last 90 days. Among its 6 catalogued features are Self-healing Extraction, Confidence Scoring, and Page Flagging. It exposes integrations via an MCP server and a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Same category — not a similarity match