PulseGateCategoriesMethodologyCompanyThe global software index— through the gate this hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets90 categories · 139 niches
PulseGate

The global index of software taking shape now.

Stores show what passed through a store. Launch sites show what launched there. Catalogs show what entered their catalog. Each sees the market through its own gate. PulseGate reads across them.

FollowGitHubX (Twitter)LinkedIn
Platform
IndexIndex by setCategoriesIndustry UpdatesMethodologySupply IndexData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutTeamDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGatePDPDFStract
Visit↗
Skip to content
  1. Index›
  2. RAG, search & retrieval›
  3. PDFStract
← Back to the index
PD

PDFStract

pdfstract.com·Infrastructure

PDFStract is a data preparation layer for RAG pipelines. It is built to take PDF documents through extraction, chunking, and embedding so they become vector-ready in one command. The page describes it as the first layer in a RAG pipeline and as a tool for AI applications.

Its conversion step uses 10+ libraries, including Marker, Docling, and PyMuPDF4LLM, with each library described as optimized for different document types. For text splitting, it offers 10+ chunking methods ranging from simple token-based approaches to advanced semantic chunking powered by AI. It also generates vector embeddings with multiple providers, including OpenAI, Sentence Transformers, and local models. The interface is described as a unified API, and switching between libraries, chunkers, and embedding providers is done by changing a single parameter.

PDFStract is delivered through a Python API, a CLI, and a Web UI. The CLI example shown on the page uses a convert-chunk-embed command. The site also links to documentation, installation, features, GitHub, PyPI, and guides for the Python API, CLI, and Web UI. It is available under the MIT license.

Open SourceApache-2.0WebEmbeddableCLIAPICloud-managed
PPDFStract preview
Visit pdfstract.com↗
151stars
13forks
10features
2025since

Overview

10 features

In the RAG, search & retrieval space, PDFStract takes a focused approach. It focuses on preparing and converting PDF documents into vector-ready data for retrieval-augmented generation pipelines. It is built as an open-source project for AI developers and data engineers. PDFStract is open source under the Apache-2.0 license. It ships for the web, embeddable surfaces, the command line, and API.

PDFStract builds and maintains PDFStract, and it first shipped in 2025. The project is developed in the open on GitHub with 151 stars. Among its 10 catalogued features are PDF extraction, text chunking, and vector embeddings. It exposes integrations via a public API.

Summary written by a language model from the project’s public pages.

  • ✓PDF extraction
  • ✓Text chunking
  • ✓Vector embeddings
  • ✓Multiple libraries
  • ✓Semantic chunking
  • ✓Python API
  • ✓CLI interface
  • ✓Web UI
  • ✓Open source
  • ✓Multi-provider support
Tags
pdf-extractionrag-pipelinesemantic-chunkingembedding-providervectorization
AI capabilities
Text
Inference: Hybrid

Built with & integrations

Framework
FastAPI
Hosting
NetlifyCloudflare
AI providers
openaimultiple
Connectors
API
Runs on
BrowserEmbeddableCLIAPI-onlyCloud-managed
Detected from
Netlify
x-nf-request-id header
Cloudflare
cf-ray header · cf-cache-status header

Trust & compliance

License
Apache-2.0
Verified signals
✓HTTPS✓Open Source✓Free tier✓GitHub · ★ 151

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed28 Jun · 17:47 UTC
    Listing verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about PDFStract

What does PDFStract do?
PDFStract focuses on preparing and converting PDF documents into vector-ready data for retrieval-augmented generation pipelines. It is catalogued under RAG, search & retrieval on PulseGate.
Who is PDFStract for?
PDFStract is an open-source project built for AI developers and data engineers.
Is PDFStract free?
Yes — PDFStract is open source under the Apache-2.0 license and free to use.
What platforms does PDFStract run on?
PDFStract runs on the web, embeddable surfaces, the command line, and API.
Is PDFStract still maintained?
Unverified. PDFStract has not been re-checked since it entered the index, so there is no finding either way — and only a positive finding would say otherwise.
Who develops PDFStract?
PDFStract is developed by PDFStract.
When did PDFStract launch?
PDFStract first shipped in 2025.
Is PDFStract open source?
Yes — PDFStract is open source under the Apache-2.0 license, developed on GitHub.

At a glance

Platforms
Web
Languages
English
Open source
Yes · ★ 151
License
Apache-2.0
Built for
AI developers and data engineers
Model
Open source
Solves
Preparing and converting PDF documents into vector-ready data for retrieval-augmented generation pipelines.

Registered as

GitHub
AKSarav/pdfstract

Developer

Small team
↗ GitHub

Open source

View on GitHub →
Stars
151
Forks
13
Open issues
6
Last commit
18 Mar 2026
Commits 90d
0
Contributors
2
Authorship
Small team
Default branch
main
Latest release
v1.1.0 · 13 Feb 2026

Index record

Identity confidence
Medium · 74.8
Indexed
28 Jun 2026
Lifecycle
Alive
Last seen
28 Jun 2026
Identity audit (12)
Slug
pdfstract-the-first-layer-in-your-rag-pipeline-pdfstract-pdfstract-com
Lifecycle last checked
8 Aug 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
28 Jun 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Derived from the project's own page and URL.
Category from
Assigned by a language model.
Summary from
Written by a language model from public pages.
Languages from
Detected by a language model, checked against the page's own declaration.
Canonical URL
https://pdfstract.com/

Ship this? Send a correction — no account, and you get a link to follow it.

Similar projects

Closest matches by what these projects do

  • UNUnstractunstract.com
  • PPPDF Parserpdfparser.co