PulseGateIndexCategoriesUpdatesLive intelligence on AI-era software— indexed in the last hour
Coverage—in the index
Freshness—newest listing
Cadence—last week average · — today
Index9 markets86 categories
PulseGate

Live intelligence on the AI-era software market — apps, models, agents and infrastructure.

Most of it never reaches an official store. PulseGate maps the whole market — every category, worldwide — and measures how it moves: what’s launching, what’s gaining, what’s going quiet.

FollowGitHubX (Twitter)LinkedIn
Platform
All AppsFull IndexCategoriesIndustry UpdatesData SourcesCoverage RulesGlossaryEmbed Widget
Support
Help CenterSubmit your projectReport an Issue
Company
AboutDimaxiaPress & DataContactPlatform Status
Legal
PrivacyTermsDisclaimer
Dimaxia · Dymaxio s.r.o. · Prague, Czechia · © 2026WatchlistSitemapSystem status
PulseGate?
Visit↗
Skip to content
  1. Index›
  2. Foundation models & chat›
← Back to the index
?

huggingface.co·Infrastructure

olmOCR-2-7B-1025 is an open-source multimodal large language model developed by the Allen Institute for AI. It specializes in optical character recognition (OCR), document layout analysis, and structured information extraction from scanned or digital documents. The model accepts both text and image inputs, making it suitable for processing complex PDFs, forms, and scanned materials in research or production pipelines.

Open SourceApache-2.0WebCLIAPI
? preview
Visit huggingface.co↗
19.2kstars
1.6kforks
4features
2024since

Overview

4 features

is a Foundation models & chat project. It focuses on converting complex document images into structured, machine-readable text and layout data at high accuracy. It is built as an open-source project for developers and researchers. The project is open source (Apache-2.0). It ships for the web, the command line, and API.

Allen Institute for AI builds and maintains , and it first shipped in 2024. The project is developed in the open on GitHub with 19.2k stars. Among its 4 catalogued features are Document OCR, Layout Analysis, and Multimodal Understanding. It exposes integrations via a public API.

Summary written by a language model from the project’s public pages.

  • ✓Document OCR
  • ✓Layout Analysis
  • ✓Multimodal Understanding
  • ✓Vision-Language Processing
Tags
document-ocrlayout-understandingmultimodal-llmvision-languagehuggingface-model
AI capabilities
Multimodal
Inference: LocalWeights: Open

Built with & integrations

Hosting
AWS
AI providers
local_oss
Connectors
API
Runs on
BrowserCLIAPI-only
Detected from
AWS
x-amz-cf-id header · x-amz-cf-pop header · via header
local_oss
bollama in the HTML · bllama in the HTML · bvllm in the HTML

Trust & compliance

License
Apache-2.0
Verified signals
✓HTTPS✓Privacy Policy✓Terms of Service✓Open Source✓Free tier✓GitHub · ★ 19.2k
Legal
Privacy Policy →Terms of Service →

Indexing history

1

What PulseGate has recorded for this listing

  1. Indexed31 Jul · 08:59 UTC
    allenai/olmOCR-2-7B-1025 verified against its public source
    Source: PulseGate · Open ↗

Frequently asked questions about

What does do?
Focuses on converting complex document images into structured, machine-readable text and layout data at high accuracy. It is catalogued under Foundation models & chat on PulseGate.
Who is for?
is an open-source project built for developers and researchers.
Does have a free plan?
Yes — is open source under the Apache-2.0 license and free to use.
What platforms does run on?
runs on the web, the command line, and API.
Is still active?
Unverified. has not been re-checked since it entered the index, so there is no finding either way — and only a positive finding would say otherwise.
What projects are similar to ?
Similar projects tracked by PulseGate include OLMo 2 0425 1B, Olmo Hybrid 7B, and Olmo 3 1025 7B.OLMo 2 0425 1BOlmo Hybrid 7BOlmo 3 1025 7B
Who develops ?
is developed by Allen Institute for AI.
When did launch?
first shipped in 2024.

At a glance

Pricing
Open Source
Platforms
Cli · Web
Languages
English
Open source
Yes · ★ 19.2k
License
Apache-2.0
Built for
developers and researchers
Model
Open source
Solves
Converting complex document images into structured, machine-readable text and layout data at high accuracy.

Registered as

GitHub
allenai/olmocr
Hugging Face
allenai/olmOCR-2-7B-1025

Developer

Allen Institute for AI
Team
↗ GitHub

Open source

View on GitHub →
Stars
19,245
Forks
1,591
Open issues
87
Last commit
25 Mar 2026
Commits 90d
0
Contributors
16
Authorship
Team
Default branch
main
Latest release
v0.4.27 · 12 Mar 2026

Live coverage

Identity confidence
Medium · 67
Indexed
31 Jul 2026
Lifecycle
Alive
Last seen
31 Jul 2026
Identity audit (13)
Slug
allenai-olmocr-2-7b-1025-huggingface-co
Lifecycle state recorded
31 Jul 2026
Verification state
Indexed for public listing
Listing state
Listed: yes
Index status
Included in index
Latest evidence snapshot
31 Jul 2026
Timeline basis
Indexed-at chronology. This listing's first-seen date was written by the catalog backfill, not observed here, so it is not treated as a sighting.
Name from
Derived from the project's own page and URL.
Category from
Derived from the model's own Hugging Face metadata.
Summary from
Written by a language model from public pages.
Languages from
Detected by a language model from page content.
Last updated
5 Aug 2026
Canonical URL
https://huggingface.co/allenai/olmOCR-2-7B-1025

Ship this? Send a correction — no account, and you get a link to follow it.

Similar apps

Closest matches by what these projects do

  • ?huggingface.co
  • O2OLMo 2 0425 1Bhuggingface.co
  • OHOlmo Hybrid 7Bhuggingface.co
  • O3Olmo 3 1025 7Bhuggingface.co
  • O1OLMoE 1B 7B 0924huggingface.co
  • OVOvisOCR2huggingface.co