HunyuanOCR is a multimodal large language model from Tencent focused on optical character recognition and document understanding. It supports both image and text inputs and is optimized for high-accuracy extraction from complex layouts, tables, and multilingual documents. The model can be used via the Hugging Face Transformers library and includes a specialized chat template for conversational document tasks.
In the Multimodal & vision space, HunyuanOCR takes a focused approach. It focuses on extracting and understanding text and structure from complex document images at scale. It is built as an open-source project for developers building document AI applications. HunyuanOCR is open source under the Open Source license. It runs on the web, the command line, and API.
It is developed by Tencent, and it first shipped in 2025. The project is developed in the open on GitHub with 1.9k stars and 36 commits in the last 90 days. Key capabilities include OCR, Document Understanding, and Multimodal Input. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do