HunyuanOCR is a multimodal large language model from Tencent focused on optical character recognition and document understanding. It supports both image and text inputs and is optimized for high-accuracy extraction from complex layouts, tables, and multilingual documents. The model can be used via the Hugging Face Transformers library and includes a specialized chat template for conversational document tasks.
In the Foundation models & chat space, HunyuanOCR takes a focused approach. It focuses on extracting and understanding text and structure from complex document images at scale. It is built as an open-source project for developers building document AI applications. HunyuanOCR is open source under the Open Source license. HunyuanOCR is available on the web, the command line, and API.
Behind HunyuanOCR is Tencent, and the product first shipped in 2025. Development happens publicly on GitHub with 1.9k stars and 36 commits in the last 90 days. Key capabilities include OCR, Document Understanding, and Multimodal Input. It exposes integrations via a public API.
Latest indexed changes and source events
tencent/HunyuanOCR verified by the PulseGate indexer
Other apps tracked under the same category.