baidu/Qianfan-OCR is a vision-language model from Baidu's Qianfan platform designed for optical character recognition and document understanding tasks. It accepts image inputs and produces structured text output, supporting both English and Chinese content. The model is hosted on Hugging Face and can be used via the Transformers library or Baidu's inference services.
Qianfan OCR sits in PulseGate's Foundation models & chat category. It focuses on extracting and understanding text from images or scanned documents using AI without building custom OCR pipelines. It is built as an open-source project for developers. Qianfan OCR is available on the web, the command line, and API.
Baidu builds and maintains Qianfan OCR, and it first shipped in 2025. The project is developed in the open on GitHub with 421 stars. Among its 3 catalogued features are Optical Character Recognition, Multimodal Input, and Chat Template Support.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do