Donut (Document understanding transformer) is an end-to-end model for visual document understanding. The base version from NAVER Clova converts document images directly into structured text outputs. It eliminates the need for separate OCR engines and is trained for tasks such as receipt parsing, form understanding, and document VQA. It is available via the Hugging Face Transformers library.
In the Foundation models & chat space, Donut Base takes a focused approach. It focuses on understanding and extracting information from document images without relying on separate OCR and layout analysis pipelines. Donut Base is an open-source project aimed at Document AI developers. The project is open source (MIT). Donut Base is available on the web and API.
NAVER Clova builds and maintains Donut Base, and the product first shipped in 2022. The project is developed in the open on GitHub with 6.9k stars. Among its 4 catalogued features are Document Understanding, image-to-Text, and OCR.
Latest indexed changes and source events
naver-clova-ix/donut-base verified by the PulseGate indexer
Other apps tracked under the same category.