LayoutLMv3 Large is a pre-trained multimodal model developed by Microsoft that integrates text, layout, and image information for document understanding tasks. It excels at layout analysis, form understanding, receipt parsing, and visual question answering on scanned or digital documents. The model is distributed via Hugging Face and is used by developers building document AI pipelines, often fine-tuned for specific enterprise document processing needs.
Layoutlmv3 Large is a Computer vision, OCR & document AI project. It focuses on understanding and extracting information from complex document images that combine text, layout, and visual elements. Layoutlmv3 Large is an open-source project aimed at developers. Layoutlmv3 Large is open source under the Apache-2.0 license. It runs on the web and API, and it can be self-hosted.
Behind Layoutlmv3 Large is Microsoft, and it first shipped in 2018. The project is developed in the open on GitHub with 162.7k stars and 758 commits in the last 90 days. Key capabilities include document layout analysis, visual question answering, and information extraction. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do