LayoutLMv3 is an open-weight multimodal document-understanding model from Microsoft that combines text, image, and document-layout information. Developers use it for tasks such as document classification, token classification, and layout analysis.
Layoutlmv3 Base is a Multimodal & vision project. It focuses on understanding and extracting structured information from documents that combine text, images, and visual layout. Layoutlmv3 Base is an open-source project aimed at machine learning developers. Layoutlmv3 Base is open source under the Apache-2.0 license. It ships for the web, and it can be self-hosted.
Behind Layoutlmv3 Base is Microsoft, and it first shipped in 2018. The project is developed in the open on GitHub with 163.4k stars and 794 commits in the last 90 days. Key capabilities include document understanding, text extraction, and image-text processing.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do