NVIDIA-Nemotron-Parse-v1.1 is an image-text-to-text model developed by NVIDIA for multimodal parsing tasks. It processes both visual and textual inputs to generate structured outputs. The model is available on Hugging Face, supports the Transformers library, and is suitable for developers building vision-language applications.
NVIDIA Nemotron Parse sits in PulseGate's Multimodal & vision category. It focuses on understanding and parsing combined image and text inputs using a single specialized model. NVIDIA Nemotron Parse is an open-source project aimed at developers. NVIDIA Nemotron Parse is open source under the Apache-2.0 license. It runs on the web, the command line, and API.
Behind NVIDIA Nemotron Parse is NVIDIA, based in the United States, and it first shipped in 2023. The project is developed in the open on GitHub with 87.1k stars and 3k commits in the last 90 days. Among its 3 catalogued features are Image-Text Parsing, Multimodal Understanding, and Transformers Compatible.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do