NVIDIA-Nemotron-Parse-v1.1 is an image-text-to-text model developed by NVIDIA for multimodal parsing tasks. It processes both visual and textual inputs to generate structured outputs. The model is available on Hugging Face, supports the Transformers library, and is suitable for developers building vision-language applications.
NVIDIA Nemotron Parse sits in PulseGate's Foundation models & chat category. It focuses on understanding and parsing combined image and text inputs using a single specialized model. It is built as an open-source project for developers. NVIDIA Nemotron Parse is open source under the Apache-2.0 license. The product ships for the web, the command line, and API.
NVIDIA builds and maintains NVIDIA Nemotron Parse, and the product first shipped in 2023. Development happens publicly on GitHub with 87.1k stars and 3k commits in the last 90 days. Key capabilities include Image-Text Parsing, Multimodal Understanding, and Transformers Compatible.
Latest indexed changes and source events
nvidia/NVIDIA-Nemotron-Parse-v1.1 verified by the PulseGate indexer
Other apps tracked under the same category.