NVLM-D-72B is NVIDIA's large vision-language model capable of processing both images and text. It supports advanced tool calling and follows a Qwen-style chat template. The model is designed for complex multimodal reasoning tasks.
In the Foundation models & chat space, NVLM D 72B takes a focused approach. It focuses on understanding and reasoning over both text and visual inputs at high performance. NVLM D 72B is an open-source project aimed at AI developers. The project is open source (Open Source). NVLM D 72B is available on the web, the command line, and API.
It is developed by NVIDIA (United States), and the product first shipped in 2019. The project is developed in the open on GitHub with 17.1k stars and 591 commits in the last 90 days. Among its 3 catalogued features are Vision Language, Tool Use, and multimodal.
Latest indexed changes and source events
nvidia/NVLM-D-72B verified by the PulseGate indexer