This is a highly specialized, NVIDIA-optimized version of the Qwen3-VL 235B-A22B vision-language model prepared for the MLPerf Inference Closed division benchmark (v6.1). It uses NVFP4 and FP8 KV cache quantization for maximum performance on NVIDIA hardware. The model supports multimodal inputs and advanced tool-calling capabilities via its chat template.
In the Foundation models & chat space, 1 Fp8 Kv takes a focused approach. It focuses on benchmarking high-performance inference of large vision-language models using optimized quantization formats. It is built as an open-source project for MLPerf benchmark participants and large model inference engineers. 1 Fp8 Kv is open source under the Open Source license. The product ships for the web and API.
Behind 1 Fp8 Kv is nvidia.
Latest indexed changes and source events
nvidia/Qwen3-VL-235B-A22B-Instruct-NVFP4-MLPerf-Inference-Closed-V6.1-FP8-KV verified by the PulseGate indexer
Other apps tracked under the same category.