An FP4 quantized and optimized version of the DeepSeek-R1 model provided by NVIDIA. It is designed for efficient inference on NVIDIA GPUs while maintaining reasoning capabilities. The model uses a custom chat template and is suitable for developers building AI applications that require strong reasoning performance with lower memory footprint.
DeepSeek R1 0528 NVFP4 sits in PulseGate's Quantised & converted weights category. It focuses on running large reasoning models efficiently on NVIDIA hardware with reduced precision. It is built as an open-source project for developers. The project is open source (Apache-2.0). It runs on the web, and it can be self-hosted.
Behind DeepSeek R1 0528 NVFP4 is NVIDIA, based in the United States, and it first shipped in 2024. Development happens publicly on GitHub with 3.3k stars and 354 commits in the last 90 days. Among its 3 catalogued features are Quantized Inference, Reasoning Model, and NVIDIA Optimized.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do