An FP8 quantized variant of the DeepSeek-V4 model published by the SGL project. It is optimized for high-throughput inference while maintaining strong performance. Intended for developers building applications that require fast generation from a very large language model.
In the Quantised & converted weights space, DeepSeek V4 Flash takes a focused approach. It focuses on enabling high-speed inference of a large DeepSeek model using reduced precision formats. It is built as an open-source project for developers. The project is open source (Open Source). It runs on the web and API.
SGL Project builds and maintains DeepSeek V4 Flash. PulseGate's similarity index places it among 11 comparable projects.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do