This Hugging Face repository provides a GGUF-format DeepSeek-V4 Flash model checkpoint for local inference. It includes guidance for llama-server and SGLang usage, configurable thinking modes, reasoning effort, and OpenAI-shaped tool calling for developers running compatible inference stacks.
DeepSeek V4 Flash 0731 Reap 150b sits in PulseGate's Foundation models & chat category. It focuses on running a DeepSeek language model locally for text generation, reasoning, coding, and tool calling. It is built as an open-source project for developers and machine learning practitioners. The project is open source (MIT). It ships for the web, the command line, and API, and it can be self-hosted.
puwaer builds and maintains DeepSeek V4 Flash 0731 Reap 150b, and it first shipped in 2026. The project is developed in the open on GitHub with 49 commits in the last 90 days. It operates in a well-populated space: PulseGate tracks 7 similar projects. Key capabilities include local inference, GGUF format, and thinking mode. It exposes integrations via a public API.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do