Step-3.5-Flash is a multimodal language model from StepFun AI that supports text, image inputs, and tool/function calling. It features a sophisticated Jinja chat template for handling system prompts, tools in JSON schema format, and mixed content types including image patches. The model is available on Hugging Face for use with Transformers or compatible inference engines.
Step 3.5 Flash is a Multimodal & vision project. It focuses on accessing a fast, multimodal LLM with native tool use without building or hosting the model yourself. It is built as an open-source project for developers. The project is open source (Apache-2.0). It ships for the command line and API.
Behind Step 3.5 Flash is StepFun AI, based in China, and it first shipped in 2026. The project is developed in the open on GitHub with 2.1k stars. Among its 4 catalogued features are Tool Calling, Multimodal Input, and Chat Template.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do