Step3-VL-10B is a 10-billion parameter vision-language model developed by StepFun AI. It supports image and text inputs and includes advanced features such as function/tool calling. The model is distributed on Hugging Face with support for pip and Docker deployment, making it suitable for local inference or integration into custom AI applications.
In the Foundation models & chat space, Step3 VL 10B takes a focused approach. It focuses on providing open multimodal AI models that understand both images and text with built-in tool use for developers. It is built as an open-source project for AI developers and researchers. Step3 VL 10B is open source under the Apache-2.0 license. Step3 VL 10B is available on the web, the command line, and API.
Behind Step3 VL 10B is StepFun AI, based in China, and the product first shipped in 2023. Development happens publicly on GitHub with 86.7k stars and 2.9k commits in the last 90 days. Key capabilities include Vision-Language Model, Tool Calling, and Multimodal Input.
Latest indexed changes and source events
stepfun-ai/Step3-VL-10B verified by the PulseGate indexer
Other apps tracked under the same category.