MidashengLM-7B is a 7-billion parameter multimodal language model hosted on Hugging Face. It can perceive auditory and visual inputs while generating both text and speech outputs. The model uses a chat template that supports mixed content types including text and audio tokens, making it suitable for researchers and developers building multimodal AI applications.
In the Other AI space, Midashenglm 7b 0804 Fp32 takes a focused approach. It focuses on running and experimenting with a specific open-weight multimodal LLM locally or via Hugging Face. It is built as an open-source project for AI researchers and developers. Midashenglm 7b 0804 Fp32 is open source under the Apache-2.0 license. The product ships for the web and API.
mispeech builds and maintains Midashenglm 7b 0804 Fp32, and the product first shipped in 2025. Development happens publicly on GitHub with 429 stars and 1 commits in the last 90 days. Key capabilities include Multimodal Input, Audio Processing, and Text Generation.
Latest indexed changes and source events
mispeech/midashenglm-7b-0804-fp32 verified by the PulseGate indexer
Other apps tracked under the same category.