Rex-Omni-AWQ is an AWQ-quantized multimodal foundation model from IDEA Research. It processes images, videos, and text within a unified architecture, supporting tasks such as visual question answering, video captioning, and multimodal reasoning. The model is designed for efficient inference on hardware with limited resources and is distributed via Hugging Face for use with Transformers and compatible inference frameworks.
In the Foundation models & chat space, Rex Omni takes a focused approach. It focuses on integrating vision, video, and language understanding into a single efficient model for multimodal AI applications. Rex Omni is an open-source project aimed at developers. The project is open source (Open Source). It ships for the web, the command line, and API.
Behind Rex Omni is IDEA-Research, and it first shipped in 2024. Among its 4 catalogued features are multimodal input, vision-language, and video understanding.
Summary written by a language model from the project’s public pages.
What PulseGate has recorded for this listing
Closest matches by what these projects do