Audiobox Aesthetics is a model for unified automatic quality assessment of speech, music, and general sound. Developed by AI at Meta, it addresses the need for a single system that can evaluate audio quality across different audio types rather than requiring separate specialized tools.
The model is provided as a PyTorch implementation with Safetensors weights. It supports prediction through a command-line interface as well as a Python API or interpreter. Installation is available via pip with the package audiobox_aesthetics, or by installing directly from source after cloning the repository. The repository specifies a requirement of Python 3.9 and PyTorch 2.2 or greater.
It is hosted on Hugging Face under the identifier facebook/audiobox-aesthetics. The associated paper appears on arXiv as 2502.05139, and the source code resides at github.com/facebookresearch/audiobox-aesthetics. The model card indicates it was pushed to the Hub using the PytorchModelHubMixin integration.
The license is cc-by-4.0. No pricing information is stated because the model is distributed as an open artifact for research and development use.
In the Voice, TTS & speech space, Audiobox Aesthetics takes a focused approach. It focuses on evaluating the quality of generated or recorded audio for speech, music, and general sound without manual review. Audiobox Aesthetics is an open-source project aimed at AI researchers and developers. The project is open source (CC-BY-4.0). Audiobox Aesthetics is available on the web, the command line, and API.
Meta AI builds and maintains Audiobox Aesthetics, and the product first shipped in 2025. The project is developed in the open on GitHub with 744 stars. Among its 4 catalogued features are Audio Quality Assessment, CLI Prediction, and Python Integration.
Latest indexed changes and source events
facebook/audiobox-aesthetics verified by the PulseGate indexer
Other apps tracked under the same category.