MMAG is a comprehensive benchmark for evaluating mixed audio generation under multiple control conditions. It assesses a model's ability to generate coherent acoustic scenes containing speech, music, and sound effects simultaneously, while supporting fine-grained control over speaker identity and temporal alignment.
- Evaluation pipeline
- Raw multi-expert model pipeline
- New version release (more precise sound event and timestamp annotations)
We evaluate the following models on MMAG:
| Model | Description |
|---|---|
| AudioDirector | Agentic orchestrator |
| LTX-2.3 | Unified audio-visual generation |
| Ovi 1.1 | Unified audio-visual generation |
| MOVA-720p | Unified audio-visual generation |
| JavisDiT++ | Unified audio-visual generation |
| UniAVGen | Unified audio-visual generation |
| Dasheng-AudioGen | Native mixed-audio generation |
| Ming-Omni-TTS | Native mixed-audio generation |
We thank the authors of MECAT for their valuable work.
This project is licensed under the CC BY 4.0 license.