Scenema Audio
An expressive speech and ambience generator built on an audio diffusion transformer, supporting action-tag performance control, scene-aware background audio, zero-shot voice cloning from short references, and multilingual synthesis across 13 languages.
Explore more models