Scenema Audio

An expressive speech and ambience generator built on an audio diffusion transformer, supporting action-tag performance control, scene-aware background audio, zero-shot voice cloning from short references, and multilingual synthesis across 13 languages.

Explore more models