AudioLDM2
A latent text-to-audio diffusion model that generates sound effects, speech, and music from text; it conditions on CLAP and Flan-T5 embeddings with a GPT-2 stage to drive a UNet LDM, and is available via the AudioLDM2 pipeline in Diffusers.
Explore more models