Stable Audio Open 1.0
An open text-to-audio latent diffusion model from Stability AI that generates up to ~47 seconds of stereo audio at 44.1 kHz from text prompts, built with an autoencoder, T5-based text conditioning, and a DiT transformer operating in latent space.
Explore more models