MOSS-TTS

A production-grade TTS foundation model with state-of-the-art zero-shot voice cloning, multilingual/code-switched synthesis, and hour-long stable generation, offering token-level duration and phoneme control via a discrete-token autoregressive architecture.

Explore more models