Higgs Audio v3
Higgs Audio v3 TTS is a 4B parameter foundation model built specifically for expressive conversational speech and zero-shot voice cloning across 100+ languages. Expression is driven by embedding <|category:value|> inline tags directly into the text sequence. Global tags like <|emotion:anger|> or <|style:whispering|> sit at the very beginning of a prompt to set the overall tone, speed, and pitch. Positional tags like <|sfx:laughter|> or <|prosody:pause|> are placed mid-sentence right where they occur, paired with explicit onomatopoeia (e.g., Haha) to tell the model exactly when and how to render realistic vocal effects. Higgs Audio v3 is a standalone release and does not depend on the code here. Just call the API.
Explore more models