Model Terminal

XTTS-v2

XTTS-v2 is a multilingual text-to-speech and zero-shot voice cloning model originally developed by Coqui, capable of generating natural speech in 17 languages from as little as 6 seconds of reference audio. Coqui Inc. shut down in January 2024; the model remains publicly available and community-maintained on Hugging Face. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Hugging Face
Openness
Modalities
Text-to-speech
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed