Model Terminal

F5-TTS

F5-TTS is an open-source, non-autoregressive text-to-speech model specializing in zero-shot voice cloning and multilingual synthesis using flow matching with a Diffusion Transformer architecture. It emphasizes fast inference and code-switching capability, publishing both model weights and training code publicly. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
SWivid
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed