Model Terminal
Fish Audio S2 Pro
Fish Audio S2 Pro is an open-source text-to-speech model supporting 80+ languages with natural-language prosody and emotion control via bracketed style cues. It features a dual-path architecture (4B Slow AR + 400M Fast AR) enabling multi-speaker dialogue, voice cloning, and streaming inference with ~100 ms time-to-first-audio latency. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- Fish Audio
- Openness
- —
- Modalities
- —
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —