Model Terminal

Fish Audio S2 Pro

Fish Audio S2 Pro is an open-source text-to-speech model supporting 80+ languages with natural-language prosody and emotion control via bracketed style cues. It features a dual-path architecture (4B Slow AR + 400M Fast AR) enabling multi-speaker dialogue, voice cloning, and streaming inference with ~100 ms time-to-first-audio latency. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Fish Audio
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed