Model Terminal

VibeVoice-Realtime-0.5B

VibeVoice-Realtime-0.5B is a compact, streaming-first text-to-speech model from Microsoft that begins generating audible speech before the full input text is complete, targeting interactive voice applications. It is English-only and optimized for single-speaker low-latency inference at roughly 0.5B parameters. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Microsoft AI
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed