Model Terminal

VibeVoice-ASR

VibeVoice-ASR is Microsoft's open-source long-form automatic speech recognition model that transcribes up to 60 minutes of continuous audio in a single pass, producing structured output with speaker labels and timestamps. It is designed for meeting transcription, podcast analysis, contact center workflows, and other long-form audio understanding tasks. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Microsoft AI
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed