Model comparison
VibeVoice-ASR vs Phi-3.5 Mini
Capability radar, value bars, and side-by-side posture.
VibeVoice-ASR
Not scored
value score
Phi-3.5 Mini
71
value score
VibeVoice-ASR ctx
—
Phi-3.5 Mini ctx
128K
Capability radar
Value · context · multimodal · openness · speed posture
Value head-to-head
- Phi-3.5 Mini71
Attribute tape
| Field | VibeVoice-ASR | Phi-3.5 Mini |
|---|---|---|
| Developer | Microsoft AI | Microsoft AI |
| Context | — | 128K |
| Modalities | Text, Code | |
| Openness | — | Open Weights |
| Speed | — | — |
| Price | — | — |
| Value score | — | 71 |
| Summary | VibeVoice-ASR is Microsoft's open-source long-form automatic speech recognition model that transcribes up to 60 minutes of continuous audio in a single pass, producing structured output with speaker labels and timestamps. It is designed for meeting transcription, podcast analysis, contact center workflows, and other long-form audio understanding tasks. | Compact Phi model for edge and cheap inference. |