Model Terminal
ERNIE 4.5 VL 424B A47B
ERNIE 4.5 VL 424B A47B is Baidu's flagship multimodal vision-language model, using a Mixture-of-Experts architecture with 424B total parameters and 47B activated per token. It accepts text and image inputs, supports a 131,072-token context window, and is instruction-tuned via SFT and reinforcement learning for visual reasoning, document understanding, and visual question answering. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- Baidu AI
- Openness
- —
- Modalities
- —
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —