Model comparison
GOT-OCR2.0 vs o3
Capability radar, value bars, and side-by-side posture.
GOT-OCR2.0
Not scored
value score
o3
93
value score
GOT-OCR2.0 ctx
—
o3 ctx
200K
Capability radar
Value · context · multimodal · openness · speed posture
Value head-to-head
- o393
Attribute tape
| Field | GOT-OCR2.0 | o3 |
|---|---|---|
| Developer | stepfun-ai | OpenAI |
| Context | — | 200K |
| Modalities | Text, Code | |
| Openness | — | Proprietary Model |
| Speed | — | — |
| Price | — | $10/1M in · $40/1M out |
| Value score | — | 93 |
| Summary | GOT-OCR2.0 is a 0.7B-parameter unified end-to-end OCR model from StepFun that extracts text from images including scene text, scanned documents, tables, math formulas, charts, geometric shapes, molecular formulas, and sheet music. It supports structured formatted output and interactive region-level OCR via prompts, going beyond plain-text extraction. | OpenAI reasoning model focused on hard problems. |