Model comparison
dree25/llama-3.2-3b-grpo-model vs o3
Capability radar, value bars, and side-by-side posture.
dree25/llama-3.2-3b-grpo-model
Not scored
value score
o3
93
value score
dree25/llama-3.2-3b-grpo-model ctx
—
o3 ctx
200K
Capability radar
Value · context · multimodal · openness · speed posture
Value head-to-head
- o393
Attribute tape
| Field | dree25/llama-3.2-3b-grpo-model | o3 |
|---|---|---|
| Developer | dree25 | OpenAI |
| Context | — | 200K |
| Modalities | Text, Code | |
| Openness | — | Proprietary Model |
| Speed | — | — |
| Price | — | $10/1M in · $40/1M out |
| Value score | — | 93 |
| Summary | A community fine-tune of Meta's Llama 3.2 3B Instruct adapted with GRPO (Group Relative Policy Optimization) reinforcement learning, targeting improved reasoning or instruction-following over the base model. Published as an open-weight model artifact on Hugging Face by individual user dree25. | OpenAI reasoning model focused on hard problems. |