Model comparison

train_llama8b_grpo_single_dartmath_fromsft vs o3

Capability radar, value bars, and side-by-side posture.

← Back to Compare

train_llama8b_grpo_single_dartmath_fromsft
Not scored
value score
o3
93
value score
train_llama8b_grpo_single_dartmath_fromsft ctx
o3 ctx
200K

Capability radar

Value · context · multimodal · openness · speed posture

Value head-to-head

  • o393

Attribute tape

Fieldtrain_llama8b_grpo_single_dartmath_fromsfto3
DeveloperUnknownOpenAI
Context200K
ModalitiesText, Code
OpennessProprietary Model
Speed
Price$10/1M in · $40/1M out
Value score93
SummaryA community fine-tuned Llama 8B model checkpoint trained with GRPO reinforcement-learning optimization on top of an SFT base, targeting math and reasoning tasks via a dartmath-style dataset. Released as an individual experiment on Hugging Face by user YukinoshitaYukino.OpenAI reasoning model focused on hard problems.
Open train_llama8b_grpo_single_dartmath_fromsftOpen o3