Model Terminal

train_llama8b_grpo_single_dartmath_fromsft

A community fine-tuned Llama 8B model checkpoint trained with GRPO reinforcement-learning optimization on top of an SFT base, targeting math and reasoning tasks via a dartmath-style dataset. Released as an individual experiment on Hugging Face by user YukinoshitaYukino. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Unknown
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed