Model Terminal
qwen2.5-3b-grpo-gsm8k
A 3-billion-parameter Qwen2.5 language model fine-tuned with Group Relative Policy Optimization (GRPO) on the GSM8K grade-school math benchmark. It is designed to generate step-by-step arithmetic and word-problem reasoning on modest hardware. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- sarimahsan101
- Openness
- —
- Modalities
- —
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —