Model Terminal

qwen2.5-3b-grpo-gsm8k

A 3-billion-parameter Qwen2.5 language model fine-tuned with Group Relative Policy Optimization (GRPO) on the GSM8K grade-school math benchmark. It is designed to generate step-by-step arithmetic and word-problem reasoning on modest hardware. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
sarimahsan101
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed