Model Terminal
Qwen2.5-Math-15B GRPO ACEMath Fine-tune (Jongbin-kr)
A community fine-tune of Alibaba's Qwen2.5-Math-15B base model, applying GRPO-style reinforcement learning post-training on the ACEMath dataset to improve step-by-step mathematical reasoning. It is not a standalone product or organization, but a research checkpoint uploaded to Hugging Face by individual user Jongbin-kr. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- Jongbin-kr
- Openness
- —
- Modalities
- —
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —