Model Terminal
Uigyu/qwen_2.5_7b-grpo_sweep_lr3
A Qwen2.5 7B-class language model checkpoint further post-trained with GRPO (Group Relative Policy Optimization) as part of a learning-rate sweep experiment. It is a community fine-tune uploaded to Hugging Face by user Uigyu, inheriting the base capabilities of the Alibaba Qwen Team's Qwen2.5-7B open-weight model. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- Uigyu
- Openness
- —
- Modalities
- —
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —