Model Terminal

Uigyu/qwen_2.5_7b-grpo_sweep_lr3

A Qwen2.5 7B-class language model checkpoint further post-trained with GRPO (Group Relative Policy Optimization) as part of a learning-rate sweep experiment. It is a community fine-tune uploaded to Hugging Face by user Uigyu, inheriting the base capabilities of the Alibaba Qwen Team's Qwen2.5-7B open-weight model. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Uigyu
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed