Model Terminal
dree25/llama-3.2-3b-grpo-model
A community fine-tune of Meta's Llama 3.2 3B Instruct adapted with GRPO (Group Relative Policy Optimization) reinforcement learning, targeting improved reasoning or instruction-following over the base model. Published as an open-weight model artifact on Hugging Face by individual user dree25. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- dree25
- Openness
- —
- Modalities
- —
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —