Model Terminal

dree25/llama-3.2-3b-grpo-model

A community fine-tune of Meta's Llama 3.2 3B Instruct adapted with GRPO (Group Relative Policy Optimization) reinforcement learning, targeting improved reasoning or instruction-following over the base model. Published as an open-weight model artifact on Hugging Face by individual user dree25. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
dree25
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed