Model Terminal

gemma-4-12B-it-8bit-mlx

An 8-bit MLX quantized conversion of Google's Gemma 4 12B instruction-tuned model, packaged by Hugging Face user salohcin714 for local inference on Apple Silicon Macs. It reduces memory requirements while maintaining close-to-original model performance for chat and text generation tasks. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
salohcin714
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed