Model Terminal

nvidia/Qwen3.6-35B-A3B-NVFP4

A quantized inference checkpoint of Alibaba's Qwen3.6-35B-A3B mixture-of-experts language model, published by NVIDIA in its NVFP4 format for efficient deployment on NVIDIA GPUs. The checkpoint targets reduced VRAM usage and faster inference on Blackwell-era hardware while preserving the original 35B-total/3B-active architecture. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Peer value bars

Identity

Developer
Hugging Face
Openness
Modalities
Text-generation
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed