Model Terminal

nanoMoE

nanoMoE is a small-scale, open-source Mixture-of-Experts (MoE) language model and training codebase forked from nanoGPT, designed to help researchers and developers understand and experiment with sparse MoE transformer architectures on modest hardware. It implements a 6-layer decoder-only transformer with 8 total experts and 2 active experts per token, trained on OpenWebText. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Unknown
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed