Model Terminal

DeepSeek R1 Distill Llama 70B

An open-weight 70B reasoning model produced by distilling DeepSeek-R1's chain-of-thought behavior into Meta's Llama-3.3-70B-Instruct backbone. It targets math, coding, and logic tasks with strong reasoning quality at a fraction of the compute cost of the full 671B DeepSeek-R1. Source: supabase.

Value score
Not scored
Context
128K
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
~deepseek
Openness
Open Weights
Modalities
Release
2025-01-20
Knowledge cutoff
Deprecation

Benchmarks

  • GPQA71.5 · 2025-01-20 · source

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed