Model comparison
lagessiehcs/ppo-LunarLander-v2 vs Claude 4 Opus
Capability radar, value bars, and side-by-side posture.
lagessiehcs/ppo-LunarLander-v2
Not scored
value score
Claude 4 Opus
95
value score
lagessiehcs/ppo-LunarLander-v2 ctx
—
Claude 4 Opus ctx
200K
Capability radar
Value · context · multimodal · openness · speed posture
Value head-to-head
- Claude 4 Opus95
Attribute tape
| Field | lagessiehcs/ppo-LunarLander-v2 | Claude 4 Opus |
|---|---|---|
| Developer | Safe Superintelligence | Anthropic |
| Context | — | 200K |
| Modalities | Text, Code, Image | |
| Openness | — | Proprietary Model |
| Speed | — | — |
| Price | — | $15/1M in · $75/1M out |
| Value score | — | 95 |
| Summary | A PPO-trained reinforcement learning policy for the LunarLander-v2 benchmark environment, where a neural-network agent learns to control a lunar lander's thrusters for safe landing. Trained with Stable-Baselines3 and hosted as a personal model artifact on the Hugging Face Hub. | Anthropic frontier Opus-class model. |