Model Terminal

lagessiehcs/ppo-LunarLander-v2

A PPO-trained reinforcement learning policy for the LunarLander-v2 benchmark environment, where a neural-network agent learns to control a lunar lander's thrusters for safe landing. Trained with Stable-Baselines3 and hosted as a personal model artifact on the Hugging Face Hub. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed