Model Terminal

rl_core-e2-d-rl3000-seed42-verified-v2

A reinforcement-learning-with-verifiable-rewards (RLVR) model checkpoint produced by the Pre-to-Post-2 project, trained for 3,000 RL steps with seed 42 and filtered to verified outputs. It represents a post-training alignment experiment in which a language model is fine-tuned using deterministic correctness signals rather than human preference labels. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Developer
Pre-to-Post-2
Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed