Model Terminal
rl_core-e2-d-rl3000-seed42-verified-v2
A reinforcement-learning-with-verifiable-rewards (RLVR) model checkpoint produced by the Pre-to-Post-2 project, trained for 3,000 RL steps with seed 42 and filtered to verified outputs. It represents a post-training alignment experiment in which a language model is fine-tuned using deterministic correctness signals rather than human preference labels. Source: supabase.
Value score
Not scored
Context
—
tokens
Max output
—
tokens
Price
—
Capability radar
Peer value bars
Identity
- Developer
- Pre-to-Post-2
- Openness
- —
- Modalities
- —
- Release
- —
- Knowledge cutoff
- —
- Deprecation
- —
- API docs
- —
Benchmarks
No benchmark scores yet.
Compare nearby
Pricing
- Input / 1M
- —
- Output / 1M
- —
- Speed
- —