native env renderGPT-5.5
Codex- Held-out
- 95.22
- Checkpoint
- #005
- Rerun
- 88 steps
Validation-selected checkpoint rerun in the original environment; case return 95.69.
Control · Core16
Continuous throttle control with delayed goal reward.
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Each card reruns the validation-selected checkpoint in the original research Environment.
native env renderValidation-selected checkpoint rerun in the original environment; case return 95.69.
native env renderValidation-selected checkpoint rerun in the original environment; case return 98.89.
native env renderValidation-selected checkpoint rerun in the original environment; case return 91.61.
native env renderValidation-selected checkpoint rerun in the original environment; case return 94.50.
Higher is better within this Environment. Raw reward scales are not comparable across tasks.
98.7795.2294.4891.88-33.36