native env renderGPT-5.5
Codex- Held-out
- -84.69
- Checkpoint
- #001
- Rerun
- 76 steps
Validation-selected checkpoint rerun in the original environment; case return -75.00.
Control · Core16
Swing-up control with a two-link underactuated arm.
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Each card reruns the validation-selected checkpoint in the original research Environment.
native env renderValidation-selected checkpoint rerun in the original environment; case return -75.00.
native env renderValidation-selected checkpoint rerun in the original environment; case return -73.00.
native env renderValidation-selected checkpoint rerun in the original environment; case return -248.0.
native env renderValidation-selected checkpoint rerun in the original environment; case return -129.0.
Higher is better within this Environment. Raw reward scales are not comparable across tasks.
-84.69-88.59-91.16-136.4-499.2