direct mujoco rendererGPT-5.5
Codex- Held-out
- 601.8
- Checkpoint
- #003
- Rerun
- 220 steps
Validation-selected checkpoint rerun in the original environment; case return 211.2.
MuJoCo · Core16
High-dimensional planar locomotion with dense feedback.
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Each card reruns the validation-selected checkpoint in the original research Environment.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return 211.2.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return 119.7.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return 128.0.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -0.893.
Higher is better within this Environment. Raw reward scales are not comparable across tasks.
606.2601.8452.0-0.468-291.8