direct mujoco rendererGPT-5.5
Codex- Held-out
- -11.81
- Checkpoint
- #007
- Rerun
- 50 steps
Validation-selected checkpoint rerun in the original environment; case return -12.00.
Robotics · Core16
Grasp, lift, and place an object at a target.
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Each card reruns the validation-selected checkpoint in the original research Environment.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -12.00.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -13.00.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -13.00.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -35.00.
Higher is better within this Environment. Raw reward scales are not comparable across tasks.
-11.81-13.22-13.34-20.56-43.75