direct mujoco rendererGPT-5.5
Codex- Held-out
- -37.11
- Checkpoint
- #005
- Rerun
- 100 steps
Validation-selected checkpoint rerun in the original environment; case return -38.36.
MuJoCo · Core16
Contact manipulation that pushes an object to a goal.
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Each card reruns the validation-selected checkpoint in the original research Environment.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -38.36.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -37.53.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -41.99.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -55.45.
Higher is better within this Environment. Raw reward scales are not comparable across tasks.
-37.11-38.52-39.25-54.18-147.5