direct mujoco rendererGPT-5.5
Codex- Held-out
- -3.473
- Checkpoint
- #012
- Rerun
- 50 steps
Validation-selected checkpoint rerun in the original environment; case return -1.880.
MuJoCo · Core16
Continuous robotic-arm control toward a target point.
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Each card reruns the validation-selected checkpoint in the original research Environment.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -1.880.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -2.670.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -2.308.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -4.613.
Higher is better within this Environment. Raw reward scales are not comparable across tasks.
-3.473-3.979-5.103-6.506-43.77