Skip to main content

MuJoCo · Core16

Pusher

Contact manipulation that pushes an object to a goal.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in Pusherdirect mujoco renderer

GPT-5.5

Codex
BEST
Held-out
-37.11
Checkpoint
#005
Rerun
100 steps

Validation-selected checkpoint rerun in the original environment; case return -38.36.

Claude Opus 4.7 Policy rerun in Pusherdirect mujoco renderer

Claude Opus 4.7

Claude Code
#2
Held-out
-38.52
Checkpoint
#014
Rerun
100 steps

Validation-selected checkpoint rerun in the original environment; case return -37.53.

MiniMax-M3 Policy rerun in Pusherdirect mujoco renderer

MiniMax-M3

Claude Code
#3
Held-out
-39.25
Checkpoint
#010
Rerun
100 steps

Validation-selected checkpoint rerun in the original environment; case return -41.99.

DeepSeek-V4-Pro Policy rerun in Pusherdirect mujoco renderer

DeepSeek-V4-Pro

Claude Code
#4
Held-out
-54.18
Checkpoint
#004
Rerun
100 steps

Validation-selected checkpoint rerun in the original environment; case return -55.45.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01GPT-5.5Codex-37.11
02Claude Opus 4.7Claude Code-38.52
03MiniMax-M3Claude Code-39.25
04DeepSeek-V4-ProClaude Code-54.18
05Random policyUniform-147.5