Skip to main content

Robotics · Core16

FetchPickAndPlace

Grasp, lift, and place an object at a target.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in FetchPickAndPlacedirect mujoco renderer

GPT-5.5

Codex
BEST
Held-out
-11.81
Checkpoint
#007
Rerun
50 steps

Validation-selected checkpoint rerun in the original environment; case return -12.00.

Claude Opus 4.7 Policy rerun in FetchPickAndPlacedirect mujoco renderer

Claude Opus 4.7

Claude Code
#3
Held-out
-13.34
Checkpoint
#008
Rerun
50 steps

Validation-selected checkpoint rerun in the original environment; case return -13.00.

MiniMax-M3 Policy rerun in FetchPickAndPlacedirect mujoco renderer

MiniMax-M3

Claude Code
#2
Held-out
-13.22
Checkpoint
#006
Rerun
50 steps

Validation-selected checkpoint rerun in the original environment; case return -13.00.

DeepSeek-V4-Pro Policy rerun in FetchPickAndPlacedirect mujoco renderer

DeepSeek-V4-Pro

Claude Code
#4
Held-out
-20.56
Checkpoint
#011
Rerun
50 steps

Validation-selected checkpoint rerun in the original environment; case return -35.00.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01GPT-5.5Codex-11.81
02MiniMax-M3Claude Code-13.22
03Claude Opus 4.7Claude Code-13.34
04DeepSeek-V4-ProClaude Code-20.56
05Random policyUniform-43.75