Skip to main content

MuJoCo · Core16

HalfCheetah

High-dimensional planar locomotion with dense feedback.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in HalfCheetahdirect mujoco renderer

GPT-5.5

Codex
#2
Held-out
601.8
Checkpoint
#003
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return 211.2.

Claude Opus 4.7 Policy rerun in HalfCheetahdirect mujoco renderer

Claude Opus 4.7

Claude Code
#3
Held-out
452.0
Checkpoint
#038
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return 119.7.

MiniMax-M3 Policy rerun in HalfCheetahdirect mujoco renderer

MiniMax-M3

Claude Code
BEST
Held-out
606.2
Checkpoint
#040
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return 128.0.

DeepSeek-V4-Pro Policy rerun in HalfCheetahdirect mujoco renderer

DeepSeek-V4-Pro

Claude Code
#4
Held-out
-0.468
Checkpoint
#017
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return -0.893.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01MiniMax-M3Claude Code606.2
02GPT-5.5Codex601.8
03Claude Opus 4.7Claude Code452.0
04DeepSeek-V4-ProClaude Code-0.468
05Random policyUniform-291.8