Skip to main content

Box2D · Core16

Bipedal

Contact-rich locomotion across uneven terrain.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in Bipedalnative env render

GPT-5.5

Codex
BEST
Held-out
248.9
Checkpoint
#010
Rerun
1600 steps

Validation-selected checkpoint rerun in the original environment; case return 273.3.

Claude Opus 4.7 Policy rerun in Bipedalnative env render

Claude Opus 4.7

Claude Code
#2
Held-out
-15.84
Checkpoint
#027
Rerun
1600 steps

Validation-selected checkpoint rerun in the original environment; case return -15.30.

MiniMax-M3 Policy rerun in Bipedalnative env render

MiniMax-M3

Claude Code
#3
Held-out
-80.87
Checkpoint
#001
Rerun
217 steps

Validation-selected checkpoint rerun in the original environment; case return -68.83.

DeepSeek-V4-Pro Policy rerun in Bipedalnative env render

DeepSeek-V4-Pro

Claude Code
#4
Held-out
-97.48
Checkpoint
#004
Rerun
113 steps

Validation-selected checkpoint rerun in the original environment; case return -97.29.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01GPT-5.5Codex248.9
02Claude Opus 4.7Claude Code-15.84
03MiniMax-M3Claude Code-80.87
04DeepSeek-V4-ProClaude Code-97.48
05Random policyUniform-101.0