Skip to main content

Control · Core16

Acrobot

Swing-up control with a two-link underactuated arm.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in Acrobotnative env render

GPT-5.5

Codex
BEST
Held-out
-84.69
Checkpoint
#001
Rerun
76 steps

Validation-selected checkpoint rerun in the original environment; case return -75.00.

Claude Opus 4.7 Policy rerun in Acrobotnative env render

Claude Opus 4.7

Claude Code
#2
Held-out
-88.59
Checkpoint
#011
Rerun
74 steps

Validation-selected checkpoint rerun in the original environment; case return -73.00.

MiniMax-M3 Policy rerun in Acrobotnative env render

MiniMax-M3

Claude Code
#4
Held-out
-136.4
Checkpoint
#010
Rerun
249 steps

Validation-selected checkpoint rerun in the original environment; case return -248.0.

DeepSeek-V4-Pro Policy rerun in Acrobotnative env render

DeepSeek-V4-Pro

Claude Code
#3
Held-out
-91.16
Checkpoint
#007
Rerun
130 steps

Validation-selected checkpoint rerun in the original environment; case return -129.0.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01GPT-5.5Codex-84.69
02Claude Opus 4.7Claude Code-88.59
03DeepSeek-V4-ProClaude Code-91.16
04MiniMax-M3Claude Code-136.4
05Random policyUniform-499.2