Skip to main content

Control · Core16

ContinuousCar

Continuous throttle control with delayed goal reward.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in ContinuousCarnative env render

GPT-5.5

Codex
#2
Held-out
95.22
Checkpoint
#005
Rerun
88 steps

Validation-selected checkpoint rerun in the original environment; case return 95.69.

Claude Opus 4.7 Policy rerun in ContinuousCarnative env render

Claude Opus 4.7

Claude Code
BEST
Held-out
98.77
Checkpoint
#010
Rerun
344 steps

Validation-selected checkpoint rerun in the original environment; case return 98.89.

MiniMax-M3 Policy rerun in ContinuousCarnative env render

MiniMax-M3

Claude Code
#4
Held-out
91.88
Checkpoint
#010
Rerun
143 steps

Validation-selected checkpoint rerun in the original environment; case return 91.61.

DeepSeek-V4-Pro Policy rerun in ContinuousCarnative env render

DeepSeek-V4-Pro

Claude Code
#3
Held-out
94.48
Checkpoint
#012
Rerun
69 steps

Validation-selected checkpoint rerun in the original environment; case return 94.50.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01Claude Opus 4.7Claude Code98.77
02GPT-5.5Codex95.22
03DeepSeek-V4-ProClaude Code94.48
04MiniMax-M3Claude Code91.88
05Random policyUniform-33.36