Skip to main content

Driving · Core16

Parking

Continuous vehicle control toward a target parking pose.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in Parkinghighway state renderer

GPT-5.5

Codex
BEST
Held-out
-30.24
Checkpoint
#012
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return -48.79.

Claude Opus 4.7 Policy rerun in Parkinghighway state renderer

Claude Opus 4.7

Claude Code
#3
Held-out
-38.18
Checkpoint
#010
Rerun
116 steps

Validation-selected checkpoint rerun in the original environment; case return -46.70.

MiniMax-M3 Policy rerun in Parkinghighway state renderer

MiniMax-M3

Claude Code
#2
Held-out
-32.71
Checkpoint
#013
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return -45.59.

DeepSeek-V4-Pro Policy rerun in Parkinghighway state renderer

DeepSeek-V4-Pro

Claude Code
#4
Held-out
-53.38
Checkpoint
#003
Rerun
54 steps

Validation-selected checkpoint rerun in the original environment; case return -24.69.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01GPT-5.5Codex-30.24
02MiniMax-M3Claude Code-32.71
03Claude Opus 4.7Claude Code-38.18
04Random policyUniform-47.45
05DeepSeek-V4-ProClaude Code-53.38