Skip to main content

Box2D · Core16

CarRacing

Pixel-observation driving on procedurally generated tracks.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in CarRacingnative env render

GPT-5.5

Codex
BEST
Held-out
604.1
Checkpoint
#009
Rerun
1000 steps

Validation-selected checkpoint rerun in the original environment; case return 499.4.

Claude Opus 4.7 Policy rerun in CarRacingnative env render

Claude Opus 4.7

Claude Code
#2
Held-out
602.3
Checkpoint
#006
Rerun
1000 steps

Validation-selected checkpoint rerun in the original environment; case return 499.4.

MiniMax-M3 Policy rerun in CarRacingnative env render

MiniMax-M3

Claude Code
#3
Held-out
233.4
Checkpoint
#008
Rerun
1000 steps

Validation-selected checkpoint rerun in the original environment; case return 171.9.

DeepSeek-V4-Pro Policy rerun in CarRacingnative env render

DeepSeek-V4-Pro

Claude Code
#4
Held-out
25.20
Checkpoint
#002
Rerun
1000 steps

Validation-selected checkpoint rerun in the original environment; case return -3.509.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01GPT-5.5Codex604.1
02Claude Opus 4.7Claude Code602.3
03MiniMax-M3Claude Code233.4
04DeepSeek-V4-ProClaude Code25.20
05Random policyUniform-29.63