Skip to main content

Driving · Core16

Roundabout

Structured traffic control through a roundabout.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in Roundabouthighway state renderer

GPT-5.5

Codex
#2
Held-out
9.818
Checkpoint
#012
Rerun
11 steps

Validation-selected checkpoint rerun in the original environment; case return 10.25.

Claude Opus 4.7 Policy rerun in Roundabouthighway state renderer

Claude Opus 4.7

Claude Code
#4
Held-out
9.331
Checkpoint
#004
Rerun
11 steps

Validation-selected checkpoint rerun in the original environment; case return 10.08.

MiniMax-M3 Policy rerun in Roundabouthighway state renderer

MiniMax-M3

Claude Code
#3
Held-out
9.521
Checkpoint
#007
Rerun
8 steps

Validation-selected checkpoint rerun in the original environment; case return 6.500.

DeepSeek-V4-Pro Policy rerun in Roundabouthighway state renderer

DeepSeek-V4-Pro

Claude Code
BEST
Held-out
10.20
Checkpoint
#009
Rerun
11 steps

Validation-selected checkpoint rerun in the original environment; case return 10.17.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01DeepSeek-V4-ProClaude Code10.20
02GPT-5.5Codex9.818
03MiniMax-M3Claude Code9.521
04Claude Opus 4.7Claude Code9.331
05Random policyUniform7.188