Skip to main content

MiniGrid · Core16

FourRooms

Partial-observation navigation through bottlenecks.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in FourRoomsnative env render

GPT-5.5

Codex
#2
Held-out
0.664
Checkpoint
#013
Rerun
38 steps

Validation-selected checkpoint rerun in the original environment; case return 0.658.

Claude Opus 4.7 Policy rerun in FourRoomsnative env render

Claude Opus 4.7

Claude Code
BEST
Held-out
0.701
Checkpoint
#009
Rerun
32 steps

Validation-selected checkpoint rerun in the original environment; case return 0.712.

MiniMax-M3 Policy rerun in FourRoomsnative env render

MiniMax-M3

Claude Code
#3
Held-out
0.403
Checkpoint
#009
Rerun
100 steps

Validation-selected checkpoint rerun in the original environment; case return 0.000.

DeepSeek-V4-Pro Policy rerun in FourRoomsnative env render

DeepSeek-V4-Pro

Claude Code
#4
Held-out
0.079
Checkpoint
#036
Rerun
100 steps

Validation-selected checkpoint rerun in the original environment; case return 0.000.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01Claude Opus 4.7Claude Code0.701
02GPT-5.5Codex0.664
03MiniMax-M3Claude Code0.403
04DeepSeek-V4-ProClaude Code0.079
05Random policyUniform0.042