跳到主要内容

MiniGrid · Core16

KeyCorridor

需要物体交互与记忆的多房间搜索。

历史研究记录

这些证据属于论文时期实验,并不表示相关 package 当前仍然可用。

01

最终 Policy 证据

每张卡片都在原研究 Environment 中重跑 validation 选出的 checkpoint。

GPT-5.5 Policy rerun in KeyCorridornative env render

GPT-5.5

Codex
#2
Held-out 分数
0.918
Checkpoint
#005
重跑
25 steps

Validation-selected checkpoint rerun in the original environment; case return 0.953.

Claude Opus 4.7 Policy rerun in KeyCorridornative env render

Claude Opus 4.7

Claude Code
BEST
Held-out 分数
0.930
Checkpoint
#003
重跑
26 steps

Validation-selected checkpoint rerun in the original environment; case return 0.951.

MiniMax-M3 Policy rerun in KeyCorridornative env render

MiniMax-M3

Claude Code
#3
Held-out 分数
0.000
Checkpoint
#012
重跑
220 steps

Validation-selected checkpoint rerun in the original environment; case return 0.000.

DeepSeek-V4-Pro Policy rerun in KeyCorridornative env render

DeepSeek-V4-Pro

Claude Code
#4
Held-out 分数
0.000
Checkpoint
#023
重跑
220 steps

Validation-selected checkpoint rerun in the original environment; case return 0.000.

02

报告分数

在该 Environment 内分数越高越好;不同任务的原始 reward scale 不可比较。

01Claude Opus 4.7Claude Code0.930
02GPT-5.5Codex0.918
03MiniMax-M3Claude Code0.000
04DeepSeek-V4-ProClaude Code0.000
05Random policyUniform0.000