native env renderGPT-5.5
Codex- Held-out
- 0.904
- Checkpoint
- #021
- Rerun
- 90 steps
Validation-selected checkpoint rerun in the original environment; case return 0.887.
MiniGrid · Core16
Maze navigation with blockers and interaction sequences.
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Each card reruns the validation-selected checkpoint in the original research Environment.
native env renderValidation-selected checkpoint rerun in the original environment; case return 0.887.
native env renderValidation-selected checkpoint rerun in the original environment; case return 0.940.
native env renderValidation-selected checkpoint rerun in the original environment; case return 0.000.
native env renderValidation-selected checkpoint rerun in the original environment; case return 0.000.
Higher is better within this Environment. Raw reward scales are not comparable across tasks.
0.9110.9040.0000.0000.000