Skip to main content

Paper companion · v0.1.0

Core16 experiment record

Held-out returns and final-Policy reruns from four coding-agent/model lanes across sixteen interactive tasks.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Held-out score matrix

Scores are comparable only within one Environment column. Highlighted values are column maxima.

ModelAcrobotContinuousCarBipedalCarRacingReacherHalfCheetahAntPusherDoorKeyKeyCorridorFourRoomsObstructedMazeParkingRoundaboutFetchPushFetchPickAndPlace
GPT-5.5Codex-84.6995.22248.9604.1-3.473601.8989.6-37.110.9860.9180.6640.904-30.249.818-18.16-11.81
Claude Opus 4.7Claude Code-88.5998.77-15.84602.3-3.979452.0990.1-38.520.9830.9300.7010.911-38.189.331-19.69-13.34
MiniMax-M3Claude Code-136.491.88-80.87233.4-5.103606.2983.4-39.250.0000.0000.4030.000-32.719.521-25.88-13.22
DeepSeek-V4-ProClaude Code-91.1694.48-97.4825.20-6.506-0.468898.0-54.180.0000.0000.0790.000-53.3810.20-27.66-20.56
Random policyUniform-499.2-33.36-101.0-29.63-43.77-291.8-34.60-147.50.0000.0000.0420.000-47.457.188-46.03-43.75
02

Agent lanes

Column-best counts summarize the matrix; they are not a cross-task aggregate score.

Codex

GPT-5.5

9 / 16column bests
Claude Code

Claude Opus 4.7

5 / 16column bests
Claude Code

MiniMax-M3

1 / 16column bests
Claude Code

DeepSeek-V4-Pro

1 / 16column bests

Behavioral evidence

Inspect the selected Policies.

The archive preserves 64 original-environment reruns—one for every model–environment lane.

Open the rerun archiveRead the paper