Paper companion · v0.1.0
Core16 experiment record
Held-out returns and final-Policy reruns from four coding-agent/model lanes across sixteen interactive tasks.
Historical research record
This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.
Held-out score matrix
Scores are comparable only within one Environment column. Highlighted values are column maxima.
| Model | Acrobot | ContinuousCar | Bipedal | CarRacing | Reacher | HalfCheetah | Ant | Pusher | DoorKey | KeyCorridor | FourRooms | ObstructedMaze | Parking | Roundabout | FetchPush | FetchPickAndPlace |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.5Codex | -84.69 | 95.22 | 248.9 | 604.1 | -3.473 | 601.8 | 989.6 | -37.11 | 0.986 | 0.918 | 0.664 | 0.904 | -30.24 | 9.818 | -18.16 | -11.81 |
| Claude Opus 4.7Claude Code | -88.59 | 98.77 | -15.84 | 602.3 | -3.979 | 452.0 | 990.1 | -38.52 | 0.983 | 0.930 | 0.701 | 0.911 | -38.18 | 9.331 | -19.69 | -13.34 |
| MiniMax-M3Claude Code | -136.4 | 91.88 | -80.87 | 233.4 | -5.103 | 606.2 | 983.4 | -39.25 | 0.000 | 0.000 | 0.403 | 0.000 | -32.71 | 9.521 | -25.88 | -13.22 |
| DeepSeek-V4-ProClaude Code | -91.16 | 94.48 | -97.48 | 25.20 | -6.506 | -0.468 | 898.0 | -54.18 | 0.000 | 0.000 | 0.079 | 0.000 | -53.38 | 10.20 | -27.66 | -20.56 |
| Random policyUniform | -499.2 | -33.36 | -101.0 | -29.63 | -43.77 | -291.8 | -34.60 | -147.5 | 0.000 | 0.000 | 0.042 | 0.000 | -47.45 | 7.188 | -46.03 | -43.75 |
Agent lanes
Column-best counts summarize the matrix; they are not a cross-task aggregate score.
Claude Opus 4.7
5 / 16column bestsMiniMax-M3
1 / 16column bestsDeepSeek-V4-Pro
1 / 16column bestsBehavioral evidence
Inspect the selected Policies.
The archive preserves 64 original-environment reruns—one for every model–environment lane.
Open the rerun archive →Read the paper ↗