Skip to main content

MuJoCo · Core16

Ant

Multi-legged locomotion with eight continuous actions.

Historical research record

This evidence belongs to the paper-era experiment and is not a guarantee of current package availability.

01

Final-Policy evidence

Each card reruns the validation-selected checkpoint in the original research Environment.

GPT-5.5 Policy rerun in Antdirect mujoco renderer

GPT-5.5

Codex
#2
Held-out
989.6
Checkpoint
#006
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return 221.1.

Claude Opus 4.7 Policy rerun in Antdirect mujoco renderer

Claude Opus 4.7

Claude Code
BEST
Held-out
990.1
Checkpoint
#002
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return 204.6.

MiniMax-M3 Policy rerun in Antdirect mujoco renderer

MiniMax-M3

Claude Code
#3
Held-out
983.4
Checkpoint
#007
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return 215.7.

DeepSeek-V4-Pro Policy rerun in Antdirect mujoco renderer

DeepSeek-V4-Pro

Claude Code
#4
Held-out
898.0
Checkpoint
#019
Rerun
220 steps

Validation-selected checkpoint rerun in the original environment; case return 192.5.

02

Reported scores

Higher is better within this Environment. Raw reward scales are not comparable across tasks.

01Claude Opus 4.7Claude Code990.1
02GPT-5.5Codex989.6
03MiniMax-M3Claude Code983.4
04DeepSeek-V4-ProClaude Code898.0
05Random policyUniform-34.60