direct mujoco rendererGPT-5.5
Codex- Held-out 分数
- -37.11
- Checkpoint
- #005
- 重跑
- 100 steps
Validation-selected checkpoint rerun in the original environment; case return -38.36.
MuJoCo · Core16
通过接触操作把物体推向目标位置。
这些证据属于论文时期实验,并不表示相关 package 当前仍然可用。
每张卡片都在原研究 Environment 中重跑 validation 选出的 checkpoint。
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -38.36.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -37.53.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -41.99.
direct mujoco rendererValidation-selected checkpoint rerun in the original environment; case return -55.45.
在该 Environment 内分数越高越好;不同任务的原始 reward scale 不可比较。
-37.11-38.52-39.25-54.18-147.5