Auditable evolution record
Run records
A Program-Evolution Run is a bounded sequence of immutable Programs evaluated on Agent-selected indices from one fixed training Episode pool.
Run model
Search, selection, and final measurement remain distinct stages.
Reusing an index across Submissions preserves the hidden Episode specification for matched comparison. Every evaluation still creates a fresh Environment and Policy runtime and consumes budget.
After the Agent Session closes, Host-only Validation selects a submitted Program and held-out Assessment measures only that selection.
Current publication state
No public v0.3 Run record yet.
The Run-record surface is implemented, but active CartPole and Balatro evolution records have not yet been published on this site.
Open historical Core16 evidence →