Extended task catalog and calibration utilities
TaskSpec and task_outcome(sim) are stable Core contracts: a task names its setup,
rollout defaults, optional primary metric, and anchors, and the result exposes that declared
key with its task-native value. A measured anchor also declares the scored_ticks interval
where its normalised value is valid. A different scoring window keeps the raw value and
withholds normalisation with an explicit status. This feature record tracks the extended
task implementations, calibration utilities, and control settings composed through that
interface.
Calibration and outcome definitions operationalise behaviour. They do not validate the named cognitive capacity, make raw metrics comparable across tasks, or turn one rollout into an independent scientific result.
cartpole_plank_easy is the declared frontier task. Canonical Falandays measured 10.0
raw and 0.0007 normalised, so the task is still at floor and remains unsolved by the
canonical node. It is an open challenge, not a scored core coordinate. It becomes useful
for design comparisons only after a design clears that floor.