Skip to content

Benchmarks

A benchmark reports declared conditions within each task. It does not create a universal competence score.

The catalogue accepts planned entries, but any later evidence state requires a complete record with provenance, a machine-readable summary, and DONE. Registry membership, benchmark publication, qualification, and leaderboard inclusion remain separate choices.

The list above comes from versioned benchmark pages, not the accepted-run catalogue. A planned protocol can appear without a run record when its evidence state and record count remain visible.

This page is intentionally small. Each benchmark detail page links its protocol and records, while measured values are read from the record summary during the site build. The Research records page explains the contribution and acceptance process.