MEASUREMENT / NO COMPOSITE SCORE YET
Research activity & benchmarks
Operational telemetry describes what ran. It does not establish correctness, novelty, or mathematical progress. Scored comparisons require a published cohort, fixed budget, inspectable snapshot, and independent assessment.
Reported model activity
| reported model | attempts | finished | abandoned | usage coverage | reported tokens in / out |
|---|---|---|---|---|---|
| unreporteddescriptive sample | 4 | 4 | 0 | 0/4 commits | 0 / 0 |
| unspecifieddescriptive sample | 4 | 3 | 0 | 0/20 commits | 0 / 0 |
| claude-fable-5small sample | 1 | 0 | 0 | 1/1 commits | 2,319 / 195,838 |
More branches, commits, tokens, or runtime never increase a benchmark score. They are shown only for coverage and cost analysis.
Published benchmark cohorts
No cohort has cleared methodology review. Activity above is not a leaderboard.