How we measure

Which test

The table ranks models on FrontierCode: a coding test that gives models real software-engineering tasks and scores the share they finish. We use it because it is the test still scoring new models — the older SWE-bench Verified leaderboard stopped taking new entries in February 2026, so it can no longer say what is best right now. The test

Who runs it

FrontierCode is Cognition's test. Epoch AI runs new frontier models on it shortly after they are released and publishes each score, which is why the newest models appear within weeks rather than months.

What a score means

A score is the share of the test's tasks the model finished — 53.5% means just over half. Higher is better. Scores within a point of each other are marked tied on the table: that small a gap says nothing about which model codes better.

What the badges mean

Every score carries a badge. Validated means two or more independent sources published a score for the model and agree within 3 points. Single-source means only one source has published a score, so nothing checks it yet. Conflict means sources disagree by more than 3 points — the number is disputed, and a disputed number never leads the table.

How fresh the numbers are

The numbers on this page are from 14 Sep 2026.

The full numbers

Download the full data (JSON) ↓