Code models. All time.

1,057 catalogued models + 82 score-only names · data as of 2026-09-07

Highest recorded score on the selected benchmark

87 / 1,139 entries scored here · partial coverage, not a global ranking

Highest recorded score

79.2%

Run 2025-12-05 · Sonar Foundation Agent · +0.4 points over Doubao-Seed-Code (78.8%)

Score source

87 scores · 87 with a run date · 0 without

Releases accelerated

13 in 2009 → 49 in 2016 → 119 in 2023. 2023 is the recorded peak; 2026 is partial, 64 to 2026-09-07.

Section 3 · this benchmark

How the recorded frontier moved

SWE-bench Verifieda 500-instance subset of SWE-bench human-confirmed as solvable, OpenAI/Princeton collaboration. Introduced 2024-08-13.

Benchmark source
0%20%40%60%80%100%2023202420252026run date → data as of 2026-09-07Claude 2 · 4.4% · run 2023-10-10 · RAG baseline (retrieval only, no agent)SWE-Llama 7B · 1.4% · run 2023-10-10 · RAG baseline (retrieval only, no agent)SWE-Llama 13B · 1.2% · run 2023-10-10 · RAG baseline (retrieval only, no agent)GPT-3.5 · 0.4% · run 2023-10-10 · RAG baseline (retrieval only, no agent)GPT-4 (1106) · 22.4% · run 2024-04-02Claude 3 Opus · 15.8% · run 2024-04-02Lingma SWE-GPT 72b (v0918) · 25% · run 2024-09-18Lingma SWE-GPT 7b (v0918) · 10.2% · run 2024-09-18Lingma SWE-GPT 72b (v0925) · 28.8% · run 2024-10-02Lingma SWE-GPT 7b (v0925) · 18.2% · run 2024-10-02Claude 3.5 Haiku · 40.6% · run 2024-10-22GPT-4o · 38.8% · run 2024-10-28CodeAct v2.1 (claude-3-5-sonnet-20241022) · 53% · run 2024-10-29Gemini 2.0 Flash (v20241212-experimental) · 52.2% · run 2024-12-12swe-search · 62.2% · run 2024-12-21o1-preview · 64.6% · run 2025-01-17Gemini 2.0 Flash (Experimental) · 44.2% · run 2025-01-18Claude 3.5-Sonnet-20241022 · 51.6% · run 2025-01-22Claude 3.5 Sonnet · 63.4% · run 2025-02-06O3 Mini · 42.4% · run 2025-02-14Claude 3.7 Sonnet w/ Review Heavy · 62.4% · run 2025-02-25Llama3-SWE-RL-70B · 41.2% · run 2025-02-26Qwen2.5 (7B + 72B) · 32.8% · run 2025-03-06o4-mini · 64.6% · run 2025-05-03SWE-agent-LM-32B · 40.2% · run 2025-05-11Claude 3.7 Sonnet · 66.4% · run 2025-05-14Undisclosed · 70.6% · run 2025-05-19DevStral Small 2505 · 46.8% · run 2025-05-20Claude 4 Opus · 73.2% · run 2025-05-22Amazon.nova Premier v1:0 · 42.4% · run 2025-05-27Co-PatcheR · 46% · run 2025-05-28TTS(Bo8) · 47% · run 2025-06-16Qwen2.5 Coder 32B Instruct · 38% · run 2025-06-16MCTS Refine 7B · 23.2% · run 2025-06-27TTS(Bo16) · 58.8% · run 2025-06-29DeepSWE-Preview · 42.2% · run 2025-06-29GPT 4.1 mini · 23.94% · run 2025-07-20GPT 4o · 21.62% · run 2025-07-20Llama 4 Maverick Instruct · 21.04% · run 2025-07-20Llama 4 Scout Instruct · 9.06% · run 2025-07-20DevStral Small 2507 · 38% · run 2025-07-25o3 · 58.4% · run 2025-07-26Gemini 2.5 Pro · 53.6% · run 2025-07-26GPT 4.1 · 39.58% · run 2025-07-26Gemini 2.5 Flash · 28.73% · run 2025-07-26Gemini 2.0 Flash · 13.52% · run 2025-07-26GLM-4.5 · 64.2% · run 2025-07-28Claude Sonnet 4 · 74.8% · run 2025-07-31Qwen3-Coder 480B/A35B Instruct · 55.4% · run 2025-08-02Qwen2.5-Coder 32B Instruct · 9% · run 2025-08-03Claude 4 Sonnet · 76.8% · run 2025-08-04Kimi K2 Instruct · 53.4% · run 2025-08-04Qwen3-Coder-480B-A35B-Instruct · 69.6% · run 2025-08-05DeepSeek V3 0324 · 42% · run 2025-08-06GPT 5 · 65% · run 2025-08-07GPT 5 mini · 59.8% · run 2025-08-07GPT 5 nano · 34.8% · run 2025-08-07gpt-oss-120b · 26% · run 2025-08-07GLM 4.5 · 54.2% · run 2025-08-22Qwen3-Coder-30B-A3B-Instruct · 60.4% · run 2025-09-01Doubao-Seed-Code · 78.8% · run 2025-09-28GLM-4.6 · 68.2% · run 2025-09-30Kimi K2 · 71.2% · run 2025-10-14GPT-5 · 74.4% · run 2025-10-15Claude 4.5 Sonnet · 74.8% · run 2025-11-03Frogboss 32B 2510 · 53.6% · run 2025-11-10Frogmini 14B 2510 · 45% · run 2025-11-10Gemini 3 Pro Preview · 77.4% · run 2025-11-20GPT 5.1 · 66% · run 2025-11-20GPT 5.1 Codex · 66% · run 2025-11-24MiniMax M2 · 61% · run 2025-11-24DeepSeek V3.2 Reasoner · 60% · run 2025-12-01GLM 4.6 · 55.4% · run 2025-12-01Claude 4.5 Opus · 79.2% · run 2025-12-05Devstral Small (2512) · 56.4% · run 2025-12-09Devstral (2512) · 53.8% · run 2025-12-09Kimi K2 Thinking · 63.4% · run 2025-12-10Gemini 3 Flash · 75.8% · run 2026-02-17MiniMax M2.5 · 75.8% · run 2026-02-17Claude 4.6 Opus · 75.6% · run 2026-02-17GLM 5 · 72.8% · run 2026-02-17GPT 5.2 · 72.8% · run 2026-02-17Kimi K2.5 · 70.8% · run 2026-02-17DeepSeek V3.2 · 70% · run 2026-02-17Claude 4.5 Haiku · 66.6% · run 2026-02-17GPT 5.2 Codex · 72.8% · run 2026-02-19Gemini 3 Pro · 69.6% · run 2026-02-26Claude 4.5 Opus 79.2%
  • best recorded submission so far — agent submissions, pooled across scaffolds; not a controlled model-only comparison
  • a new record
  • another dated score
  • RAG baseline (retrieval only, no agent) — a different setup, not on the line
  • x: leaderboard run date · y: score (%)
Every frontier step · 14
Frontier steps on SWE-bench Verified, oldest first
Run dateModelScoreRiseScaffoldLabSource
2024-04-0222.4%firstSWE-agentlab not recordedScore
2024-09-1825%+2.6Lingma Agentlab not recordedScore
2024-10-0228.8%+3.8Lingma Agentlab not recordedScore
2024-10-2240.6%+11.8Toolslab not recordedScore
2024-10-2953%+12.4OpenHandslab not recordedScore
2024-12-2162.2%+9.2CodeStory Midwit Agentlab not recordedScore
2025-01-1764.6%+2.4W&B Programmer O1 crosscheck5OpenAILab sourceScore
2025-05-1466.4%+1.8Aime-coder v1AnthropicLab sourceScore
2025-05-1970.6%+4.2TRAElab not recordedScore
2025-05-2273.2%+2.6Toolslab not recordedScore
2025-07-3174.8%+1.6Harness AIAnthropicLab sourceScore
2025-08-0476.8%+2EPAM AI/Run Developer Agentlab not recordedScore
2025-09-2878.8%+2TRAElab not recordedScore
2025-12-0579.2%+0.4Sonar Foundation Agentlab not recordedScore
All scores with a run date · 87
Every SWE-bench Verified score with a run date, oldest first
Run dateModelScoreSetupRecordSource
2023-10-104.4%RAG baseline (retrieval only, no agent)Score
2023-10-101.4%RAG baseline (retrieval only, no agent)Score
2023-10-101.2%RAG baseline (retrieval only, no agent)Score
2023-10-100.4%RAG baseline (retrieval only, no agent)Score
2024-04-0222.4%agent submissions · SWE-agentnew recordScore
2024-04-0215.8%agent submissions · SWE-agentScore
2024-09-1825%agent submissions · Lingma Agentnew recordScore
2024-09-1810.2%agent submissions · Lingma AgentScore
2024-10-0228.8%agent submissions · Lingma Agentnew recordScore
2024-10-0218.2%agent submissions · Lingma AgentScore
2024-10-2240.6%agent submissions · Toolsnew recordScore
2024-10-2838.8%agent submissions · Agentless-1.5Score
2024-10-2953%agent submissions · OpenHandsnew recordScore
2024-12-1252.2%agent submissions · Google JulesScore
2024-12-2162.2%agent submissions · CodeStory Midwit Agentnew recordScore
2025-01-1764.6%agent submissions · W&B Programmer O1 crosscheck5new recordScore
2025-01-1844.2%agent submissions · CodeShellAgentScore
2025-01-2251.6%agent submissions · AutoCodeRover-v2.1Score
2025-02-0663.4%agent submissions · AgentScopeScore
2025-02-1442.4%agent submissions · Agentless LiteScore
2025-02-2562.4%agent submissions · SWE-agentScore
2025-02-2641.2%agent submissions · Agentless MiniScore
2025-03-0632.8%agent submissions · SWE-FixerScore
2025-05-0364.6%agent submissions · PatchPilot-v1.1Score
2025-05-1140.2%agent submissions · SWE-agentScore
2025-05-1466.4%agent submissions · Aime-coder v1new recordScore
2025-05-1970.6%agent submissions · TRAEnew recordScore
2025-05-2046.8%agent submissions · OpenHandsScore
2025-05-2273.2%agent submissions · Toolsnew recordScore
2025-05-2742.4%agent submissions · Amazon Nova Premier 1.0Score
2025-05-2846%agent submissions · PatchPilotScore
2025-06-1647%agent submissions · Skywork-SWE-32BScore
2025-06-1638%agent submissions · Skywork-SWE-32BScore
2025-06-2723.2%agent submissions · MCTS-Refine-7BScore
2025-06-2958.8%agent submissions · DeepSWE-PreviewScore
2025-06-2942.2%agent submissions · R2E-GymScore
2025-07-2023.94%agent submissions · mini-SWE-agentScore
2025-07-2021.62%agent submissions · mini-SWE-agentScore
2025-07-2021.04%agent submissions · mini-SWE-agentScore
2025-07-209.06%agent submissions · mini-SWE-agentScore
2025-07-2538%agent submissions · SWE-agentScore
2025-07-2658.4%agent submissions · mini-SWE-agentScore
2025-07-2653.6%agent submissions · mini-SWE-agentScore
2025-07-2639.58%agent submissions · mini-SWE-agentScore
2025-07-2628.73%agent submissions · mini-SWE-agentScore
2025-07-2613.52%agent submissions · mini-SWE-agentScore
2025-07-2864.2%agent submissions · UndisclosedScore
2025-07-3174.8%agent submissions · Harness AInew recordScore
2025-08-0255.4%agent submissions · mini-SWE-agentScore
2025-08-039%agent submissions · mini-SWE-agentScore
2025-08-0476.8%agent submissions · EPAM AI/Run Developer Agentnew recordScore
2025-08-0453.4%agent submissions · CodeSweep - SWE-agentScore
2025-08-0569.6%agent submissions · OpenHandsScore
2025-08-0642%agent submissions · SWE-ExpScore
2025-08-0765%agent submissions · mini-SWE-agentScore
2025-08-0759.8%agent submissions · mini-SWE-agentScore
2025-08-0734.8%agent submissions · mini-SWE-agentScore
2025-08-0726%agent submissions · mini-SWE-agentScore
2025-08-2254.2%agent submissions · mini-SWE-agentScore
2025-09-0160.4%agent submissions · EntroPO + R2EScore
2025-09-2878.8%agent submissions · TRAEnew recordScore
2025-09-3068.2%agent submissions · UndisclosedScore
2025-10-1471.2%agent submissions · Lingxi v1.5Score
2025-10-1574.4%agent submissions · Prometheus-v1.2.1Score
2025-11-0374.8%agent submissions · Sonar Foundation AgentScore
2025-11-1053.6%agent submissions · FrogBoss-32B-2510Score
2025-11-1045%agent submissions · FrogMini-14B-2510Score
2025-11-2077.4%agent submissions · live-SWE-agentScore
2025-11-2066%agent submissions · mini-SWE-agentScore
2025-11-2466%agent submissions · mini-SWE-agentScore
2025-11-2461%agent submissions · mini-SWE-agentScore
2025-12-0160%agent submissions · mini-SWE-agentScore
2025-12-0155.4%agent submissions · mini-SWE-agentScore
2025-12-0579.2%agent submissions · Sonar Foundation Agentnew recordScore
2025-12-0956.4%agent submissions · mini-SWE-agentScore
2025-12-0953.8%agent submissions · mini-SWE-agentScore
2025-12-1063.4%agent submissions · mini-SWE-agentScore
2026-02-1775.8%agent submissions · mini-SWE-agentScore
2026-02-1775.8%agent submissions · mini-SWE-agentScore
2026-02-1775.6%agent submissions · mini-SWE-agentScore
2026-02-1772.8%agent submissions · mini-SWE-agentScore
2026-02-1772.8%agent submissions · mini-SWE-agentScore
2026-02-1770.8%agent submissions · mini-SWE-agentScore
2026-02-1770%agent submissions · mini-SWE-agentScore
2026-02-1766.6%agent submissions · mini-SWE-agentScore
2026-02-1972.8%agent submissions · mini-SWE-agentScore
2026-02-2669.6%agent submissions · mini-SWE-agentScore
No run date on record · 0none — every score here carries a leaderboard run date; 37 of them have no model date in the source export

Nothing to list: all 87 scores are plotted above.

Section 4 · every model

Explore every year

1,057 catalogue models across 77 calendar years, 1950–2026, and 82 score-only names. Scores are SWE-bench Verified.

2026 64 models
2025 106 models
2024 98 models
2023 119 models
2022 91 models
2021 79 models
2020 50 models
2019 68 models
2018 37 models
2017 55 models
2016 49 models
2015 29 models
2014 33 models
2013 19 models
2012 13 models
2011 11 models
2010 10 models
2009 13 models
2008 6 models
2007 6 models
2006 9 models
19502005 · 92 models Open early history
82 score-only names · on a leaderboard, not in the Epoch catalogue Open

Section 5 · the whole dataset

What this view covers

1,057 entries from Epoch AI’s notable-models catalogue — a maintained sample, not every model ever trained — and 27 of them have a score. 82 more names exist only as leaderboard scores. 63 of the 214 scores are undated in the source export. Closed-model benchmark results (GPQA, MATH and the rest for GPT, Claude, Gemini), LMArena Elo and ARC-AGI are not covered by this snapshot.

Data as of 2026-09-07 · fetched by HTTP GET only

Read all five source gaps and the method

Source gaps, verbatim

  1. LMArena (Chatbot Arena) Elo — only export found was a pickle file; unpickling an untrusted file can run arbitrary code, so it was not loaded. Left out.
  2. ARC-AGI leaderboard — no downloadable JSON/CSV endpoint found without a browser. Left out.
  3. HumanEval / GPQA / MATH as standalone numbers for closed frontier models (GPT, Claude, Gemini, etc.) — each lives in a separate model card or paper with no bulk feed; not pulled in this pass. GPQA/MATH/MMLU-PRO/BBH/MUSR/IFEval are present here for the 11 open-weight models whose base weights matched Epoch AI's notable-models list exactly, sourced from the Hugging Face Open LLM Leaderboard v2.
  4. SWE-bench Lite and SWE-bench Multilingual 'introduced' dates are the earliest submission date in the published leaderboard data, not a separate primary announcement — documented as a data-floor proxy in raw/SOURCES.md.
  5. 1057 models is Epoch AI's notable-models list as of the 2026-09-07 CSV pull, not literally every model ever trained — it is the closest thing to a maintained, sourced, comprehensive spine that exists.

How this page counts

  • 63 of 214 scores carry no date in the source export’s date field, and that field is a model date, not the day a result was recorded. The frontier chart places a score only by its leaderboard run date; 66 scores have none and are listed beside the chart, never plotted.
  • 82 of the 109 scored names are not in the Epoch catalogue: they are searchable score-only entries with no lab, release date or size on record. Only 27 of the 1,057 catalogue models carry any score. A benchmark’s own counts are stated with its result; these are counts for the whole dataset.

The 11 benchmarks

Every benchmark, what it measures and its source
BenchmarkWhat it measuresIntroducedScoresSource
SWE-benchresolving real GitHub issues end-to-end in a live repo, as a coding agent2023-10-1013Source
SWE-bench Litea 300-instance curated subset of SWE-bench for cheaper agent evaluation2023-10-1025Source
SWE-bench Verifieda 500-instance subset of SWE-bench human-confirmed as solvable, OpenAI/Princeton collaboration2024-08-1387Source
SWE-bench Multimodalresolving GitHub issues in visual/frontend software domains (JS/TS UI repos), needs image understanding2024-10-0410Source
SWE-bench Multilingualresolving GitHub issues across non-Python languages; date is the earliest submission in the published leaderboard data, not a separate announcement post2026-02-1313Source
IFEvalfollowing verifiable natural-language instructions (formatting, length, keywords)2023-11-1411Source
BBHBIG-Bench Hard: 23 challenging multi-step reasoning tasks2022-10-1711Source
MATH Lvl 5hardest tier (level 5) of the MATH competition-mathematics benchmark2021-03-0511Source
GPQAgraduate-level Google-proof multiple-choice science questions2023-11-2011Source
MUSRmultistep soft reasoning over long narrative text (murder mysteries, object placement, team allocation)2023-10-2411Source
MMLU-PROharder, 10-option-per-question revision of MMLU across 57 subjects2024-06-0311Source

vyke score (proposed)

Unpublished, not a benchmark.

For each (model, benchmark) score, the model’s percentile rank among every model scored on that same benchmark within ±12 months of its own release date. A model’s vyke score is the weighted average of those percentiles, weighted by how many models sit in each cohort. 64 of the 109 scored names receive one (27 of them catalogue models). It is shown only in a model’s profile and takes no part in the highest recorded score or the frontier.