Skip to main content
Skip to main content

Artificial Analysis Coding Agent Index: September 2026

19 min read
by Divanshu Chauhan
AI Claude Fable 5 Claude Opus GPT 5.6 Grok 4.5 Muse Spark Kimi K2.7 SWE-1.7 MiniMax M3 GLM 5.2 Gemini DeepSeek V4 AI Models AI Comparison Coding 2026 Best AI Model
Artificial Analysis Coding Agent Index v1.5 leaderboard, September 2026

TL;DR

The September 28, 2026 snapshot ranks Claude Code — Opus 5.5 (max) first with 66 and Codex — GPT-6 Sol (max) sixth with 57. The index compares agent configurations across three benchmarks, not standalone models.

Key Takeaways

  • The September 28 snapshot ranks Claude Code — Opus 5.5 (max) first with 66 and Codex — GPT-6 Sol (max) sixth with 57.
  • Version 1.5 equally weights DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA.
  • Each component score averages pass@1 across three attempts per task.
  • The chart displayed 14 of 20 models; similar composite scores can hide different strengths across the component benchmarks.

Snapshot checked: September 28, 2026. The live Artificial Analysis Coding Agent Index displayed 14 of 20 models. The v1.5 index equally weights DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA (methodology and version history).

Short answer: Claude Code — Opus 5.5 (max) leads the displayed chart at 66; three configurations tie at 62. This scores agent configurations, not standalone model quality.

Model / versionRank / scoreSnapshot dateSource
Claude Code — Opus 5.5 (max)1 / 662026-09-28Index
Claude Code — Fable 5.1 (max) (with fallback)T-2 / 622026-09-28Index
Devin Fusion CLI — Claude Fable 5.1 XHigh + SWE-2 MediumT-2 / 622026-09-28Index
Codex — GPT-6 Astra (max)T-2 / 622026-09-28Index
Devin Fusion CLI — GPT-6 Astra XHigh + SWE-2 Medium5 / 592026-09-28Index
Codex — GPT-6 Sol (max)6 / 572026-09-28Index
Grok Build — Grok 4.7 (xhigh)7 / 562026-09-28Index
Muse Code — Muse Spark 1.3 (max)T-8 / 542026-09-28Index
Opencode — GLM-5.3T-8 / 542026-09-28Index
Kimi Code CLI — Kimi K310 / 522026-09-28Index
Claude Code — Qwen3.8 MaxT-11 / 432026-09-28Index
Codex — DeepSeek V4 Pro 0813 (max)T-11 / 432026-09-28Index
Antigravity SDK — Gemini 3.8 Flash (high)13 / 422026-09-28Index
Codex — GPT-6 Luna (max)14 / 412026-09-28Index

Ranks use standard competition ranking; equal scores share the same rank. These are the 14 configurations shown in the chart, not a universal ranking of standalone models.

General-purpose model comparison: July 2026

Historical July 2026 context: The comparisons below preserve July-era benchmark results, model availability, and pricing. They were not rechecked for this September index update; treat the figures and recommendations as historical, not current.

July 2026 summary: Fable 5 if you need peak coding/planning and can get access without torching the budget. Opus 4.8 for normal professional coding. GPT-5.6 Sol for coding-agent work. Grok 4.5 when volume and unit economics matter. On the open side: K2.7-Code, SWE-1.7, GLM 5.2, or MiniMax M3 depending on whether you want raw weights, a Cognition post-train, value, or cheap tokens. Muse Spark 1.1 if you want Meta’s agent API. Gemini for long multimodal research.

Anyone selling one model for every job is selling a subscription.

What changed from April to July 2026

April’s list: GPT 5.5, Opus 4.7, Kimi K2.6, MiniMax M2.7, Muse Spark with no API. Most of those names are previous-gen now.

What actually moved the needle:

  • Claude Fable 5. Anthropic’s tier above Opus. Tops SWE-bench Pro and AA-Briefcase in several reports. Expensive. Mid-June export controls briefly knocked Fable offline for customers.
  • Grok 4.5 (July 8). Trained with Cursor. The independent AA numbers that matter are cost and speed on messy agent work, not another “we’re #1” tweet.
  • GPT-5.6 (public around July 8–9 after a government-access hold). Sol, Terra, and Luna are different lanes. Sol is the coding/Codex one. Joint #1 with Fable on Arena Code Frontend, cheaper than Fable.
  • Claude Opus 4.8 (late May). What a lot of pros still run every day when Fable is too expensive or flaky.
  • Muse Spark 1.1 (July 9). Real Meta Model API. About $1.25 / $4.25 per million tokens. 1M context, sub-agents, computer use.
  • Kimi K2.7-Code. Open coding model. NVIDIA shipped NVFP4 builds.
  • Cognition SWE-1.7. K2.7-Code base plus Cognition RL. Coding near Opus-class, alignment cleaned up vs raw Kimi.
  • GLM 5.2. Keeps showing up next to Sol and Grok. Cheap. Not a laptop model.
  • MiniMax M3. Replaces M2.7. People leave it on as a token-cheap coding agent. mlx already supports it.

Still useful, no longer the headline: GPT 5.5, Opus 4.7, Kimi K2.6, GLM-5.1, MiniMax M2.7, Grok 4.20 beta.

Quick comparison: July 2026 model snapshot

ModelMakerBest forWeaknessWhat the numbers sayVerdict
Claude Fable 5AnthropicPeak coding, planning, hard agentsPrice; access/policy riskSWE Pro ~80.4% (max); AA-Briefcase ~1390; DeepSWE/Terminal top-tierCeiling when available
Claude Opus 4.8AnthropicDaily professional coding, writingPrice and token burn on long agents~69.2% SWE-bench Pro (max); AA-Briefcase Elo 1354 (max)Coding daily driver
GPT-5.6 SolOpenAICoding agents, Codex, hard reasoningOverthinks; ultra modes can be slowLeads AA Coding Agent Index (max); Arena Code Frontend joint #1 with Fable; ~$5/$30 MTokAgent coding leader
Grok 4.5SpaceXAI / xAICheap fast agents, automationPresentation polish weaker; can tool-loopAA-Briefcase Elo 1328; AutomationBench-AA 51% at ~$0.34/task; SWE Pro ~64.7%Best ROI agent
Cognition SWE-1.7CognitionSpecialized coding agentsProduct access; not a free open baseK2.7-Code + Cognition RL; coding near Opus / GPT-5.5 class; cleaner alignment than raw KimiSpecialist coding stack
Muse Spark 1.1MetaCheap agent API, computer use, long contextNot the SWE Pro kingAgentic + coding focus; AA-Omniscience jump mostly from fewer hallucinationsPrice disruptor
Kimi K2.7-CodeMoonshotOpen coding agentsHardware; base alignment concernsOpen 1T-class coding model; NVFP4; base for SWE-1.7Open coding pick
GLM 5.2Z.AIOpen coding valueSelf-host is heavy (hundreds of GB at 4-bit)Competitive vs Sol/Grok; SWE Pro ~62.1%; AA-Briefcase ~$1.71/task (max)Open value pick
MiniMax M3MiniMaxToken-efficient open coding agentsTrails Fable/Opus/Sol on peak SWEStrong coding agent in community use; low token burn; mlx supportToken-sipper open pick
Gemini 3.1 Pro (and Flash siblings)GoogleResearch, multimodal, long contextWriting and pure SWE lag coding specialistsStill the default for long docs / audio / video for a lot of research workResearch leader
DeepSeek V4 ProDeepSeekCheap near-frontier coding via APIHosting frictionStill one of the best $/coding deals from springValue API

These mix Artificial Analysis runs, Arena, vendor charts (especially Grok 4.5’s launch deck), and lab notes. If you are shipping on a score, re-check the live board that day.

Meet the contenders: July 2026 edition

Closed frontier

Claude Fable 5 is what people mean by “Anthropic still has the top.” On the Grok 4.5 launch charts and AA writeups it leads hard coding and agent knowledge work: SWE-bench Pro around 80.4% (max), AA-Briefcase around 1390, DeepSWE 1.0 ~66.1%, DeepSWE 1.1 ~70%, Terminal-Bench 2.1 ~84.3% (vendor/AA-cited; re-verify). Arena ties it with GPT-5.6 Sol at #1 on Code Frontend, at a higher price. In practice a lot of people use Fable for specs and review, then shove implementation to Grok, Sol, or an open agent. The limiting factor is not smarts. It is the invoice and whether your account can call the model. Mid-June 2026, a US export-control order made Anthropic suspend Fable 5 (and Mythos 5) for a stretch. Everyone fell back to Opus 4.8. I would not make Fable the team default until access is boring again.

Claude Opus 4.8 is what I open for a refactor that has to land clean. SWE-bench Pro around 69.2% (max) is the always-on Anthropic number. Writing is still good. Long agent loops still hurt on price, just less “call finance” than Fable max.

GPT-5.6 Sol is OpenAI’s coding lane in the Sol / Terra / Luna split. Sol (max) leads the AA Coding Agent Index and often finishes tasks cheaper than Opus max or Fable max. Arena has it joint #1 on Code Frontend with Fable at about $5/$30 per million tokens, roughly half Fable’s sticker in that report. Early notes from people I trust: strong on hard agent work, ultra modes can crawl, and it will rewrite code that was fine.

Grok 4.5 (July 8) is why I reopened the spreadsheet. SpaceXAI trained it with Cursor. AA-Briefcase: Elo 1328 (best non-Anthropic), about $1.12 per task and 12.4 minutes, vs Opus 4.8 max around $8.26 and 23.9 minutes. AutomationBench-AA: #1 at 51% of workflow objectives without rule breaks, at about a quarter of Claude’s cost. Vendor coding chart (treat as marketing until confirmed): DeepSWE 1.0 ~62%, Terminal-Bench 2.1 ~83.3%, SWE-bench Pro ~64.7%. On AA-Omniscience, accuracy went up and hallucinations went up with it. Cheap and fast is not the same as careful.

Muse Spark 1.1 is Meta finally selling a frontier model with a real developer price list. Zuck’s own pitch: agents, tools, computer use, 1M context, parallel sub-agents. Community pricing around $1.25 input / $0.15 cached / $4.25 output per million tokens. AA-Omniscience score jumped mostly because it refuses more (hallucinations down hard), not because it knows more. That is a product trait. Sometimes “I don’t know” is the right answer.

Gemini still owns long multimodal research for me. Flash models show up when people care about agent cost. If you live in PDFs, video, and giant contexts, ignore the SWE Pro flame war.

Open weights and specialized coding stacks

Kimi K2.7-Code is the open coding headline now. Moonshot’s 1T-class coding model is what people plug into agent runtimes. NVIDIA posted NVFP4 builds for Blackwell. Capability is real. Alignment on the raw base has also been dinged hard on Cognition’s public evals. Treat raw K2.7-Code as “smart first, policy second.”

Cognition SWE-1.7 is not a chat bot with a new name. It is Cognition’s coding product on Kimi K2.7-Code with their RL on top. Community notes put coding near Opus 4.8 / GPT-5.5. On cost/performance coding rollouts it sits near GPT-5.5 while Opus still leads pure score. Cognition also published the alignment story: raw Kimi bases scored very high on their misalignment evals (~89–100% on some checkpoints); SWE-1.7 lands around ~20%, near Opus/GPT. Same family of weights, different product after a US lab finishes it. Use SWE-1.7 when you want that coding level with Cognition’s harness. Do not download base K2.7 and pretend you shipped SWE-1.7.

GLM 5.2 is the other open name in the July pile-ons. Cheap API, solid coding demos, SWE-bench Pro around 62.1% on the Grok launch chart, AA-Briefcase around $1.71 per task at max. Self-hosting is a rack problem. 4-bit footprints get quoted in the hundreds of GB.

MiniMax M3 is the open coding agent people actually leave running. Successor to M2.7. mlx-lm and agent UIs already list it next to DeepSeek V4, GLM 5.2, and K2.7-Code. Nobody serious claims it beats Fable on SWE Pro. The pitch is high-volume coding with a small token bill. If your agents loop all day, M3 can be the workhorse while Fable/Sol/Opus do review and hard cuts.

DeepSeek V4 Pro still makes sense if you already live on DeepSeek pricing. It is no longer the only open story, which is the real change from April.

The July fight: Fable 5, Opus 4.8, Sol, and Grok 4.5

April was two models arguing about “smartest.” July is four answers to different questions.

Fable 5 when money and access are not the bottleneck. Specs, architecture, adversarial review. Pair it with a cheaper implementer if you want to stay solvent.

Opus 4.8 for careful software work you run every day. Among always-available flagships on hard GitHub-style SWE Pro work, this is still the Anthropic model I trust when a bad edit burns an afternoon.

GPT-5.6 Sol when the job looks like Codex or multi-tool agents. Coding Agent Index and Arena Code Frontend are the relevant scoreboards.

Grok 4.5 when you care about finishing agent work fast and cheap. AA-Briefcase and AutomationBench are the point. It is not secretly better at everything. It is better at cost per messy completed task.

How I route work:

  • Peak coding / planning, access ok, budget ok → Fable 5
  • Daily multi-file coding that has to be careful → Opus 4.8
  • Coding agents / Codex-style loops → GPT-5.6 Sol
  • High volume agent / automation where $8 per task is a joke → Grok 4.5
  • K2.7 lineage with Cognition’s post-train → SWE-1.7
  • Token-cheap open coding agents → MiniMax M3
  • Cheap Meta agent API with computer use → Muse Spark 1.1

Open and specialized coding stacks: July 2026 context

Open models are no longer “almost there on one chart.” They show up in the same side-by-sides as Sol and Grok. Post-trains also change the product enough that the brand name alone lies.

StackRoleWhen I would try it
Kimi K2.7-CodeOpen base coding modelYou want weights and will own hosting or a host
Cognition SWE-1.7K2.7 + Cognition RLYou want coding near Opus-class with Cognition’s product
GLM 5.2Open value codingPrice first; API first, cluster second
MiniMax M3Token-efficient open agentHigh-volume loops where tokens dominate the bill
DeepSeek V4 ProCheap near-frontier APIExisting DeepSeek stack, backlog cleanup

Quote the checkpoint. Kimi and SWE-1.7 are related and not interchangeable. MiniMax M3 is not M2.7 with a new version string.

Meta Muse Spark: July 2026 context

April Muse Spark was a product story. July Muse Spark 1.1 is a price list.

Meta is not claiming SWE Pro. They are selling agents, computer use, long context, and sub-agents at roughly a quarter of premium Anthropic/OpenAI sticker shock. If you need agents more than perfect repo surgery, that is enough reason to test. Read the Meta Model API docs for limits. WhatsApp distribution is not a quality guarantee.

Coding and software engineering: July 2026 comparisons

Peak coding when available: Fable 5.
Daily careful coding: Opus 4.8.
Coding-agent indexes: GPT-5.6 Sol.
Open base: Kimi K2.7-Code.
Specialized K2.7 product: SWE-1.7.
Token-efficient open agents: MiniMax M3.

Approximate SWE-bench Pro snapshot around the Grok 4.5 launch (vendor + independent citations; re-verify):

ModelSWE-bench Pro (approx.)Notes
Claude Fable 5 (max)~80.4%Peak when available
Claude Opus 4.8 (max)~69.2%Best always-on Anthropic default
Grok 4.5~64.7%Strong for cost; not the Pro leader
Claude Opus 4.7 (max)~64.3%April champion, now previous gen
GLM 5.2~62.1%Open, competitive
GPT 5.5 (xhigh)~58.6%Prior OpenAI flagship on this chart
Cognition SWE-1.7near Opus / GPT-5.5 class (lab/community)K2.7 base + Cognition RL; check live harness scores
MiniMax M3competitive open agent (community)Strength is token efficiency more than top SWE Pro

DeepSWE and Terminal-Bench reorder this list. Grok and Sol look better on some agent setups than raw SWE Pro suggests. Scaffold choice — which AI coding tools you wrap around the model — still moves scores by double digits.

Reasoning and hard problems

For the hardest agent reasoning I start with Fable 5 or GPT-5.6 Sol. For careful, lower-chaos answers I stay on Opus 4.8.

I trust task-level evidence more than a frozen intelligence index from April: Fable on peak engineering and briefcase-style work, Sol on coding-agent suites, Grok on automation ROI, Opus on disciplined multi-step engineering, Gemini on weird multimodal research.

Pure science Q&A? Re-check GPQA live. I am not reprinting April’s GPQA table as if it still ranks today’s models.

Agents and tool use

Cost-efficient agents: Grok 4.5.
Premium agent quality: Fable 5, then Opus 4.8 max if budget allows.
Cheap Meta agents: Muse Spark 1.1.
Open/specialized coding agents: SWE-1.7 or MiniMax M3, depending on product access vs token bill.

AA-Briefcase (private agentic knowledge-work set) around Grok 4.5 launch:

ModelAA-Briefcase Elo (approx.)Cost / taskTime / task
Claude Fable 5~1390highefficient turns relative to peers
Claude Sonnet 5 (max)~1390highslow (many turns)
Claude Opus 4.8 (max)~1354~$8.26~23.9 min
Grok 4.5~1328~$1.12~12.4 min
GLM 5.2 (max)lower~$1.71more turns

AutomationBench-AA: Grok 4.5 at 51% headline score and about $0.34 per task. Fable and Opus were close on score and several times more expensive.

Capability gaps got smaller. Price-per-finished-task gaps got rude.

Writing and communication

Still Claude. Opus 4.8 for daily prose; Fable if you are already paying for it.

I have not seen Grok or Sol unseat Claude for writing I would ship without a heavy edit pass. Grok is better when the output is a spreadsheet, mock, or automation. Muse Spark is not a writing product in my head. SWE-1.7 and MiniMax M3 are coding tools.

Multimodal: vision, audio, video

Gemini.

The July coding race does not replace Gemini for long audio/video plus huge document sets. Muse Spark and the GPT family got better at vision and computer use. Claude is better at high-res image-heavy coding than it was. For “watch this hour of video and reconcile it with a 400-page PDF,” I still start with Gemini.

Benchmark scorecard: July 2026

SignalApprox. leaderRunner-upOpen / specialist contender
SWE-bench ProClaude Fable 5 (~80.4%)Opus 4.8 (~69.2%)Grok 4.5 (~64.7%) / GLM 5.2 (~62.1%)
AA Coding Agent IndexGPT-5.6 Sol (max)Grok 4.5 (ties on some axes)K2.7-Code / SWE-1.7 / GLM 5.2
Arena Code FrontendFable 5 = GPT-5.6 Sol (joint #1)—open agents (task-dependent)
AA-BriefcaseClaude Fable 5 / Sonnet 5Opus 4.8, then Grok 4.5GLM 5.2 (cost-competitive)
AutomationBench-AAGrok 4.5 (51%)Fable 5 / Opus 4.8—
Cost per hard agent taskGrok 4.5Muse Spark 1.1 / MiniMax M3 / GLMDeepSeek / Kimi hosted
Multimodal long contextGeminiMuse Spark 1.1 / GPT familyopen VLMs (task-dependent)

Every leader here will move. Private benches matter more for agent product work than public quiz scores, and they can still overfit harness style.

Caveat: Scores come from Artificial Analysis posts, Arena, vendor launch charts (especially Grok 4.5), Cognition notes on SWE-1.7, and community replications. Agent scaffolds swing coding numbers by double digits. See “The Leaderboard Illusion”. Data as of July 10, 2026.

What the benches still hide: July 2026 context

Access is a feature. Fable 5’s ceiling is worthless if your org cannot call it. GPT-5.6 just left a restricted rollout. Free Grok Build access will not last. SWE-1.7 is a product, not Cognition’s full stack dumped on Hugging Face.

Cost is a trajectory. Token sticker prices lie. Turns, tool calls, and retries decide the invoice. Grok and MiniMax M3 win on efficiency as much as on raw scores.

Open is not free. Weights you cannot host are a PDF with a download button. Budget GPUs, pay a host, or run models locally on a Mac.

Base model is not the post-train. K2.7-Code and SWE-1.7 share a lineage and diverge hard on behavior. Quote the checkpoint.

Hallucinations scale with confidence. Grok 4.5 got more accurate and more willing to invent on AA-Omniscience. Muse Spark 1.1 climbed partly by abstaining. For legal, medical, or finance work, prefer models that fail loudly.

Scaffold beats model more often than launch day admits.

As of July 2026: I leaned on Claude for coding (Opus most days; Fable when the task was difficult and access was available). High-volume agent work was shifting toward Grok 4.5, while Sol was used in Codex-style trials. K2.7-Code, SWE-1.7, GLM 5.2, MiniMax M3, and DeepSeek were also in my cost-and-control comparison.

Which AI model should you actually use? (July 2026)

Hardest coding and planning, access and budget ok: Claude Fable 5. Keep a fallback. Do not wire the whole company to one flaky ceiling model.

Production code every day: Claude Opus 4.8. Lowest-regret Anthropic default.

Coding agents / Codex loops: GPT-5.6 Sol. Benchmark latency and edit quality on your own repos before you standardize.

Lots of AI agent tasks, unit economics matter: Grok 4.5. The AA cost and automation numbers are the reason.

Specialized coding on the Kimi lineage: Cognition SWE-1.7. Expect coding near Opus-class with post-train alignment, not a general chat model.

Open weights for coding agents: Kimi K2.7-Code. GLM 5.2 for value. MiniMax M3 when token burn is the enemy. Hosted endpoints unless you enjoy rack work.

Cheap Meta agent API: Muse Spark 1.1. Check SWE depth yourself.

Cost is the main constraint: DeepSeek V4 Pro, MiniMax M3, or hosted GLM/Kimi. Paying Fable rates for backlog cleanup is how side projects die.

Huge multimodal research: Gemini.

One default because decision fatigue is real: Opus 4.8 if you optimize for quality, Grok 4.5 if you optimize for volume. Add Sol for agent harnesses. Keep Fable for the hardest slices when access is stable.

Nothing wins every category in July 2026. Match the model to the job and re-check every few weeks. This post will age in public.

If you care more about the tools around these models (IDE agents, terminal agents, autocomplete) than the base checkpoints, see Claude Code vs Cursor vs Copilot (2026).

Frequently Asked Questions

Which coding agent leads the Artificial Analysis Coding Agent Index in September 2026?

The September 28, 2026 snapshot ranks Claude Code — Opus 5.5 (max) first with a score of 66. This is a score for the complete agent configuration, not just its model.

Where does Codex — GPT-6 Sol (max) rank?

Codex — GPT-6 Sol (max) ranks sixth with a score of 57 in the September 28, 2026 snapshot.

What does the Artificial Analysis Coding Agent Index measure?

Version 1.5 combines DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA with equal weight. Each component score averages pass@1 across three attempts per task.

Sources & References

Divanshu Chauhan

Divanshu Chauhan (@divkix)

Software Engineer Intern at Cloudflare. MS CS from Arizona State University. Currently open to full-time software engineering roles. Based in Tempe, Arizona, USA.

Expertise: AI, Claude Fable 5, Claude Opus, GPT 5.6. More about divkix