State of the AI Model Landscape

Introducing Artificial Analysis and Its Role in AI Benchmarking

George, co-founder of Artificial Analysis, introduces the company as an independent AI benchmarking platform used by millions to compare AI technologies. He notes the company's growing influence, being referenced by major labs like NVIDIA, Anthropic, and Google during recent model releases, and outlines that they benchmark across the full AI stack including agents, models, inference providers, and hardware.

AI Progress Has Not Slowed Down

George presents a chart of leading AI lab releases over recent years, arguing that claims of AI progress slowing down are exaggerated. He points to accelerating release frequency and rising intelligence, citing Opus 4.6 and improved agent long-horizon performance as evidence that progress continues unabated.

The Artificial Analysis Intelligence Index and Global Model Landscape

George introduces the Artificial Analysis Intelligence Index, a synthesis of 10 benchmarks used to compare language model intelligence, noting Claude Opus 4.8 recently overtook GPT 5.5 as the leading model. He then discusses the geographic landscape of AI labs, highlighting the dominance of the US and China, with notable contributions from France, South Korea, and the UAE, while pointing out Australia's absence from frontier model development.

Open Weights vs Proprietary Models

George compares the intelligence trajectory of open-weight models against proprietary models since mid-2023, showing that open-weight models have consistently trailed by roughly three to nine months but have kept pace over time. He explains the incentives sustaining open-weight development, such as flexibility and fine-tuning, and notes that current open models like Kimi K2.6 and DeepSeek V4 Pro offer intelligence comparable to older proprietary models like Opus 4.5 or GPT 5.2.

Validating Benchmarks Against Real-World Agentic Performance

George explains how Artificial Analysis ensures benchmarks align with real-world usability, showcasing their GDPval AA agentic benchmark built from OpenAI's dataset. He demonstrates year-over-year progress in knowledge work outputs comparing Claude 4 Sonnet to GPT 5.5, and shows a creative task example (a music video mood board) illustrating clear capability differences between smaller and larger models like Gemma and GPT 5.5.

The Paradox of Cheaper Intelligence and Rising Spending

George addresses the paradox that while AI intelligence has become cheaper than ever, companies are spending more due to increased usage of premium plans like Claude Code and Codex. He begins outlining six key drivers behind this dynamic, starting with smaller models achieving greater intelligence, increased model sparsity, and software efficiency gains in inference stacks like VLLM and SGLang.

Six Drivers of Cost and Efficiency in AI Inference

Continuing the cost analysis, George details hardware efficiency gains (e.g., NVL72 nodes lowering per-query costs) alongside three factors driving costs upward: growing demand for larger, more capable frontier models, increased reasoning tokens in reasoning models, and the multi-turn nature of agentic workflows that can involve 20 to 100 turns per task, amplifying token costs.

Falling Costs of Intelligence Over Time

George presents data showing that the cost of achieving a given level of intelligence falls 10 to 100 times within six to eighteen months, using bucketed intelligence index data and cost-to-run benchmark comparisons. He emphasizes that selecting slightly older or cheaper models can yield massive cost savings, illustrated by the Pareto curve of intelligence versus benchmarking cost, where frontier models like Opus 4.8 cost over $4,000 to fully benchmark.

Hardware Improvements Underpinning Efficiency Gains

George briefly highlights how hardware advances, such as the NVIDIA B200 node, improve both output speed and system throughput, enabling better amortization of hardware costs across more users and further driving down inference costs. This sets up the transition into discussing agent categories and their real-world applications.

Seven Fast-Growing Agent Categories

George outlines seven agent categories that Artificial Analysis tracks as having reached product-market fit and rapid growth: coding agents, general work agents, chatbots, presentation agents, OCR, data analysis, and customer support. He notes these categories represent mature, real-world use cases for AI agents beyond simple chatbot interactions.

Benchmarking Coding Agents: Models and Harnesses

George explains that coding agent performance depends on both the underlying model and its harness (e.g., Claude Code, Cursor CLI, OpenCode), presenting benchmark results showing Opus 4.8 and GPT 5.5 leading in intelligence but costing over $4 per task. He highlights significant cost-performance trade-offs, noting cheaper options like Cursor CLI with Composer 2.5 under 50 cents per task, suggesting cost-sensitive users consider alternatives depending on the task type.

Key Trends Shaping AI Through 2026

George closes by outlining major trends expected to continue through 2026: general work agents becoming mainstream tools for knowledge workers (e.g., Claude Cowork, Codex Desktop), the blurring line between coding agents and general agents as code becomes the universal task-completion paradigm, and the growing maturity of continual learning in agents. He ends with a call to action, noting Artificial Analysis is hiring in Melbourne, across Australia, and in San Francisco.

Alrighty. Are we, there we go. Great. Thanks for having me here. So I'm George, one of the co founders of Artificial Analysis. Very quickly, about us. We're an independent AI benchmarking company. We help millions of users with our website, artificialanalysis.ai, understand what's happening in AI and choose between all the different technologies with our independent benchmarks. And we're also quite widely referred to in the industry at the start of this week.

Jensen referred to us for the Nemotron three Ultra release last week, Anthropic with Opus 4.8, and our GDP Val AA benchmark the week prior to that, Sundar with a g p t sorry, Gemini 3.5 Flash release. And so very happy to be here and share a bit about what we do and give an overview of how we see the state of AI in June 2026.

We benchmark across the AI stacks. We benchmark agents, models, inference providers, and hardware, and we also benchmark kinda text focused language models, but also image, video, speech, and music models as well. I think to start off with, or provided essentially this chart, it's a bit hectic, which shows the leading releases from the leading AI labs over the last few years.

And it's a bit hectic, but I think that's probably no surprise to us in this room, because it's been a hectic few years. There's been a lot of releases and a lot of progress, and I think why I like to start with this chat is it shows that AI progress has not slowed down. I think claims of that have been greatly greatly exaggerated.

I heard a lot of that in late twenty five, but then Opus 4.6 came out. Agents started working for much longer horizons, and people quieted down a bit. And I think this chart shows that really well, is that there's more dots on this chart than ever. If you look at the last three months here, which shows there's been more leading releases, and it's going up and to the right pretty quickly, which is showing progress in intelligence.

I'll first introduce our artificial analysis intelligence index metric. Now this is a synthesis metric of 10 benchmarks that we run to test language model intelligence, which provides a high level synthesis overview and relative comparison of the intelligence of these models. We have the current set of leading models on this chart here, and I think to note, Claude Opus 4.8 last week took the mantle from GPT 5.5 as the leading language model in terms of intelligence. But I think rather than kind of stopping there if if we were able to stop there, it'd make our jobs a lot easier, but I think there's a the the story is there's there's a reason still to use other models, considering the cost, the speed, and other trade offs at play amongst language models today.

I'll first take a look at different perspectives on thinking around the current state of models. You can see here that it's very much a US and China story when looking at where the labs making these leading language models are based. Also present on this chart is France with Mistral, South Korea with a number of labs and a very successful sovereign AI initiative, and then also The United Arab Emirates.

I think to us in this room, notably missing is Australia. Australia doesn't have kind of language models that are competitive with the frontier in terms of intelligence. And from my perspective, I I don't think that's going to happen soon either. Another perspective is looking at open weights or colloquially open source models compared to that of proprietary models and proprietary intelligence achieved.

On the y axis, we have our intelligence index, and then on the x axis, release date, and what these lines plot is the leading model that is kind of proprietary compared to OpenWeights. And OpenWeights has always trailed proprietary intelligence. This chart is really going back to kind of mid mid twenty three with the Lama two seventy b release, and you can see that, yes, OpenWaits has trailed open sorry, proprietary intelligence, but it's in a sense kind of kept up roughly three to nine months behind that of proprietary intelligence.

This was not a given. It was an open question. Years ago, or only a few years ago, people asking around the commercial models that could support open weights, but I think what we've seen has kept up for a number of reasons, and I think our perspective is this is going to continue with there being sufficient incentive there, particularly for labs wanting to serve those looking for open weights models for their flexibility, the ability to fine tune, and other reasons. And then secondly, I think another perspective on this is if open weights is kind of three to nine months behind, we're looking at OPUS maybe 4.5 territory, GBT 5.2 for the intelligence that you can essentially get with open weights models, with a Kimi k two point Kimi k 2.6 or the the latest kind of DeepSeg v four Pro. And so it remains an option.

You could do a lot with with with OPUS 4.5 or GBD 5.2. And so it remains an option, and it looks like that's going to continue based on the last few years and our house perspective. We benchmark using quantitative benchmarks, but we try and essentially ensure that those benchmarks are aligned from real world use. You don't have models hill climbing on a number that doesn't relate to increasing in, essentially, the usability of these models the use cases we're exploring with AI now, and so we like to sense check, essentially the progress. And I think what we tell here with our gdpval.a agentic benchmark that uses an OpenAI data set, we turned it into an agentic benchmark, is a perspective of, okay, a leading model a year ago with Claude four O Claude four Sonnet, sorry, and then now with GBD five point GBD 5.5, and you can see progress is being made in terms of real world knowledge work output.

On the left, we have a table useful. On the right, we have essentially additional synthesis, executive summary provided, very useful totals, and further analysis than we had a year ago. So agents have progressed on economically valuable tasks, and we also like to look at kind of maybe less knowledge work y, economically valuable tasks as well. Progress has been made there as well.

And so this is a kind of music video mood board task within GDP val a a, and you can see that this is comparing to kind of some open weight models, but on the left, we have what Gemma e4b. I mean, it's amazing that it even did this, but I think as a music video mood board, maybe I'm not creative enough, but I don't think I could use the one on my right here.

But if we compare that with a GBD 5.5, maybe I could kind of do something with this. And so I think it shows that there are still real differences between the models, and the kind of benchmarks do correlate to real world capabilities. I think one, kind of two things that are hard to reconcile in AI today, is one, that we have cheaper intelligence.

So it's cheaper than ever to access GPT four level of intelligence. Kind of going back to what I was saying earlier, you can get kind of Kimi K 2.6 now cheaper than you can get Opus 4.5, you know, six months ago, and so I think this speaks to you can get intelligence cheaper than ever, but we're also spending more than ever within our companies.

We're kind of upgraded to the $200 a month Claude code or Codex plans. And kind of why is that is? And I kind of break it down into six drivers here that we see as the most critical. I'll kind of breeze through them. Each of these could be a talk, but, like, smaller models able to achieve more intelligence is the first.

Lowering cost, increased sparsity of the models, lower proportion of active parameters to total, kind of when inference compute is at scale, the kind of number of active parameters is really the driver of cost, how you amortize hardware. Next is kind of software efficiency gains to call out here. In the inference stack, in kind of VLLM or SGLANG.

Optimizations are happening every day, that of flash attention. Also, kind of closer to the model, exploring different quantizations. We're no longer at BF 16. Now we're at four bit precisions, whether that's INFOR with Moonshot models or whether that's NVFP4 with other models. We're making better efficiency trade offs essentially to offer cheaper intelligence and increase efficiency. Third is hardware efficiency.

Hardware is kind of costing more, but is also offering more. And so a NVL 72 node is able to offer a lower cost when serving at scale than your h 100 node, even though it costs more, and so that higher efficiency can translate to lower costs offered. On what's raising costs, larger models.

And how do you reconcile this with the first? It's that everybody wants or not every use case, but there's an insatiable demand for frontier intelligence. We see that continuing. I think for yourself, I'd love to be more intelligent. I'd love the people I work with to be more intelligent, and so I think that translates to models, and so there's an insatiable demand for kind of larger models increases cost.

Next, reasoning models. Labs are increasing intelligence through more reasoning tokens. Everybody knows, we pay per token. And so, higher cost. Lastly, agents. So it used to be you, you know, in ChatGPT, you would ask a question or send a request to the API, and you get the response back, maybe make that, maybe that's, put that into a word through code.

Now you're having agents doing exploration of other files, of checking its own work, and it's improving the output, but we're dealing with multiple turns for the same task. And so in our benchmarks, we're commonly seeing 60 turns for a GDP valet A task, and so kind of 20 to 100 turns is quite sensible for many knowledge work tasks, and it acts as a multiplier on the cost.

This is a chart showing kind of model, language model inference price falling. So here, we've bucketed models in their intelligence, so in our intelligence index, zero to ten, ten to 20, etcetera. And we've looked at, okay, when was that model, when was that intelligence achieved, which is the first dot of each line, and then where are we now in terms of the latest release, or the cheapest release that is able to access that intelligence?

And what we see here is really in kind of six to eighteen month periods, you have the cost of that intelligence falling, often 10 to 100 times. And this is something that we can all think about and take advantage of. If a task was able to you could do it with OPUS 4.5, odds are now that you can there's orders of magnitude at play. It's not two x cheaper, as you can go 10 x cheaper, in many cases, by choosing a cheaper model that has more recently been released.

And so I think both insatiable demand, but for, like, kind of increased intelligence, but especially for, like, real world, more defined use cases, often there's 10 to a 100 times cheaper options at play in terms of model selection. And this is that played out again. So here we have a chart of intelligence, and then the cost to us of running benchmarks.

And so here's the on the x axis, we have the cost to run our intelligence index, those 10 benchmarks, and I don't know if we would be able to afford it back in when we kind of started at kind of in 2324, but now it's costing above 4,000 US to run these 10 benchmarks for frontier models that are pushing what AI intelligence is able to achieve.

And so you can see OPUS 4.8 costs over $4,000, but there's a clear kind of Pareto curve that one should think about when selecting models for use cases, thinking about the cost sensitivity, and there's options at play amongst the Pareto curve that really does span a 100 x plus in terms of cost differences.

Again, could talk about it in a lot more detail, but this is a chart just showing what's underpinning a lot of the efficiency gains, and that's hardware improvements. So you can see that the kind of a B200 node is able to offer greater output speed per query, but also greater system throughput. It can scale to more users, greater concurrency.

And so this allows amortizing the cost of that hardware over more users, which is supporting lowering of intelligence costs and inference costs. I think agents are a big big category, it requires different ways to think about each of the agent categories.

We track seven fast growing agent categories that we see as having achieved product market fit, in a sense, and are fast scaling. So that's coding agents, general work agents, not just ChatGPT, but your co your local codex, your Claude co work that have really reached a level of maturity, a lot more to go, but a level of maturity for real world use.

Next, chatbots, presentation agents, OCR, data analysis, and customer support as seven agent categories that have really reached a level of maturity and product market fit, and they're continuing to grow very quickly. Double clicking into that, which is probably most interesting to our AI engineering audience here, and that is coding agents.

So coding agents, they're not the model, they're the model and the harness. Both are drivers of performance. Because when we're using coding agents, we're not just using OPUS 4.7 or OPUS 4.8, we're using OPUS 4.8 in Claude code, in Cursor CLI, in OpenCode, and so we benchmark both to give a view of how those interact and the intelligence of the model and the harness together. I think what we see here is that we have OPUS 4.8, I think, coming today.

We're finishing up benchmarks, but that's going to lead in terms of intelligence in our coding agent index. GPD 5.5 from OpenAI is close by, and what's interesting is that there's a bit of a drop off when looking to the other models the performance that they offer. And so I think probably not a huge surprise to a lot here. I think it shows the importance of kind of continuing to stay on top of things and upgrading to the latest models.

But again, with coding, like when using language model APIs directly or inferencing yourself, I think you can see that there are still big trade offs at play when we think about speed, when we think about cost. And so here is the kind of coding agent index score, and then on the x axis, we have cost per task.

You can see that definitely Claude Code with Opus 4.7 and Codex with GBD 5.5 offering leading intelligence but costing over $4 per task. Then, if you come way back here, you can see that models that are in line of sight there, but not quite achieving that level of intelligence, are under 50¢, notably Cursor CLI with Composer 2.5.

And so, again, orders of magnitude, an order of magnitude trade off at play here, and it might still make sense, probably for a lot in this room, for your regular coding activities, use the best model achievable and pay the price, but for other tasks, coding agents, I'll go into it a little bit later, but coding agents are becoming agents, agents are becoming coding agents.

For other tasks where you're using coding, think that worth considering other models. Lastly, I'll walk through some of the key trends that we're seeing across the AI stack that we're monitoring and we think will play out across the rest of 2026.

So starting with agents, I think the first trend is kinda general work agents starting to roll out to the masses. Kinda our experience six or nine months ago when using Claude Code for the first time and getting an immediate productivity uplift from agents that can do long horizon tasks, debug themselves, what's going wrong, fix problems, before coming back to us is coming to knowledge workers more broadly with Claude Cowork, Codex Desktop, and other work general work agents that essentially I mean, they're pretty thin somewhat like thin wrappers in a sense. You can think about them this way, over Cloak Code or over the model, but they're able to do real world tasks, and that's something that we see kind of rolling out to the masses, and the masses are going to have our experience that we had nine months ago.

Next, agents are coding agents, and also coding agents are agents. So we're seeing it mature, essentially, becoming more mature, this paradigm of agents just using code to complete work tasks, even if they're not coding tasks as we think about them. And it's because, essentially, the labs have focused on code.

It's a flexible paradigm that they can use to complete tasks, and it works. And so I think we're seeing coding agent performance become more important, not just for doing coding tasks, but for doing general work tasks as well. And so I think, practically for us, when we're looking at models and thinking about how it's gonna complete tasks, it's not just about, okay, kind of design, how well is it integrated, like, essentially first party support of harnesses to our libraries.

Think about, okay, how is this model going to do tasks? Often it's through running code. How good is it at running code? Third, continual learning working natively is something that is maturing a lot as well and a focus of the labs. There's a few more there.

Might post this on our Twitter artificial analysis. But I guess we're out of time. Thanks very much. We're hiring in Melbourne and everywhere across Australia as well as SF. Please reach out. Thank you.

State of the AI Landscape

George Cameron
Co-Founder, Artificial Analysis

An Artificial Analysis logo is displayed in the top right corner.

Artificial Analysis

Artificial Analysis is an independent AI benchmarking company trusted by millions of users and leading players across AI

Entities that have publicly referenced Artificial Analysis

  • AI/Tech Companies: OpenAI, Google, Meta, xAI, Amazon, Microsoft, IBM, ByteDance, Mistral, Anthropic, Stability AI, ServiceNow, Alibaba, Salesforce, VERITAS, LG AI, Perplexity, KllingAI, Together AI, PixVerse, ElevenLabs, Cartesia, Databricks, Zapier
  • Hardware/Chip Companies: NVIDIA, Groq, AMD, Cerebras, Sambanova, Qualcomm
  • AI Infrastructure/Platforms: Baseten, Clarifai, Fireworks AI, Nebula, Replicate
  • Media/Publications: Wall Street Journal, Bloomberg, Reuters, The Economist, CNBC, The Washington Post, Business Insider, TechCrunch, Financial Times, VentureBeat, WIRED, All In
  • Financial/Consulting/Institutions: J.P. Morgan, Citi, McKinsey & Company, BCG, Deutsche Bank, ECO, MIT, Stanford University, CSAIL, HAI, OECD
A presentation slide displaying text and a large grid of over 60 company and institutional logos.

Artificial Analysis benchmarks intelligence, performance & cost across the AI stack

Agents

  • Coding Agent Index
  • 6+ other agent categories
  • Performance and cost

Language model intelligence

  • Intelligence Index
  • GDPval-AA
  • Performance and cost

Image, video, speech & music

  • Arenas
  • Speech benchmarks
  • Performance and cost

Hardware & inference

  • AA-AgentPerf
  • AA-SLT (System Load Test)
  • NVIDIA, AMD, Cerebras, Groq
The slide displays four panels, each representing a different area of AI benchmarking. The 'Agents' panel shows colored blocks, likely representing different agent types or performance levels. The 'Language model intelligence' panel features a line graph depicting the progression of language model intelligence over time. The 'Image, video, speech & music' panel shows a screenshot of a user interface with various vehicles and other objects, suggesting image or video analysis. The 'Hardware & inference' panel includes a line graph illustrating throughput versus latency for system load tests.

The language model intelligence race is as fast and competitive as ever

Frontier Language Model Intelligence, Over Time

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, r², Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IFBench, Humanity's Last Exam, GPQA Diamond, CritPT

  • OpenAI
  • xAI
  • Meta
  • Google
  • Anthropic
  • Mistral
  • DeepSeek
  • Upstage
  • MiniMax
  • Kimi
  • Xiaomi
  • Alibaba
  • MBZUAI Institute of Foundation Models
  • Z AI

A line chart titled "Artificial Analysis Intelligence Index" tracking the intelligence of various language models released by different AI labs over time, from November 2022 to May 2026.

Our Intelligence Index synthesises leading evaluations; Claude Opus 4.8 has recently taken the lead from GPT-5.5

Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, r²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, IfBench, Humanity's Last Exam, GPQA Diamond, CritPI

  • Claude Opus 4.8 (max): 61.4
  • GPT-5.5 (xhigh): 60.2
  • Claude Opus 4.7 (max): 57.3
  • Gemini 1.0 Pro Preview: 57.2
  • GPT-5.4 (xhigh): 56.8
  • Owen3.7 Max: 56.6

Reasoning models are indicated by a lightbulb icon.

A horizontal bar chart titled 'Artificial Analysis Intelligence Index' displays the performance scores of 25 different AI models, ranked from highest to lowest. Claude Opus 4.8 leads with a score of 61.4, followed by GPT-5.5 at 60.2, and Claude Opus 4.7 at 57.3. The scores range down to 24.1 for K2 Think V2. Each bar includes a numerical score and an icon representing the model's developer. Lightbulb icons next to some models indicate they are reasoning models.

The US leads in frontier intelligence, with China based labs are very present; few other countries build competitive models

Leading Models by Country

Artificial Analysis Intelligence Index - Leading models
  • United States
  • China
  • France
  • South Korea
  • United Arab Emirates
  • Claude Opus (United States): 81.0
  • GPT-4 (0.5 Pro) (United States): 80.2
  • Claude Sonnet (United States): 72.3
  • GPT-4 (0.1 Pro) (United States): 67.2
  • Gemini 1.5 Pro (United States): 66.8
  • Claude 3.5 Sonnet (United States): 59.8
  • GPT-4 Omni (United States): 58.4
  • Qwen 2.5 (China): 56.8
  • Gemini 1.0 Pro (United States): 56.6
  • Claude Haiku (United States): 55.9
  • Qwen 2.1 (China): 54.0
  • MMNU V2.3 Pro (China): 53.8
  • Sense Core 4.0 Pro (China): 53.2
  • Mixtral (France): 52.2
  • Claude Sonnet V4 Pro (dev) (United States): 51.7
  • GLM 5.1 (China): 51.4
  • Minimax MT-7 (China): 49.6
  • GPT-4 (0.4 Pro) (United States): 48.0
  • DeepSeek V4-7B (dev) (China): 46.5
  • Gemini 1.0 Ultra (United States): 46.0
  • Mistral Medium 2.0 (France): 45.0
  • Gamma 4-70B (South Korea): 39.3
  • Claude 4.5 Haiku (United States): 38.2
  • MNNU V2.0 Pro (China): 36.9
  • Minimax MT-6 (dev) (China): 36.7
  • K2 Pro (South Korea): 33.3
  • A2 Trace V2 (United Arab Emirates): 34.1
  • Super Pro 3 (United Arab Emirates): 28.9
  • K2 Haiku (South Korea): 24.5
Reasoning Models are indicated by a lightbulb icon
A bar chart titled "Leading Models by Country" displays the Artificial Analysis Intelligence Index for various AI models. The bars are colored to represent the country of origin: dark blue for United States, red for China, orange for France, light blue for South Korea, and green for United Arab Emirates. US-based models generally show the highest index scores, followed by China-based models. A lightbulb icon next to some model names signifies them as "Reasoning Models."

Open-weights models trail the proprietary frontier by around 3-9 months, but remain a viable option

Progress in Open Weights vs. Proprietary Intelligence

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GQPer-AA, 7-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, FBench, Humanity's Last Exam, GPQA Diamond, CRISP.

A line graph titled 'Progress in Open Weights vs. Proprietary Intelligence'. The y-axis represents the Artificial Analysis Intelligence Index, and the x-axis represents the Release Date from November 2022 to May 2026. Two lines are plotted: 'Proprietary' (black squares) and 'Open Weights' (blue circles). The 'Proprietary' line generally maintains a higher intelligence index than the 'Open Weights' line. Both lines show an overall upward trend over time, indicating increasing intelligence, with proprietary models consistently leading open-weight models.

Artificial Analysis

Agents have dramatically improved in economically valuable tasks

Progression in 'Inventory Management' task in GDPval-AA

MAY 2025

Number of Incidents per Supplier

Supplier Number of Incidents Percentage
Apexian Defense 10 31.3%
Orion Systems 7 21.9%
Veracity 6 18.8%
AeroNexx Instruments 4 12.5%
Creston-Road 2 6.3%
Nuvolla Technologies 1 3.1%
Synergetic 1 3.1%
Monarch Defense 1 3.1%

Claude 4 Sonnet

MAY 2026

Inventory Incident Analysis - 2025

GPT-5.5 (xhigh)

The slide compares AI agent capabilities in inventory management over time. On the left, a screenshot of a simple table generated in May 2025 by Claude 4 Sonnet, showing incident counts and percentages per supplier. On the right, a screenshot of a more advanced dashboard generated in May 2026 by GPT-5.5 (xhigh), titled "Inventory Incident Analysis - 2025", displaying key metrics such as total incidents, estimated cost, average resolution time, and resolution rate, along with a donut chart illustrating incident categorization.

... and less-economically valuable tasks

Progression in 'Music Video Mood Board' task in GDPval-AA

  • GEMMA 4 E4B

    • Moodboard Concept: Ballad of Duality
    • Style: Theatrical, Dramatic, Vulnerable
  • GEMMA 4 31B

    • Music Video Moodboard: Baroque Masquerade & Surreal Gothic Fantasy
    • Colors: Ruby Crimson, Jet Black, Goldenrod, Candlelight Beige, Dusty Rose, Umber Brown
  • GPT-5.5

    • MUSIC VIDEO MOODBOARD
    • Colors: Gold, Pink, Red, Brown, Yellow, Peach, Black
The slide compares the output of three AI models (GEMMA 4 E4B, GEMMA 4 31B, and GPT-5.5) for a 'Music Video Mood Board' generation task. GEMMA 4 E4B displays a minimalist mood board with abstract geometric shapes: a black square, a gold square, a grey circle, and a red triangle. GEMMA 4 31B shows a mood board featuring a black and white image of a clown, a blue abstract image, a small room interior, and a color palette. GPT-5.5 presents a rich, multi-image mood board grid with ornate masks, clothing, chandeliers, people, and a more diverse color palette.

While efficiency gains have been made...

GPT-4 level intelligence is now 100x cheaper than original GPT-4

...compute demand continues to increase

New applications continue to demand more compute: a single deep research query can cost >10x an original GPT-4 query

  • LOWERS COST A: Smaller Models and Sparsity

    ~1/10x compute

    Algorithmic and training data improvements have allowed smaller models to get smarter

  • LOWERS COST B: Software Efficiency

    ~1/3x compute

    Inference optimizations (e.g. Flash Attention) improve efficiency

  • LOWERS COST C: Hardware Efficiency

    ~1/3x costs

    Next generation accelerators offer more compute efficiency

  • RAISES COST D: Larger Models

    ~5x compute/query

    Scaling laws continue to demand higher parameter counts for greater intelligence

  • RAISES COST E: Reasoning Models

    ~10x tokens/query

    Significant increase in output tokens when models 'think' before answering

  • RAISES COST F: AI Agents

    ~20x requests/use

    Agents chain multiple requests to LLMs to complete tasks autonomously across long conversations

Figures are highly indicative and serve to illustrate the directional impact of each factor impacting cost.

A line graph illustrates the changing cost of AI, showing an initial decrease labeled 'Cheaper Intelligence' followed by an increase labeled 'Higher spend than ever'. Below the graph, six distinct cards are presented: three highlighting factors that 'LOWERS COST' (Smaller Models and Sparsity, Software Efficiency, Hardware Efficiency) and three highlighting factors that 'RAISES COST' (Larger Models, Reasoning Models, AI Agents).

Token costs are falling for all levels of intelligence...

Language Model Inference Price

Price: USD per 1M tokens

  • Intelligence Index < 10
  • 10 ≤ Intelligence Index < 20
  • 20 ≤ Intelligence Index < 30
  • 30 ≤ Intelligence Index < 40
  • 40 ≤ Intelligence Index < 50
  • 50 ≤ Intelligence Index < 60
  • Intelligence Index ≥ 60

Artificial Analysis

A line chart titled "Language Model Inference Price" plots the price in USD per million tokens on a logarithmic y-axis against release date on the x-axis. Multiple colored lines represent different intelligence index ranges, showing a general trend of decreasing costs over time for various intelligence levels. The x-axis ranges from November 2022 to May 2026, and the y-axis ranges from approximately 0.00625 to 64 USD per million tokens.

... but accessing frontier intelligence is costing more than ever

Intelligence vs. Cost to Run Artificial Analysis Intelligence Index

Artificial Analysis Intelligence Index · Cost to run Intelligence Index

Most attractive quadrant

Legend: OpenAI (black), Google (green), Anthropic (red), Mistral (orange), DeepSeek (dark blue), xAI (light blue), Amazon (purple), MiniMax (maroon), NVIDIA (light green), Kimi (teal), Xiaomi (dark red), Alibaba (brown).

  • gpt-oss-20B (high)
  • gpt-oss-120B (high)
  • Gemini 3.1 Preview
  • MiMo-V2.5-Pro
  • DeepSeek V4 Pro (Max)
  • MiniMax-M2.7
  • DeepSeek V4 Flash (Max)
  • Kimi K2.6
  • Grok 4.3 (high)
  • Qwen3.5 397B A17B
  • NVIDIA Nemotron 3 Super
  • Nova 2.0 Pro Preview (medium)
  • Mistral Medium 3.5
  • Claude 4.5 Haiku
  • Gemini 3.5 Flash
  • Qwen3.7 Max
  • GPT-5.4 (xhigh)
  • GPT-5.5 (xhigh)
  • GPT-5.4 mini (xhigh)
  • Claude Opus 4.8 (max)
  • Claude Opus 4.7 (max)
  • Claude Sonnet 4.6 (max)
  • Gemini 3.5 (Max)
A scatter plot showing "Artificial Analysis Intelligence Index" on the Y-axis and "Cost to Run Intelligence Index (USD, Log Scale)" on the X-axis. A green shaded area in the upper-left quadrant is labeled "Most attractive quadrant," indicating models with high intelligence and low cost. Various AI models are plotted as points, color-coded by their respective developers (OpenAI, Google, Anthropic, Mistral, DeepSeek, xAI, Amazon, MiniMax, NVIDIA, Kimi, Xiaomi, Alibaba).

Hardware is a key driver of inference performance improvement

Output Speed per Query vs. Concurrency

(gpt-neo 120B (fp8) - Output speed per query (tokens per second))

A line graph displays output speed per query (tokens per second) on the Y-axis against concurrency on a log scale on the X-axis, evaluating different hardware and software configurations for gpt-neo 120B (fp8).

The X-axis, "Concurrency (log scale)", ranges from 1 to 2048. The Y-axis, "Output Speed per Query", ranges from 0 to 400 tokens per second.

Four distinct lines represent the following configurations:

  • A blue line for 8xB200 - TensorRT-LLM - Optimal.
  • A purple line for 8xH100 - vLLM.
  • A red/orange line for 8xM300X - vLLM.
  • A green line for 8xH200 - vLLM.

All configurations show a decrease in output speed per query as concurrency increases. The 8xB200 with TensorRT-LLM - Optimal configuration consistently achieves the highest output speed per query across all concurrency levels, demonstrating superior performance compared to the other configurations.

Frontier labs are moving up the stack into agents: We track 7 fast-growing agent categories

Coding Agents
Plan, write and ship production code

  • Claude Code
  • Codex
  • Cursor
  • Copilot
  • Devin

General Work
Multi-step knowledge work across apps

  • ChatGPT Agent
  • Claude Cowork
  • Copilot
  • Gemini Enterprise
  • Manus

Chatbots
Everyday Q&A; and conversation

  • ChatGPT
  • Claude
  • Gemini
  • Perplexity
  • Grok

Presentations
Turn a prompt into a polished deck

  • Gamma
  • Canva
  • Copilot
  • Gemini
  • Beautiful.ai

OCR
Documents and scans into structured data

  • Mistral OCR
  • Document AI
  • Textract
  • Azure DI
  • ABBYY

Data Analysis
From raw data to charts and insight

  • Power BI
  • Tableau
  • Databricks
  • Hex
  • Julius

Customer Support
Resolve tickets across chat, email, voice

  • Sierra
  • Decagon
  • Fin
  • Agentforce
  • Zendesk

Our Coding Agent Index also shows the tradeoff between Intelligence and Cost

Artificial Analysis Coding Agent Index vs. Cost per Task

Artificial Analysis Coding Agent Index vs. mean pay-per-token API cost per task (USD)

Legend: Anthropic, OpenAI, Cursor, Z.ai, Moonshot AI, DeepSeek, Google

The chart plots the Artificial Analysis Coding Agent Index on the Y-axis against the Cost per Task (USD) on the X-axis. A green shaded area highlights the region of higher index and lower cost.

  • Low Cost (~$0.25 - $0.50):
    • Claude Code - Composer 2 (Cursor)
    • Claude Code - DeepSeek V4 Pro (high) (DeepSeek)
    • Claude Code - Kimmi K2.6 (Moonshot AI)
  • Medium Cost (~$0.50 - $2.75):
    • Cursor CLI - Composer 2.5 Fast (Cursor)
    • Cursor CLI - GPT-5.5 (medium) (Cursor)
    • Claude Code - Opus 4.7 (medium) (Anthropic)
    • Codex - GPT-5.5 (medium) (OpenAI)
    • Claude Code - GLM-5.1 (FriendAI)
    • Claude Code - Opus 4.6 (medium) (Anthropic)
    • Gemini CLI - Gemini 3.3 Pro (high) (Gemini)
  • High Cost (~$4.00 - $5.00):
    • Claude Code - Opus 4.7 (max) (Anthropic)
    • Codex - GPT-5.5 (xhigh) (OpenAI)

A scatter plot titled "Artificial Analysis Coding Agent Index vs. Cost per Task" displays various AI coding agents. The Y-axis represents the Artificial Analysis Coding Agent Index, and the X-axis represents the Cost per Task in USD. Data points are colored according to their respective companies, as indicated in the legend. A green shaded area in the lower-left quadrant highlights models with a high index score and low cost, indicating a favorable performance-to-cost ratio.

In 2026 we anticipate continued progress across the stack

  • Agents
    • Generalist work agents roll out to the masses
    • Agents are coding agents
    • Continual learning works natively
    • Computer use becomes usable
  • Models
    • Interactivity improves and models become proactive, not reactive
    • Cost efficiency becomes a top priority
    • Models become natively multimodal including speech
    • World models get worked out
  • Inference & hardware
    • Caching innovates and becomes critical for the cost & speed story
    • Hardware goes rack-scale (NVL-72), Jensen no longer holds chips at conferences
    • Speed-focused systems are offered at a premium (Groq+NVIDIA, Cerebras, etc.)

People

  • Jensen Huang
  • Sundar Pichai

Technologies & Tools

  • B200
  • BF16
  • ChatGPT
  • Claude 4 Sonnet
  • Claude Code
  • Claude Cowork
  • Codex
  • Codex Desktop
  • Composer 2.5
  • Cursor CLI
  • DeepSeek V4 Pro
  • Flash Attention
  • Gemini 3.5 Flash
  • Gemma 4B
  • GPT-5.5
  • H100
  • Kimi K2.6
  • Llama 2 70B
  • Nemotron 3 Ultra
  • NVFP4
  • NVL72
  • OpenCode
  • Opus 4.5
  • Opus 4.6
  • Opus 4.7
  • Opus 4.8
  • SGLang
  • VLLM

Concepts & Methods

  • Artificial Analysis Intelligence Index
  • Coding Agent Index
  • Continual learning
  • GDP Val AA
  • Open weights models
  • Pareto curve
  • Sovereign AI

Organisations & Products

  • Anthropic
  • Mistral
  • Moonshot AI
  • NVIDIA