State of the AI Model Landscape
Introducing Artificial Analysis and Its Role in AI Benchmarking
George, co-founder of Artificial Analysis, introduces the company as an independent AI benchmarking platform used by millions to compare AI technologies. He notes the company's growing influence, being referenced by major labs like NVIDIA, Anthropic, and Google during recent model releases, and outlines that they benchmark across the full AI stack including agents, models, inference providers, and hardware.
AI Progress Has Not Slowed Down
George presents a chart of leading AI lab releases over recent years, arguing that claims of AI progress slowing down are exaggerated. He points to accelerating release frequency and rising intelligence, citing Opus 4.6 and improved agent long-horizon performance as evidence that progress continues unabated.
The Artificial Analysis Intelligence Index and Global Model Landscape
George introduces the Artificial Analysis Intelligence Index, a synthesis of 10 benchmarks used to compare language model intelligence, noting Claude Opus 4.8 recently overtook GPT 5.5 as the leading model. He then discusses the geographic landscape of AI labs, highlighting the dominance of the US and China, with notable contributions from France, South Korea, and the UAE, while pointing out Australia's absence from frontier model development.
Open Weights vs Proprietary Models
George compares the intelligence trajectory of open-weight models against proprietary models since mid-2023, showing that open-weight models have consistently trailed by roughly three to nine months but have kept pace over time. He explains the incentives sustaining open-weight development, such as flexibility and fine-tuning, and notes that current open models like Kimi K2.6 and DeepSeek V4 Pro offer intelligence comparable to older proprietary models like Opus 4.5 or GPT 5.2.
Validating Benchmarks Against Real-World Agentic Performance
George explains how Artificial Analysis ensures benchmarks align with real-world usability, showcasing their GDPval AA agentic benchmark built from OpenAI's dataset. He demonstrates year-over-year progress in knowledge work outputs comparing Claude 4 Sonnet to GPT 5.5, and shows a creative task example (a music video mood board) illustrating clear capability differences between smaller and larger models like Gemma and GPT 5.5.
The Paradox of Cheaper Intelligence and Rising Spending
George addresses the paradox that while AI intelligence has become cheaper than ever, companies are spending more due to increased usage of premium plans like Claude Code and Codex. He begins outlining six key drivers behind this dynamic, starting with smaller models achieving greater intelligence, increased model sparsity, and software efficiency gains in inference stacks like VLLM and SGLang.
Six Drivers of Cost and Efficiency in AI Inference
Continuing the cost analysis, George details hardware efficiency gains (e.g., NVL72 nodes lowering per-query costs) alongside three factors driving costs upward: growing demand for larger, more capable frontier models, increased reasoning tokens in reasoning models, and the multi-turn nature of agentic workflows that can involve 20 to 100 turns per task, amplifying token costs.
Falling Costs of Intelligence Over Time
George presents data showing that the cost of achieving a given level of intelligence falls 10 to 100 times within six to eighteen months, using bucketed intelligence index data and cost-to-run benchmark comparisons. He emphasizes that selecting slightly older or cheaper models can yield massive cost savings, illustrated by the Pareto curve of intelligence versus benchmarking cost, where frontier models like Opus 4.8 cost over $4,000 to fully benchmark.
Hardware Improvements Underpinning Efficiency Gains
George briefly highlights how hardware advances, such as the NVIDIA B200 node, improve both output speed and system throughput, enabling better amortization of hardware costs across more users and further driving down inference costs. This sets up the transition into discussing agent categories and their real-world applications.
Seven Fast-Growing Agent Categories
George outlines seven agent categories that Artificial Analysis tracks as having reached product-market fit and rapid growth: coding agents, general work agents, chatbots, presentation agents, OCR, data analysis, and customer support. He notes these categories represent mature, real-world use cases for AI agents beyond simple chatbot interactions.
Benchmarking Coding Agents: Models and Harnesses
George explains that coding agent performance depends on both the underlying model and its harness (e.g., Claude Code, Cursor CLI, OpenCode), presenting benchmark results showing Opus 4.8 and GPT 5.5 leading in intelligence but costing over $4 per task. He highlights significant cost-performance trade-offs, noting cheaper options like Cursor CLI with Composer 2.5 under 50 cents per task, suggesting cost-sensitive users consider alternatives depending on the task type.
Key Trends Shaping AI Through 2026
George closes by outlining major trends expected to continue through 2026: general work agents becoming mainstream tools for knowledge workers (e.g., Claude Cowork, Codex Desktop), the blurring line between coding agents and general agents as code becomes the universal task-completion paradigm, and the growing maturity of continual learning in agents. He ends with a call to action, noting Artificial Analysis is hiring in Melbourne, across Australia, and in San Francisco.
Alrighty. Are we, there we go. Great. Thanks for having me here. So I'm George, one of the co founders of Artificial Analysis. Very quickly, about us. We're an independent AI benchmarking company. We help millions of users with our website, artificialanalysis.ai, understand what's happening in AI and choose between all the different technologies with our independent benchmarks. And we're also quite widely referred to in the industry at the start of this week.
Jensen referred to us for the Nemotron three Ultra release last week, Anthropic with Opus 4.8, and our GDP Val AA benchmark the week prior to that, Sundar with a g p t sorry, Gemini 3.5 Flash release. And so very happy to be here and share a bit about what we do and give an overview of how we see the state of AI in June 2026.
We benchmark across the AI stacks. We benchmark agents, models, inference providers, and hardware, and we also benchmark kinda text focused language models, but also image, video, speech, and music models as well. I think to start off with, or provided essentially this chart, it's a bit hectic, which shows the leading releases from the leading AI labs over the last few years.
And it's a bit hectic, but I think that's probably no surprise to us in this room, because it's been a hectic few years. There's been a lot of releases and a lot of progress, and I think why I like to start with this chat is it shows that AI progress has not slowed down. I think claims of that have been greatly greatly exaggerated.
I heard a lot of that in late twenty five, but then Opus 4.6 came out. Agents started working for much longer horizons, and people quieted down a bit. And I think this chart shows that really well, is that there's more dots on this chart than ever. If you look at the last three months here, which shows there's been more leading releases, and it's going up and to the right pretty quickly, which is showing progress in intelligence.
I'll first introduce our artificial analysis intelligence index metric. Now this is a synthesis metric of 10 benchmarks that we run to test language model intelligence, which provides a high level synthesis overview and relative comparison of the intelligence of these models. We have the current set of leading models on this chart here, and I think to note, Claude Opus 4.8 last week took the mantle from GPT 5.5 as the leading language model in terms of intelligence. But I think rather than kind of stopping there if if we were able to stop there, it'd make our jobs a lot easier, but I think there's a the the story is there's there's a reason still to use other models, considering the cost, the speed, and other trade offs at play amongst language models today.
I'll first take a look at different perspectives on thinking around the current state of models. You can see here that it's very much a US and China story when looking at where the labs making these leading language models are based. Also present on this chart is France with Mistral, South Korea with a number of labs and a very successful sovereign AI initiative, and then also The United Arab Emirates.
I think to us in this room, notably missing is Australia. Australia doesn't have kind of language models that are competitive with the frontier in terms of intelligence. And from my perspective, I I don't think that's going to happen soon either. Another perspective is looking at open weights or colloquially open source models compared to that of proprietary models and proprietary intelligence achieved.
On the y axis, we have our intelligence index, and then on the x axis, release date, and what these lines plot is the leading model that is kind of proprietary compared to OpenWeights. And OpenWeights has always trailed proprietary intelligence. This chart is really going back to kind of mid mid twenty three with the Lama two seventy b release, and you can see that, yes, OpenWaits has trailed open sorry, proprietary intelligence, but it's in a sense kind of kept up roughly three to nine months behind that of proprietary intelligence.
This was not a given. It was an open question. Years ago, or only a few years ago, people asking around the commercial models that could support open weights, but I think what we've seen has kept up for a number of reasons, and I think our perspective is this is going to continue with there being sufficient incentive there, particularly for labs wanting to serve those looking for open weights models for their flexibility, the ability to fine tune, and other reasons. And then secondly, I think another perspective on this is if open weights is kind of three to nine months behind, we're looking at OPUS maybe 4.5 territory, GBT 5.2 for the intelligence that you can essentially get with open weights models, with a Kimi k two point Kimi k 2.6 or the the latest kind of DeepSeg v four Pro. And so it remains an option.
You could do a lot with with with OPUS 4.5 or GBD 5.2. And so it remains an option, and it looks like that's going to continue based on the last few years and our house perspective. We benchmark using quantitative benchmarks, but we try and essentially ensure that those benchmarks are aligned from real world use. You don't have models hill climbing on a number that doesn't relate to increasing in, essentially, the usability of these models the use cases we're exploring with AI now, and so we like to sense check, essentially the progress. And I think what we tell here with our gdpval.a agentic benchmark that uses an OpenAI data set, we turned it into an agentic benchmark, is a perspective of, okay, a leading model a year ago with Claude four O Claude four Sonnet, sorry, and then now with GBD five point GBD 5.5, and you can see progress is being made in terms of real world knowledge work output.
On the left, we have a table useful. On the right, we have essentially additional synthesis, executive summary provided, very useful totals, and further analysis than we had a year ago. So agents have progressed on economically valuable tasks, and we also like to look at kind of maybe less knowledge work y, economically valuable tasks as well. Progress has been made there as well.
And so this is a kind of music video mood board task within GDP val a a, and you can see that this is comparing to kind of some open weight models, but on the left, we have what Gemma e4b. I mean, it's amazing that it even did this, but I think as a music video mood board, maybe I'm not creative enough, but I don't think I could use the one on my right here.
But if we compare that with a GBD 5.5, maybe I could kind of do something with this. And so I think it shows that there are still real differences between the models, and the kind of benchmarks do correlate to real world capabilities. I think one, kind of two things that are hard to reconcile in AI today, is one, that we have cheaper intelligence.
So it's cheaper than ever to access GPT four level of intelligence. Kind of going back to what I was saying earlier, you can get kind of Kimi K 2.6 now cheaper than you can get Opus 4.5, you know, six months ago, and so I think this speaks to you can get intelligence cheaper than ever, but we're also spending more than ever within our companies.
We're kind of upgraded to the $200 a month Claude code or Codex plans. And kind of why is that is? And I kind of break it down into six drivers here that we see as the most critical. I'll kind of breeze through them. Each of these could be a talk, but, like, smaller models able to achieve more intelligence is the first.
Lowering cost, increased sparsity of the models, lower proportion of active parameters to total, kind of when inference compute is at scale, the kind of number of active parameters is really the driver of cost, how you amortize hardware. Next is kind of software efficiency gains to call out here. In the inference stack, in kind of VLLM or SGLANG.
Optimizations are happening every day, that of flash attention. Also, kind of closer to the model, exploring different quantizations. We're no longer at BF 16. Now we're at four bit precisions, whether that's INFOR with Moonshot models or whether that's NVFP4 with other models. We're making better efficiency trade offs essentially to offer cheaper intelligence and increase efficiency. Third is hardware efficiency.
Hardware is kind of costing more, but is also offering more. And so a NVL 72 node is able to offer a lower cost when serving at scale than your h 100 node, even though it costs more, and so that higher efficiency can translate to lower costs offered. On what's raising costs, larger models.
And how do you reconcile this with the first? It's that everybody wants or not every use case, but there's an insatiable demand for frontier intelligence. We see that continuing. I think for yourself, I'd love to be more intelligent. I'd love the people I work with to be more intelligent, and so I think that translates to models, and so there's an insatiable demand for kind of larger models increases cost.
Next, reasoning models. Labs are increasing intelligence through more reasoning tokens. Everybody knows, we pay per token. And so, higher cost. Lastly, agents. So it used to be you, you know, in ChatGPT, you would ask a question or send a request to the API, and you get the response back, maybe make that, maybe that's, put that into a word through code.
Now you're having agents doing exploration of other files, of checking its own work, and it's improving the output, but we're dealing with multiple turns for the same task. And so in our benchmarks, we're commonly seeing 60 turns for a GDP valet A task, and so kind of 20 to 100 turns is quite sensible for many knowledge work tasks, and it acts as a multiplier on the cost.
This is a chart showing kind of model, language model inference price falling. So here, we've bucketed models in their intelligence, so in our intelligence index, zero to ten, ten to 20, etcetera. And we've looked at, okay, when was that model, when was that intelligence achieved, which is the first dot of each line, and then where are we now in terms of the latest release, or the cheapest release that is able to access that intelligence?
And what we see here is really in kind of six to eighteen month periods, you have the cost of that intelligence falling, often 10 to 100 times. And this is something that we can all think about and take advantage of. If a task was able to you could do it with OPUS 4.5, odds are now that you can there's orders of magnitude at play. It's not two x cheaper, as you can go 10 x cheaper, in many cases, by choosing a cheaper model that has more recently been released.
And so I think both insatiable demand, but for, like, kind of increased intelligence, but especially for, like, real world, more defined use cases, often there's 10 to a 100 times cheaper options at play in terms of model selection. And this is that played out again. So here we have a chart of intelligence, and then the cost to us of running benchmarks.
And so here's the on the x axis, we have the cost to run our intelligence index, those 10 benchmarks, and I don't know if we would be able to afford it back in when we kind of started at kind of in 2324, but now it's costing above 4,000 US to run these 10 benchmarks for frontier models that are pushing what AI intelligence is able to achieve.
And so you can see OPUS 4.8 costs over $4,000, but there's a clear kind of Pareto curve that one should think about when selecting models for use cases, thinking about the cost sensitivity, and there's options at play amongst the Pareto curve that really does span a 100 x plus in terms of cost differences.
Again, could talk about it in a lot more detail, but this is a chart just showing what's underpinning a lot of the efficiency gains, and that's hardware improvements. So you can see that the kind of a B200 node is able to offer greater output speed per query, but also greater system throughput. It can scale to more users, greater concurrency.
And so this allows amortizing the cost of that hardware over more users, which is supporting lowering of intelligence costs and inference costs. I think agents are a big big category, it requires different ways to think about each of the agent categories.
We track seven fast growing agent categories that we see as having achieved product market fit, in a sense, and are fast scaling. So that's coding agents, general work agents, not just ChatGPT, but your co your local codex, your Claude co work that have really reached a level of maturity, a lot more to go, but a level of maturity for real world use.
Next, chatbots, presentation agents, OCR, data analysis, and customer support as seven agent categories that have really reached a level of maturity and product market fit, and they're continuing to grow very quickly. Double clicking into that, which is probably most interesting to our AI engineering audience here, and that is coding agents.
So coding agents, they're not the model, they're the model and the harness. Both are drivers of performance. Because when we're using coding agents, we're not just using OPUS 4.7 or OPUS 4.8, we're using OPUS 4.8 in Claude code, in Cursor CLI, in OpenCode, and so we benchmark both to give a view of how those interact and the intelligence of the model and the harness together. I think what we see here is that we have OPUS 4.8, I think, coming today.
We're finishing up benchmarks, but that's going to lead in terms of intelligence in our coding agent index. GPD 5.5 from OpenAI is close by, and what's interesting is that there's a bit of a drop off when looking to the other models the performance that they offer. And so I think probably not a huge surprise to a lot here. I think it shows the importance of kind of continuing to stay on top of things and upgrading to the latest models.
But again, with coding, like when using language model APIs directly or inferencing yourself, I think you can see that there are still big trade offs at play when we think about speed, when we think about cost. And so here is the kind of coding agent index score, and then on the x axis, we have cost per task.
You can see that definitely Claude Code with Opus 4.7 and Codex with GBD 5.5 offering leading intelligence but costing over $4 per task. Then, if you come way back here, you can see that models that are in line of sight there, but not quite achieving that level of intelligence, are under 50¢, notably Cursor CLI with Composer 2.5.
And so, again, orders of magnitude, an order of magnitude trade off at play here, and it might still make sense, probably for a lot in this room, for your regular coding activities, use the best model achievable and pay the price, but for other tasks, coding agents, I'll go into it a little bit later, but coding agents are becoming agents, agents are becoming coding agents.
For other tasks where you're using coding, think that worth considering other models. Lastly, I'll walk through some of the key trends that we're seeing across the AI stack that we're monitoring and we think will play out across the rest of 2026.
So starting with agents, I think the first trend is kinda general work agents starting to roll out to the masses. Kinda our experience six or nine months ago when using Claude Code for the first time and getting an immediate productivity uplift from agents that can do long horizon tasks, debug themselves, what's going wrong, fix problems, before coming back to us is coming to knowledge workers more broadly with Claude Cowork, Codex Desktop, and other work general work agents that essentially I mean, they're pretty thin somewhat like thin wrappers in a sense. You can think about them this way, over Cloak Code or over the model, but they're able to do real world tasks, and that's something that we see kind of rolling out to the masses, and the masses are going to have our experience that we had nine months ago.
Next, agents are coding agents, and also coding agents are agents. So we're seeing it mature, essentially, becoming more mature, this paradigm of agents just using code to complete work tasks, even if they're not coding tasks as we think about them. And it's because, essentially, the labs have focused on code.
It's a flexible paradigm that they can use to complete tasks, and it works. And so I think we're seeing coding agent performance become more important, not just for doing coding tasks, but for doing general work tasks as well. And so I think, practically for us, when we're looking at models and thinking about how it's gonna complete tasks, it's not just about, okay, kind of design, how well is it integrated, like, essentially first party support of harnesses to our libraries.
Think about, okay, how is this model going to do tasks? Often it's through running code. How good is it at running code? Third, continual learning working natively is something that is maturing a lot as well and a focus of the labs. There's a few more there.
Might post this on our Twitter artificial analysis. But I guess we're out of time. Thanks very much. We're hiring in Melbourne and everywhere across Australia as well as SF. Please reach out. Thank you.
People
- Jensen Huang
- Sundar Pichai
Technologies & Tools
- B200
- BF16
- ChatGPT
- Claude 4 Sonnet
- Claude Code
- Claude Cowork
- Codex
- Codex Desktop
- Composer 2.5
- Cursor CLI
- DeepSeek V4 Pro
- Flash Attention
- Gemini 3.5 Flash
- Gemma 4B
- GPT-5.5
- H100
- Kimi K2.6
- Llama 2 70B
- Nemotron 3 Ultra
- NVFP4
- NVL72
- OpenCode
- Opus 4.5
- Opus 4.6
- Opus 4.7
- Opus 4.8
- SGLang
- VLLM
Concepts & Methods
- Artificial Analysis Intelligence Index
- Coding Agent Index
- Continual learning
- GDP Val AA
- Open weights models
- Pareto curve
- Sovereign AI
Organisations & Products
- Anthropic
- Mistral
- Moonshot AI
- NVIDIA
An analysis of the leading AI models and the real trade-offs between them across intelligence, speed, price, token usage, and beyond, grounded in Artificial Analysis’ independent benchmarking. Includes a forward-looking read on the trends shaping where AI is heading.














