Token Town (why compute strategy is product strategy)

The Challenge of Serving AI Without Fortune 50 Leverage

The speaker opens by polling the audience about company scale and AI product delivery, framing the core challenge: how smaller companies can deliver great AI experiences without the negotiating power of Fortune 50 firms. They introduce the team behind Notion's AI work before diving into real-world pricing scenarios that illustrate the problem.

Real Pricing Traps: Reasoning Upgrades and Model Deprecation

Using anonymized but recognizable examples, the speaker walks through two scenarios: a reasoning model upgrade that triples output tokens at the same per-token price, and a successor model that's 40% more expensive while the predecessor gets deprecated. They highlight how companies without dedicated AI negotiation teams are left with no leverage in this market.

Representing the 'Fortune 5,000,000' at Scale

The speaker explains Notion's role in negotiating on behalf of smaller customers who lack scale, leveraging its own massive token usage and customer base. They discuss how usage-based pricing determines AI accessibility and note that while $200/month plans can offer huge compute value for frontier tasks, this isn't appropriate for all workloads.

When Your Supplier Becomes Your Competitor

The speaker discusses the risky dynamic where AI providers like OpenAI and Anthropic can become competitors, describing two failure modes: companies locked into exclusive frontier lab deals with no exit, and companies with undefensible value built as thin API wrappers. They argue the solution is building product moats through data flywheels, RL on open weight models, and strong evals rather than trying to out-train frontier labs.

Notion's Managed Agents and the Cost-Per-Capability Mindset

The speaker showcases Notion's newly launched managed agents product, which lets users combine multiple AI agents (Decagon, Claude Code, Codex) for different steps of a workflow, emphasizing that the value lies in orchestration and experience, not raw model access. They stress evaluating cost per capability per second rather than capability alone, citing recent viral examples of companies caught off guard by cost/latency tradeoffs.

Matching Model Tiers to Task Complexity

The speaker explains how Notion segments tasks by complexity—simple tasks like database edits or email triage use cheaper or open-weight models, while complex tasks like deep research get frontier models. They critique the current provider duopoly for not pricing according to task needs, and warn against locking into a single provider given how fast the frontier model landscape shifts (Notion changes its default model every 3-4 weeks).

Optionality as Leverage: Notion's Auto Model Picker

The speaker describes Notion's 'Auto' model picker, which selects the best model for each task by latency, quality, and cost, while still letting users override the choice 25% of the time. This approach demonstrates how maintaining optionality across providers protects margins and user experience rather than committing to one vendor.

The Playbook: Build for Multi-Model, Evaluate on Value

The speaker presents a concrete playbook: build for multi-model flexibility, evaluate providers based on end-to-end task value rather than raw token price, and use competitive dynamics among providers to negotiate better deals. They use a web search provider example to show why per-request cost alone is misleading, and advocate for choosing open weight and RL'd models to gain financial independence.

Open Weight Models Closing the Capability Gap

The speaker highlights how open weight models like Kimi 2.6 are now matching frontier models like GPT-5.2 on Notion-specific evals, offering negotiation leverage and inference provider choice. They stress the importance of tracking token consumption and error/retry costs, not just headline pricing, recommending Philip Keeley's 'Inference Engineering' book for deeper understanding.

Architecture Over Model Choice: Harness Engineering

The speaker explains that harness and architecture decisions can impact price by up to 3x more than model selection alone, discussing options like native harnesses (Codex, Claude APIs), open source tools like Py, or custom-built harnesses tailored to specific needs like prompt caching.

Leaving Token Town: Workers and Deterministic CPU Tasks

The speaker argues that not every task needs an LLM, illustrating with a simple CSV-to-PDF webhook example that shouldn't require reasoning tokens. They introduce Notion's Workers platform (built with Vercel), which offloads repeated deterministic tasks to CPU-based code actions, reducing token costs by up to 80% for some customers.

Closing Thoughts: Collective Leverage and Responsible AI

The speaker closes by urging the audience to advocate collectively for better AI pricing and accessibility, noting AI's surprisingly poor public perception in California. They frame Notion's mission as making AI valuable and accessible for smaller businesses ('the Fortune 5,000,000'), and invite attendees to connect via Twitter or email before wrapping up the talk.

Hi. Whoo. Hello. We're good? Yeah. Okay. Thank you for having me. If I was in the same cinema, I would be giving artificial intelligence a gigantic high five. So I'll just assume that you see me and you're doing it. The reason is because I completely agree with everything that was just presented. And I think what I want to talk about is the challenge of actually serving that and making those decisions.

Raise your hand if you work for a Fortune 50 company. Okay. Good for you, three of you, four of you. Raise your hand if you don't. And raise your hand if you're serving AI products to your customers. Okay. What about us? Right? What does it mean to work at when you don't have the scale and the negotiating power, but you wanna be delivering what's best to your customers?

But first, like any engineering manager, it's never right to take credit for everything that you're doing. This is just a sample of the team that actually built everything I'm about to talk about. So I always like to really include kind of just a sample of everything we're building. So I wanna walk through some real scenarios. And again, I we didn't share notes.

So these will feel very familiar with what you just saw. And these are real. I won't name names, but you can quickly Google. So exhibit a, a reasoning model gets upgraded. But don't worry. It's the same price. Okay? It might be the same price per token, but you run it on the same exact task. And what happens? It's three times as many output tokens.

Okay. Maybe that makes it better. It reasons more. But what do you do? Right? Let's look at exhibit b. There's a successor model. So increment your random decimal by point one. Okay? It's 40% more expensive than its predecessor. But breaking news, the predecessor is being deprecated in four months and you've built your whole product on top of it.

Are you increasing your prices by 40%? No. Hopefully not. So what do you do? Well, if you're one of the four people that raised your hand, it's fine. Your CEO gets to go on Bloomberg and complain about it and you probably get a great deal. But what about everyone else? Right? The Fortune five hundred, they have these dedicated AI teams.

They get to kind of negotiate, and they negotiate alone, and they negotiate with incredible leverage. And I have a hypothesis that that's leaving behind a lot of the market, and it's not an efficient market. You're basically left today. I see this all the time with our customers, and I see this with my partners in the industry. You're left behind with no leverage.

My job at Notion is to represent that Fortune 5,000,000 because there are customers. Especially with usage based pricing, the accessibility of AI is determined by the price, by the price of those workflows. You know, we're handling tens and tens of trillion tokens a month, but we also have over a 100,000,000 customers. What does that mean?

It means that I get to play that job of negotiating at scale, but bringing it to the rest of you. And there's a couple lessons I learned that I don't think you need those trillions of tokens a month to use. And I want to share that with all of you because I actually think in this economy, there's a lot to be gained and I think the market's not fair to us.

So we see this. Right? For $200 a month, you can consume $5,000 in compute. How is that possible? Because Anthropic has the cost of good serve for themselves. And by the way, we're very close partners with them. And I think for very frontier reasoning tasks, this is appropriate. Okay? For really hard tasks, I wanna pay more because we're not there yet. But for a lot of situations, that's that's not where we are. So the supplier is your competitor.

Right? How many of you see in the next ten years that you could be competing with OpenAI or Anthropic on some product that you're building? Okay. I feel like not enough of you are raising your hands, which is interesting. But this is the way that it's working right now. And I see this happening in two variations. One, I see that people are locked in.

They have no exit. When you see very large applied AI companies with extremely, extremely vocal partnerships with Frontier Labs, there's a very high chance that they're locked into that one vendor, that they've taken all of their spend and they've committed it to that one Frontier lab in exchange for exorbitant discounts, but they're stuck. Okay? That means that tomorrow, if open weight or another Frontier model comes out, they don't have the optionality to leave because they've committed $20,000,000 to one of them.

The second place I see this happening is when you have value that you can't defend. You know, you're obviously getting a bad deal on tokens. And if you're a 2022 era wrapper around the API, you have nothing durable in your business. You have no financial independence. And we'll talk through both of those. You win on the product.

You can't be winning on tokens. You need a data flywheel. We saw about open weight, and we'll talk about it more. You need the ability to either reinforcement learn on that open weight model, increase your own products intelligence with the right evals to understand what models are appropriate, and reduce your financial dependence on frontier pricing. And the second one are those moats.

You need a compelling UI. You need orchestration, architecture, integrations to really justify the cost that you're spending. Your job here is not to train the best model. It's not. And if it is, you probably should work at one of those three companies. Your job is to build the best product that uses the best model available, whatever it is.

So for instance, at Notion, we just launched this. This is our manage agents product. We believe that there's value in optionality. You'll see in this example, you can use Decagon agents, move them so Claude Code can write a fix, and then perhaps have Codex review it, and then put it in a task list for humans to review. Right? We can charge you sticker price if not a small discount on these models, but the benefit that you're getting is the experience around it.

And it's not just capability. We see this all the time now. Think kind of I've read a Twitter article that a lot of people identified with a couple months ago now. I think the pot's boiling on capability alone. You should always be thinking about cost per capability per second. Artificial analysis is a wonderful resource to do that if you don't have your own resources.

And, like, it's it's happening. Right? We see these crazy stories. This is just in the past week, headlines that have come out, where if you only focus on capability and you don't think about latency or cost, you've put yourself in a rough position. It's pretty funny, honestly. I think I love these tweets. If you're on AI Twitter, which half of you are, it's it's a funny time to be alive.

But it's real. Right? But you don't want this to be your customers. Okay? And it's not appropriate to put your customers in this position of whatever this guy's doing with his cigar. Not all traffic is equal. We heard a little bit in that graph where we saw, you know, where different models are living. In Notion, it looks like this.

Changing a database field, triaging an email, you know, looking like this dude. Okay. Let's get him on mini max as soon as possible. Summarizing meeting notes. These are all things that we actually don't wanna pay 40% more when GPT six comes out or whatever is next. We wanna either be using open weight or we wanna be RL ing. But for the frontier tasks, we need to be giving that frontier to our customers, data analysis, deep research.

The point is that our customers need us to choose for them. Otherwise, this will be their headlines. Okay? And your product is inaccessible. The problem again is that this isn't how our providers are working. Right now, they're incentivized basically with two paths. Again, won't name names. Use your critical reasoning or your favorite model to figure it out.

Either you are the best reasoning model, you are the example in all of these tweets. Right? You are the example of great, and no one really questions your price because you're the first one that passed whatever benchmark you said you passed. The second is you're slightly worse, but that's okay. You just need to be about thirty seconds per million tokens cheaper, and you have the rest of the market to you.

What? What about everyone else? Right? We're in a duopoly right now where things aren't priced for their needs. Complex tasks, though, I think are appropriate. I wanna keep highlighting this. We see a bifurcation at least in notion usage, and I want you all to be investigating and thinking about. There is a place for complex, expensive reasoning models.

The trap is putting everything there. And the best frontier model is changing fast. This is an older slide you saw brand new from artificial analysis. Okay? But the point is that it's changing constantly. We're changing our default model for our customers probably every three to four weeks. And it's a lot of work and we have very advanced evals and teams to do it, but you shouldn't need that.

I mean, are a lot of benchmarks you can be using. If you're switching all the time, you can't be locking your product into a particular provider. Because if you are, the second you go from January to February, you're giving your users a worse experience. Okay? It's not worth the discount. I don't think so. I think you're putting a big risk on your company and you're and you're willing to fall behind.

And depending on the business, for that kind of saturated knowledge work land, that's fine. But if you wanna be offering frontier, you shouldn't be doing this. Optionality is leverage. You should be ready to walk at all times for a durable business. You should not be locking yourself into a single provider and you should be comfortable knowing what the landscape of models are so that you're ready to maintain your margins and maintain your business in a way that you're confident brings your users the best experience.

We do this with our auto model. So today, we have a model at Notion that's called Auto. This is our model picker. It's due for a refresh. The number of models is getting long. But we let users choose because sometimes that latency, quality, cost, trade off is not something you can assume for your customers. You'd be surprised how many people really want to spend a ton of money on email triage and how many people really don't care about how accurate the research tasks are.

I mean, it's a phenomenal user research exercise that we could have a whole other conversation on. And so this is how we do it. We choose for our customers the majority of the time, but 75% of the time, they they stick there. But that 25% is valid, and we give them opportunities to leave. With that auto model, we have the opportunity to actually think about the task at hand and give them the most quality that they need at the best price and latency.

Okay? So if you're setting up an email triage agent, it's very unlikely that auto will be opus because you're gonna get your first usage based pricing bill and say, no way, Notion, I'm out of here. Okay? This is the playbook. Okay? All of you that love the taking your phone out with slides, this is the time. This will be online, but I I get it.

Okay. Build for multi model. Be ready to switch. Understand what the model providers are. Evaluate on value, not tokens. What does that mean? This is an example. We tweeted this a couple months ago on web search providers. It might be that a certain web search provider is cheaper per token or per request. How many requests is your agent doing?

How accurate are the results? You should be looking at the entire task when evaluating. You should not be thinking about a particular API call. That's where they get you. Okay? Think about your use case. You are the expert on quality for your use case. Switch fast, switch often, give them something back.

Do you know why these evolves were awesome? I love competitive dynamics in a market. I love it. If we say parallel, we chose you, but it's close. Stay up there. Everyone else on the list, here's exactly why we didn't choose you. Please fix it. A rising tide lives all ships. Position yourself in a way where you're getting what you want from your providers. Forego discounts for optionality.

I think we talked about that already. And there's a third option, which I'm sensing is the theme of the day, which is open weight. That moderate task, we're very heavily considering now. And we have live in production a lot of open weight traffic as well as reinforcement learned models. They're strong enough to handle workloads. I don't think that they are going to be in the upper right quadrant of capability soon, but the gap is closing.

And it also gives you negotiation leverage, complete financial independence, choose the inference provider of your choice. Right? I would say 2.6 was the most was the first time this really happened for us. There's been a lot since then. But 2.6 was the first moment where we actually saw it compare to GPT 5.2 in quality.

So here's that eval. We saw we see scores, and these are on Notion specific tasks. Again, you own your product. You get to decide what works for your product. For Notion specific tasks that we thought the auto model needed to do for a subset of them. We score just fine.

Kimmy two six is great. And errors are important, by the way. Errors are what you end up paying for. Just as a side note, you're still paying token rates on errors and retries. Think about that in another presentation. But it's not just the token costs. We look at the average number of tokens. There's some crazy things on this, like Opus four seven, $4.06 versus Sonnet. Look at the token consumption.

Forget the price on Opus versus Sonnet. Look at the token consumption. Right? This is why understanding your whole trajectory is really important for understanding your model. I love Philip Keeley. He has a great book called Inference Engineering. Even if you're not an inference nerd, it's good to read. We don't need to be frontier with open weight.

It's just a gap in time. Right? There's product that we served six months ago that our customers love, and they don't want to randomly pay for more tomorrow. Right? We're trusting that open weight is closing that gap, but we need to have the right evals and stay on top of it to know when it's there. And that's why investing in your evals are important.

How do you win these negotiations? Well, the first thing that you need to do to be ready is think about your architecture and not just your model choice. For us, we've noticed that harness engineering and architecture decisions can account for about three x the change in price as model selection. Sometimes that means using a native harness like the Codex and the Cloud APIs. Sometimes it means using open source like Py.

And sometimes it means creating your own harness because there are particular capabilities you want and you understand the implications of how prompt caching, etcetera, might affect how you work. Okay. But what if you don't need an LLM at all? I know we said that this is welcome to Token Town, but we're leaving. Okay? We're now departing Token Town.

And I wanna take us back to an old world, an old world where we used CPUs to do our jobs. This is my nana banana, I think, generated image of an engineer trying to turn a CSV into a PDF and post it on a webhook. Why would he ever need an LLM to do that?

Why would you want that repeated task constantly using reasoning tokens to navigate MCPs? Okay. Well, because Frontier Labs want you to token max. They want you to do it in a way that they control their capacity. It's its own economic issue. But they want you to token max. Your users don't. Your users want you to outcome max.

That's why we launched our developer platform and something called Workers. At Notion, we believe that many tasks do better work on CPUs. Determinism is a valuable thing. State machines existed. Right? The pendulum has gone so far because it can, but that's not durable software, and it's not a formal software. So we've partnered with Vercel on the capabilities to launch what we call workers, where you can actually call on internal APIs, computer sandboxes, host small, small code as an action that your LLM calls.

We've seen this decrease token cost by up to 80% for some of our customers on repeated tasks. It's wild out there. I mean, I feel like if I go on a week vacation and I do a good job not looking at Twitter, it's like I hibernated for two years. Okay. I get it. It's exhausting. Take breaks.

Take care of yourself. But it's crazy. And the market is so young. It's so opaque. It's moving so fast, and you are a player in it. I find it very sad that we're all kind of rolling over, but we're all together and our negotiating power is so much stronger if we all advocate for ourselves. And I also think that you owe it to your customers.

I'm really proud. I'm so proud. And I come to work every day because of the way that I think we make AI accessible and valuable for the Fortune 5,000,000. But not everyone uses notion and I want all of you to do that too. I want us to make AI something that isn't a meme. I mean in California, I read a statistic that AI is more unpopular than ice which is our immigration controls.

It's crazy. We don't have the, you know, it's not a popular institution, ice. So that means AI is really unpopular. Why? Because we're not creating valuable work and we're not doing it at the right price and people aren't seeing the value and there's a ton of Twitter memes, but there's real work to be done. And that's our obligation.

Right? I think it's a powerful technology and I think it's our duty to bring it to our customers in a way that's responsible, durable, and appropriate. So I'm on Twitter a lot as you saw. Please tweet at me. You can always email me. I'll be here in this hemisphere till the end of the week in Melbourne. Thank you for your time.

Thank you for hosting me. I love Australia, and have a great day.

Token Town

Why compute strategy is product strategy

Sarah Sachs, AI Lead at Notion

A Notion logo, a stylized 'N' inside a square, is in the top left. The right side of the slide features a black and white cartoon illustration of a town with several buildings, referred to as "Token Town". Signs on the buildings read: "THE BEST!", "TOKENS", "Cheap TOKENS FOR SALE!", "TOKENS SOLD HERE", "TOKENS", "100% SMART", "SMART!", "TOKEN TAVERN", and "FASTEST IN TOWN!".
The slide displays two photographs side-by-side. The left image shows a group of people working on laptops around a long conference table in an office with large windows. The right image depicts a larger group of smiling people gathered around a conference table in a bright office space, posing for a group picture.

The token market feels structured against buyers

Exhibit A

A reasoning model gets upgraded at identical per-token pricing.

Upgrade uses ~3x more output tokens for certain tasks.

Exhibit B

A successor model makes significant steps in reasoning.

Costs 40% more than its predecessor.

FORTUNE

500 GLOBAL

The Fortune 500 has dedicated AI teams

Everyone else negotiates alone, with no leverage

A gold-colored cover of Fortune magazine. The number '500' is stylized with the zeroes as tall, golden pillars, and the word 'GLOBAL' is centered within the two zeroes. Smaller text at the top reads "AUGUST/SEPTEMBER 2023 FORTUNE.COM" and at the bottom "PLUS THE MOST POWERFUL PEOPLE IN BUSINESS".

Notion represents the Fortune 5 Million

An illustration of a person holding a glowing cube with the letter 'N' on it, representing the Notion product.

Forbes:

$200/month Claude Code plan consumes up to $5,000 in compute

25x subsidy by Anthropic

Your supplier is also your competitor

Two ways this goes wrong for applied AI companies

Lock in with no exit

  • Tie yourself to cheaper provider today. Prices change and your left with no where to go.

Value you can't defend

  • You must provide additional value to supersede the obvious "bad deal" you have on tokens.

Win on the product, not the token

Data flywheels

  • Reinforcement fine-tuning on open weight.
  • Increase your product's own intelligence
  • Reduce dependence on frontier pricing.

Product moats

  • Compelling UI, orchestration, architecture, and integrations to justify the cost.
An illustration featuring two overlapping computer screens. One screen shows wavy blue lines, and a gear icon is next to it. The other screen displays a blue bar chart showing an upward trend, with a dotted line and a blue circle indicating progress.

Train the best model

Build the best product that uses many models

An illustration of a stylized person holding an open book and gesturing with one hand.

Tasks

A screenshot of the Notion application interface, displaying a task management board integrated with AI agents like Claude Agent and Codex Agent, showing lists for tasks, proposed fixes, and automated code reviews.

Cost-per-capability-per-second Tradeoff

An illustration of a speedometer with its needle indicating high speed, surrounded by lines and starbursts suggesting quickness.

One company spent half a billion dollars on Claude in a single month: Report comes as AI costs climb

05-29-2026 | NEWS

Uber burned through its entire 2026 AI budget in four months. Now its COO is questioning whether it's worth it

By Jake Angelo, News Fellow

May 26, 2026, 2:03 PM ET

Microsoft reports are exposing AI's real cost problem: Using the tech is more expensive than paying human employees

By Jake Angelo, News Fellow

May 22, 2026, 12:56 PM ET

Screenshots of three news articles discussing the high costs and budget implications of artificial intelligence usage. Headlines include topics like a company spending half a billion dollars on Claude, Uber's 2026 AI budget being used in four months, and Microsoft reports showing AI is more expensive than human employees.

Uber employees using Claude tokens

Uber employees using $100k of Claude tokens to respond to an email

Turner Novak (@TurnerNovak)
Screenshot of a tweet by Turner Novak. An image on the right shows a man in a light suit using a large torch to light a cigar, producing blue and pink flames.

Not all traffic is equal

Moderate tasks

  • Changing database field
  • Triage email inbox
  • Summarize meeting notes

Frontier tasks

  • Autonomous agent paths
  • Large-scale data analysis
  • Deep research journeys

Providers are incentivized to pursue one of two paths

  • Best reasoning model
  • Worse model, priced as close to the top as possible
A diagram illustrates two distinct paths. On the left, associated with the "Best reasoning model," are black arrows showing a path that curves upwards and loops downwards around a vertical line. On the right, associated with the "Worse model, priced as close to the top as possible," are red arrows depicting a path with a wavy upward movement and a straight downward movement, also around a vertical line.

complex tasks

An illustration shows a stylized person pointing with a stick to a whiteboard. The whiteboard has the text "complex tasks" and three colored rectangular blocks (red, orange, blue). Wires connect the whiteboard to a simple drawing of a head with red glasses, depicting processing or thought.

The "best" frontier model changes fast

Frontier Language Model Intelligence, Over Time

Artificial Analysis Intelligence Index v4.0 incorporates 10 evaluations: GDPval-AA, T²-Bench Telecom, Terminal-Bench Hard, SciCode, AA-LCR, AA-Omniscience, Ibench, Humanity's Last Exam, GPQA Diamond, CritPt

  • Alibaba
  • Anthropic
  • DeepSeek
  • Google
  • Kimi
  • KwaiKAT
  • LG AI Research
  • MBZUAI Institute of Foundation Models
  • Meta
  • MiniMax
  • Mistral
  • OpenAI
  • TII UAE
  • xAI
  • Xiaomi
  • Z AI
A line graph titled "Frontier Language Model Intelligence, Over Time" showing the Artificial Analysis Intelligence Index score (Y-axis, from 0 to 60) against Release Date (X-axis, from November '22 to March '26). Multiple colored lines, each representing a different language model from various organizations, show their performance on the index over time. The graph indicates a general upward trend, with models from companies like OpenAI, Anthropic, Google, and others showing significant increases in their intelligence index scores, particularly from late 2024 to early 2026, reaching scores between 30 and 50. The graph illustrates the rapid advancement and changing landscape of frontier language model intelligence.

Optionality is leverage

Preserving optionality is how you get the best price

An illustration shows a stack of blue and white rectangular blocks, with a blue block being added to the top of the stack by a white arrow, symbolizing building or adding options.

Notion's Auto Model

  • Ensures every customer is always on a state-of-the-art model
  • Handles 75% of AI traffic

Model Picker

Currently selected: Auto

Select a model (Beta)

  • Sonnet 4.6
  • Opus 4.7
  • Opus 4.8
  • Gemini 3.1 Pro
  • GPT-5.2
  • GPT-5.4
  • GPT-5.5
  • Grok 4.3
  • Grok Build 0.1

Open models (US-provider hosted) (Beta)

  • Kimi K2.6
  • DeepSeek V4 Pro

A screenshot of a Notion user interface element, resembling a dropdown or modal for selecting an AI model. It shows "Auto" as the currently selected option. Below this, a list of models includes Sonnet, Opus, Gemini, GPT, and Grok models, each preceded by a distinct icon, followed by a section for "Open models (US-provider hosted)" which lists Kimi and DeepSeek models.

Here's a high-level summary of the findings from Google [SOT] Web Search Eval Readout.

ProviderAgent prod-log Recall@10
(% queries w/ GT in top-10 snippets)
Agent prod-log One-hop correct@10
(% queries answered correctly using only top-10 snippets)
Avg tokens (K=10)
(avg total tokens across 10 snippets)
Freshness Recall / Correct
(on 10 "freshness" queries)
Key takeawaysLatency notes
Parallel96.5%
URL overlap vs Serp: 48.5%
89.5%192670%/70%Best overall on prod-log style queries; strong URL overlap vs Google; good snippet density (lower avg token count vs many providers).Avg latency
2291ms.
[Redacted Provider]78.9%
URL overlap vs Serp: 44.1%
70.2%246190%/50%Improves recall vs [redacted text] but freshness is much worse than peers in the 10-query Freshness slice (despite high freshness recall in the overall Recall@10 table).Avg latency
4252ms (slow).
[Redacted Provider]82.5%
URL overlap vs Serp: 35.9%
73.7%2991100%/70%Great freshness in the 10-query Freshness slice; solid overall but below Parallel on prod-log recall.Fast (751ms).
[Redacted Provider]93.0%
(Serp reference)
86.0%38870%/50%Very strong baseline; used as a stand-in for results.Avg latency
2362ms.
[Redacted Provider]75.4%68.4%748370%/70%Fast and strong on SimpleQA/synthetic/multilingual, but 0% on walled-garden aggregate in this eval and notably worse than Parallel on prod-log recall.Very fast
(666ms).
[Redacted Provider]96.5%84.2%362070%/60%Ties Parallel on prod-log recall; [redacted text].Avg latency
1147ms.

Here's a high-level summary of the findings from [SOT] Web Search Eval Readout.

ProviderAgent prod-log Recall@10
(% queries w/ GT in top-10 snippets)
Agent prod-log One-hop correct@10
(% queries answered correctly
using only top-10 snippets)
Avg tokens
(K=10)
(avg total tokens across
10 "freshness" queries)
Freshness Recall / Correct
(on 10 "freshness"
queries)
Key takeawaysLatency notes
Parallel96.5%
URL overlap vs
Serp: 48.5%
89.5%192670% / 70%Best overall on prod-log style queries; strong URL overlap vs Google; good snippet density (lower avg token count vs many providers).Avg latency
2291ms.
[Obscured]78.9%
URL overlap vs
Serp: 44.1%
70.2%246190% / 50%Improves recall vs [obscured]
but freshness is much
worse than peers in the 10-
query Freshness slice (despite
high freshness recall in the
overall Recall@10 table).
Avg latency
4252ms (slow).
[Obscured]82.5%
URL overlap vs
Serp: 35.9%
73.7%2991100% / 70%Great freshness in the 10-
query Freshness slice; solid
overall but below Parallel on
prod-log recall.
Fast (751ms).
[Obscured]93.0%
(Serp
reference)
86.0%38870% / 50%Very strong baseline; used as a
stand-in for [obscured]
results.
Avg latency
2362ms.
[Obscured]75.4%68.4%748370% / 70%Fast and strong on
SimpleQA/synthetic/multilingual
[obscured], but 0% on walled-garden
aggregate in this eval and
notably worse than Parallel on
prod-log recall.
Very fast
(666ms).
[Obscured]96.5%
URL overlap vs
Serp: [obscured]
84.2%362070% / 60%Ties Parallel on prod-log recall;
[obscured]
Avg latency
1147ms.

Third option: Open-weight

Moderate tasks

  • Changing database field
  • Triage email inbox
  • Summarise meeting notes

Frontier tasks

  • Autonomous agent paths
  • Large-scale data analysis
  • Deep research journeys

A diagram categorizes tasks into "Moderate tasks" and "Frontier tasks". "Moderate tasks" is highlighted with a red circle and an arrow.

Open-weight models are now strong enough to handle these workloads

Cost lever

Stop overpaying for capabilities you don't need on routine tasks

Negotiating leverage

A credible alternative puts downward pressure on frontier pricing

Kimi 2.6

Gigantic industry shift

  • Open-weight is closing the gap
  • Plays like top-class agentic model
  • Cheaper direct inference provider pricing
  • Outperforming GPT 5.2
  • Reliable for quality + price + latency
Screenshot of an application interface showing a social media post by Akshay Kothari (@akothari) announcing Kimi K2.6. The post reads, "Kimi K2.6 just landed in @NotionHQ. Open-weight, but absolutely a heavyweight." Below the post, a dropdown menu lists various AI models, including Sonnet 4.6 Beta, Opus 4.6 Beta, Opus 4.7 Beta, Gemini 3.1 Pro Beta, GPT-5.2 Beta, GPT-5.4 Beta, and Kimi K2.6 Beta, with Kimi K2.6 Beta selected.

eli(as) @earlierism · Apr 25

We've often been asked about how we eval each new model at Notion, since we're one of the few apps deploying both frontier lab models and leading open source models for general knowledge work.

Read on for what we share with the model providers after we evaluate their models.

Modelav score (%)av tool calls (n)total tool errors (n)av tokens (n)total tokens (n)av time (s)total time (s)
Kimi k2.644.4232213,265145,91724.2265.7
Sonnet 4.635.321139006117,08114.3186.0
Opus 4.648.3291710,869141,29220.5266.8
Opus 4.762.425610,481136,25921.0273.1
GPT 5.238.32454432368,5556.673.1
GPT 5.443.82117331736,4836.673.0
GPT 5.562.91812582364,05214.9164.0
Screenshot of a tweet containing a table of model evaluation metrics.

MODEL CAPABILITIES FOR SOFTWARE ENGINEERING TASKS

Philip Keily

Head of AI Education

baseten

  • Closed Models
  • Open Models
  • Reasoning Agent
  • Code Generation
  • Tab Completion

Source: Inference Engineering, 2026

A line graph titled "Model Capabilities for Software Engineering Tasks" shows "Time" on the x-axis and "Capabilities" on the y-axis. The y-axis indicates increasing capability levels from bottom to top: Tab Completion, Code Generation, Reasoning Agent, Open Models, and Closed Models.

Two upward-sloping lines illustrate the progression of capabilities over time. The "Open Models" line starts at a lower capability level and rises, while the "Closed Models" line starts at a higher capability level and also rises, consistently staying above the "Open Models" line.

A circular headshot shows Philip Keily, a man with glasses and a collared shirt, smiling. Below his name and title is the baseten logo, a stylized 'b' or 't' shape.

How to win in negotiations with model labs

  1. Architecture over model choice
  2. Open weight optionality doesn't mean offering every model
  3. Build product value that transcends tokens

What if you don't need a LLM at all?

An illustration of a minimalist face with a skeptical or confused expression inside a circle.

YOU'RE NOW DEPARTING TOKENTOWN

An illustration of a woman with a backpack flying away from a town called Tokentown on an airplane. The airplane is trailing a banner with the text "YOU'RE NOW DEPARTING TOKENTOWN". The buildings in Tokentown have various signs related to "TOKENS" such as "TOKENS", "THE BEST! IN TOWN", "100% SMART", "SMART!", "Cheap TOKENS FOR SALE!", "TOKEN TAVERN", "TOKENS SOLD HERE", "FASTEST IN TOWN!".

GenAI models vs. One shell command

$$$$

csvtojson data.csv | curl -X POST 
  -H "Content-Type: application/json" 
  -d @- https://api.example.com/data

$

A slide comparing the complexity and cost of using Generative AI models versus a simple shell command for a data processing task. On the left, under 'GenAI models', there are logos for OpenAI, Google AI, and a colorful diamond, alongside a detailed diagram illustrating the various components and interactions within an AI agent's workflow, such as planning, memory, tool calling, observation, and evaluation. An illustration of an engineer working at a desk with multiple monitors is also present. Below this, four dollar signs ($$$$) indicate high cost. On the right, under 'One shell command', a terminal window displays a concise `bash` command to convert a CSV to JSON and post it to an API. Below this, a single dollar sign ($) indicates low cost.

System Architecture Diagram

  • An Agent component connects to both Worker + Tools and MCP.
  • The Worker + Tools component interacts with the Internal API, which consists of multiple distinct API endpoints.
  • The MCP component interacts with the API, which also consists of multiple distinct API endpoints.
A block diagram illustrating a system architecture. An 'Agent' block is central, with arrows pointing to 'Worker + Tools' and 'MCP' blocks. The 'Worker + Tools' block has multiple arrows pointing to an 'Internal API' section, which contains several empty rectangular sub-blocks representing API endpoints. The 'MCP' block has multiple arrows pointing to an 'API' section, also containing several empty rectangular sub-blocks representing API endpoints.

It's the wild west.

The market is young, opaque and moving fast.

We owe it to the customers we serve to get it right.

@sarahmsachs ssachs@makenotion.com

An illustration featuring a block with the letter 'N' on top, and a cartoon head with a prominent nose and a cowboy hat.

People

  • Philip Keeley

Technologies & Tools

  • Auto model
  • Claude Code
  • Codex
  • GPT-5.2
  • GPT-6
  • Kimi 2.6
  • MCP
  • Opus 4
  • Sonnet
  • Workers

Concepts & Methods

  • open weight models
  • prompt caching
  • reinforcement learning
  • state machines
  • usage based pricing

Organisations & Products

  • Anthropic
  • Artificial Analysis
  • Decagon
  • MiniMax
  • Notion
  • OpenAI
  • Vercel

Works

  • Inference Engineering