Token Town (why compute strategy is product strategy)
The Challenge of Serving AI Without Fortune 50 Leverage
The speaker opens by polling the audience about company scale and AI product delivery, framing the core challenge: how smaller companies can deliver great AI experiences without the negotiating power of Fortune 50 firms. They introduce the team behind Notion's AI work before diving into real-world pricing scenarios that illustrate the problem.
Real Pricing Traps: Reasoning Upgrades and Model Deprecation
Using anonymized but recognizable examples, the speaker walks through two scenarios: a reasoning model upgrade that triples output tokens at the same per-token price, and a successor model that's 40% more expensive while the predecessor gets deprecated. They highlight how companies without dedicated AI negotiation teams are left with no leverage in this market.
Representing the 'Fortune 5,000,000' at Scale
The speaker explains Notion's role in negotiating on behalf of smaller customers who lack scale, leveraging its own massive token usage and customer base. They discuss how usage-based pricing determines AI accessibility and note that while $200/month plans can offer huge compute value for frontier tasks, this isn't appropriate for all workloads.
When Your Supplier Becomes Your Competitor
The speaker discusses the risky dynamic where AI providers like OpenAI and Anthropic can become competitors, describing two failure modes: companies locked into exclusive frontier lab deals with no exit, and companies with undefensible value built as thin API wrappers. They argue the solution is building product moats through data flywheels, RL on open weight models, and strong evals rather than trying to out-train frontier labs.
Notion's Managed Agents and the Cost-Per-Capability Mindset
The speaker showcases Notion's newly launched managed agents product, which lets users combine multiple AI agents (Decagon, Claude Code, Codex) for different steps of a workflow, emphasizing that the value lies in orchestration and experience, not raw model access. They stress evaluating cost per capability per second rather than capability alone, citing recent viral examples of companies caught off guard by cost/latency tradeoffs.
Matching Model Tiers to Task Complexity
The speaker explains how Notion segments tasks by complexity—simple tasks like database edits or email triage use cheaper or open-weight models, while complex tasks like deep research get frontier models. They critique the current provider duopoly for not pricing according to task needs, and warn against locking into a single provider given how fast the frontier model landscape shifts (Notion changes its default model every 3-4 weeks).
Optionality as Leverage: Notion's Auto Model Picker
The speaker describes Notion's 'Auto' model picker, which selects the best model for each task by latency, quality, and cost, while still letting users override the choice 25% of the time. This approach demonstrates how maintaining optionality across providers protects margins and user experience rather than committing to one vendor.
The Playbook: Build for Multi-Model, Evaluate on Value
The speaker presents a concrete playbook: build for multi-model flexibility, evaluate providers based on end-to-end task value rather than raw token price, and use competitive dynamics among providers to negotiate better deals. They use a web search provider example to show why per-request cost alone is misleading, and advocate for choosing open weight and RL'd models to gain financial independence.
Open Weight Models Closing the Capability Gap
The speaker highlights how open weight models like Kimi 2.6 are now matching frontier models like GPT-5.2 on Notion-specific evals, offering negotiation leverage and inference provider choice. They stress the importance of tracking token consumption and error/retry costs, not just headline pricing, recommending Philip Keeley's 'Inference Engineering' book for deeper understanding.
Architecture Over Model Choice: Harness Engineering
The speaker explains that harness and architecture decisions can impact price by up to 3x more than model selection alone, discussing options like native harnesses (Codex, Claude APIs), open source tools like Py, or custom-built harnesses tailored to specific needs like prompt caching.
Leaving Token Town: Workers and Deterministic CPU Tasks
The speaker argues that not every task needs an LLM, illustrating with a simple CSV-to-PDF webhook example that shouldn't require reasoning tokens. They introduce Notion's Workers platform (built with Vercel), which offloads repeated deterministic tasks to CPU-based code actions, reducing token costs by up to 80% for some customers.
Closing Thoughts: Collective Leverage and Responsible AI
The speaker closes by urging the audience to advocate collectively for better AI pricing and accessibility, noting AI's surprisingly poor public perception in California. They frame Notion's mission as making AI valuable and accessible for smaller businesses ('the Fortune 5,000,000'), and invite attendees to connect via Twitter or email before wrapping up the talk.
Hi. Whoo. Hello. We're good? Yeah. Okay. Thank you for having me. If I was in the same cinema, I would be giving artificial intelligence a gigantic high five. So I'll just assume that you see me and you're doing it. The reason is because I completely agree with everything that was just presented. And I think what I want to talk about is the challenge of actually serving that and making those decisions.
Raise your hand if you work for a Fortune 50 company. Okay. Good for you, three of you, four of you. Raise your hand if you don't. And raise your hand if you're serving AI products to your customers. Okay. What about us? Right? What does it mean to work at when you don't have the scale and the negotiating power, but you wanna be delivering what's best to your customers?
But first, like any engineering manager, it's never right to take credit for everything that you're doing. This is just a sample of the team that actually built everything I'm about to talk about. So I always like to really include kind of just a sample of everything we're building. So I wanna walk through some real scenarios. And again, I we didn't share notes.
So these will feel very familiar with what you just saw. And these are real. I won't name names, but you can quickly Google. So exhibit a, a reasoning model gets upgraded. But don't worry. It's the same price. Okay? It might be the same price per token, but you run it on the same exact task. And what happens? It's three times as many output tokens.
Okay. Maybe that makes it better. It reasons more. But what do you do? Right? Let's look at exhibit b. There's a successor model. So increment your random decimal by point one. Okay? It's 40% more expensive than its predecessor. But breaking news, the predecessor is being deprecated in four months and you've built your whole product on top of it.
Are you increasing your prices by 40%? No. Hopefully not. So what do you do? Well, if you're one of the four people that raised your hand, it's fine. Your CEO gets to go on Bloomberg and complain about it and you probably get a great deal. But what about everyone else? Right? The Fortune five hundred, they have these dedicated AI teams.
They get to kind of negotiate, and they negotiate alone, and they negotiate with incredible leverage. And I have a hypothesis that that's leaving behind a lot of the market, and it's not an efficient market. You're basically left today. I see this all the time with our customers, and I see this with my partners in the industry. You're left behind with no leverage.
My job at Notion is to represent that Fortune 5,000,000 because there are customers. Especially with usage based pricing, the accessibility of AI is determined by the price, by the price of those workflows. You know, we're handling tens and tens of trillion tokens a month, but we also have over a 100,000,000 customers. What does that mean?
It means that I get to play that job of negotiating at scale, but bringing it to the rest of you. And there's a couple lessons I learned that I don't think you need those trillions of tokens a month to use. And I want to share that with all of you because I actually think in this economy, there's a lot to be gained and I think the market's not fair to us.
So we see this. Right? For $200 a month, you can consume $5,000 in compute. How is that possible? Because Anthropic has the cost of good serve for themselves. And by the way, we're very close partners with them. And I think for very frontier reasoning tasks, this is appropriate. Okay? For really hard tasks, I wanna pay more because we're not there yet. But for a lot of situations, that's that's not where we are. So the supplier is your competitor.
Right? How many of you see in the next ten years that you could be competing with OpenAI or Anthropic on some product that you're building? Okay. I feel like not enough of you are raising your hands, which is interesting. But this is the way that it's working right now. And I see this happening in two variations. One, I see that people are locked in.
They have no exit. When you see very large applied AI companies with extremely, extremely vocal partnerships with Frontier Labs, there's a very high chance that they're locked into that one vendor, that they've taken all of their spend and they've committed it to that one Frontier lab in exchange for exorbitant discounts, but they're stuck. Okay? That means that tomorrow, if open weight or another Frontier model comes out, they don't have the optionality to leave because they've committed $20,000,000 to one of them.
The second place I see this happening is when you have value that you can't defend. You know, you're obviously getting a bad deal on tokens. And if you're a 2022 era wrapper around the API, you have nothing durable in your business. You have no financial independence. And we'll talk through both of those. You win on the product.
You can't be winning on tokens. You need a data flywheel. We saw about open weight, and we'll talk about it more. You need the ability to either reinforcement learn on that open weight model, increase your own products intelligence with the right evals to understand what models are appropriate, and reduce your financial dependence on frontier pricing. And the second one are those moats.
You need a compelling UI. You need orchestration, architecture, integrations to really justify the cost that you're spending. Your job here is not to train the best model. It's not. And if it is, you probably should work at one of those three companies. Your job is to build the best product that uses the best model available, whatever it is.
So for instance, at Notion, we just launched this. This is our manage agents product. We believe that there's value in optionality. You'll see in this example, you can use Decagon agents, move them so Claude Code can write a fix, and then perhaps have Codex review it, and then put it in a task list for humans to review. Right? We can charge you sticker price if not a small discount on these models, but the benefit that you're getting is the experience around it.
And it's not just capability. We see this all the time now. Think kind of I've read a Twitter article that a lot of people identified with a couple months ago now. I think the pot's boiling on capability alone. You should always be thinking about cost per capability per second. Artificial analysis is a wonderful resource to do that if you don't have your own resources.
And, like, it's it's happening. Right? We see these crazy stories. This is just in the past week, headlines that have come out, where if you only focus on capability and you don't think about latency or cost, you've put yourself in a rough position. It's pretty funny, honestly. I think I love these tweets. If you're on AI Twitter, which half of you are, it's it's a funny time to be alive.
But it's real. Right? But you don't want this to be your customers. Okay? And it's not appropriate to put your customers in this position of whatever this guy's doing with his cigar. Not all traffic is equal. We heard a little bit in that graph where we saw, you know, where different models are living. In Notion, it looks like this.
Changing a database field, triaging an email, you know, looking like this dude. Okay. Let's get him on mini max as soon as possible. Summarizing meeting notes. These are all things that we actually don't wanna pay 40% more when GPT six comes out or whatever is next. We wanna either be using open weight or we wanna be RL ing. But for the frontier tasks, we need to be giving that frontier to our customers, data analysis, deep research.
The point is that our customers need us to choose for them. Otherwise, this will be their headlines. Okay? And your product is inaccessible. The problem again is that this isn't how our providers are working. Right now, they're incentivized basically with two paths. Again, won't name names. Use your critical reasoning or your favorite model to figure it out.
Either you are the best reasoning model, you are the example in all of these tweets. Right? You are the example of great, and no one really questions your price because you're the first one that passed whatever benchmark you said you passed. The second is you're slightly worse, but that's okay. You just need to be about thirty seconds per million tokens cheaper, and you have the rest of the market to you.
What? What about everyone else? Right? We're in a duopoly right now where things aren't priced for their needs. Complex tasks, though, I think are appropriate. I wanna keep highlighting this. We see a bifurcation at least in notion usage, and I want you all to be investigating and thinking about. There is a place for complex, expensive reasoning models.
The trap is putting everything there. And the best frontier model is changing fast. This is an older slide you saw brand new from artificial analysis. Okay? But the point is that it's changing constantly. We're changing our default model for our customers probably every three to four weeks. And it's a lot of work and we have very advanced evals and teams to do it, but you shouldn't need that.
I mean, are a lot of benchmarks you can be using. If you're switching all the time, you can't be locking your product into a particular provider. Because if you are, the second you go from January to February, you're giving your users a worse experience. Okay? It's not worth the discount. I don't think so. I think you're putting a big risk on your company and you're and you're willing to fall behind.
And depending on the business, for that kind of saturated knowledge work land, that's fine. But if you wanna be offering frontier, you shouldn't be doing this. Optionality is leverage. You should be ready to walk at all times for a durable business. You should not be locking yourself into a single provider and you should be comfortable knowing what the landscape of models are so that you're ready to maintain your margins and maintain your business in a way that you're confident brings your users the best experience.
We do this with our auto model. So today, we have a model at Notion that's called Auto. This is our model picker. It's due for a refresh. The number of models is getting long. But we let users choose because sometimes that latency, quality, cost, trade off is not something you can assume for your customers. You'd be surprised how many people really want to spend a ton of money on email triage and how many people really don't care about how accurate the research tasks are.
I mean, it's a phenomenal user research exercise that we could have a whole other conversation on. And so this is how we do it. We choose for our customers the majority of the time, but 75% of the time, they they stick there. But that 25% is valid, and we give them opportunities to leave. With that auto model, we have the opportunity to actually think about the task at hand and give them the most quality that they need at the best price and latency.
Okay? So if you're setting up an email triage agent, it's very unlikely that auto will be opus because you're gonna get your first usage based pricing bill and say, no way, Notion, I'm out of here. Okay? This is the playbook. Okay? All of you that love the taking your phone out with slides, this is the time. This will be online, but I I get it.
Okay. Build for multi model. Be ready to switch. Understand what the model providers are. Evaluate on value, not tokens. What does that mean? This is an example. We tweeted this a couple months ago on web search providers. It might be that a certain web search provider is cheaper per token or per request. How many requests is your agent doing?
How accurate are the results? You should be looking at the entire task when evaluating. You should not be thinking about a particular API call. That's where they get you. Okay? Think about your use case. You are the expert on quality for your use case. Switch fast, switch often, give them something back.
Do you know why these evolves were awesome? I love competitive dynamics in a market. I love it. If we say parallel, we chose you, but it's close. Stay up there. Everyone else on the list, here's exactly why we didn't choose you. Please fix it. A rising tide lives all ships. Position yourself in a way where you're getting what you want from your providers. Forego discounts for optionality.
I think we talked about that already. And there's a third option, which I'm sensing is the theme of the day, which is open weight. That moderate task, we're very heavily considering now. And we have live in production a lot of open weight traffic as well as reinforcement learned models. They're strong enough to handle workloads. I don't think that they are going to be in the upper right quadrant of capability soon, but the gap is closing.
And it also gives you negotiation leverage, complete financial independence, choose the inference provider of your choice. Right? I would say 2.6 was the most was the first time this really happened for us. There's been a lot since then. But 2.6 was the first moment where we actually saw it compare to GPT 5.2 in quality.
So here's that eval. We saw we see scores, and these are on Notion specific tasks. Again, you own your product. You get to decide what works for your product. For Notion specific tasks that we thought the auto model needed to do for a subset of them. We score just fine.
Kimmy two six is great. And errors are important, by the way. Errors are what you end up paying for. Just as a side note, you're still paying token rates on errors and retries. Think about that in another presentation. But it's not just the token costs. We look at the average number of tokens. There's some crazy things on this, like Opus four seven, $4.06 versus Sonnet. Look at the token consumption.
Forget the price on Opus versus Sonnet. Look at the token consumption. Right? This is why understanding your whole trajectory is really important for understanding your model. I love Philip Keeley. He has a great book called Inference Engineering. Even if you're not an inference nerd, it's good to read. We don't need to be frontier with open weight.
It's just a gap in time. Right? There's product that we served six months ago that our customers love, and they don't want to randomly pay for more tomorrow. Right? We're trusting that open weight is closing that gap, but we need to have the right evals and stay on top of it to know when it's there. And that's why investing in your evals are important.
How do you win these negotiations? Well, the first thing that you need to do to be ready is think about your architecture and not just your model choice. For us, we've noticed that harness engineering and architecture decisions can account for about three x the change in price as model selection. Sometimes that means using a native harness like the Codex and the Cloud APIs. Sometimes it means using open source like Py.
And sometimes it means creating your own harness because there are particular capabilities you want and you understand the implications of how prompt caching, etcetera, might affect how you work. Okay. But what if you don't need an LLM at all? I know we said that this is welcome to Token Town, but we're leaving. Okay? We're now departing Token Town.
And I wanna take us back to an old world, an old world where we used CPUs to do our jobs. This is my nana banana, I think, generated image of an engineer trying to turn a CSV into a PDF and post it on a webhook. Why would he ever need an LLM to do that?
Why would you want that repeated task constantly using reasoning tokens to navigate MCPs? Okay. Well, because Frontier Labs want you to token max. They want you to do it in a way that they control their capacity. It's its own economic issue. But they want you to token max. Your users don't. Your users want you to outcome max.
That's why we launched our developer platform and something called Workers. At Notion, we believe that many tasks do better work on CPUs. Determinism is a valuable thing. State machines existed. Right? The pendulum has gone so far because it can, but that's not durable software, and it's not a formal software. So we've partnered with Vercel on the capabilities to launch what we call workers, where you can actually call on internal APIs, computer sandboxes, host small, small code as an action that your LLM calls.
We've seen this decrease token cost by up to 80% for some of our customers on repeated tasks. It's wild out there. I mean, I feel like if I go on a week vacation and I do a good job not looking at Twitter, it's like I hibernated for two years. Okay. I get it. It's exhausting. Take breaks.
Take care of yourself. But it's crazy. And the market is so young. It's so opaque. It's moving so fast, and you are a player in it. I find it very sad that we're all kind of rolling over, but we're all together and our negotiating power is so much stronger if we all advocate for ourselves. And I also think that you owe it to your customers.
I'm really proud. I'm so proud. And I come to work every day because of the way that I think we make AI accessible and valuable for the Fortune 5,000,000. But not everyone uses notion and I want all of you to do that too. I want us to make AI something that isn't a meme. I mean in California, I read a statistic that AI is more unpopular than ice which is our immigration controls.
It's crazy. We don't have the, you know, it's not a popular institution, ice. So that means AI is really unpopular. Why? Because we're not creating valuable work and we're not doing it at the right price and people aren't seeing the value and there's a ton of Twitter memes, but there's real work to be done. And that's our obligation.
Right? I think it's a powerful technology and I think it's our duty to bring it to our customers in a way that's responsible, durable, and appropriate. So I'm on Twitter a lot as you saw. Please tweet at me. You can always email me. I'll be here in this hemisphere till the end of the week in Melbourne. Thank you for your time.
Thank you for hosting me. I love Australia, and have a great day.
People
- Philip Keeley
Technologies & Tools
- Auto model
- Claude Code
- Codex
- GPT-5.2
- GPT-6
- Kimi 2.6
- MCP
- Opus 4
- Sonnet
- Workers
Concepts & Methods
- open weight models
- prompt caching
- reinforcement learning
- state machines
- usage based pricing
Organisations & Products
- Anthropic
- Artificial Analysis
- Decagon
- MiniMax
- Notion
- OpenAI
- Vercel
Works
- Inference Engineering
Token pricing is a noisy headline. What matters in production is what you pay per task once outputs get longer, retries creep in, and “good enough” models get deprecated.
This talk breaks down the current dynamics of the frontier model market from the perspective of Notion AI product team on embarking this with Notion custom agents model, and shares a practical playbook for keeping leverage: stay genuinely multi-provider, invest in product value that is hard to undercut, and use open weight models as an escape hatch for everyday workloads that do not need frontier reasoning.














