Designing Inference-Native Systems
Introduction: Designing Inference-Native Systems
The speaker opens by polling the audience about system design and chatbot usage, then introduces the core thesis: designing 'inference native' systems requires starting greenfield rather than retrofitting old apps with chatbots. He frames the talk around how software must evolve to handle probability instead of certainty.
From Certainty to Probability: A 70-Year Software Shift
The speaker contrasts 70 years of deterministic software (e.g., simple age-based rules) with human decision-making, which involves observing, updating beliefs, and adapting. He argues the next era of software must embrace probability and context the way humans naturally do.
Standing on the Shoulders of HCI Pioneers
The speaker honors HCI visionaries like Vannevar Bush, Licklider, Douglas Engelbart, and Alan Kay, describing their dreams of memory expansion, human-machine partnership, and personal intelligence devices. He notes that despite decades of progress, true symbiosis between humans and computers still feels distant, though we're now closer than ever.
New Math, New Interfaces: The Shift to Probabilistic Design
The speaker discusses how the mathematical foundations of software are shifting from boolean logic and set theory to probability, enabling new interface paradigms. He explains how design is moving from hand-drawn forms and pixel-perfect UIs (like Figma) toward intent-based systems, using healthcare examples like reducing knee pain instead of navigating multiple screens.
Defining Inference-Native Systems and the iCARV Framework
The speaker defines inference-native systems as ones that update beliefs, decide, and act on behalf of users, contrasting this with today's rigid input-process-output and CRUD-based systems. He introduces 'iCARV' (Intent, Context, Act, Reconcile, Verify) as a proposed new primitive to replace CRUD in this new era.
Three C's of Inference-Native Systems: Composability, Cognition, Creativity
The speaker outlines three key design principles: composability (interfaces generated dynamically rather than hand-drawn), cognition (embedding intelligence so users don't have to think, illustrated by a surgical template tool), and creativity (avoiding simply bolting AI features onto old paradigms, using Uber's GPS reinvention as a model for genuine innovation).
Reimagining User Stories Through Observability and Intent
The speaker proposes redefining 'user stories' beyond Jira tickets to capture real human goals and intentions. He describes using OpenTelemetry and observability stacks not just for system monitoring but to reverse-engineer user intent, aiming to reduce the latency between user actions and truly understanding what they want.
Case Study: Onset Health and Designing for Surgeons
The speaker shares a real-world example from his company, Onset Health, describing how they build inference-native systems for surgeons—a demanding user group. The system learns individual preferences and billing styles over time, aiming to reduce reliance on large administrative staff and let surgeons focus on patient care.
Rethinking Design Sprints for the Inference Era
The speaker argues that traditional design methodologies like Google Design Sprints and wireframing need to be rewritten for inference-native systems. He emphasizes that future design must center on users' evolving goals, values, and real-time feedback loops, contrasting this with healthcare software's notoriously slow, multi-year feedback cycles.
Closing Vision: Toward True Human-Machine Symbiosis
In closing, the speaker calls for genuine innovation beyond superficial AI chatbot integrations, urging designers to build composable cognition layers, shared and personal memory systems, and adaptive decision engines. He ends by inviting the audience to connect and continue exploring this evolving space together.
Who designs systems for a living? Put your hands up. Alright. Quite a few. Who's designing inference native systems? Now confess, who's put a chatbot on their apps? There's a few of you. TLDR, if you want to go back to your Codex or Claude sessions, don't build a chatbot.
The world's changing, start greenfield, it's really difficult when you've got a Titanic to move, speedboats, get things across faster and design around decisions. That's sort of the TLDR of our thesis and my talk around how do you design inference native systems.
There's a few of you who design systems for living in the room. And for the last seventy years we've built software with, you know, certainty at its core. In the next seventy years, we're gonna have to deal with probability. And if you look at software today, that's how we've built.
If age is over 18, let's go approve and reject them. We don't have a lot of like context, but as humans what we do is we observe, we'll update belief, we'll decide and act, and depending on the situation, you might be able to bend rules and let somebody who's below 18 get the job done.
But that's not how systems work today. But we're at a changing arc. Now in the last seventy years, though, we we here stand today standing on the work of legends in the past. Any HCI nerds in the in the house? Couple of them. So, you know, these folks have been thinking about human computer interaction for a long time, whether it was Vannevar Bush with the Memex and memory expansion, all the way to Allen Kay's Dynabook, the iPhones or the smartphones that you hold in your pocket today, is really, you know, some of the work that, you know, these some of these guys have worked in the last seven years.
So, Bush talked about, computer systems with that allow us to expand human memory, and navigate universe's knowledge. We had, Lake Litter talk about just this partnership of humans and machines. But today, let alone partnership, we're debugging minified JavaScript code, and, you know, navigating documentation.
So it truly doesn't feel like a symbiotic relationship, but see if some of these guys have been talking about these topics for a while. We had Douglas Engelbart gave us the mouse, some of us don't like using the mouse, but he conceptualized it, but ultimately their thought was how do you make humans smarter, amplify intelligence.
We had Alan Kay talk about the Dynavoc, and what would happen if personal intelligence were in our pockets, and today we're actually at the glimpse of it. There's a long way to go, but we're we're just getting started and and some of these guys have been conceptualizing these thoughts.
I think they're smiling down from heaven looking at all the possibilities today. So the abstractions have changed. You know, thanks to the MT for the intro, and, you know, I've kinda gotten started around, I was a toddler around the web era, and then early career mobile, but it's really exciting today, I get to catch this abstraction around Inferences.
Because Inferences help us, you know, follow, chase our decisions, whether what we want for lunch to where we wanna go, and systems today hopefully get us to this world of enabling better life decisions. And the Math has changed, and, you know, self confessed Math Nerd, that's what I studied.
Most systems have been based on pretty basic math or on boolean logic, a bit of set theory, state machines. The math is completely I mean, the systems of today are based on like new math primitives around probability. So certainty is moving away to probability with just some of these things, and that's enabled us to think about new interfaces.
We've spent a lifetime designing forms, and web pages and data structures and it's all and historically it's been, let's hand draw them, then came Figma, let's get pixel perfect UI, that's we we just shoved software down users whether they were willing or not.
But now, we have the option to design systems through intent. And that intent could be, I work in healthcare, we're building a healthcare startup. So the intents like, I couldn't care less to click through six screens, I just want my knee pain relieved. I don't wanna go play golf over the weekend and I want clear vision. Get me something that just helps me play golf with my friends. So that's kind of the evolution of new systems and our thought around, how do you design new systems for the inference era. And what's infer?
Infer is, you know, there's rooted on probabilistic thinking, your beliefs get updated, you decide, and then you act on things. And an inference state of system should be able to do that for you. It updates your updates its belief, ideally autonomously, it decides on your behalf, and then it just gets it done.
It acts. Whereas the systems we experience today are all like, in per process output. The script is being flipped. CRUD, who's who's built CRUD apps? Almost everyone on the house, that's that's the thing we love.
Let's go build CRUD apps, let's, you know, the database is being primitive, so I this is not popular, then I think about, like, what's the new primitive? CRUD's not gonna cut it, like how how difficult is it to put together a CRUD app with a Clot or a Codec session? So we've been thinking about iCarve being the new CRUD, which is you start off with the intent, the context is given, extracted, found somewhere, it acts on your behalf, it reconciles and verifies.
Now, that's a quite quite a bit much. I don't think I can cover it on eighteen minutes and we're learning real time as we build our products. I'll touch a little bit little bit more about the intent side of today, but follow the journey as we think about the evolution of inference native systems, but the way we've been thinking about it from this iCarve lens, which is how do you go from, that was my intent and you verify, you close the loop by verifying that it got it done to my users liking or the users liking. And the three c's, we'll just leave you out with and there's a lot more to add, it's early days of our thoughts, are what could inference native systems consist of?
Well, one it could be composable. Gone should be the day where we've hand drawn and created these prescriptive user interfaces. That's where the table goes, that's where the columns go. Two, let the user interface generate on the fly. Perhaps your system or application is entirely headless, it doesn't have an interface, it's got logic built into it, but no two users will have the same interface, they could see anything.
So one thing to think about is really this composability, and I a lot of the foundational models have been teasing and experimenting with a couple things now. If you open up Clot or OpenAI, it's not just a text stream. You might find text editor, you might find a a tab table, like there's different kinds of interfaces that are quite composable and they get generated on the fly.
So, as you think about designing infras native systems, they get composable. Cognition, there's a popular book called Don't Make Me a Thing, and I think that applies today more than ever, which is, let's not make the users think, let them do what they wish, embed cognition in your tool. So this is our interface where doctors get to design templates for surgeries, as opposed to like long forms and fields, and let's put this structure, let's put that, we've just given them a scratch pad. Chuck whatever text you want in, we'll go structure it for you. Historically, this would have required two people, two days, give me your text dump, give me your brain dump, I'll go organize it, format it, design it, send you back the PDF.
Today this happens on the run time. Dump your thoughts, out comes this beautiful PDF at the end. So, embed cognition into your applications. And the third being creativity. When the mobile phone first came out, there's this app called I'm rich. All it did was glow this red light.
And there was a whole bunch of people that were also trying to shove websites on a native mobile applications. And that's where we are in this like new era of inference. We're all just trying to shove like glowing red light into our software as opposed to rethinking some of these primitives because think what Uber did. Uber really took advantage of GPS coordinates and reimagined the experience. Very few people use Uber on on their web.
They use it on their mobile application. So it's a that's where I went and said, start greenfield, because if you're trying to shove a chatbot into your existing application, you're not gonna make it. This is an opportunity to really rethink what the possibilities are. So iCarve's I, intent kinda linking to stories.
Any kind of product manager, like, heard of user stories, Jira, everyday language? Let's change that narrative. I mean, today we've gone from like countless interviews, founders influencing, random thoughts to let's structure everything into user stories. But if you peel back the layer users like stories are what moves us humans. The decisions we make, the life choices we want, the intents we have.
So let's think what could like the user stories of the future look like. So this might be a bit of an overview of how we've been thinking about building, which is the new stories don't necessarily map back to like records tables and rows. They're your goals, they're your intentions, your decisions, your narratives, like what does the user really want?
And how do you capture them? And our sort of backdoor approach of trying to create this narrative stack is taking advantage of OpenTelemetry. Now, OpenTelemetry, some of you in the room probably use it, for others it's it's just a it's a it's a stack that allows us to look at what did the user do, how's the system performing, whole bunch of logs, things that people who are conceptualizing and designing systems are like, oh, this is an afterthought, let's just like, it's it's an afterthought to have to have stable systems and looking at usage metrics. But this stack can be flipped around to actually aid in developing these new inference native systems where how do you go from observability to intent?
Because the real opportunity set is really uncovering the stories through your observability stack to try to figure out what the user was trying to do. We use a combination of some of these, we're talking to the click house hosts in the the exhibition, so we wanna start working towards that, but it's we're trying to weave the stories together to figure out what is our user really want.
And for us going from user just clicked a button to like, their intention was I just wanted to schedule a surgery or submit a payment claim and ultimately just relieve that nasty knee pain. So for us it's looking at how do we take that but create these narratives that allow us to get to a world where we're not shoving user stories in Jira boards, we're looking at user stories and converting in run time to the intent the user had.
So the latency for going from here's a thought bubble, the user or there's a cluster of users who could want this feature all the way to like they wanted this, we've got the data, we've got the tool calls, we've got the integrations, let's get them to what they want in the fastest possible way.
That might happen inside your application, that might be somewhere else and you've helped orchestrate it. So there's a real opportunity to just say, how do you lower the latency across the stack? And that's what's gonna give you, like, inference native systems. And we've been doing it, so we're Onset Health, we we're building inference native systems for a really difficult user group, surgeons.
They've they've spent a lot of time trying to figure out like the human body, so they're a really difficult bunch to get to adopt something new because they've got their set ways, but we've been thinking, alright, very difficult group to work with, but they want the best of tools that feel invisible. So our system listens to surgeons, updates its belief over time, everybody's got a memory index, know Rob prefers working this way, this is a style, suggest some of his billing stuff, ultimately, you know, extracts what surgery they've done, the system's learning over time and the byproduct is Rob just wants the patient relieved, get them out of knee pain. Today Rob does this with the support of six human admin staff. So it's a good good opportunity for us to be like, how do we just make Rob's life a lot easier by helping him harness the best of, like, the inference stack?
And the kind of primitive is around that, is like how do you even start thinking around designing these systems and, you know, we've spent a career thinking through the Google design script and those like, let's sketch out eight interfaces on pieces of paper and try user studies of let's what happens when you click through here, that screens leads to there, you've got forms and at the end you get back to your crude apps.
That all all that literature we feel like has to be re rewritten, re rethought in. Just lots of companies building at a really rapid clip, but I think, you know, now the element that's sort of under hyped is taking that step back and just thinking through what what what are the new design sprints of the future? How do you build systems that people will, you know, use and will enhance their lives and they're not gonna happen through designing new forms and curd apps.
It's gonna be around the decisions they wanna take, align to their preferences and beliefs and values, what are their goals? Their goals change over time and the feedback loops have to be as real time as possible. We've all heard of software legacy or not, where the feedback loops can pass In healthcare, it's not uncommon to have software that's feedback loop is six years.
By the time you've expressed your frustration to software being used, couple of years, six years to get it to production. That world cannot exist. So we really think the kind of design sprint process really ought to be reimagined. So the next generation of software, for those who are interested, should be reimagined around inference.
Everybody's trying to think about AI, the boards are pushing an AI strategy, people are using their cloud skills and shoving chatbots into wherever they see fit, that's not really the kind of uber equivalent of the inference era. There's a real opportunity to rethink what this is gonna be. And some of the tenants of these inference native systems are gonna be, how do you create these like composable cognition layers in your systems? How do you have shared memory around your application but your users will have their memory as well vis a vis the use case they have, it could be healthcare, finance, meal delivery, whatever it may be, but like everybody's got life choices.
So the n is equal to one is gonna matter extensively for the systems you develop. What are gonna be your decision engines of the the thought process of how your systems evolve. So it cannot be, you know, boolean based logic systems that are just like, let's just just give them these rules because your systems are gonna have to adapt and the decision engines have to enable that for users to experience that.
And ultimately ultimately that's get that gets us to a world of the true human machine symbiosis. Thank you very much for engaging in this conversation. That's me. Feel free to connect on LinkedIn as we try to figure things out.
People
- Alan Kay
- Douglas Engelbart
- J.C.R. Licklider
- Vannevar Bush
Technologies & Tools
- Claude
- ClickHouse
- Codex
- Figma
- GPS
- iPhone
- Jira
- OpenTelemetry
Concepts & Methods
- Boolean Logic
- CRUD
- Google Design Sprint
- HCI
- iCarve
Organisations & Products
- I'm Rich
- Onset Health
- OpenAI
- Uber
Works
- Don't Make Me Think
- Dynabook
- Memex
For a long time, the world has run on systems built on logic. You put something in, follow a set of rules, and you get an output. Now we have systems that can run on inference: systems that can update belief, decide, and act. That changes how we should think about building systems. We don’t need to keep forcing everything into rigid workflows. We can start designing systems that are built around inference from the onset. This talk is a thought process on designing these systems, drawing from principles in human-computer interaction, mathematics, and software design.














