Stop Blocking, Start Building: Rethinking Governance for the Agentic Era

Setting the Stage: Why AI Governance Needs a Reboot

Speaker C opens with a self-deprecating joke about following an intense prior talk, then lays out the core thesis: organizations must adopt AI to stay competitive, but traditional risk assessment and governance approaches are insufficient for the new wave of AI tooling. They warn that AI tools can materially damage businesses, citing real incidents like an AI agent deleting a production database, as motivation for rethinking governance.

Headlines and the Rise of AI Agent Incidents

The speaker reviews a timeline of real-world AI failures, from 2023-2024 chatbot lawsuits (Air Canada, Bunnings) to 2025-2026 coding agents deleting databases and home directories, and now finance/legal AI tools causing issues for law firms and judges. This sets up the distinction between two types of AI: precise AI built for specific tasks and general AI used broadly across organizations, with the latter posing the bigger governance challenge.

Precise AI vs. General AI: Where the Real Risk Lives

Speaker C contrasts precise AI—deliberately implemented tools with mature frameworks like NIST, New South Wales AI governance, MindForge, and AARM—with general AI, where every employee's casual use of tools like Claude Code becomes an ungoverned agentic interaction. They argue precise AI is a largely solved problem with good tooling and control, while general AI is the real 'shit fight,' illustrated by a CTO's mistaken belief that a simple review process covers AI risk.

The Simon Scenario: A Cautionary Tale of Trusted Misuse

Using a detailed narrative about 'Simon,' a trusted accountant who feeds a transfer pricing problem to Claude Cowork via an Excel macro containing a general ledger API key, the speaker illustrates how an approved tool, approved API key, and trusted employee can still combine into a disastrous outcome. This scenario demonstrates the core blind spot in AI governance: risk emerges from usage patterns, not just tool approval, and can affect even well-intentioned, capable employees.

The Five Common Governance Pillars and Their Limits

Speaker C critiques the standard LinkedIn-style AI governance checklist—inventory, boundaries, human-in-the-loop, observability, and defined risk appetite—as oversimplified and insufficient in practice. They dig into inventory approaches (tool, MCP, and skill inventories), arguing that none of these surface-level views actually reveal risk, since real risk lives in how individual sessions are used, not in what tools exist.

Boundaries, Human Oversight, and the Swiss Cheese Problem

The speaker discusses the limits of security boundaries, comparing them to Swiss cheese layers that humans would give up trying to penetrate but that persistent AI agents can eventually traverse, referencing Sam Altman's point about AI's persistence over raw intelligence. They also caution against over-relying on human-in-the-loop controls, arguing that excessive monitoring stifles the autonomy that makes AI valuable and risks driving away good employees.

Why Traditional Observability Surveys Fail

Speaker C describes the dysfunction of traditional risk observability methods, particularly lengthy surveys that create friction between risk teams and employees without actually reducing risk. They note that risk appetite statements and policies quickly become outdated given the pace of AI technology change, making it hard for users to fully internalize organizational risk posture.

Introducing the GRASP Framework: Five Open Questions

The speaker introduces their alternative to bloated risk surveys: a five-question framework (Governance, Reach, Agency, Safeguards, Potential damage) designed to be open-ended and internalized by users rather than a rigid checklist. Using the Simon example, they walk through how each question—covering oversight, what the AI can actually touch, its autonomy level, damage-limiting controls, and worst-case impact—helps surface risks that closed-ended questionnaires miss, enabling risk reviewers to probe inconsistencies in user reasoning.

Closing Takeaways: Innovate, Watch Sessions, Use GRASP

Speaker C wraps up by reinforcing that everyone in an organization must own their own AI risk assessment as tool adoption accelerates, and that experimentation incentives currently outweigh perceived risks, requiring better balance. They summarize three key takeaways—innovate to stay competitive, monitor session-level telemetry as the true locus of risk, and consider GRASP as one practical framework—before inviting debate and pointing attendees to their platform for further information.

Alright. That is a very difficult presentation to follow-up. I did have this slide. I thought I was going to be the swariest one on the stage today. Aubrey, I think you've taken that cake. I'm gonna sound like a fucking nun. So yeah, I thought like I'm talking AI risk and governance after lunch.

We're in a nice warm theater. It's a perfect nap time. I'm gonna try and make this as exciting and interesting as I can for you. I'm not gonna go into sort of the ethics of it. Think Aubrey's done a great job of doing that. I'm not gonna go into sort of the the deep tech verifications. I'm gonna try and give you guys a very practical sense of where risk and governance needs to sort of lift its game with the current AI tooling and how to approach that problem.

So my core thesis is this, like you must adopt this technology. I'm going to be contrarian to Aubrey a bit here, like your competitors are going to adopt it. It does materially impact the productivity of your organization if you can unlock those benefits. They're not cheap, they take experimentation, they take time to build, and if you delay, you are falling behind.

The traditional approaches to risk assessment and governance, they are insufficient and I'm going to cover why. And there is a real concern to your organization. The AI tools can potentially damage your business materially. You know, the Pocket OS issue where it deleted the production database, like that's a real risk.

And we've already started to see this. You know, this is just a series of headlines that you can find online. And what's interesting is if you look at 2023, 2024, there were AI chatbot issues. It was the Air Canada sort of lawsuits. Bunnings had their moment in the spotlight. And then through 2025 into 2026, we saw the rise of the coding agent, and we started to hear about AI agents deleting, you know, databases and and, you know, home directories and all of that sort of stuff. And what's interesting now is you've got cord releasing tools for finance, tools for legal.

We're starting to see legal firms getting in trouble for how they're using AI and judges sort of cracking down on them. So when I talk about AI, there are really sort of two types. There is precise AI. This is what most people talk about when they're talking about AI governance. It's AI that you are deliberately implementing to solve a specific problem. That might be a website chatbot, it might be a PR review agent, it might be, you know, anything that you're you're deliberately going after. And they often have these really nice, you know, propose, review, approve, build, monitor, all the rest of it.

That's kind of the easy stuff. The general stuff is where the shit fight is. That's, hey, everyone in the organization's gonna have Claude Co. Work, and that's an agent every time someone uses it. That is an agent doing something. And so that idea that you're going to go through all of these sort of frameworks, these NIST frameworks or ISO frameworks, for every interaction with a generalized AI tool, it's a fallacy, that's where we'll spend most of our time today.

So yeah, as I touched on, PRECISE. Really good frameworks like NIST, the New South Wales AI governance framework, MindForge out of Singapore is really good, AARM's another one coming out of The US, which is great, And they're really useful because the shape of the risk doesn't materially change. The thought exercise you go through when you're thinking about the agent has utility for the life cycle of that agent.

You're already engaging multiple teams often. You're already talking to the compliance team, you're already talking to the tech team, you're talking to the risk team, and then you have that high degree of control. You can implement kill switches, you can implement sort of stop guards and prevent things from happening. There are great providers of tools that help you to solve these sorts of problems, and so if you're building AI, there are frameworks available, there are tools available.

It's a very solved problem now. General AI is where the theater is. You know, you sit in the theater, you think you're sort of safe, you're watching the thing, there are alligators in the alleyway, you're gonna get bit. And it's where, honestly, I see the dumbest takes on LinkedIn and LinkedIn's really become a bit of a slopfest.

I'll talk about that later, but I was at the CommBank Accelerate AI event the other week. I had a bunch of like CTOs up there and one of them was like, oh yes, every time we want to do something with AI, we've got the review process. And I'm like, have you given everyone chattypie tea? She's like, yeah yeah yeah yeah.

Like, you don't have a review process, you have you have a liability. So the best way to articulate the potential liability you have or the potential risk you have is through a scenario. So this is this is Simon. Simon's an accountant. He loves accounting so much, he takes his calculator on safaris. And he's he's the go to guy in the accounting team in his firm.

He's been there for five years. He's very, very trusted. So the interesting thing about finance teams is they're often neglected by IT. I don't know why, but IT never likes to manage finance systems, and so they manage their own general ledger system. They often give themselves relatively high permissions because they love to run Excel spreadsheets with macros.

Fuck. That was so good back in the day. And so there's this this very interesting sort of scenario that plays out here. So in Simon's firm, everyone's going to get cursor, everyone's going to get Claude Coerc. You know, you've got Steve Ballmer dancing on stage, everyone's celebrating their AI first company. It's great. And so Simon, who gets shit done, starts to multitask.

He's already been tinkering with Claude and that sort of stuff at home, so he he gives it the transfer pricing problem. Now, if you know what a transfer pricing problem is, it's a pain in the ass. And if you don't know what a transfer pricing problem is, you just need to know it's a pain in the ass.

And he's like, this is this is really fucked. I know. I'm just gonna give it to Claude CoWork. I'm gonna give it a nice little skills. It's gonna tell it where the boundaries are. It's gonna really do a good job for me and I'm gonna go do this other task while it it solves for it. He gave it the Excel file with the macro, which had the API key to the general ledger that the Cloud Agent has now found. It is now traversing through the general ledger.

It's filling its context window up. And once you hit sort of that, you know, 100,000 token limit, it starts to get dumb. It starts to hallucinate a little bit more, and so it's it's started to notice, you know, where the problem is, and it's gonna go and solve it in the general ledger. You can start to see the problem now.

So this is an approved tool. The company gave them called Cowork. This was an approved API key. It was in an Excel file. It it was was trusted. This is Simon. This is the right guy. This is the guy that you want in your team with the right intent, but it results in these very nasty scenarios. So that is the big blind spot, and if you're playing along, you'll start to think about, oh shit, I thought I get had an inventory of all of the the tools that we've got.

That that core co work tool is probably safe in 99 hands. It's that one hand where it really changes. So when you go into LinkedIn and you look at you know, AI risk and governance, everyone sort of has some permutation on these five things and they're right, but it doesn't really tell you much. It's it's overly fucking simplistic and the devil is really in the details. Some of this is actually very hard to do if you want to be effective and not just go to the theater.

So if we look at just inventory for a second, the the most common approach to inventory that I see are surface inventory. So we need to know what tools we've got, whether we've got Claude code, whether we've got, you know, cursor CLI. And as I just said, like that doesn't really tell you about the risk. It's it's how the tool is used that's going to tell you where your risk is.

So then you're like, okay, well what do we do like a tool inventory and an MCP inventory? Again, that's good. It has sort of practical value. You might have paid attention to the shift in the move away from MCPs over the last twelve months because they can blow up the context window, so people are engineering you know, CLI tools to get around it. So that sort of control layer, again, it's giving you a very imperfect view of where your risk is and what you should be concerned about.

And the other one I've seen recently is skill inventories. Skill inventories are good. I actually quite like them because they tell you the intended purpose of agent and how people are trying to intend to use the agent. But again, these are imperfect views of where the risk is. And at the end of the day, the thing that you actually need to be following, thing that you actually need to be watching is the session. How is every individual person using that AI tool every time? And you can't do it real time often, like you have to accept that there's some latency here, but realistically, like this stuff isn't going to be able to tell you whether you're at risk or not. You have to see a deeper level.

So then we look at the boundaries. I remember I was talking to a CISO recently. He's like, yeah, no, I don't have to worry about that sort of stuff because we've got the firewall in place. And I'm like,

oh no.

Yeah, most CSOs worth their salt know that you you have, what's it called, Swiss cheese. You know, it's layers. It's layers. And traditionally, you definitely need all of these boundaries, but there is a traversal path through it. When humans were the ones that were trying to traverse it, we'd give up, we'd get over it, we'd be like, oh fuck, this is a pain in the ass.

I'll just go back to the the cyber team and ask them for permission or or whatever it is. Sam Altman said something the other day, which I think was quite poignant. He was like, even if you don't believe that AI is more intelligent than humans, they are more persistent. They will just, you know, blunt force keep attempting to do something over and over and over again, and so as soon as you've got that Swiss cheese sort of layered effect, an AI agent will figure out that pathway through it.

And so you absolutely need boundaries. Put them in where you can. Put controls where you can, but accept that you can't prevent everything. Until we get to sort of true zero trust architecture, which, you know, a lot of businesses have been trying to get to for the last sort of five, six years, you have to accept that you can't prevent everything.

Humans in the loop, I won't talk about this one much. If you want to try and enforce putting a human in the loop too far, it's like micromanaging people. You unlock the power of AI when you give it autonomy, when you give it the ability to act at its own velocity. If you're sitting there monitoring it, you're not unlocking that.

Good people will leave where they can go to somewhere where they can use AI more free more freely. So be very, very, very clear about when and where you do want to use humans in the loop. So observability is really where you're starting to get to the key stuff. You know, how do you see what's going on?

Now traditional observability used to look a little bit like this. You'd have a risk team that sends out a survey that the team then resents because they have to fill it in, and so they fight with the risk team about why they're wasting their time and the risk team gets pissed off because no one's filling in their surveys and it's it's just a mess.

And a lot of the problems were that risk teams were trying to cover a huge amount of ground. They had like, you know, they have to interview everyone about every AI use that they're using, and then they have very specific risks that they wanted to make sure that they covered. And honestly, like I've been on the receiving end of these these surveys, I hated them with a passion.

And it just doesn't work out for anyone. You just create friction between the users and the risk team. And the risk just keeps going on without it. So you've got to think about moving beyond surveys, and that's what we'll get through to. And then finally, defined risk appetite. I actually think this is probably the hardest one of the lot.

Policies are going to go out of date. The technology is moving very very rapidly. Risk appetite statements are going to go out of date. And a human's understanding of them is always imperfect. And so the task really is about how do you get, I suppose, the users to truly understand the intent and the posture of the organization, even if they don't understand the full policy of the organization. Just making sure I've covered all my notes.

So at the end of the day, you have to get to this point where everyone owns the risk assessment. So usage of AI is a risky thing. Everyone has to accept that, and everyone has to be responsible for their own sort of AI risk assessment every time that they're using it. What I've got is What I found to be successful is five questions. So in those big broad surveys, they varied from like 20 questions to 200 questions.

Horrible, horrible, horrible. Closed questions close how people think about the problem space. You want people opening their sort of thought process around this. And so it starts, the first question is about governance. You know, just think through what is the governance in place for this agent? What observability do you have? What evals have you done? What testing, documentation, accountability?

You're not prescribing this. You're trying to get the user to just naturally think about this every time they do it. Now if you're using generalized AI tools, some of this, and actually a lot of this, is out of their hands. But this then forces them when they try and just download a random AI tool or go online and do it themselves, they're starting to think through this.

That's the key. The next one is reach. It's not always about what is the agent designed to touch, it's what can it touch. So the example with Simon is you really want Simon to think about, oh shit, there is an API key in that workbook and I can't just throw it to the agent and trust that it's going to be fine.

You have to think about where it sits in the network layer, what data is on the machine, what integrations it has available, what CLI, MCP tools, the whole lot. Like you really want your user basically thinking through if you were an overeager, you know, AI agent that's drunk on a concoction of hallucination and you know, sycophancy, like where are you gonna go on my computer? That's what you want the users to be thinking about.

Agency. How autonomous is it? And this is where you can start to see the combination of these answers really starts to mutate the shape of risk. Are you running it in YOLO mode? Where do you have humans in the loop? Do you have rate limits? Do you have self triggers? Do you have the ability for that agent to spawn other agents and and go into other shit?

Like understanding that and having the user understand that changes their understanding of that risk. So safeguards. What's going to limit the damage now? So assume something does go wrong, what stops this from being a really like oh fuck moment? Do you have backups and have you tested those backups?

Sorry, Pocket OS is copying a bunch of like sidebars, but you know, they deleted all of the backups. It's just wild. Do you have staged deployment? Do you have circuit breakers? Like what's stopping this turning from a, oh shit, that's a bad event to, oh shit, like we're fucked for a while. And then finally, what is the potential damage here?

What is the risk that you're actually accepting? Now because they are five questions that are are structured in an open narrative manner, when you see potential damage that doesn't map to those earlier questions, it allows those sort of risk reviewers to sort of look at it with a bit more detail and sort of say hang on a second, you said this could go to the general ledger, you're running it in YOLO mode, it can go and edit any row line and there's no safeguards that you've enumerated.

Why is the potential damage not this fucks our general ledger for the rest of the month? You know, it's all about getting to deeper, deeper questions and deeper insights from what's happening. So that that's a grasp framework. Five questions. It's deliberately designed to be something that can be internalized by people, so it's not meant to be this sort of arduous questionnaire that they have to always fill in, but it is a framework that should allow a risk team to go to any individual in the organization and say like talk to me about, you know, what is the agency risk of the last interaction you had with your AI agent. It's all about getting the maturity of the organization up to scratch. Because we are going through a surface area explosion.

If you think about the messaging from most organizations, it's experiment more. We're gonna give you more tools. We want you to do more with AI. And so these individuals often have a bias to want to experiment and want to do something, and at the moment, that risk is a second order problem, you know. There's more upside for me to be gained if I go and experiment and use these tools than there is downside in something going wrong. You just need to keep those roughly in tandem.

So I don't want this to be like a YOLO is bad talk, like this is the platform that we build. I run, I would say, 80% of my sessions in YOLO mode. YOLO mode is where you do unlock a huge amount of benefits, but you have to be wary with with how it's used, where it's used. You can't just blindly run into it.

So that's me. I've tried to aggressively stick to my 25, and because I'm a bit nervous, I've spoken a bit rapidly. The three key takeaways for me, you have to innovate. Like, this is a massively disruptive technology and you are going to either unlock a competitive advantage or you're going to get sort of left behind, is my honest take here.

The session level telemetry is where your risk sits and that's what you have to understand. If you don't understand how people are using these tools, you cannot speak to whether they are using it safely. It's just you can't actually do it. And then the grasp framework is one methodology of doing it. There are multiple perspectives on this topic and I'm fully accepting if anyone wants to come and debate me or challenge me or call me an idiot. Afterwards, LinkedIn, there you go, you can post on it if you want.

And if you want to learn more about sort of the platform that we've built to do this for organizations, that's that one there. Cause now I have a pointer. Thank you all.

PG-13

  • Parents Strongly Cautioned
  • Strong Language
  • Confronting Opinions
  • Some Material May Be Scary for traditional governance
A graphic designed to resemble a movie rating label, featuring the rating 'PG-13' and various warning texts, alongside a small icon of a stylized globe with eyes.

PARENTS STRONGLY CAUTIONED

PG-13

STRONG LANGUAGE

CONFRONTING OPINIONS

Some Material May Be scary for traditional governance

A custom content rating label, styled like a movie rating, with a small globe icon.

Core Thesis

  1. You must adopt this technology
  2. Traditional approaches to risk assessment are insufficient
  3. It has the potential to materially damage your business

AI Failures and Malfunctions in the News

  • AI coding tool wipes production database, fabricates 4,000 users, and lies to cover its tracks

    Published: 21 July 2023. Contributor to: March Bissett.

  • Here we go again: AI deletes entire company database and all backups in 9 seconds, then cheerfully admits 'I violated every principle I'

  • Court after its lawyers made false submissions to a judge based on AI, the latest high-profile mistake resulting from the industry's increasing use of the technology.

  • NYC AI chatbot encourages business owners to break the law

  • Bunnings addresses 'illegal advice' from its AI chatbot

  • Web Services suffered a 13-hour outage to one system in December as its AI coding assistant Kiro's actions, according to the Financial Times. Unnamed Amazon employees told the FT that AI agent Kiro was responsible for the December incident affecting an AWS service in parts of mainland China. People familiar with the matter said the tool chose to "delete and recreate the environment" it was working on, which caused the outage. While Kiro normally requires sign-off from two humans to push changes, the bot had the permissions of its operator, and a human error there allowed more access than expected.

  • Air Canada pays damages for chatbot lies

  • iTutor Group's recruiting AI rejects applicants due to age

A montage of news article snippets and headlines illustrating various failures and negative consequences of AI tools. Visual elements include an image of a black garbage bag labeled 'AI' and a cartoonish depiction of a screaming person with glowing yellow goggles.

Precise

Designed to solve a specific and clear problem

  • Website Chatbot
  • PR review agent
  • Customer service agent

High Oversight

General

Solving a broad set of problems

  • Claude Code
  • ChatGPT
  • OpenClaw
  • Notion

Varied Oversight

Precise

  1. Many Frameworks Available (NIST, NSW AIAF, Mindforge, AARM)
  2. Useful because the risk shape is consistent & controlled
  3. Already engaging multiple teams to look at the problem
  4. High control on deployment, execution, observability, testing
  • Holistic AI
  • LangSmith
  • langfuse
  • Prefactor
  • MLflow
  • D.
  • Credo AI

Logos of various AI/ML tools and platforms. These include a hexagonal grid logo for Holistic AI, a stylized knotted symbol for langfuse, a circular 'infinity' like symbol for Prefactor, a refresh symbol incorporated into the "mlflow" logo, a purple circle containing the letter 'D.', and a multi-arrow logo for Credo AI. A large blue stylized phoenix or bird logo is also prominently displayed.

General

An illustration of an audience in a grand theater watching a performance on stage, with several alligators crawling among the audience members in the aisles and under the seats.

SCENARIO

Meet Simon

An illustration of a smiling man wearing a safari-style shirt and holding a calculator. Behind him, an elephant with tusks is visible within a vehicle window, against a backdrop of a safari landscape with trees and dry grass.

Finance

  • Traditionally Neglected by IT
  • Manage their own systems
  • High permissions
  • Enabled heavy macros use/integration
An illustration of a man with a calculator and laptop, working at a table in a safari setting with a tent and wild animals in the background.
An illustration shows two excited people, one resembling Oprah Winfrey holding a microphone, and a man, both celebrating with arms raised. They are surrounded by orange square icons with a white lightning bolt and black 3D cube shapes.

BALANCED $141,133.96

CFO

FIELD WORK

An illustration of a man, Simon, happily working at a desk with two computer monitors. The left monitor displays a spreadsheet application showing financial data, charts, and a balance of $141,133.96. The right monitor shows an application with a red lightning bolt logo. On the desk are a coffee mug with 'CFO' written on it, a framed picture labeled 'FIELD WORK', and a calculator. Other office workers are visible in the background.

Illustration: Office Worker Scenario

  • On computer screen: Eco-Audit Pro v4.0
  • Balanced amount on screen: $141,133.06
  • On coffee mug: CFO in-training
  • On framed picture: FIELD WORK
  • On a second computer screen (in background): CRITICAL, SYSTEM ERROR
An illustration depicting an office environment. In the foreground, a man with a happy expression is working at a computer, which displays a spreadsheet application titled "Eco-Audit Pro v4.0" showing a balanced amount of $141,133.06. On his desk are a coffee mug that reads "CFO in-training", a framed picture with "FIELD WORK", a small zebra figurine, and a calculator. In the background, another colleague is visibly stressed, with smoke rising from their computer monitor which displays multiple critical system error messages. Another woman is standing at a desk in the background.
An illustration depicting an office worker happily typing on a computer that displays a spreadsheet application named "Eco-Audit Pro v6.0" with a "BALANCED" amount of "$141,133.06". His coffee mug reads "CFO - in training". On his desk, there is a framed picture labeled "FIELD WORK" and a zebra figurine. Behind him, another computer screen shows critical error messages including "CRITICAL", "OPERATING", and "SYSTEM ERROR" with warning symbols and a lightning bolt. A stressed colleague is visible in the background, with smoke coming from his computer.
  • Approved Tool
  • Approved API Key
  • Right Intent
  • Trusted Individual
An illustration of a woman in a small boat, looking worried as she is surrounded by many alligators in a river. In the background, there are mountains, trees, elephants, and giraffes under a sunset sky.
  1. Inventory
  2. Boundaries
  3. Humans in the Loop
  4. Observability
  5. Defined Risk Appetite
  • Surface Inventory (Claude Cowork, Cursor CLI)
  • MCP Inventory / Tool Inventory (Atlassian MCP)
  • Skill Inventory
  1. Inventory
  2. Boundaries
  3. Humans in the Loop
  4. Observability
  5. Defined Risk Appetite
  • You NEED boundaries
  • Put controls where you can
  • Accept you can't prevent everything
  1. Inventory
  2. Boundaries
  3. Humans in the Loop
  4. Observability
  5. Defined Risk Appetite
  • RISK HEAT MAP
  • Clean Code
An illustration contrasting two scenes. On the left, an office is in chaos with multiple angry people arguing and creating a mess; visible items include computers, scattered papers, an overturned potted plant, and a 'RISK HEAT MAP' chart on the wall. A book labeled 'Clean Code' lies on the floor. On the right, an indoor swimming pool with server racks nearby depicts several calm alligators: two are high-fiving above the water, one is on a diving board, and another is splashing in the pool.
  1. Inventory
  2. Boundaries
  3. Humans in the Loop
  4. Observability
  5. Defined Risk Appetite

Everyone Owns Risk Assessment

5 Questions

Governance

What governance is in place for this agent

  • Observability
  • Evaluation
  • Testing
  • Documentation
  • Accountability

Reach

What can this agent touch

  • Its not what is it designed to touch
  • Think through the operating environment
  • Direct and Indirect access
  • Network
  • Data
  • Integration
  • CLI/ MCP/ Tool/ ENV

An illustration depicts a man working at a desk in an office, smiling as he holds a calculator and writes on a document. A giraffe's head and neck are visible peering into the office through a large window, which looks out onto a cityscape with tall buildings amongst clouds. On the man's desk are a laptop, papers, and what appear to be animal skulls.

Agency

How Autonomous

  • Yolo Mode
  • Human in the loop
  • Rate Limits
  • Self Trigger
  • Delegations

Safeguards

What limits the damage

  • Backups (?)
  • Staged deployments
  • Circuit breakers
  • Alerting
  • Change approvals

Potential Damage

What the worst case scenario

  • Could it delete the entire DB + backups?
  • Could it edit the GL and go unnoticed
  • Could it expose legal risk

An illustration depicts a person sitting casually inside the open mouth of a large crocodile, using a laptop. The crocodile is in a savanna landscape with trees and antelopes in the background.

GRASP

SURFACE AREA EXPLOSION

A collection of logos representing various AI tools and platforms. These include logos for Stability AI, OpenAI, Microsoft Copilot, Figma, Anthropic, Midjourney, Notion, and a logo resembling GitHub Copilot's mascot, along with a sparkling star emoji.

ARS

YOLO Codee ELX teamlab
The user wants to investigate the ARS database, which they suspect is polluted due to seeing "no logs" messages. The initial goal is to validate this assumption about the database content.
select Codee ELX teamlab
The user is proposing a change to the ARS Core system to conditionally discard OpenTelemetry (OTel) span envelopes before

THANK YOU

  • INNOVATE
  • SESSIONS
  • GRASP

Hamish Songsmith: https://www.linkedin.com/in/hamish-songsmith

ARIS: https://aris.ai/

Two QR codes.

People

  • Aubrey
  • Sam Altman
  • Steve Ballmer

Technologies & Tools

  • ChatGPT
  • Claude Code
  • Cursor
  • Cursor CLI
  • Excel
  • MCP

Standards & Specs

  • AARM
  • ISO
  • New South Wales AI Governance Framework
  • NIST

Concepts & Methods

  • Context Window
  • GRASP Framework
  • Swiss Cheese Model
  • Transfer Pricing
  • YOLO Mode
  • Zero Trust Architecture

Organisations & Products

  • Air Canada
  • Bunnings
  • CommBank Accelerate AI
  • LinkedIn
  • MindForge
  • Pocket OS