Engineering for the Agentic Web When 50% of Your Traffic is Robots
The Mysterious Traffic Spike: Setting the Scene
The speaker opens by describing a real client scenario where a marketing website targeting Australian audiences received an alarming notification: 75% of daily traffic and 48% of monthly traffic were bots. Comparing normal Australian traffic patterns (daytime peaks, mobile-heavy, newsletter spikes) against actual data, the speaker reveals anomalous spikes from Ireland (over 50% of traffic on April 15), constant pounding from Norway, and suspicious activity from the United States, setting up the investigation into what's really happening.
Good Bots vs. Evil Bots: Know Your Enemy
The speaker distinguishes between legitimate 'good' bots (search crawlers, AI assistants, archivers, social media bots) that follow robots.txt and have predictable, identifiable behavior, versus malicious 'evil' bots that scrape data, hide their identity, ignore robots.txt, and exhibit unusual traffic patterns. Concrete examples illustrate evil bot indicators, such as outdated 2014 user agents still active today and bots hammering websites with thousands of requests every twelve hours.
Mitigation Basics: CDN and Managed Challenges
The speaker explains why a CDN is essential for bot mitigation, introducing the 'managed challenge' as the primary defense tool against bad bots—a lightweight, user-side verification method offered by CDNs like Fastly and Cloudflare, or self-hosted options like Anubis. A case study shows how implementing managed challenges against Norwegian bot traffic reduced malicious requests from 15,000 to 1,000, though the speaker notes this is an ongoing cat-and-mouse game as bots adapt and relocate.
Catching Bots: Canary Tricks and Traffic Management
The speaker demonstrates the 'canary trick'—creating a hidden sitemap page invisible to humans and search engines but detectable by scrapers, allowing site owners to identify and ban malicious bots. The discussion expands to managing search bot traffic, avoiding crawler traps from faceted search pages, and using Cloudflare's bot analytics and 'super bot fighting mode' to selectively allow or block specific AI bots like Amazon bot or ByteDance based on their value versus resource drain.
Embracing AI Bots: LLMs.txt and Markdown Serving
The speaker introduces LLMs.txt, an emerging standard providing AI models with a clean, concise summary of site content without navigation clutter or JavaScript, and explains how to implement it by updating robots.txt to whitelist the roughly 150 known AI bots. The segment also covers serving markdown versions of pages to AI bots via Cloudflare's automatic HTML-to-markdown conversion, reducing processing overhead for both search and AI crawlers.
Protecting Content: Emerging Standards and Limitations
The speaker discusses proposed (not yet standardized) HTML meta tags like 'no AI' and 'no image AI' for signaling content restrictions to AI crawlers, alongside image protection tools like Glaze and Nightshade that obscure images from AI training. The speaker candidly notes this is ultimately 'a losing battle' since there's more financial incentive in using AI than in defending against it.
Action Plan for Content Providers
The speaker presents a practical strategy for website owners to handle AI ingestion: convert content to markdown, publish LLMs.txt, regularly review and update robots.txt, audit third-party JavaScript libraries to avoid unexpected costs, mitigate crawler traps, enable managed challenges, and continuously monitor traffic patterns—potentially using automated agents to flag anomalies alongside techniques like the canary trick.
Action Plan for Content Consumers and Closing Remarks
The speaker closes by addressing AI agent and crawler builders directly, urging them to 'be a good bot' by creating identifiable user agents, seeking out LLMs.txt, and following robots.txt to receive clean, ad-free content in return. The talk ends with a lighthearted warning that non-compliant bots risk being redirected to unwanted content, followed by a wish for a harmonious coexistence with AI and bots.
And welcome. I hope you had a really nice afternoon tea, and let's get into it. So today, we'll try to find the robots, identify the robots, know the robots, and mitigate the bots, and we'll create an action plan of what to do. Have you received an email like this?
Your traffic spike. Your 75% of today's traffic was just robots, and 48% of all your traffic for the last thirty days is bots. What do you do? Must be those AI bots. Yeah. But before pointing finger, let's investigate. What is it? Let's give a little bit of context.
So this notification came for my client who's got two marketing websites targeting Australian audience, may mainly around Southwest Queensland. Tourists are welcome, but probably 90% is Australian traffic. So the usual for Australian traffic, we've got most traffic during the day, a bit of activity during the night, very minimal activity during the night. There is a spike after email newsletter.
There is mostly mobile traffic. There is most traffic is from Australia. So this is how it looks like in Google Analytics or Cloudflare. So you got spikes, you got your time to go to sleep time. And this is our mobile traffic.
This is for one day. You got mostly mobile devices. But this was the actual pattern. Ireland? Hello? We've got unknown browser. What's happening? What's going on? Let's, like, dig into it. Also, Analytics showed up.
It's like, oh, what happened on April 15? Ireland decided to take over. And it's 50%. It's more than 50% of the traffic. Another one that didn't show up on Google Analytics was Norway, and you can see that from Cloudflare. We've got just constant pounding 500 requests a minute.
Like, no. Norwegian people are not that interested in what's happening in Brisbane. United States, yeah, no. Something is fishy there as well. So let's get to know our bots. What's actually happening? So there are good bots that you need to let go and crawl your web applications and your websites.
You've got search crawlers, AI assistants, archivers, social media, application bots. If you have a monitoring bot, you probably know it hits your website every minute and you know its user agent. So now with bots, yes, you can easily identify them. You can see their user agent.
It's a BIN bot or is it Google bot or Facebook, and things on. You can reverse look up where they're coming from, and they're coming from big data centers. They do follow robots. Txt, which tells which pages on your web application, web portal they're allowed to go to. And they have predictable behavior.
They usually don't flood your website and take it down. The evil bots, those ones that are really difficult to identify, scrapers, spiders, whatever, they just steal your data. They usually come with direct traffic, no refers. They usually have unusual location and time, unusual to your usual traffic.
They have unusual request volumes. They can just spike every twelve hours, or they can just constantly download all your data. It's difficult to identify them because they don't present themselves of anything. They actually try to hide who they are. They don't follow robots, and they don't follow any other no follows or no index attributes when you for search engines.
And they also fall into the crawler traps. If you have lots of facets on your search page, they'll just max it out. You'll have your your application might even fail just because it will try to get all the all the data for the search that a regular human will never achieve. So here is a bit of an example of the evil bot indication. So this one, user agent is quite valid in 2014.
That's what they were using. And they also have unusual patterns. So they would just keep going all day, all night. They still do. They've just been thrown away. And the other one is every twelve hours, let's just hit this web application with thousands of requests.
Thanks, but no thanks. So let's mitigate them once we found them. Use CDN. If you don't, well, get ready. Get a team of cloud and system engineers because it's quite difficult without CDN. And the main sort of baseball bat against those bad bots is the mana challenge at the moment. So this is something that works on the user side. It's not it doesn't take your server resources.
It just gathers a bit of signals about who the user is. It's provided by CDN. So there is a couple of examples from Fastly and from Cloudflare. Or you can install your own if you don't have a CDN. So this is Anubis, and it just challenges your users if they're real users or not.
So it's not a capture, but it's slightly different. So after implementing managed challenge for this Norwegian traffic, we still have they still keep asking for all this information, but now Cloudflare just kicks them off because they do not resolve the manage challenge.
And if there are any legitimate users from Norway from this particular area, and there are a couple of other indications you can say which traffic you want to get in or which traffic you want to challenge, they can get in if they're humans. But it is a work them all game, so you need to find them, you need to ban them, and then they try to get in again, and then they probably change where they're trying to get into your system.
So we had 15,000 requests, but now it's down to 1,000 requests. And they might have changed location where they're hitting this particular system. You can try to catch a bot. Here's a canary trick. Create in your site map, create one page that's only visible to the bots, obviously, because they'll check out your site map.
Put a no index on it so the search engines will not hit it. Humans will not find it, but scrapers will and then ban them. Then you can manage your search bot traffic as well. Some of them, like BIN just keeps hammering all the pages.
Avoid crawl traps. You can get the crawl delay. Also, look out if you have any third party paper use libraries. Search bots might actually execute them, and it might become quite expensive. And tools like Cloudflare, they actually provide you with a bit of more assistance in analyzing your AI bot traffic.
So as you can see in this particular example, Microsoft's BIN bot is just keeps coming at it. Just analyze what's happening, which bots are actually visiting your platform. Cloudflare provides a super bot fighting mode. So you can turn on turn off things and see how it goes, which bots you want to keep on your site and which bots won't. So for example, we turned on Amazon bot because there is not too many referral from them or ByteDance.
They scrape a lot of data, but they don't bring traffic. So it's like, oh, no. Only Google and Bing allowed and maybe some others. And another way to mitigate the traffic is to point to the LLMs. TXT. So this is an emerging standard.
It's plain text or markdown. So you provide AI models and agents a curate, concise summary of the site content. It doesn't have navigation ads or JavaScript, so it's it's cheap for AI models to actually read this content so they don't have to process it all. But they don't seek out for it.
So you actually need to push all the bots into it. And how you do it? You just update your robots. Txt. So in robots. Txt, you have to white list which robots you actually want to point to LLMs. And there is 150 AI bots at the moment. So if you follow this GitHub repository, just two weeks ago, it was 141. So it just keeps going, you'll have to update your robots. There is another way to manage traffic.
So AI bots don't really need all this fancy HTML or JavaScript. They just need your text. So serve markdown. In Cloudflare, again, we have automatically converting HTML to markdown if bots ask asking for text markdown in the headers.
So it just cleans semantic text, again, with no pollution. And you can enable it at either on your CDN level or your application now will have a user agent check that will serve it to all those 150 bots or make sure that all the links in LLMs are pointing to MD version of the pages.
Or if we go back to 2021, what can humans do to protect their assets? There are some talks about standards on the HTML level. They are not standards. There are some proposals. This one particular comes from Dev and Art. So just like search engines have no index, no follow, This is no AI, no image AI, meta tags, metadata, HTTP headers.
So you can try to use that, but it's not the standard. It's something that may not come up. And again, we can try to protect our images with some sort of algorithm that will obscure it from LLMs. So you have glaze that protects the image style imitations, or you have nightshades that actually prompt specific poisoning.
And, of course, it's a losing battle because there's more money in actually using AI than protecting from AI. And this in the self explanatory, of course, you can't really defend from AI, but you can embrace it.
So the action plan for content providers. So for you, it's you need to create a strategy for AI ingestion. So you will get AI bots, and you will it's the same as you are going to get search engine bots. So make sure you're converting your data, your content into markdown, publish LLMs. TXT, and review your robots.
TXT, which robots you are accepting, which ones blacklisted. Mitigate your third party libraries so that you avoid any bill shock because they do execute JavaScript and not only AI bots, but search bots as well, and mitigate the crawler traps and facets crawling.
So make sure your your server is not being overloaded by facets that's query strings. That just doesn't make sense. Enable manage challenges and monitor your traffic. And it's really difficult to say what's actually your normal traffic because sometimes you might have an event that lots of people will come at once. So it's actually you can try to write some sort of agent that will monitor your traffic and tell you, oh, there's something fishy happening. And try the canary trick and tell me how it goes. And for content consumers, so when you build your AI agent and crawlers, be a good bot.
So create your own user agent, even if it's your own, so that we can identify you and we can serve you good data. So seek out LLMs. TXT. Follow robots. TXT. Don't be then you'll get clean, ad free semantic content. Well, if not, it's really easy to send you to Moldbook, and you will be talking to your friends.
Otherwise, thank you, and let's have a really good live together with AI and with all the pots.
Technologies & Tools
- Anubis
- CDN
- Glaze
- Google Analytics
- HTML
- Markdown
- Nightshade
- Super Bot Fight Mode
Standards & Specs
- LLMs.txt
- No AI meta tag
- robots.txt
Concepts & Methods
- Canary trick
- Crawler traps
- Managed Challenge
Organisations & Products
- Amazon Bot
- Bing
- ByteDance
- Cloudflare
- Fastly
- GitHub
- Google Bot
- Microsoft
Over the last two years, our customer web traffic changed: today around 50% of visitors were unknown browsers and AI agents. The era of aligning with the traditional search engine crawlers with Core Web Vitals is shifting; the new challenge is feeding focused low-noise context to autonomous agents and Large Language Models (LLMs).
Traditional Search Engine Optimisation (SEO) relies on techniques such as keyword density, backlink tracking, and human-readable formatting, but when a significant part of your traffic suddenly becomes AI agents, how do you ensure your content is being parsed correctly by machines?
Our team will share the strategic and architectural shifts the organisations are facing to embrace this new web reality. This isn’t about meta-tags; it’s about re-architecting how a business presents itself and its content in the AI-driven internet, including
1. Identifying Agent Traffic: How to identify & separate agent traffic.
2. Realigning Context: Trade-offs in modifying traditional website structure vs. serving structured markdown, llms.txt files, and specialised API endpoints.
3. Omni Channel Content: Serving dual-experiences with web applications for humans versus data streams for agents.
4. Lessons Learned: To block or embrace agent traffic, and how embracing “LLM Instructions” can increase content’s reach.














