Engineering for the Agentic Web When 50% of Your Traffic is Robots

The Mysterious Traffic Spike: Setting the Scene

The speaker opens by describing a real client scenario where a marketing website targeting Australian audiences received an alarming notification: 75% of daily traffic and 48% of monthly traffic were bots. Comparing normal Australian traffic patterns (daytime peaks, mobile-heavy, newsletter spikes) against actual data, the speaker reveals anomalous spikes from Ireland (over 50% of traffic on April 15), constant pounding from Norway, and suspicious activity from the United States, setting up the investigation into what's really happening.

Good Bots vs. Evil Bots: Know Your Enemy

The speaker distinguishes between legitimate 'good' bots (search crawlers, AI assistants, archivers, social media bots) that follow robots.txt and have predictable, identifiable behavior, versus malicious 'evil' bots that scrape data, hide their identity, ignore robots.txt, and exhibit unusual traffic patterns. Concrete examples illustrate evil bot indicators, such as outdated 2014 user agents still active today and bots hammering websites with thousands of requests every twelve hours.

Mitigation Basics: CDN and Managed Challenges

The speaker explains why a CDN is essential for bot mitigation, introducing the 'managed challenge' as the primary defense tool against bad bots—a lightweight, user-side verification method offered by CDNs like Fastly and Cloudflare, or self-hosted options like Anubis. A case study shows how implementing managed challenges against Norwegian bot traffic reduced malicious requests from 15,000 to 1,000, though the speaker notes this is an ongoing cat-and-mouse game as bots adapt and relocate.

Catching Bots: Canary Tricks and Traffic Management

The speaker demonstrates the 'canary trick'—creating a hidden sitemap page invisible to humans and search engines but detectable by scrapers, allowing site owners to identify and ban malicious bots. The discussion expands to managing search bot traffic, avoiding crawler traps from faceted search pages, and using Cloudflare's bot analytics and 'super bot fighting mode' to selectively allow or block specific AI bots like Amazon bot or ByteDance based on their value versus resource drain.

Embracing AI Bots: LLMs.txt and Markdown Serving

The speaker introduces LLMs.txt, an emerging standard providing AI models with a clean, concise summary of site content without navigation clutter or JavaScript, and explains how to implement it by updating robots.txt to whitelist the roughly 150 known AI bots. The segment also covers serving markdown versions of pages to AI bots via Cloudflare's automatic HTML-to-markdown conversion, reducing processing overhead for both search and AI crawlers.

Protecting Content: Emerging Standards and Limitations

The speaker discusses proposed (not yet standardized) HTML meta tags like 'no AI' and 'no image AI' for signaling content restrictions to AI crawlers, alongside image protection tools like Glaze and Nightshade that obscure images from AI training. The speaker candidly notes this is ultimately 'a losing battle' since there's more financial incentive in using AI than in defending against it.

Action Plan for Content Providers

The speaker presents a practical strategy for website owners to handle AI ingestion: convert content to markdown, publish LLMs.txt, regularly review and update robots.txt, audit third-party JavaScript libraries to avoid unexpected costs, mitigate crawler traps, enable managed challenges, and continuously monitor traffic patterns—potentially using automated agents to flag anomalies alongside techniques like the canary trick.

Action Plan for Content Consumers and Closing Remarks

The speaker closes by addressing AI agent and crawler builders directly, urging them to 'be a good bot' by creating identifiable user agents, seeking out LLMs.txt, and following robots.txt to receive clean, ad-free content in return. The talk ends with a lighthearted warning that non-compliant bots risk being redirected to unwanted content, followed by a wish for a harmonious coexistence with AI and bots.

And welcome. I hope you had a really nice afternoon tea, and let's get into it. So today, we'll try to find the robots, identify the robots, know the robots, and mitigate the bots, and we'll create an action plan of what to do. Have you received an email like this?

Your traffic spike. Your 75% of today's traffic was just robots, and 48% of all your traffic for the last thirty days is bots. What do you do? Must be those AI bots. Yeah. But before pointing finger, let's investigate. What is it? Let's give a little bit of context.

So this notification came for my client who's got two marketing websites targeting Australian audience, may mainly around Southwest Queensland. Tourists are welcome, but probably 90% is Australian traffic. So the usual for Australian traffic, we've got most traffic during the day, a bit of activity during the night, very minimal activity during the night. There is a spike after email newsletter.

There is mostly mobile traffic. There is most traffic is from Australia. So this is how it looks like in Google Analytics or Cloudflare. So you got spikes, you got your time to go to sleep time. And this is our mobile traffic.

This is for one day. You got mostly mobile devices. But this was the actual pattern. Ireland? Hello? We've got unknown browser. What's happening? What's going on? Let's, like, dig into it. Also, Analytics showed up.

It's like, oh, what happened on April 15? Ireland decided to take over. And it's 50%. It's more than 50% of the traffic. Another one that didn't show up on Google Analytics was Norway, and you can see that from Cloudflare. We've got just constant pounding 500 requests a minute.

Like, no. Norwegian people are not that interested in what's happening in Brisbane. United States, yeah, no. Something is fishy there as well. So let's get to know our bots. What's actually happening? So there are good bots that you need to let go and crawl your web applications and your websites.

You've got search crawlers, AI assistants, archivers, social media, application bots. If you have a monitoring bot, you probably know it hits your website every minute and you know its user agent. So now with bots, yes, you can easily identify them. You can see their user agent.

It's a BIN bot or is it Google bot or Facebook, and things on. You can reverse look up where they're coming from, and they're coming from big data centers. They do follow robots. Txt, which tells which pages on your web application, web portal they're allowed to go to. And they have predictable behavior.

They usually don't flood your website and take it down. The evil bots, those ones that are really difficult to identify, scrapers, spiders, whatever, they just steal your data. They usually come with direct traffic, no refers. They usually have unusual location and time, unusual to your usual traffic.

They have unusual request volumes. They can just spike every twelve hours, or they can just constantly download all your data. It's difficult to identify them because they don't present themselves of anything. They actually try to hide who they are. They don't follow robots, and they don't follow any other no follows or no index attributes when you for search engines.

And they also fall into the crawler traps. If you have lots of facets on your search page, they'll just max it out. You'll have your your application might even fail just because it will try to get all the all the data for the search that a regular human will never achieve. So here is a bit of an example of the evil bot indication. So this one, user agent is quite valid in 2014.

That's what they were using. And they also have unusual patterns. So they would just keep going all day, all night. They still do. They've just been thrown away. And the other one is every twelve hours, let's just hit this web application with thousands of requests.

Thanks, but no thanks. So let's mitigate them once we found them. Use CDN. If you don't, well, get ready. Get a team of cloud and system engineers because it's quite difficult without CDN. And the main sort of baseball bat against those bad bots is the mana challenge at the moment. So this is something that works on the user side. It's not it doesn't take your server resources.

It just gathers a bit of signals about who the user is. It's provided by CDN. So there is a couple of examples from Fastly and from Cloudflare. Or you can install your own if you don't have a CDN. So this is Anubis, and it just challenges your users if they're real users or not.

So it's not a capture, but it's slightly different. So after implementing managed challenge for this Norwegian traffic, we still have they still keep asking for all this information, but now Cloudflare just kicks them off because they do not resolve the manage challenge.

And if there are any legitimate users from Norway from this particular area, and there are a couple of other indications you can say which traffic you want to get in or which traffic you want to challenge, they can get in if they're humans. But it is a work them all game, so you need to find them, you need to ban them, and then they try to get in again, and then they probably change where they're trying to get into your system.

So we had 15,000 requests, but now it's down to 1,000 requests. And they might have changed location where they're hitting this particular system. You can try to catch a bot. Here's a canary trick. Create in your site map, create one page that's only visible to the bots, obviously, because they'll check out your site map.

Put a no index on it so the search engines will not hit it. Humans will not find it, but scrapers will and then ban them. Then you can manage your search bot traffic as well. Some of them, like BIN just keeps hammering all the pages.

Avoid crawl traps. You can get the crawl delay. Also, look out if you have any third party paper use libraries. Search bots might actually execute them, and it might become quite expensive. And tools like Cloudflare, they actually provide you with a bit of more assistance in analyzing your AI bot traffic.

So as you can see in this particular example, Microsoft's BIN bot is just keeps coming at it. Just analyze what's happening, which bots are actually visiting your platform. Cloudflare provides a super bot fighting mode. So you can turn on turn off things and see how it goes, which bots you want to keep on your site and which bots won't. So for example, we turned on Amazon bot because there is not too many referral from them or ByteDance.

They scrape a lot of data, but they don't bring traffic. So it's like, oh, no. Only Google and Bing allowed and maybe some others. And another way to mitigate the traffic is to point to the LLMs. TXT. So this is an emerging standard.

It's plain text or markdown. So you provide AI models and agents a curate, concise summary of the site content. It doesn't have navigation ads or JavaScript, so it's it's cheap for AI models to actually read this content so they don't have to process it all. But they don't seek out for it.

So you actually need to push all the bots into it. And how you do it? You just update your robots. Txt. So in robots. Txt, you have to white list which robots you actually want to point to LLMs. And there is 150 AI bots at the moment. So if you follow this GitHub repository, just two weeks ago, it was 141. So it just keeps going, you'll have to update your robots. There is another way to manage traffic.

So AI bots don't really need all this fancy HTML or JavaScript. They just need your text. So serve markdown. In Cloudflare, again, we have automatically converting HTML to markdown if bots ask asking for text markdown in the headers.

So it just cleans semantic text, again, with no pollution. And you can enable it at either on your CDN level or your application now will have a user agent check that will serve it to all those 150 bots or make sure that all the links in LLMs are pointing to MD version of the pages.

Or if we go back to 2021, what can humans do to protect their assets? There are some talks about standards on the HTML level. They are not standards. There are some proposals. This one particular comes from Dev and Art. So just like search engines have no index, no follow, This is no AI, no image AI, meta tags, metadata, HTTP headers.

So you can try to use that, but it's not the standard. It's something that may not come up. And again, we can try to protect our images with some sort of algorithm that will obscure it from LLMs. So you have glaze that protects the image style imitations, or you have nightshades that actually prompt specific poisoning.

And, of course, it's a losing battle because there's more money in actually using AI than protecting from AI. And this in the self explanatory, of course, you can't really defend from AI, but you can embrace it.

So the action plan for content providers. So for you, it's you need to create a strategy for AI ingestion. So you will get AI bots, and you will it's the same as you are going to get search engine bots. So make sure you're converting your data, your content into markdown, publish LLMs. TXT, and review your robots.

TXT, which robots you are accepting, which ones blacklisted. Mitigate your third party libraries so that you avoid any bill shock because they do execute JavaScript and not only AI bots, but search bots as well, and mitigate the crawler traps and facets crawling.

So make sure your your server is not being overloaded by facets that's query strings. That just doesn't make sense. Enable manage challenges and monitor your traffic. And it's really difficult to say what's actually your normal traffic because sometimes you might have an event that lots of people will come at once. So it's actually you can try to write some sort of agent that will monitor your traffic and tell you, oh, there's something fishy happening. And try the canary trick and tell me how it goes. And for content consumers, so when you build your AI agent and crawlers, be a good bot.

So create your own user agent, even if it's your own, so that we can identify you and we can serve you good data. So seek out LLMs. TXT. Follow robots. TXT. Don't be then you'll get clean, ad free semantic content. Well, if not, it's really easy to send you to Moldbook, and you will be talking to your friends.

Otherwise, thank you, and let's have a really good live together with AI and with all the pots.

Agenda

  • Finding robots
  • Know your bots
  • Mitigating bots
  • Self-explanatory scene
  • Action plan

Have you ever received that email...

Spike in automated traffic detected

Configure Super Bot Fight Mode

Cloudflare scores every request on our network to determine the likelihood the request came from a bot or a human. We've detected an increase in your automated traffic that may indicate malicious bot activity. Automated traffic typically makes up 48.0% of your traffic. On 2026-03-15, automated traffic increased by 56.0% for...

  • 48.0% traffic that was automated over a recent 30-day period
  • 75.0% Traffic that was automated on 2026-03-15
  • 56.0% Increase in automated traffic
A screenshot of an alert notification, likely from Cloudflare, titled 'Spike in automated traffic detected'. It displays a button labeled 'Configure Super Bot Fight Mode' and text explaining how Cloudflare detects bots. Below this, three data boxes show statistics: '48.0% traffic that was automated over a recent 30-day period', '75.0% Traffic that was automated on 2026-03-15', and '56.0% Increase in automated traffic'.

Must be evil AI bots!

Image credit: Futurama / 20th Century Studios

An illustration of Bender from Futurama, shown from the chest up and slightly angled, with a menacing smirk and glowing yellow eyes. His mouth is open, revealing his teeth.

Let's investigate!

An illustration featuring Fry from Futurama, pointing his finger upwards, with a surprised expression. Next to him is a cartoon frog wearing a collar and a small medal. The image is credited to Futurama / 20th Century Studios.

Let's investigate!

Context:

  • Two marketing websites
  • Target audience is in Australia (around SE QLD), but tourists are welcome too.

Usual traffic pattern (AEST):

  • Most traffic is during the day
  • Minimal activity at night
  • Spike after sending email newsletter
  • Most traffic is from mobile devices
  • Most traffic is from Australia

Usual traffic pattern: visitors timeline

Total visits: 31.64k

Australia visits: 31.64k

A line graph titled 'Visitors timeline' displays visit data for Australia over three days, from Sunday 10th to Tuesday 12th, with time in GMT+12. The Y-axis represents 'Visits' up to 1.1k. The blue line shows a distinct daily pattern of high visitor numbers during the day, peaking around 10 AM (e.g., 1.02k on May 10, 2026), and significantly lower activity during early morning hours.

Usual traffic pattern: devices

  • Mobile: 25,305
  • Desktop: 4,977
  • Tablet: 1,356
A bar chart shows device types and their counts: Mobile with 25,305, Desktop with 4,977, and Tablet with 1,356, illustrating that mobile devices constitute the majority of traffic.

Actual traffic pattern

Summary Data

  • Visits: 13.91k (up 40.79%)
  • Page views: 15.78k (up 43.19%)
  • Page load time: 4,540ms (down 17.04%)

Traffic by Device Type

  • ChromeMobile: 5.97k
  • Unknown: 3.95k
  • ChromeMobileWebview: 1.27k

Traffic by Countries

  • Ireland: 8.89k
  • Australia: 52
  • United States: 13
  • Singapore: 4
  • Belgium: 2
A screenshot of an analytics dashboard, likely from Google Analytics or Cloudflare, displaying website traffic patterns. It shows summary metrics for visits, page views, and page load time, along with percentage changes. Detailed breakdowns include traffic by device type (ChromeMobile, Unknown, ChromeMobileWebview) and by country (Ireland, Australia, United States, Singapore, Belgium).

Unusual traffic pattern from Google Analytics

Country Period 1 Total (31,628) Period 2 Total (31,177) Period 3 Total (8,353)
Total 100% of total 100% of total 100% of total
1 Ireland 14,719 (46.54%) 14,627 (46.92%) 16 (0.19%)
2 Australia 14,470 (45.75%) 13,999 (44.9%) 7,653 (91.62%)
3 Singapore 1,039 (3.29%) 1,040 (3.34%) 66 (0.79%)
4 United States 620 (1.96%) 618 (1.98%) 204 (2.44%)
A line graph displays website traffic trends from April 13th to April 27th, illustrating total traffic and specific traffic patterns for Australia, Ireland, Singapore, United States, and New Zealand. Below the graph, a table provides detailed traffic data by country, showing the number of hits and their percentage contribution for three distinct data collection periods.

Unusual traffic pattern (24 hrs)

HTTP traffic analytics in Cloudflare

  • Total: 106.01k
  • Norway: 45.98k
  • Australia: 36.02k
  • United States: 21.43k
  • Singapore: 1.64k
  • New Zealand: 948

A line graph displaying HTTP traffic analytics from Cloudflare for different countries over a 24-hour period, from 18:00 to 15:00 GMT+10. The y-axis represents Page views, ranging from 0 to 800, and the x-axis represents Time. The graph shows that traffic from Norway (green line) is consistently high, hovering around 500 page views. Australia (blue line) shows fluctuating traffic with a dip during the local night and a rise during the day. The United States (yellow line) has lower, fluctuating traffic. Singapore and New Zealand (purple line) show very low traffic throughout the period.

Know your bots!

Image credit: Futurama / 20th Century Studios

An illustration from Futurama depicting a diverse group of robots, including the large, armored Destructor, a police robot, and several other distinctive robotic characters, standing together in a spacious room.

Known bots

  • Search crawlers
  • AI Assistants/Crawlers
  • Archivers
  • Social media crawlers
  • Application bots/crawlers:
    • Marketing/analytics crawlers
    • Monitoring (uptime) bots
    • Security scanning
An illustration of three robot characters from the animated TV show Futurama. On the left is Bender, in the middle is a smaller, green robot, and on the right is a large purple robot who is eating a miniature ship model.

Known bots

  • Identifiable user agent
    Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible;
    bingbot/2.0; +http://www.bing.com/bingbot.htm)
    Chrome/116.0.1938.76 Safari/537.36
  • Reverse DNS Lookup
  • Follow robots.txt
  • Predictable behavior

Evil bots

  • Scrapers, spiders
  • Direct traffic (no referer)
  • Unusual location/time
  • Unusual request volume
  • Difficult to define/identify
  • Ignore /robots.txt and rel="nofollow" attributes
  • Deep facet scanning or “crawler trap”
An illustration depicting three robots from the animated series Futurama, identified as Bender, a gray robot, and a yellow robot, all firing ray guns. The image credit states "Futurama / 20th Century Studios".

Evil bots

  • Bots use weird (valid) user agents (valid user agent from 2014!)
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_9_2)
AppleWebKit/537.75.14 (KHTML, like Gecko)
Version/7.0.3 Safari/537.75.14
  • Unusual traffic pattern
A line graph titled "Requests" shows fluctuating request volume over a 24-hour period (GMT+10), with a sharp decline at the end. Another line graph shows website traffic with two significant spikes, one around 3 AM and another around 4 PM, indicating unusual activity.

Mitigate bad bots

  • Edge-level: use CDN
    • Cloudflare, Cloudfront, F5, Fastly, etc
  • Origin-level: Get a team of Cloud/System Engineers
    • If you can't have a CDN :(

Managed challenge, eg proof-of-work

  • small non-interactive JavaScript challenges
  • gathering more signals about the visitor/browser environment.
  • in-browser detections for visitor characteristics
  • provided by CDN
  • Or TecharoHQ/anubis (open source)

Details from Anubis screenshot: Protected by Anubis From Techaro. Made with ❤️ in 🇯🇵. Mascot design by CELPHASE. This website is running Anubis version v1.22.0-sciolly.org-1984.

Three screenshots illustrating different security verification challenges. The first shows a Fastly "is verifying your browser..." screen. The second shows a Cloudflare "Performing security verification" screen with a "Verify you are human" checkbox. The third shows an Anubis "Making sure you're not a bot!" screen with an anime-style character, displaying "Calculating..." and "Difficulty: 4, Speed: 0kH/s."

Ban bots: before vs after

Added a managed challenge to Norwegian traffic:

  • 7 day analytics
  • Bot keeps scraping but fails managed challenge
  • Legitimate users (possible)
A dashboard displaying request analytics, featuring a line graph tracking requests over approximately 7 days. The graph shows a high volume of '403 Forbidden' requests (blue line) consistently around 1.5k-2k, and a very low volume of '200 OK' requests (orange line) near zero. A summary box above the graph indicates a 'Total' of 281.48k requests, '403 Forbidden' at 280.73k, and '200 OK' at 751.

Whac-A-Mole

  • Ban bot (or challenge) based on their properties - goodbye Norway and Ireland, this is NOT Eurovision!
  • 15K requests without mitigation
  • 1K requests mitigated by Cloudflare - bot stopped
A line graph titled with overall statistics: Total 17.21k, Mitigated by Cloudflare 1.06k, Served by Cloudflare 15.38k, Served by origin 765. The graph shows request count on the y-axis (0-16k) versus time (GMT+12) on the x-axis, from 15:00 to 12:00 the next day. A large spike in 'Served by Cloudflare' requests is visible around 15:00 on May 12, 2026, reaching almost 15k, while 'Served by origin' and 'Mitigated by Cloudflare' remain very low. A tooltip for May 12, 2026, 3:00 PM shows 14.9k requests served by Cloudflare, 708 served by origin, and 0 mitigated by Cloudflare.

How to catch a bot?

The "Canary" Trick

  • Create one high-value article or page that is only appears in your sitemap.
  • Real humans will rarely find it, but scrapers will find it fast!
  • *hide this page from search bots with "noindex" meta tag
An illustration of the robot Bender from Futurama holding a white chicken upside down by its legs. Image credit: Futurama / 20th Century Studios.

Search bots traffic management

  • Search crawlers, archivers, social media crawlers:
    • Let them crawl, it's your discoverability!
    • Look out for any 3rd party pay-per-use libraries to avoid bill shock:
      • Google Maps API calls may become very expensive
  • Update robots.txt
    • Avoid crawl traps:
      Disallow: /search?
    • Slow down bots:
      User-agent: *
      Crawl-delay: 10

* unofficial directive, it may be ignored by Google, but Bing takes it into consideration

AI bot traffic management

AI Assistants/Crawlers

  • Analyse AI bot traffic
The slide includes a dashboard titled 'Crawlers' displaying traffic data for various AI bots and web crawlers. The dashboard shows 'Allowed requests' and 'Total referrals' with associated numerical values and percentages for the following bots: Microsoft BingBot, Google Googlebot, OpenAI ChatGPT-User, ByteDance Bytespider, Amazon Amazonbot, DuckDuckGo DuckAssistBot, Internet Archive archive.org_bot, and Meta FacebookBot.

Bot traffic management

Configure superbot fighting mode

Screenshot of a web interface showing "Super Bot fight mode" configuration settings with options for JS Detections, Static resource protection, Verified bots, Optimize for wordpress, and Definitely automated traffic.

AI bot traffic management

Allow bots which bring traffic, ban others

A screenshot of a bot traffic management interface showing a table with entries for BingBot, Googlebot, and Amazonbot. For each bot, it displays its category (Search Engine Crawler or AI Crawler), bytes transferred, request statistics (allowed vs. unsuccessful), and a toggle to block the crawler. Amazonbot is shown with its 'Block Crawler' toggle enabled.

AI bot traffic management

Enable /llms.txt

  • an emerging standard
  • plain text or Markdown file
  • provide AI models and agents with a curated, concise summary of site content
  • no navigation/ads/js/etc

* AI bots don't use llms.txt by default, they have to be pushed into it.

AI bot traffic management

Update /robots.txt to force AI bots to use /llms.txt

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: PerplexityBot
Allow: /llms.txt
Disallow: /

150 AI bots and counting

https://github.com/ai-robots-txt/ai.robots.txt

AI bot traffic management

Serve Markdown (not HTML) to AI bots

  • Clean semantic text
  • No javascript
  • No ads
  • No repeated content

How:

  • Enable conversion on CDN level based on HTTP header
  • Convert output based on user-agent on application level
  • All links in /llms.txt should point to *.md version of a page

Markdown for Agents

Automatically convert HTML to Markdown for requests that use content negotiation headers (Accept: text/markdown)

A screenshot of a user interface setting labeled "Markdown for Agents" with an active toggle switch.

NOAI

If only we could go back to 2021...

(before ChatGPT)

Image credit: Futurama / 20th Century Studios

An illustration of Bender from Futurama, depicted as if made of wood, standing outdoors by a body of water with trees and giving a thumbs-up gesture.

Anti-AI offensive

"Image credit: Futurama / 20th Century Studios"

An illustration from Futurama depicting a line of five characters in green military-style outfits with pointy helmets and yellow goggles, holding futuristic purple weapons. The character in the foreground has wide, bugged-out eyes and a surprised expression.

HTML Standard proposals

“DeviantArt offers a “NoAI” setting that unambiguously communicates that artwork with this flag is not authorized for inclusion in third-party datasets for training AI models.”

  • Add noai attribute to <img> tags, eg
    <img src="..." noai>
  • NOAI metatags
    <meta name="robots" content="noimageai" />
    <meta name="robots" content="noai" />
  • NOAI HTTP header
    X-Robots-Tag: noai

Anti-AI offensive

  • The challenge is, without visible image distortion, make it difficult for LLMs to understand context/imitate style
  • Research by Dr Ben Zhao and team from University of Chicago
  • * Visual arts only, sorry dancers, writers and music producers and other creatives

https://glaze.cs.uchicago.edu/what-is-glaze.html

Glaze

Style imitation protection

  • Distorts the style of art seen by AI
  • Adds little visible artefacts
  • Works after the image was through post processing or screenshots

* as of 2023

Glaze: Protecting Artists from Style Mimicry by Text-to-Image Models (https://doi.org/10.48550/arXiv.2302.04222)

A comparison image demonstrating Glaze. On the left, two identical "Original" illustrations of Alice from Wonderland by John Tenniel. On the top right, under "AI Generated (No Glaze)", an AI-generated image of a fluffy white dog in a style closely mimicking the original illustration. On the bottom right, under "AI Generated (With Glaze)", an AI-generated image of a bulldog-like dog, but its style is noticeably distorted and does not match the original illustration, indicating Glaze's protection.

Nightshade

changes the subject matter. For example, an image of a cow may instead be seen as a purse through an AI lens, prompt-specific poisoning.

* as of 2023

Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models (https://doi.org/10.48550/arXiv.2310.13828)

A 2x2 grid of four stylized images of what appear to be cat-like faces, each with distorted or glowing features, illustrating the concept of prompt-specific poisoning in generative AI models.

Losing Battle

Overlai

  • App on your phone to cloak your photos before you upload them to socials/internet
  • No updates since 2024

Mist

  • AI systems trained on "Misted" images typically output images with an ugly full-bleed watermark that renders the image useless to bad faith users.
  • No updates since 2023

Anti-Glaze, Nightshade was defeated by LightShed

  • Glaze and Nightshade are hard to use
  • It's like DRM in early 2000s

This scene is self explanatory

More investment goes to AI than to NOAI or AI defence 😟

An illustration from Futurama depicts a yellow robot, possibly Bender, kneeling in the rain with arms outstretched, looking upwards with an expression of despair. Image credit: Futurama / 20th Century Studios.

TL;DR; for content providers

  • Define your content strategy for AI ingestion
    • Convert content into markdown
    • Publish /llms.txt
    • Review /robots.txt
    • White-list AI bots, block low value bots
    • Mitigate 3rd party libraries loading to avoid bill-shock
    • Mitigate crawler traps or facets crawling
  • Enable managed challenges
  • Monitor your traffic
  • Try the "Canary" trick

TL;DR; for content consumers

Build your agents and AI crawlers to become good bots:

  • Use identifiable `user-agent`
  • Seek out `/llms.txt`
  • Follow `/robots.txt`

You'll get clean, ad-free, semantic content

Or we'll redirect you to Moltbook!

Thank you!

This presentation features fair-use commentary utilising imagery from Futurama (TM & © 20th Television Animation / Matt Groening). No robots were harmed, and no copyright infringement is intended. Please don't send Robot Santa's lawyers.

An illustration from Futurama depicting Fry happily hugging Bender, who appears surprised.

Technologies & Tools

  • Anubis
  • CDN
  • Glaze
  • Google Analytics
  • HTML
  • Markdown
  • Nightshade
  • Super Bot Fight Mode

Standards & Specs

  • LLMs.txt
  • No AI meta tag
  • robots.txt

Concepts & Methods

  • Canary trick
  • Crawler traps
  • Managed Challenge

Organisations & Products

  • Amazon Bot
  • Bing
  • ByteDance
  • Cloudflare
  • Facebook
  • Fastly
  • GitHub
  • Google Bot
  • Microsoft