Blocking The Bad Bots and Welcoming the Good Ones header graphic - business owner's guide to AI safety series

Blocking the Bad Bots, Welcoming the Good Ones

|

Our client’s website was down. It looked like a denial-of-service attack, but a malicious actor wasn’t responsible for the outage. A bot was.

A poorly controlled AI crawler was making hundreds of requests a second for the pages on the site. That’s what crawlers do – they collect data from websites to feed search engines or large language models like ChatGPT or Claude. But this one was asking for too much, too quickly.

Wellington Street Consulting blocked the bot and updated the client’s robots.txt to set clearer boundaries with crawlers.

The problem was solved, but the damage was done. While our client’s site was unavailable, its existing customers couldn’t use the platform, and new customers couldn’t sign up either. Outages damage businesses’ reputations, hit their bottom line, and waste their time.

This could happen to your business’s website. But it’s not as simple as blocking all bots. You need bots to crawl your site for visibility, so you can’t block them all. At the same time, you need to protect your site from excessive non-human traffic, so you can’t let all the bots in. Let’s talk about what small business owners actually need to know about keeping their websites discoverable and safe from being swamped.

AI bot traffic is a business operations issue

You’ve probably heard about AI bots in the context of search engine optimization (SEO) or artificial intelligence optimization (AIO), since they’re the programs that index your website’s pages to ingest their data and suggest them in response to user queries. But bots can also affect your site performance, analytics, content control, search visibility, and customer experience.

Bot traffic is growing. The Cloudflare Radar 2025 Year in Review reported that AI bots averaged 4.2 percent of HTML requests across its customer base, with bot purpose varying across training, search and user-action crawling.

Signs like these point to a potential bot issue:

  • Traffic spikes that appear not to match real customer interest
  • Pages crawled heavily, but little referral traffic and few leads
  • Server strain, strange analytics or unexpected bandwidth use

Excessive bot traffic can make you or your marketing team think you’re seeing more pageviews than you are, leading you to optimize for the wrong signals. That traffic can also increase your bandwidth and web infrastructure costs, or take your site down entirely. Meanwhile, the bots overrunning your site will ingest the intellectual property on your site to use as source data, but may not cite you or refer visitors back to you.

How do you know if AI bots crawling your site are good or bad?

There’s not a binary of good AI bots and bad ones. All bots are useful to someone – the question is whether they’re useful for your business. Your website needs to be able to distinguish between useful, useless or harmful bots so it can reject the ones that don’t serve your goals. This means having a site policy to distinguish between bot behaviors.

These behaviors include:

  • Search: Search or answer-engine crawlers may help prospects find your expertise.
  • Training: Training crawlers copy content at scale without a clear business return.
  • User-requested retrieval: User-action agents may visit the site because a real person asked them to.
  • Abuse: Undeclared or spoofed bots may ignore your robots.txt’s bot behavior rules.

To set your website up to differentiate bot behaviors and block the unhelpful ones, start by creating a list of which behaviors you’d like to allow, block, and/or monitor. You’ll want to block abuse, and you may want to block training depending on the nature of the content on your website; allowing user-requested retrieval and search is important for brand visibility and your customers’ research.

Next, talk to your hosting provider, content delivery network (CDN) or managed service provider (MSP) partners about rate limiting and bot controls.

Finally, when making decisions about blocking bots, make sure to consider SEO and AI visibility. You don’t want to hamstring your inbound marketing strategy while trying to keep your site safe. You can, in fact, block malicious bots while keeping your site friendly and readable for beneficial bots that will cite your expertise to potential customers.

Robots.txt is a signal to bots, not a security boundary

Your website’s robots.txt communicates preferences to bots like “I only want you to make one request every twenty seconds.” But robots.txt can’t actually control the bots, just suggest to them how the site prefers to be interacted with.

Bots can choose to ignore those preferences. A “nice” bot will respect your robots.txt, but malicious or poorly programmed bots can ignore it. If you assume your site’s robots.txt file controls bots, prevents scraping, and keeps non-human traffic to a reasonable level, you may be in for an unpleasant surprise.

You need a way to observe violations of your robots.txt, and a way to enforce consequences. That means using tools like your CDN, a firewall, rate limiting, bot management, and log review to take action, not just passively rely on a robots.txt.

A simple governance cycle is all a small or medium-size business needs for bot security

If your organization relies on an MSP instead of having a dedicated IT department – or if your IT department consists of a lone employee who’s good with computers – you don’t need to start a huge project or a complicated program to handle bots on your website.

Bot governance at your organization can be as simple as a light quarterly review.

Designate one person in your organization to be in charge of bot issues on your website. They don’t have to know everything about bots, but they should have expertise about your website and the ability to escalate issues that they observe.

This person should:

  • Ask hosting or MSP partners for recent bot traffic and rate-limit data and review those logs
  • Work with your security and marketing team to decide what to allow search, training, or user-requested retrieval agents to do on your site
  • Work with your marketing team to make important public content clear, structured, and easy for AI bots to cite
  • Classify known crawlers by business purpose: search, AI training, AI answer retrieval, user action, suspicious, or unknown
  • Ensure your robots.txt, sitemap, schema, CDN, and firewall settings are up to date and match what your security and marketing team chose to allow and disallow – and that you’re not relying solely on robots.txt to block unwanted traffic
  • Revisit the bot policy quarterly, as well as when the site, search strategy, traffic or AI crawler behavior changes

How to prepare your business to deal with unwanted AI crawler traffic

To know if your business is set up to handle AI bots of all kinds, ask yourself the following questions:

  • What would it cost us if bot traffic slowed or crashed the site, especially during a marketing or ad campaign?
  • Which content do we want AI systems to understand and cite to users? Which content should they never reach?
  • Are we blocking traffic because it is harmful, or because we have not decided what good AI visibility looks like?
  • Who at our organization owns the tradeoff between SEO visibility, AI discoverability, content control and uptime?
  • What evidence would tell us that our current bot policy is working?

If you need help answering those questions, or you’re not sure whether your site is both optimized for AI citation and shielded from excessive AI traffic, reach out to Wellington Street Consulting for an AI security consultation. We will review your current bot security setup, then provide you with a list of action items to meet your security and marketing goals.

Similar Posts