Blog

AI crawlers are the new scrapers. Should you block them?

LLM training bots and AI answer engines now make up a real share of traffic. Here is how to tell them apart and decide what to allow.

If your traffic graph has crept upward this year without a matching rise in signups or sales, a chunk of that climb is probably AI crawlers. Bots that scrape the web to train language models, or to fetch pages live when someone asks an AI assistant a question, now make up a measurable slice of requests to most sites.

Three kinds of AI bot

Training crawlers pull your content to feed a model. They can be heavy, requesting large parts of your site in bulk.

Live retrieval bots fetch a page in the moment because a user asked an assistant something your page answers. This one can actually send you visitors.

Impersonators that claim to be a well-known AI crawler in their user-agent string but are really someone scraping under a trusted name.

The first two are mostly honest and announce themselves. The third is the reason you cannot decide based on the user-agent alone.

Verify, do not trust the label

A user-agent string is just text; anyone can set it to anything. The reliable check is the network: a genuine crawler from a major AI company comes from that company's own address ranges. A request claiming to be that crawler but arriving from a random residential proxy or a cheap datacenter is lying about who it is.

So the useful question is not merely what does this bot call itself, but does the network it came from match the identity it claims.

Deciding what to allow

  • Allow verified live-retrieval bots if you want the referral traffic they can send.
  • Rate-limit verified training crawlers so they do not run up your bandwidth bill, rather than blocking outright.
  • Block or challenge anything claiming an AI identity from a network that does not back it up. That is not an AI company; it is a scraper in a costume.

Keep your analytics honest

Whatever you decide, tag this traffic so it does not pollute your real visitor numbers. Fraudex gives you the network and datacenter verdict behind each request, so you can confirm whether a self-described crawler actually comes from where it claims and treat it accordingly.