August 3, 2026
New federal bill could help websites take control of their content back from stealth bots
We last wrote about two recent laws with similar goals: providing websites with more control over who accesses their site. New York passed the Stealth Crawler Prohibition Act (though it awaits signature by the governor), and the UK's CMA places requirements on Google.
In July, a bipartisan group in the U.S. House introduced the Stealth Bot Prohibition Act (H.R.9915). It has the same key requirement as the New York bill, requiring bots to identify themselves, and goes two steps further than a state law could reach: it obviously applies nationwide, and it covers every type of website, not just news publishers.
What the bill covers
This bipartisan federal bill was introduced by Representatives Laurel Lee (FL), Valerie Foushee (NC), and Gus Bilirakis (FL). It makes it illegal to run a "stealth bot," which the text defines as a bot that accesses, retrieves, scans, indexes, scrapes, or otherwise interacts with a website, digital platform, or online service without prior disclosure of its identity and purpose. In practice, that means bots need to identify themselves with a valid, accurate user-agent string and clearly communicate their intent for accessing the site and using its content.
It isn't enough for a bot to identify itself; it also needs to declare its purpose when it requests access to a page, according to the bill. Purposes may include: text and data mining, search indexing, inferencing, artificial intelligence development, support, or operations (such as training, fine-tuning, retrieval, augmented generation). A bot that shows up honestly about who it is but does not disclose how it’s using your content is not compliant if this bill becomes law.
The bill has two main focuses. A person can't deploy a bot in a way that's reasonably likely to damage, impair, or burden a site's technical or commercial operation. And a person can't intentionally disguise a bot as a human user for use with a generative AI model. The word "commercial" is important there. This bill is focused on the business harm of having your content taken and resold out from under you.
"Bot," in this bill, is defined about as widely as it can be: crawlers, spiders, fetchers, clients, user agents, AI agents, and equivalent tools. Agents are named directly, which matters given where the web is heading.
On enforcement, the Federal Trade Commission (FTC) can bring civil actions with penalties of up to $53,000 per violation, adjusted for inflation each year, and state attorneys general can sue on behalf of their residents. That's a different model from New York, which handed news organizations a private right of action to sue on their own. Here, the teeth are the FTC and the states.
Expansion: Protecting more sites
The New York bill was a good start, but it worked for a limited group. A publisher had to produce original journalism, publish on a schedule, run a corrections process, and clear 1,000 monthly readers in New York before the law applied to them.
The federal bill protects "a website, digital platform, or online service." This is the gap we flagged in our last blog post about this bill. The bot scraping problem is not unique to media. Any site whose content is worth taking faces a bot that would rather not be seen taking it.
Strava is an example of a site beyond traditional news that has faced this problem. In a recent TechCrunch article, Strava’s CEO described the negative effects that AI scraping was having on their site and mentioned a specific AI company routing their scraping through an aggregator service to hide its identity. Strava is not a news site, but it is still facing the same issue of disguised scrapers stealing their valuable content that news publishers are. Under this federal bill, a company like Strava would be protected the same as news sites.
Complying is technically very easy
The objection tech and scraping companies will raise is that this is a burden. It isn't. Identifying yourself is one of the easiest things a bot can do. It is far less complicated and expensive than the extensive efforts these bots use today to hide themselves.
We know what lengths scraping companies go to remain hidden, because we’ve cataloged it for nearly 40 vendors. The evasion techniques these scraping vendors use today include residential proxy networks that route requests through real homes (the same technique foreign hackers use), rotating IP addresses, user agent spoofing, and they charge up to $22.50 per thousand pages.
All this effort and cost is deployed to avoid detection. Instead of investing in that, this bill would simply require identification. It's a basic ask that the well-behaved bots, like Googlebot, already abide by.
The News/Media Alliance and the Campaign to Stop Bad Bots
The News/Media Alliance (NMA) has been the engine behind the "stop bad bots" push from the start. The NMA helped draft and promote the New York bill alongside the New York News Publishers Association, and it's been out front on the federal version too, with a resource center full of talking points and advocacy tools for anyone who wants to weigh in.
The media industry has already shown strong support for this bill. Executives from News Corp, Hearst Magazines, Condé Nast, and the Tampa Bay Times all publicly supported the bill via a press release from the NMA. In the press release, NMA CEO Danielle Coffey called the bill "a sorely-needed, common-sense solution to a problem plaguing news publishers and other industries across the internet."
What's next, and how to get involved
The wild west scraping era is coming to an end. More government officials are agreeing: bots and agents shouldn't get to mimic humans to take your content, and you deserve to know who's at the door and what they want. A federal bill is a significant step in the right direction.
The bill has been introduced and referred to a committee. To become law, it needs to move through the House and the Senate, so this is still early. What makes it more than a one-off is that it's the leading attempt to pull a scattered set of state efforts into one national standard instead of fifty different ones. A single federal rule is far easier for a site to rely on, and far harder for a bot to route around.
If you want it to pass, the most direct thing you can do is tell your representatives it matters. The NMA’s Stealth Bot Prohibition Act resource center has the talking points and toolkit to help show you how to support.
While this bill goes through the legal process, you can start taking action today. Want to test your site against stealth scrapers, see how they evade your cybersecurity, or monitor what AI traffic is accessing your content? Sign up for TollBit for free.