Community Strategy Insights

The latest insights on community strategy, technology, and value by FeverBee’s founder, Richard Millington

Community Platform Overage Fees Are Becoming A Big Problem

Richard Millington
Richard Millington

Founder of FeverBee

Have You Been Hit With An Unexpected Invoice? What Should You Do?

Over the past six months or so, several organizations have contacted me about billing errors from one specific platform vendor. A handful are being charged $100k+ for a single quarter in additional, unexpected bot traffic.

This wasn’t the result of a new campaign, a favorable change in algorithms, or an influx of new customers. It was sharp spikes that arrived without warning and vanished just as fast.

In recent months I’ve reviewed the data of three clients and came to the same conclusion. The traffic which presents itself as human most definitely isn’t.

But this presents a problem. What should you do in this scenario?

Bot Traffic Now Exceeds Human Traffic

Over the past year or so, most communities reached a point where bot traffic exceeded human traffic.

Chart showing bot traffic overtaking human traffic in online communities

I analysed a few of our clients and discovered some have seen bot traffic reach 90%+ of all traffic to their community. As best as I can tell, the majority sit somewhere between 50% and 80%. That’s a staggering increase in usage for the vast majority of communities.

And it comes at the exact time that human visitors are declining because of the collapse of search traffic.

For major sites like Reddit and Stack Overflow, the problem became so extreme they’ve struck paid licensing deals with the AI companies, and Wikipedia’s parent now sells a paid enterprise API for bulk access.

Communities don’t have that leverage. And they don’t want to block everything anyway, because they still want the useful crawlers: Google, so members can find them in search, and the AI engines that cite their sources and send readers back.

What they don’t want is to pay to be scraped by the extractive crawlers that take everything and return nothing.

This is the catch with blocking. It isn’t free. Block the wrong crawler and you lose your search visibility or your place in AI answers. So the goal isn’t to shut bots out, it’s to stop paying for the ones that give you nothing back.

Part of the challenge is that building and running crawlers has become cheaper and easier than ever. On top of that, AI agents that fetch pages live, on demand, have multiplied the number of requests hitting your community. User-triggered crawling grew more than 15x last year alone.

Part of it is just the incentive to grab this data before more organisations lock it down.

Two things are happening here.

  1. The baseline is rising permanently. Every new model generation re-crawls the archives, and agent traffic climbs as more people use AI to answer questions, so the floor under your usage keeps moving up and won’t come back down.
  2. Spikes are becoming a problem. They’re the thing clients keep reporting, and the thing that actually blows the budget. These are sharp, sudden surges of traffic pretending to be human, arriving without warning and pushing them over their usage allowance in a matter of days.

What Are These Bots Exactly?

A bot is essentially just software that requests pages in the same way a person does – but there’s no human reading the other end. For the best part of two decades, they were a good thing. Googlebot would crawl your community, index valuable content, and then surface them when people requested relevant information.

Over the past few years, however, training crawlers like GPTBot, ClaudeBot, and CCBot pulls entire archives in bulk to feed their language models. Worse yet, retrieval and agent bots will fetch pages in real time when someone asks a question, rather than just sending them a link to the potential answer. The numbers vary, but AI might be retrieving 30+ pages to effectively answer a single query.

The Real Scourge Isn’t The Good Bots

But the real scourge here isn’t the clearly labelled LLMs or search bots, although they do make a lot of data requests. The real challenge is the growing share of bots, sometimes as high as 30%, that pretend to be humans (or other bots).

This means they use generic browser identifiers and datacenter connections to look like ordinary visitors.

Their goal is to consume knowledge and information without honouring blocks or, in some cases, being forced to pay Reddit, Stack Overflow, or Wikipedia. They don’t honour robots.txt, because that would defeat their entire objective.

And this is where the real problem begins.

Because if you are on a usage-based platform package, as most of you are, and this dark bot traffic causes you to go over your agreed usage allowance, you may get hit with significant overage fees.

There’s a nasty irony here, and it drives everything below. The bots that are honest about being bots are the easy ones to deal with. It’s the ones that lie about being human that end up on your invoice.

Be fair to the vendors here. Most of them aren’t villains and most want to help. But they get charged for this traffic too, and that charge is the same whether the visitor is a human or a bot. When a vendor passes an overage on to you, it’s often a bill that landed on them first.

Historically, most smart vendors, when confronted with obvious proof that traffic is non-human, waive the overage charge for the benefit of the relationship. The vendor, in essence, eats the cost and tries to do a better job filtering and blocking it in future. They do this for the benefit of the client relationship (or, ideally, because of carefully negotiated contractual terms, more on that later).

But not all vendors are alike.

Today a growing number of organisations are being threatened with legal action (or even closure of their community) if they refuse to pay the overage fee. And this creates a big problem, because it messes with budgets and calls the community’s viability into question.

What Should You Do If You’re Hit With An Overage Charge?

The most important thing to do when you get an overage request is not to rely on the platform vendor’s analytics or tools like Google Analytics. Instead, you need to request access to the raw server logs. This is a reasonable request to make.

The reason is simple. Your analytics and your server logs show very different things. Google Analytics and similar tools run on JavaScript tags, and most crawlers don’t execute JavaScript, so they’re invisible to Google Analytics. Your server logs, however, or your CDN logs, see every single request. Every time a bot requests a page or data, it shows up in the logs.

Separate Out The Expected Bot Traffic

Now within this, there are some things you can count exactly.

GPTBot, ClaudeBot, and CCBot all publish their user agent string, so you can filter the logs for them and have a hard floor.

Most importantly, you can verify them against the published IP ranges.

OpenAI, Anthropic, and Google all publish those, or you can use reverse DNS. That separates the real bots from the ones trying to pretend to be them.

Create A Range For Dark Bot Traffic

The hard part is the undeclared ones. These you infer from behaviour rather than counting them, but there are some very specific tells to look out for.

  1. Request rate. A single IP pulling hundreds of pages a minute isn’t a person.
  2. Hitting deep archive threads no human would navigate to.
  3. Fetching only HTML and skipping the CSS and images a browser would load.
  4. Requests coming from datacenter IP ranges (AWS, GCP, Azure) instead of residential ISPs.

So, for example, if you see a datacenter hammering your old forum posts at machine speed, that’s a bot, even if it claims to be Safari. This gives you an estimate, not a clean number.

An even better route, when available, is to have a CDN (content delivery network) or WAF (web application firewall) classify the traffic for you.

Tools like Cloudflare and Fastly are already pretty good at labelling verified bots, AI crawlers, and traffic that is likely automated, and giving you a breakdown. It’s probably the most credible, independent number you can use.

Then you can compare the metering number the vendor charges you against your human analytics. For example, if your invoice is built on one million sessions and Google Analytics shows 400,000, that gap is your automated traffic. It’s not exact, but it gives you a rough estimate to work with.

Three Terms In The Contract Are Critical Here

In the contract (or Master Services Agreement) there are three things to pay attention to.

  1. What is the usage allowance? This is typically how many page views and API calls are allowed per month or quarter. This number is often in the low millions to begin with, but can rise significantly for larger organisations.
  2. How is usage defined and calculated? What counts as a view? Many contracts define a view in a way that is meant to exclude bots, often with a phrase along these lines: Traffic not detected as originating from a browser is not counted as page views.
  3. What happens when that usage is exceeded? Do you simply move up to the next tier indefinitely? Or do you get billed separately (often excessively) for the overage?

Of the three, it’s the definition of a ‘view’ that matters most, and it’s worth understanding why it doesn’t protect you as well as it looks.

A definition like the one above sounds reassuring. Bots are excluded, so you only pay for humans. But look at what it actually does. The honest, declared crawlers, GPTBot and the like, announce themselves as bots rather than browsers, so they’re easy to detect and get excluded.

The disguised ones, the dark bot traffic wearing a browser identity, and the AI agents like ChatGPT’s and Perplexity’s browsers that literally are browsers rendering the page, are all detected as browsers. So they count.

In other words, the standard definition protects you from the bots that tell the truth and bills you for the bots that lie. Which is exactly backwards, because the lying ones are the fastest-growing share of the problem.

The definition also puts the burden of proof on you. By this wording, it’s on the client to demonstrate a spike is dark bot traffic and not real humans. And that usually requires the server logs. If the vendor won’t give you access to them, you can’t make your case.

The vendor also has leverage. If you don’t pay your bills, they can threaten to close your community down.

Negotiate These Conditions At Contract Signing And Renewal

All of which means the critical thing to do at renewal is to negotiate all three of the terms above.

Before the specifics, one shift matters more than any single clause: what counts as proof. You will almost never be able to demonstrate, request by request, that a spike was automated. What you can produce is an independent CDN or WAF classification. So the real win at renewal is getting the vendor to accept that classification as sufficient evidence. That moves you off an impossible standard, proving every request was a bot, and onto a fair one, where the independent tooling makes the call.

  1. Negotiate what counts as a view. It’s not enough to rely on ‘non-browser’ traffic when the worst bots are pretending to be browsers. You need a fair way to determine whether a spike or increase is bots posing as humans, and to have that traffic excluded. Language worth pushing for looks something like this:

Page views shall exclude any traffic the Customer can reasonably demonstrate, through server logs or CDN classification, to originate from automated or non-human sources, whether or not that traffic presents as a browser. Where a sudden increase in traffic is shown to be predominantly automated, it shall be excluded from usage calculations and shall not, by itself, trigger overage fees.

  1. Secure timely access to server logs. On request, you should be able to obtain the raw server logs within a defined window, say [x] business days. Without this you can’t prove a spike is automated, and the definition above is worthless to you.
  2. Agree what happens when you exceed the allowance. Pin down whether you simply move up a tier or whether you’re billed per unit at a punitive rate. Cap the overage rate, and build in a waiver or review process for spikes that are shown to be automated rather than human.

Summary

Overage invoices from community platforms are rarely a billing error. They’re the result of bot traffic, most of it now non-human, and increasingly of dark bots that disguise themselves as human visitors to get past blocks and avoid paying for what they take.

If you’re hit with one, don’t argue from the vendor’s dashboard or Google Analytics. Get the raw server logs, separate the declared bots you can count from the disguised ones you have to infer, and use a CDN or WAF classification as your independent number.

Then fix it in the contract. Negotiate the usage allowance, the definition of a view, and what happens when you exceed it. These are the three terms that decide who pays for the bots.

Right now, the standard ‘non-browser’ definition has you pay the cost; renewal should be the time to change that.

Subscribe for regular insights

Subscribe for regular insights