Meta's AI Crawler Ran Up a $1,312 Vercel Bill on a Store With Zero Customers
meta-externalagent, Meta's AI-training crawler, sent 100,000 requests a day to a 35-product shop nobody uses, pulled 10 TB in three months, kept going 43 days after robots.txt told it to stop, and left me the bandwidth bill. The data, the charts, and how to block it.
On October 6 I opened my Vercel usage page expecting the usual $25 and found $583 in on-demand charges, almost all of it bandwidth. The project responsible was not my production app. It was a storefront prototype I had shelved months ago, with 35 products, a handful of test orders and no customers at all. Its only visitor, at roughly 100,000 requests a day, was Meta's AI-training crawler.
This post is the full accounting: what hit the site, how I traced it, what it cost across three billing cycles, why robots.txt did not stop it, and what I think a crawler at Meta's scale owes the sites it reads. All numbers come from Vercel's usage and firewall dashboards and from the project's git history; the raw daily data sits behind every chart.
A site nobody uses, which makes the measurement clean
The site is lnobeautysupply.com, a Next.js shop I built in July 2026 as a prototype for a beauty-supply line and then abandoned. It holds 35 products, 16 MB of images and a cart that eight test orders ever went through. Nothing links to it. It has never been advertised. Google indexes it and sends a couple of hundred requests a day.
That emptiness is the point. On a real site, crawler cost hides inside human traffic and you argue about attribution. Here there is nothing to subtract. Every gigabyte in the charts below was served to a bot, and 99.4% of the bandwidth on my entire Vercel team was this one project.
Finding it
The team usage page said 4 TB of Fast Data Transfer against a 1 TB allowance. Grouping by project put 4.09 TB on the prototype and 23 GB on the production app that real people use all day.
The project's Firewall → Traffic view for the previous 24 hours then answered the rest:
- 118,000 requests allowed. 115,400 of them to one path,
/shop. - 103,400 from Meta's network, the autonomous system named "Facebook, Inc.", from a block of IPs at 57.141.24.x.
- 17,300 from Alibaba's US cloud, a second crawler that never identifies itself.
- 266 from Google. Googlebot is the only crawler with a reason to be there, and it is the quietest.
The user agent on Meta's requests is worth reading in full, because it is built to look like a browser:
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko)
Chrome/145.0.0.0 Safari/537.36
(compatible; meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/crawler))
It rotates between Windows, macOS, Linux and Edge variants, so a log grouped by user agent shows four "Chrome" entries rather than one bot. The meta-externalagent/1.1 token is there, but at the end, after a full desktop Chrome signature.
View the data as a table
| Network | Requests / 24 h | Share |
|---|---|---|
| Meta (Facebook, Inc.) | 103,400 | 85.1% |
| Alibaba (US) Technology | 17,300 | 14.2% |
| Techoff Srv Limited | 436 | 0.4% |
| Google LLC | 266 | 0.2% |
| CustodianDC Limited | 53 | 0.0% |
What meta-externalagent is, in Meta's words
Meta documents its crawlers on its web crawlers page. Three matter here:
- facebookexternalhit fetches a page when someone shares it on Facebook or Instagram, to build the link preview.
- meta-externalads checks the landing pages behind ads.
- meta-externalagent "crawls the web for use cases such as training foundation AI models or improving products by indexing content directly."
So the bot on my site was not building a preview and not reviewing an ad. It was collecting training data. The same page says Meta's crawlers honor robots.txt, that a disallow for the relevant crawler is how you opt out, and that changes take effect within 24 hours because the file may be cached that long. Keep those three claims in mind.
What it did to the site
The shop page is the trap. /shop is a faceted listing: brand, category, search term, sort and page number all come from the query string, and brand and category are comma-joined multi-selects. Over 35 products that is a combinatorial space of near-identical URLs, and the crawler walked it. I know it was walking the combinations rather than refetching one page, because in August I tried caching the listing per filter combination and measured no effect at all: every request carried a combination nothing had asked for before, so every request missed the cache. Caching the whole 35-row table instead fixed the database load in one commit.
Two properties of the site turned that crawl into money. The page is rendered on demand, so every hit is a serverless function invocation. And each product card carries a 500 KB PNG, which the crawlers pulled alongside the HTML.
View the data as a table
| Day | lno-supply cost | CDN requests |
|---|---|---|
| 2026-07-07 | $0.01 | 5,463 |
| 2026-07-08 | $0.01 | 8,663 |
| 2026-07-09 | $0.00 | 16,735 |
| 2026-07-10 | $0.00 | 54,413 |
| 2026-07-11 | $0.01 | 48,606 |
| 2026-07-12 | $0.04 | 87,215 |
| 2026-07-13 | $0.09 | 118,238 |
| 2026-07-14 | $0.11 | 172,491 |
| 2026-07-15 | $0.10 | 179,141 |
| 2026-07-16 | $0.10 | 176,294 |
| 2026-07-17 | $0.09 | 157,551 |
| 2026-07-18 | $0.07 | 172,683 |
| 2026-07-19 | $0.11 | 180,822 |
| 2026-07-20 | $0.11 | 158,091 |
| 2026-07-21 | $0.10 | 184,907 |
| 2026-07-22 | $0.29 | 234,654 |
| 2026-07-23 | $0.33 | 241,141 |
| 2026-07-24 | $0.88 | 340,420 |
| 2026-07-25 | $1.38 | 362,105 |
| 2026-07-26 | $2.21 | 604,282 |
| 2026-07-27 | $2.25 | 642,948 |
| 2026-07-28 | $2.57 | 770,062 |
| 2026-07-29 | $2.77 | 878,614 |
| 2026-07-30 | $1.19 | 808,940 |
| 2026-07-31 | $1.26 | 866,497 |
| 2026-08-01 | $1.24 | 836,614 |
| 2026-08-02 | $1.29 | 885,668 |
| 2026-08-03 | $1.86 | 1,040,761 |
| 2026-08-04 | $3.45 | 1,107,252 |
| 2026-08-05 | $14.35 | 1,127,459 |
| 2026-08-06 | $26.00 | 1,200,498 |
| 2026-08-07 | $1.54 | 1,216,416 |
| 2026-08-08 | $1.58 | 1,225,725 |
| 2026-08-09 | $1.61 | 1,342,711 |
| 2026-08-10 | $1.65 | 1,361,035 |
| 2026-08-11 | $1.55 | 1,284,234 |
| 2026-08-12 | $1.60 | 1,280,322 |
| 2026-08-13 | $22.77 | 1,030,745 |
| 2026-08-14 | $28.43 | 1,245,393 |
| 2026-08-15 | $27.23 | 1,098,561 |
| 2026-08-16 | $1.67 | 283,587 |
| 2026-08-17 | $24.73 | 1,062,560 |
| 2026-08-18 | $11.86 | 697,448 |
| 2026-08-19 | $1.21 | 373,517 |
| 2026-08-20 | $25.18 | 979,302 |
| 2026-08-21 | $31.77 | 1,184,565 |
| 2026-08-22 | $31.96 | 1,224,640 |
| 2026-08-23 | $31.90 | 1,219,096 |
| 2026-08-24 | $31.42 | 1,198,631 |
| 2026-08-25 | $30.19 | 1,053,600 |
| 2026-08-26 | $31.32 | 1,069,962 |
| 2026-08-27 | $26.80 | 936,222 |
| 2026-08-28 | $28.56 | 960,579 |
| 2026-08-29 | $29.55 | 996,810 |
| 2026-08-30 | $32.03 | 1,088,312 |
| 2026-08-31 | $29.69 | 1,045,873 |
| 2026-09-01 | $28.81 | 1,043,804 |
| 2026-09-02 | $28.59 | 1,025,939 |
| 2026-09-03 | $29.60 | 1,003,989 |
| 2026-09-04 | $30.25 | 1,028,645 |
| 2026-09-05 | $33.97 | 1,104,628 |
| 2026-09-06 | $36.62 | 1,219,231 |
| 2026-09-07 | $2.24 | 1,266,895 |
| 2026-09-08 | $2.50 | 1,284,075 |
| 2026-09-09 | $2.30 | 978,386 |
| 2026-09-10 | $2.56 | 1,273,734 |
| 2026-09-11 | $27.50 | 1,553,180 |
| 2026-09-12 | $47.68 | 1,548,701 |
| 2026-09-13 | $47.60 | 1,530,612 |
| 2026-09-14 | $49.53 | 1,567,566 |
| 2026-09-15 | $50.21 | 1,587,921 |
| 2026-09-16 | $49.24 | 1,588,437 |
| 2026-09-17 | $24.32 | 793,759 |
| 2026-09-18 | $10.48 | 380,832 |
| 2026-09-19 | $9.89 | 491,096 |
| 2026-09-20 | $10.83 | 877,087 |
| 2026-09-21 | $30.18 | 925,690 |
| 2026-09-22 | $52.34 | 1,475,221 |
| 2026-09-23 | $53.97 | 1,445,057 |
| 2026-09-24 | $9.37 | 352,942 |
| 2026-09-25 | $7.31 | 355,262 |
| 2026-09-26 | $9.82 | 403,143 |
| 2026-09-27 | $9.31 | 397,863 |
| 2026-09-28 | $4.86 | 355,226 |
| 2026-09-29 | $5.26 | 354,615 |
| 2026-09-30 | $7.76 | 382,070 |
| 2026-10-01 | $9.60 | 408,965 |
| 2026-10-02 | $9.27 | 401,406 |
| 2026-10-03 | $9.92 | 413,279 |
| 2026-10-04 | $9.25 | 398,979 |
| 2026-10-05 | $6.18 | 339,253 |
| 2026-10-06 | $1.31 | 162,504 |
The timeline, read off the daily cost:
- July 7 to 23. The project costs a few cents a day. This is what a 35-product site with no users should cost.
- July 24. Fast Data Transfer jumps from under 1 GB a day to 17 GB, then 42, then 87. The crawl has started.
- August 5 and 6. $14 and $26 in a day. The Vercel cycle closes at $64 for the project.
- August 13 onward. A steady $25 to $37 every day for the rest of the cycle. The August 7 to September 6 cycle closes at $675.63.
- September 11 to 16. Six consecutive days at 300 GB and about $50 a day.
- September 22 to 23. The peak: 269 and 281 GB, $52.34 and $53.97.
- September 24 to October 6. It eases to 20 to 55 GB and $5 to $10 a day. The cycle closes at $572.61.
View the data as a table
| Day | Fast Data Transfer | CDN requests |
|---|---|---|
| 2026-07-07 | 0 GB | 5,463 |
| 2026-07-08 | 0 GB | 8,663 |
| 2026-07-09 | 0 GB | 16,735 |
| 2026-07-10 | 0 GB | 54,413 |
| 2026-07-11 | 0 GB | 48,606 |
| 2026-07-12 | 0 GB | 87,215 |
| 2026-07-13 | 0 GB | 118,238 |
| 2026-07-14 | 1 GB | 172,491 |
| 2026-07-15 | 1 GB | 179,141 |
| 2026-07-16 | 1 GB | 176,294 |
| 2026-07-17 | 0 GB | 157,551 |
| 2026-07-18 | 0 GB | 172,683 |
| 2026-07-19 | 1 GB | 180,822 |
| 2026-07-20 | 1 GB | 158,091 |
| 2026-07-21 | 1 GB | 184,907 |
| 2026-07-22 | 4 GB | 234,654 |
| 2026-07-23 | 7 GB | 241,141 |
| 2026-07-24 | 17 GB | 340,420 |
| 2026-07-25 | 23 GB | 362,105 |
| 2026-07-26 | 42 GB | 604,282 |
| 2026-07-27 | 47 GB | 642,948 |
| 2026-07-28 | 69 GB | 770,062 |
| 2026-07-29 | 87 GB | 878,614 |
| 2026-07-30 | 82 GB | 808,940 |
| 2026-07-31 | 85 GB | 866,497 |
| 2026-08-01 | 85 GB | 836,614 |
| 2026-08-02 | 100 GB | 885,668 |
| 2026-08-03 | 121 GB | 1,040,761 |
| 2026-08-04 | 140 GB | 1,107,252 |
| 2026-08-05 | 158 GB | 1,127,459 |
| 2026-08-06 | 150 GB | 1,200,498 |
| 2026-08-07 | 150 GB | 1,216,416 |
| 2026-08-08 | 146 GB | 1,225,725 |
| 2026-08-09 | 171 GB | 1,342,711 |
| 2026-08-10 | 182 GB | 1,361,035 |
| 2026-08-11 | 176 GB | 1,284,234 |
| 2026-08-12 | 174 GB | 1,280,322 |
| 2026-08-13 | 144 GB | 1,030,745 |
| 2026-08-14 | 182 GB | 1,245,393 |
| 2026-08-15 | 164 GB | 1,098,561 |
| 2026-08-16 | 9 GB | 283,587 |
| 2026-08-17 | 148 GB | 1,062,560 |
| 2026-08-18 | 68 GB | 697,448 |
| 2026-08-19 | 2 GB | 373,517 |
| 2026-08-20 | 151 GB | 979,302 |
| 2026-08-21 | 191 GB | 1,184,565 |
| 2026-08-22 | 192 GB | 1,224,640 |
| 2026-08-23 | 191 GB | 1,219,096 |
| 2026-08-24 | 187 GB | 1,198,631 |
| 2026-08-25 | 173 GB | 1,053,600 |
| 2026-08-26 | 180 GB | 1,069,962 |
| 2026-08-27 | 153 GB | 936,222 |
| 2026-08-28 | 164 GB | 960,579 |
| 2026-08-29 | 170 GB | 996,810 |
| 2026-08-30 | 184 GB | 1,088,312 |
| 2026-08-31 | 170 GB | 1,045,873 |
| 2026-09-01 | 165 GB | 1,043,804 |
| 2026-09-02 | 164 GB | 1,025,939 |
| 2026-09-03 | 172 GB | 1,003,989 |
| 2026-09-04 | 176 GB | 1,028,645 |
| 2026-09-05 | 199 GB | 1,104,628 |
| 2026-09-06 | 216 GB | 1,219,231 |
| 2026-09-07 | 227 GB | 1,266,895 |
| 2026-09-08 | 231 GB | 1,284,075 |
| 2026-09-09 | 163 GB | 978,386 |
| 2026-09-10 | 244 GB | 1,273,734 |
| 2026-09-11 | 300 GB | 1,553,180 |
| 2026-09-12 | 300 GB | 1,548,701 |
| 2026-09-13 | 299 GB | 1,530,612 |
| 2026-09-14 | 300 GB | 1,567,566 |
| 2026-09-15 | 297 GB | 1,587,921 |
| 2026-09-16 | 291 GB | 1,588,437 |
| 2026-09-17 | 144 GB | 793,759 |
| 2026-09-18 | 56 GB | 380,832 |
| 2026-09-19 | 40 GB | 491,096 |
| 2026-09-20 | 5 GB | 877,087 |
| 2026-09-21 | 137 GB | 925,690 |
| 2026-09-22 | 269 GB | 1,475,221 |
| 2026-09-23 | 281 GB | 1,445,057 |
| 2026-09-24 | 53 GB | 352,942 |
| 2026-09-25 | 39 GB | 355,262 |
| 2026-09-26 | 60 GB | 403,143 |
| 2026-09-27 | 56 GB | 397,863 |
| 2026-09-28 | 19 GB | 355,226 |
| 2026-09-29 | 20 GB | 354,615 |
| 2026-09-30 | 44 GB | 382,070 |
| 2026-10-01 | 55 GB | 408,965 |
| 2026-10-02 | 52 GB | 401,406 |
| 2026-10-03 | 50 GB | 413,279 |
| 2026-10-04 | 50 GB | 398,979 |
| 2026-10-05 | 29 GB | 339,253 |
| 2026-10-06 | 2 GB | 162,504 |
For scale: the entire site, every page and every image, is about 20 MB. Three hundred gigabytes a day is the whole site fifteen thousand times over, every day, for a week.
robots.txt did not stop it
On August 24, while fixing the database load, I added a robots.txt the site had never had:
User-Agent: *
Allow: /
Disallow: /shop?
Disallow: /admin
Disallow: /account
Disallow: /cart
Disallow: /api/
The plain /shop listing stays open, because it is the page that should rank. Every filtered or paginated variant, which is the entire space the crawler was walking, is disallowed. This is standard, well-formed syntax; Googlebot reads it as intended.
Meta says its crawlers honor robots.txt within 24 hours. Forty-three days later, on October 6, meta-externalagent was still sending 100,000 requests a day to that path. The cost chart has a vertical line on August 24 so you can look for the effect yourself. There is none. Of the $1,312 this project has been billed, $1,000 was billed after the rule went in.
I can think of charitable explanations. Perhaps Meta's parser does not treat Disallow: /shop? as a prefix the way the standard and every other major crawler do. Perhaps it only honors groups addressed to it by name, not the wildcard group. Neither is a defense. A crawler that reads a site's robots.txt, finds a rule it does not understand, and resolves the ambiguity by continuing at a hundred thousand requests a day has chosen the wrong default.
The bill
| Billing cycle | Data served | CDN requests | Bandwidth charge | Project total |
|---|---|---|---|---|
| Jul 7 – Aug 6 | 1.22 TB | 13.7M | $33.38 | $64.28 |
| Aug 7 – Sep 6 | 4.92 TB | 32.9M | $586.06 | $675.63 |
| Sep 7 – Oct 6 | 4.11 TB | 25.3M | $463.91 | $572.61 |
| Total | 10.25 TB | 71.8M | $1,083.35 | $1,312.52 |
Bandwidth dominates because of how Vercel prices it: the Pro plan includes 1 TB of Fast Data Transfer a month and bills the rest per gigabyte, at $0.15 per GB for US regions. Four terabytes is about $465. The rest of each cycle is CDN requests ($2 per million past the allowance), function CPU for rendering /shop on demand, and invocations.
Vercel is not the villain here. The pricing is published, the dashboards are what let me find the cause in an hour, and the firewall that will fix it is included in the plan. The prototype was also, by my own earlier accounting, the single largest line on my Neon Postgres bill in August, when it kept the database awake around the clock at 1.8 queries a second. That cost is real but harder to isolate, so it is not in the $1,312.
Where Meta falls short
I want to be precise about the complaint, because "a crawler visited my site" is not one. Crawling is how the web works. The complaint is about proportion, consent and who pays.
Proportion. The site has 35 products. A complete crawl of everything on it is a few hundred requests. Meta sent that many every few minutes, around the clock, for eleven weeks. Google, which runs the largest crawler on earth and actually indexes the store, sends 266 requests a day. There is no crawl budget logic in meta-externalagent that notices it has fetched the same 35 products from a hundred thousand angles.
Consent. Meta's page says robots.txt is how you tell its crawlers what you prefer. I told it. It did not listen, and nothing in the documentation describes a rate limit, a Crawl-delay, or a way to report a runaway crawl beyond a generic webmaster mailbox.
Identification. The user agent leads with a full desktop Chrome string and rotates operating systems. The honest token is present, but anyone filtering logs by browser family will count this bot as four browsers. Meta's documentation also offers no way to verify the crawler's IP ranges; you have to know to look up the AS32934 route object yourself.
Who pays. Every one of those 10 TB was egress billed to me at retail cloud prices. The crawl benefits Meta's models. The invoice goes to the site owner, and the smaller the site, the larger the bill is relative to anything the site earns. This one earns nothing. Across the independent web, a crawler that costs a thousand dollars a quarter per small site is a tax collected by default.
None of this requires malice. It looks exactly like a crawler with no per-site budget, a loose robots.txt parser and no feedback loop, running at the scale of Meta's infrastructure. That is the problem: at that scale, a bug in crawl politeness is a cost imposed on thousands of sites that will never work out why their bill went up.
How to block meta-externalagent
Two layers, because the first one alone did not work for me.
1. A named robots.txt group. Address the crawler by its own token rather than relying on the wildcard group:
User-agent: meta-externalagent
Disallow: /
2. A firewall rule. On Vercel, go to the project's Firewall → Rules, add a rule where the User-Agent header contains meta-externalagent, and set the action to Deny. Cloudflare, AWS WAF and nginx can all do the same match. This is enforced at the edge, costs nothing per request, and does not depend on the crawler's goodwill.
3. Stop paying for the crawl you still get. Make listing pages cacheable (on Next.js, drop force-dynamic or add a short revalidate), serve images as WebP at the size the page actually displays, and put a canonical tag on faceted listings pointing at the plain page.
Blocking meta-externalagent does not affect Facebook or Instagram link previews, which use facebookexternalhit, or ad reviews, which use meta-externalads. If you run Meta ads to the site, leave those two alone.
Recovering the money: every route, and the one I chose
I am not a lawyer, and nothing here is legal advice. It is the list I worked through for my own $1,312.52, with the reasoning, so that the next site owner in this position does not have to start from zero. One thing clears away a lot of fog first: the website has no agreement with Meta. The terms I accepted for a personal Facebook login or an ad account govern my use of Meta's products. They say nothing about Meta's crawler visiting an unrelated site, and the site itself never agreed to anything. So there is no arbitration clause to route around; what is left is ordinary tort law and ordinary courts.
- A demand letter. Before any filing: a dated letter to Meta's legal department and to the crawler team's published mailbox (webmasters@meta.com), with the invoices, the
robots.txthistory and the request logs, asking for reimbursement within 30 days. Cheap, required in spirit by every court, and the point at which most documented small claims against large companies quietly get paid. - A credit from the host. Vercel does not owe me anything; it delivered exactly the bytes it billed. But hosts routinely credit traffic that was plainly abusive, and the ask costs one support ticket. This recovers money without touching Meta at all.
- Small claims court. Up to $7,500 in Colorado, no lawyer needed, a $55 filing fee. The claim is trespass to chattels, the theory courts have applied to crawlers since eBay v. Bidder's Edge in 2000: using someone's server capacity without permission, in a way that causes real harm. Later cases (Intel v. Hamidi, hiQ v. LinkedIn) narrowed it to require actual impairment or cost. I have the cost, itemised to the cent, and a written opt-out that was ignored. This is the route I chose.
- A regular civil suit. County court up to $25,000, district court above that, lawyers on both sides. For $1,312 the fees exceed the claim before the first hearing. Rejected on arithmetic.
- A federal Computer Fraud and Abuse Act claim. The statute allows a civil action once losses pass $5,000 in a year, which, strictly, three months of this would nearly reach. But it requires access "without authorization", and the Ninth Circuit held in hiQ v. LinkedIn that fetching public web pages is not that.
robots.txtis not an access control. Rejected on the merits. - Regulators. A complaint to the FTC and to the Colorado Attorney General's consumer protection office. Neither returns money to me, but both keep a record, and a pattern of complaints is what turns one site's problem into an enforcement matter.
- Collective action. If this crawler behaves the same way on thousands of small sites, the aggregate is a class claim, not a small one. I cannot bring that alone, but I can publish the data and the method, which is this post.
Why Colorado small claims, specifically
Everything the court needs is public, and I checked each item before writing it down.
- Venue exists. Colorado's small claims rule (C.R.C.P. 503) requires filing in a county where the defendant "has an office for the transaction of business". An out-of-state company with no Colorado office cannot be sued in small claims at all. Meta has leased office space at 1900 16th Street in downtown Denver since 2018 and renewed it in 2024, so Denver County Court, small claims division is the venue.
- Meta can be served. Meta Platforms, Inc. is registered with the Colorado Secretary of State as a foreign corporation in good standing (ID 20171961374). Its registered agent for service of process is Corporation Service Company, 1900 W Littleton Blvd, Littleton, CO 80120.
- The amount fits. $1,312.52 plus the filing fee and service costs is well under the $7,500 cap. The filing fee for a claim over $500 is $55, on form JDF 250, "Notice, Claim and Summons to Appear for Trial".
- No lawyers by default. A corporation appears through an officer or full-time employee. If Meta wants counsel it has to file notice first, and then I may bring a lawyer too. Either way the hearing is a trial to a magistrate, no jury, usually within a couple of months.
- The evidence is the kind this court likes. Three invoices, a dated
robots.txtcommit, daily billing data showing the rule changed nothing, Meta's own page promising compliance within 24 hours, and a firewall export naming the user agent and the network. No expert witness is needed to read a bar chart. - Time is not a problem. Colorado gives two years for property-damage and trespass claims (C.R.S. 13-80-102). The first crawl day was July 24, 2026.
- The downside is bounded. If I lose, I am out the filing fee and an afternoon. If Meta does not show up, the court enters a default judgment for the amount claimed.
The order of operations is the demand letter first, the Vercel ticket in parallel, and the JDF 250 filed the day the 30-day deadline passes. I will update this post with what happens at each step.
What I would ask Meta to change
- Honor
robots.txtas written, including wildcard groups and prefix rules with query strings, and say so in the documentation with examples. - Scale the crawl to the site. A crawler that has fetched the same 35 products a thousand times should notice. Googlebot has had adaptive crawl budgets for over a decade.
- Support
Crawl-delayor publish a hard per-host rate limit. - Lead with the bot token in the user agent, as Googlebot and Bingbot do, instead of appending it to a desktop Chrome string.
- Publish a report path for runaway crawls, with a response time, on the crawler page itself.
If you work on this crawler and want the raw logs, the firewall exports or the per-day data behind these charts, email me at the address in the footer. Everything in this post is reproducible from the Vercel dashboards of one account.
Frequently asked questions
What is meta-externalagent? meta-externalagent is Meta's web crawler for, in Meta's own words, "training foundation AI models or improving products by indexing content directly." It is a different bot from facebookexternalhit, which builds link previews when a page is shared on Facebook or Instagram, and from meta-externalads, which checks ad landing pages.
How do I block meta-externalagent?
Add a robots.txt group addressed to it by name (User-agent: meta-externalagent, Disallow: /). In my case that was not enough on its own, so also add a firewall or WAF rule that denies requests whose User-Agent header contains "meta-externalagent". On Vercel that is one rule in Firewall → Rules.
Does blocking meta-externalagent break Facebook and Instagram link previews? No. Link previews are fetched by facebookexternalhit, and ad reviews by meta-externalads. Blocking meta-externalagent only stops Meta's AI-training and indexing crawler.
How much does Vercel charge for bandwidth? On the Pro plan, Fast Data Transfer beyond the plan's allowance is billed per gigabyte at a regional rate, $0.15 per GB for US regions at the time of writing. Four terabytes a month works out to roughly $465 on top of the plan.
Can you sue Meta for the bandwidth its crawler cost you? Possibly, as trespass to chattels: using someone's server capacity without permission in a way that causes measurable cost. The practical venue for a bill this size is small claims court, which in Colorado covers claims up to $7,500 for a $55 filing fee with no lawyer required. Meta has a Denver office, which satisfies Colorado's venue rule, and a registered agent in Littleton who can be served. A demand letter to Meta's legal department and webmasters@meta.com comes first.
Was any of this traffic from real people? No. The site is an abandoned prototype with zero customers. In the 24 hours I measured, Meta's network sent 103,400 requests, an Alibaba cloud crawler sent 17,300, and Google, which actually indexes the store, sent 266.
Methodology. Daily cost and transfer figures are Vercel's own billing data at daily granularity for the three cycles from July 7 to October 6, 2026, read from the team usage dashboard; cost is attributed to the lno-supply project, and the transfer and request series are team-wide totals of which that project was 99.4%. Request-source and user-agent figures are the Vercel Firewall traffic view for the 24 hours ending 3:15 PM Mountain Time on October 6, 2026. Crawler descriptions and the robots.txt claims are quoted from Meta's web crawlers documentation as published on that date.