+1 (276) 265-7197

When Crawlers Are Your Heaviest User: Taming Bot Traffic on a Legacy PHP App

A client called about a PHP 7.4 catalogue application that had been comfortable on the same 4-core VM for six years and was now returning 502s most afternoons. Nothing had been deployed in five months. Traffic in analytics was flat. The Apache access log told a different story: 71% of requests over the previous week came from automated clients, most of them hitting faceted search URLs that run a six-join query and are never cached.

That shape of problem has become common. Between AI training and retrieval crawlers, SEO tools, price scrapers, uptime checkers, and vulnerability scanners, a lot of older sites now serve more machine traffic than human traffic. The application is not slower than it was. It is just being asked to do far more work, and legacy PHP apps tend to have the exact endpoints crawlers love: parameterised URLs that produce infinite variations of an expensive query.

Before you size up the server, spend an hour with the logs.

Step 1: measure it from the access log

Analytics will not show you this — JavaScript beacons only fire for clients that run JavaScript. The Apache access log is the source of truth. Assuming a combined log format:

# top user agents for one day
awk -F'"' '{print $6}' /var/log/apache2/access.log | sort | uniq -c | sort -rn | head -30

# top client IPs
awk '{print $1}' /var/log/apache2/access.log | sort | uniq -c | sort -rn | head -30

# requests per hour, to find the pattern
awk '{print substr($4,2,14)}' /var/log/apache2/access.log | uniq -c

Two numbers matter. First, the share of requests from non-browser agents — anything with bot, crawler, spider, http, python, curl, or a vendor name in the UA string:

total=$(wc -l < /var/log/apache2/access.log)
bots=$(grep -ciE 'bot|crawler|spider|slurp|python-requests|curl/|wget|scrapy|http-client' /var/log/apache2/access.log)
echo "$bots / $total"

Second, where that traffic lands. Group the bot requests by URL path, ignoring the query string, and then by path with the query string. If the second list is enormously longer than the first, you have found the problem: a handful of endpoints multiplied into millions of URL variants.

grep -iE 'bot|crawler|spider' /var/log/apache2/access.log \
  | awk '{print $7}' | cut -d? -f1 | sort | uniq -c | sort -rn | head -20

Cross-reference with response time if your log format records it (%D in microseconds — add it if it is missing; it costs nothing and you will want it forever). Crawler hits on a cached static page are irrelevant. Crawler hits averaging 1.8 seconds of PHP and MySQL are the whole bill.

Step 2: decide which bots you actually want

This is a business decision, not a technical one, and it is worth ten minutes with the site owner. Three buckets:

  • Wanted. Search engines that send customers, the payment provider's webhook retries, your own monitoring. These get served, but they do not need unlimited crawl rate.
  • Indifferent. SEO tooling, archive crawlers, AI retrieval agents. Whether these are worth serving depends on whether the business wants to appear in those products. Some clients care a great deal; some do not. Ask rather than assume.
  • Unwanted. Scrapers copying your catalogue, credential stuffers, and the constant background scan for /wp-login.php, /.env, /phpmyadmin on a site that has none of those.

Verify before you trust a user agent. Anyone can claim to be Googlebot; the real one resolves. Confirm with a reverse-then-forward DNS check on the client IP before granting it any special treatment.

Step 3: the cheap controls, in order

robots.txt and crawl-delay. Honest crawlers respect it, and the well-behaved majority of your automated traffic is honest. Disallow the URL patterns that explode — faceted search, sort parameters, session-ID URLs, calendar pages that go forward forever:

User-agent: *
Disallow: /search
Disallow: /*?sort=
Disallow: /*?filter=

This one file often removes more load than any server change, because it stops the combinatorial crawl at the source. It does nothing about dishonest clients, which is why it is step one of several.

Cache what crawlers hit. A crawler requesting the same product page as a logged-out visitor should be served from cache, not from PHP. If anonymous responses are cacheable and you are not caching them yet, that is the highest-leverage change available — and it helps human visitors identically.

Rate-limit per client. mod_ratelimit throttles bandwidth, which helps with large file scraping but not with query load. For request-rate control, mod_qos or mod_evasive are the Apache-native options; a reverse proxy or CDN in front is better still if the client already has one. The useful policy is usually not a global cap but a per-IP cap on the expensive paths:

<Location "/search">
  QS_LocRequestLimitMatch "^/search" 4
</Location>

Start permissive, log what would have been blocked, and tighten. A rate limit that blocks the payment provider at month end is a worse outage than the one you were fixing.

Block the clearly unwanted. Deny by user agent or network for scrapers and scanners. Keep the list short and in one file with a comment saying why each entry is there, or it will become unmaintainable folklore that nobody dares touch.

Stop serving 200s to junk. Scanner requests for paths that do not exist should be a fast 404 from Apache, not a PHP bootstrap that loads the framework, opens a database connection, and then renders a 404 page. On many legacy apps the front controller handles every miss, so a scanner sweeping 2,000 paths costs 2,000 full application boots. An Apache-level rule that short-circuits known-bogus paths turns that into nothing.

Step 4: measure again, then decide about the server

Rerun the same log analysis a week later. On the catalogue app above, the numbers went from 71% automated traffic and daily 502s to 34% and none, with no application code changed and no hardware added: a robots.txt that stopped the facet explosion, anonymous page caching, a per-IP limit on /search, and a fast 404 for scanner paths. About a day of work.

Sometimes the honest answer is different. If the traffic is legitimate, wanted, and already efficiently served, then the server really is too small and you should buy a bigger one — that is a fine outcome, and it is worth knowing it is the real reason rather than a guess. But we have not yet seen a legacy LAMP app where measuring the traffic first was the wrong order.

If your access logs are already telling you something like this and you would rather not spend the week on it, send us the numbers — the stack, the versions, and what the log says — and we will tell you whether this is a day of tuning or something larger.