Key points
SEO log file analysis means reading your server’s logs. It shows which pages Googlebot actually visits, how often and with which response code. It is the only source that describes how robots really behave. According to practitioners, it becomes useful beyond about 1,000 pages. Logs are read alongside a crawl and Search Console.
- A user-agent can be faked: verify Googlebot with a reverse DNS lookup or against the IP ranges published by Google.
- The ratio between crawled pages and pages that receive traffic measures how cost-effective the crawl is.
- Logs replace neither Search Console nor your analytics tool: they complement them.
What is SEO log file analysis?
A log file, or access log, records every request received by the server: human visitors, robots, pages, images, scripts. SEO log file analysis filters this log down to search engine robots. It observes their actual path through the site.
A crawler like Screaming Frog simulates a robot. Search Console summarizes what Google is willing to show. Logs, on the other hand, record what happened, request by request. They are the most factual building block of a technical SEO audit.
Definition
SEO log file analysis: the study of a web server’s access logs, filtered on search engine robots. It measures what they crawl, how often and with which responses.
When is log file analysis useful?
The exercise takes time and requires access to the server. It is mainly justified in three cases:
- A site with more than 1,000 pages, or an e-commerce site whose URLs change often. Below that, Search Console is usually enough. This is a practitioners’ benchmark, not a Google rule.
- Important pages that struggle to get indexed, or many URLs marked “Discovered – currently not indexed”.
- A redesign or a migration, to check that Googlebot drops the old URLs and discovers the new ones.
Google’s guide to crawl budget (updated on July 22, 2026) targets two profiles. Sites with about 1 million pages updated every week. Sites with about 10,000 pages whose content changes every day. Google states that these figures are rough estimates, not thresholds. Our article on crawl budget details these notions.
1,000
Screaming Frog Log File Analyser analyzes 1,000 log lines for free on a single project. The license that removes this limit costs £99 a year.
Screaming Frog, official Log File Analyser page, accessed on September 24, 2026.
What does a log file contain?
The most common format is Apache’s “Combined Log Format”, also used by Nginx. The Apache documentation gives this example:
127.0.0.1 - frank [10/Oct/2000:13:55:36 -0700] "GET /apache_pb.gif HTTP/1.0" 200 2326 "http://www.example.com/start.html" "Mozilla/4.08 [en] (Win98; I ;Nav)"
| Field | Example | SEO use |
|---|---|---|
| IP address | 127.0.0.1 | Check that the robot is genuine |
| Timestamp | [10/Oct/2000:13:55:36 -0700] | Frequency and regularity of visits |
| Request | GET /apache_pb.gif | URL crawled, parameters included |
| HTTP code | 200 | What the robot received: 200, 301, 404, 5xx |
| Size | 2326 | Bytes transferred, heavy resources |
| Referer | http://www.example.com/start.html | Page the request came from |
| User-agent | Mozilla/4.08… | Declared robot: Googlebot Smartphone, Bingbot, GPTBot |
Some servers add the response time, which is very useful for spotting slow sections. On Apache, logs are often found in /var/log/. With a shared hosting provider, you download them from the customer area.
Warning
Check the retention period before planning the analysis. Some hosting providers purge logs after a few days or a few weeks. Ask for longer retention several weeks in advance.
How do you analyze logs, step by step?
- Collect the logs from all servers over a representative period, ideally 30 days or more.
- Normalize formats and URLs: protocol, case, trailing slash, parameters, so you can match them with a crawl.
- Filter the relevant robots: Googlebot first, then Bingbot, then AI search engine robots if the subject concerns you.
- Verify that the robots are genuine (see the next section). Fake Googlebots skew every measurement.
- Separate pages and static resources: images and scripts also consume requests, to be measured separately.
- Cross-reference with a full crawl and Search Console data, then segment by page type.
As for tools, Screaming Frog Log File Analyser reads Apache and W3C formats (IIS, Nginx, AWS ELB). It verifies robots automatically. GoAccess, which is free, produces reports from the command line. Oncrawl and Botify offer full platforms, priced on quotation. The Screaming Frog guide explains how to prepare the crawl to cross-reference.
How do you check that Googlebot is genuine?
Any script can present itself as Googlebot. Google documents two methods for verifying its robots (page updated on March 20, 2026):
- Reverse then forward DNS: the
hostcommand on the IP must return a name ingooglebot.com,google.comorgoogleusercontent.com. A forward lookup on that name must return the same IP. - IP ranges: compare the IP with the JSON files published by Google (common crawlers, special-case crawlers, user-triggered fetchers).
These files moved on March 31, 2026, according to the Search Central blog. Google moved them to developers.google.com/crawling/ipranges/. It says the old locations will be redirected within six months. Update your verification scripts.
The list of Google’s common crawlers lets you identify each user-agent (Googlebot Smartphone, Googlebot Image, etc.). OpenAI likewise publishes its robots (GPTBot, OAI-SearchBot, ChatGPT-User) and their IP ranges on its dedicated page. How to manage them is covered in our article on AI search engine robots.
Which indicators should you track in the logs?
| Indicator | How to read it | Typical action |
|---|---|---|
| Distinct URLs crawled | Scope actually crawled | Compare with the number of useful pages |
| Active / crawled pages ratio | Share of the crawl that reaches pages receiving organic traffic | Reduce URLs with no value |
| Crawl frequency by segment | Neglected or over-crawled sections | Internal linking, depth, performance |
| HTTP codes of robot hits | Each 301 or 404 is a wasted request | Fix internal links and redirects |
| Share of Googlebot Smartphone | Consistency with mobile-first indexing | Check mobile parity |
| Response time by segment | Slow sections are crawled less | Server performance |
The most telling comparison sets logs against the crawl. A page found in the crawl but missing from the logs is too deep or poorly linked. A page found in the logs but missing from the crawl is often an orphan. It can also be a remnant of an old version or a junk URL.
What should you do with the results?
The findings translate into actions on the site, rarely on the server alone:
- move rarely crawled strategic pages higher in the site structure and internal linking;
- fix internal links that point to 301s or 404s;
- deal with junk URLs (parameters, filters) that soak up the crawl, then monitor indexing in Google;
- reintegrate or redirect orphan pages that still receive robot visits;
- speed up slow templates, then measure again a month later.
A log file analysis never promises a traffic gain. It establishes what robots do, and therefore the order of fixes. It is part of the broader technical SEO workstream.
Frequently asked questions
How do you analyze logs?
Collect the server logs over at least one month and filter the search engine robots. Verify that they are genuine with a reverse DNS lookup. Then cross-reference the crawled URLs with a crawl and Search Console. A tool like Screaming Frog Log File Analyser automates most of these steps.
How do you read logs?
Each line describes a request: IP address, date and time, requested URL, response code, size, referring page and user-agent. For SEO, you mainly read the URL, the HTTP code and the robot, aggregated by page type.
Does log file analysis replace Search Console?
No. The Search Console Crawl stats report gives an aggregated view, useful but summarized. Logs give the exhaustive detail, without click or position data. The two complement each other.
Should you analyze AI search engine robots?
Yes, if visibility in ChatGPT or other answer engines is a goal. Logs show whether GPTBot or OAI-SearchBot actually access the site. The robots.txt file alone does not guarantee it.
How often should you repeat an SEO log file analysis?
After every structural change (redesign, migration, opening up filters), then at regular intervals on large sites. A one-off SEO log file analysis gives a snapshot; monthly monitoring reveals trends.
Sources
- Google Search Central, Verifying Googlebot and other Google crawlers, updated on March 20, 2026. Accessed on September 24, 2026.
- Google Search Central Blog, New Location for the Google Crawlers’ IP Range Files, published on March 31, 2026. Accessed on September 24, 2026.
- Google Search Central, Crawl Budget Management For Large Sites, updated on July 22, 2026. Accessed on September 24, 2026.
- Google Search Central, Google’s common crawlers. Accessed on September 24, 2026.
- Apache Software Foundation, Log Files, Apache HTTP Server 2.4. Accessed on September 24, 2026.
- Screaming Frog, SEO Log File Analyser, prices displayed on September 24, 2026. Accessed on September 24, 2026.
- OpenAI, Overview of OpenAI Crawlers. Accessed on September 24, 2026.
Cite this article
, . (2026, September 26). SEO log file analysis: method and tools. Elev8 Lab. https://elev8-lab.fr/en/seo/log-file-analysis/