FR, version française
Home / Blog / SEO / SEO log file analysis: method and tools

SEO log file analysis: method and tools

SEO log file analysis is the only method that shows what Googlebot actually does on a site: pages crawled, frequency of visits, response codes received. It complements the crawl and Search Console, without replacing them. This article explains what a log file contains, when the analysis becomes useful, how to carry it out step by step, how to check that a robot really is Googlebot with the methods documented by Google (IP files moved in March 2026), which indicators to track and how to turn the results into concrete fixes. AI search engine robots, such as GPTBot, are also covered.

Key points

SEO log file analysis means reading your server’s logs. It shows which pages Googlebot actually visits, how often and with which response code. It is the only source that describes how robots really behave. According to practitioners, it becomes useful beyond about 1,000 pages. Logs are read alongside a crawl and Search Console.

  • A user-agent can be faked: verify Googlebot with a reverse DNS lookup or against the IP ranges published by Google.
  • The ratio between crawled pages and pages that receive traffic measures how cost-effective the crawl is.
  • Logs replace neither Search Console nor your analytics tool: they complement them.

What is SEO log file analysis?

A log file, or access log, records every request received by the server: human visitors, robots, pages, images, scripts. SEO log file analysis filters this log down to search engine robots. It observes their actual path through the site.

A crawler like Screaming Frog simulates a robot. Search Console summarizes what Google is willing to show. Logs, on the other hand, record what happened, request by request. They are the most factual building block of a technical SEO audit.

Definition

SEO log file analysis: the study of a web server’s access logs, filtered on search engine robots. It measures what they crawl, how often and with which responses.

When is log file analysis useful?

The exercise takes time and requires access to the server. It is mainly justified in three cases:

  • A site with more than 1,000 pages, or an e-commerce site whose URLs change often. Below that, Search Console is usually enough. This is a practitioners’ benchmark, not a Google rule.
  • Important pages that struggle to get indexed, or many URLs marked “Discovered – currently not indexed”.
  • A redesign or a migration, to check that Googlebot drops the old URLs and discovers the new ones.

Google’s guide to crawl budget (updated on July 22, 2026) targets two profiles. Sites with about 1 million pages updated every week. Sites with about 10,000 pages whose content changes every day. Google states that these figures are rough estimates, not thresholds. Our article on crawl budget details these notions.

1,000

Screaming Frog Log File Analyser analyzes 1,000 log lines for free on a single project. The license that removes this limit costs £99 a year.

Screaming Frog, official Log File Analyser page, accessed on September 24, 2026.

What does a log file contain?

The most common format is Apache’s “Combined Log Format”, also used by Nginx. The Apache documentation gives this example:

127.0.0.1 - frank [10/Oct/2000:13:55:36 -0700] "GET /apache_pb.gif HTTP/1.0" 200 2326 "http://www.example.com/start.html" "Mozilla/4.08 [en] (Win98; I ;Nav)"
FieldExampleSEO use
IP address127.0.0.1Check that the robot is genuine
Timestamp[10/Oct/2000:13:55:36 -0700]Frequency and regularity of visits
RequestGET /apache_pb.gifURL crawled, parameters included
HTTP code200What the robot received: 200, 301, 404, 5xx
Size2326Bytes transferred, heavy resources
Refererhttp://www.example.com/start.htmlPage the request came from
User-agentMozilla/4.08…Declared robot: Googlebot Smartphone, Bingbot, GPTBot
Fields of the Combined Log Format, based on the Apache HTTP Server 2.4 documentation.

Some servers add the response time, which is very useful for spotting slow sections. On Apache, logs are often found in /var/log/. With a shared hosting provider, you download them from the customer area.

Warning

Check the retention period before planning the analysis. Some hosting providers purge logs after a few days or a few weeks. Ask for longer retention several weeks in advance.

How do you analyze logs, step by step?

  1. Collect the logs from all servers over a representative period, ideally 30 days or more.
  2. Normalize formats and URLs: protocol, case, trailing slash, parameters, so you can match them with a crawl.
  3. Filter the relevant robots: Googlebot first, then Bingbot, then AI search engine robots if the subject concerns you.
  4. Verify that the robots are genuine (see the next section). Fake Googlebots skew every measurement.
  5. Separate pages and static resources: images and scripts also consume requests, to be measured separately.
  6. Cross-reference with a full crawl and Search Console data, then segment by page type.

As for tools, Screaming Frog Log File Analyser reads Apache and W3C formats (IIS, Nginx, AWS ELB). It verifies robots automatically. GoAccess, which is free, produces reports from the command line. Oncrawl and Botify offer full platforms, priced on quotation. The Screaming Frog guide explains how to prepare the crawl to cross-reference.

How do you check that Googlebot is genuine?

Any script can present itself as Googlebot. Google documents two methods for verifying its robots (page updated on March 20, 2026):

  1. Reverse then forward DNS: the host command on the IP must return a name in googlebot.com, google.com or googleusercontent.com. A forward lookup on that name must return the same IP.
  2. IP ranges: compare the IP with the JSON files published by Google (common crawlers, special-case crawlers, user-triggered fetchers).

These files moved on March 31, 2026, according to the Search Central blog. Google moved them to developers.google.com/crawling/ipranges/. It says the old locations will be redirected within six months. Update your verification scripts.

The list of Google’s common crawlers lets you identify each user-agent (Googlebot Smartphone, Googlebot Image, etc.). OpenAI likewise publishes its robots (GPTBot, OAI-SearchBot, ChatGPT-User) and their IP ranges on its dedicated page. How to manage them is covered in our article on AI search engine robots.

Which indicators should you track in the logs?

IndicatorHow to read itTypical action
Distinct URLs crawledScope actually crawledCompare with the number of useful pages
Active / crawled pages ratioShare of the crawl that reaches pages receiving organic trafficReduce URLs with no value
Crawl frequency by segmentNeglected or over-crawled sectionsInternal linking, depth, performance
HTTP codes of robot hitsEach 301 or 404 is a wasted requestFix internal links and redirects
Share of Googlebot SmartphoneConsistency with mobile-first indexingCheck mobile parity
Response time by segmentSlow sections are crawled lessServer performance
Elev8 Lab grid (SEO knowledge base verified on September 9, 2026). Thresholds to set for each site: Google does not publish any.

The most telling comparison sets logs against the crawl. A page found in the crawl but missing from the logs is too deep or poorly linked. A page found in the logs but missing from the crawl is often an orphan. It can also be a remnant of an old version or a junk URL.

What should you do with the results?

The findings translate into actions on the site, rarely on the server alone:

  • move rarely crawled strategic pages higher in the site structure and internal linking;
  • fix internal links that point to 301s or 404s;
  • deal with junk URLs (parameters, filters) that soak up the crawl, then monitor indexing in Google;
  • reintegrate or redirect orphan pages that still receive robot visits;
  • speed up slow templates, then measure again a month later.

A log file analysis never promises a traffic gain. It establishes what robots do, and therefore the order of fixes. It is part of the broader technical SEO workstream.

Frequently asked questions

How do you analyze logs?

Collect the server logs over at least one month and filter the search engine robots. Verify that they are genuine with a reverse DNS lookup. Then cross-reference the crawled URLs with a crawl and Search Console. A tool like Screaming Frog Log File Analyser automates most of these steps.

How do you read logs?

Each line describes a request: IP address, date and time, requested URL, response code, size, referring page and user-agent. For SEO, you mainly read the URL, the HTTP code and the robot, aggregated by page type.

Does log file analysis replace Search Console?

No. The Search Console Crawl stats report gives an aggregated view, useful but summarized. Logs give the exhaustive detail, without click or position data. The two complement each other.

Should you analyze AI search engine robots?

Yes, if visibility in ChatGPT or other answer engines is a goal. Logs show whether GPTBot or OAI-SearchBot actually access the site. The robots.txt file alone does not guarantee it.

How often should you repeat an SEO log file analysis?

After every structural change (redesign, migration, opening up filters), then at regular intervals on large sites. A one-off SEO log file analysis gives a snapshot; monthly monitoring reveals trends.

Sources

  1. Google Search Central, Verifying Googlebot and other Google crawlers, updated on March 20, 2026. Accessed on September 24, 2026.
  2. Google Search Central Blog, New Location for the Google Crawlers’ IP Range Files, published on March 31, 2026. Accessed on September 24, 2026.
  3. Google Search Central, Crawl Budget Management For Large Sites, updated on July 22, 2026. Accessed on September 24, 2026.
  4. Google Search Central, Google’s common crawlers. Accessed on September 24, 2026.
  5. Apache Software Foundation, Log Files, Apache HTTP Server 2.4. Accessed on September 24, 2026.
  6. Screaming Frog, SEO Log File Analyser, prices displayed on September 24, 2026. Accessed on September 24, 2026.
  7. OpenAI, Overview of OpenAI Crawlers. Accessed on September 24, 2026.
claude-editeur Avatar

Digital marketing, SEO and GEO consultant

More about the author

Article checked and updated by the author. Sources consulted on the date shown.

Cite this article

, . (2026, September 26). SEO log file analysis: method and tools. Elev8 Lab. https://elev8-lab.fr/en/seo/log-file-analysis/