Metricum Lab

Evidence workbench

SEO Log File Analysis: Verify Googlebot and Find Crawl Gaps

Yurii PekachPublished Updated 16 min read

Define the inventory first. Count verified requests second. Check index inclusion in a different evidence source.

A log export can contain thousands of Googlebot-looking requests and still fail to answer whether your important pages were reached. User-agents can be spoofed; an origin log can miss edge-served responses; a report built only from visited URLs has no room for the missing ones. This guide builds a coverage audit from a fixed page inventory and verified requests. You will choose the capture layer, reproduce a small report, and carry a defensible finding into a crawl or indexation investigation. It is a request-evidence workflow, not a way to infer the Google index from status codes.

Working materials

Script · synthetic CSV · 10 tests · handoff template

Download the Python audit kit

Decide what a log entry can prove

A request log answers a narrow question: what reached this logging layer, and what response did that layer record? It does not show whether Google rendered the page, chose it as canonical, or kept it in the index. Even a 200 response is only evidence of serving, not successful indexing.

Sources:[5] Google Crawling Infrastructure[6] Google Search Console Help

Start with a decision you could change. For a catalog release, that might be whether newly published category pages are being requested at all. For a migration, it might be whether a retired route still receives crawler requests. “Find crawl waste” is too loose: it encourages a large count to become a verdict before you have decided which pages should be crawled.

Use each source for the question it can answer
EvidenceUseful questionWhat remains unknown
Request logsDid a verified request appear in this capture window?Rendering, index inclusion, and missing capture layers.
Crawl StatsDid Google report host problems or a change in crawl composition?A complete per-URL request history.
URL InspectionWhat does the indexed record say, or can the live URL be fetched now?Whether a live test will lead to index inclusion.
Owned URL inventoryWhich URLs were intended to be available in this cohort?Whether Google knows or wants to crawl every URL.

Sources:[3] Google Search Console Help[6] Google Search Console Help

A request is one link in the evidence chain

Keep the requested URL as the join key until you deliberately reconcile it with a canonical URL.

  1. Capture

    One logging layer; explicit UTC window.

  2. Verify

    Client IP snapshot plus the crawler product token.

  3. Join

    Distinct requested URLs versus the owned inventory.

  4. Inspect

    Index and canonical evidence checked separately.

Explanatory diagram. The sequence separates evidence sources; it is not a model of Google’s internal processing time.

Sources:[1] Google Crawling Infrastructure[6] Google Search Console Help

Use Search Console → Settings → Crawl stats as a cross-check, not as the source CSV for this audit. Its URL examples are not exhaustive. A mismatch with your logs is a reason to check scope, dates and capture completeness before calling either dataset wrong.

Sources:[3] Google Search Console Help

Choose the capture layer before counting requests

When a CDN sits in front of the origin, decide which layer represents the question. For crawler-facing responses, use an edge export when it is available and sufficiently complete. In Cloudflare’s HTTP requests dataset, EdgeResponseStatus is the status returned to the client; OriginResponseStatus describes a different leg. Do not concatenate both as though they were separate visits.

Sources:[7] Cloudflare Docs

An origin-only report can still be useful, but label it that way. A URL absent from that report may have been served at the edge. Equally, an edge export with sampling or gaps cannot support a claim that a particular URL received no requests. Keep a small acquisition note alongside the file: host scope, logging layer, UTC interval, sampling policy, missing periods and export filters.

Normalized input for the downloadable audit
FieldRequired meaningPreparation check
timestamp_utcISO 8601 with Z or an explicit offsetUse one documented timestamp convention; do not mix request-start and request-end exports.
client_ipOriginal client address from a trusted logging layerA proxy address or an untrusted forwarded header is not bot evidence.
user_agentComplete user-agent valueKeep the product token; do not pre-label every Google-related agent as Search Googlebot.
urlAbsolute requested HTTP(S) URLPreserve path case, query, scheme and trailing slash. Use the same representation in the inventory.
status / methodHTTP response code and request methodThe audit uses GET for coverage; HEAD is counted separately.
resource_kindhtml, asset, or otherClassify upstream. An extensionless path is not sufficient evidence that a resource is HTML.
request_idOne globally scoped ID for this logging layer, or blankDeduplicate only repeated exports of the same request; leave blank if no trustworthy ID exists.

Sources:[7] Cloudflare Docs

Keep duration fields out of the first coverage pass unless you have verified their units and boundaries. NGINX $request_time, for example, is measured in seconds from the first client bytes to the log write after the last response bytes are sent. Calling it “TTFB in milliseconds” would change both the quantity and its unit.

Sources:[8] NGINX

Prepare a defensible sample

  1. Fix the cohort

    Export intended indexable URLs that were available at the start of the audit window. Give newly published URLs a separate cohort or a later window.

    Expected result: A URL missing before its publication date is not treated as a crawl gap.

  2. Reconcile raw records

    Check one known request from source export to normalized CSV, including timestamp, host, query, status and ID. Document any redaction consistently.

    Expected result: The prepared row still identifies the same request and joins to the intended inventory entry.

  3. Confirm acquisition coverage

    Ask the owner of logging whether the export is sampled, whether a cache layer is omitted, and whether collection failed during the period.

    Expected result: A written scope note accompanies the export; unexplained gaps block claims of zero crawling.

Verify Googlebot without trusting the user-agent

A Googlebot-looking user-agent is a claim made by the client. Google warns that it can be spoofed. For batch work, match the recorded client IP against the published common-crawler ranges and then identify the crawler product from the user-agent. Keep Search Googlebot separate from image/video crawlers, AdsBot and user-triggered inspection tools.

Sources:[1] Google Crawling Infrastructure[4] Google Search Central

The verification documentation checked on 29 September 2026 points to common-crawlers.json. Save a copy with its retrieval date and hash. The audit records the snapshot hash and its creationTime, so a second analyst can identify the same input instead of silently using a newer list.

Sources:[1] Google Crawling Infrastructure

The IP check used by the audit

python · VERIFIED
def matches_ranges(address: str, networks: list[Any]) -> bool:
    ip = ipaddress.ip_address(address)
    return any(ip.version == net.version and ip in net for net in networks)

Extract from log_audit.py; requires ipaddress and typing.Any imports, plus networks loaded by load_ranges(). VERIFIED in the included Python tests against IPv4, IPv6 and a deceptive string-prefix case. This is not a standalone Googlebot classifier.

Sources:[9] Python documentation

For one suspicious address, follow the official reverse-DNS procedure: resolve the IP to a hostname, check the permitted Google domain boundary, resolve that hostname forward, and confirm that the original address is returned. A hostname merely containing the word “google” is not sufficient. Keep the commands and responses with the case rather than replacing the evidence with a boolean.

Sources:[1] Google Crawling Infrastructure

Join verified HTML requests to the URL inventory

Request volume and URL coverage answer different questions. Repeated requests to one page increase volume without expanding coverage. Use a left join from the owned inventory to the filtered requests, keeping URLs with no match. Starting with the log and joining outward makes the absent pages disappear from your report.

For this workflow, observed coverage is distinct eligible URLs with at least one verified HTML GET / all eligible URLs in the fixed cohort. A separate serving measure requires a 200 or 304 response. These are audit definitions, not Google thresholds. A requested page that returned 503 is observed but does not pass the serving check.

The fixture: three URLs observed, two served successfully

Four intended URLs form the denominator. Repeated requests to A do not create extra covered URLs.

  1. 200 / 304A · observed

    Three requests: 200, 304, 200. Serving check passed.

  2. 503B · observed

    One request returned 503. Serving check failed.

  3. 200C · observed

    One request returned 200. Duplicate export removed.

  4. —D · not observed

    No qualifying request in this capture window.

Synthetic dataset shipped with the kit. Observed coverage = 3/4 = 75%; serving coverage = 2/4 = 50%. Neither percentage is an indexing rate.

Keep 304 separate from redirects. It indicates that the cached representation can be reused; lumping every 3xx into a “redirect waste” bucket misclassifies it. At the same time, a 200 can still contain a soft error page, so inspect response content before calling an affected template healthy.

Sources:[5] Google Crawling Infrastructure

Treat exact URL identity as a join contract. Do not strip all query strings or merge trailing-slash variants merely to increase the match rate. If the export and inventory use different encodings or host representations, reconcile them with a documented mapping. Preserve the original URL so a suspicious group can be traced back to individual requests.

Read the pattern before changing crawl rules

My priority order is serving failures, missed important pages, then unwanted URL families. It puts a failing destination ahead of a large but possibly harmless request bucket. The finding should name the page cohort and the competing explanation, not just the largest bar in a chart.

From log finding to a falsifiable next action
FindingCheck before actingNext action and verification
Verified HTML requests receive 5xx or 429Same host, capture layer and time period; edge versus origin behavior.Investigate serving/rate-limiting failures. Recheck affected URLs and subsequent verified responses.
Important URLs are absentPublication date, inventory identity, complete capture, crawlable links and current access rules.Test one discoverability or access hypothesis; compare the same cohort in a subsequent window.
A parameter family dominates requestsDoes it serve a distinct user/search purpose? Is the denominator requests or unique URLs?Decide the URL policy first. Validate both the unwanted family and the valuable destination pages.
Old URLs still receive redirect requestsIntended permanent mapping and whether links still point through old hops.Update avoidable internal hops; retain necessary redirects. Follow the target to a valid response.
A large 304 shareCorrect conditional requests and validators for unchanged content.Do not call it redirect waste. Confirm changed content is no longer treated as unchanged.

Sources:[2] Google Crawling Infrastructure[5] Google Crawling Infrastructure

A reduction in requests to unimportant URLs is not proof that Google will spend the difference on your preferred pages. Google describes both crawl capacity and crawl demand, and explicitly warns against treating temporary robots.txt blocking as a way to reallocate budget. Diagnose the constraint before setting a crawl-volume target.

Sources:[2] Google Crawling Infrastructure

Likewise, noindex is not a crawl-saving instruction: Google must request a crawlable page to encounter it. When the issue is filter URL proliferation, use the separate faceted navigation policy guide. When requests succeed but index inclusion remains the concern, move to the indexation investigation rather than stretching log evidence beyond its scope.

Sources:[2] Google Crawling Infrastructure

Run the audit on the reproducible dataset

The SEO log file analysis kit contains a standard-library Python script, a normalized request CSV, a four-URL inventory, a deliberately synthetic range list, tests and a handoff template. No database account or external service is required. It does not fetch your site or modify crawl controls.

Reproduce the fixture report

bash · VERIFIED
python log_audit.py \
  --logs requests.csv \
  --inventory inventory.csv \
  --ranges fixture-ranges.json \
  --start 2026-09-01T00:00:00Z \
  --end 2026-09-08T00:00:00Z \
  --out results-fixture \
  --allow-synthetic-ranges

python -m unittest -v

Run inside the unpacked kit using Python 3.10+; use python3 if that is your interpreter name. VERIFIED on the supplied fixture in Python 3.13. The output folder must not already exist. The interval includes the start and excludes the end. Synthetic ranges are for the exercise only.

Expected results: 4 eligible URLs, 3 observed URLs, 2 URLs served with 200 or 304. The six verified HTML GETs include one search URL outside the inventory. The asset request is kept out of HTML coverage; the spoofed user-agent, inspection-tool request, duplicate export and out-of-window row do not inflate the result. The HEAD request is reported separately.

Read summary.json before opening coverage.csv. The summary records filters and discarded-row categories; the CSV preserves every inventory URL, its observed and successful request counts, first/last observations and statuses. If the fixture differs, stop and resolve the input or code change. A pleasing percentage is not a substitute for reproducing the known result.

For production, replace all three inputs with your normalized log export, owned inventory and the saved official range snapshot. Remove --allow-synthetic-ranges; choose your real interval and a new output directory. The tool intentionally rejects duplicate inventory URLs, malformed rows and conflicting request IDs instead of silently repairing them. Resolve those issues at the source and rerun.

Validate the intervention on the same page cohort

Before changing anything, write a one-sentence hypothesis. For example: “The new category URLs are linked only after an unavailable interaction, which may explain the absence of qualifying requests in this complete edge window.” The absence is an observation; the explanation is a hypothesis. The next step is to inspect the links and access conditions, not to state that Google has rejected the pages.

Close the finding with an evidence packet

  1. Save the baseline

    Record the log layer, window, range hash, inventory version and affected URLs. Keep the exact script command with the output.

    Expected result: A second analyst can recreate the counts and distinguish “not observed” from “failed response”.

  2. Make one attributable change

    Change the supported cause: a blocked route, an avoidable redirect hop, a serving failure or a missing crawlable link. Keep the previous rule or deployment available for rollback.

    Expected result: Direct checks confirm the intended response or link change; unrelated page families still work.

  3. Repeat and inspect

    Compare the same cohort with a documented follow-up interval. Check request presence and response health, then inspect index/canonical evidence separately where relevant.

    Expected result: The report states what changed, what did not, and what logs still cannot establish.

Do not require equal crawl volume in the before and after periods. New content, demand and site changes can alter the population. Compare response failure rates and coverage with their denominators, and describe any change in eligibility. If a new window is incomplete, mark it incomplete instead of converting its missing requests into an apparent improvement.

The first useful deliverable is a small list of important URLs whose observed behavior conflicts with their intended role. Start there. Logs can show that a serving or discovery problem deserves attention; closing the SEO question may still require rendering checks and index evidence. Keep those conclusions separate in the handoff.

Sources and documentation

Verification dates are listed with the sources. Code examples state the scope of their testing.

  1. Verify requests from Google crawlers and fetchersGoogle Crawling Infrastructure · Checked September 29, 2026
  2. Crawl budget managementGoogle Crawling Infrastructure · Checked September 29, 2026
  3. Crawl Stats reportGoogle Search Console Help · Checked September 29, 2026
  4. GooglebotGoogle Search Central · Checked September 29, 2026
  5. HTTP status codes and network errorsGoogle Crawling Infrastructure · Checked September 29, 2026
  6. URL Inspection toolGoogle Search Console Help · Checked September 29, 2026
  7. HTTP requests datasetCloudflare Docs · Checked September 29, 2026
  8. Module ngx_http_log_moduleNGINX · Checked September 29, 2026
  9. ipaddress — IPv4/IPv6 manipulation libraryPython documentation · Checked September 29, 2026

Need to connect crawl evidence with an implementation fix?

Bring the URL cohort, capture scope and reproducible finding. The next step is to verify the cause and change the right part of the system.

Discuss a crawl and indexation audit