Evidence workbench
SEO Log File Analysis: Verify Googlebot and Find Crawl Gaps
Define the inventory first. Count verified requests second. Check index inclusion in a different evidence source.
A log export can contain thousands of Googlebot-looking requests and still fail to answer whether your important pages were reached. User-agents can be spoofed; an origin log can miss edge-served responses; a report built only from visited URLs has no room for the missing ones. This guide builds a coverage audit from a fixed page inventory and verified requests. You will choose the capture layer, reproduce a small report, and carry a defensible finding into a crawl or indexation investigation. It is a request-evidence workflow, not a way to infer the Google index from status codes.
Script · synthetic CSV · 10 tests · handoff template
Decide what a log entry can prove
A request log answers a narrow question: what reached this logging layer, and what response did that layer record? It does not show whether Google rendered the page, chose it as canonical, or kept it in the index. Even a 200 response is only evidence of serving, not successful indexing.
Sources:[5] Google Crawling Infrastructure[6] Google Search Console Help
Start with a decision you could change. For a catalog release, that might be whether newly published category pages are being requested at all. For a migration, it might be whether a retired route still receives crawler requests. “Find crawl waste” is too loose: it encourages a large count to become a verdict before you have decided which pages should be crawled.
| Evidence | Useful question | What remains unknown |
|---|---|---|
| Request logs | Did a verified request appear in this capture window? | Rendering, index inclusion, and missing capture layers. |
| Crawl Stats | Did Google report host problems or a change in crawl composition? | A complete per-URL request history. |
| URL Inspection | What does the indexed record say, or can the live URL be fetched now? | Whether a live test will lead to index inclusion. |
| Owned URL inventory | Which URLs were intended to be available in this cohort? | Whether Google knows or wants to crawl every URL. |
Sources:[3] Google Search Console Help[6] Google Search Console Help
A request is one link in the evidence chain
Keep the requested URL as the join key until you deliberately reconcile it with a canonical URL.
- Capture
One logging layer; explicit UTC window.
- Verify
Client IP snapshot plus the crawler product token.
- Join
Distinct requested URLs versus the owned inventory.
- Inspect
Index and canonical evidence checked separately.
Sources:[1] Google Crawling Infrastructure[6] Google Search Console Help
Use Search Console → Settings → Crawl stats as a cross-check, not as the source CSV for this audit. Its URL examples are not exhaustive. A mismatch with your logs is a reason to check scope, dates and capture completeness before calling either dataset wrong.
Sources:[3] Google Search Console Help
Choose the capture layer before counting requests
When a CDN sits in front of the origin, decide which layer represents the question. For crawler-facing responses, use an edge export when it is available and sufficiently complete. In Cloudflare’s HTTP requests dataset, EdgeResponseStatus is the status returned to the client; OriginResponseStatus describes a different leg. Do not concatenate both as though they were separate visits.
Sources:[7] Cloudflare Docs
An origin-only report can still be useful, but label it that way. A URL absent from that report may have been served at the edge. Equally, an edge export with sampling or gaps cannot support a claim that a particular URL received no requests. Keep a small acquisition note alongside the file: host scope, logging layer, UTC interval, sampling policy, missing periods and export filters.
| Field | Required meaning | Preparation check |
|---|---|---|
| timestamp_utc | ISO 8601 with Z or an explicit offset | Use one documented timestamp convention; do not mix request-start and request-end exports. |
| client_ip | Original client address from a trusted logging layer | A proxy address or an untrusted forwarded header is not bot evidence. |
| user_agent | Complete user-agent value | Keep the product token; do not pre-label every Google-related agent as Search Googlebot. |
| url | Absolute requested HTTP(S) URL | Preserve path case, query, scheme and trailing slash. Use the same representation in the inventory. |
| status / method | HTTP response code and request method | The audit uses GET for coverage; HEAD is counted separately. |
| resource_kind | html, asset, or other | Classify upstream. An extensionless path is not sufficient evidence that a resource is HTML. |
| request_id | One globally scoped ID for this logging layer, or blank | Deduplicate only repeated exports of the same request; leave blank if no trustworthy ID exists. |
Sources:[7] Cloudflare Docs
Keep duration fields out of the first coverage pass unless you have verified their units and boundaries. NGINX $request_time, for example, is measured in seconds from the first client bytes to the log write after the last response bytes are sent. Calling it “TTFB in milliseconds” would change both the quantity and its unit.
Sources:[8] NGINX
Prepare a defensible sample
Fix the cohort
Export intended indexable URLs that were available at the start of the audit window. Give newly published URLs a separate cohort or a later window.
Expected result: A URL missing before its publication date is not treated as a crawl gap.
Reconcile raw records
Check one known request from source export to normalized CSV, including timestamp, host, query, status and ID. Document any redaction consistently.
Expected result: The prepared row still identifies the same request and joins to the intended inventory entry.
Confirm acquisition coverage
Ask the owner of logging whether the export is sampled, whether a cache layer is omitted, and whether collection failed during the period.
Expected result: A written scope note accompanies the export; unexplained gaps block claims of zero crawling.
Verify Googlebot without trusting the user-agent
A Googlebot-looking user-agent is a claim made by the client. Google warns that it can be spoofed. For batch work, match the recorded client IP against the published common-crawler ranges and then identify the crawler product from the user-agent. Keep Search Googlebot separate from image/video crawlers, AdsBot and user-triggered inspection tools.
Sources:[1] Google Crawling Infrastructure[4] Google Search Central
The verification documentation checked on 29 September 2026 points to common-crawlers.json. Save a copy with its retrieval date and hash. The audit records the snapshot hash and its creationTime, so a second analyst can identify the same input instead of silently using a newer list.
The IP check used by the audit
python · VERIFIEDdef matches_ranges(address: str, networks: list[Any]) -> bool:
ip = ipaddress.ip_address(address)
return any(ip.version == net.version and ip in net for net in networks)Extract from log_audit.py; requires ipaddress and typing.Any imports, plus networks loaded by load_ranges(). VERIFIED in the included Python tests against IPv4, IPv6 and a deceptive string-prefix case. This is not a standalone Googlebot classifier.
Sources:[9] Python documentation
For one suspicious address, follow the official reverse-DNS procedure: resolve the IP to a hostname, check the permitted Google domain boundary, resolve that hostname forward, and confirm that the original address is returned. A hostname merely containing the word “google” is not sufficient. Keep the commands and responses with the case rather than replacing the evidence with a boolean.
Join verified HTML requests to the URL inventory
Request volume and URL coverage answer different questions. Repeated requests to one page increase volume without expanding coverage. Use a left join from the owned inventory to the filtered requests, keeping URLs with no match. Starting with the log and joining outward makes the absent pages disappear from your report.
For this workflow, observed coverage is distinct eligible URLs with at least one verified HTML GET / all eligible URLs in the fixed cohort. A separate serving measure requires a 200 or 304 response. These are audit definitions, not Google thresholds. A requested page that returned 503 is observed but does not pass the serving check.
The fixture: three URLs observed, two served successfully
Four intended URLs form the denominator. Repeated requests to A do not create extra covered URLs.
- 200 / 304A · observed
Three requests: 200, 304, 200. Serving check passed.
- 503B · observed
One request returned 503. Serving check failed.
- 200C · observed
One request returned 200. Duplicate export removed.
- —D · not observed
No qualifying request in this capture window.
Keep 304 separate from redirects. It indicates that the cached representation can be reused; lumping every 3xx into a “redirect waste” bucket misclassifies it. At the same time, a 200 can still contain a soft error page, so inspect response content before calling an affected template healthy.
Treat exact URL identity as a join contract. Do not strip all query strings or merge trailing-slash variants merely to increase the match rate. If the export and inventory use different encodings or host representations, reconcile them with a documented mapping. Preserve the original URL so a suspicious group can be traced back to individual requests.
Read the pattern before changing crawl rules
My priority order is serving failures, missed important pages, then unwanted URL families. It puts a failing destination ahead of a large but possibly harmless request bucket. The finding should name the page cohort and the competing explanation, not just the largest bar in a chart.
| Finding | Check before acting | Next action and verification |
|---|---|---|
| Verified HTML requests receive 5xx or 429 | Same host, capture layer and time period; edge versus origin behavior. | Investigate serving/rate-limiting failures. Recheck affected URLs and subsequent verified responses. |
| Important URLs are absent | Publication date, inventory identity, complete capture, crawlable links and current access rules. | Test one discoverability or access hypothesis; compare the same cohort in a subsequent window. |
| A parameter family dominates requests | Does it serve a distinct user/search purpose? Is the denominator requests or unique URLs? | Decide the URL policy first. Validate both the unwanted family and the valuable destination pages. |
| Old URLs still receive redirect requests | Intended permanent mapping and whether links still point through old hops. | Update avoidable internal hops; retain necessary redirects. Follow the target to a valid response. |
| A large 304 share | Correct conditional requests and validators for unchanged content. | Do not call it redirect waste. Confirm changed content is no longer treated as unchanged. |
Sources:[2] Google Crawling Infrastructure[5] Google Crawling Infrastructure
A reduction in requests to unimportant URLs is not proof that Google will spend the difference on your preferred pages. Google describes both crawl capacity and crawl demand, and explicitly warns against treating temporary robots.txt blocking as a way to reallocate budget. Diagnose the constraint before setting a crawl-volume target.
Likewise, noindex is not a crawl-saving instruction: Google must request a crawlable page to encounter it. When the issue is filter URL proliferation, use the separate faceted navigation policy guide. When requests succeed but index inclusion remains the concern, move to the indexation investigation rather than stretching log evidence beyond its scope.
Run the audit on the reproducible dataset
The SEO log file analysis kit contains a standard-library Python script, a normalized request CSV, a four-URL inventory, a deliberately synthetic range list, tests and a handoff template. No database account or external service is required. It does not fetch your site or modify crawl controls.
Reproduce the fixture report
bash · VERIFIEDpython log_audit.py \
--logs requests.csv \
--inventory inventory.csv \
--ranges fixture-ranges.json \
--start 2026-09-01T00:00:00Z \
--end 2026-09-08T00:00:00Z \
--out results-fixture \
--allow-synthetic-ranges
python -m unittest -vRun inside the unpacked kit using Python 3.10+; use python3 if that is your interpreter name. VERIFIED on the supplied fixture in Python 3.13. The output folder must not already exist. The interval includes the start and excludes the end. Synthetic ranges are for the exercise only.
Expected results: 4 eligible URLs, 3 observed URLs, 2 URLs served with 200 or 304. The six verified HTML GETs include one search URL outside the inventory. The asset request is kept out of HTML coverage; the spoofed user-agent, inspection-tool request, duplicate export and out-of-window row do not inflate the result. The HEAD request is reported separately.
Read summary.json before opening coverage.csv. The summary records filters and discarded-row categories; the CSV preserves every inventory URL, its observed and successful request counts, first/last observations and statuses. If the fixture differs, stop and resolve the input or code change. A pleasing percentage is not a substitute for reproducing the known result.
For production, replace all three inputs with your normalized log export, owned inventory and the saved official range snapshot. Remove --allow-synthetic-ranges; choose your real interval and a new output directory. The tool intentionally rejects duplicate inventory URLs, malformed rows and conflicting request IDs instead of silently repairing them. Resolve those issues at the source and rerun.
Validate the intervention on the same page cohort
Before changing anything, write a one-sentence hypothesis. For example: “The new category URLs are linked only after an unavailable interaction, which may explain the absence of qualifying requests in this complete edge window.” The absence is an observation; the explanation is a hypothesis. The next step is to inspect the links and access conditions, not to state that Google has rejected the pages.
Close the finding with an evidence packet
Save the baseline
Record the log layer, window, range hash, inventory version and affected URLs. Keep the exact script command with the output.
Expected result: A second analyst can recreate the counts and distinguish “not observed” from “failed response”.
Make one attributable change
Change the supported cause: a blocked route, an avoidable redirect hop, a serving failure or a missing crawlable link. Keep the previous rule or deployment available for rollback.
Expected result: Direct checks confirm the intended response or link change; unrelated page families still work.
Repeat and inspect
Compare the same cohort with a documented follow-up interval. Check request presence and response health, then inspect index/canonical evidence separately where relevant.
Expected result: The report states what changed, what did not, and what logs still cannot establish.
Do not require equal crawl volume in the before and after periods. New content, demand and site changes can alter the population. Compare response failure rates and coverage with their denominators, and describe any change in eligibility. If a new window is incomplete, mark it incomplete instead of converting its missing requests into an apparent improvement.
The first useful deliverable is a small list of important URLs whose observed behavior conflicts with their intended role. Start there. Logs can show that a serving or discovery problem deserves attention; closing the SEO question may still require rendering checks and index evidence. Keep those conclusions separate in the handoff.
Sources and documentation
Verification dates are listed with the sources. Code examples state the scope of their testing.
- Verify requests from Google crawlers and fetchersGoogle Crawling Infrastructure · Checked September 29, 2026
- Crawl budget managementGoogle Crawling Infrastructure · Checked September 29, 2026
- Crawl Stats reportGoogle Search Console Help · Checked September 29, 2026
- GooglebotGoogle Search Central · Checked September 29, 2026
- HTTP status codes and network errorsGoogle Crawling Infrastructure · Checked September 29, 2026
- URL Inspection toolGoogle Search Console Help · Checked September 29, 2026
- HTTP requests datasetCloudflare Docs · Checked September 29, 2026
- Module ngx_http_log_moduleNGINX · Checked September 29, 2026
- ipaddress — IPv4/IPv6 manipulation libraryPython documentation · Checked September 29, 2026
Need to connect crawl evidence with an implementation fix?
Bring the URL cohort, capture scope and reproducible finding. The next step is to verify the cause and change the right part of the system.
Discuss a crawl and indexation audit