release states in the Page Eligibility Contract
Metricum Lab model: indexable, hold, merge, or do-not-generate. It is an engineering decision aid, not a Google-defined taxonomy.
Technical SEO · Programmatic SEO · Automated QA
Programmatic SEO is safe to scale only when page generation is downstream of explicit product and quality decisions. Treat source data as a versioned contract, decide whether each candidate deserves a standalone search page before rendering it, and make publishing depend on automated QA instead of trusting a template to create value.

The short answer
Programmatic SEO should be designed as an eligibility-controlled publishing pipeline, not as a bulk page generator. Start with a repeatable user/search task and reliable structured data, evaluate every candidate against explicit release rules, render only supported states, and fail the build when technical or value contracts break. Scale the number of pages only after a representative cohort behaves as expected in rendered HTML, crawl evidence and Search Console.
Metricum Lab model: indexable, hold, merge, or do-not-generate. It is an engineering decision aid, not a Google-defined taxonomy.
Google limits one sitemap to 50,000 URLs or 50 MB uncompressed, so large programmatic sections need deterministic sitemap partitioning.
The current per-site URL Inspection quota makes stratified sampling more practical than attempting to inspect every large-scale generated URL through the API.
Stage 1 · Model the problem
A programmatic page family is justified when the same information architecture solves many distinct, real user tasks with page-specific data. A spreadsheet of modifiers is not enough.
Google's current spam policy does not prohibit templates, databases or automation. The policy boundary is purpose and value: scaled content abuse is large-scale generation primarily for manipulating rankings where pages add little or no value, regardless of whether the pages were produced by AI, scripts or people. The doorway-abuse policy separately warns against substantially similar pages created mainly to rank for similar queries and funnel visitors elsewhere. That makes the first design question: what independent task does each generated URL complete?
Describe the repeating task without listing thousands of keywords: for example, “Does product A integrate with product B?”, “What is available in location X?”, or “How do plan A and plan B differ?”
PathSearch demand / product research → page-family hypothesis
Pass criterionA reader can explain why two sibling URLs answer different real tasks rather than the same page with swapped nouns.
List the fields or observations that will materially differ per URL: availability, compatibility, price, specifications, local inventory, first-party usage data, verified attributes, reviews or other supported facts.
Pass criterionEvery candidate page can identify information that is specific to that entity/combination and useful even if Search did not exist.
Specify what the visitor can decide or do on the page itself. If every generated page merely routes users to the same generic tool or category without meaningful intermediate value, re-check doorway risk.
Pass criterionThe page has a useful endpoint: answer, comparison, inventory, workflow or direct next action—not only a search-engine entry point.
Custom diagram
A six-stage system moves from source data to a typed contract, page eligibility, template rendering, automated QA, controlled release and monitoring. Failed candidates loop back instead of becoming public URLs.
Stage 2 · Make the data enforceable
The template should consume validated records, not interpret an unstructured spreadsheet at render time. The contract needs identity, provenance, freshness and relationship rules as well as display fields.
A useful pSEO dataset is more than columns that fit placeholders. I recommend a Data Contract: a typed specification for what each record means, where each field came from, when it becomes stale, which fields are required for a page family, and how records relate to one another. This turns “content quality” from a final editorial feeling into properties that can be checked before a URL exists.
| Contract area | Required question | Example field | Fail behavior |
|---|---|---|---|
| Identity | What exactly is this entity/combination? | entity_id, locale, relation_id | Do not generate ambiguous records |
| User task | Which standalone task does this record answer? | user_job, intent_family | Hold if the task is missing/unvalidated |
| Required facts | Which facts must exist for a useful page? | price, compatibility, inventory, attributes | Hold when mandatory facts are absent |
| Provenance | Where did each material fact come from? | source_url, source_type, observed_at | Hold unsupported high-impact facts |
| Freshness | When should this record be refreshed or retired? | updated_at, expires_at | Hold or retire stale records |
| Relationships | Which other entities are genuinely related? | parent_id, related_ids, category_ids | Do not synthesize arbitrary internal links |
| Eligibility inputs | What evidence decides release? | demand_validated, unique_value, duplicate_of | Evaluate before routing/sitemap generation |
type PageStatus = 'indexable' | 'hold' | 'merge' | 'do-not-generate';
type PageRecord = {
key: string;
userJob: string;
canonicalKey?: string;
requiredFactsComplete: boolean;
hasPageSpecificValue: boolean;
demandValidated: boolean;
sourceFresh: boolean;
duplicateOf?: string;
};
type EligibilityDecision = {
status: PageStatus;
reasons: string[];
};
export function evaluateEligibility(row: PageRecord): EligibilityDecision {
const reasons: string[] = [];
if (row.duplicateOf) {
return { status: 'merge', reasons: [`duplicate of ${row.duplicateOf}`] };
}
if (!row.userJob.trim() || !row.demandValidated) {
return { status: 'do-not-generate', reasons: ['no validated standalone search/user task'] };
}
if (!row.requiredFactsComplete) reasons.push('required facts incomplete');
if (!row.hasPageSpecificValue) reasons.push('page-specific value missing');
if (!row.sourceFresh) reasons.push('source data stale');
if (reasons.length) return { status: 'hold', reasons };
return { status: 'indexable', reasons: ['all release gates passed'] };
}
const sample: PageRecord[] = [
{
key: 'integration/slack',
userJob: 'Understand whether the product integrates with Slack and how',
requiredFactsComplete: true,
hasPageSpecificValue: true,
demandValidated: true,
sourceFresh: true,
},
{
key: 'integration/unknown-tool',
userJob: '',
requiredFactsComplete: false,
hasPageSpecificValue: false,
demandValidated: false,
sourceFresh: false,
},
];
for (const row of sample) {
console.log(row.key, evaluateEligibility(row));
}VERIFIED on Node.js 22.16.0 + TypeScript 5.8.3 on 2026-09-01. The sample compiles and executes. Replace the example fields with your real data contract; the four statuses are Metricum Lab's engineering model, not Google statuses.
Stage 3 · Decide before rendering
Generation and indexability are separate decisions. A candidate can have valid data and still fail to justify its own search page.
The Page Eligibility Contract is the release policy between data and routing. It prevents the common anti-pattern where every row or Cartesian product automatically becomes a public URL and the team tries to clean up thin/duplicate inventory later with canonical or noindex. Google explicitly says indexing is not guaranteed, and its quality guidance asks whether content adds substantial value beyond what is already available. Eligibility is therefore where product usefulness, search intent and duplicate identity should meet.
| Outcome | When to use it | Search behavior | Implementation |
|---|---|---|---|
| Indexable | Distinct task + complete facts + page-specific value + valid identity | Eligible to be discovered and indexed | Create stable route, self-canonical, crawlable internal links, include in intended sitemap cohort |
| Hold | Potentially useful but data/provenance/freshness/QA is incomplete | Not released to Search yet | Keep out of public route/sitemap until failing gates pass |
| Merge | The candidate is materially the same task/content as another canonical record | Consolidate identity instead of multiplying pages | Map to preferred record; redirect existing obsolete duplicate when appropriate |
| Do not generate | No independent task, impossible combination, insufficient data or keyword-only permutation | No search URL should be created | Do not create route, internal link or sitemap entry |
Custom diagram
A decision tree checks standalone user task, required facts, page-specific value, freshness and duplication before routing a candidate to indexable, hold, merge or do-not-generate states.
Stage 4 · Keep the template honest
Conditional components should expose the facts a record actually has. Boilerplate cannot compensate for a thin record, and structured data must describe visible content.
A strong template is mostly an information architecture for page-specific evidence: comparison rows, availability, specifications, examples, calculations, maps, screenshots, reviews or other supported artifacts. It should remove sections when the underlying data is absent rather than generating filler. If content is JavaScript-driven, the primary information should still be available to crawlers through reliable rendering; Google notes that server-side or pre-rendering remains a good idea. Structured data should be generated from the same validated record and must represent what users can actually see on the page.
Define prerequisites for each block. If a comparison requires five fields and the record has three, omit or hold the block/page instead of filling empty space with generic prose.
Pass criterionA sparse record cannot accidentally produce a visually complete but informationally empty page.
For JavaScript frameworks, inspect both the HTTP response and rendered DOM. Do not require a search crawler to execute a user-only interaction just to reveal the primary answer.
Pathcurl/server HTML → rendered DOM → page-specific main content
Pass criterionThe route returns 200 when valid, and the main page-specific content is present in the intended crawler-visible/rendered state.
Do not let a separate structured-data pipeline invent reviews, prices, availability or entities that are not represented on the visible page.
Pass criterionStructured data validates technically and describes visible, current content on that URL.
For JavaScript-heavy page families, use the same server-vs-rendered checks described in the JavaScript SEO debugging guide. Programmatic generation magnifies a rendering bug because one template failure can affect every route in the family.
Stage 5 · Make identity deterministic
At scale, small routing ambiguities become duplicate inventories. URL construction, canonical signals and sitemap membership should all derive from the same identity rule.
Define a deterministic key → URL function and treat it as part of the contract. Google recommends crawlable URL structures, and canonicalization signals can stack: redirects and rel=canonical are stronger signals, while sitemap inclusion is weaker. For a clean programmatic family, internal links, self-canonical and sitemap should consistently point at the preferred indexable URL instead of exposing several spellings or parameter orders for the same record.
The same record and locale must always resolve to the same normalized route. Version migration rules if identity changes.
Pass criterionRepeated builds produce identical URLs for unchanged records; collisions fail the build.
Indexable records should normally declare the intended preferred URL. Merge states should resolve to the preferred record instead of shipping thousands of near-identical self-canonicals.
Pass criterionDeclared canonical, internal links and route identity agree for every sampled record.
Segment sitemaps by page family or cohort so QA and Search Console monitoring can isolate rollout behavior. Split files before Google's 50,000-URL or 50-MB limit.
Pass criterionEvery sitemap URL is indexable by policy, canonical by intent, absolute, and belongs to exactly the expected cohort.
If the page family also exposes combinatorial filters or sort states, keep that URL-space problem separate from the core entity routes. The faceted navigation SEO guide shows how to prevent useful generated landing pages from becoming an uncontrolled filter graph.
Stage 6 · Generate discovery from relationships
A programmatic section needs a browseable graph. Relations in the data contract should determine which pages link to each other and why.
Google uses links to discover pages and recommends that every page you care about have a link from at least one other page on the site. For programmatic systems, that makes internal linking a data-model problem: parent categories, compatible entities, nearby locations, alternatives, previous/next states and other relations should be explicit fields or reproducible rules. Do not create thousands of “related” links solely because two keywords share a token; that produces link noise and can push the architecture toward doorway-like collections.
| Relationship | Link purpose | Generation rule | QA question |
|---|---|---|---|
| Parent / hub | Hierarchy and discovery | Every indexable child links to the canonical hub and the hub links to eligible children | Can a crawler reach the page through normal links? |
| Sibling / alternative | Decision support | Link only when the relation is meaningful to the user task | Would the target still be useful if Search did not exist? |
| Entity relation | Deeper exploration | Derive from verified compatibility/category/location data | Is the relation present in the source data? |
| Editorial guide | Explain concepts the template cannot | Contextual link from page family to stable expert resources | Does the anchor explain why the user should follow it? |
For graph extraction, donor/target prioritization and template-level link rules, use the data-driven internal linking guide. The same logic becomes more valuable as the number of generated routes grows.
Stage 7 · Fail before release
The generator should output a manifest that can be tested before deployment. A failed contract must stop the release rather than create a cleanup ticket for thousands of live URLs.
Preflight QA should inspect the generated result, not only the source rows. A record can pass the data contract yet fail after routing, templating or metadata generation. At minimum, test route uniqueness, eligibility state, HTTP/render expectations, title/H1, canonical, noindex, sitemap membership, crawlable internal inlinks, rendered main content and any structured data required by the page type.
import { readFile } from 'node:fs/promises';
const file = process.argv[2] || 'generated-pages.json';
const pages = JSON.parse(await readFile(file, 'utf8'));
const seenUrls = new Set();
const failures = [];
for (const page of pages) {
const issues = [];
if (page.status !== 'indexable') continue;
if (!page.url?.startsWith('/')) issues.push('invalid URL');
if (!page.title || page.title.length < 10) issues.push('missing/weak title');
if (!page.h1) issues.push('missing H1');
if (!page.canonical || page.canonical !== page.url) issues.push('canonical mismatch');
if (page.noindex) issues.push('indexable page has noindex');
if (!page.inSitemap) issues.push('missing from sitemap cohort');
if (!page.renderedMainText || page.renderedMainText.length < 80) issues.push('main content absent/sparse in rendered HTML');
if (!page.internalInlinks || page.internalInlinks < 1) issues.push('no crawlable internal inlink');
if (seenUrls.has(page.url)) issues.push('duplicate URL');
seenUrls.add(page.url);
if (issues.length) failures.push({ url: page.url, issues });
}
if (failures.length) {
console.error(JSON.stringify(failures, null, 2));
process.exit(1);
}
console.log(`PASS: ${pages.length} generated page records checked`);VERIFIED on Node.js 22.16.0 on 2026-09-01 against a fixture containing passing and failing records. In production, feed this script the manifest produced by your actual generator and extend it with framework-specific SSR/HTML checks.
Custom diagram
A layered gate moves through data validation, eligibility, route identity, rendered HTML, SEO signals and release cohort checks. Failures return to the owning layer instead of being patched page by page.
Stage 8 · Expand observably
A controlled rollout is an engineering practice for catching systemic mistakes. It should not be framed as a secret indexing-speed rule.
The supplied research materials include practitioners describing their own publishing cadence. Treat those numbers as project-specific observations, not platform limits. The current Google sources used for this article define quality, crawl, sitemap and API constraints, but they do not prescribe a universal “N programmatic pages per week” threshold. I would still stage a large release because a template, data or canonical bug can replicate instantly across the whole inventory—and Google does not guarantee that every crawled page will be indexed.
Include dense and sparse records, high- and low-demand variants, edge-case slugs, different relationship counts and all template branches. Do not cherry-pick only perfect rows.
Pass criterionThe cohort exercises every important conditional path and known risk in the page family.
Avoid changing templates, canonicals, internal-link modules, sitemap logic and page copy independently during the same validation window. Keep the experiment attributable.
Pass criterionIf a cohort changes behavior, the team can identify which release changed the system.
Compare generated manifest → production HTTP/rendered HTML → crawl evidence → Search Console sample. If a shared defect appears, stop the next cohort and fix the owning rule.
Pass criterionThe next release is triggered by evidence, not by a calendar quota.
Stage 9 · Measure cohorts
At scale, individual URL checks are samples. Use sitemap segmentation, logs and Search Console to compare page families and release cohorts over time.
Search Console's URL Inspection API can report indexed status for a URL, including canonical information, but the API currently returns the version in Google's index rather than performing a live test. The per-site quota is 2,000 inspections per day and 600 per minute. For a large programmatic inventory, treat inspection as a stratified sample and combine it with sitemap cohorts, Page Indexing/Search Analytics trends, crawler checks and verified search-bot logs where available.
Inspect pages across data density, age, route pattern, locale, demand tier and template branch instead of choosing only top-performing URLs.
Pathrelease manifest → stratified sample → URL Inspection / rendered checks
Pass criterionThe sample represents the conditions that can fail, not just the pages the team expects to succeed.
Track response status, indexability, declared canonical, Google-selected canonical when available, sitemap membership and internal-link reachability as separate fields.
Pass criterionCanonical or indexation anomalies can be grouped by shared route/template/data cause.
Monitor discovery/crawl, index coverage and search impressions/clicks by template and release cohort. Do not interpret “not indexed” as one universal root cause.
Pass criterionA decline or failure can be localized to a page family, release version or eligibility condition before changing the whole site.
Custom diagram
A monitoring loop joins the release manifest with production rendering, crawl/log evidence and Search Console samples, then sends recurring failures back to the data contract, eligibility rules or template instead of hand-editing URLs.
If a healthy technical cohort is still reported as Crawled — currently not indexed, use a separate index-selection diagnosis rather than weakening the eligibility or canonical rules just to increase an index count.
Stage 10 · Keep automation subordinate to evidence
AI can help classify records, draft constrained text and rank hypotheses. It should not invent missing facts, decide search eligibility without evidence, or bulk-change indexation controls without review.
Google's current guidance is consistent on the important distinction: automation and generative AI can assist content production, but generating many pages without adding user value can violate scaled-content policy. For a programmatic system, the safest AI role is bounded: transform approved data, flag anomalies, draft from explicit fields, cluster failed QA rows, or propose falsification tests. The source data and production observations remain the evidence layer.
<context>
You are reviewing a failed programmatic SEO release cohort. The supplied files are evidence; your output is not evidence.
</context>
<goal>
Classify each failed page by the smallest falsifiable cause and recommend the next validation step. Do not rewrite pages automatically.
</goal>
<input>
{{PAGE_ELIGIBILITY_EXPORT}}
{{PREFLIGHT_FAILURES}}
{{RENDERED_HTML_SAMPLE}}
{{CRAWL_OR_LOG_SAMPLE}}
{{GSC_INSPECTION_SAMPLE}}
{{DATA_PROVENANCE_NOTES}}
</input>
<constraints>
1. Separate observed evidence from hypothesis.
2. Never infer search demand from a keyword string alone.
3. Never invent missing source data, product facts, prices, locations, reviews or availability.
4. Do not change canonical, noindex, robots, redirects, schemas or routing without human review.
5. Group failures by shared template/data cause before proposing a patch.
6. Prefer fixing the data contract or generator over hand-editing generated pages.
</constraints>
<required_output>
For each cohort return:
- observed evidence;
- failed gate;
- likely owning layer: data | eligibility | template | routing | SEO signals | monitoring;
- hypothesis;
- falsification test;
- smallest proposed change;
- expected result;
- regression checks;
- remaining uncertainty.
</required_output>
<stop_conditions>
Stop and request human review when evidence is contradictory, provenance is missing, the change affects more than one page family, or the proposed action can publish/redirect/noindex pages in bulk.
</stop_conditions>REUSABLE PROMPT, not evidence. Use only with sanitized exports and keep human approval for bulk routing, canonical, noindex, robots, schema or publishing changes.
| Signal | Why it blocks scale | Next action |
|---|---|---|
| Missing provenance for required facts | The system cannot distinguish a real fact from generated filler | Resolve the source or keep candidates on hold |
| High duplicate/merge rate | The page-family hypothesis may be too granular | Revisit identity and consolidate the model |
| Shared canonical/rendering failure | One template defect can affect the whole inventory | Stop rollout, patch the owning layer, rerun preflight |
| Sparse records dominate failures | The data contract may not support the promised page type | Narrow eligibility or enrich data from approved sources |
| Search behavior differs sharply by cohort | A single global rule is hiding meaningful subtypes | Split the page family and define separate eligibility/QA rules |
Primary sources and documentation
Every changing search, browser, interface or technical-behavior claim in this guide is tied to a current primary source.
Need help diagnosing and implementing the fix?
Metricum Lab can help design or audit the data contract, page eligibility rules, template architecture, indexation governance, automated QA and monitoring for a programmatic SEO system.
Explore scalable SEO website generation