A published page is not automatically a searchable page. Between publishing and appearing in results, a search engine has to discover a URL, retrieve its content, interpret what it contains and decide how to handle it. Keeping those stages separate makes technical SEO much easier to diagnose.
This guide follows a hypothetical new service page from launch through investigation. The examples are diagnostic workflows, not a promise that every eligible page will be indexed. Start with the exact URL and the intended behavior rather than a broad assumption that “Google cannot see the site.”
Discovery: how the URL enters the picture
A website’s internal links give both visitors and crawlers routes to its pages. An XML sitemap can also identify URLs the site wants search engines to consider. Google describes discovery, crawling and indexing in its Search process documentation. Discovery alone does not establish that a page was fetched, indexed or selected for a query.
Imagine a contractor publishes a detailed service page but links to it only from a confirmation email. The page may be publicly accessible, yet absent from the browsing paths on the site. Before requesting repeated recrawls, add a useful contextual link from the services overview and check that the sitemap contains the intended address.
A useful internal link describes the destination. “Drain inspection service” gives more context than “click here.” Do not add a dense directory of unrelated links just to make a crawl report look complete. Navigation should express a structure people can understand.
Crawling: check the actual response
When investigating a URL, record its response status and any redirects. A successful page normally returns a 200 response. A permanent move should lead to the relevant replacement. A missing page should not present a full error message while pretending to be a successful resource.
Check the final destination as well as the first response. A redirect can point to another redirect, a login screen or an unrelated homepage. Write the chain down before changing it. That simple habit prevents a team from correcting one URL while overlooking the actual failure at the end.
Temporary server failures are a different problem from deliberate exclusions. Compare the timing of failures with deployments, hosting incidents and cache changes. If the page works in your browser, inspect whether that result depends on a logged-in session or a warm cache.
Robots rules are not a privacy system
A robots.txt file controls crawler access for compliant crawlers; it does not protect confidential information. Google also explains that a disallowed URL can still appear in results without its content being crawled. A crawler generally needs access to a page to see its noindex instruction. Review Google’s robots.txt guidance before combining those controls.
For private material, require proper authentication. For a public page that should stay out of search, choose an indexing control appropriate to the content type and verify it. For a page that should appear in search, check both HTML directives and HTTP headers for unintended restrictions.
Rendering: compare what visitors actually receive
Some websites deliver their main text in the initial HTML. Others depend heavily on JavaScript to retrieve and display it. When a page appears incomplete to a crawler, compare the initial response, the rendered view and the version a normal visitor can use.
For a service page, check the service description, location information and primary navigation. If those appear only after an interaction, an API request or a consent choice, document that dependency. A content placeholder and a successful status code do not prove that the meaningful page content was available.
This is also an accessibility and reliability exercise. Content that disappears when a script fails can affect people on unstable mobile connections. Prefer simple delivery for essential information whenever the site’s application requirements allow it.
Indexing: eligible does not mean selected
Once content has been processed, the question becomes how it fits into the search engine’s index. Similar pages, weak differentiation and inconsistent canonical signals can complicate the result. A canonical is a preference signal for duplicate or closely related versions; it should point to the version the site intends to represent that content.
Audit signals together. The sitemap, internal links, redirects and canonical references should not disagree about the preferred address. If all navigation points to one version while the canonical points elsewhere, resolve the underlying publishing rule instead of editing pages one by one indefinitely.
For Google’s treatment of these signals, consult its canonicalization documentation. Do not use canonical tags as a substitute for a clear information architecture.
A focused launch checklist
| Check | What to verify | Evidence to keep |
|---|---|---|
| URL delivery | The intended page returns successfully over HTTPS | Status and final URL |
| Indexing controls | No unintended noindex or access restriction | HTML and response headers |
| Canonical | Preferred URL is deliberate and consistent | Canonical value |
| Discovery | Useful internal links and a valid sitemap entry exist | Linking page and sitemap URL |
| Mobile content | Essential text and links remain available | Rendered mobile inspection |
| Structured data | Any markup matches visible content | Validation result and sample page |
Investigate with a small representative sample
For a large site, do not begin by manually examining every URL. Select examples from each template and each reported problem. A product template, location template and article template may fail for entirely different reasons. Grouping by template helps identify a shared code or publishing issue.
Use Search Console’s URL Inspection for the specific page, then compare its available information with the live behavior you observed. Record the inspection date. A report about an earlier crawl and a live test answer different questions; apparent disagreement may simply reflect a change made between them.
Prioritize by impact and reversibility
Fix an accidental sitewide exclusion before adjusting individual descriptions. Restore broken primary navigation before adding more pages. Correct a template that misstates every canonical before polishing a minor layout issue. State the expected result of each change so the team knows what to recheck.
After deployment, verify the change on a fresh public request and on more than one example page. Keep the old configuration or a rollback route. Technical SEO often involves shared rules, so an apparently small setting can affect a much larger set of URLs.
Move beyond the access problem
If a page is accessible and its signals are coherent, investigate its purpose and distinct value. Another indexing request cannot supply missing substance. Use our search intent framework to check whether the page answers a real need, and the measurement guide to track useful outcomes once it receives visits.
Maintain a short technical log: affected URLs, observed issue, change made, date and next verification. That record turns a one-off investigation into a repeatable process and makes future regressions much easier to spot.

