# Technical SEO Checklist 2026: 13 Checks to Get Indexed

> A technical SEO checklist for 2026: 13 checks from crawl budget to AI crawler access, each with the exact tool and fix to get your site fully indexed.

- URL: https://missiongrowth.io/blog/technical-seo-checklist-2026
- Published: 2026-07-04 · Updated: 2026-09-19
- Author: Furkan Aktaş, Co-Founder, Mission Growth
- Publisher: Mission Growth. Company facts: https://missiongrowth.io/llms.txt

A technical SEO checklist for 2026 runs thirteen checks, from HTTP status codes up to AI crawler access. Each one decides whether a crawler reaches a page and whether the index keeps it once it does.

Trust and ranking come later, and only for pages that already cleared those gates.

::dataset{key="checklist-areas-2026" name="Technical SEO checklist areas and primary tools, 2026"}

| Area | Check | Primary tool |
|---|---|---|
| HTTP status & HTTPS | Redirect chains, soft 404s, HTTP-to-HTTPS coverage | Screaming Frog, Search Console, curl |
| Robots.txt | Directive errors and 5xx failures | curl, Search Console |
| Indexability | noindex tags, canonical conflicts, coverage status | Screaming Frog, Search Console |
| Duplicate content | Canonical consolidation and signal conflicts | Screaming Frog, Search Console |
| XML sitemap | Size limits, parse errors, lastmod accuracy | curl, xmllint, Search Console |
| URL structure | Fragment routing, parameters, faceted navigation | Screaming Frog |
| Crawl budget | Crawl efficiency and bot traps | Search Console, server logs |
| Internal linking | Orphan pages and click depth | Screaming Frog |
| Mobile parity | Content, link, and schema parity | Search Console, dual-user-agent crawl |
| JavaScript rendering | Raw HTML versus rendered DOM | Screaming Frog, URL Inspection |
| Core Web Vitals | LCP, INP, CLS, TTFB | PageSpeed Insights, Search Console |
| Structured data | Schema validity and mobile parity | Rich Results Test, Search Console |
| AI crawler access | Raw HTML response to bot user agents | curl |

::figure{src="/blog/figures/technical-seo-checklist-2026-1.svg" alt="The thirteen technical SEO checklist sections group into three tiers: foundation, standard, and advanced, ordered by which gates block everything else." caption="The checklist puts foundation-tier fixes like status codes, robots.txt and canonical consolidation ahead of standard and advanced items such as Core Web Vitals and AI crawler access." width="720" height="396"}

Most teams start with content and assume the technical layer sorts itself out. It doesn't. On a brand-new domain the order is stricter still: [this startup SEO guide](https://missiongrowth.io/blog/startup-seo) covers which technical items gate a first launch.

A page with a broken canonical tag or a 3.2s LCP never gets the chance to rank: it never gets fully indexed.

That order matters: you fix the gates that block everything before you spend a day on schema.

::figure{src="/blog/figures/technical-seo-checklist-2026-2.svg" alt="Crawl, index, and render are the three stages that every technical SEO checklist section maps back to, from HTTP status to JavaScript rendering." caption="Every checklist item intervenes at one of three stages: crawl (HTTP status, robots.txt), index (canonicals, noindex), or render (JavaScript rendering, which can break both)." width="720" height="252"}

One mental model holds the whole checklist together. A search engine crawls a URL, decides whether to index it, and for JavaScript sites, renders it somewhere in between.

Every check here works on one of those three stages. Robots.txt and HTTP status govern the crawl. Canonicals and noindex govern the index. JavaScript rendering sits in the middle, and it quietly breaks both.

Keep that spine in mind: each section names its stage.

> [!TAKEAWAY]
> Crawl and index gates (1-6) decide whether a page is even in the running. Crawl budget, links, and rendering (7-10) decide how fast it gets there. Core Web Vitals, schema, and AI crawler access add ranking and citation upside once the foundation holds.

## 1. HTTP status codes, redirects, HTTPS and soft-404s

HTTP status codes, redirects, and HTTPS decide whether a crawler can even reach your content, before ranking is a question.

Crawl your site with Screaming Frog and open the Response Codes report. Cross-check Google Search Console's Pages report and spot-check individual URLs with `curl -I`.

You're hunting four things:

1. **Broken internal links.** They leak link equity into dead ends.
2. **Soft 404s.** A 200 status with no real content, so Google may swap in a better URL for the query.
3. **Redirect chains longer than one hop.** Every extra hop burns crawl budget and weakens the PageRank passed through.
4. **302s on permanent moves.** They keep Google canonicalizing the old URL, so link equity never fully transfers.

The fix is mechanical:

- Collapse every redirect chain to a single hop (A to C, not A to B to C).
- Change permanent 302s to 301s.
- Serve a real 410 for content you removed for good. A 410 drops from the index faster than a 404.
- Confirm every HTTP URL 301-redirects to its HTTPS version, and add an HSTS header.

This stopped being optional maintenance. Google confirmed in an October 2025 Security Blog post that Chrome will warn before loading public HTTP sites by default.

That starts with Chrome 147 in April 2026, for the billion-plus Enhanced Safe Browsing users, and reaches everyone with Chrome 154 in October 2026. Stray HTTP URLs become a visible trust problem this year.

## 2. Robots.txt and crawl directives

Robots.txt is the front door Googlebot reads before every crawl request, so one wrong line can wall off your whole site.

Fetch it with `curl -sI https://yourdomain.com/robots.txt` and confirm a 200. Then open the robots.txt report in Search Console to see the version Google actually cached.

Two failure modes matter most:

1. **A 5xx on the robots file.** Google treats it as a temporary full disallow and stops crawling. The Web Almanac 2024 found about 84% of sites return a clean 200, so a non-200 status is a rare, fixable problem worth chasing down fast.
2. **Blocked CSS and JavaScript.** Google can't render the page without them, so JavaScript-dependent content drops from the index. Check this against the render test in section 10.

The single most common misuse is treating `Disallow` as a way to remove a page from Google. It isn't:

> [!WARNING]
> `Disallow` blocks crawling; it does not block indexing. A blocked URL that earns an external link can still appear in results as a bare link with no snippet, because Google never crawls it to see any noindex tag. To actually deindex a page, remove the Disallow so Google can crawl it and see the tag, then add a `noindex` directive.

Google also ignores `Crawl-delay`, so anyone relying on it for server load is protecting nothing.

## 3. Indexability: noindex, canonical and GSC index coverage

Indexability separates a page Google crawled from a page Google actually keeps in the index.

Run Screaming Frog's Directives tab to list every `noindex` and `X-Robots-Tag`. Then open Search Console's Pages report and read the "Why pages aren't indexed" breakdown.

The distinction that matters is between two similar-sounding statuses, and they need opposite fixes:

::dataset{key="index-coverage-status-fixes" name="Google Search Console index coverage statuses and fixes, 2026"}

| GSC status | What Googlebot did | Root cause | Fix |
|---|---|---|---|
| Crawled, currently not indexed | Fetched the page, chose not to index it | Quality: thin, duplicate, or low-value content | Improve or consolidate the page, or add a deliberate `noindex` |
| Discovered, currently not indexed | Knows the URL, has not fetched it yet | Crawl budget: too many low-value URLs ahead of it in the queue | Cut crawl waste, add internal links so the page rises in priority |
| Excluded by noindex tag | Saw a noindex directive | Intentional, or a stray tag on a page you want indexed | Confirm each one is deliberate, remove any stray noindex |
| Duplicate, Google chose different canonical | Ignored your canonical tag | Conflicting signals across tag, sitemap, and links | Align every signal to one URL (see section 4) |

"Crawled, not indexed" is a content problem. "Discovered, not crawled" is a crawl budget problem. Fix them the same way and you'll waste your week.

Also watch for the trap where a page carries both a `noindex` and a `canonical` to another URL. Google ignores the canonical on a noindexed page, so pick one intent.

## 4. Duplicate content and canonical consolidation

Duplicate handling decides which URL competes and where link equity pools.

Crawl with Screaming Frog and check the Canonicals report for pages whose canonical tag doesn't point at itself. Then look in Search Console for "Duplicate, Google chose different canonical than user." That warning means Google overrode your tag, which almost always traces back to a signal conflict.

Google treats `rel=canonical` as a hint it can override. It weighs your tag against your sitemap, internal link counts, and backlinks. The fix is alignment: point the canonical tag, the sitemap entry, your internal links, and any redirects at the same single URL.

When GSC reports "Google chose a different canonical than user," diagnose it in order:

1. Check URL Inspection to see which URL Google actually picked.
2. Check which signal it's following: the sitemap entry, the majority of internal links, or external backlinks.
3. Fix that signal to point at your preferred URL before you touch the canonical tag again.

Consolidate the obvious duplicates too. Make sure `http` redirects to `https`, `www` and non-`www` resolve to one host, and trailing slash variants don't both return 200.

Two specifics for 2026. Pagination no longer uses `rel=next`/`rel=prev`, since Google deprecated it in 2019, so every paginated page should carry a canonical that points at itself.

And as an internal check, treat any page that shares most of its wording with another of your pages as a duplicate risk.

Google will pick one and drop the rest. Programmatic and template pages are the usual offenders, so measure content overlap, excluding boilerplate, before you publish a large batch of them. On sites with thousands of templated pages this turns into a crawl budget question as well; [enterprise SEO](https://missiongrowth.io/blog/enterprise-seo) covers where that threshold sits.

## 5. XML sitemap health

Your XML sitemap is the clean list of URLs you want crawled, plus a freshness signal.

Fetch it with `curl`, validate the XML with `xmllint --noout sitemap.xml`, and check the Sitemaps report in Search Console for parse errors and the gap between submitted and indexed URLs.

Respect the hard limit the sitemap protocol and Google set: a single sitemap holds a maximum of 50,000 URLs and 50MB uncompressed. Past either ceiling, split into a sitemap index that references child sitemaps by category or content type.

Keep the file accurate:

- List only canonical URLs that return 200. No noindex pages, no redirects, no 404s: every dirty URL in the sitemap is a crawl instruction Google can't act on.
- Make `lastmod` mean something. Update it when the content genuinely changed, not on a CSS tweak, or Google learns to distrust it and crawls the file less often.

When submitted greatly exceeds indexed, don't blame the sitemap first. Filter each URL by its Coverage category instead. The real cause is usually content quality, a duplicate cluster, or an orphan page with no internal links, and each lives in a different section of this checklist. For the order to run all of it in, from measurement to authority, follow [the full SEO audit process](https://missiongrowth.io/blog/how-to-do-seo-audit).

## 6. URL structure, parameters and faceted navigation

Messy URLs multiply the number of pages Google thinks you have.

Crawl with Screaming Frog and review the URL and Parameters reports. Then check Search Console for a swollen "Discovered, currently not indexed" count, which often signals a URL explosion the crawler can't keep up with.

Three checks carry the weight:

1. **Fragment routing.** A route after a hash fragment, like `/#/product/123`, is invisible to Google. Route with the History API instead, so it serves a clean path like `/product/123`.
2. **Parameters that don't add value.** Sort, color, and tracking parameters create duplicate URLs. Point them at the canonical version with a `rel=canonical` tag.
3. **Faceted navigation.** Filter combinations breed thousands of thin URLs. Canonicalize the ones with no unique value and disallow the rest in robots.txt.

Standardize the boring details across the whole site too: one trailing slash convention, all lowercase paths, hyphens instead of underscores (Google reads `seo-guide` as two words and `seo_guide` as one), and short slugs that carry the keyword.

None of these move rankings on their own. But each variant you leave loose is another duplicate URL splitting your signals.

## 7. Crawl budget and log file analysis

Server log analysis is the check that separates a real technical audit from a plugin scan.

Search Console coverage tells you what Google says it indexed. Your server logs tell you what Googlebot actually requested, and the two disagree more often than teams expect.

Pull the Crawl Stats report in Search Console for the overall trend. Then parse raw access logs with pandas or a tool like Screaming Frog's Log File Analyser for the detail.

Crawl efficiency is the core metric: the share of Googlebot requests that returned a 200, calculated as 200-responses divided by total Googlebot requests. We flag anything below roughly 85% as a sign budget is leaking to low-value URLs, 404s, redirects, and 5xx errors instead of your content.

From the same logs, run three checks:

1. **Group requests by URL section.** Compare crawl share by `/blog/`, `/product/`, `/tag/` against value. If tag or parameter URLs out-pull product pages, fix robots.txt first.
2. **Sort by daily hit count.** A URL crawled ten or more times a day usually means a session ID, an infinite paginator, or a redirect loop.
3. **Confirm the hits are real.** Run a reverse DNS lookup (`host -t PTR <ip>`) and check that it resolves to a genuine googlebot.com hostname; a spoofed user agent skews every number above.

This is the unglamorous discipline behind results, the opposite of a growth hack. The same crawl and content rigor drove [Pozitif Teknoloji's +225K organic clicks in six months](https://missiongrowth.io/case-studies/pozitif-teknoloji), and it starts with reading the logs instead of guessing.

## 8. Internal linking and site architecture

Internal links are both the crawler's map and the pipes that move PageRank around your site.

Crawl with Screaming Frog and read two reports: Inlinks, to find orphan pages with zero internal links, and Crawl Depth, to find pages buried too far from the homepage.

Fix the structural faults in order:

1. **Orphan pages.** A page with zero inlinks can't be discovered and earns no internal authority. Link to it from a relevant hub.
2. **Click depth over three.** Pages four or more clicks deep get crawled less often, so updates take longer to register. Keep important pages within three clicks of the homepage.
3. **Broken internal links.** A 4xx target is a dead end that blocks equity flow. Clear every one you find.
4. **Links that only render via JavaScript.** They enter the crawl queue late, if at all, starving deeper pages of discovery. Keep hub links as real `<a href>` elements in the static HTML.

## 9. Mobile-first parity

Google indexes the mobile version of your pages, so anything present only on desktop doesn't count.

Confirm parity with the URL Inspection tool in Search Console, which shows the rendered mobile HTML Googlebot Smartphone actually sees. Back it up with a crawl in Screaming Frog using both user agents: once as Googlebot Smartphone, once as Googlebot Desktop, then diff the two.

Check three things for parity:

1. **Content and links.** A heading hidden on mobile doesn't get indexed. A desktop-only internal link breaks equity flow and can orphan the target.
2. **Structured data.** Markup missing from the mobile HTML makes the page ineligible for the matching rich result.
3. **Viewport tag.** Confirm every template carries a `<meta name="viewport" content="width=device-width, initial-scale=1">`. Without it, mobile rendering breaks and page experience signals suffer.

Parity gaps are almost always at the template level, so fix the template once and it propagates across every page built from it.

## 10. JavaScript rendering SEO and render parity

JavaScript rendering SEO is the practice of shipping content so Google and AI crawlers see the same page your browser does, before any script runs.

Google processes a JavaScript page in three phases: crawl, then a separate render queue, then indexing. Its own documentation puts it plainly: rendering "may stay on this queue for a few seconds, but it can take longer than that."

A fix that depends on JavaScript, a noindex tag added by script, a canonical set after render, only applies once that queue clears.

GPTBot, ClaudeBot, PerplexityBot and the other major AI crawlers skip the queue entirely. In Vercel's December 2024 measurement none of them executed JavaScript, so they read the raw HTML as if the fix never happened.

Every indexing directive that matters has to live in that raw HTML response. Whatever JavaScript builds afterward arrives too late for a crawler that never renders.

### Where Googlebot and AI crawlers read the same fix differently

Four common JavaScript fixes land differently depending on which crawler reads them:

| Directive | Googlebot (after render) | AI crawlers (raw HTML only) | What to do |
|---|---|---|---|
| noindex added via JavaScript | Sees and applies it once rendering finishes | Never sees it; treats the page as indexable | Add `noindex` in the initial HTML response |
| Canonical set via JavaScript | Applies it if it matches the HTML value | Reads whatever canonical is already in the raw HTML | Keep the JS canonical identical to the HTML one, or set it in the initial response |
| Soft-404 fixed with a JS redirect | Follows the redirect after render | Never redirects; reads the original page as real content | Return the actual HTTP status or redirect in the initial response |
| Content injected client-side | Indexes it once rendered | Never sees it | Render the content on the initial response, not after load |

Get the raw HTML right, and every row resolves the same way for every crawler.

### How much AI crawl traffic can't render at all

GPTBot, ClaudeBot and PerplexityBot account for 75.4% of the four AI-linked crawlers' combined monthly volume, per Vercel's December 2024 network data.

That's roughly 3 to 1 against AppleBot, the one crawler in that group that renders like Googlebot does (full breakdown in section 13).

None of that volume executes a script, so for those crawlers the raw HTML is the whole page.

### Test which version crawlers actually see

Three checks confirm what GPTBot, ClaudeBot and PerplexityBot actually get from a page:

1. Fetch it with the bot's user agent (`curl -A "GPTBot" ...`) and count the words in the response.
2. Compare that count against the rendered DOM in a browser or Screaming Frog.
3. Confirm the hits in your logs against the crawler's published IP ranges; a spoofed user agent proves nothing.

If any of those three already hit a template in your logs, its critical content has to be in the raw HTML. A fix that only satisfies Googlebot leaves those visits reading an empty shell.

Content can also disappear before it reaches the raw HTML in quieter ways: a value fetched inside a Web Worker, content behind an API call that returns a 429, text added through `::before`/`::after` CSS, or, for Googlebot, anything past the 2MB it fetches per resource. None of it reliably reaches what a crawler reads.

### Ship server-rendered content so everyone sees it

Screaming Frog and URL Inspection confirm whether your own render matches what ships:

- Compare the raw HTML against the rendered DOM using Screaming Frog's rendered-page crawl mode.
- Confirm against reality with URL Inspection's "View Crawled Page," which shows the HTML Googlebot genuinely rendered.
- Check that your H1, meta description, canonical, body text, and JSON-LD are already in the raw HTML, before any script runs.

When they're not, render on the server (SSR) or generate the page statically (SSG), so the content ships in the first response. The [SaaS SEO playbook](https://missiongrowth.io/blog/saas-seo) covers the sharpest case: comparison pages built in the same JavaScript framework as the product.

We hit this on our own site. Mission Growth's marketing site began as a React single-page app, so a crawler that skipped JavaScript received almost nothing. We migrated 20 marketing pages to prerendered static HTML for exactly that reason, and they now return their full text on the first request with no JavaScript executed.

::figure{src="/blog/figures/technical-seo-checklist-2026-3.svg" alt="A React single-page-app shell returns an empty div before migration, while the same page returns full prerendered HTML with content and schema after it." caption="After our migration to prerendered static HTML, 20 marketing pages serve their full text before any JavaScript runs." width="720" height="263"}

You don't have to rebuild your stack to match it, but skip this and most AI-linked crawl traffic never reads past the shell. That's the floor the rest of GEO builds on: [how that visibility turns into AI citations](https://missiongrowth.io/blog/generative-engine-optimization) is the next layer once crawlers can actually read your content.

## 11. Core Web Vitals and page experience

Core Web Vitals work as a tiebreaker rather than a ranking multiplier, and they're measured on real users rather than lab tests.

Use PageSpeed Insights for field data, the Experience report in Search Console to find "Poor" URL groups, and Chrome DevTools to diagnose a single page.

The one rule people forget: a Lighthouse lab improvement doesn't automatically move your field data. Field data comes from CrUX on a 28-day rolling window, so expect four to eight weeks before a fix shows in Search Console.

The thresholds haven't changed for 2026. Google's "good" bar stays at LCP 2.5s or faster, INP under 200ms, and CLS under 0.1, all measured at the 75th percentile of real Chrome users.

INP replaced FID as a Core Web Vital in March 2024. Google's web.dev targets a TTFB under 800ms, the server response time that feeds into LCP.

::dataset{key="core-web-vitals-thresholds-2026" name="Core Web Vitals thresholds and fixes, 2026"}

| Metric | Good at CrUX p75 | Most common cause | Fix |
|---|---|---|---|
| LCP (Largest Contentful Paint) | 2.5s or faster | Slow TTFB, late-discovered hero image, render-blocking CSS/JS | Preload the LCP image, inline critical CSS, cache at a CDN, and never lazy-load the hero |
| INP (Interaction to Next Paint) | under 200ms | Long JavaScript tasks blocking the main thread | Break long tasks with `scheduler.yield()`, defer non-critical JS, remove layout thrashing |
| CLS (Cumulative Layout Shift) | under 0.1 | Images with no dimensions, injected banners, font swap | Set `width`/`height` on media, reserve space for ad slots, use `font-display: optional` |
| TTFB (feeds LCP) | under 800ms | Slow server, no caching | Add server-side cache and a CDN, optimize slow database queries |

One frequent self-inflicted wound: applying `loading="lazy"` to the hero image. Google confirmed that lazy loading an above-the-fold image delays its paint and shows up directly in LCP. Set `fetchpriority="high"` and explicit dimensions on the hero, and save lazy loading for what sits below the fold.

## 12. Structured data and schema

Structured data doesn't move rankings directly, but valid markup makes your page eligible for rich results, and rich results change click-through at the same rank.

Validate every type with Google's Rich Results Test and monitor the Enhancements reports in Search Console for errors and warnings.

Three rules keep schema working:

1. **Serve JSON-LD in the raw HTML response.** Markup injected by JavaScript, or by Google Tag Manager, enters the render queue and indexes late.
2. **Never mark up data that isn't visible on the page.** A fabricated `aggregateRating` with no on-page reviews risks a manual action, and pulling in reviews from a third-party service is explicitly disallowed.
3. **Repeat the markup in your mobile HTML.** Mobile-first indexing ignores schema that only exists on desktop.

Schema earns its place in AI answers too. AirOps's "The 2026 State of AI Search" found pages with three or more schema types are about 13% more likely to be cited by AI engines.

Roughly 61% of cited pages already use three or more types, while FAQ schema appears on only 10.5% of cited pages, so it's under-used rather than saturated.

::figure{src="/blog/figures/technical-seo-checklist-2026-5.svg" alt="AI-cited pages use three or more schema types 61% of the time, but FAQ schema appears on just 10.5% of them." caption="FAQ schema is under-used on AI-cited pages, not saturated: only 10.5% carry it against 61% using three or more schema types." width="720" height="212"}

The upside compounds with the rest of the checklist. Google's own case study on Saramin reports that fixing crawl errors, consolidating canonicals, and adding JobPosting and Breadcrumb schema doubled organic traffic year over year, lifted signups 93%, and raised conversion 9%.

For how the same markup discipline feeds AI citations specifically, see our guide to [generative engine optimization](https://missiongrowth.io/blog/generative-engine-optimization).

::figure{src="/blog/figures/technical-seo-checklist-2026-6.svg" alt="The same fixes, crawl-error cleanup, canonical consolidation, and added schema, lifted signups 93% and conversion rate 9% year over year." caption="In Google's own case study on Saramin, the same fixes lifted signups 93% and conversion 9% year over year." width="720" height="212"}

## 13. AI crawler access: GPTBot, PerplexityBot, Google-Extended and llms.txt

AI crawler access turns the rendering problem from section 10 into a citation problem.

Test what an AI crawler sees with a spoofed user-agent: `curl -A "GPTBot" https://yourdomain.com/your-page` and check whether your key content is in the response. The mechanical fact underneath everything: the major AI crawlers don't execute JavaScript. AppleBot is the exception.

In Vercel's December 2024 network data, monthly request volumes were Googlebot at 4.5B, GPTBot at 569M, ClaudeBot at 370M, AppleBot at 314M, and PerplexityBot at 24.4M.

If your content renders client-side, most of the engines feeding AI answers receive an empty shell: the same failure as section 10, but with citations at stake instead of rankings.

::figure{src="/blog/figures/technical-seo-checklist-2026-4.svg" alt="Googlebot’s rendered DOM shows full article text, while GPTBot, ClaudeBot and PerplexityBot receive only the raw, unrendered HTML shell of the same page." caption="Googlebot sees the rendered page; most AI crawlers see only the raw HTML." width="720" height="263"}

::figure{src="/blog/figures/technical-seo-checklist-2026-7.svg" alt="Monthly AI crawler requests: GPTBot at 569M, ClaudeBot at 370M, AppleBot at 314M and PerplexityBot at 24.4M." caption="In Vercel's network data, GPTBot sent more monthly requests than any other AI crawler." width="720" height="292"}

The most common misconfiguration is conflating two different decisions:

> [!IMPORTANT]
> Training crawlers (GPTBot, Google-Extended, ClaudeBot in its training role) feed model weights. Retrieval crawlers (OAI-SearchBot, PerplexityBot) fetch live at query time to build a cited answer. Blocking a retrieval bot kills your citations. Blocking a training bot doesn't, and it may be a deliberate rights decision. Manage them with one blanket rule, and you'll either give away training data you meant to protect or lose the retrieval access you wanted.

::figure{src="/blog/figures/technical-seo-checklist-2026-8.svg" alt="Training AI crawlers such as GPTBot feed model weights, while retrieval crawlers such as PerplexityBot fetch live to build a cited answer." caption="Blocking a retrieval crawler loses citations immediately, while blocking a training crawler mainly withholds data used for model training." width="720" height="319"}

For the full allow-or-block call per crawler role, see [when to block GPTBot and other AI crawlers](https://missiongrowth.io/blog/gptbot-ai-crawlers).

Three more traps to check:

1. **A `nosnippet` directive.** It suppresses your page from AI Overviews as well as regular snippets, so confirm it's intentional.
2. **A `max-snippet:0` directive.** It has the same suppressing effect, so treat it the same way.
3. **A WAF or CDN rule.** Cloudflare's AI bot blocking can override robots.txt entirely, so an allowed bot may still be blocked at the edge.

llms.txt is hygiene with no proven growth effect. Frame it accordingly.

The major AI crawlers rarely fetch it: public llms.txt-adoption monitoring logged only 408 fetches across a window of 500 million AI bot requests.

Google confirmed in July 2025 it doesn't use llms.txt as a signal, and adoption among the top 10,000 sites sits around 5.6% in 2026. It's cheap and it won't hurt you.

Two resources handle the rest:

- Our own [llms.txt](https://missiongrowth.io/tools/llms-txt-checker) and a fuller llms-full.txt: a curated plain-text map for AI crawlers, plus a free [llms.txt checker](https://missiongrowth.io/tools/llms-txt-checker) to validate yours before you ship.
- The full playbook: [how to optimize for AI search engines](https://missiongrowth.io/blog/how-to-optimize-for-ai-search-engines), and for the AI Overviews channel specifically, [how to show up in Google AI Overviews](https://missiongrowth.io/blog/how-to-show-up-in-google-ai-overviews).

## Your 13-point technical SEO checklist recap

1. HTTP status, redirects (one hop), HTTPS with HSTS, no soft-404s.
2. Robots.txt returns 200, never blocks CSS/JS, and you use noindex (not Disallow) to deindex.
3. Split "Crawled, not indexed" (quality) from "Discovered, not crawled" (crawl budget).
4. Align canonical tag, sitemap, links, and redirects to one URL.
5. Sitemap under 50,000 URLs and 50MB, canonical 200s only, honest lastmod.
6. Clean URLs, History API routing, canonicalized parameters, disallowed junk facets.
7. Read the logs: watch crawl efficiency, kill bot traps and crawl waste.
8. No orphans, click depth three or less, no broken internal links, links in static HTML.
9. Mobile-first parity for content, links, and schema, with a viewport tag.
10. Ship every indexing directive, and AI-crawler-critical content, in the raw HTML; JavaScript-only fixes never reach a crawler that can't render.
11. Core Web Vitals at CrUX p75: LCP 2.5s, INP under 200ms, CLS under 0.1, TTFB under 800ms.
12. Valid server-rendered JSON-LD, no markup for invisible data, mobile parity.
13. AI crawler access decided per bot, retrieval versus training kept separate.

Run the foundation checks (1 through 6) every quarter and after any migration or major deploy. A single bad release can undo three months of cleanup in a week, which is why the log file, not the dashboard, is the last thing you should trust.

When you hand the results to a client or team, an [SEO report template](https://missiongrowth.io/blog/seo-reporting) turns each fix into what changed and what's next.

## Frequently asked questions

### What is technical SEO and why does it matter in 2026?

Technical SEO is the work of making sure search engines can crawl, render, index, and trust your pages. It matters in 2026 because the gates got stricter: Chrome starts warning on plain HTTP sites, Core Web Vitals stay a ranking input, and AI crawlers that don't run JavaScript now feed a growing share of discovery. Weak technical foundations cap everything content can earn. Once they hold, [how to rank higher on Google](https://missiongrowth.io/blog/how-to-rank-higher-on-google) turns on intent, information gain and clicks.

### How often should you run a technical SEO audit?

Run the foundation checks (HTTP status, robots.txt, index coverage, canonicals, sitemaps, URLs) once a quarter, and again immediately after any site migration, redesign, or large deploy, since those are when technical debt appears fastest.

Monitor Search Console's Pages and Experience reports continuously so a bad release surfaces within days. To see them in one place, [build an SEO dashboard to track these checks over time](https://missiongrowth.io/blog/seo-dashboard) with a technical-health tab. Deep log-file analysis fits a twice-yearly cadence for most sites.

### Do AI crawlers like GPTBot and PerplexityBot execute JavaScript?

No. The major AI crawlers fetch raw HTML and don't run JavaScript, with AppleBot the exception. That's why a client-rendered page that looks complete in a browser can arrive nearly empty to GPTBot, ClaudeBot, or PerplexityBot, which makes it uncitable. The fix is server-side rendering or static generation, so your key content ships in the first HTML response.

### What is the difference between robots.txt Disallow and noindex?

`Disallow` in robots.txt blocks crawling; `noindex` blocks indexing. They are not interchangeable. A Disallowed page can still appear in results as a bare link if others link to it, because Google never crawls it to see any noindex tag. To remove a page from the index, remove the Disallow so Google can crawl it, then add a `noindex` directive.

### Is llms.txt necessary for AI SEO in 2026?

No. Treat llms.txt as low-impact hygiene. Major AI crawlers rarely fetch it, and Google confirmed in July 2025 that it does not use llms.txt as a signal, with adoption near 5.6% of the top 10,000 sites.

It's cheap and harmless to add, but crawler access and parseable HTML do the real work. You can check any file with [our free validator](https://missiongrowth.io/tools/llms-txt-checker).

### What tools do you need for a technical SEO audit?

A crawler (Screaming Frog) for status codes, canonicals, directives, and render diffs. Google Search Console for index coverage, crawl stats, and URL inspection. PageSpeed Insights for Core Web Vitals field data, the Rich Results Test for schema, and `curl` for quick header and user-agent checks. For crawl budget, add server log analysis with pandas or a dedicated log-file tool. That stack covers all thirteen checks.
