The five technical SEO checks I run before anything else
Great content on a site that can't be crawled is invisible. These are the five foundation checks I run before any content or link-building work, now covering AI crawlers as well as Google.

Quick answer
Before you invest in content or links, run five technical SEO checks: can search and AI crawlers reach your pages (robots.txt and your CDN or firewall), can the right pages be indexed (noindex, canonicals, Search Console), does key content exist without JavaScript, do pages pass Core Web Vitals including INP, and is your structured data valid. Fix these first, or everything built on top leaks.
Great content on a site that can't be crawled is invisible. I have watched businesses spend months on blog calendars and link outreach while a single line in robots.txt, or a firewall rule nobody remembered switching on, quietly stopped search engines and AI assistants from ever reading the work. After seven years of running growth, SEO and performance programmes, first inside B2B firms and now through my agency, Paradigm Media Networks, my rule hasn't changed: before any strategy work, I check the foundations.
Most people treat technical SEO as a hygiene task you outsource once and forget. I see it as the layer that decides whether your marketing spend compounds or evaporates, and in 2026 it has to work for two audiences: traditional search engines, and AI systems such as ChatGPT search and Perplexity. This technical SEO checklist is the five checks I run, in the order I run them.
Key takeaways
- Crawlability now includes AI crawlers. Blocking the wrong user agent, or letting your CDN block it for you, can remove you from ChatGPT or Perplexity answers.
- Indexability is about the right pages being indexed, not all pages. Google itself says you should not expect every URL to be indexed.
- If your key content only appears after JavaScript runs, assume some bots will never see it. Check the raw HTML.
- Core Web Vitals are a user and conversion problem first. INP replaced FID in 2024 and is the metric most sites now fail.
- Structured data makes you eligible for richer results and easier to understand. It never guarantees them.
Check 1: Can search and AI crawlers reach your pages?
Crawlability is the first gate. If a bot cannot fetch a page, nothing else in this checklist matters for that page. The classic version of this check is reviewing robots.txt and confirming that important pages return a 200 status rather than redirect chains, soft 404s or server errors. What has changed is the cast of crawlers you need to think about.
OpenAI, Perplexity and Google each now publish separate user agents for separate jobs, and they behave differently. Treating them as one "AI bot" category is how sites end up blocking the crawler that drives visibility while letting through the one they actually wanted to stop.
| Token / user agent | Owner | What it does | Follows robots.txt? | Block it if… |
|---|---|---|---|---|
| Googlebot | Crawls for Google Search | Yes | Almost never | |
| Google-Extended | A robots.txt control token for whether crawled content may train future Gemini models; no separate HTTP user agent | Yes (control token only) | You object to Gemini training. Google states it does not affect Search inclusion or ranking | |
| OAI-SearchBot | OpenAI | Surfaces sites in ChatGPT search features | Yes | You do not want to appear in ChatGPT search answers |
| GPTBot | OpenAI | Collects content for training OpenAI foundation models | Yes | You object to model training |
| ChatGPT-User | OpenAI | Fetches pages when a user asks ChatGPT to | OpenAI says robots.txt rules may not apply | Not reliably controllable via robots.txt |
| PerplexityBot | Perplexity | Surfaces and links sites in Perplexity search results; Perplexity says it is not used to train models | Yes | You do not want to appear in Perplexity |
| Perplexity-User | Perplexity | Visits pages in response to individual user questions | Generally ignores robots.txt, per Perplexity | Not reliably controllable via robots.txt |
The practical consequence: disallowing GPTBot is a training decision, not a visibility decision. If you want to be cited in ChatGPT search, it is OAI-SearchBot that has to be allowed, and OpenAI notes it can take around 24 hours for its systems to adjust after you change robots.txt. Many site owners copied a "block all AI" snippet in 2023 or 2024 and never revisited it. If you care about visibility in AI search, that snippet deserves a second look.
The block you didn't write
The operational reality is that robots.txt is no longer the only place crawling gets decided. On 1 July 2025, Cloudflare announced it was changing its default to block AI crawlers unless they pay for content. Many sites sit behind Cloudflare or a similar CDN, web application firewall or security plugin, and bot-management rules there can return a 403 or a challenge page to a crawler your robots.txt explicitly welcomes. Perplexity's own documentation now includes guidance on allowlisting its user agents and published IP ranges in firewalls such as Cloudflare and AWS WAF, which tells you how common this problem has become.
Pro tip
Open your CDN or firewall dashboard and search the bot or security settings for "AI", "crawler" and "bot fight". Then decide crawler by crawler, using the table above, which ones you want. Write the decision down, so the next developer doesn't silently reverse it.
Internal links belong in this check too. Every important page should be reachable within a few clicks from the home page through descriptive, crawlable <a href> links. Orphan pages that exist only in a sitemap are discovered slowly and treated as unimportant. In audits I run, service pages linked only from a JavaScript mega-menu are a recurring find; plain text links from the home page and relevant posts are often the cheapest fix on the list.
Check 2: Are the right pages indexable, and actually indexed?
Crawling and indexing are different events. A page can be crawled and still excluded from the index, and a page blocked from crawling can still appear in Google. Google's documentation is explicit that robots.txt manages crawler traffic and is not a mechanism for keeping a page out of Google; a disallowed URL can still be indexed if other sites link to it. To keep a page out, you use a noindex directive or password protection.
Common mistake
Adding noindex to a page and also disallowing it in robots.txt. If Google cannot crawl the page, it never sees the noindex, so the URL can linger in results. Let it be crawled until it drops out, then decide whether to block.
The second lever is canonicalisation. Google treats rel="canonical" as a strong signal, not a command, alongside redirects (strong) and sitemaps (weak). When your canonical points to URL A, internal links to URL B and the sitemap to URL C, Google picks for you.
The place to see the outcome is the Page indexing report in Search Console. Compare what is indexed with the pages you actually want ranking. Gaps in either direction are clues. The statuses I look at first:
- Crawled, currently not indexed: Google fetched the page and chose not to index it yet. On important pages this is usually a quality or duplication signal, not a technical bug.
- Discovered, currently not indexed: Google knows the URL but hasn't crawled it. Weak internal linking is a frequent cause on smaller sites.
- Duplicate, Google chose different canonical than user: your canonical was overruled. Find out why before you fight it.
- Excluded by noindex: check every URL here was meant to be excluded. Staging-site
noindextags that survived launch live here.
What most people miss
A shrinking index is not automatically bad news. Google states you should not expect all URLs on your site to be indexed, only the canonical ones. If your index count drops because filter pages, tag archives and thin duplicates have been consolidated, that is often the healthiest change you can make. Chasing a bigger number is the wrong goal.
Check 3: Does your key content exist without JavaScript?
Google processes JavaScript pages in three phases: crawling, rendering and indexing. Pages are queued for rendering in a headless version of Chromium, and Google says a page may stay in that queue for a few seconds but it can take longer.
Google's own JavaScript SEO guidance still recommends server-side rendering or pre-rendering, partly for speed and partly because, in its words, not all bots can run JavaScript. That is the line I keep coming back to. If your product descriptions, prices, reviews or article body only appear after client-side scripts run, assume some crawlers, including many AI crawlers, will see an empty shell.
Technical SEO is not the strategy. It's the permission for the strategy to work.
In practice, the test takes two minutes. Open the page, view source (not the inspector, which shows the rendered DOM), and search for a sentence from your main content. If it isn't in the raw HTML, you have a rendering dependency. Then use Search Console's URL Inspection tool to see what Google rendered.
Two subtle failures are worth knowing. Google notes that if it finds a noindex tag in the original HTML, it may skip rendering, so removing that tag with JavaScript may not work. And "links" that are really buttons with click handlers may not be followed; navigation and pagination should be real anchor elements.
Don't rebuild a whole site because someone said "JavaScript is bad for SEO". Make sure revenue-carrying templates deliver their core content in the initial HTML; interactive extras can stay client-side.
Check 4: Is it fast enough, including INP?
Core Web Vitals measure three things: loading (Largest Contentful Paint), responsiveness (Interaction to Next Paint) and visual stability (Cumulative Layout Shift). The thresholds web.dev publishes for a "good" experience are an LCP within 2.5 seconds, an INP of 200 milliseconds or less, and a CLS of 0.1 or less, measured at the 75th percentile of page loads, split by mobile and desktop.
INP became a stable Core Web Vital in 2024, replacing First Input Delay. That shift matters because FID only measured the delay before the first interaction was handled, while INP looks at responsiveness across interactions during the visit. Sites heavy with tag managers, chat widgets and page-builder scripts that sailed through FID often struggle with INP, because the main thread is busy when the visitor taps a button.
Here is my contrarian view: Core Web Vitals won't win rankings on their own, and I have seen teams burn a quarter chasing a perfect Lighthouse score while their positioning stayed mediocre. Fix them because slow pages lose users and conversions, especially on mid-range phones. Treat any ranking benefit as a bonus.
How to approach it: start with field data (the Core Web Vitals report in Search Console or PageSpeed Insights' real-user section), not lab scores, and work template by template rather than URL by URL. A tactical example: if your blog template fails INP, audit third-party scripts first. Removing one unused heatmap or chat script often does more than a week of image optimisation.
Check 5: Is your structured data valid and honest?
Structured data is the explicit layer: it tells search engines what a page is about in a machine-readable vocabulary, usually schema.org. Google recommends JSON-LD where your setup allows it, because it sits in a script tag and is easier to maintain at scale than markup woven through your HTML.
Two points get lost. First, structured data makes a page eligible for rich results; it does not guarantee them. You need the required properties for each feature, and Google still decides what to show. Second, the markup must describe what is visible on the page. Review markup on pages without reviews, or FAQ markup for questions that don't appear, is the kind of shortcut that creates problems rather than solving them.
For most businesses, a sensible baseline is Organization (or LocalBusiness) on the home page, Article with a clear author on blog posts, BreadcrumbList across the site, and Product markup on product pages with accurate price and availability. Validate with Google's Rich Results Test, and watch Search Console for errors after theme or plugin updates.
For AI search, I would not promise that schema alone earns citations; nothing I have read from OpenAI or Perplexity says so. It is part of being unambiguous, which is the point of this whole checklist.
How to run the five checks in one afternoon
You don't need an enterprise crawler for a useful first pass. This is my sequence on a new site; as with any marketing system, a repeatable process beats a heroic one-off audit.
- Read robots.txt and your CDN settings togetherOpen
/robots.txt, list every user agent rule, then check your CDN, firewall and security plugin for bot blocking. Decide on Googlebot, OAI-SearchBot, PerplexityBot, GPTBot and Google-Extended individually. - Spot-check status codes and internal linksTake your 20 most important URLs and confirm each returns 200, isn't redirected, and is linked from the home page or a main hub with a plain text link.
- Review the Page indexing reportExport indexed and non-indexed URLs from Search Console. Flag important pages that are missing and junk pages that are present. Check canonicals on both lists.
- View source on each key templateFor your home, service, product and article templates, confirm the main content and links exist in the raw HTML, and use URL Inspection to compare with what Google rendered.
- Pull field Core Web Vitals by templateUse Search Console's Core Web Vitals report to find failing URL groups, then investigate the worst template first, starting with third-party scripts for INP.
- Validate structured dataRun one URL per template through the Rich Results Test, fix errors, and confirm markup matches what users can see.
Give every finding an owner and a date. A list of problems with no owner is a document, not an audit.
Fix the floor, then build
Content, links and AI visibility work all multiply whatever foundation they sit on. When the five checks pass, every new article and earned link has a fair chance of being found, indexed and cited. When they don't, you pay full price for a fraction of the return.
Start with Check 1 this week, because the AI crawler and CDN issues are the newest and the least likely to have been reviewed. If you'd like a structured way to work through the rest, the free audit checklist covers all five areas and more, and my other SEO and AI search resources pick up where it leaves off.
Frequently asked questions
What should a technical SEO checklist include?
At minimum, five areas. Crawlability: robots.txt rules, CDN or firewall bot settings, status codes and internal links. Indexability: noindex tags, canonicals and Search Console's Page indexing report. Rendering: whether key content and links exist in the raw HTML without JavaScript. Performance: Core Web Vitals, including INP, measured from field data. Structured data: valid JSON-LD that matches visible content. Run them in that order, because a failure early in the chain makes later fixes irrelevant for the affected pages.
Should I block GPTBot in robots.txt?
That depends on whether you object to your content being used for model training, because that is what GPTBot is for, according to OpenAI. Blocking it is a training decision. Visibility in ChatGPT search is governed by a different crawler, OAI-SearchBot, which OpenAI says must be allowed for your pages to appear in ChatGPT search answers. Many sites block both by accident with a blanket rule. Decide on each one separately, and remember that OpenAI says robots.txt rules may not apply to ChatGPT-User, which fetches pages on a user's request.
Does blocking Google-Extended hurt my Google rankings?
According to Google, no. Google-Extended is a robots.txt control token that lets publishers manage whether content Google crawls may be used to train future generations of Gemini models. Google states that it does not affect a site's inclusion in Google Search and is not used as a ranking signal. It also has no separate HTTP user agent; crawling still happens with Google's existing user agents. So disallowing it is a policy choice about AI training, not an SEO risk to your search listings.
Why is my page "Crawled, currently not indexed" in Search Console?
It means Google fetched the page but chose not to add it to the index, at least for now. Google says it may index it later. On important pages, this is usually a signal about quality, uniqueness or duplication rather than a technical fault. Check whether the page adds something distinct, whether a similar page on your site competes with it, and whether it has meaningful internal links. Resubmitting the URL without improving it rarely changes the outcome for long.
Does JavaScript hurt SEO?
Not automatically. Google renders JavaScript using a headless version of Chromium, though pages wait in a render queue first. The risk is with other crawlers and with content that only appears after scripts run. Google itself recommends server-side rendering or pre-rendering, noting that not all bots can run JavaScript. The practical approach is to make sure your core content and navigation links are present in the initial HTML of revenue-carrying templates, and check it with view source and the URL Inspection tool.
What is a good INP score?
web.dev defines a good Interaction to Next Paint as 200 milliseconds or less, measured at the 75th percentile of page loads and assessed separately for mobile and desktop. INP became a stable Core Web Vital in 2024, replacing First Input Delay. Poor INP is commonly caused by heavy JavaScript keeping the browser's main thread busy, so third-party scripts such as chat widgets, heatmaps and tag managers are the first place to look when a template fails.
Sources
- Overview of OpenAI Crawlers — OpenAI, 2026
- Perplexity Crawlers — Perplexity, 2026
- Google's common crawlers (incl. Google-Extended) — Google Search Central, 2026
- Content Independence Day: no AI crawl without compensation — Cloudflare, 2025
- Page indexing report — Google Search Console Help, 2026
- Understand JavaScript SEO basics — Google Search Central, 2026
- Web Vitals — web.dev (Google), 2026

