field note
first published 2026-10-09
Technical SEO audit checklist: the one I actually run
The technical SEO audit checklist I run on every engagement — access, crawl, render, index, speed, structured data, AI-search readiness, logs — with what each item costs you when it's wrong, and the order to do it in.
1,668 words · 8 min read · technical seo
Every audit I've been handed by a new client has the same shape: a long export from a crawling tool, severity colours assigned by the tool, and a dev team that stopped reading at page four. Most of the items were real. Almost none of them mattered. The two things that did matter were not in the export at all, because the tool could not see them.
This is the checklist I run instead. It is ordered by what a finding costs you, not by how many pages it touches, and every item names where the evidence comes from. If you want the short version, it is at the bottom.
Before any item: freeze the inventory
An audit changes things. If you do not know what the site looked like before, you cannot say what the audit did. So the first hour is spent making a record: every URL the sitemap claims, every URL the crawler finds, every URL that received a click or an impression in Search Console in the last sixteen months, and the current ranking positions for the queries that pay the bills.
That inventory is the baseline every later readout is measured against. On a recent migration file it was also the thing that proved the client's collapse started seven months before the migration they were blaming. Nobody had looked at the exports side by side.
1. Access: can the machines get in?
Everything downstream depends on this, which is why it comes first and why it is so often skipped.
robots.txt: what is disallowed, and whether any of it was meant to be. I have seen an entire product catalogue blocked by a rule written for a staging server three years earlier.- Status codes from the origin, not from behind a CDN. A CDN can serve a cached 200 for a page the origin now 500s.
- Bot-management and firewall rules. The quickest way to lose rankings in 2026 is a security product that challenges Googlebot, and the second quickest is one that challenges the AI search crawlers. A UK garden-machinery retailer I worked with had a CDN firewall take 1,647 of 1,752 pages offline for crawlers in thirty seconds, and a different storefront app I built had to be paced to crawl its own customers without tripping the same kind of rule.
- The AI crawlers, decided on purpose: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended. Blocking training bots is a legitimate choice. Blocking the search bots by accident, while trying to block the training ones, is a common mistake with a measurable cost.
Evidence: the robots file, a fetch sweep against the origin with each user agent, the firewall's own event log.
2. Crawl and discovery
- The sitemap tells the truth: only canonical, indexable, 200 URLs, with real
lastmodvalues rather than a timestamp that updates on every build. - Internal links reach every page that matters within a few clicks, and the anchor text says what the page is.
- Faceted navigation, filters and parameters have a written policy. Which combinations are pages, which are parameters, which are
noindex, which canonical they point to. If nobody can state the policy, the crawler is writing it for you. - Redirect chains and loops, measured from the first hop to the last. One hop is the target.
- Orphans: pages in the sitemap or the logs that no internal link reaches.
Evidence: a full crawl, the sitemap diffed against the crawl, Search Console's crawl stats.
3. Rendering: what the crawler actually reads
Fetch the raw HTML with no JavaScript and compare it with the rendered page. Every important element should be in the raw response: the title, the headings, the body copy, the product data, the internal links. If the content only exists after a script runs, Google will probably still see it eventually, and almost every AI crawler will see nothing at all.
This item has grown from a footnote to a gate. The answer engines mostly do not run your JavaScript. A page whose answer lives in a client-rendered component is, to them, an empty page with a nice title.
Evidence: raw fetch against rendered DOM, the crawler's JavaScript-rendering mode, and the Agent-Ready Grader on this site, which tests extractability as one of its three gates.
4. Indexation and canonicals
- Search Console's Pages report, read page by page for the "Crawled, currently not indexed" and "Duplicate, Google chose different canonical" buckets. Those two lists are a map of where Google disagrees with you.
- Self-referencing canonicals on every unique page; cross-domain and protocol consistency; no canonical pointing at a redirect or a
noindexpage. - Hreflang, if the site has locales: reciprocal, self-referencing, valid codes, agreeing with the canonicals.
- Thin and near-duplicate pages, counted honestly. A UK tile retailer had 196 shadow pages in the index and the sitemap that nobody had meant to publish.
- Soft 404s: pages that return 200 with no real content, which Google quietly treats as missing.
Evidence: Search Console exports, the crawl's canonical report, URL Inspection on a sample.
5. Speed, found at the root
Core Web Vitals field data tells you whether users are suffering. It does not tell you why. The why is almost always one of a short list: slow server response, render-blocking resources, uncompressed or oversized images, a layout that shifts as fonts and ads land, and JavaScript that runs before the first paint.
Measure the origin directly, with the CDN bypassed, because that is the number you can actually change. That tile retailer's server response was 2.2 seconds at the first reading and 47 milliseconds after the causes were fixed. The causes were not in any front-end tool: a caching layer doing nothing, and an error log at 113 MB growing daily.
Evidence: origin timing, PageSpeed Insights field and lab data, the browser's performance panel, the server's own logs.
6. Structured data as a graph, not islands
Most sites have structured data because a plugin emitted some. The questions that matter are different: does it describe the real entities, do the entities link to each other with @id, does it validate with zero errors, and does it agree with the visible page? A Product block with a price the page does not show is worse than no block at all.
Three small businesses on three different platforms went to zero validator errors on exactly this work, and the pattern was the same each time: delete the islands, write one connected graph.
Evidence: the Rich Results Test, the schema.org validator, and a read of the page against its own markup.
7. AI-search readiness
This is the newest section of the checklist and the one most audits still leave out. The items are small and the consequences are not.
- Each AI crawler's access, decided and documented.
- Extractability: the answer is present in the raw HTML, near the top, in plain prose.
- Snippet eligibility: no
max-snippetornosnippetdirectives set by accident. - An
llms.txtfile that tells a model what the site is and where the machine-readable surfaces are. - Markdown mirrors or an equivalent clean surface for the pages that matter, and a
Link rel="alternate"header that points at them.
This site is the reference build for the list, and it grades itself: 100/A on the instrument it runs for clients. The scoring model is published on the lab page, with the three gates that cap a grade at C if any one of them fails.
Evidence: the grader, the robots file, a raw fetch, the headers.
8. The logs, which are the only source of truth
Everything above can be estimated from tools. Only the logs tell you what the crawlers actually did. What Googlebot fetched, how often, what it got back, which dead paths it is still visiting a year after a migration. On one migration file there were 192 of those dead paths, each one a small tax on crawl budget that the redirect map had never been told about.
If the hosting makes logs hard to get, say so in the audit, because the client will need them again. The site you are reading keeps its own: a guestbook on the lab page counts every known crawler by product, and at the time of writing Googlebot reads this site around 165 times a day.
Evidence: the raw access logs, parsed by user agent and status, over at least thirty days.
What a finding needs before it goes in the report
Each item that makes the report carries four things: the evidence (where I read it), the consequence (what it costs in crawl, index, rank or revenue), the fix (what changes, where, and whether it is reversible), and the readout (how we will know it worked). A finding missing any of those is a note, not a finding.
Severity follows the consequence. A broken canonical on the page that earns a third of the revenue outranks four thousand missing alt attributes on images nobody searches for. Tools cannot make that judgement, which is the whole reason the audit is done by a person.
The checklist, in one block
00 Freeze the inventory: sitemap, crawl, Search Console 16 months, rankings
01 Access: robots.txt · origin status codes · firewall/bot rules · AI crawlers
02 Crawl: sitemap truth · internal links · facet policy · redirect chains · orphans
03 Render: raw HTML vs rendered DOM for every element that matters
04 Index: GSC Pages buckets · canonicals · hreflang · thin/duplicate · soft 404s
05 Speed: origin timing · CWV field + lab · images · fonts · blocking JS
06 Structured data: one connected graph · zero validator errors · agrees with page
07 AI search: crawler access · extractability · snippet directives · llms.txt · mirrors
08 Logs: what the crawlers actually did, 30 days minimum
++ Every finding: evidence · consequence · fix · readout
Run it in that order. The early items decide whether the later ones matter, and the logs at the end tell you whether any of it was true.
asked straight — answered straight
01How long does a technical SEO audit take?
It scales with the site and with what the logs show. A small site can be read quickly; a catalogue with hundreds of thousands of URLs and a year of logs is a different job. The useful answer comes from the brief, and the audit opens against a frozen inventory so you see progress from the first day.
02What tools do you need for a technical SEO audit?
A crawler that can render JavaScript, access to Search Console and the server logs, a way to time the origin directly, a structured-data validator, and a browser with the network panel open. Everything else is convenience. The logs are the one source no tool can replace.
03What's the difference between a technical SEO audit and a site audit tool report?
A tool counts symptoms and ranks them by how often they occur. An audit explains causes and ranks them by what they cost you. Tools are a fine input to an audit and a poor substitute for one.
04Should the audit include AI search?
Yes, and most still don't. AI crawler access, content extractability without JavaScript, and a connected structured-data graph decide whether ChatGPT, Perplexity and Google's AI features can read and cite the page. The same audit that checks Googlebot can check them in the same pass.
next step — technical seo audit
Technical SEO audit finished as fixes.
An audit that ends in shipped fixes, with the readouts to prove it. Send a brief — a few lines about the site, the problem and the evidence you have — and the reply is a straight read on fit and shape, from the person who does the work.
no forms · no funnels · jamie@jamiemckaye.com

written by
Technical SEO consultant and full-stack developer in Hersham, Surrey, in practice since 2007. One person, no handoffs: the audits, the code and the writing come from the same pair of hands. This site is the working proof — it grades itself on the same instrument it runs for clients.