Google can only rank the pages it can find. Most websites accidentally hide a few from it. Broken links and pages nobody links to quietly drain the attention Google gives your site. A site crawl finds them before Google does, using the same method a search engine uses to explore your pages. It only takes an afternoon, but it usually dredges up something your developers swore was fixed.

What is a site crawl?

A site crawl is an automated sweep of your website by a tool called a crawler. It starts at your homepage, follows every internal link it can find and records what it sees on each page. The result is a map of your whole site, with the status code and key SEO details for every URL.

Google’s own crawler, Googlebot, works in much the same way to discover pages. Running a crawl yourself shows you your site the way a search engine sees it, which is rarely the way it looks from the homepage.

Crawling and indexing are different jobs

Crawling is how search engines find pages, while indexing is when Google stores a page so it can appear in search results. A page has to be crawled before it can be indexed, but being crawled doesn’t guarantee it will be. Pages being crawled but never indexed is just one example of the common technical SEO issues we come across in our work.

How often does Google crawl a site?

There’s no fixed schedule, as Google crawls each site at its own pace based on how often the content changes and how quickly the server responds. You can’t ask Google to crawl more often. On larger sites, that pace is tied to your crawl budget.

You can see Google’s activity in the Crawl Stats report, found under Settings in Google Search Console. Google says sites with fewer than about a thousand pages generally don’t need to watch it closely. On bigger sites, it’s a useful early warning, as a sudden drop often points to server errors or a new robots.txt rule blocking pages.

Choosing a website crawler

Most SEO crawling is done with either a desktop crawler or a cloud-based audit tool. Both do the same core job, so the right choice depends on your site size and budget.

Screaming Frog

Screaming Frog’s SEO Spider is the tool most SEOs reach for first. The free version crawls up to 500 URLs, while a licence to remove that limit costs £199 a year. It runs on your own computer, so that large sites can test the patience of both your laptop and its fan.

Screaming Frog alternatives

If you’d rather not install anything, Ahrefs Webmaster Tools includes its site audit free for websites you can verify you own. It gives you 5,000 crawl credits per verified project each month and checks for more than 170 common issues. Paid tools such as Sitebulb and Lumar suit larger or more complex sites.

How to crawl a website

We’re using Screaming Frog for our layout. Most crawlers have the same settings under similar names:

Check the crawler can get in.

Make sure your robots.txt file isn’t blocking crawlers from the pages you want to check. If the site sits behind a firewall or security plugin, ask your developer to allow the crawler through. Otherwise, you’ll get a remarkably quick crawl of exactly one page.

Crawl as Googlebot Smartphone

Google uses the mobile version of your pages for indexing and ranking, a setup known as mobile-first indexing. Set your crawler’s user agent to Googlebot Smartphone, so you’re auditing the version Google cares about. This also catches content that only appears on desktop, which Google may never see.

Turn on JavaScript rendering if you need it.

If your site builds its content with JavaScript, switch on JavaScript rendering so the crawler sees what a browser sees. Without it, the crawler may only find an empty shell of a page and report it as thin. Rendering makes the crawl slower, so leave it off for simple sites that don’t need it. Our glossary entry on JavaScript SEO explains when it matters.

Connect analytics & your sitemap.

Connecting Google Search Console and GA4 adds clicks and traffic to every URL in the crawl. Adding your XML sitemap fills in more of the gaps. Together, they reveal orphan pages that get visits but have no internal links pointing to them. A crawler can’t find those by following links alone, because there aren’t any to follow.

Run the crawl

Enter your homepage URL to begin. Depending on the size of your site, this could be minutes or hours. Keep the crawl speed modest on smaller servers, as flooding your own site with requests can slow it down for customers.

What to look for in your crawl results

Broken pages and server errors

Filter by status code and look for 4xx errors (pages that don’t exist) and 5xx errors (server failures). Fix internal links that point to broken pages or redirect the old URL to the closest relevant page. Expired pages and 404s can quickly add up on sites with changing stock.

Redirect chains and loops

A redirect chain sends visitors from page A to page B to page C before they arrive anywhere useful. Each hop wastes crawl time and slows the page down. Update internal links to point straight at the final URL. Then break any loops, where two pages redirect to each other as two people insisting the other goes through the door first.

Pages blocked from indexing

Check which pages carry a noindex tag or are blocked by robots.txt. Then check canonical tags, which tell Google which version of a page to index. Some blocks will be deliberate, such as basket pages or internal search results. Others will be important pages someone blocked during a redesign and forgot about, which happens more often than anyone admits.

Duplicate and thin content

Look for pages that share the same title tag or H1, as these often signal duplicate pages competing against each other. Filtered category URLs and old versions of a service page are common examples. Low word counts can also flag thin pages that give Google little reason to rank them.

Pages buried too deep

Crawl depth shows how many clicks a page sits from the homepage. Important pages buried several clicks deep get less attention from both visitors and Google. Bring them closer with links from menus or related content, which also passes on PageRank.

Turn the crawl into a fix list.

The crawl itself doesn’t fix anything. Sort issues by the traffic and revenue of the pages they affect. A broken link on your best-selling product beats a missing alt tag on a 2017 blog post. Fix template-level problems first, as one change can clear hundreds of URLs at once. Crawl again to confirm the fixes worked and nothing new broke along the way.

Crawls matter most around big changes. During Corndel’s site migration, we audited the site before launch and monitored it afterwards, finding zero critical issues once the new site went live. If you’re planning a rebuild, our website migration guide explains where crawls fit in.

Crawl little and often

A site crawl shows you your website the way Google sees it, warts and all. Running one every month or two catches problems while they’re still small and cheap to fix. Start with the broken and blocked pages on your most valuable URLs. If a quick site crawl doesn’t cover everything, we can perform a full technical SEO audit.

If you’d rather have someone else read the results, contact us for an SEO site audit. We’ll crawl your site, rank the issues by commercial impact and tell you which ones are worth fixing first.

Jozef Raczka

Joe is a Search Executive at Circulate, with over 6 years of experience in SEO content writing and strategy. Having achieved Masters' in script writing & magazine journalism, he brings the best of both disciplines to his writing. Providing creativity, a facts-led approach to finding the story and a rigid commitment to good structure. His articles have been featured in UNILAD, being cited in multiple sources for his SEO expertise and has led to participation in nationally broadcast media events. He's written for multiple publications and media, contributing to short-form audio play content and podcast hosting & production. Outside of work, he enjoys watching a good film and making a good soup.