A web crawler, also called a spider or bot, is a program that systematically browses the web, following links from page to page and downloading what it finds so a search engine can index it. Googlebot is Google’s crawler and Bingbot is Bing’s. “Spider” and “crawler” mean the same thing.
On this page
Spiders and crawlers: is there a difference?
No. “Spider”, “crawler”, “bot” and “robot” are used interchangeably. “Spider” is the older nickname, from a program that moves across the web. Most documentation today says “crawler”.
How a crawler works
- Discovery. The crawler starts from a list of known URLs, plus URLs from XML sitemaps and links found on pages it has already visited.
- Permission. It checks the site’s
robots.txtfile to see which paths it may request. - Fetching. It requests the page and records the HTTP status code and the content.
- Rendering. Google runs the page’s JavaScript in a headless browser so it can see content that is added after the initial load.
- Extraction. New links go into the queue, and the content is passed on for indexing.
Crawling and indexing are separate steps. A page can be crawled and still not be indexed if Google judges it duplicate, thin or blocked by a noindex rule.
Common crawlers
| Crawler | Operator | Purpose |
|---|---|---|
| Googlebot | Google Search (smartphone and desktop versions) | |
| Bingbot | Microsoft | Bing and services that use Bing’s index |
| Applebot | Apple | Siri and Spotlight suggestions |
| GPTBot, OAI-SearchBot | OpenAI | AI model training and ChatGPT search |
| ClaudeBot | Anthropic | AI model training |
| PerplexityBot | Perplexity | Perplexity’s answer engine index |
| AhrefsBot, SemrushBot | SEO tools | Backlink and site audit data |
How to control crawlers
robots.txttells crawlers which paths they may not request. It controls crawling, not indexing.- Meta robots tags and the X-Robots-Tag header control indexing and how a page is shown.
- XML sitemaps list the URLs you want crawled.
- Canonical tags point crawlers to the preferred version of duplicate pages.
- Internal links decide how easily crawlers reach each page. See website structure.
Do not block a page in robots.txt if you want it removed from results. A crawler that cannot fetch the page cannot see its noindex tag.
Crawl budget
Crawl budget is the number of URLs a search engine is willing and able to crawl on a site in a given period. It matters mainly for large sites with tens of thousands of URLs or more. Wasted budget usually comes from faceted navigation, URL parameters, redirect chains, duplicate pages and slow servers.
How to see what crawlers do on your site
- The Crawl Stats report in Google Search Console shows requests per day, response codes and file types.
- The URL Inspection tool shows when Googlebot last crawled a page and what it saw.
- Server log files show every crawler request, including AI crawlers.
- A desktop crawler such as Screaming Frog simulates a spider so you can find broken links and blocked pages before search engines do.
Frequently asked questions
What are spiders and crawlers?
They are programs that search engines use to discover and read web pages by following links. The two words mean the same thing.
What is the difference between crawling and indexing?
Crawling is fetching a page. Indexing is analysing it and storing it so it can appear in search results. A crawled page is not always indexed.
How do I stop a crawler accessing my site?
Add a disallow rule for that crawler’s user agent in robots.txt. Well-behaved crawlers follow it; for others you need server-level blocking.
Should I block AI crawlers?
It depends on your goals. Blocking training crawlers keeps content out of model training, but blocking AI search crawlers can remove you from AI answers and citations.
Want experts to handle this for you?
Our specialists turn these fundamentals into rankings, leads and revenue. Start with a free, no-obligation audit of your site.