A search engine crawls the web to discover pages, indexes what it finds, then ranks and serves the most relevant, highest-quality results for each query. Google's own documentation, the Search Central guide to how Search works, last updated 18 December 2025, gives the process three labels: crawling, indexing and serving search results. Discovery happens before you search. Ranking happens when you search.
That sequence is the thread of this page. When we place a Rhino Rank Curated Link or Guest Post on a page, the link is live for a reader the moment that page is live. Google sees it only when it next crawls the host page and extracts the link. The rest of this page is the three stages, the engine's own words at each one, one measured test of how long a link can take to be followed, and a diagnostic to run in order when a page will not show.
What a search engine is
A search engine is a retrieval system for the web. It helps you move from a query to a set of pages, images, videos, products, local results or AI-generated summaries that may answer that query. The important point is that a search engine does not start from a blank internet each time you type. It has already discovered and processed a large set of URLs. When you search, the engine looks inside that stored index and chooses which results to show. That is why Google's three-stage model is useful for marketers and site owners. It separates discovery from storage, and storage from ranking.
A page can be live but not crawled. It can be crawled but not indexed. It can be indexed but not served for the query you care about. Those three states are the reason the order of the stages matters, and the diagnostic at the end of this page follows them.

Crawling: how search engines discover URLs
Crawling is the discovery stage. Googlebot and other web crawlers fetch URLs so the search engine can see what exists, what has changed and which links point elsewhere. The same Search Central guide describes how new URLs are found. On links: "Other pages are discovered when Google extracts a link from a known page to a new page: for example, a hub page, such as a category page, links to a new blog post." On sitemaps: "Still other pages are discovered when you submit a list of pages (a sitemap) for Google to crawl." Two routes, then: links from known pages, and a sitemap.
Bing describes the same mechanism in its own words. Fabrice Canel, Principal Program Manager for Webmaster Tools at Microsoft, put it on the Bing Webmaster Blog on 16 October 2018, his spelling kept: "We also crawl content specifically to discovery links to new URLs that have yet to be discovered. Sitemaps and RSS/Atom feeds are examples of URLs fetched primarily to discovery new links." Two engines, one mechanism: a new URL is found when a crawler reads a page that links to it.
A sitemap is the second route, and Google's sitemaps documentation, last updated 10 December 2025, is precise about what it can and cannot do. Think of a sitemap as a clean list of URLs you want crawlers to know about. The guide's own baseline is: "If your site's pages are properly linked, Google can usually discover most of your site." One case where a sitemap earns its keep, in its words: "Your site is new and has few external links to it." It is not a command. A URL in a sitemap can still be ignored, crawled later, canonicalised to another URL or left out of the index.
robots.txt sits beside discovery but controls a different thing. robots.txt is a crawl-control file, not a private vault. It tells compliant crawlers which paths they may access, and it can help manage crawl traffic on sections that do not need crawling. Google's introduction to robots.txt, last updated 10 December 2025, states the limit plainly: "A page that's disallowed in robots.txt can still be indexed if linked to from other sites." Use robots.txt to manage access for crawlers. Use indexing controls or authentication when appearance in search is the real issue.
Crawl budget is the part of crawling most often overstated. Google's crawl-budget guide, last updated 22 July 2026, is written for sites of "1 million+ unique pages" with content that changes about weekly, sites of "10,000+ unique pages" with content that changes daily, and sites with a large share of their URLs sitting in the Search Console state "Discovered - currently not indexed". The guide calls its own numbers "a rough estimate". For everything else, clear internal links and a clean sitemap do the work. The full scope is in the Rhino Rank guide to crawl budget.
In practice, when we place a link on an existing page, discovery depends heavily on that host page. I've seen a busy, often-updated page get recrawled quickly, while a stale page can sit untouched for a long stretch. That is why crawl frequency matters: the link exists for people as soon as it is live, but Google still has to revisit the page before it can discover and process it.
There is measured data on how long that can take for a link. Andy Zhao, writing on andyaeo.com on 11 April 2026, tracked 1,247 backlinks across eight sites, three of his own and five of his clients', checking each link every seven days. His links were a mix of types, including forum comments and directory submissions. He reported: "Only 64% of all backlinks were indexed within 90 days. After 6 months, that number only climbed to 73%." On timing, he wrote: "In my data, the median time for a backlink to be discovered was 16 days." One of the five reasons he gives for links that never get indexed is a host page Google is not crawling. That is why our checks after publication cover the host page as well as the link: we check that the page is live, that the link sits in the body copy, that the anchor and the URL match the order, and that the page is indexable. Only then does the URL go on the report. The checks are listed on our link building quality assurance page. How that wait shows up in results is a separate question, answered in the Rhino Rank guide to how long link building takes.
Indexing: how search engines store and interpret pages
Indexing is where a search engine analyses a crawled page and decides what to store. The index is not a simple copy of the page. It includes content, metadata, signals about the page, relationships to other URLs and Google's view of which version should represent a set of duplicates. The Search Central guide to how Search works opens this stage in one sentence: "Indexing: Google analyzes the text, images, and video files on the page, and stores the information in the Google index, which is a large database."
Rendering is part of this stage for modern pages. Google's JavaScript SEO basics, last updated 4 March 2026, says Google runs JavaScript with an "evergreen version of Chromium", and that the work queues: "The page may stay on this queue for a few seconds, but it can take longer than that." The evergreen renderer was announced in a Google Search Central blog post on 7 May 2019. The cleaner model is crawling, rendering and indexing, and there is a queue inside it.
How long that queue can run is measured, not guessed. Ziemek Bućko ran the test with Marcin Gorczyca and wrote it up on Onely's site, published 9 November 2022: seven pages in each of two folders, six of them reachable only through an internal link, one folder linked in plain HTML and the other with links injected by JavaScript. "It took Google 313 hours to get to the final, seventh page of the JavaScript folder." "With HTML, it took just 36 hours. That’s nearly 9 times faster." And for the first link, not only the last page: "Even with the first JavaScript-injected internal link, it took Googlebot twice as long to follow it as opposed to the HTML link (52 hours vs 25 hours)." A link that only exists after a script runs is a link Google has to render before it can follow, so a link belongs in the page's own HTML.
During indexing, Google tries to understand the main content of the page, the title, headings, media context, language, visible links and metadata. It also has to separate the main page from boilerplate such as navigation, sidebars and repeated footer content. A clean template helps, because the useful content is easier to identify.
Canonicalisation is the indexing decision most likely to surprise a site owner. When Google finds duplicate or very similar pages, it chooses a canonical URL to represent that content. Google's canonical documentation, last updated 10 July 2026, weighs the signals in its own words: redirects are "A strong signal", rel="canonical" link annotations are "A strong signal", and sitemap inclusion is "A weak signal". The final choice stays with the engine: "Google will identify which version of the URL is objectively the best version to show to users in Search." The goal is to consolidate signals to the right version rather than split them across duplicates. Canonicalisation is not only for exact duplicates. It can also affect filter URLs, tracking-parameter URLs, print versions, tag pages and product variants with very similar copy.
Mobile-first indexing closes the stage. Google's mobile-first documentation, last updated 10 December 2025, states it in two sentences: "Google uses the mobile version of a site's content, crawled with the smartphone agent, for indexing and ranking. This is called mobile-first indexing." The practical requirement is parity: the mobile page should carry the same main content, metadata, structured data and useful internal links as the desktop version. Indexing itself is not guaranteed. The how-Search-works guide says so directly: "Indexing isn't guaranteed; not every page that Google processes will be indexed." The crawl log tells you Google reached the URL. It does not prove the URL has earned a place in the index. That is why index coverage is a better diagnostic than raw crawl activity. Search Console can show whether Google selected a canonical, indexed the page, excluded it or discovered it without indexing.
Ranking and serving search results
Serving is the query-time stage. When you search, Google looks through its index and uses automated ranking systems to order results. It is more accurate to think in terms of ranking systems than one single algorithm. Google's guide to its ranking systems, last updated 10 December 2025, is built as a list of named systems, each with its own job.

The first job is understanding the query. The how-Search-works guide's own example is a search for bicycle repair shops, which shows different results to a user in Paris than to a user in Hong Kong. A query with a specific brand or URL pattern may need navigational results.
Matching pages to that meaning comes next. Page relevance still matters, so clear titles, headings and body copy help search systems understand what a page covers. An informational query may reward a clear explanation, while a commercial query may need comparison or product pages. Ranking systems are trying to satisfy the task behind the words, not just match the words themselves.
Links enter here, in the ranking systems guide's own words: "We have various systems that understand how pages link to each other as a way to determine" what a page is about and which pages best answer a query. The guide names "PageRank, one of our core ranking systems used when Google first launched." And it brings that up to date: "How PageRank works has evolved a lot since then, and it continues to be part of our core" ranking systems. A relevant editorial link can also help discovery and topic understanding. When you are earning quality backlinks, the link is strongest when the source page, anchor, surrounding context and destination all make sense together. What a page's links say about it over time is the Rhino Rank guide to link authority.
Quality has its own named history. Google's guide says of its helpful content system that "In March 2024, it evolved and became part of" the core ranking systems, and the Search Central blog post of 5 March 2024 said that month's core update marked "an evolution in how we identify the" helpfulness of content. The system's stated purpose, in Google's framing, is to favour original, helpful content written for people over content made mainly to gain search traffic. This is why organic SEO and better rankings depend on the whole page, not a single ranking factor. What to do with an indexed page that ranks below where it should is the Rhino Rank guide to improving Google rankings.
Serving is never guaranteed for any page. Being indexed gives a page the chance to appear. Freshness depends on the query: the guide's example of a search that needs fresh results is a film that has just been released, while other queries are stable, so an accurate older page can still be served.
How modern search results surface answers
Classic blue links are still part of search, but they now sit beside richer result formats. The mechanics still start with crawling and indexing. The presentation layer has changed. None of the richer formats changes the entry condition: a page has to be crawled and indexed before it can appear in any of them.
Featured snippets are the clearest case. Google's featured snippets documentation, last updated 10 December 2025, defines them: "Featured snippets are special boxes where the format of a regular search result is reversed, showing" the descriptive snippet first, ahead of the link. To the question of how to mark a page as a snippet, Google's answer begins "You can't." The systems decide which result, if any, goes in the box. You can only make the answer clear enough that Google's systems may select it.
AI Overviews and AI Mode follow the same entry rule. Google's AI features documentation, last updated 10 December 2025, states that "To be eligible to be shown as a supporting link in AI Overviews or AI Mode, a page must be indexed" and eligible to appear in Search with a snippet. Both features may use what Google calls "query fan-out", and both may show "a wider and more diverse set of helpful links" than a classic result. On markup the same page is blunt: "There's also no special schema.org structured data that you need to add." It also says: "There are no additional technical requirements."
In a Google post on 20 May 2025, Hema Budaraju wrote that "AI Overviews are now available in over 200 countries and territories, and more than 40 languages." Google's en-GB AI Overviews page says "Currently available in over 120 countries and territories, and 11 languages." The two Google pages give different counts.
Search engines beyond Google
Google is where most searches happen. The Statcounter chart headed "Search Engine Market Share Worldwide - August 2026" puts Google at "91.1%" of worldwide share.
Bing follows the same discover, index and rank pattern set out in the crawling section above, and DuckDuckGo's help page describes its traditional links and images as results "which we largely source from Bing".
DuckDuckGo is more than a relay, though. Its privacy page says its search "never tracks your searches", and its help page adds: "We also maintain our own crawler (DuckDuckBot) and many indexes to support our results."
Brave Search states the independent route on brave.com: "Brave delivers results from its own, built-from-scratch index." That sets it apart from engines that source their results from another index.
Ecosia keeps the retrieval shape and changes where the money goes, stating on ecosia.org: "We use all our profits for climate action, with the majority going into tree-planting projects around the world."
Statcounter's August 2026 chart shows Yandex at "0.99%" and Baidu at "0.62%" of worldwide share; their weight inside Russia and China is a different measure, and it is not on that chart.
How to use the three-stage model
The three-stage model is useful when a page will not show, or when it ranks far below where you expect. Work through the system in order. Do not start with ranking theories if Google has not found or indexed the page.

Start by checking crawl access. Can Google reach the URL? Is the page linked from the site? Is it in the sitemap? Does the server return a clean page? Is robots.txt blocking a path that should be open? Before a Rhino Rank order goes out we check the HTTP status code of the target URL, so a link never points at a dead page or a redirect chain.
Then check what Google indexed. Has Google selected the URL you want, or has it chosen another canonical? Is there a noindex tag? Does the mobile page carry the same main copy and links as desktop? Can the main content render without a failed script? Bućko's Onely test, 313 hours for the JavaScript folder against 36 for HTML, is the reason the last question is on the list.
Only then move to ranking. Does the page match the task behind the query? Is the answer clear enough? Is the page useful enough compared with the results that already rank?
This order saves wasted work. A crawl issue needs a crawl fix. An index issue needs an index fix. A ranking issue needs a better reason for Google to serve the page.
Frequently asked questions
What is a search engine?
A search engine is a system that helps users find information on the web. It discovers URLs, stores information about pages in an index, then returns results that its systems judge relevant to the query.
How does a search engine work?
A search engine crawls pages, indexes the content it can process, then ranks and serves results from that index when someone searches. Results adapt to context such as location, language and device. Discovery and storage happen before the query; ranking happens at the moment of it.
What is the difference between a browser and a search engine?
A browser is the software you use to visit websites, such as Chrome, Safari, Edge or Firefox. A search engine is a service you use inside a browser or app to find pages and answers.
How do search engines discover new URLs?
Search engines discover new URLs through links from pages they already know and through sitemaps that site owners submit. Site owners can also send URLs and sitemaps through search-engine tools. Discovery does not guarantee indexing: a discovered URL can still be crawled late, canonicalised to another address or left out of the index.
What is the role of web crawlers in search engines?
Web crawlers fetch pages so search engines can discover content, links and changes over time. The crawler supplies raw material for rendering and indexing; ranking happens later when a query is served.
related Blog Posts

Join 7,000+ Clients Growing with Rhino Rank
Sign UpStay ahead of the SEO curve
Get the latest link building strategies, SEO tips and industry insights delivered straight to your inbox.







