One of the things I love most about SEO is that it requires you to build mental models, based on combining insights from lots of different topics.

Links for example, can be spammy or not; and spamminess can be determined based on the source domain’s content, link profile and user experience.

One of the more interesting new models I’ve been working on is how indexation works, a topic which is arguably even more important due to AI grounding.

With AI now incentivising businesses to produce industrial levels of content, search engines are facing an exponentially growing web - which means their associated costs will be going up proportionally.

Assuming creating content at huge rates does become the norm, the question then becomes, how do they handle the increase?

Whilst not directly connected, it is noteworthy that as AI has begun to increase the amount of content being produced, Google’s “Crawled - currently not indexed” report seems to be becoming busier and busier.

This report in Google’s Page Indexing report shows site owners how many URLs have been found by Google, but for some reason they have decided not to index them.

Unfortunately, Google does not explain in their documentation for the report, what is the reasoning behind their decision - they instead simply state that there is “no need to resubmit this URL for crawling”.

This, as you can imagine, left my mental model for getting and keeping pages indexed wanting; so I decided to undergo some research into what is known publicly, and what dots can be connected from other known systems we are aware of running behind the scenes at Google.

The 3 steps to getting a page in Google’s SERPs

Showing in Search Engine Results Pages (SERPs) is obviously an integral part of being successful in SEO, and more recently, is also important for being included as part of grounding by AI platforms.

Before we begin getting into the details of Crawled - currently not indexed, it’s worth quickly establishing the steps to getting a page in Google’s results - so we can understand precisely at what point the issue is happening.

Crawling, Indexing and Serving

Search engines by and large all work in the same way. They start by crawling the web, using links to find new websites and URLs. They use their newly discovered list of URLs to build an index, so that they can organise, score and eventually serve URLs to users.

Click for a more in-depth explanation

Crawling: The first step is to be discovered and crawled. URL discovery happens because you either submit the URL to the search engine, or a link elsewhere on the internet exists that directs the search engines to your website/page. Once found, your URL is queued and eventually crawled.

Indexing: Once your page has been discovered and crawled, the content throughout is used to assess the topic, quality and originality of the page. If the page passes, it will be included in a table of eligible URLs called an index, and potentially used in results.

Serving: The last step is for the results to be sent down as part of a list of results. The order and position of a listing is dependent on hundreds of considerations that are scored during indexing, with some additional changes happening at the time of serving such as swapping out hreflang URL versions.


The crawled - currently not indexed report is related to the indexing stage of the process. Essentially, the search engine has found the URL, but has decided that it will not be including it within its indexes.

Is Crawled - currently not indexed a problem?

Seeing large numbers of pages which have been crawled but are not indexed is of course worrying, and seems at first glance to suggest something needs solving.

In an episode of Search Off the Record podcast (How to read the Indexing Report), Gary and John explain that indexing reports are intended as a tool to alert you to changes.

The podcast explains that the reports are meant to help with discerning patterns or trends, which you then must apply your own judgement as to whether they are an issue or not. So for example, a page with a meta robots tag instructing Google to not index the page, could potentially show in this report as the page has been discovered, crawled and not indexed. If this tag was added deliberately, then there isn’t an issue to fix.

Conversely, if lots of pages are being shown, some of which are valuable to the business and supposed to be indexed, then it may suggest an issue that means Google is unhappy with the site at large.

Known causes of pages becoming Crawled - currently not indexed?

After reading quotes from Googlers, and as many posts as I could get my hands on, I’ve collected what looks to be the most popular reasons for what causes pages to become crawled but not indexed.

  • When the count of pages is large, includes URLs that are important to the business and does not seem to be specifically tied to a given type of page - then the issue may be due to overall site quality.
  • When the count of pages are small, or the list is all of a similar type, this is potentially due to those pages being seen as low quality.
  • When Google’s perception of the domains has changed, for example. the technical performance has slipped or the content is less valuable (e.g. pages haven’t been updated).
  • There are issues with content quality; too many thin pages, duplicate content or poorly written copy which is thin on value
  • There is a technical issue that is causing Google to be concerned, the example being given that the same content is showing for every URL (think bot protection “Are you a bot” challenge screens).
  • The structure of the website and the discoverability of URLs throughout the website.
  • Connected to the structure of the website, the internal linking to the page.
  • Another, higher quality, page is being preferred, be that on the same sub-domain or in the SERPs.
  • The team working on the website have done something to prevent the page from being indexed, e.g. included a meta robots tag with a noindex value.

Reflecting on the above, and discounting instances where pages are being excluded because of an action taken by the site owner, the common theme is an issue which broadly can be described as related to the site’s quality.

Considering what we know about Google’s metrics and systems related scoring site quality, gleaned from the DOJ hearings and their content warehouse leak, it’s worth exploring further what specific considerations map to fixing and improving yours.

Google’s Authority System

Through the Content Warehouse leak and statements made during Google’s DOJ hearing, we have by and large confirmed that Google has two systems which are used for assessing URLs.

The Authority System is used to provide a query-independent signal of the sub-domain’s authority - providing a score based on aggregating the data it receives from various sources and systems.

The Relevance System is focused on establishing the topical connectedness and authority of both the URL and wider domain.

The metric associated with authority is Q* (pronounced ‘Q Star’). This metric is described as being query-independent (it doesn’t change based on query), and as receiving data from various quality and trust related systems, with the most cited and relevant being contentEffort, OriginalContentScore, chard (believed to be related to integrity and trust) and predictedDefaultNsr (a metric that tracks performance over time and is believed to drive confidence as to whether a domain is good, bad or in-between).

The aim, according to testimony taken from Googlers in DOJ hearings, of Q* is to score the credibility and utility of the domain, by aggregating inputs from algorithms and systems like Page Rank, and those that assess content quality and trust signals.

When asked, Google’s engineering and Product teams have described Q* as “…incredibly important” and “…hugely important even today” in court hearings.

Indexing tiers & Mustang

In addition to the above quality metrics, we also are aware of systems that are believed to be involved in classifying URLs into indexes; based on quality signals.

“Base”, “Zepplins” and “Landfills” are different index tiers which are believed to enforce different hard drive requirements based on a policy of how often they are to be updated.

For Base domains, they are believed to sit on Flash Memory and can be updated frequently. Zepplins are believed to be on SSD hard drives, with updates and changes made less frequently (for example at weekly intervals). Lastly, landfills are standard hard drive and rarely updated.

The systems responsible for tiering pages are called SegIndexer and TeraGoogle, and it is believed that they help with organising data into the most performant form of storage based on the quality scores associated with each URL.

So where does this leave us?

Assuming the above checks out, we know that Google has scores they collect and widely distribute that focus on quality at the URL level; and these scores are contributors to where pages are stored in their indexes - with each scoring level indicating how often a page is adjusted.

Bringing together the above anecdotal feedback on causes of URLs becoming not indexed, the information we have about how Google collects and scores quality metrics and how they affect indexing - it seems more than likely that optimising for Q*, is the way forward.

The REVIVE framework

After digesting all of the above information, I began putting together a framework with which to support our clients.

Our REVIVE framework has as its aim to step us through a conditional process for finding, qualifying and remediating issues that we expect contribute to Q* related sources.

The aim with the framework is to systematically remove issues which we believe due to data and testimony, are key to improving URL and site quality scoring.

The steps are outlined below.

Research: Due to the Crawled - currently not indexed report using data which is less frequently updated than the URL Inspection API, we start by crawling the domain and testing each URL’s indexation status.

N.b. We use a tool that we have developed which uses 5 different strategies for establishing whether a URL is indexed, so we can scale our efforts quickly - see Indexation Monitor.

Evaluate: Being included within the report does not necessarily mean something is wrong. For this reason, our second step is to cohort URLs and label them based on template and whether they are “core” or “Miscellaneous”. We review the pages found, assess to what extent they indicate an issue likely to be problematic for Google, and provide a mini-report indicating the commonalities amongst the affected pages and expected impact.

Verify: With affected URLs shortlisted and organised, we then begin gathering evidence and creating a scorecard to measure the issues present on the domain.

Typically this involves a review of the following:

  • Technical issues on site
  • Content quality issues on site
  • Spamminess and relative power of their link profile
  • Commonalities across affected pages
  • Drops in traffic that relate to relative algorithm updates
  • contentEffort scores (we have a tool we’ve built specifically to measure and score contentEffort).

Improve: Based on our verify step, we shortlist, categorise and create tickets and briefs to resolve the issues found. Each ticket is related back to either a metric, patent or specific quote that corroborates the topic as being potentially relevant and related to improving Q* scoring.

Validate: Because signals relating to overall site quality can take a significant amount of time to be reassessed, we install and setup indexation monitoring after the improvements are implemented.

Quote from John Mueller

N.b. in the SEO Community recently John Mueller shared that it can take more than a year for a domain to recover.

We track whether the number of unindexed pages is increasing, decreasing or remaining stable. When a change occurs, we investigate the likely cause and determine whether it can be linked to the work happening or an external update.

Embed: Many of the metrics contributing to site authority and quality evaluate performance over time. Sustainable improvement therefore requires addressing the processes and ways of working that allowed the issues to develop in the first place.

Training, regular auditing and the creation of clear standard operating procedures are worthwhile to prevent recurring problems, as well as to protect and develop your algorithmic momentum.