What to expect

  • A case is put forward that optimisation consistency is a signal or data point used by Google to build confidence that a result is a good response
  • We go through a quick deep-dive into how Machine Learning thinks, and how what we know might transfer to SEO
  • Assuming the above is correct, I explore what does this change?

The author WIll Durant, whilst explaining a quote from Aristotle, suggested that “We are what we repeatedly do…therefore excellence is not an act, but a habit”.

I’m a fan of this quote because it reinforces the idea that grittiness is a better predictor of success than being smart or having anything that resembles an innate advantage. As someone with an internal locus of control, this appeals because it agrees with the idea I have agency over how things go.

During my time working on SEO for clients and brands, I’ve for as long as I can remember always framed the work as we fix and improve what we inherit, and then look to build engines that keep their website ahead of competitors.

I like the language of engines because it strips away the geeky details, leaving only the idea that the goal is to have systems, rituals and tooling in place to keep the website winning. Engines bring to mind out-competing your competitors by having a more powerful, better maintained, system - which makes the “winning formula” less about titles and meta descriptions and more about working happening and teams being cohesive.

The aim with this post is to make a case that engine building is not just a shorthand for being organised and proactive, but it’s also an optimisation that creates resilience because of metrics we know Google use for scoring site quality.

Crawled / Discovered - Currently Not Indexed

Whilst helping a client with understanding the cause of, and how to reduce the number of, their pages which were crawled/discovered and not indexed, I went looking into what is the common wisdom and lore about why Google is dropping pages.

Broadly speaking, what came from my research is that it’s related to quality and trust - both of which come together to inform whether a URL is worthwhile crawling or not.

This tracks as it makes sense to avoid indexing URLs which are low quality, especially considering that lots of URLs and speed are inversely correlated.

Whilst helpful, I got the feeling from what I was reading that the advice was generic and parroted from official lines from Google’s public facing teams, so I decided to go looking through DOJ and leaked Content Warehouse docs to see what first principles could be gleaned that I could build my own world-view from.

Composite Docs

Thanks to the exceptional work done by Shaun Anderson over on Hobo-Web, I was able to quickly find out that according to leaked documents from Google, pages are organised into composite docs, with each doc having associated attributes that are essentially scores, values or labels.

As I understand it, algorithms and systems are either consumers or contributors to these scores, the provided data indicating things like the anchors pointing to the page, indexing directives and snippets usable in feature snippets.

predictedDefaultNsr

One attribute that caught my eye was predictedDefaultNsr, which is described as being a quality score that scores a page using the trended performance of that page over time.

This metric is interesting because it suggests that a part of the quality score attributed to a page or domain, could be judged based on how you perform over a set timeframe - either allowing for some leeway should something break temporarily or requiring longstanding good scoring to justify sitting in higher positions.

Curiously, the docs describe this as creating ‘algorithmic momentum’, a term which does have a meaning outside of Google. Curiously, one of the main topics associated algorithmic momentum is Machine Learning.

Gradient Descent?

When training AI models, algorithms are built through repeatedly testing which combination of inputs results in an output that is most closely similar to a control group.

If for example you have 20 URLs, 10 which were agreed to be Spam and 10 that were agreed as not spam. You also gave the algorithm several tests that can be used to detect spam, such as content quality, author scores or how relevant the topic is to the wider domain.

The system would essentially begin attempting each of the tests using different weightings, mixing and matching until it has either found a combination that results in 100% correct or the smallest possible difference. As the matching is happening, the systems will be attempting to reduce the amount of “loss” it experiences - which in our example means they would be trying to reduce the number of URLs that erroneously labelled. When loss is small, they will continue iterating, comparing previous tests to current, and adjusting depending on the result. When Loss is large, they will look for a new strategy with a smaller amount of loss.

The process of reducing loss to find the right balance of weights and parameters is called gradient descent, and is itself a well defined algorithm that comes from mathematics.

Algorithmic momentum

Whilst gradient descent is a relatively robust approach to finding answers, it is potentially very wasteful; as time is spent checking different approaches without the benefit of the context from previous tests.

Put another way, each test is essentially only aware of itself, so it cannot test and determine how well things are going by looking at how a line of testing is trending.

Ideally, the algorithm would be clever enough to either double down when an approach is showing promise, or conversely stop when the loss is too high and doesn’t seem to be getting significantly closer.

Here is where algorithmic momentum helps - the trending data that assesses performance based on an average over time.

In plain English, algorithmic momentum suggests that the losses of previous attempts can be used to determine if an algorithm is miles-off, almost there, or somewhere in between. Depending on the momentum of the scoring, it can diverge the direction the algorithm is taking - reducing wastage and growing confidence in the suggested changes.

Algorithmic momentum and Spam

Continuing with our spam example, algorithmic momentum could be used to determine from 3 different strategies - which looks to be both within the right score range and most dependably scoring with minimal loss.

This could then allow the algorithm to focus more efforts on making small improvements to the winning algorithm and/or changing tact to another line of testing when results are too far off.

What does this have to do with SEO?

Assuming that PredictedDefaultNSR is a score/attribute (and that it is used) which represents a quality score that takes into consideration your performance over time, this then means that my ‘engines’ are actually speaking directly to a quality signal (i.e. confidence of quality).

This tracks with how systems like Navboost and Glue are described in other resources, and makes sense as a heuristic for how trustworthy is this URL’s/domain’s score.

Additionally, algorithmic momentum makes sense in the negative too, as a way to confidently determine that a website is spam. If for example signals coming from the last 30 days of a website’s publishing is poor and irrelevant, and that tracks with the scores seen in the previous 30 days, then the inference that the URL/domain is spam should stronger or more confident.

Ultimately, algorithmic momentum could be seen as a ranking factor/signal that brings to the fore the importance of consistency.

What does this change?

It is my belief that creating momentum is more about how you work, over doing new things.

One notable change however is that momentum justifies the priority of maintenance, legitimising it as a score that must be worked towards every day/week/month to avoid momentum dissipation. A growing list of outdated pages or screaming frog issues now, can be seen as a slowing force that will hurt momentum.

Framed positively, momentum also is a useful motivator for teams; that translates in an accessible way into everything from technical debt, content optimisation, internal linking, offpage and more. Framing winning as the small stuff being done better than your competitors is not always that sexy, so gamifying it into speed makes the business feel high performance and like it is speeding ahead.

Whilst we can never know our official scores that are internal to Google, it’s my belief that just being aware that this is how your efforts are scored changes the game you’re playing. It changes the game from being all about the new, to how well you do the things that matter, or, how robust and efficient your engines are.

Categories:
Tags: