SpamBrain: How Google Says Its AI Neutralizes Link Spam

What Google claims about its link-spam AI — and what two decades of Google's own patents suggest it is actually made of.

Ask a business owner what they fear from bad backlinks and they'll describe a punishment: a penalty, a manual action, a site that gets "slapped." That fear isn't baseless — Google still issues manual link penalties, as recently as last year. But it describes the exception, not the rule. The common fate of a bad link today is quieter and stranger: Google simply refuses to count it. No penalty. No notice. The link just sits there, doing nothing.

The system Google credits for most of that quiet work is SpamBrain. And here is the first honest thing to say about it: almost everything written about how it works is a guess — including the accounts that sound most certain. SpamBrain has no patent. It has no research paper. It never surfaced in the DOJ antitrust trial, where Google was put under oath about how Search works. Its entire public existence rests on a handful of Google blog posts and one or two field names in the 2024 Content Warehouse API leak.

So this guide does two things. First, it pins down exactly what Google has said — with the receipts, because a claim without a source is just a rumor. Then it does the more useful thing: it reasons about what a link-spam classifier is most plausibly made of, using the one body of evidence Google can't walk back — the link-evaluation systems it spent twenty years patenting and shipping. That is where a defensible picture of SpamBrain actually comes from.

Sober Summary

On a system with no patent and no court record, the honest thing is to separate what Google asserts from what we can actually substantiate. Three tiers, kept apart on purpose.

What Google Claims (its own words)

That SpamBrain exists, is AI/ML-based, launched in 2018, and — since 14 December 2022 — is used to "neutralize the impact of unnatural links," detecting "both sites buying links, and sites used for the purpose of passing outgoing links," in all languages. That it caught "50 times more link spam sites" than the previous update, and that ranking benefit from neutralized links "cannot be regained." These are Google's figures and Google's framing. None of them are independently audited.

What Our Own Research Substantiates (patents + leak)

The components a link-spam classifier would need are real, granted, and in several cases production-confirmed: per-link click-probability weighting, link-context fingerprinting, click-validated link quality, seed-distance trust, link velocity, topical-relatedness checks, and independence detection. Systems theory says a working complex classifier is composed from simpler working systems — and Google has patented all of these. The 2024 leak also carries at least one SpamBrain-named attribute (spambrainLavcScore, on the NSR site-data).

What Nobody Can Verify

That SpamBrain is actually built from those components — that mapping is our reasoned inference, not a Google disclosure. Its true architecture, inputs, thresholds, and whether "neutralize" means devalue-to-zero or something more. And the largest unknown: how far to trust the narrative at all, given Google spent years denying it used clicks in ranking until the DOJ trial put a Google VP under oath and the clicks came out. Link spam never got that scrutiny.



The Frame

SpamBrain is a name, not a blueprint. Google talks about it the way you'd talk about a black box — here's what it does, never here's how. But a spam classifier that spots bought and sold links can't have come from nowhere. Complex systems that work are always assembled from simpler systems that already worked. Google has spent twenty years building and patenting those simpler systems in plain sight. The most honest way to understand SpamBrain is not to parse Google's press releases — it's to read the parts list Google already published.


What Google Actually Said (With Receipts)

Google's first-party record on SpamBrain is small and specific. It lives in a few Search Central blog posts and one documentation page. Every quote below is verbatim, and every one has a capture beneath it — because on a topic this thin on evidence, "Google said so" only counts if you can see where.

The naming (April 2022)

SpamBrain was named publicly for the first time in Google's 2021 webspam report:

Google Search Central — 21 April 2022

"SpamBrain was launched in 2018 and we've been continuously improving its performance. In 2021, SpamBrain identified nearly six times more spam sites than in 2020."

Screenshot from Google Search Central's 2021 webspam report stating SpamBrain was launched in 2018.
Google Search Central, "How we fought Search spam on Google in 2021" (21 April 2022). Note the framing: SpamBrain is described, never diagrammed.

This post is not about links. The link-spam update it references is the July 2021 one, which pre-dates SpamBrain's involvement with links entirely. In April 2022, SpamBrain was a general spam system.

The link pivot (December 2022)

Eight months later, Google extended it to the link graph. The December 2022 link spam update is the single most-cited document on this subject:

Google Search Central — 14 December 2022

"Today, we're leveraging the power of SpamBrain to neutralize the impact of unnatural links on search results."

Screenshot from Google's December 2022 link spam update announcement using the word neutralize.
Google Search Central, "December 2022 link spam update" (14 December 2022). The operative verb is neutralize — not penalize.
Google Search Central — 14 December 2022

"SpamBrain is our AI-based spam-prevention system. Besides using it to detect spam directly, it can now detect both sites buying links, and sites used for the purpose of passing outgoing links."

Screenshot from Google stating SpamBrain can detect both sites buying links and sites passing outgoing links.
Both sides of the transaction — buyers and sellers. The seller half is the one the industry keeps forgetting.

Two things carry the weight here. Google names both sides of the transaction — a site "used for the purpose of passing outgoing links" is in scope, so a site that sells links can see its whole outbound profile devalued. And the verb is neutralize: the credit a link passed is "lost," the link stops counting, and — Google is explicit — this is not framed as a site-level penalty.

The scaling claim (April 2023)

In its 2022 webspam report, Google attached a number:

Google Search Central — 11 April 2023

"Thanks to SpamBrain's learning capability, we detected 50 times more link spam sites compared to the previous link spam update."

Screenshot from Google's 2022 webspam report claiming 50 times more link spam sites detected.
"50 times more" — measured against the pre-SpamBrain July 2021 update, self-reported, and never externally audited. Read it as a marketing figure, not a benchmark.

"50 times more" is the number everyone repeats. It compares against the July 2021 update, which ran before SpamBrain touched links — an AI-versus-pre-AI comparison, Google's own, with no outside audit. It is also the last SpamBrain scaling figure Google published; the annual "How we fought spam" report series ended with this edition.

The permanence (standing documentation)

The claim that turns this from a news item into a strategy problem lives in Google's maintained spam-updates documentation:

Google Search Console documentation

"When our systems remove the effects spammy links may have, any ranking benefit the links may have previously generated for your site is lost. Any potential ranking benefits generated by those links cannot be regained."

Screenshot from Google documentation stating ranking benefits from neutralized links cannot be regained.
Google Search Central documentation, "Google Search spam updates and your site." The phrase to sit with: cannot be regained.

The Correction: Penalties Aren't Dead, They're Rare

A lot of SEO writing gets this wrong, claiming link penalties are a thing of the past. They aren't. Google still polices links; what changed is the default response — from a penalty to silent devaluation.

The proof is in Google's own Manual actions report documentation, which still lists both "Unnatural links to your site" and "Unnatural links from your site" as live categories, with the reconsideration process intact. Google still issues them — recoveries from "unnatural links" manual actions continue to be documented. And John Mueller, asked in 2024 and again in 2025 when the disavow tool is worth using, keeps naming the same single scenario: when you have a manual action for links you actually placed.

The accurate version

Manual link penalties still exist and are still issued — for blatant, obvious, or reported schemes. But since 2016, the standard response to a bad link is silent algorithmic devaluation, not a penalty. A manual action is now the exception a human reviewer reserves for the egregious cases; being quietly ignored is what happens to everyone else.


The Decade-Long Shift in the Default

SpamBrain didn't invent silent link handling. It inherited it. Google's default response to a manipulated link moved through three eras, and the decisive move happened years before SpamBrain touched a link.

Stylized three-stage timeline in a dark editorial style titled 'Penalty to Devaluation to Neutralization,' tracing Google's default response to a bad link across three eras: Penalty (2012–2016), when a bad link could demote the whole site and recovery took months; Devaluation (2016+), when Penguin folded into core and bad links were ignored rather than punished, in real time; and Neutralization (2022 onward), when SpamBrain targets both link buyers and sellers and the passed credit is lost and cannot be regained.
Google's default response to a manipulative link across a decade — whole-site penalty, then silent devaluation, then SpamBrain neutralization. The decisive shift to devaluation happened in 2016, years before SpamBrain touched a link.

The 2012 update — Matt Cutts's "Another step to reward high-quality sites" — announced "an important algorithm change targeted at webspam" that would "decrease rankings for sites." Worth a small honesty note: that post never used the word "Penguin." The nickname was applied afterward by the SEO press. Google shipped a webspam update; the industry named the animal.

The turn that matters came on 23 September 2016, when Gary Illyes announced "Penguin is now part of our core algorithm":

Google Search Central — 23 September 2016

"Penguin is now more granular. Penguin now devalues spam by adjusting ranking based on spam signals, rather than affecting ranking of the whole site."

Screenshot from Google's 2016 announcement that Penguin now devalues spam rather than affecting the whole site.
Google Search Central, "Penguin is now part of our core algorithm" (23 September 2016) — the moment devaluation, not punishment, became the default. The same post stated, "we're not going to comment on future refreshes," and there hasn't been a named Penguin update since.

That is the real pivot. From 2016 on, link-spam enforcement went quiet and continuous. SpamBrain, in 2022, is best read as the next step in that same direction — a learned system generalizing the ignore-by-default model and pointing it at both ends of the link trade. Same instinct. Bigger engine.


What SpamBrain Is Most Likely Made Of

Here is the part almost every SpamBrain explainer skips, because it requires having done the homework. If SpamBrain is a machine-learning system that spots link buyers and sellers, then by definition it is a complex system. And complex systems have a law: the working ones are always assembled from simpler systems that already worked on their own. You don't train a model to "detect all link spam" from a blank slate. You compose it from narrow, individually-validated signal generators.

Google spent two decades building exactly those generators — and patenting them in the open. Our research library is, in effect, a catalogue of SpamBrain's plausible parts. None of this is a Google disclosure that SpamBrain uses them; it is the systems-theory argument that a working classifier would be built from precisely the components Google already had on the shelf.

Start with the strongest single piece of evidence: more than a decade ago Google patented a per-link quality model — the Reasonable Surfer model — that weights each individual link by how likely it is to actually be followed, judged from the link's own features and user behavior rather than treating every link equally. A learned model that judges link quality one link at a time is not something SpamBrain would need to invent. It's something it would inherit.

The rest sort cleanly by the signal family a spam classifier would need — and each one is a system already sitting in our research library:

The signal a link-spam classifier needs Where Google already built it
Anchor text & velocity
How anchors are distributed, and how fast they appear.
Historical Data — timing, spikes, rotating links · Click-Validated Quality — anchor n-gram spam
Link context & relatedness
The editorial text around a link, and whether the anchor matches the destination's topic.
Reference Contexts — link-context fingerprinting · Phrase-Based Indexing — anchor-to-page topic match
Source, resource & seed quality
A page's worth as a link source vs. destination, and its distance from trusted seeds.
Seed Distance — trust proximity · Entity Trust — the endorser's trust · Click-Validated Quality — source scoring
Referral traffic & click validation
Whether anyone actually clicks the link.
Click-Validated Quality — genuine long-dwell clicks · Reasonable Surfer — click-probability weighting
Independence & networks
Whether links are genuinely separate, or one operator in disguise.
Entity Trust — anti-collusion · Reference Contexts — context diversity · Implied Links — profile vs. brand demand
Time
The tempo of acquisition, and domain legitimacy.
Historical Data — velocity, freshness, DNS & registration
Stylized systems-theory diagram in a dark editorial style titled 'What SpamBrain Is Most Likely Made Of.' Eight documented, patented Google systems in solid gold boxes — Reasonable Surfer (per-link click-probability weighting), Historical Data (velocity, freshness, spikes), Reference Contexts (link-context fingerprinting), Phrase-Based Indexing (anchor-to-topic match), Seed Distance (trust proximity), Entity Trust (endorser trust and anti-collusion), Click-Validated Quality (genuine long-dwell clicks) and Implied Links (profile versus brand demand) — feed a dashed box marked 'inferred: learned composition layer,' which yields one probabilistic judgment: genuine editorial endorsement versus manipulative link. A legend marks solid boxes as documented and patented, and the dashed box as reasoned inference, not a Google disclosure.
The systems-theory reading: documented components (solid) composed by an inferred learned layer (dashed) into a single probabilistic judgment — a reasoned inference, not a documented SpamBrain architecture.
The Core Inference

SpamBrain doesn't need to catch the payment. It reads the shape of the transaction in the link graph — and every feature it would need to read that shape is a system Google already built. The most defensible picture of SpamBrain is a neural layer composing these documented signals into one probabilistic judgment, where Google once used hand-tuned thresholds. That is an inference — a well-grounded one, but an inference, not a wiring diagram.


A careful point, because it's routinely mangled. None of these systems detects whether money changed hands. They estimate something else entirely: whether a link is a genuine editorial endorsement that produces real signal. That distinction is the whole game.

Run the parts list back and the pattern is obvious. A link that sits in relevant editorial context, on a page with real traffic, from an independent and trusted source, that people actually click — passes every one of those documented checks. Whether someone paid for that placement is invisible to the machine, because none of the signals measure payment. What gets neutralized is the manipulative link: irrelevant, unread, dropped into a template on a site with no independent trust. That link fails the checks whether it was bought, traded, or begged for free.

Stylized two-column comparison in a dark editorial style titled 'The Link That Survives — and the One That Gets Subtracted.' The Survives column (green) lists: sits in relevant editorial context; on a real, trafficked page; from an independent, trusted source; earns genuine, long-dwell clicks — nothing artificial to subtract. The Neutralized column (red) lists: irrelevant to the destination; unread and never clicked; dropped into a template; on a trustless, isolated site — it fails the checks whether paid, traded, or free. The dividing line is genuineness, not money.
None of the systems measure payment — they estimate genuine editorial endorsement. A link with real relatedness, traffic, independence, and clicks has nothing artificial for SpamBrain to subtract; a manipulative one fails whether it was bought or free.
What survives

The durable link isn't the free one — it's the genuine one. A placement that would earn its keep across relatedness, source trust, real clicks, and independence has nothing artificial for SpamBrain to subtract, because everything it's made of is real. The failure mode isn't "you paid for it." The failure mode is "nobody would have linked to it on the merits."


The Skeptic's Ledger: Claimed vs. Verified

Because SpamBrain is described but never documented, precision about who is telling us what matters more than usual. Sort every common claim into one of three bins.

Google-claimed (unaudited)
  • SpamBrain exists and is AI/ML-based
  • It neutralizes link spam at scale
  • "50× more link spam"; "99% spam-free"
  • Buyers and sellers both detected
VS
Independently substantiated
  • The component patents are real and granted
  • Several are production-confirmed in the leak
  • One SpamBrain-named leak field (spambrainLavcScore)
  • Manual link actions still documented & issued

Why the skepticism isn't cynicism: Google has a track record. For years its representatives said clicks weren't used in ranking. Then the DOJ antitrust trial put a Google VP, Pandu Nayak, under oath, and NavBoost's use of click data since roughly 2005 entered the record. The public story and the sworn story diverged. Link spam never got that test — the DOJ case was about monopoly maintenance through default-placement contracts and the data-scale advantage they create, not about how the ranking algorithm polices links. SpamBrain was categorically outside the case, so the record is simply silent on it. It has never been cross-examined.

The honest position

SpamBrain's existence and ML nature are Google-stated. Its scale numbers are Google's own, unaudited. The component composition in this article is our reasoned inference from patents and the leak — not a documented architecture. And independent, third-party verification of how it actually works does not exist. Believe the parts list, because Google published it. Treat the black-box narrative as a claim, because that's all it is.


The Recovery Trap

Neutralization has a quiet cruelty a penalty doesn't. A penalty is visible and, usually, reversible. Silent devaluation is neither.

Manual action (the exception)
  • Shows up in Search Console
  • Reversible — clean up, request reconsideration
  • Disavow is a real part of the fix
  • You know what happened
VS
Algorithmic neutralization (the rule)
  • No notice, anywhere
  • Credit "cannot be regained" — Google's own words
  • Disavow does nothing; there's no penalty to lift
  • You often can't confirm it happened

The disavow tool is the casualty of the confusion. It's a manual-action instrument — for when you bought links, got a Search Console notice, and can't remove them. It does not reverse algorithmic devaluation, because there's nothing to clean up when Google has simply stopped counting a link. That's the trap: a business owner refreshing Search Console for a warning that never comes, quietly funding authority that stopped existing months ago. No villain to point at, no letter to appeal. Just a slow leak nobody mentioned — and understanding the mechanism is the one thing that hands that person back some control.

Stylized decision diagram in a dark editorial style titled 'Should You Disavow?' A central question — do you have a Search Console manual action for links you actually placed? — branches two ways. Yes, a manual action (the exception): a human reviewer flagged your link scheme, so remove what you can, disavow the rest, and request reconsideration; it is recoverable and disavow is part of the fix. No, algorithmic devaluation (the rule): the link was quietly neutralized with no notice and already passes no value, so there is nothing to clean up; disavow does nothing and the credit cannot be regained. Reserve disavow for a real manual action on links you placed.
The disavow tool fixes exactly one situation — a manual action for links you placed. Against silent algorithmic devaluation, there is nothing to disavow.

What Most SpamBrain Guides Get Wrong

The genre treats SpamBrain as a revealed truth — "Google's AI does X, Y, Z" — and sells you protection from it. Three errors run through nearly all of it.

It overstates the certainty: there is no patent, no paper, no testimony, so any confident account of SpamBrain's mechanism is guesswork dressed as fact. It gets the verb wrong: the everyday risk is silent devaluation, not penalty, and the two demand opposite responses — one you fix, the other you can only out-earn. And it skips the homework: it never connects SpamBrain to the documented systems that would actually compose it, so it can't tell you anything durable about how to build links that survive.

A smaller error shows how the confident-guessing spreads: a patent application, US20240020476A1 — since granted to Pinterest as US12572741B2 — gets passed around as "SpamBrain's patent." It's assigned to Pinterest, not Google. The lesson isn't that someone slipped — it's how badly we want to bolt an official-looking artifact onto a system that has none. The credible writers cite no patent, because there isn't one.

The frame that survives contact with the evidence is quieter than the hype: Google mostly stopped punishing bad links a decade ago and started ignoring them, and the AI it credits for that is real but undocumented — a black box we can only understand by the parts Google already published. Which lands on the one discipline this whole piece is built to defend. Clarity here isn't a style. It's the practical difference between knowing what you're doing and repeating what you were told.


The Link-Spam Reality Checklist

Pulled from everything above — sorted by the belief each item corrects.

Get the model right

  • Assume a bad link's default fate is to be ignored, not penalized.
  • Know that manual link penalties still exist — reserved for blatant, obvious, or reported schemes.
  • Treat every "SpamBrain does X" claim (including Google's) as a hypothesis until it maps to a documented system.

Judge links by genuineness, not by whether they're paid

  • Ask of any link: would this exist if the machine were reading it as an editorial endorsement?
  • Require topical relatedness between the linking context and the destination.
  • Require a real, independent, trafficked source — not a template on a trust-less site.
  • Favor placements that earn genuine clicks; unread links validate nothing.

Respect the documented systems

  • Keep acquisition gradual and un-patterned — velocity spikes are the oldest tell in the book.
  • Diversify the editorial contexts your links live in; homogeneous contexts read as machine-made.
  • Keep your link profile proportional to genuine brand demand.

Don't fall for the recovery trap

  • Don't disavow to fix silent devaluation — it only addresses manual actions.
  • Don't expect neutralized credit to return; Google says it can't.
  • Reserve disavow for a real Search Console manual action on links you placed.

Frequently Asked Questions

What is SpamBrain, exactly?

The name Google gives its AI spam-prevention system. Google says it launched in 2018, was named publicly in April 2022, and was extended to link spam on 14 December 2022. There is no patent, no paper, and no court record for it, so everything about its internals is inference. Its existence rests on Google's own statements plus a field name or two in the 2024 API leak.

Does Google penalize you for spammy backlinks?

Rarely. Manual link penalties still exist — Search Console still lists "Unnatural links to your site" and "from your site," and Google still issues them for blatant schemes. But since Penguin 4.0 in 2016, the default response is silent algorithmic devaluation: the link is ignored, passes no value, and you get no notice. Penalty is the exception; being ignored is the rule.

Can buying links still work?

Yes. The systems don't detect payment — they estimate whether a link is a genuine editorial endorsement that produces real signal. A paid placement on a relevant, trafficked, independently-run site can pass value; a cheap, irrelevant, unread link gets ignored whether it was paid for or free. The line is quality and genuineness, not money.

How does SpamBrain actually work?

Nobody outside Google knows, and anyone who says otherwise is guessing. The most defensible answer comes from systems theory: a working complex classifier is assembled from simpler working systems, and Google spent two decades patenting exactly the ones a link-spam classifier would need — per-link click weighting, context fingerprinting, click validation, seed-distance trust, velocity, independence. SpamBrain is most plausibly the learned layer composing those. That's a reasoned inference, not a documented fact.

Is there a patent for SpamBrain?

No. There is no publicly identified Google patent for SpamBrain. The application sometimes cited as its patent — US20240020476A1, since granted to Pinterest as US12572741B2 — is assigned to Pinterest, not Google. Credible analysts cite no patent, because none exists.

Can I recover ranking lost to a link spam update?

Not the part that came from neutralized links. Google's documentation states that once its systems remove the effect of spammy links, "any potential ranking benefits generated by those links cannot be regained." You recover by earning new, genuine signals — not by scrubbing the old ones. Disavow only helps against a manual action.

Why be skeptical of Google's SpamBrain claims?

Because Google's public statements about ranking have diverged from reality before. For years it said clicks weren't a ranking factor; the DOJ antitrust trial put a Google VP under oath and the opposite came out. Link spam never faced that scrutiny — SpamBrain has no patent, no paper, and no court record. Its scale figures are self-reported and unaudited. Trust the documented parts; treat the black-box story as a claim.