Ask a business owner what they fear from bad backlinks and they'll describe a punishment: a penalty, a manual action, a site that gets "slapped." That fear isn't baseless — Google still issues manual link penalties, as recently as last year. But it describes the exception, not the rule. The common fate of a bad link today is quieter and stranger: Google simply refuses to count it. No penalty. No notice. The link just sits there, doing nothing.
The system Google credits for most of that quiet work is SpamBrain. And here is the first honest thing to say about it: almost everything written about how it works is a guess — including the accounts that sound most certain. SpamBrain has no patent. It has no research paper. It never surfaced in the DOJ antitrust trial, where Google was put under oath about how Search works. Its entire public existence rests on a handful of Google blog posts and one or two field names in the 2024 Content Warehouse API leak.
So this guide does two things. First, it pins down exactly what Google has said — with the receipts, because a claim without a source is just a rumor. Then it does the more useful thing: it reasons about what a link-spam classifier is most plausibly made of, using the one body of evidence Google can't walk back — the link-evaluation systems it spent twenty years patenting and shipping. That is where a defensible picture of SpamBrain actually comes from.
Sober Summary
On a system with no patent and no court record, the honest thing is to separate what Google asserts from what we can actually substantiate. Three tiers, kept apart on purpose.
That SpamBrain exists, is AI/ML-based, launched in 2018, and — since 14 December 2022 — is used to "neutralize the impact of unnatural links," detecting "both sites buying links, and sites used for the purpose of passing outgoing links," in all languages. That it caught "50 times more link spam sites" than the previous update, and that ranking benefit from neutralized links "cannot be regained." These are Google's figures and Google's framing. None of them are independently audited.
The components a link-spam classifier would need are real, granted, and in several cases production-confirmed: per-link click-probability weighting, link-context fingerprinting, click-validated link quality, seed-distance trust, link velocity, topical-relatedness checks, and independence detection. Systems theory says a working complex classifier is composed from simpler working systems — and Google has patented all of these. The 2024 leak also carries at least one SpamBrain-named attribute (spambrainLavcScore, on the NSR site-data).
That SpamBrain is actually built from those components — that mapping is our reasoned inference, not a Google disclosure. Its true architecture, inputs, thresholds, and whether "neutralize" means devalue-to-zero or something more. And the largest unknown: how far to trust the narrative at all, given Google spent years denying it used clicks in ranking until the DOJ trial put a Google VP under oath and the clicks came out. Link spam never got that scrutiny.
SpamBrain is a name, not a blueprint. Google talks about it the way you'd talk about a black box — here's what it does, never here's how. But a spam classifier that spots bought and sold links can't have come from nowhere. Complex systems that work are always assembled from simpler systems that already worked. Google has spent twenty years building and patenting those simpler systems in plain sight. The most honest way to understand SpamBrain is not to parse Google's press releases — it's to read the parts list Google already published.
What Google Actually Said (With Receipts)
Google's first-party record on SpamBrain is small and specific. It lives in a few Search Central blog posts and one documentation page. Every quote below is verbatim, and every one has a capture beneath it — because on a topic this thin on evidence, "Google said so" only counts if you can see where.
The naming (April 2022)
SpamBrain was named publicly for the first time in Google's 2021 webspam report:
"SpamBrain was launched in 2018 and we've been continuously improving its performance. In 2021, SpamBrain identified nearly six times more spam sites than in 2020."
This post is not about links. The link-spam update it references is the July 2021 one, which pre-dates SpamBrain's involvement with links entirely. In April 2022, SpamBrain was a general spam system.
The link pivot (December 2022)
Eight months later, Google extended it to the link graph. The December 2022 link spam update is the single most-cited document on this subject:
"Today, we're leveraging the power of SpamBrain to neutralize the impact of unnatural links on search results."
"SpamBrain is our AI-based spam-prevention system. Besides using it to detect spam directly, it can now detect both sites buying links, and sites used for the purpose of passing outgoing links."
Two things carry the weight here. Google names both sides of the transaction — a site "used for the purpose of passing outgoing links" is in scope, so a site that sells links can see its whole outbound profile devalued. And the verb is neutralize: the credit a link passed is "lost," the link stops counting, and — Google is explicit — this is not framed as a site-level penalty.
The scaling claim (April 2023)
In its 2022 webspam report, Google attached a number:
"Thanks to SpamBrain's learning capability, we detected 50 times more link spam sites compared to the previous link spam update."
"50 times more" is the number everyone repeats. It compares against the July 2021 update, which ran before SpamBrain touched links — an AI-versus-pre-AI comparison, Google's own, with no outside audit. It is also the last SpamBrain scaling figure Google published; the annual "How we fought spam" report series ended with this edition.
The permanence (standing documentation)
The claim that turns this from a news item into a strategy problem lives in Google's maintained spam-updates documentation:
"When our systems remove the effects spammy links may have, any ranking benefit the links may have previously generated for your site is lost. Any potential ranking benefits generated by those links cannot be regained."
The Correction: Penalties Aren't Dead, They're Rare
A lot of SEO writing gets this wrong, claiming link penalties are a thing of the past. They aren't. Google still polices links; what changed is the default response — from a penalty to silent devaluation.
The proof is in Google's own Manual actions report documentation, which still lists both "Unnatural links to your site" and "Unnatural links from your site" as live categories, with the reconsideration process intact. Google still issues them — recoveries from "unnatural links" manual actions continue to be documented. And John Mueller, asked in 2024 and again in 2025 when the disavow tool is worth using, keeps naming the same single scenario: when you have a manual action for links you actually placed.
Manual link penalties still exist and are still issued — for blatant, obvious, or reported schemes. But since 2016, the standard response to a bad link is silent algorithmic devaluation, not a penalty. A manual action is now the exception a human reviewer reserves for the egregious cases; being quietly ignored is what happens to everyone else.
The Decade-Long Shift in the Default
SpamBrain didn't invent silent link handling. It inherited it. Google's default response to a manipulated link moved through three eras, and the decisive move happened years before SpamBrain touched a link.
The 2012 update — Matt Cutts's "Another step to reward high-quality sites" — announced "an important algorithm change targeted at webspam" that would "decrease rankings for sites." Worth a small honesty note: that post never used the word "Penguin." The nickname was applied afterward by the SEO press. Google shipped a webspam update; the industry named the animal.
The turn that matters came on 23 September 2016, when Gary Illyes announced "Penguin is now part of our core algorithm":
"Penguin is now more granular. Penguin now devalues spam by adjusting ranking based on spam signals, rather than affecting ranking of the whole site."
That is the real pivot. From 2016 on, link-spam enforcement went quiet and continuous. SpamBrain, in 2022, is best read as the next step in that same direction — a learned system generalizing the ignore-by-default model and pointing it at both ends of the link trade. Same instinct. Bigger engine.
What SpamBrain Is Most Likely Made Of
Here is the part almost every SpamBrain explainer skips, because it requires having done the homework. If SpamBrain is a machine-learning system that spots link buyers and sellers, then by definition it is a complex system. And complex systems have a law: the working ones are always assembled from simpler systems that already worked on their own. You don't train a model to "detect all link spam" from a blank slate. You compose it from narrow, individually-validated signal generators.
Google spent two decades building exactly those generators — and patenting them in the open. Our research library is, in effect, a catalogue of SpamBrain's plausible parts. None of this is a Google disclosure that SpamBrain uses them; it is the systems-theory argument that a working classifier would be built from precisely the components Google already had on the shelf.
Start with the strongest single piece of evidence: more than a decade ago Google patented a per-link quality model — the Reasonable Surfer model — that weights each individual link by how likely it is to actually be followed, judged from the link's own features and user behavior rather than treating every link equally. A learned model that judges link quality one link at a time is not something SpamBrain would need to invent. It's something it would inherit.
The rest sort cleanly by the signal family a spam classifier would need — and each one is a system already sitting in our research library:
| The signal a link-spam classifier needs | Where Google already built it |
|---|---|
| Anchor text & velocity How anchors are distributed, and how fast they appear. |
Historical Data — timing, spikes, rotating links · Click-Validated Quality — anchor n-gram spam |
| Link context & relatedness The editorial text around a link, and whether the anchor matches the destination's topic. |
Reference Contexts — link-context fingerprinting · Phrase-Based Indexing — anchor-to-page topic match |
| Source, resource & seed quality A page's worth as a link source vs. destination, and its distance from trusted seeds. |
Seed Distance — trust proximity · Entity Trust — the endorser's trust · Click-Validated Quality — source scoring |
| Referral traffic & click validation Whether anyone actually clicks the link. |
Click-Validated Quality — genuine long-dwell clicks · Reasonable Surfer — click-probability weighting |
| Independence & networks Whether links are genuinely separate, or one operator in disguise. |
Entity Trust — anti-collusion · Reference Contexts — context diversity · Implied Links — profile vs. brand demand |
| Time The tempo of acquisition, and domain legitimacy. |
Historical Data — velocity, freshness, DNS & registration |
SpamBrain doesn't need to catch the payment. It reads the shape of the transaction in the link graph — and every feature it would need to read that shape is a system Google already built. The most defensible picture of SpamBrain is a neural layer composing these documented signals into one probabilistic judgment, where Google once used hand-tuned thresholds. That is an inference — a well-grounded one, but an inference, not a wiring diagram.
What This Means for Buying Links
A careful point, because it's routinely mangled. None of these systems detects whether money changed hands. They estimate something else entirely: whether a link is a genuine editorial endorsement that produces real signal. That distinction is the whole game.
Run the parts list back and the pattern is obvious. A link that sits in relevant editorial context, on a page with real traffic, from an independent and trusted source, that people actually click — passes every one of those documented checks. Whether someone paid for that placement is invisible to the machine, because none of the signals measure payment. What gets neutralized is the manipulative link: irrelevant, unread, dropped into a template on a site with no independent trust. That link fails the checks whether it was bought, traded, or begged for free.
The durable link isn't the free one — it's the genuine one. A placement that would earn its keep across relatedness, source trust, real clicks, and independence has nothing artificial for SpamBrain to subtract, because everything it's made of is real. The failure mode isn't "you paid for it." The failure mode is "nobody would have linked to it on the merits."
The Skeptic's Ledger: Claimed vs. Verified
Because SpamBrain is described but never documented, precision about who is telling us what matters more than usual. Sort every common claim into one of three bins.
- SpamBrain exists and is AI/ML-based
- It neutralizes link spam at scale
- "50× more link spam"; "99% spam-free"
- Buyers and sellers both detected
- The component patents are real and granted
- Several are production-confirmed in the leak
- One SpamBrain-named leak field (
spambrainLavcScore) - Manual link actions still documented & issued
Why the skepticism isn't cynicism: Google has a track record. For years its representatives said clicks weren't used in ranking. Then the DOJ antitrust trial put a Google VP, Pandu Nayak, under oath, and NavBoost's use of click data since roughly 2005 entered the record. The public story and the sworn story diverged. Link spam never got that test — the DOJ case was about monopoly maintenance through default-placement contracts and the data-scale advantage they create, not about how the ranking algorithm polices links. SpamBrain was categorically outside the case, so the record is simply silent on it. It has never been cross-examined.
SpamBrain's existence and ML nature are Google-stated. Its scale numbers are Google's own, unaudited. The component composition in this article is our reasoned inference from patents and the leak — not a documented architecture. And independent, third-party verification of how it actually works does not exist. Believe the parts list, because Google published it. Treat the black-box narrative as a claim, because that's all it is.
The Recovery Trap
Neutralization has a quiet cruelty a penalty doesn't. A penalty is visible and, usually, reversible. Silent devaluation is neither.
- Shows up in Search Console
- Reversible — clean up, request reconsideration
- Disavow is a real part of the fix
- You know what happened
- No notice, anywhere
- Credit "cannot be regained" — Google's own words
- Disavow does nothing; there's no penalty to lift
- You often can't confirm it happened
The disavow tool is the casualty of the confusion. It's a manual-action instrument — for when you bought links, got a Search Console notice, and can't remove them. It does not reverse algorithmic devaluation, because there's nothing to clean up when Google has simply stopped counting a link. That's the trap: a business owner refreshing Search Console for a warning that never comes, quietly funding authority that stopped existing months ago. No villain to point at, no letter to appeal. Just a slow leak nobody mentioned — and understanding the mechanism is the one thing that hands that person back some control.
What Most SpamBrain Guides Get Wrong
The genre treats SpamBrain as a revealed truth — "Google's AI does X, Y, Z" — and sells you protection from it. Three errors run through nearly all of it.
It overstates the certainty: there is no patent, no paper, no testimony, so any confident account of SpamBrain's mechanism is guesswork dressed as fact. It gets the verb wrong: the everyday risk is silent devaluation, not penalty, and the two demand opposite responses — one you fix, the other you can only out-earn. And it skips the homework: it never connects SpamBrain to the documented systems that would actually compose it, so it can't tell you anything durable about how to build links that survive.
A smaller error shows how the confident-guessing spreads: a patent application, US20240020476A1 — since granted to Pinterest as US12572741B2 — gets passed around as "SpamBrain's patent." It's assigned to Pinterest, not Google. The lesson isn't that someone slipped — it's how badly we want to bolt an official-looking artifact onto a system that has none. The credible writers cite no patent, because there isn't one.
The frame that survives contact with the evidence is quieter than the hype: Google mostly stopped punishing bad links a decade ago and started ignoring them, and the AI it credits for that is real but undocumented — a black box we can only understand by the parts Google already published. Which lands on the one discipline this whole piece is built to defend. Clarity here isn't a style. It's the practical difference between knowing what you're doing and repeating what you were told.
The Link-Spam Reality Checklist
Pulled from everything above — sorted by the belief each item corrects.
Get the model right
- Assume a bad link's default fate is to be ignored, not penalized.
- Know that manual link penalties still exist — reserved for blatant, obvious, or reported schemes.
- Treat every "SpamBrain does X" claim (including Google's) as a hypothesis until it maps to a documented system.
Judge links by genuineness, not by whether they're paid
- Ask of any link: would this exist if the machine were reading it as an editorial endorsement?
- Require topical relatedness between the linking context and the destination.
- Require a real, independent, trafficked source — not a template on a trust-less site.
- Favor placements that earn genuine clicks; unread links validate nothing.
Respect the documented systems
- Keep acquisition gradual and un-patterned — velocity spikes are the oldest tell in the book.
- Diversify the editorial contexts your links live in; homogeneous contexts read as machine-made.
- Keep your link profile proportional to genuine brand demand.
Don't fall for the recovery trap
- Don't disavow to fix silent devaluation — it only addresses manual actions.
- Don't expect neutralized credit to return; Google says it can't.
- Reserve disavow for a real Search Console manual action on links you placed.
Frequently Asked Questions
What is SpamBrain, exactly?
The name Google gives its AI spam-prevention system. Google says it launched in 2018, was named publicly in April 2022, and was extended to link spam on 14 December 2022. There is no patent, no paper, and no court record for it, so everything about its internals is inference. Its existence rests on Google's own statements plus a field name or two in the 2024 API leak.
Does Google penalize you for spammy backlinks?
Rarely. Manual link penalties still exist — Search Console still lists "Unnatural links to your site" and "from your site," and Google still issues them for blatant schemes. But since Penguin 4.0 in 2016, the default response is silent algorithmic devaluation: the link is ignored, passes no value, and you get no notice. Penalty is the exception; being ignored is the rule.
Can buying links still work?
Yes. The systems don't detect payment — they estimate whether a link is a genuine editorial endorsement that produces real signal. A paid placement on a relevant, trafficked, independently-run site can pass value; a cheap, irrelevant, unread link gets ignored whether it was paid for or free. The line is quality and genuineness, not money.
How does SpamBrain actually work?
Nobody outside Google knows, and anyone who says otherwise is guessing. The most defensible answer comes from systems theory: a working complex classifier is assembled from simpler working systems, and Google spent two decades patenting exactly the ones a link-spam classifier would need — per-link click weighting, context fingerprinting, click validation, seed-distance trust, velocity, independence. SpamBrain is most plausibly the learned layer composing those. That's a reasoned inference, not a documented fact.
Is there a patent for SpamBrain?
No. There is no publicly identified Google patent for SpamBrain. The application sometimes cited as its patent — US20240020476A1, since granted to Pinterest as US12572741B2 — is assigned to Pinterest, not Google. Credible analysts cite no patent, because none exists.
Can I recover ranking lost to a link spam update?
Not the part that came from neutralized links. Google's documentation states that once its systems remove the effect of spammy links, "any potential ranking benefits generated by those links cannot be regained." You recover by earning new, genuine signals — not by scrubbing the old ones. Disavow only helps against a manual action.
Why be skeptical of Google's SpamBrain claims?
Because Google's public statements about ranking have diverged from reality before. For years it said clicks weren't a ranking factor; the DOJ antitrust trial put a Google VP under oath and the opposite came out. Link spam never faced that scrutiny — SpamBrain has no patent, no paper, and no court record. Its scale figures are self-reported and unaudited. Trust the documented parts; treat the black-box story as a claim.