# Google Leak Site Authority: What the Documents Prove

> Google leak site authority, checked against the documents: siteAuthority is a real field, but no formula, weight or live status is confirmed. Here's the proof.

- URL: https://missiongrowth.io/blog/google-leak-site-authority
- Published: 2026-07-12 · Updated: 2026-09-24
- Author: Ömer Furkan Aktaş, Founder, Mission Growth
- Publisher: Mission Growth. Company facts: https://missiongrowth.io/llms.txt

Google's leaked Content Warehouse API documentation shows one real, named, stored field called `siteAuthority`. Google's own engineer used the word "authority" to describe the score it feeds, on the record, in a February 2025 DOJ call. That's the extent of what's confirmed: the field exists, it's stored, and Google ties it to the whole site rather than to any single query.

Nothing in the documents or the testimony discloses siteAuthority's formula, its weight, or whether it's live in today's ranking system. The Google leak site authority story usually skips past that gap: the field exists, so a Domain Authority-style score must be real, case closed.

That's not the same claim. A field existing and a field deciding rankings are two different things, and they get collapsed into one more often than they should.

In this guide:
- What the leaked field's own text says about siteAuthority, quoted directly
- The other site-level fields next to it, and the per-page field two write-ups stretch into a site signal
- The DOJ exhibit where a Google engineer calls the underlying score a "notion of authority"
- A three-question test for the next leaked-field claim you see

## What is the leaked Content Warehouse API documentation?

The Content Warehouse API leak is a set of internal Google documentation, more than 2,500 pages describing Google's search API, that reached the public in late May 2024.

It circulated among a handful of SEO practitioners before going public. This piece calls it the Content Warehouse leak or the Google API leak, the two names most SEOs use for the same documents.

It still matters in 2026 because it remains the only primary view of the field names Google stores, and claims built on it keep circulating. Reading a field's own text is still the fastest way to check one.

News of the Google Search Algorithm leak spread fast. The Verge broke the story on May 28, 2024: Rand Fishkin says a source shared 2,500 pages of documents with him.

Google confirmed the documents' authenticity the next day. That confirmation is narrower than most retellings treat it. Spokesperson Davis Thompson told The Verge: "We would caution against making inaccurate assumptions about Search based on out-of-context, outdated, or incomplete information." He declined to comment on specific fields.

Two points on the record, nine months apart, confirm the same picture from different directions:

- May 28, 2024: The Verge breaks the story; Rand Fishkin says a source shared 2,500 pages with him.
- The next day: Google confirms the documents are authentic and cautions against out-of-context conclusions.
- February 18, 2025: In a DOJ call, Google engineer Hyung-Jin Kim ties the site-level quality score to the word "authority," on the record.

::figure{src="/blog/figures/google-leak-site-authority-2.svg" alt="A Google leak site authority timeline: The Verge’s May 2024 report, Google’s confirmation, and February 2025 DOJ testimony naming PageRank as a Quality input." caption="Nine months apart, a journalist’s report and a DOJ court call independently put Google’s own site-level quality signals on the record." width="720" height="184"}

The leak also documents a separate family of click-based signals, the ones NavBoost draws on to rank pages inside a fixed time window. That's its own mechanism with its own fields; see [google navboost](https://missiongrowth.io/blog/navboost) for how it works.

## What does the leak actually say about siteAuthority?

siteAuthority is an integer field inside Google's `CompressedQualitySignals` module, a subproto of `PerDocData`. Its own doc comment reads: "converted from quality_nsr.SiteAuthority, applied in Qstar."

That Q* is the same site-level quality system Kim describes in the DOJ call below. The leak's own text writes it "Qstar"; this piece uses the DOJ exhibit's "Q*" everywhere except inside that exact quotation.

The field's own text confirms exactly three things: it exists, it's an integer, and it feeds into Q*. It doesn't say how the number gets calculated, what other inputs shape it, or how much it counts once it's in the system, and no other part of the documentation fills that gap either.

That's a narrower claim than "Google has a Domain Authority score," which is the version most retellings default to.

The module's other compressed fields state their own encoding range: "we convert the float values in [0, 1] to integers in range [0, 1023] (use 10 bits)." That line belongs to the neighboring QualityBoost-derived entries; siteAuthority's own entry states no range at all. Treat the 0-1023 figure as module-level context, and leave it out of any claim about siteAuthority's own documented range.

## What else does the leak show about site-level signals?

The leak documents several other site-level fields alongside siteAuthority, including `authorityPromotion` and `unauthoritativeScore`. Here's the rest of what sits in that module, plus one per-page field some write-ups promote to a site-level signal:

- **authorityPromotion.** Documented as "authority promotion: converted from QualityBoost.authority.boost." A separate field, a separate conversion source, same module.
- **unauthoritativeScore.** Documented as "Unauthoritative score. Used as one of the web page quality qstar signals." The text stops there. It doesn't say whether this field works as a penalty distinct from siteAuthority's own role, and a field's name doesn't hand you that answer either; reading a penalty function into "unauthoritative" because of the word alone is the kind of inference the documents don't support.
- **Links and anchor text.** The leak's PageRank-derived signals sit in a different bucket from siteAuthority. The DOJ trial exhibit quoted below names PageRank as an input to the Quality score and never mentions siteAuthority.
- **Click behavior.** Tracked by a separate system, already covered above.
- **Content quality.** A separate per-document field, `contentEffort`, covers this territory on its own. See [contentEffort field](https://missiongrowth.io/blog/content-effort) for how Google's raters and vendors treat it.

Two of the pages that cover this topic, trapilot.ai and szymonslowik.com, treat OriginalContentScore as a broad originality signal, and one of them rolls it up into how the whole site reads. The field is real, but it lives in the per-document `PerDocData` module, and its own description narrows it sharply: "Only pages with little content have this field."

Nothing in that description says the field measures originality across a site, feeds siteAuthority or carries a weight. It's a page-level value recorded for thin pages, and that's all the documentation states.

szymonslowik.com is the more careful of the two on nearly every other point; it's the page that draws the existence-versus-weight distinction this piece rests on. That's the clearest argument for reading a field's own description yourself rather than trusting how a careful source summarizes it.

## Google leak site authority vs. Domain Authority: is it the same score?

No. siteAuthority is Google's own internal field, built from Google's own quality model. Moz's Domain Authority and Ahrefs' Domain Rating are separate, third-party systems built from each vendor's own link index.

That also answers how to tell whether a website has authority: the only scores you can actually look up are those third-party estimates. Google's own value isn't exposed anywhere, so treat DA or DR as a proxy for link strength, never as a reading of siteAuthority.

Is Domain Authority a Google ranking factor, then? Not by Google's own account, and Google has said so twice, five years apart, in two different contexts.

In June 2019, in a Google Webmaster Central hangout, John Mueller said: "In general, Google doesn't evaluate a site's authority. So it's not something where we would give you a score on authority..." Google's own recording of that June 2019 hangout is still on its Search Central YouTube channel; Search Engine Journal's contemporaneous report is a secondary account built from it.

Then, in May 2024, after the leak, Google gave a second statement, cautioning against out-of-context assumptions and declining to comment on specific fields. Those are two different statements: a flat denial in 2019, and a careful non-answer in 2024.

Reading the second one as reversing the first, as though the leak flipped a denial into a confession, is where most coverage of this claim goes wrong. Four of six pages that describe siteAuthority state or imply the leak confirms or contradicts Domain Authority outright; two correctly hold the claim to existence alone.

Domain Authority and Domain Rating count backlinks and estimate a score from crawl data neither vendor gets from Google. siteAuthority counts something else: an internal number converted from Google's own quality_nsr system. The leak gives no reason to treat the two as interchangeable.

::figure{src="/blog/figures/google-leak-site-authority-3.svg" alt="siteAuthority is Google’s own field feeding Q*, while Domain Authority and Domain Rating are third-party scores Moz and Ahrefs build from backlink data." caption="Google’s siteAuthority converts the company’s own quality data; Domain Authority and Domain Rating are estimated from backlink data Google never shares." width="720" height="339"}

## What does Google's own engineer say connects PageRank, quality and authority?

In a February 2025 DOJ call, Google engineer Hyung-Jin Kim described the Quality signal as "largely static and largely related to the site rather than the query," fed in part by PageRank.

He added that if competitors saw the ranking logs, "they have a notion of 'authority' for a given site."

The exchange runs across two pages of the same exhibit, PXR0356. On page 3, under a bullet labeled "Quality," Kim opens with: "Generally static across multiple queries and not connected to a specific query."

A few lines later, closing the same paragraph, he adds: "Q is largely static and largely related to the site rather than the query." Two bullets down, under "PageRank," he states plainly: "This is a single signal relating to distance from a known good source, and it is used as an input to the Quality score."

The "notion of authority" line sits earlier on the same page, opening a paragraph that begins "Q* (page quality (i.e., the notion of trustworthiness)) is incredibly important." Kim's full sentence: "If competitors see the logs, then they have a notion of 'authority' for a given site."

That's the primary confirmation this piece exists to add: not a blog's paraphrase of "Google has authority," but Google's own engineer, under DOJ questioning, using the word for a site-level score that PageRank feeds into.

::figure{src="/blog/figures/google-leak-site-authority-4.svg" alt="PageRank feeds Q*, the site-level quality score Google engineer Hyung-Jin Kim calls a notion of authority under DOJ questioning." caption="PageRank feeds Q*, the same static, site-level score Google’s own engineer described under DOJ questioning as a notion of authority." width="720" height="216"}

Kim also names the boundary of what any of this proves, on pages 4 and 5 of the same exhibit. Opening section VII: "There was a leak of Google documents which named certain components of Google's ranking system, but the documents don't go into specifics of the curves and thresholds."

His closing line: "The documents alone do not give you enough details to figure it out, but the data likely does." Components named, weights withheld, straight from the person who'd know.

## How do you tell if a leaked field is really a ranking factor?

Ask three questions before treating any leaked field as a ranking factor: is it named with its own description, does that description state a formula or weight, and is its current, live use corroborated outside the leak itself.

Run the test against every field named above:

- **siteAuthority.** Passes question one outright: it has a doc comment naming what it converts from and what it feeds. Partially passes question three, because Kim's DOJ testimony corroborates a static, site-level quality signal he himself calls a "notion of authority." Fails question two by design: Kim says directly that the documents "don't go into specifics of the curves and thresholds," so no field in this leak discloses a formula or a weight.
- **authorityPromotion and unauthoritativeScore.** Pass question one the same way siteAuthority does. Fail question two the same way. Neither gets the outside corroboration siteAuthority has from the DOJ exhibit.
- **OriginalContentScore.** Passes question one: it's named with a description. Fails question two, and its description sets a scope instead of a weight: only pages with little content carry it. Nothing outside the leak corroborates its live use.

::figure{src="/blog/figures/google-leak-site-authority-1.svg" alt="siteAuthority, authorityPromotion and OriginalContentScore each pass the naming test, but no leaked field states a formula or a weight." caption="Every field here is named and described; none discloses a formula, and only siteAuthority has outside corroboration of live use." width="720" height="325"}

Any Google leak ranking factors claim you run into from here forward gets the same three questions, not a verdict borrowed from whoever posted it first.

For the practical side, the tactics that actually move rankings today live in [google ranking factors](https://missiongrowth.io/blog/how-to-rank-higher-on-google); this piece sticks to what the leak proves instead of a generic authority-building checklist.

Q* itself is an emergent score built from many granular attributes rather than a single field. The same pattern shows up in [e-e-a-t ai content](https://missiongrowth.io/blog/is-ai-content-bad-for-seo) rather than one single eeat_score line.

Treat a leaked-field claim the way you'd treat a claim inside a [technical seo audit](https://missiongrowth.io/blog/how-to-do-seo-audit): verify it against the primary source before you act on it.

Google's leaked documents prove one thing about site authority: siteAuthority is real, named and stored, and it fits into Q* the way Kim describes it under DOJ questioning. They don't prove a formula, a weight or a live status, and neither does anyone summarizing them for you.

Next time a claim like this crosses your feed, open the primary document yourself and run it through the three questions above before you repeat it.

## FAQ

### What does Google's own engineer say the leaked documents can't tell you?
They name components, like PageRank feeding into Quality, but don't disclose the curves, thresholds or weights that decide how much each one counts. Kim states this directly in the same DOJ call that names PageRank as an input to Q*.

### Is siteAuthority the same as Moz's Domain Authority?
No. They're different systems built from different data. Google's siteAuthority is documented as its own field inside CompressedQualitySignals; Domain Authority is calculated from Moz's own link index, and Google has twice said it doesn't hand out an authority score.

### Did Google confirm the leaked documents are real?
Yes. A Google spokesperson confirmed their authenticity the day after The Verge's report, while cautioning against out-of-context conclusions and declining to comment on specific fields.

### Is "OriginalContentScore" a real field in the Google leak?
Yes, but it's narrower than most write-ups suggest. It sits in the per-document PerDocData module, and its own description says only pages with little content carry it; nothing in the leak ties it to siteAuthority or to site-level quality.

### Can you check your own site's siteAuthority score?
No public tool reads Google's internal systems. The field is known only from the 2024 documentation leak itself, and nothing published since has given outside access to its value.
