Cybersecurity & Tech

Invisible Source Skew

Kate Klonick, Renée DiResta
Friday, August 28, 2026, 3:10 PM
Our information flows aren’t just threatened by slop or foreign actors. They're being reshaped by invisible decisions from AI companies.
Graphic showing lit up brain
Machine Learning & Artificial Intelligence (Mike MacKenzie, https://www.flickr.com/photos/mikemacmarketing/42271822770, CC BY 2.0, https://creativecommons.org/licenses/by/2.0/)

Plenty has been written about foreign actors manipulating information ecosystems with artificial intelligence (AI), and about the proliferation of AI-generated slop polluting feeds and search results. But those stories focus on what content is being added to the information ecosystem. Far less attention is paid to a different set of forces that shapes what we see: what information AI systems choose to retrieve, which sources are reachable and current enough to be useful, and what data is increasingly available only to particular companies.

AI answer engines assemble their picture of the world through at least two distinct information pipelines. One is training: the enormous corpus a model learns before deployment. The other is retrieval: the current information a system searches for and incorporates while generating an answer. Retrieval-augmented generation (RAG) was developed in part to solve the problem of models hallucinating or not understanding events following their last training cutoff; RAG lets systems consult outside sources while generating an answer.

What a model learned from, therefore, is not necessarily the same as what it looks up. Paywalls, crawler restrictions, commercial arrangements, and product choices shape AI-generated responses. Both training and retrieval are shaped by consequential forces largely invisible to the public, which increasingly relies on the answers.

We call this phenomenon invisible source skew: the gradual shaping of an information environment by routine product decisions made at AI companies about which sources are retrieved, which are maintained, and even which datasets are acquired. Any one of those decisions may make technical sense, be commercially rational, or even be entirely unremarkable. But together, they create a public epistemic infrastructure whose inputs can change dramatically without users—or often anyone outside the companies—being able to see, understand, or contest those changes.

Invisible source skew happens in at least three ways. First, at the retrieval layer, AI systems can silently and abruptly stop drawing on sources they once relied on heavily, with no announcement or explanation. This appears to have happened more than once with Reddit and ChatGPT.

Second, at the source layer, sites that have become foundational reference material for downstream systems can quietly stop being updated for accuracy or timeliness while continuing to appear authoritative. Grokipedia, constructed by an AI itself, appears to have done exactly that when its editorial pipeline froze in April.

And third, at the training-data layer, AI companies are acquiring exclusive, idiosyncratic datasets: nonpublic views of the world that may exert outsized, unrepresentative influence on what models learn. Google’s bid, in bankruptcy court, for the corporate remains of Spirit Airlines’ data offers a striking example.

None of these stories are about malicious actors poisoning the information environment. But that’s precisely why they matter.

How Source Retrieval Shapes the AI Answers

Let’s start with retrieval and Reddit’s precipitous decline as a cited source by ChatGPT.

Reddit is enormously valuable for large language models. In an internet increasingly saturated with synthetic text, it is one of the last large, open repositories of mostly human conversation: people troubleshooting their dishwashers, comparing chemotherapy experiences, arguing about zoning, asking if they’re the asshole. That kind of organic, authentic content has made it one of the most heavily cited domains in AI-generated answers. As a testament to its value, Reddit struck data licensing deals with both OpenAI and Google, each reportedly in the neighborhood of $60 to $70 million.

Then, on Aug. 18, AI engine optimization and analytics company Promptwatch reported that Reddit’s citation footprint largely vanished from ChatGPT in the space of six days. According to its research, Reddit held a steady average of 3.83 percent of ChatGPT Search citations—among the largest shares of any domain—from July 18 through Aug. 7. On Aug. 8, the same day ChatGPT changed aspects of how it fans out search queries to retrieve sources, Reddit’s share began sliding. On Aug. 14, it collapsed to under 1 percent, where it has stayed; the Aug. 14–17 average of 0.52 percent represents an 86 percent relative drop.

No one outside OpenAI knows for sure why this happened.

This is not the first abrupt shift in Reddit citations. In September 2025, Reddit’s share of ChatGPT citations cratered from roughly 10 percent to around 2 percent. This shift wiped more than 10 percent off Reddit’s stock as investors tried to interpret what it meant. Without any public explanation for the shift, outside analysts attempted to reconstruct what happened. One widely discussed explanation focused on an obscure Google Search change. Around Sept. 11, 2025, Google stopped honoring the “num=100” parameter, which allowed users and automated tools to retrieve 100 results at once. The change, which limited providers to seeing only the top 10 or 20 results, disrupted rank-tracking and search-data providers, and its timing roughly coincided with the citation decline.

But other analysis complicated that explanation. Semrush suggested that the parameter change alone probably could not account for the size of Reddit’s decline, and suggested that ChatGPT may have also deliberately rebalanced sources it had previously cited heavily.

In other words, an obscure change by one company became a leading explanation for why a second company’s chatbot largely stopped citing a third company’s content—and it took weeks of outside detective work to even identify that possibility.

Users and the public who don’t follow AI product news, or read Answer Engine Optimization boards, are largely unaware that a source their answers depended on yesterday has become dramatically less influential overnight. Brands and creators actively work to ensure RAG picks up their content, even as there is little visibility into why particular sources rise or fall.

A single unannounced product decision at one company can sharply reduce the role of one of the internet’s largest archives of human conversation in the answers that hundreds of millions of people receive, with almost no way for those users to know it happened. And there’s really not much they can do about it besides leave the answer engine and visit Reddit directly themselves.

Out-of-Date Sources Can Keep Their Authority

A second form of invisible source skew occurs when a source remains available but quietly falls into disrepair or stasis.

Grokipedia, the AI-generated encyclopedia xAI launched in October 2025 as Elon Musk’s answer to Wikipedia, is a key case study here. As one of us reported in Lawfare earlier this month, Grokipedia appears not to have updated a single article since April 24. Its model-driven editorial pipeline—which had once processed suggested edits with a median decision time of roughly three minutes—simply stopped, with no announcements. An analysis of more than 34,000 pages with over 225,000 suggested edits found no accepted or rejected corrections in more than three months. More than 13,000 user-submitted suggestions now sit trapped “in review,” in a queue that apparently neither human nor AI is monitoring. In Grokipedia’s universe, SpaceX still has not yet had its initial public offering, and Vincent Pastore is still alive.

If Grokipedia were a hobby project or isolated site, this would be a curiosity. It isn’t. Independent analysis estimated that Grokipedia URLs surfaced in roughly 356,000 citations across AI systems—most often ChatGPT and Google’s AI Mode. A frozen encyclopedia can therefore remain upstream of other AI products, supplying apparently authoritative information long after its own updating mechanism has stopped working.

The comparison with Wikipedia is instructive. Wikipedia is also used heavily in model training and retrieval for much the same reason as Reddit: It’s human-authored. Its editorial process is radically legible: every edit, reversion, and talk-page fight is public, and even its failures or biases are therefore contestable. While AI-created information products can be very useful, Grokipedia demonstrates what happens when a product stops functioning as a live reference while continuing to occupy that space—and how downstream systems treat it as if nothing has changed.

The danger here is not simply that an encyclopedia can contain outdated information. Encyclopedias have always contained errors. The important distinction is that the state of the source itself has become difficult to perceive. Users of a downstream AI system may not know that an answer relies on Grokipedia at all, much less that Grokipedia's mechanisms for correcting itself stopped functioning months ago.

When Training Data Reflects Only What Companies Can Acquire

Finally, we come to the question of private, narrow data—and its influence on training. Here, the private source is not deteriorating, but its contents are both potentially exclusive and unrepresentative: an unusual slice of the world that becomes available to a single AI developer through channels having little to do with the design of representative training corpora.

On Aug. 14, Spirit Airlines designated Google the winner of a bankruptcy auction in the Southern District of New York for a trove of the defunct airline’s internal data. Google’s $10 million bid, which the company says it made to help improve its products and AI models, covered roughly 100 million emails, 500 million Teams messages, 7.2 billion competitor-flight pricing records, 7.5 billion passenger transaction records dating to 2008, 30 million lines of code, and more than 175,000 employee records—plus data on marketing, human resources, strategy, audits, and fraud. Passenger profiles, loyalty-program records, privileged material, and personally identifying information are supposed to be excluded or removed; Google has promised to scrub the dataset of all personally identifiable information before it is transferred.

The proposed sale is not yet final (Spirit’s flight-attendant union objected on employee privacy grounds). But whatever ultimately happens to this particular sale, the auction illustrates a new channel by which unique bodies of real-world information can become proprietary AI assets.

The complete internal record of a real, failed American company—its competition strategy, email history of all employees, operations decisions—is a genuinely rare artifact. Economists, regulators, or historians might have studied it. Instead, Spirit’s records may become a proprietary training asset for exactly one AI developer, acquired for the price of a large Super Bowl ad, through a bankruptcy proceeding whose purpose is maximizing creditor recovery—not considering the downstream effects of concentrating a unique informational resource.

The deeper issue is selection. AI developers have strong incentives to acquire high-quality, nonsynthetic, authentic data. But an interesting asymmetry exists in what is available to them: No functioning companies will hand over 100 million internal emails to external AI firms, but a dead one, sold for parts, certainly will. That means the enterprise data that does reach training pipelines arrives through a survivorship filter in reverse: Models may disproportionately learn “how businesses work” from the internal records of businesses that collapsed, and from whichever idiosyncratic corporate cultures happened to hit the auction block.

The claim is not that Spirit’s emails will somehow teach a model that all companies operate like Spirit. We do not know how Google—or any eventual buyer—would incorporate or weigh such material. The broader point is that bankruptcy proceedings, acquisitions, licensing negotiations, and distressed-asset sales are becoming channels to assemble proprietary AI corpora. None of the institutions involved in those channels are designed to ask whether the resulting corpus is representative, whether access to it should be exclusive, or what it means for one company to privately absorb a piece of social or economic history unavailable to everyone else. Bankruptcy courts, it turns out, are now a channel of AI data policy, but nobody is thinking of them that way.

Most end users will remain largely unaware of this, too. There is no ingredient label, no way to ask what a model’s picture of the world is built from, or how much a particular corpus mattered. But consequential choices about the construction of public knowledge are now occurring in unexpected places.

The Sum of Small Decisions

None of these three episodes involves a malign actor intentionally polluting the information ecosystem. There’s no troll farm, no coordinated inauthentic behavior, no deepfake. There is a tweak to retrieval engineering, a seemingly abandoned AI reference product, and a routine asset sale. But that mundanity is exactly what makes them important.

The prevailing frames for AI’s threat to the information ecosystem assume the underlying problem is bad content flowing in. But an information ecosystem can also be transformed by ordinary decisions about selection: what gets retrieved, what gets maintained, and what gets acquired.

Each decision can be individually rational: Google changes a search parameter, OpenAI rebalances the sources its chatbot retrieves. Together, they are reshaping a public epistemic infrastructure that nobody outside the companies can meaningfully inspect or contest.

The old information ecosystem had its gatekeepers too: Newspapers decided what to print, libraries what they held, and encyclopedias what qualified for an entry. These institutions weren’t neutral either. But these choices were comparatively visible. A reader could see the source and, at least in principle, identify the institution making the judgment. AI-generated answers increasingly obscure both.

The striking thing about the current system is how much informational power is concentrated, and how seemingly small, opaque decisions can have significant downstream effects. A source can disappear from an answer engine, a reference work can stop updating, or a strange new corpus can enter a model’s training data without users ever knowing anything happened. Increasingly, that is how the information environment changes: not through dramatic acts of manipulation, but through proprietary decisions almost no one on the outside can see.


Kate Klonick is an Associate Professor at St. John’s University Law School, a fellow at the Brookings Institution, Yale Law School’s Information Society Project, Harvard Berkman Klein Center and a Distinguished Scholar at the Institute for Humane Studies. Her writing on online speech, freedom of expression, and private internet platform governance has appeared in the Harvard Law Review, Yale Law Journal, The New Yorker, the New York Times, The Atlantic, the Washington Post and numerous other publications. For the 2023-2024 academic year, she was a Fulbright Schuman Innovation Scholar in the European Union where she was a Visiting Professor at SciencesPo and University of Amsterdam researching and writing about the Digital Services Act and Digital Markets Act.
Renée DiResta is an Associate Research Professor at the McCourt School of Public Policy at Georgetown. She is a contributing editor at Lawfare.
}

Subscribe to Lawfare