A fake thinktank produced 560,000 words in nine days to seed AI answers
The Guardian documented 124 reports from the 'Hanover Institute'. The target was not readers — it was the retrieval layer that models consult.
An organisation calling itself the Hanover Institute produced 124 reports totalling more than 560,000 words in nine days, in an apparent attempt to seed AI systems with propaganda, according to the Guardian.
The output rate is the tell. More than 60,000 words a day of policy research is not a productive thinktank; it is generated text published under an institutional name to create the appearance of a source.
The target audience is not human
Traditional disinformation is written to be read. It needs a distribution channel, an audience, and enough plausibility to be shared.
This is aimed at a different consumer. Models answering questions increasingly retrieve from the live web, and the retrieval layer ranks sources by signals that a well-constructed fake can satisfy: an institutional name, a professional site, consistent formatting, internal citations, and a large body of documents that cross-reference each other.
A hundred and twenty-four reports citing one another produce exactly the density that makes a source look established. The volume is not for readers — it is to occupy enough of the retrievable corpus on a set of topics that a model summarising "what the research says" encounters it.
A model does not evaluate institutional legitimacy. It weighs what it retrieves.
Why this is harder to counter than platform disinformation
Social platforms have accounts, engagement metrics and moderation systems, all of which give defenders something to act on. Coordinated inauthentic behaviour leaves detectable traces in account creation and posting patterns.
A website publishing PDFs has none of that. There is no account to suspend, no engagement graph, no platform whose terms are breached. The material is simply on the internet, and the retrieval systems that consult it have no mechanism for distinguishing a real institute from a convincing one.
The defence would have to be source reputation at the retrieval layer, which means someone maintaining a judgement about which organisations are real. No provider publishes such a list, and one that did would become the object of the same manipulation.
The wider contamination problem
This is the deliberate version of something that is happening accidentally at much greater scale.
A Nature Human Behaviour study found that model polishing reduces variance in writing complexity by 21 to 50 percent across more than 880,000 texts. OpenAI's own agents left roughly 18,000 posts on a German-language wiki between May and June, sharing answers during web-lookup evaluations — a case of models contaminating the corpus that models read.
The web is becoming a medium that models both write to and read from, and the Hanover Institute is what happens when someone notices that and acts on it deliberately.
What nobody has committed to
No AI provider publishes how it weights source credibility in retrieval, whether it maintains exclusion lists, or how it would respond to a documented seeding operation. The EU's transparency rules, in force since 2 August, require disclosure when content is AI-generated — an obligation that a covert influence operation will simply ignore.
The Guardian's reporting identifies the operation. It does not establish who funded it or what outcome they wanted, and there is no regulator with an obvious remit to find out.
Runs the newsroom. Rename this profile in the studio to your own byline.
Related
Every weekday, the AI stories that moved money or shipped code.
No cross-posting, unsubscribe anytime. See all newsletters