The Citation Problem Was Already Human

AI didn’t invent fake references. It made an old bad habit faster, cheaper, and harder to notice.

A few days ago, I opened a newsletter warning educators about AI-generated fake citations in academic work. The article claimed researchers had audited more than 100 million references and uncovered nearly 150,000 fabricated citations embedded in scientific literature. Strong claim. Serious claim.

Then I clicked the “Fake Reference Study” button.

fakecit 1

It redirected through a tracking-heavy marketing link and ultimately landed on an unrelated article about AI funding in New York City grade schools. Not the study. Not even close. Ironically, an article warning readers not to trust citations failed to properly link to its own source, though this appears more careless than malicious. (The actual study is here: https://arxiv.org/abs/2605.07723.) That doesn’t mean the underlying issue is fake. The study appears legitimate. Researchers auditing large academic repositories estimated roughly 146,932 hallucinated citations across more than 111 million references. That works out to roughly 0.132% of the references in that audit.

Over the past two years or so, headlines have increasingly framed AI as uniquely deceptive, dishonest, or prone to “lying.” The problem with this framing is that scientific integrity issues long predate large language models (LLMs). In a widely cited PLOS ONE meta-analysis from 2009, roughly 1.97% of researchers admitted to fabrication or falsification of data. Broader questionable research practices ranged from 12.5% to 33.7% depending on the definition and methodology. That isn’t a one-to-one comparison: the AI figure measures references, while the Fanelli figures measure researcher self-report. But the contrast still matters. Importantly, those numbers were self-reported, meaning the true figures may be higher. Humans were padding references, manipulating data, selectively citing allies, and laundering weak claims through reputable systems long before AI entered the picture.

AI changes the scale. A hallucinated reference can now be generated in seconds, formatted correctly, inserted plausibly into a paper, and repeated downstream by people who never verify the source material. The danger is not that AI invented citation fraud. The danger is that it has industrialized it. At the same time, AI is exposing existing weaknesses in human systems. Researchers, writers, journalists, and students have always relied heavily on authority signals: publication names, formatting, institutional branding, database reputation, citation count, and social trust. AI didn’t create that habit – it exploits it. And that’s the line people keep missing. If society frames this entirely as an “AI problem,” then people reach for panic, prohibition, or moral outrage aimed at the tool. The real issue is that too many people do not verify enough. Weak verification existed before AI and will exist after it. The answer is not blind trust in AI. Nor is it reflexive fear of it.

The answer is going back to the basics that should have existed all along: verify the source, trace the citation, read beyond the headline, and do not outsource critical thinking to either machines or humans.

References

Zhao et al., 2026. “LLM Hallucinations in the Wild: Large-Scale Evidence from Non-Existent Citations.” arXiv. https://arxiv.org/abs/2605.07723

Fanelli, Daniele. 2009. “How Many Scientists Fabricate and Falsify Research? A Systematic Review and Meta-Analysis of Survey Data.” PLOS ONE. https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0005738

Consulting: Need independent analysis or security support? See AI & Cybersecurity Consulting.

Scroll to Top