The Rise of 'Hallucinated' References
Generative AI tools like ChatGPT have become popular assistants for researchers, helping to summarise complex topics and streamline the writing process. But these large language models (LLMs) come with a significant flaw: they have a tendency to "hallucinate."
In the academic world, this often manifests as fabricated citations. These are references that look entirely legitimate, complete with authors, plausible titles, and proper formatting, but point to articles or even entire journals that do not exist. This happens because AI models are designed to predict the next most likely word, not to check facts against a database. They learn the patterns of academic writing so well that they can create convincing fakes, sometimes combining a real author with a non-existent paper or a real journal with incorrect volume numbers. This issue isn't rare; some studies have found that a significant percentage of AI-generated references can be partially or completely fabricated.
Why Fake Citations Are a Serious Problem
A fabricated citation isn't just a minor error; it undermines the very foundation of scholarly research. Academic work is built upon a verifiable chain of evidence, where each claim is supported by prior studies. When that chain includes non-existent sources, it pollutes the scientific record, sends other researchers on wild goose chases for phantom papers, and can lead to flawed hypotheses being taken seriously. The consequences can be severe. Researchers who unknowingly include these hallucinated references in their work risk having their papers retracted, facing accusations of academic misconduct, and damaging their professional credibility. In fields like medicine, the stakes are even higher, as research based on faulty evidence could have real-world consequences. The problem has become significant enough that legal and academic professionals have been penalised for failing to verify AI-generated sources.
Fighting AI with AI: The New Verification Tools
In response to this growing crisis of academic integrity, a new category of software has emerged: AI-powered citation verification engines. Rather than using AI to generate content, these tools use it to police content. Scholars are increasingly turning to these platforms as a necessary final step before submitting their work. Tools with names like CiteTrue, Citely, and Scite are designed specifically to tackle the problem of hallucinated references. They function as digital detectives, cross-referencing every citation in a manuscript against vast academic databases like CrossRef, PubMed, and Google Scholar. This automated process allows researchers to quickly identify which sources are legitimate and which are phantoms created by a language model.
How the AI Verifiers Work
These verification engines employ sophisticated methods to check a citation's authenticity. It's more than just checking if a link works. The software analyses multiple data points, including the author names, article title, journal, publication year, and Digital Object Identifier (DOI). A key function is checking for mismatches that humans might miss, such as a real DOI that points to a completely different paper than the one cited. Some tools generate a "confidence score" for each citation, flagging those that are likely to be fabricated. Others, like Scite, go a step further by analysing the context of a citation, showing how a specific paper has been cited by others and whether it was to support or contrast a finding. This creates a multi-layered defence against the increasingly plausible fictions that generative AI can produce.
A New Step in the Research Workflow
The rise of these verification tools signals a fundamental shift in the academic workflow. Just as plagiarism checkers became a standard part of the writing process, citation verifiers are becoming an essential guardrail in the age of AI. While these tools are powerful, they are not a silver bullet. Experts caution that researchers still hold the final responsibility. A tool can confirm a source exists, but it cannot always determine if the source truly supports the specific claim being made in the text. This means scholars must continue to engage critically with their sources. The dynamic has been described as a new 'arms race'—as generative AI gets better at creating convincing fakes, verification AI must become smarter at detecting them. For the modern researcher, relying on AI for assistance now requires an equal reliance on another form of AI for validation.














