A new preprint from Samsung and the University of Warsaw has identified a disturbing contamination of academic publishing: hundreds of papers, books, and articles are co-authored by AI-generated 'ghost' personas—names like Elena Vasquez and Marcus Chen that do not correspond to real people but are consistently produced by large language models (LLMs) when asked to generate expert names. The phenomenon, detailed in a paper on arXiv, highlights how AI models' tendency to default to the same fictitious identities is polluting the scholarly record, with one of these ghost names even appearing as the founder of a company found to be entirely AI-generated.
Researchers have documented a systematic flaw in large language models that is undermining the integrity of academic publishing. In a preprint titled 'The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing,' scientists from Samsung and the University of Warsaw show that certain names—most notably Elena Vasquez and Marcus Chen—are repeatedly generated by AI models as experts in fields ranging from volcanology to medicine. These personas have never existed, yet they appear as co-authors in hundreds of independently produced AI-generated documents, including academic papers, books, and even podcast transcripts.
The study builds on earlier observations: Last month, 404 Media reported that chatbots like ChatGPT, Gemini, and Claude frequently generate the name 'Elias Thorne' for fictional lighthouse keepers. Similarly, users noted that ChatGPT often names a software developer 'Marcus Chen.' The new research reveals that these names form 'correlated character ensembles'—meaning certain names are more likely to appear together, creating the illusion of a consistent set of fictional experts. Other recurring names include Elena Amara Okafor (Claude), Aris Thorne and Lena Petrova (Gemini), and Elara Voss (ChatGPT).
The contamination is not limited to academic papers. The lead author, Michał Brzozowskim, told 404 Media that searching for these names on Google turned up AI-generated content about real-world events, including a false rumor spread on Facebook after the killing of Alex Pretti by U.S. Border Patrol agents. The rumor, which Snopes debunked, was attributed to an AI-generated persona. In a separate case, 404 Media revealed that a company called Research Gold, which claimed its medical research was '100% human written, never AI,' was entirely AI-generated. The company's founder and lead methodologist was listed as Elena Vasquez—exactly one of the ghost names. Research Gold removed the name after the story was published.
The implications for academia are serious. Peer review relies on the assumption that authors are real researchers with verifiable credentials. The presence of ghost authors could erode trust in published findings, waste reviewer time, and potentially influence policy or funding decisions based on fabricated research. The researchers urge journals and preprint servers to implement better screening for AI-generated content and to be aware of these name priors.
Analysis
Why This Matters
- Erosion of trust: If readers cannot distinguish real researchers from AI-generated personas, confidence in academic literature collapses, affecting everything from medical guidelines to climate policy.
- Waste of resources: Peer reviewers and editors are already overburdened; ghost papers consume time and delay legitimate science.
- Real-world harm: AI-generated research could be cited in policy documents or clinical trials, leading to dangerous decisions based on fabricated data.
Background
AI-generated content has plagued academic publishing for years, with 'paper mills' selling fake manuscripts for profit. But the advent of generative LLMs has made producing convincing fake papers far easier and cheaper. This new research reveals an unexpected vector: the models' tendency to reuse the same fictional names, creating a detectable signature. The phenomenon was first noticed in creative writing (the lighthouse keeper 'Elias Thorne') but turns out to contaminate expert databases, academic profiles, and even published research. The paper from Samsung and the University of Warsaw is the first systematic study of this 'name prior' problem, showing that different LLMs have distinct favorite name sets. The problem is compounded by the fact that many academic papers are now submitted in a 'publish or perish' environment where quantity often trumps quality, making ghost-authored papers more likely to slip through.
Key Perspectives
Academic publishers and journals: Institutions like Elsevier and arXiv are racing to deploy AI detection tools, but these often flag legitimate work. Publishers want a scalable solution that does not burden reviewers but have been slow to act. The ghost names provide a simple filter—any paper listing Elena Vasquez or Marcus Chen as an author should be scanned for AI generation.
Researchers and scientists: Legitimate scientists worry that the crisis will lead to blanket skepticism of all research, making it harder to publish genuine work or secure funding. Some also fear that their own names could be harvested and used in fake papers without consent.
AI developers (OpenAI, Google, Anthropic, etc.): These companies are aware of the character ensemble problem but have not publicly announced fixes. They face pressure to reduce name priors, but doing so may introduce new biases or reduce utility. Critics argue that the onus should be on AI companies to prevent their models from generating content that can be passed off as human expert work.
Critics and skeptics: Some argue that the scale of the problem is overstated—most journals still catch obvious AI-generated submissions. Others note that the ghost names are a symptom, not the cause, and that the real issue is the incentive system in academia that rewards publication volume over quality. Additionally, there are ethical questions about whether papers with ghost authors should be retracted or merely corrected.
What to Watch
- Adoption of name-based filters: Whether publishers and preprint servers automatically screen for known LLM name priors and flag suspicious submissions.
- Response from AI companies: Updates from OpenAI, Google, and Anthropic on model adjustments to reduce correlated name generation.
- Retraction or correction notices: How many previously published papers are found to contain ghost authors will reveal the true scope of contamination.
- Regulatory or policy changes: Possible moves by research councils or governments to mandate AI disclosure in academic submissions.