A Scene from the AI Flooding of Academic Journals
A Scene from the AI Flooding of Academic Journals
âWe recently received a double-digit number of submissions from authors in the Far East⌠where we had strong evidence of machine generated content. Nothing made much sense in these papers, the authors, with their yahoo and aol type email addresses, frequently supplied false university affiliations, among other glaring problems. I suspected that these submissions were a test of our detection systems more so than anything else.â
That is Udo Schuklenk, professor of philosophy at Queenâs University, writing in his capacity as one of the editors of the journal, Bioethics.
His editorial was prompted by his journalâs publication of a critique of an article in the Journal of Medical Ethics (JME) for, among other things, multiple citations to âbibliographic fictionsâ (seemingly the result of unacknowledged AI use), which led to the latter articleâs retraction.
The now-retracted JME article is âPanem, corticoids and circenses: the ethical fallout of Enhanced Games,â by Alexis Demas (who, Schuklenk notes, used a yahoo email address and whose stated institutional affiliation could not be verified).
The critique of the JME article, published (open access) in Bioethics, is by Ognjen ArandjeloviÄ (University of St. Andrews), and is titled âAgainst Moral Panic and Citation Fiction: A Critique of âPanem, Corticoids and Circensesâ and a Proposal for Editorial Gatekeeping on Reference Integrityâ.
ArandjeloviÄ discusses the reference failures and what to do about them in section 4 of his article, and I recommend people look at it. Hereâs the introduction to that section:
It is tempting to treat reference failures as a mechanistic problem that could be addressed through technology and bureaucracy by installing more automated checks, tightening and enforcing checklists, reminding authors of their obligations. Not only banal, that response misdiagnoses the failure mode. In a publication culture where fast-turnaround commentaries are rewarded, reviewers often treat references as decorative rather than load-bearing, and journals rarely suffer meaningful consequences for bibliographic unreliability, low-grade fabrication, and high-grade sloppiness become rational strategies.
The core issue, then, is not attempting to help authors avoid mistakes. Authors know this. Rather, the challenge is in making it costly (in probability and consequence) to submit work with citations that do not exist or do not support what they are cited for. In other words, what is needed is deterrence, not assistance; accountability, not coaching.
Schuklenk describes how the case looks from the point of view of the editor:
Given the number of made-up references in the Journal of Medical Ethics paper, it is reasonable to assume that these references were the result of an AI hallucinating, a known phenomenon of AI-written content that adds non-existent references to the content the AI generated, and it is not unreasonable to suggest that the author of this manuscript was careless enough to leave a plethora of fake references in the manuscript. This doesnât happen by accident. During the proof corrections stage of the article the author would have had ample time to correct any errors but chose not to. Your guess about how much else of the manuscript was generated using AI is as good as mine. We donât currently have detection tools available to assist us in making reliable determinations. AI use also wasnât disclosed by this author, if this is what is (very likely) at the heart of the fake references. Iâm pleased to say that, apparently unlike the BMJ group of journals, Wiley, the publisher of this journal, has in place a highly sophisticated automated reference check that is available to the editorial team. This manuscript would not have gone out for peer review if it had been submitted to Bioethics, because it would have been eliminated after the reference (and possibly the AI generated content) screening. Surprisingly, the BMJ group of journals doesnât seem to possess this sort of capacity, or it hasnât been deployed in this instance.
I donât know how the reader views the fact that the reviewers and editors of the Journal of Medical Ethics didnât undertake a reference check for this manuscript given some of the already mentioned other âred flagsâ that were in place. ArandeloviÄ thinks that they failed the readers of the journal, and ultimately that the journal failed as an institution on this occasion. It is easy to point fingers here, but truth be told, it has become exceedingly difficult to find willing reviewers for the ever-increasing number of papers that are prima facie worthy of peer review, and those reviewers are still expected to work pro bono, because publishers donât wish to pay for their services. I am not surprised that reviewers do not spend their volunteer time undertaking detailed (or any) reference checks, especially given that the publisher could have systems in place that undertake that sort of task automatically.
Schuklenkâs editorial is here.
(via Johann Go)
A little digging suggested that the author didnât even exist. The paper appeared to have been generated by AI and attributed to a fabricated scholar. That immediately brought to mind something Iâd recently read about: the possibility of creating âacademic avatarsââŚ. basically invented researchers with plausible institutional affiliations, email addresses, publication histories,etc and online profiles. In principle, AI now makes it cheap to generate a steady stream of competent-looking papers under such an identity.. so you can gradually start building what appears to be a genuine scholarly record. Then, as the scam goes, once the fictitious researcher has accumulated enough publications, citations, and digital presence, someone could plausibly âstep intoâ that identity, â they do that by presenting themselves as the person behind the established record . giving one public talk under the avatar is the coming out party. The real value of the scheme would come later: the laundered academic reputation could then be hired out to lend apparent expert authority to opinion pieces (as well as consultancy reports, policy interventions etc.) or even expert testimony on behalf of whoever was willing to pay.
And AI makes it cheap to detect as well. See my comment.
As someone working on the use of AI in journal production processes (to bring down production costs for in-house journal production for non-profit academic publishing), I have used Claude Code to write a reference validation script (so it runs independently of AI to prevent further hallucinations) as part of a production pipeline.
The application of this script immediately flagged 6 suspicious references from this article. I then tested this by using Claude Cowork to double-check. Same 6. This does not require any proprietary and âhighly sophisticated automated reference checkâ from Wiley. This can be done easily by the editorial staff or even automated as part of the submission process with little or no effort.
In an important sense, the failure here isnât AI at all. Rather itâs the workplaceâs. AI has made fabricated references cheap to produce, but the duty to check them was never technologically hard: as I said, a trivial non-AI script flagged all six in under a minute. What has broken is an incentive structure economists call moral hazard (whoever skips the check doesnât bear the reputational cost) and adverse selection (production is tendered on price, and diligence is invisible at the point of decision). Editorial rigour is a credence good, so no one downstream can see it was skipped. For whatever reason the workflow is not being adhered to. Why? Maybe the incentives of the corporate publishing model. Just consider that in fiscal 2024 Wileyâs Research division â the journals business that runs on the voluntary labour of academic authors, reviewers, and editors â turned revenue of US$1,043M into US$331M of adjusted earnings, a margin of 31.8%. With returns like that resting on unpaid academic work, the problem is not AI but the failure to apply HI â Human Intelligence â to designing publication models whose incentives are resolved in favour of academic integrity rather than the bottom line.
The problem is not AI and hallucinated references, but the business model that the affected journals are tied into. Anyone following the postings on Daily Nous will find that a host of high ranking philosophy journals have flipped from Wiley. That is the way forward.
Check here:
https://s27.q4cdn.com/812717746/files/doc_financials/2024/q4/Q424-Earnings-Presentation.pdf
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content â general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached â you'll always get the same 5 for this article.