If You Think You’re Refereeing an AI
If You Think You’re Refereeing an AI-Written Submission, What Should You Do?
You’re refereeing a paper and you come to suspect it was substantially and/or illicily written by an LLM. What should you do?
My suggestion:
- Keep in mind the possibility that you are mistaken as well as the possibility that the AI use falls within the journal’s rules and was disclosed during the submission process.
- Pause your work refereeing the piece, as your suspicions of illicit AI use may unduly negatively influence your assessment of the submission.
- Gather a few examples of passages that strike you as AI-written, briefly explaining why. (Do not feed the manuscript into an AI-detector, for two reasons: first, doing so may violate the norms of confidentiality referees are expected to abide by, and second, such detectors appear to currently be of questionable accuracy.)
- Write to the journal’s editor or managing editor expressing your concerns and sharing the evidence you’ve gathered.
With luck, the editor will look into the matter and either agree with you and relieve you of refereeing the piece, or the editor will convince you that you are mistaken and you can proceed to finish refereeing the paper.
But what if you aren’t so lucky?
One associate professor of philosophy wrote in with the following:
I recently accepted a referee invitation for a good generalist journal. I was completely convinced after about 20 minutes with the submission that it was mostly (if not entirely) written using AI tools of some sort. The paper was in my areas of research. The editor disagreed (or at least did not share my level of confidence), and so I was left with the awkward task of writing (brief) comments for the “author” to go along with my rejection. For obvious reasons, this seemed to me to be an obvious waste of everyone’s time. Nothing like this has happened to me before, and I referee fairly regularly.
The professor is curious about the frequency with which reviewers are finding themselves in disagreement with journal editors and/or other reviewers on particular cases. Has this happened to you? What did you end up doing? How should referees and editors navigate disagreement on this?
and so I was left with the awkward task of writing (brief) comments for the “author” to go along with my rejection.
Or maybe the awkward task of telling the editor that your verdict is Reject but that you won’t be providing comments to be fed back into the model to generate the next version of the paper.
Have another LLM write the referee report.
As an editor, I’ve received a bunch of these papers. As people may be able to tell from my comments here, I’m ideologically committed to the idea that there *could* be a worthwhile paper whose text is substantially AI-composed, but so far only one of the dozen or so papers I’ve had this suspicion about seemed promising enough to send out to a referee. (There was one other that I unfortunately identified *after* getting referee comments, and I apologized to the referee after the fact.)
As far as comments, I don’t think anything that is actually AI-authored is getting past the initial desk reject phase. If the paper looks like a real paper throughout, and makes a point and gives some sort of argument for it, there’s almost certainly a real human there who is thinking about the project. In some cases, it’s a person working across disciplinary boundaries, or across languages, relying on the AI to get the disciplinary norms and language right (but unfortunately, these AI systems aren’t actually good enough tools to be effective for this purpose yet).
I’ve been writing fairly substantial comments on a bunch of these in my desk rejections, pointing out formalisms they take pains setting up but never use, or ideas that they say they will argue for but no argument comes. I’ve been doing this because there’s often an interesting idea there that I really want to see someone write a good paper about!
But I’ll probably stop doing this when it’s just another paper written by AI, arguing that AI has constitute epistemic limitations and therefore can’t be a testifier or a knower (which seems to be the claim of about half of these AI-written papers).
My attitude is similar to Kenny’s. So far I’ve seen some as an editor, and some as a referee. So far I’ve never thought: “this is a good paper, but it looks like it’s AI, so what do I do?” Rather, the kind of papers that are striking me as likely AI written are ones that go on for pages saying very little, or which have other serious substantive problems. So I’ll recommend rejection in a way that doesn’t have to cite their having been AI-written.
I suppose, like Kenny, I’m uncomfortable with the idea that an otherwise publishable paper–one that makes an interesting, worthwhile contribution to philosophy–should be rejected on grounds of AI authorship. Though that’s a bigger conversation. (Very roughly, my view is that just as a mathematical proof of a significant result is valuable because of the understanding it provides its readers, regardless of whether it’s mainly human or mainly AI-authored, likewise with a philosophy paper. If we get to the point that they’ve already gotten in math, where AI is playing a major role in generating important results, those results should absolutely be publicized, and the journal system is our discipline’s primary way of publicizing what we take to be worthwhile contributions that advance our collective understanding of philosophy.) But that’s also very much not what I’m seeing so far.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content — general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached — you'll always get the same 5 for this article.