Saturday, September 5, 2026
NewsWhite
ArXiv will ban researchers who upload papers full of AI slop
TECHNOLOGY

ArXiv will ban researchers who upload papers full of AI slop

By Jay PetersMay 15, 2026·Source: The Verge·17 views

ArXiv, the widely used preprint server where researchers share scientific papers before formal peer review, is moving to ban researchers who submit work containing clear signs of unchecked AI-generated content, according to The Verge. The platform will reportedly remove papers that carry what it describes as incontrovertible evidence of AI carelessness — including hallucinated references and stray meta-commentary left over from large language model outputs — and bar the researchers responsible from uploading further work.

To understand why this matters, it helps to understand what ArXiv actually is and why it occupies such a peculiar position in the research ecosystem. Founded in 1991, it predates the modern internet as most people use it, and for decades it served as a kind of honor-system repository where scientists — initially physicists, later mathematicians, computer scientists, economists and biologists — could share findings quickly without waiting months or years for formal journal publication. The implicit bargain was always that ArXiv was not peer-reviewed, that readers were sophisticated enough to evaluate what they found there, and that researchers were honest enough not to abuse the platform's openness. That bargain is now under considerable strain.

The arrival of capable large language models changed the economics of academic paper production in ways that are still being absorbed. Writing a plausible-sounding research paper, complete with abstract, methodology section and citations, became dramatically cheaper in time and effort. The problem is that these systems confabulate with great fluency. They invent citations that do not exist, attribute claims to authors who never made them, and produce text that reads as authoritative while containing no actual knowledge. A researcher who uses such a tool carelessly and submits the output without checking it is not simply cutting corners — they are potentially poisoning a shared resource that the scientific community depends on to stay current.

ArXiv's importance here is not trivial. For fields like machine learning and artificial intelligence, it is effectively the primary literature. Results appear there first, sometimes months before any journal would publish them, and researchers, journalists, and engineers treat those papers as meaningful signals about where the field is heading. If the platform fills with AI-generated noise — papers that cite sources that do not exist, or that contain the unmistakable fingerprints of unreviewed LLM output — then the signal degrades for everyone. The likely reading is that ArXiv has concluded it cannot afford to let the problem metastasize.

What makes the policy notable is its framing. ArXiv is not banning the use of AI tools outright, which would be both unenforceable and arguably counterproductive given that many researchers use such tools legitimately for tasks like editing or translation. Instead, the target is negligence — the failure to actually check what the model produced. Hallucinated references are the clearest tell, because they represent a specific kind of error that only arises when someone submits machine output they have not read carefully. A researcher who uses a language model and verifies the results would catch invented citations before submission. One who does not is essentially laundering AI output as their own work.

The consequences will fall unevenly across the research community. For established researchers at well-resourced institutions, the policy is unlikely to change much. They have reputational incentives to be careful and support staff who can help verify submissions. The more complicated question involves researchers working under intense pressure to publish — those early in their careers, those at institutions where output metrics drive funding and promotion, and those operating in fields where the pace of publication has accelerated to a point that careful scholarship is genuinely difficult. This group faces the greatest temptation to cut corners, and they are also the ones for whom a ban from ArXiv could be most professionally damaging.

There is also a detection problem that ArXiv has not publicly solved. Identifying hallucinated references is relatively tractable — citations can in principle be checked against databases of real publications. But meta-comments, the other example The Verge reports the platform citing, are more idiosyncratic and may be harder to catch systematically. This suggests the policy may initially depend heavily on community reporting, with other researchers flagging suspect papers, which introduces its own dynamics around competitive incentive and personal grievance.

The broader industry to watch here is the preprint infrastructure itself. BioRxiv, MedRxiv and other discipline-specific repositories face versions of the same problem, and whatever enforcement approach ArXiv develops will likely influence how those platforms respond. Journal publishers, who already operate AI-detection policies of varying rigor, will be watching to see whether a preprint-level crackdown shifts submission patterns. And within the AI research community specifically, there is a certain irony in a field that built the tools now being used to flood its own primary literature having to police that literature against those same tools. How ArXiv implements, communicates and enforces this policy over the coming months will be the real test of whether it can hold the line.

Originally reported by The Verge. Read the original article

Related Articles