The Flood of AI Junk Papers: The Academic Crisis Behind Nearly 600 Daily arXiv Submissions

Surge in suspected AI-generated papers on arXiv exposes a growing academic quality crisis.
A single academic field saw nearly 600 arXiv submissions in one day, many suspected of being AI-generated low-quality papers. This article examines how AI slop is overwhelming review systems, eroding academic trust, and threatening to contaminate future AI training data through model collapse, while exploring possible platform and community responses.
A Quiet Academic Flood
Recently, an observation that sparked heated discussion on Hacker News revealed a concerning phenomenon: a single academic field saw nearly 600 submissions to arXiv in a single day, with the majority suspected of being "AI slop" — low-quality content generated by AI. While the discussion was modest in scale, it strikes at a deep-seated vulnerability in today's academic publishing ecosystem: generative AI is diluting the overall quality of academic papers at an unprecedented pace.
As one of the world's most important preprint platforms, arXiv has long been the go-to venue for researchers in physics, mathematics, computer science, and other fields to share their work. Founded in 1991 by physicist Paul Ginsparg at Los Alamos National Laboratory, the platform originally served only the high-energy physics community for preprint sharing. It has since grown into a massive knowledge repository spanning mathematics, computer science, quantitative biology, statistics, and more, hosting over 2.4 million papers. Operated by Cornell University on a nonprofit basis, arXiv allows researchers to upload and access papers at no cost. The core value of preprints lies in making research available to the global community before formal peer review, dramatically accelerating knowledge dissemination — a benefit that became especially evident during the COVID-19 pandemic, when critical research was first shared through preprint platforms. However, it is precisely this openness and low barrier to entry that is now being exploited by mass AI-generated content, turning it into a double-edged sword.

What Is AI Slop: From Internet Garbage to Academic Padding
"AI slop" is a pejorative term that has gained traction in English-language tech communities over the past two years. It broadly refers to low-quality content mass-produced by large language models — content that lacks substantive value, is loosely reasoned, or contains misinformation. The word borrows the imagery of "slop" (pig feed, kitchen waste), implying that while the content may appear superficially "consumable," it has zero nutritional value. Originally used to describe AI-generated articles and images flooding the internet — degrading information quality in search engine results, social media feeds, and content farms — the phenomenon has now spread into academia, where its harm far exceeds ordinary internet content pollution because academic papers carry the serious mission of advancing human knowledge.
Why Academic Papers Have Become a Prime Target for AI Padding
Generative AI has dramatically lowered the bar for producing something that "looks like a paper." A structurally complete article with professional terminology and properly formatted citations can now be drafted by AI in minutes. Frontier large language models like GPT-4 and Claude, trained on vast bodies of academic literature, have mastered the writing conventions, terminology systems, and argumentation patterns of various disciplines, enabling them to generate academic text that is nearly flawless in form. This leads to several direct consequences:
- Formally compliant but substantively hollow: Papers feature all the standard sections — abstract, methods, experiments, conclusions — but lack genuine innovation or reliable experimental support. AI excels at mimicking the "shell" of academic writing — the appropriate passive voice, cautious phrasing, well-structured paragraphs — but cannot independently design meaningful experiments or produce genuine scholarly insights.
- Fabricated references and data: The "hallucination" problem inherent in large models can lead papers to cite nonexistent literature or present experimental results that cannot be reproduced. Hallucination is a fundamental technical flaw of large language models — the model is essentially a probabilistic next-token prediction system that doesn't truly "understand" the truth or falsity of facts, but instead generates the most "plausible" text continuation based on statistical patterns in its training data. When asked to list references, it generates citation entries that appear real based on common patterns of author names, journal titles, and paper titles — the authors may be real, the journals may exist, but the specific papers were never published. Similarly, models can generate numerical results that fit expected distributions but are entirely fabricated, making this type of fraud harder to detect than traditional data manipulation.
- Mass padding to inflate metrics: Some researchers or institutions use AI to rapidly produce large volumes of similar papers to artificially boost publication metrics. The root cause lies in structural flaws in the global academic evaluation system — many countries and institutions still rely heavily on quantitative metrics such as publication counts, h-index, and Impact Factor to assess researchers' academic performance, directly linking these metrics to promotions, funding, and institutional rankings. Under the enormous pressure to "publish or perish," AI padding becomes a low-cost, high-reward "rational choice," especially in academic environments where paper count serves as a core evaluation criterion.
Nearly 600 submissions in a single day is itself an anomalous signal. For any specialized field, this production rate far exceeds the normal pace of research, making it hard not to suspect that a significant portion consists of machine-generated content.
The Multifaceted Impact of AI Junk Papers on the Academic Ecosystem
Surging Review and Screening Pressure
Although arXiv is a preprint platform that doesn't conduct rigorous peer review, it does maintain a basic moderation mechanism. Specifically, arXiv's review system consists of automated screening tools and approximately 180 volunteer moderators working in tandem. The automated system checks submissions for basic formatting compliance, potential plagiarism, and similar issues, while volunteer moderators determine whether a paper belongs to its claimed subject category and meets minimum academic standards. It's important to emphasize that this moderation is not meant to judge the quality of a paper or the correctness of its conclusions — it merely ensures that a submission "looks like a serious piece of academic work." When submission volumes explode, this already thin line of defense faces collapse. Neither volunteer moderators nor automated systems can effectively distinguish genuine research from AI padding — precisely because AI-generated papers are getting increasingly better at "looking like serious academic work." The massive influx of low-quality content significantly increases the risk that truly valuable research gets buried, creating a "bad money drives out good" effect in academic communication.
Erosion of Academic Trust
Trust is the foundation of academic exchange. When readers cannot be confident that a paper is backed by real research, the credibility of the entire preprint system suffers. Researchers must spend ever more effort distinguishing authentic from fabricated work when searching the literature, effectively raising the "information noise" cost for the entire scientific community. The cascading effects of this trust erosion may far exceed expectations: if researchers begin to view arXiv papers with general suspicion, they may shift their attention exclusively to formally published versions in traditional journals. This would fundamentally undermine the core value of the preprint model in accelerating knowledge dissemination, pushing academic communication back toward a more closed and slower paradigm.
Self-Contamination of Training Data and Model Collapse
A more far-reaching concern is that these AI-generated papers are very likely to be scraped as training data for the next generation of large models. When models train on low-quality AI-generated content, a vicious cycle of "model collapse" emerges: errors are continuously amplified and solidified, and the reliability of academic knowledge degrades layer by layer.
The concept of model collapse was formally proposed in 2023 by a research team from the University of Oxford and the University of Cambridge and published in Nature. Through theoretical analysis and experiments, they demonstrated that when language models are repeatedly trained on data generated by themselves or similar models, a systematic degradation occurs: the diversity of model outputs decreases with each generation, tail information in the distribution (i.e., rare but important knowledge) is gradually lost, and eventually model outputs converge to a narrow and distorted distribution. In layman's terms, it's like repeatedly photocopying a document — each copy becomes more blurry and distorted than the last. In the academic context, this means that if AI papers flooding arXiv are absorbed by future models as "high-quality academic text," the flawed reasoning, fabricated data, and hollow arguments in those papers will be "learned" by models and reproduced in subsequent generations, forming a self-reinforcing chain of contamination.
This means the harm of AI junk papers extends beyond the present and could have profound negative effects on future AI systems. "High-quality human-generated data" on the internet is becoming an increasingly scarce resource, while the proportion of AI-generated content is growing exponentially. This diverging trend makes the urgency of training data contamination grow by the day.
The Platform and Community Response Dilemma
Faced with the trend of AI junk paper proliferation, preprint platforms like arXiv find themselves in a bind. On one hand, tightening review standards risks penalizing genuine independent researchers and betraying the platform's mission of open sharing; on the other hand, a laissez-faire approach will accelerate the decline in content quality.
Possible responses currently being discussed include:
- Introducing AI content detection mechanisms: However, existing AI text detectors have limited accuracy and can be easily bypassed by carefully edited text. Current mainstream AI text detection methods fall roughly into two categories: statistical feature-based detection, which analyzes metrics like perplexity and word frequency distribution (since AI-generated text often exhibits specific patterns in these metrics); and watermarking-based methods, which embed statistical signals invisible to the human eye but identifiable by algorithms during text generation. However, both approaches face fundamental limitations. Statistical detection methods struggle to reduce both false positive rates (misidentifying human text as AI-generated) and false negative rates (misidentifying AI text as human-written) to acceptable levels, especially when humans edit and polish AI drafts, which significantly reduces detection accuracy. Watermarking methods require all AI service providers to uniformly embed watermarks — practically impossible in a reality where commercial competition and open-source models coexist. OpenAI attempted to develop an AI text classifier but took it offline in 2023 due to insufficient accuracy.
- Strengthening submitter identity verification and endorsement systems: Raising the cost of padding to curb abuse, but also raising the submission threshold. arXiv already has an endorsement system — new users submitting to a subject category for the first time need endorsement from a researcher who already has submission privileges in that category. In practice, however, this system has limited constraining power and may create unfair access barriers for researchers from developing countries or non-mainstream institutions.
- Community annotation and rating systems: Relying on readers' collective intelligence to assess and flag paper quality. Similar efforts have appeared on some platforms, such as PubPeer's anonymous peer commentary system and Semantic Scholar's citation context analysis features.
However, all these solutions have clear limitations. Detection technology will always lag behind the evolution of generation technology, while manual review is difficult to scale. This arms race is fundamentally asymmetric — producing junk is extremely cheap (generating a 10,000-word paper using a commercial API may cost less than a dollar), while identifying and cleaning up junk is extremely expensive (requiring domain-expert human reviewers to invest significant time).
A Measured Perspective: The Legitimate Boundaries of AI in Academic Writing
It's worth adding that this Hacker News discussion was not particularly popular (only 13 points and 7 comments), and the claim that "most are AI slop" is more of a subjective observation than a rigorously validated statistical conclusion. We should avoid amplifying panic unnecessarily.
Generative AI also has legitimate and valuable uses in academic writing:
- Helping non-native speakers polish their language — this is particularly important because over 90% of high-impact academic papers are published in English, and non-native English-speaking researchers have long faced an unfair competitive disadvantage in writing quality. AI writing assistance tools have substantially leveled this playing field, allowing researchers to focus more on scientific questions rather than language barriers.
- Assisting with literature organization and review writing
- Accelerating draft writing and improving research efficiency
The tool itself is neutral; the problem lies in its misuse. The real challenge is not whether AI is being used, but how to distinguish "AI-assisted genuine research" from "AI-fabricated fake research." This requires the coordinated evolution of platform mechanisms, academic norms, and technological tools. Notably, top journals including Nature and Science have successively issued AI usage policies, generally permitting AI as a writing assistance tool but requiring authors to explicitly disclose the manner and extent of AI use in their papers, and emphasizing that AI cannot be listed as a co-author because it cannot bear academic responsibility for the research content.
Rebuilding Trust Mechanisms for Academic Content
The phenomenon of nearly 600 daily submissions to arXiv is a microcosm of the challenges facing academic publishing in the generative AI era. It reminds us that: when the marginal cost of content production approaches zero, "quantity" is no longer proof of value — it may even become an obstacle to it.
The healthy development of the future academic ecosystem may require a shift from an evaluation system that rewards quantity toward a quality-oriented mechanism that emphasizes reproducibility, genuine contributions, and long-term impact. In fact, this conversation began well before the AI wave — the 2012 San Francisco Declaration on Research Assessment (DORA) called on academia to stop using single metrics like journal Impact Factor to evaluate individual research outputs; the open science movement in recent years has also been promoting practices like study pre-registration, data sharing, and code sharing to increase research transparency. The explosion of AI padding may serve precisely as the catalyst to accelerate these reforms. In an era when AI can easily fabricate form, how to safeguard the substantive value of academic research is a question the entire scientific community must answer together.
Key Takeaways
Related articles

A New Framework for Evaluating Phonetic Encoding Algorithms: Rand Index, Discordance Scores, and Orthographic Transparency Quantification
An in-depth analysis of a new phonetic encoding evaluation framework based on the generalized Rand index, covering the Hüllermeier-Rifqi index, normalized edit distance, random baseline correction, and orthographic transparency quantification.

EXAONE Finance: A Deep Dive into the Time Series Foundation Model Built for Finance
Deep analysis of LG AI Research's EXAONE Finance model: how its attention-free architecture, masked context augmentation, and multi-asset financial corpus solve efficiency, missing data, and domain adaptation challenges to achieve SOTA on the FinVerse benchmark.

How Do Large Models Integrate External Evidence? Distributional Theory Reveals the Mechanism Behind LLM Evidence Integration
New research using 10M+ experiments reveals how LLMs integrate external evidence via distributional control, finding verification and integration are dissociated — with key implications for RAG and multi-agent systems.