Why Don't Top ML Venues Cap Submission Counts? The Root Causes of the Peer Review Quality Crisis

Why ML top conferences don't limit submissions — and why fixing the review quality crisis is so hard.
Explosive submission growth at top ML venues like NeurIPS and ARR is driving a peer review quality crisis, yet the community has been slow to adopt the kind of per-author submission caps used in security and architecture. This article examines the cultural, career, and authorship-related barriers to reform, and surveys practical solutions including submission quotas, mandatory reviewing contributions, and AI-assisted desk rejection.
An Overlooked Systemic Problem
The machine learning research community is facing an increasingly serious challenge: an explosive growth in submission volume. Recent reviewing cycles — including those run through ACL Rolling Review (ARR) — have sparked widespread discussion about declining review quality.
Background on ACL Rolling Review (ARR): ARR is a rolling review system officially launched for the NLP community in October 2021, designed to decouple paper reviewing from acceptance decisions. Built on the OpenReview platform, it operates on monthly submission windows with approximately two-month review cycles. Under the traditional conference model, authors submit separately to each venue and wait for results; ARR lets authors submit to a shared review pool and, once reviewed, "commit" their results to multiple top venues such as ACL, EMNLP, and NAACL. This "review-then-commit" design was inspired by journal reviewing and intended to reduce duplicated effort and improve efficiency. However, because a single review can be committed to multiple venues, the marginal cost of each submission dropped significantly — and the openness of ARR produced an unintended consequence: researchers became far more willing to submit, causing the review pool to balloon. Since 2023, ARR has received thousands of submissions per cycle, drastically increasing reviewer load and drawing mounting complaints about review quality, making it a prime example of the current ML review crisis.
A researcher active across multiple fields posed a sharp question on Reddit: Why doesn't the ML community limit the number of submissions per author, the way other fields do?
The question may seem simple, but it cuts to the heart of a core tension in today's AI academic ecosystem — the imbalance between chasing publication counts and maintaining review quality.

Other Fields Already Have Precedents
Submission limits are not a new idea. In several mature sub-fields of computer science, such mechanisms have been in place for years and have proven effective.
How the Security Community Does It
Take CCS (ACM Conference on Computer and Communications Security), a top venue in cybersecurity. By placing reasonable caps on the number of submissions from any single author, the conference effectively controls reviewer workload and prevents review resources from being spread too thin.
CCS is one of the four flagship security conferences, with an annual acceptance rate of roughly 15–20%. Its submission limit policy stipulates that the same author may not appear as corresponding or first author on more than a set number of papers (typically 2–3) within a single review cycle; papers exceeding that limit are desk-rejected outright. The core logic is that a top conference's review resources are, in essence, a scarce public good. Without limits, a small number of high-output teams can consume a disproportionate share of those resources, creating an uneven playing field for other authors. Empirical data from security venues show measurable improvements in review depth and specificity after implementing such limits.
Practices in Computer Architecture
In computer architecture, conferences such as DAC (Design Automation Conference) have adopted similar strategies. DAC is the flagship venue for electronic design automation and chip design, and its submission limit policy follows the same reasoning — by constraining the number of submissions from core authors, it nudges researchers to focus their energy on more mature and polished work. These communities broadly believe that capping submissions not only eases reviewer burden but also encourages researchers to prioritize paper quality over quantity.
These success stories raise a critical question: if this mechanism works well elsewhere, why has the ML community chosen a different path?
The Deep Cultural Roots in the ML Community
Understanding this requires a close look at the unique ecosystem of ML research.
A Culture of Extreme Publication Speed
Machine learning is renowned for its breathtaking pace of iteration. A new idea can go from conception to a top-conference paper in a timeframe far shorter than in traditional disciplines. In this "move fast" culture, any move to cap submissions risks being seen as a brake on research vitality — researchers worry that artificial limits will cause them to miss the chance to stake a claim early in a rapidly evolving field.
The growth trajectory of NeurIPS (Conference on Neural Information Processing Systems) submissions is the best quantitative illustration of the ML publishing crisis. Around 1,500 submissions in 2012 corresponded to the first deep learning wave triggered by AlexNet; roughly 4,800 in 2018 reflected the explosion of GANs, reinforcement learning, and related directions; more than 10,000 in 2022, over 13,000 in 2023 — nearly a ninefold increase in a decade. Acceptance rates have been compressed from roughly 25% in earlier years to around 25% today, but the absolute number of accepted papers has grown from a few hundred to over 3,000. This means the total demand for reviewers is expanding even faster — NeurIPS 2023 mobilized approximately 15,000 reviewers, a substantial fraction of whom were doctoral students with limited experience, producing a low-quality equilibrium of "submit papers to get reviewing assignments." This exponential growth is both driven by the "publish or perish" culture and, in turn, reinforces it — a self-reinforcing cycle that is hard to break.
Fierce Career Pressure
In ML, whether competing for academic faculty positions or industrial research roles, publication count is often a key evaluation metric. The "publish or perish" pressure is especially extreme in AI: since NeurIPS, ICML, ICLR, and similar top venues have become important screening criteria for research hiring at Google, Meta, Microsoft, and other industry labs, top-conference paper counts directly affect salary negotiation leverage. PhD students, postdocs, and even senior researchers all face enormous pressure to keep publishing. In this incentive structure, capping submission counts would directly threaten the interests of many researchers, making it very difficult to build broad support.
The Reality of Large-Scale Collaboration
Modern ML research — especially in the era of large models and large-scale experiments — often involves massive collaborative teams. The GPT-4 technical report lists over 200 named authors; Google's PaLM paper similarly carries dozens of co-authors. In large industrial labs, a single researcher may be involved in dozens of projects simultaneously, appearing at different positions in different author lists.
In large-scale ML collaborations, determining authorship has become the central technical obstacle to implementing submission limits. Current ML papers exhibit three common authorship contribution patterns: "infrastructure authorship" common in industrial labs (those who provide compute, data, or engineering support); "supervisory authorship" in academic collaborations (advisors listed due to mentoring relationships); and genuine intellectual contributors. CRediT (Contributor Roles Taxonomy) offers 14 standardized contribution roles (such as Conceptualization, Methodology, Software, Data Curation, etc.) and has been adopted by Nature, PLOS, and other journals, but ML top conferences have not yet systematically introduced it. The four authorship criteria proposed by the International Committee of Medical Journal Editors (ICMJE) — substantial contribution, drafting or revising the manuscript, final approval, and accountability for the work — are similarly rarely enforced in the ML community. The absence of standardized contribution disclosure mechanisms means that "limiting submissions by lead author" has an inherent loophole at the operational level — researchers can circumvent limits by adjusting author ordering. How to distinguish core authors from auxiliary contributors — for example, separating the researcher who contributed the core algorithmic innovation from one who merely provided compute resources or data annotation support — remains without industry consensus. These complexities create genuine technical difficulties for any "limit submissions by author" policy, making its design quite thorny.
The Cascading Effects of the Review Quality Crisis
The most immediate consequence of unchecked submission volume is a sustained decline in review quality.
When each reviewer is assigned too many papers, thoroughly reading and rigorously evaluating each one becomes nearly impossible, triggering several vicious cycles:
- Superficial reviews: Reviewers have no time to dig into methodological details, and evaluations become cursory.
- Strong papers may get buried: In a sea of submissions, genuinely valuable work may be overlooked by reviewers stretched too thin.
- Reviewer burden grows, participation declines: Excessive workload makes senior researchers reluctant to take on reviewing duties, further degrading the quality of the reviewer pool.
The problems exposed in recent ARR cycles are a concentrated manifestation of this systemic pressure.
Possible Solutions
Faced with this dilemma, the ML community is not without options. Beyond directly capping submission counts, several other potential approaches exist.
Introduce Submission Quotas
Drawing on the experience of security and architecture communities, it would be possible to set a reasonable submission cap per author — especially for corresponding or lead authors. This requires the community to reach consensus on the definition of "lead author" and to establish operational rules by referencing standardized contribution disclosure frameworks such as CRediT.
Mandatory Reviewing Contribution Mechanisms
Linking submission rights to reviewing obligations directly balances supply and demand at the source. This mechanism is essentially an "internalization" of review resources — using institutional design to make submitters bear the social costs their behavior generates. The OpenReview platform has experimented with a "review credit" system, where completing high-quality reviews earns credits that can be used to offset future submission fees; the "authors must provide reviewing commitments" rule introduced at CVPR 2024 is the closest the ML community has come to such a mechanism in practice. Behavioral economics research suggests that the key design parameters for such systems include: the measurability of review quality (which affects incentive calibration), the enforcement cost of penalizing those who evade their obligations (which affects compliance rates), and exemption clauses for junior researchers (which affects fairness).
Improve Desk Reject Efficiency
A desk reject refers to a paper being rejected by the editorial board or area chairs before entering the peer review process. Common triggers include: failure to comply with submission formatting requirements, clearly being out of scope for the conference, substantial overlap with published work (exceeding plagiarism thresholds), or paper quality obviously falling below the conference's baseline standards. In ICLR 2023's practice, approximately 5–8% of submissions were desk-rejected before entering full review, effectively conserving reviewer resources.
In recent years, some conferences have begun exploring the use of large language models (LLMs) to assist with preliminary format checks and topic matching: ICLR and NeurIPS have used BERT/GPT-based text similarity models internally to detect duplicate submissions; the Semantic Scholar S2ORC corpus has been used to train paper topic classifiers to help assess whether a submission is out of scope. However, using LLMs to assess paper "quality" remains far more controversial — a 2023 ICLR experiment found that GPT-4's predictions of paper acceptance rates agreed with human reviewers roughly 30% of the time, slightly better than a random baseline but far from reliable. This suggests that AI-assisted desk rejection has practical value for format and topic matching, but quality judgment still requires human involvement. By implementing stricter initial screening to filter out clearly unqualified submissions before formal review, the pressure on the subsequent review process can be substantially reduced.
Conclusion: The Necessity of a Cultural Shift
Ultimately, the ML community's long reluctance to impose submission limits reflects deep cultural and incentive-structure problems, not mere policy oversight. The culture of rapid iteration, intense career competition, and the realities of large-scale collaboration together form the resistance to reform.
Yet as the review quality crisis becomes increasingly salient, more and more researchers are questioning the sustainability of the current model. The success stories from other fields demonstrate that moderate constraints need not stifle innovation — in fact, they may benefit the community by raising overall quality. Whether the ML community is willing to learn from those experiences, and how it might do so, will be a defining question for the health of its academic ecosystem going forward.
Key Takeaways
Related articles

Disaster and Glory of the Apollo Program: The History We Must Revisit Before Returning to the Moon
From the fatal Apollo 1 fire to Apollo 8's daring lunar orbit to Apollo 11's successful landing—revisiting the disasters, fears, and compromises of the Apollo program and their lessons for today's return to the Moon.

Netflix Trust Exercise Turns Into Firing Trap: Where Are the Boundaries of Corporate Trust?
A Netflix employee was fired after sharing private info in a trust exercise. We analyze the risks of corporate trust exercises and how employees can protect themselves.

AMD CDNA5 Architecture Deep Dive: Technical Evolution and the AI Computing Competition Landscape
Deep analysis of AMD's CDNA5 architecture covering Chiplet packaging upgrades, HBM memory evolution, and low-precision compute optimization, examining how AMD challenges NVIDIA's AI chip dominance.