The NeurIPS Review Dilemma: Why Do Reviewers Refuse to Raise Scores After Rebuttal Resolves Their Concerns?

Exploring why NeurIPS reviewers refuse to raise scores even after acknowledging rebuttals addressed their concerns.
A viral Reddit discussion highlights a pervasive problem in top AI conference peer review: reviewers who admit their concerns were resolved during rebuttal still refuse to update their scores. This "vibe review" phenomenon undermines the self-correcting function of the review process, distorts quality signals for Area Chairs, and compresses the diversity of scientific exploration. The post examines counterarguments, systemic causes tied to submission volume explosion, and proposes accountability reforms.
A Public Call for Academic Fairness
As the NeurIPS 2026 review season approaches, a post tagged as "Discussion" ([D]) on the Reddit machine learning community (r/MachineLearning) has struck a deep chord. r/MachineLearning is Reddit's largest academic ML discussion community, with over 3 million subscribers ranging from PhD students to prominent professors, industry researchers to independent scholars. Thanks to its anonymity and openness, many researchers are willing to discuss views here that would be difficult to express publicly in formal settings, making it a significant part of academia's informal public discourse.
The poster raised a question that seems basic yet has long troubled the academic community: When reviewers acknowledge that their concerns have been fully addressed during the rebuttal phase, why do so many still refuse to adjust their scores?
The poster admitted this might be an "unpopular opinion" (hot take), but the core appeal is straightforward and reasonable: if you list specific issues in your review and those issues are resolved one by one during the rebuttal, then regardless of whether you personally "like" the paper or its methodology, you should raise your score accordingly.

This post hit a nerve because it exposes a pervasive structural contradiction in the peer review mechanism of top AI conferences — the subjective drift of scoring criteria. "Subjective drift" refers to the gradual shift of review standards from objective quality indicators (such as technical correctness of methods, sufficiency of experiments, reproducibility of conclusions) toward reviewers' personal subjective judgments (such as whether the research direction is "hot," whether the method is "elegant," or whether the contribution is "exciting"). This drift is particularly severe against the backdrop of surging submission volumes — when reviewers need to evaluate multiple papers within limited time, they often rely on intuition and first impressions rather than systematically verifying each technical contribution. Research has shown that scores for the same paper can vary by 2-3 points (on a 10-point scale) between different reviewers, and this high variance is direct evidence of the subjectivity problem.
The Design Intent vs. Reality of the Rebuttal Mechanism
Rebuttal Was Meant to Be a Review Correction Channel
In the review process of top conferences like NeurIPS, ICML, and ICLR, the rebuttal phase was designed as a critical "error-correction mechanism." Taking NeurIPS as an example, the complete review process typically includes: initial assignment (where Area Chairs assign papers to 3-4 reviewers), independent review (reviewers provide scores and detailed comments within approximately 3 weeks), author rebuttal (usually a one-week window), reviewer discussion and score updates, and Area Chair synthesis for final decisions. The entire process uses double-blind review (authors and reviewers are unaware of each other's identities), aiming to ensure objectivity.
Reviewers first provide initial scores and comments, then authors have the opportunity to respond specifically: supplementing experiments, clarifying misunderstandings, correcting expressions, and providing additional evidence. The underlying logic of this design is that review is a dynamic correction process, not a one-shot verdict.
In theory, if a reviewer's core concerns are proven to be based on misreadings, or if supplemental experimental data from authors fills previous gaps, then the basis for the initial score no longer holds, and the score should be updated accordingly.
The "Vibe Review" Phenomenon in Practice
However, in reality, many reviewers fall into exactly the trap described by the poster: on one hand, they publicly acknowledge in their responses "Thank you for the clarification, my concerns have been addressed," while on the other hand maintaining their original score, with vague justifications — "I'm still not sold on this direction" or "it doesn't feel impressive enough overall."
The poster used an apt phrase to capture this phenomenon: reviewers simply "don't vibe with the paper." In other words, scores are no longer determined by whether the paper is rigorous or its contributions are valid, but by the reviewer's personal preferences and first impressions.
Why Reviewers Not Raising Scores Is a Serious Systemic Problem
Distorted Review Signals Affect Paper Acceptance Quality
When rebuttals cannot effectively influence scores, the entire review system loses its capacity for self-correction. Authors spend enormous effort writing rebuttals and running additional experiments, only to receive responses of "you're right, but I'm not changing my score" — this not only discourages research motivation but also reduces review outcomes to noise.
More seriously, this behavior undermines the credibility of review scores as "quality signals." Area Chairs (ACs) rely precisely on reviewer scores and their trends when making final decisions. Each AC typically oversees the review process for 20-30 papers, with responsibilities including: ensuring review quality, guiding discussions between reviewers, making adjudicative judgments when reviewers disagree, and submitting acceptance recommendations to Senior Area Chairs. When making decisions, ACs primarily reference initial reviewer scores, post-rebuttal score changes (score increases are typically viewed as positive signals of paper quality), the quality of inter-reviewer discussion, and the paper content itself. When reviewer scores cannot accurately reflect a paper's actual state after rebuttal, ACs must spend significant extra time personally reading papers and rebuttals — something nearly impossible to do for every paper given the massive volume. If scores don't reflect the paper's true state after rebuttal, AC judgments will also be misled.
The Diversity of Scientific Exploration Gets Compressed
The poster's final sentence touches on a deeper value: "The beauty of scientific research is that each of us can explore ideas we find meaningful, and the value of these ideas may not be obvious to every reviewer."
This is essentially a critique of review hegemony — reviewers using personal taste as a threshold to exclude research directions that don't match their aesthetic preferences. Historically, countless breakthrough works were not favored by the mainstream when first conceived. If the review mechanism tacitly permits "I don't like it so I'll suppress the score," then truly pioneering, non-consensus research will face even greater difficulty passing review gates.
Do Reviewers Have the Right to Maintain Their Score After Rebuttal?
It's worth noting that this issue is not without controversy. The counterargument holds that resolving "specifically listed concerns" in the rebuttal doesn't mean the paper overall meets the acceptance threshold.
A common defense is: the reviewer's initial score may have already generously given the author a chance to respond — meaning "I scored assuming the problems would be resolved; if you don't resolve them the score should drop, if you do it stays the same." Under this framework, the rebuttal's role is to "defend" rather than "improve" the score.
Additionally, there are cases where reviewers fail to fully list all their concerns in the initial review — the rebuttal addresses the three listed issues, but upon deeper reading, the reviewer discovers new, more fundamental flaws. Maintaining a low score in such cases has legitimacy, but the prerequisite is that the reviewer should clearly communicate the new reasons, rather than brushing it off with "don't vibe."
Implications and Improvement Suggestions for the Research Community
The value of this discussion extends far beyond NeurIPS 2026 itself. It reflects a review quality crisis in the AI academic community against the backdrop of explosive growth in paper submissions. AI conference submission volumes have experienced exponential growth over the past decade — taking NeurIPS as an example, submissions were approximately 1,700 in 2014 and surged to over 13,300 by 2023, nearly an 8-fold increase. ICML and ICLR show similar trends. The direct consequences of this growth are: the supply of qualified reviewers is severely insufficient, with many junior PhD students and even master's students being recruited into the reviewer pool; individual reviewer workloads increase, with each person typically needing to review 4-6 papers within 3-4 weeks; review quality becomes difficult to guarantee, and "fast-food reviewing" becomes the norm. In this context, the importance of the rebuttal mechanism is actually heightened — it is the last line of defense for correcting hasty initial reviews.
For reviewers, professionalism means: review comments should be specific and actionable; once raised issues are resolved, judgments should be honestly updated; if maintaining a low score, clear, debatable reasons must be given rather than appealing to vague personal feelings.
For authors, this also reminds us to precisely "close" each concern in the rebuttal, responding with data and logic rather than emotion, leaving reviewers no room to suppress scores based on "vibes."
For conference organizers, perhaps stronger accountability mechanisms need to be introduced in process design. In fact, several top conferences have begun exploring related reforms in recent years: ICLR pioneered a review quality scoring system that allows authors and ACs to provide feedback on reviewer quality, with low-quality reviewers potentially being deprioritized or excluded from the reviewer pool in the future; NeurIPS introduced a "consistency experiment" starting in 2021, randomly assigning some papers to two independent groups of reviewers to quantify the noise level of the review system — the 2022 experiment results showed that approximately 50% of papers received completely different accept/reject decisions between the two reviewer groups, a striking finding that directly drove subsequent reform discussions; additionally, the adoption of the OpenReview platform has made the review process more transparent, with some conferences now publishing complete review records for accepted papers. Going forward, it may also be necessary to require reviewers to explicitly state the specific reasons a rebuttal failed to address their concerns when maintaining their score, making "not changing the score" carry an explanation cost.
Ultimately, peer review is the cornerstone of academic community self-governance. Making scores return to evidence and the paper's intrinsic quality — rather than reviewers' personal taste — is a necessary prerequisite for maintaining the credibility of this institution.
Key Takeaways
Related articles

After Being Laid Off by AI, a Programmer Open-Sourced an AI CEO: Who Should the Automation Axe Really Fall On?
A CEO used AI as a reason to fire developers. They responded by open-sourcing an AI CEO, exposing the power bias in automation narratives and who really should be replaced.

A 4-Year Engineering Study Plan: The Path from Zero to Landing Your First Offer
A systematic 4-year engineering study plan covering foundation building, specialization, interview prep, and job hunting to help students build an actionable technical growth path.

Roc 0.1.0 Preview: A Fast, Friendly, and Functional New Programming Language
Roc language nears its first numbered release 0.1.0, transitioning from experimental to usable. Explore its platform architecture, core features, and toolchain.