Low Initial Review Scores for NeurIPS 2026 Theory Papers? Community Discussion and In-Depth Analysis

Community discusses why NeurIPS 2026 theory papers receive systematically lower initial review scores.
A Reddit discussion reveals that NeurIPS 2026 main track theory papers are receiving lower-than-expected initial review scores. This article explores structural reasons including reviewer-paper mismatches and difficulties in evaluating theoretical contributions, examines whether scores are generally lower this year, and provides practical rebuttal strategies for theory researchers.
A Community Discussion on NeurIPS Initial Review Scores
Every year during the NeurIPS (Neural Information Processing Systems) review season, AI researchers worldwide find themselves in a state of heightened attention and anxiety. NeurIPS is one of the most prestigious academic conferences in AI and machine learning, standing alongside ICML and ICLR as the "Big Three" ML conferences. The conference began in 1987, originally named NIPS, and was renamed NeurIPS in 2018. Submission volumes have surged in recent years, exceeding 15,000 papers in 2024, with acceptance rates typically around 25%. The review process uses a double-blind system, with each paper typically assigned 3-4 reviewers. The workflow encompasses initial scoring, author rebuttal, reviewer discussion, and Area Chair (AC) decision-making across multiple stages.
Recently, a discussion thread appeared on Reddit's machine learning community (r/MachineLearning) about NeurIPS 2026 Main Track Theory Papers, sparking collective resonance among theory-focused researchers.
The original poster shared their initial review results: three reviewers gave scores of 4/3/3, with confidence levels uniformly at 3/3/3. On NeurIPS's commonly used 1-6 scale, each score has a clear semantic definition: 6 means "strong accept," 5 means "lean accept," 4 means "borderline accept," 3 means "borderline reject," 2 means "lean reject," and 1 means "strong reject." Thus, scores of 4/3/3 indicate one reviewer with a mildly positive stance while the other two lean slightly negative—placing the paper squarely in the borderline zone between acceptance and rejection. A confidence score of 3 typically means the reviewer is "somewhat confident but not entirely certain" about their assessment, which is particularly common for theory papers—reviewers may not have fully verified all proof details and thus hesitate to make high-confidence strong judgments.
Based on experience from previous years, the author raised a noteworthy observation—theory papers tend to receive more conservative scores during initial review compared to other directions, and this year the initial scores across all areas seem to be generally lower.

The post's core request was straightforward: calling on other authors with theory submissions to anonymously share their initial scores and confidence levels to determine whether this is a systemic trend or merely an isolated individual experience.
Structural Reasons Behind Low Initial Scores for Theory Papers
Theory papers have long faced unique challenges in top-conference review processes. Several structural factors are worth examining.
Reviewer Matching Misalignment
As a conference that leans empirical, NeurIPS has a relatively limited pool of reviewers with solid mathematical theory backgrounds. Reviewer assignment at large academic conferences is a complex optimization problem—in recent years, NeurIPS has adopted automated matching systems based on the Toronto Paper Matching System (TPMS) and the OpenReview platform, computing match quality by analyzing reviewers' publication histories, self-reported research areas, and paper keywords. However, this system often underperforms for theory papers: the number of active researchers in theoretical directions (such as learning theory, optimization theory, and statistical learning theory) is far smaller than those in deep learning experimental work. Combined with the dilution of the reviewer pool caused by continuously growing submission volumes, the probability of mismatches increases significantly.
When a paper filled with theorem proofs, convergence analysis, or generalization bound derivations is assigned to an experimentally-focused reviewer, that reviewer often struggles to fully evaluate its technical contributions and tends to assign conservative middle scores (such as 3) paired with moderate-to-low confidence. This also explains why the poster's confidence scores were uniformly 3—this itself signals "I can't fully understand it but don't dare score it high." A paper exploring PAC learning lower bounds might be assigned to a reviewer who primarily works on computer vision experiments—this asymmetry in domain expertise is one of the root causes of the theory paper review dilemma.
The Difficulty of Evaluating Contribution Format
Empirical papers can persuade reviewers with intuitive metrics like SOTA (state-of-the-art) performance and benchmark improvements, while theory papers derive their value from insight, rigor, and generality—precisely the qualities hardest to recognize in a short timeframe. SOTA refers to the best performance level achieved on a specific task and dataset, while benchmarks are standardized evaluation bases for measuring algorithm performance (such as ImageNet, GLUE, MMLU, etc.). In current machine learning research, many empirical papers present their core contributions through statements like "surpassing the previous best method by Y percentage points on X benchmarks"—this evaluation paradigm is simple, intuitive, and easy for reviewers to quickly assess incremental contributions.
In contrast, theory papers offer mathematical theorems and proofs—for example, a tighter generalization bound or a new convergence rate—whose value lies in revealing fundamental reasons behind algorithmic behavior. But this type of contribution is difficult to measure with simple numerical metrics. A beautiful theoretical result might require reviewers to spend hours verifying the proof, which is a luxury under tight review timelines.
Are NeurIPS Review Scores Generally Lower This Year?
The other observation mentioned by the poster carries broader significance: initial review scores this year across all directions seem lower than in previous years.
If this observation holds true, multiple explanations are possible:
- Continued submission volume inflation: NeurIPS submission numbers have repeatedly hit new highs in recent years (from approximately 6,700 papers in 2019 to over 15,000 in 2024), increasing reviewer burden. Each reviewer may need to evaluate 6-8 or more papers, and with limited time and energy, this may lead to more "defensive" scoring strategies—assigning middle scores is the safest choice, neither risking scrutiny for recommending acceptance of a problematic paper nor facing challenges during discussion for a strong reject.
- Changes in score calibration: Conference organizers sometimes adjust review guidelines or explanations of score semantics, making reviewers more cautious about high scores. For example, explicitly emphasizing that "scores of 5 and above should represent the top 10% of work in the field" could cause reviewers to collectively calibrate downward.
- Anchoring effect: When the community generally expects scores to be low, authors' psychological baselines for scores also shift downward, amplifying the subjective perception that "scores are low." The anchoring effect is a classic cognitive bias from behavioral economics, where people over-rely on the first piece of information they encounter when making judgments.
It's worth emphasizing that NeurIPS initial review scores are not final decisions. After author rebuttals, reviewer discussions, and AC comprehensive evaluation, scores often change significantly. History has no shortage of cases where papers initially scored 3/3/3 were accepted after a strong rebuttal. Therefore, an initial score of 4/3/3 is far from a final verdict.
The Practical Value of "Score Tracker Posts" for the Community
Every review season, the machine learning community spontaneously forms various "score tracker" discussion threads. These threads serve not merely to satisfy curiosity but to provide anxious researchers with a lateral frame of reference.
When an author discovers that their 3/3/3 score is actually quite common among concurrent theory submissions, the self-doubt of "is my paper particularly bad" can be alleviated. Similarly, by aggregating large samples, the community can roughly sketch score distribution profiles across different directions—this has practical reference value for assessing a paper's relative competitiveness, formulating rebuttal strategies, and planning next submissions.
However, it's important to maintain perspective: such self-reported data suffers from serious survivorship bias and self-selection bias. Survivorship bias is a classic statistical concept where people tend to focus only on samples that "survived" a filtering process while ignoring eliminated ones, leading to biased conclusions. Self-selection bias refers to the fact that people who participate in surveys or discussions are not randomly sampled. In the context of score tracker threads, both biases are particularly prominent: authors with very high scores may not bother sharing because acceptance feels secure, those with very low scores may stay silent out of embarrassment, and ultimately those willing to post tend to cluster in the middle range (3-5), potentially systematically distorting the true score distribution. The original poster specifically emphasized "share only if you're willing" and "note that it's a theory paper for same-category comparison"—a deliberate effort to improve data comparability.
Rebuttal Strategy Suggestions for Theory Researchers
For theory researchers currently in similar situations, here are several practical considerations:
-
Don't let initial scores hold you hostage. Theory papers typically have larger score fluctuation ranges than empirical papers—a clear and compelling rebuttal can completely turn things around. Rebuttal is a critical component of modern academic conference review—in the NeurIPS process, authors typically have about one week to write their response after receiving initial reviews, usually with a character limit of around 5,000. Effective rebuttals need to precisely address each reviewer's specific concerns rather than offering generic responses.
-
Leverage low reviewer confidence. When reviewer confidence is generally low, the rebuttal should focus on explaining core contributions more intuitively, reducing the reviewer's comprehension cost. Common strategies for theory papers include: using intuitive examples or diagrams to explain the significance of core theorems (rather than repeating proof details), supplementing precise comparison tables with existing results to highlight the magnitude of improvement, pointing out and politely correcting potential misunderstandings by reviewers, and when necessary, adding supplementary experiments or simulations that reviewers requested to substantiate the practical significance of theoretical results.
-
Participate in community discussions but maintain independent judgment. Tracker posts provide reference points, but a paper's fate ultimately depends on specific interactions between the AC and reviewers, not group averages. After rebuttals comes the reviewer discussion stage, where the Area Chair guides reviewers to exchange opinions on controversial points, ultimately making acceptance or rejection recommendations—the dynamic interplay at this stage determines a paper's fate far more than initial scores alone.
Ultimately, this Reddit discussion reflects the enduring tension that AI theory research faces in an increasingly empirical field. Scores are merely a momentary snapshot, while the value of genuine theoretical contributions often requires a much longer timescale to be fully appreciated.
Key Takeaways
Related articles

A Practical Guide to Switching from Cursor to Claude Code: Pitfalls and Solutions
A practical guide to migrating from Cursor to Claude Code covering habit adaptation, context rebuilding, the Plan-Check-Apply risk control method, and debugging strategies for developers.

Why Are AI Coding Assistants So Expensive? The Real Bill Behind the Harness
Deep dive into the hidden cost structure of AI coding assistants like Claude Code, Cursor, and Cline — revealing how system prompts, Agent round trips, and Prompt Caching impact your bill.

monolog: The AI Note-Taking App That Requires No Organization — Semantic Search Finds Everything
monolog is an AI note-taking app that eliminates folders and tags. Just chat to yourself, and AI understands your content and retrieves it via semantic search. Syncs across iOS, Android, Web, and more.