ACL ARR Score Underwhelming? Should You Withdraw and Submit to a Workshop? A Practical Guide

How to decide whether to withdraw a borderline ACL ARR paper and submit it to a workshop like BlackboxNLP.
Facing a mediocre ACL ARR score with no path to the main conference? This guide explains how Rolling Review works, why interpretability papers get misjudged, and offers a three-step framework to decide between waiting out the rebuttal or withdrawing to submit to a targeted workshop like BlackboxNLP.
The Real Dilemma of a First-Year PhD Student
Researchers who have just entered the NLP academic community often find the ACL ARR (ACL Rolling Review) review process confusing. A first-year PhD student once shared a classic dilemma on Reddit: he submitted an interpretability paper to EMNLP and received mediocre scores of 2.5/3, 3/4, and 2.5/4 in a certain ACL ARR cycle.
The reviewers had no major objections to the paper's methodology or overall quality. The crux of the problem was that they seemed unable to grasp the paper's core value—the "so what" of the research. The author did his best to clarify during the rebuttal phase, but the reviewers showed little willingness to actively engage in follow-up discussion.
With his current scores, he could get into neither the main conference nor even Findings. So he began to consider: withdraw the paper, polish the presentation, and resubmit to the BlackboxNLP workshop, whose deadline was the following week. Was this a wise decision?
Understanding How ACL ARR Works
The Core Logic and Historical Background of Rolling Review
ACL ARR (ACL Rolling Review) was officially launched in 2021 as a systematic reform by the Association for Computational Linguistics (ACL) to the traditional conference review model. Before this, the major conferences in the NLP field (ACL, EMNLP, NAACL, etc.) each ran their own independent review systems. This meant that when the same paper was repeatedly submitted to different conferences, reviewers had to be recruited anew and the review process repeated, resulting in a huge waste of community resources.
ARR borrowed the "rolling" concept from journal review, decoupling review from publication decisions: a paper first receives review comments and scores in the ARR system, and then the author decides which specific conference to "commit" the results to. Review results can be reused within a certain period. After receiving scores, the author can choose to "commit" the paper to a specific conference or Findings, or withdraw it and re-enter a later cycle carrying the existing reviews.
The original intent of this design was to reduce duplicate labor, but critics also point out that in practice ARR increases process complexity, and the acceptance standards of different conferences still vary. For newcomers, it likewise brings strategic complexity: when should you persist, and when should you cut your losses in time?
What the Scores Actually Mean
The range of 2.5 to 3 falls into the "borderline-below" position in the ACL review system. There are no fundamental flaws in the methodology, but "reviewers didn't understand the contribution" is itself a dangerous signal—against the backdrop of fierce competition for main conference acceptance, a paper that fails to clearly convey its value has a hard time standing out.
More realistically, the probability of reviewers actively raising their scores after the rebuttal has always been low. Multiple statistical analyses of conferences like ACL, EMNLP, and ICLR show that the probability of the rebuttal leading to a substantial score increase by reviewers is usually below 20%, and score improvements tend to concentrate on papers that already had relatively high initial scores (above 3.5). For borderline papers, the rebuttal more often plays a "damage control" role—preventing reviewers from further lowering scores rather than turning the situation around. The reasons for this phenomenon include the cognitive anchoring effect after reviewers submit initial scores, the limited discussion time window, and the inherently low engagement of some reviewers with the rebuttal phase. Understanding this statistical reality helps PhD students enter the rebuttal phase with more rational expectations.
Withdraw and Submit to a Workshop: A Comprehensive Analysis of the Pros and Cons
The Disciplinary Position of Interpretability Research
Before delving into the withdrawal strategy, it is necessary to understand the structural challenges that interpretability papers face in academic review. Interpretability research is one of the fastest-growing directions in the NLP field in recent years. Its core goal is to reveal the internal working mechanisms and decision-making basis of neural network models—especially large language models. The "so what" of such research often lies not in improving a benchmark metric for a certain task, but in advancing understanding of model behavior, which is essentially an epistemological contribution.
For this very reason, reviewers accustomed to "performance-oriented" review standards sometimes struggle to quickly recognize the core value of interpretability papers—this is also the structural reason why such papers are prone to being "misjudged" in the review of comprehensive top-tier conferences, rather than simply a matter of paper quality.
The Core Advantages of Submitting to a Workshop
For papers in the interpretability direction, BlackboxNLP is a highly relevant top-tier workshop. Since its first edition at EMNLP in 2018, it has become one of the most influential academic exchange platforms in this niche field. The word "Blackbox" in the workshop's name directly points to its research motivation: the internal mechanisms of neural network models are like a black box to the outside world, and this workshop brings together a community of researchers dedicated to opening this black box.
BlackboxNLP's paper acceptance rate is usually higher than the main conference, but this does not mean its academic standards are lax—workshop papers also undergo peer review, and the reviewers are often experts deeply engaged in the direction, making their judgment of contributions more precise. Importantly, workshop papers are formally archived in the ACL Anthology and are fully citable.
Several substantive benefits of submitting to a workshop:
- Better reviewer match: Specialized review in a niche field better understands the paper's true "so what" and can recognize the epistemological value of interpretability research.
- Higher acceptance probability: Workshop acceptance rates are usually much higher than the main conference, making the entry threshold more reasonable.
- Establish a publication record early: For a first-year PhD student, a citable result can significantly boost confidence and momentum for subsequent research.
- Precise community exposure: Workshops often gather the most active researchers in the direction, and poster and oral presentation sessions are efficient channels for building academic connections.
Potential Costs That Cannot Be Ignored
Withdrawing also means giving up the existing review results and the potential Findings opportunity. Several points are worth carefully weighing:
- Findings is still possible: Borderline scores may still be accepted into Findings in certain years, depending on the overall quality distribution of that cycle.
- The gap in CV value: Main conference or Findings papers usually carry more weight in resumes, job applications, and even postdoc applications than workshop papers, and this gap cannot be ignored in certain scenarios. However, for researchers focused on interpretability, BlackboxNLP's community recognition can to some extent make up for this brand gap.
- Extremely compressed time window: If there is only one week left until the deadline, rushed revisions may actually lower the paper's quality and backfire.
A Three-Step Decision Framework for First-Year PhD Students
Step 1: Objectively Assess the Score Trajectory
Before making the final decision, first complete the entire rebuttal and observe whether any reviewers respond positively. As long as one reviewer is willing to raise their score or shows understanding, it is worth waiting for the meta-review. The ACL ARR mechanism allows you to retain all your options before committing—don't give up this card prematurely.
Combined with the statistical reality mentioned earlier—that the probability of a score increase from rebuttal is generally low—if reviewers have clearly indicated during the rebuttal that they have no intention of engaging in deeper discussion, this itself is an important signal that can appropriately speed up your decision-making pace.
Step 2: Distinguish Between a "Presentation Problem" and a "Contribution Problem"
The reviewers "not getting the point" is actually an actionable signal. Cognitive science research shows that a reader's judgment of a paper's core contribution is often largely completed within reading the abstract and the first two paragraphs of the introduction—the subsequent methods and experimental details play more of a "verification" role than a "persuasion" one. This means that the quality of the paper's framing narrative largely determines the reviewers' first judgment of the work's value.
If the problem truly lies only at the presentation level, then no matter where you ultimately submit, the top priority is to clearly articulate the "so what." The way the core contribution is expressed often determines the paper's fate more than the method itself.
When revising, try to directly answer these three levels of questions: What real problem does this research solve? Why are existing methods insufficient? What essential difference does your solution bring, and what does this difference imply for a broader group of researchers? For interpretability papers, the last question is especially critical—"soft contributions" require the author to proactively build an evaluation framework for the reader, rather than expecting reviewers to figure it out on their own. Write the answers to these three questions clearly into the abstract and the first two paragraphs of the introduction.
Step 3: Set a Clear Decision Bottom Line
A rational and feasible path is:
- Complete the entire rebuttal and observe whether the final scores show substantial improvement;
- If the scores remain unchanged and it is confirmed that neither the main conference nor Findings is within reach, decisively withdraw the paper;
- If time permits, submit with the improved presentation to a highly relevant workshop such as BlackboxNLP.
For a first-year PhD student, "done" is often more important than "perfect." Rather than keeping the paper in a drawer waiting for the next full cycle, it is better to publish it in an appropriate form first, paving the way for subsequent research and building your own confidence.
Final Thoughts
This case reflects a reality faced by many young researchers: academic publication is not just a contest of research ability, but also a game of strategy and expression. The rolling mechanism of ACL ARR gives authors greater flexibility, but it also requires them to more soberly assess the actual position of their papers.
Interpretability papers face the structural dilemma of "contribution-type mismatch" at comprehensive top-tier conferences. This is not a failure of the research itself, but a friction that objectively exists in the academic publication ecosystem. For interpretability papers that are methodologically solid but disadvantaged only in presentation, submitting to a precisely relevant workshop like BlackboxNLP is often a pragmatic and wise choice—there, your work is more likely to meet people who can truly evaluate it.
There is only one key point: don't waste any opportunity for improvement, and truly make the value of your research clear.
Key Takeaways
Related articles

The Truth Behind Codex 'Build a Website in 5 Minutes': AI Isn't Creating Sites—It's Helping You Copy Them
Exposing the truth behind viral Codex 5-minute website videos: creators aren't building original sites with AI—they're copying shared prompts or scraping others' work. Learn AI coding tools' real limits.

Getting Started with AI Agent Development: A Complete Guide from Concept to Practice
A comprehensive guide to AI Agent architecture and development, covering automated marketing, intelligent customer service, and investment analysis scenarios with single and multi-agent collaboration.

The Truth Behind Codex 'Build a Website in 5 Minutes': AI Isn't Creating Sites — It's Helping You Copy Them
Exposing the truth behind viral Codex 5-minute website videos: creators aren't building original sites with AI — they're copying shared prompts or scraping others' work.