TMLR's Bold Experiment: Asking Authors to Explain Their Own Papers — The Results Are Alarming

TMLR asked authors of 10 desk-rejected papers to explain their work — none passed the test.
TMLR ran an unusual experiment: contacting authors of 10 desk-rejected submissions and asking them to explain their own papers. The results were troubling — one withdrew, one declined, one no-showed, three couldn't answer basic questions, three could only discuss high-level ideas, and the one author who answered everything still had a major flaw identified by the editor. Not a single paper survived this most basic check. The experiment reflects a growing crisis in the LLM era: generative AI has lowered the barrier to producing formally convincing papers, but cannot give listed authors genuine understanding of the research. TMLR's move signals clearly — authorship means accountability.
TMLR's Bold Experiment: Asking Authors to Explain Their Own Papers — The Results Are Alarming
TMLR (Transactions on Machine Learning Research), one of the leading journals in machine learning, recently ran a provocative yet revealing experiment: for 10 papers flagged for desk rejection, the editorial team reached out to the authors and asked them to explain their own submissions in person. This seemingly straightforward approach exposed a growing crisis in academia — who actually wrote the paper?
According to an official post published by TMLR on Medium, the findings were deeply unsettling.
Ten Papers, Ten Ways to Dodge

TMLR's Co-Editor-in-Chief personally reached out to the authors of all 10 papers, inviting each to a brief interview. The results were dramatic:
- 1 paper: The author withdrew the submission outright.
- 1 paper: The author declined, citing scheduling conflicts.
- 1 paper: The author scheduled a meeting but never showed up.
- 3 papers: The authors could not answer basic questions about their own work.
- 3 papers: The authors could discuss high-level ideas but fell apart when pressed on technical details.
- 1 paper: The author answered all questions — but the interviewer (the Co-EiC) identified a major flaw in the paper during the conversation.
In other words, not a single one of the 10 papers passed this most basic test: can the author explain what they submitted? Either authors actively avoided the conversation, or they couldn't articulate the research they had put their name on. The original post described this as "concerning" — a sentiment widely echoed across the academic community.
Beyond Desk Rejection: What This Really Points To
Desk rejection is a standard editorial mechanism for quickly filtering out clearly unqualified submissions before they go out for peer review. TMLR's decision to follow up with interviews wasn't about giving these papers a second chance — it was about verifying an increasingly common suspicion: were these papers genuinely written and understood by the people who submitted them?
The answer, based on the results, is largely no. When a named author can't answer basic questions about their own paper, or can only recite high-level talking points without being able to explain the technical implementation, several possibilities come to mind: the paper was ghostwritten by a third party, it was assembled using generative AI, or the listed author was a nominal contributor with no real involvement in the research. Any of these scenarios crosses a fundamental line of academic integrity.
The one case where the author answered everything fluently is also instructive — even then, the editor identified a significant technical flaw during the interview. This reminds us that being able to explain a paper is merely the baseline; the scientific quality of the work itself is a separate hurdle entirely.
Academic authorship norms are governed by clear international ethical guidelines. For example, the International Committee of Medical Journal Editors (ICMJE) requires that every author make substantive contributions to the design, data collection, or analysis of a study, and be accountable for its content. "Honorary" and "ghost" authorship have long been considered forms of research misconduct. The rise of generative AI pushes this problem to a new extreme — when a paper's core content is generated by a large language model and the listed author merely submits it, traditional authorship ethics face a fundamental challenge. Most top journals now require authors to disclose AI tool usage, but in practice there is still no universally enforceable standard for distinguishing "light editing assistance" from "substantive ghostwriting."
The Flood of AI-Generated Papers and the Peer Review Crisis
The broader context of this experiment matters. As large language models have become widely accessible, the barrier to producing a "convincingly formatted" machine learning paper has dropped dramatically. Generative tools can quickly produce text with proper structure, accurate terminology, and even formulas and experiment tables — yet none of this necessarily reflects real research, and the listed authors may not genuinely understand any of it.
For a journal like TMLR — which uses a rolling submission model and emphasizes rigorous review — the sharp increase in submission volume combined with declining quality creates a serious tension. Traditional peer review relies on reviewers reading submissions carefully; it's costly, slow, and easily overwhelmed when confronted with a flood of low-quality or AI-padded papers. TMLR's experiment is essentially an exploration of a low-cost "authenticity check" — just talk to the author, and use the most human approach possible to verify whether the paper genuinely belongs to them.
TMLR (Transactions on Machine Learning Research) is published by the Machine Learning Foundation (MLF) and launched in 2022 as a fully open-access, continuously rolling journal — distinct from traditional publications that release batched issues several times a year. Its review model emphasizes completing reviews within fixed deadlines and encourages transparent interaction between authors and reviewers. Because it uses a rolling submission system with no natural deadline-driven submission cycles, TMLR can face submission pressure at any time. This makes it less buffered against the wave of AI-assisted writing than traditional venues like NeurIPS or ICML, which adds urgency to its motivation for running this authenticity experiment.
Can This Approach Scale?
"Asking authors to explain their own papers" sounds simple, but turning it into a standard part of the review process faces real challenges.
Interviews require significant time investment from editors or reviewers. Ten papers can be handled one by one, but this clearly doesn't scale to thousands of submissions. There are also fairness concerns: language proficiency, real-time verbal expression, time zone differences, and other individual factors can all affect performance in an interview, and none of these are perfect proxies for the academic merit of the work. Additionally, establishing consistent standards for what counts as "unable to answer," and deciding whether authors should be given preparation time, would require formalized protocols.
That said, as a targeted spot-check or appeals mechanism, the value of this approach is clear. It sends a signal to the academic community: authorship implies accountability, and authors should be able to stand behind and explain their own research. In an era where AI-generated content is increasingly indistinguishable from genuine scholarship, this oldest of verification principles feels more essential than ever.
A Wake-Up Call for Researchers
For individuals and teams working in machine learning research, TMLR's experiment is a stark warning. Using AI to assist with writing, organizing literature, or polishing language is entirely legitimate — but if the core contributions, technical details, and experimental conclusions of a paper are things the author doesn't truly understand, no amount of polished formatting will get them through an author interview.
Academic integrity has never been just about formal compliance. It's about researchers genuinely understanding and taking ownership of their work. As journals and conferences increasingly introduce authenticity verification mechanisms, papers produced through shortcuts will find it harder and harder to slip through the cracks.
Related articles

LLM Selection Strategy for Multi-Agent SOC Applications: Rule-Based Routing vs. LLM-Driven Decisions
Should multi-agent SOC apps on LangGraph use rule-based routing or LLM-driven model selection? This article analyzes both approaches and recommends a hybrid strategy for security operations.

Snap Pushes Its $2,200 Smart Glasses Again — Can It Convince the Market?
Snap launched new features for its $2,200 smart glasses, doubling down on AR. We break down the pricing dilemma, its rivalry with Meta Ray-Ban, and what it means for the AR glasses race.

Vercel AI SDK Update: Multi-Turn Reasoning Preservation for Alibaba Models
Vercel AI SDK releases @ai-sdk/alibaba@1.0.55, enabling reasoning preservation by default in multi-turn requests for supported Alibaba models like Qwen.