NeurIPS decisions are coming soon, and I think back on many discussions with friends and peers on whether paper acceptance carries meaning anymore, given the ridiculous numbers of (accepted) submissions at these traditionally selective venues. For example, NeurIPS 2025 received more than 20,000 submissions, about a quarter of which were accepted.11. NeurIPS, “NeurIPS 2025 – Fact Sheet,” 2025. ↩ This is perhaps more than I can read in a lifetime.

Paper acceptance is a familiar event in a researcher’s career, with known immediate implications. It adds a publication to the CV, can make the work easier to publicize, and can help with hiring and funding. However, what happens after that notification is less predictable and likely more important for the community, as other researchers may build on the result or discover its limitations. But the latter has no equally clear place in how we evaluate the work compared to paper acceptance.

I think this imbalance deserves more attention in discussions of peer review. In these, we often ask how to make reviewers more careful and decisions more reliable. Even if they often come out of frustration, those are worthwhile questions! I argue that we should also ask what authors have reasons to optimize for, mirroring discussions of research incentives in older fields, such as psychology.22. Brian A. Nosek, Jeffrey R. Spies, and Matt Motyl, “Scientific Utopia: II. Restructuring Incentives and Practices to Promote Truth Over Publishability,” Perspectives on Psychological Science 7, no. 6 (2012): 615–631. ↩ This can be illuminating for the earlier question of peer review quality as well.

Consider an author deciding how to spend the week before a deadline. Improving the presentation may help reviewers understand the contribution. Checking a stronger baseline might reveal that the improvement is smaller than it first appeared. Documenting the experimental setup better could save another researcher days of work. The concern is that work which makes a paper easier to accept can be more rewarding for the author than work which makes its claims easier to trust.

Better reviewing can help align those rewards. For example, reviewers can insist on appropriate comparisons and ask authors to narrow unsupported claims. I would complement that effort by making evidence gathered after acceptance matter more to the credit a paper receives.

Making later evidence matter

I stumbled upon a study in recommender systems (the example is intentionally in the pre-AI era), where Ferrari Dacrema et al. (2019) revisited published “neural methods” and found that simpler baselines often outperformed the (few) state-of-the-art methods they could reproduce.33. Maurizio Ferrari Dacrema, Paolo Cremonesi, and Dietmar Jannach, “Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches,” Proceedings of the 13th ACM Conference on Recommender Systems (2019): 101–109. ↩

Now, generative AI may make it easier to search for analyses that support a preferred conclusion and persuade reviewers. For example, Bertran, Fogliato, and Wu (2026) found that AI analysts reached different conclusions from the same data and hypothesis, and that prompts to seek supporting evidence shifted the results.44. Martin Bertran, Riccardo Fogliato, and Zhiwei Steven Wu, “Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse,” Proceedings of the National Academy of Sciences 123, no. 29 (2026): e2606495123. ↩ Yet the same tools could help us examine, even before submission, how much a conclusion depends on unexplained analytical choices.

These studies show the value of looking more closely at published claims. However, my concern is how unevenly that scrutiny is distributed, and how uncertain its consequences are. A prominent paper may receive several independent investigations, while a less visible paper receives none. If authors expect little follow-up, or expect their paper’s findings to have little effect on how their work is evaluated, the incentive to invest in additional checking remains weak.

An audit lottery

One way to encourage this change would be for conferences to support independent checks on a randomly selected sample of accepted papers. The program would be announced before submission, so authors would know that their work might be examined.

Random selection matters because it extends the possibility of scrutiny beyond work that is already famous or controversial. Targeted investigations would still be valuable, particularly for consequential claims. But a lottery would give ordinary papers a chance of being checked, which is necessary if we want the possibility of verification to influence how ordinary papers are prepared. It is a bit like speed enforcement: police cannot check every driver, but if drivers know they might be checked, that possibility can virtuously affect how they drive. Likewise, researchers would not need to expect that every paper will be scrutinized for the prospect of scrutiny to influence how carefully they prepare their work.

Still in the interest of scalability, each audit would only examine a limited set of claims. Suppose a paper reports that a new training method outperforms existing approaches on several tasks. An independent team might reproduce a central comparison, examine whether the baselines received comparable tuning, and run additional seeds. The report would explain what those checks establish about the claimed improvement. If the available resources were insufficient to resolve the question, it would also say so.

The methods and supporting results should be public where possible, with space for an author response and another independent assessment of consequential disputes. Auditors can make mistakes too. The point would be to give readers evidence they can inspect, rather than another unexplained verdict. Selection would not automatically question acceptance, although findings could warrant corrections or other action through the appropriate procedures.

Fortunately, there are already efforts to build on in the machine learning community. ACM artifact badging provides an established mechanism for recognizing evaluated artifacts and reproduced or replicated results,66. Association for Computing Machinery, “Artifact Review and Badging,” Version 1.1, August 24, 2020. ↩ while TMLR recognizes reproducibility studies as publishable contributions and offers a dedicated reproducibility certification.77. Transactions on Machine Learning Research, “Accepted Papers: Reproducibility Certification.” ↩ More recently, MLRC became an official NeurIPS track in 2026, giving reproducibility work an additional path to conference-level recognition.88. NeurIPS, “MLRC 2026: Reproducibility as an Official Track at NeurIPS,” NeurIPS Blog May 4, 2026. ↩ OpenReview likewise provides infrastructure for keeping submissions, reviews, responses, and revisions linked in a persistent record.99. OpenReview, “About OpenReview.” ↩ Along similar lines, Schaeffer et al. (2025) propose a dedicated “Refutations and Critiques” track at ML conferences to give researchers recognition for critically examining published work.1010. Rylan Schaeffer et al., “Position: Machine Learning Conferences Should Establish a ‘Refutations and Critiques’ Track,” Advances in Neural Information Processing Systems 38 (NeurIPS 2025), Position Paper Track. ↩ Barnett et al. (2018) have already explored random audits in the context of health research, using simulations to study how they could encourage more careful work.1111. Adrian G. Barnett, Pauline Zardo, and Nicholas Graves, “Randomly Auditing Research Labs Could Be an Affordable Way to Improve Research Quality: A Simulation Study,” PLOS ONE 13, no. 4 (2018): e0195613. ↩ We could adapt their proposal to conference-supported checks of accepted papers, with findings attached to the work and considered in its subsequent evaluation.

Would it change behavior?

An audit lottery would only affect incentives if authors expected both a meaningful chance of being checked and meaningful consequences from the findings. Again, those consequences need not involve withdrawing acceptance. They could include greater confidence in a result, recognition for careful work, or reduced confidence in a claim that does not survive examination. But if audit reports are rarely read or used, the mechanism would be weak. Indeed, conferences can make the evidence accessible; they cannot, by themselves, ensure that the rest of the field takes it seriously. The program would certainly need real resources to work. For example, qualified auditors can receive credit through citable reports and support for their time and compute.

It is reasonable to imagine starting with a small funded pilot and examining both what it discovers and what happens to those findings. Do reports lead to corrections, better comparisons, or useful confirmations? Do subsequent researchers and reviewers use them? Is that value worth the effort, compared with supporting researcher-selected reproduction studies or improving the original reviews?

It seems that an audit lottery is worth trying. It actually seems “trivial” after the fact: if later scrutiny mattered more to the credit a paper receives, authors would have more reason to check and document their work before submission.

Acknowledgements

Thanks to Gautam Kamath, Sanmi Koyejo, Mahdi Haghifam, John Duchi, Andreas Haupt, Rylan Schaeffer, and Lydia Zakynthinou for stimulating discussions.