koyejo
Position: Machine Learning Conferences Should Establish a "Refutations and Critiques" Track
Schaeffer, Rylan, Kazdan, Joshua, Denisov-Blanch, Yegor, Miranda, Brando, Gerstgrasser, Matthias, Zhang, Susan, Haupt, Andreas, Gupta, Isha, Obbad, Elyas, Dodge, Jesse, Forde, Jessica Zosa, Orabona, Francesco, Koyejo, Sanmi, Donoho, David
Science progresses by iteratively advancing and correcting humanity's understanding of the world. In machine learning (ML) research, rapid advancements have led to an explosion of publications, but have also led to misleading, incorrect, flawed or perhaps even fraudulent studies being accepted and sometimes highlighted at ML conferences due to the fallibility of peer review. While such mistakes are understandable, ML conferences do not offer robust processes to help the field systematically correct when such errors are made. This position paper argues that ML conferences should establish a dedicated "Refutations and Critiques" (R&C) Track. This R&C Track would provide a high-profile, reputable platform to support vital research that critically challenges prior research, thereby fostering a dynamic self-correcting research ecosystem. We discuss key considerations including track design, review principles, potential pitfalls, and provide an illustrative example submission concerning a recent ICLR 2025 Oral. We conclude that ML conferences should create official, reputable mechanisms to help ML research self-correct.
The Collapse of GPT
Ever since ChatGPT was released to the public in November 2022, people have been using it to generate text, from emails to blog posts to bad poetry, much of which they post online. Since that release, the companies that build the large language models (LLMs) on which such chatbots are based--such as OpenAI's GPT 3.5, the technology underlying ChatGPT--have also continued to put out newer versions of their models, training them with new text data, some of which they scraped off the Web. That means, inevitably, that some of the training data used to create LLMs did not come from humans, but from the LLMs themselves. That has led computer scientists to worry about a phenomenon they call model collapse. Basically, model collapse happens when the training data no longer matches real-world data, leading the new LLM to produce gibberish, in a 21st-century version of the classic computer aphorism "garbage in, garbage out."
Is It Possible to Truly Understand Performance in LLMs?
The lightning-like growth of large language models (LLMs) has taken the world by storm. Generative artificial intelligence (AI) is radically reshaping business, education, government, academia and other parts of society. Yet, for all the remarkable capabilities these systems deliver--and they are clearly impressive--a major question emerges: how can data scientists measure model performance and fully understand how they gain abilities and skills? It is far from an abstract question. These criteria, in turn, require an understanding of what constitutes correctness.