Goto

Collaborating Authors

 Personal


MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

arXiv.org Artificial Intelligence

We present MultiChallenge, a pioneering benchmark evaluating large language models (LLMs) on conducting multi-turn conversations with human users, a crucial yet underexamined capability for their applications. MultiChallenge identifies four categories of challenges in multi-turn conversations that are not only common and realistic among current human-LLM interactions, but are also challenging to all current frontier LLMs. All 4 challenges require accurate instruction-following, context allocation, and in-context reasoning at the same time. We also develop LLM as judge with instance-level rubrics to facilitate an automatic evaluation method with fair agreement with experienced human raters. Despite achieving near-perfect scores on existing multi-turn evaluation benchmarks, all frontier models have less than 50% accuracy on MultiChallenge, with the top-performing Claude 3.5 Sonnet (June 2024) achieving just a 41.4% average accuracy.


Review for NeurIPS paper: Emergent Reciprocity and Team Formation from Randomized Uncertain Social Preferences

Neural Information Processing Systems

Four knowledgeable referees reviewed this paper. After conducting initial reviews, reading the authors' rebuttal (which resolved some concerns, but not the core concerns of two of the reviewers), and discussing the paper, the reviewers did not agree on an outcome. Two of the reviewers came to the conclusion that this is a ground-breaking paper (simple and elegant). The other two reviewers were perhaps somewhat intrigued, but did not feel the paper was yet ready for publication. For example, during the discussion phase, R4 (a very accomplished and well-respected research in the field) made very valid points about the papers weaknesses: "So all this leads me to suggest that there needs to be a better context, more related work and a better way to situate the paper in related arenas, e.g., provide some sort of a framework to back up the findings. I understand the issue of limited space, but given the amount of literature in this area, I feel that the paper doesnt do a good enough job explaining its findings in context."


DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models

arXiv.org Artificial Intelligence

Most of the world's languages and dialects are low-resource, and lack support in mainstream machine translation (MT) models. However, many of them have a closely-related high-resource language (HRL) neighbor, and differ in linguistically regular ways from it. This underscores the importance of model robustness to dialectical variation and cross-lingual generalization to the HRL dialect continuum. We present DialUp, consisting of a training-time technique for adapting a pretrained model to dialectical data (M->D), and an inference-time intervention adapting dialectical data to the model expertise (D->M). M->D induces model robustness to potentially unseen and unknown dialects by exposure to synthetic data exemplifying linguistic mechanisms of dialectical variation, whereas D->M treats dialectical divergence for known target dialects. These methods show considerable performance gains for several dialects from four language families, and modest gains for two other language families. We also conduct feature and error analyses, which show that language varieties with low baseline MT performance are more likely to benefit from these approaches.


Review for NeurIPS paper: Forget About the LiDAR: Self-Supervised Depth Estimators with MED Probability Volumes

Neural Information Processing Systems

Weaknesses: I have no major concerns, but only remarks and suggestions for improvements. Although this is unambiguous in the experimental section, the abstract and introduction should clarify that the method is self-supervised from stereo pairs. There is a lot of confusion in the literature, because all monocular methods predict depth from a single image (by definition) but can be trained in different ways: from lidar supervision (full or partial), from stereo pairs (as is the case here), or from videos (a.k.a. Some of the authors' critique of related works (e.g., regarding dynamic objects) are only applicable to the SfM self-supervised scenario, as in the case of stereo-based self-supervised learning pairs of images are captured at the same time. Furthermore, the SfM case requires estimating the camera's ego-motion, which vastly complicates the self-supervised learning task (hence why the comparison is not entirely fair in my opinion).


Reviews: Piecewise Strong Convexity of Neural Networks

Neural Information Processing Systems

Originality: I am not convinced that the contributions of this paper are more significant than that of [1], which have been cited in this paper already. Specifically, in comparison with [1] in Line 82, the authors state that these conclusions apply to a smaller set in weight space. I would appreciate it if the authors could quantify the difference here and have a discussion section to show the comparison with some form of mathematical comparison. Further, there have been quite a few papers that show convergence of GD on neural networks using something like strong convexity. Clarity The paper is written quite clearly and it is easy enough to follow the paper.


Why are comedians trending toward Catholicism? One quirky comic offers a surprising explanation

FOX News

Comedian Anthony Rodia discusses the comedy industry and talks about the inspiration behind his jokes on'One Nation.' Though he may be covered in tattoos from head to toe -- quite literally -- the only thing more obvious than comedian Shayne Smith's body art lately might be his newfound Catholicism. And the former motorcycle gang member is certainly in good company. Jim Gaffigan, Kevin James, Stephen Colbert, Tom Leopold, Russell Brand, and Rob Schneider are just a few other comedians who share in the same faith -- the latter half of the boisterous bunch having converted to Catholicism in their adulthood. The former half has been just as busy keeping Catholicism alive: Gaffigan recently performed at The Sheen Center for Thought & Culture, at which Cardinal Timothy Dolan is a board member; Kevin James reportedly hosted a Catholic retreat before the pandemic; and Stephen Colbert is known for teaching Sunday school.


Reviews: DINGO: Distributed Newton-Type Method for Gradient-Norm Optimization

Neural Information Processing Systems

In this paper, the authors propose a distributed Newton method for gradient-norm optimization. The method does not impose any specific form on the underlying objective function. The authors present convergence analysis for the method and illustrate the performance of the method on a convex problem (in the main paper). Originality: The topic of the paper, in my opinion, is very interesting. The paper presents an efficient Newton method that is motivated via the optimization of the norm of the gradient.


Reviews: Adaptive Density Estimation for Generative Models

Neural Information Processing Systems

Summary: The authors propose a hybrid method that combines VAEs with adversarial training and flow based models. In particular, they derive an explicit density function p(x) where the likelihood can be evaluated, the corresponding components p(x z) are more flexible than the standard VAE that utilizes diagonal Gaussians, and the generated samples have better quality than a standard VAE. The basic idea of the proposed model is that the VAE is defined between a latent space and an intermediate representation space, and then, the representation space is connected with the data space through an invertible non-linear flow. In general, I think the paper is quite well written, but on the same time I believe that there is a lot of compressed information, and the consequence is that in some parts it is not even clear what the authors want to say (see Clarity comments). The proposed idea of the paper seems quite interesting, but on the same time I have some doubts (see Quality comments).


Review for NeurIPS paper: Continual Learning in Low-rank Orthogonal Subspaces

Neural Information Processing Systems

Weaknesses: Despite having a novel core idea, I think this paper is not ready for publication and needs substantial improvement before publication: 1. Currently it seems that you need to know T because projection matrices P_t should be constructed before starting continual learning. This is a huge limitation because the very notion of "continual learning" implies that T is not known a priori because the learning agent supposedly is learning over unlimited time periods (i.e., we may even have T\rightarrow\infty) . Currently, learning task T 1 is going to invalidate your core idea because building an orthogonal P_{T 1} does not seem to be trivial. In my opinion, this constraint should be removed. But I think this is a highly slippery assumption.


Reviews: Infra-slow brain dynamics as a marker for cognitive function and decline

Neural Information Processing Systems

The authors provide a new integrated analysis approach (allowing for simultaneous dimensionality reduction and the possibility of de-noising/artifact correction) to assess slow and infra-slow fluctuations of functional MRI data. They evaluate their approach in a very representative sample and show its potential utility by decoding the task that participants were asked to perform, while being scanned, as well as by predicting behavioral scores from the newly derived latent components as well as clinically-relevant outcomes in a clinical sample. In the following sections, I provide specific feedback with respect to originality, quality, clarity and significance. I hope you will find my comments helpful and constructive. Originality To my knowledge the proposed approach is a novel and innovative way of assessing (task-related or task-free) functional connectivity in the brain in a data-driven manner.