Goto

Collaborating Authors

 Instructional Material


Persistence diagrams of random matrices via Morse theory: universality and a new spectral diagnostic

arXiv.org Machine Learning

We prove that the persistence diagram of the sublevel set filtration of the quadratic form f(x) = x^T M x restricted to the unit sphere S^{n-1} is analytically determined by the eigenvalues of the symmetric matrix M. By Morse theory, the diagram has exactly n-1 finite bars, with the k-th bar living in homological dimension k-1 and having length equal to the k-th eigenvalue spacing s_k = λ_{k+1} - λ_k. This identification transfers random matrix theory (RMT) universality to persistence diagram universality: for matrices drawn from the Gaussian Orthogonal Ensemble (GOE), we derive the closed-form persistence entropy PE = log(8n/π) - 1, and verify numerically that the coefficient of variation of persistence statistics decays as n^{-0.6}. Different random matrix ensembles (GOE, GUE, Wishart) produce distinct universal persistence diagrams, providing topological fingerprints of RMT universality classes. As a practical consequence, we show that persistence entropy outperforms the standard level spacing ratio \langle r \rangle for discriminating GOE from GUE matrices (AUC 0.978 vs. 0.952 at n = 100, non-overlapping bootstrap 95% CIs), and detects global spectral perturbations in the Rosenzweig-Porter model to which \langle r \rangle is blind. These results establish persistence entropy as a new spectral diagnostic that captures complementary information to existing RMT tools.


Neurally-Guided Procedural Models: Amortized Inference for Procedural Graphics Programs using Neural Networks

Neural Information Processing Systems

Probabilistic inference algorithms such as Sequential Monte Carlo (SMC) provide powerful tools for constraining procedural models in computer graphics, but they require many samples to produce desirable results. In this paper, we show how to create procedural models which learn how to satisfy constraints. We augment procedural models with neural networks which control how the model makes random choices based on the output it has generated thus far. We call such models neurally-guided procedural models. As a pre-computation, we train these models to maximize the likelihood of example outputs generated via SMC. They are then used as efficient SMC importance samplers, generating high-quality results with very few samples. We evaluate our method on L-system-like models with imagebased constraints. Given a desired quality threshold, neurally-guided models can generate satisfactory results up to 10x faster than unguided models.


Avoiding Imposters and Delinquents: Adversarial Crowdsourcing and Peer Prediction

Neural Information Processing Systems

We consider a crowdsourcing model in which nworkers are asked to rate the quality of nitems previously generated by other workers. An unknown set of αnworkers generate reliable ratings, while the remaining workers may behave arbitrarily and possibly adversarially. The manager of the experiment can also manually evaluate the quality of a small number of items, and wishes to curate together almost all of the high-quality items with at most anfraction of low-quality items.


Simple and Effective Masked Diffusion Language Models

Neural Information Processing Systems

While diffusion models excel at generating high-quality images, prior work reports a significant performance gap between diffusion and autoregressive (AR) methods in language modeling.In this work, we show that simple masked discrete diffusion is more performant than previously thought.We apply an effective training recipe that improves the performance of masked diffusion models and derive a simplified, Rao-Blackwellized objective that results in additional improvements.Our objective has a simple form--it is a mixture of classical masked language modeling losses--and can be used to train encoder-only language models that admit efficient samplers, including ones that can generate arbitrary lengths of text semi-autoregressively like a traditional language model.On language modeling benchmarks, a range of masked diffusion models trained with modern engineering practices achieves a new state-of-the-art among diffusion models, and approaches AR perplexity. We provide the code, along with a blog post and video tutorial on the project page: https://s-sahoo.com/mdlm


IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos

Neural Information Processing Systems

Shape assembly is a ubiquitous task in daily life, integral for constructing complex 3D structures like IKEA furniture. While significant progress has been made in developing autonomous agents for shape assembly, existing datasets have not yet tackled the 4D grounding of assembly instructions in videos, essential for a holistic understanding of assembly in 3D space over time. We introduce IKEA Video Manuals, a dataset that features 3D models of furniture parts, instructional manuals, assembly videos from the Internet, and most importantly, annotations of dense spatio-temporal alignments between these data modalities. To demonstrate the utility of IKEA Video Manuals, we present five applications essential for shape assembly: assembly plan generation, part-conditioned segmentation, part-conditioned pose estimation, video object segmentation, and furniture assembly based on instructional video manuals. For each application, we provide evaluation metrics and baseline methods. Through experiments on our annotated data, we highlight many challenges in grounding assembly instructions in videos to improve shape assembly, including handling occlusions, varying viewpoints, and extended assembly sequences.


F-OAL: Forward-only Online Analytic Learning with Fast Training and Low Memory Footprint in Class Incremental Learning

Neural Information Processing Systems

Online Class Incremental Learning (OCIL) aims to train models incrementally, where data arrive in mini-batches, and previous data are not accessible. A major challenge in OCIL is Catastrophic Forgetting, i.e., the loss of previously learned knowledge. Among existing baselines, replay-based methods show competitive results but requires extra memory for storing exemplars, while exemplar-free (i.e., data need not be stored for replay in production) methods are resource friendly but often lack accuracy. In this paper, we propose an exemplar-free approach--Forward-only Online Analytic Learning (F-OAL). Unlike traditional methods, F-OAL does not rely on back-propagation and is forward-only, significantly reducing memory usage and computational time. Cooperating with a pre-trained frozen encoder with Feature Fusion, F-OAL only needs to update a linear classifier by recursive least square. This approach simultaneously achieves high accuracy and low resource consumption. Extensive experiments on bench mark datasets demonstrate F-OAL's robust performance in OCIL scenarios.


Overcoming Core Engineering Barriers in Humanoid Robotics Development

IEEE Spectrum Robotics

Register now free-of-charge to explore this white paper This Whitepaper offers engineers and researchers a technical examination of the key design barriers in humanoid robotics and the component-level strategies emerging to address them, from sensing and motion control to power systems and thermal management. What you will learn about:   The core engineering challenges — complex motion control, safe human-robot interaction, and hardware cost constraints — that currently limit practical humanoid robot deployment. Sensing system architectures: how IMUs, gyroscopes, accelerometers, tactile sensors, and AMR magnetic sensors support real-time posture estimation, perception fusion, and environmental awareness. Motion and actuation design considerations including actuator-level power delivery, motor noise mitigation, PCB bend-stress resistance, and dexterous hand integration. Power and thermal system trade-offs: battery chemistry selection (LFP vs. NCA), BMS design, DC/DC converter topologies, and thermistor-based protection for operational reliability. Click 'LOOK INSIDE' to Download Now.


Gamified math. Video read-alouds. Why parents are saying no to screens in class

Los Angeles Times

Things to Do in L.A. Kate Brody's 7-year-old son plays at home in North Hollywood on March 14. This is read by an automated voice. Please report any issues or inconsistencies here . Early childhood experts say excessive screen time displaces hands-on learning and peer interaction critical to development. At least 11 states have considered legislation limiting technology in the classroom this year.


Virtual Class Enhanced Discriminative Embedding Learning

Neural Information Processing Systems

Recently, learning discriminative features to improve the recognition performances gradually becomes the primary goal of deep learning, and numerous remarkable works have emerged. In this paper, we propose a novel yet extremely simple method Virtual Softmax to enhance the discriminative property of learned features by injecting a dynamic virtual negative class into the original softmax. Injecting virtual class aims to enlarge inter-class margin and compress intra-class distribution by strengthening the decision boundary constraint. Although it seems weird to optimize with this additional virtual class, we show that our method derives from an intuitive and clear motivation, and it indeed encourages the features to be more compact and separable. This paper empirically and experimentally demonstrates the superiority of Virtual Softmax, improving the performances on a variety of object classification and face verification tasks.


AI is nearly exclusively designed by men – here's how to fix it

New Scientist

AI is nearly exclusively designed by men - here's how to fix it With the Trump administration's attacks on so-called woke AI it is becoming even harder to make the technology we use fairer and more diverse. It's day two of the conference at the Royal Society in London, but I'm finding it increasingly hard to concentrate on the speakers because my AI transcription software - which is supposed to make my life easier - keeps insisting on mistyping someone's name. The irony isn't lost on me: this is the session about artificial intelligence, and specifically about how women are being erased from the latest AI technologies. This is much bigger than the now-familiar idea that AI algorithms carry the biases of the datasets they are trained on, including gender bias. Instead, the focus of the conference session, chaired by computer scientist Wendy Hall, is seeking to address a more fundamental issue: the fact that new AI technologies, which will have a transformative effect on all of society, are being designed almost exclusively by men.