Goto

Collaborating Authors

 Industry


TETRIS: TilE-matching the TRemendous Irregular Sparsity

Neural Information Processing Systems

Compressing neural networks by pruning weights with small magnitudes can significantly reduce the computation and storage cost. Although pruning makes the model smaller, it is difficult to get a practical speedup in modern computing platforms such as CPU and GPU due to the irregularity.


Roundtables: Surviving the New Age of Conspiracies

MIT Technology Review

Watch a subscriber-only conversation unpacking our new series, "The New Conspiracy Age," and how this moment is changing science and technology. Everything is a conspiracy theory now. Watch a discussion with our editors and Mike Rothschild, journalist and conspiracy theory expert, about how we can make sense of them all. What it's like to be in the middle of a conspiracy theory (according to a conspiracy theory expert) It's surprisingly easy to stumble into a relationship with an AI chatbot Rhiannon Williams OpenAI's new LLM exposes the secrets of how AI really works Will Douglas Heaven It's surprisingly easy to stumble into a relationship with an AI chatbot The idea that machines will be as smart as--or smarter than--humans has hijacked an entire industry. But look closely and you'll see it's a myth that persists for many of the same reasons conspiracies do. The experimental model won't compete with the biggest and best, but it could tell us why they behave in weird ways--and how trustworthy they really are.







The committee machine: Computational to statistical gaps in learning a two-layers neural network

Neural Information Processing Systems

Heuristic tools from statistical physics have been used in the past to locate the phase transitions and compute the optimal learning and generalization errors in the teacher-student scenario in multi-layer neural networks. In this contribution, we provide a rigorous justification of these approaches for a two-layers neural network model called the committee machine.



Horizon-Independent Minimax Linear Regression

Neural Information Processing Systems

We consider online linear regression: at each round, an adversary reveals a covariate vector, the learner predicts a real value, the adversary reveals a label, and the learner suffers the squared prediction error. The aim is to minimize the difference between the cumulative loss and that of the linear predictor that is best in hindsight.