Deep Learning
Are You There God? Lightweight Narrative Annotation of Christian Fiction with LMs
Hicke, Rebecca M. M., Haggard, Brian W., Ferrante, Mia, Khanna, Rayhan, Mimno, David
In addition to its more widely studied cultural movements, American Evangelicalism has a well-developed but less externally visible literary side. Christian Fiction, however, has been little studied, and what scholarly attention there is has focused on the explosively popular Left Behind series. In this work, we use computational tools to provide both a broad topical overview of Christian Fiction as a genre and a more directed exploration of how its authors depict divine acts. Working with human annotators, we first developed a codebook for identifying "acts of God." We then adapted the codebook for use by a recent, lightweight LM with the assistance of a much larger model. The laptop-scale LM is largely capable of matching human annotations, even when the task is subtle and challenging. Using these annotations, we show that significant and meaningful differences exist between divine acts depicted by the Left Behind books and Christian Fiction more broadly.
Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
Qin, Yulu, Varghese, Dheeraj, Lindström, Adam Dahlgren, Donatelli, Lucia, Misra, Kanishka, Kim, Najoung
Does vision-and-language (VL) training change the linguistic representations of language models in meaningful ways? Most results in the literature have shown inconsistent or marginal differences, both behaviorally and representationally. In this work, we start from the hypothesis that the domain in which VL training could have a significant effect is lexical-conceptual knowledge, in particular its taxonomic organization. Through comparing minimal pairs of text-only LMs and their VL-trained counterparts, we first show that the VL models often outperform their text-only counterparts on a text-only question-answering task that requires taxonomic understanding of concepts mentioned in the questions. Using an array of targeted behavioral and representational analyses, we show that the LMs and VLMs do not differ significantly in terms of their taxonomic knowledge itself, but they differ in how they represent questions that contain concepts in a taxonomic relation vs. a non-taxonomic relation. This implies that the taxonomic knowledge itself does not change substantially through additional VL training, but VL training does improve the deployment of this knowledge in the context of a specific task, even when the presentation of the task is purely linguistic.
Learning World Models for Interactive Video Generation
Chen, Taiye, Hu, Xun, Ding, Zihan, Jin, Chi
Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities due to two main challenges: compounding errors and insufficient memory mechanisms. We enhance image-to-video models with interactive capabilities through additional action conditioning and autoregressive framework, and reveal that compounding error is inherently irreducible in autoregressive video generation, while insufficient memory mechanism leads to incoherence of world models. We propose video retrieval augmented generation (VRAG) with explicit global state conditioning, which significantly reduces long-term compounding errors and increases spatiotemporal consistency of world models. In contrast, naive autoregressive generation with extended context windows and retrieval-augmented generation prove less effective for video generation, primarily due to the limited in-context learning capabilities of current video models. Our work illuminates the fundamental challenges in video world models and establishes a comprehensive benchmark for improving video generation models with internal world modeling capabilities.
Deep sequence models tend to memorize geometrically; it is unclear why
Noroozizadeh, Shahriar, Nagarajan, Vaishnavh, Rosenfeld, Elan, Kumar, Sanjiv
In sequence modeling, the parametric memory of atomic facts has been predominantly abstracted as a brute-force lookup of co-occurrences between entities. We contrast this associative view against a geometric view of how memory is stored. We begin by isolating a clean and analyzable instance of Transformer reasoning that is incompatible with memory as strictly a storage of the local co-occurrences specified during training. Instead, the model must have somehow synthesized its own geometry of atomic facts, encoding global relationships between all entities, including non-co-occurring ones. This in turn has simplified a hard reasoning task involving an $\ell$-fold composition into an easy-to-learn 1-step geometric task. From this phenomenon, we extract fundamental aspects of neural embedding geometries that are hard to explain. We argue that the rise of such a geometry, despite optimizing over mere local associations, cannot be straightforwardly attributed to typical architectural or optimizational pressures. Counterintuitively, an elegant geometry is learned even when it is not more succinct than a brute-force lookup of associations. Then, by analyzing a connection to Node2Vec, we demonstrate how the geometry stems from a spectral bias that -- in contrast to prevailing theories -- indeed arises naturally despite the lack of various pressures. This analysis also points to practitioners a visible headroom to make Transformer memory more strongly geometric. We hope the geometric view of parametric memory encourages revisiting the default intuitions that guide researchers in areas like knowledge acquisition, capacity, discovery and unlearning.
Uncertainty-Aware Diagnostics for Physics-Informed Machine Learning
Daniels, Mara, Hodgkinson, Liam, Mahoney, Michael
Physics-informed machine learning (PIML) integrates prior physical information, often in the form of differential equation constraints, into the process of fitting machine learning models to physical data. Popular PIML approaches, including neural operators, physics-informed neural networks, neural ordinary differential equations, and neural discrete equilibria, are typically fit to objectives that simultaneously include both data and physical constraints. However, the multi-objective nature of this approach creates ambiguity in the measurement of model quality. This is related to a poor understanding of epistemic uncertainty, and it can lead to surprising failure modes, even when existing statistical metrics suggest strong fits. Working within a Gaussian process regression framework, we introduce the Physics-Informed Log Evidence (PILE) score. Bypassing the ambiguities of test losses, the PILE score is a single, uncertainty-aware metric that provides a selection principle for hyperparameters of a PIML model. We show that PILE minimization yields excellent choices for a wide variety of model parameters, including kernel bandwidth, least squares regularization weights, and even kernel function selection. We also show that, even prior to data acquisition, a special 'data-free' case of the PILE score identifies a priori kernel choices that are 'well-adapted' to a given PDE. Beyond the kernel setting, we anticipate that the PILE score can be extended to PIML at large, and we outline approaches to do so.
Here's How the AI Crash Happens
The U.S. is becoming an Nvidia-state. Listen to more stories on the Noa app. The AI boom is visible from orbit. Satellite photos of New Carlisle, Indiana, show greenish splotches of farmland transformed into unmistakable industrial parks in less than a year's time. There are seven rectangular data centers there, with 23 more on the way.
WIRED Roundup: AI Psychosis, Missing FTC Files, and Google Bedbugs
In this episode of, we run through the top stories of the week and look closely at people's complaints to the FTC alleging that ChatGPT led them or loved ones into AI psychosis. In today's episode, Zoë Schiffer is joined by senior editor Louise Matsakis to run through five stories that you need to know about this week--from how SEO is changing in the era of AI to how frogs became a protest symbol. Then, Zoë and Louise dive into why some people have been filing complaints to the FTC about ChatGPT, arguing it has led them to AI psychosis. People Who Say They're Experiencing AI Psychosis Beg the FTC for Help The FTC Is Disappearing Blog Posts About AI Published During Lina Khan's Tenure Write to us at uncannyvalley@wired.com . You can always listen to this week's podcast through the audio player on this page, but if you want to subscribe for free to get every episode, here's how: If you're on an iPhone or iPad, open the app called Podcasts, or just tap this link . Today on the show, we're bringing you five stories that you need to know about this week. And later, we'll dive into our main story about how several people have filed complaints to the FTC claiming OpenAI's ChatGPT led them or people they love into supposed AI psychosis. I'm joined today by WIRED's senior business editor, Louise Matsakis. It's great to be here. So Louise, our first story this week is actually one that we worked on together, part of our ongoing collaboration with Model Behavior, and it's all about how this holiday season, more shoppers are expected to use chatbots to figure out what to buy.
OpenAI thought to be preparing for 1tn stock market float
A float would support Sam Altman's ambitions to splash trillions of dollars on building datacentres. A float would support Sam Altman's ambitions to splash trillions of dollars on building datacentres. OpenAI is reportedly gearing up for a stock market listing valuing the company at $1tn (£760bn) as soon as next year, in what would be one of the biggest ever initial public offerings. The developer behind the hit AI chatbot ChatGPT is considering whether to file for an IPO as soon as the second half of 2026, according to Reuters, which cited people familiar with the matter. The company is thought to be looking to raise at least $60bn.
Copilot AI's latest trick? A secure sandbox for its agentic activity
When you purchase through links in our articles, we may earn a small commission. Microsoft 365 users can now test Researcher with Computer Use, an autonomous agent that can access files that it couldn't before. Microsoft Copilot is tapping a key feature from Windows 11 Pro to enable Copilot's AI to dig even further than it already has. It's part of an update to Microsoft 365 Copilot called Researcher with Computer Use, debuting today for a limited subset of Microsoft 365 Copilot users. LLMs that engage in deep research, like Copilot, face a problem: some content is locked away behind an authentication process, like requiring a password.
The Download: Introducing: the new conspiracy age
Everything is a conspiracy theory now. Conspiracists are all over the White House, turning fringe ideas into dangerous policy. America's institutions are crumbling under the weight of deep suspicion and the lasting effects of covid isolation. Online echo chambers are getting harder to escape, and generative AI is altering the fabric of truth. A mix of technology and politics has given an unprecedented boost to once-fringe ideas--but they are pretty much the same fantasies that have been spreading for hundreds of years. MIT Technology Review helps break down how this moment is changing science and technology--and how we can make it through.