Goto

Collaborating Authors

 Technology


Scalable Early Childhood Reading Performance Prediction

Neural Information Processing Systems

Models for student reading performance can empower educators and institutions to proactively identify at-risk students, thereby enabling early and tailored instructional interventions. However, there are no suitable publicly available educational datasets for modeling and predicting future reading performance. In this work, we introduce the Enhanced Core Reading Instruction (ECRI) dataset, a novel large-scale longitudinal tabular dataset collected across 44 schools with 6,916 students and 172 teachers. We leverage the dataset to empirically evaluate the ability of state-of-the-art machine learning models to recognize early childhood educational patterns in multivariate and partial measurements. Specifically, we demonstrate a simple self-supervised strategy in which a Multi-Layer Perception (MLP) network is pre-trained over masked inputs to outperform several strong baselines while generalizing over diverse educational settings. To facilitate future developments in precise modeling and responsible use of models for individualized and early intervention strategies, our data and code are available at https://ecri-data.github.io/.


Secret Collusion among AI Agents: Multi-Agent Deception via Steganography

Neural Information Processing Systems

Recent advancements in generative AI suggest the potential for large-scale interaction between autonomous agents and humans across platforms such as the internet. While such interactions could foster productive cooperation, the ability of AI agents to circumvent security oversight raises critical multi-agent security problems, particularly in the form of unintended information sharing or undesirable coordination. In our work, we establish the subfield of secret collusion, a form of multi-agent deception, in which two or more agents employ steganographic methods to conceal the true nature of their interactions, be it communicative or otherwise, from oversight. We propose a formal threat model for AI agents communicating steganographically and derive rigorous theoretical insights about the capacity and incentives of large language models (LLMs) to perform secret collusion, in addition to the limitations of threat mitigation measures. We complement our findings with empirical evaluations demonstrating rising steganographic capabilities in frontier single and multi-agent LLM setups and examining potential scenarios where collusion may emerge, revealing limitations in countermeasures such as monitoring, paraphrasing, and parameter optimization. Our work is the first to formalize and investigate secret collusion among frontier foundation models, identifying it as a critical area in AI Safety and outlining a comprehensive research agenda to mitigate future risks of collusion between generative AI systems.


Graph Edit Distance with General Costs Using Neural Set Divergence

Neural Information Processing Systems

Graph Edit Distance (GED) measures the (dis-)similarity between two given graphs in terms of the minimum-cost edit sequence, which transforms one graph to the other.GED is related to other notions of graph similarity, such as graph and subgraph isomorphism, maximum common subgraph, etc. However, the computation of exact GED is NP-Hard, which has recently motivated the design of neural models for GED estimation.However, they do not explicitly account for edit operations with different costs. In response, we propose $\texttt{GraphEdX}$, a neural GED estimator that can work with general costs specified for the four edit operations, viz., edge deletion, edge addition, node deletion, and node addition.We first present GED as a quadratic assignment problem (QAP) that incorporates these four costs.Then, we represent each graph as a set of node and edge embeddings and use them to design a family of neural set divergence surrogates. We replace the QAP terms corresponding to each operation with their surrogates. Computing such neural set divergence requires aligning nodes and edges of the two graphs.We learn these alignments using a Gumbel-Sinkhorn permutation generator, additionally ensuring that the node and edge alignments are consistent with each other. Moreover, these alignments are cognizant of both the presence and absence of edges between node pairs.Through extensive experiments on several datasets, along with a variety of edit cost settings, we show that $\texttt{GraphEdX}$ consistently outperforms state-of-the-art methods and heuristics in terms of prediction error.


Self-Play Fine-tuning of Diffusion Models for Text-to-image Generation

Neural Information Processing Systems

Fine-tuning Diffusion Models remains an underexplored frontier in generative artificial intelligence (GenAI), especially when compared with the remarkable progress made in fine-tuning Large Language Models (LLMs). While cutting-edge diffusion models such as Stable Diffusion (SD) and SDXL rely on supervised fine-tuning, their performance inevitably plateaus after seeing a certain volume of data. Recently, reinforcement learning (RL) has been employed to fine-tune diffusion models with human preference data, but it requires at least two images ( loser'' images) for each text prompt.In this paper, we introduce an innovative technique called self-play fine-tuning for diffusion models (SPIN-Diffusion), where the diffusion model engages in competition with its earlier versions, facilitating an iterative self-improvement process. Our approach offers an alternative to conventional supervised fine-tuning and RL strategies, significantly improving both model performance and alignment. Our experiments on the Pick-a-Pic dataset reveal that SPIN-Diffusion outperforms the existing supervised fine-tuning method in aspects of human preference alignment and visual appeal right from its first iteration. By the second iteration, it exceeds the performance of RLHF-based methods across all metrics, achieving these results with less data.


On Affine Homotopy between Language Encoders

Neural Information Processing Systems

Pre-trained language encoders---functions that represent text as vectors---are an integral component of many NLP tasks. We tackle a natural question in language encoder analysis: What does it mean for two encoders to be similar? We contend that a faithful measure of similarity needs to be \emph{intrinsic}, that is, task-independent, yet still be informative of \emph{extrinsic} similarity---the performance on downstream tasks. It is common to consider two encoders similar if they are \emph{homotopic}, i.e., if they can be aligned through some transformation. In this spirit, we study the properties of \emph{affine} alignment of language encoders and its implications on extrinsic similarity. We find that while affine alignment is fundamentally an asymmetric notion of similarity, it is still informative of extrinsic similarity. We confirm this on datasets of natural language representations. Beyond providing useful bounds on extrinsic similarity, affine intrinsic similarity also allows us to begin uncovering the structure of the space of pre-trained encoders by defining an order over them.


Revealed: The LEAST scenic places in the UK, according to science - including a spot in the usually picturesque Cornwall

Daily Mail - Science & tech

Trump administration'unlocks' 140MILLION barrels of precious Iranian oil with major policy change to fight back against'hoarding' China... here's what it means for your wallet Buffy the Vampire Slayer star Nicholas Brendon dead at 54 as'heartbroken' family reveal cause of death Joseph Duggar's wife Kendra is arrested for allegedly endangering welfare of a minor as he faces new charges Behind closed doors, the Duggar family's next nightmare began long before Joseph's arrest: Insiders reveal what they knew and how they plan to recover America is about to be torn apart by a financial tsunami - and it's not just an oil crisis to fear. However, it seems not every corner of Britain is quite so beautiful - as a survey has revealed the least scenic locations. Voters on the Scenic Or Not survey awarded the top spot to Basingstoke's Newbury Road. This unappealing location received the lowest possible score, with just one out of 10 for'scenicness'. And while Cornwall might be renowned for its beautiful scenery, a rather less attractive part of the county - the Electricity Station in Landulph - joins Basingstoke at the bottom of the pile.


Drone strike near Iraqi intelligence headquarters in Baghdad kills officer

Al Jazeera

Will Gulf states join war? One police officer has been killed in a drone strike by "outlaw groups" on the headquarters of the Iraqi National Intelligence Service in the heart of capital Baghdad. "A drone targeted the headquarters of the Iraqi National Intelligence Service in the Mansour district" at about 10am local time (07:00 GMT), General Saad Maan, head of the Iraqi government's security media unit, said in a brief statement on Saturday. Another drone, filming the operation, crashed into a private members ' sports club popular with the Iraqi elite and foreign diplomats, according to the same source. The drone attack on the headquarters of the National Intelligence Service came hours after another attack on the US military complex.


Russian drone attack kills two in Ukraine ahead of talks in US, officials say

BBC News

Two people were killed in a Russian drone attack on a home in the Ukrainian city of Zaporizhzhia, local authorities say. Two children, 11 and 15, were also injured in the attack which took place on the eve of new talks between Ukrainian and American negotiators in the US. Negotiations on ending the war have been on hold since the start of the latest conflict in Iran. President Volodymy Zelensky wants his negotiators to discuss the US decision to ease sanctions on Russian oil - implemented by Washington to help keep down global energy prices. Talks mediated by the US have so far failed to stop the fighting in Ukraine or change Russia's demands, and there is little hope of a breakthrough.


OmniJARVIS: Unified Vision-Language-Action Tokenization Enables Open-World Instruction Following Agents

Neural Information Processing Systems

This paper presents OmniJARVIS, a novel Vision-Language-Action (VLA) model for open-world instruction-following agents in Minecraft. Compared to prior works that either emit textual goals to separate controllers or produce the control command directly, OmniJARVIS seeks a different path to ensure both strong reasoning and efficient decision-making capabilities via unified tokenization of multimodal interaction data. First, we introduce a self-supervised approach to learn a behavior encoder that produces discretized tokens for behavior trajectories $\tau = \{o_0, a_0, \dots\}$ and an imitation learning policy decoder conditioned on these tokens. These additional behavior tokens will be augmented to the vocabulary of pretrained Multimodal Language Models. With this encoder, we then pack long-term multimodal interactions involving task instructions, memories, thoughts, observations, textual responses, behavior trajectories, etc into unified token sequences and model them with autoregressive transformers. Thanks to the semantically meaningful behavior tokens, the resulting VLA model, OmniJARVIS, can reason (by producing chain-of-thoughts), plan, answer questions, and act (by producing behavior tokens for the imitation learning policy decoder). OmniJARVIS demonstrates excellent performances on a comprehensive collection of atomic, programmatic, and open-ended tasks in open-world Minecraft. Our analysis further unveils the crucial design principles in interaction data formation, unified tokenization, and its scaling potentials. The dataset, models, and code will be released at https://craftjarvis.org/OmniJARVIS.


Physical Consistency Bridges Heterogeneous Data in Molecular Multi-Task Learning

Neural Information Processing Systems

In recent years, machine learning has demonstrated impressive capability in handling molecular science tasks. To support various molecular properties at scale, machine learning models are trained in the multi-task learning paradigm. Nevertheless, data of different molecular properties are often not aligned: some quantities, e.g.