Media
HistoryBankQA: Multilingual Temporal Question Answering on Historical Events
Mandal, Biswadip, Khandelwal, Anant, Gupta, Manish
Temporal reasoning about historical events is a critical skill for NLP tasks like event extraction, historical entity linking, temporal question answering, timeline summarization, temporal event clustering and temporal natural language inference. Yet efforts on benchmarking temporal reasoning capabilities of large language models (LLMs) are rather limited. Existing temporal reasoning datasets are limited in scale, lack multilingual coverage and focus more on contemporary events. To address these limitations, we present HistoryBank, a multilingual database of 10M+ historical events extracted from Wikipedia timeline pages and article infoboxes. Our database provides unprecedented coverage in both historical depth and linguistic breadth with 10 languages. Additionally, we construct a comprehensive question answering benchmark for temporal reasoning across all languages. This benchmark covers a diverse set of 6 temporal QA reasoning tasks, and we evaluate a suite of popular language models (LLaMA-3-8B, Mistral-7B, Gemma-2-9b, Qwen3-8B, GPT4o) to assess their performance on these tasks. As expected GPT4o performs best across all answer types and languages; Gemma-2 outperforms the other small language models. Our work aims to provide a comprehensive resource for advancing multilingual and temporally-aware natural language understanding of historical events. To facilitate further research, we will make our code and datasets publicly available upon acceptance of this paper.
Joint AoI and Handover Optimization in Space-Air-Ground Integrated Network
Lang, Zifan, Liu, Guixia, Sun, Geng, Li, Jiahui, Wang, Jiacheng, Yuan, Weijie, Niyato, Dusit, Kim, Dong In
Despite the widespread deployment of terrestrial networks, providing reliable communication services to remote areas and maintaining connectivity during emergencies remains challenging. Low Earth orbit (LEO) satellite constellations offer promising solutions with their global coverage capabilities and reduced latency, yet struggle with intermittent coverage and limited communication windows due to orbital dynamics. This paper introduces an age of information (AoI)-aware space-air-ground integrated network (SAGIN) architecture that leverages a high-altitude platform (HAP) as intelligent relay between the LEO satellites and ground terminals. Our three-layer design employs hybrid free-space optical (FSO) links for high-capacity satellite-to-HAP communication and reliable radio frequency (RF) links for HAP-to-ground transmission, and thus addressing the temporal discontinuity in LEO satellite coverage while serving diverse user priorities. Specifically, we formulate a joint optimization problem to simultaneously minimize the AoI and satellite handover frequency through optimal transmit power distribution and satellite selection decisions. This highly dynamic, non-convex problem with time-coupled constraints presents significant computational challenges for traditional approaches. To address these difficulties, we propose a novel diffusion model (DM)-enhanced dueling double deep Q-network with action decomposition and state transformer encoder (DD3QN-AS) algorithm that incorporates transformer-based temporal feature extraction and employs a DM-based latent prompt generative module to refine state-action representations through conditional denoising. Simulation results highlight the superior performance of the proposed approach compared with policy-based methods and some other deep reinforcement learning (DRL) benchmarks.
LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
Vujanic, Robin, Rueckstiess, Thomas
We present LEAF ("Lightweight Embedding Alignment Framework"), a knowledge distillation framework for text embedding models. A key distinguishing feature is that our distilled leaf models are aligned to their teacher. In the context of information retrieval, this allows for flexible asymmetric architectures where documents are encoded with the larger teacher model, while queries can be served with the smaller leaf models. We also show that leaf models automatically inherit MRL and robustness to output quantization whenever these properties are present in the teacher model, without explicitly training for them. To demonstrate the capability of our framework we publish leaf-ir, a 23M parameters information retrieval oriented text embedding model trained using LEAF, which sets a new state-of-the-art (SOTA) on BEIR, ranking #1 on the public leaderboard for this benchmark and for models of its size. When run in asymmetric mode, its retrieval performance is further increased. Our scheme is however not restricted to the information retrieval setting, and we demonstrate its wider applicability by synthesizing the multi-task leaf-mt model. This also sets a new SOTA, ranking #1 on the public MTEB v2 (English) leaderboard for its size. LEAF is applicable to black-box models and in contrast to other embedding model training frameworks, it does not require judgments nor hard negatives, and training can be conducted using small batch sizes. Thus, dataset and training infrastructure requirements for our framework are modest. We make our models publicly available under a permissive Apache 2.0 license.
A Traditional Approach to Symbolic Piano Continuation
Zhou-Zheng, Christian, Backsund, John, Chan, Dun Li, Coventry, Alex, Eslami, Avid, Goel, Jyotin, Han, Xingwen, Soomro, Danysh, Wei, Galen
Recent developments in sequence modeling have allowed continuation to be viewed as an autore-gressive task, to be modeled with a suitable tokenization scheme and a powerful sequence model like the ubiquitous Transformer [1]. A nonexhaustive list of prior work in this vein includes the Music Transformer [2], Museformer [3], FIGARO [4], and MuseCoco [5]. Most research in symbolic music modeling has so far focused on generalizing these techniques to--and improving performance on--long-sequence, multitrack, multi-instrument, and/or text-or attribute-controllable generative tasks. Typically, specialized techniques must be developed for these foundation models to handle these harder tasks, such as fine-and coarse-grained attention for long sequences [3], and text feature extraction techniques [4] and attribute augmentation [5] for controllability.
New Kid in the Classroom: Exploring Student Perceptions of AI Coding Assistants
The arrival of AI coding assistants in educational settings presents a paradigm shift, introducing a "new kid in the classroom" for both students and instructors. Thus, understanding the perceptions of these key actors about this new dynamic is critical. This exploratory study contributes to this area by investigating how these tools are shaping the experiences of novice programmers in an introductory programming course. Through a two-part exam, we investigated student perceptions by first providing access to AI support for a programming task and then requiring an extension of the solution without it. We collected Likert-scale and open-ended responses from 20 students to understand their perceptions on the challenges they faced. Our findings reveal that students perceived AI tools as helpful for grasping code concepts and boosting their confidence during the initial development phase. However, a noticeable difficulty emerged when students were asked to work unaided, pointing to potential overreliance and gaps in foundational knowledge transfer. These insights highlight a critical need for new pedagogical approaches that integrate AI effectively while effectively enhancing core programming skills, rather than impersonating them.
TAPS: Tool-Augmented Personalisation via Structured Tagging
Taktasheva, Ekaterina, Dalton, Jeff
Recent advancements in tool-augmented large language models have enabled them to interact with external tools, enhancing their ability to perform complex user tasks. However, existing approaches overlook the role of personalisation in guiding tool use. This work investigates how user preferences can be effectively integrated into goal-oriented dialogue agents. Through extensive analysis, we identify key weaknesses in the ability of LLMs to personalise tool use. To this end, we introduce TAPS, a novel solution that enhances personalised tool use by leveraging a structured tagging tool and an uncertainty-based tool detector. TAPS significantly improves the ability of LLMs to incorporate user preferences, achieving the new state-of-the-art for open source models on the NLSI task.
De-risking investment in AI agents
AI agents thrive when trust is designed in from the start, says vice president of product management at NICE, Neeraj Verma. Automation has become a defining force in the customer experience. Between the chatbots that answer our questions and the recommendation systems that shape our choices, AI-driven tools are now embedded in nearly every interaction. But the latest wave of so-called "agentic AI"--systems that can plan, act, and adapt toward a defined goal--promises to push automation even further. Every single person that I've spoken to has at least spoken to some sort of GenAI bot on their phones. They expect experiences to be not scripted.
Matthew Prince Wants AI Companies to Pay for Their Sins
The Cloudflare CEO joined to talk about standing up to content scraping, the internet's potential futures, and his company's relationship to Trump. Matthew Prince may not be a household name, but the world most certainly knows his work. Prince is the cofounder and CEO of Cloudflare . Launched in 2010, the internet infrastructure company has found itself increasingly in the position of serving as the web's bodyguard. It filters out bad traffic, keeps sites safe, and stops them from crashing when too many people visit. Its tools defend against DDoS attacks. In 2017, Cloudflare made headlines when it dropped white supremacist site The Daily Stormer . Cloudflare's severing of ties with The Daily Stormer marked a momentous shift, one that came after years of claiming a neutral stance. Prince continues to evolve the way Cloudflare works. In July, the company rolled out a new tool tasked with blocking unauthorized AI scraping. It effectively creates a pay-per-crawl model requiring AI platforms to shell out money if they want access to a site's content. On this episode of, I talked to Prince about publishing, the old internet, and how his ideal version of the future web means that OpenAI just might become the Netflix of content. KATIE DRUMMOND: Good to have you here, Matthew. You should have been warned ahead of time, but you probably weren't.