South America
VesselGPT: Autoregressive Modeling of Vascular Geometry
Feldman, Paula, Sinnona, Martin, Delrieux, Claudio, Siless, Viviana, Iarussi, Emmanuel
Anatomical trees are critical for clinical diagnosis and treatment planning, yet their complex and diverse geometry make accurate representation a significant challenge. Motivated by the latest advances in large language models, we introduce an autoregressive method for synthesizing anatomical trees. Our approach first embeds vessel structures into a learned discrete vocabulary using a VQ-VAE architecture, then models their generation autoregressively with a GPT-2 model. This method effectively captures intricate geometries and branching patterns, enabling realistic vascular tree synthesis. Comprehensive qualitative and quantitative evaluations reveal that our technique achieves high-fidelity tree reconstruction with compact discrete representations. Moreover, our B-spline representation of vessel cross-sections preserves critical morphological details that are often overlooked in previous' methods parameterizations. To the best of our knowledge, this work is the first to generate blood vessels in an autoregressive manner. Code is available at https://github.com/LIA-DiTella/VesselGPT-MICCAI.
Four killed in Kyiv in new Russian aerial attack
Four killed in Kyiv in new Russian aerial attack 12 minutes agoShareSaveJaroslav LukivBBC NewsShareSaveUkraine's emergencies service DSNSRescuers from Ukraine's emergencies service DSNS tackle fire in a residential building destroyed in the latest Russian attack on Kyiv At least four people have been killed in an overnight Russian missile and drone attack on Ukraine's capital Kyiv, the interior minister says. In a post on social media, Ihor Klymenko says residential areas, hospitals and sports infrastructure were hit. "An entire section of a residential high-rise building was destroyed" in the worst-hit Shevchenkivskyi district, he says, adding that some people are trapped under the rubble. In the Kyiv region, a woman was killed and another two people injured in the Russian aerial attack, regional head Mykola Kalashnyk says. The Russian military has not commented on the issue.
Tower+: Bridging Generality and Translation Specialization in Multilingual LLMs
Rei, Ricardo, Guerreiro, Nuno M., Pombal, José, Alves, João, Teixeirinha, Pedro, Farajian, Amin, Martins, André F. T.
Fine-tuning pretrained LLMs has been shown to be an effective strategy for reaching state-of-the-art performance on specific tasks like machine translation. However, this process of adaptation often implies sacrificing general-purpose capabilities, such as conversational reasoning and instruction-following, hampering the utility of the system in real-world applications that require a mixture of skills. In this paper, we introduce Tower+, a suite of models designed to deliver strong performance across both translation and multilingual general-purpose text capabilities. We achieve a Pareto frontier between translation specialization and multilingual general-purpose capabilities by introducing a novel training recipe that builds on Tower (Alves et al., 2024), comprising continued pretraining, supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. At each stage of training, we carefully generate and curate data to strengthen performance on translation as well as general-purpose tasks involving code generation, mathematics problem solving, and general instruction-following. We develop models at multiple scales: 2B, 9B, and 72B. Our smaller models often outperform larger general-purpose open-weight and proprietary LLMs (e.g., Llama 3.3 70B, GPT-4o). Our largest model delivers best-in-class translation performance for high-resource languages and top results in multilingual Arena Hard evaluations and in IF-MT, a benchmark we introduce for evaluating both translation and instruction-following. Our findings highlight that it is possible to rival frontier models in general capabilities, while optimizing for specific business domains, such as translation and localization.
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
Wang, Fei, Wan, Xingchen, Sun, Ruoxi, Chen, Jiefeng, Arık, Sercan Ö.
Inference-time scaling has proven effective in boosting large language model (LLM) performance through increased test-time computation. Yet, its practical application is often hindered by reliance on external verifiers or a lack of optimization for realistic computational constraints. We propose DynScaling, which addresses these limitations through two primary innovations: an integrated parallel-sequential sampling strategy and a bandit-based dynamic budget allocation framework. The integrated sampling strategy unifies parallel and sequential sampling by constructing synthetic sequential reasoning chains from initially independent parallel responses, promoting diverse and coherent reasoning trajectories. The dynamic budget allocation framework formulates the allocation of computational resources as a multi-armed bandit problem, adaptively distributing the inference budget across queries based on the uncertainty of previously sampled responses, thereby maximizing computational efficiency. By combining these components, DynScaling effectively improves LLM performance under practical resource constraints without the need for external verifiers. Experimental results demonstrate that DynScaling consistently surpasses existing verifier-free inference scaling baselines in both task performance and computational cost.
On Path to Multimodal Historical Reasoning: HistBench and HistAgent
Qiu, Jiahao, Xiao, Fulian, Wang, Yimin, Mao, Yuchen, Chen, Yijia, Juan, Xinzhe, Zhang, Shu, Wang, Siran, Qi, Xuan, Zhang, Tongcheng, Yao, Zixin, Guo, Jiacheng, Lu, Yifu, Argon, Charles, Cui, Jundi, Chen, Daixin, Zhou, Junran, Zhou, Shuyao, Zhou, Zhanpeng, Yang, Ling, Liu, Shilong, Wang, Hongru, Huang, Kaixuan, Jiang, Xun, Cao, Yuming, Chen, Yue, Chen, Yunfei, Chen, Zhengyi, Dai, Ruowei, Deng, Mengqiu, Fu, Jiye, Gu, Yunting, Guan, Zijie, Huang, Zirui, Ji, Xiaoyan, Jiang, Yumeng, Kong, Delong, Li, Haolong, Li, Jiaqi, Li, Ruipeng, Li, Tianze, Li, Zhuoran, Lian, Haixia, Lin, Mengyue, Liu, Xudong, Lu, Jiayi, Lu, Jinghan, Luo, Wanyu, Luo, Ziyue, Pu, Zihao, Qiao, Zhi, Ren, Ruihuan, Wan, Liang, Wang, Ruixiang, Wang, Tianhui, Wang, Yang, Wang, Zeyu, Wang, Zihua, Wu, Yujia, Wu, Zhaoyi, Xin, Hao, Xing, Weiao, Xiong, Ruojun, Xu, Weijie, Shu, Yao, Xiao, Yao, Yang, Xiaorui, Yang, Yuchen, Yi, Nan, Yu, Jiadong, Yu, Yangyuxuan, Zeng, Huiting, Zhang, Danni, Zhang, Yunjie, Zhang, Zhaoyu, Zhang, Zhiheng, Zheng, Xiaofeng, Zhou, Peirong, Zhong, Linyan, Zong, Xiaoyin, Zhao, Ying, Chen, Zhenxin, Ding, Lin, Gao, Xiaoyu, Gong, Bingbing, Li, Yichao, Liao, Yang, Ma, Guang, Ma, Tianyuan, Sun, Xinrui, Wang, Tianyi, Xia, Han, Xian, Ruobing, Ye, Gen, Yu, Tengfei, Zhang, Wentao, Wang, Yuxi, Gao, Xi, Wang, Mengdi
Recent advances in large language models (LLMs) have led to remarkable progress across domains, yet their capabilities in the humanities, particularly history, remain underexplored. Historical reasoning poses unique challenges for AI, involving multimodal source interpretation, temporal inference, and cross-linguistic analysis. While general-purpose agents perform well on many existing benchmarks, they lack the domain-specific expertise required to engage with historical materials and questions. To address this gap, we introduce HistBench, a new benchmark of 414 high-quality questions designed to evaluate AI's capacity for historical reasoning and authored by more than 40 expert contributors. The tasks span a wide range of historical problems-from factual retrieval based on primary sources to interpretive analysis of manuscripts and images, to interdisciplinary challenges involving archaeology, linguistics, or cultural history. Furthermore, the benchmark dataset spans 29 ancient and modern languages and covers a wide range of historical periods and world regions. Finding the poor performance of LLMs and other agents on HistBench, we further present HistAgent, a history-specific agent equipped with carefully designed tools for OCR, translation, archival search, and image understanding in History. On HistBench, HistAgent based on GPT-4o achieves an accuracy of 27.54% pass@1 and 36.47% pass@2, significantly outperforming LLMs with online search and generalist agents, including GPT-4o (18.60%), DeepSeek-R1(14.49%) and Open Deep Research-smolagents(20.29% pass@1 and 25.12% pass@2). These results highlight the limitations of existing LLMs and generalist agents and demonstrate the advantages of HistAgent for historical reasoning.
Large Language Models as Psychological Simulators: A Methodological Guide
Large language models (LLMs) offer emerging opportunities for psychological and behavioral research, but methodological guidance is lacking. This article provides a framework for using LLMs as psychological simulators across two primary applications: simulating roles and personas to explore diverse contexts, and serving as computational models to investigate cognitive processes. For simulation, we present methods for developing psychologically grounded personas that move beyond demographic categories, with strategies for validation against human data and use cases ranging from studying inaccessible populations to prototyping research instruments. For cognitive modeling, we synthesize emerging approaches for probing internal representations, methodological advances in causal interventions, and strategies for relating model behavior to human cognition. We address overarching challenges including prompt sensitivity, temporal limitations from training data cutoffs, and ethical considerations that extend beyond traditional human subjects review. Throughout, we emphasize the need for transparency about model capabilities and constraints. Together, this framework integrates emerging empirical evidence about LLM performance--including systematic biases, cultural limitations, and prompt brittleness--to help researchers wrangle these challenges and leverage the unique capabilities of LLMs in psychological research.
CP$^2$: Leveraging Geometry for Conformal Prediction via Canonicalization
van der Linden, Putri A., Timans, Alexander, Bekkers, Erik J.
We study the problem of conformal prediction (CP) under geometric data shifts, where data samples are susceptible to transformations such as rotations or flips. While CP endows prediction models with post-hoc uncertainty quantification and formal coverage guarantees, their practicality breaks under distribution shifts that deteriorate model performance. To address this issue, we propose integrating geometric information--such as geometric pose--into the conformal procedure to reinstate its guarantees and ensure robustness under geometric shifts. In particular, we explore recent advancements on pose canonicalization as a suitable information extractor for this purpose. Evaluating the combined approach across discrete and continuous shifts and against equivariant and augmentation-based baselines, we find that integrating geometric information with CP yields a principled way to address geometric shifts while maintaining broad applicability to black-box predictors.
ContextBench: Modifying Contexts for Targeted Latent Activation
Graham, Robert, Stevinson, Edward, Richter, Leo, Chia, Alexander, Miller, Joseph, Bloom, Joseph Isaac
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of generating targeted, linguistically fluent inputs that activate specific latent features or elicit model behaviours. We formalise this approach as context modification and present ContextBench -- a benchmark with tasks assessing core method capabilities and potential safety applications. Our evaluation framework measures both elicitation strength (activation of latent features or behaviours) and linguistic fluency, highlighting how current state-of-the-art methods struggle to balance these objectives. We enhance Evolutionary Prompt Optimisation (EPO) with LLM-assistance and diffusion model inpainting, and demonstrate that these variants achieve state-of-the-art performance in balancing elicitation effectiveness and fluency.
Identifiability of Deep Polynomial Neural Networks
Usevich, Konstantin, Dérand, Clara, Borsoi, Ricardo, Clausel, Marianne
Polynomial Neural Networks (PNNs) possess a rich algebraic and geometric structure. However, their identifiability -- a key property for ensuring interpretability -- remains poorly understood. In this work, we present a comprehensive analysis of the identifiability of deep PNNs, including architectures with and without bias terms. Our results reveal an intricate interplay between activation degrees and layer widths in achieving identifiability. As special cases, we show that architectures with non-increasing layer widths are generically identifiable under mild conditions, while encoder-decoder networks are identifiable when the decoder widths do not grow too rapidly. Our proofs are constructive and center on a connection between deep PNNs and low-rank tensor decompositions, and Kruskal-type uniqueness theorems. This yields both generic conditions determined by the architecture, and effective conditions that depend on the network's parameters. We also settle an open conjecture on the expected dimension of PNN's neurovarieties, and provide new bounds on the activation degrees required for it to reach its maximum.
SlepNet: Spectral Subgraph Representation Learning for Neural Dynamics
Viswanath, Siddharth, Singh, Rahul, Zhang, Yanlei, Noah, J. Adam, Hirsch, Joy, Krishnaswamy, Smita
Graph neural networks have been useful in machine learning on graph-structured data, particularly for node classification and some types of graph classification tasks. However, they have had limited use in representing patterning of signals over graphs. Patterning of signals over graphs and in subgraphs carries important information in many domains including neuroscience. Neural signals are spatiotemporally patterned, high dimensional and difficult to decode. Graph signal processing and associated GCN models utilize the graph Fourier transform and are unable to efficiently represent spatially or spectrally localized signal patterning on graphs. Wavelet transforms have shown promise here, but offer non-canonical representations and cannot be tightly confined to subgraphs. Here we propose SlepNet, a novel GCN architecture that uses Slepian bases rather than graph Fourier harmonics. In SlepNet, the Slepian harmonics optimally concentrate signal energy on specifically relevant subgraphs that are automatically learned with a mask. Thus, they can produce canonical and highly resolved representations of neural activity, focusing energy of harmonics on areas of the brain which are activated. We evaluated SlepNet across three fMRI datasets, spanning cognitive and visual tasks, and two traffic dynamics datasets, comparing its performance against conventional GNNs and graph signal processing constructs. SlepNet outperforms the baselines in all datasets. Moreover, the extracted representations of signal patterns from SlepNet offers more resolution in distinguishing between similar patterns, and thus represent brain signaling transients as informative trajectories. Here we have shown that these extracted trajectory representations can be used for other downstream untrained tasks. Thus we establish that SlepNet is useful both for prediction and representation learning in spatiotemporal data.