afd
Eight German state premiers back Merz amid speculation over future
The leaders of eight German states have backed the nation's chancellor amid growing speculation about his political future. The premiers - all of whom are part of Friedrich Merz's Christian Democratic Union (CDU) - wrote that restoring trust in their party lay in listening and taking political action, not in debates over personnel. Merz, who has only held the post for 16 months, is facing poor approval ratings while the anti-immigration, far-right Alternative for Germany (AfD) party is surging, prompting questions about whether he is an electoral liability. He has vowed to carry on but two more state elections on Sunday have been seen as a key test of his leadership. Rumours of a Kanzlertausch - or Chancellor swap - have been around for months, but they have increased since the conservative CDU tumbled to a distant second behind the AfD in the eastern state of Saxony-Anhalt at the start of September. The AfD shot up to 43.8% of the vote while the CDU plummeted to 17.2%, marking the first time since World War Two that a far-right party has come close to controlling a German state.
Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding
StepFun, null, :, null, Wang, Bin, Wang, Bojun, Wan, Changyi, Huang, Guanzhe, Hu, Hanpeng, Jia, Haonan, Nie, Hao, Li, Mingliang, Chen, Nuo, Chen, Siyu, Yuan, Song, Xie, Wuxun, Song, Xiaoniu, Chen, Xing, Yang, Xingping, Zhang, Xuelin, Yu, Yanbo, Wang, Yaoyu, Zhu, Yibo, Jiang, Yimin, Zhou, Yu, Lu, Yuanwei, Li, Houyi, Hu, Jingcheng, Lo, Ka Man, Huang, Ailin, Jiao, Binxing, Li, Bo, Chen, Boyu, Miao, Changxin, Lou, Chang, Hu, Chen, Xu, Chen, Yu, Chenfeng, Yao, Chengyuan, Lv, Daokuan, Shi, Dapeng, Sun, Deshan, Huang, Ding, Hu, Dingyuan, Pang, Dongqing, Liu, Enle, Zhang, Fajie, Wan, Fanqi, Yan, Gulin, Zhang, Han, Zhou, Han, Wu, Hanghao, Guo, Hangyu, Chen, Hanqi, Zhang, Hanshan, Wu, Hao, Zhang, Haocheng, Yan, Haolong, Lv, Haoran, Wei, Haoran, Zhou, Hebin, Wang, Heng, Wang, Heng, Li, Hongxin, Zhou, Hongyu, Wang, Hongyuan, Guo, Huiyong, Wang, Jia, Gong, Jiahao, Xie, Jialing, Zhou, Jian, Sun, Jianjian, Wu, Jiaoren, Zhang, Jiaran, Liu, Jiayu, Cheng, Jie, Luo, Jie, Yan, Jie, Yang, Jie, Hou, Jieyi, Zhang, Jinguang, Cao, Jinlan, Yin, Jisheng, Liu, Junfeng, Huang, Junhao, Lin, Junzhe, Tan, Kaijun, Li, Kaixiang, An, Kang, Lin, Kangheng, Liu, Kenkun, Yang, Lei, Zhao, Liang, Chen, Liangyu, Shi, Lieyu, Tan, Liguo, Lin, Lin, Zhang, Lin, Chen, Lina, Huang, Liwen, Shi, Liying, Gu, Longlong, Chen, Mei, Ren, Mengqiang, Li, Ming, Chen, Mingzhe, Wang, Na, Wu, Nan, Han, Qi, Zhao, Qian, Zhang, Qiang, Liu, Qianni, Chen, Qiaohui, Wu, Qiling, He, Qinglin, Tan, Qinyuan, Wang, Qiufeng, Wu, Qiuping, Liang, Qiuyan, Sun, Quan, Li, Rui, Miao, Ruihang, Wan, Ruosi, Guo, Ruyan, Zhong, Shangwu, Pang, Shaoliang, Fan, Shengjie, Shang, Shijie, Jiang, Shilei, Yang, Shiliang, Hao, Shiming, Gao, Shuli, Huang, Siming, Liu, Siqi, Cao, Tiancheng, Cheng, Tianhao, Peng, Tianhao, You, Wang, Ji, Wei, Sun, Wen, Deng, Wenjin, He, Wenqing, Zheng, Wenzhen, Chen, Xi, Kong, Xiangwen, Luo, Xianzhen, Yang, Xiaobo, Liu, Xiaojia, Ren, Xiaoxiao, Han, Xin, Li, Xin, Wu, Xin, Zhao, Xu, Wei, Yanan, Li, Yang, Li, Yangguang, Xu, Yangshijie, Xu, Yanming, Shi, Yaqiang, Shen, Yeqing, Yang, Yi, Yang, Yifei, Gong, Yifeng, Chen, Yihan, Yang, Yijing, Zhang, Yinmin, Zhou, Yizhuang, Ding, Yuanhao, Fan, Yuantao, Yang, Yuanzhen, Luo, Yuchu, Peng, Yue, Lu, Yufan, Deng, Yuhang, Yin, Yuhe, Liu, Yujie, Chen, Yukun, Zhao, Yuling, Mou, Yun, Li, Yunlong, Ju, Yunzhou, Li, Yusheng, Yang, Yuxiang, Zhang, Yuxiang, Chen, Yuyang, Weng, Zejia, Xie, Zhe, Ge, Zheng, Gong, Zheng, Lu, Zhenyi, Huang, Zhewei, Chang, Zhichao, Huang, Zhiguo, Wang, Zhirui, Yang, Zidong, Wang, Zili, Wang, Ziqi, Zhang, Zixin, Jiao, Binxing, Jiang, Daxin, Shum, Heung-Yeung, Zhang, Xiangyu
Large language models (LLMs) face low hardware efficiency during decoding, especially for long-context reasoning tasks. This paper introduces Step-3, a 321B-parameter VLM with hardware-aware model-system co-design optimized for minimizing decoding costs. Step-3 innovates in two key dimensions: (1) A novel Multi-Matrix Factorization Attention (MFA) mechanism that significantly reduces both KV cache size and computation while maintaining high attention expressiveness, and (2) Attention-FFN Disaggregation (AFD), a distributed inference system that decouples attention and Feed-Forward Network (FFN) layers into specialized subsystems. This co-design achieves unprecedented cost efficiency: Step-3 significantly reduces theoretical decoding costs compared with models like DeepSeek-V3 and Qwen3 MoE 235B, with the gains widening at longer context. Step-3 achieves low cost while activating 38B parameters per token (more than DeepSeek-V3 and Qwen3 MoE 235B), demonstrating that hardware-aligned attention arithmetic intensity, MoE sparsity, and AFD are critical to cost-effectiveness. We perform a head-to-head comparison with DeepSeek-V3 in its favorable scenarios. Our implementation on Hopper GPUs achieves a decoding throughput of up to 4,039 tokens per second per GPU under 50ms TPOT SLA (4K context, FP8, no MTP). It is higher than DeepSeek-V3's 2,324 in the same setup and sets a new Pareto frontier for LLM decoding.
Alignment Helps Make the Most of Multimodal Data
Arnold, Christian, Küpfer, Andreas
When studying political communication, combining the information from text, audio, and video signals promises to reflect the richness of human communication more comprehensively than confining it to individual modalities alone. However, its heterogeneity, connectedness, and interaction are challenging to address when modeling such multimodal data. We argue that aligning the respective modalities can be an essential step in entirely using the potential of multimodal data because it informs the model with human understanding. Taking care of the data-generating process of multimodal data, our framework proposes four principles to organize alignment and, thus, address the challenges of multimodal data. We illustrate the utility of these principles by analyzing how German MPs address members of the far-right AfD in their speeches and predicting the tone of video advertising in the context of the 2020 US presidential race. Our paper offers important insights to all keen to analyze multimodal data effectively.
Inverse-RLignment: Inverse Reinforcement Learning from Demonstrations for LLM Alignment
Sun, Hao, van der Schaar, Mihaela
Aligning Large Language Models (LLMs) is crucial for enhancing their safety and utility. However, existing methods, primarily based on preference datasets, face challenges such as noisy labels, high annotation costs, and privacy concerns. In this work, we introduce Alignment from Demonstrations (AfD), a novel approach leveraging high-quality demonstration data to overcome these challenges. We formalize AfD within a sequential decision-making framework, highlighting its unique challenge of missing reward signals. Drawing insights from forward and inverse reinforcement learning, we introduce divergence minimization objectives for AfD. Analytically, we elucidate the mass-covering and mode-seeking behaviors of various approaches, explaining when and why certain methods are superior. Practically, we propose a computationally efficient algorithm that extrapolates over a tailored reward model for AfD. We validate our key insights through experiments on the Harmless and Helpful tasks, demonstrating their strong empirical performance while maintaining simplicity.
Mitigating Feature Gap for Adversarial Robustness by Feature Disentanglement
Zhou, Nuoyan, Zhou, Dawei, Liu, Decheng, Gao, Xinbo, Wang, Nannan
Deep neural networks are vulnerable to adversarial samples. Adversarial fine-tuning methods aim to enhance adversarial robustness through fine-tuning the naturally pre-trained model in an adversarial training manner. However, we identify that some latent features of adversarial samples are confused by adversarial perturbation and lead to an unexpectedly increasing gap between features in the last hidden layer of natural and adversarial samples. To address this issue, we propose a disentanglement-based approach to explicitly model and further remove the latent features that cause the feature gap. Specifically, we introduce a feature disentangler to separate out the latent features from the features of the adversarial samples, thereby boosting robustness by eliminating the latent features. Besides, we align features in the pre-trained model with features of adversarial samples in the fine-tuned model, to further benefit from the features from natural samples without confusion. Empirical evaluations on three benchmark datasets demonstrate that our approach surpasses existing adversarial fine-tuning methods and adversarial training baselines.
Fault-Tolerant Offline Multi-Agent Path Planning
Okumura, Keisuke, Tixeuil, Sébastien
We study a novel graph path planning problem for multiple agents that may crash at runtime, and block part of the workspace. In our setting, agents can detect neighboring crashed agents, and change followed paths at runtime. The objective is then to prepare a set of paths and switching rules for each agent, ensuring that all correct agents reach their destinations without collisions or deadlocks, despite unforeseen crashes of other agents. Such planning is attractive to build reliable multi-robot systems. We present problem formalization, theoretical analysis such as computational complexities, and how to solve this offline planning problem.