Large Language Model
Scaling Decentralized Learning with FLock
Cheng, Zehua, Sun, Rui, Sun, Jiahao, Guo, Yike
Fine-tuning the large language models (LLMs) are prevented by the deficiency of centralized control and the massive computing and communication overhead on the decentralized schemes. While the typical standard federated learning (FL) supports data privacy, the central server requirement creates a single point of attack and vulnerability to poisoning attacks. Generalizing the result in this direction to 70B-parameter models in the heterogeneous, trustless environments has turned out to be a huge, yet unbroken bottleneck. This paper introduces FLock, a decentralized framework for secure and efficient collaborative LLM fine-tuning. Integrating a blockchain-based trust layer with economic incentives, FLock replaces the central aggregator with a secure, auditable protocol for cooperation among untrusted parties. We present the first empirical validation of fine-tuning a 70B LLM in a secure, multi-domain, decentralized setting. Our experiments show the FLock framework defends against backdoor poisoning attacks that compromise standard FL optimizers and fosters synergistic knowledge transfer. The resulting models show a >68% reduction in adversarial attack success rates. The global model also demonstrates superior cross-domain generalization, outperforming models trained in isolation on their own specialized data.
Doc2Chart: Intent-Driven Zero-Shot Chart Generation from Documents
Jain, Akriti, Ramu, Pritika, Garimella, Aparna, Saxena, Apoorv
Large Language Models (LLMs) have demonstrated strong capabilities in transforming text descriptions or tables to data visualizations via instruction-tuning methods. However, it is not straightforward to apply these methods directly for a more real-world use case of visualizing data from long documents based on user-given intents, as opposed to the user pre-selecting the relevant content manually. We introduce the task of intent-based chart generation from documents: given a user-specified intent and document(s), the goal is to generate a chart adhering to the intent and grounded on the document(s) in a zero-shot setting. We propose an unsupervised, two-staged framework in which an LLM first extracts relevant information from the document(s) by decomposing the intent and iteratively validates and refines this data. Next, a heuristic-guided module selects an appropriate chart type before final code generation. To assess the data accuracy of the generated charts, we propose an attribution-based metric that uses a structured textual representation of charts, instead of relying on visual decoding metrics that often fail to capture the chart data effectively. To validate our approach, we curate a dataset comprising of 1,242 $<$intent, document, charts$>$ tuples from two domains, finance and scientific, in contrast to the existing datasets that are largely limited to parallel text descriptions/ tables and their corresponding charts. We compare our approach with baselines using single-shot chart generation using LLMs and query-based retrieval methods; our method outperforms by upto $9$ points and $17$ points in terms of chart data accuracy and chart type respectively over the best baselines.
Apple Intelligence Foundation Language Models: Tech Report 2025
Li, Ethan, Larsen, Anders Boesen Lindbo, Zhang, Chen, Zhou, Xiyou, Qin, Jun, Yap, Dian Ang, Raghavan, Narendran, Chang, Xuankai, Bowler, Margit, Yildiz, Eray, Peebles, John, Coleman, Hannah Gillis, Ronchi, Matteo, Gray, Peter, You, Keen, Spalvieri-Kruse, Anthony, Pang, Ruoming, Li, Reed, Yang, Yuli, Soroush, Emad, Lu, Zhiyun, Xiao, Crystal, Situ, Rong, Huffaker, Jordan, Griffiths, David, Ahmed, Zaid, Zhang, Peng, Parilla, Daniel, Liberman, Asaf, Mallalieu, Jennifer, Mazaheri, Parsa, Chen, Qibin, Bilkhu, Manjot, Zhang, Aonan, Wang, Eric, Nelson, Dave, FitzMaurice, Michael, Voice, Thomas, Liu, Jeremy, Shaffer, Josh, Zhao, Shiwen, Yadla, Prasanth, Rasteh, Farzin, Guo, Pengsheng, Farooq, Arsalan, Snow, Jeremy, Murphy, Stephen, Lei, Tao, Cho, Minsik, Horrell, George, Dodge, Sam, Hislop, Lindsay, Singh, Sumeet, Dombrowski, Alex, Raghavan, Aiswarya, Sirovica, Sasha, Saebi, Mandana, Lao, Faye, Lam, Max, Lu, TJ, Xu, Zhaoyang, Singh, Karanjeet, Kirchner, Marc, Mizrahi, David, Arora, Rajat, Zhang, Haotian, Mason, Henry, Zhou, Lawrence, Hua, Yi, Jain, Ankur, Bai, Felix, Astrauskas, Joseph, Weers, Floris, Gardner, Josh, Chiang, Mira, Zhang, Yi, Agrawal, Pulkit, Sun, Tony, Keunebroek, Quentin, Hopkins, Matthew, Wu, Bugu, Jia, Tao, Chen, Chen, Zhou, Xingyu, Wang, Nanzhu, Liu, Peng, Hou, Ruixuan, Rauch, Rene, Gao, Yuan, Dehghan, Afshin, Janke, Jonathan, Wang, Zirui, Chen, Cha, Ren, Xiaoyi, Nan, Feng, Elman, Josh, Yin, Dong, Goren, Yusuf, Lai, Jeff, Fei, Yiran, Evans, Syd, Yu, Muyang, Yin, Guoli, Qin, Yi, Feldman, Erin, Garg, Isha, Rajamani, Aparna, Vega, Karla, Cheng, Walker, Collins, TJ, Han, Hans, Menacho, Raul Rea, Yeung, Simon, Lee, Sophy, Mutyala, Phani, Cheng, Ying-Chang, Gan, Zhe, Chu, Sprite, Lazarow, Justin, Pappalardo, Alessandro, Scozzafava, Federico, Lu, Jing, Daxberger, Erik, Duchesne, Laurent, Liu, Jen, Gรผera, David, Ligas, Stefano, Kery, Mary Beth, Ramerth, Brent, Sannino, Ciro, Eichner, Marcin, Huang, Haoshuo, Qian, Rui, Schwarzer-Becker, Moritz, Riazati, David, Gao, Mingfei, Wang, Bailin, Cackler, Jack, Lu, Yang, Niu, Ransen, Dennison, John, Klein, Guillaume, Bigham, Jeffrey, Gopinath, Deepak, Shiee, Navid, Botten, Darren, Tartavel, Guillaume, Garcia, Alex Guillen, Xu, Sam, Haladjian, Victoria MรถnchJuan, Dou, Zi-Yi, Paulik, Matthias, Mendez, Adolfo Lopez, Li, Zhen, Chen, Hong-You, Jia, Chao, Doshi, Dhaval, Zhang, Zhengdong, Manjani, Raunak, Franklin, Aaron, Ren, Zhile, Chen, David, Peshko, Artsiom, Raghuram, Nandhitha, Hao, Hans, Shan, Jiulong, Nerella, Kavya, Tantawi, Ramsey, Kumar, Vivek, Wang, Saiwen, Wershing, Brycen, Dhingra, Bhuwan, Shah, Dhruti, Adaranijo, Ob, Zheng, Xin, Madsen, Tait, Kotek, Hadas, Liu, Chang, Xia, Yin, Li, Hanli, Jayaram, Suma, Sun, Yanchao, Fakhry, Ahmed, Saveris, Vasileios, Withers, Dustin, Li, Yanghao, Aygar, Alp, Teran, Andres Romero Mier Y, Huang, Kaiwei, Lee, Mark, Li, Xiujun, Li, Yuhong, Johnson, Tyler, Tang, Jay, Cheng, Joseph Yitan, Peng, Futang, Walkingshaw, Andrew, Guibert, Lucas, Sharma, Abhishek, Shen, Cheng, Maj, Piotr, Tanaka, Yasutaka, Jhang, You-Cyuan, Ma, Vivian, Vehvilainen, Tommi, Zou, Kelvin, Nichols, Jeff, Lei, Matthew, Qiu, David, Qian, Yihao, Santhanam, Gokul, Wu, Wentao, Han, Yena, Moritz, Dominik, Fu, Haijing, Xu, Mingze, Rathod, Vivek, Liu, Jian, D'hauwe, Louis, Ba, Qin, Sun, Haitian, Yan, Haoran, Dufter, Philipp, Nguyen, Anh, Feng, Yihao, Wang, Emma, He, Keyu, Nair, Rahul, Shah, Sanskruti, Lu, Jiarui, Sonnenberg, Patrick, Warner, Jeremy, Li, Yuanzhi, Pan, Bowen, Zhong, Ziyi, Zhou, Joe, Davarnia, Sam, Saarikivi, Olli, Belousova, Irina, Burger, Rachel, Wu, Shang-Chen, Feng, Di, Straathof, Bas, Chou, James, Zhang, Yuanyang, Zuliani, Marco, Jimenez, Eduardo, Sundararajan, Abhishek, Du, Xianzhi, Lan, Chang, Shahdadpuri, Nilesh, Grasch, Peter, Sima, Sergiu, Newnham, Josh, Paidi, Varsha, Wang, Jianyu, Haag, Kaelen, Braunstein, Alex, Molinari, Daniele, Wei, Richard, Yang, Brenda, Lusskin, Nicholas, Arreaza-Taylor, Joanna, Cao, Meng, Seidl, Nicholas, Wang, Simon, Hu, Jiaming, Ma, Yiping, Li, Mengyu, Liu, Kieran, Su, Hang, Ravi, Sachin, Wang, Chong, Wang, Xin, Smith, Kevin, You, Haoxuan, Karimzadeh, Binazir, Li, Rui, Lei, Jinhao, Fang, Wei, Doane, Alec, Wiseman, Sam, Fernandez, Ismael, Li, Jane, Hansen, Andrew, Movellan, Javier, Neubauer, Christopher, Zhou, Hanzhi, Chaney, Chris, Kamaldin, Nazir, Wolf, Valentin, Bermรบdez-Medina, Fernando, Pelemans, Joris, Fu, Peter, Xing, Howard, Kong, Xiang, Shan, Wayne, Jacoby-Cooper, Gabriel, Shen, Dongcai, Gunter, Tom, Seguin, Guillaume, Shi, Fangping, Li, Shiyu, Xu, Yang, Kamal, Areeba, Masi, Dan, Guha, Saptarshi, Zhu, Qi, Thibodeau, Jenna, Zhang, Changyuan, Callahan, Rebecca, Maalouf, Charles, Tsao, Wilson, Li, Boyue, Cao, Qingqing, Sabo, Naomy, Leong, Cheng, Wang, Yi, Anupama, Anupama Mann, Reed, Colorado, Jung, Kenneth, Chen, Zhifeng, Moorthy, Mohana Prasad Sathya, He, Yifei, Hornberger, Erik, Krishna, Devi, Tong, Senyu, Michael, null, Lee, null, Haldimann, David, Zhao, Yang, Zhang, Bowen, Gao, Chang, Bartels, Chris, Rao, Sushma, Tran, Nathalie, Lehnerer, Simon, Giang, Co, Dong, Patrick, Pan, Junting, Wang, Biyao, Li, Dongxu, Farajtabar, Mehrdad, Hwang, Dongseong, Duanmu, Grace, Verma, Eshan, Reddy, Sujeeth, Shan, Qi, Gao, Hongbin, Du, Nan, Sridhar, Pragnya, Huang, Forrest, Wang, Yingbo, Bhendawade, Nikhil, Zhu, Diane, Aitharaju, Sai, Hohman, Fred, Gardiner, Lauren, Chiu, Chung-Cheng, Yang, Yinfei, Kokmen, Alper, Chu, Frank, Ye, Ke, Elgin, Kaan, Levy, Oron, Park, John, Zhang, Donald, Schoop, Eldon, Wenzel, Nina, Booker, Michael, Kim, Hyunjik, Erdenebileg, Chinguun, Dun, Nan, Yang, Eric Liang, Chhatrapati, Priyal, Mahtani, Vishaal, Gang, Haiming, Chia, Kohen, Seshadri, Deepa, Yu, Donghan, Meng, Yan, Peterson, Kelsey, Yang, Zhen, Wang, Yongqiang, Peng, Carina, Kang, Doug, Agarwal, Anuva, Antony, Albert, Tebar, Juan Lao, Jose, Albin Madappally, Poston, Regan, De Wang, Andy, Casamayor, Gerard, Amirloo, Elmira, Yao, Violet, Kryscinski, Wojciech, Duan, Kun, L, Lezhi
We introduce two multilingual, multimodal foundation language models that power Apple Intelligence features across Apple devices and services: i a 3B-parameter on-device model optimized for Apple silicon through architectural innovations such as KV-cache sharing and 2-bit quantization-aware training; and ii a scalable server model built on a novel Parallel-Track Mixture-of-Experts PT-MoE transformer that combines track parallelism, mixture-of-experts sparse computation, and interleaved global-local attention to deliver high quality with competitive cost on Apple's Private Cloud Compute platform. Both models are trained on large-scale multilingual and multimodal datasets sourced via responsible web crawling, licensed corpora, and high-quality synthetic data, then further refined with supervised fine-tuning and reinforcement learning on a new asynchronous platform. The resulting models support several additional languages while understanding images and executing tool calls. In public benchmarks and human evaluations, both the server model and the on-device model match or surpass comparably sized open baselines. A new Swift-centric Foundation Models framework exposes guided generation, constrained tool calling, and LoRA adapter fine-tuning, allowing developers to integrate these capabilities with a few lines of code. The latest advancements in Apple Intelligence models are grounded in our Responsible AI approach with safeguards like content filtering and locale-specific evaluation, as well as our commitment to protecting our users' privacy with innovations like Private Cloud Compute.
PyVision: Agentic Vision with Dynamic Tooling
Zhao, Shitian, Zhang, Haoquan, Lin, Shaoheng, Li, Ming, Wu, Qilong, Zhang, Kaipeng, Wei, Chen
LLMs are increasingly deployed as agents, systems capable of planning, reasoning, and dynamically calling external tools. However, in visual reasoning, prior approaches largely remain limited by predefined workflows and static toolsets. In this report, we present PyVision, an interactive, multi-turn framework that enables MLLMs to autonomously generate, execute, and refine Python-based tools tailored to the task at hand, unlocking flexible and interpretable problem-solving. We develop a taxonomy of the tools created by PyVision and analyze their usage across a diverse set of benchmarks. Quantitatively, PyVision achieves consistent performance gains, boosting GPT-4.1 by +7.8% on V* and Claude-4.0-Sonnet by +31.1% on VLMsAreBlind-mini. These results point to a broader shift: dynamic tooling allows models not just to use tools, but to invent them, advancing toward more agentic visual reasoning.
Refining Czech GEC: Insights from a Multi-Experiment Approach
Pechman, Petr, Straka, Milan, Strakovรก, Jana, Nรกplava, Jakub
We present a grammar error correction (GEC) system that achieves state of the art for the Czech language. Our system is based on a neural network translation approach with the Transformer architecture, and its key feature is its real-time synthetic generation pipeline, which dynamically augments sentences with artificial errors by introducing both language-agnostic and Czech-specific errors. We conduct a comprehensive series of experiments, investigating the Czech GEC corpora as bases for synthetic error introduction, several error generation strategies, domain balancing, tokenization granularity, model size, and data scaling during fine-tuning. Additionally, we evaluate the performance of large language models (LLMs) on Czech GEC in both end-user and expert fine-tuning scenarios. Our best-performing model is superior both in performance and computational efficiency. The source code and the trained model links are available on https://github.com/ufal/tsd2025-gec.
MEraser: An Effective Fingerprint Erasure Approach for Large Language Models
Zhang, Jingxuan, Xu, Zhenhua, Hu, Rui, Xing, Wenpeng, Zhang, Xuhong, Han, Meng
Large Language Models (LLMs) have become increasingly prevalent across various sectors, raising critical concerns about model ownership and intellectual property protection. Although backdoor-based fingerprinting has emerged as a promising solution for model authentication, effective attacks for removing these fingerprints remain largely unexplored. Therefore, we present Mismatched Eraser (MEraser), a novel method for effectively removing backdoor-based fingerprints from LLMs while maintaining model performance. Our approach leverages a two-phase fine-tuning strategy utilizing carefully constructed mismatched and clean datasets. Through extensive evaluation across multiple LLM architectures and fingerprinting methods, we demonstrate that MEraser achieves complete fingerprinting removal while maintaining model performance with minimal training data of fewer than 1,000 samples. Furthermore, we introduce a transferable erasure mechanism that enables effective fingerprinting removal across different models without repeated training. In conclusion, our approach provides a practical solution for fingerprinting removal in LLMs, reveals critical vulnerabilities in current fingerprinting techniques, and establishes comprehensive evaluation benchmarks for developing more resilient model protection methods in the future.
CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval
Geng, Jiahui, Cai, Fengyu, Cui, Shaobo, Li, Qing, Chen, Liangwei, Lyu, Chenyang, Li, Haonan, Zhu, Derui, Pretschner, Walter, Koeppl, Heinz, Karray, Fakhri
Code retrieval is essential in modern software development, as it boosts code reuse and accelerates debugging. However, current benchmarks primarily emphasize functional relevance while neglecting critical dimensions of software quality. Motivated by this gap, we introduce CoQuIR, the first large-scale, multilingual benchmark specifically designed to evaluate quality-aware code retrieval across four key dimensions: correctness, efficiency, security, and maintainability. CoQuIR provides fine-grained quality annotations for 42,725 queries and 134,907 code snippets in 11 programming languages, and is accompanied by two quality-centric evaluation metrics: Pairwise Preference Accuracy and Margin-based Ranking Score. Using CoQuIR, we benchmark 23 retrieval models, covering both open-source and proprietary systems, and find that even top-performing models frequently fail to distinguish buggy or insecure code from their more robust counterparts. Furthermore, we conduct preliminary investigations into training methods that explicitly encourage retrievers to recognize code quality. Using synthetic datasets, we demonstrate promising improvements in quality-aware metrics across various models, without sacrificing semantic relevance. Downstream code generation experiments further validate the effectiveness of our approach. Overall, our work highlights the importance of integrating quality signals into code retrieval systems, laying the groundwork for more trustworthy and robust software development tools.
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
Wang, Qianli, Nguyen, Van Bach, Feldhus, Nils, Villa-Arenas, Luis Felipe, Seifert, Christin, Mรถller, Sebastian, Schmitt, Vera
Counterfactual examples are widely employed to enhance the performance and robustness of large language models (LLMs) through counterfactual data augmentation (CDA). However, the selection of the judge model used to evaluate label flipping, the primary metric for assessing the validity of generated counterfactuals for CDA, yields inconsistent results. To decipher this, we define four types of relationships between the counterfactual generator and judge models: being the same model, belonging to the same model family, being independent models, and having an distillation relationship. Through extensive experiments involving two state-of-the-art LLM-based methods, three datasets, four generator models, and 15 judge models, complemented by a user study (n = 90), we demonstrate that judge models with an independent, non-fine-tuned relationship to the generator model provide the most reliable label flipping evaluations. Relationships between the generator and judge models, which are closely aligned with the user study for CDA, result in better model performance and robustness. Nevertheless, we find that the gap between the most effective judge models and the results obtained from the user study remains considerably large. This suggests that a fully automated pipeline for CDA may be inadequate and requires human intervention.
Convert Language Model into a Value-based Strategic Planner
Wang, Xiaoyu, Zhao, Yue, Gu, Qingqing, Jiang, Zhonglin, Chen, Xiaokai, Chen, Yong, Ji, Luo
Emotional support conversation (ESC) aims to alleviate the emotional distress of individuals through effective conversations. Although large language models (LLMs) have obtained remarkable progress on ESC, most of these studies might not define the diagram from the state model perspective, therefore providing a suboptimal solution for long-term satisfaction. To address such an issue, we leverage the Q-learning on LLMs, and propose a framework called straQ*. Our framework allows a plug-and-play LLM to bootstrap the planning during ESC, determine the optimal strategy based on long-term returns, and finally guide the LLM to response. Substantial experiments on ESC datasets suggest that straQ* outperforms many baselines, including direct inference, self-refine, chain of thought, finetuning, and finite state machines.
GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data
Deng, Shengliang, Yan, Mi, Wei, Songlin, Ma, Haixin, Yang, Yuxin, Chen, Jiayi, Zhang, Zhiqi, Yang, Taoyu, Zhang, Xuheng, Zhang, Wenhao, Cui, Heming, Zhang, Zhizheng, Wang, He
Embodied foundation models are gaining increasing attention for their zero-shot generalization, scalability, and adaptability to new tasks through few-shot post-training. However, existing models rely heavily on real-world data, which is costly and labor-intensive to collect. Synthetic data offers a cost-effective alternative, yet its potential remains largely underexplored. To bridge this gap, we explore the feasibility of training Vision-Language-Action models entirely with large-scale synthetic action data. We curate SynGrasp-1B, a billion-frame robotic grasping dataset generated in simulation with photorealistic rendering and extensive domain randomization. Building on this, we present GraspVLA, a VLA model pretrained on large-scale synthetic action data as a foundational model for grasping tasks. GraspVLA integrates autoregressive perception tasks and flow-matching-based action generation into a unified Chain-of-Thought process, enabling joint training on synthetic action data and Internet semantics data. This design helps mitigate sim-to-real gaps and facilitates the transfer of learned actions to a broader range of Internet-covered objects, achieving open-vocabulary generalization in grasping. Extensive evaluations across real-world and simulation benchmarks demonstrate GraspVLA's advanced zero-shot generalizability and few-shot adaptability to specific human preferences. We will release SynGrasp-1B dataset and pre-trained weights to benefit the community.