Government
Towards a Design Guideline for RPA Evaluation: A Survey of Large Language Model-Based Role-Playing Agents
Chen, Chaoran, Yao, Bingsheng, Zou, Ruishi, Hua, Wenyue, Lyu, Weimin, Li, Toby Jia-Jun, Wang, Dakuo
Role-Playing Agent (RPA) is an increasingly popular type of LLM Agent that simulates human-like behaviors in a variety of tasks. However, evaluating RPAs is challenging due to diverse task requirements and agent designs. This paper proposes an evidence-based, actionable, and generalizable evaluation design guideline for LLM-based RPA by systematically reviewing 1,676 papers published between Jan. 2021 and Dec. 2024. Our analysis identifies six agent attributes, seven task attributes, and seven evaluation metrics from existing literature. Based on these findings, we present an RPA evaluation design guideline to help researchers develop more systematic and consistent evaluation methods.
Whose story is it? Personalizing story generation by inferring author styles
Kumar, Nischal Ashok, Pham, Chau Minh, Iyyer, Mohit, Lan, Andrew
Personalization has become essential for improving user experience in interactive writing and educational applications, yet its potential in story generation remains largely unexplored. In this work, we propose a novel two-stage pipeline for personalized story generation. Our approach first infers an author's implicit story-writing characteristics from their past work and organizes them into an Author Writing Sheet, inspired by narrative theory. The second stage uses this sheet to simulate the author's persona through tailored persona descriptions and personalized story writing rules. To enable and validate our approach, we construct Mythos, a dataset of 590 stories from 64 authors across five distinct sources that reflect diverse story-writing settings. A head-to-head comparison with a non-personalized baseline demonstrates our pipeline's effectiveness in generating high-quality personalized stories. Our personalized stories achieve a 75 percent win rate (versus 14 percent for the baseline and 11 percent ties) in capturing authors' writing style based on their past works. Human evaluation highlights the high quality of our Author Writing Sheet and provides valuable insights into the personalized story generation task. Notable takeaways are that writings from certain sources, such as Reddit, are easier to personalize than others, like AO3, while narrative aspects, like Creativity and Language Use, are easier to personalize than others, like Plot.
AI-Assisted Decision Making with Human Learning
Noti, Gali, Donahue, Kate, Kleinberg, Jon, Oren, Sigal
AI systems increasingly support human decision-making. In many cases, despite the algorithm's superior performance, the final decision remains in human hands. For example, an AI may assist doctors in determining which diagnostic tests to run, but the doctor ultimately makes the diagnosis. This paper studies such AI-assisted decision-making settings, where the human learns through repeated interactions with the algorithm. In our framework, the algorithm -- designed to maximize decision accuracy according to its own model -- determines which features the human can consider. The human then makes a prediction based on their own less accurate model. We observe that the discrepancy between the algorithm's model and the human's model creates a fundamental tradeoff. Should the algorithm prioritize recommending more informative features, encouraging the human to recognize their importance, even if it results in less accurate predictions in the short term until learning occurs? Or is it preferable to forgo educating the human and instead select features that align more closely with their existing understanding, minimizing the immediate cost of learning? This tradeoff is shaped by the algorithm's time-discounted objective and the human's learning ability. Our results show that optimal feature selection has a surprisingly clean combinatorial characterization, reducible to a stationary sequence of feature subsets that is tractable to compute. As the algorithm becomes more "patient" or the human's learning improves, the algorithm increasingly selects more informative features, enhancing both prediction accuracy and the human's understanding. Notably, early investment in learning leads to the selection of more informative features than a later investment. We complement our analysis by showing that the impact of errors in the algorithm's knowledge is limited as it does not make the prediction directly.
tn4ml: Tensor Network Training and Customization for Machine Learning
Puljak, Ema, Sanchez-Ramirez, Sergio, Masot-Llima, Sergi, Vallรจs-Muns, Jofre, Garcia-Saez, Artur, Pierini, Maurizio
Tensor Networks have emerged as a prominent alternative to neural networks for addressing Machine Learning challenges in foundational sciences, paving the way for their applications to real-life problems. This paper introduces tn4ml, a novel library designed to seamlessly integrate Tensor Networks into optimization pipelines for Machine Learning tasks. Inspired by existing Machine Learning frameworks, the library offers a user-friendly structure with modules for data embedding, objective function definition, and model training using diverse optimization strategies. We demonstrate its versatility through two examples: supervised learning on tabular data and unsupervised learning on an image dataset. Additionally, we analyze how customizing the parts of the Machine Learning pipeline for Tensor Networks influences performance metrics.
Asymptotic Optimism of Random-Design Linear and Kernel Regression Models
We derived the closed-form asymptotic optimism of linear regression models under random designs, and generalizes it to kernel ridge regression. Using scaled asymptotic optimism as a generic predictive model complexity measure, we studied the fundamental different behaviors of linear regression model, tangent kernel (NTK) regression model and three-layer fully connected neural networks (NN). Our contribution is two-fold: we provided theoretical ground for using scaled optimism as a model predictive complexity measure; and we show empirically that NN with ReLUs behaves differently from kernel models under this measure. With resampling techniques, we can also compute the optimism for regression models with real data.
South Korea pauses downloads of DeepSeek AI over privacy concerns
DeepSeek, the massively popular Chinese AI assistant, has been temporarily unavailable from app stores in South Korea since February 15. A press release from the country's data protection authority, the Personal Information Protection Commission (PIPC), stated that downloads will resume once the Chinese AI company complies with local data protection laws, while those with the app can still use it. DeepSeek is also blocked on South Korean government and military devices. DeepSeek only established a local presence in South Korea on February 10. The company also acknowledged that it didn't fully consider South Korea's data protection laws when launching the service globally.
Protecting Americans' data from China is central to an America First agenda
In January, President Donald Trump announced the 500 billion Stargate project, which will accelerate the buildout of America's digital infrastructure while creating hundreds of thousands of U.S. jobs. This investment in American data centers โ which are key to nearly every aspect of how we live in an increasingly digital world โ will be a game-changer for the U.S. far into the future. It is a reflection of the business community's recognition of the benefits of investing in the United States under President Trump. But more importantly, Stargate shows the commitment of President Trump and those in his administration to finding every opportunity to protect American citizens and our nation's digital sovereignty by combating Chinese aggression. China's pursuit of global technological dominance is a direct assault on our freedoms and sovereignty.
Chatbot vs. national security? Why DeepSeek is raising concerns
Chinese artificial intelligence chatbot DeepSeek upended the global industry and wiped billions off U.S. tech stocks when it unveiled its R1 program, which it claims was built on cheap, less sophisticated Nvidia semiconductors. But governments from Rome to Seoul are cracking down on the user-friendly Chinese app, saying they need to prevent potential leaks of sensitive information through generative AI services. Here is a look at what's going on:
DeepSeek Not Available for Download in South Korea as Authorities Address Privacy Concerns
DeepSeek, a Chinese artificial intelligence startup, has temporarily paused downloads of its chatbot apps in South Korea while it works with local authorities to address privacy concerns, according to South Korean officials on Monday. South Korea's Personal Information Protection Commission said DeepSeek's apps were removed from the local versions of Apple's App Store and Google Play on Saturday evening and that the company agreed to work with the agency to strengthen privacy protections before relaunching the apps. Read More: Is the DeepSeek Panic Overblown? The action does not affect users who have already downloaded DeepSeek on their phones or use it on personal computers. Nam Seok, director of the South Korean commission's investigation division, advised South Korean users of DeepSeek to delete the app from their devices or avoid entering personal information into the tool until the issues are resolved.
Xi-Jack Ma chat seen as next catalyst for blistering China rally
A potential encounter this week between Chinese President Xi Jinping and e-commerce icon Jack Ma, coming after a blistering run by tech shares, could be the next catalyst to extend the rally in China's stocks. Prominent entrepreneurs including Ma have been invited to meet the nation's top leaders, people familiar with the matter said last week. The potential show of support for the private sector coincides with the recent surge in equities in Hong Kong, driven by growing capabilities in artificial intelligence. The Hang Seng China Enterprises Index jumped 4.1% on Friday to its highest since February 2022, exceeding an October peak spurred by a stimulus blitz. A tech gauge in Hong Kong entered a bull market earlier this month, fueled by Chinese startup DeepSeek's AI model that's hailed as a game-changer.