Asia
OpenAI floats idea of global AI governance body with U.S. and China
OpenAI floats idea of global AI governance body with U.S. and China The U.S. has an opportunity to use its lead in artificial intelligence technology to create a global governance mechanism to ensure safer, more resilient systems, OpenAI's vice president of global affairs, Chris Lehane, said. OpenAI would support the creation of a global governance body for artificial intelligence led by the U.S. and including China as a member, a top company executive said, hours before the start of U.S. President Donald Trump's high-stakes meeting with Chinese President Xi Jinping. When asked about the China summit, OpenAI's vice president of global affairs, Chris Lehane, said Wednesday that the U.S. has an opportunity to use its lead in AI technology to create a global governance mechanism resulting in safer, more resilient systems. "AI, in some level, transcends a lot of the prevailing or traditional trade type of issues," Lehane told reporters during a briefing at the company's offices in Washington. "There is an opportunity to really start to build something up globally, and have countries around the world, including China, potentially participate." In a time of both misinformation and too much information, quality journalism is more crucial than ever.
Amazon puts Alexa inside the shopping search bar in AI push
Amazon is hoping that AI-powered answers will help keep shoppers from defecting to other sites or chatbots. Artificial intelligence algorithms are coming to some of the most valuable real estate in retail: the Amazon.com Queries typed into Amazon's website and mobile app will soon reply, depending on the context, with product comparisons or suggestions generated by AI large language models, the online retailer said Wednesday. The new tool -- called Alexa for Shopping -- supplants Rufus, the shopping assistant bot that summarized product reviews and suggested purchases. To invoke Rufus, users had to click a blue and orange icon. The new search experience will appear by default, beginning this week for users in the U.S. In a time of both misinformation and too much information, quality journalism is more crucial than ever.
Poor planning fuels Bangladesh contraceptive crisis
A worker arranges packets of condoms at a pharmacy in Dhaka. Bangladesh's family planning system is buckling under severe contraceptive shortages. DHAKA - Bangladesh's once-praised family planning system is buckling under severe contraceptive shortages, raising fears of a rise in unplanned pregnancies in one of the world's most densely populated countries. For decades, the South Asian nation was hailed as a success for slashing birth rates through an expansive state-backed family planning program that sent field workers door to door with pills, condoms and advice on birth spacing. But that system is now faltering, with government clinics across the country of 170 million people running out of basic contraceptives after procurement failures and administrative disruption left supplies depleted in nearly a third of districts. In a time of both misinformation and too much information, quality journalism is more crucial than ever.
TDK ready to step up investment to ride AI wave
TDK CEO Noboru Saito says the firm is prepared to add investments to ride the global boom in generative artificial intelligence. Electronics component linchpin TDK is prepared to add to what is already its biggest capital spending campaign ever in a push to ride the global boom in generative artificial intelligence. The company has added ¥100 billion ($640 million) to its multiyear investment plan each year since it rolled it out in 2024, and now CEO Noboru Saito says the effort may accelerate to match an expected surge in orders and demand. "Should promising prospects arise, our commitment is to make timely and opportunistic investments," Saito, 59, said in an interview. "If we don't sow the seeds for medium-to long-term growth now, we won't be able to reap the harvest later." In a time of both misinformation and too much information, quality journalism is more crucial than ever.
Japan megabanks set to win Mythos access after Bessent visit
MUFG Bank, Mizuho Bank and Sumitomo Mitsui Banking are all likely to gain access to Anthropic's artificial intelligence model, Mythos. Japan's three megabanks are set to secure access to Anthropic's artificial intelligence model, Mythos, according to a person familiar with the matter, after its limited release last month sparked fears of a new age of cybersecurity risks. MUFG Bank, Sumitomo Mitsui Banking Corp. and Mizuho Bank are all likely to gain access to the artificial intelligence model developed by the U.S. firm, the person said, asking not to be identified because the information is private. The planned access was earlier reported by Nikkei. The move comes as financial institutions around the world grow alarmed about the risks created by Mythos, which has an unprecedented ability to detect software vulnerabilities. That has raised concerns that hackers could use Mythos to disrupt critical infrastructure, and access has so far been limited to a small number of U.S. companies and organizations.
Musk's xAI races to get Wall Street firms to use Grok chatbot
Musk's xAI races to get Wall Street firms to use Grok chatbot A chat window for chatbot Grok. Musk's artificial intelligence venture, xAI, is moving with urgency to boost revenue by selling chatbot subscriptions and access to its computing resources before SpaceX's expected IPO next month. Billionaire Elon Musk's xAI has recruited multiple Wall Street firms with ties to his business empire to test its Grok chatbot, according to people familiar with the matter, part of a push to bolster revenue ahead of parent company SpaceX's initial public offering. Apollo Global Management and Morgan Stanley have begun using Grok internally alongside software from other AI model makers, said the people, who spoke on condition of anonymity as the information is not public. Valor Equity Partners is also using Grok, the people said. Despite some banks signing up for Grok, financiers are rarely using the chatbot for work, some of the people said.
From Generalist to Specialist Representation
Zheng, Yujia, Feng, Fan, Li, Yuke, Xie, Shaoan, Murphy, Kevin, Zhang, Kun
Given a generalist model, learning a task-relevant specialist representation is fundamental for downstream applications. Identifiability, the asymptotic guarantee of recovering the ground-truth representation, is critical because it sets the ultimate limit of any model, even with infinite data and computation. We study this problem in a completely nonparametric setting, without relying on interventions, parametric forms, or structural constraints. We first prove that the structure between time steps and tasks is identifiable in a fully unsupervised manner, even when sequences lack strict temporal dependence and may exhibit disconnections, and task assignments can follow arbitrarily complex and interleaving structures. We then prove that, within each time step, the task-relevant latent representation can be disentangled from the irrelevant part under a simple sparsity regularization, without any additional information or parametric constraints. Together, these results establish a hierarchical foundation: task structure is identifiable across time steps, and task-relevant latent representations are identifiable within each step. To our knowledge, each result provides a first general nonparametric identifiability guarantee, and together they mark a step toward provably moving from generalist to specialist models.
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
Du, Zhehang, He, Hangfeng, Su, Weijie
Large language models (LLMs) are pretrained by minimizing the cross-entropy loss for next-token prediction. In this paper, we study whether this optimization strategy can induce geometric structure in the learned model weights and context embeddings. We approach this problem by analyzing a constrained layer-peeled optimization program, which serves as a mathematically tractable surrogate for LLMs by treating the output projection matrix and last-layer context embeddings as optimization variables. Our analysis of this nonconvex optimization program demonstrates that symmetries in the target next-token distributions are transferred to the global minimizers of the layer-peeled model in a precise group-theoretic sense. Specifically, we prove that when the target tokens exhibit a cyclic-shift symmetry (such as the seven days of the week or the twelve months of the year), the optimal logit matrix is exactly circulant, and the Gram matrices of both the output projections and the context embeddings form circulant geometries as well. Next, for exchangeable target distributions invariant under the symmetric group and, more generally, under two-transitive group actions, we show that the global optimal output projection matrix forms a simplex equiangular tight frame, while the optimal logit matrix and context embeddings inherit the permutation symmetries present in the input data. A key technical step is to reduce the constrained nonconvex factorized problem to an explicit logit-level convex characterization for cyclic symmetry and to a symmetry-based lower bound for permutation symmetry, together with a sharp characterization of the optimal factorization. Finally, we empirically demonstrate that open-source LLMs naturally exhibit symmetries consistent with our theoretical predictions, despite being trained without any explicit regularization promoting such geometric structure.
Robust Sequential Experimental Design for A/B Testing
Wen, Qianglin, Wu, Xiangkun, Shi, Chengchun, Li, Ting, Tang, Niansheng, Zhang, Yingying, Zhu, Hongtu
Experimental design has emerged as a powerful approach for improving the sample efficiency of A/B testing, yet existing designs rely critically on correctly specified models. We study robust sequential experimental design under model misspecification and develop a unified framework that covers both contextual bandit and dynamic settings. Theoretically, we prove that our design bounds the worst-case mean squared error of the estimated treatment effect. Empirically, we demonstrate the effectiveness of the proposed approach using synthetic and real-world datasets from a leading technology company.
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
Weak-to-strong (W2S) generalization, in which a strong model is fine-tuned on outputs of a weaker, task-specialized model, has been proposed as an approach to aligning superhuman AI systems. Existing theoretical analyses either fix the student's representations or operate in restricted settings. Whether multi-step SGD can succeed in feature learning while preserving diverse pre-trained capabilities remains open. We study W2S in the setting of reward-model learning with two-layer neural networks. The strong model has pre-trained representations organized into low-dimensional subspaces $V_k$, and is fine-tuned under the supervision of a weak model specialized on task $κ$. We prove that the strong model efficiently learns task $κ$, eliciting its pre-trained knowledge while retaining general capabilities. This establishes W2S generalization in the feature-learning regime, in the sense that the strong model acquires the target feature direction through W2S training, rather than having it given a priori. Moreover, W2S preserves pre-trained off-target features, whereas standard supervised fine-tuning causes catastrophic forgetting when off-target feature directions are correlated with the target's. Numerical experiments on synthetic data confirm our theoretical results.