Industry
Kioxia ships samples of new flash memory for AI data centers
Hiroo Ota (center left), CEO of Kioxia Holdings, and others unveil Kioxia's new 3D flash memory chip at its Kitakami plant in Kitakami, Iwate Prefecture, on Friday. Kioxia Holdings has started shipping samples of its next-generation flash memory chips to artificial-intelligence data center operators, seeking to gain ground in the lucrative business against rivals. The Tokyo-based chipmaker's latest high-density 3D flash memory chips aim to better meet AI data center needs with better efficiency and transmission speeds. The 332-layer 10th-generation chips pack more data into silicon and can store 59% more data compared with its previous flagship 8th-generation chip, the company said Friday. Production will take place at the company's second manufacturing facility at its Kitakami plant in Iwate Prefecture, which began operating in September last year.
Australia news live: shadow arts minister Angie Bell, a former musician, says AI giants must pay for content
Follow the day's latest updates Court approves $23.5m fine and costs order against ASX Shadow arts minister says AI companies need to do what everyone else does: 'ask permission and pay for it' Albanese defends gambling reforms, says he's'not against someone having a punt' Pocock says it's'tragic' gambling reforms don't go nearly far enough Shadow arts minister says AI companies need to do what everyone else does: 'ask permission and pay for it' If AI companies want to use Australian creative work, they should do what everyone else does: ask permission and pay for it. Australian creativity is one of our greatest national assets - not a free resource for multinational tech companies. The Coalition will always back the right of artists to control their work and be fairly compensated when others profit from it. This is about consent, fairness and respect for Australian creativity. Court approves $23.5m fine and costs order against ASX Shadow arts minister says AI companies need to do what everyone else does: 'ask permission and pay for it' Albanese defends gambling reforms, says he's'not against someone having a punt' Pocock says it's'tragic' gambling reforms don't go nearly far enough Court approves $23.5m fine and costs order against ASX A federal court judge has ordered the ASX operator to pay $23.5m in penalties and costs after the company admitted to making a misleading statement about a troubled upgrade for technology required to run the stock exchange.
OpenAI proposes handing U.S. government a 5% stake, report says
OpenAI proposes handing U.S. government a 5% stake, report says OpenAI has discussed giving the U.S. government a 5% stake as artificial intelligence firms face scrutiny in Washington. OpenAI has discussed giving the U.S. government a 5% stake, the Financial Times reported on Thursday, as artificial intelligence firms face scrutiny in Washington over the likely misuse of advanced models and whether Americans would benefit from the industry's massive valuations. The ChatGPT creator has proposed that other U.S. AI firms also give Washington similar stakes, although it is unclear whether they would agree, the report said, citing two people familiar with the talks. The move follows growing public backlash in the U.S. over AI's potential to cause economic upheaval, including layoffs, and could help OpenAI sweeten ties with an administration that is increasingly taking an active role in regulating the technology. In a time of both misinformation and too much information, quality journalism is more crucial than ever.
A new, inexpensive Chinese AI model is catching up with Anthropic, OpenAI on their home turf
Zhipu's AI service on the web, dubbed Z.ai. BEIJING/BENGALURU - Since DeepSeek shocked markets early last year with its cheap but powerful artificial intelligence model, global consumers have been faced with a choice: Chinese offerings with lower prices and less capability or OpenAI or Anthropic, which have poured billions into development. A model called GLM-5.2, launched last month by Beijing-based startup Z.ai, may finally be closing that gap in terms of Western interest. GLM-5.2 has Silicon Valley buzzing with its coding and agent capabilities, or the ability to execute complex tasks with minimal prompting, that almost rival leading U.S. offerings at a fraction of the cost, in what some experts are calling a "mini DeepSeek moment." In a time of both misinformation and too much information, quality journalism is more crucial than ever.
Learning Effective Soliton Dynamics from Scattering Data
Minor, Seth, Dukic, Vanja, Bortz, David M.
In such settings, the inverse scattering transform (IST) of Ablowitz, Kaup, Newell, and Segur [2] has enjoyed a rich and successful history, and is now the standard theoretical framework for deriving reduced-order evolution equations for soliton dynamics. Although these derivations are traditionally of an analytical - rather than data-driven - nature, recent work has employed the IST formalism as a tool for experimental data analysis, using the technique to analyze soliton content from empirical measurements [8, 15, 24]. Moreover, recent approaches using alternative parameterization techniques have demonstrated that the learning of reduced-order, interpretable equations of motion for solitons is tenable in a data-driven setting [6, 26, 27]. Despite the success of this recent work, however, little effort has been devoted to developing a data-driven modeling approach based on the IST itself, most likely due to the fact that the framework is fundamentally problem-specific. In this paper, we address the question of whether effective soliton dynamics can be inferred directly from observed scattering data (as opposed to being derived or approximated analytically).
How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size
We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law). Fitting the proposed law on a large set of training runs, we find that it correctly recovers the scaling of the optimal batch size. Moreover, because it makes use of training runs with suboptimal batch size, our proposed law can be robustly fit with a significantly smaller amount of training runs. We further show that the three-term law can be used to derive scaling laws for suboptimal batch sizes, and that it matches previous empirical findings related to the critical batch size.
Unveiling the Non-Monotonic Effect of Privacy on Generalization under Byzantine Robustness
Boudou, Thomas, Bars, Batiste Le, Gupta, Nirupam, Bellet, Aurélien
Recent work has established a fundamental trilemma between Byzantine robustness, local differential privacy (LDP), and optimization error in distributed learning. We show that this trilemma does not universally extend to generalization error, but instead depends critically on the privacy regime. Specifically, in the high-noise regime (strong privacy), we prove that increasing privacy reduces the generalization error, i.e., there is no tension between robustness and privacy. In the low-noise regime (weaker privacy), however, the tension between robustness and privacy reappears and increasing privacy indeed degrades generalization. Our theory explains this surprising non-monotonic behavior of the generalization error via matching lower and upper bounds on the algorithmic stability of Byzantine-robust distributed learning under LDP constraints. We corroborate and further analyze these theoretical findings with empirical evaluations.
Cross-Audit Projection for Model Risk Prediction
For training-data-based model risk prediction, $K$-fold cross-validation~(CV) is widely used to mitigate the well-known over-optimism of the empirical risk and is often regarded as reliable. However, for binary classification via empirical risk minimization, our numerical studies reveal a surprising phenomenon: $K$-fold CV may perform poorly in estimating class-specific risks, even worse than the empirical estimator. We perform a higher-order asymptotic analysis showing that $K$-fold CV may converge at a slower rate, whereas the empirical estimator exhibits a second-order asymptotic bias that explains its over-optimism. These findings motivate a novel two-step procedure for model risk prediction, termed cross-audit projection (CAP). The cross-audit step adopts the same resampling scheme as $K$-fold CV to estimate over-optimism in subsamples, while the asymptotic-theory-informed projection step adjusts for the reduced sample size in bias correction of the empirical risk. The resulting CAP estimator is first-order asymptotically equivalent to the empirical risk while achieving second-order asymptotic unbiasedness. An accompanying inference procedure is also developed. Simulation studies support theoretical advantages of CAP and demonstrate favorable finite-sample performance. An application to breast cancer detection further illustrates the proposed method.
Sequential Structure-Sensitive Residual Diagnostics for PDE Inverse Problems
Computational models in science and engineering are often assessed by checking whether the residual norm is consistent with the assumed noise level. This can be misleading in smoothing inverse problems: structured model errors may be attenuated in observation space, leaving residual magnitudes below practitioner discrepancy thresholds while coherent residual patterns remain. As a result, residual-norm diagnostics can accept fitted models that still give biased parameters, predictions, or quantities of interest. We propose a structure-sensitive sequential diagnostic based on e-processes. The method uses a portfolio of spatial residual-pattern experts, updates their likelihood-ratio wealth as observations are processed, and rejects the fitted model when the aggregate wealth crosses a prescribed threshold, giving anytime-valid type-I error control for a fixed fitted model. We compare the method with Morozov discrepancy checks, fixed-sample residual tests, and batch projection tests. Across three inverse problems (elliptic diffusion, two-dimensional Stokes flow, and a glaciological ice-stream inversion implemented in the community finite-element model icepack) we demonstrate how standard discrepancy checks accept misspecified fits that produce materially wrong quantities of interest. Structure-sensitive batch tests detect these failures using the full dataset, while the e-process detects them earlier from a fraction of the observations. After rejection, the expert wealth attributes the evidence to residual patterns in the chosen dictionary and provides a basis for exploratory model correction.