Goto

Collaborating Authors

 Asia


OpenAI makes move to go public one week after rival Anthropic

The Japan Times

OpenAI, founded in San Francisco in 2015 as a nonprofit research lab, burst into the mainstream with the launch of ChatGPT in November 2022. It has since restructured as a for-profit corporation. SAN FRANCISCO, UNITED STATES - ChatGPT-maker OpenAI on Monday took the first step toward going public, one week after archrival Anthropic announced its own filing, as both companies look to raise the massive sums needed to expand. In a social media post, the Sam Altman-led company said it had confidentially submitted an S-1 registration statement to U.S. securities regulators but had "not decided on timing yet" for any potential debut. OpenAI's move follows a confidential filing by Anthropic, the maker of the Claude chatbot, which announced last Monday that it had taken the same step. In a time of both misinformation and too much information, quality journalism is more crucial than ever.


Causal Longitudinal Prior-Fitted Networks for Counterfactual Outcome Prediction

arXiv.org Machine Learning

Longitudinal treatment decisions from multivariate time-series data require predicting potential outcomes under future treatment sequences in the presence of timevarying confounding, heterogeneous patient dynamics, and limited domain-specific data. Existing longitudinal causal estimators typically address this problem by training a new model for each cohort or simulator. We introduce Causal Longitudinal Prior-Fitted Networks (CAUSALLONGPFN), a prior-fitted network for time-series causal inference in longitudinal treatment-response data and zero-shot in-context counterfactual outcome prediction. To our knowledge, CAUSALLONGPFN is the first PFN-style model for history-conditional potential-outcome prediction under planned longitudinal treatment sequences, with systematic comparison against established longitudinal causal baselines on branchable counterfactual treatmentresponse benchmarks and factual real-world clinical data. The model is pretrained entirely on synthetic episodes sampled from a broad prior over temporal structural causal models, exposing it to treatment-confounder feedback, latent heterogeneity, nonlinear state evolution, delayed effects, and cumulative treatment responses. At test time, CAUSALLONGPFN remains frozen and is used zero-shot: it conditions on support trajectories, a query history, and a planned future treatment sequence, and returns a predictive distribution over future outcomes without gradient updates or propensity-model fitting. Multi-step predictions are obtained by recursively applying the one-step predictor under the specified treatment sequence. We evaluate the model on branchable cancer, HIV, and warfarin benchmarks with ground-truth counterfactual labels, and on factual-only rolling-origin prediction in MIMIC-III ICU trajectories. CAUSALLONGPFN is competitive with domain-trained longitudinal baselines on counterfactual benchmarks and performs strongly on factual MIMIC-III prediction, suggesting that broad synthetic causal pretraining can provide a frozen, amortized alternative for zero-shot longitudinal treatment-response prediction when repeated domain-specific training is costly or impractical.


Nonparametric undirected graphical model selection using diffusion models

arXiv.org Machine Learning

Undirected graphical models provide a fundamental framework for representing conditional independence structures among high-dimensional random variables. While undirected graphical model selection has become a central problem in high-dimensional statistics, most existing methods are restricted to parametric settings. In this paper, we develop a nonparametric approach to undirected graphical model selection based on diffusion models. Recent work has shown that diffusion models can adapt to the unknown graph structure of the underlying distribution, yet utilizing these models for explicit graph estimation remains unexplored. To bridge this gap, we introduce a novel diffusion-based method for nonparametric undirected graphical model selection. We establish the model selection consistency of the proposed method and demonstrate its empirical performance through extensive simulations and two real data analyses.


Identifiability and Estimation for Unlabeled Finite Mixtures under Marginal Independence

arXiv.org Machine Learning

We study component recovery and mixing-matrix estimation from unlabeled finite mixtures whose observable distributions share the same latent components but have unknown mixing weights. The main identifying signal is marginal independence: each component is assumed to be independent on at least one coordinate pair, but no labels, clean component samples, or mixing weights are observed. We first prove a structural result for product components: under linear independence of the univariate marginals, any independent affine combination of the components must coincide with a single component. We then extend this principle to observable mixtures and show that, under full-rank and no-cancellation conditions, marginally independent affine combinations recover the corresponding latent components. When every component is independent on some coordinate pair, all components are identifiable, and the mixing matrix is recoverable under the stated completion conditions. Finally, we propose a Product-Marginal Maximum Mean Discrepancy (PM-MMD) estimator over affine combinations of the observable mixtures and prove uniform convergence and stability under approximate marginal independence. This framework also separates the empirical roles of the assumptions: irreducibility is, in general, not directly testable from the unlabeled mixtures alone, whereas marginal independence yields a candidate-level diagnostic through held-out PM-MMD. Controlled and flow-cytometry experiments show when marginal independence provides a useful recovery signal. In the reported multi-component comparisons, condition-aware representative selection stabilizes PM-MMD and improves recovery relative to clustering, factorization, and pairwise mixture-proportion baselines using the same unlabeled mixtures.


Boundary Variance Inflation Causes Acquisition Bias in Gaussian Processes

arXiv.org Machine Learning

Gaussian processes with stationary kernels on bounded domains exhibit inflated posterior variance near the boundary. Despite being a long-recognized artifact in geostatistics and a source of over-exploration in Bayesian optimization, the causes and effects of boundary-induced acquisition bias are underexplored. We trace the root cause to a simple geometric mechanism: the truncation of the kernel correlation neighborhood at the domain boundary creates an observation-independent distortion that worsens with dimensionality. We show how this distortion manifests across three acquisition classes: variance maximization concentrates selections at the corners, whereas negative integrated posterior variance and expected predictive information gain move selections inward to axis-aligned interior shells. These patterns arise without reference to any objective function, meaning that acquisition behavior can be dominated by kernel geometry rather than the desired task-specific uncertainty. To quantify this, we introduce a function-free selection-profile diagnostic for arbitrary acquisitions, kernels, and bounded-domain geometries.


LOTTERY: Learning from Reference-Only Samples in Two-Sample Testing under Size Asymmetry

arXiv.org Machine Learning

Data-adaptive two-sample testing assesses if two samples come from the same distribution, using a discrepancy learned from the data (e.g., via kernel-based feature representations). Such methods typically rely on data splitting to decouple learning from testing and control type I error. However, this paradigm is ill-suited to few-shot settings with severe sample-size imbalance: abundant reference samples are available, while only a handful of query samples arrive. In this paper, we show how this imbalance can be leveraged constructively. Using abundant reference data, we learn reference-dependent representations that summarize salient structure of the reference distribution and provide informative signals for detecting departures. We incorporate a collection of representation families that capture both global and local structure, and adaptively weight them using only reference samples via an uncertainty-guided principle. Theoretically, we establish permutation-based type I error control and show consistency of the aggregated test: as the sample sizes grow, the test power converges to one whenever the representation set contains at least one consistent representation. Empirically, our aggregation achieves strong performance across a range of benchmarks while retaining type I error control.


CP-factorization for high dimensional tensor time series and double projection iterations

arXiv.org Machine Learning

We adopt the canonical polyadic (CP) decomposition to model high-dimensional tensor time series. Our primary goal is to identify and estimate the factor loadings in the CP decomposition. We propose a one-pass estimation procedure through standard eigen-analysis for a matrix constructed based on the serial dependence structure of the data. The asymptotic properties of the proposed estimator are established under a general setting as long as the factor loading vectors are linearly independent, allowing the factors to be correlated and the factor loading vectors to be not nearly orthogonal. The procedure adapts to the sparsity of the factor loading vectors, accommodates weak factors, and demonstrates strong performance across a wide range of scenarios. To further reduce estimation errors, we also introduce an iterative algorithm based on a novel double projection approach. We theoretically justify the improved convergence rate of the iterative estimator, and derive the associated limiting distribution. A consistent estimator of the asymptotic variance is also provided, which plays a key role in the related inference problems. All results are validated through extensive simulations and two real data applications.


Beyond Additivity: Causal Discovery in Location-Scale Noise Models with Hidden Variables

arXiv.org Machine Learning

We study causal discovery from observational data when some variables are hidden and the data-generating process follows a location-scale noise model (LSNM). Existing methods that handle hidden confounders typically assume additive noise, but in practice, causes often modulate not just the mean but also the variance of their effects. We prove that acyclic directed mixed graphs (ADMGs) satisfying a bow-free condition are identifiable under LSNM with hidden variables, establishing the first identifiability result for causally insufficient models beyond noise additivity. We further provide sufficient conditions for identifying causal direction even when the bow-free assumption is violated. Our two-stage algorithm, LSNM-UV, is sound and complete, and experiments demonstrate improved performance over additive baselines on heteroscedastic data.


Could humanoid robots be heading for the battlefield?

BBC News

Could humanoid robots be heading for the battlefield? I've come to an industrial space in a tech-heavy area of San Francisco expecting to see a menacing humanoid robot solider doing something combat-like: the future of land-based warfare, perhaps. Instead, the black shiny faceless Phantom robot is engaged in free play, manipulating a bunch of coloured kids blocks. We need data from it just interacting with its environment [and] this is today's menu, explains Sankaet Pathak, co-founder and CEO of two-year-old start-up Foundation Robotics, which is developing Phantom for military and civilian applications. Later he pushes its 80kg steel-covered body around the room to demonstrate its stability and shows me how it walks.


Watch: Intel's Tom Petersen talks Arc G3 and handheld gaming at Computex

PCWorld

PCWorld interviewed Intel Fellow Tom Petersen at Computex 2026 about Intel's new Arc G3 Extreme chipset designed for handheld gaming PCs. Intel claims the Arc G3 Extreme delivers a 42% performance advantage over AMD's competing Ryzen Z2 Extreme processor. The chipset targets sustained high performance in thin and light portable devices, potentially revolutionizing the handheld gaming market. Intel's new Arc G3 Extreme chipset could be a game-changer for handheld gaming PCs, and we got our first glimpse of how that'd be on the show floor at Computex 2026. PCWorld's Adam Patrick Murray was there in Taiwan all week checking out the future of PCs, and that included a quick chat with Intel Fellow Tom Petersen about the gaming potential of the Arc G3 Extreme.