regime
Thousands of North Korean IT workers are infiltrating corporate America
This material may not be published, broadcast, rewritten, or redistributed. Quotes displayed in real-time or delayed by at least 15 minutes. Market data provided by Factset . Powered and implemented by FactSet Digital Solutions . Mutual Fund and ETF data provided by LSEG . World's first solar-powered ambulance brings healthcare off-grid The scammer at your doctor's office may already know who you are Shared VPN vs dedicated IP: Which one is right for you? Is cyberbullying hiding in your child's group chat?
Intermittent swimming promotes the energy efficiency of fish-like robot movements
Improving energy performance can effectively extend the time a robot can operate and reduce battery load, enabling lighter, more flexible, and more durable robotic systems. Nature has evolved optimal energy-saving locomotion strategies through billions of years of natural selection, providing unparalleled blueprints for robotic optimization. Among diverse modes of aquatic locomotion, intermittent swimming, also called bout-and-glide swimming, is a widespread adaptive behavior in aquatic organisms of a wide range of sizes, including larval zebrafish, red-nose tetra, koi carp, and even whales. This natural bout-and-glide gait features alternating motion phases: short periods of active body and tail undulation for propulsion, followed by passive gliding with a streamlined, straight body posture. It is widely recognized that this intermittent swimming gait is closely associated with optimizing biological energy, making it of great research value to transplant and explore such natural motion mechanisms into robotic control systems.
The Quiet Miracles of Ordinary Life in New Syria
Follow this section to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens. Follow this tag to personalize your feed and get instant alerts. Follow Go to your personalized feed WHY FOLLOW? Smart Alerts: Get notified about major news as it happens.
Learning Effective Soliton Dynamics from Scattering Data
Minor, Seth, Dukic, Vanja, Bortz, David M.
In such settings, the inverse scattering transform (IST) of Ablowitz, Kaup, Newell, and Segur [2] has enjoyed a rich and successful history, and is now the standard theoretical framework for deriving reduced-order evolution equations for soliton dynamics. Although these derivations are traditionally of an analytical - rather than data-driven - nature, recent work has employed the IST formalism as a tool for experimental data analysis, using the technique to analyze soliton content from empirical measurements [8, 15, 24]. Moreover, recent approaches using alternative parameterization techniques have demonstrated that the learning of reduced-order, interpretable equations of motion for solitons is tenable in a data-driven setting [6, 26, 27]. Despite the success of this recent work, however, little effort has been devoted to developing a data-driven modeling approach based on the IST itself, most likely due to the fact that the framework is fundamentally problem-specific. In this paper, we address the question of whether effective soliton dynamics can be inferred directly from observed scattering data (as opposed to being derived or approximated analytically).
Unveiling the Non-Monotonic Effect of Privacy on Generalization under Byzantine Robustness
Boudou, Thomas, Bars, Batiste Le, Gupta, Nirupam, Bellet, Aurélien
Recent work has established a fundamental trilemma between Byzantine robustness, local differential privacy (LDP), and optimization error in distributed learning. We show that this trilemma does not universally extend to generalization error, but instead depends critically on the privacy regime. Specifically, in the high-noise regime (strong privacy), we prove that increasing privacy reduces the generalization error, i.e., there is no tension between robustness and privacy. In the low-noise regime (weaker privacy), however, the tension between robustness and privacy reappears and increasing privacy indeed degrades generalization. Our theory explains this surprising non-monotonic behavior of the generalization error via matching lower and upper bounds on the algorithmic stability of Byzantine-robust distributed learning under LDP constraints. We corroborate and further analyze these theoretical findings with empirical evaluations.
Dead-Direction Conditioners: Gauge-Equivariant Preconditioning for Deep Networks
A deep network's loss is invariant to continuous symmetries of its parameters: the logit shift, the ReLU rescaling, the LayerNorm scale, the per-head attention rotation. Adam's per-coordinate preconditioner drifts along each symmetry orbit, which pulls the trajectory off the symmetry quotient where the optimization lives and blurs the singular-learning rate the quotient makes readable. We build DDC, a Dead-Direction Conditioner that lifts a base optimizer into a $G$-equivariant one: it conditions the optimizer's state in the orbit decomposition of a $G$-invariant metric, so the trajectory stays a preconditioned gradient flow on the quotient $\barΘ= Θ/G$. The construction carries four architectural gauges (cross-entropy shift, ReLU and SwiGLU rescaling, LayerNorm and RMSNorm scale, and a per-head $O(d_{\rm head})$ attention rotation matched to RoPE), proves exactly equivariant on an Adam base, and composes with a Muon base through a gauge-equivariant orthogonaliser. Respecting the symmetry changes both the minimum the optimizer reaches and what it leaves measurable there. On a language model trained past the point of fit, DDCAdam resists the over-training collapse AdamW falls into, holding a validation-train loss gap of 0.67 against 5.88, and reads the dead-direction rate in 32 of 65 layer-by-observable cells where AdamW reads it in 7. A vision transformer trained from scratch reaches lower validation loss (1.71 against 2.12) while compressing spare feed-forward capacity a matched AdamW leaves intact. On a Muon base, where the rotation gauge composes exactly, DDCMuon groks ten of eleven seeds at depth 24 that a plain Muon never reaches. Built into the optimizer, a network's gauge symmetry sharpens the minimum it finds and turns that minimum's geometry into something the trajectory can measure.
Disentangling Continuous-Time Latent Dynamics: Identifiability of Latent SDEs via Diffusion Shifts
Wang, Yuanyuan, Wang, Wenjie, Li, Haoxuan, Gong, Mingming, Zhang, Kun
Causal representation learning for time series has developed strong identifiability results in discrete-time latent causal models, but identifiability in continuous-time latent stochastic differential equation (SDE) models remains largely open. We address this gap using environment-induced shifts in diffusion covariance. We study additive-noise latent SDEs observed through an unknown nonlinear diffeomorphism, with shared drift but environment-specific diffusion covariance. We show that two diagonal diffusion regimes with pairwise distinct coordinate-wise variance ratios identify the latent coordinates up to permutation and scaling, without any sparsity assumption on the drift. We first prove this result for linear Ornstein--Uhlenbeck systems and then extend it to general additive-noise latent SDEs. Under mild smoothness, the instantaneous drift-Jacobian causal graph is identifiable up to the same permutation. We propose a two-stage estimator for latent disentanglement and optional graph recovery; experiments on synthetic systems confirm the predicted identifiability boundary, and an application to Hardanger Bridge monitoring data illustrates the approach on real sensor trajectories.
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
Girardin, Julius, Troiani, Emanuele, Xu, Yizhou, Erba, Vittorio, Krzakala, Florent, Zdeborová, Lenka
Understanding how performance scales jointly with model size and data is a central problem in modern machine learning. Existing theoretical works on scaling laws typically describe generalization as a function of data or compute, often in fixed-feature or infinite-width regimes and for online SGD. Here, we instead study how generalization scales with the number of trainable parameters and the number of samples in a feature-learning model. We analyze $\ell_2$-regularized empirical test error minimization in a quadratic two-layer network in a finite-sample setting with structured data. This setting allows for an explicit characterization of the generalization error as a function of the number of samples, model width, and regularization. Our results reveal a phase diagram with distinct scaling regimes as the number of parameters varies. In particular, the generalization error follows data-dependent power laws controlled by the spectral structure of the target. We further characterize the transitions between regimes, including the onset of interpolation, and their impact on generalization.