amari
Sutton's predictions v Crookhaven stars Amari Bacchus & Genesis Lynea
Two of the teams fighting relegation meet on Sunday when Tottenham host Nottingham Forest, but are there more than just points at stake? If we do get a winner here, it is a huge boost for that team psychologically going into the international break, said BBC Sport football expert Chris Sutton. But, for the losing manager, it could mean the sack. That applies to Forest's Vitor Pereira as well as Igor Tudor at Spurs - this is a classic game where triumph or disaster awaits both clubs. Sutton is making predictions for all 380 Premier League games this season, against AI, BBC Sport readers and a variety of guests. His guests for week 31 are Amari Bacchus and Genesis Lynea, stars of new CBBC drama series Crookhaven. Crookhaven begins with a double bill on Sunday, 22 March at 15:05 GMT on BBC One and BBC iPlayer, and at 17:25 on CBBC. The full series will be available to watch on BBC iPlayer from this date.
Legendre Decomposition for Tensors
Mahito Sugiyama, Hiroyuki Nakahara, Koji Tsuda
CP decomposition compresses an input tensor into a sum of rank-one components, and Tucker decomposition approximates an input tensor by a core tensor multiplied by matrices. To date, matrix and tensor decomposition has been extensively analyzed, and there are a number of variations of such decomposition (Kolda and Bader, 2009), where the common goal is to approximate a given tensor by a smaller number of components, or parameters,inanefficientmanner. However, despite the recent advances of decomposition techniques, a learning theory that can systematically define decomposition for any order tensors including vectors and matrices is still under development. Moreover, it is well known that CP and Tucker tensor decomposition include non-convex optimization and that the global convergence is not guaranteed.
Legendre Decomposition for Tensors
Mahito Sugiyama, Hiroyuki Nakahara, Koji Tsuda
We present a novel nonnegative tensor decomposition method, called Legendre decomposition, which factorizes an input tensor into a multiplicative combination of parameters. Thanks to the well-developed theory of information geometry, the reconstructed tensor is unique and always minimizes the KL divergence from an input tensor. We empirically show that Legendre decomposition can more accurately reconstruct tensors than other nonnegative tensor decomposition methods.
Generalized Power Priors for Improved Bayesian Inference with Historical Data
Kimura, Masanari, Bondell, Howard
The power prior is a class of informative priors designed to incorporate historical data alongside current data in a Bayesian framework. It includes a power parameter that controls the influence of historical data, providing flexibility and adaptability. A key property of the power prior is that the resulting posterior minimizes a linear combination of KL divergences between two pseudo-posterior distributions: one ignoring historical data and the other fully incorporating it. We extend this framework by identifying the posterior distribution as the minimizer of a linear combination of Amari's $ฮฑ$-divergence, a generalization of KL divergence. We show that this generalization can lead to improved performance by allowing for the data to adapt to appropriate choices of the $ฮฑ$ parameter. Theoretical properties of this generalized power posterior are established, including behavior as a generalized geodesic on the Riemannian manifold of probability distributions, offering novel insights into its geometric interpretation.
Duality induced by an embedding structure of determinantal point process
Specifically, we clarify the embedding structure of a DPP model in the exponential family of log-linear models (c.f., Agresti, 1990; Amari, 2001) in Theorem 1. Models embedded in exponential families are called curved exponential families. Information geometry (Amari, 1985) provides a measure, the e-embedding curvature tensor (Efron, 1975; Reeds, 1975; Amari, 1982; Sei, 2011), to quantify the extent to which a curved exponential family deviates from an exponential family. To check the e-embedding curvature as well as the Fisher information matrix, we apply the diagonal scaling (Marshall and Olkin, 1968), also known as the quality vs. diversity decomposition in the DPP literature (Kulesza and Taskar, 2012), to an L-ensemble kernel of a DPP model and then evaluate them, which clarifies that the subset of parameters related to the item-wise effects (quality terms) has zero e-embedding curvature (Corollary 1).
On the Fisher-Rao Gradient of the Evidence Lower Bound
This article studies the Fisher-Rao gradient, also referred to as the natural gradient, of the evidence lower bound, the ELBO, which plays a crucial role within the theory of the Variational Autonecoder, the Helmholtz Machine and the Free Energy Principle. The natural gradient of the ELBO is related to the natural gradient of the Kullback-Leibler divergence from a target distribution, the prime objective function of learning. Based on invariance properties of gradients within information geometry, conditions on the underlying model are provided that ensure the equivalence of minimising the prime objective function and the maximisation of the ELBO.
Clustering above Exponential Families with Tempered Exponential Measures
Amid, Ehsan, Nock, Richard, Warmuth, Manfred
The link with exponential families has allowed $k$-means clustering to be generalized to a wide variety of data generating distributions in exponential families and clustering distortions among Bregman divergences. Getting the framework to work above exponential families is important to lift roadblocks like the lack of robustness of some population minimizers carved in their axiomatization. Current generalisations of exponential families like $q$-exponential families or even deformed exponential families fail at achieving the goal. In this paper, we provide a new attempt at getting the complete framework, grounded in a new generalisation of exponential families that we introduce, tempered exponential measures (TEM). TEMs keep the maximum entropy axiomatization framework of $q$-exponential families, but instead of normalizing the measure, normalize a dual called a co-distribution. Numerous interesting properties arise for clustering such as improved and controllable robustness for population minimizers, that keep a simple analytic form.
q-Paths: Generalizing the Geometric Annealing Path using Power Means
Masrani, Vaden, Brekelmans, Rob, Bui, Thang, Nielsen, Frank, Galstyan, Aram, Steeg, Greg Ver, Wood, Frank
Many common machine learning methods involve the geometric annealing path, a sequence of intermediate densities between two distributions of interest constructed using the geometric average. While alternatives such as the moment-averaging path have demonstrated performance gains in some settings, their practical applicability remains limited by exponential family endpoint assumptions and a lack of closed form energy function. In this work, we introduce $q$-paths, a family of paths which is derived from a generalized notion of the mean, includes the geometric and arithmetic mixtures as special cases, and admits a simple closed form involving the deformed logarithm function from nonextensive thermodynamics. Following previous analysis of the geometric path, we interpret our $q$-paths as corresponding to a $q$-exponential family of distributions, and provide a variational representation of intermediate densities as minimizing a mixture of $\alpha$-divergences to the endpoints. We show that small deviations away from the geometric path yield empirical gains for Bayesian inference using Sequential Monte Carlo and generative model evaluation using Annealed Importance Sampling.
Japan must work with TSMC to rebuild chipmaking base, ex-economy minister says
Japan can't build a cutting-edge chip development and manufacturing base on its own, and must seek to cooperate with Taiwan Semiconductor Manufacturing Co. (TSMC), according to Akira Amari, a senior lawmaker from the ruling Liberal Democratic Party. Amari, a former economy minister who heads an LDP working group on semiconductor strategy, added that the government must be prepared to spend trillions of yen to keep up with the U.S. and Europe. Both have plans to pour money into the industry amid a global shortage of semiconductors, including advanced logic chips that are essential for everything from artificial intelligence to autonomous driving. "Unlike the purely domestic, independent way it was done in the past, I think we need to cooperate with overseas counterparts," Amari said in an interview in Tokyo on Monday. "The world's top logic chipmaker is TSMC, so we must think about how to cooperate with them."
Japan shouldn't ignore potential TikTok data risks, top LDP official says
Japan shouldn't ignore the data security risks posed by the Chinese video app TikTok, a senior ruling party official said. "Not only President Trump but also other countries such as the U.K. and India, are gradually becoming aware of the risks," Akira Amari, the ruling Liberal Democratic Party's tax panel chief, said Sunday on Fuji Television Network. "Since there are so many countries pointing out the risks, Japan cannot just stand by and watch." U.S. President Donald Trump on Friday ordered ByteDance Ltd., TikTok's Chinese owner, to sell its U.S. assets. Trump cited national security grounds, delivering the latest salvo in his standoff with Beijing.