Fisher information flow in artificial neural networks
Weimar, Maximilian, Rachbauer, Lukas M., Starshynov, Ilya, Faccio, Daniele, Adilova, Linara, Bouchet, Dorian, Rotter, Stefan
–arXiv.org Artificial Intelligence
In physics, artificial neural networks (ANNs) are used in a wide variety of research areas, ranging from optics [1] to high-energy physics [2], the computation of material parameters [3], the design of experiments [4], and the solution of complex inverse problems [5]. In the context of ANN training, one of the central issues is to understand how models process the information that they receive as input [6-9]. One particularly successful approach is based on the notion of mutual information (MI) and the information bottleneck method [10-12], which provide a peek into the training process of the model. MI is also used for regularizing the optimization [13, 14] and for deriving generalization bounds [15, 16]. Nevertheless, this line of research is severely hindered by the challenges of estimating MI in high-dimensional spaces [17] and by the issue that the MI diverges for continuous random variables when they are connected by a deterministic transformation [18-21]. The starting point for the current work is the observation that a whole class of estimation problems in physics centers around the concept of Fisher information (FI) [22, 23]. Arising from the field of (statistical) estimation theory [22], FI is the key quantity when dealing with estimating continuous parameters from noisy data. The amount of FI one has available on a given parameter ultimately determines how precisely the value of this parameter can be estimated and, thus, bounds the achievable performance of the estimating model.
arXiv.org Artificial Intelligence
Sep-25-2025
- Country:
- North America > United States (1.00)
- Europe (1.00)
- Asia > Middle East
- Israel (0.28)
- Genre:
- Research Report (1.00)
- Technology: