Fidelity Isn't Accuracy: When Linearly Decodable Functions Fail to Match the Ground Truth
Neural networks have revolutionized supervised learning, but their internal complexity often renders their learned functions opaque. In effect, they behave as black boxes. This lack of interpretability hinders trust, accountability, and transparency--especially in high-stakes domains such as healthcare, finance, and criminal justice, where understanding a model's decisions can be as important as their accuracy. While many techniques exist for interpreting classification networks--such as saliency maps, feature attribution methods like LIME [10] and SHAP [6], and probing approaches [1]--the interpretability of regression networks remains comparatively underexplored, especially at the level of input-output function behavior [6, 14]. One common strategy in model interpretability is to approximate a complex system with a simpler, interpretable surrogate, such as a linear model or decision tree [3, 12]. In classification settings, surrogate models are often applied to internal representations to evaluate the linear accessibility of information at each layer [1, 5]. However, in regression tasks, there has been limited work focused on assessing the linearity of a network's full input-output function. In this paper, we introduce a simple yet powerful idea: we quantify how linearly decodable a trained regression network's output function is.
Jun-17-2025
- Country:
- North America > United States
- California (0.07)
- New York > New York County
- New York City (0.04)
- Massachusetts > Middlesex County
- Cambridge (0.04)
- North America > United States
- Genre:
- Research Report > New Finding (0.68)
- Industry:
- Health & Medicine (0.50)
- Technology: