Balancing Sparse RNNs with Hyperparameterization Benefiting Meta-Learning
Hershey, Quincy, Paffenroth, Randy
–arXiv.org Artificial Intelligence
Abstract--This paper develops alternative hyperparameters for specifying sparse Recurrent Neural Networks (RNNs). These hyperparameters allow for varying sparsity within the trainable weight matrices of the model while improving overall performance. This architecture enables the definition of a novel metric, hidden proportion, which seeks to balance the distribution of unknowns within the model and provides significant explanatory power of model performance. T ogether, the use of the varied sparsity RNN architecture combined with the hidden proportion metric generates significant performance gains while improving performance expectations on an a priori basis. This combined approach provides a path forward towards generalized meta-learning applications and model optimization based on intrinsic characteristics of the data set, including input and output dimensions. Selection and specification of neural networks remains an actively studied arena [1]-[3]. Practitioners traditionally follow a pathway of first selecting a model architecture suited to the general characteristics of the data set. Afterwards, model hyperparameters are often optimized through a process involving repeated training runs and cross-validation [4]. The search for optimal hyperparameters is often costly as the total potential combinations grow exponentially with the number of hyperparameters and model effectiveness can often only be judged after training.
arXiv.org Artificial Intelligence
Sep-19-2025