Network Inversion of Binarised Neural Nets
Suhail, Pirzada, Chakraborty, Supratik, Sethi, Amit
–arXiv.org Artificial Intelligence
The need for demystifying the internal workings of neural networks has led to the development of techniques like network inversion that unravel the black-box nature of the network by reconstructing the input space from the model's learned internal representations. The challenge of network inversion is underscored by the complex many-to-one mappings inherent in neural networks, exacerbated by the activation functions employed, making the inversion process non-trivial. This paper proposes a novel approach to invert a Binarised Neural Network (BNN) by encoding the trained BNN into a Conjunctive Normal Form (CNF) propositional formula that comprehensively captures the network's structure, encompassing the values in the input, hidden, and output layers of the network. The CNF formula precisely encodes the computation in each neuron of the BNN, thereby making it possible to mimic the overall inference task with 100% precision. Interestingly, the same encoding can be used for the inversion problem simply by constraining the propositional variables corresponding to network outputs, and invoking a satisfiability (SAT) solver to find satisfying assignments for the propositional variables corresponding to network inputs. This approach unlike other optimization-based techniques, circumvents the need for any careful hyper-parameter tuning. Additionally, the deterministic nature of CNF encodings provides finegrained control over inverted samples. Specifically, sampling the satisfying assignments (near)uniformly guarantees diverse input generation during inversion, addressing concerns like mode collapse commonly associated with generative models.
arXiv.org Artificial Intelligence
Feb-20-2024