Weobtain them directly from the low-rank Gaussian distribution for the logits in the network head of SSNs, based on a previously unconsidered view of this distribution as a factor model.
Inmany cases, increasing model capacity beyond the memory limit of a single acceleratorhas required developing special algorithms orinfrastructure. These solutions are often architecture-specific and do not transfer to other tasks.
For example, suppose that a chemical contaminant has accidentally been released and is rapidly spreading; we need to quickly discover its unknown source.
To demonstrate the utility ofPREF-SHAP, we apply our method to a variety of synthetic and real-worlddatasets andshowthatricher andmoreinsightful explanations canbe obtainedoverthebaseline.