Goto

Collaborating Authors

 Country


1 Appendix

Neural Information Processing Systems

L(fk(xi),yi), (1) wherefk()andθk are the local model and model parameter,respectively. However,theyintroduced apublic dataset to enhance training, which is not practical. Overall, none of the above methods can be practically applied. Zhu et al. [14] also proposed a data-free knowledge distillation approach for FL, which learns a generator derived from the prediction of local models. In2019 IEEE 25Th international conference on parallel and distributed systems (ICPADS),pages985-989.IEEE,2019.



What Matters in Graph Class Incremental Learning An Information Preservation Perspective

Neural Information Processing Systems

Graph class incremental learning (GCIL) requires the model to classify emerging nodes of new classes while remembering old classes. Existing methods are designed to preserve effective information of old models or graph data to alleviate forgetting, but there is no clear theoretical understanding of what matters in information preservation.


ConceptEmbeddingModels: BeyondtheAccuracy-ExplainabilityTrade-Off

Neural Information Processing Systems

To address this, we propose Concept Embedding Models, a novel family of concept bottleneck models which goes beyond the current accuracy-vs-interpretability trade-off by learning interpretable highdimensional conceptrepresentations.




PAC-BayesianBoundfortheConditionalValueat Risk

Neural Information Processing Systems

Standard concentration inequalities are well suited for learning problems where the goal is to minimize the expected riskE[`(h,X)]. However, the expected risk--the mean performance of an algorithm--might fail to capture the underlying phenomenon of interest.




SupplementaryMaterialforthePaper: Digraph InceptionConvolutionalNetworks

Neural Information Processing Systems

Meanwhile,adding self-loops makes the greatest common divisor of the lengths of graph'scycles is 1. Clearly,πappr is upper bounded by πappr 1. To support the reproducibility of the results in this paper, we detail datasets, the baseline settings pseudocode and model implementation in experiments. In this paper, we usemean as its aggregator since it performs best [7].