Review for NeurIPS paper: Fourier-transform-based attribution priors improve the interpretability and stability of deep learning models for genomics

Neural Information Processing Systems 

Weaknesses: The most cited methods in this space use a different pooling architecture and train multi-task for dozens of epochs. In contrast, the authors train single task and choose the model achieved after one or two epochs as best. These differences may contribute to the rapid overfitting and saliency variance. Do multi-task models trained for longer improve motif annotation? Does your method also improve motif annotation in that framework?