Goto

Collaborating Authors

 Deep Learning


2052b3e0617ecb2ce9474a6feaf422b3-Paper-Datasets_and_Benchmarks.pdf

Neural Information Processing Systems

Textual backdoor attacks are a kind of practical threat to NLP systems. By injecting a backdoor in the training phase, the adversary could control model predictions via predefined triggers.



BidirectionalConvolutionalPoissonGamma DynamicalSystems

Neural Information Processing Systems

Incorporating the natural document-sentence-word structure into hierarchical Bayesian modeling, we propose convolutional Poisson gamma dynamical systems (PGDS) that introduce not only word-level probabilistic convolutions, but alsosentence-levelstochastic temporaltransitions.


TowardsCrowdsourcedTrainingofLargeNeural NetworksusingDecentralizedMixture-of-Experts SupplementaryMaterial

Neural Information Processing Systems

With this data structure, DMoE can use beam search toselect the best experts. Manypopular architectures, including Transformers, can train entirely in that precision mode [7]. In addition, the deep learning architectures discussed in this work rely on backpropagation for training.


TowardsCrowdsourcedTrainingofLargeNeural NetworksusingDecentralizedMixture-of-Experts

Neural Information Processing Systems

Many recent breakthroughs in deep learning were achieved by training increasingly larger models on massivedatasets. However,training such models can be prohibitively expensive. For instance, the cluster used to train GPT-3 costs over $250 million2. Asaresult, most researchers cannot afford totrain state oftheart models and contribute to their development.




Cross-ScaleSelf-SupervisedBlindImageDeblurring viaImplicitNeuralRepresentation

Neural Information Processing Systems

Blind image deblurring (BID) is an important yet challenging image recovery problem. Most existing deep learning methods require supervised training with ground truth (GT) images. This paper introduces a self-supervised method for BID that does not require GT images. The key challenge is to regularize the training to prevent over-fitting due to the absence of GT images. By leveraging an exact relationship among the blurred image, latent image, and blur kernel across consecutive scales, we propose an effective cross-scale consistency loss.