Goto

Collaborating Authors

 Deep Learning



AsCAN: AsymmetricConvolution-AttentionNetworks forEfficientRecognitionandGeneration

Neural Information Processing Systems

Tosatisfy that, architectures must provide promising latency and performance trade-offs, support a variety of tasks, scale efficiently with respect to the amounts of data and compute, leverage available data from other tasks, and efficiently support various hardware.


CiD 2: Accelerating Asynchronous Communication in Decentralized Deep Learning

Neural Information Processing Systems

Distributed training of Deep Learning models has been critical to many recent successes in the field. Current standard methods primarily rely on synchronous centralized algorithms which induce major communication bottlenecks and synchronization locks at scale.