Communication-Efficient Multi-Device Inference Acceleration for Transformer Models

Open in new window