habana accelerator
Supercharge your training with zero code changes using Intel's Habana Accelerator
We recently added support for Habana's Gaudi AI Processors, which can be used to accelerate deep learning training workloads. Habana Gaudi was designed from the ground up to maximize training throughput and efficiency. The processors are built on a heterogeneous architecture with a cluster of fully programmable Tensor Processing Cores (TPC), along with its associated development tools and libraries and a configurable Matrix Math engine. The TPC core is a VLIW SIMD processor with an instruction set and hardware tailored to efficiently serve training workloads. The Gaudi memory architecture includes on-die SRAM and local memories in each TPC, and Gaudi is the first DL training processor that has integrated RDMA over Converged Ethernet (RoCE v2) engines on-chip. On the software side, the PyTorch-Habana bridge interfaces between the framework and the SynapseAI software stack to enable the execution of deep learning models on the Habana Gaudi device.