Among all the contributing factors, the quality and selection of data is becoming increasingly recognized for its importance in training LLMs effectively.
The rapid scaling of language models is motivating research using low-bitwidth quantization. In this work, we propose a novel binarization technique for Transformers applied to machine translation (BMT), the first of its kind.