Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters

Open in new window