SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices

Open in new window