Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–1 of 1 results for author: Delfin, G

Searching in archive cs. Search in all archives.
.
  1. arXiv:2508.08192  [pdf, ps, other

    cs.CL

    Efficient Speculative Decoding for Llama at Scale: Challenges and Solutions

    Authors: Bangsheng Tang, Carl Chengyan Fu, Fei Kou, Grigory Sizov, Haoci Zhang, Jason Park, Jiawen Liu, Jie You, Qirui Yang, Sachin Mehta, Shengyong Cai, Xiaodong Wang, Xingyu Liu, Yunlu Li, Yanjun Zhou, Wei Wei, Zhiwei Zhao, Zixi Qi, Adolfo Victoria, Aya Ibrahim, Bram Wasti, Changkyu Kim, Daniel Haziza, Fei Sun, Giancarlo Delfin , et al. (13 additional authors not shown)

    Abstract: Speculative decoding is a standard method for accelerating the inference speed of large language models. However, scaling it for production environments poses several engineering challenges, including efficiently implementing different operations (e.g., tree attention and multi-round speculative decoding) on GPU. In this paper, we detail the training and inference optimization techniques that we h… ▽ More

    Submitted 11 August, 2025; originally announced August 2025.

    Comments: 15 pages