Teacher Retains Full Tokens, Student Merges Efficiently: TM20K for E-Commerce Sequence Modeling in Ad Recommendation
X Li, D Zheng, W Zhao, Z Zhou, J Tan, H Yang… - arXiv preprint arXiv …, 2026 - arxiv.org
X Li, D Zheng, W Zhao, Z Zhou, J Tan, H Yang, L Chen, D Wang, D Wang, X Li, H Mao…
arXiv preprint arXiv:2608.07055, 2026•arxiv.orgBenefiting from ultra-long behavior sequence modeling, existing recommender systems
bring users a better experience via simultaneously considering their long-term and short-
term interests. Nevertheless, extended sequence lengths introduce substantial burdens on
training efficiency and serving throughput. Prior approaches typically utilize search-based or
cluster-based compression on ultra-long sequences at the cost of fine-grained information,
or rely on various lightweight target attention structures incapable of sufficient sequential …
bring users a better experience via simultaneously considering their long-term and short-
term interests. Nevertheless, extended sequence lengths introduce substantial burdens on
training efficiency and serving throughput. Prior approaches typically utilize search-based or
cluster-based compression on ultra-long sequences at the cost of fine-grained information,
or rely on various lightweight target attention structures incapable of sufficient sequential …
Benefiting from ultra-long behavior sequence modeling, existing recommender systems bring users a better experience via simultaneously considering their long-term and short-term interests. Nevertheless, extended sequence lengths introduce substantial burdens on training efficiency and serving throughput. Prior approaches typically utilize search-based or cluster-based compression on ultra-long sequences at the cost of fine-grained information, or rely on various lightweight target attention structures incapable of sufficient sequential feature extraction. In this paper, we balance the effectiveness and efficiency for ultra-long sequence modeling via full transformer modeling accompanied with a two-stage knowledge distillation framework. First, both teacher and student models take the full attention mechanism rather than pure target-sequence attention for effective sequence scaling. For student models, we propose several simple yet well-motivated token merge approaches, significantly compressing the sequence length while maintaining an acceptable performance. Then, a one-time teacher is heavily trained with full sequence tokens, further boosting the performance of student models via knowledge distillation. The proposed paradigm named TM20K has been successfully deployed in ByteDance's e-commerce advertising recommender system that extends the e-commerce sequence length to 20K, delivering substantial improvements in key business metrics (e.g., ADSS +1.036\%) while keeping the training and serving cost nearly the same as the online state-of-the-art model (e.g., serving latency only +5.6\%).
arxiv.org