Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–26 of 26 results for author: Chi, P

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.18016  [pdf, ps, other

    cs.RO

    Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation

    Authors: Peng Ren, Haoyang Ge, Jiang Zhao, Cong Huang, Yukun Shi, Pei Chi, Kai Chen

    Abstract: Vision-language-action policies are a promising foundation for general robot control, but long-horizon humanoid loco-manipulation requires the robot to treat task objects as persistent physical entities across movement, contact, occlusion, and recovery. We study this problem as object-state divergence: the object state used to condition a whole-body action can differ from the state used to decide… ▽ More

    Submitted 26 July, 2026; v1 submitted 20 July, 2026; originally announced July 2026.

  2. arXiv:2607.15907  [pdf, ps, other

    cs.LG

    A Semiparametric Framework for Stochastic Fundamental Diagram Modeling

    Authors: Pengnan Chi, Xiaoliang Ma, Magnus Jansson, Magnus Nordenvaad

    Abstract: The stochastic fundamental diagram (SFD) provides a probabilistic description of the relationship between traffic density and flow or speed, enabling uncertainty-aware traffic modeling. However, existing stochastic models frequently struggle to accommodate rigorous physical constraints while retaining sufficient flexibility to capture complex nonlinear patterns. To address this, we propose a novel… ▽ More

    Submitted 17 July, 2026; originally announced July 2026.

    Comments: 22 pages, 11 figures, 5 tables

  3. arXiv:2606.21851  [pdf, ps, other

    cs.CL

    TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding Distillation

    Authors: Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van, Nguyen Thi Ngoc Diep, Thien Huu Nguyen, Trung Le

    Abstract: Knowledge Distillation (KD) has established itself as a pivotal technique for compressing large pre-trained language models. However, existing methods that force a student to strictly mimic the teacher's sentence embeddings or internal features often incur prohibitive computational costs and yield suboptimal performance due to the inherent capacity gap. To address these challenges, we propose TALA… ▽ More

    Submitted 19 June, 2026; originally announced June 2026.

    Comments: ACL 2026

  4. arXiv:2605.01374  [pdf, ps, other

    cs.CL

    MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation

    Authors: Pham Khanh Chi, Quoc Phong Dao, Thuat Nguyen, Linh Ngo Van, Trung Le, Thanh Hong Nguyen

    Abstract: Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers or token-level outputs, ignoring how representations evolve across depth. As a result, the student is only weakly guided to capture the teacher's internal relational structure during distillation, which limits knowledge transfer. To address this li… ▽ More

    Submitted 2 June, 2026; v1 submitted 2 May, 2026; originally announced May 2026.

    Comments: ACL 2026

  5. arXiv:2605.01205  [pdf, ps, other

    cs.CL

    SRA: Span Representation Alignment for Large Language Model Distillation

    Authors: Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen, Linh Ngo Van, Nguyen Thi Ngoc Diep, Trung Le

    Abstract: Cross-Tokenizer Knowledge Distillation (CTKD) enables knowledge transfer between a large language model and a smaller student, even when they employ different tokenizers. While existing approaches mainly focus on token-level alignment strategies, which are often brittle and sensitive to discrepancies between tokenizers, we argue that the method of aggregating tokens into more robust representation… ▽ More

    Submitted 2 June, 2026; v1 submitted 1 May, 2026; originally announced May 2026.

    Comments: ACL 2026

  6. arXiv:2604.15336  [pdf, ps, other

    cs.HC cs.AI

    Facial-Expression-Aware Prompting for Empathetic LLM Tutoring

    Authors: Shuangquan Feng, Laura Fleig, Ruisen Tu, Philip Chi, Edmund Bu, Melinda Ozel, Junhua Ma, Teng Fei, Virginia R. de Sa

    Abstract: Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone. Facial expressions provide immediate and practical cues of confusion, frustration, or engagement, but remain underexplored in LLM-driven tutoring. We investigate whether facial-expression-aware sign… ▽ More

    Submitted 28 July, 2026; v1 submitted 10 March, 2026; originally announced April 2026.

  7. arXiv:2603.14605  [pdf, ps, other

    cs.RO

    CyboRacket: A Perception-to-Action Framework for Humanoid Racket Sports

    Authors: Peng Ren, Chuan Qi, Haoyang Ge, Qiyuan Su, Xuguo He, Cong Huang, Pei Chi, Jiang Zhao, Kai Chen

    Abstract: Dynamic ball-interaction tasks remain challenging for robots because they require tight perception-action coupling under limited reaction time. This challenge is especially pronounced in humanoid racket sports, where successful interception depends on accurate visual tracking, trajectory prediction, coordinated stepping, and stable whole-body striking. Existing robotic racket-sport systems often r… ▽ More

    Submitted 15 March, 2026; originally announced March 2026.

  8. arXiv:2603.10675  [pdf, ps, other

    cs.RO

    Cybo-Waiter: A Physical Agentic Framework for Humanoid Whole-Body Locomotion-Manipulation

    Authors: Peng Ren, Haoyang Ge, Chuan Qi, Cong Huang, Hong Li, Jiang Zhao, Pei Chi, Kai Chen

    Abstract: Robots are increasingly expected to execute open ended natural language requests in human environments, which demands reliable long horizon execution under partial observability. This is especially challenging for humanoids because locomotion and manipulation are tightly coupled through stance, reachability, and balance. We present a humanoid agent framework that turns VLM plans into verifiable ta… ▽ More

    Submitted 11 March, 2026; originally announced March 2026.

  9. arXiv:2603.07644  [pdf, ps, other

    cs.RO

    PanoDP: Learning Collision-Free Navigation with Panoramic Depth and Differentiable Physics

    Authors: Hao Zhong, Pei Chi, Jiang Zhao, Shenghai Yuan, Xuyang Gao, Thien-Minh Nguyen, Lihua Xie

    Abstract: Autonomous collision-free navigation in cluttered environments requires safe decision-making under partial observability with both static structure and dynamic obstacles. We present \textbf{PanoDP}, a communication-free learning framework that combines four-view panoramic depth perception with differentiable-physics-based training signals. PanoDP encodes panoramic depth using a lightweight CNN and… ▽ More

    Submitted 8 March, 2026; originally announced March 2026.

  10. arXiv:2602.21669  [pdf, ps, other

    cs.CL

    DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillation

    Authors: Duc Trung Vu, Pham Khanh Chi, Dat Phi Van, Linh Ngo Van, Sang Dinh, Trung Le

    Abstract: Knowledge Distillation (KD) has emerged as a crucial technique for compressing Large Language Models (LLMs). Although existing cross-tokenizer KD methods have made notable progress, their effectiveness remains constrained by suboptimal alignment across sequence and vocabulary levels. To address these limitations, we introduce Dual-Space Weighting and Time-Warped Alignment (DWA-KD), a novel cross-t… ▽ More

    Submitted 25 February, 2026; originally announced February 2026.

    Comments: EACL Findings

  11. arXiv:2510.00129  [pdf, ps, other

    cs.LG cond-mat.mtrl-sci cs.AI physics.comp-ph

    BigBang-Proton Technical Report: Next-Word-Prediction is Scientific Multitask Learner

    Authors: Hengkui Wu, Liujiang Liu, Jihua He, Qihao Wang, Keke Zhao, Shuyang Hu, Renle Fu, Dahao Liang, Lingyu Zeng, Bruce Liu, Yuan Liu, Jin Zhan, Jiaqiang Niu, Xinglong Jia, Yaqin Hu, Wenjun Ji, Panpan Chi, Ken Chen, Hengyuan Wu, Yingsi Xin, Yongfeng Zhu, Yuexin Wang, Manqi Ruan, Ningtao Bian, Xiaohua Wu , et al. (1 additional authors not shown)

    Abstract: We introduce BigBang-Proton, a unified sequence-based architecture for auto-regressive language modeling pretrained on cross-scale, cross-structure, cross-discipline real-world scientific tasks to construct a scientific multi-task learner. BigBang-Proton incorporates three fundamental innovations compared to mainstream general-purpose LLMs: Theory-Experiment Learning paradigm aligns large-scale nu… ▽ More

    Submitted 30 September, 2025; originally announced October 2025.

    Comments: 93 pages, 39 figures

    MSC Class: 68T05; 68T50; 00A69; 94A99 ACM Class: I.2.6; I.2.7; J.2; I.6.3; K.4.1

  12. arXiv:2412.00129  [pdf, other

    cs.LG hep-ex physics.data-an

    Scaling Particle Collision Data Analysis

    Authors: Hengkui Wu, Panpan Chi, Yongfeng Zhu, Liujiang Liu, Shuyang Hu, Yuexin Wang, Chen Zhou, Qihao Wang, Yingsi Xin, Bruce Liu, Dahao Liang, Xinglong Jia, Manqi Ruan

    Abstract: For decades, researchers have developed task-specific models to address scientific challenges across diverse disciplines. Recently, large language models (LLMs) have shown enormous capabilities in handling general tasks; however, these models encounter difficulties in addressing real-world scientific problems, particularly in domains involving large-scale numerical data analysis, such as experimen… ▽ More

    Submitted 9 December, 2024; v1 submitted 28 November, 2024; originally announced December 2024.

  13. arXiv:2404.09385  [pdf, other

    eess.AS cs.CL eess.SP

    A Large-Scale Evaluation of Speech Foundation Models

    Authors: Shu-wen Yang, Heng-Jui Chang, Zili Huang, Andy T. Liu, Cheng-I Lai, Haibin Wu, Jiatong Shi, Xuankai Chang, Hsiang-Sheng Tsai, Wen-Chin Huang, Tzu-hsun Feng, Po-Han Chi, Yist Y. Lin, Yung-Sung Chuang, Tzu-Hsien Huang, Wei-Cheng Tseng, Kushal Lakhotia, Shang-Wen Li, Abdelrahman Mohamed, Shinji Watanabe, Hung-yi Lee

    Abstract: The foundation model paradigm leverages a shared foundation model to achieve state-of-the-art (SOTA) performance for various tasks, requiring minimal downstream-specific modeling and data annotation. This approach has proven crucial in the field of Natural Language Processing (NLP). However, the speech processing community lacks a similar setup to explore the paradigm systematically. In this work,… ▽ More

    Submitted 29 May, 2024; v1 submitted 14 April, 2024; originally announced April 2024.

    Comments: The extended journal version for SUPERB and SUPERB-SG. Published in IEEE/ACM TASLP. The Arxiv version is preferred

  14. arXiv:2310.02383  [pdf, other

    cs.HC

    Automatic Multi-Path Web Story Creation from a Structural Article

    Authors: Daniel Nkemelu, Peggy Chi, Daniel Castro Chin, Krishna Srinivasan, Irfan Essa

    Abstract: Web articles such as Wikipedia serve as one of the major sources of knowledge dissemination and online learning. However, their in-depth information--often in a dense text format--may not be suitable for mobile browsing, even in a responsive UI. We propose an automatic approach that converts a structural article of any length into a set of interactive Web Stories that are ideal for mobile experien… ▽ More

    Submitted 3 October, 2023; originally announced October 2023.

  15. ConceptEVA: Concept-Based Interactive Exploration and Customization of Document Summaries

    Authors: Xiaoyu Zhang, Jianping Li, Po-Wei Chi, Senthil Chandrasegaran, Kwan-Liu Ma

    Abstract: With the most advanced natural language processing and artificial intelligence approaches, effective summarization of long and multi-topic documents -- such as academic papers -- for readers from different domains still remains a challenge. To address this, we introduce ConceptEVA, a mixed-initiative approach to generate, evaluate, and customize summaries for long and multi-topic documents. Concep… ▽ More

    Submitted 31 March, 2023; originally announced March 2023.

    Comments: 16 pages, 7 figures

  16. arXiv:2206.03931  [pdf, other

    cs.CL cs.AI cs.LG

    Learning to Generate Prompts for Dialogue Generation through Reinforcement Learning

    Authors: Hsuan Su, Pohan Chi, Shih-Cheng Huang, Chung Ho Lam, Saurav Sahay, Shang-Tse Chen, Hung-yi Lee

    Abstract: Much literature has shown that prompt-based learning is an efficient method to make use of the large pre-trained language model. Recent works also exhibit the possibility of steering a chatbot's output by plugging in an appropriate prompt. Gradient-based methods are often used to perturb the prompts. However, some language models are not even available to the public. In this work, we first explore… ▽ More

    Submitted 13 October, 2022; v1 submitted 8 June, 2022; originally announced June 2022.

  17. arXiv:2112.00344  [pdf, other

    q-bio.QM cs.AI cs.LG q-bio.BM

    Leveraging Sequence Embedding and Convolutional Neural Network for Protein Function Prediction

    Authors: Wei-Cheng Tseng, Po-Han Chi, Jia-Hua Wu, Min Sun

    Abstract: The capability of accurate prediction of protein functions and properties is essential in the biotechnology industry, e.g. drug development and artificial protein synthesis, etc. The main challenges of protein function prediction are the large label space and the lack of labeled training data. Our method leverages unsupervised sequence embedding and the success of deep convolutional neural network… ▽ More

    Submitted 1 December, 2021; originally announced December 2021.

    Comments: Published in NeurIPS 2018 Machine Learning for Molecules and Materials Workshop

  18. HelpViz: Automatic Generation of Contextual Visual MobileTutorials from Text-Based Instructions

    Authors: Mingyuan Zhong, Gang Li, Peggy Chi, Yang Li

    Abstract: We present HelpViz, a tool for generating contextual visual mobile tutorials from text-based instructions that are abundant on the web. HelpViz transforms text instructions to graphical tutorials in batch, by extracting a sequence of actions from each text instruction through an instruction parsing model, and executing the extracted actions on a simulation infrastructure that manages an array of A… ▽ More

    Submitted 6 August, 2021; originally announced August 2021.

    Comments: Accepted to UIST'21

  19. arXiv:2105.06988  [pdf, other

    cs.CV

    Automatic Non-Linear Video Editing Transfer

    Authors: Nathan Frey, Peggy Chi, Weilong Yang, Irfan Essa

    Abstract: We propose an automatic approach that extracts editing styles in a source video and applies the edits to matched footage for video creation. Our Computer Vision based techniques considers framing, content type, playback speed, and lighting of each input video segment. By applying a combination of these features, we demonstrate an effective method that automatically transfers the visual and tempora… ▽ More

    Submitted 14 May, 2021; originally announced May 2021.

    Comments: Published to AI for Content Creation Workshop at CVPR 2021

    Journal ref: AI for Content Creation Workshop at CVPR 2021

  20. arXiv:2105.03070  [pdf, other

    cs.CL cs.LG cs.SD eess.AS

    SpeechNet: A Universal Modularized Model for Speech Processing Tasks

    Authors: Yi-Chen Chen, Po-Han Chi, Shu-wen Yang, Kai-Wei Chang, Jheng-hao Lin, Sung-Feng Huang, Da-Rong Liu, Chi-Liang Liu, Cheng-Kuang Lee, Hung-yi Lee

    Abstract: There is a wide variety of speech processing tasks ranging from extracting content information from speech signals to generating speech signals. For different tasks, model networks are usually designed and tuned separately. If a universal model can perform multiple speech processing tasks, some tasks might be improved with the related abilities learned from other tasks. The multi-task learning of… ▽ More

    Submitted 31 May, 2021; v1 submitted 7 May, 2021; originally announced May 2021.

  21. arXiv:2105.01051  [pdf, ps, other

    cs.CL cs.SD eess.AS

    SUPERB: Speech processing Universal PERformance Benchmark

    Authors: Shu-wen Yang, Po-Han Chi, Yung-Sung Chuang, Cheng-I Jeff Lai, Kushal Lakhotia, Yist Y. Lin, Andy T. Liu, Jiatong Shi, Xuankai Chang, Guan-Ting Lin, Tzu-Hsien Huang, Wei-Cheng Tseng, Ko-tik Lee, Da-Rong Liu, Zili Huang, Shuyan Dong, Shang-Wen Li, Shinji Watanabe, Abdelrahman Mohamed, Hung-yi Lee

    Abstract: Self-supervised learning (SSL) has proven vital for advancing research in natural language processing (NLP) and computer vision (CV). The paradigm pretrains a shared model on large volumes of unlabeled data and achieves state-of-the-art (SOTA) for various tasks with minimal adaptation. However, the speech processing community lacks a similar setup to systematically explore the paradigm. To bridge… ▽ More

    Submitted 15 October, 2021; v1 submitted 3 May, 2021; originally announced May 2021.

    Comments: To appear in Interspeech 2021

  22. arXiv:2006.05174  [pdf, other

    eess.AS cs.CL cs.SD

    Input-independent Attention Weights Are Expressive Enough: A Study of Attention in Self-supervised Audio Transformers

    Authors: Tsung-Han Wu, Chun-Chen Hsieh, Yen-Hao Chen, Po-Han Chi, Hung-yi Lee

    Abstract: In this paper, we seek solutions for reducing the computation complexity of transformer-based models for speech representation learning. We evaluate 10 attention algorithms; then, we pre-train the transformer-based model with those attention algorithms in a self-supervised fashion and treat them as feature extractors on downstream tasks, including phoneme classification and speaker classification.… ▽ More

    Submitted 3 November, 2020; v1 submitted 9 June, 2020; originally announced June 2020.

  23. arXiv:2005.08575  [pdf, other

    eess.AS cs.CL cs.SD

    Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation

    Authors: Po-Han Chi, Pei-Hung Chung, Tsung-Han Wu, Chun-Cheng Hsieh, Yen-Hao Chen, Shang-Wen Li, Hung-yi Lee

    Abstract: For self-supervised speech processing, it is crucial to use pretrained models as speech representation extractors. In recent works, increasing the size of the model has been utilized in acoustic model training in order to achieve better performance. In this paper, we propose Audio ALBERT, a lite version of the self-supervised speech representation model. We use the representations with two downstr… ▽ More

    Submitted 3 May, 2021; v1 submitted 18 May, 2020; originally announced May 2020.

    Comments: Accepted by IEEE Spoken Language Technology Workshop 2021

  24. arXiv:2001.09309  [pdf, other

    cs.CL cs.LG

    BERT's output layer recognizes all hidden layers? Some Intriguing Phenomena and a simple way to boost BERT

    Authors: Wei-Tsung Kao, Tsung-Han Wu, Po-Han Chi, Chun-Cheng Hsieh, Hung-Yi Lee

    Abstract: Although Bidirectional Encoder Representations from Transformers (BERT) have achieved tremendous success in many natural language processing (NLP) tasks, it remains a black box. A variety of previous works have tried to lift the veil of BERT and understand each layer's functionality. In this paper, we found that surprisingly the output layer of BERT can reconstruct the input sentence by directly t… ▽ More

    Submitted 15 February, 2021; v1 submitted 25 January, 2020; originally announced January 2020.

    Comments: 7 pages, 8 figures, 3 tables

  25. arXiv:1910.12638  [pdf, other

    eess.AS cs.CL cs.LG cs.SD

    Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders

    Authors: Andy T. Liu, Shu-wen Yang, Po-Han Chi, Po-chun Hsu, Hung-yi Lee

    Abstract: We present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech. Previous speech representation methods learn through conditioning on past frames and predicting information about future frames. Whereas Mockingjay is designed to predict the current frame through jointly conditioning on both past a… ▽ More

    Submitted 2 February, 2020; v1 submitted 24 October, 2019; originally announced October 2019.

    Comments: Accepted by ICASSP 2020, Lecture Session

    Journal ref: ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

  26. arXiv:1304.3081  [pdf

    cs.AI

    Predicting The Performance of Minimax and Product in Game-Tree

    Authors: Ping-Chung Chi, Dana Nau

    Abstract: The discovery that the minimax decision rule performs poorly in some games has sparked interest in possible alternatives to minimax. Until recently, the only games in which minimax was known to perform poorly were games which were mainly of theoretical interest. However, this paper reports results showing poor performance of minimax in a more common game called kalah. For the kalah games tested, a… ▽ More

    Submitted 27 March, 2013; originally announced April 2013.

    Comments: Appears in Proceedings of the Second Conference on Uncertainty in Artificial Intelligence (UAI1986)

    Report number: UAI-P-1986-PG-49-56