Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–29 of 29 results for author: Zong, S

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.08749  [pdf, ps, other

    cs.RO

    OnEvoMemory: Evolving Memory through Online Robot Rollouts for Pretrained Robot Policies

    Authors: Zhongxi Chen, Shenqi Zong

    Abstract: Long-horizon robot manipulation requires policies to track completed subtasks and critical interaction events. However, existing memory mechanisms heavily rely on external models or predefined update rules. To address this, we propose OnEvoMemory, a value-guided memory module for pretrained robot policies. It maintains recent context, high-value experiences, and salient transitions, while learning… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

    Comments: 6 pages, 1 figure. Accepted as a poster at the ECCV 2026 Workshop on Embodied Multimodal Reasoning in Physical Environments (EMR)

  2. arXiv:2606.12240  [pdf, ps, other

    cs.LG cs.AI

    Multi-Rate Mixture of Experts for Accelerating Liquid Neural Network Training

    Authors: Shilong Zong, Almuatazbellah Boker, Hoda Eldardiry

    Abstract: Multivariate time-series data often exhibit complex temporal dependencies, irregular sampling, and heterogeneous dynamics across multiple time scales, making accurate sequence modeling particularly challenging. Traditional recurrent neural networks (RNNs), such as Long Short-Term Memory (LSTM) networks, operate in discrete time and may struggle to effectively capture continuous and irregular tempo… ▽ More

    Submitted 10 June, 2026; originally announced June 2026.

  3. arXiv:2606.05868  [pdf, ps, other

    cs.CL

    YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition

    Authors: PSBC LLM Team, Huawei LLM Team, Ruihan Long, Junjie Wu, Tianan Zhang, Duo Zhang, Yaozong Wu, Jinbin Fu, Chang Liu, Zhentao Tang, Wenshuang Yang, Xin Wang, Zhihao Song, Ning Huang, Wenjing Xu, Shuai Zong, Shupei Sun, Sen Wang, Jing Hu, Bin Wang, Xinyu Wang, Junkui Ju, Zequn Ding, Jie Ran, Man Luo , et al. (34 additional authors not shown)

    Abstract: Large language models (LLMs) drive significant financial innovations, yet their high-concurrency deployment is severely bottlenecked by KV cache memory overhead, which inflates infrastructure costs and throttles scalability. To address this, we propose YouZhi-LLM, a highly efficient financial LLM empowered by a comprehensive structural transition and training pipeline natively built on the Huawei… ▽ More

    Submitted 4 June, 2026; originally announced June 2026.

  4. arXiv:2512.07882  [pdf

    physics.med-ph cs.AI

    Referenceless Proton Resonance Frequency Thermometry Using Deep Learning with Self-Attention

    Authors: Yueran Zhao, Chang-Sheng Mei, Nathan J. McDannold, Shenyan Zong, Guofeng Shen

    Abstract: Background: Accurate proton resonance frequency (PRF) MR thermometry is essential for monitoring temperature rise during thermal ablation with high intensity focused ultrasound (FUS). Conventional referenceless methods such as complex field estimation (CFE) and phase finite difference (PFD) tend to exhibit errors when susceptibility-induced phase discontinuities occur at tissue interfaces.

    Submitted 27 November, 2025; originally announced December 2025.

  5. arXiv:2510.23163  [pdf, ps, other

    cs.CL cs.AI

    Beyond Direct Generation: A Decomposed Approach to Well-Crafted Screenwriting with LLMs

    Authors: Hang Lei, Shengyi Zong, Zhaoyan Li, Ziren Zhou, Hao Liu, Liang Yu

    Abstract: The screenplay serves as the foundation for television production, defining narrative structure, character development, and dialogue. While Large Language Models (LLMs) show great potential in creative writing, direct end-to-end generation approaches often fail to produce well-crafted screenplays. We argue this failure stems from forcing a single model to simultaneously master two disparate capabi… ▽ More

    Submitted 7 January, 2026; v1 submitted 27 October, 2025; originally announced October 2025.

    ACM Class: I.2.0

  6. arXiv:2510.07578  [pdf, ps, other

    cs.LG cs.AI eess.SY

    Accuracy, Memory Efficiency and Generalization: A Comparative Study on Liquid Neural Networks and Recurrent Neural Networks

    Authors: Shilong Zong, Alex Bierly, Almuatazbellah Boker, Hoda Eldardiry

    Abstract: This review aims to conduct a comparative analysis of liquid neural networks (LNNs) and traditional recurrent neural networks (RNNs) and their variants, such as long short-term memory networks (LSTMs) and gated recurrent units (GRUs). The core dimensions of the analysis include model accuracy, memory efficiency, and generalization ability. By systematically reviewing existing research, this paper… ▽ More

    Submitted 8 October, 2025; originally announced October 2025.

    Comments: 13 pages, 12 figures. Submitted to IEEE Transactions on Neural Networks and Learning Systems (TNNLS)

    ACM Class: I.2.6; I.2.8

  7. arXiv:2509.21732  [pdf, ps, other

    cs.CL

    How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts?

    Authors: Xiliang Zhu, Shi Zong, David Rossouw

    Abstract: Deploying Large Language Models (LLMs) for question answering (QA) over lengthy contexts is a significant challenge. In industrial settings, this process is often hindered by high computational costs and latency, especially when multiple questions must be answered based on the same context. In this work, we explore the capabilities of LLMs to answer multiple questions based on the same conversatio… ▽ More

    Submitted 25 September, 2025; originally announced September 2025.

    Comments: Accepted by EMNLP 2025 Industry Track

  8. arXiv:2407.03308  [pdf, other

    physics.med-ph cs.AI eess.IV

    Accelerated Proton Resonance Frequency-based Magnetic Resonance Thermometry by Optimized Deep Learning Method

    Authors: Sijie Xu, Shenyan Zong, Chang-Sheng Mei, Guofeng Shen, Yueran Zhao, He Wang

    Abstract: Proton resonance frequency (PRF) based MR thermometry is essential for focused ultrasound (FUS) thermal ablation therapies. This work aims to enhance temporal resolution in dynamic MR temperature map reconstruction using an improved deep learning method. The training-optimized methods and five classical neural networks were applied on the 2-fold and 4-fold under-sampling k-space data to reconstruc… ▽ More

    Submitted 3 July, 2024; originally announced July 2024.

  9. arXiv:2406.18762  [pdf, other

    cs.CL

    Categorical Syllogisms Revisited: A Review of the Logical Reasoning Abilities of LLMs for Analyzing Categorical Syllogism

    Authors: Shi Zong, Jimmy Lin

    Abstract: There have been a huge number of benchmarks proposed to evaluate how large language models (LLMs) behave for logic inference tasks. However, it remains an open question how to properly evaluate this ability. In this paper, we provide a systematic overview of prior works on the logical reasoning ability of LLMs for analyzing categorical syllogisms. We first investigate all the possible variations f… ▽ More

    Submitted 11 December, 2024; v1 submitted 26 June, 2024; originally announced June 2024.

    Comments: camera-ready version

  10. arXiv:2308.01143  [pdf, other

    cs.CV cs.CL

    ADS-Cap: A Framework for Accurate and Diverse Stylized Captioning with Unpaired Stylistic Corpora

    Authors: Kanzhi Cheng, Zheng Ma, Shi Zong, Jianbing Zhang, Xinyu Dai, Jiajun Chen

    Abstract: Generating visually grounded image captions with specific linguistic styles using unpaired stylistic corpora is a challenging task, especially since we expect stylized captions with a wide variety of stylistic patterns. In this paper, we propose a novel framework to generate Accurate and Diverse Stylized Captions (ADS-Cap). Our ADS-Cap first uses a contrastive learning module to align the image an… ▽ More

    Submitted 2 August, 2023; originally announced August 2023.

    Comments: Accepted at Natural Language Processing and Chinese Computing (NLPCC) 2022

  11. arXiv:2305.08271  [pdf, other

    cs.CL

    $SmartProbe$: A Virtual Moderator for Market Research Surveys

    Authors: Josh Seltzer, Jiahua Pan, Kathy Cheng, Yuxiao Sun, Santosh Kolagati, Jimmy Lin, Shi Zong

    Abstract: Market research surveys are a powerful methodology for understanding consumer perspectives at scale, but are limited by depth of understanding and insights. A virtual moderator can introduce elements of qualitative research into surveys, developing a rapport with survey participants and dynamically asking probing questions, ultimately to elicit more useful information for market researchers. In th… ▽ More

    Submitted 14 May, 2023; originally announced May 2023.

  12. arXiv:2301.07006  [pdf, other

    cs.CL

    Which Model Shall I Choose? Cost/Quality Trade-offs for Text Classification Tasks

    Authors: Shi Zong, Josh Seltzer, Jiahua, Pan, Kathy Cheng, Jimmy Lin

    Abstract: Industry practitioners always face the problem of choosing the appropriate model for deployment under different considerations, such as to maximize a metric that is crucial for production, or to reduce the total cost given financial concerns. In this work, we focus on the text classification task and present a quantitative analysis for this challenge. Using classification accuracy as the main metr… ▽ More

    Submitted 17 January, 2023; originally announced January 2023.

  13. arXiv:2210.09550  [pdf, other

    cs.CL

    Probing Cross-modal Semantics Alignment Capability from the Textual Perspective

    Authors: Zheng Ma, Shi Zong, Mianzhi Pan, Jianbing Zhang, Shujian Huang, Xinyu Dai, Jiajun Chen

    Abstract: In recent years, vision and language pre-training (VLP) models have advanced the state-of-the-art results in a variety of cross-modal downstream tasks. Aligning cross-modal semantics is claimed to be one of the essential capabilities of VLP models. However, it still remains unclear about the inner working mechanism of alignment in VLP models. In this paper, we propose a new probing method that is… ▽ More

    Submitted 17 October, 2022; originally announced October 2022.

    Comments: Findings of EMNLP2022

  14. arXiv:2210.00434  [pdf, other

    eess.AS cs.AI cs.CL cs.MM cs.SD

    Music-to-Text Synaesthesia: Generating Descriptive Text from Music Recordings

    Authors: Zhihuan Kuang, Shi Zong, Jianbing Zhang, Jiajun Chen, Hongfu Liu

    Abstract: In this paper, we consider a novel research problem: music-to-text synaesthesia. Different from the classical music tagging problem that classifies a music recording into pre-defined categories, music-to-text synaesthesia aims to generate descriptive texts from music recordings with the same sentiment for further understanding. As existing music-related datasets do not contain the semantic descrip… ▽ More

    Submitted 7 May, 2023; v1 submitted 2 October, 2022; originally announced October 2022.

  15. Characterizing player's playing styles based on Player Vectors for each playing position in the Chinese Football Super League

    Authors: Yuesen Li, Shouxin Zong, Yanfei Shen, Zhiqiang Pu, Miguel-Ángel Gómez, Yixiong Cui

    Abstract: Characterizing playing style is important for football clubs on scouting, monitoring and match preparation. Previous studies considered a player's style as a combination of technical performances, failing to consider the spatial information. Therefore, this study aimed to characterize the playing styles of each playing position in the Chinese Football Super League (CSL) matches, integrating a rece… ▽ More

    Submitted 7 July, 2022; v1 submitted 5 May, 2022; originally announced May 2022.

    Comments: 40 pages, 5 figures, already published on Journal of Sports Sciences

    ACM Class: I.2.1

  16. arXiv:2204.09366  [pdf, other

    cs.CL

    Analyzing the Intensity of Complaints on Social Media

    Authors: Ming Fang, Shi Zong, Jing Li, Xinyu Dai, Shujian Huang, Jiajun Chen

    Abstract: Complaining is a speech act that expresses a negative inconsistency between reality and human expectations. While prior studies mostly focus on identifying the existence or the type of complaints, in this work, we present the first study in computational linguistics of measuring the intensity of complaints from text. Analyzing complaints from such perspective is particularly useful, as complaints… ▽ More

    Submitted 20 April, 2022; originally announced April 2022.

    Comments: NAACL 2022 (Findings)

  17. arXiv:2203.02932  [pdf, other

    cs.CL cs.CY

    Doctor Recommendation in Online Health Forums via Expertise Learning

    Authors: Xiaoxin Lu, Yubo Zhang, Jing Li, Shi Zong

    Abstract: Huge volumes of patient queries are daily generated on online health forums, rendering manual doctor allocation a labor-intensive task. To better help patients, this paper studies a novel task of doctor recommendation to enable automatic pairing of a patient to a doctor with relevant expertise. While most prior work in recommendation focuses on modeling target users from their past behavior, we ca… ▽ More

    Submitted 14 March, 2022; v1 submitted 6 March, 2022; originally announced March 2022.

    Comments: Accepted to ACL 2022 main conference

  18. arXiv:2112.08152  [pdf, other

    cs.CL

    Faster Nearest Neighbor Machine Translation

    Authors: Shuhe Wang, Jiwei Li, Yuxian Meng, Rongbin Ouyang, Guoyin Wang, Xiaoya Li, Tianwei Zhang, Shi Zong

    Abstract: $k$NN based neural machine translation ($k$NN-MT) has achieved state-of-the-art results in a variety of MT tasks. One significant shortcoming of $k$NN-MT lies in its inefficiency in identifying the $k$ nearest neighbors of the query representation from the entire datastore, which is prohibitively time-intensive when the datastore size is large. In this work, we propose \textbf{Faster $k… ▽ More

    Submitted 15 December, 2021; originally announced December 2021.

  19. arXiv:2110.08743  [pdf, other

    cs.CL

    GNN-LM: Language Modeling based on Global Contexts via GNN

    Authors: Yuxian Meng, Shi Zong, Xiaoya Li, Xiaofei Sun, Tianwei Zhang, Fei Wu, Jiwei Li

    Abstract: Inspired by the notion that ``{\it to copy is easier than to memorize}``, in this work, we introduce GNN-LM, which extends the vanilla neural language model (LM) by allowing to reference similar contexts in the entire training corpus. We build a directed heterogeneous graph between an input context and its semantically related neighbors selected from the training corpus, where nodes are tokens in… ▽ More

    Submitted 4 May, 2022; v1 submitted 17 October, 2021; originally announced October 2021.

    Comments: To appear at ICLR 2022

  20. arXiv:2110.07379  [pdf

    cs.CV eess.IV

    Towards Safer Transportation: a self-supervised learning approach for traffic video deraining

    Authors: Shuya Zong, Sikai Chen, Samuel Labi

    Abstract: Video monitoring of traffic is useful for traffic management and control, traffic counting, and traffic law enforcement. However, traffic monitoring during inclement weather such as rain is a challenging task because video quality is corrupted by streaks of falling rain on the video image, and this hinders reliable characterization not only of the road environment but also of road-user behavior du… ▽ More

    Submitted 11 October, 2021; originally announced October 2021.

    Comments: Under review for presentation at TRB 2022 Annual Meeting

  21. arXiv:2110.06775  [pdf

    cs.RO cs.AI eess.IV

    Using UAVs for vehicle tracking and collision risk assessment at intersections

    Authors: Shuya Zong, Sikai Chen, Majed Alinizzi, Yujie Li, Samuel Labi

    Abstract: Assessing collision risk is a critical challenge to effective traffic safety management. The deployment of unmanned aerial vehicles (UAVs) to address this issue has shown much promise, given their wide visual field and movement flexibility. This research demonstrates the application of UAVs and V2X connectivity to track the movement of road users and assess potential collisions at intersections. T… ▽ More

    Submitted 11 October, 2021; originally announced October 2021.

    Comments: Under review for presentation at TRB 2022 Annual Meeting

  22. arXiv:2110.05559  [pdf

    cs.CV

    Development and testing of an image transformer for explainable autonomous driving systems

    Authors: Jiqian Dong, Sikai Chen, Shuya Zong, Tiantian Chen, Mohammad Miralinaghi, Samuel Labi

    Abstract: In the last decade, deep learning (DL) approaches have been used successfully in computer vision (CV) applications. However, DL-based CV models are generally considered to be black boxes due to their lack of interpretability. This black box behavior has exacerbated user distrust and therefore has prevented widespread deployment DLCV models in autonomous driving tasks even though some of these mode… ▽ More

    Submitted 11 October, 2021; originally announced October 2021.

    Comments: Under review for presentation at TRB 2022 Annual Meeting

  23. arXiv:2006.07425  [pdf, other

    cs.CL

    Measuring Forecasting Skill from Text

    Authors: Shi Zong, Alan Ritter, Eduard Hovy

    Abstract: People vary in their ability to make accurate predictions about the future. Prior studies have shown that some individuals can predict the outcome of future events with consistently better accuracy. This leads to a natural question: what makes some forecasters better than others? In this paper we explore connections between the language people use to describe their predictions and their forecastin… ▽ More

    Submitted 16 June, 2020; v1 submitted 12 June, 2020; originally announced June 2020.

    Comments: Accepted at ACL 2020

  24. arXiv:2006.02567  [pdf, other

    cs.CL cs.SI

    Extracting a Knowledge Base of COVID-19 Events from Social Media

    Authors: Shi Zong, Ashutosh Baheti, Wei Xu, Alan Ritter

    Abstract: In this paper, we present a manually annotated corpus of 10,000 tweets containing public reports of five COVID-19 events, including positive and negative tests, deaths, denied access to testing, claimed cures and preventions. We designed slot-filling questions for each event type and annotated a total of 31 fine-grained slots, such as the location of events, recent travel, and close contacts. We s… ▽ More

    Submitted 9 September, 2022; v1 submitted 3 June, 2020; originally announced June 2020.

    Comments: Accepted at COLING 2022

  25. arXiv:1902.10680  [pdf, other

    cs.CL cs.CR

    Analyzing the Perceived Severity of Cybersecurity Threats Reported on Social Media

    Authors: Shi Zong, Alan Ritter, Graham Mueller, Evan Wright

    Abstract: Breaking cybersecurity events are shared across a range of websites, including security blogs (FireEye, Kaspersky, etc.), in addition to social media platforms such as Facebook and Twitter. In this paper, we investigate methods to analyze the severity of cybersecurity threats based on the language that is used to describe them online. A corpus of 6,000 tweets describing software vulnerabilities is… ▽ More

    Submitted 3 May, 2019; v1 submitted 27 February, 2019; originally announced February 2019.

    Comments: Accepted at NAACL 2019

  26. arXiv:1701.08716  [pdf, other

    cs.CY cs.LG

    Does Weather Matter? Causal Analysis of TV Logs

    Authors: Shi Zong, Branislav Kveton, Shlomo Berkovsky, Azin Ashkan, Nikos Vlassis, Zheng Wen

    Abstract: Weather affects our mood and behaviors, and many aspects of our life. When it is sunny, most people become happier; but when it rains, some people get depressed. Despite this evidence and the abundance of data, weather has mostly been overlooked in the machine learning and data science research. This work presents a causal analysis of how weather affects TV watching patterns. We show that some wea… ▽ More

    Submitted 24 March, 2017; v1 submitted 25 January, 2017; originally announced January 2017.

    Comments: Companion of the 26th International World Wide Web Conference

  27. arXiv:1606.00540  [pdf

    cs.NE cs.LG

    Multi-pretrained Deep Neural Network

    Authors: Zhen Hu, Zhuyin Xue, Tong Cui, Shiqiang Zong, Chenglong He

    Abstract: Pretraining is widely used in deep neutral network and one of the most famous pretraining models is Deep Belief Network (DBN). The optimization formulas are different during the pretraining process for different pretraining models. In this paper, we pretrained deep neutral network by different pretraining models and hence investigated the difference between DBN and Stacked Denoising Autoencoder (S… ▽ More

    Submitted 2 June, 2016; originally announced June 2016.

  28. Detecting Localized Categorical Attributes on Graphs

    Authors: Siheng Chen, Yaoqing Yang, Shi Zong, Aarti Singh, Jelena Kovačević

    Abstract: Do users from Carnegie Mellon University form social communities on Facebook? Do signal processing researchers from tightly collaborate with each other? Do Chinese restaurants in Manhattan cluster together? These seemingly different problems share a common structure: an attribute that may be localized on a graph. In other words, nodes activated by an attribute form a subgraph that can be easily se… ▽ More

    Submitted 11 February, 2017; v1 submitted 3 April, 2016; originally announced April 2016.

    Comments: Accepted for publication in the IEEE Transactions on Signal Processing

  29. arXiv:1603.05359  [pdf, ps, other

    cs.LG stat.ML

    Cascading Bandits for Large-Scale Recommendation Problems

    Authors: Shi Zong, Hao Ni, Kenny Sung, Nan Rosemary Ke, Zheng Wen, Branislav Kveton

    Abstract: Most recommender systems recommend a list of items. The user examines the list, from the first item to the last, and often chooses the first attractive item and does not examine the rest. This type of user behavior can be modeled by the cascade model. In this work, we study cascading bandits, an online learning variant of the cascade model where the goal is to recommend $K$ most attractive items f… ▽ More

    Submitted 30 June, 2016; v1 submitted 17 March, 2016; originally announced March 2016.

    Comments: Accepted to UAI 2016