Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 335 results for author: Jin, Q

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.18076  [pdf, ps, other

    cs.CV cs.AI

    From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

    Authors: Xingjian Wang, Zhao Wang, Taihang Hu, Jun Zheng, Qing Jin, Qinye Zhou, Zhengtao Wu, Yongchao Du, Zuan Gao, Chao Lin, Yefeng Shen, Xiaoli Xu, Zhengze Xu, Hao Yan, Yuhang Yu, Mingzhou Zhang, Mengting Chen

    Abstract: Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines typically optimize task-specific datasets in isolation. A central challenge is not only how to curate each task-specific corpus, but also how to organize heterogeneous supervision according to the dependencies among generative capabilities. We present a \textbf… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

    Comments: 19 pages, 10 figures

  2. A Survey of Large Models in Sports

    Authors: Yichen Xu, Jianzhe Ma, Chuhan Wang, Zhonghao Cao, Liangyu Chen, Wenxuan Wang, Qin Jin

    Abstract: Sports have witnessed growing global enthusiasm in recent years, serving as a vital force for physical health, cultural exchange, social connection, and economic growth. The rapid advancement of large models, particularly (multimodal) large language models (M)LLMs, has demonstrated transformative potential to reshape sports understanding, analysis, and interaction across diverse domains. This pape… ▽ More

    Submitted 14 August, 2026; originally announced August 2026.

    Comments: 36 pages, 4 figures, 6 tables. Accepted to Findings of ACL 2026

  3. arXiv:2608.13786  [pdf, ps, other

    cs.IR cs.AI cs.CL

    Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

    Authors: Qingfang Liu, Qiao Jin, Joe D. Menke, Thorsten Kahnt, Zhiyong Lu

    Abstract: Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatG… ▽ More

    Submitted 13 August, 2026; originally announced August 2026.

  4. arXiv:2608.10562  [pdf, ps, other

    cs.LG

    MARCO: Click-Intent Decomposition for Calibrated Ads Conversion Prediction

    Authors: Shiwen Shen, Xiru Huang, Liang Luo, Jianbo Sun, He Lyu, Zihang Fu, Ivonne Xu, Zhizhuo Li, Zhengyu Zhang, Pei-Ju Sung, Yunmiao Wang, Zixuan Wang, Zhengli Zhao, Qiang Jin, Mike Jermann, Mingda Li, Yang Xiao, Bhavana Challa, Brooke Bian, Yang Li, Ashish Chamoli, Bibek Bhusal, Danning Di, Yuan Jin, Meet Raval , et al. (10 additional authors not shown)

    Abstract: Not all clicks are equal. Industrial ads ranking decouples conversion probability into click-through rate (CTR) and post-click conversion rate (CVR), yet treats every click as the same event. In reality, users provide a free, self-generated signal of intent through their physical UI interactions. Different click types on the same ad exhibit a 4-fold difference in actual conversion rates. By confla… ▽ More

    Submitted 11 August, 2026; originally announced August 2026.

  5. arXiv:2608.09819  [pdf, ps, other

    cs.LG cs.CL

    Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

    Authors: Mind Lab, :, Vin Bo, Asher Cai, Jingwei Cao, Song Cao, Vic Cao, Amelia Chen, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Pyke Han, Nolan Ho, Ori Hong, Hailee Hou, Piers Hua, Charles Huang, Miles Jiang, Nora Jiang , et al. (52 additional authors not shown)

    Abstract: Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals. Adaptation is pursued through recursive improvement of versioned model-harness pairs, where experience from one configuration is evaluated under an external contract and used to construct its success… ▽ More

    Submitted 10 August, 2026; originally announced August 2026.

    Comments: 49 pages, technical report

  6. arXiv:2608.08487  [pdf, ps, other

    cs.CV

    RenderMatte: Exact-Alpha Rendering and Group-Relative Alignment for Image Matting

    Authors: Zecheng Ren, Yafei Hu, Jianing Zhao, Ruichen Cong, Qun Jin, Yiren Song

    Abstract: Image matting is an essential enabling technology for modern visual content production, where foreground extraction determines the realism and editability of downstream creation workflows. However, precise alpha estimation in open-world scenes remains challenging because real foregrounds exhibit highly diverse appearances and opacity patterns. This makes existing methods struggle with semantic amb… ▽ More

    Submitted 9 August, 2026; originally announced August 2026.

  7. arXiv:2608.06931  [pdf, ps, other

    cs.AI

    Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery

    Authors: Taolin Han, Yuchen Zhang, Jinghang Wang, Yun Wu, Wai Yuet Chiu, Zhaohai Li, Yifei Zhang, Jinxin Wang, Yuhao Zhou, Chen Zhao, Jiajia Li, Jiaxin Li, Qile Jin, Kewei Sun, Shuang Wu, Weiqi Zhai, Renquan Lv, Junchao Li, Ruodan Chen, Qingteng Chen, Zhibo Yang, Hu Wei, Lin Qu, Shuai Bai, Bing Zhao

    Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science. Here we introduce Science Edge Evaluation (SEE), a multimodal benchmark of expert-curated questions grounded in peer-reviewed literature and experimental practice in chemistry, biology, and materials science. Evaluation of 19 multimodal la… ▽ More

    Submitted 7 August, 2026; originally announced August 2026.

  8. arXiv:2608.02580  [pdf, ps, other

    cs.RO

    Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

    Authors: Ye Wang, Pei Lin, Xiong-Hui Chen, Haoqi Yuan, Zhixuan Liang, Yiyang Huang, Anzhe Chen, Zixing Lei, Jie Zhang, Tao Zhang, Haoyang Li, Tong Zhang, Chenxi Xiao, Ziyuan Jiao, Qin Jin

    Abstract: Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such videos into robot-format data can yield effective per-task policies at small scale. However, whether this approach can provide pretraining benefits for vision-la… ▽ More

    Submitted 3 August, 2026; originally announced August 2026.

  9. arXiv:2607.18786  [pdf, ps, other

    cs.IR

    Beyond Noisy Signals: Dual-Level Denoising for Multi-modal Sequential Recommendation

    Authors: Jie Luo, Qi Jin, Xinming Zhang

    Abstract: Multi-modal Sequential Recommendation (SR) incorporates rich side information (e.g., textual and visual features) to enhance dynamic user preference modeling. However, existing frameworks inevitably suffer from a Dual-Noise Dilemma: (1) Feature-level redundancy stemming from the semantic gap between generic pre-trained representations and fine-grained recommendation intent; and (2) Sequence-level… ▽ More

    Submitted 14 August, 2026; v1 submitted 21 July, 2026; originally announced July 2026.

    Comments: Accepted by ACM MM 2026. 12 Pages

  10. arXiv:2607.13056  [pdf, ps, other

    cs.RO cs.LG

    HRIBench: Benchmarking Interaction-Centric Human-Robot Collaboration

    Authors: Chang Liu, Jiawei Zhang, Tao Zhang, Ye Wang, Hongyu Zhou, Qin Jin

    Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled. However, real-world collaboration fundamentally requires coordination under shared agency, including intent understanding, temporal synchronization, protocol adherence, and safe interaction in dynamic environments. To address this gap, w… ▽ More

    Submitted 5 July, 2026; originally announced July 2026.

  11. arXiv:2607.09165  [pdf

    cs.LG cs.AI

    A Personalized Computational Framework for Assessing the Sufficiency of Partially Observed Data in Healthcare AI models

    Authors: Qingchu Jin, Felistas Mazhude, Jamie B. Rabb, Robert S. Kramer, Douglas B. Sawyer, Raimond L. Winslow

    Abstract: Achieving early and timely diagnosis and treatment for disease is a major challenge. Recent applications of machine learning (ML) algorithms trained on patient data have shown promise in many different settings for predicting the patient health state. A challenge often faced when applying these ML algorithms is that at any given time, not all clinical variables (features) needed as input to perfor… ▽ More

    Submitted 10 July, 2026; originally announced July 2026.

  12. arXiv:2607.06986  [pdf, ps, other

    cs.SD

    MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

    Authors: Wenhao Feng, Yuxun Tang, Jiatong Shi, Qin Jin

    Abstract: Singing voice synthesis (SVS) has progressed rapidly, yet its ability to generalize across diverse musical genres remains underexplored. Existing benchmarks are heavily biased toward pop music, limiting systematic analysis of genre-dependent behavior. We introduce MMGenre, a benchmark for multi-genre SVS diagnosis, supported by an automatic pipeline for constructing genre-aligned music scores. MMG… ▽ More

    Submitted 8 July, 2026; originally announced July 2026.

    Comments: Accepted by Interspeech 2026. Camera-ready version. 4 pages, 5 figures.Project page: https://fengjin1117.github.io/mmgenre-demo/

  13. arXiv:2607.02770  [pdf, ps, other

    cs.CL cs.AI

    Gemma 4 Technical Report

    Authors: Gemma Team, Sherif El Abd, Vaibhav Aggarwal, Robin Algayres, Alek Andreev, Olivier Bachem, Ian Ballantyne, Cormac Brick, Victor Cărbune, Michelle Casbon, Mayank Chaturvedi, Aditya Chawla, Victor Cotruta, Alice Coucke, Phil Culliton, Robert Dadashi, Lucas Dixon, Mohamed Elhawaty, Utku Evci, Clément Farabet, Johan Ferret, Filippo Galgani, Sertan Girgin, Jean-Bastien Grill, Maarten Grootendorst , et al. (298 additional authors not shown)

    Abstract: We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite features dense and Mixture-of-Experts architectures, ranging from 2.3B to 31B parameters. Alongside improved vision and audio encoders for all model sizes, we propose a unified, encoder-free architecture… ▽ More

    Submitted 24 July, 2026; v1 submitted 2 July, 2026; originally announced July 2026.

    Comments: 17 pages, 2 figures, technical report, updated

  14. arXiv:2607.02606  [pdf, ps, other

    cs.SE

    ChainSWE: Benchmarking Coding Agents on Multi-Bug Software Maintenance

    Authors: Qirui Jin, Lingching Tung, Kenan Li, Qiyang Shi, Yushi She, Huanzhong Jia, Harrison Zhao, Kejing Xia, Zhenbang Du, Yikai Zhang, Jiaxin Pei, Zhenyu Zhang, Zhen Qi, Yuyan Duan, Wenke Lee, Zijian Jin

    Abstract: Language model (LM) agents are increasingly deployed to maintain codebases over extended periods, fixing streams of related defects while carrying context from one fix to the next. Yet existing software engineering (SWE) benchmarks evaluate models one bug at a time: the repository is reset, the codebase is re-read, and a single self-contained issue is graded in isolation. This setting collapses a… ▽ More

    Submitted 1 July, 2026; originally announced July 2026.

  15. arXiv:2606.28714  [pdf, ps, other

    cs.HC

    "If I Can See You": Understanding Spatially Situated Virtual Embodiment in Close Human-AI Relationships

    Authors: Yulin Chen, Yang Zhan, Qiao Jin

    Abstract: AI companions are increasingly used for emotional support, companionship, and intimate interaction. While prior work has examined text- and voice-based AI companionship and emerging XR companion designs, less is known about how users with existing close AI companion relationships expect those relationships to change when companions become virtually embodied and spatially situated in everyday envir… ▽ More

    Submitted 26 June, 2026; originally announced June 2026.

    Comments: 17 pages, 3 figures

  16. arXiv:2606.26713  [pdf, ps, other

    cs.AI

    LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography

    Authors: Yuqi Jiang, Yumeng Liu, Zimu Li, Jinyuan Deng, Qian Jin, Yucheng Cui, Yu Li, Xunzhao Yin, Qi Sun, Cheng Zhuo

    Abstract: As semiconductor technology nodes scale, computational lithography is essential for ensuring yield and performance. However, lithography is a continuous physical process involving mask optimization, optical imaging, resist exposure, and development, which existing models fail to capture. To overcome this limitation, we present LithoDreamer, the first physics-informed World Model (WM) framework for… ▽ More

    Submitted 25 June, 2026; originally announced June 2026.

    Comments: Correspondence to: Qi Sun <qisunchn \at zju \dot edu \dot cn>

    Journal ref: Proceedings of the 43rd International Conference on Machine Learning (ICML), Jul. 6-11, 2026

  17. arXiv:2606.17520  [pdf, ps, other

    cs.RO cs.CV

    GASE: Gaussian Splatting-Based Automated System for Reconstructing Embodied-Simulation Environments

    Authors: Jiawei Zhang, Yiming Yan, Chao Liang, Nuo Xu, Seson Sun, Qichen Zhang, Yuhao Xu, Yantai Yang, Yingqiao Wang, Qin Jin, Zhipeng Zhang

    Abstract: Training embodied agents in the real world requires skilled operators and expensive hardware. Simulation environments offer a compelling alternative by enabling large-scale, cost-effective data augmentation. Consequently, rapidly constructing high-fidelity simulation scenes with a minimal sim-to-real gap has become a critical objective in robot learning. While reconstruction-based methods provide… ▽ More

    Submitted 16 June, 2026; originally announced June 2026.

  18. arXiv:2606.16215  [pdf, ps, other

    cs.CL cs.AI cs.LG

    PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents

    Authors: Zhenbang Du, Jun Luo, Zhiwei Zheng, Xiangchi Yuan, Kejing Xia, Dachuan Shi, Qirui Jin, Qijia He, Shaofeng Zou, Yingbin Liang, Wenke Lee

    Abstract: Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents is challenging, as reinforcement learning often suffers from sparse rewards and weak credit assignment despite matching the prompt-only inference setting, while supervised fine-tuning on expert traces provides dense process supervision but can over-constrain the… ▽ More

    Submitted 15 June, 2026; originally announced June 2026.

    Comments: Project page: https://zhenbangdu.github.io/pact-project-page/

  19. arXiv:2606.05405  [pdf, ps, other

    cs.AI cs.CL cs.LG

    Agents' Last Exam

    Authors: Yiyou Sun, Xinyang Han, Weichen Zhang, Yuanbo Pang, Tianyu Wang, Yuhan Cao, Yixiao Huang, Chris Duroiu, Haoyun Zhang, Jeffrey Lin, Weishu Zhang, Tyler Zeng, Ying Yan, Bo Liu, Hanson Wen, Mingyang Xu, Xiaoyuan Liu, Zimeng Chen, Weiyan Shi, Amanda Dsouza, Vincent Sunn Chen, Patrick Bryant, Carl Boettiger, Yamini Rangan, Bradley Rothenberg , et al. (285 additional authors not shown)

    Abstract: Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a… ▽ More

    Submitted 11 June, 2026; v1 submitted 3 June, 2026; originally announced June 2026.

    Comments: Project website: https://agents-last-exam.org Code: https://github.com/rdi-berkeley/agents-last-exam

  20. arXiv:2606.02437  [pdf, ps, other

    cs.LG cs.CL

    On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters

    Authors: Mind Lab, :, Vin Bo, Song Cao, Vic Cao, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Nolan Ho, Mutian Hong, Hailee Hou, Peixuan Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin , et al. (42 additional authors not shown)

    Abstract: Parameter-efficient fine-tuning (PEFT) is usually treated as a cheaper alternative to full fine-tuning. We study a broader role: small trainable adapters as persistent local state on top of strong shared foundation models. In this framing, the base model provides shared competence while adapters carry instance-specific behavior such as preferences, skills, tool habits, and memory-like updates. We… ▽ More

    Submitted 2 June, 2026; v1 submitted 1 June, 2026; originally announced June 2026.

  21. arXiv:2605.29280  [pdf, ps, other

    cs.LG cs.AI cs.IR

    LoopFM: Learning frOm HistOrical RePresentations of Foundation Model for Recommendation

    Authors: Shali Jiang, Hua Zheng, Boyang Liu, Laming Chen, Kenny Lov, Chuanqi Xu, Lisang Ding, Qinghai Zhou, Can Cui, Xiaolong Liu, Xiaoyi Liu, Yasmine Badr, Xin Xu, Jiyan Yang, Ellie Dingqiao Wen, Gerard Jonathan Mugisha Akkerhuis, Chenxiao Guan, Rong Jin, Ruichao Qiu, Xian Chen, Shifu Xu, Zhehui Zhou, Ping Chen, Rui Yang, Haicheng Chen , et al. (18 additional authors not shown)

    Abstract: Knowledge distillation (KD) transfers a single scalar prediction from a large foundation model (FM) to compact vertical models (VMs), suffering from diminishing transfer ratio -- the fraction of FM improvement captured by the VM -- as a single scalar cannot convey the rich intermediate knowledge that larger FMs learn. To address this bottleneck, we propose LoopFM (Learning frOm HistOrical RePresen… ▽ More

    Submitted 2 June, 2026; v1 submitted 27 May, 2026; originally announced May 2026.

    Comments: Shali Jiang, Hua Zheng, Boyang Liu contributed equally to this work

  22. arXiv:2605.28073  [pdf, ps, other

    cs.CL cs.AI

    StoryLens: Preference-Aligned Story Rewriting via Context-Aware Narrative Enrichment

    Authors: Hanwen Cui, Yuting Mei, Yuhang Fu, Dingyi Yang, Qin Jin

    Abstract: Story rewriting aims to adapt existing narratives to diverse reader preferences while preserving plot consistency and narrative coherence. Unlike conventional work on style transfer, we argue that effective story rewriting demands context-aware narrative enrichment beyond surface-level stylistic adaptation. Our pilot human study shows that style adaptation alone provides only marginal gains in rea… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

    Comments: 16 pages, 7 figures, 15 tables

  23. arXiv:2605.21392  [pdf, ps, other

    cs.CR

    VIPER-MCP: Detecting and Exploiting Taint-Style Vulnerabilities in Model Context Protocol Servers

    Authors: Pengyu Sun, Zifeng Kang, Qishu Jin, Enhao Huang, Xin Liu, Dakun Shen, Song Li

    Abstract: Model Context Protocol (MCP) has emerged as a standard interface for connecting LLM agents to external tools. Because MCP servers expose privileged operations such as shell execution, network access, and file-system manipulation to agent-driven invocation, implementation flaws in tool handlers can create a direct path from natural-language input to security-sensitive sinks, potentially granting at… ▽ More

    Submitted 12 August, 2026; v1 submitted 20 May, 2026; originally announced May 2026.

  24. arXiv:2605.20158  [pdf, ps, other

    cs.CV cs.AI cs.CL

    Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models

    Authors: Guangzhi Xiong, Qiao Jin, Sanchit Sinha, Zhiyong Lu, Aidong Zhang

    Abstract: Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clinical trustworthiness. While visual attribution methods are widely used to explain LVLM predictions, whether these explanations actually reflect the visual evidence underlying the model's decision is largely unverified, si… ▽ More

    Submitted 19 May, 2026; originally announced May 2026.

  25. arXiv:2605.16381  [pdf, ps, other

    cs.CV cs.AI

    StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video

    Authors: Ao Li, Zihan Xiao, Zihao Yue, Boshen Xu, Linli Yao, Jiaze Li, Pei Fu, Jianzhong Ju, Jian Luan, Qin Jin

    Abstract: Proactive streaming video understanding requires models to continuously process video streams and decide when to respond, rather than merely what to respond. This naturally introduces a decision-making problem under partial observations, where models must balance early prediction against sufficient evidence. However, existing benchmarks largely follow a "see-then-answer" paradigm, where responses… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  26. arXiv:2605.13779  [pdf, ps, other

    cs.LG cs.AI cs.DC

    MinT: Managed Infrastructure for Training and Serving Millions of LLMs

    Authors: Mind Lab, :, Song Cao, Vic Cao, Andrew Chen, Kaijie Chen, Cleon Cheng, Steven Chiang, Kaixuan Fan, Hera Feng, Huan Feng, Arthur Fu, Jun Gao, Hongquan Gu, Aaron Guan, Nolan Ho, Mutian Hong, Hailee Hou, Peixuan Hua, Charles Huang, Miles Jiang, Nora Jiang, Yuyi Jiang, Qiuyu Jin, Fancy Kong , et al. (38 additional authors not shown)

    Abstract: We present MindLab Toolkit (MinT), a managed infrastructure system for Low-Rank Adaptation (LoRA) post-training and online serving. MinT targets a setting where many trained policies are produced over a small number of expensive base-model deployments. Instead of materializing each policy as a merged full checkpoint, MinT keeps the base model resident and moves exported LoRA adapter revisions thro… ▽ More

    Submitted 26 May, 2026; v1 submitted 13 May, 2026; originally announced May 2026.

    Comments: 30 pages, technical report

  27. arXiv:2605.13045  [pdf, ps, other

    cs.LG cs.CL

    Large Language Models Lack Temporal Awareness of Medical Knowledge

    Authors: Zihan Guan, Qiao Jin, Guangzhi Xiong, Fangyuan Chen, Mengxuan Hu, Qingyu Chen, Yifan Peng, Zhiyong Lu, Anil Vullikanti

    Abstract: The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, while in reality, medical knowledge is inherently dynamic and continuously evolves as new evidence emerges and treatments are approved. Consequently, evaluating medical knowledge without a temporal context may provide an incomplete assessment of whe… ▽ More

    Submitted 13 May, 2026; originally announced May 2026.

    Comments: 35 pages, 18 figures

  28. arXiv:2605.12361  [pdf

    cs.CL cs.AI cs.IR

    MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering

    Authors: Rezarta Islamaj, Robert Leaman, Joey Chan, Nicholas Wan, Qiao Jin, Natalie Xie, John Wilbur, Shubo Tian, Lana Yeganova, Po-Ting Lai, Chih-Hsuan Wei, Yifan Yang, Yao Ge, Qingqing Zhu, Zhizheng Wang, Zhiyong Lu

    Abstract: Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabilities improve. Existing biomedical question answering (QA) benchmarks are limited in this respect. Multiple-choice formats can allow models to succeed through answer elimination rather than inference, while widely circul… ▽ More

    Submitted 12 May, 2026; originally announced May 2026.

  29. SpatialPrompt: XR-Based Spatial Intent Expression as Executable Constraints for AI Generative 3D Design

    Authors: Yichen Andy Yu, Wanru Li, Qiaoran Wang, Jymon Ross, Gavin Johnson, Mandy Lui, Qiao Jin

    Abstract: We present SpatialPrompt, an Extended Reality(XR) system that turns spatial sketches into executable constraints for controllable 3D generation. Users draw rough structures with a 3D pen and add voice prompts for semantic and stylistic intent. The system supports iterative refinement and synchronous co-creation in shared space with color-coded contributions. Implemented on Apple Vision Pro with Lo… ▽ More

    Submitted 8 May, 2026; originally announced May 2026.

    Journal ref: Proc. DIS Companion 2026, 4 pages

  30. arXiv:2604.26102  [pdf, ps, other

    cs.SE cs.CL

    SWE-Edit: Rethinking Code Editing for Efficient SWE-Agent

    Authors: Yikai Zhang, Jiaxin Pei, Kenan Li, Qirui Jin, Maoquan Wang, Jin Pan, Yu Kang, Shengyu Fu, Elsie Nallipogu, Junjie Hu, Yufan Huang, Zijian Jin

    Abstract: Large language model agents have made strong progress on software engineering, yet current systems suffer from a context coupling problem: the standard code editing interface conflates code inspection, modification planning, and edit execution within a single context window, forcing agents to interleave exploratory viewing with strictly formatted edit generation. Irrelevant context accumulates and… ▽ More

    Submitted 26 May, 2026; v1 submitted 28 April, 2026; originally announced April 2026.

  31. arXiv:2604.25721  [pdf, ps, other

    cs.HC

    Designing and Evaluating Next-Generation Learning Interfaces: Linking AI, HCI, and the Learning Sciences

    Authors: Meng Xia, Yan Chen, Qiao Jin, Yang Shi, Paul Denny, Tiffany Barnes, Qingsong Wen, Vincent Aleven

    Abstract: This workshop addresses this gap by bringing together researchers and practitioners from AI, HCI, and the learning sciences to explore how interactive systems can better support learning. We focus on the design and evaluation of human-AI collaborative learning interfaces that are technically robust, human-centered, and pedagogically grounded. By fostering interdisciplinary dialogue, the workshop a… ▽ More

    Submitted 28 April, 2026; originally announced April 2026.

  32. DARC-CLIP: Dynamic Adaptive Refinement with Cross-Attention for Meme Understanding

    Authors: Qiyuan Jin

    Abstract: Memes convey meaning through the interaction of visual and textual signals, often combining humor, irony, and offense in subtle ways. Detecting harmful or sensitive content in memes requires accurate modeling of these multimodal cues. Existing CLIP-based approaches rely on static fusion, which struggles to capture fine grained dependencies between modalities. We propose DARC-CLIP, a CLIP-based fra… ▽ More

    Submitted 28 April, 2026; v1 submitted 25 April, 2026; originally announced April 2026.

    Comments: Accepted to IEEE ICASSP 2026. 5 pages, 3 figures, 4 tables

  33. arXiv:2604.18356  [pdf, ps, other

    cs.CL

    ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

    Authors: Zhaopei Huang, Yanfeng Jia, Jiayi Zhao, Xinjie Zhang, Wenxuan Wang, Qin Jin

    Abstract: Developing compassionate interactive systems requires agents to not only understand user emotions but also provide diverse, substantive support. While recent works explore empathetic dialogue generation, they remain limited in response form and content, struggling to satisfy diverse needs across users and contexts. To address this, we explore empowering agents with external tools to execute divers… ▽ More

    Submitted 20 April, 2026; originally announced April 2026.

  34. arXiv:2604.15736  [pdf, ps, other

    cs.CV cs.CL

    RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

    Authors: Yichen Xu, Yuanhang Liu, Chuhan Wang, Zihan Zhao, jinghan luo, Jianzhe Ma, Wenxuan Wang, Qin Jin

    Abstract: While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-making remains insufficiently explored. In this paper, we introduce RefereeBench, the first large-scale benchmark for evaluating MLLMs as automatic sports referees. Spanning 11 sports with 925 curated videos and 6,475 QA pairs, RefereeBench evaluates fiv… ▽ More

    Submitted 17 April, 2026; originally announced April 2026.

    Comments: Work in Progress

  35. arXiv:2604.15127  [pdf, ps, other

    cs.MM

    MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production

    Authors: Huanran Hu, Zihui Ren, Dingyi Yang, Liangyu Chen, Qixiang Gao, Tiezheng Ge, Qin Jin

    Abstract: Real-world video creation often involves a complex reasoning workflow of selecting relevant shots from noisy materials, planning missing shots for narrative completeness, and organizing them into coherent storylines. However, existing benchmarks focus on isolated sub-tasks and lack support for evaluating this full process. To address this gap, we propose Multimodal Context-to-Script Creation (MCSC… ▽ More

    Submitted 16 April, 2026; v1 submitted 16 April, 2026; originally announced April 2026.

  36. arXiv:2604.12320  [pdf, ps, other

    cs.CV cs.AI cs.MM

    EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports

    Authors: Jianzhe Ma, Zhonghao Cao, Shangkui Chen, Yichen Xu, Wenxuan Wang, Qin Jin

    Abstract: While video large language models (Video-LLMs) excel in understanding slow-paced, real-world egocentric videos, their capabilities in high-velocity, information-dense virtual environments remain under-explored. Existing benchmarks focus on daily activities, yet lack a rigorous testbed for evaluating fast, rule-bound reasoning in virtual scenarios. To fill this gap, we introduce EgoEsportsQA, a pio… ▽ More

    Submitted 20 April, 2026; v1 submitted 14 April, 2026; originally announced April 2026.

    Comments: Work in progress

  37. arXiv:2604.12208  [pdf, ps, other

    cs.RO cs.AI

    Unveiling the Surprising Efficacy of Navigation Understanding in End-to-End Autonomous Driving

    Authors: Zhihua Hua, Junli Wang, Pengfei LI, Qihao Jin, Bo Zhang, Kehua Sheng, Yilun Chen, Zhongxue Gan, Wenchao Ding

    Abstract: Global navigation information and local scene understanding are two crucial components of autonomous driving systems. However, our experimental results indicate that many end-to-end autonomous driving systems tend to over-rely on local scene understanding while failing to utilize global navigation information. These systems exhibit weak correlation between their planning capabilities and navigatio… ▽ More

    Submitted 13 April, 2026; originally announced April 2026.

    Comments: 8 pages, 6 figures. ICRA 2026. Code available at https://fudan-magic-lab.github.io/SNG-VLA-web

  38. arXiv:2604.09550  [pdf, ps, other

    cs.IR cs.DB

    HyEm: Query-Adaptive Hyperbolic Retrieval for Biomedical Ontologies via Euclidean Vector Indexing

    Authors: Ou Deng, Shoji Nishimura, Atsushi Ogihara, Qun Jin

    Abstract: Retrieval-augmented generation (RAG) for biomedical knowledge faces a hierarchy-aware ontology grounding challenge: resources like HPO, DO, and MeSH use deep ``is-a" taxonomies, yet production stacks rely on Euclidean embeddings and ANN indexes. While hyperbolic embeddings suit hierarchical representation, they face two barriers: (i) lack of native vector database support, and (ii) risk of underpe… ▽ More

    Submitted 26 January, 2026; originally announced April 2026.

  39. arXiv:2604.07789  [pdf, ps, other

    cs.MA cs.CL cs.SE

    ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents

    Authors: Kenan Li, Qirui Jin, Liao Zhu, Xiaosong Huang, Yijia Wu, Yikai Zhang, Xin Zhang, Zijian Jin, Yufan Huang, Elsie Nallipogu, Chaoyun Zhang, Yu Kang, Saravan Rajmohan, Qingwei Lin, Wenke Lee, Dongmei Zhang

    Abstract: Recent advances in language model (LM) agents have significantly improved automated software engineering (SWE). Prior work has proposed various agentic workflows and training strategies as well as analyzed failure modes of agentic systems on SWE tasks, focusing on several contextual information signals: Reproduction Test, Regression Test, Edit Location, Execution Context, and API Usage. However, t… ▽ More

    Submitted 28 May, 2026; v1 submitted 9 April, 2026; originally announced April 2026.

    Comments: Under peer review; 37 pages, 10 figures, 5 tables

    ACM Class: I.2.7; I.2.5

  40. arXiv:2604.00601  [pdf, ps, other

    cs.CV

    KG-CMI: Knowledge graph enhanced cross-Mamba interaction for medical visual question answering

    Authors: Xianyao Zheng, Hong Yu, Hui Cui, Changming Sun, Xiangyu Li, Ran Su, Leyi Wei, Jia Zhou, Junbo Wang, Qiangguo Jin

    Abstract: Medical visual question answering (Med-VQA) is a crucial multimodal task in clinical decision support and telemedicine. Recent methods fail to fully leverage domain-specific medical knowledge, making it difficult to accurately associate lesion features in medical images with key diagnostic criteria. Additionally, classification-based approaches typically rely on predefined answer sets. Treating Me… ▽ More

    Submitted 1 April, 2026; originally announced April 2026.

  41. arXiv:2603.21904  [pdf, ps, other

    cs.CV cs.AI

    SHAPE: Structure-aware Hierarchical Unsupervised Domain Adaptation with Plausibility Evaluation for Medical Image Segmentation

    Authors: Linkuan Zhou, Yinghao Xia, Yufei Shen, Xiangyu Li, Wenjie Du, Cong Cong, Leyi Wei, Ran Su, Qiangguo Jin

    Abstract: Unsupervised Domain Adaptation (UDA) is essential for deploying medical segmentation models across diverse clinical environments. Existing methods are fundamentally limited, suffering from semantically unaware feature alignment that results in poor distributional fidelity and from pseudo-label validation that disregards global anatomical constraints, thus failing to prevent the formation of global… ▽ More

    Submitted 23 March, 2026; originally announced March 2026.

  42. arXiv:2603.20804  [pdf, ps, other

    cs.CV cs.RO

    Does Peer Observation Help? Vision-Sharing Collaboration for Vision-Language Navigation

    Authors: Qunchao Jin, Yiliao Song, Qi Wu

    Abstract: Vision-Language Navigation (VLN) systems are fundamentally constrained by partial observability, as an agent can only accumulate knowledge from locations it has personally visited. As multiple robots increasingly coexist in shared environments, a natural question arises: can agents navigating the same space benefit from each other's observations? In this work, we introduce Co-VLN, a minimalist, mo… ▽ More

    Submitted 21 March, 2026; originally announced March 2026.

  43. arXiv:2603.13818  [pdf, ps, other

    cs.AI cs.CV cs.LG

    PA-Net: Precipitation-Adaptive Mixture-of-Experts for Long-Tail Rainfall Nowcasting

    Authors: Xinyu Xiao, Sen Lei, Eryun Liu, Shiming Xiang, Hao Li, Cheng Yuan, Yuan Qi, Qizhao Jin

    Abstract: Precipitation nowcasting is vital for flood warning, agricultural management, and emergency response, yet two bottlenecks persist: the prohibitive cost of modeling million-scale spatiotemporal tokens from multi-variate atmospheric fields, and the extreme long-tailed rainfall distribution where heavy-to-torrential events -- those of greatest societal impact -- constitute fewer than 0.1% of all samp… ▽ More

    Submitted 14 March, 2026; originally announced March 2026.

  44. arXiv:2603.13752  [pdf, ps, other

    cs.AI cs.LG

    MeTok: An Efficient Meteorological Tokenization with Hyper-Aligned Group Learning for Precipitation Nowcasting

    Authors: Qizhao Jin, Xianhuang Xu, Yong Cao, Shiming Xiang, Xinyu Xiao

    Abstract: Recently, Transformer-based architectures have advanced meteorological prediction. However, this position-centric tokenizer conflicts with the core principle of meteorological systems, where the weather phenomena undoubtedly involve synergistic interactions among multiple elements while positional information constitutes merely a component of the boundary conditions. This paper focuses primarily o… ▽ More

    Submitted 14 March, 2026; originally announced March 2026.

  45. arXiv:2603.05308  [pdf, ps, other

    cs.CL cs.AI

    Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

    Authors: Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu

    Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification. While large language models (LLMs) have the potential to automate this task, achieving strong performance requires frontier models such as GPT-5 that are prohibitively expensive to deploy at scale. To efficiently perform biomedical evidence attribution, we present Med-V1, a family of… ▽ More

    Submitted 31 May, 2026; v1 submitted 5 March, 2026; originally announced March 2026.

  46. arXiv:2603.05026  [pdf, ps, other

    cs.SE cs.LG cs.MA

    RepoLaunch: Automating Build and Management of Code Repositories across Languages and Platforms

    Authors: Kenan Li, Rongzhi Li, Linghao Zhang, Qirui Jin, Liao Zhu, Xiaosong Huang, Geng Zhang, Yikai Zhang, Shilin He, Chengxing Xie, Xin Zhang, Zijian Jin, Bowen Li, Chaoyun Zhang, Yu Kang, Yufan Huang, Elsie Nallipogu, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang

    Abstract: Language model (LM) agents have driven substantial progress in automated software engineering (SWE), yet building and testing software repositories at scale remains a largely manual and labor-intensive bottleneck. In this work, we introduce RepoLaunch, a novel agentic framework that automatically resolves dependencies, compiles source code, and extracts test results across diverse programming lang… ▽ More

    Submitted 6 June, 2026; v1 submitted 5 March, 2026; originally announced March 2026.

    Comments: Under peer review. 22 pages, 5 figures, 9 tables

    ACM Class: I.2.5; I.2.7

  47. arXiv:2603.01331  [pdf, ps, other

    cs.CL cs.AI cs.LG

    MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models

    Authors: Kejing Xia, Mingzhe Li, Lixuan Wei, Zhenbang Du, Xiangchi Yuan, Dachuan Shi, Qirui Jin, Wenke Lee

    Abstract: Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each denoising step solely on the current hard-masked sequence, while intermediate continuous representations are discarded after sampling and remasking. We term this bottleneck the \textbf{Information Island} issue: continuous information remains isolated within i… ▽ More

    Submitted 10 July, 2026; v1 submitted 1 March, 2026; originally announced March 2026.

  48. arXiv:2603.01121  [pdf, ps, other

    cs.AI

    HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis

    Authors: Shuo Tang, Jiadong Zhang, Gengxian Zhou, Qizhao Jin, Qinxuan Wang, Yi Hu, Ning Hu, Hongchang Ren, Lingli He, Shiming Xiang, Jingtao Ding, Jian Xu, Jiaolan Fu, Cheng-Lin Liu

    Abstract: While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a formidable challenge. This gap exists primarily because the diagnostic process demands sophisticated multi-step logical reasoning, dynamic tool invocation, and expert-level prior judgment. Although agents possess inherent advantages in task decomposition and auton… ▽ More

    Submitted 3 July, 2026; v1 submitted 1 March, 2026; originally announced March 2026.

  49. arXiv:2602.23333  [pdf, ps, other

    cs.SD

    SemanticVocoder: Bridging Audio Generation and Audio Understanding via Semantic Latents

    Authors: Zeyu Xie, Chenxing Li, Qiao Jin, Xuenan Xu, Guanrou Yang, Wenfu Wang, Mengyue Wu, Dong Yu, Yuexian Zou

    Abstract: Recent audio generation models typically rely on Variational Autoencoders (VAEs) and perform generation within the VAE latent space. Although VAEs excel at compression and reconstruction, their latents inherently encode low-level acoustic details rather than semantically discriminative information, leading to entangled event semantics and complicating the training of generative models. To address… ▽ More

    Submitted 26 February, 2026; originally announced February 2026.

    Comments: Demo: https://zeyuxie29.github.io/SemanticVocoder/

    MSC Class: 68Txx ACM Class: I.2

  50. arXiv:2602.18232  [pdf, ps, other

    cs.CL cs.AI

    Thinking by Subtraction: Confidence-Driven Contrastive Decoding for LLM Reasoning

    Authors: Lexiang Tang, Weihao Gao, Bingchen Zhao, Lu Ma, Qiao jin, Bang Yang, Yuexian Zou

    Abstract: Recent work on test-time scaling for large language model (LLM) reasoning typically assumes that allocating more inference-time computation uniformly improves correctness. However, prior studies show that reasoning uncertainty is highly localized: a small subset of low-confidence tokens disproportionately contributes to reasoning errors and unnecessary output expansion. Motivated by this observati… ▽ More

    Submitted 20 February, 2026; originally announced February 2026.