Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 80 results for author: Si, X

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.15432  [pdf, ps, other

    cs.AI

    Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs

    Authors: Tadd Mao, Tianjun Zhong, Dhruva Arekar, Yuming Feng, One An, Jiani Huang, Xujie Si, Ziyang Li

    Abstract: In formal verification, both the autoformalization of statements and automated proof search have been studied extensively. While automated proof search can produce a formal proof that compiles, the generated proof does not necessarily reflect how the natural-language argument arrives at its conclusion--a property we refer to as faithfulness. With faithfully formalized proofs, one can check the rea… ▽ More

    Submitted 15 August, 2026; originally announced August 2026.

    Comments: Preprint. 18 pages, 14 figures, 7 tables

  2. arXiv:2608.14585  [pdf, ps, other

    cs.AI

    Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry

    Authors: Zhaoyu Li, Hangrui Bi, Youyuan Zhang, Wenjie Ma, Zenan Li, Zhaolei Zhang, Xujie Si, Kaiyu Yang

    Abstract: Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic deduction, and algebraic computation. Yet, existing approaches typically address only a subset of these abilities or struggle with competition-level problems. We introduce \textit{Euclid-Omni}, a unified neuro-symbolic framework that couples a formal geometry sys… ▽ More

    Submitted 17 June, 2026; originally announced August 2026.

  3. arXiv:2607.18801  [pdf, ps, other

    cs.CV

    ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting

    Authors: Jiayu Ding, Meilu Song, Xiaoyi Zhang, Hongbo Jin, Yichen Jin, Xiangtian Si

    Abstract: Recent advancements in 3D Gaussian Splatting (3DGS) have enabled language-guided scene understanding. However, existing Referring 3D Gaussian Splatting (R3DGS) methods are fundamentally restricted to single-target queries. To reflect the ambiguity of real-world instructions, we introduce the Generalized Referring 3D Gaussian Splatting Segmentation (GR3DGS) task, which requires dynamically segmenti… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

    Comments: Accepted to ECCV 2026

  4. arXiv:2607.13292  [pdf, ps, other

    cs.AI cs.CL cs.LG cs.PL

    Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases

    Authors: Marcus J. Min, Mike He, Zhaoyu Li, Zixuan Yi, Sharad Malik, Aarti Gupta, Xujie Si, Osbert Bastani

    Abstract: Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual statements, real formalization efforts are inherently theory-level: they require an entire web of axioms, definitions, and lemmas before target theorems can even be stated. In this position paper, we argue for theory-level autoformalization: formalizing complete t… ▽ More

    Submitted 5 August, 2026; v1 submitted 14 July, 2026; originally announced July 2026.

    Comments: ICML 2026 Spotlight

    MSC Class: 68 ACM Class: F.4; I.2

  5. arXiv:2606.26525  [pdf, ps, other

    cs.LG cs.LO cs.PL

    Theory-Scale Auto-Formalization of Logics for Computer Science

    Authors: Yuming Feng, Frederick Pu, One An, Osbert Bastani, Li Zhang, Jiani Huang, Xujie Si, Ziyang Li

    Abstract: Auto-formalization is critical for scalable formal verification, but existing progress largely focuses on isolated statements, while theory-scale auto-formalization, which coherently translates hundreds of interdependent definitions, lemmas, and theorems, remains open due to challenges in consistency, faithfulness, scalability, and correctness. In this paper, we introduce LCS-Bench, a stand-alone,… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

  6. arXiv:2606.08947  [pdf, ps, other

    cs.AR

    NeuDW-CIM: a 65-nm 0.8-pJ/Sop Reconfigurable Neuromorphic Compute-in-Memory Macro with Nonlinear Dendrites and K-Winners

    Authors: Junyi Yang, Yahan Yang, Shuai Dong, Biyan Zhou, Ye Ke, Zhengnan Fu, Xin Si, An Guo, Peng Zhou, Arindam Basu

    Abstract: This work presents NeuDW-CIM, a highly efficient neuromorphic Compute-in-Memory (CIM) macro for Spiking Neural Networks (SNNs) implemented in 65 nm CMOS. The design introduces a custom twin 9T bit-cell for ternary in-puts/weights and a reconfigurable non-linear In-Memory ADC (IMA). The macro supports two specialized modes: 1) Nonlinear Dendrite (NLD) mode, which utilizes reconfigurable IMA to emul… ▽ More

    Submitted 7 June, 2026; originally announced June 2026.

  7. arXiv:2606.03240  [pdf, ps, other

    cs.RO

    GeoAlign: Beyond Semantics with State-Guided Spatial Alignment in VLA Models

    Authors: Yizhi Chen, Zhanxiang Cao, Xinyi Peng, Yixiao Zheng, Xiaxi Si, Yiheng Li, Liyun Yan, Keqi Zhu, Xueyun Chen, Shengcheng Fu, Tianyue Zhan, Yufei Jia, Jinming Yao, Yan Xie, Kun Wang, Cewu Lu, Yue Gao

    Abstract: Current Vision--Language--Action (VLA) models often optimize for semantic grounding, whereas executable manipulation requires geometry-aware spatial alignment and dynamic affordance selection. We introduce GeoAlign, a state-guided spatial alignment architecture for VLA policy learning. GeoAlign post-trains an RGB geometry branch with robot-domain RGB-D supervision, yielding RGB-derived Geometry-En… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: 20 pages, 9 figures, 8 tables, including appendix

  8. arXiv:2605.00270  [pdf, ps, other

    cs.CL cs.AI cs.CY cs.HC

    Are You the A-hole? A Fair, Multi-Perspective Ethical Reasoning Framework

    Authors: Sheza Munir, Ahanaf Rodoshi, Sumin Lee, Feiran Chang, Xujie Si, Syed Ishtiaque Ahmed

    Abstract: Standard methods for aggregating natural language judgments, such as majority voting, often fail to produce logically consistent results when applied to high-conflict domains, treating differing opinions as noise. We propose a neuro-symbolic aggregation framework that formalizes conflict resolution through Weighted Maximum Satisfiability (MaxSAT). Our pipeline utilizes a language model to map unst… ▽ More

    Submitted 30 April, 2026; originally announced May 2026.

  9. arXiv:2604.26311  [pdf, ps, other

    cs.AI

    DreamProver: Evolving Transferable Lemma Libraries via a Wake-Sleep Theorem-Proving Agent

    Authors: Youyuan Zhang, Jialiang Sun, Hangrui Bi, Chuqin Geng, Wenjie Ma, Zhaoyu Li, Xujie Si

    Abstract: We introduce DreamProver, an agentic framework that leverages a "wake-sleep" program induction paradigm to discover reusable lemmas for formal theorem proving. Existing approaches either rely on fixed lemma libraries, which limit adaptability, or synthesize highly specific intermediate lemmas tailored to individual theorems, thereby lacking generality. DreamProver addresses this gap through an ite… ▽ More

    Submitted 29 April, 2026; originally announced April 2026.

  10. arXiv:2604.17692  [pdf, ps, other

    cs.AR

    AccelCIM: Systematic Dataflow Exploration for SRAM Compute-in-Memory Accelerator

    Authors: Chenhao Xue, Yukun Wang, An Guo, Yuhui Shi, Jinwei Zhou, Xiping Dong, Yihan Yin, Yuanpeng Zhang, Tianyu Jia, Wei Gao, Qiang Wu, Xin Si, Jun Yang, Guangyu Sun

    Abstract: SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip data movement overhead for large DNN models. Existing CIM accelerator studies typically assume that DNN models fit entirely on-chip, leaving efficient dataflow design largely untapped. This paper introduces AccelCIM, a sy… ▽ More

    Submitted 19 April, 2026; originally announced April 2026.

    Comments: Accepted by DAC'26

  11. arXiv:2604.01489  [pdf, ps, other

    cs.LG cs.AI cs.DC cs.PF cs.SE

    CuTeGen: An LLM-Based Agentic Framework for Generation and Optimization of High-Performance GPU Kernels using CuTe

    Authors: Tara Saba, Zhiyang Chen, Jikai Jason Li, Anne Ouyang, Xujie Si, Fan Long

    Abstract: High-performance GPU kernels are critical to modern machine learning systems, yet developing them remains a manual, expert-driven process. Recent work has explored using LLMs to automate kernel generation, but generated kernels still fall short of carefully tuned references on standardized benchmarks. We present CuTeGen, an agentic GPU kernel synthesis framework that treats kernel development as a… ▽ More

    Submitted 3 June, 2026; v1 submitted 1 April, 2026; originally announced April 2026.

  12. Evaluating LLMs in the Context of a Functional Programming Course: A Comprehensive Study

    Authors: Yihan Zhang, Brigitte Pientka, Xujie Si

    Abstract: Large-Language Models (LLMs) are changing the way learners acquire knowledge outside the classroom setting. Previous studies have shown that LLMs seem effective in generating to short and simple questions in introductory CS courses using high-resource programming languages such as Java or Python. In this paper, we evaluate the effectiveness of LLMs in the context of a low-resource programming lang… ▽ More

    Submitted 5 March, 2026; originally announced March 2026.

    Journal ref: The Art, Science, and Engineering of Programming, 2026, Vol. 11, Issue 1, Article 5

  13. arXiv:2602.16954  [pdf, ps, other

    cs.LG

    Neural Proposals, Symbolic Guarantees: Neuro-Symbolic Graph Generation with Hard Constraints

    Authors: Chuqin Geng, Li Zhang, Mark Zhang, Haolin Ye, Ziyu Zhao, Xujie Si

    Abstract: We challenge black-box purely deep neural approaches for molecules and graph generation, which are limited in controllability and lack formal guarantees. We introduce Neuro-Symbolic Graph Generative Modeling (NSGGM), a neurosymbolic framework that reapproaches molecule generation as a scaffold and interaction learning task with symbolic assembly. An autoregressive neural model proposes scaffolds a… ▽ More

    Submitted 24 February, 2026; v1 submitted 18 February, 2026; originally announced February 2026.

    Comments: 18 pages, 6 figures

  14. arXiv:2602.16947  [pdf, ps, other

    cs.LG cs.AI

    Beyond Message Passing: A Symbolic Alternative for Expressive and Interpretable Graph Learning

    Authors: Chuqin Geng, Li Zhang, Haolin Ye, Ziyu Zhao, Yuhe Jiang, Tara Saba, Xinyu Wang, Xujie Si

    Abstract: Graph Neural Networks (GNNs) have become essential in high-stakes domains such as drug discovery, yet their black-box nature remains a significant barrier to trustworthiness. While self-explainable GNNs attempt to bridge this gap, they often rely on standard message-passing backbones that inherit fundamental limitations, including the 1-Weisfeiler-Lehman (1-WL) expressivity barrier and a lack of f… ▽ More

    Submitted 23 February, 2026; v1 submitted 18 February, 2026; originally announced February 2026.

    Comments: 23 pages, 9 pages

  15. arXiv:2601.18070  [pdf, ps, other

    cs.AR

    CIM-Tuner: Balancing the Compute and Storage Capacity of SRAM-CIM Accelerator via Hardware-mapping Co-exploration

    Authors: Jinwu Chen, Yuhui Shi, He Wang, Zhe Jiang, Jun Yang, Xin Si, Zhenhua Zhu

    Abstract: As an emerging type of AI computing accelerator, SRAM Computing-In-Memory (CIM) accelerators feature high energy efficiency and throughput. However, various CIM designs and under-explored mapping strategies impede the full exploration of compute and storage balancing in SRAM-CIM accelerator, potentially leading to significant performance degradation. To address this issue, we propose CIM-Tuner, an… ▽ More

    Submitted 25 January, 2026; originally announced January 2026.

  16. arXiv:2512.16465  [pdf, ps, other

    cs.AI

    cuPilot: A Strategy-Coordinated Multi-agent Framework for CUDA Kernel Evolution

    Authors: Jinwu Chen, Qidie Wu, Bin Li, Lin Ma, Xin Si, Yang Hu, Shouyi Yin, Jun Yang

    Abstract: Optimizing CUDA kernels is a challenging and labor-intensive task, given the need for hardware-software co-design expertise and the proprietary nature of high-performance kernel libraries. While recent large language models (LLMs) combined with evolutionary algorithms show promise in automatic kernel optimization, existing approaches often fall short in performance due to their suboptimal agent de… ▽ More

    Submitted 23 December, 2025; v1 submitted 18 December, 2025; originally announced December 2025.

  17. arXiv:2511.04321  [pdf, ps, other

    cs.AR cs.AI cs.LG

    AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIM

    Authors: Yuanpeng Zhang, Xing Hu, Xi Chen, Zhihang Yuan, Cong Li, Jingchen Zhu, Zhao Wang, Chenguang Zhang, Xin Si, Wei Gao, Qiang Wu, Runsheng Wang, Guangyu Sun

    Abstract: SRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision. However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues. Severe IR-drop can significantly degrade chip perfo… ▽ More

    Submitted 6 November, 2025; originally announced November 2025.

    Comments: 18 pages, 22 figures, accepted by ISCA 2025

  18. arXiv:2510.09710  [pdf, ps, other

    cs.CL cs.AI

    SeCon-RAG: A Two-Stage Semantic Filtering and Conflict-Free Framework for Trustworthy RAG

    Authors: Xiaonan Si, Meilin Zhu, Simeng Qin, Lijia Yu, Lijun Zhang, Shuaitong Liu, Xinfeng Li, Ranjie Duan, Yang Liu, Xiaojun Jia

    Abstract: Retrieval-augmented generation (RAG) systems enhance large language models (LLMs) with external knowledge but are vulnerable to corpus poisoning and contamination attacks, which can compromise output integrity. Existing defenses often apply aggressive filtering, leading to unnecessary loss of valuable information and reduced reliability in generation. To address this problem, we propose a two-stag… ▽ More

    Submitted 15 October, 2025; v1 submitted 9 October, 2025; originally announced October 2025.

    Comments: Accepted at NeurIPS 2025

  19. arXiv:2510.07315  [pdf, ps, other

    cs.CL cs.AI cs.LG cs.SE

    SWE-IF: Aligning Code Evaluation with Human Preference

    Authors: Ming Zhong, Xiang Zhou, Ting-Yun Chang, Qingze Wang, Nan Xu, Xiance Si, Dan Garrette, Shyam Upadhyay, Jeremiah Liu, Jiawei Han, Benoit Schillings, Jiao Sun

    Abstract: Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural language interactions until it passes their vibe check. Vibe check reflects human preference and goes beyond functionality: the solution should feel right, read cleanly, preserve intent, and remain correct. However, current code evaluation remains anchored to p… ▽ More

    Submitted 4 June, 2026; v1 submitted 8 October, 2025; originally announced October 2025.

    Comments: ICML 2026

  20. arXiv:2509.25197  [pdf, ps, other

    cs.SE cs.AI cs.PL

    Towards Repository-Level Program Verification with Large Language Models

    Authors: Si Cheng Zhong, Xujie Si

    Abstract: Recent advancements in large language models (LLMs) suggest great promises in code and proof generations. However, scaling automated formal verification to real-world projects requires resolving cross-module dependencies and global contexts, which are crucial challenges overlooked by existing LLM-based methods with a special focus on targeting isolated, function-level verification tasks. To system… ▽ More

    Submitted 30 August, 2025; originally announced September 2025.

    Comments: Accepted to LMPL 2025

  21. arXiv:2509.02372  [pdf, ps, other

    cs.CR cs.AI cs.SE

    Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs

    Authors: Zhiyang Chen, Tara Saba, Xun Deng, Xujie Si, Fan Long

    Abstract: Large Language Models have become critical to modern software development, but their reliance on uncurated web-scale datasets for training introduces a significant security risk: the absorption and reproduction of malicious content. This risk materialized in November 2024, when a user suffered a 2,500 USD financial loss after executing code generated by ChatGPT that contained a live scam phishing… ▽ More

    Submitted 8 May, 2026; v1 submitted 2 September, 2025; originally announced September 2025.

  22. A Case Study on the Effectiveness of LLMs in Verification with Proof Assistants

    Authors: Barış Bayazıt, Yao Li, Xujie Si

    Abstract: Large language models (LLMs) can potentially help with verification using proof assistants by automating proofs. However, it is unclear how effective LLMs are in this task. In this paper, we perform a case study based on two mature Rocq projects: the hs-to-coq tool and Verdi. We evaluate the effectiveness of LLMs in generating proofs by both quantitative and qualitative analysis. Our study finds t… ▽ More

    Submitted 25 August, 2025; originally announced August 2025.

    Comments: Accepted by LMPL 2025

  23. arXiv:2507.22086  [pdf, ps, other

    cs.SE cs.AI cs.PL

    TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories

    Authors: Honghua Dong, Jiacheng Yang, Xun Deng, Yuhe Jiang, Gennady Pekhimenko, Fan Long, Xujie Si

    Abstract: Type inference for dynamic languages like Python is a persistent challenge in software engineering. While large language models (LLMs) have shown promise in code understanding, their type inference capabilities remain underexplored. We introduce TypyBench, a benchmark designed to evaluate LLMs' type inference across entire Python repositories. TypyBench features two novel metrics: TypeSim, which c… ▽ More

    Submitted 28 July, 2025; originally announced July 2025.

    Journal ref: Proceedings of the 42nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025

  24. An Automated Classifier of Harmful Brain Activities for Clinical Usage Based on a Vision-Inspired Pre-trained Framework

    Authors: Yulin Sun, Xiaopeng Si, Runnan He, Xiao Hu, Peter Smielewski, Wenlong Wang, Xiaoguang Tong, Wei Yue, Meijun Pang, Kuo Zhang, Xizi Song, Dong Ming, Xiuyun Liu

    Abstract: Timely identification of harmful brain activities via electroencephalography (EEG) is critical for brain disease diagnosis and treatment, which remains limited application due to inter-rater variability, resource constraints, and poor generalizability of existing artificial intelligence (AI) models. In this study, a convolutional neural network model, VIPEEGNet, was developed and validated using E… ▽ More

    Submitted 9 July, 2025; originally announced July 2025.

  25. arXiv:2507.06261  [pdf, ps, other

    cs.CL cs.AI

    Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

    Authors: Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor, Nan-Jiang Jiang, Krishna Haridasan, Ahmed Omran, Nikunj Saunshi, Dara Bahri, Gaurav Mishra, Eric Chu , et al. (3410 additional authors not shown)

    Abstract: In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our most capable model yet, achieving SoTA performance on frontier coding and reasoning benchmarks. In addition to its incredible coding and reasoning skills, Gemini 2.5 Pro is a thinking model that excels at multimodal unde… ▽ More

    Submitted 19 December, 2025; v1 submitted 7 July, 2025; originally announced July 2025.

    Comments: 72 pages, 17 figures

  26. arXiv:2506.07982  [pdf, ps, other

    cs.AI cs.CL

    $τ^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

    Authors: Victor Barres, Honghua Dong, Soham Ray, Xujie Si, Karthik Narasimhan

    Abstract: Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce… ▽ More

    Submitted 9 June, 2025; originally announced June 2025.

  27. arXiv:2506.00563  [pdf, ps, other

    cs.LG cs.AI

    Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments

    Authors: Ziyan Luo, Tianwei Ni, Pierre-Luc Bacon, Doina Precup, Xujie Si

    Abstract: A key approach to state abstraction is approximating behavioral metrics (notably, bisimulation metrics) in the observation space and embedding these learned distances in the representation space. While promising for robustness to task-irrelevant noise, as shown in prior work, accurately estimating these metrics remains challenging, requiring various design choices that create gaps between theory a… ▽ More

    Submitted 8 September, 2025; v1 submitted 31 May, 2025; originally announced June 2025.

  28. arXiv:2505.19271  [pdf, ps, other

    cs.SE

    VerifyThisBench: Generating Code, Specifications, and Proofs All at Once

    Authors: Xun Deng, Sicheng Zhong, Barış Bayazıt, Andreas Veneris, Fan Long, Xujie Si

    Abstract: Large language models (LLMs) have demonstrated remarkable progress in code generation, but many existing benchmarks are approaching saturation and offer little guarantee on the trustworthiness of the generated programs. To improve visibility into model reasoning on formal correctness, we introduce VerifyThisBench, a new benchmark that evaluates end-to-end program verification from natural language… ▽ More

    Submitted 6 October, 2025; v1 submitted 25 May, 2025; originally announced May 2025.

  29. arXiv:2504.17384  [pdf, other

    physics.geo-ph cs.AI

    On the workflow, opportunities and challenges of developing foundation model in geophysics

    Authors: Hanlin Sheng, Xinming Wu, Hang Gao, Haibin Di, Sergey Fomel, Jintao Li, Xu Si

    Abstract: Foundation models, as a mainstream technology in artificial intelligence, have demonstrated immense potential across various domains in recent years, particularly in handling complex tasks and multimodal data. In the field of geophysics, although the application of foundation models is gradually expanding, there is currently a lack of comprehensive reviews discussing the full workflow of integrati… ▽ More

    Submitted 25 April, 2025; v1 submitted 24 April, 2025; originally announced April 2025.

  30. arXiv:2504.03048  [pdf, other

    cs.LG cs.CL

    LLM Library Learning Fails: A LEGO-Prover Case Study

    Authors: Ian Berlot-Attwell, Frank Rudzicz, Xujie Si

    Abstract: Recent advancements in the coding, reasoning, and tool-using abilities of LLMs have spurred interest in library learning (i.e., online learning through the creation, storage, and retrieval of reusable and composable functions, knowledge, checklists, or lemmas). Such systems often promise improved task performance through the automatic creation of broadly applicable tools, as well as superior compu… ▽ More

    Submitted 3 April, 2025; originally announced April 2025.

    Comments: 24 pages, 5 figures

  31. arXiv:2503.19476  [pdf, ps, other

    cs.LG

    LogicXGNN: Grounded Logical Rules for Explaining Graph Neural Networks

    Authors: Chuqin Geng, Ziyu Zhao, Zhaoyue Wang, Haolin Ye, Yuhe Jiang, Xujie Si

    Abstract: Existing rule-based explanations for Graph Neural Networks (GNNs) provide global interpretability but often optimize and assess fidelity in an intermediate, uninterpretable concept space, overlooking grounding quality for end users in the final subgraph explanations. This gap yields explanations that may appear faithful yet be unreliable in practice. To this end, we propose LogicXGNN, a post-hoc f… ▽ More

    Submitted 16 March, 2026; v1 submitted 25 March, 2025; originally announced March 2025.

    Comments: Accepted at ICLR 2026

  32. arXiv:2503.10547  [pdf, ps, other

    cs.CV

    VISIONLOGIC: From Neuron Activations to Causally Grounded Concept Rules for Vision Models

    Authors: Chuqin Geng, Yuhe Jiang, Ziyu Zhao, Haolin Ye, Anqi Xing, Li Zhang, Xujie Si

    Abstract: While concept-based explanations improve interpretability over local attributions, they often rely on correlational signals and lack causal validation. We introduce VisionLogic, a novel neural-symbolic framework that produces faithful, hierarchical explanations as global logical rules over causally validated concepts. VisionLogic first learns activation thresholds that abstract neuron activations… ▽ More

    Submitted 23 February, 2026; v1 submitted 13 March, 2025; originally announced March 2025.

    Comments: 27 pages, 18 figures

  33. arXiv:2502.13834  [pdf, other

    cs.AI

    Proving Olympiad Inequalities by Synergizing LLMs and Symbolic Reasoning

    Authors: Zenan Li, Zhaoyu Li, Wen Tang, Xian Zhang, Yuan Yao, Xujie Si, Fan Yang, Kaiyu Yang, Xiaoxing Ma

    Abstract: Large language models (LLMs) can prove mathematical theorems formally by generating proof steps (\textit{a.k.a.} tactics) within a proof system. However, the space of possible tactics is vast and complex, while the available training data for formal proofs is limited, posing a significant challenge to LLM-based tactic generation. To address this, we introduce a neuro-symbolic tactic generator that… ▽ More

    Submitted 26 February, 2025; v1 submitted 19 February, 2025; originally announced February 2025.

    Comments: Published as a conference paper at ICLR 2025. Code is available at https://github.com/Lizn-zn/NeqLIPS/

  34. arXiv:2502.07829  [pdf, other

    cs.CV cs.LG

    Preference Alignment on Diffusion Model: A Comprehensive Survey for Image Generation and Editing

    Authors: Sihao Wu, Xiaonan Si, Chi Xing, Jianhong Wang, Gaojie Jin, Guangliang Cheng, Lijun Zhang, Xiaowei Huang

    Abstract: The integration of preference alignment with diffusion models (DMs) has emerged as a transformative approach to enhance image generation and editing capabilities. Although integrating diffusion models with preference alignment strategies poses significant challenges for novices at this intersection, comprehensive and systematic reviews of this subject are still notably lacking. To bridge this gap,… ▽ More

    Submitted 10 February, 2025; originally announced February 2025.

  35. arXiv:2502.05344  [pdf, other

    cs.SE cs.AI

    RAG-Verus: Repository-Level Program Verification with LLMs using Retrieval Augmented Generation

    Authors: Sicheng Zhong, Jiading Zhu, Yifang Tian, Xujie Si

    Abstract: Scaling automated formal verification to real-world projects requires resolving cross-module dependencies and global contexts, which are challenges overlooked by existing function-centric methods. We introduce RagVerus, a framework that synergizes retrieval-augmented generation with context-aware prompting to automate proof synthesis for multi-module repositories, achieving a 27% relative improvem… ▽ More

    Submitted 7 February, 2025; originally announced February 2025.

  36. arXiv:2501.08281  [pdf, ps, other

    cs.LG

    NEUROLOGIC: From Neural Representations to Interpretable Logic Rules

    Authors: Chuqin Geng, Anqi Xing, Li Zhang, Ziyu Zhao, Yuhe Jiang, Xujie Si

    Abstract: Rule-based explanation methods offer rigorous and globally interpretable insights into neural network behavior. However, existing approaches are mostly limited to small fully connected networks and depend on costly layerwise rule extraction and substitution processes. These limitations hinder their generalization to more complex architectures such as Transformers. Moreover, existing methods produc… ▽ More

    Submitted 15 October, 2025; v1 submitted 14 January, 2025; originally announced January 2025.

    Comments: 16 pages, 9 figures

  37. arXiv:2412.20801  [pdf, other

    cs.CV

    Generalize Your Face Forgery Detectors: An Insertable Adaptation Module Is All You Need

    Authors: Xiaotian Si, Linghui Li, Liwei Zhang, Ziduo Guo, Kaiguo Yuan, Bingyu Li, Xiaoyong Li

    Abstract: A plethora of face forgery detectors exist to tackle facial deepfake risks. However, their practical application is hindered by the challenge of generalizing to forgeries unseen during the training stage. To this end, we introduce an insertable adaptation module that can adapt a trained off-the-shelf detector using only online unlabeled test data, without requiring modifications to the architectur… ▽ More

    Submitted 30 December, 2024; originally announced December 2024.

    Comments: ICASSP2025 accepted

  38. arXiv:2411.12773  [pdf, other

    cs.CV

    Decoupling Training-Free Guided Diffusion by ADMM

    Authors: Youyuan Zhang, Zehua Liu, Zenan Li, Zhaoyu Li, James J. Clark, Xujie Si

    Abstract: In this paper, we consider the conditional generation problem by guiding off-the-shelf unconditional diffusion models with differentiable loss functions in a plug-and-play fashion. While previous research has primarily focused on balancing the unconditional diffusion model and the guided loss through a tuned weight hyperparameter, we propose a novel framework that distinctly decouples these two co… ▽ More

    Submitted 18 November, 2024; originally announced November 2024.

  39. arXiv:2411.00773  [pdf, other

    cs.AI

    LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation

    Authors: Bowen Li, Zhaoyu Li, Qiwei Du, Jinqi Luo, Wenshan Wang, Yaqi Xie, Simon Stepputtis, Chen Wang, Katia P. Sycara, Pradeep Kumar Ravikumar, Alexander G. Gray, Xujie Si, Sebastian Scherer

    Abstract: Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, they are usually constrained by fixed and simplistic logical rules over limited entities, making them… ▽ More

    Submitted 3 April, 2025; v1 submitted 1 November, 2024; originally announced November 2024.

    Comments: 25 pages, 8 figures, In Advances in Neural Information Processing Systems (NeurIPS) 37 D&B Track (2024): 69840-69864

    Journal ref: Advances in Neural Information Processing Systems, 37, 69840-69864 (2024)

  40. arXiv:2410.20274  [pdf, other

    cs.LG cs.CL cs.SC

    Library Learning Doesn't: The Curious Case of the Single-Use "Library"

    Authors: Ian Berlot-Attwell, Frank Rudzicz, Xujie Si

    Abstract: Advances in Large Language Models (LLMs) have spurred a wave of LLM library learning systems for mathematical reasoning. These systems aim to learn a reusable library of tools, such as formal Isabelle lemmas or Python programs that are tailored to a family of tasks. Many of these systems are inspired by the human structuring of knowledge into reusable and extendable concepts, but do current method… ▽ More

    Submitted 26 October, 2024; originally announced October 2024.

    Comments: 24 pages, 7 figures. Accepted to the 4th MATH-AI Workshop at NeurIPS'24

  41. arXiv:2409.14779  [pdf, other

    cs.AR

    Hardware/Algorithm Co-design for Real-Time I/O Control with Improved Timing Accuracy and Robustness

    Authors: Zhe Jiang, Shuai Zhao, Ran Wei, Xin Si, Gang Chen, Nan Guan

    Abstract: In safety-critical systems, timing accuracy is the key to achieving precise I/O control. To meet such strict timing requirements, dedicated hardware assistance has recently been investigated and developed. However, these solutions are often fragile, due to unforeseen timing defects. In this paper, we propose a robust and timing-accurate I/O co-processor, which manages I/O tasks using Execution Tim… ▽ More

    Submitted 23 September, 2024; originally announced September 2024.

    Comments: Accepted at the 2024 IEEE Real-Time Systems Symposium (RTSS)

    ACM Class: C.3; D.4.7

  42. arXiv:2409.04962  [pdf, other

    physics.geo-ph cs.LG

    A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys

    Authors: Hang Gao, Xinming Wu, Luming Liang, Hanlin Sheng, Xu Si, Gao Hui, Yaxing Li

    Abstract: Seismic geobody interpretation is crucial for structural geology studies and various engineering applications. Existing deep learning methods show promise but lack support for multi-modal inputs and struggle to generalize to different geobody types or surveys. We introduce a promptable foundation model for interpreting any geobodies across seismic surveys. This model integrates a pre-trained visio… ▽ More

    Submitted 13 September, 2024; v1 submitted 7 September, 2024; originally announced September 2024.

  43. arXiv:2408.09034  [pdf, ps, other

    cs.PL

    Modernizing SMT-Based Type Error Localization

    Authors: Max Kopinsky, Brigitte Pientka, Xujie Si

    Abstract: Traditional implementations of strongly-typed functional programming languages often miss the root cause of type errors. As a consequence, type error messages are often misleading and confusing - particularly for students learning such a language. We describe Tyro, a type error localization tool which determines the optimal source of an error for ill-typed programs following fundamental ideas by P… ▽ More

    Submitted 16 August, 2024; originally announced August 2024.

    Comments: 10 pages, 7 figures. About Tyro, available at https://github.com/JKTKops/tyro. To be published in FMCAD 2024

    ACM Class: D.3.3

  44. arXiv:2407.05411  [pdf, other

    cs.SE

    Assessing Code Generation with Intermediate Languages

    Authors: Xun Deng, Sicheng Zhong, Honghua Dong, Jingyu Hu, Sidi Mohamed Beillahi, Xujie Si, Fan Long

    Abstract: Intermediate step methodologies like chain of thoughts (COT) have demonstrated effectiveness in enhancing the performance of Large Language Models (LLMs) on code generation. This study explores the utilization of intermediate languages, including various programming languages, natural language solutions, and pseudo-code, and systematically evaluates their impact on the performance of LLMs in code… ▽ More

    Submitted 7 July, 2024; originally announced July 2024.

  45. arXiv:2406.13161  [pdf, other

    cs.AI cs.CL cs.LG cs.PL

    APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts

    Authors: Honghua Dong, Qidong Su, Yubo Gao, Zhaoyu Li, Yangjun Ruan, Gennady Pekhimenko, Chris J. Maddison, Xujie Si

    Abstract: Large Language Models (LLMs) have become increasingly capable of handling diverse tasks with the aid of well-crafted prompts and integration of external tools, but as task complexity rises, the workflow involving LLMs can be complicated and thus challenging to implement and maintain. To address this challenge, we propose APPL, A Prompt Programming Language that acts as a bridge between computer pr… ▽ More

    Submitted 18 June, 2024; originally announced June 2024.

  46. Contextual Distillation Model for Diversified Recommendation

    Authors: Fan Li, Xu Si, Shisong Tang, Dingmin Wang, Kunyan Han, Bing Han, Guorui Zhou, Yang Song, Hechang Chen

    Abstract: The diversity of recommendation is equally crucial as accuracy in improving user experience. Existing studies, e.g., Determinantal Point Process (DPP) and Maximal Marginal Relevance (MMR), employ a greedy paradigm to iteratively select items that optimize both accuracy and diversity. However, prior methods typically exhibit quadratic complexity, limiting their applications to the re-ranking stage… ▽ More

    Submitted 14 August, 2024; v1 submitted 13 June, 2024; originally announced June 2024.

    Comments: accepted by KDD 2024 v2

  47. arXiv:2405.17503  [pdf, other

    cs.SE cs.AI cs.CL cs.PL

    Code Repair with LLMs gives an Exploration-Exploitation Tradeoff

    Authors: Hao Tang, Keya Hu, Jin Peng Zhou, Sicheng Zhong, Wei-Long Zheng, Xujie Si, Kevin Ellis

    Abstract: Iteratively improving and repairing source code with large language models (LLMs), known as refinement, has emerged as a popular way of generating programs that would be too complex to construct in one shot. Given a bank of test cases, together with a candidate program, an LLM can improve that program by being prompted with failed test cases. But it remains an open question how to best iteratively… ▽ More

    Submitted 29 October, 2024; v1 submitted 26 May, 2024; originally announced May 2024.

  48. arXiv:2405.17216  [pdf, other

    cs.LG cs.AI cs.LO stat.ML

    Autoformalizing Euclidean Geometry

    Authors: Logan Murphy, Kaiyu Yang, Jialiang Sun, Zhaoyu Li, Anima Anandkumar, Xujie Si

    Abstract: Autoformalization involves automatically translating informal math into formal theorems and proofs that are machine-verifiable. Euclidean geometry provides an interesting and controllable domain for studying autoformalization. In this paper, we introduce a neuro-symbolic framework for autoformalizing Euclidean geometry, which combines domain knowledge, SMT solvers, and large language models (LLMs)… ▽ More

    Submitted 27 May, 2024; originally announced May 2024.

    Comments: Accepted to ICML 2024. The first two authors contributed equally

  49. arXiv:2404.09939  [pdf, other

    cs.AI

    A Survey on Deep Learning for Theorem Proving

    Authors: Zhaoyu Li, Jialiang Sun, Logan Murphy, Qidong Su, Zenan Li, Xian Zhang, Kaiyu Yang, Xujie Si

    Abstract: Theorem proving is a fundamental aspect of mathematics, spanning from informal reasoning in natural language to rigorous derivations in formal systems. In recent years, the advancement of deep learning, especially the emergence of large language models, has sparked a notable surge of research exploring these techniques to enhance the process of theorem proving. This paper presents a comprehensive… ▽ More

    Submitted 21 August, 2024; v1 submitted 15 April, 2024; originally announced April 2024.

  50. arXiv:2404.04731  [pdf, other

    cs.PL cs.SE

    SAT-DIFF: A Tree Diffing Framework Using SAT Solving

    Authors: Chuqin Geng, Haolin Ye, Yihan Zhang, Brigitte Pientka, Xujie Si

    Abstract: Computing differences between tree-structured data is a critical but challenging problem in software analysis. In this paper, we propose a novel tree diffing approach called SatDiff, which reformulates the structural diffing problem into a MaxSAT problem. By encoding the necessary transformations from the source tree to the target tree, SatDiff generates correct, minimal, and type safe low-level e… ▽ More

    Submitted 6 April, 2024; originally announced April 2024.

    Comments: 23 pages, 7 figures