-
Beyond Local Berry Geometry: A First-Principles Finite-Momentum Theory of Electronic Position
Authors:
M. S. Si,
Y. Q. Li,
G. P. Zhang
Abstract:
Electronic position controls how a crystal polarizes and responds to an external field. In crystals, it is usually described through local changes of electronic states in momentum space. This Berry framework has reshaped modern solid-state physics, but strong fields drive electrons across a finite momentum range, where coherence between different momenta becomes part of the response. Here we estab…
▽ More
Electronic position controls how a crystal polarizes and responds to an external field. In crystals, it is usually described through local changes of electronic states in momentum space. This Berry framework has reshaped modern solid-state physics, but strong fields drive electrons across a finite momentum range, where coherence between different momenta becomes part of the response. Here we establish a first-principles theory of electronic position at finite momentum that retains this missing information. We show that unequal-momentum coherence can cancel under spatial averaging and still produce polarization, forming a coherence dipole. We obtain the matrix directly from material wave functions, without model bands or fitted transition elements. In Si, the finite-momentum geometry sets a material momentum scale. Comparing this scale with the momentum change driven by the field predicts when finite-momentum physics becomes active. Crossing the scale strongly reorganizes the fifth and higher harmonics, showing that momentum-space geometry, rather than emitted photon energy alone, controls the nonlinear response. HHG is the first demonstration, but the theory applies whenever driven electrons explore a finite momentum range. It therefore extends quantum geometry beyond the local Berry limit and provides a general basis for predicting field-driven phenomena in real materials.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
Observability inequality for the wave equation
Authors:
Suliang Si
Abstract:
In this paper, only Carleman estimates are used, without energy estimates, we derive observability inequality. The main tool consists in the use of a new Carleman estimate.
In this paper, only Carleman estimates are used, without energy estimates, we derive observability inequality. The main tool consists in the use of a new Carleman estimate.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
EnvRL: Learn from Environment Dynamics in Agentic Reinforcement Learning
Authors:
Zhitong Wang,
Songze Li,
Hao Peng,
Shuzheng Si,
Yi Wang,
Maosong Sun,
Juanzi Li
Abstract:
Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents. However, conventional RL methods for long-horizon agentic tasks often struggle with sparse outcome rewards. Intuitively, this overlooks the rich environment dynamics information contained in rollout interaction trajectories. We argue that the interaction experience inherently serves…
▽ More
Reinforcement learning (RL) has emerged as a powerful paradigm for training Large Language Models (LLMs) as agents. However, conventional RL methods for long-horizon agentic tasks often struggle with sparse outcome rewards. Intuitively, this overlooks the rich environment dynamics information contained in rollout interaction trajectories. We argue that the interaction experience inherently serves as an implicit supervision signal, reveals the underlying transition mechanisms of the environment, and enables the agent to construct a more accurate internal model of the environment.. Therefore, in this work, we investigate how to leverage this additional signal to improve policy learning. Specifically, we propose EnvRL, a framework that incorporates environment dynamics learning into agentic RL via two auxiliary objectives: state prediction and inverse dynamics. By jointly optimizing with the primary RL objective, we encourage the agent to internalize environment dynamics from its own interaction experience. Extensive experiments on two long-horizon agentic benchmarks demonstrate that EnvRL achieves significant improvements on success-rates over RL-only baselines, e.g., when trained with GRPO, lifting Qwen-2.5-1.5B-Instruct from 72.8% to 77.4% on ALFWorld, and from 56.8% to 67.0% on WebShop.
△ Less
Submitted 16 June, 2026;
originally announced June 2026.
-
APEX: Automated Prompt Engineering eXpert with Dynamic Data Selection
Authors:
Fei Wang,
Si Si,
Cho-Jui Hsieh,
Inderjit S. Dhillon
Abstract:
Large Language Models are highly sensitive to prompt formulation, necessitating automatic prompt optimization to unlock their full potential. While evolutionary algorithms have emerged as the dominant paradigm, they suffer from a critical bottleneck: data efficiency. Current methods treat the development dataset as a static benchmark, wasting significant compute budget on uninformative data. In th…
▽ More
Large Language Models are highly sensitive to prompt formulation, necessitating automatic prompt optimization to unlock their full potential. While evolutionary algorithms have emerged as the dominant paradigm, they suffer from a critical bottleneck: data efficiency. Current methods treat the development dataset as a static benchmark, wasting significant compute budget on uninformative data. In this work, we introduce APEX (Automatic Prompt Engineering eXpert), a novel framework that optimizes the data usage alongside the prompt search. APEX dynamically stratifies the dataset into Easy, Hard, and Mixed tiers based on the optimization lineage. By prioritizing the Mixed tier, which identifies the data where the LLM has mixed performance, we identify two high-leverage subsets: the addressable frontier for generating informative mutations and the rank-sensitive frontier for distinguishing candidate quality. We evaluate APEX across three diverse benchmarks: IFBench, SimpleQA Verified, and FACTS Grounding. Under a fixed budget of 5,000 evaluation calls, due to its data efficiency, APEX outperforms the initial prompt by an average of 11.2% on Gemini 2.5 Flash and 6.8% on Gemma 3 27B, demonstrating that a data-centric approach is key to efficient and effective prompt optimization.
△ Less
Submitted 9 June, 2026;
originally announced June 2026.
-
MAAM: Anchor-Preserving Compression and Contextual Calibration for Chinese Discriminatory Language Detection
Authors:
Yuxin Fu,
Shijing Si
Abstract:
Chinese discriminatory-language detection is challenging because harmful intent is often implicit and context-dependent. We propose MAAM (Myopia--Astigmatism Anchor Mechanism), a lightweight, model-agnostic framework inspired by functional visual blur: rather than preserving every token equally, MAAM retains discrimination-relevant semantic anchors and calibrates them with C--I--S contextual prior…
▽ More
Chinese discriminatory-language detection is challenging because harmful intent is often implicit and context-dependent. We propose MAAM (Myopia--Astigmatism Anchor Mechanism), a lightweight, model-agnostic framework inspired by functional visual blur: rather than preserving every token equally, MAAM retains discrimination-relevant semantic anchors and calibrates them with C--I--S contextual priors (Contextual Tone, Group Identity, and Stance Polarity). We also introduce ChLGBT, to our knowledge the first Chinese LGBT-focused discriminatory-language dataset, with 8,120 manually annotated samples and three ordinal labels: explicit bias, implicit bias, and emotional intensity. Across strong encoder baselines, MAAM improves all three prediction dimensions, with consistent gains in accuracy, F1, Brier score, and expected calibration error. Compared with frontier LLM baselines under zero-shot and few-shot prompting protocols, MAAM remains competitive while offering stronger compactness and stability. These results suggest that interpretable anchor preservation and contextual calibration provide a practical alternative to heavier model scaling for Chinese discriminatory-language assessment.
△ Less
Submitted 8 June, 2026;
originally announced June 2026.
-
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
Authors:
Haozhe Zhao,
Shuzheng Si,
Zhenhailong Wang,
Zheng Wang,
Liang Chen,
Xiaotong Li,
Zhixiang Liang,
Maosong Sun,
Minjia Zhang
Abstract:
Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains one of the most labor-intensive parts of paper preparation. Existing automated systems each target a single figure type under text-only input, leaving the diversity of types and conditions researchers actually use unaddressed; their raster outputs f…
▽ More
Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains one of the most labor-intensive parts of paper preparation. Existing automated systems each target a single figure type under text-only input, leaving the diversity of types and conditions researchers actually use unaddressed; their raster outputs further cannot be locally revised. Because scientific figures are structured compositions of discrete semantic components, the localized errors generators produce on such layouts demand not a stronger backbone but a harness. We instantiate this harness in two complementary systems: Crafter, a multi-agent harness for figure generation that generalizes across figure types and input conditions without architectural changes, and CraftEditor, which applies the same pattern to convert raster outputs into editable SVGs. Moreover, we introduce CraftBench, a benchmark spanning three figure types and four input conditions with human quality annotation. Experiments show that Crafter substantially outperforms both standalone generators and the agentic baseline on PaperBanana-Bench and CraftBench, with ablations confirming each component's independent contribution; CraftEditor faithfully converts outputs into editable SVGs that surpass all baselines. Our code and benchmark are available at https://github.com/HaozheZhao/Crafter.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
Authors:
Shengyu Si,
Yuanzhuo Lu,
Ruimeng Yang,
Ziyi Ye,
Zuxuan Wu,
Yu-Gang Jiang
Abstract:
Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories…
▽ More
Vision-Language-Action~(VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories at training time and transferring these memories during inference. Specifically, VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk. Experiments on RoboTwin, RLBench, and real-world manipulation tasks show that VLA-Pro consistently improves cross-task generalization across multiple backbones, achieving up to a 207% relative improvement in simulation and increasing real-world success rate from 5.8% to 65.0%. These results suggest that procedural memory retrieval and adaptation provide an effective mechanism for transferring manipulation experience to novel tasks while preserving modularity and execution stability.
△ Less
Submitted 28 May, 2026;
originally announced May 2026.
-
Pitfall of Precision in Noisy Signaling
Authors:
Shuhua Si,
Yangfan Zhou
Abstract:
A principal decides whether to approve an agent based on a noisy signal (e.g., test scores) generated by the agent. High-quality agents can produce high signals on average at lower cost, but the realizations are subject to noise that depends on the screening technology's precision. We uncover a paradoxical "pitfall of precision": when precision is already high, further improvements reduce screenin…
▽ More
A principal decides whether to approve an agent based on a noisy signal (e.g., test scores) generated by the agent. High-quality agents can produce high signals on average at lower cost, but the realizations are subject to noise that depends on the screening technology's precision. We uncover a paradoxical "pitfall of precision": when precision is already high, further improvements reduce screening accuracy and lower the principal's welfare. This occurs because greater precision incentivizes strategic signaling from more low-quality agents, outweighing the direct benefit from improved precision. The pitfall of precision also has implications for statistical discrimination: groups with noisier technologies face lower approval rates yet may be favored ex ante -- a reversal of discrimination. We also examine how commitment power helps mitigate the pitfall.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Is Complexity the Problem? Testing Random Choice with Heterogeneity
Authors:
Shuhua Si
Abstract:
Economic choices are often stochastic: the same person may make a different choice when facing the same alternatives repeatedly. Standard models assume that the degree of randomness reflects the size of utility differences, but choice inconsistencies could also reflect difficulty comparing alternatives. Recent studies estimate such comparison difficulty (or "complexity") by fitting functional form…
▽ More
Economic choices are often stochastic: the same person may make a different choice when facing the same alternatives repeatedly. Standard models assume that the degree of randomness reflects the size of utility differences, but choice inconsistencies could also reflect difficulty comparing alternatives. Recent studies estimate such comparison difficulty (or "complexity") by fitting functional forms to aggregate choice data under a representative agent assumption. However, aggregate data could violate standard models of random choice simply because of heterogeneity in preferences, even in the absence of variation in comparison difficulty. This paper develops a revealed preference framework, collective rationalizability, that tests for variation in comparison difficulty from aggregate data while explicitly accounting for heterogeneity. The framework characterizes whether violations of standard models can be explained by comparison difficulty alone, heterogeneity alone, or require both. I provide a statistical test with finite-sample inference and apply the method to two existing experiments. In both cases, heterogeneity alone explains observed failures of stochastic transitivity well, demonstrating that comparison difficulty can be not only theoretically but also empirically confused with heterogeneity in aggregate data.
△ Less
Submitted 3 May, 2026;
originally announced May 2026.
-
From Context to Skills: Can Language Models Learn from Context Skillfully?
Authors:
Shuzheng Si,
Haozhe Zhao,
Yu Lei,
Qingyi Wang,
Dingwei Chen,
Zhitong Wang,
Zhenhailong Wang,
Kangyang Luo,
Zheng Wang,
Gang Chen,
Fanchao Qi,
Minjia Zhang,
Maosong Sun
Abstract:
Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for context learning, where LMs directly learn relevant knowledge from the given context. An intuitive solution is inference-time skill augmentation: extracting the rules and procedures from context into natural-language skills. However, constructing such skills fo…
▽ More
Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for context learning, where LMs directly learn relevant knowledge from the given context. An intuitive solution is inference-time skill augmentation: extracting the rules and procedures from context into natural-language skills. However, constructing such skills for context learning scenarios faces two challenges: the prohibitive cost of manual skill annotation for long, technically dense contexts, and the lack of external feedback for automated skill construction. In this paper, we propose Ctx2Skill, a self-evolving framework that autonomously discovers, refines, and selects context-specific skills without human supervision or external feedback. At its core, a multi-agent self-play loop has a Challenger that generates probing tasks and rubrics, a Reasoner that attempts to solve them guided by an evolving skill set, and a neutral Judge that provides binary feedback. Crucially, both the Challenger and the Reasoner evolve through accumulated skills: dedicated Proposer and Generator agents analyze failure cases and synthesize them into targeted skill updates for both sides, enabling automated skill discovery and refinement. To prevent adversarial collapse caused by increasingly extreme task generation and over-specialized skill accumulation, we further introduce a Cross-time Replay mechanism that identifies the skill set achieving the best balance across representative cases for the Reasoner side, ensuring robust and generalizable skill evolution. The resulting skills can be plugged into any language model to obtain better context learning capability. Evaluated on four context learning tasks from CL-bench, Ctx2Skill consistently improves solving rates across backbone models.
△ Less
Submitted 27 July, 2026; v1 submitted 30 April, 2026;
originally announced April 2026.
-
KARL: Mitigating Hallucinations in LLMs via Knowledge-Boundary-Aware Reinforcement Learning
Authors:
Cheng Gao,
Cheng Huang,
Kangyang Luo,
Ziqing Qiao,
Shuzheng Si,
Huimin Chen,
Chaojun Xiao,
Maosong Sun
Abstract:
Enabling large language models (LLMs) to appropriately abstain from answering questions beyond their knowledge is crucial for mitigating hallucinations. While existing reinforcement learning methods foster autonomous abstention, they often compromise answer accuracy because their static reward mechanisms, agnostic to models' knowledge boundaries, drive models toward excessive caution. In this work…
▽ More
Enabling large language models (LLMs) to appropriately abstain from answering questions beyond their knowledge is crucial for mitigating hallucinations. While existing reinforcement learning methods foster autonomous abstention, they often compromise answer accuracy because their static reward mechanisms, agnostic to models' knowledge boundaries, drive models toward excessive caution. In this work, we propose KARL, a novel framework that continuously aligns an LLM's abstention behavior with its evolving knowledge boundary. KARL introduces two core innovations: a Knowledge-Boundary-Aware Reward that performs online knowledge boundary estimation using within-group response statistics, dynamically rewarding correct answers or guided abstention; and a Two-Stage RL Training Strategy that first explores the knowledge boundary and bypasses the "abstention trap", and subsequently converts incorrect answers beyond the knowledge boundary into abstentions without sacrificing accuracy. Extensive experiments on multiple benchmarks demonstrate that KARL achieves a superior accuracy-hallucination trade-off, effectively suppressing hallucinations while maintaining high accuracy across both in-distribution and out-of-distribution scenarios.
△ Less
Submitted 3 April, 2026;
originally announced April 2026.
-
GroupDPO: Memory efficient Group-wise Direct Preference Optimization
Authors:
Jixuan Leng,
Si Si,
Hsiang-Fu Yu,
Vinod Raman,
Inderjit S. Dhillon
Abstract:
Preference optimization is widely used to align Large Language Models (LLMs) with preference feedback. However, most existing methods train on a single positive-negative pair per prompt, discarding additional supervision available in preference datasets that typically contain multiple candidate responses. Motivated by this limitation, recent work explores group-wise preference optimization, which…
▽ More
Preference optimization is widely used to align Large Language Models (LLMs) with preference feedback. However, most existing methods train on a single positive-negative pair per prompt, discarding additional supervision available in preference datasets that typically contain multiple candidate responses. Motivated by this limitation, recent work explores group-wise preference optimization, which jointly contrasts multiple responses for the same prompt, but its empirical behavior and scalability remain underexplored due to the memory overhead of group-coupled objectives. In this work, we introduce a memory-efficient group-wise preference optimization algorithm that preserves gradients while decoupling samples during backpropagation, substantially reducing peak memory usage, which enables scalable training with larger group sizes. Across both offline and online alignment settings, we show that leveraging multiple responses consistently outperforms single-pair training. Furthermore, incorporating a negative log-likelihood (NLL) term on positive responses is critical for both performance gains and training stability.
△ Less
Submitted 16 April, 2026;
originally announced April 2026.
-
Arefinement of the Bukhgeim-Klibanov method
Authors:
Suliang Si
Abstract:
In this article, we improve the classical Bukhgeim-Klibanov method presented in [1],which can be used to prove the conditional stability of inverse source problem for a hyperbolic equation from the measurement on the subboundary. A major ingredient of our proof is a novel Carleman estimate. This inequality eliminates the need to extend the solution in time, therefore simplifies the existing proofs…
▽ More
In this article, we improve the classical Bukhgeim-Klibanov method presented in [1],which can be used to prove the conditional stability of inverse source problem for a hyperbolic equation from the measurement on the subboundary. A major ingredient of our proof is a novel Carleman estimate. This inequality eliminates the need to extend the solution in time, therefore simplifies the existing proofs, which is widely applicable to various evolution equations.
△ Less
Submitted 26 March, 2026;
originally announced March 2026.
-
InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs
Authors:
Yuzhuo Bai,
Shuzheng Si,
Kangyang Luo,
Qingyi Wang,
Wenhao Li,
Gang Chen,
Fanchao Qi,
Maosong Sun
Abstract:
Large language models (LLMs) often hallucinate, yet most existing fact-checking methods treat factuality evaluation as a binary classification problem, offering limited interpretability and failing to capture fine-grained error types. In this paper, we introduce InFi-Check, a framework for interpretable and fine-grained fact-checking of LLM outputs. Specifically, we first propose a controlled data…
▽ More
Large language models (LLMs) often hallucinate, yet most existing fact-checking methods treat factuality evaluation as a binary classification problem, offering limited interpretability and failing to capture fine-grained error types. In this paper, we introduce InFi-Check, a framework for interpretable and fine-grained fact-checking of LLM outputs. Specifically, we first propose a controlled data synthesis pipeline that generates high-quality data featuring explicit evidence, fine-grained error type labels, justifications, and corrections. Based on this, we further construct large-scale training data and a manually verified benchmark InFi-Check-FG for fine-grained fact-checking of LLM outputs. Building on these high-quality training data, we further propose InFi-Checker, which can jointly provide supporting evidence, classify fine-grained error types, and produce justifications along with corrections. Experiments show that InFi-Checker achieves state-of-the-art performance on InFi-Check-FG and strong generalization across various downstream tasks, significantly improving the utility and trustworthiness of factuality evaluation.
△ Less
Submitted 10 January, 2026;
originally announced January 2026.
-
BabyVision: Visual Reasoning Beyond Language
Authors:
Liang Chen,
Weichu Xie,
Yiyan Liang,
Hongfeng He,
Hans Zhao,
Zhibo Yang,
Zhiqi Huang,
Haoning Wu,
Haoyu Lu,
Y. charles,
Yiping Bao,
Yuantao Fan,
Guopeng Li,
Haiyang Shen,
Xuanzhong Chen,
Wendong Xu,
Shuzheng Si,
Zefan Cai,
Wenhao Chai,
Ziqi Huang,
Fangfu Liu,
Tianyu Liu,
Baobao Chang,
Ming Wu,
Xiaobo Hu
, et al. (5 additional authors not shown)
Abstract:
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introdu…
▽ More
While humans develop core visual skills long before acquiring language, contemporary Multimodal LLMs (MLLMs) still rely heavily on linguistic priors to compensate for their fragile visual understanding. We uncovered a crucial fact: state-of-the-art MLLMs consistently fail on basic visual tasks that humans, even 3-year-olds, can solve effortlessly. To systematically investigate this gap, we introduce BabyVision, a benchmark designed to assess core visual abilities independent of linguistic knowledge for MLLMs. BabyVision spans a wide range of tasks, with 388 items divided into 22 subclasses across four key categories. Empirical results and human evaluation reveal that leading MLLMs perform significantly below human baselines. Gemini3-Pro-Preview scores 49.7, lagging behind 6-year-old humans and falling well behind the average adult score of 94.1. These results show despite excelling in knowledge-heavy evaluations, current MLLMs still lack fundamental visual primitives. Progress in BabyVision represents a step toward human-level visual perception and reasoning capabilities. We also explore solving visual reasoning with generation models by proposing BabyVision-Gen and automatic evaluation toolkit. Our code and benchmark data are released at https://github.com/UniPat-AI/BabyVision for reproduction.
△ Less
Submitted 7 July, 2026; v1 submitted 10 January, 2026;
originally announced January 2026.
-
Constraining Primordial Black Holes via p-wave annihilation in light of CMB Spectral Distortion and 21-cm global signal
Authors:
Shibsankar Si,
Pravin Kumar Natwariya,
Alekha C. Nayak
Abstract:
Primordial black holes (PBHs) can form spike density halos through the accretion of weakly interacting massive particles (WIMPs). In these halos, the enhanced density significantly boosts the annihilation rate of WIMPs. For Majorana dark matter annihilation into light fermions, the s-wave part of the annihilation cross section is helicity-suppressed, making the p-wave contribution dominant. We stu…
▽ More
Primordial black holes (PBHs) can form spike density halos through the accretion of weakly interacting massive particles (WIMPs). In these halos, the enhanced density significantly boosts the annihilation rate of WIMPs. For Majorana dark matter annihilation into light fermions, the s-wave part of the annihilation cross section is helicity-suppressed, making the p-wave contribution dominant. We study the velocity-dependent p-wave annihilation case, whose resulting energy injection can modify the thermal and ionization history of the Universe, leaving observable imprints on the cosmic microwave background (CMB) spectrum and the global 21-cm signal. From the predicted energy injection into the plasma, we derive stringent upper limits on the fraction of dark matter in form of PBHs for p-wave annihilation models, based on the observational constraints of the CMB spectral distortions ($y$-type), and from the measurement of the 21-cm absorption signal at cosmic dawn. Our results highlight that accounting for the p-wave nature of annihilation is crucial for deriving robust constraints on the PBH abundance.
△ Less
Submitted 1 January, 2026;
originally announced January 2026.
-
MEIC-DT: Memory-Efficient Incremental Clustering for Long-Text Coreference Resolution with Dual-Threshold Constraints
Authors:
Kangyang Luo,
Shuzheng Si,
Yuzhuo Bai,
Cheng Gao,
Zhitong Wang,
Cheng Huang,
Yingli Shen,
Yufeng Han,
Wenhao Li,
Cunliang Kong,
Maosong Sun
Abstract:
In the era of large language models (LLMs), supervised neural methods remain the state-of-the-art (SOTA) for Coreference Resolution. Yet, their full potential is underexplored, particularly in incremental clustering, which faces the critical challenge of balancing efficiency with performance for long texts. To address the limitation, we propose \textbf{MEIC-DT}, a novel dual-threshold, memory-effi…
▽ More
In the era of large language models (LLMs), supervised neural methods remain the state-of-the-art (SOTA) for Coreference Resolution. Yet, their full potential is underexplored, particularly in incremental clustering, which faces the critical challenge of balancing efficiency with performance for long texts. To address the limitation, we propose \textbf{MEIC-DT}, a novel dual-threshold, memory-efficient incremental clustering approach based on a lightweight Transformer. MEIC-DT features a dual-threshold constraint mechanism designed to precisely control the Transformer's input scale within a predefined memory budget. This mechanism incorporates a Statistics-Aware Eviction Strategy (\textbf{SAES}), which utilizes distinct statistical profiles from the training and inference phases for intelligent cache management. Furthermore, we introduce an Internal Regularization Policy (\textbf{IRP}) that strategically condenses clusters by selecting the most representative mentions, thereby preserving semantic integrity. Extensive experiments on common benchmarks demonstrate that MEIC-DT achieves highly competitive coreference performance under stringent memory constraints.
△ Less
Submitted 6 May, 2026; v1 submitted 31 December, 2025;
originally announced December 2025.
-
FaithLens: Detecting and Explaining Faithfulness Hallucination
Authors:
Shuzheng Si,
Qingyi Wang,
Haozhe Zhao,
Yuzhuo Bai,
Guanqiao Chen,
Kangyang Luo,
Gang Chen,
Fanchao Qi,
Minjia Zhang,
Baobao Chang,
Maosong Sun
Abstract:
Recognizing whether outputs from large language models (LLMs) contain faithfulness hallucination is crucial for real-world applications, e.g., retrieval-augmented generation and summarization. In this paper, we introduce FaithLens, a cost-efficient and effective faithfulness hallucination detection model that can jointly provide binary predictions and corresponding explanations to improve trustwor…
▽ More
Recognizing whether outputs from large language models (LLMs) contain faithfulness hallucination is crucial for real-world applications, e.g., retrieval-augmented generation and summarization. In this paper, we introduce FaithLens, a cost-efficient and effective faithfulness hallucination detection model that can jointly provide binary predictions and corresponding explanations to improve trustworthiness. To achieve this, we first synthesize training data with explanations via advanced LLMs and apply a well-defined data filtering strategy to ensure label correctness, explanation quality, and data diversity. Subsequently, we fine-tune the model on these well-curated training data as a cold start and further optimize it with rule-based reinforcement learning, using rewards for both prediction correctness and explanation quality. Results on 12 diverse tasks show that the 8B-parameter FaithLens outperforms advanced models such as GPT-5.2 and o3. Also, FaithLens can produce high-quality explanations, delivering a distinctive balance of trustworthiness, efficiency, and effectiveness.
△ Less
Submitted 21 April, 2026; v1 submitted 23 December, 2025;
originally announced December 2025.
-
From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition
Authors:
Yiqing Zhou,
Yu Lei,
Shuzheng Si,
Qingyan Sun,
Wei Wang,
Yifei Wu,
Hao Wen,
Gang Chen,
Fanchao Qi,
Maosong Sun
Abstract:
Managing extensive context remains a critical bottleneck for Large Language Models (LLMs), particularly in applications like long-document question answering and autonomous agents where lengthy inputs incur high computational costs and introduce noise. Existing compression techniques often disrupt local coherence through discrete token removal or rely on implicit latent encoding that suffers from…
▽ More
Managing extensive context remains a critical bottleneck for Large Language Models (LLMs), particularly in applications like long-document question answering and autonomous agents where lengthy inputs incur high computational costs and introduce noise. Existing compression techniques often disrupt local coherence through discrete token removal or rely on implicit latent encoding that suffers from positional bias and incompatibility with closed-source APIs. To address these limitations, we introduce the EDU-based Context Compressor, a novel explicit compression framework designed to preserve both global structure and fine-grained details. Our approach reformulates context compression as a structure-then-select process. First, our LingoEDU transforms linear text into a structural relation tree of Elementary Discourse Units (EDUs) which are anchored strictly to source indices to eliminate hallucination. Second, a lightweight ranking module selects query-relevant sub-trees for linearization. To rigorously evaluate structural understanding, we release StructBench, a manually annotated dataset of 248 diverse documents. Empirical results demonstrate that our method achieves state-of-the-art structural prediction accuracy and significantly outperforms frontier LLMs while reducing costs. Furthermore, our structure-aware compression substantially enhances performance across downstream tasks ranging from long-context tasks to complex Deep Search scenarios.
△ Less
Submitted 5 January, 2026; v1 submitted 16 December, 2025;
originally announced December 2025.
-
RhinoInsight: Improving Deep Research through Control Mechanisms for Model Behavior and Context
Authors:
Yu Lei,
Shuzheng Si,
Wei Wang,
Yifei Wu,
Gang Chen,
Fanchao Qi,
Maosong Sun
Abstract:
Large language models are evolving from single-turn responders into tool-using agents capable of sustained reasoning and decision-making for deep research. Prevailing systems adopt a linear pipeline of plan to search to write to a report, which suffers from error accumulation and context rot due to the lack of explicit control over both model behavior and context. We introduce RhinoInsight, a deep…
▽ More
Large language models are evolving from single-turn responders into tool-using agents capable of sustained reasoning and decision-making for deep research. Prevailing systems adopt a linear pipeline of plan to search to write to a report, which suffers from error accumulation and context rot due to the lack of explicit control over both model behavior and context. We introduce RhinoInsight, a deep research framework that adds two control mechanisms to enhance robustness, traceability, and overall quality without parameter updates. First, a Verifiable Checklist module transforms user requirements into traceable and verifiable sub-goals, incorporates human or LLM critics for refinement, and compiles a hierarchical outline to anchor subsequent actions and prevent non-executable planning. Second, an Evidence Audit module structures search content, iteratively updates the outline, and prunes noisy context, while a critic ranks and binds high-quality evidence to drafted content to ensure verifiability and reduce hallucinations. Our experiments demonstrate that RhinoInsight achieves state-of-the-art performance on deep research tasks while remaining competitive on deep search tasks.
△ Less
Submitted 23 November, 2025;
originally announced November 2025.
-
Bag of Tricks for Subverting Reasoning-based Safety Guardrails
Authors:
Shuo Chen,
Zhen Han,
Haokun Chen,
Bailan He,
Shengyun Si,
Jingpei Wu,
Philip Torr,
Volker Tresp,
Jindong Gu
Abstract:
Recent reasoning-based safety guardrails for Large Reasoning Models (LRMs), such as deliberative alignment, have shown strong defense against jailbreak attacks. By leveraging LRMs' reasoning ability, these guardrails help the models to assess the safety of user inputs before generating final responses. The powerful reasoning ability can analyze the intention of the input query and will refuse to a…
▽ More
Recent reasoning-based safety guardrails for Large Reasoning Models (LRMs), such as deliberative alignment, have shown strong defense against jailbreak attacks. By leveraging LRMs' reasoning ability, these guardrails help the models to assess the safety of user inputs before generating final responses. The powerful reasoning ability can analyze the intention of the input query and will refuse to assist once it detects the harmful intent hidden by the jailbreak methods. Such guardrails have shown a significant boost in defense, such as the near-perfect refusal rates on the open-source gpt-oss series. Unfortunately, we find that these powerful reasoning-based guardrails can be extremely vulnerable to subtle manipulation of the input prompts, and once hijacked, can lead to even more harmful results. Specifically, we first uncover a surprisingly fragile aspect of these guardrails: simply adding a few template tokens to the input prompt can successfully bypass the seemingly powerful guardrails and lead to explicit and harmful responses. To explore further, we introduce a bag of jailbreak methods that subvert the reasoning-based guardrails. Our attacks span white-, gray-, and black-box settings and range from effortless template manipulations to fully automated optimization. Along with the potential for scalable implementation, these methods also achieve alarmingly high attack success rates (e.g., exceeding 90% across 5 different benchmarks on gpt-oss series on both local host models and online API services). Evaluations across various leading open-source LRMs confirm that these vulnerabilities are systemic, underscoring the urgent need for stronger alignment techniques for open-sourced LRMs to prevent malicious misuse. Code is open-sourced at https://chenxshuo.github.io/bag-of-tricks.
△ Less
Submitted 22 October, 2025; v1 submitted 13 October, 2025;
originally announced October 2025.
-
ImCoref-CeS: An Improved Lightweight Pipeline for Coreference Resolution with LLM-based Checker-Splitter Refinement
Authors:
Kangyang Luo,
Yuzhuo Bai,
Shuzheng Si,
Cheng Gao,
Zhitong Wang,
Yingli Shen,
Wenhao Li,
Zhu Liu,
Yufeng Han,
Jiayi Wu,
Cunliang Kong,
Maosong Sun
Abstract:
Coreference Resolution (CR) is a critical task in Natural Language Processing (NLP). Current research faces a key dilemma: whether to further explore the potential of supervised neural methods based on small language models, whose detect-then-cluster pipeline still delivers top performance, or embrace the powerful capabilities of Large Language Models (LLMs). However, effectively combining their s…
▽ More
Coreference Resolution (CR) is a critical task in Natural Language Processing (NLP). Current research faces a key dilemma: whether to further explore the potential of supervised neural methods based on small language models, whose detect-then-cluster pipeline still delivers top performance, or embrace the powerful capabilities of Large Language Models (LLMs). However, effectively combining their strengths remains underexplored. To this end, we propose \textbf{ImCoref-CeS}, a novel framework that integrates an enhanced supervised model with LLM-based reasoning. First, we present an improved CR method (\textbf{ImCoref}) to push the performance boundaries of the supervised neural method by introducing a lightweight bridging module to enhance long-text encoding capability, devising a biaffine scorer to comprehensively capture positional information, and invoking a hybrid mention regularization to improve training efficiency. Importantly, we employ an LLM acting as a multi-role Checker-Splitter agent to validate candidate mentions (filtering out invalid ones) and coreference results (splitting erroneous clusters) predicted by ImCoref. Extensive experiments demonstrate the effectiveness of ImCoref-CeS, which achieves superior performance compared to existing state-of-the-art (SOTA) methods.
△ Less
Submitted 6 May, 2026; v1 submitted 11 October, 2025;
originally announced October 2025.
-
A Goal Without a Plan Is Just a Wish: Efficient and Effective Global Planner Training for Long-Horizon Agent Tasks
Authors:
Shuzheng Si,
Haozhe Zhao,
Kangyang Luo,
Gang Chen,
Fanchao Qi,
Minjia Zhang,
Baobao Chang,
Maosong Sun
Abstract:
Agents based on large language models (LLMs) struggle with brainless trial-and-error and generating hallucinatory actions due to a lack of global planning in long-horizon tasks. In this paper, we introduce a plan-and-execute framework and propose EAGLET, an efficient and effective planner training method to enhance the executor agent's planning abilities without human effort. Specifically, we trai…
▽ More
Agents based on large language models (LLMs) struggle with brainless trial-and-error and generating hallucinatory actions due to a lack of global planning in long-horizon tasks. In this paper, we introduce a plan-and-execute framework and propose EAGLET, an efficient and effective planner training method to enhance the executor agent's planning abilities without human effort. Specifically, we train a plug-and-play global planner through a two-step process: we first synthesize high-quality plans from an advanced LLM using our proposed homologous consensus filtering strategy, and apply fine-tuning as a cold start. Moreover, we further improve the planner with a rule-based reinforcement learning stage using a novel executor capability gain reward, ensuring it can handle task instructions of varying difficulty. Experiments on three long-horizon agent tasks show that executor agents equipped with our planner outperform existing methods, achieving new state-of-the-art performance. Meanwhile, EAGLET reduces training costs by 8x compared to RL-based baselines, and it does not require manual effort or extra training data, offering an efficient and effective solution.
△ Less
Submitted 21 April, 2026; v1 submitted 7 October, 2025;
originally announced October 2025.
-
BIRD-INTERACT: Re-imagining Text-to-SQL Evaluation for Large Language Models via Lens of Dynamic Interactions
Authors:
Nan Huo,
Xiaohan Xu,
Jinyang Li,
Per Jacobsson,
Shipei Lin,
Bowen Qin,
Binyuan Hui,
Xiaolong Li,
Ge Qu,
Shuzheng Si,
Linheng Han,
Edward Alexander,
Xintong Zhu,
Rui Qin,
Ruihan Yu,
Yiyao Jin,
Feige Zhou,
Weihao Zhong,
Yun Chen,
Hongyu Liu,
Chenhao Ma,
Fatma Ozcan,
Yannis Papakonstantinou,
Reynold Cheng
Abstract:
Large language models (LLMs) have demonstrated remarkable performance on single-turn text-to-SQL tasks, but real-world database applications predominantly require multi-turn interactions to handle ambiguous queries, execution errors, and evolving user requirements. Existing multi-turn benchmarks fall short by treating conversation histories as static context or limiting evaluation to read-only ope…
▽ More
Large language models (LLMs) have demonstrated remarkable performance on single-turn text-to-SQL tasks, but real-world database applications predominantly require multi-turn interactions to handle ambiguous queries, execution errors, and evolving user requirements. Existing multi-turn benchmarks fall short by treating conversation histories as static context or limiting evaluation to read-only operations, failing to reflect production-grade database assistant challenges. We introduce BIRD-INTERACT, a benchmark that restores this realism through: (1) a comprehensive interaction environment coupling each database with a hierarchical knowledge base, metadata files, and a function-driven user simulator, enabling models to solicit clarifications, retrieve knowledge, and recover from errors without human supervision; (2) two evaluation settings consisting of a pre-defined conversational protocol (c-Interact) and an open-ended agentic setting (a-Interact) where models autonomously decide when to query the user simulator or explore the environment; (3) a challenging task suite covering the full CRUD spectrum for business-intelligence and operational use cases, guarded by executable test cases. Each task features ambiguous and follow-up sub-tasks requiring dynamic interaction. The suite comprises BIRD-INTERACT-FULL (600 tasks, up to 11,796 interactions) for comprehensive performance assessment, and BIRD-INTERACT-LITE (300 tasks with simplified databases) for detailed behavioral analysis and rapid method development. Our empirical results highlight BIRD-INTERACT's difficulty: GPT-5 completes only 8.67% of tasks in c-Interact and 17.00% in a-Interact. Analysis via memory grafting and Interaction Test-time Scaling validates the importance of effective interaction for complex, dynamic text-to-SQL tasks.
△ Less
Submitted 24 March, 2026; v1 submitted 6 October, 2025;
originally announced October 2025.
-
Lattice Boltzmann Model for Learning Real-World Pixel Dynamicity
Authors:
Guangze Zheng,
Shijie Lin,
Haobo Zuo,
Si Si,
Ming-Shan Wang,
Changhong Fu,
Jia Pan
Abstract:
This work proposes the Lattice Boltzmann Model (LBM) to learn real-world pixel dynamicity for visual tracking. LBM decomposes visual representations into dynamic pixel lattices and solves pixel motion states through collision-streaming processes. Specifically, the high-dimensional distribution of the target pixels is acquired through a multilayer predict-update network to estimate the pixel positi…
▽ More
This work proposes the Lattice Boltzmann Model (LBM) to learn real-world pixel dynamicity for visual tracking. LBM decomposes visual representations into dynamic pixel lattices and solves pixel motion states through collision-streaming processes. Specifically, the high-dimensional distribution of the target pixels is acquired through a multilayer predict-update network to estimate the pixel positions and visibility. The predict stage formulates lattice collisions among the spatial neighborhood of target pixels and develops lattice streaming within the temporal visual context. The update stage rectifies the pixel distributions with online visual representations. Compared with existing methods, LBM demonstrates practical applicability in an online and real-time manner, which can efficiently adapt to real-world visual tracking tasks. Comprehensive evaluations of real-world point tracking benchmarks such as TAP-Vid and RoboTAP validate LBM's efficiency. A general evaluation of large-scale open-world object tracking benchmarks such as TAO, BFT, and OVT-B further demonstrates LBM's real-world practicality.
△ Less
Submitted 31 October, 2025; v1 submitted 20 September, 2025;
originally announced September 2025.
-
Mapping Innovation Networks: A Network-Based Approach to Actor Heterogeneity in National Innovation Systems
Authors:
Dawoon Jeong,
Taewon Kang,
Saerom Si,
Sangnam Lee,
Wonsub Eum
Abstract:
The Triple Helix model has provided a foundational framework for analyzing National Innovation Systems by highlighting the roles of universities, industries, and government research institutes. However, increasing heterogeneity within these actor groups limits the explanatory power of typological approaches. This study introduces a capability-based network methodology that maps the structural rela…
▽ More
The Triple Helix model has provided a foundational framework for analyzing National Innovation Systems by highlighting the roles of universities, industries, and government research institutes. However, increasing heterogeneity within these actor groups limits the explanatory power of typological approaches. This study introduces a capability-based network methodology that maps the structural relationships among innovation actors based on the similarity of their research and development (R&D) capabilities. Drawing on Economic Complexity Theory, we measure each actor's revealed comparative advantage (RCA) across scientific and technological fields and construct an R&D Actor Space - a proximity-based network that reflects the relational configuration of innovation capacities. Applying this method to Korean R&D data, we uncover a stratified system in which central, highly diversified universities coexist with more specialized firms and government institutes. Network analysis reveals assortative and unequal structures, and hierarchical clustering further highlights layered subgroupings. By moving beyond categorical classification, this capability-based network approach provides a scalable and generalizable tool for analyzing structural complexity within national innovation systems.
△ Less
Submitted 5 August, 2025;
originally announced August 2025.
-
MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
Authors:
Haozhe Zhao,
Zefan Cai,
Shuzheng Si,
Liang Chen,
Jiuxiang Gu,
Wen Xiao,
Minjia Zhang,
Junjie Hu
Abstract:
Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex multimodal image generation. To address these limitations, we propose MENTOR, a novel autoregressive (AR) framework for efficient Multimodal-conditioned Tuning for Autoregressive multimodal image generation. MENTOR combin…
▽ More
Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex multimodal image generation. To address these limitations, we propose MENTOR, a novel autoregressive (AR) framework for efficient Multimodal-conditioned Tuning for Autoregressive multimodal image generation. MENTOR combines an AR image generator with a two-stage training paradigm, enabling fine-grained, token-level alignment between multimodal inputs and image outputs without relying on auxiliary adapters or cross-attention modules. The two-stage training consists of: (1) a multimodal alignment stage that establishes robust pixel- and semantic-level alignment, followed by (2) a multimodal instruction tuning stage that balances the integration of multimodal inputs and enhances generation controllability. Despite modest model size, suboptimal base components, and limited training resources, MENTOR achieves strong performance on the DreamBench++ benchmark, outperforming competitive baselines in concept preservation and prompt following. Additionally, our method delivers superior image reconstruction fidelity, broad task adaptability, and improved training efficiency compared to diffusion-based methods. Dataset, code, and models are available at: https://github.com/HaozheZhao/MENTOR
△ Less
Submitted 28 May, 2026; v1 submitted 13 July, 2025;
originally announced July 2025.
-
Constraining self-interacting ultrahigh-energy muon neutrinos by cosmic microwave background spectral distortion
Authors:
Pravin Kumar Natwariya,
Shibsankar Si,
Alekha C. Nayak,
Tripurari Srivastava
Abstract:
The neutrino telescopes have firmly established the existence of ultrahigh-energy neutrinos. Observations of these neutrinos offer a unique probe of neutrino self-interactions. This work investigates how the self-interacting neutrinos, mediated by scalar bosons, inject energy into the medium through radiative scattering with the cosmic neutrino background, leaving an imprint on the cosmic microwav…
▽ More
The neutrino telescopes have firmly established the existence of ultrahigh-energy neutrinos. Observations of these neutrinos offer a unique probe of neutrino self-interactions. This work investigates how the self-interacting neutrinos, mediated by scalar bosons, inject energy into the medium through radiative scattering with the cosmic neutrino background, leaving an imprint on the cosmic microwave background (CMB) spectrum. The energy injection into plasma in redshift ranges, $5\times10^4\lesssim z\lesssim2\times10^6$ and $ z\lesssim5\times10^4$, leads to $μ$-type and $y$-type CMB spectral distortions, respectively. Using observational constraints from Cosmic Background Explorer/Far Infrared Absolute Spectrophotometer (COBE/FIRAS) and projected sensitivities from Primordial Inflation Explorer (PIXIE) experiments for $μ$-type and $y$-type CMB distortions, we derive the stringent upper bounds on the self-interaction coupling strength as a function of mediator mass for neutrino interactions. We focus on flavor-specific self-interaction related to muon neutrinos and sub-GeV mass mediators ($m_φ$). We find the upper bound on the self-interaction coupling strength to be $\sim 2.8\times 10^{-4}$ for the muon neutrino, considering ultrahigh-energy muon neutrino energy to be 1 PeV and PIXIE projected upper bounds on $y$-type CMB spectral distortion. The bound remains constant till the mediator mass reaches the center-of-mass energy, and after that, it gets relaxed and becomes proportional to the mediator mass. We have also compared our results with existing bounds in the literature. Our findings indicate that CMB spectral distortion could play a decisive role in exploring neutrino physics beyond the standard model of particle physics, and future missions like PIXIE can provide valuable insights.
△ Less
Submitted 17 June, 2026; v1 submitted 30 June, 2025;
originally announced June 2025.
-
Determination of the potential by a fixed angle scattering data
Authors:
Suliang Si
Abstract:
In this paper, we show that a compactly supported potential is uniquely determined by the far field pattern at a fixed angle. Our method is based on a new Carleman estimate and the ideas introduced by Bukhgeim and Klibanov on the use of Carleman estimates for inverse problems.
In this paper, we show that a compactly supported potential is uniquely determined by the far field pattern at a fixed angle. Our method is based on a new Carleman estimate and the ideas introduced by Bukhgeim and Klibanov on the use of Carleman estimates for inverse problems.
△ Less
Submitted 26 June, 2025;
originally announced June 2025.
-
SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications
Authors:
Jinyang Li,
Xiaolong Li,
Ge Qu,
Per Jacobsson,
Bowen Qin,
Binyuan Hui,
Shuzheng Si,
Nan Huo,
Xiaohan Xu,
Yue Zhang,
Ziwei Tang,
Yuanshuai Li,
Florensia Widjaja,
Xintong Zhu,
Feige Zhou,
Yongfeng Huang,
Yannis Papakonstantinou,
Fatma Ozcan,
Chenhao Ma,
Reynold Cheng
Abstract:
Resolution of complex SQL issues persists as a significant bottleneck in real-world database applications. Current Large Language Models (LLMs), while adept at text-to-SQL translation, have not been rigorously evaluated on the more challenging task of debugging SQL issues. To address this gap, we introduce BIRD-CRITIC, a new SQL issue debugging benchmark comprising 530 PostgreSQL tasks (BIRD-CRITI…
▽ More
Resolution of complex SQL issues persists as a significant bottleneck in real-world database applications. Current Large Language Models (LLMs), while adept at text-to-SQL translation, have not been rigorously evaluated on the more challenging task of debugging SQL issues. To address this gap, we introduce BIRD-CRITIC, a new SQL issue debugging benchmark comprising 530 PostgreSQL tasks (BIRD-CRITIC-PG) and 570 multi-dialect tasks (BIRD-CRITIC-Multi), distilled from authentic user issues and replayed within new environments to facilitate rigorous evaluation. Baseline evaluations underscore the task's complexity, with the leading reasoning model O3-Mini achieving only 38.87% success rate on BIRD-CRITIC-PG and 33.33% on BIRD-CRITIC-Multi. Meanwhile, advancing open-source models for database tasks is crucial for empowering local development while safeguarding data privacy. Therefore, we present Six-Gym (Sql-fIX-Gym), a training environment for elevating open-source model capabilities for SQL issue debugging. This environment leverages SQL-Rewind strategy, which automatically generates executable issue-solution datasets by reverse-engineering issues from verified SQLs. However, popular trajectory-based fine-tuning methods do not explore substantial supervisory signals. We further propose f-Plan Boosting, which extracts high-level debugging plans from SQL solutions, enabling teacher LLMs to produce 73.7% more successful trajectories for training. We integrate these components into an open-source agent, Bird-Fixer. Based on Qwen-2.5-Coder-14B, Bird-Fixer achieves 38.11% success rate on BIRD-CRITIC-PG and 29.65% on BIRD-CRITIC-Multi, surpassing leading proprietary models such as Claude-3.7-Sonnet and GPT-4.1, marking a significant step toward democratizing sophisticated SQL-debugging capabilities. The leaderboard and source code are available: https://bird-critic.github.io/
△ Less
Submitted 24 January, 2026; v1 submitted 23 June, 2025;
originally announced June 2025.
-
Inverse source problem for a hyperbolic equation by Carleman estimates
Authors:
Suliang Si
Abstract:
In this article, we provide a modified argument for proving the conditional stability of inverse source problem for a hyperbolic equation. Our method does not require any extension of solution with respect to time and therefore simplifies the existing proofs, which is widely applicable to various evolution equations.
In this article, we provide a modified argument for proving the conditional stability of inverse source problem for a hyperbolic equation. Our method does not require any extension of solution with respect to time and therefore simplifies the existing proofs, which is widely applicable to various evolution equations.
△ Less
Submitted 14 June, 2025;
originally announced June 2025.
-
Teaching Large Language Models to Maintain Contextual Faithfulness via Synthetic Tasks and Reinforcement Learning
Authors:
Shuzheng Si,
Haozhe Zhao,
Cheng Gao,
Yuzhuo Bai,
Zhitong Wang,
Bofei Gao,
Kangyang Luo,
Wenhao Li,
Yufei Huang,
Gang Chen,
Fanchao Qi,
Minjia Zhang,
Baobao Chang,
Maosong Sun
Abstract:
Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable information-seeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. Specifically, we first synthesize short-form question-answering (QA) data with four diverse tasks to…
▽ More
Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable information-seeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. Specifically, we first synthesize short-form question-answering (QA) data with four diverse tasks to construct high-quality and easily verifiable training data without human annotation. Also, we propose Dual-GRPO, a rule-based reinforcement learning method that includes three tailored rule-based rewards derived from synthesized short-form QA data, while simultaneously optimizing both short-form and long-form response generation. Notably, Dual-GRPO eliminates the need to manually label preference data to train reward models and avoids over-optimizing short-form generation when relying only on the synthesized short-form QA data. Experimental results show that CANOE greatly improves the faithfulness of LLMs across 11 different tasks, even outperforming the most advanced LLMs, e.g., GPT-4o and OpenAI o1.
△ Less
Submitted 11 November, 2025; v1 submitted 22 May, 2025;
originally announced May 2025.
-
Well-posedness and large time behavior of a size-structured growth-coagulation-fragmentation model
Authors:
Saroj Si,
Ankik Kumar Giri
Abstract:
The existence and uniqueness of weak solutions to a size-structured growth-coagulation-fragmentation (GCF) equation with a renewal boundary condition are shown for a class of unbounded coagulation and fragmentation kernels. The existence proof is based on a weak compactness framework in the weighted $L^1$-space. This result extends the existence results of Banasiak and Lamb [14] and Ackleh et al.…
▽ More
The existence and uniqueness of weak solutions to a size-structured growth-coagulation-fragmentation (GCF) equation with a renewal boundary condition are shown for a class of unbounded coagulation and fragmentation kernels. The existence proof is based on a weak compactness framework in the weighted $L^1$-space. This result extends the existence results of Banasiak and Lamb [14] and Ackleh et al. [2,4]. Furthermore, we establish a stability result and derive uniqueness as a direct consequence of it. Moreover, this study explores the large time behavior of weak solutions.
△ Less
Submitted 7 April, 2025;
originally announced April 2025.
-
Revisiting constraints on superconducting cosmic strings in light of Dark Ages global 21-cm signal
Authors:
Shibsankar Si,
Vivekanand Mohapatra,
Pravin kumar Natwariya,
Alekha Chandra Nayak
Abstract:
The Superconducting Cosmic Strings (SCS) are a special case of cosmic strings that have a core carrying a charged field. When SCS passes through magnetized regions, the charged particles in the string experience a Lorentz force, which can produce radiation on the entire electromagnetic spectrum. This radiation can inject energy into the surrounding plasma, resulting in a modification of the therma…
▽ More
The Superconducting Cosmic Strings (SCS) are a special case of cosmic strings that have a core carrying a charged field. When SCS passes through magnetized regions, the charged particles in the string experience a Lorentz force, which can produce radiation on the entire electromagnetic spectrum. This radiation can inject energy into the surrounding plasma, resulting in a modification of the thermal and ionization evolution of the intergalactic medium (IGM) and, subsequently, the global 21-cm signal. The signatures of SCS in the post-recombination era have been primarily studied in the low-frequency (radio) regime, which does not impact the state of the IGM. In this work, we study the effect of decaying SCS on the dark ages global 21-cm signal $(δT_b)$, considering both the ionizing and radio radiation. The dark ages signal can provide pristine cosmological information free from astrophysical uncertainties, as the universe was primarily homogeneous during this era in the absence of baryonic structure formation. Considering a change in the $δT_b$ at redshift $z\sim 89$ from the $Λ\rm CDM$ framework, we derive an upper bound on the decay efficiency parameter, $g\equiv g(I,~Gμ_s)$, to be $\lesssim 5.1\times10^{14}\, \rm GeV^2$, where, $I$ and $Gμ_s$ represent the loop current and string tension of SCS, respectively.
△ Less
Submitted 21 October, 2025; v1 submitted 3 April, 2025;
originally announced April 2025.
-
Think Before Refusal : Triggering Safety Reflection in LLMs to Mitigate False Refusal Behavior
Authors:
Shengyun Si,
Xinpeng Wang,
Guangyao Zhai,
Nassir Navab,
Barbara Plank
Abstract:
Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in f…
▽ More
Recent advancements in large language models (LLMs) have demonstrated that fine-tuning and human alignment can render LLMs harmless. In practice, such "harmlessness" behavior is mainly achieved by training models to reject harmful requests, such as "Explain how to burn down my neighbor's house", where the model appropriately declines to respond. However, this approach can inadvertently result in false refusal, where models reject benign queries as well, such as "Tell me how to kill a Python process". In this work, we demonstrate that prompting safety reflection before generating a response can mitigate false refusal behavior. Building on this finding, we introduce the Think-Before-Refusal (TBR) schema and conduct safety-aware instruction fine-tuning incorporating safety reflection. In an ablation study across 15 pre-trained models, we show that models fine-tuned with safety reflection significantly reduce false refusal behavior while maintaining safety and overall performance compared to those fine-tuned without safety reflection.
△ Less
Submitted 22 March, 2025;
originally announced March 2025.
-
GLTW: Joint Improved Graph Transformer and LLM via Three-Word Language for Knowledge Graph Completion
Authors:
Kangyang Luo,
Yuzhuo Bai,
Cheng Gao,
Shuzheng Si,
Yingli Shen,
Zhu Liu,
Zhitong Wang,
Cunliang Kong,
Wenhao Li,
Yufei Huang,
Ye Tian,
Xuantang Xiong,
Lei Han,
Maosong Sun
Abstract:
Knowledge Graph Completion (KGC), which aims to infer missing or incomplete facts, is a crucial task for KGs. However, integrating the vital structural information of KGs into Large Language Models (LLMs) and outputting predictions deterministically remains challenging. To address this, we propose a new method called GLTW, which encodes the structural information of KGs and merges it with LLMs to…
▽ More
Knowledge Graph Completion (KGC), which aims to infer missing or incomplete facts, is a crucial task for KGs. However, integrating the vital structural information of KGs into Large Language Models (LLMs) and outputting predictions deterministically remains challenging. To address this, we propose a new method called GLTW, which encodes the structural information of KGs and merges it with LLMs to enhance KGC performance. Specifically, we introduce an improved Graph Transformer (iGT) that effectively encodes subgraphs with both local and global structural information and inherits the characteristics of language model, bypassing training from scratch. Also, we develop a subgraph-based multi-classification training objective, using all entities within KG as classification objects, to boost learning efficiency.Importantly, we combine iGT with an LLM that takes KG language prompts as input.Our extensive experiments on various KG datasets show that GLTW achieves significant performance gains compared to SOTA baselines.
△ Less
Submitted 30 May, 2025; v1 submitted 17 February, 2025;
originally announced February 2025.
-
Aligning Large Language Models to Follow Instructions and Hallucinate Less via Effective Data Filtering
Authors:
Shuzheng Si,
Haozhe Zhao,
Gang Chen,
Cheng Gao,
Yuzhuo Bai,
Zhitong Wang,
Kaikai An,
Kangyang Luo,
Chen Qian,
Fanchao Qi,
Baobao Chang,
Maosong Sun
Abstract:
Training LLMs on data containing unfamiliar knowledge during the instruction tuning stage can encourage hallucinations. To address this challenge, we introduce NOVA, a novel framework designed to identify high-quality data that aligns well with the LLM's learned knowledge to reduce hallucinations. NOVA includes Internal Consistency Probing (ICP) and Semantic Equivalence Identification (SEI) to mea…
▽ More
Training LLMs on data containing unfamiliar knowledge during the instruction tuning stage can encourage hallucinations. To address this challenge, we introduce NOVA, a novel framework designed to identify high-quality data that aligns well with the LLM's learned knowledge to reduce hallucinations. NOVA includes Internal Consistency Probing (ICP) and Semantic Equivalence Identification (SEI) to measure how familiar the LLM is with instruction data. Specifically, ICP evaluates the LLM's understanding of the given instruction by calculating the tailored consistency among multiple self-generated responses. SEI further assesses the familiarity of the LLM with the target response by comparing it to the generated responses, using the proposed semantic clustering and well-designed voting strategy. Finally, to ensure the quality of selected samples, we introduce an expert-aligned reward model, considering characteristics beyond just familiarity. By considering data quality and avoiding unfamiliar data, we can utilize the selected data to effectively align LLMs to follow instructions and hallucinate less.
△ Less
Submitted 25 May, 2025; v1 submitted 11 February, 2025;
originally announced February 2025.
-
First experimental proof of PET imaging based on multi-anode MCP-PMTs with Cherenkov radiator-integrated window
Authors:
Weiyan Pan,
Lingyue Chen,
Guorui Huang,
Jun Hu,
Wei Hou,
Xianchao Huang,
Xiaorou Han,
Xiaoshan Jiang,
Zhen Jin,
Daowu Li,
Jingwen Li,
Shulin Liu,
Zehong Liang,
Lishuang Ma,
Zhe Ning,
Sen Qian,
Ling Ren,
Jianning Sun,
Shuguang Si,
Yunhua Sun,
Long Wei,
Ning Wang,
Qing Wei,
Qi Wu,
Tianyi Wang
, et al. (11 additional authors not shown)
Abstract:
Improving the coincidence time resolution (CTR) of time-of-flight positron emission tomography (TOF-PET) systems to achieve a higher signal-to-noise ratio (SNR) gain or even direct positron emission imaging (dPEI) is of paramount importance for many advanced new clinical applications of PET imaging. This places higher demands on the timing performance of all aspects of PET systems. One effective a…
▽ More
Improving the coincidence time resolution (CTR) of time-of-flight positron emission tomography (TOF-PET) systems to achieve a higher signal-to-noise ratio (SNR) gain or even direct positron emission imaging (dPEI) is of paramount importance for many advanced new clinical applications of PET imaging. This places higher demands on the timing performance of all aspects of PET systems. One effective approach is to use microchannel plate photomultiplier tubes (MCP-PMTs) for prompt Cherenkov photon detection. In this study, we developed a dual-module Cherenkov PET imaging experimental platform, utilising our proprietary 8 * 8-anode Cherenkov radiator-integrated window MCP-PMTs in combination with custom-designed multi-channel electronics, and designed a specific calibration and correction method for the platform. Using this platform, a CTR of 103 ps FWHM was achieved. We overcame the limitations of single-anode detectors in previous experiments, significantly enhanced imaging efficiency and achieved module-level Cherenkov PET imaging for the first time. Imaging experiments involving radioactive sources and phantoms of various shapes and types were conducted, which preliminarily validated the feasibility and advancement of this imaging method. In addition, the effects of normalisation correction and the interaction probability between the gamma rays and the MCP on the images and experimental results were analysed and verified.
△ Less
Submitted 14 October, 2025; v1 submitted 10 February, 2025;
originally announced February 2025.
-
UltraIF: Advancing Instruction Following from the Wild
Authors:
Kaikai An,
Li Sheng,
Ganqu Cui,
Shuzheng Si,
Ning Ding,
Yu Cheng,
Baobao Chang
Abstract:
Instruction-following made modern large language models (LLMs) helpful assistants. However, the key to taming LLMs on complex instructions remains mysterious, for that there are huge gaps between models trained by open-source community and those trained by leading companies. To bridge the gap, we propose a simple and scalable approach UltraIF for building LLMs that can follow complex instructions…
▽ More
Instruction-following made modern large language models (LLMs) helpful assistants. However, the key to taming LLMs on complex instructions remains mysterious, for that there are huge gaps between models trained by open-source community and those trained by leading companies. To bridge the gap, we propose a simple and scalable approach UltraIF for building LLMs that can follow complex instructions with open-source data. UltraIF first decomposes real-world user prompts into simpler queries, constraints, and corresponding evaluation questions for the constraints. Then, we train an UltraComposer to compose constraint-associated prompts with evaluation questions. This prompt composer allows us to synthesize complicated instructions as well as filter responses with evaluation questions. In our experiment, for the first time, we successfully align LLaMA-3.1-8B-Base to catch up with its instruct version on 5 instruction-following benchmarks without any benchmark information, using only 8B model as response generator and evaluator. The aligned model also achieved competitive scores on other benchmarks. Moreover, we also show that UltraIF could further improve LLaMA-3.1-8B-Instruct through self-alignment, motivating broader use cases for the method. Our code is available at https://github.com/kkk-an/UltraIF.
△ Less
Submitted 28 September, 2025; v1 submitted 6 February, 2025;
originally announced February 2025.
-
Revisiting the matrix elements of the position operator in the crystal momentum representation
Authors:
M. S. Si,
G. P. Zhang
Abstract:
Fewer operators are more fundamental than the position operator in a crystal. But since it is not translationally invariant in crystal momentum representation (CMR), how to properly represent it is nontrivial. Over half a century, various methods have been proposed, but they often lead to either highly singular derivatives or extremely arcane expressions. Here we propose a resolution to this probl…
▽ More
Fewer operators are more fundamental than the position operator in a crystal. But since it is not translationally invariant in crystal momentum representation (CMR), how to properly represent it is nontrivial. Over half a century, various methods have been proposed, but they often lead to either highly singular derivatives or extremely arcane expressions. Here we propose a resolution to this problem by directly computing their matrix elements between two Bloch states. We show that the position operator is a full matrix in CMR, where the off-diagonal elements in crystal momentum $\bf k$ only appear along the direction of the position vector. Our formalism, free of singular derivative and degeneracy difficulties, can describe an array of physical properties, from intraband transitions, polarization with or without spin-orbit coupling, orbital angular momentum, to susceptibilities.
△ Less
Submitted 3 January, 2025;
originally announced January 2025.
-
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance
Authors:
Haozhe Zhao,
Shuzheng Si,
Liang Chen,
Yichi Zhang,
Maosong Sun,
Mingjia Zhang,
Baobao Chang
Abstract:
Large vision-language models (LVLMs) have achieved impressive results in various vision-language tasks. However, despite showing promising performance, LVLMs suffer from hallucinations caused by language bias, leading to diminished focus on images and ineffective visual comprehension. We identify two primary reasons for this bias: 1. Different scales of training data between the pretraining stage…
▽ More
Large vision-language models (LVLMs) have achieved impressive results in various vision-language tasks. However, despite showing promising performance, LVLMs suffer from hallucinations caused by language bias, leading to diminished focus on images and ineffective visual comprehension. We identify two primary reasons for this bias: 1. Different scales of training data between the pretraining stage of LLM and multimodal alignment stage. 2. The learned inference bias due to short-term dependency of text data. Therefore, we propose LACING, a systemic framework designed to address the language bias of LVLMs with muLtimodal duAl-attention meChanIsm (MDA) aNd soft-image Guidance (IFG). Specifically, MDA introduces a parallel dual-attention mechanism that enhances the integration of visual inputs across the model. IFG introduces a learnable soft visual prompt during training and inference to replace visual inputs, designed to compel LVLMs to prioritize text inputs. Then, IFG further proposes a novel decoding strategy using the soft visual prompt to mitigate the model's over-reliance on adjacent text inputs. Comprehensive experiments demonstrate that our method effectively debiases LVLMs from their language bias, enhancing visual comprehension and reducing hallucinations without requiring additional training resources or data. The code and model are available at [lacing-lvlm.github.io](https://lacing-lvlm.github.io).
△ Less
Submitted 28 May, 2026; v1 submitted 21 November, 2024;
originally announced November 2024.
-
Existence and Non-existence for Exchange-Driven Growth Model
Authors:
Saroj Si,
Ankik Kumar Giri
Abstract:
The exchange-driven growth (EDG) model describes the evolution of clusters through the exchange of single monomers between pairs of interacting clusters. The dynamics of this process are primarily influenced by the interaction kernel $K_{j,k}$. In this paper, the global existence of classical solutions to the EDG equations is established for non-negative, symmetric interaction kernels satisfying…
▽ More
The exchange-driven growth (EDG) model describes the evolution of clusters through the exchange of single monomers between pairs of interacting clusters. The dynamics of this process are primarily influenced by the interaction kernel $K_{j,k}$. In this paper, the global existence of classical solutions to the EDG equations is established for non-negative, symmetric interaction kernels satisfying $K_{j,k} \leq C(j^μk^ν + j^νk^μ) $, where $μ, ν\leq 2$, $μ+ ν\leq 3$, and $C>0$, with a broader class of initial data. This result extends the previous existence results obtained by Esenturk [10], Schlichting [23], and Eichenberg \& Schlichting [7].
Furthermore, the local existence of classical solutions to the EDG equations is demonstrated for symmetric interaction kernels that satisfy $K_{j,k} \leq C j^{2} k^{2}$ with $C > 0$, considering a broader class of initial data. In the intermediate regime $3 < μ+ ν\leq 4$, the occurrence of finite-time gelation is established for symmetric interaction kernels satisfying $C_{1}\left(j^{2}k^α+j^αk^{2}\right)\leq K_{j,k}\leq Cj^{2}k^{2}$, where $1 < α\leq 2$, $C>0$, and $C_{1} > 0$, as conjectured in [10]. In this case, the non-existence of the global solutions is ensured by the occurrence of finite-time gelation. Finally, the occurrence of instantaneous gelation of the solutions to EDG equations for symmetric interaction kernels satisfying $K_{j,k}\geq C\left(j^β+k^β\right)$ ($β>2, C>0)$ is shown, which also implies the non-existence of solutions in this case.
△ Less
Submitted 21 November, 2024;
originally announced November 2024.
-
Increasing stability for inverse acoustic source problems
Authors:
Suliang Si
Abstract:
In this paper, we show the increasing stability of the inverse source problems for the acoustic wave equation in the full space R3.The goal is to understand increasing stability for wave equation in the time domain. If the time and spatial variables of the source term can be separated with compact support, the increasing stability estimates of the $L^2$-norm of the acoustic source function can be…
▽ More
In this paper, we show the increasing stability of the inverse source problems for the acoustic wave equation in the full space R3.The goal is to understand increasing stability for wave equation in the time domain. If the time and spatial variables of the source term can be separated with compact support, the increasing stability estimates of the $L^2$-norm of the acoustic source function can be established. The stability estimates consist of two parts: the Lipschitz type data discrepancy and the high time tail of the source functions. As the time increases, the latter decreases and thus becomes negligible.
△ Less
Submitted 7 November, 2024;
originally announced November 2024.
-
LoRA Done RITE: Robust Invariant Transformation Equilibration for LoRA Optimization
Authors:
Jui-Nan Yen,
Si Si,
Zhao Meng,
Felix Yu,
Sai Surya Duvvuri,
Inderjit S. Dhillon,
Cho-Jui Hsieh,
Sanjiv Kumar
Abstract:
Low-rank adaption (LoRA) is a widely used parameter-efficient finetuning method for LLM that reduces memory requirements. However, current LoRA optimizers lack transformation invariance, meaning the actual updates to the weights depends on how the two LoRA factors are scaled or rotated. This deficiency leads to inefficient learning and sub-optimal solutions in practice. This paper introduces LoRA-…
▽ More
Low-rank adaption (LoRA) is a widely used parameter-efficient finetuning method for LLM that reduces memory requirements. However, current LoRA optimizers lack transformation invariance, meaning the actual updates to the weights depends on how the two LoRA factors are scaled or rotated. This deficiency leads to inefficient learning and sub-optimal solutions in practice. This paper introduces LoRA-RITE, a novel adaptive matrix preconditioning method for LoRA optimization, which can achieve transformation invariance and remain computationally efficient. We provide theoretical analysis to demonstrate the benefit of our method and conduct experiments on various LLM tasks with different models including Gemma 2B, 7B, and mT5-XXL. The results demonstrate consistent improvements against existing optimizers. For example, replacing Adam with LoRA-RITE during LoRA fine-tuning of Gemma-2B yielded 4.6\% accuracy gain on Super-Natural Instructions and 3.5\% accuracy gain across other four LLM benchmarks (HellaSwag, ArcChallenge, GSM8K, OpenBookQA).
△ Less
Submitted 16 July, 2025; v1 submitted 27 October, 2024;
originally announced October 2024.
-
GATEAU: Selecting Influential Samples for Long Context Alignment
Authors:
Shuzheng Si,
Haozhe Zhao,
Gang Chen,
Yunshui Li,
Kangyang Luo,
Chuancheng Lv,
Kaikai An,
Fanchao Qi,
Baobao Chang,
Maosong Sun
Abstract:
Aligning large language models to handle instructions with extremely long contexts has yet to be fully investigated. Previous studies have attempted to scale up the available data volume by synthesizing long instruction-following samples, as constructing such a dataset tends to be challenging for annotators. However, a lack of a well-defined strategy for ensuring data quality may introduce low-qua…
▽ More
Aligning large language models to handle instructions with extremely long contexts has yet to be fully investigated. Previous studies have attempted to scale up the available data volume by synthesizing long instruction-following samples, as constructing such a dataset tends to be challenging for annotators. However, a lack of a well-defined strategy for ensuring data quality may introduce low-quality samples and restrict the model's performance. Thus, we propose GATEAU, a novel framework to address the unique challenge of long context alignment by identifying the influential samples enriched with long-range dependency relations. Specifically, GATEAU measures the long-range dependencies from two essential aspects: the difficulty of generating target responses due to the long-range dependencies, and the difficulty of understanding long inputs due to such dependencies. Comprehensive experiments indicate that GATEAU effectively identifies influential samples, and the model trained on these selected samples exhibits better instruction-following and long-context understanding capabilities.
△ Less
Submitted 15 September, 2025; v1 submitted 21 October, 2024;
originally announced October 2024.
-
Autonomous Driving in Unstructured Environments: How Far Have We Come?
Authors:
Chen Min,
Shubin Si,
Xu Wang,
Hanzhang Xue,
Weizhong Jiang,
Zitong Chen,
Mengmeng Li,
Jilin Mei,
Erke Shang,
Zhipeng Xiao,
Bin Dai,
Qi Zhu,
Hao Fu,
Dawei Zhao,
Liang Xiao,
Yiming Nie,
Yu Hu
Abstract:
Research on autonomous driving in unstructured outdoor environments is less advanced than in structured urban settings due to challenges like environmental diversities and scene complexity. These environments-such as rural areas and rugged terrains-pose unique obstacles that are not common in structured urban areas. Despite these difficulties, autonomous driving in unstructured outdoor environment…
▽ More
Research on autonomous driving in unstructured outdoor environments is less advanced than in structured urban settings due to challenges like environmental diversities and scene complexity. These environments-such as rural areas and rugged terrains-pose unique obstacles that are not common in structured urban areas. Despite these difficulties, autonomous driving in unstructured outdoor environments is crucial for applications in agriculture, mining, and military operations. Our survey reviews over 250 papers for autonomous driving in unstructured outdoor environments, covering offline mapping, pose estimation, environmental perception, path planning, end-to-end autonomous driving, datasets, and relevant challenges. We also discuss emerging trends and future research directions. This review aims to consolidate knowledge and encourage further research for autonomous driving in unstructured environments. To support ongoing work, we maintain an active repository with up-to-date literature and open-source projects at: https://github.com/chaytonmin/Survey-Autonomous-Driving-in-Unstructured-Environments.
△ Less
Submitted 12 January, 2026; v1 submitted 10 October, 2024;
originally announced October 2024.
-
Rethinking Semantic Parsing for Large Language Models: Enhancing LLM Performance with Semantic Hints
Authors:
Kaikai An,
Shuzheng Si,
Helan Hu,
Haozhe Zhao,
Yuchi Wang,
Qingyan Guo,
Baobao Chang
Abstract:
Semantic Parsing aims to capture the meaning of a sentence and convert it into a logical, structured form. Previous studies show that semantic parsing enhances the performance of smaller models (e.g., BERT) on downstream tasks. However, it remains unclear whether the improvements extend similarly to LLMs. In this paper, our empirical findings reveal that, unlike smaller models, directly adding sem…
▽ More
Semantic Parsing aims to capture the meaning of a sentence and convert it into a logical, structured form. Previous studies show that semantic parsing enhances the performance of smaller models (e.g., BERT) on downstream tasks. However, it remains unclear whether the improvements extend similarly to LLMs. In this paper, our empirical findings reveal that, unlike smaller models, directly adding semantic parsing results into LLMs reduces their performance. To overcome this, we propose SENSE, a novel prompting approach that embeds semantic hints within the prompt. Experiments show that SENSE consistently improves LLMs' performance across various tasks, highlighting the potential of integrating semantic information to improve LLM capabilities.
△ Less
Submitted 27 May, 2025; v1 submitted 22 September, 2024;
originally announced September 2024.
-
Well-posedness of the growth-coagulation equation with singular kernels
Authors:
Ankik Kumar Giri,
Philippe Laurençot,
Saroj Si
Abstract:
The well-posedness of the growth-coagulation equation is established for coagulation kernels having singularity near the origin and growing atmost linearly at infinity. The existence of weak solutions is shown by means of the method of the characteristics and a weak $L_1$-compactness argument. For the existence result, we also show our gratitude to Banach fixed point theorem and a refined version…
▽ More
The well-posedness of the growth-coagulation equation is established for coagulation kernels having singularity near the origin and growing atmost linearly at infinity. The existence of weak solutions is shown by means of the method of the characteristics and a weak $L_1$-compactness argument. For the existence result, we also show our gratitude to Banach fixed point theorem and a refined version of the Arzelá-Ascoli theorem. In addition, the continuous dependence of solutions upon the initial data is shown with the help of the DiPerna-Lions theory, Gronwall's inequality and moment estimates. Moreover, the uniqueness of solution follows from the continuous dependence. The results presented in this article extend the contributions made in earlier literature.
△ Less
Submitted 5 August, 2024;
originally announced August 2024.
-
UltraEdit: Instruction-based Fine-Grained Image Editing at Scale
Authors:
Haozhe Zhao,
Xiaojian Ma,
Liang Chen,
Shuzheng Si,
Rujie Wu,
Kaikai An,
Peiyu Yu,
Minjia Zhang,
Qing Li,
Baobao Chang
Abstract:
This paper presents UltraEdit, a large-scale (approximately 4 million editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image editing datasets like InstructPix2Pix and MagicBrush, and provide a systematic approach to producing massive and high-quality image editing samples. UltraEdit offers several distinct a…
▽ More
This paper presents UltraEdit, a large-scale (approximately 4 million editing samples), automatically generated dataset for instruction-based image editing. Our key idea is to address the drawbacks in existing image editing datasets like InstructPix2Pix and MagicBrush, and provide a systematic approach to producing massive and high-quality image editing samples. UltraEdit offers several distinct advantages: 1) It features a broader range of editing instructions by leveraging the creativity of large language models (LLMs) alongside in-context editing examples from human raters; 2) Its data sources are based on real images, including photographs and artworks, which provide greater diversity and reduced bias compared to datasets solely generated by text-to-image models; 3) It also supports region-based editing, enhanced by high-quality, automatically produced region annotations. Our experiments show that canonical diffusion-based editing baselines trained on UltraEdit set new records on MagicBrush and Emu-Edit benchmarks. Our analysis further confirms the crucial role of real image anchors and region-based editing data. The dataset, code, and models can be found in https://ultra-editing.github.io.
△ Less
Submitted 18 December, 2024; v1 submitted 7 July, 2024;
originally announced July 2024.
-
FFN: a Fine-grained Chinese-English Financial Domain Parallel Corpus
Authors:
Yuxin Fu,
Shijing Si,
Leyi Mai,
Xi-ang Li
Abstract:
Large Language Models (LLMs) have stunningly advanced the field of machine translation, though their effectiveness within the financial domain remains largely underexplored. To probe this issue, we constructed a fine-grained Chinese-English parallel corpus of financial news called FFN. We acquired financial news articles spanning between January 1st, 2014, to December 31, 2023, from mainstream med…
▽ More
Large Language Models (LLMs) have stunningly advanced the field of machine translation, though their effectiveness within the financial domain remains largely underexplored. To probe this issue, we constructed a fine-grained Chinese-English parallel corpus of financial news called FFN. We acquired financial news articles spanning between January 1st, 2014, to December 31, 2023, from mainstream media websites such as CNN, FOX, and China Daily. The dataset consists of 1,013 main text and 809 titles, all of which have been manually corrected. We measured the translation quality of two LLMs -- ChatGPT and ERNIE-bot, utilizing BLEU, TER and chrF scores as the evaluation metrics. For comparison, we also trained an OpenNMT model based on our dataset. We detail problems of LLMs and provide in-depth analysis, intending to stimulate further research and solutions in this largely uncharted territory. Our research underlines the need to optimize LLMs within the specific field of financial translation to ensure accuracy and quality.
△ Less
Submitted 26 June, 2024;
originally announced June 2024.