-
ReForce: Learning Force-aware Retargeting for Dexterous Manipulation
Authors:
Yuhang Wu,
Lingqi Zeng,
Changwei Jing,
Jianglong Ye,
Xiaolong Wang
Abstract:
Human demonstrations offer a scalable data source for dexterous manipulation, but transferring them to robot actions remains challenging due to the embodiment gap. Today's retargeting is mostly kinematic, yet manipulation is decided by force, which governs how the hand interacts with the object and how the object moves. In this paper, we present ReForce, a Force-aware Retargeting method that turns…
▽ More
Human demonstrations offer a scalable data source for dexterous manipulation, but transferring them to robot actions remains challenging due to the embodiment gap. Today's retargeting is mostly kinematic, yet manipulation is decided by force, which governs how the hand interacts with the object and how the object moves. In this paper, we present ReForce, a Force-aware Retargeting method that turns human motion and forces into robot actions that reproduce the intended contact. ReForce predicts a residual on the kinematically retargeted action to reach the desired force, using a general force tracker trained on large-scale simulation interactions. It supports both online force-aware teleoperation and offline data translation. In simulation and on real hardware, ReForce achieves lower force-tracking error and stronger multi-finger contact engagement on contact-rich tasks such as paper-cup grasping and tongs manipulation.
△ Less
Submitted 16 August, 2026;
originally announced August 2026.
-
MobileMem: Learning from a Year of Mobile Experiences
Authors:
Xinle Deng,
Yida Xue,
Xiangyuan Ru,
Yijun Chen,
Buqiang Xu,
Mingjun Mao,
Xinjie Liu,
Haoming Xu,
Shuofei Qiao,
Mengru Wang,
Chen Jiang,
Yuchen Eleanor Jiang,
Lizhong Wang,
Jason Wang,
Li Zeng,
Haofen Wang,
Guilin Qi,
Huajun Chen,
Ningyu Zhang
Abstract:
The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. Such assistants require long-term memory to accumulate and leverage user-specific experiences over time, yet existing benchmarks remain inadequate for realistic mobile settings, whe…
▽ More
The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. Such assistants require long-term memory to accumulate and leverage user-specific experiences over time, yet existing benchmarks remain inadequate for realistic mobile settings, where experiences are heterogeneous, multimodal, evolving, and deeply personal. We introduce MobileMem, a benchmark and framework for studying on-device long-term memory, grounded in a year-scale collection of mobile experiences. MobileMem employs a knowledge-grounded synthesis pipeline to construct coherent and temporally consistent long-horizon trajectories from user-app sessions. It provides complementary text and multimodal settings covering multi-hop and temporal reasoning, knowledge updating, and implicit preference inference. Specifically, MobileMem enables agents to remember the past, understand the present, and adapt to the future. By modeling experiences rather than isolated facts, MobileMem moves memory beyond information retrieval toward experiential intelligence for continuous personal learning.
△ Less
Submitted 17 August, 2026; v1 submitted 11 August, 2026;
originally announced August 2026.
-
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
Authors:
Kaichao Liang,
Yuqi Cui,
Hao Kong,
Xinyuan Huang,
Guohaotian Hou,
Qingcan Kang,
Liang Chen,
Yiyang Yin,
Ke Ye,
Jiaquan Guo,
Da Chen,
Lingan Zeng,
Yixing Peng,
Rong Yao,
Shixiong Kai,
Mingxuan Yuan
Abstract:
Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt their memory models, organization strategies, and procedural knowledge through continued use. We present MindMemOS, a portable and self-evolving memory…
▽ More
Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interactions. However, existing memory systems often remain fixed after development, limiting their ability to adapt their memory models, organization strategies, and procedural knowledge through continued use. We present MindMemOS, a portable and self-evolving memory operating layer that organizes open-world information using a unified entity property timestructure. MindMemOS supports scenario-adaptive memory modeling, higher-order pattern discovery, autonomous memory refinement, and continuous skill evolution. Its MindMemEvolve algorithm employs validation-driven evolutionary search to optimize memory schemas for target scenarios, whiledreaming consolidates accumulated memories by merging redundant records and resolving conflicts. In addition, implicit corrective feedback serves as a human-in-the-loop signal for identifying and revising potentially inaccurate or misaligned memories. Its MindSkillEvolve algorithm further transforms agent execution trajectories into reusable and progressively refined skills. MindMemOS achieves 94.03% accuracy on LOCOMO and 70.63% on PersonaMem. MindSkillEvolve improves SpreadsheetBench success by 9.2 percentage points over the initial-skill baseline.
△ Less
Submitted 12 August, 2026;
originally announced August 2026.
-
Infrared Spectroscopy and Photochemistry of Aromatic Nitriles in Para-Hydrogen Matrices
Authors:
Sam McGrath,
Vincent J. Esposito,
Linshan Zeng,
Thomas H. Speak,
Brendan Moore,
Pavle Djuricanin,
Jun Miyazaki,
Takamasa Momose,
Ilsa R. Cooke
Abstract:
Motivated by recent detections of several aromatic nitriles in Taurus Molecular Cloud-1, we report laboratory and theoretical investigations of the vibrational spectroscopy and photochemistry of singly and doubly cyano-substituted benzene in solid para-hydrogen matrices. We compare the photochemistry of cyanobenzene (benzonitrile) and three dicyanobenzene isomers initiated by excitations at 193 nm…
▽ More
Motivated by recent detections of several aromatic nitriles in Taurus Molecular Cloud-1, we report laboratory and theoretical investigations of the vibrational spectroscopy and photochemistry of singly and doubly cyano-substituted benzene in solid para-hydrogen matrices. We compare the photochemistry of cyanobenzene (benzonitrile) and three dicyanobenzene isomers initiated by excitations at 193 nm. In addition, we report the photochemistry of deuterated cyanobenzene (d$_5$-cyanobenzene), enabling us to determine the major products produced during the cyanobenzene photodissociation. The major products observed in the photolysis of all the nitriles are HCN and HNC, which are likely produced by hydrogen abstraction from para-H$_2$ by the CN radical. This indicates that the major photodissociation channel involves cleavage of the bond between the ring and the nitrile group, forming the phenyl (or cyanophenyl) radical + CN. We observe secondary photoproducts similar to those found during benzene photolysis. Our findings may aid the interpretation of recent JWST mid-infrared observations of aromatics in photodissociation regions.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
SR-OPSD: Self-Referenced On-Policy Self-Distillation
Authors:
Zhuo Sun,
Entong Li,
Yanlong Zhao,
Xiaoyuan Cheng,
Wenxuan Yuan,
Kaiyu Li,
Che Liu,
Huihang Liu,
Harrison Bo Hua Zhu,
Li Zeng
Abstract:
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome rewards. However, the self-teacher policy used in OPSD is typically a stop-gradient or exponential-moving-average copy of the policy conditioned on additional context information,…
▽ More
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optimized, providing a useful complement to reinforcement learning with sparse outcome rewards. However, the self-teacher policy used in OPSD is typically a stop-gradient or exponential-moving-average copy of the policy conditioned on additional context information, and thus co-evolves with both the student policy and its on-policy context distribution. Directly matching such a moving target with a fixed projection objective can lead to unstable optimization or excessive distributional concentration. This nature of OPSD motivates the proposed \emph{Self-Referenced On-Policy Self-Distillation (SR-OPSD)}. At fixed student-generated contexts, a token-level variational characterization identifies the effective distillation target as a geometric interpolation between the self-teacher policy and a reference policy. Meanwhile, we use the Rényi divergence family to generalize the projection geometry. This formulation separates \emph{where} the adaptive target is placed from \emph{how} the student is projected toward it: the interpolation coefficient controls underlying target, while the Rényi order controls the projection geometry and its sensitivity to token-level density ratios. Extensive experiments across scientific evaluation, mathematical reasoning, and coding generation tasks with multiple large language models show that SR-OPSD achieves the state-of-the-art or competitive performance across various settings.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
Spin-Orbital Hall Nano-Oscillators using PtCr/NiFe
Authors:
Utkarsh Shashank,
Akash Kumar,
Daegeun Jo,
Thi Ngoc Anh Nguyen,
Jong-Guk Choi,
Sambit Ghosh,
Michal Strach,
Lunjie Zeng,
Andrew B. Yankovich,
Roman Khymyn,
Ahmad A. Awad,
Eva Olsson,
Peter M. Oppeneer,
Johan Åkerman
Abstract:
The orbital Hall effect provides a promising route for generating angular-momentum currents beyond conventional spin Hall physics. PtCr alloys exhibit unusually large current-induced torques, but the contribution of orbital transport and the ability of these torques to sustain coherent nonlinear magnetization dynamics remain unresolved. Here we demonstrate spin-orbital Hall nano-oscillators by exp…
▽ More
The orbital Hall effect provides a promising route for generating angular-momentum currents beyond conventional spin Hall physics. PtCr alloys exhibit unusually large current-induced torques, but the contribution of orbital transport and the ability of these torques to sustain coherent nonlinear magnetization dynamics remain unresolved. Here we demonstrate spin-orbital Hall nano-oscillators by exploiting a homogeneous heavy-metal/light-metal alloy in which orbital Hall currents generated by Cr are converted by Pt into spin currents, producing giant spin-orbit torques. Using PtCr/NiFe heterostructures, the effective torque efficiency increases from ~0.14 in Pt/NiFe to ~0.40 in Pt0.38Cr0.62/NiFe despite substantial Pt dilution, enabling coherent auto-oscillations with the threshold current density reduced from ~ 1.07 x 10^12 to ~ 4.4 x 10^11 A m^-2. First-principles calculations show that Cr alloying suppresses the intrinsic spin Hall conductivity while enhancing the orbital Hall conductivity, and reproduce the observed torque enhancement only when orbital transport is included. Our combined experimental and first-principles results show that alloy engineering enables giant spin-orbit torques through an intrinsic orbital-mediated contribution, enabling coherent auto-oscillations without engineered multilayers and establishing a scalable materials platform for low-power nonlinear spintronic and orbitronic devices.
△ Less
Submitted 9 August, 2026;
originally announced August 2026.
-
Overview and status of BICEP Array's BA4-90/150 CMB polarimeter
Authors:
M. A. Petroff,
P. A. R. Ade,
Z. Ahmed,
M. Amiri,
D. Barkats,
R. Basu Thakur,
C. A. Bischoff,
D. Beck,
J. J. Bock,
V. Buza,
B. Cantrall,
J. R. Cheshire IV,
J. Connors,
J. Cornelison,
M. Crumrine,
A. J. Cukierman,
E. Denison,
L. Duband,
M. A. Echter,
M. Eiben,
B. D. Elwood,
S. Fatigoni,
J. P. Filippini,
A. Fortes,
M. Gao
, et al. (61 additional authors not shown)
Abstract:
The inflation paradigm postulates a period of rapid expansion in the early Universe, which would generate gravitational waves. These tensor perturbations would produce a faint B-mode signature in the polarization of the cosmic microwave background (CMB), but this signal is orders of magnitude weaker than that from the CMB's other anisotropy and that from astrophysical foregrounds. Placing more-str…
▽ More
The inflation paradigm postulates a period of rapid expansion in the early Universe, which would generate gravitational waves. These tensor perturbations would produce a faint B-mode signature in the polarization of the cosmic microwave background (CMB), but this signal is orders of magnitude weaker than that from the CMB's other anisotropy and that from astrophysical foregrounds. Placing more-stringent upper limits on this signal or making a definitive detection thus requires exceptional control over instrument and measurement systematics, in addition to extremely-deep maps. The fourth BICEP Array receiver, BA4-90/150, aims to build and improve upon the heritage of the field-leading BICEP series of small-aperture CMB experiments with a dichroic instrument observing in 90 and 150 GHz bands, to advance the search for the inflationary B-mode signal. The instrument will utilize transition-edge-sensor bolometers, which will be read out using a new two-level time-division-multiplexed system and be fed via feedhorn-coupled orthomode transducers and refined cold refractive optics, with the goal of both improving systematics control and sensitivity over existing receivers. With a planned deployment to the South Pole in the 2026-27 austral summer, the instrument will occupy the fourth and final remaining slot in the BICEP Array mount, completing the phaseout of Keck Array receivers. An overview of the BA4-90/150 receiver will be presented, along with a discussion of its current status and future plans for the instrument.
△ Less
Submitted 7 August, 2026;
originally announced August 2026.
-
DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning
Authors:
Le Xiang,
Zhicheng Guan,
Hong Chen,
Xiaocong Lin,
Zhenghua Lei,
Teng Hu,
Bolei He,
Long Zeng
Abstract:
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, including end-to-end MLLMs, retrieval-augmented generation (RAG) pipelines, and document agents, often lack explicit mechanisms to represent and verify how grounded eviden…
▽ More
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason over heterogeneous document elements distributed across multiple pages. Existing approaches, including end-to-end MLLMs, retrieval-augmented generation (RAG) pipelines, and document agents, often lack explicit mechanisms to represent and verify how grounded evidence is progressively composed during reasoning, limiting both answer accuracy and traceability. In this paper, we cast LongDocVQA as an explicit evidence graph reasoning problem rather than implicit answer prediction. To this end, we propose DocTrace, a hierarchical framework that progressively performs evidence localization, structured document parsing, and evidence graph reasoning to enable explicit evidence provenance. To effectively learn these capabilities, we develop a two-stage training framework: joint Supervised Fine-Tuning (SFT) first initializes evidence localization and graph reasoning abilities, followed by task-specific Group Relative Policy Optimization (GRPO) with dedicated rewards to further optimize these capabilities. Extensive experiments on MMLongBench-Doc, LongDocURL, and SlideVQA demonstrate that DocTrace consistently outperforms both existing open-source baselines and proprietary MLLMs. Compared with the Qwen3-VL-8B-Instruct backbone, DocTrace achieves absolute improvements of 14.4, 11.3, and 11.7 points on the three benchmarks, respectively. Beyond competitive performance, DocTrace constructs traceable evidence graphs with explicit node-level provenance, enabling transparent and verifiable reasoning for long document understanding.
△ Less
Submitted 4 August, 2026;
originally announced August 2026.
-
Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing
Authors:
Changhao Zhao,
Haoxiang Li,
Yuke Li,
Hai Liu,
LingLin Zeng
Abstract:
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts D…
▽ More
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recently, the recent release of DINOv3 brings DINO.txt, which equips the standalone DINO backbone with image-text contrastive learning and thus opens up the possibility of open-vocabulary segmentation. We propose DinoSplat-OV, a training-free framework that adapts DINOv3 to remote sensing without fine-tuning or additional pretraining. Targeting the dense distribution, multi-scale nature, and large size of remote sensing imagery, we design two core modules. Its Text-aware Laplacian Propagation module de-noises patch-level predictions by combining textual semantic affinities with local visual similarity, improving regional consistency while preserving boundaries. Its Gaussian Splatting Upsampling module reconstructs pixel-level features through RGB-guided anisotropic aggregation and test-time optimization. A global-anchor sliding-window strategy further supports large-scale imagery. Experiments on UDD5, DOTA, and LoveDA demonstrate competitive or superior performance over existing training-free methods, effectively filling the gap of DINO-series models in training-free open-vocabulary segmentation and providing a viable new path for further advances in this direction.
△ Less
Submitted 3 August, 2026;
originally announced August 2026.
-
Homebot: A Personal AI Agent for Conversational Home Assistance and Automation
Authors:
Shengyuan Ye,
Yixin Zhang,
Han Liang,
Liekang Zeng,
Jiangsu Du,
Mu Yuan
Abstract:
\texttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-messaging requests through a shared runtime that combines language-model responses with registered tools and task-specific skills. The design separates common request processing from session ownership: messaging history remains scoped to a channel and chat, whereas…
▽ More
\texttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-messaging requests through a shared runtime that combines language-model responses with registered tools and task-specific skills. The design separates common request processing from session ownership: messaging history remains scoped to a channel and chat, whereas voice interaction is bounded by wake-word activation. For hands-free use, \texttt{Homebot} combines local wake-word detection, streaming speech recognition and synthesis, and an explicit dialogue-state protocol for ending, following up, or continuing a conversation. Clear channel, tool, and skill contracts support practical customization for household use.
△ Less
Submitted 7 August, 2026; v1 submitted 3 August, 2026;
originally announced August 2026.
-
Design and Performance of 220 and 270 GHz Bandpass Filters for BICEP Array
Authors:
A. Steiger,
The BICEP/Keck Collaboration,
P. A. R. Ade,
Z. Ahmed,
M. Amiri,
D. Barkats,
R. Basu Thakur,
C. A. Bischoff,
D. Beck,
J. J. Bock,
H. Boenish,
V. Buza,
K. Carter,
J. R. Cheshire IV,
J. Connors,
J. Cornelison,
L. Corrigan,
M. Crumrine,
S. Crystian,
A. J. Cukierman,
E. Denison,
L. Duband,
M. Echter,
M. Eiben,
B. D. Elwood
, et al. (68 additional authors not shown)
Abstract:
The BICEP Array (BA) is the latest in the BICEP/ Keck series of experiments that aim to measure the polarization of the cosmic microwave background (CMB) with small aperture polarimeters located at the South Pole. To constrain the frequency response of these receivers, each detector is serially coupled to a band-pass filter (BPF). The electric circuits of these BPFs utilize series and shunt capaci…
▽ More
The BICEP Array (BA) is the latest in the BICEP/ Keck series of experiments that aim to measure the polarization of the cosmic microwave background (CMB) with small aperture polarimeters located at the South Pole. To constrain the frequency response of these receivers, each detector is serially coupled to a band-pass filter (BPF). The electric circuits of these BPFs utilize series and shunt capacitors as well as series inductors, but critically do not include shunt inductors which simplifies fabrication. The filters are designed and simulated with Sonnet, and optimized for noise by considering loading from the atmosphere and the CMB. Multiple 220 GHz detector modules have had their frequency response measured at the South Pole. The 270 GHz detector modules have recently begun testing in a lab setting, and their performance in a BA receiver will be measured this winter.
△ Less
Submitted 31 July, 2026;
originally announced August 2026.
-
Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
Authors:
Xinke Tong,
Xuanming Zhang,
Tianyi Tang,
An Yang,
Jiatu Hu,
Guojie Lin,
Zhenzhen Shi,
Lingfeng Zeng,
Boyu Yang,
Bing Zhao,
Hu Wei,
Lin Qu,
Dayiheng Liu
Abstract:
Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture real-world industrial complexity, predominantly relying on multiple-choice questions or single-hop QA over cropped tables while ig…
▽ More
Do Large Language Models (LLMs) possess genuine structural reasoning, or merely rely on surface-level pattern matching? The financial domain, demanding numerical precision and multi-step logic over long contexts, is an ideal testbed. Existing benchmarks fail to capture real-world industrial complexity, predominantly relying on multiple-choice questions or single-hop QA over cropped tables while ignoring intricate cross-statement dynamics and temporal de-cumulation.
To bridge this gap, we introduce FinIndices, a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens). Utilizing an automated synthesis pipeline with adversarial traps, FinIndices encompasses Single-Index computation and Table-Index tabulation to test complex domain, temporal, and caliber reasoning.
Our evaluation reveals two severe LLM vulnerabilities. First, a "Knowledge Bottleneck": despite memorizing formulas during pre-training, models demonstrate fragile pattern matching. Removing explicit formula hints causes performance to collapse (e.g., Gemini-3.1-Pro drops from 70.70% to 38.22% on table tasks), exposing fatal flaws in temporal de-cumulation and stock-flow caliber mismatch. Second, a "Structural Bottleneck": the intense cognitive load of generating multi-metric, multi-period tables actively drains reasoning capacity. Under structural pressure, LLMs that flawlessly execute isolated derivations regress to shallow heuristics, such as fetching incorrect adjacent columns or substituting deep accounting adjustments with lazy literal arithmetic. Finally, Supervised Fine-Tuning (SFT) yields substantial zero-hint gains (+8.54% Single, +3.82% Table), validating that structured logic can be partially restored via data-centric alignment.
△ Less
Submitted 22 July, 2026;
originally announced July 2026.
-
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
Authors:
Lingyang Zeng,
Guangze Chen,
Kaichen Yu,
Zhicheng Pan,
Siyang Weng,
Zirui Hu,
Xiangyun Du,
Hailin He,
Rong Zhang,
Chengcheng Yang,
Kai Huang,
Xuan Zhou
Abstract:
Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversation…
▽ More
Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.
△ Less
Submitted 3 August, 2026; v1 submitted 29 July, 2026;
originally announced July 2026.
-
FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents
Authors:
Lingteng Zeng,
Yifan Jin
Abstract:
Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse can remove GPU-bound generation work, yet response caches require dependency consistency when filings, evidence chunks, and tool outputs change. FinCacheServe treats each generated answer as a serving object indexed by enterprise intent and guarded by…
▽ More
Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse can remove GPU-bound generation work, yet response caches require dependency consistency when filings, evidence chunks, and tool outputs change. FinCacheServe treats each generated answer as a serving object indexed by enterprise intent and guarded by document versions, evidence fingerprints, tool fingerprints, model identity, and decoding configuration. A vLLM implementation evaluates SEC-derived financial-document workloads with Qwen2.5 models. On a 2,230-request hosted 7B trace, FinCacheServe skips 53.27% of LLM calls with zero observed dependency-stale outputs. Across three hosted 32B operator-suite seeds, it skips 53.31% of 544 requests, compared with 38.97% for versioned semantic caching and 22.43% for grounded-style reuse. Capacity, backend, and SLO replays show oracle-bounded cache management, 100k-entry transactional metadata behavior, and 44.30% lower estimated Wh per dependency-fresh 2s-SLO success than versioned semantic caching.
△ Less
Submitted 13 July, 2026;
originally announced July 2026.
-
TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models
Authors:
Meng Xie,
Li Zeng,
Hangtao Zhang,
Xianlong Wang,
Ziqi Zhou,
Pengpeng Qiao,
Zhetao Li
Abstract:
Recent commercial image-generation models can generate high-quality images with readable text (e.g., posters, infographics, and manuals), attracting considerable attention. Yet we first show that this same capability also introduces a previously unreported safety vulnerability: these systems may refuse to generate harmful text directly, yet permit the same content when rendered as text within gene…
▽ More
Recent commercial image-generation models can generate high-quality images with readable text (e.g., posters, infographics, and manuals), attracting considerable attention. Yet we first show that this same capability also introduces a previously unreported safety vulnerability: these systems may refuse to generate harmful text directly, yet permit the same content when rendered as text within generated images, i.e., safety alignment does not reliably transfer from textual outputs to text embedded in images. In this paper, unlike existing visual jailbreaks against image-generation models, which primarily induce models to generate harmful visual objects or scenes, we introduce the concept of instruction-dense visual jailbreaks, in which image-generation models produce detailed, readable, and actionable harmful instructions within images. Such outputs can amplify harm because the rendered instructions can be readily read and widely spread. To instantiate this threat, we propose TYPO, a black-box framework that exploits this safety gap by automatically generating adversarial TYPOgraphy prompts, which covertly steer image-generation models to express harmful intent as highly legible, typographically structured text. Specifically, TYPO decomposes prompt generation into two channels: a textual channel for reframing the target intent, and a visual channel for specifying its presentation form. We formulate these two channels as a dual-channel textual-visual strategy space and optimize candidate strategy combinations through an adaptive combinatorial search. Extensive experiments across four commercial models (i.e., GPT-Image-2, Nano Banana Pro, Qwen-Image-2, and Seedream 5.0 Lite) show that TYPO substantially outperforms nine representative jailbreak attacks by 50.2% in ASR on average, while incurring an average query cost of only $0.04.
△ Less
Submitted 27 July, 2026;
originally announced July 2026.
-
A Gaussian smoothing-based zeroth-order method for Goldstein second-order stationarity
Authors:
Ming Lei,
Ting Kei Pong,
Man-Chung Yue,
Liaoyuan Zeng,
Hao Zhang
Abstract:
We introduce a new generalized Hessian, called the Goldstein second-order $δ$-subdifferential, and an associated notion of $(ε_1,ε_2,δ)$-second-order stationary point for continuously differentiable functions with locally Lipschitz gradients. We propose a zeroth-order algorithm based on cubic regularization and Gaussian smoothing with homotopy to find such approximate second-order stationary point…
▽ More
We introduce a new generalized Hessian, called the Goldstein second-order $δ$-subdifferential, and an associated notion of $(ε_1,ε_2,δ)$-second-order stationary point for continuously differentiable functions with locally Lipschitz gradients. We propose a zeroth-order algorithm based on cubic regularization and Gaussian smoothing with homotopy to find such approximate second-order stationary points for Lipschitz differentiable functions, and derive the iteration complexity under a mild coercivity-type assumption on the objective function.
△ Less
Submitted 23 July, 2026;
originally announced July 2026.
-
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Authors:
Dongfang Li,
Xiaodong Luo,
Ruoyu Sun,
Xuhui Chen,
Linyuan Qiu,
Jian Meng,
Zhengxuan Lu,
Yiting Wang,
Yucheng Xie,
Tao Guo,
Tianxiang Fang,
Jing Li,
Sihang Chen,
Shihao Hong,
Chang Liu,
Weihua Dai,
Zirong Zeng,
Ziwei Zhu,
Zhuohan Wang,
Zhengjun Yue,
Igor Vasilyev,
Min Liu,
Weijian Sun,
Xin Chen,
Yingmeng Gao
, et al. (40 additional authors not shown)
Abstract:
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on…
▽ More
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distributed training, including severe memory pressure, non-overlapped communication overhead, and inefficient kernel execution. While most large-scale LLM training systems are built around GPU-based clusters, this report presents an end-to-end optimization practice on the Ascend NPU SuperPOD. Using the DeepSeek-V4 model family as the target workload, we develop a hierarchical optimization framework spanning model-level parallelism, computation-communication orchestration, and low-level kernel execution. The resulting system achieves 34.22% Model FLOPs Utilization (MFU) with a 2.93x improvement over the open-source baseline recipe while maintaining training stability. Building on this optimized infrastructure, we further establish a CPT and SFT workflow for complex Operations Research (OR) tasks. We refer to the integrated framework as SLAI T-Rex. Using DeepSeek-V4-Flash, we develop OR-oriented CPT and SFT data pipelines that combine collected domain resources with solver-verified synthetic optimization documents. The resulting dataset contains 10K high-quality SFT samples spanning four task categories and three problem representations. The specialized model achieves the highest average zero-shot Pass@1 score among the evaluated models, reaching 71.81% and outperforming GPT-5.4-Mini and the base DeepSeek-V4-Flash model by 3.98 and 11.27 percentage points, respectively. Overall, this work demonstrates a full-stack pathway from efficient trillion-parameter model post-training on Ascend infra to domain-specialized Flash models for solver-grounded mathematical modeling, advancing frontier-model systems for complex reasoning.
△ Less
Submitted 19 August, 2026; v1 submitted 22 July, 2026;
originally announced July 2026.
-
GhostPrompt: Cross-Image Adversarial Prompt for Vision-Language Models
Authors:
Li Zeng,
Zeyu Ye,
Meng Xie,
Hangtao Zhang,
Xianlong Wang,
Yanchun Li,
Zhetao Li
Abstract:
Vision-Language Models (VLMs) are known to be vulnerable to adversarial attacks, where subtle perturbations to images or texts induce erroneous outputs. However, most text-based attacks are adapted from language-model-centric methods, in which the visual input is fixed during optimization, resulting in adversarial prompts that are tied to specific images and thus limiting their attack effectivenes…
▽ More
Vision-Language Models (VLMs) are known to be vulnerable to adversarial attacks, where subtle perturbations to images or texts induce erroneous outputs. However, most text-based attacks are adapted from language-model-centric methods, in which the visual input is fixed during optimization, resulting in adversarial prompts that are tied to specific images and thus limiting their attack effectiveness. To this end, we first introduce a new research perspective: cross-image transferability for adversarial prompts. We then propose GhostPrompt, an adversarial prompt that is optimized once and reused to steer VLM outputs toward attacker-specified responses across diverse images. GhostPrompt employs a joint optimization that distills image-invariant adversarial features into the prompt by "worst-case" generation. Specifically, it alternates between constructing hard visual conditions for the current prompt and updating the prompt to remain effective under these conditions. Extensive experiments on prevalent VLMs verify that \ourmethod achieves an improvement of over 30% in attack success rates compared to state-of-the-art (SoTA) baselines, while reducing computation time by ~70%. Our code is avalable at https://github.com/Ye-ze-yu/GhostPrompt.
△ Less
Submitted 21 July, 2026;
originally announced July 2026.
-
Biodegradable, Millimeter-Scale Light-Emitting Sensors for Distributed Environmental Monitoring-Functional Pixie Dust
Authors:
Zhiming Hu,
Danzhen Zhang,
Janghun Ko,
Haohui Zhang,
Jiale Chen,
Chanho Park,
Jiatong Zhang,
Qiuna Zhuang,
Shiwei Xu,
Xiaoran Yang,
Dain Son,
Taehoon Kim,
Uikang Joo,
Zhaojian Xu,
Hyunsoo Kim,
Richard Chai,
Gwangmin Bae,
Wooyoul Maeng,
Qiong Wang,
Sangmin Lim,
Liangsong Zeng,
Un-Seong Baik,
Kaiqing Zhang,
Liming Yuan,
Yonggang Huang
, et al. (2 additional authors not shown)
Abstract:
Methods for large-area, precise monitoring across natural environments are of growing interest due to pressing needs for sustainable management of rapidly increasing anthropogenic activities. Established approaches involve sparse spatial sampling and/or sequential measurements, while emerging techniques exploit miniaturized electronics or passive optical methods. Various constraints in scalability…
▽ More
Methods for large-area, precise monitoring across natural environments are of growing interest due to pressing needs for sustainable management of rapidly increasing anthropogenic activities. Established approaches involve sparse spatial sampling and/or sequential measurements, while emerging techniques exploit miniaturized electronics or passive optical methods. Various constraints in scalability, costs, robustness, operational range and other factors create a need for alternatives. Here, we introduce a concept that overcomes many of these limitations through the combined use of chemically induced light emission and chemically responsive optical filter elements in millimeter-scale systems that we refer to as functional pixie dust (fPD) sensors, designed specifically for monitoring natural water systems during nighttime to eliminate background optical interference and to enhance remote analysis. These floating devices act as Lagrangian tracers to follow surface flows and to simultaneously measure the concentrations of key chemical species along their trajectories. Optimized designs exploit environmentally compatible constituent materials that are also degradable through natural processes to benign end products, thereby eliminating the need for recovery. Spatially and spectrally resolved ratiometric measurement schemes ensure robust operation and ability to address practical requirements in range, operational lifetime, time response and sensitivity. Demonstrations include distributed measurements of pH, Hg2+, and NO2-, each of relevance to industrial discharge, toxic metal contamination, and nitrogen-rich runoff, adapted for static concentration gradients, flow-driven transport conditions, and outdoor aquatic settings. The results establish a framework for environmental sensing using degradable, self-powered microsystems capable of scalable deployment and remote readout.
△ Less
Submitted 20 July, 2026;
originally announced July 2026.
-
StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
Authors:
Sizhong Qin,
Yi Gu,
Yao Jiang,
Ao Cai,
Changjian Zhou,
Shaoxuan Shuai,
Jiachang Wang,
Tianhao Shen,
Yueqiang Li,
Xinhao Li,
Li Zeng,
Yueshi Chen,
Dachen Gao,
Genrong Xu,
Wenjie Liao,
Xinzheng Lu
Abstract:
Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, applicable engineering checks, and a final report. Evaluations centered on question answering or script generation may therefore reward fluent outputs even when the underlying workflow is i…
▽ More
Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, applicable engineering checks, and a final report. Evaluations centered on question answering or script generation may therefore reward fluent outputs even when the underlying workflow is incomplete, inconsistent, or non-executable. We present StructureClaw, an artifact-centered workbench in which LLM agents operate through governed engineering skills, typed tools, shared artifact state, and local analysis backends, together with StructureClaw-Bench, an executable benchmark of 150 controlled scenarios spanning standard workflows, interactive robustness, and multimodal structural-model reconstruction. Its analyzable standard and multimodal cases require both strict one-to-one structural-model matching and numerical-response agreement with frozen reference responses from the selected analysis engine; interactive cases instead require positive clarification or recovery evidence together with safe non-execution when appropriate. A trial succeeds only when every fixture-required assertion passes. Across nine text-agent configurations, generic-only execution passed the model-artifact check in 87.0% of retained outcomes but achieved only 22.0% E2E Success, whereas automatic StructureClaw reached 82.9%. Interactive and multimodal evaluations further identify semantic state consistency and executable model reconstruction as the dominant remaining bottlenecks. The code and benchmark are available at https://github.com/structureclaw/structureclaw.
△ Less
Submitted 3 August, 2026; v1 submitted 16 July, 2026;
originally announced July 2026.
-
PVDetector: Detecting Prompt Injection Attacks on Purpose-Specific LLM Agents through Policy-Violation Concept Analysis
Authors:
Junhui Wang,
Hangtao Zhang,
Zhirun Zheng,
Li Zeng,
Jiejun Xiao,
Xi Luo,
Lihua Yin,
Saiqin Long
Abstract:
Large language models (LLMs) are increasingly deployed as purpose-specific agents to handle domain-specific tasks such as customer service and code generation. These agents are expected to comply with not only generic safety guardrails but also purpose-specific restrictions tailored to their designated roles. Such additional restrictions enlarge the attack surface, particularly to prompt injection…
▽ More
Large language models (LLMs) are increasingly deployed as purpose-specific agents to handle domain-specific tasks such as customer service and code generation. These agents are expected to comply with not only generic safety guardrails but also purpose-specific restrictions tailored to their designated roles. Such additional restrictions enlarge the attack surface, particularly to prompt injection (PI) attacks. To defend against such attacks, existing detection methods primarily rely on analyzing input-output patterns, yet yield limited effectiveness. To address this limitation, we turn to analyzing the hidden activation space and discover that LLMs inherently retain latent policy-violation (PV) concepts when prompted with requests beyond their designated purpose. Particularly, PV concepts capture the semantics of conflicts between user queries and predefined restrictions, implicitly reflecting LLMs' intrinsic awareness of recognizing policy violations. Building on this insight, we propose PVDetector, a training-free framework that detects PI attacks during LLM inference by measuring hidden-state alignment with PV concepts, which are derived offline from the contrastive pairs of policy-violating and policy-compliant prompts. Experiments across multiple LLMs and datasets show that PVDetector achieves <1\% false negative rate with minimal auxiliary overhead, consistently outperforming state-of-the-art methods. Our code is available at https://github.com/Claresigle/PVDetector .
△ Less
Submitted 14 July, 2026;
originally announced July 2026.
-
PRISM Edit: One Vector for All Temporal Answers
Authors:
Chen Huang,
Qi Zheng,
Ruiqin Zheng,
Long Zeng,
Yuantong Xu
Abstract:
Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer should become current while the old answer may remain correct in historical time contexts. Building on this insight, we use causal tracing to show that LLMs alrea…
▽ More
Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing locate-and-edit paradigm: an update is not always a replacement. When a fact changes, the new answer should become current while the old answer may remain correct in historical time contexts. Building on this insight, we use causal tracing to show that LLMs already support this distinction via a two-stage internal computation: early MLP layers retrieve a time-agnostic subject representation, and later layers modulate it with temporal context to yield the time-correct answer. Motivated by this finding, we introduce PRISM Edit, which optimizes a single polysemous representation across temporal contexts and leverages the model's inherent modulation pathway to route it to temporally correct predictions, without any architectural modification. We evaluate on TimeConflict, a new temporal editing benchmark we introduce, and on temporally augmented CounterFact. PRISM Edit improves over the best baseline by +23.3 Temporal Consistency (TC) and +33.7 Current Relative-time Score (CRS) on average while being more than 2x faster. Code and data are publicly available at https://github.com/AnonymousStudy972/PRISM-Edit.
△ Less
Submitted 13 July, 2026; v1 submitted 13 July, 2026;
originally announced July 2026.
-
Active Learning for Channel Knowledge Map Construction via Bayesian Inference Diffusion Models
Authors:
Yunzhe Zhu,
Xuewen Liao,
Zhenzhen Gao,
Linzhou Zeng,
Yong Zeng
Abstract:
Channel knowledge maps (CKMs) are regarded as key enablers of environment-aware communications in future wireless networks, as they provide location-specific channel information by establishing an explicit connection between wireless devices and the physical propagation environment. As a representative CKM, the channel gain map (CGM) characterizes the spatial distributions of large-scale fading to…
▽ More
Channel knowledge maps (CKMs) are regarded as key enablers of environment-aware communications in future wireless networks, as they provide location-specific channel information by establishing an explicit connection between wireless devices and the physical propagation environment. As a representative CKM, the channel gain map (CGM) characterizes the spatial distributions of large-scale fading to support wireless environment awareness and network optimization. Existing CGM construction methods generally lack a well-defined sampling-point acquisition strategy, which may result in a limited number of sampling points being allocated to spatially redundant or highly predictable regions, thereby degrading CGM reconstruction performance in complex propagation environments. In this paper, we propose an active-learning-based diffusion framework for efficient CGM construction. By combining Bayesian inference with the diffusion model, the proposed method estimates epistemic uncertainty without retraining the model. Two uncertainty quantification algorithms are further developed along the reverse diffusion process to generate element-wise epistemic uncertainty maps. Furthermore, an uncertainty-aware sampling strategy is designed to determine new observation locations by jointly considering epistemic uncertainty and spatial distribution uniformity. Experimental results on both static and dynamic CGM datasets demonstrate that the proposed method achieves better reconstruction performance than baseline methods. These results indicate that the proposed method can effectively improve the utilization efficiency of limited sampling points and enhance the accuracy of CGM construction in complex wireless propagation environments.
△ Less
Submitted 29 June, 2026;
originally announced June 2026.
-
Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning
Authors:
Yanzhe Tang,
Xinyu Shao,
Yuxuan Hu,
Siyu Chen,
Bowen Yang,
Yajun Gao,
Tongtong Cao,
Xiu Li,
Long Zeng
Abstract:
While end-to-end Vision-Language-Action (VLA) models show promise in robotic manipulation, their monolithic paradigm inherently couples semantic reasoning and spatial control. This creates a severe alignment bottleneck, limiting precise target disambiguation in data-constrained imitation learning. To overcome this, we propose SVP-IL, a decoupled architecture that explicitly extracts spatial visual…
▽ More
While end-to-end Vision-Language-Action (VLA) models show promise in robotic manipulation, their monolithic paradigm inherently couples semantic reasoning and spatial control. This creates a severe alignment bottleneck, limiting precise target disambiguation in data-constrained imitation learning. To overcome this, we propose SVP-IL, a decoupled architecture that explicitly extracts spatial visual grounding from the action generation loop. By leveraging vision-language foundation models, we parse instructions into zero-shot geometric masks, translating language into explicit Spatial Visual Prompts (SVP). These priors are injected into a continuous action generator via a lightweight direct feature-level fusion mechanism. This integration provides explicit and uncorrupted spatial gradient guidance while ensuring highly stable optimization under low-data regimes. Extensive experiments demonstrate that SVP-IL significantly outperforms state-of-the-art VLAs and pure visuomotor baselines. Trained on as few as 50 to 100 demonstrations, SVP-IL improves average success rates on highly ambiguous language-conditioned tasks from 24.0% to 39.5%, achieving 67.8% on standard benchmarks. Real-world robotic experiments further validate its robustness and data efficiency in unstructured physical environments.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Addressing the lightest $S$-wave strange $K_0^*(700)/κ$ resonance in four-body semileptonic $\bar{B}_s^0 \to K^0π^+ \ell^-\barν_\ell$ decays
Authors:
Dong Huang,
Sheng-Bo Wu,
Fang-Ping Peng,
Long Zeng,
Hai-Bing Fu
Abstract:
As the lightest strange scalar resonance, $K_0^*(700)$ (also called $κ$) has a large width and resides close to the $Kπ$ threshold, leading to a longstanding debate about its internal structure. Within the framework of the conventional quark-antiquark ($q\bar{q}$) picture, this paper attempts to research the behaviors of $K_0^*(700)$ resonance in the four-body final state decays. Firstly, we inves…
▽ More
As the lightest strange scalar resonance, $K_0^*(700)$ (also called $κ$) has a large width and resides close to the $Kπ$ threshold, leading to a longstanding debate about its internal structure. Within the framework of the conventional quark-antiquark ($q\bar{q}$) picture, this paper attempts to research the behaviors of $K_0^*(700)$ resonance in the four-body final state decays. Firstly, we investigates the $K_0^*(700)$ resonance one twist-2 and two twist-3 light-cone distribution amplitudes (LCDAs), {\it i.e.} $φ_{2;K_0^*(700)}(x, μ)$ and $φ_{3;K_0^*(700)}^{p;σ}(x, μ)$ by constructing light-cone harmonic oscillator model. For the longitudinal distribution characteristics of the twist-2 LCDA, two typical parametrization schemes, {\it i.e.} $\varphi_{2;K_0^*(700)}^{\rm (S1)}(x)$ and $\varphi_{2;K_0^*(700)}^{\rm (S2)}(x)$, are adopted in this paper. Meanwhile, it has enriched theoretical predictions for the first ten-order LCDAs $ξ$-moments. Subsequently, we apply the QCD light-cone sum rules approach to compute the $\bar B_s^0 \to K_0^*(700)$ transition form factors (TFFs) $f_\pm^{\bar B_s^0 K_0^*(700)}(q^2)$ and $f_{\rm T}^{\bar B_s^0 K_0^*(700)}(q^2)$ by taking into account the perturbative $\mathcal{O}(α_s)$ corrections to the twist-2 LCDAs terms. By using the simplified series expansion parameterization, these TFFs are extrapolated over the full physical region. Then, we calculate the differential decay widths and branching fractions for four-body semileptonic decay $\bar{B}_s^0 \to K^0π^+ \ell^-\barν_\ell$ with $\ell= (e, μ, τ)$ via the $S$-wave $K_0^*(700)$ resonance, respectively. We expect that the obtained prediction results will provide valuable insights for future experiments and theoretical research.
△ Less
Submitted 23 June, 2026;
originally announced June 2026.
-
Hamiltonian formulation of Carrollian Maxwell theory in Deformed Light-cone Kaluza-Klein-like Null reduction
Authors:
Limin Zeng
Abstract:
We construct magnetic and electric Carrollian Maxwell theories by performing Kaluza-Klein-like null reduction of a complex Maxwell field in a Bargmann deformed light-cone background with manifest gauge symmetry. The procedure preserves a first-class U(1) Gauss constraint throughout the Carrollian limit. Gauge invariance is therefore maintained in our Hamiltonian formulation. By choosing different…
▽ More
We construct magnetic and electric Carrollian Maxwell theories by performing Kaluza-Klein-like null reduction of a complex Maxwell field in a Bargmann deformed light-cone background with manifest gauge symmetry. The procedure preserves a first-class U(1) Gauss constraint throughout the Carrollian limit. Gauge invariance is therefore maintained in our Hamiltonian formulation. By choosing different scalings, we obtain standard magnetic Carrollian theory and electric Carrollian theory. However, a scalar field could appear in the Carrollian theory in a coupled or decoupled way, which has not been found by previous methods. This result fully reveals the diversity of Carrollian theories accessible through the deformed light-cone Kaluza-Klein-like null reduction method. Furthermore, our work provides an explicit example of the correct application of this approach, thereby broadening the scope of its applicability to gauge theories.
△ Less
Submitted 25 June, 2026; v1 submitted 20 June, 2026;
originally announced June 2026.
-
ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters
Authors:
Jiasheng Zhou,
Longbin Zeng,
Clavis Chen,
Ruiming Lu,
Qinwei Yang,
Leyi Ye,
Ray Ying,
Key Zhang
Abstract:
Large-scale LLM training requires always-on, fine-grained observability for effective performance diagnosis at scale. Coarse resource monitors alone cannot localize root causes, and fine-grained profilers incur prohibitive (5%-30%) overheads and massive trace volumes, making always-on deployment impractical in large production clusters.
We propose ARGUS, a low-overhead, fine-grained, always-on t…
▽ More
Large-scale LLM training requires always-on, fine-grained observability for effective performance diagnosis at scale. Coarse resource monitors alone cannot localize root causes, and fine-grained profilers incur prohibitive (5%-30%) overheads and massive trace volumes, making always-on deployment impractical in large production clusters.
We propose ARGUS, a low-overhead, fine-grained, always-on tracing and real-time analysis system for training workloads in 10,000+ GPU-scale production clusters. ARGUS decomposes observation along the training call hierarchy into CPU call stacks, framework semantics, and GPU kernel execution, with always-on collection under a combined overhead of less than 2%. It builds a unified data pipeline and compresses raw kernel events by approximately 3,700x from 10 MB to 2.7 KB per rank per step. Its progressive diagnosis framework automatically isolates anomalous windows, straggler ranks, and degraded kernels through iteration-time, phase-level, and kernel-level analysis. Deployed for over six months on a 10,000+ GPU production cluster, ARGUS has supported continuous fail-slow detection and performance optimization. Our case studies further demonstrate its effectiveness across representative anomalies, including compute stragglers, link degradation, pipeline-bubble amplification, FlashAttention JIT stalls, and compute stragglers masked by communication symptoms.
△ Less
Submitted 8 July, 2026; v1 submitted 18 June, 2026;
originally announced June 2026.
-
Accelerating Network-Agent Dispersion: Territorial Behavior and Directionally Biased Lazy Random Walks
Authors:
Li Zeng,
Steve Alpern
Abstract:
Territorial behavior can greatly accelerate decentralized agent dispersion on networks. This paper studies a network-agent dispersion problem in which m autonomous agents move in discrete time on a connected graph and seek a configuration in which no two agents occupy the same node. We focus on the dispersion case m = n, where successful configurations contain exactly one agent per node. In the ba…
▽ More
Territorial behavior can greatly accelerate decentralized agent dispersion on networks. This paper studies a network-agent dispersion problem in which m autonomous agents move in discrete time on a connected graph and seek a configuration in which no two agents occupy the same node. We focus on the dispersion case m = n, where successful configurations contain exactly one agent per node. In the baseline model, each agent follows a lazy random walk with a common laziness parameter p. This process defines a finite absorbing Markov chain, and the expected absorption time is used to measure dispersion efficiency. We introduce two local behavioral extensions: territorial behavior, in which an agent that is alone at a node claims that node and repels later arrivals, and directional bias, in which agents share a preferred direction of movement on paths and cycles. Exact calculations on three-agent path and cycle networks and Monte Carlo simulations on larger instances show that territorial behavior substantially reduces expected dispersion time, with larger relative reductions as network size increases. Directional bias alone has limited effect in most small-network cases, but when combined with territorial behavior it can produce large additional speedups. In particular, the simulations show reductions of 99.22% on L100 and 97.48% on C100 when all agents start from one node. These results show how simple local movement rules can strongly affect global dispersion time in decentralized networked multi-agent systems.
△ Less
Submitted 17 June, 2026;
originally announced June 2026.
-
ReAge3D: Re-Aging 3D Faces with View Consistency
Authors:
Libing Zeng,
Li Ma,
Mingming He,
Ning Yu,
Paul Debevec,
Nima Khademi Kalantari
Abstract:
We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity-preserving results. Existing 3D editing methods, while effective for coarse semantic changes, are not well suited for re-aging, as even small inconsistencies across re-aged 2D views can lead to over-smoothing of subtle but perceptually important age-related details. To address this…
▽ More
We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity-preserving results. Existing 3D editing methods, while effective for coarse semantic changes, are not well suited for re-aging, as even small inconsistencies across re-aged 2D views can lead to over-smoothing of subtle but perceptually important age-related details. To address this challenge, we first introduce a 2D diffusion-based re-aging model, DiffReaging, trained on synthetically generated image pairs. We further propose a center-out editing propagation strategy that leverages this re-aging model to reconstruct multi-view-consistent re-aged images. Specifically, starting from a re-aged frontal pivot view, we reconstruct the remaining views through warping and our proposed Masked-DiffReaging process. By injecting existing content at every step of the diffusion process, Masked-DiffReaging ensures that the reconstructed regions remain coherent with existing pixels. The resulting consistent set of re-aged views supervises the optimization of the re-aged 3D representation. Our method outperforms existing 3D editing techniques both visually and quantitatively, enabling smooth, fine-grained control over age transformations in 3D face models.
△ Less
Submitted 16 June, 2026;
originally announced June 2026.
-
MiniMax Sparse Attention
Authors:
Xunhao Lai,
Weiqi Xu,
Yufeng Yang,
Qiaorui Chen,
Yang Xu,
Lunbin Zeng,
Xiaolong Li,
Haohai Sun,
Haichao Zhu,
Vito Zhang,
Jinkai Hu,
Jiayao Li,
Rui Gao,
Zekun Li,
Songquan Zhu,
Jingkai Zhou,
Pengyu Zhao
Abstract:
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale. We introduce MiniMax Sparse Attention (MSA), a blockwise sparse attention b…
▽ More
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundreds of thousands to millions of tokens, yet the quadratic cost of softmax attention makes this untenable at deployment scale. We introduce MiniMax Sparse Attention (MSA), a blockwise sparse attention built upon Grouped Query Attention (GQA). A lightweight Index Branch scores key-value blocks and independently selects a Top-k subset for each GQA group, enabling group-specific sparse retrieval while maintaining efficient block-level execution; the Main Branch then performs exact block-sparse attention over only the selected blocks. Designed around a principle of simplicity and scalability, MSA is deliberately streamlined, making it straightforward to deploy efficiently across a broad range of GPUs. To translate sparsity into practical speedups, we co-design MSA with a GPU execution path that uses exp-free Top-k selection and KV-outer sparse attention to improve tensor-core utilization under block-granular access. On a 109B-parameter model with native multimodal training, MSA performs on par with GQA while reducing per-token attention compute by 28.4x at 1M context. Paired with our co-designed kernel, MSA achieves 14.2x prefill and 7.6x decoding wall-clock speedups on H800. Our inference kernel is available at: https://github.com/MiniMax-AI/MSA. A production-grade natively multimodal model powered by MSA has been publicly released at: https://huggingface.co/MiniMaxAI/MiniMax-M3.
△ Less
Submitted 12 June, 2026; v1 submitted 11 June, 2026;
originally announced June 2026.
-
Chirp Parameter Optimization and Distributed Detection for Cooperative RSMA-AFDM Systems
Authors:
Qingyu Li,
Guanghui Liu,
Yusha Liu,
Fuchen Xu,
Chengxiang Liu,
Hongjun Liu,
Liaoyuan Zeng
Abstract:
Affine frequency division multiplexing (AFDM) exhibits excellent Doppler robustness and the ability to characterize doubly selective channels. However, its signal dispersion characteristics make it challenging to directly adopt traditional time-frequency multiple access schemes. To address this issue, we introduce cooperative rate splitting multiple access (RSMA) for AFDM systems. The flexible con…
▽ More
Affine frequency division multiplexing (AFDM) exhibits excellent Doppler robustness and the ability to characterize doubly selective channels. However, its signal dispersion characteristics make it challenging to directly adopt traditional time-frequency multiple access schemes. To address this issue, we introduce cooperative rate splitting multiple access (RSMA) for AFDM systems. The flexible configuration of AFDM chirp parameters can reduce the correlation between users' equivalent channels, which decreases the interference from RSMA private streams. We conduct a theoretical analysis of the cooperative RSMA-AFDM system and demonstrate that minimizing the overlap in the channel column spaces among users can effectively enhance the system performance. Guided by this analysis, we design a chirp parameter optimization scheme that reduces multi-user interference and maximizes diversity gain. To fully exploit the diversity gain brought by the proposed chirp parameter optimization, two expectation propagation (EP)-based distributed cooperative detection schemes are proposed. First, a decision-fusion-based method is developed, where local information and cooperative information are fused by maximum ratio combining, achieving a globally consistent estimate of the common stream. Second, we develop a belief-consensus EP-based detection scheme. In each iteration, user nodes exchange and fuse the first- and second-order statistics of the common stream, and the resulting beliefs gradually converge to a consistent global decision, which significantly improves the overall reliability.
△ Less
Submitted 10 June, 2026;
originally announced June 2026.
-
Gauge Symmetry Degeneration in Lorentzian Deformed Light-Cone Null Reduction
Authors:
Limin Zeng
Abstract:
In this work, we apply deformed light-cone null reduction method to a complex Maxwell theory in a manifestly gauge-invariant formulation. We show that the local U(1) gauge structure degenerates in the $c\to 0$ limit: the Gauss law constraint reduces from a restriction on initial data to a conservation law, releasing the longitudinal gauge mode as an independent degree of freedom (d.o.f). This rais…
▽ More
In this work, we apply deformed light-cone null reduction method to a complex Maxwell theory in a manifestly gauge-invariant formulation. We show that the local U(1) gauge structure degenerates in the $c\to 0$ limit: the Gauss law constraint reduces from a restriction on initial data to a conservation law, releasing the longitudinal gauge mode as an independent degree of freedom (d.o.f). This raises the physical field count from $2(d-1)$ to $2d$. We prove a no-go theorem: under the single-mode Kaluza-Klein(KK)-like ansatz, no scaling of the field components can simultaneously preserve nontrivial dynamics and a first-class Gauss law, due to an inherent mismatch between velocity-type and constraint-type contributions in the parent action. Rather than representing the Carrollian electrodynamics derived via group contraction, the free complex scalar theory that emerges is merely an artifact of the truncation procedure at $c\to0$.
△ Less
Submitted 15 June, 2026; v1 submitted 10 June, 2026;
originally announced June 2026.
-
DexPIE: Stable Dexterous Policy Improvement from Real-World Experience
Authors:
Ruizhe Liao,
Wenrui Chen,
Liangji Zeng,
Haoran Lin,
Fan Yang,
Kailun Yang,
Yaonan Wang
Abstract:
Dexterous manipulation presents substantial challenges for imitation learning due to its high-dimensional action space and complex contact-rich dynamics. Policies trained purely from demonstrations often suffer from compounding errors during deployment and require large amounts of expert data to achieve reliable performance. To move beyond the limitations of demonstration data, in this work, we pr…
▽ More
Dexterous manipulation presents substantial challenges for imitation learning due to its high-dimensional action space and complex contact-rich dynamics. Policies trained purely from demonstrations often suffer from compounding errors during deployment and require large amounts of expert data to achieve reliable performance. To move beyond the limitations of demonstration data, in this work, we propose DexPIE, a post-training framework for dexterous policy improvement from experience collected through real-world deployment. First, DexPIE enables effective exploration coverage through a dexterous-hand-adapted intervention system and multi-stage DAgger-style data collection across initial and intermediate task stages, providing reliable supervision for accurate policy evaluation. To reduce temporal noise between post-training rollouts and demonstration data, we introduce asynchronous inference in the relative action space, which better aligns rollout data with demonstrated behavior and allows the critic to learn a value function induced by a more consistent underlying policy. Finally, DexPIE improves the policy through conditioning on a continuous optimality indicator, allowing the policy to leverage the quality of data in a more fine-grained manner. Across three challenging real-world dexterous manipulation tasks, DexPIE achieves a 37% improvement in success rate over the demonstration-based reference policy, outperforming all baseline methods and demonstrating stronger robustness. The source code and dataset will be made publicly available.
△ Less
Submitted 8 June, 2026;
originally announced June 2026.
-
A transition-density-based operator learning method for Fokker-Planck equations with various initial conditions
Authors:
Li Zeng,
Xiaoliang Wan,
Yaobin Wang,
Fabio Nobile,
Tao Zhou
Abstract:
Solving Fokker-Planck equations (FPEs) for multiple initial conditions typically requires repeated computations, leading to substantial computational costs. In this work, we propose a transition-density-based operator learning method to efficiently approximate the solution operator of FPEs with various initial conditions. The core idea is to learn the transition probability density function (PDF)…
▽ More
Solving Fokker-Planck equations (FPEs) for multiple initial conditions typically requires repeated computations, leading to substantial computational costs. In this work, we propose a transition-density-based operator learning method to efficiently approximate the solution operator of FPEs with various initial conditions. The core idea is to learn the transition probability density function (PDF) of the underlying stochastic differential equation (SDE), from which the solution associated with a new initial distribution can be obtained through the Chapman-Kolmogorov equation without retraining the model. A major challenge in learning the transition PDF lies in the singular behavior induced by the Dirac initial condition. To address it, we introduce a conditional normalizing flow whose base distribution is given by the explicit transition PDF of a linearized SDE. This base distribution captures the short-time behavior of the target transition PDF and allows the normalizing flow to learn a near-identity transformation at small times. We further incorporate a time-weighted loss function to stabilize training near the initial time and develop an importance-sampling strategy for evaluating solutions associated with general initial conditions. A variety of numerical experiments are presented to illustrate the effectiveness and robustness of the proposed method.
△ Less
Submitted 22 July, 2026; v1 submitted 8 June, 2026;
originally announced June 2026.
-
Reinforcement learning in linear embedding space unlocks generalizable control across soft robot configurations
Authors:
Xinglong Zhang,
Cong Li,
Hangjie Mo,
Yue Jiang,
Xin Xu,
Wei Jiang,
Zhenshan Bing,
Yihe Yang,
Xiaojian Li,
Yueneng Yang,
Huimin Lu,
Ling-li Zeng,
Alois Knoll,
Dewen Hu,
Li Wen,
Wei Pan
Abstract:
Soft-bodied organisms such as octopuses and elephant trunks exhibit remarkable morphological adaptability, dynamically reconfiguring body shape and stiffness, and flexibly adjusting their control strategies to enable versatile behaviors. Inspired by these biological systems, various soft robots have emerged in recent decades, featuring diverse materials, stiffnesses, and morphologies tailored to s…
▽ More
Soft-bodied organisms such as octopuses and elephant trunks exhibit remarkable morphological adaptability, dynamically reconfiguring body shape and stiffness, and flexibly adjusting their control strategies to enable versatile behaviors. Inspired by these biological systems, various soft robots have emerged in recent decades, featuring diverse materials, stiffnesses, and morphologies tailored to specific tasks. Despite substantial advances in the materials and structural designs of soft robots, developing a generalizable control framework capable of rapid adaptation across diverse configurations remains a long-standing challenge. Existing controllers are limited to fixed configurations, demanding laborious configuration-specific remodelling and policy redesign for new configurations. Here, we introduce a generalizable control system that enables rapid adaptation across diverse soft robot configurations via reinforcement learning in a shared linear Koopman embedding space. By encoding robot dynamics into this embedding space, our method decouples control policies from specific morphologies, allowing real-time, model-free policy adaptation across diverse configurations without retraining from scratch. We validate our system across 33 distinct robot configurations. Our system achieves a 75 times reduction in transfer samples across configurations, while sustaining robust performance under high-speed motion, heavy payloads, and multiactuator faults, and achieving real-world skills previously unattainable in soft robotics. This work establishes a unified and adaptable control paradigm for diverse soft robot configurations, bridging mechanical reconfigurability with control flexibility, and may offer broader insights for generalizable control in complex physical systems.
△ Less
Submitted 6 June, 2026;
originally announced June 2026.
-
Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning
Authors:
Yijin Zhou,
Linqian Zeng,
Xiaoya Lu,
Wenyuan Xie,
Dongrui Liu,
Junchi Yan,
Jing Shao
Abstract:
Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions. Existing approaches mainly improve these behaviors through inference-time correction or coarse-grained reward signals based on decision outcomes and structured checklists, leaving t…
▽ More
Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions. Existing approaches mainly improve these behaviors through inference-time correction or coarse-grained reward signals based on decision outcomes and structured checklists, leaving the uncertainty characteristics of agent decisions underexplored. We observe that decision-oriented reinforcement learning tends to weaken the uncertainty separation between correct and incorrect actions, resulting in overconfident mistakes and weaker exploration signals. Therefore, we propose TRUST, which incorporates uncertainty quantification into reward design as a repulsive force for maintaining uncertainty separation, and labels lightweight key-turn annotations for unified post-training of multi-turn trajectories. Experimental results across diverse tool-use benchmarks show that TRUST consistently enhances both decision quality and agent performance while maintaining more reliable uncertainty estimates during optimization.
△ Less
Submitted 5 June, 2026;
originally announced June 2026.
-
Agents' Last Exam
Authors:
Yiyou Sun,
Xinyang Han,
Weichen Zhang,
Yuanbo Pang,
Tianyu Wang,
Yuhan Cao,
Yixiao Huang,
Chris Duroiu,
Haoyun Zhang,
Jeffrey Lin,
Weishu Zhang,
Tyler Zeng,
Ying Yan,
Bo Liu,
Hanson Wen,
Mingyang Xu,
Xiaoyuan Liu,
Zimeng Chen,
Weiyan Shi,
Amanda Dsouza,
Vincent Sunn Chen,
Patrick Bryant,
Carl Boettiger,
Yamini Rangan,
Bradley Rothenberg
, et al. (285 additional authors not shown)
Abstract:
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a…
▽ More
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.
△ Less
Submitted 11 June, 2026; v1 submitted 3 June, 2026;
originally announced June 2026.
-
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
Authors:
Aili Chen,
Aonian Li,
Baichuan Zhou,
Bangwei Gong,
Binyang Jiang,
Boji Dan,
Changhao Zhang,
Changqing Yu,
Chao Wang,
Cheng Ma,
Cheng Zhong,
Cheng Zhu,
Chengjun Xiao,
Chengyi Yang,
Chengyu Du,
Chenyang Zhang,
Chi Zhang,
Chuangyi Huang,
Chunhao Zhang,
Chunhui Du,
Chunyu Zhao,
Congchao Guo,
Da Chen,
Deming Ding,
Dianjun Sun
, et al. (193 additional authors not shown)
Abstract:
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale…
▽ More
We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B activated per token. Designed end-to-end for agentic deployment, the M2 series rests on three components: (i) agent-driven data pipelines producing large-scale, verifiable trajectories across agentic coding and agentic cowork, each grounded in an executable workspace and an artifact-aligned reward; (ii) Forge, a scalable agent-native RL system that adapts to long-horizon agent trajectories, paired with windowed-FIFO scheduling, prefix-tree merging, inference optimization, and a clean training-inference-agent decoupling that supports both white-box and black-box agents; (iii) the latest M2.7 checkpoint takes an early step toward self-evolution -- autonomously debugging training runs and modifying its own scaffold. Across M2 through M2.7, this combination translates a mini-activation footprint into frontier-tier performance on agentic coding, deep search, office-task, and reasoning benchmarks.
△ Less
Submitted 30 July, 2026; v1 submitted 25 May, 2026;
originally announced May 2026.
-
Superconductivity in Al-based high-entropy alloys TiHfNbTaAl and TaNbHfZrAl
Authors:
Junjin Huang,
Wenbo Sun,
Longfu Li,
Shuangyue Wang,
Jingjun Qin,
Rui Chen,
Zaichen Xiang,
Yucheng Li,
Lingyong Zeng,
Huixia Luo
Abstract:
Since the first report of a high entropy alloy (HEA) superconductor in 2014, HEAs have continued to captivate the interest of superconducting researchers. Owing to the significant degree of disorder inherent in these systems, they serve as exemplary models for examining the properties of materials that exist in states intermediate between crystalline and amorphous structures. Here we present the s…
▽ More
Since the first report of a high entropy alloy (HEA) superconductor in 2014, HEAs have continued to captivate the interest of superconducting researchers. Owing to the significant degree of disorder inherent in these systems, they serve as exemplary models for examining the properties of materials that exist in states intermediate between crystalline and amorphous structures. Here we present the superconductivity properties and crystal structure of TaNbHfZrAl and TiHfNbTaAl HEAs, which both have the VEC of 4.2 and body centered cubic (BCC) structure. Through resistivity, magnetic, and specific heat measurements, we prove that both samples are the bulk type-II superconductors with a critical temperature Tc of 5.5 K for TaNbHfZrAl and Tc of 3.2 K for TiHfNbTaAl. The Tc of HEA superconductors is influenced by the VEC and the element composition. And the incorporation of Al in high disorder HEA superconductors causes a more crystallinelike Tc dependence.
△ Less
Submitted 24 May, 2026;
originally announced May 2026.
-
From Schema to Signal: Retrieval-Augmented Modeling for Relational Data Analytics
Authors:
Lingze Zeng,
Shaofeng Cai,
Changshuo Liu,
Zhongle Xie,
Yuncheng Wu,
Beng Chin Ooi
Abstract:
Relational data stored in RDBMS is foundational to many real-world applications across domains such as e-commerce, finance, and sociality. While deep neural networks (DNNs) have achieved strong performance on tabular data with a single table, extending these models to relational databases is challenging due to the normalized multi-table structure and complex inter-table relationships. Existing app…
▽ More
Relational data stored in RDBMS is foundational to many real-world applications across domains such as e-commerce, finance, and sociality. While deep neural networks (DNNs) have achieved strong performance on tabular data with a single table, extending these models to relational databases is challenging due to the normalized multi-table structure and complex inter-table relationships. Existing approaches often rely strictly on schema-defined graphs, which overlook implicit semantic signals embedded in tuple attributes and suffer from rigid connectivity.
In this work, we propose Retrieval-Augmented Modeling (RAM), a novel framework that combines graph structure with attribute semantics for relational data analytics. RAM treats tuple attributes as tokens and uses random walks to construct contextual documents, enabling the use of information retrieval techniques to estimate semantic relevance between tuples. Building on these documents, we introduce two retrieval-based augmentations: ATRA, which leverages intra-table relevance for contrastive learning, and ETRA, which links semantically related tuples across tables to enhance graph connectivity. Then, we propose a layer-wise model architecture tailored for relational data, which involves attribute embedding, feature integration, and graph aggregation layers to enable expressive and flexible representation learning. Extensive experiments on five real-world relational databases demonstrate that RAM consistently outperforms existing baselines in diverse prediction tasks, establishing a state-of-the-art for relational data analytics.
△ Less
Submitted 14 May, 2026;
originally announced May 2026.
-
How LLMs Are Persuaded: A Few Attention Heads, Rerouted
Authors:
Xiangkun Sun,
Lingkai Kong,
Aoqi Zhang,
Liang Zeng,
Tonghan Wang
Abstract:
Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly understood. We uncover a compact causal mechanism for persuasion-induced factual errors. A small set of mid-layer attention heads almost entirely determines the model's answer. These heads write answer options into a low-dimensional polyhedron, with o…
▽ More
Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly understood. We uncover a compact causal mechanism for persuasion-induced factual errors. A small set of mid-layer attention heads almost entirely determines the model's answer. These heads write answer options into a low-dimensional polyhedron, with options occupying distinct vertices. Persuasion does not blur belief or merely reduce confidence; it causes a discrete latent jump from the correct-answer vertex to the persuasion-target vertex. We show that decision heads are not reasoning over evidence. Instead, they copy whichever option token their attention selects. Persuasion works by redirecting attention. We isolate a rank-one evidence-routing feature that controls the route. Directly modifying this feature steers the model's choice, and removing it blocks persuasion. We then trace the feature back to a band of shallower attention heads that build it from persuasive keywords in the input. Every step is validated by intervention. This mechanism appears across open-source LLMs and realistic poisoning scenarios such as Generative Engine Optimization, revealing persuasion as a narrow, monitorable circuit.
△ Less
Submitted 10 May, 2026;
originally announced May 2026.
-
AssemPlanner: A Multi-Agent Based Task Planning Framework for Flexible Assembly System
Authors:
Chenhao Zhang,
Chaoran Zhang,
Zhaobo Xu,
Yongbo Yang,
Pingfa Feng,
Long Zeng
Abstract:
In flexible assembly systems, existing task planning methods require a time-consuming configuration process by multiple experts to establish a production line for a new product. To address this challenge, we propose a multi-agent based task planning framework for flexible assembly systems, denoted as AssemPlanner. It takes tasks described in natural language as input, which are then converted into…
▽ More
In flexible assembly systems, existing task planning methods require a time-consuming configuration process by multiple experts to establish a production line for a new product. To address this challenge, we propose a multi-agent based task planning framework for flexible assembly systems, denoted as AssemPlanner. It takes tasks described in natural language as input, which are then converted into actionable sequential production operations. It comprises several specialized agents, including SchedAgent , KnowledgeAgent, LineBalanceAgent, and a scene graph. Within the proposed framework, SchedAgent serves as the central reasoning engine. Departing from traditional static pipelines, AssemPlanner utilizes a ReAct-based SchedAgent to adaptively adjust actions via multi-agent feedback. By observing the feedback from KnowledgeAgent, LineBalanceAgent, and the scene graph, it autonomously resolves complex industrial process constraints. To facilitate reproducibility, all code and datasets are released at https://github.com/chz332/Assemplanner.
△ Less
Submitted 9 May, 2026;
originally announced May 2026.
-
LLM Advertisement based on Neuron Auctions
Authors:
Peiran Yun,
Wenxin Xu,
Jiayuan Liu,
Yihang Zhang,
Liang Zeng,
Lingkai Kong,
Tonghan Wang
Abstract:
As Large Language Models (LLMs) transition into conversational agents, generative advertising emerges as a crucial monetization strategy. However, embedding advertisements within unstructured LLM outputs introduces a critical trilemma: balancing advertiser payoffs, platform revenue, and user experience. Existing methods, such as prompt injection or rigid position slots, disrupt semantic coherence…
▽ More
As Large Language Models (LLMs) transition into conversational agents, generative advertising emerges as a crucial monetization strategy. However, embedding advertisements within unstructured LLM outputs introduces a critical trilemma: balancing advertiser payoffs, platform revenue, and user experience. Existing methods, such as prompt injection or rigid position slots, disrupt semantic coherence and lack a parametric framework for independent control, rendering rigorous mechanism design intractable. To bridge this gap, we introduce Neuron Auctions, a novel paradigm that shifts the auction object from the surface text space to the LLM's internal representations. Leveraging mechanistic interpretability, we identify brand-specific feed-forward network (FFN) neurons and demonstrate that competing brands activate within approximately orthogonal subspaces. This near-perfect independence allows us to define continuous, disentangled intervention budgets (specifically, neuron counts and amplification factors) as auctionable commodities. Building on this computational carrier, we design a continuous menu-based auction mechanism that naturally guarantees strategy-proofness and optimizes revenue for the platform. By explicitly incorporating a user utility penalty into the platform's optimization objective, our framework dynamically prices out overly aggressive interventions. Extensive experiments demonstrate that Neuron Auctions effectively preserve natural discourse quality while achieving an optimal alignment between commercial incentives and user satisfaction.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
UniISP: A Unified ISP Framework for Both Human and Machine Vision
Authors:
Hanxi Li,
Yao Cheng,
Bo Zhang,
Li Zeng
Abstract:
Compared to RGB images, raw sensor data provides a richer representation of information, which is crucial for accurate recognition, particularly under challenging conditions such as low-light environments. The traditional Image Signal Processing (ISP) pipeline generates visually pleasing RGB images for human perception through a series of steps, but some of these operations may adversely impact th…
▽ More
Compared to RGB images, raw sensor data provides a richer representation of information, which is crucial for accurate recognition, particularly under challenging conditions such as low-light environments. The traditional Image Signal Processing (ISP) pipeline generates visually pleasing RGB images for human perception through a series of steps, but some of these operations may adversely impact the information integrity by introducing compression and loss. Furthermore, in computer vision tasks that directly utilize raw camera data, most existing methods integrate minimal ISP processing with downstream networks, yet the resulting images are often difficult to visualize or do not align with human aesthetic preferences. This paper proposes UniISP, a novel ISP framework designed to simultaneously meet the requirements of both human visual perception and computer vision applications. By incorporating a carefully designed Hybrid Attention Module (HAM) and employing supervised learning, the proposed method ensures that the generated images are visually appealing. Additionally, a Feature Adapter module is introduced to effectively propagate informative features from the ISP stage to subsequent downstream networks. Extensive experiments demonstrate that our approach achieves state-of-the-art performance across various scenarios and multiple datasets, proving its generalizability and effectiveness.
△ Less
Submitted 8 May, 2026;
originally announced May 2026.
-
Saliency-Aware Regularized Quantization Calibration for Large Language Models
Authors:
Yanlong Zhao,
Xiaoyuan Cheng,
Huihang Liu,
Baihua He,
Xinyu Zhang,
Harrison Bo Hua Zhu,
Wenlong Chen,
Li Zeng,
Zhuo Sun
Abstract:
Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most existing PTQ methods determine quantization parameters by minimizing a layer-wise reconstruction error on a predetermined calibration dataset, typically optimized via either scale search or Gram-based methods. However, from the perspective of generalizatio…
▽ More
Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most existing PTQ methods determine quantization parameters by minimizing a layer-wise reconstruction error on a predetermined calibration dataset, typically optimized via either scale search or Gram-based methods. However, from the perspective of generalization risk, existing PTQ calibration objectives based solely on empirical reconstruction error over limited or unrepresentative calibration data may move the quantized weights away from the original floating-point weights, potentially degrading downstream performance. To address this issue, we propose \emph{Regularized Quantization Calibration} (RQC), a unified framework that augments standard PTQ objectives with a regularizer that explicitly controls weight deviation from the original weights. We further generalize this framework to incorporate a saliency-aware regularizer, resulting in \emph{Saliency-Aware Regularized Quantization Calibration} (SARQC). The proposed regularization encourages quantized weights to remain close to the original weights during calibration, leading to improved generalization at inference time. SARQC integrates seamlessly into existing PTQ pipelines and enhances both scale-search-based and Gram-based methods under a unified formulation. Extensive experiments on dense and Mixture-of-Experts LLMs demonstrate consistent improvements in perplexity and zero-shot accuracy, without introducing additional inference overhead.
△ Less
Submitted 8 May, 2026; v1 submitted 7 May, 2026;
originally announced May 2026.
-
Dual-branch Robust Unlearnable Examples
Authors:
Xianlong Wang,
Hangtao Zhang,
Wenbo Pan,
Ziqi Zhou,
Changsong Jiang,
Li Zeng,
Xiaohua Jia
Abstract:
Unlearnable examples (UEs) aim to compromise model training by injecting imperceptible perturbations to clean samples. However, existing UE schemes exhibit limited robustness against advanced defenses due to their heuristic design or narrowly scoped domain perturbations. To address this, we propose \texttt{DUNE}, a \underline{\textbf{D}}ual-branch \underline{\textbf{UN}}learnable \underline{\textb…
▽ More
Unlearnable examples (UEs) aim to compromise model training by injecting imperceptible perturbations to clean samples. However, existing UE schemes exhibit limited robustness against advanced defenses due to their heuristic design or narrowly scoped domain perturbations. To address this, we propose \texttt{DUNE}, a \underline{\textbf{D}}ual-branch \underline{\textbf{UN}}learnable \underline{\textbf{E}}nsemble perturbation optimization approach. Specifically, \texttt{DUNE} separately optimizes perturbations in the spatial and color domains to establish the mapping between perturbations and shift-induced labels. This design extends the perturbation domain to increase noise intensity for improving robustness and drives the models to learn perturbation-oriented features with degraded generalization, thereby achieving unlearnability. To strengthen \texttt{DUNE}'s performance, we further propose an unlearnability-enhancing ensemble strategy that aggregates diverse pre-trained models during the dual-branch optimization. Extensive experiments on benchmark datasets CIFAR-10 and ImageNet verify that \texttt{DUNE}'s robustness outperforms 12 SOTA UE schemes under 7 mainstream defenses, yielding a lower average test accuracy of 14.95% to 50.82%.
△ Less
Submitted 25 June, 2026; v1 submitted 3 May, 2026;
originally announced May 2026.
-
Analysis and Compensation of Tx and Rx IQ Imbalances in AFDM System
Authors:
Hongjun Liu,
Liaoyuan Zeng,
Junhao Tian,
Qingyu Li,
Fuchen Xu,
Chengxiang Liu,
Guanghui Liu
Abstract:
Affine frequency division multiplexing (AFDM) is a recently proposed multicarrier waveform whose bit error rate (BER) performance in doubly selective channels is comparable to that of orthogonal time-frequency space (OTFS) and superior to that of orthogonal frequency division multiplexing (OFDM). In this paper, the impacts of joint transmitter (Tx) and receiver (Rx) in-phase and quadrature imbalan…
▽ More
Affine frequency division multiplexing (AFDM) is a recently proposed multicarrier waveform whose bit error rate (BER) performance in doubly selective channels is comparable to that of orthogonal time-frequency space (OTFS) and superior to that of orthogonal frequency division multiplexing (OFDM). In this paper, the impacts of joint transmitter (Tx) and receiver (Rx) in-phase and quadrature imbalance (IQI) on AFDM signals are investigated, where we show that AFDM suffers more severe IQI than OFDM and OTFS due to the inherent feature of complicated chirp-assisted modulation. We further derive analytical expressions for the pairwise and average bit error probability as a function of the IQI parameters. These indicate that such distortions significantly limit the achievable operating signal-to-noise ratio at the receiver side and data rates. To this end, we propose a cascade compensation scheme to mitigate these effects. Specifically, we first compensate for Rx IQI to convert the improper Gaussian noise into additive white Gaussian noise, and then apply a judicious design to eliminate the Tx IQI. Both analytical and simulation results reveal that joint Tx and Rx IQI introduce an error floor in the BER performance of AFDM systems, whereas the proposed approach effectively compensates such impairments.
△ Less
Submitted 2 May, 2026;
originally announced May 2026.
-
PhysiGen: Integrating Collision-Aware Physical Constraints for High-Fidelity Human-Human Interaction Generation
Authors:
Nan Lei,
Yuan-Ming Li,
Ling-An Zeng,
Liang Xu,
Zhi-Wei Xia,
Hui-Wen Huang,
Fa-Ting Hong,
Wei-Shi Zheng
Abstract:
Despite substantial progress in text-driven 3D human motion synthesis, generating realistic multi-person interaction sequences remains challenging. Notably, body inter-penetration is a pervasive issue from both data acquisition to the generated results, which significantly undermines the realism and usability. Previous generative models either ignored this issue or introduced computationally expen…
▽ More
Despite substantial progress in text-driven 3D human motion synthesis, generating realistic multi-person interaction sequences remains challenging. Notably, body inter-penetration is a pervasive issue from both data acquisition to the generated results, which significantly undermines the realism and usability. Previous generative models either ignored this issue or introduced computationally expensive mesh-level loss functions to alleviate inter-body collisions. In this paper, we propose a general-purpose and computationally efficient optimization strategy named PhysiGen to explicitly integrate collision-aware physical constraints for human-human interaction generation. Specifically, we simplify the high-resolution human body mesh into geometric primitives to greatly reduce the cost of inter-person collision detection. Moreover, we identify the collision regions as the guidance of the optimization directions. PhysiGen is plug-and-play and can be readily integrated into existing human interaction generation models. Extensive cross-dataset and cross-model experiments show that our method can effectively reduce interpenetration and significantly improve visual coherence and physical plausibility compared to the state-of-the-art methods.
△ Less
Submitted 1 May, 2026;
originally announced May 2026.
-
$\text{PKS}^4$:Parallel Kinematic Selective State Space Scanners for Efficient Video Understanding
Authors:
Lingjie Zeng,
Hailun Zhang,
Xiwen Wang,
Qijun Zhao
Abstract:
Temporal modeling remains a fundamental challenge in video understanding, particularly as sequence lengths scale. Traditional video models relying on dense spatiotemporal attention suffer from quadratic computational costs for long videos. To circumvent these costs, recent approaches adapt image models for videos via Parameter-Efficient Fine-Tuning (PEFT) methods such as adapters. However, deeply…
▽ More
Temporal modeling remains a fundamental challenge in video understanding, particularly as sequence lengths scale. Traditional video models relying on dense spatiotemporal attention suffer from quadratic computational costs for long videos. To circumvent these costs, recent approaches adapt image models for videos via Parameter-Efficient Fine-Tuning (PEFT) methods such as adapters. However, deeply inserting these modules incurs prohibitive activation memory overhead during back-propagation. While recent efficient State Space Models (SSMs) introduce linear complexity, they disrupt 2D spatial relationships and rely on extensive masked pre-training to recover spatial awareness.
To overcome these limitations, we propose Parallel Kinematic Selective State Space Scanners (PKS$^4$). We retain a standard 2D vision backbone for spatial semantics and insert a single plug-and-play PKS$^4$ module with linear-complexity temporal scanning, avoiding temporal attention and multi-layer adapters. We first extract kinematic priors via a Kinematic Prior Encoder, which captures local displacements and motion boundaries through inter-frame correlations and differences. These priors drive linear-complexity SSMs to track underlying kinematic states, adaptively modulating update speeds and read-write strategies at each time step.
Instead of global scanning, we deploy parallel scanners along the temporal dimension for each spatial location, preserving spatial structures while reducing overhead. Experiments on spatial-heavy and temporal-heavy action recognition benchmarks show that PKS$^4$ achieves state-of-the-art performance. Remarkably, our method converges in merely $20$ epochs, achieving approximately $10\times$ lower training compute than pure video SSMs, establishing a new paradigm for efficient video understanding.
△ Less
Submitted 29 April, 2026;
originally announced April 2026.
-
MotionHiFlow: Text-to-motion via hierarchical flow matching
Authors:
Heng Li,
Xiaotong Lin,
Ling-An Zeng,
Yulei Kang,
Shuai Li,
Jian-Fang Hu
Abstract:
Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which limits both semantic alignment and temporal coherence. Inspired by the fact that complex motions are…
▽ More
Text-to-motion generation aims to generate 3D human motions that are tightly aligned with the input text while remaining physically plausible and rich in fine-grained detail. Although recent approaches can produce complex and natural movements, they usually operate at only one temporal scale, which limits both semantic alignment and temporal coherence. Inspired by the fact that complex motions are conceptualized hierarchically rather than at a single temporal scale in the human cognitive system, we propose \textit{MotionHiFlow}, a hierarchical flow matching framework to generate motion progressively by constructing flow path from low to high temporal scales. The flows at lower scales capture high-level semantics and coarse motion structures, while flows at higher scales refine temporal details. To link the flows across scales, we introduce a novel cross-scale transition process, ensuring continuity and preserving noise consistency. Furthermore, by integrating a Text-Motion Diffusion Transformer and a topology-aware Motion VAE, MotionHiFlow explicitly models structural dependencies among joints via joint-aware positional encoding and skeletal topology, enabling precise semantic alignment alongside fine-grained motion details. Extensive experiments on HumanML3D and KIT-ML benchmarks demonstrate state-of-the-art performance, with ablation studies confirming the effectiveness of the hierarchical design and key components. Code is available at https://github.com/ai-lh/MotionHiFlow.
△ Less
Submitted 25 April, 2026;
originally announced April 2026.