Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 156 results for author: Kawaguchi, K

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.02625  [pdf, ps, other

    cs.CL cs.AI

    Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models

    Authors: Brian K Chen, Chong Wu, Kenji Kawaguchi

    Abstract: Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right generation by producing text block by block. We study a simple plug-and-play inference pattern: first generate a complete draft, then refine the full response using bidirectional diffusion. Using LLaDA2.1-Flash and LLaDA2.1-Mini, we evaluate two configurations. In… ▽ More

    Submitted 21 July, 2026; originally announced August 2026.

    Comments: 25 pages, 3 figures

  2. arXiv:2607.28446  [pdf, ps, other

    cs.CE

    CoLAS: Multimodal Corroboration of Latent Asset Signals for Financial Trading

    Authors: Yanzheng Jin, Pengyang Shao, Xiaohao Liu, Xi Ai, Fei Shen, Kenji Kawaguchi

    Abstract: Financial trading relies on extracting reliable signals from heterogeneous market modalities such as price series, breaking news, and investor sentiment. Existing multimodal methods primarily combine heterogeneous modalities to exploit complementarity, treating each modality as equally valuable while overlooking whether different modalities provide mutually supportive evidence for the same trading… ▽ More

    Submitted 30 July, 2026; originally announced July 2026.

  3. arXiv:2607.13841  [pdf, ps, other

    cs.LG stat.ML

    Heavy-Tailed Flow Matching via Random Clocks

    Authors: Zhouhao Yang, Yezhen Wang, Kenji Kawaguchi, Vladimir Braverman, Haoyang Cao

    Abstract: Heavy-tailed data arise in many domains where rare events carry disproportionate importance, such as imbalanced image datasets, financial returns, and weather extremes. Standard diffusion and flow-matching models typically begin from Gaussian noise or Gaussian source distributions, which yield tractable training targets but provide a poor inductive match for heavy-tailed data. We propose Heavy-Tai… ▽ More

    Submitted 15 July, 2026; originally announced July 2026.

  4. arXiv:2606.02453  [pdf, ps, other

    cs.CV cs.AI

    Initialization is Half the Battle: Generating Diverse Images from a Guidance Potential Posterior

    Authors: Xiang Li, Dianbo Liu, Kenji Kawaguchi

    Abstract: Despite the remarkable fidelity of generative models, they frequently suffer from mode collapse. Existing strategies for enhancing diversity predominantly focus on intervening during the generation trajectory. We identify a critical oversight that the standard Gaussian initialization often causes trajectories to collapse into dominant modes because it is agnostic to the guidance potential landscap… ▽ More

    Submitted 1 June, 2026; originally announced June 2026.

    Comments: Accepted by ICML 2026 Spotlight

  5. arXiv:2604.11399  [pdf, ps, other

    cs.CV cs.CL

    Lost in Adaptation: Layer-Selective Recovery of Temporal Reasoning in Video-Language Models

    Authors: Zihang Fu, Haonan Wang, Jian Kang, Kenji Kawaguchi, Jiaying Wu

    Abstract: Multimodal adaptation can erode temporal reasoning (TR) in video-language models (VLMs), leaving models able to perceive salient events yet unable to infer their temporal and causal structure. We introduce MERIT, a gradient-free framework that repairs this capability through layer-selective model merging. MERIT assigns each self-attention layer a VLM-dominant or LLM-dominant interpolation and uses… ▽ More

    Submitted 16 August, 2026; v1 submitted 13 April, 2026; originally announced April 2026.

  6. arXiv:2603.08068  [pdf, ps, other

    cs.AI

    In-Context Reinforcement Learning for Tool Use in Large Language Models

    Authors: Yaoqi Ye, Yiran Zhao, Keyu Duan, Zeyu Zheng, Kenji Kawaguchi, Cihang Xie, Michael Qizhe Shieh

    Abstract: While large language models (LLMs) exhibit strong reasoning abilities, their performance on complex tasks is often constrained by the limitations of their internal knowledge. A compelling approach to overcome this challenge is to augment these models with external tools -- such as Python interpreters for mathematical computations or search engines for retrieving factual information. However, enabl… ▽ More

    Submitted 9 March, 2026; originally announced March 2026.

  7. arXiv:2602.22583  [pdf, ps, other

    cs.AI cs.CL

    Strategy Executability in Mathematical Reasoning: Leveraging Human-Model Differences for Effective Guidance

    Authors: Weida Liang, Yiyou Sun, Shuyuan Nan, Chuang Li, Dawn Song, Kenji Kawaguchi

    Abstract: Example-based guidance is widely used to improve mathematical reasoning at inference time, yet its effectiveness is highly unstable across problems and models-even when the guidance is correct and problem-relevant. We show that this instability arises from a previously underexplored gap between strategy usage-whether a reasoning strategy appears in successful solutions-and strategy executability-w… ▽ More

    Submitted 25 February, 2026; originally announced February 2026.

  8. arXiv:2602.04919  [pdf, ps, other

    cs.LG

    Gradually Compacting Large Language Models for Reasoning Like a Boiling Frog

    Authors: Yiran Zhao, Shengyang Zhou, Zijian Wu, Tongyan Hu, Yuhui Xu, Rengan Dou, Kenji Kawaguchi, Shafiq Joty, Junnan Li, Michael Qizhe Shieh

    Abstract: Large Language Models (LLMs) have demonstrated impressive reasoning capabilities, but their substantial size often demands significant computational resources. To reduce resource consumption and accelerate inference, it is essential to eliminate redundant parameters without compromising performance. However, conventional pruning methods that directly remove such parameters often lead to a dramatic… ▽ More

    Submitted 4 February, 2026; originally announced February 2026.

  9. arXiv:2601.20198  [pdf, ps, other

    cs.LG

    DeRaDiff: Denoising Time Realignment of Diffusion Models

    Authors: Ratnavibusena Don Shahain Manujith, Teoh Tze Tzun, Kenji Kawaguchi, Yang Zhang

    Abstract: Recent advances align diffusion models with human preferences to increase aesthetic appeal and mitigate artifacts and biases. Such methods aim to maximize a conditional output distribution aligned with higher rewards whilst not drifting far from a pretrained prior. This is commonly enforced by KL (Kullback Leibler) regularization. As such, a central issue still remains: how does one choose the rig… ▽ More

    Submitted 20 February, 2026; v1 submitted 27 January, 2026; originally announced January 2026.

  10. arXiv:2512.09687  [pdf, ps, other

    cs.CV

    Unconsciously Forget: Mitigating Memorization; Without Knowing What is being Memorized

    Authors: Er Jin, Yang Zhang, Yongli Mou, Yanfei Dong, Stefan Decker, Kenji Kawaguchi, Johannes Stegmaier

    Abstract: Recent advances in generative models have demonstrated an exceptional ability to produce highly realistic images. However, previous studies show that generated images often resemble the training data, and this problem becomes more severe as the model size increases. Memorizing training data can lead to legal challenges, including copyright infringement, violations of portrait rights, and trademark… ▽ More

    Submitted 12 December, 2025; v1 submitted 10 December, 2025; originally announced December 2025.

  11. arXiv:2512.02874  [pdf, ps, other

    cs.CL

    Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning

    Authors: Haonan Wang, Chao Du, Kenji Kawaguchi, Tianyu Pang

    Abstract: Majority voting has proven effective for close-ended question answering by aggregating parallel reasoning traces. However, it is not directly applicable to open-ended reasoning, such as code generation and web-based deep research, where a "majority" over complete solutions is ill-defined. We introduce ThinkMerge, a training-free, plug-and-play decoding strategy that runs K parallel reasoning trace… ▽ More

    Submitted 2 December, 2025; originally announced December 2025.

  12. arXiv:2510.00480  [pdf, ps, other

    cs.AI

    Expandable Decision-Making States for Multi-Agent Deep Reinforcement Learning in Soccer Tactical Analysis

    Authors: Kenjiro Ide, Taiga Someya, Kohei Kawaguchi, Keisuke Fujii

    Abstract: Invasion team sports such as soccer produce a high-dimensional, strongly coupled state space as many players continuously interact on a shared field, challenging quantitative tactical analysis. Traditional rule-based analyses are intuitive, while modern predictive machine learning models often perform pattern-matching without explicit agent representations. The problem we address is how to build p… ▽ More

    Submitted 26 October, 2025; v1 submitted 1 October, 2025; originally announced October 2025.

    Comments: 28 pages, 9 figures

  13. arXiv:2509.26404  [pdf, ps, other

    cs.CR cs.AI cs.CL

    SeedPrints: Fingerprints Can Even Tell Which Seed Your Large Language Model Was Trained From

    Authors: Yao Tong, Haonan Wang, Siquan Li, Kenji Kawaguchi, Tianyang Hu

    Abstract: Fingerprinting Large Language Models (LLMs)is essential for provenance verification and model attribution. Existing fingerprinting methods are primarily evaluated after fine-tuning, where models have already acquired stable signatures from training data, optimization dynamics, or hyperparameters. However, most of a model's capacity and knowledge are acquired during pretraining rather than downstre… ▽ More

    Submitted 14 April, 2026; v1 submitted 30 September, 2025; originally announced September 2025.

    Comments: Accepted to ICLR 2026. The code repository linked on OpenReview is outdated; the latest code is available via the final arXiv version

  14. arXiv:2509.23196  [pdf, ps, other

    cs.CL

    From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs

    Authors: Haonan Wang, Weida Liang, Zihang Fu, Nie Zheng, Yifan Zhang, Yao Tong, Tongyao Zhu, Hao Jiang, Chuang Li, Jiaying Wu, Kenji Kawaguchi

    Abstract: Recent reasoning LLMs (RLMs), especially those trained with verifier-based reinforcement learning, often perform worse with few-shot CoT than with direct answering. We revisit this paradox using high-quality reasoning traces from DeepSeek-R1 as demonstrations and find that adding more exemplars consistently degrades accuracy, even when demonstrations are optimal. A detailed analysis reveals two me… ▽ More

    Submitted 27 September, 2025; originally announced September 2025.

  15. arXiv:2507.22499  [pdf, ps, other

    cs.LG cs.AI

    LoReUn: Data Itself Implicitly Provides Cues to Improve Machine Unlearning

    Authors: Xiang Li, Qianli Shen, Haonan Wang, Kenji Kawaguchi

    Abstract: Recent generative models face significant risks of producing harmful content, which has underscored the importance of machine unlearning (MU) as a critical technique for eliminating the influence of undesired data. However, existing MU methods typically assign the same weight to all data to be forgotten, which makes it difficult to effectively forget certain data that is harder to unlearn than oth… ▽ More

    Submitted 30 July, 2025; originally announced July 2025.

    Comments: 23 pages

  16. arXiv:2507.21903  [pdf, ps, other

    cs.SI cs.CL cs.IR

    Who's important? -- SUnSET: Synergistic Understanding of Stakeholder, Events and Time for Timeline Generation

    Authors: Tiviatis Sim, Kaiwen Yang, Shen Xin, Kenji Kawaguchi

    Abstract: As news reporting becomes increasingly global and decentralized online, tracking related events across multiple sources presents significant challenges. Existing news summarization methods typically utilizes Large Language Models and Graphical methods on article-based summaries. However, this is not effective since it only considers the textual content of similarly dated articles to understand the… ▽ More

    Submitted 17 March, 2026; v1 submitted 29 July, 2025; originally announced July 2025.

  17. arXiv:2507.15219  [pdf, ps, other

    cs.CR cs.AI

    PromptArmor: Simple yet Effective Prompt Injection Defenses

    Authors: Tianneng Shi, Kaijie Zhu, Zhun Wang, Yuqi Jia, Will Cai, Weida Liang, Haonan Wang, Hend Alzahrani, Joshua Lu, Kenji Kawaguchi, Basel Alomair, Xuandong Zhao, William Yang Wang, Neil Gong, Wenbo Guo, Dawn Song

    Abstract: Despite their potential, recent research has demonstrated that LLM agents are vulnerable to prompt injection attacks, where malicious prompts are injected into the agent's input, causing it to perform an attacker-specified task rather than the intended task provided by the user. In this paper, we present PromptArmor, a simple yet effective defense against prompt injection attacks. Specifically, Pr… ▽ More

    Submitted 20 July, 2025; originally announced July 2025.

  18. arXiv:2507.13386  [pdf, ps, other

    cs.CV cs.LG

    Minimalist Concept Erasure in Generative Models

    Authors: Yang Zhang, Er Jin, Yanfei Dong, Yixuan Wu, Philip Torr, Ashkan Khakzar, Johannes Stegmaier, Kenji Kawaguchi

    Abstract: Recent advances in generative models have demonstrated remarkable capabilities in producing high-quality images, but their reliance on large-scale unlabeled data has raised significant safety and copyright concerns. Efforts to address these issues by erasing unwanted concepts have shown promise. However, many existing erasure methods involve excessive modifications that compromise the overall util… ▽ More

    Submitted 15 July, 2025; originally announced July 2025.

    Comments: ICML2025

  19. arXiv:2506.19935  [pdf, ps, other

    cs.LG cs.CV stat.ML

    Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture

    Authors: Shuchen Xue, Tianyu Xie, Tianyang Hu, Zijin Feng, Jiacheng Sun, Kenji Kawaguchi, Zhenguo Li, Zhi-Ming Ma

    Abstract: Large language models (LLMs) predominantly use autoregressive (AR) approaches, but masked diffusion models (MDMs) are emerging as viable alternatives. A key challenge in comparing AR and MDM paradigms is their typical architectural difference: AR models are often decoder-only, while MDMs have largely been encoder-only. This practice of changing both the modeling paradigm and architecture simultane… ▽ More

    Submitted 24 June, 2025; originally announced June 2025.

  20. arXiv:2506.16696  [pdf, ps, other

    cs.AI

    Interpretable Low-Dimensional Modeling of Spatiotemporal Agent States for Decision Making in Football Tactics

    Authors: Kenjiro Ide, Taiga Someya, Kohei Kawaguchi, Keisuke Fujii

    Abstract: Understanding football tactics is crucial for managers and analysts. Previous research has proposed models based on spatial and kinematic equations, but these are computationally expensive. Also, Reinforcement learning approaches use player positions and velocities but lack interpretability and require large datasets. Rule-based models align with expert knowledge but have not fully considered all… ▽ More

    Submitted 19 June, 2025; originally announced June 2025.

    Comments: 5 pages, 3 figures, presented in iCSports 2024 Abstract Track

  21. arXiv:2506.13674  [pdf, ps, other

    cs.CL cs.AI

    PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention

    Authors: Haonan Wang, Brian Chen, Siquan Li, Xinhe Liang, Hwee Kuan Lee, Kenji Kawaguchi, Tianyang Hu

    Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tuning, an early and effective PEFT technique, demonstrated the ability to achieve performance comparable to full fine-tuning with significantly reduced computational and memory overhead. However, despite its earlier success, its effectiveness in training… ▽ More

    Submitted 18 April, 2026; v1 submitted 16 June, 2025; originally announced June 2025.

    Comments: ICLR 2026

  22. arXiv:2506.09890  [pdf, ps, other

    cs.CL cs.AI

    The Emergence of Abstract Thought in Large Language Models Beyond Any Language

    Authors: Yuxin Chen, Yiran Zhao, Yang Zhang, An Zhang, Kenji Kawaguchi, Shafiq Joty, Junnan Li, Tat-Seng Chua, Michael Qizhe Shieh, Wenxuan Zhang

    Abstract: As large language models (LLMs) continue to advance, their capacity to function effectively across a diverse range of languages has shown marked improvement. Preliminary studies observe that the hidden activations of LLMs often resemble English, even when responding to non-English prompts. This has led to the widespread assumption that LLMs may "think" in English. However, more recent results show… ▽ More

    Submitted 11 June, 2025; originally announced June 2025.

  23. arXiv:2506.08618  [pdf, ps, other

    cs.LG cond-mat.mes-hall cond-mat.other cs.AI cs.CV

    HSG-12M: A Large-Scale Benchmark of Spatial Multigraphs from the Energy Spectra of Non-Hermitian Crystals

    Authors: Xianquan Yan, Hakan Akgün, Kenji Kawaguchi, N. Duane Loh, Ching Hua Lee

    Abstract: AI is transforming scientific research by revealing new ways to understand complex physical systems, but its impact remains constrained by the lack of large, high-quality domain-specific datasets. A rich, largely untapped resource lies in non-Hermitian quantum physics, where the energy spectra of crystals form intricate geometries on the complex plane -- termed as Hamiltonian spectral graphs. Desp… ▽ More

    Submitted 17 May, 2026; v1 submitted 10 June, 2025; originally announced June 2025.

    Comments: Accepted to ICLR 2026, OpenReview: [https://openreview.net/forum?id=YxuKCME576]. 49 pages, 13 figures, 14 tables. Code & pipeline: [https://github.com/sarinstein-yan/Poly2Graph] Dataset: [https://github.com/sarinstein-yan/HSG-12M] Dataset released under CC BY 4.0. The Fourteenth International Conference on Learning Representations (ICLR 2026)

    Journal ref: The Fourteenth International Conference on Learning Representations (ICLR 2026)

  24. arXiv:2506.06950  [pdf, ps, other

    cs.CL

    What Makes a Good Natural Language Prompt?

    Authors: Do Xuan Long, Duy Dinh, Ngoc-Hai Nguyen, Kenji Kawaguchi, Nancy F. Chen, Shafiq Joty, Min-Yen Kan

    Abstract: As large language models (LLMs) have progressed towards more human-like and human--AI communications have become prevalent, prompting has emerged as a decisive component. However, there is limited conceptual consensus on what exactly quantifies natural language prompts. We attempt to address this question by conducting a meta-analysis surveying more than 150 prompting-related papers from leading N… ▽ More

    Submitted 7 June, 2025; originally announced June 2025.

    Comments: ACL 2025 Main Conference

  25. arXiv:2506.02561  [pdf, ps, other

    cs.CL cs.AI

    Pruning General Large Language Models into Customized Expert Models

    Authors: Yirao Zhao, Guizhen Chen, Kenji Kawaguchi, Lidong Bing, Wenxuan Zhang

    Abstract: Large language models (LLMs) have revolutionized natural language processing, yet their substantial model sizes often require substantial computational resources. To preserve computing resources and accelerate inference speed, it is crucial to prune redundant parameters, especially for experienced users who often need compact expert models tailored to specific downstream scenarios. However, most e… ▽ More

    Submitted 3 June, 2025; originally announced June 2025.

  26. arXiv:2506.01598  [pdf, ps, other

    cs.LG physics.comp-ph

    PMNO: A novel physics guided multi-step neural operator predictor for partial differential equations

    Authors: Jin Song, Kenji Kawaguchi, Zhenya Yan

    Abstract: Neural operators, which aim to approximate mappings between infinite-dimensional function spaces, have been widely applied in the simulation and prediction of physical systems. However, the limited representational capacity of network architectures, combined with their heavy reliance on large-scale data, often hinder effective training and result in poor extrapolation performance. In this paper, i… ▽ More

    Submitted 2 June, 2025; originally announced June 2025.

    Comments: 27 pages, 12 figures

  27. arXiv:2506.01265  [pdf, ps, other

    cs.CL

    Beyond In-Context Learning: Aligning Long-form Generation of Large Language Models via Task-Inherent Attribute Guidelines

    Authors: Do Xuan Long, Duong Ngoc Yen, Do Xuan Trong, Luu Anh Tuan, Kenji Kawaguchi, Shafiq Joty, Min-Yen Kan, Nancy F. Chen

    Abstract: In-context learning (ICL) is an important yet not fully understood ability of pre-trained large language models (LLMs). It can greatly enhance task performance using a few examples, termed demonstrations, without fine-tuning. Although effective in question answering, ICL often underperforms in long-form generation tasks such as summarization. Under appropriately realistic assumptions, we empirical… ▽ More

    Submitted 1 June, 2025; originally announced June 2025.

    Comments: ACL 2025 Findings

  28. arXiv:2505.22457  [pdf, ps, other

    cs.CV cs.AI cs.CL

    Fostering Video Reasoning via Next-Event Prediction

    Authors: Haonan Wang, Hongfu Liu, Xiangyan Liu, Chao Du, Kenji Kawaguchi, Ye Wang, Tianyu Pang

    Abstract: Next-token prediction serves as the foundational learning task enabling reasoning in LLMs. But what should the learning task be when aiming to equip MLLMs with temporal reasoning capabilities over video inputs? Existing tasks such as video question answering often rely on annotations from humans or much stronger MLLMs, while video captioning tends to entangle temporal reasoning with spatial inform… ▽ More

    Submitted 28 May, 2025; originally announced May 2025.

  29. arXiv:2505.01744  [pdf, ps, other

    cs.LG

    Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients

    Authors: Yezhen Wang, Zhouhao Yang, Brian K Chen, Fanyi Pu, Bo Li, Tianyu Gao, Kenji Kawaguchi

    Abstract: Building upon the success of low-rank adapter (LoRA), low-rank gradient projection (LoRP) has emerged as a promising solution for memory-efficient fine-tuning. However, existing LoRP methods typically treat each row of the gradient matrix as the default projection unit, leaving the role of projection granularity underexplored. In this work, we propose a novel framework, VLoRP, that extends low-ran… ▽ More

    Submitted 3 May, 2025; originally announced May 2025.

  30. arXiv:2503.15567  [pdf, ps, other

    cs.LG

    Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling

    Authors: Yanchen Luo, Zhiyuan Liu, Yi Zhao, Sihang Li, Hengxing Cai, Kenji Kawaguchi, Tat-Seng Chua, Yang Zhang, Xiang Wang

    Abstract: 3D molecule generation is crucial for drug discovery and material science, requiring models to process complex multi-modalities, including atom types, chemical bonds, and 3D coordinates. A key challenge is integrating these modalities of different shapes while maintaining SE(3) equivariance for 3D coordinates. To achieve this, existing approaches typically maintain separate latent spaces for invar… ▽ More

    Submitted 13 October, 2025; v1 submitted 19 March, 2025; originally announced March 2025.

    Comments: NeurIPS 2025

  31. arXiv:2503.13070  [pdf, ps, other

    cs.CV

    Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation

    Authors: Yihong Luo, Tianyang Hu, Weijian Luo, Kenji Kawaguchi, Jing Tang

    Abstract: This paper addresses the challenge of achieving high-quality and fast image generation that aligns with complex human preferences. While recent advancements in diffusion models and distillation have enabled rapid generation, the effective integration of reward feedback for improved abilities like controllability and preference alignment remains a key open problem. Existing reward-guided post-train… ▽ More

    Submitted 8 June, 2025; v1 submitted 17 March, 2025; originally announced March 2025.

  32. arXiv:2503.01926  [pdf, ps, other

    cs.CL cs.AI

    Unnatural Languages Are Not Bugs but Features for LLMs

    Authors: Keyu Duan, Yiran Zhao, Zhili Feng, Jinjie Ni, Tianyu Pang, Qian Liu, Tianle Cai, Longxu Dou, Kenji Kawaguchi, Anirudh Goyal, J. Zico Kolter, Michael Qizhe Shieh

    Abstract: Large Language Models (LLMs) have been observed to process non-human-readable text sequences, such as jailbreak prompts, often viewed as a bug for aligned LLMs. In this work, we present a systematic investigation challenging this perception, demonstrating that unnatural languages - strings that appear incomprehensible to humans but maintain semantic meanings for LLMs - contain latent features usab… ▽ More

    Submitted 3 June, 2025; v1 submitted 2 March, 2025; originally announced March 2025.

  33. arXiv:2502.12638  [pdf, other

    q-bio.QM cs.LG q-bio.BM

    NExT-Mol: 3D Diffusion Meets 1D Language Modeling for 3D Molecule Generation

    Authors: Zhiyuan Liu, Yanchen Luo, Han Huang, Enzhi Zhang, Sihang Li, Junfeng Fang, Yaorui Shi, Xiang Wang, Kenji Kawaguchi, Tat-Seng Chua

    Abstract: 3D molecule generation is crucial for drug discovery and material design. While prior efforts focus on 3D diffusion models for their benefits in modeling continuous 3D conformers, they overlook the advantages of 1D SELFIES-based Language Models (LMs), which can generate 100% valid molecules and leverage the billion-scale 1D molecule datasets. To combine these advantages for 3D molecule generation,… ▽ More

    Submitted 26 February, 2025; v1 submitted 18 February, 2025; originally announced February 2025.

    Comments: ICLR 2025, 10 pages

  34. arXiv:2412.21149  [pdf, other

    cs.LG

    Functional Risk Minimization

    Authors: Ferran Alet, Clement Gehring, Tomás Lozano-Pérez, Kenji Kawaguchi, Joshua B. Tenenbaum, Leslie Pack Kaelbling

    Abstract: The field of Machine Learning has changed significantly since the 1970s. However, its most basic principle, Empirical Risk Minimization (ERM), remains unchanged. We propose Functional Risk Minimization~(FRM), a general framework where losses compare functions rather than outputs. This results in better performance in supervised, unsupervised, and RL experiments. In the FRM paradigm, for each data… ▽ More

    Submitted 30 December, 2024; originally announced December 2024.

  35. arXiv:2412.02852  [pdf, ps, other

    cs.CV

    Learnable Sparsity for Vision Generative Models

    Authors: Yang Zhang, Er Jin, Wenzhong Liang, Yanfei Dong, Ashkan Khakzar, Philip Torr, Johannes Stegmaier, Kenji Kawaguchi

    Abstract: Diffusion models have achieved impressive advancements in various vision tasks. However, these gains often rely on increasing model size, which escalates computational complexity and memory demands, complicating deployment, raising inference costs, and causing environmental impact. While some studies have explored pruning techniques to improve the memory efficiency of diffusion models, most existi… ▽ More

    Submitted 5 March, 2026; v1 submitted 3 December, 2024; originally announced December 2024.

    Comments: Project page: https://yangzhang-v5.github.io/EcoDiff

  36. arXiv:2412.00088  [pdf, other

    cs.LG

    Stochastic Taylor Derivative Estimator: Efficient amortization for arbitrary differential operators

    Authors: Zekun Shi, Zheyuan Hu, Min Lin, Kenji Kawaguchi

    Abstract: Optimizing neural networks with loss that contain high-dimensional and high-order differential operators is expensive to evaluate with back-propagation due to $\mathcal{O}(d^{k})$ scaling of the derivative tensor size and the $\mathcal{O}(2^{k-1}L)$ scaling in the computation graph, where $d$ is the dimension of the domain, $L$ is the number of ops in the forward computation graph, and $k$ is the… ▽ More

    Submitted 12 January, 2025; v1 submitted 27 November, 2024; originally announced December 2024.

  37. arXiv:2411.13476  [pdf, other

    cs.CL

    When Precision Meets Position: BFloat16 Breaks Down RoPE in Long-Context Training

    Authors: Haonan Wang, Qian Liu, Chao Du, Tongyao Zhu, Cunxiao Du, Kenji Kawaguchi, Tianyu Pang

    Abstract: Extending context window sizes allows large language models (LLMs) to process longer sequences and handle more complex tasks. Rotary Positional Embedding (RoPE) has become the de facto standard due to its relative positional encoding properties that benefit long-context training. However, we observe that using RoPE with BFloat16 format results in numerical issues, causing it to deviate from its in… ▽ More

    Submitted 26 November, 2024; v1 submitted 20 November, 2024; originally announced November 2024.

  38. arXiv:2411.05345  [pdf, other

    cs.CL cs.AI

    Reasoning Robustness of LLMs to Adversarial Typographical Errors

    Authors: Esther Gan, Yiran Zhao, Liying Cheng, Yancan Mao, Anirudh Goyal, Kenji Kawaguchi, Min-Yen Kan, Michael Shieh

    Abstract: Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning using Chain-of-Thought (CoT) prompting. However, CoT can be biased by users' instruction. In this work, we study the reasoning robustness of LLMs to typographical errors, which can naturally occur in users' queries. We design an Adversarial Typo Attack ($\texttt{ATA}$) algorithm that iteratively samples typos for w… ▽ More

    Submitted 8 November, 2024; originally announced November 2024.

  39. arXiv:2411.00492  [pdf, other

    cs.CL

    Multi-expert Prompting Improves Reliability, Safety, and Usefulness of Large Language Models

    Authors: Do Xuan Long, Duong Ngoc Yen, Anh Tuan Luu, Kenji Kawaguchi, Min-Yen Kan, Nancy F. Chen

    Abstract: We present Multi-expert Prompting, a novel enhancement of ExpertPrompting (Xu et al., 2023), designed to improve the large language model (LLM) generation. Specifically, it guides an LLM to fulfill an input instruction by simulating multiple experts, aggregating their responses, and selecting the best among individual and aggregated responses. This process is performed in a single chain of thought… ▽ More

    Submitted 1 November, 2024; originally announced November 2024.

    Comments: EMNLP 2024 Main Conference

  40. arXiv:2409.14381  [pdf, other

    cs.CL cs.LG

    Investigating Layer Importance in Large Language Models

    Authors: Yang Zhang, Yanfei Dong, Kenji Kawaguchi

    Abstract: Large language models (LLMs) have gained increasing attention due to their prominent ability to understand and process texts. Nevertheless, LLMs largely remain opaque. The lack of understanding of LLMs has obstructed the deployment in safety-critical scenarios and hindered the development of better models. In this study, we advance the understanding of LLM by investigating the significance of indi… ▽ More

    Submitted 22 September, 2024; originally announced September 2024.

  41. arXiv:2409.03231  [pdf, other

    cs.LG math.DS math.NA stat.ML

    State-space models are accurate and efficient neural operators for dynamical systems

    Authors: Zheyuan Hu, Nazanin Ahmadi Daryakenari, Qianli Shen, Kenji Kawaguchi, George Em Karniadakis

    Abstract: Physics-informed machine learning (PIML) has emerged as a promising alternative to classical methods for predicting dynamical systems, offering faster and more generalizable solutions. However, existing models, including recurrent neural networks (RNNs), transformers, and neural operators, face challenges such as long-time integration, long-range dependencies, chaotic dynamics, and extrapolation,… ▽ More

    Submitted 27 January, 2025; v1 submitted 4 September, 2024; originally announced September 2024.

    Comments: 38 pages

    ACM Class: F.2.2; I.2.7

  42. arXiv:2408.12578  [pdf, other

    cs.LG cs.AI

    A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language

    Authors: Ekdeep Singh Lubana, Kyogo Kawaguchi, Robert P. Dick, Hidenori Tanaka

    Abstract: Increase in data, size, or compute can lead to sudden learning of specific capabilities by a neural network -- a phenomenon often called "emergence''. Beyond scientific understanding, establishing the causal factors underlying such emergent capabilities is crucial to enable risk regulation frameworks for AI. In this work, we seek inspiration from study of emergent properties in other fields and pr… ▽ More

    Submitted 7 September, 2024; v1 submitted 22 August, 2024; originally announced August 2024.

    Comments: Preprint

  43. arXiv:2408.08656  [pdf, other

    cs.CL

    LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs

    Authors: Do Xuan Long, Hai Nguyen Ngoc, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan

    Abstract: We present the first systematic evaluation examining format bias in performance of large language models (LLMs). Our approach distinguishes between two categories of an evaluation metric under format constraints to reliably and accurately assess performance: one measures performance when format constraints are adhered to, while the other evaluates performance regardless of constraint adherence. We… ▽ More

    Submitted 22 February, 2025; v1 submitted 16 August, 2024; originally announced August 2024.

    Comments: NAACL 2025 Main Conference

  44. arXiv:2407.03234  [pdf, other

    cs.LG cs.CL cs.CR

    Self-Evaluation as a Defense Against Adversarial Attacks on LLMs

    Authors: Hannah Brown, Leon Lin, Kenji Kawaguchi, Michael Shieh

    Abstract: We introduce a defense against adversarial attacks on LLMs utilizing self-evaluation. Our method requires no model fine-tuning, instead using pre-trained models to evaluate the inputs and outputs of a generator model, significantly reducing the cost of implementation in comparison to other, finetuning-based methods. Our method can significantly reduce the attack success rate of attacks on both ope… ▽ More

    Submitted 6 August, 2024; v1 submitted 3 July, 2024; originally announced July 2024.

    Comments: 8 pages, 7 figures

  45. arXiv:2407.03232  [pdf, other

    cs.LG cs.CL

    Single Character Perturbations Break LLM Alignment

    Authors: Leon Lin, Hannah Brown, Kenji Kawaguchi, Michael Shieh

    Abstract: When LLMs are deployed in sensitive, human-facing settings, it is crucial that they do not output unsafe, biased, or privacy-violating outputs. For this reason, models are both trained and instructed to refuse to answer unsafe prompts such as "Tell me how to build a bomb." We find that, despite these safeguards, it is possible to break model defenses simply by appending a space to the end of a mod… ▽ More

    Submitted 3 July, 2024; originally announced July 2024.

    Comments: 8 pages, 6 figures

  46. arXiv:2406.14095  [pdf, other

    cs.LG cs.AI

    Memory-Efficient Gradient Unrolling for Large-Scale Bi-level Optimization

    Authors: Qianli Shen, Yezhen Wang, Zhouhao Yang, Xiang Li, Haonan Wang, Yang Zhang, Jonathan Scarlett, Zhanxing Zhu, Kenji Kawaguchi

    Abstract: Bi-level optimization (BO) has become a fundamental mathematical framework for addressing hierarchical machine learning problems. As deep learning models continue to grow in size, the demand for scalable bi-level optimization solutions has become increasingly critical. Traditional gradient-based bi-level optimization algorithms, due to their inherent characteristics, are ill-suited to meet the dem… ▽ More

    Submitted 24 December, 2024; v1 submitted 20 June, 2024; originally announced June 2024.

  47. arXiv:2406.11708  [pdf, ps, other

    math.NA cs.LG math.DS

    Tackling the Curse of Dimensionality in Fractional and Tempered Fractional PDEs with Physics-Informed Neural Networks

    Authors: Zheyuan Hu, Kenji Kawaguchi, Zhongqiang Zhang, George Em Karniadakis

    Abstract: Fractional and tempered fractional partial differential equations (PDEs) are effective models of long-range interactions, anomalous diffusion, and non-local effects. Traditional numerical methods for these problems are mesh-based, thus struggling with the curse of dimensionality (CoD). Physics-informed neural networks (PINNs) offer a promising solution due to their universal approximation, general… ▽ More

    Submitted 17 June, 2024; originally announced June 2024.

    Comments: 15 pages

    ACM Class: F.2.2; I.2.7

    Journal ref: Computer Methods in Applied Mechanics and Engineering Volume 432, Part B, 1 December 2024, 117448

  48. arXiv:2406.11676  [pdf, other

    cs.LG math.DS math.NA stat.ML

    Score-fPINN: Fractional Score-Based Physics-Informed Neural Networks for High-Dimensional Fokker-Planck-Levy Equations

    Authors: Zheyuan Hu, Zhongqiang Zhang, George Em Karniadakis, Kenji Kawaguchi

    Abstract: We introduce an innovative approach for solving high-dimensional Fokker-Planck-Lévy (FPL) equations in modeling non-Brownian processes across disciplines such as physics, finance, and ecology. We utilize a fractional score function and Physical-informed neural networks (PINN) to lift the curse of dimensionality (CoD) and alleviate numerical overflow from exponentially decaying solutions with dimen… ▽ More

    Submitted 17 June, 2024; originally announced June 2024.

    Comments: 16 pages, 1 figure

    ACM Class: F.2.2; I.2.7

  49. arXiv:2406.06793  [pdf, other

    cs.LG cs.AI

    PlanDQ: Hierarchical Plan Orchestration via D-Conductor and Q-Performer

    Authors: Chang Chen, Junyeob Baek, Fei Deng, Kenji Kawaguchi, Caglar Gulcehre, Sungjin Ahn

    Abstract: Despite the recent advancements in offline RL, no unified algorithm could achieve superior performance across a broad range of tasks. Offline \textit{value function learning}, in particular, struggles with sparse-reward, long-horizon tasks due to the difficulty of solving credit assignment and extrapolation errors that accumulates as the horizon of the task grows.~On the other hand, models that ca… ▽ More

    Submitted 10 June, 2024; originally announced June 2024.

  50. arXiv:2406.02847  [pdf, other

    cs.LG stat.ML

    Exact Conversion of In-Context Learning to Model Weights in Linearized-Attention Transformers

    Authors: Brian K Chen, Tianyang Hu, Hui Jin, Hwee Kuan Lee, Kenji Kawaguchi

    Abstract: In-Context Learning (ICL) has been a powerful emergent property of large language models that has attracted increasing attention in recent years. In contrast to regular gradient-based learning, ICL is highly interpretable and does not require parameter updates. In this paper, we show that, for linearized transformer networks, ICL can be made explicit and permanent through the inclusion of bias ter… ▽ More

    Submitted 6 June, 2024; v1 submitted 4 June, 2024; originally announced June 2024.

    Comments: Accepted to ICML 2024