Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 267 results for author: Roy, B

Searching in archive cs. Search in all archives.
.
  1. A Unified Model for Cross-Domain Clone Detection via Model Merging

    Authors: Palash R. Roy, Banani Roy, Kevin A. Schneider, Chanchal K. Roy

    Abstract: The growing diversity of code clone types, from syntactic copies to cross-language semantic clones to AI-generated duplicates, has created a fragmentation crisis in clone detection. Current deep learning detectors are domain specialists that degrade significantly outside their training distribution, with F1 drops exceeding 70% across domains. Deploying multiple specialized models is impractical, y… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: Accepted at ASE 2026

    Journal ref: Proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE '26), October 12--16, 2026, Munich, Germany

  2. MergeSE: Post-Hoc Model Merging for Software Engineering Tasks Without Retraining

    Authors: Palash R. Roy, Banani Roy, Kevin A. Schneider, Chanchal K. Roy

    Abstract: Fine-tuned code models often behave as domain specialists and can degrade sharply under distribution shift: in our clone-detection setting, a model trained on same-language clones drops 71\% F1 on cross-language clones, while multi-task training falls to 0.151 F1 on unseen AI-generated clones. Our companion study shows that post-hoc model merging can address this fragmentation, achieving 93\% of m… ▽ More

    Submitted 4 August, 2026; originally announced August 2026.

    Comments: Accepted at ASE 2026

  3. arXiv:2607.26208  [pdf, ps, other

    cs.ET cond-mat.mes-hall quant-ph

    MEDA: Measurement-Efficient Disorder-Aware Majorana Zero Mode Detection in Realistic Devices

    Authors: Nathan Jones, Binayyak Roy, Valentine Mohaugen, Ian Lewis, Toby Cox, Sumanta Tewari, Rong Ge

    Abstract: Fault-tolerant topological quantum computing relies on identifying Majorana zero modes (MZMs), but reliable detection in realistic devices remains challenging. Conventional topological indicators are inherently biased in finite, disordered systems, blurring the distinction between true MZMs and trivial states. Furthermore, attempts to map these indicators to real observables via machine learning r… ▽ More

    Submitted 16 August, 2026; v1 submitted 28 July, 2026; originally announced July 2026.

    Comments: Accepted to 2026 IEEE International Conference on Quantum Computing and Engineering

  4. arXiv:2607.23909  [pdf, ps, other

    cs.LG cs.RO

    WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

    Authors: Sen Wang, R. Gnana Praveen, Bidhan Roy, Marcos Villagra

    Abstract: Many recent robot policies pursue stronger control by using large pretrained vision-language models (VLMs) as the action backbone. We introduce WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone. During training, a single diffusion transformer generates conti… ▽ More

    Submitted 31 July, 2026; v1 submitted 26 July, 2026; originally announced July 2026.

    Comments: 9 pages, 4 figures

    ACM Class: I.2.9; I.2.10

  5. arXiv:2607.21997  [pdf, ps, other

    cs.SE

    "Go Home Copilot, You're Drunk": Understanding Developer Responses to Agent-Generated Code Review Comments

    Authors: Shamse Tasnim Cynthia, Ratnadira Widyasari, Banani Roy, Ting Zhang, David Lo

    Abstract: Code review is a critical quality assurance practice in software engineering development, and AI coding agents are increasingly generating review comments on pull requests. However, little is known about how developers actually respond to such agent-generated feedback. In this paper, we present the first large-scale empirical study on the resolution of agent-generated code review comments. We anal… ▽ More

    Submitted 29 July, 2026; v1 submitted 24 July, 2026; originally announced July 2026.

  6. arXiv:2607.19621  [pdf, ps, other

    cs.SE cs.AI

    Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub

    Authors: Sahand Saed, Khairul Alam, Banani Roy

    Abstract: Federated Learning (FL) enables collaborative model training without centralizing raw data, but building and operating FL systems remains difficult due to distributed execution, rapidly evolving frameworks, and privacy and governance requirements. In this paper, we present an empirical study of FL developer challenges by independently analyzing 495 Stack Overflow posts and 9,116 GitHub issues and… ▽ More

    Submitted 21 July, 2026; originally announced July 2026.

    Comments: 44 pages

  7. arXiv:2607.10839  [pdf, ps, other

    cs.SE

    Maintenance and Support in Community-Driven Scientific Pipeline Ecosystems: A Cross-Platform Empirical Study of nf-core

    Authors: Khairul Alam, Kowsik Roy, Md Shamimur Rahman, Banani Roy

    Abstract: Community-driven scientific pipeline ecosystems are increasingly important for reproducible data-intensive research, but their sustainability depends on more than workflow engines, templates, and testing infrastructure. It also depends on how communities maintain pipelines, integrate contributions, and support users across heterogeneous execution environments. This paper presents a cross-platform… ▽ More

    Submitted 12 July, 2026; originally announced July 2026.

    Comments: 42 pages

  8. arXiv:2607.08990  [pdf, ps, other

    cs.SE

    From Generic to Personalized: Exploring Persona-Aware Code Review Explanations

    Authors: Shamse Tasnim Cynthia, Ratnadira Widyasari, Banani Roy, Italo Santos, David Lo

    Abstract: Code review is essential for ensuring software quality and supporting collaboration, yet prior work shows that developers can interpret code review comments differently. These differences can hinder effective communication, particularly in collaborative settings. To address this challenge, we explore the potential of personified code review explanations. We report initial findings from an ongoing… ▽ More

    Submitted 9 July, 2026; originally announced July 2026.

    Comments: Presented at Journal Ahead Workshop (JAWs) 2026

  9. arXiv:2607.08177  [pdf, ps, other

    cs.AI cs.MA

    ASMR: Agentic Schema Generation for Ship Maintenance Report Writing

    Authors: Sohrab Namazi Nia, Amogh Dalal, Ning Sa, Peter Ly, Marti Zentmaier, Tomek Strzalkowski, Jay Miller, Rishi Singh, Senjuti Basu Roy

    Abstract: In this paper, we study the automatic schema generation problem: given a collection of historical ship maintenance and operational reports across multiple form categories, automatically discover compact and informative schemas that capture the essential information requirements of each report type. To address this challenge, we propose ASMR, a modular agentic framework consisting of two specialize… ▽ More

    Submitted 9 July, 2026; originally announced July 2026.

    Comments: Accepted at the DASHSys 2026 workshop (Systems for Data-centric Agents with Human-in-the-loop), co-located with VLDB 2026

  10. arXiv:2606.21346  [pdf, ps, other

    cs.GT cs.MA

    Simultaneously Efficient Allocation of Indivisible Items Across Multiple Dimensions

    Authors: Yasushi Kawase, Bodhayan Roy, Mohammad Azharuddin Sanpui

    Abstract: Many allocation problems are intrinsically multidimensional, since an item may contribute differently to several criteria, and optimizing a single aggregate objective can hide severe losses in other dimensions. We study how much efficiency can be guaranteed simultaneously when indivisible items have multiple attributes. To this end, we introduce the \emph{multidimensional efficient allocation} (MD… ▽ More

    Submitted 19 June, 2026; originally announced June 2026.

  11. arXiv:2606.08369  [pdf, ps, other

    cs.LG cs.AI

    An Information-Theoretic Definition for Open-Ended Learning

    Authors: Wanqiao Xu, Yifan Zhu, Benjamin Van Roy

    Abstract: A growing body of work points to the great promise of AI systems that can continually expand their capabilities as they operate in an open-ended environment. But yet there is no coherent definition of open-endedness or theory about how an agent ought to explore an open-ended environment. We introduce an information-theoretic definition based on a new concept -- the ${\textit bit-equivalent}$ -- wh… ▽ More

    Submitted 6 June, 2026; originally announced June 2026.

  12. arXiv:2606.00367  [pdf, ps, other

    cs.LG cs.AI

    Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems

    Authors: Jonathan Colaço Carr, Prakash Panangaden, Doina Precup, Benjamin Van Roy

    Abstract: Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise preferences are often more natural for users to specify than scalar rewards, and they express certain goals that scalar rewards cannot. Methods for reinforcement learning with pairwise preferences have thus received growing interest. Unfortunately, these methods are… ▽ More

    Submitted 13 August, 2026; v1 submitted 29 May, 2026; originally announced June 2026.

    Comments: Accepted for ICML 2026. v2 has an updated abstract and introduction. Results and conclusions are unchanged

  13. arXiv:2605.26064  [pdf, ps, other

    cs.CV cs.LG

    Paris 2.0: A Decentralized Diffusion Model for Video Generation

    Authors: Ali Rouzbayani, Bidhan Roy, Marcos Villagra, Zhiying Jiang

    Abstract: We present Paris 2.0, the first video generation model pre-trained through decentralized computation. Its training recipe builds upon Paris 1.0 (arXiv:2510.03434), the first ever open-weight Decentralized Diffusion Model (DDM), which showed that image generation can be trained without a monolithic GPU cluster. However, temporally coherent video generation had remained an open problem under decentr… ▽ More

    Submitted 28 May, 2026; v1 submitted 25 May, 2026; originally announced May 2026.

    Comments: 6 pages, 5 figures

    ACM Class: I.2.10; I.2.11

  14. TuniQ: Autotuning Compilation Passes for Quantum Workloads at Scale for Effectiveness and Efficiency

    Authors: Mohammad Abrarul Hasanat, Jason Ludmir, Tirthak Patel, Rohan Basu Roy

    Abstract: Quantum processors are being integrated into HPC ecosystems as co-processors, where compilation of quantum circuits into hardware-executable form determines both output fidelity and runtime. Current compilers use a fixed pass sequence and ignore the fact that optimal pass selection varies with circuit, hardware, and noise conditions. We present TuniQ, a reinforcement learning-based system that sel… ▽ More

    Submitted 11 May, 2026; originally announced May 2026.

  15. arXiv:2605.01592  [pdf, ps, other

    cs.CG

    Witness Set: A Visibility Problem in $NP\cap XP$

    Authors: Satyabrata Jana, Debabrata Pal, Bodhayan Roy, Sasanka Roy

    Abstract: We study the Witness Set problem, a natural dual to the classical Art Gallery problem. In the Witness Set problem, we are given a polygon $P$ and an integer $k$ as input, and the objective is to determine whether $P$ has a witness set of size at least $k$. A point set $X$ in $P$ is called a witness set if every point in $P$ is visible from at most one point in $X$. For simple polygons, we show t… ▽ More

    Submitted 2 May, 2026; originally announced May 2026.

    Comments: 24 pages, 17 figures

  16. arXiv:2605.00260  [pdf, ps, other

    cs.LG

    NLPOpt-Net: A Learning Method for Nonlinear Optimization with Feasibility Guarantees

    Authors: Bimol Nath Roy, Rahul Golder, MM Faruque Hasan

    Abstract: Nonlinear Parametric Optimization Network (NLPOpt-Net) is an unsupervised learning architecture to solve constrained nonlinear programs (NLP). Given the structure of an NLP, it learns the parametric solution maps with guaranteed constraint satisfaction. The architecture consists of a backbone neural network (NN) followed by a multilayer ($k$-layered) projection. While the NN drives toward optimali… ▽ More

    Submitted 30 April, 2026; originally announced May 2026.

  17. arXiv:2604.25903  [pdf

    cs.SE cs.LG

    Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models

    Authors: Ajmain Inqiad Alam, Palash Roy, Chanchal K. Roy, Banani Roy, Kevin A. Schneider

    Abstract: The accelerating adoption of Large Language Models (LLMs) in software engineering (SE) has brought with it a silent crisis: unsustainable computational cost. While these models demonstrate remarkable capabilities in different SE tasks, they are unmanageably large, slow to deploy, memory-intensive, and carbon-heavy. This reality threatens not only the scalability and accessibility of AI-powered SE,… ▽ More

    Submitted 28 April, 2026; originally announced April 2026.

    Journal ref: Proceedings of ACM Software Engineering 3, FSE, Article FSE047, 2026

  18. arXiv:2604.17093  [pdf, ps, other

    cs.CR

    HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking

    Authors: Zeng Wang, Minghao Shao, Weimin Fu, Prithwish Basu Roy, Xiaolong Guo, Ramesh Karri, Muhammad Shafique, Johann Knechtel, Ozgur Sinanoglu

    Abstract: The integration of large language models (LLMs) into electronic design automation (EDA) workflows has introduced powerful capabilities for RTL generation, verification, and design optimization, but also raises critical security concerns. Malicious LLM outputs in this domain pose hardware-level threats, including hardware Trojan insertion, side-channel leakage, and intellectual property theft, that… ▽ More

    Submitted 18 April, 2026; originally announced April 2026.

  19. arXiv:2604.15375  [pdf, ps, other

    cs.AR cs.AI cs.CR

    VeriCWEty: Embedding enabled Line-Level CWE Detection in Verilog

    Authors: Prithwish Basu Roy, Zeng Wang, Anatolii Chuvashlov, Weihua Xiao, Johann Knechtel, Ozgur Sinanoglu, Ramesh Karri

    Abstract: Large Language Models (LLMs) have shown significant improvement in RTL code generation. Despite the advances, the generated code is often riddled with common vulnerabilities and weaknesses (CWEs) that can slip by untrained eyes. Attackers can often exploit these weaknesses to fulfill their nefarious motives. Existing RTL bug-detection techniques rely on rule-based checks, formal properties, or coa… ▽ More

    Submitted 15 April, 2026; originally announced April 2026.

  20. arXiv:2603.17378  [pdf, ps, other

    cs.LG cs.AI

    Efficient Exploration at Scale

    Authors: Seyed Mohammad Asghari, Chris Chute, Vikranth Dwaracherla, Xiuyuan Lu, Mehdi Jafarnia, Victor Minden, Zheng Wen, Benjamin Van Roy

    Abstract: We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF). Our algorithm incrementally updates reward and language models as choice data is received. The reward model is fit to the choice data, while the language model is updated by a variation of reinforce, with reinforcement signals provided by the reward model. Se… ▽ More

    Submitted 20 August, 2026; v1 submitted 18 March, 2026; originally announced March 2026.

  21. arXiv:2603.15017  [pdf, ps, other

    cs.AI cs.LG

    Consequentialist Objectives and Catastrophe

    Authors: Henrik Marklund, Alex Infanger, Benjamin Van Roy

    Abstract: Because human preferences are too complex to codify, AIs operate with misspecified objectives. Optimizing such objectives often produces undesirable outcomes; this phenomenon is known as reward hacking. Such outcomes are not necessarily catastrophic. Indeed, most examples of reward hacking in previous literature are benign. And typically, objectives can be modified to resolve the issue. We study… ▽ More

    Submitted 24 April, 2026; v1 submitted 16 March, 2026; originally announced March 2026.

  22. arXiv:2603.06741  [pdf, ps, other

    cs.LG cs.AI cs.CV

    Heterogeneous Decentralized Diffusion Models

    Authors: Zhiying Jiang, Raihan Seraj, Marcos Villagra, Bidhan Roy

    Abstract: Training frontier-scale diffusion models often requires substantial computational resources concentrated in tightly-coupled clusters, limiting participation to well-resourced institutions. While Decentralized Diffusion Models (DDM) enable training multiple experts in isolation, existing approaches require 1176 GPU-days and homogeneous training objectives across all experts. We present an efficient… ▽ More

    Submitted 31 July, 2026; v1 submitted 6 March, 2026; originally announced March 2026.

    Comments: Accepted to CVPR2026

  23. arXiv:2603.05542  [pdf, ps, other

    cs.DB cs.AI cs.ET cs.GR cs.HC cs.MM

    Human-Data Interaction, Exploration, and Visualization in the AI Era: Challenges and Opportunities

    Authors: Jean-Daniel Fekete, Yifan Hu, Dominik Moritz, Arnab Nandi, Senjuti Basu Roy, Eugene Wu, Nikos Bikakis, George Papastefanatos, Panos K. Chrysanthis, Guoliang Li, Lingyun Yu

    Abstract: The rapid advancement of AI is transforming human-centered systems, with profound implications for human-AI interaction, human-data interaction, and visual analytics. In the AI era, data analysis increasingly involves large-scale, heterogeneous, and multimodal data that is predominantly unstructured, as well as foundation models such as LLMs and VLMs, which introduce additional uncertainty into an… ▽ More

    Submitted 4 March, 2026; originally announced March 2026.

  24. XMENTOR: A Rank-Aware Aggregation Approach for Human-Centered Explainable AI in Just-in-Time Software Defect Prediction

    Authors: Saumendu Roy, Banani Roy, Chanchal Roy, Richard Bassey

    Abstract: Machine learning (ML)-based defect prediction models can improve software quality. However, their opaque reasoning creates an HCI challenge because developers struggle to trust models they cannot interpret. Explainable AI (XAI) methods such as LIME, SHAP, and BreakDown aim to provide transparency, but when used together, they often produce conflicting explanations that increase confusion, frustrat… ▽ More

    Submitted 26 March, 2026; v1 submitted 25 February, 2026; originally announced February 2026.

    Comments: 10 pages, 14 figures, conference (FORGE '26: 2026 IEEE/ACM Third International Conference on AI Foundation Models and Software Engineering, Rio de Janeiro, Brazil, April 2026)

  25. arXiv:2602.17858  [pdf, ps, other

    cs.DB

    Multi-Attribute Group Fairness in $k$-NN Queries on Vector Databases

    Authors: Thinh On, Senjuti Basu Roy, Baruch Schieber

    Abstract: We initiate the study of multi-attribute group fairness in $k$-nearest neighbor ($k$-NN) search over vector databases. Unlike prior work that optimizes efficiency or query filtering, fairness imposes count constraints to ensure proportional representation across groups defined by protected attributes. When fairness spans multiple attributes, these constraints must be satisfied simultaneously, maki… ▽ More

    Submitted 19 February, 2026; originally announced February 2026.

  26. arXiv:2602.02685  [pdf, ps, other

    cs.LG

    Expert-Data Alignment Governs Generation Quality in Decentralized Diffusion Models

    Authors: Marcos Villagra, Bidhan Roy, Raihan Seraj, Zhiying Jiang

    Abstract: Decentralized Diffusion Models (DDMs) route denoising through experts trained independently on disjoint data clusters, which can strongly disagree in their predictions. What governs the quality of generations in such systems? We present the first ever systematic investigation of this question. A priori, the expectation is that minimizing denoising trajectory sensitivity -- minimizing how perturbat… ▽ More

    Submitted 31 July, 2026; v1 submitted 2 February, 2026; originally announced February 2026.

    Comments: 15 pages, 4 figures. DeLTa@ICLR2026 and Sci4DL@ICLR2026

    ACM Class: I.2.11

  27. arXiv:2602.02564  [pdf, ps, other

    cs.LG cs.MA

    Label Curation Using Agentic AI

    Authors: Subhodeep Ghosh, Bayan Divaaniaazar, Md Ishat-E-Rabban, Spencer Clarke, Senjuti Basu Roy

    Abstract: Data annotation is essential for supervised learning, yet producing accurate, unbiased, and scalable labels remains challenging as datasets grow in size and modality. Traditional human-centric pipelines are costly, slow, and prone to annotator variability, motivating reliability-aware automated annotation. We present AURA (Agentic AI for Unified Reliability Modeling and Annotation Aggregation), an… ▽ More

    Submitted 30 January, 2026; originally announced February 2026.

  28. arXiv:2602.00164  [pdf, ps, other

    cs.SE cs.AI

    Why Are AI Agent Involved Pull Requests (Fix-Related) Remain Unmerged? An Empirical Study

    Authors: Khairul Alam, Saikat Mondal, Banani Roy

    Abstract: Autonomous coding agents (e.g., OpenAI Codex, Devin, GitHub Copilot) are increasingly used to generate fix-related pull requests (PRs) in real world software repositories. However, their practical effectiveness depends on whether these contributions are accepted and merged by project maintainers. In this paper, we present an empirical study of AI agent involved fix related PRs, examining both thei… ▽ More

    Submitted 29 January, 2026; originally announced February 2026.

    Comments: 5 pages

  29. Beyond Bug Fixes: An Empirical Investigation of Post-Merge Code Quality Issues in Agent-Generated Pull Requests

    Authors: Shamse Tasnim Cynthia, Al Muttakin, Banani Roy

    Abstract: The increasing adoption of AI coding agents has increased the number of agent-generated pull requests (PRs) merged with little or no human intervention. Although such PRs promise productivity gains, their post-merge code quality remains underexplored, as prior work has largely relied on benchmarks and controlled tasks rather than large-scale post-merge analyses. To address this gap, we analyze 1,2… ▽ More

    Submitted 27 January, 2026; originally announced January 2026.

  30. Are We All Using Agents the Same Way? An Empirical Study of Core and Peripheral Developers Use of Coding Agents

    Authors: Shamse Tasnim Cynthia, Joy Krishan Das, Banani Roy

    Abstract: Autonomous AI agents are transforming software development and redefining how developers collaborate with AI. Prior research shows that the adoption and use of AI-powered tools differ between core and peripheral developers. However, it remains unclear how this dynamic unfolds in the emerging era of autonomous coding agents. In this paper, we present the first empirical study of 9,427 agentic PRs,… ▽ More

    Submitted 27 January, 2026; originally announced January 2026.

  31. arXiv:2601.09612  [pdf, ps, other

    cs.SE

    Analyzing GitHub Issues and Pull Requests in nf-core Pipelines: Insights into nf-core Pipeline Repositories

    Authors: Khairul Alam, Banani Roy

    Abstract: Scientific Workflow Systems (SWSs) such as Nextflow have become essential software frameworks for conducting reproducible, scalable, and portable computational analyses in data-intensive fields like genomics, transcriptomics, and proteomics. Building on Nextflow, the nf-core community curates standardized, peer-reviewed pipelines that follow strict testing, documentation, and governance guidelines… ▽ More

    Submitted 10 February, 2026; v1 submitted 14 January, 2026; originally announced January 2026.

    Comments: 12 pages

  32. arXiv:2601.08976  [pdf, ps, other

    cs.LG cs.CY cs.DS

    Continuous Fairness On Data Streams

    Authors: Subhodeep Ghosh, Zhihui Du, Angela Bonifati, Manish Kumar, David Bader, Senjuti Basu Roy

    Abstract: We study the problem of enforcing continuous group fairness over windows in data streams. We propose a novel fairness model that ensures group fairness at a finer granularity level (referred to as block) within each sliding window. This formulation is particularly useful when the window size is large, making it desirable to enforce fairness at a finer granularity. Within this framework, we address… ▽ More

    Submitted 13 January, 2026; originally announced January 2026.

  33. arXiv:2601.03159  [pdf, ps, other

    cs.LG cs.AI cs.PF

    Rapid Augmentations for Time Series (RATS): A High-Performance Library for Time Series Augmentation

    Authors: Wadie Skaf, Felix Kern, Aryamaan Basu Roy, Tejas Pradhan, Roman Kalkreuth, Holger Hoos

    Abstract: Time series augmentation is critical for training robust deep learning models, particularly in domains where labelled data is scarce and expensive to obtain. However, existing augmentation libraries for time series, mainly written in Python, suffer from performance bottlenecks, where running time grows exponentially as dataset sizes increase -- an aspect limiting their applicability in large-scale… ▽ More

    Submitted 6 January, 2026; originally announced January 2026.

  34. arXiv:2601.02022  [pdf, ps, other

    cs.LG

    Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

    Authors: Yifan Zhu, John C. Duchi, Benjamin Van Roy

    Abstract: We prove that Thompson sampling exhibits $\tilde{O}(σd \sqrt{T} + d r \sqrt{\mathrm{Tr}(Σ_0)})$ Bayesian regret in the linear-Gaussian bandit with a $\mathcal{N}(μ_0, Σ_0)$ prior distribution on the coefficients, where $d$ is the dimension, $T$ is the time horizon, $r$ is the maximum $\ell_2$ norm of the actions, and $σ^2$ is the noise variance. In contrast to existing regret bounds, this shows th… ▽ More

    Submitted 3 July, 2026; v1 submitted 5 January, 2026; originally announced January 2026.

  35. arXiv:2601.00231  [pdf, ps, other

    cs.LG cs.AI

    GRIT -- Geometry-Aware PEFT with K-FACPreconditioning, Fisher-Guided Reprojection, andDynamic Rank Adaptation

    Authors: Pritish Saha, Chandrav Rajbangshi, Rudra Goyal, Mohit Goyal, Anurag Deo, Biswajit Roy, Ningthoujam Dhanachandra Singh, Raxit Goswami, Amitava Das

    Abstract: Parameter-efficient fine-tuning (PEFT) is the default way to adapt LLMs, but widely used LoRA and QLoRA are largely geometry-agnostic: they optimize in fixed, randomly oriented low-rank subspaces with first-order descent, mostly ignoring local loss curvature. This can inflate the effective update budget and amplify drift along weakly constrained directions. We introduce GRIT, a dynamic, curvature-… ▽ More

    Submitted 1 January, 2026; originally announced January 2026.

  36. arXiv:2512.18852  [pdf, ps, other

    cs.SE

    What Drives Issue Resolution Speed? An Empirical Study of Scientific Workflow Systems on GitHub

    Authors: Khairul Alam, Banani Roy

    Abstract: Scientific Workflow Systems (SWSs) play a vital role in enabling reproducible, scalable, and automated scientific analysis. Like other open-source software, these systems depend on active maintenance and community engagement to remain reliable and sustainable. However, despite the importance of timely issue resolution for software quality and community trust, little is known about what drives issu… ▽ More

    Submitted 16 January, 2026; v1 submitted 21 December, 2025; originally announced December 2025.

    Comments: 7

  37. arXiv:2512.15386  [pdf, ps, other

    cs.CV

    See It Before You Grab It: Deep Learning-based Action Anticipation in Basketball

    Authors: Arnau Barrera Roy, Albert Clapés Sintes

    Abstract: Computer vision and video understanding have transformed sports analytics by enabling large-scale, automated analysis of game dynamics from broadcast footage. Despite significant advances in player and ball tracking, pose estimation, action localization, and automatic foul recognition, anticipating actions before they occur in sports videos has received comparatively little attention. This work in… ▽ More

    Submitted 17 December, 2025; originally announced December 2025.

  38. arXiv:2512.09423  [pdf, ps, other

    cs.CV

    FunPhase: A Periodic Functional Autoencoder for Motion Generation via Phase Manifolds

    Authors: Marco Pegoraro, Evan Atherton, Bruno Roy, Aliasghar Khani, Arianna Rampini

    Abstract: Learning natural body motion remains challenging due to the strong coupling between spatial geometry and temporal dynamics. Embedding motion in phase manifolds, latent spaces that capture local periodicity, has proven effective for motion prediction; however, existing approaches are tied to fixed skeletons and narrow motion distributions, limiting their applicability across diverse settings. We in… ▽ More

    Submitted 3 July, 2026; v1 submitted 10 December, 2025; originally announced December 2025.

    Comments: Accepted at ICML26

  39. arXiv:2512.05881  [pdf, ps, other

    cs.LG

    DAE-HardNet: A Physics Constrained Neural Network Enforcing Differential-Algebraic Hard Constraints

    Authors: Rahul Golder, Bimol Nath Roy, M. M. Faruque Hasan

    Abstract: Traditional physics-informed neural networks (PINNs) do not always satisfy physics based constraints, especially when the constraints include differential operators. Rather, they minimize the constraint violations in a soft way. Strict satisfaction of differential-algebraic equations (DAEs) to embed domain knowledge and first-principles in data-driven models is generally challenging. This is becau… ▽ More

    Submitted 5 December, 2025; originally announced December 2025.

  40. arXiv:2512.00833  [pdf, ps, other

    cs.CR cs.AR

    Revisiting Logic Encryption

    Authors: Rupesh Raj Karn, Lakshmi Likhitha Mankali, Zeng Wang, Saideep Sreekumar, Prithwish Basu Roy, Ozgur Sinanoglu, Lilas Alrahis, Johann Knechtel

    Abstract: Modern circuits face various threats like reverse engineering, theft of intellectual property (IP), side-channel attacks, etc. Here, we present a novel approach for IP protection based on logic encryption (LE). Unlike established schemes for logic locking, our work obfuscates the circuit's structure and functionality by encoding and encrypting the logic itself. We devise an end-to-end method for p… ▽ More

    Submitted 21 August, 2026; v1 submitted 30 November, 2025; originally announced December 2025.

  41. arXiv:2511.18531  [pdf, ps, other

    cs.CR cs.PL

    LockForge: Automating Paper-to-Code for Logic Locking with Multi-Agent Reasoning LLMs

    Authors: Akashdeep Saha, Zeng Wang, Prithwish Basu Roy, Johann Knechtel, Ozgur Sinanoglu, Ramesh Karri

    Abstract: Despite rapid progress in logic locking (LL), reproducibility remains a challenge as codes are rarely made public. We present LockForge, a first-of-its-kind, multi-agent large language model (LLM) framework that turns LL descriptions in papers into executable and tested code. LockForge provides a carefully crafted pipeline realizing forethought, implementation, iterative refinement, and a multi-st… ▽ More

    Submitted 28 November, 2025; v1 submitted 23 November, 2025; originally announced November 2025.

  42. arXiv:2511.18187  [pdf, ps, other

    cs.SE

    Establishing Traceability Links between Release Notes & Software Artifacts: Practitioners' Perspectives

    Authors: Sristy Sumana Nath, Banani Roy, Munima Jahan

    Abstract: Maintaining traceability links between software release notes and corresponding development artifacts, e.g., pull requests (PRs), commits, and issues, is essential for managing technical debt and ensuring maintainability. However, in open-source environments where contributors work remotely and asynchronously, establishing and maintaining these links is often error-prone, time-consuming, and frequ… ▽ More

    Submitted 22 November, 2025; originally announced November 2025.

    Journal ref: In 2025 IEEE 35th International Conference on Collaborative Advances in Software and COmputiNg (CASCON '25)

  43. arXiv:2511.01757  [pdf, ps, other

    cs.SE

    Towards LLM-Powered Task-Aware Retrieval of Scientific Workflows for Galaxy

    Authors: Shamse Tasnim Cynthia, Banani Roy

    Abstract: Scientific Workflow Management Systems (SWfMSs) such as Galaxy have become essential infrastructure in bioinformatics, supporting the design, execution, and sharing of complex multi-step analyses. Despite hosting hundreds of reusable workflows across domains, Galaxy's current keyword-based retrieval system offers limited support for semantic query interpretation and often fails to surface relevant… ▽ More

    Submitted 3 November, 2025; originally announced November 2025.

  44. arXiv:2510.10810  [pdf, ps, other

    cs.LG cs.DB

    Aegis: A Correlation-Based Data Masking Advisor for Data Sharing Ecosystems

    Authors: Omar Islam Laskar, Fatemeh Ramezani Khozestani, Ishika Nankani, Sohrab Namazi Nia, Senjuti Basu Roy, Kaustubh Beedkar

    Abstract: Data sharing ecosystems connect providers, consumers, and intermediaries to facilitate the exchange and use of data for a wide range of downstream tasks. In sensitive domains such as healthcare, privacy is enforced as a hard constraint, any shared data must satisfy a minimum privacy threshold. However, among all masking configurations that meet this requirement, the utility of the masked data can… ▽ More

    Submitted 4 November, 2025; v1 submitted 12 October, 2025; originally announced October 2025.

    Comments: Accepted at SIGMOD 2026

  45. arXiv:2510.03434  [pdf, ps, other

    cs.GR cs.DC cs.LG

    Paris: A Decentralized Trained Open-Weight Diffusion Model

    Authors: Zhiying Jiang, Raihan Seraj, Marcos Villagra, Bidhan Roy

    Abstract: We present Paris, the first publicly released diffusion model pre-trained entirely through decentralized computation. Paris demonstrates that high-quality text-to-image generation can be achieved without centrally coordinated infrastructure. Paris is open for research and commercial use. Paris required implementing our Distributed Diffusion Training framework from scratch. The model consists of 8… ▽ More

    Submitted 31 July, 2026; v1 submitted 3 October, 2025; originally announced October 2025.

  46. ThirstyFLOPS: Water Footprint Modeling and Analysis Toward Sustainable HPC Systems

    Authors: Yankai Jiang, Raghavendra Kanakagiri, Rohan Basu Roy, Devesh Tiwari

    Abstract: High-performance computing (HPC) systems are becoming increasingly water-intensive due to their reliance on water-based cooling and the energy used in power generation. However, the water footprint of HPC remains relatively underexplored-especially in contrast to the growing focus on carbon emissions. In this paper, we present ThirstyFLOPS - a comprehensive water footprint analysis framework for H… ▽ More

    Submitted 30 September, 2025; originally announced October 2025.

  47. arXiv:2509.25754  [pdf

    cs.SE

    Are Classical Clone Detectors Good Enough For the AI Era?

    Authors: Ajmain Inqiad Alam, Palash Roy, Farouq Al-omari, Chanchal Roy, Banani Roy, Kevin Schneider

    Abstract: The increasing adoption of AI-generated code has reshaped modern software development, introducing syntactic and semantic variations in cloned code. Unlike traditional human-written clones, AI-generated clones exhibit systematic syntactic patterns and semantic differences learned from large-scale training data. This shift presents new challenges for classical code clone detection (CCD) tools, whic… ▽ More

    Submitted 30 September, 2025; originally announced September 2025.

    Journal ref: 41th International Conference on Software Maintenance and Evolution (ICSME), 2025

  48. arXiv:2509.25094  [pdf, ps, other

    cs.GR cs.CV

    Unsupervised Representation Learning for 3D Mesh Parameterization with Semantic and Visibility Objectives

    Authors: AmirHossein Zamani, Bruno Roy, Arianna Rampini

    Abstract: Recent 3D generative models produce high-quality textures for 3D mesh objects. However, they commonly rely on the heavy assumption that input 3D meshes are accompanied by manual mesh parameterization (UV mapping), a manual task that requires both technical precision and artistic judgment. Industry surveys show that this process often accounts for a significant share of asset creation, creating a m… ▽ More

    Submitted 27 February, 2026; v1 submitted 29 September, 2025; originally announced September 2025.

  49. arXiv:2509.25090  [pdf, ps, other

    cs.PF

    DarwinGame: Playing Tournaments for Tuning Applications in Noisy Cloud Environments

    Authors: Rohan Basu Roy, Vijay Gadepally, Devesh Tiwari

    Abstract: This work introduces a new subarea of performance tuning -- performance tuning in a shared interference-prone computing environment. We demonstrate that existing tuners are significantly suboptimal by design because of their inability to account for interference during tuning. Our solution, DarwinGame, employs a tournament-based design to systematically compare application executions with differen… ▽ More

    Submitted 29 September, 2025; originally announced September 2025.

  50. arXiv:2509.16464  [pdf, ps, other

    cs.CL cs.CY

    Computational Analysis of Conversation Dynamics through Participant Responsivity

    Authors: Margaret Hughes, Brandon Roy, Elinor Poole-Dayan, Deb Roy, Jad Kabbara

    Abstract: Growing literature explores toxicity and polarization in discourse, with comparatively less work on characterizing what makes dialogue prosocial and constructive. We explore conversational discourse and investigate a method for characterizing its quality built upon the notion of ``responsivity'' -- whether one person's conversational turn is responding to a preceding turn. We develop and evaluate… ▽ More

    Submitted 19 September, 2025; originally announced September 2025.

    Journal ref: Proc. EMNLP (2025) 35500--35519