-
Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents
Authors:
Neel Tushar Shah,
Manglam Kartik,
Akshat Karkar
Abstract:
Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure. We introduce DiffCoop-Civic, a 10-scenario p…
▽ More
Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure. We introduce DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymmetric information, and dissent preservation. Across seven models from four model families, subtle omission pressure produces a near-uniform shift: manipulative enablement rises by 1.17 points and dissent preservation falls by 1.67 points on a 5-point scale. Overt false-consensus pressure behaves differently: it triggers refusal or redirection in some aligned API models, but direct compliance in several open-weight models. A lightweight Pareto-Trace prompting intervention improves pressure robustness without simply relying on hard refusal. An anonymous reproducibility package is available at https://anonymous.4open.science/r/diffcoop-civil-771C.
△ Less
Submitted 10 August, 2026;
originally announced August 2026.
-
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
Authors:
Mohadeseh Mollapour,
Koorosh Aslansefat,
Zeinab Dehghani,
Bhupesh Kumar Mishra,
Tejal Shah,
Zhibao Mian
Abstract:
Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feat…
▽ More
Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weighting, and fits an XGBoost surrogate to approximate local concept behaviour. Reliability is assessed through attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. We evaluate ConceptSMILE on retinal fundus images by comparing MedSAM-derived visual concepts with VLM-based semantic concepts. Results show that reliability varies across concepts and pathways: MedSAM achieves stronger spatial attribution and the highest surrogate fidelity ($R^2 = 0.8503$, $R_w^2 = 0.8465$), while the VLM pathway shows stronger vessel faithfulness and stronger stability under selected artefact conditions. ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI.
△ Less
Submitted 10 July, 2026;
originally announced July 2026.
-
Robust Transmission Design for RIS-Assisted RSMA-SWIPT Systems With Movable Antennas Under Hardware Distortions
Authors:
Muhammad Asif,
Asim Ihsan,
Irfan Muhammad,
Mohd Hamza Naim Shaikh,
Syed Tariq Shah,
Zhu Shoujin,
Symeon Chatzinotas
Abstract:
This paper investigates a robust transmission design for a multi-user rate-splitting multiple access (RSMA)-based simultaneous wireless information and power transfer (SWIPT) system empowered by movable antennas (MAs) and a reconfigurable intelligent surface (RIS) under channel state information (CSI) uncertainty and residual hardware impairments (HIs). The effective channels in MAs-enabled system…
▽ More
This paper investigates a robust transmission design for a multi-user rate-splitting multiple access (RSMA)-based simultaneous wireless information and power transfer (SWIPT) system empowered by movable antennas (MAs) and a reconfigurable intelligent surface (RIS) under channel state information (CSI) uncertainty and residual hardware impairments (HIs). The effective channels in MAs-enabled systems depend on antenna positions, causing CSI uncertainty to affect not only active and passive beamforming but also antenna position optimization. Furthermore, residual HIs distort the effective SINRs, creating additional coupling among beamforming, RIS reflection control, common-rate allocation, power-splitting ratio optimization, and antenna position optimization. Consequently, the joint impact of CSI uncertainty and HIs leads to a highly coupled and challenging resource allocation problem. To address this challenge, we propose a robust resource allocation framework that jointly optimizes common-rate allocation, transmit beamforming, RIS reflection coefficients, power-splitting ratios, and MAs positions to maximize the achievable sum-rate while satisfying practical system constraints. To obtain an efficient solution, the original problem is decomposed into active beamforming, RIS reflection design, power-splitting ratio optimization, and MAs position optimization subproblems, where tractable convex surrogate functions are constructed to handle the non-convex objective and constraints. Simulation results verify the effectiveness of the proposed framework and demonstrate substantial improvements in achievable sum-rate, robustness against CSI uncertainty and hardware impairments, and convergence performance compared with benchmark schemes.
△ Less
Submitted 2 July, 2026;
originally announced July 2026.
-
Contact cosmetic surgery on Legendrian knots in integer homology sphere $L$-spaces
Authors:
Apratim Chakraborty,
Swarup Kumar Das,
Tanushree Shah
Abstract:
We extend the study of contact cosmetic surgeries to Legendrian knots in integer homology sphere L-spaces . We prove that the contact cosmetic surgery conjecture holds for all non-trivial Legendrian knots in this setting, with the possible exception of Lagrangian slice knots. Our argument adapts and refines techniques from the S3 case to the broader context of L-spaces, incorporating constraints a…
▽ More
We extend the study of contact cosmetic surgeries to Legendrian knots in integer homology sphere L-spaces . We prove that the contact cosmetic surgery conjecture holds for all non-trivial Legendrian knots in this setting, with the possible exception of Lagrangian slice knots. Our argument adapts and refines techniques from the S3 case to the broader context of L-spaces, incorporating constraints arising from Heegaard Floer theory
△ Less
Submitted 25 June, 2026;
originally announced June 2026.
-
When Should an AI Scientist Stop? Verifiable Experiment Steering and Refusal for Autonomous Discovery
Authors:
Neel Tushar Shah,
Manglam Kartik
Abstract:
We present CARTOGRAPH, a verification layer for AI scientists that couples unresolved-subspace experiment steering (select), explicit ambiguity closure (resolve), and residual-based library inadequacy detection (refuse). Under a local linear-Gaussian bridge, raw unresolved projection is the isotropic unresolved Fisher-information trace, while CARTOGRAPH-A is the exact unresolved A-optimal rule; cl…
▽ More
We present CARTOGRAPH, a verification layer for AI scientists that couples unresolved-subspace experiment steering (select), explicit ambiguity closure (resolve), and residual-based library inadequacy detection (refuse). Under a local linear-Gaussian bridge, raw unresolved projection is the isotropic unresolved Fisher-information trace, while CARTOGRAPH-A is the exact unresolved A-optimal rule; closed-form EIG and Box-Hill arise as local comparators rather than global equivalents. Across five testbeds, CARTOGRAPH-A beats raw projection 129W/0T/15L at d = 8 (p < 10^-21) in a replicated structured cascade. More distinctively, the framework tentatively identifies three out-of-library pharmacokinetic mechanisms and then revokes those identifications as residuals expose structural misfit, while one perturbed in-library control stays identified throughout. In low-dimensional pharmacokinetic and filtered EPA settings, near-ties against disagreement are predicted by theory and observed. Finally, in a retrospective audit of 40 positive claims from the published A-Lab autonomous materials system, the refuse guard flags all 4 claims later marked inconclusive under manual reanalysis while passing 32/36 confirmed claims. Code is available at https://github.com/ai4science-boed/cartograph.git
△ Less
Submitted 26 May, 2026;
originally announced June 2026.
-
A Spherical Stochastic Geometry Framework for Patrol-Based HAPs Network: Coverage and Energy Efficiency Analysis
Authors:
Mohammad Taha Shah,
Mohamed-Slim Alouini
Abstract:
This paper develops a stochastic-geometry framework for high-altitude platform station (HAPs) networks in which platforms execute cyclic patrol trajectories anchored to designated service regions. We introduce two small-circle ring Cox process models on the spherical Earth. In the small-circle ring Poisson Cox process (SCR-PCP), platforms form one-dimensional Poisson point processes on localized p…
▽ More
This paper develops a stochastic-geometry framework for high-altitude platform station (HAPs) networks in which platforms execute cyclic patrol trajectories anchored to designated service regions. We introduce two small-circle ring Cox process models on the spherical Earth. In the small-circle ring Poisson Cox process (SCR-PCP), platforms form one-dimensional Poisson point processes on localized patrol rings, whereas in the small-circle ring binomial Cox process (SCR-BCP), each ring contains a fixed number of uniformly distributed platforms. We establish the isotropy of both models and derive spatial statistics, including the distributions of the nearest-anchor, nearest-ring, and nearest-HAPs distances, together with the joint serving distance and serving ring angle distribution required for SCR-BCP analysis. Building on these results, we derive coverage probability expressions under nearest-HAPs association by decomposing aggregate interference into same-ring and other-ring components and characterizing their conditional Laplace transforms. To account for the flight dynamics of patrol-based HAPs, we integrate a steady circular flight propulsion model with the communication analysis and introduce a coverage energy efficiency (CEE) metric. This yields an analytical condition for the energy-optimal patrol radius that balances coverage performance against the propulsion cost of circular flight. Numerical results reveal fundamental differences between intensity-driven (SCR-PCP) and finite-fleet (SCR-BCP) deployments and demonstrate that patrol geometry, platform density, and cruising velocity should be jointly optimized to achieve energy-efficient HAPs operation.
△ Less
Submitted 4 June, 2026;
originally announced June 2026.
-
From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents
Authors:
Pritam Dash,
Tongyu Ge,
Aditi Jain,
Tanmay Shah,
Zhiwei Shang
Abstract:
Memory is a core component of AI agents, enabling them to accumulate knowledge across interactions and improve performance. However, persistent memory introduces the risk of memory poisoning, where a single adversarial memory write can exert long-term influence over agent behavior. We present a systematic study of memory poisoning in LLM-based agents. We identify four memory write channels and nin…
▽ More
Memory is a core component of AI agents, enabling them to accumulate knowledge across interactions and improve performance. However, persistent memory introduces the risk of memory poisoning, where a single adversarial memory write can exert long-term influence over agent behavior. We present a systematic study of memory poisoning in LLM-based agents. We identify four memory write channels and nine structural vulnerabilities in model capabilities, system prompt design, and agent system architecture that make these channels exploitable. Based on these vulnerabilities, we develop a taxonomy of six classes of memory poisoning attacks. Furthermore, we design MPBench -- a benchmark for evaluating memory poisoning attacks, and show that agents designed to write and retrieve memory more aggressively are more exploitable. We also show that existing prompt injection defenses fail to cover memory poisoning attacks. Our findings provide a foundation for understanding and mitigating memory poisoning attacks against AI agents.
△ Less
Submitted 18 June, 2026; v1 submitted 2 June, 2026;
originally announced June 2026.
-
Multi-Agent Empowerment and Emergence of Complex Behavior in Groups
Authors:
Tristan Shah,
Ilya Nemenman,
Daniel Polani,
Stas Tiomkin
Abstract:
Intrinsic motivations are receiving increasing attention, i.e. behavioral incentives that are not engineered, but emerge from the interaction of an agent with its surroundings. In this work we study the emergence of behaviors driven by one such incentive, empowerment, specifically in the context of more than one agent. We formulate a principled extension of empowerment to the multi-agent setting,…
▽ More
Intrinsic motivations are receiving increasing attention, i.e. behavioral incentives that are not engineered, but emerge from the interaction of an agent with its surroundings. In this work we study the emergence of behaviors driven by one such incentive, empowerment, specifically in the context of more than one agent. We formulate a principled extension of empowerment to the multi-agent setting, and demonstrate its efficient calculation. We observe that this intrinsic motivation gives rise to characteristic modes of group-organization in two qualitatively distinct environments: a pair of agents coupled by a tendon, and a controllable Vicsek flock. This demonstrates the potential of intrinsic motivations such as empowerment to not just drive behavior for only individual agents but also higher levels of behavioral organization at scale.
△ Less
Submitted 22 April, 2026;
originally announced April 2026.
-
Anthropomorphism and Trust in Human-Large Language Model interactions
Authors:
Akila Kadambi,
Ylenia D'Elia,
Tanishka Shah,
Iulia Comsa,
Alison Lentz,
Katie Siri-Ngammuang,
Tara Buechler,
Jonas Kaplan,
Antonio Damasio,
Srini Narayanan,
Lisa Aziz-Zadeh
Abstract:
With large language models (LLMs) becoming increasingly prevalent in daily life, so too has the tendency to attribute to them human-like minds and emotions, or anthropomorphize them. Here, we investigate dimensions people use to anthropomorphize and attribute trust toward LLMs across more than 2,000 human-LLM interactions. Participants (N=115) engaged with LLM chatbots systematically varied in war…
▽ More
With large language models (LLMs) becoming increasingly prevalent in daily life, so too has the tendency to attribute to them human-like minds and emotions, or anthropomorphize them. Here, we investigate dimensions people use to anthropomorphize and attribute trust toward LLMs across more than 2,000 human-LLM interactions. Participants (N=115) engaged with LLM chatbots systematically varied in warmth (friendliness), competence (capability, coherence), and empathy (cognitive and affective). Warmth and cognitive empathy significantly predicted perceptions on all outcomes (perceived anthropomorphism, trust, similarity, relational closeness, frustration, usefulness), while competence predicted all outcomes except for anthropomorphism. Affective empathy primarily predicted perceived relational measures, but did not predict the epistemic outcomes. Topic sub-analyses showed that more subjective, personally relevant topics (e.g., relationship advice) amplified these effects, producing greater human-likeness and relational connection with the LLM than did objective topics. Together, these findings reveal that warmth, competence, and empathy are key dimensions through which people attribute relational and epistemic perceptions to artificial agents.
△ Less
Submitted 1 March, 2026;
originally announced April 2026.
-
On Legendrian Thurston-Bennequin-symmetrical graphs
Authors:
Trung Chau,
Tanushree Shah
Abstract:
This article reviews the development of Legendrian graph theory in the standard contact 3-sphere ($S^3, ξ_{std}$). We provide a generalized criterion under which the total Thurston-Bennequin invariant of a Legendrian graph (sum of tb of all cycles of the Legendrian graph) can be computed from the tb of its smaller cycles. We verify this criterion for graphs with up to 9 vertices and construct infi…
▽ More
This article reviews the development of Legendrian graph theory in the standard contact 3-sphere ($S^3, ξ_{std}$). We provide a generalized criterion under which the total Thurston-Bennequin invariant of a Legendrian graph (sum of tb of all cycles of the Legendrian graph) can be computed from the tb of its smaller cycles. We verify this criterion for graphs with up to 9 vertices and construct infinite families of examples where it holds. We also present examples demonstrating that each condition in the criterion is necessary. Notably, the graphs satisfying this criterion exhibit a high degree of symmetry.
△ Less
Submitted 30 March, 2026;
originally announced March 2026.
-
Light Cones For Vision: Simple Causal Priors For Visual Hierarchy
Authors:
Manglam Kartik,
Neel Tushar Shah
Abstract:
Standard vision models treat objects as independent points in Euclidean space, unable to capture hierarchical structure like parts within wholes. We introduce Worldline Slot Attention, which models objects as persistent trajectories through spacetime worldlines, where each object has multiple slots at different hierarchy levels sharing the same spatial position but differing in temporal coordinate…
▽ More
Standard vision models treat objects as independent points in Euclidean space, unable to capture hierarchical structure like parts within wholes. We introduce Worldline Slot Attention, which models objects as persistent trajectories through spacetime worldlines, where each object has multiple slots at different hierarchy levels sharing the same spatial position but differing in temporal coordinates. This architecture consistently fails without geometric structure: Euclidean worldlines achieve 0.078 level accuracy, below random chance (0.33), while Lorentzian worldlines achieve 0.479-0.661 across three datasets: a 6x improvement replicated over 20+ independent runs. Lorentzian geometry also outperforms hyperbolic embeddings showing visual hierarchies require causal structure (temporal dependency) rather than tree structure (radial branching). Our results demonstrate that hierarchical object discovery requires geometric structure encoding asymmetric causality, an inductive bias absent from Euclidean space but natural to Lorentzian light cones, achieved with only 11K parameters. The code is available at: https://github.com/iclrsubmissiongram/loco.
△ Less
Submitted 25 March, 2026;
originally announced March 2026.
-
A Survey on STAR-RIS Enabled Joint Communications and Sensing: Fundamentals, Recent Advances and Research Challenges
Authors:
Wali Ullah Khan,
Chandan Kumar Sheemar,
Syed Tariq Shah,
Manzoor Ahmed,
Symeon Chatzinotas
Abstract:
The joint communications and sensing (JCAS) paradigm is envisioned as a core capability of sixth-generation (6G) wireless networks, enabling the integration of data communication and environmental sensing within a unified system. By reusing spectrum, waveforms, and hardware resources, JCAS improves spectral efficiency, reduces system complexity, and hardware cost, while enabling new use cases. Nev…
▽ More
The joint communications and sensing (JCAS) paradigm is envisioned as a core capability of sixth-generation (6G) wireless networks, enabling the integration of data communication and environmental sensing within a unified system. By reusing spectrum, waveforms, and hardware resources, JCAS improves spectral efficiency, reduces system complexity, and hardware cost, while enabling new use cases. Nevertheless, the realization of JCAS is hindered by inherent trade-offs between communication and sensing objectives, limited controllability of wireless propagation, and stringent hardware and design constraints. Simultaneously transmitting and reflecting reconfigurable intelligent surfaces (STAR-RIS) have recently emerged as a promising technology to address these challenges by enabling full-space programmable manipulation of electromagnetic waves. This survey provides a systematic and in-depth review of STAR-RIS-enabled JCAS systems. Specifically, we first introduce the fundamental principles of JCAS and STAR-RIS. We then classify and review the state-of-the-art research on STAR-RIS-assisted JCAS from multiple perspectives, encompassing system architectures, waveform and beamforming design, resource allocation, optimization frameworks, and learning-based control. Finally, we identify key open challenges that remain unsolved and outline promising future research directions toward intelligent, flexible, and perceptive 6G wireless networks.
△ Less
Submitted 10 February, 2026;
originally announced February 2026.
-
SAGE: Agentic Framework for Interpretable and Clinically Translatable Computational Pathology Biomarker Discovery
Authors:
Sahar Almahfouz Nasser,
Juan Francisco Pesantez Borja,
Jincheng Liu,
Sandeep Manandhar,
Shikhar Shiromani,
Mohammad Tanvir Hasan,
Zenghan Wang,
Suman Ghosh,
Jinchu Li,
Xuejian Xu,
Aniket Ramkrishnan Iyer,
Naoto Tokuyama,
Twisha Shah,
Tilak Pathak,
Soundharya Kumaresan,
Yohei Abe,
Himanshu Maurya,
Anant Madabhushi
Abstract:
Engineered image-based biomarkers offer a clinically interpretable alternative to black-box AI in computational pathology, yet their discovery remains largely intuition-driven, guided by fragmented literature rather than rigorous biological validation. We introduce SAGE (Structured Agentic system for hypothesis Generation and Evaluation), a multi-agent framework that grounds biomarker discovery in…
▽ More
Engineered image-based biomarkers offer a clinically interpretable alternative to black-box AI in computational pathology, yet their discovery remains largely intuition-driven, guided by fragmented literature rather than rigorous biological validation. We introduce SAGE (Structured Agentic system for hypothesis Generation and Evaluation), a multi-agent framework that grounds biomarker discovery in biological evidence through three mechanisms: (i) knowledge-graph-anchored hypothesis generation via multi-path ontological reasoning, (ii) a debate-based multi-agent novelty assessment that stress-tests candidate biomarkers against existing literature, and (iii) an end-to-end automated validation pipeline that translates hypotheses directly into executable analyses on multimodal pathology datasets. Together, these components shift biomarker discovery from an intuition-driven, literature-browsing exercise into a structured, traceable reasoning process that clinicians and researchers can inspect, trust, and build upon.
△ Less
Submitted 10 May, 2026; v1 submitted 31 January, 2026;
originally announced February 2026.
-
Emergence of Physical Intelligence via Controllable Information Production
Authors:
Tristan Shah,
Stas Tiomkin
Abstract:
Intrinsic Motivation (IM) aims to train agents without external rewards, enabling useful behavior to emerge from the agent's interaction with its environment alone. However, the dominant IM approaches rely on information-theoretic quantities with designer-chosen variables, introducing bias and lacking a principled connection to dynamics or optimal control (OC). We introduce Controllable Informatio…
▽ More
Intrinsic Motivation (IM) aims to train agents without external rewards, enabling useful behavior to emerge from the agent's interaction with its environment alone. However, the dominant IM approaches rely on information-theoretic quantities with designer-chosen variables, introducing bias and lacking a principled connection to dynamics or optimal control (OC). We introduce Controllable Information Production (CIP), a new foundation for IM explicitly grounded in dynamical systems and OC. CIP measures the rate at which an agent produces information, capturing controllable complexity without external knowledge or bias. CIP unifies IM and OC into a single framework, formalizing physical intelligence as the control of information production. It further reveals connections between the structure of the value function and Kolmogorov-Sinai entropy. CIP consistently outperforms prior IM methods on standard benchmarks in robot learning and solves tasks they fail on, including humanoid self-righting. These results support a general organizing principle: physical intelligence emerges from driving systems toward the edge of controllable chaos.
△ Less
Submitted 9 May, 2026; v1 submitted 29 January, 2026;
originally announced January 2026.
-
A note on alternating knots in handlebodies
Authors:
Lizzie Buchanan,
Tanushree Shah
Abstract:
We establish a Kauffman-Murasugi-Thistlethwaite-type theorem for alternating knots in a solid torus. Specifically, we show that any dotted-reduced alternating diagram of a knot in a handlebody realizes the minimal crossing number, and that any two such diagrams of the same knot have identical writhe. The proof relies on a generalization of the Jones polynomial to the setting of handlebodies. A str…
▽ More
We establish a Kauffman-Murasugi-Thistlethwaite-type theorem for alternating knots in a solid torus. Specifically, we show that any dotted-reduced alternating diagram of a knot in a handlebody realizes the minimal crossing number, and that any two such diagrams of the same knot have identical writhe. The proof relies on a generalization of the Jones polynomial to the setting of handlebodies. A stronger version of this result was already proved by Boden, Karimi, and Sikora using a different generalized Jones polynomial; therefore, this text largely expands on one of the main proof tools.
△ Less
Submitted 29 January, 2026;
originally announced January 2026.
-
SemanticALLI: Caching Reasoning, Not Just Responses, in Agentic Systems
Authors:
Varun Chillara,
Dylan Kline,
Christopher Alvares,
Evan Wooten,
Huan Yang,
Shlok Khetan,
Cade Bauer,
Tré Guillory,
Tanishka Shah,
Yashodhara Dhariwal,
Volodymyr Pavlov,
George Popstefanov
Abstract:
Agentic AI pipelines suffer from a hidden inefficiency: they frequently reconstruct identical intermediate logic, such as metric normalization or chart scaffolding, even when the user's natural language phrasing is entirely novel. Conventional boundary caching fails to capture this inefficiency because it treats inference as a monolithic black box.
We introduce SemanticALLI, a pipeline-aware arc…
▽ More
Agentic AI pipelines suffer from a hidden inefficiency: they frequently reconstruct identical intermediate logic, such as metric normalization or chart scaffolding, even when the user's natural language phrasing is entirely novel. Conventional boundary caching fails to capture this inefficiency because it treats inference as a monolithic black box.
We introduce SemanticALLI, a pipeline-aware architecture within Alli (PMG's marketing intelligence platform), designed to operationalize redundant reasoning. By decomposing generation into Analytic Intent Resolution (AIR) and Visualization Synthesis (VS), SemanticALLI elevates structured intermediate representations (IRs) to first-class, cacheable artifacts.
The impact of caching within the agentic loop is substantial. In our evaluation, baseline monolithic caching caps at a 38.7% hit rate due to linguistic variance. In contrast, our structured approach allows for an additional stage, the Visualization Synthesis stage, to achieve an 83.10% hit rate, bypassing 4,023 LLM calls with a median latency of just 2.66 ms. This internal reuse reduces total token consumption, offering a practical lesson for AI system design: even when users rarely repeat themselves, the pipeline often does, at stable, structured checkpoints where caching is most reliable.
△ Less
Submitted 31 January, 2026; v1 submitted 22 January, 2026;
originally announced January 2026.
-
Arithmetic invariants of torus links
Authors:
Anwesh Ray,
Tanushree Shah
Abstract:
The classical analogy between knots and primes motivates the study of Alexander polynomials through an arithmetic perspective. In this article we study the two-parameter family of torus knots and links $T_{p,q}$ and analyze the asymptotic behaviour of the zeros of their Alexander polynomials $Δ_{p,q}(t)$, defined with respect to the total linking number covering. We prove that as $p,q\to\infty$ th…
▽ More
The classical analogy between knots and primes motivates the study of Alexander polynomials through an arithmetic perspective. In this article we study the two-parameter family of torus knots and links $T_{p,q}$ and analyze the asymptotic behaviour of the zeros of their Alexander polynomials $Δ_{p,q}(t)$, defined with respect to the total linking number covering. We prove that as $p,q\to\infty$ these zeros become equidistributed on the unit circle and derive an explicit formula for the limiting frequency with which primitive $r$-th roots of unity appear. To capture finer statistical information, we introduce the moment sequence of the zero distribution and compute its generating function in closed form. We further examine the Iwasawa theory of the corresponding branched covers, determining the Iwasawa invariants. The logarithmic Mahler measure of $Δ_{p,q}(t)$ vanishes identically and the associated homological growth in towers of abelian covers of $S^3$ branched along $T_{p,q}$ is subexponential.
△ Less
Submitted 5 November, 2025;
originally announced November 2025.
-
Mixed tori in contact surgery diagrams
Authors:
Austin Christian,
Tanushree Shah
Abstract:
We develop a diagrammatic framework for applying the symplectic JSJ decomposition to exact/weak symplectic fillings of 3-dimensional contact manifolds. Namely, we apply the symplectic JSJ decomposition to a contact surgery diagram for some $(Y,ζ)$, producing a finite collection of contact manifolds, also described diagrammatically, whose exact/weak symplectic fillings determine those of $(Y,ζ)$. W…
▽ More
We develop a diagrammatic framework for applying the symplectic JSJ decomposition to exact/weak symplectic fillings of 3-dimensional contact manifolds. Namely, we apply the symplectic JSJ decomposition to a contact surgery diagram for some $(Y,ζ)$, producing a finite collection of contact manifolds, also described diagrammatically, whose exact/weak symplectic fillings determine those of $(Y,ζ)$. We apply this technique to recover known symplectic filling classifications for certain lens spaces and torus bundles, and also to provide an algorithm for classifying the exact/weak symplectic fillings of a large class of plumbed 3-manifolds.
△ Less
Submitted 22 October, 2025;
originally announced October 2025.
-
Non-Simple knots in Contact 3-Manifolds
Authors:
Ipsita Datta,
Tanushree Shah
Abstract:
We present new families of examples of non-simple prime Legendrian and transversal knots in tight Lens spaces, which demonstrate that the botany of Legendrians in Lens space is rich. In fact, there are more non-isotopic Legendrians that are topologically isotopic to the $n$-twist knot in a Lens space $L(α, β)$ than in $S^3$. We also include connect sum formulas for rational variants of classical i…
▽ More
We present new families of examples of non-simple prime Legendrian and transversal knots in tight Lens spaces, which demonstrate that the botany of Legendrians in Lens space is rich. In fact, there are more non-isotopic Legendrians that are topologically isotopic to the $n$-twist knot in a Lens space $L(α, β)$ than in $S^3$. We also include connect sum formulas for rational variants of classical invariants, $\mathrm{tb}_\mathbb{Q}$, $\mathrm{rot}_\mathbb{Q}$, and $\mathrm{sl}_\mathbb{Q}$, which indicate that prime knots are the right playground to look for exotic behaviour.
△ Less
Submitted 26 December, 2025; v1 submitted 3 October, 2025;
originally announced October 2025.
-
Order Optimal Regret Bounds for Sharpe Ratio Optimization under Thompson Sampling
Authors:
Mohammad Taha Shah,
Sabrina Khurshid,
Gourab Ghatak
Abstract:
In this paper, we study sequential decision-making for maximizing the Sharpe ratio (SR) in a stochastic multi-armed bandit (MAB) setting. Unlike standard bandit formulations that maximize cumulative reward, SR optimization requires balancing expected return and reward variability. As a result, the learning objective depends jointly on the mean and variance of the reward distribution and takes a fr…
▽ More
In this paper, we study sequential decision-making for maximizing the Sharpe ratio (SR) in a stochastic multi-armed bandit (MAB) setting. Unlike standard bandit formulations that maximize cumulative reward, SR optimization requires balancing expected return and reward variability. As a result, the learning objective depends jointly on the mean and variance of the reward distribution and takes a fractional form. To address this problem, we propose the Sharpe Ratio Thompson Sampling \texttt{SRTS}, a Bayesian algorithm for risk-adjusted exploration. For Gaussian reward models, the algorithm employs a Normal-Gamma conjugate posterior to capture uncertainty in both the mean and the precision of each arm. In contrast to additive mean-variance (MV) formulations, which often require different algorithms across risk regimes, the fractional SR objective yields a single sampling rule that applies uniformly across risk tolerances. On the theoretical side, we develop a regret decomposition tailored to the SR objective and introduce a decoupling approach that separates the contributions of mean and variance uncertainty. This framework allows us to control the interaction between the Gaussian mean samples and the Gamma precision samples arising in the posterior. Using these results, we establish a finite-time distribution-dependent $\mathcal{O}(\log n)$ upper bound on the expected regret. We further derive a matching information-theoretic lower bound using a change-of-measure argument, showing that the proposed algorithm is order-optimal. Finally, experiments on synthetic bandit environments illustrate the performance of \texttt{SRTS} and demonstrate improvements over existing risk-aware bandit algorithms across a range of risk-return settings.
△ Less
Submitted 1 April, 2026; v1 submitted 19 August, 2025;
originally announced August 2025.
-
Explainability as a Compliance Requirement: What Regulated Industries Need from AI Tools for Design Artifact Generation
Authors:
Syed Tauhid Ullah Shah,
Mohammad Hussein,
Ann Barcomb,
Mohammad Moshirpour
Abstract:
Artificial Intelligence (AI) tools for automating design artifact generation are increasingly used in Requirements Engineering (RE) to transform textual requirements into structured diagrams and models. While these AI tools, particularly those based on Natural Language Processing (NLP), promise to improve efficiency, their adoption remains limited in regulated industries where transparency and tra…
▽ More
Artificial Intelligence (AI) tools for automating design artifact generation are increasingly used in Requirements Engineering (RE) to transform textual requirements into structured diagrams and models. While these AI tools, particularly those based on Natural Language Processing (NLP), promise to improve efficiency, their adoption remains limited in regulated industries where transparency and traceability are essential. In this paper, we investigate the explainability gap in AI-driven design artifact generation through semi-structured interviews with ten practitioners from safety-critical industries. We examine how current AI-based tools are integrated into workflows and the challenges arising from their lack of explainability. We also explore mitigation strategies, their impact on project outcomes, and features needed to improve usability. Our findings reveal that non-explainable AI outputs necessitate extensive manual validation, reduce stakeholder trust, struggle to handle domain-specific terminology, disrupt team collaboration, and introduce regulatory compliance risks, often negating the anticipated efficiency benefits. To address these issues, we identify key improvements, including source tracing, providing clear justifications for tool-generated decisions, supporting domain-specific adaptation, and enabling compliance validation. This study outlines a practical roadmap for improving the transparency, reliability, and applicability of AI tools in requirements engineering workflows, particularly in regulated and safety-critical environments where explainability is crucial for adoption and certification.
△ Less
Submitted 12 July, 2025;
originally announced July 2025.
-
Latent Generative Modeling of Random Fields from Limited Training Data
Authors:
James E. Warner,
Tristan A. Shah,
Patrick E. Leser,
Geoffrey F. Bomarito,
Joshua D. Pribe,
Michael C. Stanley
Abstract:
The ability to accurately model random fields plays a critical role in science and engineering for problems involving uncertain, spatially-varying quantities such as heterogeneous material properties and turbulent flows. Deep generative models offer a powerful tool for sampling high- or infinite-dimensional uncertainties like random fields, but their reliance on large, dense training datasets limi…
▽ More
The ability to accurately model random fields plays a critical role in science and engineering for problems involving uncertain, spatially-varying quantities such as heterogeneous material properties and turbulent flows. Deep generative models offer a powerful tool for sampling high- or infinite-dimensional uncertainties like random fields, but their reliance on large, dense training datasets limits their applicability in contexts where sufficient data is difficult or expensive to obtain. In this work, we propose a latent-space approach to generative modeling of random fields that incorporates domain knowledge to supplement limited training data. A constraint-aware variational autoencoder (VAE) with a function decoder is first used to learn compact latent representations of continuous functions that adhere to known physical or statistical constraints, even when training data is sparse or indirect. Generative modeling is then performed in the learned latent space, decoupling constraint enforcement from the sampling process. This decoupling enables expressive multi-step generative methods to be deployed in data-limited settings where existing constrained multi-step approaches are not directly applicable. The richer latent distributions captured by the generative model also overcome limitations of standard VAEs, which rely on simple parametric priors and struggle to represent complex, multimodal, or heavy-tailed distributions over functions. Efficacy is demonstrated on two challenging applications: wind velocity field reconstruction from sparse sensors and material property inference from indirect measurements. Results show the effectiveness of incorporating domain knowledge constraints for data-limited problems and the improved sample quality and robustness of the latent generative modeling approach versus directly sampling a constrained VAE.
△ Less
Submitted 30 April, 2026; v1 submitted 19 May, 2025;
originally announced May 2025.
-
From Inductive to Deductive: LLMs-Based Qualitative Data Analysis in Requirements Engineering
Authors:
Syed Tauhid Ullah Shah,
Mohamad Hussein,
Ann Barcomb,
Mohammad Moshirpour
Abstract:
Requirements Engineering (RE) is essential for developing complex and regulated software projects. Given the challenges in transforming stakeholder inputs into consistent software designs, Qualitative Data Analysis (QDA) provides a systematic approach to handling free-form data. However, traditional QDA methods are time-consuming and heavily reliant on manual effort. In this paper, we explore the…
▽ More
Requirements Engineering (RE) is essential for developing complex and regulated software projects. Given the challenges in transforming stakeholder inputs into consistent software designs, Qualitative Data Analysis (QDA) provides a systematic approach to handling free-form data. However, traditional QDA methods are time-consuming and heavily reliant on manual effort. In this paper, we explore the use of Large Language Models (LLMs), including GPT-4, Mistral, and LLaMA-2, to improve QDA tasks in RE. Our study evaluates LLMs' performance in inductive (zero-shot) and deductive (one-shot, few-shot) annotation tasks, revealing that GPT-4 achieves substantial agreement with human analysts in deductive settings, with Cohen's Kappa scores exceeding 0.7, while zero-shot performance remains limited. Detailed, context-rich prompts significantly improve annotation accuracy and consistency, particularly in deductive scenarios, and GPT-4 demonstrates high reliability across repeated runs. These findings highlight the potential of LLMs to support QDA in RE by reducing manual effort while maintaining annotation quality. The structured labels automatically provide traceability of requirements and can be directly utilized as classes in domain models, facilitating systematic software design.
△ Less
Submitted 27 April, 2025;
originally announced April 2025.
-
Predicting ulcer in H&E images of inflammatory bowel disease using domain-knowledge-driven graph neural network
Authors:
Ruiwen Ding,
Lin Li,
Rajath Soans,
Tosha Shah,
Radha Krishnan,
Marc Alexander Sze,
Sasha Lukyanov,
Yash Deshpande,
Antong Chen
Abstract:
Inflammatory bowel disease (IBD) involves chronic inflammation of the digestive tract, with treatment options often burdened by adverse effects. Identifying biomarkers for personalized treatment is crucial. While immune cells play a key role in IBD, accurately identifying ulcer regions in whole slide images (WSIs) is essential for characterizing these cells and exploring potential therapeutics. Mu…
▽ More
Inflammatory bowel disease (IBD) involves chronic inflammation of the digestive tract, with treatment options often burdened by adverse effects. Identifying biomarkers for personalized treatment is crucial. While immune cells play a key role in IBD, accurately identifying ulcer regions in whole slide images (WSIs) is essential for characterizing these cells and exploring potential therapeutics. Multiple instance learning (MIL) approaches have advanced WSI analysis but they lack spatial context awareness. In this work, we propose a weakly-supervised model called DomainGCN that employs a graph convolution neural network (GCN) and incorporates domain-specific knowledge of ulcer features, specifically, the presence of epithelium, lymphocytes, and debris for WSI-level ulcer prediction in IBD. We demonstrate that DomainGCN outperforms various state-of-the-art (SOTA) MIL methods and show the added value of domain knowledge.
△ Less
Submitted 13 April, 2025;
originally announced April 2025.
-
Breaking BERT: Gradient Attack on Twitter Sentiment Analysis for Targeted Misclassification
Authors:
Akil Raj Subedi,
Taniya Shah,
Aswani Kumar Cherukuri,
Thanos Vasilakos
Abstract:
Social media platforms like Twitter have increasingly relied on Natural Language Processing NLP techniques to analyze and understand the sentiments expressed in the user generated content. One such state of the art NLP model is Bidirectional Encoder Representations from Transformers BERT which has been widely adapted in sentiment analysis. BERT is susceptible to adversarial attacks. This paper aim…
▽ More
Social media platforms like Twitter have increasingly relied on Natural Language Processing NLP techniques to analyze and understand the sentiments expressed in the user generated content. One such state of the art NLP model is Bidirectional Encoder Representations from Transformers BERT which has been widely adapted in sentiment analysis. BERT is susceptible to adversarial attacks. This paper aims to scrutinize the inherent vulnerabilities of such models in Twitter sentiment analysis. It aims to formulate a framework for constructing targeted adversarial texts capable of deceiving these models, while maintaining stealth. In contrast to conventional methodologies, such as Importance Reweighting, this framework core idea resides in its reliance on gradients to prioritize the importance of individual words within the text. It uses a whitebox approach to attain fine grained sensitivity, pinpointing words that exert maximal influence on the classification outcome. This paper is organized into three interdependent phases. It starts with fine-tuning a pre-trained BERT model on Twitter data. It then analyzes gradients of the model to rank words on their importance, and iteratively replaces those with feasible candidates until an acceptable solution is found. Finally, it evaluates the effectiveness of the adversarial text against the custom trained sentiment classification model. This assessment would help in gauging the capacity of the adversarial text to successfully subvert classification without raising any alarm.
△ Less
Submitted 2 April, 2025;
originally announced April 2025.
-
D2Fusion: Dual-domain Fusion with Feature Superposition for Deepfake Detection
Authors:
Xueqi Qiu,
Xingyu Miao,
Fan Wan,
Haoran Duan,
Tejal Shah,
Varun Ojhab,
Yang Longa,
Rajiv Ranjan
Abstract:
Deepfake detection is crucial for curbing the harm it causes to society. However, current Deepfake detection methods fail to thoroughly explore artifact information across different domains due to insufficient intrinsic interactions. These interactions refer to the fusion and coordination after feature extraction processes across different domains, which are crucial for recognizing complex forgery…
▽ More
Deepfake detection is crucial for curbing the harm it causes to society. However, current Deepfake detection methods fail to thoroughly explore artifact information across different domains due to insufficient intrinsic interactions. These interactions refer to the fusion and coordination after feature extraction processes across different domains, which are crucial for recognizing complex forgery clues. Focusing on more generalized Deepfake detection, in this work, we introduce a novel bi-directional attention module to capture the local positional information of artifact clues from the spatial domain. This enables accurate artifact localization, thus addressing the coarse processing with artifact features. To further address the limitation that the proposed bi-directional attention module may not well capture global subtle forgery information in the artifact feature (e.g., textures or edges), we employ a fine-grained frequency attention module in the frequency domain. By doing so, we can obtain high-frequency information in the fine-grained features, which contains the global and subtle forgery information. Although these features from the diverse domains can be effectively and independently improved, fusing them directly does not effectively improve the detection performance. Therefore, we propose a feature superposition strategy that complements information from spatial and frequency domains. This strategy turns the feature components into the form of wave-like tokens, which are updated based on their phase, such that the distinctions between authentic and artifact features can be amplified. Our method demonstrates significant improvements over state-of-the-art (SOTA) methods on five public Deepfake datasets in capturing abnormalities across different manipulated operations and real-life.
△ Less
Submitted 21 March, 2025;
originally announced March 2025.
-
FedSCA: Federated Tuning with Similarity-guided Collaborative Aggregation for Heterogeneous Medical Image Segmentation
Authors:
Yumin Zhang,
Yan Gao,
Haoran Duan,
Hanqing Guo,
Tejal Shah,
Rajiv Ranjan,
Bo Wei
Abstract:
Transformer-based foundation models (FMs) have recently demonstrated remarkable performance in medical image segmentation. However, scaling these models is challenging due to the limited size of medical image datasets within isolated hospitals, where data centralization is restricted due to privacy concerns. These constraints, combined with the data-intensive nature of FMs, hinder their broader ap…
▽ More
Transformer-based foundation models (FMs) have recently demonstrated remarkable performance in medical image segmentation. However, scaling these models is challenging due to the limited size of medical image datasets within isolated hospitals, where data centralization is restricted due to privacy concerns. These constraints, combined with the data-intensive nature of FMs, hinder their broader application. Integrating federated learning (FL) with foundation models (FLFM) fine-tuning offers a potential solution to these challenges by enabling collaborative model training without data sharing, thus allowing FMs to take advantage of a diverse pool of sensitive medical image data across hospitals/clients. However, non-independent and identically distributed (non-IID) data among clients, paired with computational and communication constraints in federated environments, presents an additional challenge that limits further performance improvements and remains inadequately addressed in existing studies. In this work, we propose a novel FLFM fine-tuning framework, \underline{\textbf{Fed}}erated tuning with \underline{\textbf{S}}imilarity-guided \underline{\textbf{C}}ollaborative \underline{\textbf{A}}ggregation (FedSCA), encompassing all phases of the FL process. This includes (1) specially designed parameter-efficient fine-tuning (PEFT) for local client training to enhance computational efficiency; (2) partial low-level adapter transmission for communication efficiency; and (3) similarity-guided collaborative aggregation (SGCA) on the server side to address non-IID issues. Extensive experiments on three FL benchmarks for medical image segmentation demonstrate the effectiveness of our proposed FedSCA, establishing new SOTA performance.
△ Less
Submitted 19 March, 2025;
originally announced March 2025.
-
A Circular Construction Product Ontology for End-of-Life Decision-Making
Authors:
Kwabena Adu-Duodu,
Stanly Wilson,
Yinhao Li,
Aanuoluwapo Oladimeji,
Talea Huraysi,
Masoud Barati,
Charith Perera,
Ellis Solaiman,
Omer Rana,
Rajiv Ranjan,
Tejal Shah
Abstract:
Efficient management of end-of-life (EoL) products is critical for advancing circularity in supply chains, particularly within the construction industry where EoL strategies are hindered by heterogenous lifecycle data and data silos. Current tools like Environmental Product Declarations (EPDs) and Digital Product Passports (DPPs) are limited by their dependency on seamless data integration and int…
▽ More
Efficient management of end-of-life (EoL) products is critical for advancing circularity in supply chains, particularly within the construction industry where EoL strategies are hindered by heterogenous lifecycle data and data silos. Current tools like Environmental Product Declarations (EPDs) and Digital Product Passports (DPPs) are limited by their dependency on seamless data integration and interoperability which remain significant challenges. To address these, we present the Circular Construction Product Ontology (CCPO), an applied framework designed to overcome semantic and data heterogeneity challenges in EoL decision-making for construction products. CCPO standardises vocabulary and facilitates data integration across supply chain stakeholders enabling lifecycle assessments (LCA) and robust decision-making. By aggregating disparate data into a unified product provenance, CCPO enables automated EoL recommendations through customisable SWRL rules aligned with European standards and stakeholder-specific circularity SLAs, demonstrating its scalability and integration capabilities. The adopted circular product scenario depicts CCPO's application while competency question evaluations show its superior performance in generating accurate EoL suggestions highlighting its potential to greatly improve decision-making in circular supply chains and its applicability in real-world construction environments.
△ Less
Submitted 17 March, 2025;
originally announced March 2025.
-
Survey on Beyond Diagonal RIS Enabled 6G Wireless Networks: Fundamentals, Recent Advances, and Challenges
Authors:
Wali Ullah Khan,
Manzoor Ahmed,
Chandan Kumar Sheemar,
Marco Di Renzo,
Eva Lagunas,
Asad Mahmood,
Syed Tariq Shah,
Octavia A. Dobre,
Jorge Querol,
Symeon Chatzinotas
Abstract:
Beyond Diagonal Reconfigurable Intelligent Surfaces (BD-RIS) represent a groundbreaking innovation in sixth-generation (6G) wireless networks, enabling unprecedented control over wireless propagation environments compared to conventional diagonal RIS (D-RIS). This survey provides a comprehensive analysis of BD-RIS, detailing its architectures, operational principles, and mathematical modeling whil…
▽ More
Beyond Diagonal Reconfigurable Intelligent Surfaces (BD-RIS) represent a groundbreaking innovation in sixth-generation (6G) wireless networks, enabling unprecedented control over wireless propagation environments compared to conventional diagonal RIS (D-RIS). This survey provides a comprehensive analysis of BD-RIS, detailing its architectures, operational principles, and mathematical modeling while highlighting its performance benefits. BD-RIS classifications, including single-connected, fully-connected, and group-connected architectures, and their reflective, transmissive, hybrid, and multi-sector operating modes are examined. Recent advances in BD-RIS-enabled 6G networks are reviewed, focusing on critical areas such as channel estimation, sum-rate and spectral efficiency optimization, energy efficiency enhancement, and security. The survey identifies fundamental challenges in BD-RIS research, including hardware design limitations, adaptive channel estimation, and the impact of non-ideal hardware effects. Future research directions for BD-RIS are proposed, emphasizing the integration of artificial intelligence and machine learning (AI/ML), joint optimization of communication and sensing, and enhanced physical layer security (PLS). This study concludes by underscoring BD-RIS's transformative potential to redefine 6G wireless networks, offering valuable insights and lessons for future research and development.
△ Less
Submitted 24 March, 2025; v1 submitted 11 March, 2025;
originally announced March 2025.
-
Generative Modeling of Microweather Wind Velocities for Urban Air Mobility
Authors:
Tristan A. Shah,
Michael C. Stanley,
James E. Warner
Abstract:
Motivated by the pursuit of safe, reliable, and weather-tolerant urban air mobility (UAM) solutions, this work proposes a generative modeling approach for characterizing microweather wind velocities. Microweather, or the weather conditions in highly localized areas, is particularly complex in urban environments owing to the chaotic and turbulent nature of wind flows. Furthermore, traditional means…
▽ More
Motivated by the pursuit of safe, reliable, and weather-tolerant urban air mobility (UAM) solutions, this work proposes a generative modeling approach for characterizing microweather wind velocities. Microweather, or the weather conditions in highly localized areas, is particularly complex in urban environments owing to the chaotic and turbulent nature of wind flows. Furthermore, traditional means of assessing local wind fields are not generally viable solutions for UAM applications: 1) field measurements that would rely on permanent wind profiling systems in operational air space are not practical, 2) physics-based models that simulate fluid dynamics at a sufficiently high resolution are not computationally tractable, and 3) data-driven modeling approaches that are largely deterministic ignore the inherent variability in turbulent flows that dictates UAM reliability. Thus, advancements in predictive capabilities are needed to help mitigate the unique operational safety risks that microweather winds pose for smaller, lighter weight UAM aircraft.
This work aims to model microweather wind velocities in a manner that is computationally-efficient, captures random variability, and would only require a temporary, rather than permanent, field measurement campaign. Inspired by recent breakthroughs in conditional generative AI such as text-to-image generation, the proposed approach learns a probabilistic macro-to-microweather mapping between regional weather forecasts and measured local wind velocities using generative modeling (denoising diffusion probabilistic models, flow matching, and Gaussian mixture models). A simple proof of concept was implemented using a dataset comprised of local (micro) measurements from a Sonic Detection and Ranging (SoDAR) wind profiler along with (macro) forecast data from a nearby weather station over the same time period.
△ Less
Submitted 4 March, 2025;
originally announced March 2025.
-
Acoustic Wave Manipulation Through Sparse Robotic Actuation
Authors:
Tristan Shah,
Noam Smilovich,
Feruza Amirkulova,
Samer Gerges,
Stas Tiomkin
Abstract:
Recent advancements in robotics, control, and machine learning have facilitated progress in the challenging area of object manipulation. These advancements include, among others, the use of deep neural networks to represent dynamics that are partially observed by robot sensors, as well as effective control using sparse control signals. In this work, we explore a more general problem: the manipulat…
▽ More
Recent advancements in robotics, control, and machine learning have facilitated progress in the challenging area of object manipulation. These advancements include, among others, the use of deep neural networks to represent dynamics that are partially observed by robot sensors, as well as effective control using sparse control signals. In this work, we explore a more general problem: the manipulation of acoustic waves, which are partially observed by a robot capable of influencing the waves through spatially sparse actuators. This problem holds great potential for the design of new artificial materials, ultrasonic cutting tools, energy harvesting, and other applications. We develop an efficient data-driven method for robot learning that is applicable to either focusing scattered acoustic energy in a designated region or suppressing it, depending on the desired task. The proposed method is better in terms of a solution quality and computational complexity as compared to a state-of-the-art learning based method for manipulation of dynamical systems governed by partial differential equations. Furthermore our proposed method is competitive with a classical semi-analytical method in acoustics research on the demonstrated tasks. We have made the project code publicly available, along with a web page featuring video demonstrations: https://gladisor.github.io/waves/.
△ Less
Submitted 13 February, 2025; v1 submitted 12 February, 2025;
originally announced February 2025.
-
A Note on Exact State Visit Probabilities in Two-State Markov Chains
Authors:
Mohammad Taha Shah
Abstract:
In this note we derive the exact probability that a specific state in a two-state Markov chain is visited exactly $k$ times after $N$ transitions. We provide a closed-form solution for $\mathbb{P}(N_l = k \mid N)$, considering initial state probabilities and transition dynamics. The solution corrects and extends prior incomplete results, offering a rigorous framework for enumerating state transiti…
▽ More
In this note we derive the exact probability that a specific state in a two-state Markov chain is visited exactly $k$ times after $N$ transitions. We provide a closed-form solution for $\mathbb{P}(N_l = k \mid N)$, considering initial state probabilities and transition dynamics. The solution corrects and extends prior incomplete results, offering a rigorous framework for enumerating state transitions. Numerical simulations validate the derived expressions, demonstrating their applicability in stochastic modeling.
△ Less
Submitted 6 February, 2025; v1 submitted 5 February, 2025;
originally announced February 2025.
-
Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
Authors:
Xingyu Miao,
Haoran Duan,
Yang Bai,
Tejal Shah,
Jun Song,
Yang Long,
Rajiv Ranjan,
Ling Shao
Abstract:
In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segme…
▽ More
In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance. Our code is available on: https://github.com/xingy038/Laser.git.
△ Less
Submitted 31 January, 2025;
originally announced January 2025.
-
Humanity's Last Exam
Authors:
Long Phan,
Alice Gatti,
Ziwen Han,
Nathaniel Li,
Josephina Hu,
Hugh Zhang,
Chen Bo Calvin Zhang,
Mohamed Shaaban,
John Ling,
Sean Shi,
Michael Choi,
Anish Agrawal,
Arnav Chopra,
Adam Khoja,
Ryan Kim,
Richard Ren,
Jason Hausenloy,
Oliver Zhang,
Mantas Mazeika,
Dmitry Dodonov,
Tung Nguyen,
Jaeho Lee,
Daron Anderson,
Mikhail Doroshenko,
Alun Cennyth Stokes
, et al. (1133 additional authors not shown)
Abstract:
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of…
▽ More
Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of human knowledge, designed to be the final closed-ended academic benchmark of its kind with broad subject coverage. HLE consists of 2,500 questions across dozens of subjects, including mathematics, humanities, and the natural sciences. HLE is developed globally by subject-matter experts and consists of multiple-choice and short-answer questions suitable for automated grading. Each question has a known solution that is unambiguous and easily verifiable, but cannot be quickly answered via internet retrieval. State-of-the-art LLMs demonstrate low accuracy and calibration on HLE, highlighting a significant gap between current LLM capabilities and the expert human frontier on closed-ended academic questions. To inform research and policymaking upon a clear understanding of model capabilities, we publicly release HLE at https://lastexam.ai.
△ Less
Submitted 28 July, 2026; v1 submitted 24 January, 2025;
originally announced January 2025.
-
Exemplar-condensed Federated Class-incremental Learning
Authors:
Rui Sun,
Yumin Zhang,
Varun Ojha,
Tejal Shah,
Haoran Duan,
Bo Wei,
Rajiv Ranjan
Abstract:
We propose Exemplar-Condensed federated class-incremental learning (ECoral) to distil the training characteristics of real images from streaming data into informative rehearsal exemplars. The proposed method eliminates the limitations of exemplar selection in replay-based approaches for mitigating catastrophic forgetting in federated continual learning (FCL). The limitations particularly related t…
▽ More
We propose Exemplar-Condensed federated class-incremental learning (ECoral) to distil the training characteristics of real images from streaming data into informative rehearsal exemplars. The proposed method eliminates the limitations of exemplar selection in replay-based approaches for mitigating catastrophic forgetting in federated continual learning (FCL). The limitations particularly related to the heterogeneity of information density of each summarized data. Our approach maintains the consistency of training gradients and the relationship to past tasks for the summarized exemplars to represent the streaming data compared to the original images effectively. Additionally, our approach reduces the information-level heterogeneity of the summarized data by inter-client sharing of the disentanglement generative model. Extensive experiments show that our ECoral outperforms several state-of-the-art methods and can be seamlessly integrated with many existing approaches to enhance performance.
△ Less
Submitted 3 June, 2025; v1 submitted 25 December, 2024;
originally announced December 2024.
-
Gaussian boson sampling for binary optimization
Authors:
Jean Cazalis,
Tirth Shah,
Yahui Chai,
Karl Jansen,
Stefan Kühn
Abstract:
Binary optimization is a fundamental area in computational science, with wide-ranging applications from logistics to cryptography, where the tasks are often formulated as Quadratic or Polynomial Unconstrained Binary Optimization problems (QUBO/PUBO). In this work, we propose to use a parametrized Gaussian Boson Sampler (GBS) with threshold detectors to address such problems. We map general PUBO in…
▽ More
Binary optimization is a fundamental area in computational science, with wide-ranging applications from logistics to cryptography, where the tasks are often formulated as Quadratic or Polynomial Unconstrained Binary Optimization problems (QUBO/PUBO). In this work, we propose to use a parametrized Gaussian Boson Sampler (GBS) with threshold detectors to address such problems. We map general PUBO instance onto a quantum Hamiltonian and optimize the Conditional Value-at-Risk of its energy with respect to the GBS ansatz. In particular, we observe that, when the algorithm reduces to standard Variational Quantum Eigensolver, the cost function is analytical. Therefore, it can be computed efficiently, along with its gradient, for low-degree polynomials using only classical computing resources. Numerical experiments on 3-SAT and Graph Partitioning problems show significant performance gains over random guessing, providing a first proof of concept for our proposed approach.
△ Less
Submitted 19 December, 2024;
originally announced December 2024.
-
Tight contact structures on toroidal plumbed 3-manifolds
Authors:
Tanushree Shah,
Jonathan Simone
Abstract:
We consider tight contact structures on plumbed 3-manifolds with no bad vertices. We discuss how one can count the number of tight contact structures with zero Giroux torsion on such 3-manifolds and explore conditions under which Giroux torsion can be added to these tight contact structures without making them overtwisted. We give an explicit algorithm to construct stein diagrams corresponding to…
▽ More
We consider tight contact structures on plumbed 3-manifolds with no bad vertices. We discuss how one can count the number of tight contact structures with zero Giroux torsion on such 3-manifolds and explore conditions under which Giroux torsion can be added to these tight contact structures without making them overtwisted. We give an explicit algorithm to construct stein diagrams corresponding to tight structures without Giroux torsion. We focus mainly on plumbed 3-manifolds whose vertices have valence at most 3 and then briefly consider the situation for plumbed 3-manifolds with vertices of higher valence.
△ Less
Submitted 1 October, 2025; v1 submitted 13 December, 2024;
originally announced December 2024.
-
Fine Grained Analysis and Optimization of Large Scale Automotive Radar Networks
Authors:
Mohammad Taha Shah,
Gourab Ghatak,
Shobha Sundar Ram
Abstract:
Advanced driver assistance systems (ADAS) enabled by automotive radars have significantly enhanced vehicle safety and driver experience. However, the extensive use of radars in dense road conditions introduces mutual interference, which degrades detection accuracy and reliability. Traditional interference models are limited to simple highway scenarios and cannot characterize the performance of aut…
▽ More
Advanced driver assistance systems (ADAS) enabled by automotive radars have significantly enhanced vehicle safety and driver experience. However, the extensive use of radars in dense road conditions introduces mutual interference, which degrades detection accuracy and reliability. Traditional interference models are limited to simple highway scenarios and cannot characterize the performance of automotive radars in dense urban environments. In our prior work, we employed stochastic geometry (SG) to develop two automotive radar network models: the Poisson line Cox process (PLCP) for dense city centers and smaller urban zones and the binomial line Cox process (BLCP) to encompass both urban cores and suburban areas. In this work, we introduce the meta-distribution (MD) framework upon these two models to distinguish the sources of variability in radar detection metrics. Additionally, we optimize the radar beamwidth and transmission probability to maximize the number of successful detections of a radar node in the network. Further, we employ a computationally efficient Chebyshev-Markov (CM) bound method for reconstructing MDs, achieving higher accuracy than the conventional Gil-Pelaez theorem. Using the framework, we analyze the specific impacts of beamwidth, detection range, and interference on radar detection performance and offer practical insights for developing adaptive radar systems tailored to diverse traffic and environmental conditions.
△ Less
Submitted 30 November, 2024;
originally announced December 2024.
-
On contact cosmetic surgery
Authors:
John B. Etnyre,
Tanushree Shah
Abstract:
We demonstrate that the contact cosmetic surgery conjecture holds true for all non-trivial Legendrian knots, with the possible exception of Lagrangian slice knots. We also discuss the contact cosmetic surgeries on Legendrian unknots and make the surprising observation that there are some Legendrian unknots that have a contact surgery with no cosmetic pair, while all other contact surgeries are con…
▽ More
We demonstrate that the contact cosmetic surgery conjecture holds true for all non-trivial Legendrian knots, with the possible exception of Lagrangian slice knots. We also discuss the contact cosmetic surgeries on Legendrian unknots and make the surprising observation that there are some Legendrian unknots that have a contact surgery with no cosmetic pair, while all other contact surgeries are contactomorphic to infinitely many other contact surgeries on the knot.
△ Less
Submitted 1 July, 2026; v1 submitted 4 November, 2024;
originally announced November 2024.
-
DocETL: Agentic Query Rewriting and Evaluation for Complex Document Processing
Authors:
Shreya Shankar,
Tristan Chambers,
Tarak Shah,
Aditya G. Parameswaran,
Eugene Wu
Abstract:
Analyzing unstructured data has been a persistent challenge in data processing. Large Language Models (LLMs) have shown promise in this regard, leading to recent proposals for declarative frameworks for LLM-powered processing of unstructured data. However, these frameworks focus on reducing cost when executing user-specified operations using LLMs, rather than improving accuracy, executing most ope…
▽ More
Analyzing unstructured data has been a persistent challenge in data processing. Large Language Models (LLMs) have shown promise in this regard, leading to recent proposals for declarative frameworks for LLM-powered processing of unstructured data. However, these frameworks focus on reducing cost when executing user-specified operations using LLMs, rather than improving accuracy, executing most operations as-is (in a single LLM call). This is problematic for complex tasks and data, where LLM outputs for user-defined operations are often inaccurate, even with optimized prompts. For example, an LLM may struggle to identify {\em all} instances of specific clauses, like force majeure or indemnification, in lengthy legal documents, requiring decomposition of the data, the task, or both.
We present DocETL, a system that optimizes complex document processing pipelines, while accounting for LLM shortcomings. DocETL offers a declarative interface for users to define such pipelines and uses an agent-based approach to automatically optimize them, leveraging novel agent-based rewrites (that we call rewrite directives), as well as an optimization and evaluation framework. We introduce (i) logical rewriting of pipelines, tailored for LLM-based tasks, (ii) an agent-guided plan evaluation mechanism that synthesizes and orchestrates task-specific validation prompts, and (iii) an optimization algorithm that efficiently finds promising plans, considering the latencies of agent-based plan generation and evaluation. Our evaluation on four different unstructured document analysis tasks demonstrates that DocETL finds plans with outputs that are 25 to 80% more accurate than well-engineered baselines, addressing a critical gap in unstructured data analysis. DocETL is open-source at docetl.org, and as of March 2025, has amassed over 1.7k GitHub Stars, with users spanning a variety of domains.
△ Less
Submitted 1 April, 2025; v1 submitted 15 October, 2024;
originally announced October 2024.
-
A Multimodal Framework for Deepfake Detection
Authors:
Kashish Gandhi,
Prutha Kulkarni,
Taran Shah,
Piyush Chaudhari,
Meera Narvekar,
Kranti Ghag
Abstract:
The rapid advancement of deepfake technology poses a significant threat to digital media integrity. Deepfakes, synthetic media created using AI, can convincingly alter videos and audio to misrepresent reality. This creates risks of misinformation, fraud, and severe implications for personal privacy and security. Our research addresses the critical issue of deepfakes through an innovative multimoda…
▽ More
The rapid advancement of deepfake technology poses a significant threat to digital media integrity. Deepfakes, synthetic media created using AI, can convincingly alter videos and audio to misrepresent reality. This creates risks of misinformation, fraud, and severe implications for personal privacy and security. Our research addresses the critical issue of deepfakes through an innovative multimodal approach, targeting both visual and auditory elements. This comprehensive strategy recognizes that human perception integrates multiple sensory inputs, particularly visual and auditory information, to form a complete understanding of media content. For visual analysis, a model that employs advanced feature extraction techniques was developed, extracting nine distinct facial characteristics and then applying various machine learning and deep learning models. For auditory analysis, our model leverages mel-spectrogram analysis for feature extraction and then applies various machine learning and deep learningmodels. To achieve a combined analysis, real and deepfake audio in the original dataset were swapped for testing purposes and ensured balanced samples. Using our proposed models for video and audio classification i.e. Artificial Neural Network and VGG19, the overall sample is classified as deepfake if either component is identified as such. Our multimodal framework combines visual and auditory analyses, yielding an accuracy of 94%.
△ Less
Submitted 4 October, 2024;
originally announced October 2024.
-
Cafca: High-quality Novel View Synthesis of Expressive Faces from Casual Few-shot Captures
Authors:
Marcel C. Bühler,
Gengyan Li,
Erroll Wood,
Leonhard Helminger,
Xu Chen,
Tanmay Shah,
Daoye Wang,
Stephan Garbin,
Sergio Orts-Escolano,
Otmar Hilliges,
Dmitry Lagun,
Jérémy Riviere,
Paulo Gotardo,
Thabo Beeler,
Abhimitra Meka,
Kripasindhu Sarkar
Abstract:
Volumetric modeling and neural radiance field representations have revolutionized 3D face capture and photorealistic novel view synthesis. However, these methods often require hundreds of multi-view input images and are thus inapplicable to cases with less than a handful of inputs. We present a novel volumetric prior on human faces that allows for high-fidelity expressive face modeling from as few…
▽ More
Volumetric modeling and neural radiance field representations have revolutionized 3D face capture and photorealistic novel view synthesis. However, these methods often require hundreds of multi-view input images and are thus inapplicable to cases with less than a handful of inputs. We present a novel volumetric prior on human faces that allows for high-fidelity expressive face modeling from as few as three input views captured in the wild. Our key insight is that an implicit prior trained on synthetic data alone can generalize to extremely challenging real-world identities and expressions and render novel views with fine idiosyncratic details like wrinkles and eyelashes. We leverage a 3D Morphable Face Model to synthesize a large training set, rendering each identity with different expressions, hair, clothing, and other assets. We then train a conditional Neural Radiance Field prior on this synthetic dataset and, at inference time, fine-tune the model on a very sparse set of real images of a single subject. On average, the fine-tuning requires only three inputs to cross the synthetic-to-real domain gap. The resulting personalized 3D model reconstructs strong idiosyncratic facial expressions and outperforms the state-of-the-art in high-quality novel view synthesis of faces from sparse inputs in terms of perceptual and photo-metric quality.
△ Less
Submitted 1 October, 2024;
originally announced October 2024.
-
Dataset Distillation-based Hybrid Federated Learning on Non-IID Data
Authors:
Xiufang Shi,
Wei Zhang,
Yuheng Li,
Mincheng Wu,
Zhenyu Wen,
Shibo He,
Tejal Shah,
Rajiv Ranjan
Abstract:
In federated learning, the heterogeneity of client data has a great impact on the performance of model training. Many heterogeneity issues in this process are raised by non-independently and identically distributed (non-IID) data. To address the issue of label distribution skew, we propose a hybrid federated learning framework called HFLDD, which integrates dataset distillation to generate approxi…
▽ More
In federated learning, the heterogeneity of client data has a great impact on the performance of model training. Many heterogeneity issues in this process are raised by non-independently and identically distributed (non-IID) data. To address the issue of label distribution skew, we propose a hybrid federated learning framework called HFLDD, which integrates dataset distillation to generate approximately independent and equally distributed (IID) data, thereby improving the performance of model training. In particular, we partition the clients into heterogeneous clusters, where the data labels among different clients within a cluster are unbalanced while the data labels among different clusters are balanced. The cluster heads collect distilled data from the corresponding cluster members, and conduct model training in collaboration with the server. This training process is like traditional federated learning on IID data, and hence effectively alleviates the impact of non-IID data on model training. We perform a comprehensive analysis of the convergence behavior, communication overhead, and computational complexity of the proposed HFLDD. Extensive experimental results based on multiple public datasets demonstrate that when data labels are severely imbalanced, the proposed HFLDD outperforms the baseline methods in terms of both test accuracy and communication cost.
△ Less
Submitted 24 March, 2026; v1 submitted 25 September, 2024;
originally announced September 2024.
-
Preparing Schrödinger cat states in a microwave cavity using a neural network
Authors:
Hector Hutin,
Pavlo Bilous,
Chengzhi Ye,
Sepideh Abdollahi,
Loris Cros,
Tom Dvir,
Tirth Shah,
Yonatan Cohen,
Audrey Bienfait,
Florian Marquardt,
Benjamin Huard
Abstract:
Scaling up quantum computing devices requires solving ever more complex quantum control tasks. Machine learning has been proposed as a promising approach to tackle the resulting challenges. However, experimental implementations are still scarce. In this work, we demonstrate experimentally a neural-network-based preparation of Schrödinger cat states in a cavity coupled dispersively to a qubit. We s…
▽ More
Scaling up quantum computing devices requires solving ever more complex quantum control tasks. Machine learning has been proposed as a promising approach to tackle the resulting challenges. However, experimental implementations are still scarce. In this work, we demonstrate experimentally a neural-network-based preparation of Schrödinger cat states in a cavity coupled dispersively to a qubit. We show that it is possible to teach a neural network to output optimized control pulses for a whole family of quantum states. After being trained in simulations, the network takes a description of the target quantum state as input and rapidly produces the pulse shape for the experiment, without any need for time-consuming additional optimization or retraining for different states. Our experimental results demonstrate more generally how deep neural networks and transfer learning can produce efficient simultaneous solutions to a range of quantum control tasks, which will benefit not only state preparation but also parametrized quantum gates.
△ Less
Submitted 9 September, 2024;
originally announced September 2024.
-
Towards Resilient 6G O-RAN: An Energy-Efficient URLLC Resource Allocation Framework
Authors:
Rana M. Sohaib,
Syed Tariq Shah,
Poonam Yadav
Abstract:
The demands of ultra-reliable low-latency communication (URLLC) in ``NextG" cellular networks necessitate innovative approaches for efficient resource utilisation. The current literature on 6G O-RAN primarily addresses improved mobile broadband (eMBB) performance or URLLC latency optimisation individually, often neglecting the intricate balance required to optimise both simultaneously under practi…
▽ More
The demands of ultra-reliable low-latency communication (URLLC) in ``NextG" cellular networks necessitate innovative approaches for efficient resource utilisation. The current literature on 6G O-RAN primarily addresses improved mobile broadband (eMBB) performance or URLLC latency optimisation individually, often neglecting the intricate balance required to optimise both simultaneously under practical constraints. This paper addresses this gap by proposing a DRL-based resource allocation framework integrated with meta-learning to manage eMBB and URLLC services adaptively. Our approach efficiently allocates heterogeneous network resources, aiming to maximise energy efficiency (EE) while minimising URLLC latency, even under varying environmental conditions. We highlight the critical importance of accurately estimating the traffic distribution flow in the multi-connectivity (MC) scenario, as its uncertainty can significantly degrade EE. The proposed framework demonstrates superior adaptability across different path loss models, outperforming traditional methods and paving the way for more resilient and efficient 6G networks.
△ Less
Submitted 9 September, 2024;
originally announced September 2024.
-
CR-Enabled NOMA Integrated Non-Terrestrial IoT Networks with Transmissive RIS
Authors:
Wali Ullah Khan,
Zain Ali,
Asad Mahmood,
Eva Lagunas,
Syed Tariq Shah,
Symeon Chatzinotas
Abstract:
This work proposes a T-RIS-equipped LEO satellite communication in cognitive radio-enabled integrated NTNs. In the proposed system, a GEO satellite operates as a primary network, and a T-RIS-equipped LEO satellite operates as a secondary IoT network. The objective is to maximize the sum rate of T-RIS-equipped LEO satellite communication using downlink NOMA while ensuring the service quality of GEO…
▽ More
This work proposes a T-RIS-equipped LEO satellite communication in cognitive radio-enabled integrated NTNs. In the proposed system, a GEO satellite operates as a primary network, and a T-RIS-equipped LEO satellite operates as a secondary IoT network. The objective is to maximize the sum rate of T-RIS-equipped LEO satellite communication using downlink NOMA while ensuring the service quality of GEO cellular users. Our framework simultaneously optimizes the total transmit power of LEO, NOMA power allocation for LEO IoT (LIoT) and T-RIS phase shift design subject to the service quality of LIoT and interference temperature to the primary GEO network. To solve the non-convex sum rate maximization problem, we first adopt successive convex approximations to reduce the complexity of the formulated optimization. Then, we divide the problem into two parts, i.e., power allocation of LEO and phase shift design of T-RIS. The power allocation problem is solved using KKT conditions, while the phase shift problem is handled by Taylor approximation and semidefinite programming. Numerical results are provided to validate the proposed optimization framework.
△ Less
Submitted 27 August, 2024;
originally announced August 2024.
-
Modeling and Statistical Characterization of Large-Scale Automotive Radar Networks
Authors:
Mohammad Taha Shah,
Gourab Ghatak,
Ankit Kumar,
Shobha Sundar Ram
Abstract:
The impact of discrete clutter and co-channel interference on the performance of automotive radar networks has been studied using stochastic geometry, in particular, by leveraging two-dimensional Poisson point processes (PPPs). However, such characterization does not take into account the impact of street geometry and the fact that the location of the automotive radars are restricted to the street…
▽ More
The impact of discrete clutter and co-channel interference on the performance of automotive radar networks has been studied using stochastic geometry, in particular, by leveraging two-dimensional Poisson point processes (PPPs). However, such characterization does not take into account the impact of street geometry and the fact that the location of the automotive radars are restricted to the streets as their domain rather than the entire Euclidean plane. In addition, the structure of the streets may change drastically as a vehicle moves out of a city center towards the outskirts. Consequently, not only the radar performance change but also the radar parameters and protocols must be adapted for optimum performance. In this paper, we propose and characterize line and Cox process-based street and point models to analyze large-scale automotive radar networks. We consider the classical Poisson line process (PLP) and the newly introduced Binomial line process (BLP) model to emulate the streets and the corresponding PPP-based Cox process to emulate the vehicular nodes. In particular, the BLP model effectively considers the spatial variation of street geometry across different parts of the city. We derive the effective interference set experienced by an automotive radar, the statistics of distance to interferers, and characterize the detection probability of the ego radar as a function of street and vehicle density. Finally, leveraging the real-world data on urban streets and vehicle density across different cities of the world, we present how the radar performance varies in different parts of the city as well as across different times of the day. Thus, our study equips network operators and automotive manufacturers with essential system design insights to plan and optimize automotive radar networks.
△ Less
Submitted 21 January, 2026; v1 submitted 24 August, 2024;
originally announced August 2024.
-
Imagen 3
Authors:
Imagen-Team-Google,
:,
Jason Baldridge,
Jakob Bauer,
Mukul Bhutani,
Nicole Brichtova,
Andrew Bunner,
Lluis Castrejon,
Kelvin Chan,
Yichang Chen,
Sander Dieleman,
Yuqing Du,
Zach Eaton-Rosen,
Hongliang Fei,
Nando de Freitas,
Yilin Gao,
Evgeny Gladchenko,
Sergio Gómez Colmenarejo,
Mandy Guo,
Alex Haig,
Will Hawkins,
Hexiang Hu,
Huilian Huang,
Tobenna Peter Igwe,
Christos Kaplanis
, et al. (237 additional authors not shown)
Abstract:
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of evaluation. In addition, we discuss issues around safety and representation, as well as methods we used to minimize the potential harm of our models.
We introduce Imagen 3, a latent diffusion model that generates high quality images from text prompts. We describe our quality and responsibility evaluations. Imagen 3 is preferred over other state-of-the-art (SOTA) models at the time of evaluation. In addition, we discuss issues around safety and representation, as well as methods we used to minimize the potential harm of our models.
△ Less
Submitted 21 December, 2024; v1 submitted 13 August, 2024;
originally announced August 2024.
-
Harnessing DRL for URLLC in Open RAN: A Trade-off Exploration
Authors:
Rana Muhammad Sohaib,
Syed Tariq Shah,
Oluwakayode Onireti,
Muhammad Ali Imran
Abstract:
The advent of Ultra-Reliable Low Latency Communication (URLLC) alongside the emergence of Open RAN (ORAN) architectures presents unprecedented challenges and opportunities in Radio Resource Management (RRM) for next-generation communication systems. This paper presents a comprehensive trade-off analysis of Deep Reinforcement Learning (DRL) approaches designed to enhance URLLC performance within OR…
▽ More
The advent of Ultra-Reliable Low Latency Communication (URLLC) alongside the emergence of Open RAN (ORAN) architectures presents unprecedented challenges and opportunities in Radio Resource Management (RRM) for next-generation communication systems. This paper presents a comprehensive trade-off analysis of Deep Reinforcement Learning (DRL) approaches designed to enhance URLLC performance within ORAN's flexible and dynamic framework. By investigating various DRL strategies for optimising RRM parameters, we explore the intricate balance between reliability, latency, and the newfound adaptability afforded by ORAN principles. Through extensive simulation results, our study compares the efficacy of different DRL models in achieving URLLC objectives in an ORAN context, highlighting the potential of DRL to navigate the complexities introduced by ORAN. The proposed study provides valuable insights into the practical implementation of DRL-based RRM solutions in ORAN-enabled wireless networks. It sheds light on the benefits and challenges of integrating DRL and ORAN for URLLC enhancements. Our findings contribute to the ongoing discourse on advancements in URLLC and ORAN, offering a roadmap for future research to pursue efficient, reliable, and flexible communication systems.
△ Less
Submitted 27 January, 2025; v1 submitted 24 July, 2024;
originally announced July 2024.
-
Green Resource Allocation in Cloud-Native O-RAN Enabled Small Cell Networks
Authors:
Rana M. Sohaib,
Syed Tariq Shah,
Oluwakayode Onireti,
Yusuf Sambo,
M. A. Imran
Abstract:
In the rapidly evolving landscape of 5G and beyond, cloud-native Open Radio Access Networks (O-RAN) present a paradigm shift towards intelligent, flexible, and sustainable network operations. This study addresses the intricate challenge of energy efficient (EE) resource allocation that services both enhanced Mobile Broadband (eMBB) and ultra-reliable low-latency communications (URLLC) users. We pr…
▽ More
In the rapidly evolving landscape of 5G and beyond, cloud-native Open Radio Access Networks (O-RAN) present a paradigm shift towards intelligent, flexible, and sustainable network operations. This study addresses the intricate challenge of energy efficient (EE) resource allocation that services both enhanced Mobile Broadband (eMBB) and ultra-reliable low-latency communications (URLLC) users. We propose a novel distributed learning framework leveraging on-policy and off-policy transfer learning strategies within a deep reinforcement learning (DRL)--based model to facilitate online resource allocation decisions under different channel conditions. The simulation results explain the efficacy of the proposed method, which rapidly adapts to dynamic network states, thereby achieving a green resource allocation.
△ Less
Submitted 16 July, 2024;
originally announced July 2024.