Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 95 results for author: Gray, A

Searching in archive cs. Search in all archives.
.
  1. arXiv:2608.17467  [pdf, ps, other

    cs.NI

    LoRIS: LoRaWAN-based IoT Platform for Sustainability Monitoring in Hotels

    Authors: Yash Pandey, Angus Gray, Reza Serati, Oscar Zhu, Emil Juvan, Anna Zinn, Danyelle Greene, Qingqing Chen, Sarah MacInnes, Siamak Layeghy, Sara Dolnicar, Marius Portmann

    Abstract: The hospitality sector is a major source of global greenhouse gas emissions, water stress, and waste generation, yet sustainability reporting in hotels remains constrained by coarse, manually collected operational data. We present LoRIS (LoRaWAN-based IoT platform for sustainability monitoring in hotels), a LoRaWAN-based sensing system that delivers high-resolution measurements of resource consump… ▽ More

    Submitted 18 August, 2026; originally announced August 2026.

  2. arXiv:2606.18328  [pdf, ps, other

    cs.RO

    Recover, Discover, Plan: Learning Skills and Concepts from Robot Failures

    Authors: Bowen Li, Mayank Mishra, Y. Isabel Liu, Stone Tao, Nishanth Kumar, Alexander G. Gray, Ruwan Wickramarachchi, Jonathan Francis, Sebastian Scherer, Tom Silver

    Abstract: Intelligent robots should not only recover from failures, but also acquire the abstract knowledge needed to avoid them in the future. While reinforcement learning (RL) can learn reactive recovery behaviors, training a separate policy for every distinct failure mode is highly inefficient. We introduce Recovery-Driven Synthesis of Relational Concepts (ReSYNC), the first approach that progressively d… ▽ More

    Submitted 16 June, 2026; originally announced June 2026.

    Comments: 9 pages, 6 figures. Website: https://jaraxxus-me.github.io/ReSYNC/

  3. arXiv:2603.17019  [pdf, ps, other

    cs.LG

    Transformers Can Learn Rules They've Never Seen: Proof of Computation Beyond Interpolation

    Authors: Andy Gray

    Abstract: A central question in the debate over large language models is whether transformers can learn rules they have never seen, or whether they can only interpolate: predict new cases from their similarity to training examples. We test this in a controlled setting where interpolation provably fails, so success can only come from computation beyond interpolation. We train small transformers to predict th… ▽ More

    Submitted 29 July, 2026; v1 submitted 17 March, 2026; originally announced March 2026.

    Comments: 28 pages, 6 figures

  4. arXiv:2602.23113  [pdf, ps, other

    cs.LG

    Learning Physical Operators using Neural Operators

    Authors: Vignesh Gopakumar, Ander Gray, Dan Giles, Lorenzo Zanisi, Matt J. Kusner, Timo Betcke, Stanislas Pamela, Marc Peter Deisenroth

    Abstract: Neural operators have emerged as promising surrogate models for solving partial differential equations (PDEs), but struggle to generalise beyond training distributions and are often constrained to a fixed temporal discretisation. This work introduces a physics-informed training framework that addresses these limitations by decomposing PDEs using operator splitting methods, training separate neural… ▽ More

    Submitted 2 April, 2026; v1 submitted 26 February, 2026; originally announced February 2026.

  5. arXiv:2512.17992  [pdf, ps, other

    cs.RO

    Unifying Deep Predicate Invention with Pre-trained Foundation Models

    Authors: Qianwei Wang, Bowen Li, Zhanpeng Luo, Yifan Xu, Alexander Gray, Tom Silver, Sebastian Scherer, Katia Sycara, Yaqi Xie

    Abstract: Long-horizon robotic tasks are hard due to continuous state-action spaces and sparse feedback. Symbolic world models help by decomposing tasks into discrete predicates that capture object properties and relations. Existing methods learn predicates either top-down, by prompting foundation models without data grounding, or bottom-up, from demonstrations without high-level priors. We introduce UniPre… ▽ More

    Submitted 19 December, 2025; originally announced December 2025.

    Comments: 18 pages, 11 figures

  6. arXiv:2512.01560  [pdf

    cs.DL

    Estimating the prevalence of LLM-assisted text in scholarly writing

    Authors: Andrew Gray

    Abstract: The use of large language models (LLMs) in scholarly publications has grown dramatically since the launch of ChatGPT in late 2022. This usage is often undisclosed, and it can be challenging for readers and reviewers to identify human written but LLM-revised or translated text, or predominantly LLM-generated text. Given the known quality and reliability issues connected with LLM-generated text, the… ▽ More

    Submitted 1 December, 2025; originally announced December 2025.

    Comments: 19 pages, 3 figures

  7. arXiv:2510.05709  [pdf, ps, other

    cs.CR cs.AI cs.CL

    Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering

    Authors: Mary Llewellyn, Isobel Thornton, James Bishop, Annie Gray

    Abstract: LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations are available for classical inference, and (ii) test prompts are independent. We propose a corrective Bayesian hierarchical model with embedding-space clustering that provides robust performance metrics in limited-data s… ▽ More

    Submitted 4 June, 2026; v1 submitted 7 October, 2025; originally announced October 2025.

    Comments: Accepted to the 1st Workshop on Combining Theory and Benchmarks, CTB@ICML 2026, Seoul, South Korea

  8. arXiv:2509.21527  [pdf, ps, other

    cs.DC cs.PF physics.comp-ph

    Redesigning GROMACS Halo Exchange: Improving Strong Scaling with GPU-initiated NVSHMEM

    Authors: Mahesh Doijade, Andrey Alekseenko, Ania Brown, Alan Gray, Szilárd Páll

    Abstract: Improving time-to-solution in molecular dynamics simulations often requires strong scaling due to fixed-sized problems. GROMACS is highly latency-sensitive, with peak iteration rates in the sub-millisecond, making scalability on heterogeneous supercomputers challenging. MPI's CPU-centric nature introduces additional latencies on GPU-resident applications' critical path, hindering GPU utilization a… ▽ More

    Submitted 25 September, 2025; originally announced September 2025.

    Comments: 17 pages, 8 figures, submitted to PAW-ATM Workshop, SC 2025

  9. arXiv:2509.03548  [pdf, ps, other

    cs.AI

    Multilinear and Linear Programs for Partially Identifiable Queries in Quasi-Markovian Structural Causal Models

    Authors: João P. Arroyo, João G. Rodrigues, Daniel Lawand, Denis D. Mauá, Junkyu Lee, Radu Marinescu, Alex Gray, Eduardo R. Laurentino, Fabio G. Cozman

    Abstract: We investigate partially identifiable queries in a class of causal models. We focus on acyclic Structural Causal Models that are quasi-Markovian (that is, each endogenous variable is connected with at most one exogenous confounder). We look into scenarios where endogenous variables are observed (and a distribution over them is known), while exogenous variables are not fully specified. This leads t… ▽ More

    Submitted 2 September, 2025; originally announced September 2025.

    Comments: Accepted at the Causal Abstractions and Representations (CAR) workshop of the 41st Conference on Uncertainty in Artificial Intelligence (UAI 2025)

  10. arXiv:2508.15152  [pdf, ps, other

    cs.HC

    Evaluating an Immersive Analytics Application at an Enterprise Business Intelligence Customer Conference

    Authors: Matthew Brehmer, Ginger Gloystein, Bailiang Zhou, Abby Gray, Sruthi Pillai, Ben Medina, Vidya Setlur

    Abstract: We reflect on an evaluation of an immersive analytics application (Tableau for visionOS) conducted at a large enterprise business intelligence (BI) conference. Conducting a study in such a context offered an opportunistic setting to gather diverse feedback. However, this setting also highlighted the challenge of evaluating usability while also assessing potential utility, as feedback straddled bet… ▽ More

    Submitted 20 August, 2025; originally announced August 2025.

    Comments: To appear at the Human Factors in Immersive Analytics (HFIA) Workshop at IEEE VIS 2025

  11. arXiv:2506.14095  [pdf, ps, other

    cs.LG

    Transformers Learn Faster with Semantic Focus

    Authors: Parikshit Ram, Kenneth L. Clarkson, Tim Klinger, Shashanka Ubaru, Alexander G. Gray

    Abstract: Various forms of sparse attention have been explored to mitigate the quadratic computational and memory cost of the attention mechanism in transformers. We study sparse transformers not through a lens of efficiency but rather in terms of learnability and generalization. Empirically studying a range of attention mechanisms, we find that input-dependent sparse attention models appear to converge fas… ▽ More

    Submitted 18 June, 2025; v1 submitted 16 June, 2025; originally announced June 2025.

  12. arXiv:2505.02077  [pdf, ps, other

    cs.CR cs.AI cs.MA

    Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI Agents

    Authors: Christian Schroeder de Witt, Klaudia Krawiecka, Igor Krawczuk, Ben Hagag, William L. Anderson, Peter Belcak, Ben Bucknall, Xiaohong Cai, Ayush Chopra, Doron Cohen, Ron F. Del Rosario, Andis Draguns, Annie Gray, Keren Katz, Vasilios Mavroudis, Jaron Mink, Sumeet Ramesh Motwani, Jonathan Petit, Leif-Sebastian Rembeck, Chandler Smith, John Sotiropoulos, Steven Young, Sarah Scheffler, Mary Llewellyn

    Abstract: AI agents are beginning to interact with each other directly and across internet platforms and physical environments, creating security challenges beyond traditional cybersecurity and AI safety frameworks. Free-form protocols are essential for AI's task generalization but enable new threats like secret collusion and coordinated swarm attacks. Network effects can rapidly spread privacy breaches, di… ▽ More

    Submitted 29 April, 2026; v1 submitted 4 May, 2025; originally announced May 2025.

  13. arXiv:2503.15549  [pdf, other

    cs.CY cs.AI cs.HC cs.IR

    Rendering Transparency to Ranking in Educational Assessment via Bayesian Comparative Judgement

    Authors: Andy Gray, Alma Rahat, Stephen Lindsay, Jen Pearson, Tom Crick

    Abstract: Ensuring transparency in educational assessment is increasingly critical, particularly post-pandemic, as demand grows for fairer and more reliable evaluation methods. Comparative Judgement (CJ) offers a promising alternative to traditional assessments, yet concerns remain about its perceived opacity. This paper examines how Bayesian Comparative Judgement (BCJ) enhances transparency by integrating… ▽ More

    Submitted 17 March, 2025; originally announced March 2025.

  14. arXiv:2503.00479  [pdf, ps, other

    cs.LG cs.IR stat.ML

    Bayesian Active Learning for Multi-Criteria Comparative Judgement in Educational Assessment

    Authors: Andy Gray, Alma Rahat, Tom Crick, Stephen Lindsay

    Abstract: Comparative Judgement (CJ) provides an alternative assessment approach by evaluating work holistically rather than breaking it into discrete criteria. This method leverages human ability to make nuanced comparisons, yielding more reliable and valid assessments. CJ aligns with real-world evaluations, where overall quality emerges from the interplay of various elements. However, rubrics remain widel… ▽ More

    Submitted 3 September, 2025; v1 submitted 1 March, 2025; originally announced March 2025.

  15. arXiv:2502.08697  [pdf, other

    cs.RO

    Bilevel Learning for Bilevel Planning

    Authors: Bowen Li, Tom Silver, Sebastian Scherer, Alexander Gray

    Abstract: A robot that learns from demonstrations should not just imitate what it sees -- it should understand the high-level concepts that are being demonstrated and generalize them to new tasks. Bilevel planning is a hierarchical model-based approach where predicates (relational state abstractions) can be leveraged to achieve compositional generalization. However, previous bilevel planning approaches depe… ▽ More

    Submitted 11 May, 2025; v1 submitted 12 February, 2025; originally announced February 2025.

    Comments: In Proceedings of Robotics, Science, and Systems (RSS 2025). See our website for details: https://jaraxxus-me.github.io/IVNTR/

  16. arXiv:2502.04406  [pdf, ps, other

    cs.LG cs.AI physics.comp-ph

    Calibrated Physics-Informed Uncertainty Quantification

    Authors: Vignesh Gopakumar, Ander Gray, Lorenzo Zanisi, Timothy Nunn, Daniel Giles, Matt J. Kusner, Stanislas Pamela, Marc Peter Deisenroth

    Abstract: Simulating complex physical systems is crucial for understanding and predicting phenomena across diverse fields, such as fluid dynamics and heat transfer, as well as plasma physics and structural mechanics. Traditional approaches rely on solving partial differential equations (PDEs) using numerical methods, which are computationally expensive and often prohibitively slow for real-time applications… ▽ More

    Submitted 10 June, 2025; v1 submitted 6 February, 2025; originally announced February 2025.

    Journal ref: ICML 2025

  17. arXiv:2501.18426  [pdf, ps, other

    cs.LG cs.AI

    Guaranteed prediction sets for functional surrogate models

    Authors: Ander Gray, Vignesh Gopakumar, Sylvain Rousseau, Sébastien Destercke

    Abstract: We propose a method for obtaining statistically guaranteed prediction sets for functional machine learning methods: surrogate models which map between function spaces, motivated by the need to build reliable PDE emulators. The method constructs nested prediction sets on a low-dimensional representation (an SVD) of the surrogate model's error, and then maps these sets to the prediction space using… ▽ More

    Submitted 19 June, 2025; v1 submitted 30 January, 2025; originally announced January 2025.

  18. arXiv:2501.11335  [pdf, other

    cs.CL cs.AI

    Few-shot Policy (de)composition in Conversational Question Answering

    Authors: Kyle Erwin, Guy Axelrod, Maria Chang, Achille Fokoue, Maxwell Crouse, Soham Dan, Tian Gao, Rosario Uceda-Sosa, Ndivhuwo Makondo, Naweed Khan, Alexander Gray

    Abstract: The task of policy compliance detection (PCD) is to determine if a scenario is in compliance with respect to a set of written policies. In a conversational setting, the results of PCD can indicate if clarifying questions must be asked to determine compliance status. Existing approaches usually claim to have reasoning capabilities that are latent or require a large amount of annotated data. In this… ▽ More

    Submitted 20 January, 2025; originally announced January 2025.

  19. arXiv:2501.00612  [pdf, other

    cs.IT

    Breaking through the classical Shannon entropy limit: A new frontier through logical semantics

    Authors: Luis A. Lastras, Barry M. Trager, Jonathan Lenchner, Wojciech Szpankowski, Chai Wah Wu, Mark S. Squillante, Alexander Gray

    Abstract: Information theory has provided foundations for the theories of several application areas critical for modern society, including communications, computer storage, and AI. A key aspect of Shannon's 1948 theory is a sharp lower bound on the number of bits needed to encode and communicate a string of symbols. When he introduced the theory, Shannon famously excluded any notion of semantics behind the… ▽ More

    Submitted 31 December, 2024; originally announced January 2025.

  20. arXiv:2411.00773  [pdf, other

    cs.AI

    LogiCity: Advancing Neuro-Symbolic AI with Abstract Urban Simulation

    Authors: Bowen Li, Zhaoyu Li, Qiwei Du, Jinqi Luo, Wenshan Wang, Yaqi Xie, Simon Stepputtis, Chen Wang, Katia P. Sycara, Pradeep Kumar Ravikumar, Alexander G. Gray, Xujie Si, Sebastian Scherer

    Abstract: Recent years have witnessed the rapid development of Neuro-Symbolic (NeSy) AI systems, which integrate symbolic reasoning into deep neural networks. However, most of the existing benchmarks for NeSy AI fail to provide long-horizon reasoning tasks with complex multi-agent interactions. Furthermore, they are usually constrained by fixed and simplistic logical rules over limited entities, making them… ▽ More

    Submitted 3 April, 2025; v1 submitted 1 November, 2024; originally announced November 2024.

    Comments: 25 pages, 8 figures, In Advances in Neural Information Processing Systems (NeurIPS) 37 D&B Track (2024): 69840-69864

    Journal ref: Advances in Neural Information Processing Systems, 37, 69840-69864 (2024)

  21. arXiv:2410.07966  [pdf, other

    cs.LG cs.AI

    Neural Reasoning Networks: Efficient Interpretable Neural Networks With Automatic Textual Explanations

    Authors: Stephen Carrow, Kyle Harper Erwin, Olga Vilenskaia, Parikshit Ram, Tim Klinger, Naweed Aghmad Khan, Ndivhuwo Makondo, Alexander Gray

    Abstract: Recent advances in machine learning have led to a surge in adoption of neural networks for various tasks, but lack of interpretability remains an issue for many others in which an understanding of the features influencing the prediction is necessary to ensure fairness, safety, and legal compliance. In this paper we consider one class of such tasks, tabular dataset classification, and propose a nov… ▽ More

    Submitted 10 October, 2024; originally announced October 2024.

    ACM Class: I.2.6; I.5.1

  22. arXiv:2409.20380  [pdf, other

    cs.CE

    Heterogeneous computing in a strongly-connected CPU-GPU environment: fast multiple time-evolution equation-based modeling accelerated using data-driven approach

    Authors: Tsuyoshi Ichimura, Kohei Fujita, Muneo Hori, Lalith Maddegedara, Jack Wells, Alan Gray, Ian Karlin, John Linford

    Abstract: We propose a CPU-GPU heterogeneous computing method for solving time-evolution partial differential equation problems many times with guaranteed accuracy, in short time-to-solution and low energy-to-solution. On a single-GH200 node, the proposed method improved the computation speed by 86.4 and 8.67 times compared to the conventional method run only on CPU and only on GPU, respectively. Furthermor… ▽ More

    Submitted 30 September, 2024; originally announced September 2024.

    Comments: 22 pages, 5 figures, accepted for Eleventh Workshop on Accelerator Programming and Directives (WACCPD 2024)

  23. arXiv:2408.09881  [pdf, ps, other

    cs.AI physics.ao-ph physics.plasm-ph

    Uncertainty Quantification of Surrogate Models using Conformal Prediction

    Authors: Vignesh Gopakumar, Ander Gray, Joel Oskarsson, Lorenzo Zanisi, Daniel Giles, Matt J. Kusner, Stanislas Pamela, Marc Peter Deisenroth

    Abstract: Data-driven surrogate models offer quick approximations to complex numerical and experimental systems but typically lack uncertainty quantification, limiting their reliability in safety-critical applications. While Bayesian methods provide uncertainty estimates, they offer no statistical guarantees and struggle with high-dimensional spatio-temporal problems due to computational costs. We present a… ▽ More

    Submitted 5 January, 2026; v1 submitted 19 August, 2024; originally announced August 2024.

  24. arXiv:2406.16106  [pdf, other

    cs.IR cs.AI

    Evaluating Ensemble Methods for News Recommender Systems

    Authors: Alexander Gray, Noorhan Abbas

    Abstract: News recommendation is crucial for facilitating individuals' access to articles, particularly amid the increasingly digital landscape of news consumption. Consequently, extensive research is dedicated to News Recommender Systems (NRS) with increasingly sophisticated algorithms. Despite this sustained scholarly inquiry, there exists a notable research gap regarding the potential synergy achievable… ▽ More

    Submitted 23 June, 2024; originally announced June 2024.

  25. arXiv:2406.14483  [pdf, other

    cs.LG

    Valid Error Bars for Neural Weather Models using Conformal Prediction

    Authors: Vignesh Gopakumar, Joel Oskarrson, Ander Gray, Lorenzo Zanisi, Stanislas Pamela, Daniel Giles, Matt Kusner, Marc Deisenroth

    Abstract: Neural weather models have shown immense potential as inexpensive and accurate alternatives to physics-based models. However, most models trained to perform weather forecasting do not quantify the uncertainty associated with their forecasts. This limits the trust in the model and the usefulness of the forecasts. In this work we construct and formalise a conformal prediction framework as a post-pro… ▽ More

    Submitted 20 June, 2024; originally announced June 2024.

  26. arXiv:2405.02350  [pdf, ps, other

    cs.LG cs.AI

    What makes Models Compositional? A Theoretical View: With Supplement

    Authors: Parikshit Ram, Tim Klinger, Alexander G. Gray

    Abstract: Compositionality is thought to be a key component of language, and various compositional benchmarks have been developed to empirically probe the compositional generalization of existing sequence processing models. These benchmarks often highlight failures of existing models, but it is not clear why these models fail in this way. In this paper, we seek to theoretically understand the role the compo… ▽ More

    Submitted 2 May, 2024; originally announced May 2024.

    Comments: Extended version of the original IJCAI 2024 paper with detailed supplementary materials (27 pages, 7 figures)

  27. arXiv:2403.16887  [pdf

    cs.DL

    ChatGPT "contamination": estimating the prevalence of LLMs in the scholarly literature

    Authors: Andrew Gray

    Abstract: The use of ChatGPT and similar Large Language Model (LLM) tools in scholarly communication and academic publishing has been widely discussed since they became easily accessible to a general audience in late 2022. This study uses keywords known to be disproportionately present in LLM-generated text to provide an overall estimate for the prevalence of LLM-assisted writing in the scholarly literature… ▽ More

    Submitted 25 March, 2024; originally announced March 2024.

    Comments: 12 pages, 6 figures

  28. arXiv:2402.13440  [pdf, other

    cs.AI cs.NE

    A Neuro-Symbolic Approach to Multi-Agent RL for Interpretability and Probabilistic Decision Making

    Authors: Chitra Subramanian, Miao Liu, Naweed Khan, Jonathan Lenchner, Aporva Amarnath, Sarathkrishna Swaminathan, Ryan Riegel, Alexander Gray

    Abstract: Multi-agent reinforcement learning (MARL) is well-suited for runtime decision-making in optimizing the performance of systems where multiple agents coexist and compete for shared resources. However, applying common deep learning-based MARL solutions to real-world problems suffers from issues of interpretability, sample efficiency, partial observability, etc. To address these challenges, we present… ▽ More

    Submitted 20 February, 2024; originally announced February 2024.

    ACM Class: I.2.6

  29. arXiv:2311.05967  [pdf, other

    physics.plasm-ph cs.LG

    Plasma Surrogate Modelling using Fourier Neural Operators

    Authors: Vignesh Gopakumar, Stanislas Pamela, Lorenzo Zanisi, Zongyi Li, Ander Gray, Daniel Brennand, Nitesh Bhatia, Gregory Stathopoulos, Matt Kusner, Marc Peter Deisenroth, Anima Anandkumar, JOREK Team, MAST Team

    Abstract: Predicting plasma evolution within a Tokamak reactor is crucial to realizing the goal of sustainable fusion. Capabilities in forecasting the spatio-temporal evolution of plasma rapidly and accurately allow us to quickly iterate over design and control strategies on current Tokamak devices and future reactors. Modelling plasma evolution using numerical solvers is often expensive, consuming many hou… ▽ More

    Submitted 18 June, 2024; v1 submitted 10 November, 2023; originally announced November 2023.

    Journal ref: Nucl. Fusion 64 056025 (2024)

  30. arXiv:2309.16467  [pdf, other

    cs.LG

    Compositional Program Generation for Few-Shot Systematic Generalization

    Authors: Tim Klinger, Luke Liu, Soham Dan, Maxwell Crouse, Parikshit Ram, Alexander Gray

    Abstract: Compositional generalization is a key ability of humans that enables us to learn new concepts from only a handful examples. Neural machine learning models, including the now ubiquitous Transformers, struggle to generalize in this way, and typically require thousands of examples of a concept during training in order to generalize meaningfully. This difference in ability between humans and artificia… ▽ More

    Submitted 18 January, 2024; v1 submitted 28 September, 2023; originally announced September 2023.

    Comments: 7 pages of text with 1 page of references

  31. arXiv:2308.13292  [pdf, other

    cs.LG cs.CY cs.IR

    A Bayesian Active Learning Approach to Comparative Judgement

    Authors: Andy Gray, Alma Rahat, Tom Crick, Stephen Lindsay

    Abstract: Assessment is a crucial part of education. Traditional marking is a source of inconsistencies and unconscious bias, placing a high cognitive load on the assessors. An approach to address these issues is comparative judgement (CJ). In CJ, the assessor is presented with a pair of items and is asked to select the better one. Following a series of comparisons, a rank is derived using a ranking model,… ▽ More

    Submitted 25 August, 2023; originally announced August 2023.

    Comments: 16 pages

  32. arXiv:2307.02689  [pdf, other

    cs.CL

    Learning Symbolic Rules over Abstract Meaning Representations for Textual Reinforcement Learning

    Authors: Subhajit Chaudhury, Sarathkrishna Swaminathan, Daiki Kimura, Prithviraj Sen, Keerthiram Murugesan, Rosario Uceda-Sosa, Michiaki Tatsubori, Achille Fokoue, Pavan Kapanipathi, Asim Munawar, Alexander Gray

    Abstract: Text-based reinforcement learning agents have predominantly been neural network-based models with embeddings-based representation, learning uninterpretable policies that often do not generalize well to unseen games. On the other hand, neuro-symbolic methods, specifically those that leverage an intermediate formal representation, are gaining significant attention in language understanding tasks. Th… ▽ More

    Submitted 5 July, 2023; originally announced July 2023.

    Comments: ACL 2023

  33. arXiv:2306.15041  [pdf

    q-bio.QM cs.DB

    A Comparison of Neuroelectrophysiology Databases

    Authors: Priyanka Subash, Alex Gray, Misque Boswell, Samantha L. Cohen, Rachael Garner, Sana Salehi, Calvary Fisher, Samuel Hobel, Satrajit Ghosh, Yaroslav Halchenko, Benjamin Dichter, Russell A. Poldrack, Chris Markiewicz, Dora Hermes, Arnaud Delorme, Scott Makeig, Brendan Behan, Alana Sparks, Stephen R Arnott, Zhengjia Wang, John Magnotti, Michael S. Beauchamp, Nader Pouratian, Arthur W. Toga, Dominique Duncan

    Abstract: As data sharing has become more prevalent, three pillars - archives, standards, and analysis tools - have emerged as critical components in facilitating effective data sharing and collaboration. This paper compares four freely available intracranial neuroelectrophysiology data repositories: Data Archive for the BRAIN Initiative (DABI), Distributed Archives for Neurophysiology Data Integration (DAN… ▽ More

    Submitted 30 August, 2023; v1 submitted 26 June, 2023; originally announced June 2023.

    Comments: 22 pages, 6 figures, 5 tables

  34. Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise?

    Authors: Jaromir Savelka, Kevin D. Ashley, Morgan A Gray, Hannes Westermann, Huihui Xu

    Abstract: We evaluated the capability of generative pre-trained transformers~(GPT-4) in analysis of textual data in tasks that require highly specialized domain expertise. Specifically, we focused on the task of analyzing court opinions to interpret legal concepts. We found that GPT-4, prompted with annotation guidelines, performs on par with well-trained law student annotators. We observed that, with a rel… ▽ More

    Submitted 24 June, 2023; originally announced June 2023.

    Journal ref: ITiCSE 2023: Proceedings of the 2023 Conference on Innovation and Technology in Computer Science Education V. 1. June 2023. Pages 117 - 123

  35. arXiv:2306.10452  [pdf, other

    cs.CL

    MISMATCH: Fine-grained Evaluation of Machine-generated Text with Mismatch Error Types

    Authors: Keerthiram Murugesan, Sarathkrishna Swaminathan, Soham Dan, Subhajit Chaudhury, Chulaka Gunasekara, Maxwell Crouse, Diwakar Mahajan, Ibrahim Abdelaziz, Achille Fokoue, Pavan Kapanipathi, Salim Roukos, Alexander Gray

    Abstract: With the growing interest in large language models, the need for evaluating the quality of machine text compared to reference (typically human-generated) text has become focal attention. Most recent works focus either on task-specific evaluation metrics or study the properties of machine-generated text captured by the existing metrics. In this work, we propose a new evaluation scheme to model huma… ▽ More

    Submitted 17 June, 2023; originally announced June 2023.

    Comments: Accepted at ACL 2023 (ACL Findings Long)

  36. arXiv:2306.09525  [pdf, other

    cs.CL cs.AI

    Explaining Legal Concepts with Augmented Large Language Models (GPT-4)

    Authors: Jaromir Savelka, Kevin D. Ashley, Morgan A. Gray, Hannes Westermann, Huihui Xu

    Abstract: Interpreting the meaning of legal open-textured terms is a key task of legal professionals. An important source for this interpretation is how the term was applied in previous court cases. In this paper, we evaluate the performance of GPT-4 in generating factually accurate, clear and relevant explanations of terms in legislation. We compare the performance of a baseline setup, where GPT-4 is direc… ▽ More

    Submitted 22 June, 2023; v1 submitted 15 June, 2023; originally announced June 2023.

  37. arXiv:2305.20018  [pdf, other

    cs.CL cs.AI

    Scalable Learning of Latent Language Structure With Logical Offline Cycle Consistency

    Authors: Maxwell Crouse, Ramon Astudillo, Tahira Naseem, Subhajit Chaudhury, Pavan Kapanipathi, Salim Roukos, Alexander Gray

    Abstract: We introduce Logical Offline Cycle Consistency Optimization (LOCCO), a scalable, semi-supervised method for training a neural semantic parser. Conceptually, LOCCO can be viewed as a form of self-learning where the semantic parser being trained is used to generate annotations for unlabeled text that are then used as new supervision. To increase the quality of annotations, our method utilizes a coun… ▽ More

    Submitted 31 May, 2023; originally announced May 2023.

  38. arXiv:2305.15022  [pdf, other

    stat.ML cs.LG

    Hierarchical clustering with dot products recovers hidden tree structure

    Authors: Annie Gray, Alexander Modell, Patrick Rubin-Delanchy, Nick Whiteley

    Abstract: In this paper we offer a new perspective on the well established agglomerative clustering algorithm, focusing on recovery of hierarchical structure. We recommend a simple variant of the standard algorithm, in which clusters are merged by maximum average dot product and not, for example, by minimum distance or within-cluster variance. We demonstrate that the tree output by this algorithm provides a… ▽ More

    Submitted 1 March, 2024; v1 submitted 24 May, 2023; originally announced May 2023.

  39. arXiv:2301.10414  [pdf, ps, other

    cs.IT cs.LO

    Towards a Unification of Logic and Information Theory

    Authors: Luis A. Lastras, Barry Trager, Jonathan Lenchner, Wojtek Szpankowski, Chai Wah Wu, Mark Squillante, Ron Fagin, Alex Gray

    Abstract: Today, the vast majority of the world's digital information is represented using the fundamental assumption, introduced by Claude Shannon in 1948, that ``...the semantic aspects of communication are irrelevant to the engineering problem (of the design of communication systems)...''. Consider, nonetheless, the observation that we often combine a message with other information in order to deduce new… ▽ More

    Submitted 12 September, 2025; v1 submitted 25 January, 2023; originally announced January 2023.

  40. arXiv:2301.05131  [pdf, other

    cs.LG

    Toward Theoretical Guidance for Two Common Questions in Practical Cross-Validation based Hyperparameter Selection

    Authors: Parikshit Ram, Alexander G. Gray, Horst C. Samulowitz, Gregory Bramble

    Abstract: We show, to our knowledge, the first theoretical treatments of two common questions in cross-validation based hyperparameter selection: (1) After selecting the best hyperparameter using a held-out set, we train the final model using {\em all} of the training data -- since this may or may not improve future generalization error, should one do this? (2) During optimization such as via SGD (stochasti… ▽ More

    Submitted 12 January, 2023; originally announced January 2023.

    Comments: Extended version of the paper appearing at the SIAM International Conference on Data Mining 2023 (SDM23)

  41. arXiv:2208.11665  [pdf, other

    stat.ME cs.LG stat.ML

    Statistical exploration of the Manifold Hypothesis

    Authors: Nick Whiteley, Annie Gray, Patrick Rubin-Delanchy

    Abstract: The Manifold Hypothesis is a widely accepted tenet of Machine Learning which asserts that nominally high-dimensional data are in fact concentrated near a low-dimensional manifold, embedded in high-dimensional space. This phenomenon is observed empirically in many real world situations, has led to development of a wide range of statistical methods in the last few decades, and has been suggested as… ▽ More

    Submitted 21 March, 2025; v1 submitted 24 August, 2022; originally announced August 2022.

    MSC Class: 62R20; 62R40; 62G05; 62G20; 62R07; 62-08; 62H25; 62H30

  42. arXiv:2204.01805  [pdf

    cs.HC

    Using Elo Rating as a Metric for Comparative Judgement in Educational Assessment

    Authors: Andy Gray, Alma Rahat, Tom Crick, Stephen Lindsay, Darren Wallace

    Abstract: Marking and feedback are essential features of teaching and learning, across the overwhelming majority of educational settings and contexts. However, it can take a great deal of time and effort for teachers to mark assessments, and to provide useful feedback to the students. Furthermore, it also creates a significant cognitive load on the assessors, especially in ensuring fairness and equity. Ther… ▽ More

    Submitted 4 April, 2022; originally announced April 2022.

    Comments: 12 pages, 4 figures, one table, pre-review version

  43. arXiv:2201.05793  [pdf, other

    cs.CL cs.AI

    A Benchmark for Generalizable and Interpretable Temporal Question Answering over Knowledge Bases

    Authors: Sumit Neelam, Udit Sharma, Hima Karanam, Shajith Ikbal, Pavan Kapanipathi, Ibrahim Abdelaziz, Nandana Mihindukulasooriya, Young-Suk Lee, Santosh Srivastava, Cezar Pendus, Saswati Dana, Dinesh Garg, Achille Fokoue, G P Shrivatsa Bhargav, Dinesh Khandelwal, Srinivas Ravishankar, Sairam Gurajada, Maria Chang, Rosario Uceda-Sosa, Salim Roukos, Alexander Gray, Guilherme Lima, Ryan Riegel, Francois Luus, L Venkata Subramaniam

    Abstract: Knowledge Base Question Answering (KBQA) tasks that involve complex reasoning are emerging as an important research direction. However, most existing KBQA datasets focus primarily on generic multi-hop reasoning over explicit facts, largely ignoring other reasoning types such as temporal, spatial, and taxonomic reasoning. In this paper, we present a benchmark dataset for temporal reasoning, TempQA-… ▽ More

    Submitted 15 January, 2022; originally announced January 2022.

    Comments: 7 pages, 2 figures, 7 tables. arXiv admin note: substantial text overlap with arXiv:2109.13430

  44. A Simple Standard for Sharing Ontological Mappings (SSSOM)

    Authors: Nicolas Matentzoglu, James P. Balhoff, Susan M. Bello, Chris Bizon, Matthew Brush, Tiffany J. Callahan, Christopher G Chute, William D. Duncan, Chris T. Evelo, Davera Gabriel, John Graybeal, Alasdair Gray, Benjamin M. Gyori, Melissa Haendel, Henriette Harmse, Nomi L. Harris, Ian Harrow, Harshad Hegde, Amelia L. Hoyt, Charles T. Hoyt, Dazhi Jiao, Ernesto Jiménez-Ruiz, Simon Jupp, Hyeongsik Kim, Sebastian Koehler , et al. (19 additional authors not shown)

    Abstract: Despite progress in the development of standards for describing and exchanging scientific information, the lack of easy-to-use standards for mapping between different representations of the same or similar objects in different databases poses a major impediment to data integration and interoperability. Mappings often lack the metadata needed to be correctly interpreted and applied. For example, ar… ▽ More

    Submitted 13 December, 2021; originally announced December 2021.

    Comments: Corresponding author: Christopher J. Mungall <cjmungall@lbl.gov>

  45. arXiv:2112.03324  [pdf, other

    cs.AI cs.LG cs.LO cs.SC

    Neuro-Symbolic Inductive Logic Programming with Logical Neural Networks

    Authors: Prithviraj Sen, Breno W. S. R. de Carvalho, Ryan Riegel, Alexander Gray

    Abstract: Recent work on neuro-symbolic inductive logic programming has led to promising approaches that can learn explanatory rules from noisy, real-world data. While some proposals approximate logical operators with differentiable operators from fuzzy or real-valued logic that are parameter-free thus diminishing their capacity to fit the data, other approaches are only loosely based on logic making it dif… ▽ More

    Submitted 6 December, 2021; originally announced December 2021.

  46. arXiv:2110.10973  [pdf, other

    cs.AI cs.CL cs.LG cs.RO

    LOA: Logical Optimal Actions for Text-based Interaction Games

    Authors: Daiki Kimura, Subhajit Chaudhury, Masaki Ono, Michiaki Tatsubori, Don Joven Agravante, Asim Munawar, Akifumi Wachi, Ryosuke Kohita, Alexander Gray

    Abstract: We present Logical Optimal Actions (LOA), an action decision architecture of reinforcement learning applications with a neuro-symbolic framework which is a combination of neural network and symbolic knowledge acquisition approach for natural language interaction games. The demonstration for LOA experiments consists of a web-based interactive platform for text-based games and visualization for acqu… ▽ More

    Submitted 21 October, 2021; originally announced October 2021.

    Comments: ACL-IJCNLP 2021 (demo paper)

  47. arXiv:2110.10963  [pdf, other

    cs.AI cs.CL cs.LG cs.RO

    Neuro-Symbolic Reinforcement Learning with First-Order Logic

    Authors: Daiki Kimura, Masaki Ono, Subhajit Chaudhury, Ryosuke Kohita, Akifumi Wachi, Don Joven Agravante, Michiaki Tatsubori, Asim Munawar, Alexander Gray

    Abstract: Deep reinforcement learning (RL) methods often require many trials before convergence, and no direct interpretability of trained policies is provided. In order to achieve fast convergence and interpretability for the policy in RL, we propose a novel RL method for text-based games with a recent neuro-symbolic framework called Logical Neural Network, which can learn symbolic and interpretable rules… ▽ More

    Submitted 21 October, 2021; originally announced October 2021.

    Comments: EMNLP 2021 (main conference)

  48. arXiv:2110.01295  [pdf, other

    cs.CL

    SPaR.txt, a cheap Shallow Parsing approach for Regulatory texts

    Authors: Ruben Kruiper, Ioannis Konstas, Alasdair Gray, Farhad Sadeghineko, Richard Watson, Bimal Kumar

    Abstract: Automated Compliance Checking (ACC) systems aim to semantically parse building regulations to a set of rules. However, semantic parsing is known to be hard and requires large amounts of training data. The complexity of creating such training data has led to research that focuses on small sub-tasks, such as shallow parsing or the extraction of a limited subset of rules. This study introduces a shal… ▽ More

    Submitted 4 October, 2021; originally announced October 2021.

    Comments: To be published in the NLLP workshop at EMNLP 2021, 9 pages (15 including reference and appendices). For the ScotReg corpus, SPaR.txt dataset and code see: http://github.com/rubenkruiper/SPaR.txt

  49. arXiv:2109.13430  [pdf, other

    cs.CL cs.AI

    SYGMA: System for Generalizable Modular Question Answering OverKnowledge Bases

    Authors: Sumit Neelam, Udit Sharma, Hima Karanam, Shajith Ikbal, Pavan Kapanipathi, Ibrahim Abdelaziz, Nandana Mihindukulasooriya, Young-Suk Lee, Santosh Srivastava, Cezar Pendus, Saswati Dana, Dinesh Garg, Achille Fokoue, G P Shrivatsa Bhargav, Dinesh Khandelwal, Srinivas Ravishankar, Sairam Gurajada, Maria Chang, Rosario Uceda-Sosa, Salim Roukos, Alexander Gray, Guilherme LimaRyan Riegel, Francois Luus, L Venkata Subramaniam

    Abstract: Knowledge Base Question Answering (KBQA) tasks that in-volve complex reasoning are emerging as an important re-search direction. However, most KBQA systems struggle withgeneralizability, particularly on two dimensions: (a) acrossmultiple reasoning types where both datasets and systems haveprimarily focused on multi-hop reasoning, and (b) across mul-tiple knowledge bases, where KBQA approaches are… ▽ More

    Submitted 27 September, 2021; originally announced September 2021.

  50. arXiv:2109.12240  [pdf, other

    cs.AI cs.LO

    Logical Credal Networks

    Authors: Haifeng Qian, Radu Marinescu, Alexander Gray, Debarun Bhattacharjya, Francisco Barahona, Tian Gao, Ryan Riegel, Pravinda Sahu

    Abstract: This paper introduces Logical Credal Networks, an expressive probabilistic logic that generalizes many prior models that combine logic and probability. Given imprecise information represented by probability bounds and conditional probability bounds of logic formulas, this logic specifies a set of probability distributions over all interpretations. On the one hand, our approach allows propositional… ▽ More

    Submitted 24 September, 2021; originally announced September 2021.