Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 85 results for author: Anderson, C

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.22807  [pdf, ps, other

    cs.SE cs.CL

    The Best Programming Language for Tokenmaxxing: An Investigation of Coding Agent Behavior Across Programming Languages

    Authors: Zixuan Wu, Carolyn Jane Anderson, Arjun Guha

    Abstract: Although coding agents are now very effective in a variety of programming languages, this paper first shows that the cost (in tokens) can very significantly by programming language. We evaluate five recent models on programming problems in Python, Java, Rust, and OCaml. We carefully control for problem difficulty, and show that there can be stark variation in token consumption that is consistent a… ▽ More

    Submitted 24 July, 2026; originally announced July 2026.

  2. arXiv:2605.17618  [pdf, ps, other

    cs.AI

    Prediction of Challenging Behaviors Associated with Profound Autism in a Classroom Setting Using Wearable Sensors

    Authors: Yadhu Kartha, Conor Anderson, Jenny Foster, Theresa Hamlin, Johanna Lantz, Ryan Lay, Juergen Hahn, Gari D. Clifford, Hyeokhyen Kwon

    Abstract: Autism Spectrum Disorder (ASD) is characterized by challenges with social interaction and communication and by restricted or repetitive patterns of thought and behavior, with significant variability in presentation. Approximately a quarter of children with ASD are classified as having profound autism, who often exhibit challenging behaviors, such as self-injurious behavior, aggression, elopement,… ▽ More

    Submitted 20 May, 2026; v1 submitted 17 May, 2026; originally announced May 2026.

  3. arXiv:2605.15547  [pdf

    cs.MS

    Correctly Rounded Functions For Vector Applications: A Performance Study

    Authors: Cristina Anderson, Marius Cornea, Andrey Stepin, Mihai Tudor Panu

    Abstract: Following recent interest in correctly rounded math library functions (as currently recommended by the IEEE 754 standard), we have designed several SIMD algorithms for one-input single precision functions and integrated them into our CPU math library; these will form the core of the first correctly rounded vector math library, to be available to users in mid-2026. To take advantage of the cross-pl… ▽ More

    Submitted 14 May, 2026; originally announced May 2026.

    Comments: 5 pages, 2 figures, 4 tables

  4. arXiv:2605.09063  [pdf, ps, other

    cs.CL

    Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

    Authors: Guijin Son, Seungone Kim, Catherine Arnett, Hyunwoo Ko, Hyein Lee, Hyeonah Kang, Jiang Longxi, Jin Yun, JungYup Lee, Kyungmin Lee, Sam Yoosuk Kim, Sang Park, Seunghyeok Hong, SeungJae Lee, Seungyeop Yi, Shinae Shin, SunHye Bok, Sunyoung Shin, Yonghoon Ji, Youngtaek Kim, Hanearl Jung, Akari Asai, Graham Neubig, Sean Welleck, Youngjae Yu , et al. (51 additional authors not shown)

    Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challenging target for measuring LLM reasoning. Whereas olympiad-style problems measure step-by-step reasoning alone, research-level problems use such reasoning to advance the frontier of mathematical knowledge itself, emerging as a compelling alternative.… ▽ More

    Submitted 19 May, 2026; v1 submitted 9 May, 2026; originally announced May 2026.

    Comments: Under review, For questions or model-evaluation requests, contact $guijin.son@snu.ac.kr$

  5. arXiv:2605.02316  [pdf, ps, other

    cs.CV cs.LG

    Open-access model for detecting openly dumped dispersed municipal solid waste from crowdsourced UAV imagery in Sub-Saharan Africa

    Authors: Steffen Knoblauch, Ram Kumar Muthusamy, Luis M. A. Bettencourt, Costas Velis, Pierre Chrzanowski, Edward Charles Anderson, Pete Masters, Innocent Maholi, Antonio Inguane, Levi Szamek, Alexander Zipf

    Abstract: Managing municipal solid waste in rapidly urbanizing Sub-Saharan Africa remains challenging due to dispersed informal dumping and limited high-resolution datasets for spatial monitoring. We present an open-access deep learning model for automated detection of openly dumped dispersed solid waste via crowdsourced UAV imagery, trained and evaluated across 29 regions in 10 countries, encompassing dive… ▽ More

    Submitted 4 May, 2026; originally announced May 2026.

  6. arXiv:2603.03206  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Understanding and Mitigating Dataset Corruption in LLM Steering

    Authors: Cullen Anderson, Narmeen Oozeer, Foad Namjoo, Remy Ogasawara, Amirali Abdullah, Jeff M. Phillips

    Abstract: Contrastive steering has been shown as a simple and effective method to adjust the generative behavior of LLMs at inference time. It uses examples of prompt responses with and without a trait to identify a direction in an intermediate activation layer, and then shifts activations in this 1-dimensional subspace. However, despite its growing use in AI safety applications, the robustness of contrasti… ▽ More

    Submitted 3 March, 2026; originally announced March 2026.

  7. arXiv:2601.18963  [pdf, ps, other

    cs.RO cs.AI

    Fauna Sprout: A lightweight, approachable, developer-ready humanoid robot

    Authors: Fauna Robotics, :, Diego Aldarondo, Ana Pervan, Daniel Corbalan, Dave Petrillo, Bolun Dai, Aadhithya Iyer, Nina Mortensen, Erik Pearson, Sridhar Pandian Arunachalam, Emma Reznick, David Weis, Jacob Davison, Samuel Patterson, Tess Carella, Michael Suguitan, David Ye, Oswaldo Ferro, Nilesh Suriyarachchi, Spencer Ling, Erik Su, Daniel Giebisch, Peter Traver, Sam Fonseca , et al. (26 additional authors not shown)

    Abstract: Recent advances in learned control, large-scale simulation, and generative models have accelerated progress toward general-purpose robotic controllers, yet the field still lacks platforms suitable for safe, expressive, long-term deployment in human environments. Most existing humanoids are either closed industrial systems or academic prototypes that are difficult to deploy and operate around peopl… ▽ More

    Submitted 26 January, 2026; originally announced January 2026.

  8. arXiv:2601.13358   

    cs.AI cs.LG

    The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models

    Authors: Samuel Cyrenius Anderson

    Abstract: Scale does not uniformly improve reasoning - it restructures it. Analyzing 25,000+ chain-of-thought trajectories across four domains (Law, Science, Code, Math) and two scales (8B, 70B parameters), we discover that neural scaling laws trigger domain-specific phase transitions rather than uniform capability gains. Legal reasoning undergoes Crystallization: 45% collapse in representational dimensiona… ▽ More

    Submitted 30 March, 2026; v1 submitted 19 January, 2026; originally announced January 2026.

    Comments: The theoretical framework has been shown to be wrong and should not be followed for future research direction

    ACM Class: I.2.6; I.2.7

  9. arXiv:2508.04865  [pdf, ps, other

    cs.LG cs.PL

    Agnostics: Learning to Code in Any Programming Language via Reinforcement with a Universal Learning Environment

    Authors: Aleksander Boruch-Gruszecki, Yangtian Zi, Zixuan Wu, Tejas Oberoi, Carolyn Jane Anderson, Joydeep Biswas, Arjun Guha

    Abstract: Large language models (LLMs) already excel at writing code in high-resource languages such as Python and JavaScript, yet stumble on low-resource languages that remain essential to science and engineering. Besides the obvious shortage of pre-training data, post-training itself is a bottleneck: every new language seems to require new datasets, test harnesses, and reinforcement-learning (RL) infrastr… ▽ More

    Submitted 23 March, 2026; v1 submitted 6 August, 2025; originally announced August 2025.

    Comments: 30 pages, 19 figures. Accepted at ICLR 2026. For data, code, artifacts, see https://agnostics.abgru.me

  10. arXiv:2505.00268  [pdf, ps, other

    cs.CL cs.AI

    Consistency in Language Models: Current Landscape, Challenges, and Future Directions

    Authors: Jekaterina Novikova, Carol Anderson, Borhane Blili-Hamelin, Domenic Rosati, Subhabrata Majumdar

    Abstract: The hallmark of effective language use lies in consistency: expressing similar meanings in similar contexts and avoiding contradictions. While human communication naturally demonstrates this principle, state-of-the-art language models (LMs) struggle to maintain reliable consistency across task- and domain-specific applications. Here we examine the landscape of consistency research in LMs, analyze… ▽ More

    Submitted 11 July, 2025; v1 submitted 30 April, 2025; originally announced May 2025.

    Comments: Accepted in ICML 2025 Workshop on Reliable and Responsible Foundation Models

  11. "I Would Have Written My Code Differently'': Beginners Struggle to Understand LLM-Generated Code

    Authors: Yangtian Zi, Luisa Li, Arjun Guha, Carolyn Jane Anderson, Molly Q Feldman

    Abstract: Large language models (LLMs) are being increasingly adopted for programming work. Prior work shows that while LLMs accelerate task completion for professional programmers, beginning programmers struggle to prompt models effectively. However, prompting is just half of the code generation process -- when code is generated, it must be read, evaluated, and integrated (or rejected). How accessible are… ▽ More

    Submitted 26 April, 2025; originally announced April 2025.

    Comments: To appear in 33rd ACM International Conference on the Foundations of Software Engineering (FSE Companion '25), June 23-28, 2025, Trondheim, Norway

  12. arXiv:2502.16329  [pdf

    cs.LG

    Generalization is not a universal guarantee: Estimating similarity to training data with an ensemble out-of-distribution metric

    Authors: W. Max Schreyer, Christopher Anderson, Reid F. Thompson

    Abstract: Failure of machine learning models to generalize to new data is a core problem limiting the reliability of AI systems, partly due to the lack of simple and robust methods for comparing new data to the original training dataset. We propose a standardized approach for assessing data similarity in a model-agnostic manner by constructing a supervised autoencoder for generalizability estimation (SAGE).… ▽ More

    Submitted 25 February, 2025; v1 submitted 22 February, 2025; originally announced February 2025.

    Comments: 10 pages, 5 figures

  13. arXiv:2502.11324  [pdf, other

    stat.ML cs.LG

    Robust High-Dimensional Mean Estimation With Low Data Size, an Empirical Study

    Authors: Cullen Anderson, Jeff M. Phillips

    Abstract: Robust statistics aims to compute quantities to represent data where a fraction of it may be arbitrarily corrupted. The most essential statistic is the mean, and in recent years, there has been a flurry of theoretical advancement for efficiently estimating the mean in high dimensions on corrupted data. While several algorithms have been proposed that achieve near-optimal error, they all rely on la… ▽ More

    Submitted 16 February, 2025; originally announced February 2025.

    Journal ref: Transactions on Machine Learning Research (TMLR), February 2025, ISSN: 2835-8856

  14. arXiv:2502.01584  [pdf, ps, other

    cs.AI cs.LG

    ReasoningWeekly: A General Knowledge and Verbal Reasoning Challenge for Large Language Models

    Authors: Zixuan Wu, Francesca Lucchetti, Aleksander Boruch-Gruszecki, Jingmiao Zhao, Carolyn Jane Anderson, Joydeep Biswas, Federico Cassano, Arjun Guha

    Abstract: Existing benchmarks for frontier models often test specialized, "PhD-level" knowledge that is difficult for non-experts to grasp. In contrast, we present a benchmark with 613 problems based on the NPR Sunday Puzzle Challenge that requires only general knowledge. Our benchmark is challenging for both humans and models; however correct solutions are easy to verify, and models' mistakes are easy to s… ▽ More

    Submitted 26 November, 2025; v1 submitted 3 February, 2025; originally announced February 2025.

  15. A generalizable 3D framework and model for self-supervised learning in medical imaging

    Authors: Tony Xu, Sepehr Hosseini, Chris Anderson, Anthony Rinaldi, Rahul G. Krishnan, Anne L. Martel, Maged Goubran

    Abstract: Current self-supervised learning methods for 3D medical imaging rely on simple pretext formulations and organ- or modality-specific datasets, limiting their generalizability and scalability. We present 3DINO, a cutting-edge SSL method adapted to 3D datasets, and use it to pretrain 3DINO-ViT: a general-purpose medical imaging model, on an exceptionally large, multimodal, and multi-organ dataset of… ▽ More

    Submitted 8 June, 2026; v1 submitted 20 January, 2025; originally announced January 2025.

    Comments: Published in npj Digital Medicine

  16. arXiv:2410.19792  [pdf, other

    cs.CY cs.LG

    Substance Beats Style: Why Beginning Students Fail to Code with LLMs

    Authors: Francesca Lucchetti, Zixuan Wu, Arjun Guha, Molly Q Feldman, Carolyn Jane Anderson

    Abstract: Although LLMs are increasing the productivity of professional programmers, existing work shows that beginners struggle to prompt LLMs to solve text-to-code tasks. Why is this the case? This paper explores two competing hypotheses about the cause of student-LLM miscommunication: (1) students simply lack the technical vocabulary needed to write good prompts, and (2) students do not understand the ex… ▽ More

    Submitted 15 October, 2024; originally announced October 2024.

  17. arXiv:2408.16131  [pdf, other

    cs.CL

    Evaluating Computational Representations of Character: An Austen Character Similarity Benchmark

    Authors: Funing Yang, Carolyn Jane Anderson

    Abstract: Several systems have been developed to extract information about characters to aid computational analysis of English literature. We propose character similarity grouping as a holistic evaluation task for these pipelines. We present AustenAlike, a benchmark suite of character similarities in Jane Austen's novels. Our benchmark draws on three notions of character similarity: a structurally defined n… ▽ More

    Submitted 28 August, 2024; originally announced August 2024.

  18. arXiv:2408.05894  [pdf, ps, other

    cs.CV cs.CL

    GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models

    Authors: Zixuan Wu, Yoolim Kim, Carolyn Jane Anderson

    Abstract: Vision-Language Models (VLMs) building upon the foundation of powerful large language models have made rapid progress in reasoning across visual and textual data. While VLMs perform well on vision tasks that they are trained on, our results highlight key challenges in abstract pattern recognition. We present GlyphPattern, a 954 item dataset that pairs 318 human-written descriptions of visual patte… ▽ More

    Submitted 24 June, 2025; v1 submitted 11 August, 2024; originally announced August 2024.

    Journal ref: Findings of the ACL 2025

  19. arXiv:2407.21691  [pdf, other

    cs.CV

    Explainable Artificial Intelligence for Quantifying Interfering and High-Risk Behaviors in Autism Spectrum Disorder in a Real-World Classroom Environment Using Privacy-Preserving Video Analysis

    Authors: Barun Das, Conor Anderson, Tania Villavicencio, Johanna Lantz, Jenny Foster, Theresa Hamlin, Ali Bahrami Rad, Gari D. Clifford, Hyeokhyen Kwon

    Abstract: Rapid identification and accurate documentation of interfering and high-risk behaviors in ASD, such as aggression, self-injury, disruption, and restricted repetitive behaviors, are important in daily classroom environments for tracking intervention effectiveness and allocating appropriate resources to manage care needs. However, having a staff dedicated solely to observing is costly and uncommon i… ▽ More

    Submitted 31 July, 2024; originally announced July 2024.

  20. arXiv:2407.01757  [pdf, other

    astro-ph.EP astro-ph.IM cs.MA physics.ao-ph physics.geo-ph

    Distributed Instruments for Planetary Surface Science: Scientific Opportunities and Technology Feasibility

    Authors: Federico Rossi, Robert C. Anderson, Saptarshi Bandyopadhyay, Erik Brandon, Ashish Goel, Joshua Vander Hook, Michael Mischna, Michaela Villarreal, Mark Wronkiewicz

    Abstract: In this paper, we assess the scientific promise and technology feasibility of distributed instruments for planetary science. A distributed instrument is an instrument designed to collect spatially and temporally correlated data from multiple networked, geographically distributed point sensors. Distributed instruments are ubiquitous in Earth science, where they are routinely employed for weather an… ▽ More

    Submitted 1 July, 2024; originally announced July 2024.

  21. arXiv:2406.16955  [pdf, other

    eess.SP cs.CV cs.LG

    SRViT: Vision Transformers for Estimating Radar Reflectivity from Satellite Observations at Scale

    Authors: Jason Stock, Kyle Hilburn, Imme Ebert-Uphoff, Charles Anderson

    Abstract: We introduce a transformer-based neural network to generate high-resolution (3km) synthetic radar reflectivity fields at scale from geostationary satellite imagery. This work aims to enhance short-term convective-scale forecasts of high-impact weather events and aid in data assimilation for numerical weather prediction over the United States. Compared to convolutional approaches, which have limite… ▽ More

    Submitted 28 June, 2024; v1 submitted 20 June, 2024; originally announced June 2024.

    Comments: Published as a workshop paper at "Machine Learning for Earth System Modeling", ICML 2024; added acknowledgements and github link

  22. arXiv:2404.18774  [pdf

    cond-mat.supr-con cs.AI

    Self-training superconducting neuromorphic circuits using reinforcement learning rules

    Authors: M. L. Schneider, E. M. Jué, M. R. Pufall, K. Segall, C. W. Anderson

    Abstract: Reinforcement learning algorithms are used in a wide range of applications, from gaming and robotics to autonomous vehicles. In this paper we describe a set of reinforcement learning-based local weight update rules and their implementation in superconducting hardware. Using SPICE circuit simulations, we implement a small-scale neural network with a learning time of order one nanosecond. This netwo… ▽ More

    Submitted 29 April, 2024; originally announced April 2024.

    Comments: 15 pages, 6 figures

    Journal ref: npj Unconventional Computing. 2, 5 (2025)

  23. arXiv:2402.19173  [pdf, other

    cs.SE cs.AI

    StarCoder 2 and The Stack v2: The Next Generation

    Authors: Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo , et al. (41 additional authors not shown)

    Abstract: The BigCode project, an open-scientific collaboration focused on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder2. In partnership with Software Heritage (SWH), we build The Stack v2 on top of the digital commons of their source code archive. Alongside the SWH repositories spanning 619 programming languages, we carefully select other high-quality data… ▽ More

    Submitted 29 February, 2024; originally announced February 2024.

  24. arXiv:2402.01969  [pdf, other

    cs.LG eess.SP

    Simulation-Enhanced Data Augmentation for Machine Learning Pathloss Prediction

    Authors: Ahmed P. Mohamed, Byunghyun Lee, Yaguang Zhang, Max Hollingsworth, C. Robert Anderson, James V. Krogmeier, David J. Love

    Abstract: Machine learning (ML) offers a promising solution to pathloss prediction. However, its effectiveness can be degraded by the limited availability of data. To alleviate these challenges, this paper introduces a novel simulation-enhanced data augmentation method for ML pathloss prediction. Our method integrates synthetic data generated from a cellular coverage simulator and independently collected re… ▽ More

    Submitted 5 February, 2024; v1 submitted 2 February, 2024; originally announced February 2024.

    Comments: 6 pages, 5 figures, Accepted at ICC 2024

  25. How Beginning Programmers and Code LLMs (Mis)read Each Other

    Authors: Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, Molly Q Feldman

    Abstract: Generative AI models, specifically large language models (LLMs), have made strides towards the long-standing goal of text-to-code generation. This progress has invited numerous studies of user interaction. However, less is known about the struggles and strategies of non-experts, for whom each step of the text-to-code problem presents challenges: describing their intent in natural language, evaluat… ▽ More

    Submitted 7 July, 2024; v1 submitted 26 January, 2024; originally announced January 2024.

    Comments: Published in CHI 2024

  26. Segmenting Messy Text: Detecting Boundaries in Text Derived from Historical Newspaper Images

    Authors: Carol Anderson, Phil Crone

    Abstract: Text segmentation, the task of dividing a document into sections, is often a prerequisite for performing additional natural language processing tasks. Existing text segmentation methods have typically been developed and tested using clean, narrative-style text with segments containing distinct topics. Here we consider a challenging text segmentation task: dividing newspaper marriage announcement l… ▽ More

    Submitted 20 December, 2023; originally announced December 2023.

    Comments: 8 pages, 4 figures

    ACM Class: I.2.7; I.7.5

    Journal ref: 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 2021, pp. 5543-5550

  27. arXiv:2312.12450  [pdf, other

    cs.SE cs.AI cs.LG cs.PL

    Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions

    Authors: Federico Cassano, Luisa Li, Akul Sethi, Noah Shinn, Abby Brennan-Jones, Jacob Ginesin, Edward Berman, George Chakhnashvili, Anton Lozhkov, Carolyn Jane Anderson, Arjun Guha

    Abstract: A significant amount of research is focused on developing and evaluating large language models for a variety of code synthesis tasks. These include synthesizing code from natural language, synthesizing tests from code, and synthesizing explanations of code. In contrast, the behavior of instructional code editing with LLMs is understudied. These are tasks in which the model is provided a block of c… ▽ More

    Submitted 23 September, 2024; v1 submitted 10 December, 2023; originally announced December 2023.

  28. arXiv:2310.13304  [pdf, ps, other

    cs.HC

    Exploring Emotional and Social Dynamics in Mobile Usage During Home Confinement

    Authors: Nan Gao, Sam Nolan, Kaixin Ji, Shakila Khan Rumi, Judith Simone Heinisch, Christoph Anderson, Klaus David, Flora D. Salim

    Abstract: Home confinement, a situation experienced by individuals for reasons ranging from medical quarantines, rehabilitation needs, disability accommodations, and remote working, is a common yet impactful aspect of modern life. While essential in various scenarios, confinement within the home environment can profoundly influence mental well-being and digital device usage. Using the COVID-19 lockdown as a… ▽ More

    Submitted 7 July, 2025; v1 submitted 20 October, 2023; originally announced October 2023.

  29. Autonomous Systems' Safety Cases for use in UK Nuclear Environments

    Authors: Christopher R. Anderson, Louise A. Dennis

    Abstract: An overview of the process to develop a safety case for an autonomous robot deployment on a nuclear site in the UK is described and a safety case for a hypothetical robot incorporating AI is presented. This forms a first step towards a deployment, showing what is possible now and what may be possible with development of tools. It forms the basis for further discussion between nuclear site licensee… ▽ More

    Submitted 3 October, 2023; originally announced October 2023.

    Comments: In Proceedings AREA 2023, arXiv:2310.00333

    Journal ref: EPTCS 391, 2023, pp. 83-88

  30. A Large Language Model Approach to Educational Survey Feedback Analysis

    Authors: Michael J. Parker, Caitlin Anderson, Claire Stone, YeaRim Oh

    Abstract: This paper assesses the potential for the large language models (LLMs) GPT-4 and GPT-3.5 to aid in deriving insight from education feedback surveys. Exploration of LLM use cases in education has focused on teaching and learning, with less exploration of capabilities in education feedback analysis. Survey analysis in education involves goals such as finding gaps in curricula or evaluating teachers,… ▽ More

    Submitted 26 June, 2024; v1 submitted 29 September, 2023; originally announced September 2023.

    Journal ref: Int J Artif Intell Educ (2024)

  31. Julia as a unifying end-to-end workflow language on the Frontier exascale system

    Authors: William F. Godoy, Pedro Valero-Lara, Caira Anderson, Katrina W. Lee, Ana Gainaru, Rafael Ferreira da Silva, Jeffrey S. Vetter

    Abstract: We evaluate Julia as a single language and ecosystem paradigm powered by LLVM to develop workflow components for high-performance computing. We run a Gray-Scott, 2-variable diffusion-reaction application using a memory-bound, 7-point stencil kernel on Frontier, the US Department of Energy's first exascale supercomputer. We evaluate the performance, scaling, and trade-offs of (i) the computational… ▽ More

    Submitted 27 September, 2023; v1 submitted 18 September, 2023; originally announced September 2023.

    Comments: 11 pages, 8 figures, accepted at the 18th Workshop on Workflows in Support of Large-Scale Science (WORKS23), IEEE/ACM The International Conference for High Performance Computing, Networking, Storage, and Analysis, SC23

  32. arXiv:2308.09895  [pdf, other

    cs.PL cs.LG

    Knowledge Transfer from High-Resource to Low-Resource Programming Languages for Code LLMs

    Authors: Federico Cassano, John Gouwar, Francesca Lucchetti, Claire Schlesinger, Anders Freeman, Carolyn Jane Anderson, Molly Q Feldman, Michael Greenberg, Abhinav Jangda, Arjun Guha

    Abstract: Over the past few years, Large Language Models of Code (Code LLMs) have started to have a significant impact on programming practice. Code LLMs are also emerging as building blocks for research in programming languages and software engineering. However, Code LLMs produce impressive results on programming languages that are well represented in their training data (e.g., Java, Python, or JavaScript)… ▽ More

    Submitted 21 September, 2024; v1 submitted 18 August, 2023; originally announced August 2023.

  33. arXiv:2307.08692  [pdf, other

    eess.SY cs.LG

    A Multiobjective Reinforcement Learning Framework for Microgrid Energy Management

    Authors: M. Vivienne Liu, Patrick M. Reed, David Gold, Garret Quist, C. Lindsay Anderson

    Abstract: The emergence of microgrids (MGs) has provided a promising solution for decarbonizing and decentralizing the power grid, mitigating the challenges posed by climate change. However, MG operations often involve considering multiple objectives that represent the interests of different stakeholders, leading to potentially complex conflicts. To tackle this issue, we propose a novel multi-objective rein… ▽ More

    Submitted 14 February, 2025; v1 submitted 17 July, 2023; originally announced July 2023.

  34. arXiv:2306.12255  [pdf, other

    cs.CL

    Solving and Generating NPR Sunday Puzzles with Large Language Models

    Authors: Jingmiao Zhao, Carolyn Jane Anderson

    Abstract: We explore the ability of large language models to solve and generate puzzles from the NPR Sunday Puzzle game show using PUZZLEQA, a dataset comprising 15 years of on-air puzzles. We evaluate four large language models using PUZZLEQA, in both multiple choice and free response formats, and explore two prompt engineering techniques to improve free response performance: chain-of-thought reasoning and… ▽ More

    Submitted 21 June, 2023; originally announced June 2023.

    Comments: To appear in the Proceedings of the 14th International Conference on Computational Creativity (ICCC)

  35. arXiv:2306.04556  [pdf, other

    cs.LG cs.HC cs.SE

    StudentEval: A Benchmark of Student-Written Prompts for Large Language Models of Code

    Authors: Hannah McLean Babe, Sydney Nguyen, Yangtian Zi, Arjun Guha, Molly Q Feldman, Carolyn Jane Anderson

    Abstract: Code LLMs are being rapidly deployed and there is evidence that they can make professional programmers more productive. Current benchmarks for code generation measure whether models generate correct programs given an expert prompt. In this paper, we present a new benchmark containing multiple prompts per problem, written by a specific population of non-expert prompters: beginning programmers. Stud… ▽ More

    Submitted 7 June, 2023; originally announced June 2023.

  36. Keep It Simple: Fault Tolerance Evaluation of Federated Learning with Unreliable Clients

    Authors: Victoria Huang, Shaleeza Sohail, Michael Mayo, Tania Lorido Botran, Mark Rodrigues, Chris Anderson, Melanie Ooi

    Abstract: Federated learning (FL), as an emerging artificial intelligence (AI) approach, enables decentralized model training across multiple devices without exposing their local training data. FL has been increasingly gaining popularity in both academia and industry. While research works have been proposed to improve the fault tolerance of FL, the real impact of unreliable devices (e.g., dropping out, misc… ▽ More

    Submitted 16 May, 2023; originally announced May 2023.

  37. arXiv:2305.06161  [pdf, other

    cs.CL cs.AI cs.PL cs.SE

    StarCoder: may the source be with you!

    Authors: Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, Logesh Kumar Umapathi, Jian Zhu , et al. (42 additional authors not shown)

    Abstract: The BigCode community, an open-scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder and StarCoderBase: 15.5B parameter models with 8K context length, infilling capabilities and fast large-batch inference enabled by multi-query attention. StarCoderBase is trained on 1 trillion tokens sourced from The Stack, a large colle… ▽ More

    Submitted 13 December, 2023; v1 submitted 9 May, 2023; originally announced May 2023.

  38. arXiv:2302.08584  [pdf, other

    eess.SP cs.RO eess.SY

    Propagation Measurements and Analyses at 28 GHz via an Autonomous Beam-Steering Platform

    Authors: Bharath Keshavamurthy, Yaguang Zhang, Christopher R. Anderson, Nicolo Michelusi, James V. Krogmeier, David J. Love

    Abstract: This paper details the design of an autonomous alignment and tracking platform to mechanically steer directional horn antennas in a sliding correlator channel sounder setup for 28 GHz V2X propagation modeling. A pan-and-tilt subsystem facilitates uninhibited rotational mobility along the yaw and pitch axes, driven by open-loop servo units and orchestrated via inertial motion controllers. A geo-pos… ▽ More

    Submitted 16 February, 2023; originally announced February 2023.

    Comments: 6 pages, 18 figures, 2 tables; Accepted at IEEE International Conference on Communications (ICC) 2023: Paper #1570867736

    Report number: ICC Paper #1570867736

  39. arXiv:2301.03988  [pdf, other

    cs.SE cs.AI cs.LG

    SantaCoder: don't reach for the stars!

    Authors: Loubna Ben Allal, Raymond Li, Denis Kocetkov, Chenghao Mou, Christopher Akiki, Carlos Munoz Ferrandis, Niklas Muennighoff, Mayank Mishra, Alex Gu, Manan Dey, Logesh Kumar Umapathi, Carolyn Jane Anderson, Yangtian Zi, Joel Lamy Poirier, Hailey Schoelkopf, Sergey Troshin, Dmitry Abulkhanov, Manuel Romero, Michael Lappert, Francesco De Toni, Bernardo García del Río, Qian Liu, Shamik Bose, Urvashi Bhattacharyya, Terry Yue Zhuo , et al. (16 additional authors not shown)

    Abstract: The BigCode project is an open-scientific collaboration working on the responsible development of large language models for code. This tech report describes the progress of the collaboration until December 2022, outlining the current state of the Personally Identifiable Information (PII) redaction pipeline, the experiments conducted to de-risk the model architecture, and the experiments investigat… ▽ More

    Submitted 24 February, 2023; v1 submitted 9 January, 2023; originally announced January 2023.

  40. arXiv:2212.04478  [pdf, other

    physics.ao-ph cs.LG

    An Interpretable Model of Climate Change Using Correlative Learning

    Authors: Charles Anderson, Jason Stock

    Abstract: Determining changes in global temperature and precipitation that may indicate climate change is complicated by annual variations. One approach for finding potential climate change indicators is to train a model that predicts the year from annual means of global temperatures and precipitations. Such data is available from the CMIP6 ensemble of simulations. Here a two-hidden-layer neural network tra… ▽ More

    Submitted 5 December, 2022; originally announced December 2022.

    Comments: NeurIPS 2022 Workshop - Tackling Climate Change with Machine Learning, 4 page limit w/ appendix

  41. arXiv:2210.12185  [pdf, other

    cs.CV cs.AI cs.LG

    Attention-Based Scattering Network for Satellite Imagery

    Authors: Jason Stock, Chuck Anderson

    Abstract: Multi-channel satellite imagery, from stacked spectral bands or spatiotemporal data, have meaningful representations for various atmospheric properties. Combining these features in an effective manner to create a performant and trustworthy model is of utmost importance to forecasters. Neural networks show promise, yet suffer from unintuitive computations, fusion of high-level features, and may be… ▽ More

    Submitted 21 October, 2022; originally announced October 2022.

    Comments: NeurIPS 2022 Workshop - Tackling Climate Change with Machine Learning, 4 page limit w/ appendix

  42. arXiv:2208.08227  [pdf, other

    cs.LG cs.PL

    MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation

    Authors: Federico Cassano, John Gouwar, Daniel Nguyen, Sydney Nguyen, Luna Phipps-Costin, Donald Pinckney, Ming-Ho Yee, Yangtian Zi, Carolyn Jane Anderson, Molly Q Feldman, Arjun Guha, Michael Greenberg, Abhinav Jangda

    Abstract: Large language models have demonstrated the ability to generate both natural language and programming language text. Such models open up the possibility of multi-language code generation: could code generation models generalize knowledge from one language to another? Although contemporary code generation models can generate semantically correct Python code, little is known about their abilities wi… ▽ More

    Submitted 19 December, 2022; v1 submitted 17 August, 2022; originally announced August 2022.

  43. arXiv:2207.03405  [pdf, other

    cs.HC

    Investigating the Effects of Mood & Usage Behaviour on Notification Response Time

    Authors: Judith S. Heinisch, Nan Gao, Christoph Anderson, Shohreh Deldari, Klaus David, Flora Salim

    Abstract: Notifications are one of the most prevailing mechanisms on smartphones and personal computers to convey timely and important information. Despite these benefits, smartphone notifications demand individuals' attention and can cause stress and frustration when delivered at inopportune timings. This paper investigates the effect of individuals' smartphone usage behavior and mood on notification respo… ▽ More

    Submitted 7 July, 2022; originally announced July 2022.

  44. arXiv:2205.06351  [pdf, other

    cs.LG

    Interpretable Climate Change Modeling With Progressive Cascade Networks

    Authors: Charles Anderson, Jason Stock, David Anderson

    Abstract: Typical deep learning approaches to modeling high-dimensional data often result in complex models that do not easily reveal a new understanding of the data. Research in the deep learning field is very actively pursuing new methods to interpret deep neural networks and to reduce their complexity. An approach is described here that starts with linear models and incrementally adds complexity only as… ▽ More

    Submitted 12 May, 2022; originally announced May 2022.

  45. arXiv:2205.03355  [pdf, other

    cs.LG

    Trainable Wavelet Neural Network for Non-Stationary Signals

    Authors: Jason Stock, Chuck Anderson

    Abstract: This work introduces a wavelet neural network to learn a filter-bank specialized to fit non-stationary signals and improve interpretability and performance for digital signal processing. The network uses a wavelet transform as the first layer of a neural network where the convolution is a parameterized function of the complex Morlet wavelet. Experimental results, on both simplified data and atmosp… ▽ More

    Submitted 6 May, 2022; originally announced May 2022.

    Comments: AI for Earth and Space Science Workshop at the International Conference on Learning Representations (ICLR), April, 2022

  46. arXiv:2201.10511  [pdf, other

    eess.IV cs.CV cs.LG

    Initial Investigations Towards Non-invasive Monitoring of Chronic Wound Healing Using Deep Learning and Ultrasound Imaging

    Authors: Maja Schlereth, Daniel Stromer, Yash Mantri, Jason Tsujimoto, Katharina Breininger, Andreas Maier, Caesar Anderson, Pranav S. Garimella, Jesse V. Jokerst

    Abstract: Chronic wounds including diabetic and arterial/venous insufficiency injuries have become a major burden for healthcare systems worldwide. Demographic changes suggest that wound care will play an even bigger role in the coming decades. Predicting and monitoring response to therapy in wound care is currently largely based on visual inspection with little information on the underlying tissue. Thus, t… ▽ More

    Submitted 25 January, 2022; originally announced January 2022.

    Comments: 6 pages, 2 figures, accepted by BVM conference proceedings 2022

  47. Statistical detection of format dialects using the weighted Dowker complex

    Authors: Michael Robinson, Letitia W. Li, Cory Anderson, Steve Huntsman

    Abstract: This paper provides an experimentally validated, probabilistic model of file behavior when consumed by a set of pre-existing parsers. File behavior is measured by way of a standardized set of Boolean "messages" produced as the files are read. By thresholding the posterior probability that a file exhibiting a particular set of messages is from a particular dialect, our model yields a practical clas… ▽ More

    Submitted 20 January, 2022; originally announced January 2022.

    Comments: 15 pages, 11 figures, 5 tables

    MSC Class: 62P30; 55U10 ACM Class: D.3.4

  48. arXiv:2111.15641  [pdf, ps, other

    cs.CL

    Automatic Extraction of Medication Names in Tweets as Named Entity Recognition

    Authors: Carol Anderson, Bo Liu, Anas Abidin, Hoo-Chang Shin, Virginia Adams

    Abstract: Social media posts contain potentially valuable information about medical conditions and health-related behavior. Biocreative VII Task 3 focuses on mining this information by recognizing mentions of medications and dietary supplements in tweets. We approach this task by fine tuning multiple BERT-style language models to perform token-level classification, and combining them into ensembles to gener… ▽ More

    Submitted 30 November, 2021; originally announced November 2021.

    Comments: Submission to the BioCreative VII challenge - Track-3

  49. arXiv:2111.15622  [pdf, other

    cs.CL

    Chemical Identification and Indexing in PubMed Articles via BERT and Text-to-Text Approaches

    Authors: Virginia Adams, Hoo-Chang Shin, Carol Anderson, Bo Liu, Anas Abidin

    Abstract: The Biocreative VII Track-2 challenge consists of named entity recognition, entity-linking (or entity-normalization), and topic indexing tasks -- with entities and topics limited to chemicals for this challenge. Named entity recognition is a well-established problem and we achieve our best performance with BERT-based BioMegatron models. We extend our BERT-based approach to the entity linking task.… ▽ More

    Submitted 30 November, 2021; originally announced November 2021.

    Comments: Submission to the BioCreative VII challenge - Track-2

  50. arXiv:2111.15617  [pdf, other

    cs.CL

    Text Mining Drug/Chemical-Protein Interactions using an Ensemble of BERT and T5 Based Models

    Authors: Virginia Adams, Hoo-Chang Shin, Carol Anderson, Bo Liu, Anas Abidin

    Abstract: In Track-1 of the BioCreative VII Challenge participants are asked to identify interactions between drugs/chemicals and proteins. In-context named entity annotations for each drug/chemical and protein are provided and one of fourteen different interactions must be automatically predicted. For this relation extraction task, we attempt both a BERT-based sentence classification approach, and a more n… ▽ More

    Submitted 30 November, 2021; originally announced November 2021.

    Comments: Submission to the BioCreative VII challenge, Track-1