Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 145 results for author: Mitchell, M

Searching in archive cs. Search in all archives.
.
  1. arXiv:2607.24891  [pdf, ps, other

    cs.IT

    High-Performance Reinforcement-Learned BP Decoding of Quantum LDPC Codes

    Authors: Mohsen Moradi, Vahid Nourozi, Taejoon Kim, Remi A. Chou, David G. M. Mitchell

    Abstract: Belief-propagation (BP) decoding is attractive for quantum low-density parity-check (QLDPC) codes because it uses local message passing on sparse Tanner graphs. However, conventional flooding BP often stalls due to stabilizer degeneracy and short cycles. Reinforcement-learning-based sequential variable-node scheduling (RL-S), which learns the update order offline, has shown that adaptive schedulin… ▽ More

    Submitted 27 July, 2026; originally announced July 2026.

  2. arXiv:2607.20130  [pdf, ps, other

    cs.IT math.QA quant-ph

    Learning to Decode Quantum LDPC Codes via Cluster-Based Sequential Belief Propagation

    Authors: Mohsen Moradi, Taejoon Kim, Rémi A. Chou, David G. M. Mitchell

    Abstract: Belief-propagation (BP) decoding for quantum low-density parity-check (QLDPC) codes is attractive due to its low complexity, but its performance is often limited by short cycles, degeneracy, and convergence failures. Recently, reinforcement-learning-based sequential variable-node (VN) scheduling (RL-S) was shown to improve BP decoding by learning state-dependent update orders. However, the VN-by-V… ▽ More

    Submitted 22 July, 2026; originally announced July 2026.

  3. arXiv:2606.20637  [pdf, ps, other

    cs.AI cond-mat.stat-mech cs.CY physics.soc-ph

    Constituency Optimisation Through Hamiltonian Representation Of Mandates (COTHROM): Algorithmic Redistricting of Irish Election Boundaries

    Authors: Ruaidhrí Campion, Matthew Fenlon, Joshua Cooney Mercedal, Casey Farren-Colloty, Eliza Somerville, Michael A. J. Mitchell

    Abstract: Electoral redistricting in Ireland's Proportional Representation Single Transferable Vote (PR-STV) system faces the challenge of selecting an optimally representative set of electoral boundaries from an enormous set of possible configurations, and where ``representative'' is a delicate balance of constitutional objectives that are often in tension with one another. We present the first computation… ▽ More

    Submitted 2 June, 2026; originally announced June 2026.

    Comments: 35 pages, 16 figures

  4. arXiv:2604.17124  [pdf, ps, other

    cs.IT

    Dynamic Parameter Scheduling in Soft-Hard BPGD for Lossy Source Coding

    Authors: Masoumeh Alinia, David G. M. Mitchell

    Abstract: We investigate lossy source coding based on a soft-decision belief propagation guided decimation (BPGD) encoder for low-density generator matrix (LDGM) codes, referred to as \emph{soft-hard BPGD}. The performance of this encoder is highly sensitive to the choice of ``softness'' parameters, typically denoted by $(β,μ)$, which are conventionally tuned via exhaustive empirical sweeps. To reduce this… ▽ More

    Submitted 18 April, 2026; originally announced April 2026.

  5. arXiv:2603.10192  [pdf, ps, other

    cs.IT

    Learning to Decode Quantum LDPC Codes Via Belief Propagation

    Authors: Mohsen Moradi, Vahid Nourozi, Salman Habib, David G. M. Mitchell

    Abstract: Belief-propagation (BP) decoding for quantum low-density parity-check (QLDPC) codes is appealing due to its low complexity, yet it often exhibits convergence issues due to quantum degeneracy and short cycles that exist in the Tanner graph. To overcome this challenge, this paper proposes a reinforcement-learning (RL) approach that learns (offline) how to decode QLDPC codes based on sequential decod… ▽ More

    Submitted 10 March, 2026; originally announced March 2026.

  6. arXiv:2602.13420  [pdf, ps, other

    cs.IT

    Sequential BP-based Decoding of QLDPC Codes

    Authors: Mohsen Moradi, Salman Habib, Vahid Nourozi, David G. M. Mitchell

    Abstract: Quantum low-density parity-check (QLDPC) codes are a leading approach to quantum error correction, yet conventional belief propagation (BP) decoders often perform poorly, primarily due to non-convergence exacerbated by stabilizer constraints, which induce short cycles and degeneracy. We propose two scheduling variants, sequential check node scheduling (SCNS) and sequential variable node scheduling… ▽ More

    Submitted 13 February, 2026; originally announced February 2026.

  7. arXiv:2601.05340  [pdf, ps, other

    cs.IT

    The Number of Cycles of Bi-regular Tanner Graphs in Terms of the Eigenvalues of the Adjacency Matrix

    Authors: Roxana Smarandache, David G. M. Mitchell

    Abstract: In this paper, we explore new connections between the cycles in the graph of low-density parity-check (LDPC) codes and the eigenvalues of the corresponding adjacency matrix. The resulting observations are used to derive fast, simple, recursive formulas for the number of cycles $N_{2k}$ of length $2k$, $k<g$, in a bi-regular graph of girth $g$. Moreover, we derive explicit formulas for $N_{2k}$,… ▽ More

    Submitted 8 January, 2026; originally announced January 2026.

    Comments: 25 pages, submitted to IT Transactions

  8. arXiv:2510.02125  [pdf, ps, other

    cs.AI cs.CL

    Do AI Models Perform Human-like Abstract Reasoning Across Modalities?

    Authors: Claas Beger, Ryan Yi, Shuhao Fu, Kaleda Denton, Arseny Moskvichev, Sarah W. Tsai, Sivasankaran Rajamanickam, Melanie Mitchell

    Abstract: OpenAI's o3-preview reasoning model exceeded human accuracy on the ARC-AGI-1 benchmark, but does that mean state-of-the-art models recognize and reason with the abstractions the benchmark was designed to test? Here we investigate abstraction abilities of AI models using the closely related but simpler ConceptARC benchmark. Our evaluations vary input modality (textual vs. visual), use of external P… ▽ More

    Submitted 2 February, 2026; v1 submitted 2 October, 2025; originally announced October 2025.

    Comments: 9 pages, 3 figures

  9. arXiv:2508.07030  [pdf, ps, other

    cs.IT

    Generalized Quasi-Cyclic LDPC Codes: Design and Efficient Encoding

    Authors: Roxana Smarandache, David G. M. Mitchell, Anthony Gómez-Fonseca

    Abstract: Generalized low-density parity-check (GLDPC) codes, where single parity-check constraints on the code bits are replaced with generalized constraints (an arbitrary linear code), are a promising class of codes for low-latency communication. The block error rate performance of the GLDPC codes, combined with a complementary outer code, has been shown to outperform a variety of state-of-the-art code an… ▽ More

    Submitted 9 August, 2025; originally announced August 2025.

  10. arXiv:2506.14652  [pdf, ps, other

    cs.CY cs.AI cs.LG

    Rigor in AI: Doing Rigorous AI Work Requires a Broader, Responsible AI-Informed Conception of Rigor

    Authors: Alexandra Olteanu, Su Lin Blodgett, Agathe Balayn, Angelina Wang, Fernando Diaz, Flavio du Pin Calmon, Margaret Mitchell, Michael Ekstrand, Reuben Binns, Solon Barocas

    Abstract: In AI research and practice, rigor remains largely understood in terms of methodological rigor -- such as whether mathematical, statistical, or computational methods are correctly applied. We argue that this narrow conception of rigor has contributed to the concerns raised by the responsible AI community, including overblown claims about the capabilities of AI systems. Our position is that a broad… ▽ More

    Submitted 25 November, 2025; v1 submitted 17 June, 2025; originally announced June 2025.

    Comments: 21 pages, 1 figure, 1 table, accepted at NeurIPS'25 position papers track

  11. arXiv:2506.11135  [pdf, ps, other

    cs.CL cs.AI cs.LG cs.NE

    Large Language Models and Emergence: A Complex Systems Perspective

    Authors: David C. Krakauer, John W. Krakauer, Melanie Mitchell

    Abstract: Emergence is a concept in complexity science that describes how many-body systems manifest novel higher-level properties, properties that can be described by replacing high-dimensional mechanisms with lower-dimensional effective variables and theories. This is captured by the idea "more is different". Intelligence is a consummate emergent property manifesting increasingly efficient -- cheaper and… ▽ More

    Submitted 10 June, 2025; originally announced June 2025.

  12. arXiv:2503.05336  [pdf, ps, other

    cs.AI cs.LG

    Toward an Evaluation Science for Generative AI Systems

    Authors: Laura Weidinger, Inioluwa Deborah Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Deep Ganguli, Sanmi Koyejo, William Isaac

    Abstract: There is an increasing imperative to anticipate and understand the performance and safety of generative AI systems in real-world deployment contexts. However, the current evaluation ecosystem is insufficient: Commonly used static benchmarks face validity challenges, and ad hoc case-by-case audits rarely scale. In this piece, we advocate for maturing an evaluation science for generative AI systems.… ▽ More

    Submitted 12 March, 2025; v1 submitted 7 March, 2025; originally announced March 2025.

    Comments: First two authors contributed equally to this work

  13. arXiv:2502.03774  [pdf, ps, other

    cs.IT

    High-Rate Spatially Coupled LDPC Codes Based on Massey's Convolutional Self-Orthogonal Codes

    Authors: Daniel J. Costello, Jr., Min Zhu, David G. M. Mitchell, Michael Lentmaier

    Abstract: In this paper, we study a new class of high-rate spatially coupled LDPC (SC-LDPC) codes based on the convolutional self-orthogonal codes (CSOCs) first introduced by Massey. The SC-LDPC codes are constructed by treating the irregular graph corresponding to the parity-check matrix of a systematic rate R = (n - 1)/n CSOC as a convolutional protograph. The protograph can then be lifted using permutati… ▽ More

    Submitted 17 February, 2025; v1 submitted 5 February, 2025; originally announced February 2025.

  14. arXiv:2502.03689  [pdf, ps, other

    cs.CY

    Stop treating `AGI' as the north-star goal of AI research

    Authors: Borhane Blili-Hamelin, Christopher Graziul, Leif Hancox-Li, Hananel Hazan, El-Mahdi El-Mhamdi, Avijit Ghosh, Katherine Heller, Jacob Metcalf, Fabricio Murai, Eryk Salvaggio, Andrew Smart, Todd Snider, Mariame Tighanimine, Talia Ringer, Margaret Mitchell, Shiri Dori-Hacohen

    Abstract: The AI research community plays a vital role in shaping the scientific, engineering, and societal goals of AI research. In this position paper, we argue that focusing on the highly contested topic of `artificial general intelligence' (`AGI') undermines our ability to choose effective goals. We identify six key traps -- obstacles to productive goal setting -- that are aggravated by AGI discourse: I… ▽ More

    Submitted 7 July, 2025; v1 submitted 5 February, 2025; originally announced February 2025.

    Comments: Position Paper accepted to ICML 2025. OpenReview: https://openreview.net/forum?id=1RlrtH6ydW

  15. arXiv:2502.02649  [pdf, ps, other

    cs.AI

    Fully Autonomous AI Agents Should Not be Developed

    Authors: Margaret Mitchell, Avijit Ghosh, Alexandra Sasha Luccioni, Giada Pistilli

    Abstract: This paper argues that fully autonomous AI agents should not be developed. In support of this position, we build from prior scientific literature and current product marketing to delineate different AI agent levels and detail the ethical values at play in each, documenting trade-offs in potential benefits and risks. Our analysis reveals that risks to people increase with the autonomy of a system:… ▽ More

    Submitted 19 October, 2025; v1 submitted 4 February, 2025; originally announced February 2025.

  16. arXiv:2411.14215  [pdf, other

    cs.CL cs.AI cs.LG

    Evaluating the Robustness of Analogical Reasoning in Large Language Models

    Authors: Martha Lewis, Melanie Mitchell

    Abstract: LLMs have performed well on several reasoning benchmarks, including ones that test analogical reasoning abilities. However, there is debate on the extent to which they are performing general abstract reasoning versus employing non-robust processes, e.g., that overly rely on similarity to pre-training data. Here we investigate the robustness of analogy-making abilities previously claimed for LLMs o… ▽ More

    Submitted 21 November, 2024; originally announced November 2024.

    Comments: 31 pages, 13 figures. arXiv admin note: text overlap with arXiv:2402.08955

  17. Traceable random numbers from a nonlocal quantum advantage

    Authors: Gautam A. Kavuri, Jasper Palfree, Dileep V. Reddy, Yanbao Zhang, Joshua C. Bienfang, Michael D. Mazurek, Mohammad A. Alhejji, Aliza U. Siddiqui, Joseph M. Cavanagh, Aagam Dalal, Carlos Abellán, Waldimar Amaya, Morgan W. Mitchell, Katherine E. Stange, Paul D. Beale, Luís T. A. N. Brandão, Harold Booth, René Peralta, Sae Woo Nam, Richard P. Mirin, Martin J. Stevens, Emanuel Knill, Lynden K. Shalm

    Abstract: The unpredictability of random numbers is fundamental to both digital security and applications that fairly distribute resources. However, existing random number generators have limitations-the generation processes cannot be fully traced, audited, and certified to be unpredictable. The algorithmic steps used in pseudorandom number generators are auditable, but they cannot guarantee that their outp… ▽ More

    Submitted 7 November, 2024; originally announced November 2024.

    Comments: 40 pages, 4 main figures, 10 supplementary figures

  18. arXiv:2411.02478  [pdf

    cs.AI cs.CY cs.HC

    Imagining and building wise machines: The centrality of AI metacognition

    Authors: Samuel G. B. Johnson, Amir-Hossein Karimi, Yoshua Bengio, Nick Chater, Tobias Gerstenberg, Kate Larson, Sydney Levine, Melanie Mitchell, Iyad Rahwan, Bernhard Schölkopf, Igor Grossmann

    Abstract: Although AI has become increasingly smart, its wisdom has not kept pace. In this article, we examine what is known about human wisdom and sketch a vision of its AI counterpart. We analyze human wisdom as a set of strategies for solving intractable problems-those outside the scope of analytic techniques-including both object-level strategies like heuristics [for managing problems] and metacognitive… ▽ More

    Submitted 7 January, 2026; v1 submitted 4 November, 2024; originally announced November 2024.

    Comments: 23 pages, 2 figures, 2 tables

  19. arXiv:2411.02348  [pdf, ps, other

    cs.AI cs.CL cs.HC

    Can Large Language Models generalize analogy solving like children can?

    Authors: Claire E. Stevenson, Alexandra Pafford, Han L. J. van der Maas, Melanie Mitchell

    Abstract: In people, the ability to solve analogies such as "body : feet :: table : ?" emerges in childhood, and appears to transfer easily to other domains, such as the visual domain "( : ) :: < : ?". Recent research shows that large language models (LLMs) can solve various forms of analogies. However, can LLMs generalize analogy solving to new domains like people can? To investigate this, we had children,… ▽ More

    Submitted 6 October, 2025; v1 submitted 4 November, 2024; originally announced November 2024.

    Comments: Accepted to Transactions of the Association for Computational Linguistics (TACL)

  20. arXiv:2410.05844  [pdf, other

    eess.SY cs.IT eess.SP

    Spectrally Efficient LDPC Codes For IRIG-106 Waveforms via Random Puncturing

    Authors: Andrew D. Cummins, David G. M. Mitchell, Erik Perrins

    Abstract: Low-density parity-check (LDPC) codes form part of the IRIG-106 standard and have been successfully deployed for the Telemetry Group version of shaped-offset quadrature phase shift keying (SOQPSK-TG) modulation. Recently, LDPC code solutions have been proposed and optimized for continuous phase modulations (CPMs), including the pulse code modulation/frequency modulation (PCM/FM) and the multi-h CP… ▽ More

    Submitted 8 October, 2024; originally announced October 2024.

    Comments: Accepted for inclusion in the 2024 International Telemetry Conference

  21. arXiv:2407.18471  [pdf, other

    cs.CL cs.IR cs.LG

    Constructing the CORD-19 Vaccine Dataset

    Authors: Manisha Singh, Divy Sharma, Alonso Ma, Bridget Tyree, Margaret Mitchell

    Abstract: We introduce new dataset 'CORD-19-Vaccination' to cater to scientists specifically looking into COVID-19 vaccine-related research. This dataset is extracted from CORD-19 dataset [Wang et al., 2020] and augmented with new columns for language detail, author demography, keywords, and topic per paper. Facebook's fastText model is used to identify languages [Joulin et al., 2016]. To establish author d… ▽ More

    Submitted 25 July, 2024; originally announced July 2024.

  22. arXiv:2406.17557  [pdf, other

    cs.CL

    The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

    Authors: Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, Thomas Wolf

    Abstract: The performance of a large language model (LLM) depends heavily on the quality and size of its pretraining dataset. However, the pretraining datasets for state-of-the-art open LLMs like Llama 3 and Mixtral are not publicly available and very little is known about how they were created. In this work, we introduce FineWeb, a 15-trillion token dataset derived from 96 Common Crawl snapshots that produ… ▽ More

    Submitted 31 October, 2024; v1 submitted 25 June, 2024; originally announced June 2024.

  23. arXiv:2405.13974  [pdf, other

    cs.CL cs.AI

    CIVICS: Building a Dataset for Examining Culturally-Informed Values in Large Language Models

    Authors: Giada Pistilli, Alina Leidinger, Yacine Jernite, Atoosa Kasirzadeh, Alexandra Sasha Luccioni, Margaret Mitchell

    Abstract: This paper introduces the "CIVICS: Culturally-Informed & Values-Inclusive Corpus for Societal impacts" dataset, designed to evaluate the social and cultural variation of Large Language Models (LLMs) across multiple languages and value-sensitive topics. We create a hand-crafted, multilingual dataset of value-laden prompts which address specific socially sensitive topics, including LGBTQI rights, so… ▽ More

    Submitted 22 May, 2024; originally announced May 2024.

  24. arXiv:2402.08955  [pdf, other

    cs.AI cs.CL

    Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models

    Authors: Martha Lewis, Melanie Mitchell

    Abstract: Large language models (LLMs) have performed well on several reasoning benchmarks, including ones that test analogical reasoning abilities. However, it has been debated whether they are actually performing humanlike abstract reasoning or instead employing less general processes that rely on similarity to what has been seen in their training data. Here we investigate the generality of analogy-making… ▽ More

    Submitted 14 February, 2024; originally announced February 2024.

  25. arXiv:2401.10376  [pdf, other

    cs.IT

    PAC Code Rate-Profile Design Using Search-Constrained Optimization Algorithms

    Authors: Mohsen Moradi, David G. M. Mitchell

    Abstract: In this paper, we introduce a novel rate-profile design based on search-constrained optimization techniques to assess the performance of polarization-adjusted convolutional (PAC) codes under Fano (sequential) decoding. The results demonstrate that the resulting PAC code offers much reduced computational complexity compared to a construction based on a conventional genetic algorithm without a perfo… ▽ More

    Submitted 18 January, 2024; originally announced January 2024.

  26. arXiv:2401.10284  [pdf, other

    eess.SP cs.AI cs.LG

    MorpheusNet: Resource efficient sleep stage classifier for embedded on-line systems

    Authors: Ali Kavoosi, Morgan P. Mitchell, Raveen Kariyawasam, John E. Fleming, Penny Lewis, Heidi Johansen-Berg, Hayriye Cagnan, Timothy Denison

    Abstract: Sleep Stage Classification (SSC) is a labor-intensive task, requiring experts to examine hours of electrophysiological recordings for manual classification. This is a limiting factor when it comes to leveraging sleep stages for therapeutic purposes. With increasing affordability and expansion of wearable devices, automating SSC may enable deployment of sleep-based therapies at scale. Deep Learning… ▽ More

    Submitted 14 January, 2024; originally announced January 2024.

    Comments: This paper was presented at the 2023 IEEE conference on Systems, Man, and Cybernetics (SMC)

  27. arXiv:2312.09323  [pdf, other

    cs.AI cs.LG

    Perspectives on the State and Future of Deep Learning - 2023

    Authors: Micah Goldblum, Anima Anandkumar, Richard Baraniuk, Tom Goldstein, Kyunghyun Cho, Zachary C Lipton, Melanie Mitchell, Preetum Nakkiran, Max Welling, Andrew Gordon Wilson

    Abstract: The goal of this series is to chronicle opinions and issues in the field of machine learning as they stand today and as they change over time. The plan is to host this survey periodically until the AI singularity paperclip-frenzy-driven doomsday, keeping an updated list of topical questions and interviewing new community members for each edition. In this issue, we probed people's opinions on inter… ▽ More

    Submitted 18 December, 2023; v1 submitted 7 December, 2023; originally announced December 2023.

  28. arXiv:2311.16896  [pdf, other

    physics.optics cs.ET physics.app-ph

    120 GOPS Photonic Tensor Core in Thin-film Lithium Niobate for Inference and in-situ Training

    Authors: Zhongjin Lin, Bhavin J. Shastri, Shangxuan Yu, Jingxiang Song, Yuntao Zhu, Arman Safarnejadian, Wangning Cai, Yanmei Lin, Wei Ke, Mustafa Hammood, Tianye Wang, Mengyue Xu, Zibo Zheng, Mohammed Al-Qadasi, Omid Esmaeeli, Mohamed Rahim, Grzegorz Pakulski, Jens Schmid, Pedro Barrios, Weihong Jiang, Hugh Morison, Matthew Mitchell, Xun Guan, Nicolas A. F. Jaeger, Leslie A. n Rusch , et al. (5 additional authors not shown)

    Abstract: Photonics offers a transformative approach to artificial intelligence (AI) and neuromorphic computing by enabling low-latency, high-speed, and energy-efficient computations. However, conventional photonic tensor cores face significant challenges in constructing large-scale photonic neuromorphic networks. Here, we propose a fully integrated photonic tensor core, consisting of only two thin-film lit… ▽ More

    Submitted 8 October, 2024; v1 submitted 28 November, 2023; originally announced November 2023.

    Comments: 21 pages, 6 figures

    MSC Class: 78A05

  29. arXiv:2311.09247  [pdf, other

    cs.AI cs.LG

    Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks

    Authors: Melanie Mitchell, Alessandro B. Palmarini, Arseny Moskvichev

    Abstract: We explore the abstract reasoning abilities of text-only and multimodal versions of GPT-4, using the ConceptARC benchmark [10], which is designed to evaluate robust understanding and reasoning with core-knowledge concepts. We extend the work of Moskvichev et al. [10] by evaluating GPT-4 on more detailed, one-shot prompting (rather than simple, zero-shot prompts) with text versions of ConceptARC ta… ▽ More

    Submitted 11 December, 2023; v1 submitted 13 November, 2023; originally announced November 2023.

    Comments: Corrected Figure 3 (extra spaces were replaced by commas, which were lost in original formatting)

    Journal ref: Proceedings of the LLM-CP Workshop, AAAI 2024

  30. arXiv:2310.01557  [pdf, other

    cs.LG cs.AI

    SmartPlay: A Benchmark for LLMs as Intelligent Agents

    Authors: Yue Wu, Xuan Tang, Tom M. Mitchell, Yuanzhi Li

    Abstract: Recent large language models (LLMs) have demonstrated great potential toward intelligent agents and next-gen automation, but there currently lacks a systematic benchmark for evaluating LLMs' abilities as agents. We introduce SmartPlay: both a challenging benchmark and a methodology for evaluating LLMs as agents. SmartPlay consists of 6 different games, including Rock-Paper-Scissors, Tower of Hanoi… ▽ More

    Submitted 17 March, 2024; v1 submitted 2 October, 2023; originally announced October 2023.

  31. arXiv:2307.14213  [pdf, other

    cs.RO

    Soft Air Pocket Force Sensors for Large Scale Flexible Robots

    Authors: Michael R. Mitchell, Ciera McFarland, Margaret M. Coad

    Abstract: Flexible robots have advantages over rigid robots in their ability to conform physically to their environment and to form a wide variety of shapes. Sensing the force applied by or to flexible robots is useful for both navigation and manipulation tasks, but it is challenging due to the need for the sensors to withstand the robots' shape change without encumbering their functionality. Also, for robo… ▽ More

    Submitted 26 July, 2023; originally announced July 2023.

    Comments: M. R. Mitchell, C. McFarland, and M. M. Coad, "Soft Air Pocket Force Sensors for Large Scale Flexible Robots," in IEEE International Conference on Soft Robotics, 2023, pp. 1-8. Video: https://youtu.be/2De0htilW74

  32. arXiv:2307.13905  [pdf, other

    cs.IT

    Reinforcement Learning for Sequential Decoding of Generalized LDPC Codes

    Authors: Salman Habib, David G. M. Mitchell

    Abstract: In this work, we propose reinforcement learning (RL) for sequential decoding of moderate length generalized low-density parity-check (GLDPC) codes. Here, sequential decoding refers to scheduling all the generalized constraint nodes (GCNs) and single parity-check nodes (SPCNs) of a GLDPC code serially in each iteration. A GLDPC decoding environment is modeled as a finite Markov decision process (MD… ▽ More

    Submitted 25 July, 2023; originally announced July 2023.

    Comments: accepted for publication at ISTC 2023. arXiv admin note: text overlap with arXiv:2112.13934

  33. arXiv:2306.05949   

    cs.CY cs.AI

    Evaluating the Social Impact of Generative AI Systems in Systems and Society

    Authors: Irene Solaiman, Zeerak Talat, William Agnew, Lama Ahmad, Dylan Baker, Su Lin Blodgett, Canyu Chen, Hal Daumé III, Jesse Dodge, Isabella Duan, Ellie Evans, Felix Friedrich, Avijit Ghosh, Usman Gohar, Sara Hooker, Yacine Jernite, Ria Kalluri, Alberto Lusoli, Alina Leidinger, Michelle Lin, Xiuzhu Lin, Sasha Luccioni, Jennifer Mickel, Margaret Mitchell, Jessica Newman , et al. (6 additional authors not shown)

    Abstract: Generative AI systems across modalities, ranging from text (including code), image, audio, and video, have broad social impacts, but there is no official standard for means of evaluating those impacts or for which impacts should be evaluated. In this paper, we present a guide that moves toward a standard approach in evaluating a base generative AI system for any modality in two overarching categor… ▽ More

    Submitted 28 June, 2024; v1 submitted 9 June, 2023; originally announced June 2023.

    Comments: This version has been removed by arXiv administrators as the submitter did not have the right to agree to the license at the time of submission

  34. Stronger Together: on the Articulation of Ethical Charters, Legal Tools, and Technical Documentation in ML

    Authors: Giada Pistilli, Carlos Munoz Ferrandis, Yacine Jernite, Margaret Mitchell

    Abstract: The growing need for accountability of the people behind AI systems can be addressed by leveraging processes in three fields of study: ethics, law, and computer science. While these fields are often considered in isolation, they rely on complementary notions in their interpretation and implementation. In this work, we detail this interdependence and motivate the necessary role of collaborative gov… ▽ More

    Submitted 9 May, 2023; originally announced May 2023.

  35. arXiv:2305.07141  [pdf, other

    cs.LG cs.AI

    The ConceptARC Benchmark: Evaluating Understanding and Generalization in the ARC Domain

    Authors: Arseny Moskvichev, Victor Vikram Odouard, Melanie Mitchell

    Abstract: The abilities to form and abstract concepts is key to human intelligence, but such abilities remain lacking in state-of-the-art AI systems. There has been substantial research on conceptual abstraction in AI, particularly using idealized domains such as Raven's Progressive Matrices and Bongard problems, but even when AI systems succeed on such problems, the systems are rarely evaluated in depth to… ▽ More

    Submitted 11 May, 2023; originally announced May 2023.

    Journal ref: Transactions on Machine Learning Research, 8/2023

  36. arXiv:2304.13626  [pdf, other

    cs.AI

    The Roles of Symbols in Neural-based AI: They are Not What You Think!

    Authors: Daniel L. Silver, Tom M. Mitchell

    Abstract: We propose that symbols are first and foremost external communication tools used between intelligent agents that allow knowledge to be transferred in a more efficient and effective manner than having to experience the world directly. But, they are also used internally within an agent through a form of self-communication to help formulate, describe and justify subsymbolic patterns of neural activit… ▽ More

    Submitted 26 April, 2023; originally announced April 2023.

    Comments: 28 pages

  37. arXiv:2303.17853  [pdf, other

    physics.pop-ph astro-ph.HE cs.CL

    Can AI Put Gamma-Ray Astrophysicists Out of a Job?

    Authors: Samuel T. Spencer, Vikas Joshi, Alison M. W. Mitchell

    Abstract: In what will likely be a litany of generative-model-themed arXiv submissions celebrating April the 1st, we evaluate the capacity of state-of-the-art transformer models to create a paper detailing the detection of a Pulsar Wind Nebula with a non-existent Imaging Atmospheric Cherenkov Telescope (IACT) Array. We do this to evaluate the ability of such models to interpret astronomical observations and… ▽ More

    Submitted 4 April, 2023; v1 submitted 31 March, 2023; originally announced March 2023.

  38. arXiv:2303.11408  [pdf, other

    cs.CY

    Stable Bias: Analyzing Societal Representations in Diffusion Models

    Authors: Alexandra Sasha Luccioni, Christopher Akiki, Margaret Mitchell, Yacine Jernite

    Abstract: As machine learning-enabled Text-to-Image (TTI) systems are becoming increasingly prevalent and seeing growing adoption as commercial services, characterizing the social biases they exhibit is a necessary first step to lowering their risk of discriminatory outcomes. This evaluation, however, is made more difficult by the synthetic nature of these systems' outputs: common definitions of diversity a… ▽ More

    Submitted 9 November, 2023; v1 submitted 20 March, 2023; originally announced March 2023.

    Comments: Accepted to NeurIPS Datasets and Benchmarks 2023 (spotlight)

  39. arXiv:2303.03915  [pdf, other

    cs.CL cs.AI

    The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset

    Authors: Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Šaško, Quentin Lhoest, Angelina McMillan-Major, Gerard Dupont, Stella Biderman, Anna Rogers, Loubna Ben allal, Francesco De Toni, Giada Pistilli, Olivier Nguyen, Somaieh Nikpoor, Maraim Masoud, Pierre Colombo, Javier de la Rosa , et al. (29 additional authors not shown)

    Abstract: As language models grow ever larger, the need for large-scale high-quality text datasets has never been more pressing, especially in multilingual settings. The BigScience workshop, a 1-year international and multidisciplinary initiative, was formed with the goal of researching and training large language models as a values-driven undertaking, putting issues of ethics, harm, and governance in the f… ▽ More

    Submitted 7 March, 2023; originally announced March 2023.

    Comments: NeurIPS 2022, Datasets and Benchmarks Track

    ACM Class: I.2.7

  40. arXiv:2302.04449  [pdf, other

    cs.LG cs.AI cs.CL

    Read and Reap the Rewards: Learning to Play Atari with the Help of Instruction Manuals

    Authors: Yue Wu, Yewen Fan, Paul Pu Liang, Amos Azaria, Yuanzhi Li, Tom M. Mitchell

    Abstract: High sample complexity has long been a challenge for RL. On the other hand, humans learn to perform tasks not only from interaction or demonstrations, but also by reading unstructured text documents, e.g., instruction manuals. Instruction manuals and wiki pages are among the most abundant data that could inform agents of valuable features and policies or task-specific environmental dynamics and re… ▽ More

    Submitted 20 July, 2024; v1 submitted 9 February, 2023; originally announced February 2023.

  41. arXiv:2212.05129  [pdf, other

    cs.AI cs.LG

    Measuring Data

    Authors: Margaret Mitchell, Alexandra Sasha Luccioni, Nathan Lambert, Marissa Gerchick, Angelina McMillan-Major, Ezinwanne Ozoani, Nazneen Rajani, Tristan Thrush, Yacine Jernite, Douwe Kiela

    Abstract: We identify the task of measuring data to quantitatively characterize the composition of machine learning data and datasets. Similar to an object's height, width, and volume, data measurements quantify different attributes of data along common dimensions that support comparison. Several lines of research have proposed what we refer to as measurements, with differing terminology; we bring some of t… ▽ More

    Submitted 13 February, 2023; v1 submitted 9 December, 2022; originally announced December 2022.

  42. arXiv:2211.15533  [pdf, other

    cs.CL cs.AI

    The Stack: 3 TB of permissively licensed source code

    Authors: Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, Harm de Vries

    Abstract: Large Language Models (LLMs) play an ever-increasing role in the field of Artificial Intelligence (AI)--not only for natural language processing but also for code understanding and generation. To stimulate open and responsible research on LLMs for code, we introduce The Stack, a 3.1 TB dataset consisting of permissively licensed source code in 30 programming languages. We describe how we collect t… ▽ More

    Submitted 20 November, 2022; originally announced November 2022.

  43. arXiv:2211.05100  [pdf, other

    cs.CL

    BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    Authors: BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît Sagot, Niklas Muennighoff, Albert Villanova del Moral, Olatunji Ruwase, Rachel Bawden, Stas Bekman, Angelina McMillan-Major , et al. (369 additional authors not shown)

    Abstract: Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed by resource-rich organizations and are frequently kept from the public. As a step towards democratizing this powerful technology, we present BLOOM, a 176B-parameter open-access… ▽ More

    Submitted 27 June, 2023; v1 submitted 9 November, 2022; originally announced November 2022.

  44. arXiv:2210.15767  [pdf

    cs.AI

    Gathering Strength, Gathering Storms: The One Hundred Year Study on Artificial Intelligence (AI100) 2021 Study Panel Report

    Authors: Michael L. Littman, Ifeoma Ajunwa, Guy Berger, Craig Boutilier, Morgan Currie, Finale Doshi-Velez, Gillian Hadfield, Michael C. Horowitz, Charles Isbell, Hiroaki Kitano, Karen Levy, Terah Lyons, Melanie Mitchell, Julie Shah, Steven Sloman, Shannon Vallor, Toby Walsh

    Abstract: In September 2021, the "One Hundred Year Study on Artificial Intelligence" project (AI100) issued the second report of its planned long-term periodic assessment of artificial intelligence (AI) and its impact on society. It was written by a panel of 17 study authors, each of whom is deeply rooted in AI research, chaired by Michael Littman of Brown University. The report, entitled "Gathering Strengt… ▽ More

    Submitted 27 October, 2022; originally announced October 2022.

    Comments: 82 pages, https://ai100.stanford.edu/gathering-strength-gathering-storms-one-hundred-year-study-artificial-intelligence-ai100-2021-study

  45. The Debate Over Understanding in AI's Large Language Models

    Authors: Melanie Mitchell, David C. Krakauer

    Abstract: We survey a current, heated debate in the AI research community on whether large pre-trained language models can be said to "understand" language -- and the physical and social situations language encodes -- in any important sense. We describe arguments that have been made for and against such understanding, and key questions for the broader sciences of intelligence that have arisen in light of th… ▽ More

    Submitted 10 February, 2023; v1 submitted 14 October, 2022; originally announced October 2022.

    Comments: Under submission as a Perspective article. Updated with additional discussion and citations

    Journal ref: Proceedings of the National Academy of Sciences 120 (13), 2023

  46. arXiv:2210.13589  [pdf, ps, other

    cs.AI cs.LG cs.RO

    Embodied, Situated, and Grounded Intelligence: Implications for AI

    Authors: Tyler Millhouse, Melanie Moses, Melanie Mitchell

    Abstract: In April of 2022, the Santa Fe Institute hosted a workshop on embodied, situated, and grounded intelligence as part of the Institute's Foundations of Intelligence project. The workshop brought together computer scientists, psychologists, philosophers, social scientists, and others to discuss the science of embodiment and related issues in human intelligence, and its implications for building robus… ▽ More

    Submitted 24 October, 2022; originally announced October 2022.

    Comments: 38 pages, workshop report

  47. arXiv:2210.05839  [pdf, other

    cs.CL cs.HC

    SEAL : Interactive Tool for Systematic Error Analysis and Labeling

    Authors: Nazneen Rajani, Weixin Liang, Lingjiao Chen, Meg Mitchell, James Zou

    Abstract: With the advent of Transformers, large language models (LLMs) have saturated well-known NLP benchmarks and leaderboards with high aggregate performance. However, many times these models systematically fail on tail data or rare groups not obvious in aggregate evaluation. Identifying such problematic data groups is even more challenging when there are no explicit labels (e.g., ethnicity, gender, etc… ▽ More

    Submitted 11 October, 2022; originally announced October 2022.

    Comments: Accepted at EMNLP 2022 demo track

  48. arXiv:2210.02667  [pdf, ps, other

    cs.AI cs.CY

    A Human Rights-Based Approach to Responsible AI

    Authors: Vinodkumar Prabhakaran, Margaret Mitchell, Timnit Gebru, Iason Gabriel

    Abstract: Research on fairness, accountability, transparency and ethics of AI-based interventions in society has gained much-needed momentum in recent years. However it lacks an explicit alignment with a set of normative values and principles that guide this research and interventions. Rather, an implicit consensus is often assumed to hold for the values we impart into our models - something that is at odds… ▽ More

    Submitted 6 October, 2022; originally announced October 2022.

    Comments: Presented as a (non-archival) poster at the 2022 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization or (EAAMO '22)

  49. arXiv:2210.01970  [pdf, other

    cs.LG

    Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model Measurements

    Authors: Leandro von Werra, Lewis Tunstall, Abhishek Thakur, Alexandra Sasha Luccioni, Tristan Thrush, Aleksandra Piktus, Felix Marty, Nazneen Rajani, Victor Mustar, Helen Ngo, Omar Sanseviero, Mario Šaško, Albert Villanova, Quentin Lhoest, Julien Chaumond, Margaret Mitchell, Alexander M. Rush, Thomas Wolf, Douwe Kiela

    Abstract: Evaluation is a key part of machine learning (ML), yet there is a lack of support and tooling to enable its informed and systematic practice. We introduce Evaluate and Evaluation on the Hub --a set of tools to facilitate the evaluation of models and datasets in ML. Evaluate is a library to support best practices for measurements, metrics, and comparisons of data and models. Its goal is to support… ▽ More

    Submitted 6 October, 2022; v1 submitted 30 September, 2022; originally announced October 2022.

  50. arXiv:2207.08939  [pdf, other

    cs.LG

    Learning Sparsity-Promoting Regularizers using Bilevel Optimization

    Authors: Avrajit Ghosh, Michael T. McCann, Madeline Mitchell, Saiprasad Ravishankar

    Abstract: We present a method for supervised learning of sparsity-promoting regularizers for denoising signals and images. Sparsity-promoting regularization is a key ingredient in solving modern signal reconstruction problems; however, the operators underlying these regularizers are usually either designed by hand or learned from data in an unsupervised way. The recent success of supervised learning (mainly… ▽ More

    Submitted 5 September, 2023; v1 submitted 18 July, 2022; originally announced July 2022.

    Journal ref: SIAM Journal on Imaging Sciences (SIIMS-2023)