Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–12 of 12 results for author: Schut, L

Searching in archive cs. Search in all archives.
.
  1. arXiv:2606.26299  [pdf, ps, other

    cs.AI

    COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami

    Authors: Tom Zahavy, Shaobo Hou, Thomas Tumiel, James Doran, Francesco Faccio, Xidong Feng, Alex Havrilla, Igor Khytryi, Chenglei Li, Lisa Schut, Vivek Veeriah, Arijan Abrashi, Michał Kosmulski, Robert J. Lang, Nick Robinson, Brandon Wong, Marcus Chiam, Gloria Fang, Satinder Singh

    Abstract: While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies both strict geometric constraints and subjective visual aesthetics remains a challenge. This paper presents an approach to tackle these difficulties in the domain of computational origami, a mathematically rigid environment that grounds artistic design within th… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

  2. arXiv:2502.15603  [pdf, other

    cs.CL cs.AI cs.LG

    Do Multilingual LLMs Think In English?

    Authors: Lisa Schut, Yarin Gal, Sebastian Farquhar

    Abstract: Large language models (LLMs) have multilingual capabilities and can solve tasks across various languages. However, we show that current LLMs make key decisions in a representation space closest to English, regardless of their input and output languages. Exploring the internal representations with a logit lens for sentences in French, German, Dutch, and Mandarin, we show that the LLM first emits re… ▽ More

    Submitted 21 February, 2025; originally announced February 2025.

    Comments: Main paper 9 pages; including appendix 48 pages

  3. Humanity's Last Exam

    Authors: Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Chen Bo Calvin Zhang, Mohamed Shaaban, John Ling, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, Adam Khoja, Ryan Kim, Richard Ren, Jason Hausenloy, Oliver Zhang, Mantas Mazeika, Dmitry Dodonov, Tung Nguyen, Jaeho Lee, Daron Anderson, Mikhail Doroshenko, Alun Cennyth Stokes , et al. (1133 additional authors not shown)

    Abstract: Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achieve over 90\% accuracy on popular benchmarks like MMLU, limiting informed measurement of state-of-the-art LLM capabilities. In response, we introduce Humanity's Last Exam (HLE), a multi-modal benchmark at the frontier of… ▽ More

    Submitted 28 July, 2026; v1 submitted 24 January, 2025; originally announced January 2025.

    Comments: 29 pages, 6 figures

  4. arXiv:2406.15927  [pdf, other

    cs.CL cs.AI cs.LG

    Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs

    Authors: Jannik Kossen, Jiatong Han, Muhammed Razzak, Lisa Schut, Shreshth Malik, Yarin Gal

    Abstract: We propose semantic entropy probes (SEPs), a cheap and reliable method for uncertainty quantification in Large Language Models (LLMs). Hallucinations, which are plausible-sounding but factually incorrect and arbitrary model generations, present a major challenge to the practical adoption of LLMs. Recent work by Farquhar et al. (2024) proposes semantic entropy (SE), which can detect hallucinations… ▽ More

    Submitted 22 June, 2024; originally announced June 2024.

    Comments: First three authors contributed equally

  5. arXiv:2404.03713  [pdf, other

    cs.LG cs.AI cs.CV cs.HC

    Explaining Explainability: Recommendations for Effective Use of Concept Activation Vectors

    Authors: Angus Nicolson, Lisa Schut, J. Alison Noble, Yarin Gal

    Abstract: Concept-based explanations translate the internal representations of deep learning models into a language that humans are familiar with: concepts. One popular method for finding concepts is Concept Activation Vectors (CAVs), which are learnt using a probe dataset of concept exemplars. In this work, we investigate three properties of CAVs: (1) inconsistency across layers, (2) entanglement with othe… ▽ More

    Submitted 13 February, 2025; v1 submitted 4 April, 2024; originally announced April 2024.

    Comments: Accepted by Transactions on Machine Learning Research (02/2025)

    ACM Class: I.2.6

  6. arXiv:2310.16410  [pdf, other

    cs.AI cs.HC cs.LG stat.ML

    Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero

    Authors: Lisa Schut, Nenad Tomasev, Tom McGrath, Demis Hassabis, Ulrich Paquet, Been Kim

    Abstract: Artificial Intelligence (AI) systems have made remarkable progress, attaining super-human performance across various domains. This presents us with an opportunity to further human knowledge and improve human expert performance by leveraging the hidden knowledge encoded within these highly performant AI systems. Yet, this knowledge is often hard to extract, and may be hard to understand or learn fr… ▽ More

    Submitted 25 October, 2023; originally announced October 2023.

    Comments: 61 pages, 29 figures

  7. arXiv:2308.09175  [pdf, other

    cs.AI cs.LG

    Diversifying AI: Towards Creative Chess with AlphaZero

    Authors: Tom Zahavy, Vivek Veeriah, Shaobo Hou, Kevin Waugh, Matthew Lai, Edouard Leurent, Nenad Tomasev, Lisa Schut, Demis Hassabis, Satinder Singh

    Abstract: In recent years, Artificial Intelligence (AI) systems have surpassed human intelligence in a variety of computational tasks. However, AI systems, like humans, make mistakes, have blind spots, hallucinate, and struggle to generalize to new situations. This work explores whether AI can benefit from creative decision-making mechanisms when pushed to the limits of its computational rationality. In par… ▽ More

    Submitted 31 July, 2024; v1 submitted 17 August, 2023; originally announced August 2023.

  8. arXiv:2111.15639  [pdf, other

    cs.CV cs.LG stat.ML

    DeDUCE: Generating Counterfactual Explanations Efficiently

    Authors: Benedikt Höltgen, Lisa Schut, Jan M. Brauner, Yarin Gal

    Abstract: When an image classifier outputs a wrong class label, it can be helpful to see what changes in the image would lead to a correct classification. This is the aim of algorithms generating counterfactual explanations. However, there is no easily scalable method to generate such counterfactuals. We develop a new algorithm providing counterfactual explanations for large image classifiers trained with s… ▽ More

    Submitted 29 November, 2021; originally announced November 2021.

    Comments: Presented at the 1st Workshop on eXplainable AI approaches for debugging and diagnosis (XAI4Debugging@NeurIPS2021)

  9. arXiv:2103.08951  [pdf, other

    cs.LG stat.AP

    Generating Interpretable Counterfactual Explanations By Implicit Minimisation of Epistemic and Aleatoric Uncertainties

    Authors: Lisa Schut, Oscar Key, Rory McGrath, Luca Costabello, Bogdan Sacaleanu, Medb Corcoran, Yarin Gal

    Abstract: Counterfactual explanations (CEs) are a practical tool for demonstrating why machine learning classifiers make particular decisions. For CEs to be useful, it is important that they are easy for users to interpret. Existing methods for generating interpretable CEs rely on auxiliary generative models, which may not be suitable for complex datasets, and incur engineering overhead. We introduce a simp… ▽ More

    Submitted 16 March, 2021; originally announced March 2021.

    Comments: 21 pages, 13 Figures

    Journal ref: Proceedings of the 24th International Conference on Artificial Intelligence and Statistics (AISTATS) 2021

  10. arXiv:2010.14499  [pdf, other

    cs.LG

    A Bayesian Perspective on Training Speed and Model Selection

    Authors: Clare Lyle, Lisa Schut, Binxin Ru, Yarin Gal, Mark van der Wilk

    Abstract: We take a Bayesian perspective to illustrate a connection between training speed and the marginal likelihood in linear models. This provides two major insights: first, that a measure of a model's training speed can be used to estimate its marginal likelihood. Second, that this measure, under certain conditions, predicts the relative weighting of models in linear model combinations trained to minim… ▽ More

    Submitted 27 October, 2020; originally announced October 2020.

    Comments: To be presented at NeurIPS 2020

  11. arXiv:2006.04492  [pdf, other

    stat.ML cs.LG

    Speedy Performance Estimation for Neural Architecture Search

    Authors: Binxin Ru, Clare Lyle, Lisa Schut, Miroslav Fil, Mark van der Wilk, Yarin Gal

    Abstract: Reliable yet efficient evaluation of generalisation performance of a proposed architecture is crucial to the success of neural architecture search (NAS). Traditional approaches face a variety of limitations: training each architecture to completion is prohibitively expensive, early stopped validation accuracy may correlate poorly with fully trained performance, and model-based estimators require l… ▽ More

    Submitted 7 June, 2021; v1 submitted 8 June, 2020; originally announced June 2020.

    Comments: 23 pages, 14 figures

  12. arXiv:2004.03553  [pdf, other

    cs.LG stat.ML

    Capsule Networks -- A Probabilistic Perspective

    Authors: Lewis Smith, Lisa Schut, Yarin Gal, Mark van der Wilk

    Abstract: 'Capsule' models try to explicitly represent the poses of objects, enforcing a linear relationship between an object's pose and that of its constituent parts. This modelling assumption should lead to robustness to viewpoint changes since the sub-object/super-object relationships are invariant to the poses of the object. We describe a probabilistic generative model which encodes such capsule assump… ▽ More

    Submitted 6 January, 2021; v1 submitted 7 April, 2020; originally announced April 2020.