Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–50 of 129 results for author: Turner, R E

.
  1. arXiv:2608.12271  [pdf, ps, other

    cs.LG physics.ao-ph

    Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling

    Authors: Pedro Sousa, Will Tebbutt, Sadiq Jaffer, Robin Young, Anil Madhavapeddy, Richard E. Turner

    Abstract: Global weather reanalyses and forecasts resolve the evolving atmospheric state on coarse grids, but site-specific applications require predictions at arbitrary locations where near-surface conditions also depend on unresolved terrain and land-surface properties. Existing probabilistic downscalers address this gap using hand-crafted topographic descriptors. We ask instead whether Earth observation… ▽ More

    Submitted 12 August, 2026; originally announced August 2026.

    Comments: 39 pages, 12 figures, 6 tables

  2. arXiv:2608.09959  [pdf, ps, other

    physics.ao-ph cs.LG

    AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting

    Authors: Anna Allen, Wessel P. Bruinsma, Michael Maier-Gerber, Harrison Cook, Matthew Chantry, Richard E. Turner

    Abstract: AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to physics-based NWP in forecasting tropical cyclone (TC) tracks, they tend to dramatically underestimate intensity. Here we present AIFS-TC, a simple correction to the AIFS-Single model that is competitive with the operational state-of-the-art for forecas… ▽ More

    Submitted 13 August, 2026; v1 submitted 24 July, 2026; originally announced August 2026.

    Comments: 6 pages, 5 figures, 2 tables

  3. arXiv:2606.26421  [pdf, ps, other

    cs.LG cs.CE

    Otter Weather: Skillful and Computationally Efficient Medium-Range Weather Forecasting

    Authors: Cristiana Diaconu, Jonas Scholz, Aliaksandra Shysheya, Stratis Markou, Payel Mukhopadhyay, Miles Cranmer, Richard E. Turner

    Abstract: State-of-the-art medium-range AI weather models can outperform traditional Numerical Weather Prediction (NWP) but require massive training budgets. This restricts usage for under-resourced groups and severely limits fast model iteration. Here we develop Otter Weather, a highly efficient spatiotemporal forecasting model designed to democratise high-performance weather prediction with AI. Evaluated… ▽ More

    Submitted 24 June, 2026; originally announced June 2026.

  4. arXiv:2606.13796  [pdf, ps, other

    stat.ML cs.LG

    Recursively Trained Diffusion Models: Limiting Collapse Distribution and Spectral Characterization

    Authors: Naïl B. Khelifa, Richard E. Turner, Ramji Venkataramanan

    Abstract: Recursive training of generative models on their own outputs can lead to model collapse, a compounding drift away from the true data distribution. Existing theoretical works bound finite-round error accumulation in the context of diffusion models, but two questions remain open:~what distribution does the recursion converge to, and how fast? We answer both, isolating a mechanism distinct from imper… ▽ More

    Submitted 11 June, 2026; originally announced June 2026.

  5. arXiv:2606.06179  [pdf, ps, other

    stat.ML cs.LG

    Diffusion Models Observe Only Gradients: A Geometric Perspective on Score Matching Errors

    Authors: Naïl B. Khelifa, Richard E. Turner, Ramji Venkataramanan

    Abstract: Score-based diffusion models are typically trained by minimizing the $L^2$ score matching error, and standard theoretical analyses rely on this quantity to bound the sampling discrepancy between the learned and target distributions. We show the $L^2$ score error is not the right intrinsic measure of marginal distributional quality: a learned diffusion model can incur arbitrarily large $L^2$ score… ▽ More

    Submitted 28 June, 2026; v1 submitted 4 June, 2026; originally announced June 2026.

  6. arXiv:2605.28153  [pdf, ps, other

    physics.ao-ph cs.LG

    Skillful high-resolution weather forecasting independent of physical models

    Authors: Pengcheng Zhao, Siqi Xiang, Weixin Jin, Zekun Ni, Jiang Bian, Zuliang Fang, Hongyu Sun, Bin Zhang, Richard E. Turner, Jonathan Weyn, Haiyu Dong, Kit Thambiratnam, Qi Zhang

    Abstract: Accurate and timely weather forecasts are critical for high-impact decisions in modern society. Machine-learning-based weather prediction is emerging as an alternative for producing initial conditions, forecasts, and even both in end-to-end systems. These methods deliver predictions faster and often with higher skill than traditional numerical weather prediction (NWP). However, even end-to-end mod… ▽ More

    Submitted 27 May, 2026; originally announced May 2026.

    Comments: 26 pages, 10 figures

  7. arXiv:2603.01949  [pdf, ps, other

    cs.LG cs.AI cs.CE

    Probabilistic Retrofitting of Learned Simulators

    Authors: Cristiana Diaconu, Miles Cranmer, Richard E. Turner, Tanya Marwah, Payel Mukhopadhyay

    Abstract: Dominant approaches for modelling Partial Differential Equations (PDEs) rely on deterministic predictions, yet many physical systems of interest are inherently chaotic and uncertain. While training probabilistic models from scratch is possible, it is computationally expensive and fails to leverage the significant resources already invested in high-performing deterministic backbones. In this work,… ▽ More

    Submitted 22 June, 2026; v1 submitted 2 March, 2026; originally announced March 2026.

    Comments: Code provided at https://github.com/cddcam/lola_crps

  8. arXiv:2602.18955  [pdf, ps, other

    cs.LG

    Incremental Transformer Neural Processes

    Authors: Philip Mortimer, Cristiana Diaconu, Tommy Rochussen, Bruno Mlodozeniec, Richard E. Turner

    Abstract: Neural Processes (NPs), and specifically Transformer Neural Processes (TNPs), have demonstrated remarkable performance across tasks ranging from spatiotemporal forecasting to tabular data modelling. However, many of these applications are inherently sequential, involving continuous data streams such as real-time sensor readings or database updates. In such settings, models should support cheap, in… ▽ More

    Submitted 3 June, 2026; v1 submitted 21 February, 2026; originally announced February 2026.

    Comments: Accepted at ICML 2026

  9. arXiv:2602.16601  [pdf, ps, other

    stat.ML cs.LG

    Quantifying Error Propagation and Model Collapse in Diffusion Models

    Authors: Nail B. Khelifa, Richard E. Turner, Ramji Venkataramanan

    Abstract: Machine learning models are increasingly trained or fine-tuned on synthetic data. Recursively training on such data has been observed to significantly degrade performance in a wide range of tasks, often characterized by a progressive drift away from the target distribution. In this work, we theoretically analyze this phenomenon in the setting of score-based diffusion models. For a realistic pipeli… ▽ More

    Submitted 28 May, 2026; v1 submitted 18 February, 2026; originally announced February 2026.

    Comments: Accepted at ICML 2026

  10. arXiv:2512.23408  [pdf, ps, other

    stat.ML cs.LG stat.ME

    Probabilistic Modelling is Sufficient for Causal Inference

    Authors: Bruno Mlodozeniec, David Krueger, Richard E. Turner

    Abstract: Causal inference is a key research area in machine learning, yet confusion reigns over the tools needed to tackle it. There are prevalent claims in the machine learning literature that you need a bespoke causal framework or notation to answer causal questions. In this paper, we want to make it clear that you \emph{can} answer any causal inference question within the realm of probabilistic modellin… ▽ More

    Submitted 29 December, 2025; originally announced December 2025.

    Journal ref: Proceedings of the 42nd International Conference on Machine Learning, PMLR 267:81810-81840, 2025

  11. arXiv:2511.21777  [pdf, ps, other

    cs.LG

    Artificial intelligence for methane detection: from continuous monitoring to verified mitigation

    Authors: Gonzalo Mateo-Garcia, Anna Allen, Itziar Irakulis-Loitxate, Manuel Montesino-San Martin, Marc Watine, Cynthia Randles, Tharwat Mokalled, Alma Raunak, Carol Castañeda-Martinez, Juan E. Jonhson, Javier Gorroño, James Requeima, Claudio Cifarelli, Luis Guanter, Richard E. Turner, Manfredi Caltagirone

    Abstract: Methane is a potent greenhouse gas, responsible for roughly 30% of warming since pre-industrial times. A small number of large point sources account for a disproportionate share of emissions, creating an opportunity for substantial reductions by targeting relatively few sites. Detection and attribution of large emissions at scale for notification to asset owners remains challenging. Here, we intro… ▽ More

    Submitted 24 April, 2026; v1 submitted 26 November, 2025; originally announced November 2025.

  12. arXiv:2509.22259  [pdf

    cs.LG cs.AI

    Rotary Position Encodings for Graphs

    Authors: Isaac Reid, Arijit Sehanobish, Cederik Höfs, Bruno Mlodozeniec, Leonhard Vulpius, Federico Barbero, Adrian Weller, Krzysztof Choromanski, Richard E. Turner, Petar Veličković

    Abstract: We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large language models (LLMs) and vision transformers (ViTs), can be applied to graph-structured data. We find that rotating tokens depending on the spectrum of the graph Laplacian efficiently injects structural information into the attention mechanism, boosting perform… ▽ More

    Submitted 24 June, 2026; v1 submitted 26 September, 2025; originally announced September 2025.

  13. arXiv:2509.14223  [pdf, ps, other

    cs.LG cs.AI cs.CL

    Fresh in memory: Training-order recency is linearly encoded in language model activations

    Authors: Dmitrii Krasheninnikov, Richard E. Turner, David Krueger

    Abstract: We show that language models' activations linearly encode when information was learned during training. Our setup involves creating a model with a known training order by sequentially fine-tuning Llama-3.2-1B on six disjoint but otherwise similar datasets about named entities. We find that the average activations of test samples corresponding to the six training datasets encode the training order:… ▽ More

    Submitted 22 September, 2025; v1 submitted 17 September, 2025; originally announced September 2025.

  14. arXiv:2509.03691  [pdf

    cs.LG

    Graph Random Features for Scalable Gaussian Processes

    Authors: Matthew Zhang, Jihao Andreas Lin, Krzysztof Choromanski, Adrian Weller, Richard E. Turner, Isaac Reid

    Abstract: We study the application of graph random features (GRFs) - a recently introduced stochastic estimator of graph node kernels - to scalable Gaussian processes on discrete input spaces. We prove that (under mild assumptions) Bayesian inference with GRFs enjoys $O(N^{3/2})$ time complexity with respect to the number of nodes $N$, compared to $O(N^3)$ for exact kernels. Substantial wall-clock speedups… ▽ More

    Submitted 25 September, 2025; v1 submitted 3 September, 2025; originally announced September 2025.

  15. arXiv:2507.09212  [pdf, ps, other

    cs.LG cs.CV stat.ML

    Warm Starts Accelerate Conditional Diffusion

    Authors: Jonas Scholz, Richard E. Turner

    Abstract: Generative models like diffusion and flow-matching create high-fidelity samples by progressively refining noise. The refinement process is notoriously slow, often requiring hundreds of function evaluations. We introduce Warm-Start Diffusion (WSD), a method that uses a simple, deterministic model to dramatically accelerate conditional generation by providing a better starting point. Instead of star… ▽ More

    Submitted 29 September, 2025; v1 submitted 12 July, 2025; originally announced July 2025.

    Comments: 10 pages, 6 figures

  16. arXiv:2507.05526  [pdf, ps, other

    cs.LG stat.ME stat.ML

    Estimating Interventional Distributions with Uncertain Causal Graphs through Meta-Learning

    Authors: Anish Dhir, Cristiana Diaconu, Valentinian Mihai Lungu, James Requeima, Richard E. Turner, Mark van der Wilk

    Abstract: In scientific domains -- from biology to the social sciences -- many questions boil down to \textit{What effect will we observe if we intervene on a particular variable?} If the causal relationships (e.g.~a causal graph) are known, it is possible to estimate the intervention distributions. In the absence of this domain knowledge, the causal structure must be discovered from the available observati… ▽ More

    Submitted 10 February, 2026; v1 submitted 7 July, 2025; originally announced July 2025.

  17. arXiv:2506.12965  [pdf

    cs.LG cs.AI stat.ML

    Distributional Training Data Attribution: What do Influence Functions Sample?

    Authors: Bruno Mlodozeniec, Isaac Reid, Sam Power, David Krueger, Murat Erdogdu, Richard E. Turner, Roger Grosse

    Abstract: Randomness is an unavoidable part of training deep learning models, yet something that traditional training data attribution algorithms fail to rigorously account for. They ignore the fact that, due to stochasticity in the initialisation and batching, training on the same dataset can yield different models. In this paper, we address this shortcoming through introducing distributional training data… ▽ More

    Submitted 25 October, 2025; v1 submitted 15 June, 2025; originally announced June 2025.

  18. arXiv:2506.03595  [pdf, ps, other

    cs.LG cs.AI stat.ML

    Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its Preconditioner

    Authors: Runa Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner, Hao-Jun Michael Shi

    Abstract: The recent success of Shampoo in the AlgoPerf contest has sparked renewed interest in Kronecker-factorization-based optimization algorithms for training neural networks. Despite its success, Shampoo relies heavily on several heuristics such as learning rate grafting and stale preconditioning to achieve performance at-scale. These heuristics increase algorithmic complexity, necessitate further hype… ▽ More

    Submitted 29 October, 2025; v1 submitted 4 June, 2025; originally announced June 2025.

    Comments: NeurIPS 2025

  19. arXiv:2502.11877  [pdf, other

    stat.ML cs.LG

    JoLT: Joint Probabilistic Predictions on Tabular Data Using LLMs

    Authors: Aliaksandra Shysheya, John Bronskill, James Requeima, Shoaib Ahmed Siddiqui, Javier Gonzalez, David Duvenaud, Richard E. Turner

    Abstract: We introduce a simple method for probabilistic predictions on tabular data based on Large Language Models (LLMs) called JoLT (Joint LLM Process for Tabular data). JoLT uses the in-context learning capabilities of LLMs to define joint distributions over tabular data conditioned on user-specified side information about the problem, exploiting the vast repository of latent problem-relevant knowledge… ▽ More

    Submitted 17 February, 2025; originally announced February 2025.

  20. arXiv:2502.06268  [pdf, other

    stat.ML cs.LG

    Spectral-factorized Positive-definite Curvature Learning for NN Training

    Authors: Wu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae, Richard E. Turner, Roger B. Grosse

    Abstract: Many training methods, such as Adam(W) and Shampoo, learn a positive-definite curvature matrix and apply an inverse root before preconditioning. Recently, non-diagonal training methods, such as Shampoo, have gained significant attention; however, they remain computationally inefficient and are limited to specific types of curvature information due to the costly matrix root computation via matrix d… ▽ More

    Submitted 28 March, 2025; v1 submitted 10 February, 2025; originally announced February 2025.

    Comments: fixed some typos in the appendix

  21. arXiv:2502.04750  [pdf, other

    stat.ML cs.LG

    Tighter sparse variational Gaussian processes

    Authors: Thang D. Bui, Matthew Ashman, Richard E. Turner

    Abstract: Sparse variational Gaussian process (GP) approximations based on inducing points have become the de facto standard for scaling GPs to large datasets, owing to their theoretical elegance, computational efficiency, and ease of implementation. This paper introduces a provably tighter variational approximation by relaxing the standard assumption that the conditional approximate posterior given the ind… ▽ More

    Submitted 13 February, 2025; v1 submitted 7 February, 2025; originally announced February 2025.

  22. arXiv:2502.04098  [pdf, other

    cs.CV cs.AI

    Efficient Few-Shot Continual Learning in Vision-Language Models

    Authors: Aristeidis Panos, Rahaf Aljundi, Daniel Olmeda Reino, Richard E. Turner

    Abstract: Vision-language models (VLMs) excel in tasks such as visual question answering and image captioning. However, VLMs are often limited by their use of pretrained image encoders, like CLIP, leading to image understanding errors that hinder overall performance. On top of that, real-world applications often require the model to be continuously adapted as new and often limited data continuously arrive.… ▽ More

    Submitted 7 February, 2025; v1 submitted 6 February, 2025; originally announced February 2025.

  23. arXiv:2501.15690  [pdf, ps, other

    physics.ao-ph cs.LG stat.ML

    Refined climatologies of future precipitation over High Mountain Asia using probabilistic ensemble learning

    Authors: Kenza Tazi, Sun Woo P. Kim, Marc Girona-Mata, Richard E. Turner

    Abstract: High Mountain Asia (HMA) holds the highest concentration of frozen water outside the polar regions, serving as a crucial water source for more than 1.9 billion people. Precipitation represents the largest source of uncertainty for future hydrological modelling in this area. In this study, we propose a probabilistic machine learning framework to combine monthly precipitation from 13 regional climat… ▽ More

    Submitted 30 June, 2025; v1 submitted 26 January, 2025; originally announced January 2025.

    Comments: 16 pages 8 figures (main text), 32 pages 14 figures (total)

  24. arXiv:2410.16415  [pdf, other

    cs.LG cs.AI

    On conditional diffusion models for PDE simulations

    Authors: Aliaksandra Shysheya, Cristiana Diaconu, Federico Bergamin, Paris Perdikaris, José Miguel Hernández-Lobato, Richard E. Turner, Emile Mathieu

    Abstract: Modelling partial differential equations (PDEs) is of crucial importance in science and engineering, and it includes tasks ranging from forecasting to inverse problems, such as data assimilation. However, most previous numerical and machine learning approaches that target forecasting cannot be applied out-of-the-box for data assimilation. Recently, diffusion models have emerged as a powerful tool… ▽ More

    Submitted 21 October, 2024; originally announced October 2024.

    Comments: Accepted at NeurIPS 2024

  25. arXiv:2410.06731  [pdf, other

    stat.ML cs.LG

    Gridded Transformer Neural Processes for Large Unstructured Spatio-Temporal Data

    Authors: Matthew Ashman, Cristiana Diaconu, Eric Langezaal, Adrian Weller, Richard E. Turner

    Abstract: Many important problems require modelling large-scale spatio-temporal datasets, with one prevalent example being weather forecasting. Recently, transformer-based approaches have shown great promise in a range of weather forecasting problems. However, these have mostly focused on gridded data sources, neglecting the wealth of unstructured, off-the-grid data from observational measurements such as t… ▽ More

    Submitted 10 October, 2024; v1 submitted 9 October, 2024; originally announced October 2024.

  26. arXiv:2410.03462  [pdf, other

    cs.LG stat.ML

    Linear Transformer Topological Masking with Graph Random Features

    Authors: Isaac Reid, Kumar Avinava Dubey, Deepali Jain, Will Whitney, Amr Ahmed, Joshua Ainslie, Alex Bewley, Mithun Jacob, Aranyak Mehta, David Rendleman, Connor Schenck, Richard E. Turner, René Wagner, Adrian Weller, Krzysztof Choromanski

    Abstract: When training transformers on graph-structured data, incorporating information about the underlying topology is crucial for good performance. Topological masking, a type of relative position encoding, achieves this by upweighting or downweighting attention depending on the relationship between the query and keys in a graph. In this paper, we propose to parameterise topological masks as a learnable… ▽ More

    Submitted 15 October, 2024; v1 submitted 4 October, 2024; originally announced October 2024.

  27. arXiv:2408.04745  [pdf, other

    cs.AI physics.ao-ph

    AI for operational methane emitter monitoring from space

    Authors: Anna Vaughan, Gonzalo Mateo-Garcia, Itziar Irakulis-Loitxate, Marc Watine, Pablo Fernandez-Poblaciones, Richard E. Turner, James Requeima, Javier Gorroño, Cynthia Randles, Manfredi Caltagirone, Claudio Cifarelli

    Abstract: Mitigating methane emissions is the fastest way to stop global warming in the short-term and buy humanity time to decarbonise. Despite the demonstrated ability of remote sensing instruments to detect methane plumes, no system has been available to routinely monitor and act on these events. We present MARS-S2L, an automated AI-driven methane emitter monitoring system for Sentinel-2 and Landsat sate… ▽ More

    Submitted 8 August, 2024; originally announced August 2024.

  28. arXiv:2407.16526  [pdf, other

    cs.CV cs.AI cs.CL cs.LG

    Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models

    Authors: Aristeidis Panos, Rahaf Aljundi, Daniel Olmeda Reino, Richard E Turner

    Abstract: Vision language models (VLMs) demonstrate impressive capabilities in visual question answering and image captioning, acting as a crucial link between visual and language models. However, existing open-source VLMs heavily rely on pretrained and frozen vision encoders (such as CLIP). Despite CLIP's robustness across diverse domains, it still exhibits non-negligible image understanding errors. These… ▽ More

    Submitted 23 July, 2024; originally announced July 2024.

  29. arXiv:2406.13493  [pdf, other

    cs.LG stat.ML

    In-Context In-Context Learning with Transformer Neural Processes

    Authors: Matthew Ashman, Cristiana Diaconu, Adrian Weller, Richard E. Turner

    Abstract: Neural processes (NPs) are a powerful family of meta-learning models that seek to approximate the posterior predictive map of the ground-truth stochastic process from which each dataset in a meta-dataset is sampled. There are many cases in which practitioners, besides having access to the dataset of interest, may also have access to other datasets that share similarities with it. In this case, int… ▽ More

    Submitted 19 June, 2024; originally announced June 2024.

  30. arXiv:2406.13488  [pdf, other

    stat.ML cs.LG

    Approximately Equivariant Neural Processes

    Authors: Matthew Ashman, Cristiana Diaconu, Adrian Weller, Wessel Bruinsma, Richard E. Turner

    Abstract: Equivariant deep learning architectures exploit symmetries in learning problems to improve the sample efficiency of neural-network-based models and their ability to generalise. However, when modelling real-world data, learning problems are often not exactly equivariant, but only approximately. For example, when estimating the global temperature field from weather station observations, local topogr… ▽ More

    Submitted 9 November, 2024; v1 submitted 19 June, 2024; originally announced June 2024.

  31. arXiv:2406.13151  [pdf, other

    stat.ML cs.LG stat.CO

    Bayesian Circular Regression with von Mises Quasi-Processes

    Authors: Yarden Cohen, Alexandre Khae Wu Navarro, Jes Frellsen, Richard E. Turner, Raziel Riemer, Ari Pakman

    Abstract: The need for regression models to predict circular values arises in many scientific fields. In this work we explore a family of expressive and interpretable distributions over circle-valued random functions related to Gaussian processes targeting two Euclidean dimensions conditioned on the unit circle. The probability model has connections with continuous spin models in statistical physics. Moreov… ▽ More

    Submitted 18 March, 2025; v1 submitted 18 June, 2024; originally announced June 2024.

    Journal ref: Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (AISTATS) 2025

  32. arXiv:2406.12409  [pdf, other

    stat.ML cs.LG

    Translation Equivariant Transformer Neural Processes

    Authors: Matthew Ashman, Cristiana Diaconu, Junhyuck Kim, Lakee Sivaraya, Stratis Markou, James Requeima, Wessel P. Bruinsma, Richard E. Turner

    Abstract: The effectiveness of neural processes (NPs) in modelling posterior prediction maps -- the mapping from data to posterior predictive distributions -- has significantly improved since their inception. This improvement can be attributed to two principal factors: (1) advancements in the architecture of permutation invariant set functions, which are intrinsic to all NPs; and (2) leveraging symmetries p… ▽ More

    Submitted 18 June, 2024; originally announced June 2024.

  33. arXiv:2406.08569  [pdf, other

    cs.LG cs.CR stat.ML

    Noise-Aware Differentially Private Regression via Meta-Learning

    Authors: Ossi Räisä, Stratis Markou, Matthew Ashman, Wessel P. Bruinsma, Marlon Tobaben, Antti Honkela, Richard E. Turner

    Abstract: Many high-stakes applications require machine learning models that protect user privacy and provide well-calibrated, accurate predictions. While Differential Privacy (DP) is the gold standard for protecting user privacy, standard DP mechanisms typically significantly impair performance. One approach to mitigating this issue is pre-training models on simulated data before DP learning on the private… ▽ More

    Submitted 8 May, 2025; v1 submitted 12 June, 2024; originally announced June 2024.

    Comments: NeurIPS 2024

  34. arXiv:2406.01801  [pdf, other

    stat.ML cs.LG

    Fearless Stochasticity in Expectation Propagation

    Authors: Jonathan So, Richard E. Turner

    Abstract: Expectation propagation (EP) is a family of algorithms for performing approximate inference in probabilistic models. The updates of EP involve the evaluation of moments -- expectations of certain functions -- which can be estimated from Monte Carlo (MC) samples. However, the updates are not robust to MC noise when performed naively, and various prior works have attempted to address this issue in d… ▽ More

    Submitted 29 October, 2024; v1 submitted 3 June, 2024; originally announced June 2024.

  35. arXiv:2405.16541  [pdf, other

    stat.ML cs.LG

    Variance-Reducing Couplings for Random Features

    Authors: Isaac Reid, Stratis Markou, Krzysztof Choromanski, Richard E. Turner, Adrian Weller

    Abstract: Random features (RFs) are a popular technique to scale up kernel methods in machine learning, replacing exact kernel evaluations with stochastic Monte Carlo estimates. They underpin models as diverse as efficient transformers (by approximating attention) to sparse spectrum Gaussian processes (by approximating the covariance function). Efficiency can be further improved by speeding up the convergen… ▽ More

    Submitted 2 October, 2024; v1 submitted 26 May, 2024; originally announced May 2024.

  36. arXiv:2405.13063  [pdf, other

    physics.ao-ph cs.LG

    A Foundation Model for the Earth System

    Authors: Cristian Bodnar, Wessel P. Bruinsma, Ana Lucic, Megan Stanley, Anna Vaughan, Johannes Brandstetter, Patrick Garvan, Maik Riechert, Jonathan A. Weyn, Haiyu Dong, Jayesh K. Gupta, Kit Thambiratnam, Alexander T. Archibald, Chun-Chieh Wu, Elizabeth Heider, Max Welling, Richard E. Turner, Paris Perdikaris

    Abstract: Reliable forecasts of the Earth system are crucial for human progress and safety from natural disasters. Artificial intelligence offers substantial potential to improve prediction accuracy and computational efficiency in this field, however this remains underexplored in many domains. Here we introduce Aurora, a large-scale foundation model for the Earth system trained on over a million hours of di… ▽ More

    Submitted 21 November, 2024; v1 submitted 20 May, 2024; originally announced May 2024.

  37. arXiv:2405.12856  [pdf, other

    stat.ML cs.CL cs.LG

    LLM Processes: Numerical Predictive Distributions Conditioned on Natural Language

    Authors: James Requeima, John Bronskill, Dami Choi, Richard E. Turner, David Duvenaud

    Abstract: Machine learning practitioners often face significant challenges in formally integrating their prior knowledge and beliefs into predictive models, limiting the potential for nuanced and context-aware analyses. Moreover, the expertise needed to integrate this prior knowledge into probabilistic modeling typically limits the application of these models to specialists. Our goal is to build a regressio… ▽ More

    Submitted 19 December, 2024; v1 submitted 21 May, 2024; originally announced May 2024.

    Journal ref: 38th Conference on Neural Information Processing Systems (NeurIPS 2024)

  38. arXiv:2404.00411  [pdf, other

    physics.ao-ph cs.LG

    Aardvark weather: end-to-end data-driven weather forecasting

    Authors: Anna Vaughan, Stratis Markou, Will Tebbutt, James Requeima, Wessel P. Bruinsma, Tom R. Andersson, Michael Herzog, Nicholas D. Lane, Matthew Chantry, J. Scott Hosking, Richard E. Turner

    Abstract: Weather forecasting is critical for a range of human activities including transportation, agriculture, industry, as well as the safety of the general public. Machine learning models have the potential to transform the complex weather prediction pipeline, but current approaches still rely on numerical weather prediction (NWP) systems, limiting forecast speed and accuracy. Here we demonstrate that a… ▽ More

    Submitted 13 July, 2024; v1 submitted 30 March, 2024; originally announced April 2024.

  39. arXiv:2403.12977  [pdf, other

    cs.CV cs.LG eess.IV stat.AP

    SportsNGEN: Sustained Generation of Realistic Multi-player Sports Gameplay

    Authors: Lachlan Thorpe, Lewis Bawden, Karanjot Vendal, John Bronskill, Richard E. Turner

    Abstract: We present a transformer decoder based sports simulation engine, SportsNGEN, trained on sports player and ball tracking sequences, that is capable of generating sustained gameplay and accurately mimicking the decision making of real players. By training on a large database of professional tennis tracking data, we demonstrate that simulations produced by SportsNGEN can be used to predict the outcom… ▽ More

    Submitted 18 November, 2024; v1 submitted 9 February, 2024; originally announced March 2024.

    Journal ref: Proceedings of the 12th International Conference on Sport Sciences Research and Technology Support (icSPORTS 2024)

  40. arXiv:2403.01946  [pdf, other

    cs.LG

    A Generative Model of Symmetry Transformations

    Authors: James Urquhart Allingham, Bruno Kacper Mlodozeniec, Shreyas Padhy, Javier Antorán, David Krueger, Richard E. Turner, Eric Nalisnick, José Miguel Hernández-Lobato

    Abstract: Correctly capturing the symmetry transformations of data can lead to efficient models with strong generalization capabilities, though methods incorporating symmetries often require prior knowledge. While recent advancements have been made in learning those symmetries directly from the dataset, most of this work has focused on the discriminative setting. In this paper, we take inspiration from grou… ▽ More

    Submitted 2 December, 2024; v1 submitted 4 March, 2024; originally announced March 2024.

    Comments: Accepted at NeurIPS 2024

  41. arXiv:2402.04384  [pdf, other

    cs.LG stat.ML

    Denoising Diffusion Probabilistic Models in Six Simple Steps

    Authors: Richard E. Turner, Cristiana-Diana Diaconu, Stratis Markou, Aliaksandra Shysheya, Andrew Y. K. Foong, Bruno Mlodozeniec

    Abstract: Denoising Diffusion Probabilistic Models (DDPMs) are a very popular class of deep generative model that have been successfully applied to a diverse range of problems including image and video generation, protein and material synthesis, weather forecasting, and neural surrogates of partial differential equations. Despite their ubiquity it is hard to find an introduction to DDPMs which is simple, co… ▽ More

    Submitted 10 February, 2024; v1 submitted 6 February, 2024; originally announced February 2024.

  42. arXiv:2402.03496  [pdf, other

    cs.LG math.OC

    Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective

    Authors: Wu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae, Richard E. Turner, Alireza Makhzani

    Abstract: Adaptive gradient optimizers like Adam(W) are the default training algorithms for many deep learning architectures, such as transformers. Their diagonal preconditioner is based on the gradient outer product which is incorporated into the parameter update via a square root. While these methods are often motivated as approximate second-order methods, the square root represents a fundamental differen… ▽ More

    Submitted 4 October, 2024; v1 submitted 5 February, 2024; originally announced February 2024.

    Comments: A long version of the ICML 2024 paper. Updated the caption of Fig 4 to emphasize the importance of the scale invariance of root-free methods

  43. arXiv:2401.01855  [pdf, other

    cs.LG

    Transformer Neural Autoregressive Flows

    Authors: Massimiliano Patacchiola, Aliaksandra Shysheya, Katja Hofmann, Richard E. Turner

    Abstract: Density estimation, a central problem in machine learning, can be performed using Normalizing Flows (NFs). NFs comprise a sequence of invertible transformations, that turn a complex target distribution into a simple one, by exploiting the change of variables theorem. Neural Autoregressive Flows (NAFs) and Block Neural Autoregressive Flows (B-NAFs) are arguably the most perfomant members of the NF… ▽ More

    Submitted 3 January, 2024; originally announced January 2024.

  44. arXiv:2312.05705  [pdf, other

    cs.LG stat.ML

    Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC

    Authors: Wu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov, Agustinus Kristiadi, Richard E. Turner, Alireza Makhzani

    Abstract: Second-order methods such as KFAC can be useful for neural net training. However, they are often memory-inefficient since their preconditioning Kronecker factors are dense, and numerically unstable in low precision as they require matrix inversion or decomposition. These limitations render such methods unpopular for modern mixed-precision training. We address them by (i) formulating an inverse-fre… ▽ More

    Submitted 23 July, 2024; v1 submitted 9 December, 2023; originally announced December 2023.

    Comments: A long version of the ICML 2024 paper, updated the text about a related work

  45. arXiv:2311.16849  [pdf, other

    stat.ML cs.LG

    Identifiable Feature Learning for Spatial Data with Nonlinear ICA

    Authors: Hermanni Hälvä, Jonathan So, Richard E. Turner, Aapo Hyvärinen

    Abstract: Recently, nonlinear ICA has surfaced as a popular alternative to the many heuristic models used in deep representation learning and disentanglement. An advantage of nonlinear ICA is that a sophisticated identifiability theory has been developed; in particular, it has been proven that the original components can be recovered under sufficiently strong latent dependencies. Despite this general theory… ▽ More

    Submitted 28 November, 2023; originally announced November 2023.

    Comments: Work under review

  46. arXiv:2311.09848  [pdf, other

    cs.LG

    Diffusion-Augmented Neural Processes

    Authors: Lorenzo Bonito, James Requeima, Aliaksandra Shysheya, Richard E. Turner

    Abstract: Over the last few years, Neural Processes have become a useful modelling tool in many application areas, such as healthcare and climate sciences, in which data are scarce and prediction uncertainty estimates are indispensable. However, the current state of the art in the field (AR CNPs; Bruinsma et al., 2023) presents a few issues that prevent its widespread deployment. This work proposes an alter… ▽ More

    Submitted 16 November, 2023; originally announced November 2023.

    Comments: Accepted to the NeurIPS 2023 Workshop on Diffusion Models

    ACM Class: I.2.6

  47. arXiv:2311.00636  [pdf, other

    cs.LG stat.ML

    Kronecker-Factored Approximate Curvature for Modern Neural Network Architectures

    Authors: Runa Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider, Philipp Hennig

    Abstract: The core components of many modern neural network architectures, such as transformers, convolutional, or graph neural networks, can be expressed as linear layers with $\textit{weight-sharing}$. Kronecker-Factored Approximate Curvature (K-FAC), a second-order optimisation method, has shown promise to speed up neural network training and thereby reduce computational costs. However, there is currentl… ▽ More

    Submitted 11 January, 2024; v1 submitted 1 November, 2023; originally announced November 2023.

    Comments: NeurIPS 2023

  48. arXiv:2310.19932  [pdf, other

    cs.LG physics.ao-ph

    Sim2Real for Environmental Neural Processes

    Authors: Jonas Scholz, Tom R. Andersson, Anna Vaughan, James Requeima, Richard E. Turner

    Abstract: Machine learning (ML)-based weather models have recently undergone rapid improvements. These models are typically trained on gridded reanalysis data from numerical data assimilation systems. However, reanalysis data comes with limitations, such as assumptions about physical laws and low spatiotemporal resolution. The gap between reanalysis and reality has sparked growing interest in training ML mo… ▽ More

    Submitted 30 October, 2023; originally announced October 2023.

    Comments: 4 pages, 3 figures, To be published in Tackling Climate Change with Machine Learning workshop at NeurIPS

  49. arXiv:2310.11837  [pdf, other

    stat.ML cs.LG

    Optimising Distributions with Natural Gradient Surrogates

    Authors: Jonathan So, Richard E. Turner

    Abstract: Natural gradient methods have been used to optimise the parameters of probability distributions in a variety of settings, often resulting in fast-converging procedures. Unfortunately, for many distributions of interest, computing the natural gradient has a number of challenges. In this work we propose a novel technique for tackling such issues, which involves reframing the optimisation as one with… ▽ More

    Submitted 4 March, 2024; v1 submitted 18 October, 2023; originally announced October 2023.

    Journal ref: PMLR 238 (2024):2224-2232

  50. arXiv:2308.05732  [pdf, other

    cs.LG cs.AI

    PDE-Refiner: Achieving Accurate Long Rollouts with Neural PDE Solvers

    Authors: Phillip Lippe, Bastiaan S. Veeling, Paris Perdikaris, Richard E. Turner, Johannes Brandstetter

    Abstract: Time-dependent partial differential equations (PDEs) are ubiquitous in science and engineering. Recently, mostly due to the high computational cost of traditional solution techniques, deep neural network based surrogates have gained increased interest. The practical utility of such neural PDE solvers relies on their ability to provide accurate, stable predictions over long time horizons, which is… ▽ More

    Submitted 21 October, 2023; v1 submitted 10 August, 2023; originally announced August 2023.

    Comments: Project website: https://phlippe.github.io/PDERefiner/