-
Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
Authors:
Batu El,
Jinhee Paeng,
Fatih Dinc,
Shiye Su,
Mete Erdogan,
Aneesh Pappu,
Haotian Ye,
Wanjia Zhao,
Surya Ganguli,
James Zou
Abstract:
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent syste…
▽ More
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.
△ Less
Submitted 17 August, 2026;
originally announced August 2026.
-
An exact information theory of generalization phase transitions in Bayesian diffusion models
Authors:
Henry Hunt,
Mason Kamb,
Surya Ganguli
Abstract:
How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model ti…
▽ More
How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models \textit{early in training}, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.
△ Less
Submitted 8 July, 2026;
originally announced July 2026.
-
Letting the neural code speak: Automated characterization of monkey visual neurons through human language
Authors:
Vedang Lad,
Katrin Franke,
Tamar Rott Shaham,
Surya Ganguli,
Andreas S. Tolias,
Sophia Sanborn,
Nikos Karantzas
Abstract:
Understanding what individual neurons encode is a core question in neuroscience. In primary visual cortex (V1), mathematical models (e.g., Gabor functions) capture neural selectivity, but no comparable framework exists for higher areas. We show that natural language can fill this role: across macaque V1 and V4, the selectivity of most neurons is captured by concise, verifiable semantic description…
▽ More
Understanding what individual neurons encode is a core question in neuroscience. In primary visual cortex (V1), mathematical models (e.g., Gabor functions) capture neural selectivity, but no comparable framework exists for higher areas. We show that natural language can fill this role: across macaque V1 and V4, the selectivity of most neurons is captured by concise, verifiable semantic descriptions. Using digital twins of V1 and V4, we develop a closed-loop framework that translates each neuron's high- and low-activating images into dense captions, generates a semantic hypothesis and synthesized images, and verifies the hypothesis in silico. Descriptions range from oriented edges and spatial frequency in V1 to conjunctions of form, color, and texture in V4. In V4, images generated from activating and suppressing hypotheses drove 96.1% of neurons above the 95th and 97.6% below the 5th percentile of natural-image responses, respectively (vs. ~10% for random images); V1 activation results matched V4, while V1 suppression was less describable in language. Representational similarity analysis reveals partial alignment between neural activity, vision embeddings, and language embeddings, with vision most aligned to neural activity; alignment lost in the text bottleneck is recovered when hypotheses are rendered back into images, showing that linguistic compression is lossy yet semantically faithful. Together, these results show that combining generative models with neural digital twins enables interpretable, testable descriptions of neural function at scale, toward agentic scientific discovery.
△ Less
Submitted 18 May, 2026; v1 submitted 12 May, 2026;
originally announced May 2026.
-
Causal Interpretation of Neural Network Computations with Contribution Decomposition
Authors:
Joshua Brendan Melander,
Zaki Alaoui,
Shenghua Liu,
Surya Ganguli,
Stephen A. Baccus
Abstract:
Understanding how neural networks transform inputs into outputs is crucial for interpreting and manipulating their behavior. Most existing approaches analyze internal representations by identifying hidden-layer activation patterns correlated with human-interpretable concepts. Here we take a direct approach to examine how hidden neurons act to drive network outputs. We introduce CODEC (Contribution…
▽ More
Understanding how neural networks transform inputs into outputs is crucial for interpreting and manipulating their behavior. Most existing approaches analyze internal representations by identifying hidden-layer activation patterns correlated with human-interpretable concepts. Here we take a direct approach to examine how hidden neurons act to drive network outputs. We introduce CODEC (Contribution Decomposition), a method that uses sparse autoencoders to decompose network behavior into sparse motifs of hidden-neuron contributions, revealing causal processes that cannot be determined by analyzing activations alone. Applying CODEC to benchmark image-classification networks, we find that contributions grow in sparsity and dimensionality across layers and, unexpectedly, that they progressively decorrelate positive and negative effects on network outputs. We further show that decomposing contributions into sparse modes enables greater control and interpretation of intermediate layers, supporting both causal manipulations of network output and human-interpretable visualizations of distinct image components that combine to drive that output. Finally, by analyzing state-of-the-art models of neural activity in the vertebrate retina, we demonstrate that CODEC uncovers combinatorial actions of model interneurons and identifies the sources of dynamic receptive fields. Overall, CODEC provides a rich and interpretable framework for understanding how nonlinear computations evolve across hierarchical layers, establishing contribution modes as an informative unit of analysis for mechanistic insights into artificial neural networks.
△ Less
Submitted 6 March, 2026;
originally announced March 2026.
-
Solving adversarial examples requires solving exponential misalignment
Authors:
Alessandro Salvatore,
Stanislav Fort,
Surya Ganguli
Abstract:
Adversarial attacks - input perturbations imperceptible to humans that fool neural networks - remain both a persistent failure mode in machine learning, and a phenomenon with mysterious origins. To shed light, we define and analyze a network's perceptual manifold (PM) for a class concept as the space of all inputs confidently assigned to that class by the network. We find, strikingly, that the dim…
▽ More
Adversarial attacks - input perturbations imperceptible to humans that fool neural networks - remain both a persistent failure mode in machine learning, and a phenomenon with mysterious origins. To shed light, we define and analyze a network's perceptual manifold (PM) for a class concept as the space of all inputs confidently assigned to that class by the network. We find, strikingly, that the dimensionalities of neural network PMs are orders of magnitude higher than those of natural human concepts. Since volume typically grows exponentially with dimension, this suggests exponential misalignment between machines and humans, with exponentially many inputs confidently assigned to concepts by machines but not humans. Furthermore, this provides a natural geometric hypothesis for the origin of adversarial examples: because a network's PM fills such a large region of input space, any input will be very close to any class concept's PM. Our hypothesis thus suggests that adversarial robustness cannot be attained without dimensional alignment of machine and human PMs, and therefore makes strong predictions: both robust accuracy and distance to any PM should be negatively correlated with the PM dimension. We confirmed these predictions across 18 different networks of varying robust accuracy. Crucially, we find even the most robust networks are still exponentially misaligned, and only the few PMs whose dimensionality approaches that of human concepts exhibit alignment to human perception. Our results connect the fields of alignment and adversarial examples, and suggest the curse of high dimensionality of machine PMs is a major impediment to adversarial robustness.
△ Less
Submitted 10 March, 2026; v1 submitted 3 March, 2026;
originally announced March 2026.
-
Deriving Neural Scaling Laws from the statistics of natural language
Authors:
Francesco Cagnetta,
Allan Raventós,
Surya Ganguli,
Matthieu Wyart
Abstract:
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of lan…
▽ More
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset. We provide the first such theory in the case of data-limited scaling laws. We isolate two key statistical properties of language that alone can predict neural scaling exponents: (i) the decay of pairwise token correlations with time separation between token pairs, and (ii) the decay of the next-token conditional entropy with the length of the conditioning context. We further derive a simple formula in terms of these statistics that predicts data-limited neural scaling exponents from first principles without any free parameters or synthetic data models. Our theory exhibits a remarkable match with experimentally measured neural scaling laws obtained from training GPT-2 and LLaMA style models from scratch on two qualitatively different benchmarks, TinyStories and WikiText.
△ Less
Submitted 2 July, 2026; v1 submitted 7 February, 2026;
originally announced February 2026.
-
From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
Authors:
Ziming Liu,
Sophia Sanborn,
Surya Ganguli,
Andreas Tolias
Abstract:
Can general-purpose AI architectures go beyond prediction to discover the physical laws governing the universe? True intelligence relies on "world models" -- causal abstractions that allow an agent to not only predict future states but understand the underlying governing dynamics. While previous "AI Physicist" approaches have successfully recovered such laws, they typically rely on strong, domain-…
▽ More
Can general-purpose AI architectures go beyond prediction to discover the physical laws governing the universe? True intelligence relies on "world models" -- causal abstractions that allow an agent to not only predict future states but understand the underlying governing dynamics. While previous "AI Physicist" approaches have successfully recovered such laws, they typically rely on strong, domain-specific priors that effectively "bake in" the physics. Conversely, Vafa et al. recently showed that generic Transformers fail to acquire these world models, achieving high predictive accuracy without capturing the underlying physical laws. We bridge this gap by systematically introducing three minimal inductive biases. We show that ensuring spatial smoothness (by formulating prediction as continuous regression) and stability (by training with noisy contexts to mitigate error accumulation) enables generic Transformers to surpass prior failures and learn a coherent Keplerian world model, successfully fitting ellipses to planetary trajectories. However, true physical insight requires a third bias: temporal locality. By restricting the attention window to the immediate past -- imposing the simple assumption that future states depend only on the local state rather than a complex history -- we force the model to abandon curve-fitting and discover Newtonian force representations. Our results demonstrate that simple architectural choices determine whether an AI becomes a curve-fitter or a physicist, marking a critical step toward automated scientific discovery.
△ Less
Submitted 6 February, 2026;
originally announced February 2026.
-
Contrastive Concept-Tree Search for LLM-Assisted Algorithm Discovery
Authors:
Timothee Leleu,
Sudeera Gunathilaka,
Federico Ghimenti,
Surya Ganguli
Abstract:
Large language Model (LLM)-assisted algorithm discovery is an iterative, black-box optimization process over programs to approximatively solve a target task, where an LLM proposes candidate programs and an external evaluator provides task feedback. Despite intense recent research on the topic and promising results, how can the LLM internal representation of the space of possible programs be maxima…
▽ More
Large language Model (LLM)-assisted algorithm discovery is an iterative, black-box optimization process over programs to approximatively solve a target task, where an LLM proposes candidate programs and an external evaluator provides task feedback. Despite intense recent research on the topic and promising results, how can the LLM internal representation of the space of possible programs be maximally exploited to improve performance is an open question. Here, we introduce Contrastive Concept-Tree Search (CCTS), which extracts a hierarchical concept representation from the generated programs and learns a contrastive concept model that guides parent selection. By reweighting parents using a likelihood-ratio score between high- and low-performing solutions, CCTS biases search toward useful concept combinations and away from misleading ones, providing guidance through an explicit concept hierarchy rather than the algorithm lineage constructed by the LLM. We show that CCTS improves search efficiency over fitness-based baselines and produces interpretable, task-specific concept trees across a benchmark of open Erdős-type combinatorics problems. Our analysis indicates that the gains are driven largely by learning which concepts to avoid. We further validate these findings in a controlled synthetic algorithm-discovery environment, which reproduces qualitatively the search dynamics observed with the LLMs.
△ Less
Submitted 3 February, 2026;
originally announced February 2026.
-
Reshaping Global Loop Structure to Accelerate Local Optimization by Smoothing Rugged Landscapes
Authors:
Timothee Leleu,
Sam Reifenstein,
Atsushi Yamamura,
Surya Ganguli
Abstract:
Probabilistic graphical models with frustration exhibit rugged energy landscapes that trap iterative optimization dynamics. These landscapes are shaped not only by local interactions, but crucially also by the global loop structure of the graph. The famous Bethe approximation treats the graph as a tree, effectively ignoring global structure, thereby limiting its effectiveness for optimization. Loo…
▽ More
Probabilistic graphical models with frustration exhibit rugged energy landscapes that trap iterative optimization dynamics. These landscapes are shaped not only by local interactions, but crucially also by the global loop structure of the graph. The famous Bethe approximation treats the graph as a tree, effectively ignoring global structure, thereby limiting its effectiveness for optimization. Loop expansions capture such global structure in principle, but are often impractical due to combinatorial explosion. The $M$-layer construction provides an alternative: make $M$ copies of the graph and reconnect edges between them uniformly at random. This provides a controlled sequence of approximations from the original graph at $M=1$, to the Bethe approximation as $M \rightarrow \infty$. Here we generalize this construction by replacing uniform random rewiring with a structured mixing kernel $Q$ that sets the probability that any two layers are interconnected. As a result, the global loop structure can be shaped without modifying local interactions. We show that, after this copy-and-reconnect transformation, there exists a regime in which layer-to-layer fluctuations decay, increasing the probability of reaching the global minimum of the energy function of the original graph. This yields a highly general and practical tool for optimization. Using this approach, the computational cost required to reach these optimal solutions is reduced across sparse and dense Ising benchmarks, including spin glasses and planted instances. When combined with replica-exchange Monte Carlo, the same construction increases the polynomial-time algorithmic threshold for the maximum independent set problem. A cavity analysis shows that structured inter-layer coupling significantly smooths rugged energy landscapes by collapsing configurational complexity and suppressing many suboptimal metastable states.
△ Less
Submitted 1 February, 2026;
originally announced February 2026.
-
Short-term plasticity recalls forgotten memories through a trampoline mechanism
Authors:
Martina Del Gaudio,
Federico Ghimenti,
Surya Ganguli
Abstract:
We analyze continuous Hopfield associative memories augmented by additional, rapid short-term associative synaptic plasticity. Through the cavity method, we determine the boundary between the retrieval and forgetting, or spin-glass phase, of the network as a function of the fraction of stored memories and the neuronal gain. We find that short-term synaptic plasticity yields marginal improvements i…
▽ More
We analyze continuous Hopfield associative memories augmented by additional, rapid short-term associative synaptic plasticity. Through the cavity method, we determine the boundary between the retrieval and forgetting, or spin-glass phase, of the network as a function of the fraction of stored memories and the neuronal gain. We find that short-term synaptic plasticity yields marginal improvements in critical memory capacity. However, through dynamical mean field theory, backed by extensive numerical simulations, we find that short-term synaptic plasticity has a dramatic impact on memory retrieval above the critical capacity. When short-term synaptic plasticity is turned on, the combined neuronal and synaptic dynamics descends a high-dimensional energy landscape over both neurons and synapses. The energy landscape over neurons alone is thus dynamic, and is lowered in the vicinity of recent neuronal patterns visited by the network, just like the surface of a trampoline is lowered in the vicinity of regions recently visited by a heavy ball. This trampoline-like reactivity of the neuronal energy landscape to short-term plasticity in synapses can lead to the recall of stored memories that would otherwise have been forgotten. This occurs because the dynamics without short-term plasticity transiently moves towards a stored memory before departing away from it. Thus short-term plasticity, operating during the transient, lowers the energy in the vicinity of the stored memory, eventually trapping the combined neuronal and synaptic dynamics at a fixed point close to the stored memory. In this manner, short-term plasticity enables the recall of memories that would otherwise be forgotten, by trapping transients that would otherwise escape. We furthermore find an optimal time constant for short-term synaptic plasticity, matched to the transient dynamics, to empower recall of forgotten memories.
△ Less
Submitted 6 February, 2026; v1 submitted 27 November, 2025;
originally announced November 2025.
-
The geometry and dynamics of annealed optimization in the coherent Ising machine with hidden and planted solutions
Authors:
Federico Ghimenti,
Adithya Sriram,
Atsushi Yamamura,
Hideo Mabuchi,
Surya Ganguli
Abstract:
The coherent Ising machine (CIM) is a nonconventional hardware architecture for finding approximate solutions to large-scale combinatorial optimization problems. It operates by annealing a laser gain parameter to adiabatically deform a high-dimensional energy landscape over a set of soft spins, going from a simple convex landscape to the more complex optimization landscape of interest. We address…
▽ More
The coherent Ising machine (CIM) is a nonconventional hardware architecture for finding approximate solutions to large-scale combinatorial optimization problems. It operates by annealing a laser gain parameter to adiabatically deform a high-dimensional energy landscape over a set of soft spins, going from a simple convex landscape to the more complex optimization landscape of interest. We address how the evolving energy landscapes guides the optimization dynamics against problems with hidden planted solutions. We study the Sherrington-Kirkpatrick spin-glass with ferromagnetic couplings that favor a hidden configuration by combining the replica method, random matrix theory, the Kac-Rice method and dynamical mean field theory. We characterize energy, number, location, and Hessian eigenspectra of global minima, local minima, and critical points as the landscape evolves. We find that low energy global minima develop soft-modes which the optimization dynamics can exploit to descend the energy landscape. Even when these global minima are aligned to the hidden configuration, there can be exponentially many higher energy local minima that are all unaligned with the hidden solution. Nevertheless, the annealed optimization dynamics can evade this cloud of unaligned high energy local minima and descend near to aligned lower energy global minima. Eventually, as the landscape is further annealed, these global minima become rigid, terminating any further optimization gains from annealing. We further consider a second optimization problem, the Wishart planted ensemble, which contains a hidden planted solution in a landscape with tunable ruggedness. We describe CIM phase transitions between recoverability and non-recoverability of the hidden solution. Overall, we find intriguing relations between high-dimensional geometry and dynamics in analog machines for combinatorial optimization.
△ Less
Submitted 26 October, 2025; v1 submitted 23 October, 2025;
originally announced October 2025.
-
CMT-Benchmark: A Benchmark for Condensed Matter Theory Built by Expert Researchers
Authors:
Haining Pan,
James V. Roggeveen,
Erez Berg,
Juan Carrasquilla,
Debanjan Chowdhury,
Surya Ganguli,
Federico Ghimenti,
Juraj Hasik,
Henry Hunt,
Hong-Chen Jiang,
Mason Kamb,
Ying-Jer Kao,
Ehsan Khatami,
Michael J. Lawler,
Di Luo,
Titus Neupert,
Xiaoliang Qi,
Michael P. Brenner,
Eun-Ah Kim
Abstract:
Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce. To fill this gap, we present CMT-Benchmark, a dataset of 50 problems covering condensed matter theory (CMT) at the level of an expert researcher. Topics span analytical and computational approaches in quantum many-body,…
▽ More
Large language models (LLMs) have shown remarkable progress in coding and math problem-solving, but evaluation on advanced research-level problems in hard sciences remains scarce. To fill this gap, we present CMT-Benchmark, a dataset of 50 problems covering condensed matter theory (CMT) at the level of an expert researcher. Topics span analytical and computational approaches in quantum many-body, and classical statistical mechanics. The dataset was designed and verified by a panel of expert researchers from around the world. We built the dataset through a collaborative environment that challenges the panel to write and refine problems they would want a research assistant to solve, including Hartree-Fock, exact diagonalization, quantum/variational Monte Carlo, density matrix renormalization group (DMRG), quantum/classical statistical mechanics, and model building. We evaluate LLMs by programmatically checking solutions against expert-supplied ground truth. We developed machine-grading, including symbolic handling of non-commuting operators via normal ordering. They generalize across tasks too. Our evaluations show that frontier models struggle with all of the problems in the dataset, highlighting a gap in the physical reasoning skills of current LLMs. Notably, experts identified strategies for creating increasingly difficult problems by interacting with the LLMs and exploiting common failure modes. The best model, GPT5, solves 30\% of the problems; average across 17 models (GPT, Gemini, Claude, DeepSeek, Llama) is 11.4\pm2.1\%. Moreover, 18 problems are solved by none of the 17 models, and 26 by at most one. These unsolved problems span Quantum Monte Carlo, Variational Monte Carlo, and DMRG. Answers sometimes violate fundamental symmetries or have unphysical scaling dimensions. We believe this benchmark will guide development toward capable AI research assistants and tutors.
△ Less
Submitted 27 February, 2026; v1 submitted 6 October, 2025;
originally announced October 2025.
-
High-capacity associative memory in a quantum-optical spin glass
Authors:
Brendan P. Marsh,
David Atri Schuller,
Yunpeng Ji,
Henry S. Hunt,
Surya Ganguli,
Sarang Gopalakrishnan,
Jonathan Keeling,
Benjamin L. Lev
Abstract:
The Hopfield model describes a neural network that stores memories using all-to-all-coupled spins. Memory patterns are recalled under equilibrium dynamics. Storing too many patterns breaks the associative recall process because frustration causes an exponential number of spurious patterns to arise as the network becomes a spin glass. Despite this, memory recall in a spin glass can be restored, and…
▽ More
The Hopfield model describes a neural network that stores memories using all-to-all-coupled spins. Memory patterns are recalled under equilibrium dynamics. Storing too many patterns breaks the associative recall process because frustration causes an exponential number of spurious patterns to arise as the network becomes a spin glass. Despite this, memory recall in a spin glass can be restored, and even enhanced, under quantum-optical nonequilibrium dynamics because spurious patterns can now serve as reliable memories. We experimentally observe associative memory with high storage capacity in a driven-dissipative spin glass made of atoms and photons. The capacity surpasses the Hopfield limit by up to seven-fold in a sixteen-spin network. Atomic motion boosts capacity by dynamically modifying connectivity akin to short-term synaptic plasticity in neural networks, realizing a precursor to learning in a quantum-optical system.
△ Less
Submitted 15 September, 2025;
originally announced September 2025.
-
Review of contact models used in Discrete Element Method (DEM)
Authors:
S Ganguli,
P S Goswami,
M Bose
Abstract:
This work presents a detailed review of the methods proposed to implement Mindlin's no-slip and partial slip model under constant normal loading and Mindlin Deresiewicz's extensional work on micro-slip under varying normal loading, for the simulation of granular flow. Various methods that followed Mindlin's and Mindlin and Deresiewicz's approaches for modeling the tangential contact between two sp…
▽ More
This work presents a detailed review of the methods proposed to implement Mindlin's no-slip and partial slip model under constant normal loading and Mindlin Deresiewicz's extensional work on micro-slip under varying normal loading, for the simulation of granular flow. Various methods that followed Mindlin's and Mindlin and Deresiewicz's approaches for modeling the tangential contact between two spherical particles are reviewed thoroughly, first, to understand the tangential load-displacement behaviour and finally to compare the tangential coefficient of restitution obtained from each of these models with the experimentally determined values reported in the literature. The tangential load-displacement behaviour obtained from the integral method is qualitatively different from that determined by the incremental method, which accounts for the loading history. Tangential and rotational coefficient of restitution in the gross sliding regime obtained from all the models show excellent agreement with the experimental result, but in the sticking regime, the agreement is only qualitative.
△ Less
Submitted 9 September, 2025;
originally announced September 2025.
-
A universal route to chiral Ising superconductivity in monolayer TaS$_2$ and NbSe$_2$
Authors:
Lucia Gibelli,
Simon Höcherl,
Julian Siegl,
Viliam Vaňo,
Somesh C. Ganguli,
Magdalena Marganska,
Milena Grifoni
Abstract:
We investigate Ising superconductivity in two archetypal intrinsic superconductors, monolayer 1H-TaS$_2$ and 1H-NbSe$_2$, in a bottom-up approach. Using ab initio-based tight-binding parameterizations for the relevant low-energy d-bands, the screened interaction is evaluated microscopically, in a scheme including Bloch overlaps. In direct space, the screened potential displays for both systems lon…
▽ More
We investigate Ising superconductivity in two archetypal intrinsic superconductors, monolayer 1H-TaS$_2$ and 1H-NbSe$_2$, in a bottom-up approach. Using ab initio-based tight-binding parameterizations for the relevant low-energy d-bands, the screened interaction is evaluated microscopically, in a scheme including Bloch overlaps. In direct space, the screened potential displays for both systems long-range Friedel oscillations alternating in sign. Upon scaling, the oscillation pattern becomes universal, with the periodic features locked to the lattice. Solving the momentum-resolved gap equations, a chiral ground state with p-like symmetry is generically found. Due to the larger Ising spin-orbit coupling, the chiral gap is more anisotropic in TaS$_2$ than in NbSe$_2$. This is reflected in tunneling spectra displaying V-shaped features for the former, in quantitative agreement with low-temperature scanning tunneling experiments on TaS$_2$. At the same time, our results reconcile the apparent discordance with hard gap tunneling spectra reported for the sibling NbSe$_2$.
△ Less
Submitted 6 September, 2025;
originally announced September 2025.
-
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
Authors:
Daniel Kunin,
Giovanni Luca Marchetti,
Feng Chen,
Dhruva Karkada,
James B. Simon,
Michael R. DeWeese,
Surya Ganguli,
Nina Miolane
Abstract:
What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dynamics of feature learning in two-layer networks trained from small initialization. Prior works have shown that gradient flow in this regime exhibits a staircase-like loss curve, alternating between plateaus where neuron…
▽ More
What features neural networks learn, and how, remains an open question. In this paper, we introduce Alternating Gradient Flows (AGF), an algorithmic framework that describes the dynamics of feature learning in two-layer networks trained from small initialization. Prior works have shown that gradient flow in this regime exhibits a staircase-like loss curve, alternating between plateaus where neurons slowly align to useful directions and sharp drops where neurons rapidly grow in norm. AGF approximates this behavior as an alternating two-step process: maximizing a utility function over dormant neurons and minimizing a cost function over active ones. AGF begins with all neurons dormant. At each iteration, a dormant neuron activates, triggering the acquisition of a feature and a drop in the loss. AGF quantifies the order, timing, and magnitude of these drops, matching experiments across several commonly studied architectures. We show that AGF unifies and extends existing saddle-to-saddle analyses in fully connected linear networks and attention-only linear transformers, where the learned features are singular modes and principal components, respectively. In diagonal linear networks, we prove AGF converges to gradient flow in the limit of vanishing initialization. Applying AGF to quadratic networks trained to perform modular addition, we give the first complete characterization of the training dynamics, revealing that networks learn Fourier features in decreasing order of coefficient magnitude. Altogether, AGF offers a promising step towards understanding feature learning in neural networks.
△ Less
Submitted 24 December, 2025; v1 submitted 6 June, 2025;
originally announced June 2025.
-
Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical Reasoning
Authors:
Feng Chen,
Allan Raventos,
Nan Cheng,
Surya Ganguli,
Shaul Druckmann
Abstract:
Recent progress in large language models (LLMs) highlights the power of scaling test-time compute to achieve strong performance on complex tasks, such as mathematical reasoning and code generation. This raises a critical question: how should model training be modified to optimize performance under a subsequent test-time compute strategy and budget? To explore this, we focus on pass@N, a simple tes…
▽ More
Recent progress in large language models (LLMs) highlights the power of scaling test-time compute to achieve strong performance on complex tasks, such as mathematical reasoning and code generation. This raises a critical question: how should model training be modified to optimize performance under a subsequent test-time compute strategy and budget? To explore this, we focus on pass@N, a simple test-time strategy that searches for a correct answer in $N$ independent samples. We show, surprisingly, that training with cross-entropy (CE) loss can be ${\it misaligned}$ with pass@N in that pass@N accuracy ${\it decreases}$ with longer training. We explain the origins of this misalignment in terms of model overconfidence induced by CE, and experimentally verify our prediction of overconfidence as an impediment to scaling test-time compute via pass@N. Furthermore we suggest a principled, modified training loss that is better aligned to pass@N by limiting model confidence and rescuing pass@N test performance. Our algorithm demonstrates improved mathematical reasoning on MATH and MiniF2F benchmarks under several scenarios: (1) providing answers to math questions; and (2) proving theorems by searching over proof trees of varying shapes. Overall our work underscores the importance of co-designing two traditionally separate phases of LLM development: training-time protocols and test-time search and reasoning strategies.
△ Less
Submitted 23 November, 2025; v1 submitted 10 February, 2025;
originally announced February 2025.
-
An analytic theory of creativity in convolutional diffusion models
Authors:
Mason Kamb,
Surya Ganguli
Abstract:
We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original images that lie far from their training data. However, optimal score-matching theory suggests that these models should only be able to produce memorized training examples. To reconcile this theory-experiment gap, we identif…
▽ More
We obtain an analytic, interpretable and predictive theory of creativity in convolutional diffusion models. Indeed, score-matching diffusion models can generate highly original images that lie far from their training data. However, optimal score-matching theory suggests that these models should only be able to produce memorized training examples. To reconcile this theory-experiment gap, we identify two simple inductive biases, locality and equivariance, that: (1) induce a form of combinatorial creativity by preventing optimal score-matching; (2) result in fully analytic, completely mechanistically interpretable, local score (LS) and equivariant local score (ELS) machines that, (3) after calibrating a single time-dependent hyperparameter can quantitatively predict the outputs of trained convolution only diffusion models (like ResNets and UNets) with high accuracy (median $r^2$ of $0.95, 0.94, 0.94, 0.96$ for our top model on CIFAR10, FashionMNIST, MNIST, and CelebA). Our model reveals a locally consistent patch mosaic mechanism of creativity, in which diffusion models create exponentially many novel images by mixing and matching different local training set patches at different scales and image locations. Our theory also partially predicts the outputs of pre-trained self-attention enabled UNets (median $r^2 \sim 0.77$ on CIFAR10), revealing an intriguing role for attention in carving out semantic coherence from local patch mosaics.
△ Less
Submitted 5 June, 2025; v1 submitted 28 December, 2024;
originally announced December 2024.
-
Role of the ratio of tangential to normal stiffness coefficient on the behaviour of vibrofluidised particles
Authors:
Alok Tiwari,
Sourav Ganguli,
Manaswita Bose,
V Kumaran
Abstract:
The selection of parameters in the contact law for inter-particle interactions affects the results of simulations of flowing granular materials. The present study aims to understand the effect of the ratio of tangential to normal spring stiffness coefficient ($κ$) on inter-particle contact behaviour in terms of the rotational coefficient of restitution determined using data obtained from multi-par…
▽ More
The selection of parameters in the contact law for inter-particle interactions affects the results of simulations of flowing granular materials. The present study aims to understand the effect of the ratio of tangential to normal spring stiffness coefficient ($κ$) on inter-particle contact behaviour in terms of the rotational coefficient of restitution determined using data obtained from multi-particle simulations. The effect of $κ$ on the profiles of the micro- and macroscopic properties of particles in a vibrofluidised bed is also investigated. The Discrete Element Method (DEM) is used to simulate a vertically vibrated fluidised bed using the open-source software LAMMPS. The inter-particle and wall-particle contact forces are determined using the linear spring-dashpot (LSD) model. The distribution of the mean co-ordination number, force during the contact, contact regimes, and rotational coefficient of restitution are determined from the data obtained from simulations. It was shown that $κ$ plays a significant role in the distribution of inter-particle contacts between different regimes and, thereby, the velocity distribution and profiles of statistically averaged properties of the vibrofluidised particles. Our results show that for particles with surface friction coefficient $μ>0.1$, the commonly used value $κ=\frac{2}{7}$ results in quantitatively different results from those obtained using $0.67 \le κ< 1$, a range consistent with the realistic values of Poisson ratios for simple materials.
△ Less
Submitted 20 December, 2024;
originally announced December 2024.
-
Fooling LLM graders into giving better grades through neural activity guided adversarial prompting
Authors:
Atsushi Yamamura,
Surya Ganguli
Abstract:
The deployment of artificial intelligence (AI) in critical decision-making and evaluation processes raises concerns about inherent biases that malicious actors could exploit to distort decision outcomes. We propose a systematic method to reveal such biases in AI evaluation systems and apply it to automated essay grading as an example. Our approach first identifies hidden neural activity patterns t…
▽ More
The deployment of artificial intelligence (AI) in critical decision-making and evaluation processes raises concerns about inherent biases that malicious actors could exploit to distort decision outcomes. We propose a systematic method to reveal such biases in AI evaluation systems and apply it to automated essay grading as an example. Our approach first identifies hidden neural activity patterns that predict distorted decision outcomes and then optimizes an adversarial input suffix to amplify such patterns. We demonstrate that this combination can effectively fool large language model (LLM) graders into assigning much higher grades than humans would. We further show that this white-box attack transfers to black-box attacks on other models, including commercial closed-source models like Gemini. They further reveal the existence of a "magic word" that plays a pivotal role in the efficacy of the attack. We trace the origin of this magic word bias to the structure of commonly-used chat templates for supervised fine-tuning of LLMs and show that a minor change in the template can drastically reduce the bias. This work not only uncovers vulnerabilities in current LLMs but also proposes a systematic method to identify and remove hidden biases, contributing to the goal of ensuring AI safety and security.
△ Less
Submitted 17 December, 2024;
originally announced December 2024.
-
Features are fate: a theory of transfer learning in high-dimensional regression
Authors:
Javan Tahir,
Surya Ganguli,
Grant M. Rotskoff
Abstract:
With the emergence of large-scale pre-trained neural networks, methods to adapt such "foundation" models to data-limited downstream tasks have become a necessity. Fine-tuning, preference optimization, and transfer learning have all been successfully employed for these purposes when the target task closely resembles the source task, but a precise theoretical understanding of "task similarity" is st…
▽ More
With the emergence of large-scale pre-trained neural networks, methods to adapt such "foundation" models to data-limited downstream tasks have become a necessity. Fine-tuning, preference optimization, and transfer learning have all been successfully employed for these purposes when the target task closely resembles the source task, but a precise theoretical understanding of "task similarity" is still lacking. While conventional wisdom suggests that simple measures of similarity between source and target distributions, such as $φ$-divergences or integral probability metrics, can directly predict the success of transfer, we prove the surprising fact that, in general, this is not the case. We adopt, instead, a feature-centric viewpoint on transfer learning and establish a number of theoretical results that demonstrate that when the target task is well represented by the feature space of the pre-trained model, transfer learning outperforms training from scratch. We study deep linear networks as a minimal model of transfer learning in which we can analytically characterize the transferability phase diagram as a function of the target dataset size and the feature space overlap. For this model, we establish rigorously that when the feature space overlap between the source and target tasks is sufficiently strong, both linear transfer and fine-tuning improve performance, especially in the low data limit. These results build on an emerging understanding of feature learning dynamics in deep linear networks, and we demonstrate numerically that the rigorous results we derive for the linear case also apply to nonlinear networks.
△ Less
Submitted 7 July, 2025; v1 submitted 10 October, 2024;
originally announced October 2024.
-
Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learning
Authors:
Daniel Kunin,
Allan Raventós,
Clémentine Dominé,
Feng Chen,
David Klindt,
Andrew Saxe,
Surya Ganguli
Abstract:
While the impressive performance of modern neural networks is often attributed to their capacity to efficiently extract task-relevant features from data, the mechanisms underlying this rich feature learning regime remain elusive, with much of our theoretical understanding stemming from the opposing lazy regime. In this work, we derive exact solutions to a minimal model that transitions between laz…
▽ More
While the impressive performance of modern neural networks is often attributed to their capacity to efficiently extract task-relevant features from data, the mechanisms underlying this rich feature learning regime remain elusive, with much of our theoretical understanding stemming from the opposing lazy regime. In this work, we derive exact solutions to a minimal model that transitions between lazy and rich learning, precisely elucidating how unbalanced layer-specific initialization variances and learning rates determine the degree of feature learning. Our analysis reveals that they conspire to influence the learning regime through a set of conserved quantities that constrain and modify the geometry of learning trajectories in parameter and function space. We extend our analysis to more complex linear models with multiple neurons, outputs, and layers and to shallow nonlinear networks with piecewise linear activation functions. In linear networks, rapid feature learning only occurs from balanced initializations, where all layers learn at similar speeds. While in nonlinear networks, unbalanced initializations that promote faster learning in earlier layers can accelerate rich learning. Through a series of experiments, we provide evidence that this unbalanced rich regime drives feature learning in deep finite-width networks, promotes interpretability of early layers in CNNs, reduces the sample complexity of learning hierarchical data, and decreases the time to grokking in modular arithmetic. Our theory motivates further exploration of unbalanced initializations to enhance efficient feature learning.
△ Less
Submitted 12 October, 2024; v1 submitted 10 June, 2024;
originally announced June 2024.
-
Geometric Dynamics of Signal Propagation Predict Trainability of Transformers
Authors:
Aditya Cowsik,
Tamra Nebabu,
Xiao-Liang Qi,
Surya Ganguli
Abstract:
We investigate forward signal propagation and gradient back propagation in deep, randomly initialized transformers, yielding simple necessary and sufficient conditions on initialization hyperparameters that ensure trainability of deep transformers. Our approach treats the evolution of the representations of $n$ tokens as they propagate through the transformer layers in terms of a discrete time dyn…
▽ More
We investigate forward signal propagation and gradient back propagation in deep, randomly initialized transformers, yielding simple necessary and sufficient conditions on initialization hyperparameters that ensure trainability of deep transformers. Our approach treats the evolution of the representations of $n$ tokens as they propagate through the transformer layers in terms of a discrete time dynamical system of $n$ interacting particles. We derive simple update equations for the evolving geometry of this particle system, starting from a permutation symmetric simplex. Our update equations show that without MLP layers, this system will collapse to a line, consistent with prior work on rank collapse in transformers. However, unlike prior work, our evolution equations can quantitatively track particle geometry in the additional presence of nonlinear MLP layers, and it reveals an order-chaos phase transition as a function of initialization hyperparameters, like the strength of attentional and MLP residual connections and weight variances. In the ordered phase the particles are attractive and collapse to a line, while in the chaotic phase the particles are repulsive and converge to a regular $n$-simplex. We analytically derive two Lyapunov exponents: an angle exponent that governs departures from the edge of chaos in this particle system, and a gradient exponent that governs the rate of exponential growth or decay of backpropagated gradients. We show through experiments that, remarkably, the final test loss at the end of training is well predicted just by these two exponents at the beginning of training, and that the simultaneous vanishing of these two exponents yields a simple necessary and sufficient condition to achieve minimal test loss.
△ Less
Submitted 4 March, 2024;
originally announced March 2024.
-
Principal Component Regression to Study the Impact of Economic Factors on Disadvantaged Communities
Authors:
Narmadha M. Mohankumar,
Milan Jain,
Heng Wan,
Sumitrra Ganguli,
Kyle D. Wilson,
David M. Anderson
Abstract:
The Council on Environmental Quality's Climate and Economic Justice Screening Tool defines "disadvantaged communities" (DAC) in the USA, highlighting census tracts where benefits of climate and energy investments are not accruing. We use a principal component generalized linear model, which addresses the intertwined nature of economic factors, income and employment and model their relationship to…
▽ More
The Council on Environmental Quality's Climate and Economic Justice Screening Tool defines "disadvantaged communities" (DAC) in the USA, highlighting census tracts where benefits of climate and energy investments are not accruing. We use a principal component generalized linear model, which addresses the intertwined nature of economic factors, income and employment and model their relationship to DAC status. Our study 1) identifies the most significant income groups and employment industries that impact DAC status, 2) provides the probability of DAC status across census tracts and compares the predictive accuracy with widely used machine learning approaches, 3) obtains historical predictions of the probability of DAC status, 4) obtains spatial downscaling of DAC status across block groups. Our study provides valuable insights for policymakers and stakeholders to develop strategies that promote sustainable development and address inequities in climate and energy investments in the USA.
△ Less
Submitted 24 January, 2024;
originally announced January 2024.
-
Temporal Embeddings: Scalable Self-Supervised Temporal Representation Learning from Spatiotemporal Data for Multimodal Computer Vision
Authors:
Yi Cao,
Swetava Ganguli,
Vipul Pandey
Abstract:
There exists a correlation between geospatial activity temporal patterns and type of land use. A novel self-supervised approach is proposed to stratify landscape based on mobility activity time series. First, the time series signal is transformed to the frequency domain and then compressed into task-agnostic temporal embeddings by a contractive autoencoder, which preserves cyclic temporal patterns…
▽ More
There exists a correlation between geospatial activity temporal patterns and type of land use. A novel self-supervised approach is proposed to stratify landscape based on mobility activity time series. First, the time series signal is transformed to the frequency domain and then compressed into task-agnostic temporal embeddings by a contractive autoencoder, which preserves cyclic temporal patterns observed in time series. The pixel-wise embeddings are converted to image-like channels that can be used for task-based, multimodal modeling of downstream geospatial tasks using deep semantic segmentation. Experiments show that temporal embeddings are semantically meaningful representations of time series data and are effective across different tasks such as classifying residential area and commercial areas. Temporal embeddings transform sequential, spatiotemporal motion trajectory data into semantically meaningful image-like tensor representations that can be combined (multimodal fusion) with other data modalities that are or can be transformed into image-like tensor representations (for e.g., RBG imagery, graph embeddings of road networks, passively collected imagery like SAR, etc.) to facilitate multimodal learning in geospatial computer vision. Multimodal computer vision is critical for training machine learning models for geospatial feature detection to keep a geospatial mapping service up-to-date in real-time and can significantly improve user experience and above all, user safety.
△ Less
Submitted 15 October, 2023;
originally announced January 2024.
-
Doped Mott phase and charge correlations in monolayer 1T-NbSe$_2$
Authors:
Xin Huang,
Jose L. Lado,
Jani Sainio,
Peter Liljeroth,
Somesh Chandra Ganguli
Abstract:
The doped Hubbard model is one of the paradigmatic platforms to engineer exotic quantum many-body states, including charge-ordered states, strange metals and unconventional superconductors. While undoped and doped correlated phases have been experimentally realized in a variety twisted van der Waals materials, experiments in monolayer materials, and in particular 1T transition metal dichalcogenide…
▽ More
The doped Hubbard model is one of the paradigmatic platforms to engineer exotic quantum many-body states, including charge-ordered states, strange metals and unconventional superconductors. While undoped and doped correlated phases have been experimentally realized in a variety twisted van der Waals materials, experiments in monolayer materials, and in particular 1T transition metal dichalcogenides, have solely reached the conventional insulating undoped regime. Correlated phases in monolayer two-dimensional materials have much higher associated energy scales than their twisted counterparts, making doped correlated monolayers an attractive platform for high temperature correlated quantum matter. Here, we demonstrate the realization of a doped Mott phase in a van der Waals dichalcogenide 1T-NbSe$_2$ monolayer. The system is electron doped due to electron transfer to a monolayer van der Waals substrate via proximity, leading to a correlated triangular lattice with both half-filled and fully-filled sites. We analyze the distribution of the half-filled and filled sites and show the arrangement is unlikely to be controlled by disorder alone, and we show that the presence of competing non-local many-body correlations would account for the charge correlations found experimentally. Our results establish 1T-NbSe$_2$ as a potential monolayer platform to explore correlated doped Mott physics in a frustrated lattice.
△ Less
Submitted 5 March, 2024; v1 submitted 16 January, 2024;
originally announced January 2024.
-
Robust Collaborative Inference with Vertically Split Data Over Dynamic Device Environments
Authors:
Surojit Ganguli,
Zeyu Zhou,
Christopher G. Brinton,
David I. Inouye
Abstract:
When each edge device of a network only perceives a local part of the environment, collaborative inference across multiple devices is often needed to predict global properties of the environment. In safety-critical applications, collaborative inference must be robust to significant network failures caused by environmental disruptions or extreme weather. Existing collaborative learning approaches,…
▽ More
When each edge device of a network only perceives a local part of the environment, collaborative inference across multiple devices is often needed to predict global properties of the environment. In safety-critical applications, collaborative inference must be robust to significant network failures caused by environmental disruptions or extreme weather. Existing collaborative learning approaches, such as privacy-focused Vertical Federated Learning (VFL), typically assume a centralized setup or that one device never fails. However, these assumptions make prior approaches susceptible to significant network failures. To address this problem, we first formalize the problem of robust collaborative inference over a dynamic network of devices that could experience significant network faults. Then, we develop a minimalistic yet impactful method called Multiple Aggregation with Gossip Rounds and Simulated Faults (MAGS) that synthesizes simulated faults via dropout, replication, and gossiping to significantly improve robustness over baselines. We also theoretically analyze our proposed approach to explain why each component enhances robustness. Extensive empirical results validate that MAGS is robust across a range of fault rates-including extreme fault rates.
△ Less
Submitted 25 April, 2025; v1 submitted 27 December, 2023;
originally announced December 2023.
-
Demonstrating Kondo behavior by temperature-dependent scanning tunneling spectroscopy
Authors:
Elia Turco,
Markus Aapro,
Somesh C. Ganguli,
Nils Krane,
Robert Drost,
Nahual Sobrino,
Annika Bernhardt,
Michal Juríček,
Roman Fasel,
Pascal Ruffieux,
Peter Liljeroth,
David Jacob
Abstract:
The Kondo effect describes the scattering of conduction electrons by magnetic impurities, manifesting as an electronic resonance at the Fermi energy with a distinctive temperature evolution. In this letter, we present a critical evaluation of the current methodology employed to demonstrate Kondo behavior in transport measurements, underscoring the limitations of established theoretical frameworks…
▽ More
The Kondo effect describes the scattering of conduction electrons by magnetic impurities, manifesting as an electronic resonance at the Fermi energy with a distinctive temperature evolution. In this letter, we present a critical evaluation of the current methodology employed to demonstrate Kondo behavior in transport measurements, underscoring the limitations of established theoretical frameworks and the influence of extrinsic broadening. We introduce a novel approach for analyzing spectroscopic indicators of the Kondo effect, employing the Hurwitz-Fano lineshape as a model for the Kondo resonance in the presence of extrinsic broadening. Through precise scanning tunneling spectroscopy measurements on an exemplary spin-1/2 Kondo system, phenalenyl on Au(111), we demonstrate the efficacy of our proposed protocol in extracting accurate intrinsic Kondo linewidths from finite-temperature measurements. The extracted linewidths exhibit a robust fit with a recently derived expression for the temperature-dependent intrinsic Kondo linewidth, providing compelling evidence for the validity of the underlying theory.
△ Less
Submitted 19 April, 2024; v1 submitted 13 October, 2023;
originally announced October 2023.
-
SeMAnD: Self-Supervised Anomaly Detection in Multimodal Geospatial Datasets
Authors:
Daria Reshetova,
Swetava Ganguli,
C. V. Krishnakumar Iyer,
Vipul Pandey
Abstract:
We propose a Self-supervised Anomaly Detection technique, called SeMAnD, to detect geometric anomalies in Multimodal geospatial datasets. Geospatial data comprises of acquired and derived heterogeneous data modalities that we transform to semantically meaningful, image-like tensors to address the challenges of representation, alignment, and fusion of multimodal data. SeMAnD is comprised of (i) a s…
▽ More
We propose a Self-supervised Anomaly Detection technique, called SeMAnD, to detect geometric anomalies in Multimodal geospatial datasets. Geospatial data comprises of acquired and derived heterogeneous data modalities that we transform to semantically meaningful, image-like tensors to address the challenges of representation, alignment, and fusion of multimodal data. SeMAnD is comprised of (i) a simple data augmentation strategy, called RandPolyAugment, capable of generating diverse augmentations of vector geometries, and (ii) a self-supervised training objective with three components that incentivize learning representations of multimodal data that are discriminative to local changes in one modality which are not corroborated by the other modalities. Detecting local defects is crucial for geospatial anomaly detection where even small anomalies (e.g., shifted, incorrectly connected, malformed, or missing polygonal vector geometries like roads, buildings, landcover, etc.) are detrimental to the experience and safety of users of geospatial applications like mapping, routing, search, and recommendation systems. Our empirical study on test sets of different types of real-world geometric geospatial anomalies across 3 diverse geographical regions demonstrates that SeMAnD is able to detect real-world defects and outperforms domain-agnostic anomaly detection strategies by 4.8-19.7% as measured using anomaly classification AUC. We also show that model performance increases (i) up to 20.4% as the number of input modalities increase and (ii) up to 22.9% as the diversity and strength of training data augmentations increase.
△ Less
Submitted 26 September, 2023;
originally announced September 2023.
-
Geometric landscape annealing as an optimization principle underlying the coherent Ising machine
Authors:
Atsushi Yamamura,
Hideo Mabuchi,
Surya Ganguli
Abstract:
Given the fundamental importance of combinatorial optimization across many diverse application domains, there has been widespread interest in the development of unconventional physical computing architectures that can deliver better solutions with lower resource costs. These architectures embed discrete optimization problems into the annealed, analog evolution of nonlinear dynamical systems. Howev…
▽ More
Given the fundamental importance of combinatorial optimization across many diverse application domains, there has been widespread interest in the development of unconventional physical computing architectures that can deliver better solutions with lower resource costs. These architectures embed discrete optimization problems into the annealed, analog evolution of nonlinear dynamical systems. However, a theoretical understanding of their performance remains elusive, unlike the cases of simulated or quantum annealing. We develop such understanding for the coherent Ising machine (CIM), a network of optical parametric oscillators that can be applied to any quadratic unconstrained binary optimization problem. Here we focus on how the CIM finds low-energy solutions of the Sherrington-Kirkpatrick spin glass. As the laser gain is annealed, the CIM interpolates between gradient descent on the soft-spin energy landscape, to optimization on coupled binary spins. By exploiting spin-glass theory, we develop a detailed understanding of the evolving geometry of the high-dimensional CIM energy landscape as the laser gain increases, finding several phase transitions, from flat, to rough, to rigid. Additionally, we develop a cavity method that provides a precise geometric interpretation of supersymmetry breaking in terms of the response of a rough landscape to specific perturbations. We confirm our theory with numerical experiments, and find detailed information about critical points of the landscape. Our extensive analysis of phase transitions provides theoretically motivated optimal annealing schedules that can reliably find near-ground states. This analysis reveals geometric landscape annealing as a powerful optimization principle and suggests many further avenues for exploring other optimization problems, as well as other types of annealed dynamics, including chaotic, oscillatory or quantum dynamics.
△ Less
Submitted 14 September, 2023;
originally announced September 2023.
-
Entanglement and replica symmetry breaking in a driven-dissipative quantum spin glass
Authors:
Brendan P. Marsh,
Ronen M. Kroeze,
Surya Ganguli,
Sarang Gopalakrishnan,
Jonathan Keeling,
Benjamin L. Lev
Abstract:
We describe simulations of the quantum dynamics of a confocal cavity QED system that realizes an intrinsically driven-dissipative spin glass. A close connection between open quantum dynamics and replica symmetry breaking is established, in which individual quantum trajectories are the replicas. We observe that entanglement plays an important role in the emergence of replica symmetry breaking in a…
▽ More
We describe simulations of the quantum dynamics of a confocal cavity QED system that realizes an intrinsically driven-dissipative spin glass. A close connection between open quantum dynamics and replica symmetry breaking is established, in which individual quantum trajectories are the replicas. We observe that entanglement plays an important role in the emergence of replica symmetry breaking in a fully connected, frustrated spin network of up to fifteen spin-1/2 particles. Quantum trajectories of entangled spins reach steady-state spin configurations of lower energy than that of semiclassical trajectories. Cavity emission allows monitoring of the continuous stochastic evolution of spin configurations, while backaction from this projects entangled states into states of broken Ising and replica symmetry. The emergence of spin glass order manifests itself through the simultaneous absence of magnetization and the presence of nontrivial spin overlap density distributions among replicas. Moreover, these overlaps reveal incipient ultrametric order, in line with the Parisi RSB solution ansatz for the Sherrington-Kirkpatrick model. A nonthermal Parisi order parameter distribution, however, highlights the driven-dissipative nature of this quantum optical spin glass. This practicable system could serve as a testbed for exploring how quantum effects enrich the physics of spin glasses.
△ Less
Submitted 19 November, 2023; v1 submitted 19 July, 2023;
originally announced July 2023.
-
Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression
Authors:
Allan Raventós,
Mansheej Paul,
Feng Chen,
Surya Ganguli
Abstract:
Pretrained transformers exhibit the remarkable ability of in-context learning (ICL): they can learn tasks from just a few examples provided in the prompt without updating any weights. This raises a foundational question: can ICL solve fundamentally $\textit{new}$ tasks that are very different from those seen during pretraining? To probe this question, we examine ICL's performance on linear regress…
▽ More
Pretrained transformers exhibit the remarkable ability of in-context learning (ICL): they can learn tasks from just a few examples provided in the prompt without updating any weights. This raises a foundational question: can ICL solve fundamentally $\textit{new}$ tasks that are very different from those seen during pretraining? To probe this question, we examine ICL's performance on linear regression while varying the diversity of tasks in the pretraining dataset. We empirically demonstrate a $\textit{task diversity threshold}$ for the emergence of ICL. Below this threshold, the pretrained transformer cannot solve unseen regression tasks, instead behaving like a Bayesian estimator with the $\textit{non-diverse pretraining task distribution}$ as the prior. Beyond this threshold, the transformer significantly outperforms this estimator; its behavior aligns with that of ridge regression, corresponding to a Gaussian prior over $\textit{all tasks}$, including those not seen during pretraining. Thus, when pretrained on data with task diversity greater than the threshold, transformers $\textit{can}$ optimally solve fundamentally new tasks in-context. Importantly, this capability hinges on it deviating from the Bayes optimal estimator with the pretraining distribution as the prior. This study also explores the effect of regularization, model capacity and task structure and underscores, in a concrete example, the critical role of task diversity, alongside data and model scale, in the emergence of ICL. Code is available at https://github.com/mansheej/icl-task-diversity.
△ Less
Submitted 8 November, 2023; v1 submitted 26 June, 2023;
originally announced June 2023.
-
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks
Authors:
Feng Chen,
Daniel Kunin,
Atsushi Yamamura,
Surya Ganguli
Abstract:
In this work, we reveal a strong implicit bias of stochastic gradient descent (SGD) that drives overly expressive networks to much simpler subnetworks, thereby dramatically reducing the number of independent parameters, and improving generalization. To reveal this bias, we identify invariant sets, or subsets of parameter space that remain unmodified by SGD. We focus on two classes of invariant set…
▽ More
In this work, we reveal a strong implicit bias of stochastic gradient descent (SGD) that drives overly expressive networks to much simpler subnetworks, thereby dramatically reducing the number of independent parameters, and improving generalization. To reveal this bias, we identify invariant sets, or subsets of parameter space that remain unmodified by SGD. We focus on two classes of invariant sets that correspond to simpler (sparse or low-rank) subnetworks and commonly appear in modern architectures. Our analysis uncovers that SGD exhibits a property of stochastic attractivity towards these simpler invariant sets. We establish a sufficient condition for stochastic attractivity based on a competition between the loss landscape's curvature around the invariant set and the noise introduced by stochastic gradients. Remarkably, we find that an increased level of noise strengthens attractivity, leading to the emergence of attractive invariant sets associated with saddle-points or local maxima of the train loss. We observe empirically the existence of attractive invariant sets in trained deep neural networks, implying that SGD dynamics often collapses to simple subnetworks with either vanishing or redundant neurons. We further demonstrate how this simplifying process of stochastic collapse benefits generalization in a linear teacher-student framework. Finally, through this analysis, we mechanistically explain why early training with large learning rates for extended periods benefits subsequent generalization.
△ Less
Submitted 28 May, 2024; v1 submitted 7 June, 2023;
originally announced June 2023.
-
Singular Vectors of Sums of Rectangular Random Matrices and Optimal Estimators of High-Rank Signals: The Extensive Spike Model
Authors:
Itamar D. Landau,
Gabriel C. Mel,
Surya Ganguli
Abstract:
Across many disciplines from neuroscience and genomics to machine learning, atmospheric science and finance, the problems of denoising large data matrices to recover signals obscured by noise, and of estimating the structure of these signals, are of fundamental importance. A key to solving these problems lies in understanding how the singular value structure of a signal is deformed by noise. This…
▽ More
Across many disciplines from neuroscience and genomics to machine learning, atmospheric science and finance, the problems of denoising large data matrices to recover signals obscured by noise, and of estimating the structure of these signals, are of fundamental importance. A key to solving these problems lies in understanding how the singular value structure of a signal is deformed by noise. This question has been thoroughly studied in the well-known spiked matrix model, in which data matrices originate from low-rank signals perturbed by additive noise, in an asymptotic limit where the size of these matrices tends to infinity but the signal rank remains finite. We first show, strikingly, that the singular value structure of large finite matrices (of size $\sim 1000$) with even moderate-rank signals, as low as $10$, is not accurately predicted by the finite-rank theory, thereby limiting the application of this theory to real data. To address these deficiencies, we analytically compute how the singular values and vectors of an arbitrary high-rank signal matrix are deformed by additive noise. We next study an asymptotic limit corresponding to an $\textit{extensive}$ spike model, in which the rank of the hidden signal is proportional to the size of the data matrix, while both tend to infinity. We map out the phase diagram of the singular value structure of the extensive spike model as a joint function of signal strength and rank. We further exploit these analytics to derive optimal rotationally invariant denoisers to recover hidden $\textit{high}$-rank signals from data, as well as optimal invariant estimators of the signal covariance structure. Overall, our results provide fundamental theory governing how high-dimensional signals are deformed by additive noise, together with practical formulas for optimal denoising and covariance estimation.
△ Less
Submitted 4 December, 2023; v1 submitted 1 June, 2023;
originally announced June 2023.
-
Self-Supervised Temporal Analysis of Spatiotemporal Data
Authors:
Yi Cao,
Swetava Ganguli,
Vipul Pandey
Abstract:
There exists a correlation between geospatial activity temporal patterns and type of land use. A novel self-supervised approach is proposed to stratify landscape based on mobility activity time series. First, the time series signal is transformed to the frequency domain and then compressed into task-agnostic temporal embeddings by a contractive autoencoder, which preserves cyclic temporal patterns…
▽ More
There exists a correlation between geospatial activity temporal patterns and type of land use. A novel self-supervised approach is proposed to stratify landscape based on mobility activity time series. First, the time series signal is transformed to the frequency domain and then compressed into task-agnostic temporal embeddings by a contractive autoencoder, which preserves cyclic temporal patterns observed in time series. The pixel-wise embeddings are converted to image-like channels that can be used for task-based, multimodal modeling of downstream geospatial tasks using deep semantic segmentation. Experiments show that temporal embeddings are semantically meaningful representations of time series data and are effective across different tasks such as classifying residential area and commercial areas.
△ Less
Submitted 25 April, 2023;
originally announced April 2023.
-
Neural networks: from the perceptron to deep nets
Authors:
Marylou Gabrié,
Surya Ganguli,
Carlo Lucibello,
Riccardo Zecchina
Abstract:
Artificial networks have been studied through the prism of statistical mechanics as disordered systems since the 80s, starting from the simple models of Hopfield's associative memory and the single-neuron perceptron classifier. Assuming data is generated by a teacher model, asymptotic generalisation predictions were originally derived using the replica method and the online learning dynamics has b…
▽ More
Artificial networks have been studied through the prism of statistical mechanics as disordered systems since the 80s, starting from the simple models of Hopfield's associative memory and the single-neuron perceptron classifier. Assuming data is generated by a teacher model, asymptotic generalisation predictions were originally derived using the replica method and the online learning dynamics has been described in the large system limit. In this chapter, we review the key original ideas of this literature along with their heritage in the ongoing quest to understand the efficiency of modern deep learning algorithms. One goal of current and future research is to characterize the bias of the learning algorithms toward well-generalising minima in a complex overparametrized loss landscapes with many solutions perfectly interpolating the training data. Works on perceptrons, two-layer committee machines and kernel-like learning machines shed light on these benefits of overparametrization. Another goal is to understand the advantage of depth while models now commonly feature tens or hundreds of layers. If replica computations apparently fall short in describing general deep neural networks learning, studies of simplified linear or untrained models, as well as the derivation of scaling laws provide the first elements of answers.
△ Less
Submitted 13 April, 2023;
originally announced April 2023.
-
SemDeDup: Data-efficient learning at web-scale through semantic deduplication
Authors:
Amro Abbas,
Kushal Tirumala,
Dániel Simig,
Surya Ganguli,
Ari S. Morcos
Abstract:
Progress in machine learning has been driven in large part by massive increases in data. However, large web-scale datasets such as LAION are largely uncurated beyond searches for exact duplicates, potentially leaving much redundancy. Here, we introduce SemDeDup, a method which leverages embeddings from pre-trained models to identify and remove semantic duplicates: data pairs which are semantically…
▽ More
Progress in machine learning has been driven in large part by massive increases in data. However, large web-scale datasets such as LAION are largely uncurated beyond searches for exact duplicates, potentially leaving much redundancy. Here, we introduce SemDeDup, a method which leverages embeddings from pre-trained models to identify and remove semantic duplicates: data pairs which are semantically similar, but not exactly identical. Removing semantic duplicates preserves performance and speeds up learning. Analyzing a subset of LAION, we show that SemDeDup can remove 50% of the data with minimal performance loss, effectively halving training time. Moreover, performance increases out of distribution. Also, analyzing language models trained on C4, a partially curated dataset, we show that SemDeDup improves over prior approaches while providing efficiency gains. SemDeDup provides an example of how simple ways of leveraging quality embeddings can be used to make models learn faster with less data.
△ Less
Submitted 22 March, 2023; v1 submitted 16 March, 2023;
originally announced March 2023.
-
Holistic Evaluation of Language Models
Authors:
Percy Liang,
Rishi Bommasani,
Tony Lee,
Dimitris Tsipras,
Dilara Soylu,
Michihiro Yasunaga,
Yian Zhang,
Deepak Narayanan,
Yuhuai Wu,
Ananya Kumar,
Benjamin Newman,
Binhang Yuan,
Bobby Yan,
Ce Zhang,
Christian Cosgrove,
Christopher D. Manning,
Christopher Ré,
Diana Acosta-Navas,
Drew A. Hudson,
Eric Zelikman,
Esin Durmus,
Faisal Ladhak,
Frieda Rong,
Hongyu Ren,
Huaxiu Yao
, et al. (25 additional authors not shown)
Abstract:
Language models (LMs) are becoming the foundation for almost all major language technologies, but their capabilities, limitations, and risks are not well understood. We present Holistic Evaluation of Language Models (HELM) to improve the transparency of language models. First, we taxonomize the vast space of potential scenarios (i.e. use cases) and metrics (i.e. desiderata) that are of interest fo…
▽ More
Language models (LMs) are becoming the foundation for almost all major language technologies, but their capabilities, limitations, and risks are not well understood. We present Holistic Evaluation of Language Models (HELM) to improve the transparency of language models. First, we taxonomize the vast space of potential scenarios (i.e. use cases) and metrics (i.e. desiderata) that are of interest for LMs. Then we select a broad subset based on coverage and feasibility, noting what's missing or underrepresented (e.g. question answering for neglected English dialects, metrics for trustworthiness). Second, we adopt a multi-metric approach: We measure 7 metrics (accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency) for each of 16 core scenarios when possible (87.5% of the time). This ensures metrics beyond accuracy don't fall to the wayside, and that trade-offs are clearly exposed. We also perform 7 targeted evaluations, based on 26 targeted scenarios, to analyze specific aspects (e.g. reasoning, disinformation). Third, we conduct a large-scale evaluation of 30 prominent language models (spanning open, limited-access, and closed models) on all 42 scenarios, 21 of which were not previously used in mainstream LM evaluation. Prior to HELM, models on average were evaluated on just 17.9% of the core HELM scenarios, with some prominent models not sharing a single scenario in common. We improve this to 96.0%: now all 30 models have been densely benchmarked on the same core scenarios and metrics under standardized conditions. Our evaluation surfaces 25 top-level findings. For full transparency, we release all raw model prompts and completions publicly for further analysis, as well as a general modular toolkit. We intend for HELM to be a living benchmark for the community, continuously updated with new scenarios, metrics, and models.
△ Less
Submitted 1 October, 2023; v1 submitted 16 November, 2022;
originally announced November 2022.
-
Heegaard Floer invariants for cyclic 3-orbifolds
Authors:
Saibal Ganguli,
Mainak Poddar
Abstract:
We define a notion of Heegaard Floer homology for three dimensional orbifolds with arbitrary cyclic singularities, generalizing the recent work of Biji Wong where the singular locus is assumed to be connected.
We define a notion of Heegaard Floer homology for three dimensional orbifolds with arbitrary cyclic singularities, generalizing the recent work of Biji Wong where the singular locus is assumed to be connected.
△ Less
Submitted 13 February, 2024; v1 submitted 22 October, 2022;
originally announced October 2022.
-
Toward Next-Generation Artificial Intelligence: Catalyzing the NeuroAI Revolution
Authors:
Anthony Zador,
Sean Escola,
Blake Richards,
Bence Ölveczky,
Yoshua Bengio,
Kwabena Boahen,
Matthew Botvinick,
Dmitri Chklovskii,
Anne Churchland,
Claudia Clopath,
James DiCarlo,
Surya Ganguli,
Jeff Hawkins,
Konrad Koerding,
Alexei Koulakov,
Yann LeCun,
Timothy Lillicrap,
Adam Marblestone,
Bruno Olshausen,
Alexandre Pouget,
Cristina Savin,
Terrence Sejnowski,
Eero Simoncelli,
Sara Solla,
David Sussillo
, et al. (2 additional authors not shown)
Abstract:
Neuroscience has long been an essential driver of progress in artificial intelligence (AI). We propose that to accelerate progress in AI, we must invest in fundamental research in NeuroAI. A core component of this is the embodied Turing test, which challenges AI animal models to interact with the sensorimotor world at skill levels akin to their living counterparts. The embodied Turing test shifts…
▽ More
Neuroscience has long been an essential driver of progress in artificial intelligence (AI). We propose that to accelerate progress in AI, we must invest in fundamental research in NeuroAI. A core component of this is the embodied Turing test, which challenges AI animal models to interact with the sensorimotor world at skill levels akin to their living counterparts. The embodied Turing test shifts the focus from those capabilities like game playing and language that are especially well-developed or uniquely human to those capabilities, inherited from over 500 million years of evolution, that are shared with all animals. Building models that can pass the embodied Turing test will provide a roadmap for the next generation of AI.
△ Less
Submitted 22 February, 2023; v1 submitted 15 October, 2022;
originally announced October 2022.
-
What does a deep neural network confidently perceive? The effective dimension of high certainty class manifolds and their low confidence boundaries
Authors:
Stanislav Fort,
Ekin Dogus Cubuk,
Surya Ganguli,
Samuel S. Schoenholz
Abstract:
Deep neural network classifiers partition input space into high confidence regions for each class. The geometry of these class manifolds (CMs) is widely studied and intimately related to model performance; for example, the margin depends on CM boundaries. We exploit the notions of Gaussian width and Gordon's escape theorem to tractably estimate the effective dimension of CMs and their boundaries t…
▽ More
Deep neural network classifiers partition input space into high confidence regions for each class. The geometry of these class manifolds (CMs) is widely studied and intimately related to model performance; for example, the margin depends on CM boundaries. We exploit the notions of Gaussian width and Gordon's escape theorem to tractably estimate the effective dimension of CMs and their boundaries through tomographic intersections with random affine subspaces of varying dimension. We show several connections between the dimension of CMs, generalization, and robustness. In particular we investigate how CM dimension depends on 1) the dataset, 2) architecture (including ResNet, WideResNet \& Vision Transformer), 3) initialization, 4) stage of training, 5) class, 6) network width, 7) ensemble size, 8) label randomization, 9) training set size, and 10) robustness to data corruption. Together a picture emerges that higher performing and more robust models have higher dimensional CMs. Moreover, we offer a new perspective on ensembling via intersections of CMs. Our code is at https://github.com/stanislavfort/slice-dice-optimize/
△ Less
Submitted 11 October, 2022;
originally announced October 2022.
-
The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural Networks
Authors:
Daniel Kunin,
Atsushi Yamamura,
Chao Ma,
Surya Ganguli
Abstract:
In this work, we explore the maximum-margin bias of quasi-homogeneous neural networks trained with gradient flow on an exponential loss and past a point of separability. We introduce the class of quasi-homogeneous models, which is expressive enough to describe nearly all neural networks with homogeneous activations, even those with biases, residual connections, and normalization layers, while stru…
▽ More
In this work, we explore the maximum-margin bias of quasi-homogeneous neural networks trained with gradient flow on an exponential loss and past a point of separability. We introduce the class of quasi-homogeneous models, which is expressive enough to describe nearly all neural networks with homogeneous activations, even those with biases, residual connections, and normalization layers, while structured enough to enable geometric analysis of its gradient dynamics. Using this analysis, we generalize the existing results of maximum-margin bias for homogeneous networks to this richer class of models. We find that gradient flow implicitly favors a subset of the parameters, unlike in the case of a homogeneous model where all parameters are treated equally. We demonstrate through simple examples how this strong favoritism toward minimizing an asymmetric norm can degrade the robustness of quasi-homogeneous models. On the other hand, we conjecture that this norm-minimization discards, when possible, unnecessary higher-order parameters, reducing the model to a sparser parameterization. Lastly, by applying our theorem to sufficiently expressive neural networks with normalization layers, we reveal a universal mechanism behind the empirical phenomenon of Neural Collapse.
△ Less
Submitted 16 February, 2023; v1 submitted 7 October, 2022;
originally announced October 2022.
-
Scalable Self-Supervised Representation Learning from Spatiotemporal Motion Trajectories for Multimodal Computer Vision
Authors:
Swetava Ganguli,
C. V. Krishnakumar Iyer,
Vipul Pandey
Abstract:
Self-supervised representation learning techniques utilize large datasets without semantic annotations to learn meaningful, universal features that can be conveniently transferred to solve a wide variety of downstream supervised tasks. In this work, we propose a self-supervised method for learning representations of geographic locations from unlabeled GPS trajectories to solve downstream geospatia…
▽ More
Self-supervised representation learning techniques utilize large datasets without semantic annotations to learn meaningful, universal features that can be conveniently transferred to solve a wide variety of downstream supervised tasks. In this work, we propose a self-supervised method for learning representations of geographic locations from unlabeled GPS trajectories to solve downstream geospatial computer vision tasks. Tiles resulting from a raster representation of the earth's surface are modeled as nodes on a graph or pixels of an image. GPS trajectories are modeled as allowed Markovian paths on these nodes. A scalable and distributed algorithm is presented to compute image-like representations, called reachability summaries, of the spatial connectivity patterns between tiles and their neighbors implied by the observed Markovian paths. A convolutional, contractive autoencoder is trained to learn compressed representations, called reachability embeddings, of reachability summaries for every tile. Reachability embeddings serve as task-agnostic, feature representations of geographic locations. Using reachability embeddings as pixel representations for five different downstream geospatial tasks, cast as supervised semantic segmentation problems, we quantitatively demonstrate that reachability embeddings are semantically meaningful representations and result in 4-23% gain in performance, as measured using area under the precision-recall curve (AUPRC) metric, when compared to baseline models that use pixel representations that do not account for the spatial connectivity between tiles. Reachability embeddings transform sequential, spatiotemporal mobility data into semantically meaningful tensor representations that can be combined with other sources of imagery and are designed to facilitate multimodal learning in geospatial computer vision.
△ Less
Submitted 6 October, 2022;
originally announced October 2022.
-
Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask?
Authors:
Mansheej Paul,
Feng Chen,
Brett W. Larsen,
Jonathan Frankle,
Surya Ganguli,
Gintare Karolina Dziugaite
Abstract:
Modern deep learning involves training costly, highly overparameterized networks, thus motivating the search for sparser networks that can still be trained to the same accuracy as the full network (i.e. matching). Iterative magnitude pruning (IMP) is a state of the art algorithm that can find such highly sparse matching subnetworks, known as winning tickets. IMP operates by iterative cycles of tra…
▽ More
Modern deep learning involves training costly, highly overparameterized networks, thus motivating the search for sparser networks that can still be trained to the same accuracy as the full network (i.e. matching). Iterative magnitude pruning (IMP) is a state of the art algorithm that can find such highly sparse matching subnetworks, known as winning tickets. IMP operates by iterative cycles of training, masking smallest magnitude weights, rewinding back to an early training point, and repeating. Despite its simplicity, the underlying principles for when and how IMP finds winning tickets remain elusive. In particular, what useful information does an IMP mask found at the end of training convey to a rewound network near the beginning of training? How does SGD allow the network to extract this information? And why is iterative pruning needed? We develop answers in terms of the geometry of the error landscape. First, we find that$\unicode{x2014}$at higher sparsities$\unicode{x2014}$pairs of pruned networks at successive pruning iterations are connected by a linear path with zero error barrier if and only if they are matching. This indicates that masks found at the end of training convey the identity of an axial subspace that intersects a desired linearly connected mode of a matching sublevel set. Second, we show SGD can exploit this information due to a strong form of robustness: it can return to this mode despite strong perturbations early in training. Third, we show how the flatness of the error landscape at the end of training determines a limit on the fraction of weights that can be pruned at each iteration of IMP. Finally, we show that the role of retraining in IMP is to find a network with new small weights to prune. Overall, these results make progress toward demystifying the existence of winning tickets by revealing the fundamental role of error landscape geometry.
△ Less
Submitted 6 October, 2022;
originally announced October 2022.
-
Disentanglement with Biological Constraints: A Theory of Functional Cell Types
Authors:
James C. R. Whittington,
Will Dorrell,
Surya Ganguli,
Timothy E. J. Behrens
Abstract:
Neurons in the brain are often finely tuned for specific task variables. Moreover, such disentangled representations are highly sought after in machine learning. Here we mathematically prove that simple biological constraints on neurons, namely nonnegativity and energy efficiency in both activity and weights, promote such sought after disentangled representations by enforcing neurons to become sel…
▽ More
Neurons in the brain are often finely tuned for specific task variables. Moreover, such disentangled representations are highly sought after in machine learning. Here we mathematically prove that simple biological constraints on neurons, namely nonnegativity and energy efficiency in both activity and weights, promote such sought after disentangled representations by enforcing neurons to become selective for single factors of task variation. We demonstrate these constraints lead to disentanglement in a variety of tasks and architectures, including variational autoencoders. We also use this theory to explain why the brain partitions its cells into distinct cell types such as grid and object-vector cells, and also explain when the brain instead entangles representations in response to entangled task factors. Overall, this work provides a mathematical understanding of why single neurons in the brain often represent single human-interpretable factors, and steps towards an understanding task structure shapes the structure of brain representation.
△ Less
Submitted 31 March, 2023; v1 submitted 30 September, 2022;
originally announced October 2022.
-
Visualization of moiré magnons in monolayer ferromagnet
Authors:
Somesh Chandra Ganguli,
Markus Aapro,
Shawulienu Kezilebieke,
Mohammad Amini,
Jose L. Lado,
Peter Liljeroth
Abstract:
Two-dimensional magnetic materials provide an ideal platform to explore collective many-body excitations associated with spin fluctuations. In particular, it should be feasible to explore, manipulate and ultimately design magnonic excitations in two-dimensional van der Waals magnets in a controllable way. Here we demonstrate the emergence of moiré magnon excitations, stemming from the interplay of…
▽ More
Two-dimensional magnetic materials provide an ideal platform to explore collective many-body excitations associated with spin fluctuations. In particular, it should be feasible to explore, manipulate and ultimately design magnonic excitations in two-dimensional van der Waals magnets in a controllable way. Here we demonstrate the emergence of moiré magnon excitations, stemming from the interplay of spin-excitations in monolayer CrBr$_3$ and the moiré pattern stemming from the lattice mismatch with the underlying substrate. The existence of moiré magnons is further confirmed via inelastic quasiparticle interference, showing the appearance of a dispersion pattern correlated with the moiré length scale. Our results provide a direct visualization in real-space of the dispersion of moiré magnons, demonstrating the versatility of moiré patterns in creating emerging many-body excitations.
△ Less
Submitted 21 January, 2023; v1 submitted 23 August, 2022;
originally announced August 2022.
-
Beyond neural scaling laws: beating power law scaling via data pruning
Authors:
Ben Sorscher,
Robert Geirhos,
Shashank Shekhar,
Surya Ganguli,
Ari S. Morcos
Abstract:
Widely observed neural scaling laws, in which error falls off as a power of the training set size, model size, or both, have driven substantial performance improvements in deep learning. However, these improvements through scaling alone require considerable costs in compute and energy. Here we focus on the scaling of error with dataset size and show how in theory we can break beyond power law scal…
▽ More
Widely observed neural scaling laws, in which error falls off as a power of the training set size, model size, or both, have driven substantial performance improvements in deep learning. However, these improvements through scaling alone require considerable costs in compute and energy. Here we focus on the scaling of error with dataset size and show how in theory we can break beyond power law scaling and potentially even reduce it to exponential scaling instead if we have access to a high-quality data pruning metric that ranks the order in which training examples should be discarded to achieve any pruned dataset size. We then test this improved scaling prediction with pruned dataset size empirically, and indeed observe better than power law scaling in practice on ResNets trained on CIFAR-10, SVHN, and ImageNet. Next, given the importance of finding high-quality pruning metrics, we perform the first large-scale benchmarking study of ten different data pruning metrics on ImageNet. We find most existing high performing metrics scale poorly to ImageNet, while the best are computationally intensive and require labels for every image. We therefore developed a new simple, cheap and scalable self-supervised pruning metric that demonstrates comparable performance to the best supervised metrics. Overall, our work suggests that the discovery of good data-pruning metrics may provide a viable path forward to substantially improved neural scaling laws, thereby reducing the resource costs of modern deep learning.
△ Less
Submitted 21 April, 2023; v1 submitted 29 June, 2022;
originally announced June 2022.
-
Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable Networks
Authors:
Mansheej Paul,
Brett W. Larsen,
Surya Ganguli,
Jonathan Frankle,
Gintare Karolina Dziugaite
Abstract:
A striking observation about iterative magnitude pruning (IMP; Frankle et al. 2020) is that $\unicode{x2014}$ after just a few hundred steps of dense training $\unicode{x2014}$ the method can find a sparse sub-network that can be trained to the same accuracy as the dense network. However, the same does not hold at step 0, i.e. random initialization. In this work, we seek to understand how this ear…
▽ More
A striking observation about iterative magnitude pruning (IMP; Frankle et al. 2020) is that $\unicode{x2014}$ after just a few hundred steps of dense training $\unicode{x2014}$ the method can find a sparse sub-network that can be trained to the same accuracy as the dense network. However, the same does not hold at step 0, i.e. random initialization. In this work, we seek to understand how this early phase of pre-training leads to a good initialization for IMP both through the lens of the data distribution and the loss landscape geometry. Empirically we observe that, holding the number of pre-training iterations constant, training on a small fraction of (randomly chosen) data suffices to obtain an equally good initialization for IMP. We additionally observe that by pre-training only on "easy" training data, we can decrease the number of steps necessary to find a good initialization for IMP compared to training on the full dataset or a randomly chosen subset. Finally, we identify novel properties of the loss landscape of dense networks that are predictive of IMP performance, showing in particular that more examples being linearly mode connected in the dense network correlates well with good initializations for IMP. Combined, these results provide new insight into the role played by the early phase training in IMP.
△ Less
Submitted 2 June, 2022;
originally announced June 2022.
-
MetaMorph: Learning Universal Controllers with Transformers
Authors:
Agrim Gupta,
Linxi Fan,
Surya Ganguli,
Li Fei-Fei
Abstract:
Multiple domains like vision, natural language, and audio are witnessing tremendous progress by leveraging Transformers for large scale pre-training followed by task specific fine tuning. In contrast, in robotics we primarily train a single robot for a single task. However, modular robot systems now allow for the flexible combination of general-purpose building blocks into task optimized morpholog…
▽ More
Multiple domains like vision, natural language, and audio are witnessing tremendous progress by leveraging Transformers for large scale pre-training followed by task specific fine tuning. In contrast, in robotics we primarily train a single robot for a single task. However, modular robot systems now allow for the flexible combination of general-purpose building blocks into task optimized morphologies. However, given the exponentially large number of possible robot morphologies, training a controller for each new design is impractical. In this work, we propose MetaMorph, a Transformer based approach to learn a universal controller over a modular robot design space. MetaMorph is based on the insight that robot morphology is just another modality on which we can condition the output of a Transformer. Through extensive experiments we demonstrate that large scale pre-training on a variety of robot morphologies results in policies with combinatorial generalization capabilities, including zero shot generalization to unseen robot morphologies. We further demonstrate that our pre-trained policy can be used for sample-efficient transfer to completely new robot morphologies and tasks.
△ Less
Submitted 22 March, 2022;
originally announced March 2022.
-
Evidence of nodal superconductivity in monolayer 1H-TaS$_2$ with hidden order fluctuations
Authors:
Viliam Vaňo,
Somesh Chandra Ganguli,
Mohammad Amini,
Linghao Yan,
Maryam Khosravian,
Guangze Chen,
Shawulienu Kezilebieke,
Jose L. Lado,
Peter Liljeroth
Abstract:
Unconventional superconductors represent one of the fundamental directions in modern quantum materials research. In particular, nodal superconductors are known to appear naturally in strongly correlated systems, including cuprate superconductors and heavy-fermion systems. Van der Waals materials hosting superconducting states are well known, yet nodal monolayer van der Waals superconductors have r…
▽ More
Unconventional superconductors represent one of the fundamental directions in modern quantum materials research. In particular, nodal superconductors are known to appear naturally in strongly correlated systems, including cuprate superconductors and heavy-fermion systems. Van der Waals materials hosting superconducting states are well known, yet nodal monolayer van der Waals superconductors have remained elusive. Here, using low-temperature scanning tunneling microscopy (STM) and spectroscopy (STS) experiments, we show that pristine monolayer 1H-TaS$_2$ realizes a nodal superconducting state. By including non-magnetic disorder, we drive the nodal superconducting state to a conventional gapped s-wave state. Furthermore, we observe the emergence of many-body excitations close to the gap edge, signalling a potential unconventional pairing mechanism. Our results demonstrate the emergence of nodal superconductivity in a van der Waals monolayer, providing a building block for van der Waals heterostructures exploiting unconventional superconducting states.
△ Less
Submitted 21 August, 2023; v1 submitted 14 December, 2021;
originally announced December 2021.