-
Exponential Capacity in Multilayer Hetero-Associative Neural Networks
Authors:
Elena Agliari,
Adriano Barra,
Andrea Ladiana,
Andrea Lepre
Abstract:
Exponential Hopfield networks store a number of patterns that grows exponentially with the number of neurons, and in their classical formulation they are auto-associative: they complete a corrupted copy of a memory into the memory itself. Many of the tasks one wants such a network to perform are instead hetero-associative, mapping a cue to a different target. We introduce and analyse an exponentia…
▽ More
Exponential Hopfield networks store a number of patterns that grows exponentially with the number of neurons, and in their classical formulation they are auto-associative: they complete a corrupted copy of a memory into the memory itself. Many of the tasks one wants such a network to perform are instead hetero-associative, mapping a cue to a different target. We introduce and analyse an exponential neural network of $L$ layers of $N$ binary neurons, each layer carrying its own dataset, whose energy is an exponential of the product of the per-layer Mattis overlaps, so that it is minimised precisely when every layer retrieves the pattern of the same index; the stored association must be a surjective function of the cue, and we show why nothing else can be stored at all. A cavity/signal-to-noise analysis, made exact at leading order by a large-deviation evaluation of the noise, shows that the aligned hetero-associative state is a fixed point of the zero-temperature dynamics up to a number of stored patterns $P_c\sim e^{Nρ_L}$, exponential in the layer size, with an explicit rate $ρ_L$ that grows like $L\log 2$; enlarging the basins of attraction lowers the rate but never destroys its exponential character. Comparing the theory with structured data we find that the exponential capacity and the predicted basins survive correlated, many-to-one patterns: the network is a near-perfect content-addressable memory. The same closed forms describe, without refitting, a synthetic manifold, real T-cell-receptor/epitope triples and natural-language intent data, so the mechanism is domain-universal. Generalisation to unseen cues, though significantly above chance, stays below memorisation, and it is the geometry of the encoding, rather than the data domain, that sets how far above chance it reaches. In this family, exponential storage and strong generalisation are distinct capabilities.
△ Less
Submitted 31 July, 2026;
originally announced July 2026.
-
Do Hopfield Networks Dream of Stored Patterns? A Statistical-Mechanical Theory of Dreaming in Multidirectional Associative Memories
Authors:
Adriano Barra,
Fabrizio Durante,
Andrea Ladiana,
Michela Marra Solazzo
Abstract:
We introduce the Dreaming $L$-directional Associative Memory (DLAM), a multi-layer Hebbian architecture in which off-line dreaming and supervised heteroassociative coupling coexist within a single energy function, placing our approach within the framework of energy-based models (EBMs). The replica-symmetric free energy, derived via the Guerra interpolation scheme, yields self-consistency equations…
▽ More
We introduce the Dreaming $L$-directional Associative Memory (DLAM), a multi-layer Hebbian architecture in which off-line dreaming and supervised heteroassociative coupling coexist within a single energy function, placing our approach within the framework of energy-based models (EBMs). The replica-symmetric free energy, derived via the Guerra interpolation scheme, yields self-consistency equations governing the order parameters across the control-parameter space. The effective local field decomposes into signal, intra-layer dreaming noise, and inter-layer noise. Dreaming improves retrieval by differentially attenuating high-eigenvalue interference modes of the empirical correlation matrix, suppressing inter-pattern crosstalk while preserving the signal. Dreaming and inter-layer coupling prove synergistic, opening retrieval regions unreachable by either mechanism alone, as confirmed by Monte Carlo simulations for $L=3$. Their interplay is most pronounced on pattern disentanglement: given a mixture state as input, the network splits the constituent patterns one-per-layer, recovering each modality-specific pattern from a common cue that simultaneously blends noisy evidence from all sensory channels. Phase diagrams are planar projections of the hyperspace $(α,β,ρ,t)$-where $α$ is the storage load, $β$ the fast-noise inverse temperature, $ρ$ the dataset entropy, and $t$ the sleeping time. In the $(ρ,t)$-plane, the diagrams reveal a data-computation trade-off: off-line consolidation substitutes for additional training data, extending to heteroassociative architectures a phenomenon previously established for autoassociative networks. Enriching the standard Hopfield model with heteroassociativity and dreaming gives rise to EBMs capable of complex tasks beyond classical pattern recognition, contributing to a modern theory of neural information processing.
△ Less
Submitted 13 May, 2026;
originally announced May 2026.
-
Partial annealing and pattern decorrelation in associative neural networks
Authors:
Linda Albanese,
Andrea Alessandrelli,
Adriano Barra,
Silvio Franz,
Federico Ricci-Tersenghi
Abstract:
Using the Hopfield model as a benchmark case, the present work focuses on the investigation of partially annealed associative neural networks, wherein neural dynamics is coupled to slowly evolving patterns within the two-temperature-two-timescale framework. This setting inherently introduces a real parameter n, reminiscent of the number of replicas in the celebrated replica trick, that tunes the s…
▽ More
Using the Hopfield model as a benchmark case, the present work focuses on the investigation of partially annealed associative neural networks, wherein neural dynamics is coupled to slowly evolving patterns within the two-temperature-two-timescale framework. This setting inherently introduces a real parameter n, reminiscent of the number of replicas in the celebrated replica trick, that tunes the separation of timescales and the effective interaction between fast (i.e. the neurons) and slow (i.e. the synapses) degrees of freedom. By adapting Guerra's interpolation to the case, we derive the free energy without relying on analytical continuation. The obtained results demonstrate that negative values of n induce a progressive decorrelation of the stored patterns, thereby effectively reducing interference, promoting orthogonal configurations and ultimately conferring to the network the maximal storage alphac=1. Numerical simulations based on a mean field Monte Carlo dynamics have been employed to confirm this scenario and prove that partial annealing restores retrieval in challenging regimes, such as in the presence of biased patterns, outperforming standard decorrelation methods. These findings underscore the notion of partial annealing as an adaptive mechanism for enhancing memory organisation and retrieval in complex systems.
△ Less
Submitted 11 May, 2026;
originally announced May 2026.
-
Magnetic-field-induced magnon portfolio in a van der Waals magnet
Authors:
T. Riccardi,
F. Le Mardélé,
L. A. Veyrat de Lachenal,
A. Pawbake,
I. Plutnarova,
Z. Sofer,
G. Jacquet,
F. Petot,
A. Saùl,
B. Grémaud,
A. L. Barra,
M. Orlita,
J. Coraux,
C. Faugeras,
B. A. Piot
Abstract:
Magnonic excitations are investigated in chromium oxychloride (CrOCl), a van der Waal (vdW) antiferromagnet prone to a multitude of magnetic phase transitions, with absorption experiments in a broad continuous energy range. At low magnetic fields, the magnon spectra show a strong bi-axial anisotropy and inform on the relative weights of the effective exchange coupling and the system anisotropies.…
▽ More
Magnonic excitations are investigated in chromium oxychloride (CrOCl), a van der Waal (vdW) antiferromagnet prone to a multitude of magnetic phase transitions, with absorption experiments in a broad continuous energy range. At low magnetic fields, the magnon spectra show a strong bi-axial anisotropy and inform on the relative weights of the effective exchange coupling and the system anisotropies. As the magnetic field increases, magnons characteristic of a canted phase are first observed, with peculiarities attributed to in-plane anisotropies and magnon-magnon coupling. Subsequently, a hysteretic magnon spectrum appears as the system transitions to a ferrimagnetic state, with two new magnon branches partly coexisting with the lower energy canted phase branch, indicating the formation of spatially separated magnetic phases. Further changes in the magnon spectrum in higher magnetic fields accompany transitions to the different canted magnetic phases previously reported. Our experiments show that competing exchange interactions and ground states broaden the options to generate different kinds of magnonic excitations in the same vdW material upon the variation of external parameters.
△ Less
Submitted 13 March, 2026;
originally announced March 2026.
-
Serial vs parallel recall in the Blume-Every-Griffiths neural networks
Authors:
Linda Albanese,
Andrea Alessandrelli,
Adriano Barra,
Emilio N. M. Cirillo
Abstract:
Fully connected Blume-Emery-Griffiths neural networks performing pattern recognition and associative memory have been heuristically studied in the past (mainly via the replica trick and under the replica symmetric assumption) as generalization of the standard Hopfield reference. In these notes, at first, by relying upon Guerra interpolation, we re-obtain the existing picture rigorously. Next we sh…
▽ More
Fully connected Blume-Emery-Griffiths neural networks performing pattern recognition and associative memory have been heuristically studied in the past (mainly via the replica trick and under the replica symmetric assumption) as generalization of the standard Hopfield reference. In these notes, at first, by relying upon Guerra interpolation, we re-obtain the existing picture rigorously. Next we show that, due to dilution in the patterns, these networks are able to switch from serial recall (where one pattern is retrieved per time) to parallel recall (where several patterns are retrieved at once) and the larger the dilution, the stronger this emerging multi-tasking capability. In particular, we inspect the regimes of mild dilution (where solely a low storage of pattern can be enabled) and extreme dilution (where a medium storage of patterns can be sustained) separately as they give rise to different outcomes: the former displays hierarchical recall (distributing the amplitudes of the retrieved signals with different amplitudes), the latter -- instead -- performs a equal-strength recall (where a O(1) fraction of all the patterns is simultaneously retrieved with the same amplitude per pattern). Finally, in order to implement graded responses in the neurons, variations on theme obtained by enlarging the possible values of neural activity these neurons may sustain are also discussed generalizing the Ghatak-Sherrington model for inverse freezing in Hebbian terms.
△ Less
Submitted 12 January, 2026;
originally announced January 2026.
-
On the supra-linear storage in dense networks of grid and place cells
Authors:
Adriano Barra,
Martino S. Centonze,
Michela Marra Solazzo,
Daniele Tantari
Abstract:
Place-cell networks, typically forced to pairwise synaptic interactions, are widely studied as models of cognitive maps: such models, however, share a severely limited storage capacity, scaling linearly with network size and with a very small critical storage. This limitation is a challenge for navigation in 3-dimensional space because, oversimplifying, if encoding motion along a one-dimensional t…
▽ More
Place-cell networks, typically forced to pairwise synaptic interactions, are widely studied as models of cognitive maps: such models, however, share a severely limited storage capacity, scaling linearly with network size and with a very small critical storage. This limitation is a challenge for navigation in 3-dimensional space because, oversimplifying, if encoding motion along a one-dimensional trajectory embedded in 2-dimensions requires $O(K)$ patterns (interpreted as bins), extending this to a 2-dimensional manifold embedded in a 3-dimensional space -- yet preserving the same resolution -- requires roughly $O(K^2)$ patterns, namely a supra-linear amount of patterns. In these regards, dense Hebbian architectures, where higher-order neural assemblies mediate memory retrieval, display much larger capacities and are increasingly recognized as biologically plausible, but have never linked to place cells so far. Here we propose a minimal two-layer model, with place cells building a layer and leaving the other layer populated by neural units that account for the internal representations (so to qualitatively resemble grid cells in the MEC of mammals): crucially, by assuming that each place cell interacts with pairs of grid cells, we show how such a model is formally equivalent to a dense Battaglia-Treves-like Hebbian network of grid cells only endowed with four-body interactions. By studying its emergent computational properties by means of statistical mechanics of disordered systems, we prove -- analytically -- that such effective higher-order assemblies (constructed under the guise of biological plausibility) can support supra-linear storage of continuous attractors; furthermore, we prove -- numerically -- that the present neural network is capable of recognition and navigation on general surfaces embedded in a 3-dimensional space.
△ Less
Submitted 21 November, 2025; v1 submitted 4 November, 2025;
originally announced November 2025.
-
Yet another exponential Hopfield model
Authors:
Linda Albanese,
Andrea Alessandrelli,
Adriano Barra,
Peter Sollich
Abstract:
We propose and analyze a new variation of the so-called {\em exponential Hopfield model}, a recently introduced family of associative neural networks with unprecedented storage capacity. Our construction is based on a cost function defined through exponentials of standard quadratic loss functions, which naturally favors configurations corresponding to perfect recall. Despite not being a mean-field…
▽ More
We propose and analyze a new variation of the so-called {\em exponential Hopfield model}, a recently introduced family of associative neural networks with unprecedented storage capacity. Our construction is based on a cost function defined through exponentials of standard quadratic loss functions, which naturally favors configurations corresponding to perfect recall. Despite not being a mean-field system, the model admits a tractable mathematical analysis of its dynamics and retrieval properties that agree with those for the original exponential model introduced by Ramsauer and coworkers. By means of a signal-to-noise approach, we demonstrate that stored patterns remain stable fixed points of the zero-temperature dynamics up to an exponentially large number of patterns in the system size. We further quantify the basins of attraction of the retrieved memories, showing that while enlarging their radius reduces the overall load, the storage capacity nonetheless retains its exponential scaling. An independent derivation within the perfect recall regime confirms these results and provides an estimate of the relevant prefactors. Our findings thus complement and extend previous studies on exponential Hopfield networks, establishing that even under robustness constraints these models preserve their exceptional storage capabilities. Beyond their theoretical interest, such networks point towards principled mechanisms for massively scalable associative memory, with potential implications for both neuroscience-inspired computation and high-capacity machine learning architectures.
△ Less
Submitted 8 September, 2025;
originally announced September 2025.
-
Supervised and Unsupervised protocols for hetero-associative neural networks
Authors:
Andrea Alessandrelli,
Adriano Barra,
Andrea Ladiana,
Andrea Lepre,
Federico Ricci-Tersenghi
Abstract:
This paper introduces a learning framework for Three-Directional Associative Memory (TAM) models, extending the classical Hebbian paradigm to both supervised and unsupervised protocols within an hetero-associative setting. These neural networks consist of three interconnected layers of binary neurons interacting via generalized Hebbian synaptic couplings that allow learning, storage and retrieval…
▽ More
This paper introduces a learning framework for Three-Directional Associative Memory (TAM) models, extending the classical Hebbian paradigm to both supervised and unsupervised protocols within an hetero-associative setting. These neural networks consist of three interconnected layers of binary neurons interacting via generalized Hebbian synaptic couplings that allow learning, storage and retrieval of structured triplets of patterns. By relying upon glassy statistical mechanical techniques (mainly replica theory and Guerra interpolation), we analyze the emergent computational properties of these networks, at work with random (Rademacher) datasets and at the replica-symmetric level of description: we obtain a set of self-consistency equations for the order parameters that quantify the critical dataset sizes (i.e. their thresholds for learning) and describe the retrieval performance of these networks, highlighting the differences between supervised and unsupervised protocols. Numerical simulations validate our theoretical findings and demonstrate the robustness of the captured picture about TAMs also at work with structured datasets. In particular, this study provides insights into the cooperative interplay of layers, beyond that of the neurons within the layers, with potential implications for optimal design of artificial neural network architectures.
△ Less
Submitted 24 May, 2025;
originally announced May 2025.
-
Guerra interpolation for inverse freezing
Authors:
Linda Albanese,
Adriano Barra,
Emilio N. M. Cirillo
Abstract:
In these short notes, we adapt and systematically apply Guerra's interpolation techniques on a class of disordered mean-field spin glasses equipped with crystal fields and multi-value spin variables. These models undergo the phenomenon of inverse melting or inverse freezing. In particular, we focus on the Ghatak-Sherrington model, its extension provided by Katayama and Horiguchi, and the disordere…
▽ More
In these short notes, we adapt and systematically apply Guerra's interpolation techniques on a class of disordered mean-field spin glasses equipped with crystal fields and multi-value spin variables. These models undergo the phenomenon of inverse melting or inverse freezing. In particular, we focus on the Ghatak-Sherrington model, its extension provided by Katayama and Horiguchi, and the disordered Blume-Emery-Griffiths-Capel model in the mean-field limit, deepened by Crisanti and Leuzzi and by Schupper and Shnerb. Once shown how all these models can be retrieved as particular limits of a unique broader Hamiltonian, we study their free energies. We provide explicit expressions of their annealed and quenched expectations, inspecting the cases of replica symmetry and (first-step) broken replica symmetry. We recover several results previously obtained via heuristic approaches (mainly the replica trick) to prove full agreement with the existing literature. As a sideline, we also inspect the onset of replica symmetry breaking by providing analytically the expression of the de Almeida-Thouless instability line for a replica symmetric description: in this general setting, the latter is new also from a physical viewpoint.
△ Less
Submitted 9 May, 2025;
originally announced May 2025.
-
Beyond Disorder: Unveiling Cooperativeness in Multidirectional Associative Memories
Authors:
Andrea Alessandrelli,
Adriano Barra,
Andrea Ladiana,
Andrea Lepre,
Federico Ricci-Tersenghi
Abstract:
By leveraging tools from the statistical mechanics of complex systems, in these short notes we extend the architecture of a neural network for hetero-associative memory (called three-directional associative memories, TAM) to explore supervised and unsupervised learning protocols. In particular, by providing entropic-heterogeneous datasets to its various layers, we predict and quantify a new emerge…
▽ More
By leveraging tools from the statistical mechanics of complex systems, in these short notes we extend the architecture of a neural network for hetero-associative memory (called three-directional associative memories, TAM) to explore supervised and unsupervised learning protocols. In particular, by providing entropic-heterogeneous datasets to its various layers, we predict and quantify a new emergent phenomenon -- that we term {\em layer's cooperativeness} -- where the interplay of dataset entropies across network's layers enhances their retrieval capabilities Beyond those they would have without reciprocal influence. Naively we would expect layers trained with less informative datasets to develop smaller retrieval regions compared to those pertaining to layers that experienced more information: this does not happen and all the retrieval regions settle to the same amplitude, allowing for optimal retrieval performance globally. This cooperative dynamics marks a significant advancement in understanding emergent computational capabilities within disordered systems.
△ Less
Submitted 6 March, 2025;
originally announced March 2025.
-
Antiferromagnetic resonance in $α$-MnTe
Authors:
J. Dzian,
P. Kubaščík,
S. Tázlarů,
M. Białek,
M. Šindler,
F. Le Mardelé,
C. Kadlec,
F. Kadlec,
M. Gryglas-Borysiewicz,
K. P. Kluczyk,
A. Mycielski,
P. Skupiński,
J. Hejtmánek,
R. Tesař,
J. Železný,
A. -L. Barra,
C. Faugeras,
J. Volný,
K. Uhlířová,
L. Nádvorník,
M. Veis,
K. Výborný,
M. Orlita
Abstract:
Antiferromagnetic resonance in a bulk $α$-MnTe crystal is investigated using both frequency-domain and time-domain THz spectroscopy techniques. At low temperatures, an excitation at the photon energy of $(3.5\pm 0.1)$~meV is observed and identified as a magnon mode through its distinctive dependence on temperature and magnetic field. This behavior is reproduced using a simplified model for antifer…
▽ More
Antiferromagnetic resonance in a bulk $α$-MnTe crystal is investigated using both frequency-domain and time-domain THz spectroscopy techniques. At low temperatures, an excitation at the photon energy of $(3.5\pm 0.1)$~meV is observed and identified as a magnon mode through its distinctive dependence on temperature and magnetic field. This behavior is reproduced using a simplified model for antiferromagnetic resonance in an easy-plane antiferromagnet. The results of our experiments, when compared to exchange constants established in the literature, allow us to extract the out-of-plane component of the single-ion magnetic anisotropy reaching $(40\pm10)$~$μ$eV.
△ Less
Submitted 11 July, 2025; v1 submitted 26 February, 2025;
originally announced February 2025.
-
Networks of neural networks: more is different
Authors:
Elena Agliari,
Andrea Alessandrelli,
Adriano Barra,
Martino Salomone Centonze,
Federico Ricci-Tersenghi
Abstract:
The common thread behind the recent Nobel Prize in Physics to John Hopfield and those conferred to Giorgio Parisi in 2021 and Philip Anderson in 1977 is disorder. Quoting Philip Anderson: "more is different". This principle has been extensively demonstrated in magnetic systems and spin glasses, and, in this work, we test its validity on Hopfield neural networks to show how an assembly of these mod…
▽ More
The common thread behind the recent Nobel Prize in Physics to John Hopfield and those conferred to Giorgio Parisi in 2021 and Philip Anderson in 1977 is disorder. Quoting Philip Anderson: "more is different". This principle has been extensively demonstrated in magnetic systems and spin glasses, and, in this work, we test its validity on Hopfield neural networks to show how an assembly of these models displays emergent capabilities that are not present at a single network level. Such an assembly is designed as a layered associative Hebbian network that, beyond accomplishing standard pattern recognition, spontaneously performs also pattern disentanglement. Namely, when inputted with a composite signal -- e.g., a musical chord -- it can return the single constituting elements -- e.g., the notes making up the chord. Here, restricting to notes coded as Rademacher vectors and chords that are their mixtures (i.e., spurious states), we use tools borrowed from statistical mechanics of disordered systems to investigate this task, obtaining the conditions over the model control-parameters such that pattern disentanglement is successfully executed.
△ Less
Submitted 15 October, 2025; v1 submitted 28 January, 2025;
originally announced January 2025.
-
The thermodynamic limit in mean field neural networks
Authors:
Elena Agliari,
Adriano Barra,
Pierluigi Bianco,
Alberto Fachechi,
Diego Pallara
Abstract:
In the last five decades, mean-field neural-networks have played a crucial role in modelling associative memories and, in particular, the Hopfield model has been extensively studied using tools borrowed from the statistical mechanics of spin glasses. However, achieving mathematical control of the infinite-volume limit of the model's free-energy has remained elusive, as the standard treatments deve…
▽ More
In the last five decades, mean-field neural-networks have played a crucial role in modelling associative memories and, in particular, the Hopfield model has been extensively studied using tools borrowed from the statistical mechanics of spin glasses. However, achieving mathematical control of the infinite-volume limit of the model's free-energy has remained elusive, as the standard treatments developed for spin-glasses have proven unfeasible. Here we address this long-standing problem by proving that a measure-concentration assumption for the order parameters of the theory is sufficient for the existence of the asymptotic limit of the model's free energy. The proof leverages the equivalence between the free energy of the Hopfield model and a linear combination of the free energies of a hard and a soft spin-glass, whose thermodynamic limits are rigorously known. Our work focuses on the replica-symmetry level of description (for which we recover the explicit expression of the free-energy found in the eighties via heuristic methods), yet, our scheme is expected to work also under (at least) the first step of replica symmetry breaking.
△ Less
Submitted 16 September, 2024;
originally announced September 2024.
-
Generalized hetero-associative neural networks
Authors:
Elena Agliari,
Andrea Alessandrelli,
Adriano Barra,
Martino Salomone Centonze,
Federico Ricci-Tersenghi
Abstract:
Auto-associative neural networks (e.g., the Hopfield model implementing the standard Hebbian prescription) serve as a foundational framework for pattern recognition and associative memory in statistical mechanics. However, their hetero-associative counterparts, though less explored, exhibit even richer computational capabilities. In this work, we examine a straightforward extension of Kosko's Bidi…
▽ More
Auto-associative neural networks (e.g., the Hopfield model implementing the standard Hebbian prescription) serve as a foundational framework for pattern recognition and associative memory in statistical mechanics. However, their hetero-associative counterparts, though less explored, exhibit even richer computational capabilities. In this work, we examine a straightforward extension of Kosko's Bidirectional Associative Memory (BAM), introducing a Three-directional Associative Memory (TAM), that is a tripartite neural network equipped with generalized Hebbian weights. Through both analytical approaches (using replica-symmetric statistical mechanics) and computational methods (via Monte Carlo simulations), we derive phase diagrams within the space of control parameters, revealing a region where the network can successfully perform pattern recognition as well as other tasks tasks. In particular, it can achieve pattern disentanglement, namely, when presented with a mixture of patterns, the network can recover the original patterns. Furthermore, the system is capable of retrieving Markovian sequences of patterns and performing generalized frequency modulation.
△ Less
Submitted 21 October, 2024; v1 submitted 12 September, 2024;
originally announced September 2024.
-
Chemical tuning of quantum spin-electric coupling in molecular nanomagnets
Authors:
Mikhail V. Vaganov,
Nicolas Suaud,
Francois Lambert,
Benjamin Cahier,
Christian Herrero,
Regis Guillot,
Anne-Laure Barra,
Nathalie Guihery,
Talal Mallah,
Arzhang Ardavan,
Junjie Liu
Abstract:
Controlling quantum spins using electric rather than magnetic fields promises significant architectural advantages for developing quantum technologies. In this context, spins in molecular nanomagnets offer tunability of spin-electric couplings (SEC) by rational chemical design. Here we demonstrate systematic control of SECs in a family of Mn(II)-containing molecules via chemical engineering. The t…
▽ More
Controlling quantum spins using electric rather than magnetic fields promises significant architectural advantages for developing quantum technologies. In this context, spins in molecular nanomagnets offer tunability of spin-electric couplings (SEC) by rational chemical design. Here we demonstrate systematic control of SECs in a family of Mn(II)-containing molecules via chemical engineering. The trigonal bipyramidal (tbp) molecular structure with C3 symmetry leads to a significant molecular electric dipole moment that is directly connected to its magnetic anisotropy. The interplay between these two features gives rise to significant experimentally observed SECs, which can be rationalised by wavefunction theoretical calculations. Our findings guide strategies for the development of electrically controllable molecular spin qubits for quantum technologies.
△ Less
Submitted 3 September, 2024;
originally announced September 2024.
-
Guerra interpolation for place cells
Authors:
Martino Salomone Centonze,
Alessandro Treves,
Elena Agliari,
Adriano Barra
Abstract:
Pyramidal cells that emit spikes when the animal is at specific locations of the environment are known as "place cells": these neurons are thought to provide an internal representation of space via "cognitive maps". Here, we consider the Battaglia-Treves neural network model for cognitive map storage and reconstruction, instantiated with McCulloch & Pitts binary neurons. To quantify the informatio…
▽ More
Pyramidal cells that emit spikes when the animal is at specific locations of the environment are known as "place cells": these neurons are thought to provide an internal representation of space via "cognitive maps". Here, we consider the Battaglia-Treves neural network model for cognitive map storage and reconstruction, instantiated with McCulloch & Pitts binary neurons. To quantify the information processing capabilities of these networks, we exploit spin-glass techniques based on Guerra's interpolation: in the low-storage regime (i.e., when the number of stored maps scales sub-linearly with the network size and the order parameters self-average around their means) we obtain an exact phase diagram in the noise vs inhibition strength plane (in agreement with previous findings) by adapting the Hamilton-Jacobi PDE-approach. Conversely, in the high-storage regime, we find that -- for mild inhibition and not too high noise -- memorization and retrieval of an extensive number of spatial maps is indeed possible, since the maximal storage capacity is shown to be strictly positive. These results, holding under the replica-symmetry assumption, are obtained by adapting the standard interpolation based on stochastic stability and are further corroborated by Monte Carlo simulations (and replica-trick outcomes for the sake of completeness). Finally, by relying upon an interpretation in terms of hidden units, in the last part of the work, we adapt the Battaglia-Treves model to cope with more general frameworks, such as bats flying in long tunnels.
△ Less
Submitted 25 August, 2024;
originally announced August 2024.
-
Hebbian Learning from First Principles
Authors:
Linda Albanese,
Adriano Barra,
Pierluigi Bianco,
Fabrizio Durante,
Diego Pallara
Abstract:
Recently, the original storage prescription for the Hopfield model of neural networks -- as well as for its dense generalizations -- has been turned into a genuine Hebbian learning rule by postulating the expression of its Hamiltonian for both the supervised and unsupervised protocols. In these notes, first, we obtain these explicit expressions by relying upon maximum entropy extremization à la Ja…
▽ More
Recently, the original storage prescription for the Hopfield model of neural networks -- as well as for its dense generalizations -- has been turned into a genuine Hebbian learning rule by postulating the expression of its Hamiltonian for both the supervised and unsupervised protocols. In these notes, first, we obtain these explicit expressions by relying upon maximum entropy extremization à la Jaynes. Beyond providing a formal derivation of these recipes for Hebbian learning, this construction also highlights how Lagrangian constraints within entropy extremization force network's outcomes on neural correlations: these try to mimic the empirical counterparts hidden in the datasets provided to the network for its training and, the denser the network, the longer the correlations that it is able to capture. Next, we prove that, in the big data limit, whatever the presence of a teacher (or its lacking), not only these Hebbian learning rules converge to the original storage prescription of the Hopfield model but also their related free energies (and, thus, the statistical mechanical picture provided by Amit, Gutfreund and Sompolinsky is fully recovered). As a sideline, we show mathematical equivalence among standard Cost functions (Hamiltonian), preferred in Statistical Mechanical jargon, and quadratic Loss Functions, preferred in Machine Learning terminology. Remarks on the exponential Hopfield model (as the limit of dense networks with diverging density) and semi-supervised protocols are also provided.
△ Less
Submitted 3 October, 2024; v1 submitted 13 January, 2024;
originally announced January 2024.
-
Unsupervised and Supervised learning by Dense Associative Memory under replica symmetry breaking
Authors:
Linda Albanese,
Andrea Alessandrelli,
Alessia Annibale,
Adriano Barra
Abstract:
Statistical mechanics of spin glasses is one of the main strands toward a comprehension of information processing by neural networks and learning machines. Tackling this approach, at the fairly standard replica symmetric level of description, recently Hebbian attractor networks with multi-node interactions (often called Dense Associative Memories) have been shown to outperform their classical pair…
▽ More
Statistical mechanics of spin glasses is one of the main strands toward a comprehension of information processing by neural networks and learning machines. Tackling this approach, at the fairly standard replica symmetric level of description, recently Hebbian attractor networks with multi-node interactions (often called Dense Associative Memories) have been shown to outperform their classical pairwise counterparts in a number of tasks, from their robustness against adversarial attacks and their capability to work with prohibitively weak signals to their supra-linear storage capacities. Focusing on mathematical techniques more than computational aspects, in this paper we relax the replica symmetric assumption and we derive the one-step broken-replica-symmetry picture of supervised and unsupervised learning protocols for these Dense Associative Memories: a phase diagram in the space of the control parameters is achieved, independently, both via the Parisi's hierarchy within then replica trick as well as via the Guerra's telescope within the broken-replica interpolation. Further, an explicit analytical investigation is provided to deepen both the big-data and ground state limits of these networks as well as a proof that replica symmetry breaking does not alter the thresholds for learning and slightly increases the maximal storage capacity. Finally the De Almeida and Thouless line, depicting the onset of instability of a replica symmetric description, is also analytically derived highlighting how, crossed this boundary, the broken replica description should be preferred.
△ Less
Submitted 15 December, 2023;
originally announced December 2023.
-
Inverse modeling of time-delayed interactions via the dynamic-entropy formalism
Authors:
Elena Agliari,
Francesco Alemanno,
Adriano Barra,
Michele Castellana,
Daniele Lotito,
Matthieu Piel
Abstract:
Although instantaneous interactions are unphysical, a large variety of maximum entropy statistical inference methods match the model-inferred and the empirically-measured equal-time correlation functions. Focusing on collective motion of active units, this constraint is reasonable when the interaction timescale is much faster than that of the interacting units, as in starling flocks, yet it fails…
▽ More
Although instantaneous interactions are unphysical, a large variety of maximum entropy statistical inference methods match the model-inferred and the empirically-measured equal-time correlation functions. Focusing on collective motion of active units, this constraint is reasonable when the interaction timescale is much faster than that of the interacting units, as in starling flocks, yet it fails in a number of counter examples, as in leukocyte coordination (where signalling proteins diffuse among two cells). Here, we relax this assumption and develop a path integral approach to maximum-entropy framework, which includes delay in signalling. Our method is able to infer the strength of couplings and fields, but also the time required by the couplings to completely transfer information among the units. We demonstrate the validity of our approach providing excellent results on synthetic datasets of non-Markovian trajectories generated by the Heisenberg-Kuramoto and Vicsek models equipped with delayed interactions. As a proof of concept, we also apply the method to experiments on dendritic migration, where matching equal-time correlations results in a significant information loss.
△ Less
Submitted 10 July, 2024; v1 submitted 3 September, 2023;
originally announced September 2023.
-
Parallel Learning by Multitasking Neural Networks
Authors:
Elena Agliari,
Andrea Alessandrelli,
Adriano Barra,
Federico Ricci-Tersenghi
Abstract:
A modern challenge of Artificial Intelligence is learning multiple patterns at once (i.e.parallel learning). While this can not be accomplished by standard Hebbian associative neural networks, in this paper we show how the Multitasking Hebbian Network (a variation on theme of the Hopfield model working on sparse data-sets) is naturally able to perform this complex task. We focus on systems process…
▽ More
A modern challenge of Artificial Intelligence is learning multiple patterns at once (i.e.parallel learning). While this can not be accomplished by standard Hebbian associative neural networks, in this paper we show how the Multitasking Hebbian Network (a variation on theme of the Hopfield model working on sparse data-sets) is naturally able to perform this complex task. We focus on systems processing in parallel a finite (up to logarithmic growth in the size of the network) amount of patterns, mirroring the low-storage level of standard associative neural networks at work with pattern recognition. For mild dilution in the patterns, the network handles them hierarchically, distributing the amplitudes of their signals as power-laws w.r.t. their information content (hierarchical regime), while, for strong dilution, all the signals pertaining to all the patterns are raised with the same strength (parallel regime). Further, confined to the low-storage setting (i.e., far from the spin glass limit), the presence of a teacher neither alters the multitasking performances nor changes the thresholds for learning: the latter are the same whatever the training protocol is supervised or unsupervised. Results obtained through statistical mechanics, signal-to-noise technique and Monte Carlo simulations are overall in perfect agreement and carry interesting insights on multiple learning at once: for instance, whenever the cost-function of the model is minimized in parallel on several patterns (in its description via Statistical Mechanics), the same happens to the standard sum-squared error Loss function (typically used in Machine Learning).
△ Less
Submitted 8 August, 2023;
originally announced August 2023.
-
Statistical Mechanics of Learning via Reverberation in Bidirectional Associative Memories
Authors:
Martino Salomone Centonze,
Ido Kanter,
Adriano Barra
Abstract:
We study bi-directional associative neural networks that, exposed to noisy examples of an extensive number of random archetypes, learn the latter (with or without the presence of a teacher) when the supplied information is enough: in this setting, learning is heteroassociative -- involving couples of patterns -- and it is achieved by reverberating the information depicted from the examples through…
▽ More
We study bi-directional associative neural networks that, exposed to noisy examples of an extensive number of random archetypes, learn the latter (with or without the presence of a teacher) when the supplied information is enough: in this setting, learning is heteroassociative -- involving couples of patterns -- and it is achieved by reverberating the information depicted from the examples through the layers of the network. By adapting Guerra's interpolation technique, we provide a full statistical mechanical picture of supervised and unsupervised learning processes (at the replica symmetric level of description) obtaining analytically phase diagrams, thresholds for learning, a picture of the ground-state in plain agreement with Monte Carlo simulations and signal-to-noise outcomes. In the large dataset limit, the Kosko storage prescription as well as its statistical mechanical picture provided by Kurchan, Peliti, and Saber in the eighties is fully recovered. Computational advantages in dealing with information reverberation, rather than storage, are discussed for natural test cases. In particular, we show how this network admits an integral representation in terms of two coupled restricted Boltzmann machines, whose hidden layers are entirely built of by grand-mother neurons, to prove that by coupling solely these grand-mother neurons we can correlate the patterns they are related to: it is thus possible to recover Pavlov's Classical Conditioning by adding just one synapse among the correct grand-mother neurons (hence saving an extensive number of these links for further information storage w.r.t. the classical autoassociative setting).
△ Less
Submitted 17 July, 2023;
originally announced July 2023.
-
Ultrametric identities in glassy models of Natural Evolution
Authors:
Elena Agliari,
Francesco Alemanno,
Miriam Aquaro,
Adriano Barra
Abstract:
Spin-glasses constitute a well-grounded framework for evolutionary models. Of particular interest for (some of) these models is the lack of self-averaging of their order parameters (e.g. the Hamming distance between the genomes of two individuals), even in asymptotic limits, much as like the behavior of the overlap between the configurations of two replica in mean-field spin-glasses. In the latter…
▽ More
Spin-glasses constitute a well-grounded framework for evolutionary models. Of particular interest for (some of) these models is the lack of self-averaging of their order parameters (e.g. the Hamming distance between the genomes of two individuals), even in asymptotic limits, much as like the behavior of the overlap between the configurations of two replica in mean-field spin-glasses. In the latter, this lack of self-averaging is related to peculiar fluctuations of the overlap, known as Ghirlanda-Guerra identities and Aizenman-Contucci polynomials, that cover a pivotal role in describing the ultrametric structure of the spin-glass landscape. As for evolutionary models, such identities may therefore be related to a taxonomic classification of individuals, yet a full investigation on their validity is missing. In this paper, we study ultrametric identities in simple cases where solely random mutations take place, while selective pressure is absent, namely in {\em flat landscape} models. In particular, we study three paradigmatic models in this setting: the {\em one parent model} (which, by construction, is ultrametric at the level of single individuals), the {\em homogeneous population model} (which is replica symmetric), and the {\em species formation model} (where a broken-replica scenario emerges at the level of species). We find analytical and numerical evidence that in the first and in the third model nor the Ghirlanda-Guerra neither the Aizenman-Contucci constraints hold, rather a new class of ultrametric identities is satisfied; in the second model all these constraints hold trivially. Very preliminary results on a real biological human genome derived by {\em The 1000 Genome Project Consortium} and on two artificial human genomes (generated by two different types neural networks) seem in better agreement with these new identities rather than the classic ones.
△ Less
Submitted 23 June, 2023;
originally announced June 2023.
-
About the de Almeida-Thouless line in neural networks
Authors:
Linda Albanese,
Andrea Alessandrelli,
Adriano Barra,
Alessia Annibale
Abstract:
In this work we present a rigorous and straightforward method to detect the onset of the instability of replica-symmetric theories in information processing systems, which does not require a full replica analysis as in the method originally proposed by de Almeida and Thouless for spin glasses. The method is based on an expansion of the free-energy obtained within one-step of replica symmetry break…
▽ More
In this work we present a rigorous and straightforward method to detect the onset of the instability of replica-symmetric theories in information processing systems, which does not require a full replica analysis as in the method originally proposed by de Almeida and Thouless for spin glasses. The method is based on an expansion of the free-energy obtained within one-step of replica symmetry breaking (RSB) around the RS value. As such, it requires solely continuity and differentiability of the free-energy and it is robust to be applied broadly to systems with quenched disorder. We apply the method to the Hopfield model and to neural networks with multi-node Hebbian interactions, as case studies. In the appendices we test the method on the Sherrington-Kirkpatrick and the Ising P-spin models, recovering the AT lines known in the literature for these models, as a special limit, which corresponds to assuming that the transition from the RS to the RSB phase can be obtained by varying continuously the order parameters. Our method provides a generalization of the AT approach, which does not rely on this limit and can be applied to systems with discontinuous phase transitions, as we show explicitly for the spherical P-spin model, recovering the known RS instability line.
△ Less
Submitted 12 November, 2023; v1 submitted 11 March, 2023;
originally announced March 2023.
-
Dense Hebbian neural networks: a replica symmetric picture of supervised learning
Authors:
Elena Agliari,
Linda Albanese,
Francesco Alemanno,
Andrea Alessandrelli,
Adriano Barra,
Fosca Giannotti,
Daniele Lotito,
Dino Pedreschi
Abstract:
We consider dense, associative neural-networks trained by a teacher (i.e., with supervision) and we investigate their computational capabilities analytically, via statistical-mechanics of spin glasses, and numerically, via Monte Carlo simulations. In particular, we obtain a phase diagram summarizing their performance as a function of the control parameters such as quality and quantity of the train…
▽ More
We consider dense, associative neural-networks trained by a teacher (i.e., with supervision) and we investigate their computational capabilities analytically, via statistical-mechanics of spin glasses, and numerically, via Monte Carlo simulations. In particular, we obtain a phase diagram summarizing their performance as a function of the control parameters such as quality and quantity of the training dataset, network storage and noise, that is valid in the limit of large network size and structureless datasets: these networks may work in a ultra-storage regime (where they can handle a huge amount of patterns, if compared with shallow neural networks) or in a ultra-detection regime (where they can perform pattern recognition at prohibitive signal-to-noise ratios, if compared with shallow neural networks). Guided by the random theory as a reference framework, we also test numerically learning, storing and retrieval capabilities shown by these networks on structured datasets as MNist and Fashion MNist. As technical remarks, from the analytic side, we implement large deviations and stability analysis within Guerra's interpolation to tackle the not-Gaussian distributions involved in the post-synaptic potentials while, from the computational counterpart, we insert Plefka approximation in the Monte Carlo scheme, to speed up the evaluation of the synaptic tensors, overall obtaining a novel and broad approach to investigate supervised learning in neural networks, beyond the shallow limit, in general.
△ Less
Submitted 2 July, 2023; v1 submitted 25 November, 2022;
originally announced December 2022.
-
Microscopic parameters of the van der Waals CrSBr antiferromagnet from microwave absorption experiments
Authors:
C. W. Cho,
A. Pawbake,
N. Aubergier,
A. L. Barra,
K. Mosina,
Z. Sofer,
M. E. Zhitomirsky,
C. Faugeras,
B. A. Piot
Abstract:
Microwave absorption experiments employing a phase-sensitive external resistive detection are performed for a topical van der Waals antiferromagnet CrSBr. The field dependence of two resonance modes is measured in an applied field parallel to the three principal crystallographic directions, revealing anisotropies and magnetic transitions in this material. To account for the observed results, we fo…
▽ More
Microwave absorption experiments employing a phase-sensitive external resistive detection are performed for a topical van der Waals antiferromagnet CrSBr. The field dependence of two resonance modes is measured in an applied field parallel to the three principal crystallographic directions, revealing anisotropies and magnetic transitions in this material. To account for the observed results, we formulate a microscopic spin model with a bi-axial single-ion anisotropy and inter-plane exchange. Theoretical calculations give an excellent description of full magnon spectra enabling us to precisely determine microscopic interaction parameters for CrSBr.
△ Less
Submitted 25 November, 2022;
originally announced November 2022.
-
Dense Hebbian neural networks: a replica symmetric picture of unsupervised learning
Authors:
Elena Agliari,
Linda Albanese,
Francesco Alemanno,
Andrea Alessandrelli,
Adriano Barra,
Fosca Giannotti,
Daniele Lotito,
Dino Pedreschi
Abstract:
We consider dense, associative neural-networks trained with no supervision and we investigate their computational capabilities analytically, via a statistical-mechanics approach, and numerically, via Monte Carlo simulations. In particular, we obtain a phase diagram summarizing their performance as a function of the control parameters such as the quality and quantity of the training dataset and the…
▽ More
We consider dense, associative neural-networks trained with no supervision and we investigate their computational capabilities analytically, via a statistical-mechanics approach, and numerically, via Monte Carlo simulations. In particular, we obtain a phase diagram summarizing their performance as a function of the control parameters such as the quality and quantity of the training dataset and the network storage, valid in the limit of large network size and structureless datasets. Moreover, we establish a bridge between macroscopic observables standardly used in statistical mechanics and loss functions typically used in the machine learning. As technical remarks, from the analytic side, we implement large deviations and stability analysis within Guerra's interpolation to tackle the not-Gaussian distributions involved in the post-synaptic potentials while, from the computational counterpart, we insert Plefka approximation in the Monte Carlo scheme, to speed up the evaluation of the synaptic tensors, overall obtaining a novel and broad approach to investigate neural networks in general.
△ Less
Submitted 2 July, 2023; v1 submitted 25 November, 2022;
originally announced November 2022.
-
Thermodynamics of bidirectional associative memories
Authors:
Adriano Barra,
Giovanni Catania,
Aurélien Decelle,
Beatriz Seoane
Abstract:
In this paper we investigate the equilibrium properties of bidirectional associative memories (BAMs). Introduced by Kosko in 1988 as a generalization of the Hopfield model to a bipartite structure, the simplest architecture is defined by two layers of neurons, with synaptic connections only between units of different layers: even without internal connections within each layer, information storage…
▽ More
In this paper we investigate the equilibrium properties of bidirectional associative memories (BAMs). Introduced by Kosko in 1988 as a generalization of the Hopfield model to a bipartite structure, the simplest architecture is defined by two layers of neurons, with synaptic connections only between units of different layers: even without internal connections within each layer, information storage and retrieval are still possible through the reverberation of neural activities passing from one layer to another. We characterize the computational capabilities of a stochastic extension of this model in the thermodynamic limit, by applying rigorous techniques from statistical physics. A detailed picture of the phase diagram at the replica symmetric level is provided, both at finite temperature and in the noiseless regimes. Also for the latter, the critical load is further investigated up to one step of replica symmetry breaking. An analytical and numerical inspection of the transition curves (namely critical lines splitting the various modes of operation of the machine) is carried out as the control parameters - noise, load and asymmetry between the two layer sizes - are tuned. In particular, with a finite asymmetry between the two layers, it is shown how the BAM can store information more efficiently than the Hopfield model by requiring less parameters to encode a fixed number of patterns. Comparisons are made with numerical simulations of neural dynamics. Finally, a low-load analysis is carried out to explain the retrieval mechanism in the BAM by analogy with two interacting Hopfield models. A potential equivalence with two coupled Restricted Boltmzann Machines is also discussed.
△ Less
Submitted 27 March, 2023; v1 submitted 17 November, 2022;
originally announced November 2022.
-
Pavlov Learning Machines
Authors:
Elena Agliari,
Miriam Aquaro,
Adriano Barra,
Alberto Fachechi,
Chiara Marullo
Abstract:
As well known, Hebb's learning traces its origin in Pavlov's Classical Conditioning, however, while the former has been extensively modelled in the past decades (e.g., by Hopfield model and countless variations on theme), as for the latter modelling has remained largely unaddressed so far; further, a bridge between these two pillars is totally lacking. The main difficulty towards this goal lays in…
▽ More
As well known, Hebb's learning traces its origin in Pavlov's Classical Conditioning, however, while the former has been extensively modelled in the past decades (e.g., by Hopfield model and countless variations on theme), as for the latter modelling has remained largely unaddressed so far; further, a bridge between these two pillars is totally lacking. The main difficulty towards this goal lays in the intrinsically different scales of the information involved: Pavlov's theory is about correlations among \emph{concepts} that are (dynamically) stored in the synaptic matrix as exemplified by the celebrated experiment starring a dog and a ring bell; conversely, Hebb's theory is about correlations among pairs of adjacent neurons as summarized by the famous statement {\em neurons that fire together wire together}. In this paper we rely on stochastic-process theory and model neural and synaptic dynamics via Langevin equations, to prove that -- as long as we keep neurons' and synapses' timescales largely split -- Pavlov mechanism spontaneously takes place and ultimately gives rise to synaptic weights that recover the Hebbian kernel.
△ Less
Submitted 2 July, 2022;
originally announced July 2022.
-
Recurrent neural networks that generalize from examples and optimize by dreaming
Authors:
Miriam Aquaro,
Francesco Alemanno,
Ido Kanter,
Fabrizio Durante,
Elena Agliari,
Adriano Barra
Abstract:
The gap between the huge volumes of data needed to train artificial neural networks and the relatively small amount of data needed by their biological counterparts is a central puzzle in machine learning. Here, inspired by biological information-processing, we introduce a generalized Hopfield network where pairwise couplings between neurons are built according to Hebb's prescription for on-line le…
▽ More
The gap between the huge volumes of data needed to train artificial neural networks and the relatively small amount of data needed by their biological counterparts is a central puzzle in machine learning. Here, inspired by biological information-processing, we introduce a generalized Hopfield network where pairwise couplings between neurons are built according to Hebb's prescription for on-line learning and allow also for (suitably stylized) off-line sleeping mechanisms. Moreover, in order to retain a learning framework, here the patterns are not assumed to be available, instead, we let the network experience solely a dataset made of a sample of noisy examples for each pattern. We analyze the model by statistical-mechanics tools and we obtain a quantitative picture of its capabilities as functions of its control parameters: the resulting network is an associative memory for pattern recognition that learns from examples on-line, generalizes and optimizes its storage capacity by off-line sleeping. Remarkably, the sleeping mechanisms always significantly reduce (up to $\approx 90\%$) the dataset size required to correctly generalize, further, there are memory loads that are prohibitive to Hebbian networks without sleeping (no matter the size and quality of the provided examples), but that are easily handled by the present "rested" neural networks.
△ Less
Submitted 17 April, 2022;
originally announced April 2022.
-
Supervised Hebbian Learning
Authors:
Francesco Alemanno,
Miriam Aquaro,
Ido Kanter,
Adriano Barra,
Elena Agliari
Abstract:
In neural network's Literature, Hebbian learning traditionally refers to the procedure by which the Hopfield model and its generalizations store archetypes (i.e., definite patterns that are experienced just once to form the synaptic matrix). However, the term "Learning" in Machine Learning refers to the ability of the machine to extract features from the supplied dataset (e.g., made of blurred exa…
▽ More
In neural network's Literature, Hebbian learning traditionally refers to the procedure by which the Hopfield model and its generalizations store archetypes (i.e., definite patterns that are experienced just once to form the synaptic matrix). However, the term "Learning" in Machine Learning refers to the ability of the machine to extract features from the supplied dataset (e.g., made of blurred examples of these archetypes), in order to make its own representation of the unavailable archetypes. Here, given a sample of examples, we define a supervised learning protocol by which the Hopfield network can infer the archetypes, and we detect the correct control parameters (including size and quality of the dataset) to depict a phase diagram for the system performance. We also prove that, for structureless datasets, the Hopfield model equipped with this supervised learning rule is equivalent to a restricted Boltzmann machine and this suggests an optimal and interpretable training routine. Finally, this approach is generalized to structured datasets: we highlight a quasi-ultrametric organization (reminiscent of replica-symmetry-breaking) in the analyzed datasets and, consequently, we introduce an additional "replica hidden layer" for its (partial) disentanglement, which is shown to improve MNIST classification from 75% to 95%, and to offer a new perspective on deep architectures.
△ Less
Submitted 7 September, 2022; v1 submitted 2 March, 2022;
originally announced March 2022.
-
Anisotropic long-range spin transport in canted antiferromagnetic orthoferrite YFeO$_3$
Authors:
Shubhankar Das,
A. Ross,
X. X. Ma,
S. Becker,
C. Schmitt,
F. van Duijn,
F. Fuhrmann,
M. -A. Syskaki,
U. Ebels,
V. Baltz,
A. -L. Barra,
H. Y. Chen,
G. Jakob,
S. X. Cao,
J. Sinova,
O. Gomonay,
R. Lebrun,
M. Kläui
Abstract:
In antiferromagnets, the efficient propagation of spin-waves has until now only been observed in the insulating antiferromagnet hematite, where circularly (or a superposition of pairs of linearly) polarized spin-waves propagate over long distances. Here, we report long-distance spin-transport in the antiferromagnetic orthoferrite YFeO$_3$, where a different transport mechanism is enabled by the co…
▽ More
In antiferromagnets, the efficient propagation of spin-waves has until now only been observed in the insulating antiferromagnet hematite, where circularly (or a superposition of pairs of linearly) polarized spin-waves propagate over long distances. Here, we report long-distance spin-transport in the antiferromagnetic orthoferrite YFeO$_3$, where a different transport mechanism is enabled by the combined presence of the Dzyaloshinskii-Moriya interaction and externally applied fields. The magnon decay length is shown to exceed hundreds of nano-meters, in line with resonance measurements that highlight the low magnetic damping. We observe a strong anisotropy in the magnon decay lengths that we can attribute to the role of the magnon group velocity in the propagation of spin-waves in antiferromagnets. This unique mode of transport identified in YFeO$_3$ opens up the possibility of a large and technologically relevant class of materials, i.e., canted antiferromagnets, for long-distance spin transport.
△ Less
Submitted 11 December, 2021;
originally announced December 2021.
-
Replica symmetry breaking in dense neural networks
Authors:
Linda Albanese,
Francesco Alemanno,
Andrea Alessandrelli,
Adriano Barra
Abstract:
Understanding the glassy nature of neural networks is pivotal both for theoretical and computational advances in Machine Learning and Theoretical Artificial Intelligence. Keeping the focus on dense associative Hebbian neural networks, the purpose of this paper is two-fold: at first we develop rigorous mathematical approaches to address properly a statistical mechanical picture of the phenomenon of…
▽ More
Understanding the glassy nature of neural networks is pivotal both for theoretical and computational advances in Machine Learning and Theoretical Artificial Intelligence. Keeping the focus on dense associative Hebbian neural networks, the purpose of this paper is two-fold: at first we develop rigorous mathematical approaches to address properly a statistical mechanical picture of the phenomenon of {\em replica symmetry breaking} (RSB) in these networks, then -- deepening results stemmed via these routes -- we aim to inspect the {\em glassiness} that they hide. In particular, regarding the methodology, we provide two techniques: the former is an adaptation of the transport PDE to the case, while the latter is an extension of Guerra's interpolation breakthrough. Beyond coherence among the results, either in replica symmetric and in the one-step replica symmetry breaking level of description, we prove the Gardner's picture and we identify the maximal storage capacity by a ground-state analysis in the Baldi-Venkatesh high-storage regime.
In the second part of the paper we investigate the glassy structure of these networks: in contrast with the replica symmetric scenario (RS), RSB actually stabilizes the spin-glass phase. We report huge differences w.r.t. the standard pairwise Hopfield limit: in particular, it is known that it is possible to express the free energy of the Hopfield neural network as a linear combination of the free energies of an hard spin glass (i.e. the Sherrington-Kirkpatrick model) and a soft spin glass (the Gaussian or "spherical" model). This is no longer true when interactions are more than pairwise (whatever the level of description, RS or RSB): for dense networks solely the free energy of the hard spin glass survives, proving a huge diversity in the underlying glassiness of associative neural networks.
△ Less
Submitted 25 November, 2021;
originally announced November 2021.
-
The emergence of a concept in shallow neural networks
Authors:
Elena Agliari,
Francesco Alemanno,
Adriano Barra,
Giordano De Marzo
Abstract:
We consider restricted Boltzmann machine (RBMs) trained over an unstructured dataset made of blurred copies of definite but unavailable ``archetypes'' and we show that there exists a critical sample size beyond which the RBM can learn archetypes, namely the machine can successfully play as a generative model or as a classifier, according to the operational routine. In general, assessing a critical…
▽ More
We consider restricted Boltzmann machine (RBMs) trained over an unstructured dataset made of blurred copies of definite but unavailable ``archetypes'' and we show that there exists a critical sample size beyond which the RBM can learn archetypes, namely the machine can successfully play as a generative model or as a classifier, according to the operational routine. In general, assessing a critical sample size (possibly in relation to the quality of the dataset) is still an open problem in machine learning. Here, restricting to the random theory, where shallow networks suffice and the grand-mother cell scenario is correct, we leverage the formal equivalence between RBMs and Hopfield networks, to obtain a phase diagram for both the neural architectures which highlights regions, in the space of the control parameters (i.e., number of archetypes, number of neurons, size and quality of the training set), where learning can be accomplished. Our investigations are led by analytical methods based on the statistical-mechanics of disordered systems and results are further corroborated by extensive Monte Carlo simulations.
△ Less
Submitted 1 September, 2021;
originally announced September 2021.
-
Robust magnetic anisotropy of a monolayer of hexacoordinate Fe( ii ) complexes assembled on Cu(111)
Authors:
Massine Kelai,
Benjamin Cahier,
Mihail Atanasov,
Frank Neese,
Yongfeng Tong,
Luqiong Zhang,
Amandine Bellec,
Olga Iasco,
Eric Rivière,
Régis Guillot,
Cyril Chacon,
Yann Girard,
Jérôme Lagoute,
Sylvie Rousset,
Vincent Repain,
Edwige Otero,
Marie-Anne Arrio,
Philippe Sainctavit,
Anne-Laure Barra,
Marie-Laure Boillot,
Talal Mallah
Abstract:
The tris pyrazolyl borate ligand imposes a rigid scaffold around Fe( ii ) ensuring a robust magnetic anisotropy when the molecules assembled as monolayers suffer from the dissymmetric environment of the substrate/vacuum interface.
The tris pyrazolyl borate ligand imposes a rigid scaffold around Fe( ii ) ensuring a robust magnetic anisotropy when the molecules assembled as monolayers suffer from the dissymmetric environment of the substrate/vacuum interface.
△ Less
Submitted 10 May, 2021;
originally announced May 2021.
-
Effective strain manipulation of the antiferromagnetic state of polycrystalline NiO
Authors:
A. Barra,
A. Ross,
O. Gomonay,
L. Baldrati,
A. Chavez,
R. Lebrun,
J. D. Schneider,
P. Shirazi,
Q. Wang,
J. Sinova,
G. P. Carman,
M. Kläui
Abstract:
As a candidate material for applications such as magnetic memory, polycrystalline antiferromagnets offer the same robustness to external magnetic fields, THz spin dynamics, and lack of stray field as their single crystalline counterparts, but without the limitation of epitaxial growth and lattice matched substrates. Here, we first report the detection of the average Neel vector orientiation in pol…
▽ More
As a candidate material for applications such as magnetic memory, polycrystalline antiferromagnets offer the same robustness to external magnetic fields, THz spin dynamics, and lack of stray field as their single crystalline counterparts, but without the limitation of epitaxial growth and lattice matched substrates. Here, we first report the detection of the average Neel vector orientiation in polycrystalline NiO via spin Hall magnetoresistance (SMR). Secondly, by applying strain through a piezo-electric substrate, we reduce the critical magnetic field required to reach a saturation of the SMR signal, indicating a change of the anisotropy. Our results are consistent with polycrystalline NiO exhibiting a positive sign of the in-plane magnetostriction. This method of anisotropy-tuning offers an energy efficient, on-chip alternative to manipulate a polycrystalline antiferromagnets magnetic state.
△ Less
Submitted 24 March, 2021;
originally announced March 2021.
-
Chemical tuning of spin clock transitions in molecular monomers based on nuclear spin-free Ni(II)
Authors:
Marcos Rubín-Osanz,
François Lambert,
Feng Shao,
Eric Rivière,
Régis Guillot,
Nicolas Suaud,
Nathalie Guihéry,
David Zueco,
Anne-Laure Barra,
Talal Mallah,
Fernando Luis
Abstract:
We report the existence of a sizeable quantum tunnelling splitting between the two lowest electronic spin levels of mononuclear Ni complexes. The level anti-crossing, or magnetic clock transition, associated with this gap has been directly monitored by heat capacity experiments. The comparison of these results with those obtained for a Co derivative, for which tunnelling is forbidden by symmetry,…
▽ More
We report the existence of a sizeable quantum tunnelling splitting between the two lowest electronic spin levels of mononuclear Ni complexes. The level anti-crossing, or magnetic clock transition, associated with this gap has been directly monitored by heat capacity experiments. The comparison of these results with those obtained for a Co derivative, for which tunnelling is forbidden by symmetry, shows that the clock transition leads to an effective suppression of intermolecular spin-spin interactions. In addition, we show that the quantum tunnelling splitting admits a chemical tuning via the modification of the ligand shell that determines the crystal field and the magnetic anisotropy. These properties are crucial to realize model spin qubits that combine the necessary resilience against decoherence, a proper interfacing with other qubits and with the control circuitry and the ability to initialize them by cooling.
△ Less
Submitted 4 March, 2021;
originally announced March 2021.
-
Long-distance spin-transport across the Morin phase transition up to room temperature in ultra-low damping single crystals of the antiferromagnet α-Fe2O3
Authors:
Romain Lebrun,
Andrew Ross,
Olena Gomonay,
Vincent Baltz,
Ursula Ebels,
Anne Laure Barra,
Alireza Qaiumzadeh,
Arne Brataas,
Jairo Sinova,
Mathias Kläui
Abstract:
Antiferromagnetic materials can host spin-waves with polarizations ranging from circular to linear depending on their magnetic anisotropies. Until now, only easy-axis anisotropy antiferromagnets with circularly polarized spin-waves were reported to carry spin-information over long distances of micrometers. In this article, we report long-distance spin-transport in the easy-plane canted antiferroma…
▽ More
Antiferromagnetic materials can host spin-waves with polarizations ranging from circular to linear depending on their magnetic anisotropies. Until now, only easy-axis anisotropy antiferromagnets with circularly polarized spin-waves were reported to carry spin-information over long distances of micrometers. In this article, we report long-distance spin-transport in the easy-plane canted antiferromagnetic phase of hematite and at room temperature, where the linearly polarized magnons are not intuitively expected to carry spin. We demonstrate that the spin-transport signal decreases continuously through the easy-axis to easy-plane Morin transition, and persists in the easy-plane phase through current induced pairs of linearly polarized magnons with dephasing lengths in the micrometer range. We explain the long transport distance as a result of the low magnetic damping, which we measure to be below 0.0001 as in the best ferromagnets. All of this together demonstrates that long-distance transport can be achieved across a range of anisotropies and temperatures, up to room temperature, highlighting the promising potential of this insulating antiferromagnet for magnon-based devices.
△ Less
Submitted 28 April, 2021; v1 submitted 29 May, 2020;
originally announced May 2020.
-
Annealing and replica-symmetry in Deep Boltzmann Machines
Authors:
Diego Alberici,
Adriano Barra,
Pierluigi Contucci,
Emanuele Mingione
Abstract:
In this paper we study the properties of the quenched pressure of a multi-layer spin-glass model (a deep Boltzmann Machine in artificial intelligence jargon) whose pairwise interactions are allowed between spins lying in adjacent layers and not inside the same layer nor among layers at distance larger than one. We prove a theorem that bounds the quenched pressure of such a K-layer machine in terms…
▽ More
In this paper we study the properties of the quenched pressure of a multi-layer spin-glass model (a deep Boltzmann Machine in artificial intelligence jargon) whose pairwise interactions are allowed between spins lying in adjacent layers and not inside the same layer nor among layers at distance larger than one. We prove a theorem that bounds the quenched pressure of such a K-layer machine in terms of K Sherrington-Kirkpatrick spin glasses and use it to investigate its annealed region. The replica-symmetric approximation of the quenched pressure is identified and its relation to the annealed one is considered. The paper also presents some observation on the model's architectural structure related to machine learning. Since escaping the annealed region is mandatory for a meaningful training, by squeezing such region we obtain thermodynamical constraints on the form factors. Remarkably, its optimal escape is achieved by requiring the last layer to scale sub-linearly in the network size.
△ Less
Submitted 21 January, 2020;
originally announced January 2020.
-
A statistical-inference approach to reconstruct inter-cellular interactions in cell-migration experiments
Authors:
Elena Agliari,
Pablo J. Sáez,
Adriano Barra,
Matthieu Piel,
Pablo Vargas,
Michele Castellana
Abstract:
Migration of cells can be characterized by two, prototypical types of motion: individual and collective migration. We propose a statistical-inference approach designed to detect the presence of cell-cell interactions that give rise to collective behaviors in cell-motility experiments. Such inference method has been first successfully tested on synthetic motional data, and then applied to two exper…
▽ More
Migration of cells can be characterized by two, prototypical types of motion: individual and collective migration. We propose a statistical-inference approach designed to detect the presence of cell-cell interactions that give rise to collective behaviors in cell-motility experiments. Such inference method has been first successfully tested on synthetic motional data, and then applied to two experiments. In the first experiment, cell migrate in a wound-healing model: when applied to this experiment, the inference method predicts the existence of cell-cell interactions, correctly mirroring the strong intercellular contacts which are present in the experiment. In the second experiment, dendritic cells migrate in a chemokine gradient. Our inference analysis does not provide evidence for interactions, indicating that cells migrate by sensing independently the chemokine source. According to this prediction, we speculate that mature dendritic cells disregard inter-cellular signals that could otherwise delay their arrival to lymph vessels.
△ Less
Submitted 4 December, 2019;
originally announced December 2019.
-
Generalized Guerra's interpolation schemes for dense associative neural networks
Authors:
Elena Agliari,
Francesco Alemanno,
Adriano Barra,
Alberto Fachechi
Abstract:
In this work we develop analytical techniques to investigate a broad class of associative neural networks set in the high-storage regime. These techniques translate the original statistical-mechanical problem into an analytical-mechanical one which implies solving a set of partial differential equations, rather than tackling the canonical probabilistic route. We test the method on the classical Ho…
▽ More
In this work we develop analytical techniques to investigate a broad class of associative neural networks set in the high-storage regime. These techniques translate the original statistical-mechanical problem into an analytical-mechanical one which implies solving a set of partial differential equations, rather than tackling the canonical probabilistic route. We test the method on the classical Hopfield model - where the cost function includes only two-body interactions (i.e., quadratic terms) - and on the "relativistic" Hopfield model - where the (expansion of the) cost function includes p-body (i.e., of degree p) contributions. Under the replica symmetric assumption, we paint the phase diagrams of these models by obtaining the explicit expression of their free energy as a function of the model parameters (i.e., noise level and memory storage). Further, since for non-pairwise models ergodicity breaking is non necessarily a critical phenomenon, we develop a fluctuation analysis and find that criticality is preserved in the relativistic model.
△ Less
Submitted 16 April, 2020; v1 submitted 28 November, 2019;
originally announced November 2019.
-
Neural networks with redundant representation: detecting the undetectable
Authors:
Elena Agliari,
Francesco Alemanno,
Adriano Barra,
Martino Centonze,
Alberto Fachechi
Abstract:
We consider a three-layer Sejnowski machine and show that features learnt via contrastive divergence have a dual representation as patterns in a dense associative memory of order P=4. The latter is known to be able to Hebbian-store an amount of patterns scaling as N^{P-1}, where N denotes the number of constituting binary neurons interacting P-wisely. We also prove that, by keeping the dense assoc…
▽ More
We consider a three-layer Sejnowski machine and show that features learnt via contrastive divergence have a dual representation as patterns in a dense associative memory of order P=4. The latter is known to be able to Hebbian-store an amount of patterns scaling as N^{P-1}, where N denotes the number of constituting binary neurons interacting P-wisely. We also prove that, by keeping the dense associative network far from the saturation regime (namely, allowing for a number of patterns scaling only linearly with N, while P>2) such a system is able to perform pattern recognition far below the standard signal-to-noise threshold. In particular, a network with P=4 is able to retrieve information whose intensity is O(1) even in the presence of a noise O(\sqrt{N}) in the large N limit. This striking skill stems from a redundancy representation of patterns -- which is afforded given the (relatively) low-load information storage -- and it contributes to explain the impressive abilities in pattern recognition exhibited by new-generation neural networks. The whole theory is developed rigorously, at the replica symmetric level of approximation, and corroborated by signal-to-noise analysis and Monte Carlo simulations.
△ Less
Submitted 28 November, 2019;
originally announced November 2019.
-
Dreaming neural networks: rigorous results
Authors:
Elena Agliari,
Francesco Alemanno,
Adriano Barra,
Alberto Fachechi
Abstract:
Recently a daily routine for associative neural networks has been proposed: the network Hebbian-learns during the awake state (thus behaving as a standard Hopfield model), then, during its sleep state, optimizing information storage, it consolidates pure patterns and removes spurious ones: this forces the synaptic matrix to collapse to the projector one (ultimately approaching the Kanter-Sompolink…
▽ More
Recently a daily routine for associative neural networks has been proposed: the network Hebbian-learns during the awake state (thus behaving as a standard Hopfield model), then, during its sleep state, optimizing information storage, it consolidates pure patterns and removes spurious ones: this forces the synaptic matrix to collapse to the projector one (ultimately approaching the Kanter-Sompolinksy model). This procedure keeps the learning Hebbian-based (a biological must) but, by taking advantage of a (properly stylized) sleep phase, still reaches the maximal critical capacity (for symmetric interactions). So far this emerging picture (as well as the bulk of papers on unlearning techniques) was supported solely by mathematically-challenging routes, e.g. mainly replica-trick analysis and numerical simulations: here we rely extensively on Guerra's interpolation techniques developed for neural networks and, in particular, we extend the generalized stochastic stability approach to the case. Confining our description within the replica symmetric approximation (where the previous ones lie), the picture painted regarding this generalization (and the previously existing variations on theme) is here entirely confirmed. Further, still relying on Guerra's schemes, we develop a systematic fluctuation analysis to check where ergodicity is broken (an analysis entirely absent in previous investigations). We find that, as long as the network is awake, ergodicity is bounded by the Amit-Gutfreund-Sompolinsky critical line (as it should), but, as the network sleeps, sleeping destroys spin glass states by extending both the retrieval as well as the ergodic region: after an entire sleeping session the solely surviving regions are retrieval and ergodic ones and this allows the network to achieve the perfect retrieval regime (the number of storable patterns equals the number of neurons in the network).
△ Less
Submitted 21 December, 2018;
originally announced December 2018.
-
A novel derivation of the Marchenko-Pastur law through analog bipartite spin-glasses
Authors:
Elena Agliari,
Francesco Alemanno,
Adriano Barra,
Alberto Fachechi
Abstract:
In this work we consider the {\em analog bipartite spin-glass} (or {\em real-valued restricted Boltzmann machine} in a neural network jargon), whose variables (those quenched as well as those dynamical) share standard Gaussian distributions. First, via Guerra's interpolation technique, we express its quenched free energy in terms of the natural order parameters of the theory (namely the self- and…
▽ More
In this work we consider the {\em analog bipartite spin-glass} (or {\em real-valued restricted Boltzmann machine} in a neural network jargon), whose variables (those quenched as well as those dynamical) share standard Gaussian distributions. First, via Guerra's interpolation technique, we express its quenched free energy in terms of the natural order parameters of the theory (namely the self- and two-replica overlaps), then, we re-obtain the same result by using the replica-trick: a mandatory tribute, given the special occasion. Next, we show that the quenched free energy of this model is the functional generator of the moments of the correlation matrix among the weights connecting the two layers of the spin-glass (i.e., the Wishart matrix in random matrix theory or the Hebbian coupling in neural networks): as weights are quenched stochastic variables, this plays as a novel tool to inspect random matrices. In particular, we find that the Stieltjes transform of the spectral density of the correlation matrix is determined by the (replica-symmetric) quenched free energy of the bipartite spin-glass model. In this setup, we re-obtain the Marchenko-Pastur law in a very simple way.
△ Less
Submitted 20 November, 2018;
originally announced November 2018.
-
Dreaming neural networks: forgetting spurious memories and reinforcing pure ones
Authors:
Alberto Fachechi,
Elena Agliari,
Adriano Barra
Abstract:
The standard Hopfield model for associative neural networks accounts for biological Hebbian learning and acts as the harmonic oscillator for pattern recognition, however its maximal storage capacity is $α\sim 0.14$, far from the theoretical bound for symmetric networks, i.e. $α=1$. Inspired by sleeping and dreaming mechanisms in mammal brains, we propose an extension of this model displaying the s…
▽ More
The standard Hopfield model for associative neural networks accounts for biological Hebbian learning and acts as the harmonic oscillator for pattern recognition, however its maximal storage capacity is $α\sim 0.14$, far from the theoretical bound for symmetric networks, i.e. $α=1$. Inspired by sleeping and dreaming mechanisms in mammal brains, we propose an extension of this model displaying the standard on-line (awake) learning mechanism (that allows the storage of external information in terms of patterns) and an off-line (sleep) unlearning$\&$consolidating mechanism (that allows spurious-pattern removal and pure-pattern reinforcement): this obtained daily prescription is able to saturate the theoretical bound $α=1$, remaining also extremely robust against thermal noise. Both neural and synaptic features are analyzed both analytically and numerically. In particular, beyond obtaining a phase diagram for neural dynamics, we focus on synaptic plasticity and we give explicit prescriptions on the temporal evolution of the synaptic matrix. We analytically prove that our algorithm makes the Hebbian kernel converge with high probability to the projection matrix built over the pure stored patterns. Furthermore, we obtain a sharp and explicit estimate for the "sleep rate" in order to ensure such a convergence. Finally, we run extensive numerical simulations (mainly Monte Carlo sampling) to check the approximations underlying the analytical investigations (e.g., we developed the whole theory at the so called replica-symmetric level, as standard in the Amit-Gutfreund-Sompolinsky reference framework) and possible finite-size effects, finding overall full agreement with the theory.
△ Less
Submitted 29 October, 2018;
originally announced October 2018.
-
The Relativistic Hopfield network: rigorous results
Authors:
Elena Agliari,
Adriano Barra,
Matteo Notarnicola
Abstract:
The relativistic Hopfield model constitutes a generalization of the standard Hopfield model that is derived by the formal analogy between the statistical-mechanic framework embedding neural networks and the Lagrangian mechanics describing a fictitious single-particle motion in the space of the tuneable parameters of the network itself. In this analogy the cost-function of the Hopfield model plays…
▽ More
The relativistic Hopfield model constitutes a generalization of the standard Hopfield model that is derived by the formal analogy between the statistical-mechanic framework embedding neural networks and the Lagrangian mechanics describing a fictitious single-particle motion in the space of the tuneable parameters of the network itself. In this analogy the cost-function of the Hopfield model plays as the standard kinetic-energy term and its related Mattis overlap (naturally bounded by one) plays as the velocity. The Hamiltonian of the relativisitc model, once Taylor-expanded, results in a P-spin series with alternate signs: the attractive contributions enhance the information-storage capabilities of the network, while the repulsive contributions allow for an easier unlearning of spurious states, conferring overall more robustness to the system as a whole. Here we do not deepen the information processing skills of this generalized Hopfield network, rather we focus on its statistical mechanical foundation. In particular, relying on Guerra's interpolation techniques, we prove the existence of the infinite volume limit for the model free-energy and we give its explicit expression in terms of the Mattis overlaps. By extremizing the free energy over the latter we get the generalized self-consistent equations for these overlaps, as well as a picture of criticality that is further corroborated by a fluctuation analysis. These findings are in full agreement with the available previous results.
△ Less
Submitted 29 October, 2018;
originally announced October 2018.
-
Free energies of Boltzmann Machines: self-averaging, annealed and replica symmetric approximations in the thermodynamic limit
Authors:
Elena Agliari,
Adriano Barra,
Brunello Tirozzi
Abstract:
Restricted Boltzmann machines (RBMs) constitute one of the main models for machine statistical inference and they are widely employed in Artificial Intelligence as powerful tools for (deep) learning. However, in contrast with countless remarkable practical successes, their mathematical formalization has been largely elusive: from a statistical-mechanics perspective these systems display the same (…
▽ More
Restricted Boltzmann machines (RBMs) constitute one of the main models for machine statistical inference and they are widely employed in Artificial Intelligence as powerful tools for (deep) learning. However, in contrast with countless remarkable practical successes, their mathematical formalization has been largely elusive: from a statistical-mechanics perspective these systems display the same (random) Gibbs measure of bi-partite spin-glasses, whose rigorous treatment is notoriously difficult. In this work, beyond providing a brief review on RBMs from both the learning and the retrieval perspectives, we aim to contribute to their analytical investigation, by considering two distinct realizations of their weights (i.e., Boolean and Gaussian) and studying the properties of their related free energies. More precisely, focusing on a RBM characterized by digital couplings, we first extend the Pastur-Shcherbina-Tirozzi method (originally developed for the Hopfield model) to prove the self-averaging property for the free energy, over its quenched expectation, in the infinite volume limit, then we explicitly calculate its simplest approximation, namely its annealed bound. Next, focusing on a RBM characterized by analogical weights, we extend Guerra's interpolating scheme to obtain a control of the quenched free-energy under the assumption of replica symmetry: we get self-consistencies for the order parameters (in full agreement with the existing Literature) as well as the critical line for ergodicity breaking that turns out to be the same obtained in AGS theory. As we discuss, this analogy stems from the slow-noise universality. Finally, glancing beyond replica symmetry, we analyze the fluctuations of the overlaps for an estimate of the (slow) noise affecting the retrieval of the signal, and by a stability analysis we recover the Aizenman-Contucci identities typical of glassy systems.
△ Less
Submitted 8 March, 2019; v1 submitted 20 October, 2018;
originally announced October 2018.
-
Voltage Control of Magnetic Monopoles in Artificial Spin Ice
Authors:
Andres C. Chavez,
Anthony Barra,
Gregory P. Carman
Abstract:
Current research on artificial spin ice (ASI) systems has revealed unique hysteretic memory effects and mobile quasi-particle monopoles controlled by externally applied magnetic fields. Here, we numerically demonstrate a strain-mediated multiferroic approach to locally control the ASI monopoles. The magnetization of individual lattice elements is controlled by applying voltage pulses to the piezoe…
▽ More
Current research on artificial spin ice (ASI) systems has revealed unique hysteretic memory effects and mobile quasi-particle monopoles controlled by externally applied magnetic fields. Here, we numerically demonstrate a strain-mediated multiferroic approach to locally control the ASI monopoles. The magnetization of individual lattice elements is controlled by applying voltage pulses to the piezoelectric layer resulting in strain-induced magnetic precession timed for 180 degree reorientation. The model demonstrates localized voltage control to move the magnetic monopoles across lattice sites, in CoFeB, Ni, and FeGa based ASI$'$s. The switching is achieved at frequencies near ferromagnetic resonance and requires energies below 620 aJ. The results demonstrate that ASI monopoles can be efficiently and locally controlled with a strain-mediated multiferroic approach.
△ Less
Submitted 22 March, 2018;
originally announced March 2018.
-
Strain-mediated spin-orbit torque switching for magnetic memory
Authors:
Qianchang Wang,
John Domann,
Guoqiang Yu,
Anthony Barra,
Kang L. Wang,
Gregory P. Carman
Abstract:
Spin-orbit torque (SOT) represents an energy efficient method to control magnetization in magnetic memory devices. However, deterministically switching perpendicular memory bits usually requires the application of an additional bias field for breaking lateral symmetry. Here we present a new approach of field-free deterministic perpendicular switching using a strain-mediated SOT switching method. T…
▽ More
Spin-orbit torque (SOT) represents an energy efficient method to control magnetization in magnetic memory devices. However, deterministically switching perpendicular memory bits usually requires the application of an additional bias field for breaking lateral symmetry. Here we present a new approach of field-free deterministic perpendicular switching using a strain-mediated SOT switching method. The strain-induced magnetoelastic anisotropy breaks the lateral symmetry, and the resulting symmetry-breaking is controllable. A finite element model and a macrospin model are used to numerically simulate the strain-mediated SOT switching mechanism. The results show that a relatively small voltage (${\pm}0.5$ V) along with a modest current ($3.5 \times 10^{7} A/cm^{2}$) can produce a 180° perpendicular magnetization reversal. The switching direction (up or down) is dictated by the voltage polarity (positive or negative) applied to the piezoelectric layer in the magnetoelastic/heavy metal/piezoelectric heterostructure. The switching speed can be as fast as 10 GHz. More importantly, this control mechanism can be potentially implemented in a magnetic random-access memory system with small footprint, high endurance and high tunnel magnetoresistance (TMR) readout ratio.
△ Less
Submitted 8 December, 2017;
originally announced February 2018.
-
Complex Reaction Kinetics in Chemistry: A unified picture suggested by Mechanics in Physics
Authors:
Elena Agliari,
Adriano Barra,
Giulio Landolfi,
Sara Murciano,
Sarah Perrone
Abstract:
Complex biochemical pathways or regulatory enzyme kinetics can be reduced to chains of elementary reactions, which can be described in terms of chemical kinetics. This discipline provides a set of tools for quantifying and understanding the dialogue between reactants, whose framing into a solid and consistent mathematical description is of pivotal importance in the growing field of biotechnology.…
▽ More
Complex biochemical pathways or regulatory enzyme kinetics can be reduced to chains of elementary reactions, which can be described in terms of chemical kinetics. This discipline provides a set of tools for quantifying and understanding the dialogue between reactants, whose framing into a solid and consistent mathematical description is of pivotal importance in the growing field of biotechnology. Among the elementary reactions so far extensively investigated, we recall the socalled Michaelis-Menten scheme and the Hill positive-cooperative kinetics, which apply to molecular binding and are characterized by the absence and the presence, respectively, of cooperative interactions between binding sites, giving rise to qualitative different phenomenologies. However, there is evidence of reactions displaying a more complex, and by far less understood, pattern: these follow the positive-cooperative scenario at small substrate concentration, yet negative-cooperative effects emerge and get stronger as the substrate concentration is increased. In this paper we analyze the structural analogy between the mathematical backbone of (classical) reaction kinetics in Chemistry and that of (classical) mechanics in Physics: techniques and results from the latter shall be used to infer properties on the former.
△ Less
Submitted 5 January, 2018;
originally announced January 2018.
-
A relativistic extension of Hopfield neural networks via the mechanical analogy
Authors:
Adriano Barra,
Matteo Beccaria,
Alberto Fachechi
Abstract:
We propose a modification of the cost function of the Hopfield model whose salient features shine in its Taylor expansion and result in more than pairwise interactions with alternate signs, suggesting a unified framework for handling both with deep learning and network pruning. In our analysis, we heavily rely on the Hamilton-Jacobi correspondence relating the statistical model with a mechanical s…
▽ More
We propose a modification of the cost function of the Hopfield model whose salient features shine in its Taylor expansion and result in more than pairwise interactions with alternate signs, suggesting a unified framework for handling both with deep learning and network pruning. In our analysis, we heavily rely on the Hamilton-Jacobi correspondence relating the statistical model with a mechanical system. In this picture, our model is nothing but the relativistic extension of the original Hopfield model (whose cost function is a quadratic form in the Mattis magnetization which mimics the non-relativistic Hamiltonian for a free particle). We focus on the low-storage regime and solve the model analytically by taking advantage of the mechanical analogy, thus obtaining a complete characterization of the free energy and the associated self-consistency equations in the thermodynamic limit. On the numerical side, we test the performances of our proposal with MC simulations, showing that the stability of spurious states (limiting the capabilities of the standard Hebbian construction) is sensibly reduced due to presence of unlearning contributions in this extended framework.
△ Less
Submitted 5 January, 2018;
originally announced January 2018.