Abstract
Feedback control of open quantum systems is of fundamental importance for practical applications in various contexts, ranging from quantum computation to quantum error correction and quantum metrology. Its use in the context of thermodynamics further enables the study of the interplay between information and energy. However, deriving optimal feedback control strategies is highly challenging, as it involves the optimal control of open quantum systems, the stochastic nature of quantum measurement, and the inclusion of policies that maximize a long-term time- and trajectory-averaged goal. In this work, we employ a reinforcement learning approach to automate and capture the role of a quantum Maxwell’s demon: the agent takes the literal role of discovering optimal feedback control strategies in qubit-based systems that maximize a trade-off between measurement-powered cooling and measurement efficiency. Considering weak or projective quantum measurements, we explore different regimes based on the ordering between the thermalization, the measurement, and the unitary feedback timescales, finding different and highly non-intuitive, yet interpretable, strategies. In the thermalization-dominated regime, we find strategies with elaborate finite-time thermalization protocols conditioned on measurement outcomes. In the measurement-dominated regime, we find that optimal strategies involve adaptively measuring different qubit observables reflecting the acquired information, and repeating multiple weak measurements until the quantum state is ‘sufficiently pure’, leading to random walks in state space. Finally, we study the case when all timescales are comparable, finding new feedback control strategies that considerably outperform more intuitive ones. We discuss a two-qubit example where we explore the role of entanglement and conclude discussing the scaling of our results to quantum many-body systems.
Original content from this work may be used under the terms of the Creative Commons Attribution 4.0 license. Any further distribution of this work must maintain attribution to the author(s) and the title of the work, journal citation and DOI.
1. Introduction
Recent years have seen remarkable progress in the realization and control of various quantum technological platforms, ranging from superconducting qubits [1–4] to ultra-cold trapped ion [5, 6] and quantum dots [7–9]. A key ingredient in achieving this feat is feedback control, whereby a set of operations is performed depending on the information acquired through measurements about the controlled system [10–12]. Crucially, as demonstrated by Maxwell’s demon thought-experiment [13, 14] and Szilard’s engine [15], feedback control represents the arch-stone bridging information theory, statistical mechanics and thermodynamics. Indeed, a Maxwell’s demon is an agent that can extract work from a thermalized system solely by acquiring information about it, leading to an apparent violation of the second law of thermodynamics. This apparent contradiction was later solved by Landauer [16] and Bennett [17], who accounted for the minimum cost of erasing the classical information acquired about the system, given by Landauer’s limit [18]. For a random bit of information, this minimum work cost amounts on average to
, kB being Boltzmann’s constant, and T the temperature. Although of fundamental importance, this limit notoriously represents a very loose estimate since many-body or macroscopic systems dissipate several orders of magnitude more. However, Maxwell’s demon has recently attracted great renewed interest due to technological leaps in miniaturization and control. Several state of the art experiments have successfully managed to probe Landauer’s limit on classical platforms, including colloidal particles [19–22], nano-magnets [23–25], superconducting flux logic cells [26] and underdamped micro-mechanical oscillators [27, 28]. This has opened the door to the investigation of thermodynamics of nano-scale quantum thermal machines [7, 29–38], and in particular to the impact of quantum information and feedback control on them [39–46]. Understanding this cost is hence crucial to optimize current and next-generation quantum devices [47–70] and ultimately minimizing their energetic footprint [71].
The role of information is even more fundamental in quantum systems, as it becomes inseparable from the measurement procedure necessary to acquire and manipulate it. In contrast to classical machines, measurements intrinsically and necessarily alter the state and the energy content of quantum systems [12, 39–41, 72, 73]. All these factors paint a very rich and complex landscape, where determining the control strategies that give the optimal trade-off between power and efficiency becomes a difficult task, even for very simple models.
The effect of measurement and feedback in the context of quantum systems and the idea of a quantum Maxwell’s demon have been studied in [39–41, 72–85]. However, their finite-time analysis and systematic optimization, crucial for actual quantum technologies, has yet to be carried out. Finite-time thermodynamic optimization of quantum systems is a notoriously complex problem, even in the absence of measurements and feedback, due to the out-of-equilibrium quantum dynamics of the system, the large optimization space given by all possible drivings, and the existence of trade-offs, e.g. between power and efficiency [86–110]. Whenever quantum measurements and feedback are taken into account, the problem becomes even richer, as the goal shifts from finding the best (fixed) protocol to finding the best feedback control strategy, i.e. the policy which, based on the information gathered about the system, selects a specific optimal sequence of operations that maximize a target figure of merit. Notice that this optimization is further complicated by the stochastic nature of quantum measurement, which gives rise to stochastic trajectories, and the optimization goal is the long-term time- and trajectory-average of the figure of merit. Crucially, if we choose as figure of merit the cooling power from a single thermal reservoir, the optimization goal evidently coincides with the prototypical behavior a Maxwell’s demon, i.e. an agent that acquires information about a system to lower its temperature, and its thermodynamic cost is characterized by the above-mentioned Landauer’s limit.
Changing gear for a moment, recent years have witnessed the rise of machine learning methods applied across all scientific fields [111], including quantum mechanics [112–114]. Among these, reinforcement learning (RL) has emerged as a powerful tool to prepare quantum states [115–121], including feedback control [122, 123], to discover efficient implementations of quantum gates [124, 125], and for quantum error correction [126]. Recently, it has been applied in the context of quantum thermodynamics to design optimal time-dependent cycles and protocols [127–132]. However, the full potential of RL as a tool to discover optimal feedback control strategies for quantum thermodynamic applications has not yet been explored.
In this work, we take a radical step in bringing together all the above ideas—quantum feedback control, methods of RL, and quantum thermodynamics – in an entirely fresh fashion: in our approach, we take the RL agent literally as a Maxwell’s demon-like entity that, at every time-step, cn decide to acquire information about a quantum system and to act on it accordingly. In this way, we give the agent in RL an ontological meaning and a thermodynamic interpretation. Concretely, we employ RL to perform a systematic and comprehensive study to discover optimal feedback control strategies to maximize the long-term time-average and trajectory-averaged heat extracted from a single thermal bath, thus finding the ‘best artificially intelligent quantum Maxwell’s demon’ without any prior knowledge about quantum mechanics, nor of thermodynamics.
As depicted in figure 1 the agent, acting effectively as a quantum Maxwell’s demon, has to choose at every time step whether it should perform a (partial) thermalization [133] of the system with a heat bath (orange line), measure the quantum system (blue line), or apply unitary feedback (green line). Additionally, it must learn the value of some continuous time-dependent control governing the dynamics of the system. The aim is to maximize the long-term average cooling power, i.e. the heat extraction from the thermal bath, exploiting invasive quantum measurements and feedback while also minimizing the thermodynamic cost of measuring the system. Notice that, without measurements, it would not be thermodynamically possible to cool a single thermal bath.
Figure 1. Schematic representation of a quantum Maxwell’s demon. Based on previous measurement outcomes, the demon can decide whether to (partially) thermalize the quantum system, perform further measurements, or perform unitary feedback. The goal is to optimize the trade-off between cooling power and the cost of measuring the system. In this manuscript, we consider a reinforcement learning agent an actual Maxwell’s demon.
Download figure:
Standard image High-resolution imageThere are three timescales that are relevant in our analysis: (1) the thermalization timescale
, representing the typical timescale for the quantum system to thermalize with the thermal bath, (2) the measurement timescale
, representing the typical timescale required to perform an approximately projective measurement, and (3) the unitary timescale
, representing the typical timescale required to implement a rotation on the quantum system while decoupled from the thermal bath. Depending on the ordering between these timescales, we can be in completely different regimes that describe different physics and are thus characterized by profoundly different control strategies. In general, this regime will be dictated by the experimental details of the device.
In this work, we wish to strike a balance between the complexity of a realistic description of the dynamics of the system and the interpretability of our results. We thus proceed incrementally: first, we consider the limiting case where only one timescale among
,
and
is relevant, the other two being much faster. This allows us to obtain interesting yet interpretable results from which we can learn general strategies. We then showcase the effectiveness of our method applying it to more complicated scenarios where multiple timescales are finite. In such cases, we find complicated control strategies that provably outperform intuitive strategies. For concreteness, we will consider one and two-qubit systems coupled to a thermal bath.
In section 2, we introduce the model and the optimization objectives. In section 3, we frame the optimization of the quantum Maxwell’s demon performance as an RL task. In the following sections, we apply our method to different physical regimes, namely:
- Thermalization dominated regime. In section 4, we consider a thermalization timescale, modeled with Lindblad dynamics, that is much slower than the other ones, i.e.
. While the strategy is simple when we are only interested in maximizing the cooling power, we find elaborate yet intuitive finite-time control strategies when we are interested in a trade-off between cooling power and measurement efficiency. The duration of the control sequence becomes gradually slower and slower as we are more interested in efficiency. We also compare our results with a more intuitive strategy inspired by the RL results, finding that the RL control performs better than a standard Maxwell’s demon. - Measurement dominated. In section 5, we consider the measurement timescale to be much slower than the other ones, i.e.
. We model the finiteness of
considering weak measurements, both discrete and continuous, and we consider both a fixed and learnable measurement basis. Here, we find elaborate yet interpretable control strategies, characterized by multiple repeated measurements, that vary abruptly depending on the measurement strength, shedding light on the relation between the generation of coherence during the measurement process, the measurement strength, the measured observable, and the performance of the Maxwell’s demon. We further find that adaptively changing the measured observable outperforms a fixed observable. - Measurement and thermalization dominated regime. In section 6, we consider comparable thermalization and measurement timescales that are much slower than the unitary timescale, i.e.
. The RL agent learns to perform thermalization strokes whose duration and line-shape adaptively depend on the level of quantum purity reached during the previous quantum measurements. - Two qubit setup: all timescales are finite. In section 7, we consider a two-qubit setup where all timescales are finite and comparable, i.e.
. Here, we model finite-time thermalization and unitary dynamics using the Lindblad equation. For the measurement, we couple the main qubit to an additional auxiliary qubit that acts as a measurement probe that is then projectively measured. This can be considered as an explicit and finite-time modelization of a positive-operator-valued measure (POVM). We consider both the case with and without counter-rotating terms in the qubit–qubit interaction. Here, we rigorously show that the strategies learned by the RL agent outperform intuitive control strategies, and we analyze the entanglement generated between the two qubits.
2. Quantum Maxwell’s demon
To realize a quantum Maxwell’s demon gedankenexperiment, we consider the setup sketched in figure 1. It consists of a quantum system, a thermal bath, a measurement probe, and time-dependent controls on the quantum system. We discretize time in steps
, where Δt is the duration of each time step. In each time step, we assume that the quantum system can either undergo (partial) thermalization with the thermal bath, a quantum measurement, or unitary feedback at discretion of the Maxwell’s demon. Because of this, strictly speaking all three steps can be considered feedback; however rotations and Hamiltonian dynamics that pertain to unitary feedback are kept distinct for consistency.
Unitary feedback: the system undergoes evolution in the absence of the bath and the measurement probe. We consider two different cases of unitary evolution: 1) unitary rotation of a qubit around the Bloch sphere given by

where
are Pauli matrices and
are a set of tunable angles, and 2) Hamiltonian dynamics of the system given by

where
is the family of Hamiltonians of the quantum system that depends on the time-dependent controls
(henceforth denoted by u(t) for simplicity),
is the reduced density matrix of the quantum system, and
is Planck’s reduced constant. The former case describes the regime where unitary dynamics is the fastest timescale, whereas the latter describes finite-time unitary dynamics induced by Hamiltonian dynamics.
Thermalization: when the system is coupled to a thermal bath with inverse temperature
, we describe its dynamics by the Lindblad master equation [134, 135]

where
describes the family of Hamiltonians of the system during thermalization (that, in principle, can be different from the family of Hamiltonians
describing the unitary evolution), and
is the dissipator describing the coupling of the system to the thermal bath. It can be expressed as

where
are the jump operators (possibly u(t)-dependent), and
are the corresponding transition rates. We employ the global master equation to guarantee thermodynamic consistency [136]. Further, we consider a bosonic heat bath with flat spectral density.
Measurement: We consider both discrete and continuous, as well as strong and weak quantum measurements. All such scenarios can be encompassed in the general framework of POVMs, i.e. a set of Kraus operators Mk—one for each measurement outcome—satisfying
for discrete measurements, and
for continuous measurements [137, 138]. The probability (probability density in the continuous case) of measuring outcome k at time t is given by

The post-measurement state
, conditioned on the observation k is given by

where we account for finite measurement time Δt. When studying a qubit system, we consider quantum measurements of the observable
, where
. Weak discrete quantum measurement are then described by the two Kraus operators
given by [138]

where I2 is the
identity, and
is an indicator of the strength of the discrete measurement [12, 139, 140]. The κ → 1 limits describe strong (projective) measurements, where the demon acquires maximum information about the system, whereas
describes the opposite limit, where no information is acquired. Intermediate values of κ describe the transition from strong to weak measurements.
In the case of continuous measurements, we have a continuum of Kraus operators
where k is now a continuous measurement outcome. They are given by [12]

for the measurement of σθ. Crucially
, which is the characteristic measurement timescale, can be seen as the inverse of the measurement strength [12, 138, 141] and is the time required to achieve unit measurement signal to noise ratio [141]. When
is large (small), the measurement is referred to as strong (weak). Following equation (8), the measurement readout is described by two Gaussian distributions with variance
and mean +1 (associated with the
measurement outcome) and −1 (associated with the
measurement outcome).
2.1. Demon’s optimization goals
Let P(t) be the instantaneous cooling power, i.e. the heat flux flowing out of the thermal bath. It will be zero when we measure or evolve unitarily. When thermalizing, it is given by [142, 143]

Notice that such quantity depends on the specific trajectory, i.e. on the stochasticity of the measurement outcome and (eventually) on the stochasticity of the choice of the feedback. Since we are interested in the average long-term performance of the demon, we define the ‘average long-term cooling power’
as

where t0 is an arbitrary initial time,
represents the average over all trajectories, and the time integral is an exponentially weighted time average with typical timescale
. Notice that, by taking such an average over all trajectories,
does not depend on t0. Ideally, we are interested in the limit
, i.e.
.
Various proposals have been put forward in the literature to quantify the thermodynamic cost of measurements [144–146]. Here, we take an approach inspired by Bennett’s solution to the Maxwell’s demon paradox. When the demon measures the quantum system, it stores the measurement outcome as classical information. Landauer showed that the minimum work necessary to erase such information is given by

where
is the Shannon entropy of the probability distribution pk, and
is the distribution of measurement outcomes if we measure at time t, given by equation (5). Since the measurement outcomes are stored in a classical memory by the demon, the entropy of the classical memory will correspond to the entropy of the
distribution. We thus define as ‘instantaneous dissipation rate’, the quantity

where
are the (stochastic) times at which we perform a measurement. As we did for the cooling power, we define the long-term average dissipation as

Applying the second law of thermodynamics to the thermal bath, the quantum system, and the heat dissipated by Landauer erasure of the classical memories and averaging over time and trajectories, we find

where the first term represents the entropy change of the thermal bath, the second term is the (minimum) entropy dissipated by Landauer erasure, and the third term is the entropy change of the quantum system. Notice that, thanks to the long-time average and the average over trajectories, the entropy change of the quantum system will be zero in expectation. This leads to

where we define η > 0 as the measurement efficiency. Indeed, if η = 1, we are converting into cooling power all of the heat that we dissipate by Landauer erasure. In this sense, we are making optimal use of all the classical information acquired about the system; if instead η < 1, we have irreversibilities.
Interestingly, one cannot simultaneously optimize the average cooling power
and dissipation
. Indeed, zero dissipation implies a vanishingly small cooling power (since the overall transformation must be reversible), while a high cooling power can be achieved through many measurements and high dissipation. We thus perform a multi-objective optimization searching for the Pareto-optimal trade-off between the two quantities. These are feedback control schemes where one objective (
or
) cannot be further improved without sacrificing the other one [131]. The collection of
points delivered by all Pareto-optimal feedback control schemes is known as the Pareto front. If the Pareto front is convex, it can be found by maximizing the figure of merit

with respect to all possible feedback control schemes and repeating such an optimization for all values of
. Solutions found for c = 1 correspond to maximum cooling power, whereas solutions for c = 0 correspond to minimum dissipation (notice the minus sign in front of
). For intermediate values of c, we will find different trade-offs between the two objectives. Using the linearity of the expectation value, notice that we can interpret equation (16) as the trajectory and time average of the instantaneous quantity

We stress that, given the monotonic dependence of η on
and
, it can be shown that feedback control schemes that are Pareto-optimal for power and efficiency are also Pareto-optimal for power and dissipation. This implies that the Pareto front between power and dissipation, found e.g. by maximizing equation (17), also provides the Pareto front between power and efficiency (see appendix
3. RL formulation
RL is a general optimization framework based on Markov decision processes [147], which consists of learning how to maximize a long-term goal by repeated interactions between an ‘agent’ and an ‘environment’. To avoid confusion with the terminology commonly used in open quantum systems’ theory, we henceforth denote the environment in the RL context with RL-E. As shown in figure 2, interactions between the agent (blue box) and the RL-E (green box) occur at every time-step ti. We choose as the RL-E a quantum system that can be coupled to a thermal bath, measured, and controlled by unitary evolution (green box). At each time-step
, the agent chooses an action ai to perform on the RL-E (lower orange box) according to the policy function
, which represents the probability of choosing action ai, given that the RL-E is in state si. We choose as
a combination of a discrete di and continuous ui action. As discrete action, we choose the three possibilities
described in section 2, corresponding respectively to performing unitary feedback, a (partial) thermalization, or a measurement. As continuous action ui we choose the value of the control
, assumed to be constant in each time-interval
, corresponding either to the value of a Hamiltonian parameter or to a rotation angle of the qubit state. The RL-E then returns as feedback what is the state
at the next time-step
, and a scalar quantity
, known as the reward.
Figure 2. Schematic representation of the RL method learning to act as an optimal quantum Maxwell’s demon. At every small time-step
, the RL agent (blue box) interacts with an open quantum system (green box) performing actions ai (lower orange box) consisting of a discrete choice (whether to thermalize, measure, or evolve unitarily), and a continuous action (representing some time-dependent control). After the quantum state has evolved for a time-step Δt, the agent receives as feedback the state of the RL environment (RL-E)
, given by the density matrix of the quantum system conditioned on the measurement outcome, and a reward
representing the optimization goal (upper orange boxes). The computer agent must learn an optimal policy
that maximizes the long-term and trajectory averaged trade-off between cooling power and cost of measurement. Through the trial and error attempt, the computer agent learns a gradually better and better policy until convergence.
Download figure:
Standard image High-resolution imageCrucially, we choose as a state of the RL-E
, i.e. the reduced density matrix of the quantum system conditioned on the measurement outcomes. Since
encodes all information previously acquired by measuring the state, and since actions are chosen based on the state of the RL-E, this allows the agent to learn feedback control strategies
that make use of all previous measurement outcomes.
The goal of RL in the so-called continuing task is to identify an optimal policy function
such that, at every time-step ti, the expectation value of the exponentially weighted sum of future rewards is maximized, i.e.

The reward
is a (possibly stochastic) scalar quantity that only depends on the last state and the last action chosen, i.e. on
and ai, whereas
is the discount factor which determines the timescale over which we optimize future rewards. The expectation value
, as previously, represents an average over all trajectories. Let us choose as the reward function

Plugging equation (19) into equation (18), and using equation (17), we see that, in the limit
,

where we choose

As we can see, the discount factor γ plays the role of the time-average timescale, with γ → 1 corresponding to
. Identifying the optimal policy
defined in equation (20) is a formidable task. Indeed, in
, we have an expectation value over time, over all possible stochastic measurement outcomes and stochastic actions, and the space of feedback schemes increases exponentially with the number of time steps.
Notice that with this choice of RL-E, state, and reward, the agent is behaving as an actual Maxwell’s demon, and identifying the optimal policy precisely corresponds to discovering the best performing Maxwell’s demon. In this work, we employ a modification of the soft actor-critic (SAC) algorithm [148–150] to solve the optimization problem stated in equation (20) with rewards given by equation (19). The method, implemented in PyTorch [151], is detailed in appendix
, and by interacting many times with the quantum system, we iteratively improve the policy until we find a (quasi) optimal policy. We either consider the case of maximum power (c = 1) or repeat the optimization process for many values of c, and for each one, we compute the corresponding values of the dissipation
and the cooling power
. Notably, the RL method naturally optimizes the sum of the reward averaging over time (with a timescale determined by γ) and over the trajectories, i.e. over the stochasticity of the measurement outcome and the policy. See appendix
4. Thermalization dominated regime
We start by studying the regime where the thermalization timescale
is the slowest one, i.e.
. We show that the main impact of the thermalization timescale is the emergence of strategies where a single measurement is performed, followed by a carefully designed and finite-time thermalization protocol, conditioned on the last measurement outcome, whose duration depends on the power-efficiency trade-off (the more we are interested in the cooling power, the shorter the thermalization and the higher the dissipation). This regime can be a realistic scenario in those setups where the coupling to the thermal baths is engineered to be slow compared to the other timescales, as in [152]. We begin for simplicity with a single qubit, and later we generalize it to a larger system. Specifically, we consider a single qubit system described by the Hamiltonian

where E0 is a constant that sets the typical energy gap of the qubit, and u(t) is a continuous time-dependent control. In the experiment described in [152], this would correspond to a time-dependent gate voltage.
We describe finite-time thermalization using the Lindblad master equation with characteristic thermalization rate Γ. We have the following rates and jump operators for
(see equation (4))

Here
, Γ is a constant that depends on the qubit-bath coupling strength, and
is the Bose–Einstein distribution function. We then choose a value of Δt that is much smaller than the thermalization timescale of the qubit. To enforce that
, we consider projective measurements in the σz basis that occur in a single time-step Δt.
Since the only control is on σz and we only perform measurements of σz, the Lindblad dynamics of the system is fully characterized by the probability
of being in the higher eigenstate
of σz, and all dependence on the coherence drops. Mathematically, this is equivalent to a classical two-level system, which can be regarded as having no unitary dynamics timescale (
). Notice that by changing the sign of u(t), we can exchange the ground and the excited state of the system; this will be exploited by the agent to apply conditional feedback to the system.
In figure 3(a), we plot the Pareto front between the long-term average of the cooling power
and of the negative dissipation
. Each black dot corresponds to a separate optimization performed with a different value of c. In figure 3(b), we plot the same points as in figure 3(a), but plotting the Pareto front between the average cooling power and the measurement efficiency η. As we can see, we find a class of policies that trade between high power and high efficiency.
Figure 3. Maxwell’s demon performance in the thermalization-dominated regime. Pareto front between the long-term average of the cooling power
and of the negative measurement dissipation
(a), and between the cooling power
and the measurement efficiency η (b). Each black dot corresponds to a separate RL optimization for different values of c, whereas the red line corresponds to the Pareto front of the interpretable policy described in section 4. Example of actions chosen by the agent along an arbitrary trajectory, as a function of time, in the c = 0.8 case (c) and c = 0.65 case (d). The corresponding points on the Pareto front are shown in (a), (b) as empty black circles. The color corresponds to the discrete action (see legend), and the value shown on the y-axis corresponds to the value of u(t) during the thermalization step. Measurements are shown as vertical lines. Parameters:
,
,
, and
for
since the thermalization time increases as c decreases.
Download figure:
Standard image High-resolution imageTo visualize the type of policies that we find, in figures 3(c) and (d), we plot an example of the chosen actions along an arbitrary trajectory as a function of time, respectively, for c = 0.8 and c = 0.65. The corresponding value on the Pareto front is shown in figures 3(c) and (d) as an empty black circle. In both cases, the agent learns the following policy:
- Perform a projective measurement (blue vertical line).
- Perform feedback: if the system is in the ground state, do nothing; if it is in the excited state, change the sign of u(t) as to have the system in the ground state.
- Perform a finite-time thermalization of the qubit (orange dots) while ramping the value of u(t) (plotted on the y-axis). The qubit, being in the ground state, absorbs heat from the bath, thus cooling it.
Feedback conditioned on the last measurement outcome is clearly visible in figures 3(c) and (d) as changes of sign of the orange thermalization ramps. Furthermore, as we would intuitively expect, the agent automatically learns to increase the duration of the thermalization as we shift our interest from high power to high efficiency. Indeed, a slow thermalization with a slow ramping of the control allows for the maximization of the heat extracted per measurement at the cost of taking a longer time, thus decreasing the cooling power but increasing the measurement efficiency.
In the c → 1 limit, in which we are only interested in power, we find that the optimal policy consists of performing measurements followed by infinitesimally short thermalization. Interestingly, this implies that the optimal duration of the thermalization (given by the duration of the orange ramps in figures 3(c) and (d)), is only determined by the value c, i.e. by how much we are interested in trading power and efficiency. We also notice that the agent learns to never perform unitary actions: indeed, in this system, a unitary evolution would not change the state p(t), nor would it exchange any heat with the thermal bath, making it always a suboptimal choice. The slight asymmetry in the thermalization ramps in figure 3(d) between positive and negative values of the control u(t) is likely due to numerical imperfections in the RL optimization and hints at the fact that the performance only weakly depends on the exact shape of the thermalization ramp. We now confirm this substantially more quantitatively.
Thanks to the intuition gained from analyzing the behavior of the RL agent, we define an ‘interpretable policy’ that behaves as in the bullets described above, but thermalizes the qubit while keeping the control u(t) fixed at some value
for some time
. Therefore, we do not perform the smooth ramping of u(t) shown in figures 3(c) and (d), but we keep it constant. We numerically optimize its performance for many values of c, with respect to the two parameters
and
, and plot the corresponding performance in the Pareto fronts in figures 3(a) and (b) as a red thin line.
Notably, the interpretable policy captures the main features of the Pareto front and the increase of τ from 0 (when we are only interested in power) to a larger and larger value as we shift our interest from power to efficiency. However, the non-trivial ramping of the control u(t) discovered by the RL agent slightly outperforms the interpretable policy in the high-efficiency regime (see the black dots in figure 3(b) that lie above the red curve for η > 0.5). This is due to the further reduction of irreversibilities during the thermalization process thanks to a smooth ramping of u(t), a phenomenon that was proven analytically in the slow-driving regime and observed experimentally [153].
5. Measurement dominated regime
We now study the regime where the measurement timescale
is the slowest one, i.e.
. We show that optimal feedback control strategies involve adaptively measuring the qubit in different bases according to the acquired information. Furthermore, the main impact of the measurement timescale is generally the emergence of weak measurements that are repeated until the state is sufficiently pure, leading to a random walk in state space, followed by short unitary and thermalization steps conditioned on the previous outcomes. The use of continuous or discrete measurements leads to qualitatively similar feedback control strategies, with the former exploring a larger portion of the state space of the qubit because of the continuous measurement outcomes. This regime is experimentally relevant since in most superconducting qubit platforms, the time to perform, e.g. a π pulse, i.e. a unitary rotation, is much faster than the timescale for a measurement readout [154, 155].
Specifically, we consider a single qubit system whose Hamiltonian is given by equation (22). We then proceed incrementally: in section 5.1, we consider discrete weak measurements of a fixed observable that may or may not generate coherence in the Hamiltonian basis. In section 5.2, we allow the agent to learn and adaptively change the observable to measure. In section 5.3, we consider continuous weak measurements.
5.1. Discrete, fixed measurements
We start by considering a fixed measurement basis. In this section, we model the three discrete actions in the following way. During thermalization, we keep the qubit’s energy gap constant at
, and we consider a qubit-bath coupling strength of
, which brings the qubit state close to a thermal state in a single time-step Δt. We consider discrete yet non-projective measurements of
with measurement strength κ as described by equation (7). When performing the unitary dynamics, in a single time-step Δt we rotate the state of qubit about the y-axis of an angle
that the RL agent can choose.
In figure 4 we show the results of the RL agent in various scenarios: each row represents a different value of the measurement strength κ, and each column the type of measurement that the agent can perform (either σx, left, or σz, right). The position of the dots in figure 4 represent the states of the qubit visited while interacting with the RL agent (their
and
components are plotted, since
), and the color of the dots represents the corresponding action chosen by the agent after visiting such states.
Figure 4. Actions chosen by the RL agent as a function of qubits state represented as a point on the Bloch sphere in the measurement-dominated regime with discrete fixed measurements. The type of action is represented by the color of a point, and the black cross represents the thermal state. We allow the RL agent to cool the thermal bath either by σx measurement (left column) or σz measurement (right column). ρx and ρz denote qubit’s x and z coordinates on the Bloch sphere. Each row represents a different measurement strength, respectively
. Each plot represents a trajectory using the RL policy for 10 000 time steps with parameters
,
.
.
Download figure:
Standard image High-resolution imageInterestingly, aside from important differences for different values of κ commented below, the following general and interpretable strategy emerges from all cases.
- 1.Let us assume that the qubit is in a thermal state (black cross in figure 4).
- 2.Perform a measurement of the qubit (blue dots).
- 3.If the purity of the state reaches a threshold value (green dots), go to step 4; otherwise, go back to step 2 (blue dots).
- 4.Perform a unitary rotation of the qubit to the negative z-axis (orange dots).
- 5.Thermalize the qubit, leading us back to step 1.
This strategy can be easily interpreted. While unitary control cannot change the purity of the state (i.e. the radius of the state on the Bloch sphere), the measurement process gives us information about the state of the system and thus tends to purify the state of the qubit. The RL agent learns to repeat measurements until the state is pure enough and then rotates the state such that we are nearest to the ground state of the qubit (corresponding to
). Then, the qubit is allowed to thermalize, absorbing heat from the bath and consequently decreasing its purity. The process is then repeated.
As we could expect, for equal measurement strength κ, measuring along σz and σx leads to similar control strategies, with the state either moving along the y or x axis (see blue and green dots in figure 4). Measuring along σz displays a small advantage since we can skip the unitary rotation if the measurement process naturally drives the state near the ground state. On the contrary, if we measure σx, we always have to rotate the state since the Hamiltonian and the measurement operator do not commute.
Notably, the optimal control strategy strongly depends on the measurement strength κ. For nearly projective measurements (κ = 0.99), shown in figures 4(a) and (b), the RL agent naturally discovers a known control strategy described in [138]. This corresponds to the limit where, after a single measurement, we are guaranteed to have a pure state. Therefore, assuming we are in a thermal state, it is always sufficient to perform one measurement step (single blue dot), then one unitary step (green dots), followed by a thermalization step (orange dot). As discussed above, the unitary step can be skipped in the σz-measurement case if the post-measurement state is already the ground state of the qubit [orange dot in panel (b)].
However, as we move to weaker measurements (see figures 4(c)–(f)), the agent starts repeating multiple measurements (blue dots) before performing a unitary rotation (green dots). In this regime, the state of the qubit experiences a random walk in state space until it reaches a target purity (which is automatically learned by the RL agent). This is clearly visible in figure 5, where a trajectory of visited states (represented by ρx) and corresponding actions (colors) are plotted as a function of time for the σx-measurement case. As in figure 4, every row corresponds to a different value of κ (qualitatively similar results are found in the σz case, see appendix
Figure 5. Plots of the state component ρx, as a function of time, corresponding to the results shown in figures 4(a), (c) and (e) relative to the σx measurement case. An arbitrary trajectory is shown. The colors of the dots indicate the actions chosen by the RL agent at a given moment.
Download figure:
Standard image High-resolution imageAnother interesting feature is the change of target purity (measured, e.g. as the norm of the Bloch vector) before rotating the state as κ varies. This can be seen noticing how the green dots move towards the center in figure 4 and towards
in figure 5 as we decrease the measurement strength. Since reaching a target purity requires more and more time as the measurement gets weaker, the RL agent learns to ‘save time’ by choosing a lower and lower target purity at which the state is rotated and then thermalized.
5.2. Discrete, adaptive measurements
In this section, we consider the same setup and model as in the previous subsection. However, we explore the possibility of learning the best observable to measure instead of fixing it to either σz or σx. As we will see, the optimal feedback control strategy involves adaptively changing the measurement basis according to the acquired information. Therefore, during each measurement step, we consider a measurement of σθ (see equation (7)) and allow the agent to choose any value of
. Furthermore, since the agent chose to rotate the qubit along the negative z-axis in all previous results, we here fix the unitary step to be a rotation to the negative z-axis and remove the freedom of choosing the rotation angle. The thermalization step is the same as in section 5.1.
In figure 6, we plot the results we find in the same style as figure 4, but now the measurement action is not plotted as a blue dot, but as a color between blue and red depending on the learned angle (see color bar). For κ = 0.99 (figure 6(a)), the RL agent only chooses σz measurements, and thus finds the same cycle as in figure 4(b).
Figure 6. Actions chosen by the RL agent as a function of qubits state represented as a point on the Bloch sphere in the measurement-dominated regime with discrete, adaptive measurements. The action is represented by the color of a point (see legend). When a measurement is performed, we allow the RL agent to choose the angle of discrete measurement (see color bar). The black cross is the thermal state. Each plot corresponds to a different measurement strength and corresponds to a trajectory of 10 000 time-steps interacting with the RL agent. The same parameters were used as in figure 4.
Download figure:
Standard image High-resolution imageHowever, as we decrease the measurement strength, we see the emergence of points whose color interpolates between blue and red. This implies that the agent is performing different measurements based on the conditional state of the qubit. It is thus not optimal to fix a single measurement basis, but we can increase the demon’s performance by varying it adaptively. To prove that this is indeed the case, in figure 7, we plot the cooling power as a function of κ, fixing σx measurements (red dots), σz measurements (blue dots), and with a tunable observable (green dots). In the inset, we further plot the performance improvement ratio
(
) i.e. the ratio between the power
delivered by the agent with learnable angles, and the power
(
) fixing the measurement basis to σz (σx). Indeed, the ratio is greater or equal to 1 in the majority of the cases, reaching up to 30% improvements at κ = 0.7 with respect to measurements of σx. Interestingly, in all cases, the power decreases roughly linearly with the measurement strength parameter κ.
Figure 7. Comparison of the cooling power for fixed measurements of σx (red dots), fixed measurements of σz (blue dots), and for learnable measurements of σθ (green dots). The inset contains the ratio between the cooling power
with a learnable observable, and the cooling power
(
) with fixed measurement of σx (σz). Every dot is computed averaging over 100 000 time steps. The agent was trained for the same set of parameters as in figures 4 and 6.
Download figure:
Standard image High-resolution image
We notice that the same general feedback control strategy, detailed as bullet points in section 5.1, also emerges in this case. This consists of a series of measurements until a target purity is reached, followed by a unitary rotation and a thermalization step. Also in this case, as the measurement strength decreases (i.e. as κ goes from 1 to
), the target purity reached when performing a unitary rotation (green dots) is progressively smaller and smaller in order to limit the duration of the measurements.
However, the novelty here is in the choice of the measurement basis. As we can see from figures 6(c)–(f), the agent chooses a measurement angle θ (blue to red dots) roughly perpendicular to the direction of the conditional qubit’s state on the Bloch sphere. Indeed, when the state only has a ρz component, we see red dots corresponding to measurements along σx. As the state acquires a ρx component, also the measurement angle changes (purple and blue dots).
This effect can be interpreted in the following way: as shown in appendix
of the qubit’s state is given by

where
,
, and α is the angle between the state on the Bloch sphere the measurement axis. As can be easily seen, the average purity
is maximized for
, which is indeed the measurement perpendicular to the qubit’s state. Interestingly, the RL agent automatically learns to choose adaptively the measurement strategy that allows to maximize the purification of the state and thus the cooling power. To visualize this, let us consider figure 6(c). When we are in the thermal state (black cross), the agent chooses a red action corresponding to a measurement of σx, which is perpendicular to the qubit state. The state then stochastically evolves along the red arrows. Then, a blue measurement is performed, corresponding to a σθ measurement, which is again roughly perpendicular to the qubit state. The state then evolves stochastically along the blue arrows. Now, the target purity is high enough, so the agent either decides to thermalize the state (orange dots and arrows) directly, or to rotate it (green dots and arrows) and then thermalize it (orange arrow). For lower values of the measurement strength, we see that the variation of the state after each measurement becomes smaller and smaller, and more measurements are performed. These measurements are roughly perpendicular to the qubit state, although we find small deviations that actually deliver marginally higher performance; see appendix
So far, we have studied the maximization of the power. We conclude this section by assessing the impact of optimizing the trade-off between cooling power and measurement efficiency. In figure 8, we show the Pareto front describing the trade-off between the average cooling power and measurement dissipation (panel (a)) and between power and efficiency (panel (b)) fixing κ = 0.90. The points tend to cluster in three different regions. It indicates that the RL agent abruptly changes the control strategy as we shift our interest between power and efficiency; these strategies are explicitly shown in appendix
Figure 8. Maxwell’s demon performance in the measurement-dominated regime with discrete, adaptive measurements. Pareto front between the average cooling power
and negative dissipation
(a), and between
and the measurement efficiency η (b). Each dot corresponds to a separate RL optimization for different values of c. The red (blue) dots correspond to RL strategies that turn out to only measure σx (σz), and the green dots that include both σx and σz measurements. We fix κ = 0.9, and the other system parameters are as in figure 4.
Download figure:
Standard image High-resolution image5.3. Continuous measurements
We now consider the influence of continuous measurements instead of discrete measurements. As we will see, we find similar results as in the discrete case, but because of the continuous measurement outcomes, many more qubit states are visited, resulting in a ‘smoothed’ version of the discrete measurement case. Specifically, we model the thermalization and unitary parts as in the previous section (section 5.2), and we model the measurements using the continuous measurement formalism described in equation (8).
As we did in figure 4 for discrete measurements, in figure 9, we plot our results for the continuous measurement case using the same plotting style. In particular, the states and actions chosen by the agent are plotted for different values of the measurement strength
by row (larger
corresponds to stronger measurements), and for the σx and σz cases by column. In appendix
Figure 9. Actions chosen by the RL agent as a function of qubit’s state represented as a point on the Bloch sphere in the measurement-dominated regime with continuous measurements. The type of action is represented by the color of a point. The black cross denotes the thermal state. The measurement is fixed and given by σx (left column) or σz (right column), and the measurement strength
is reported on the plot. Each plot is generated by a trajectory with 10 000 time steps. The optimization was performed with parameters
,
,
.
Download figure:
Standard image High-resolution imageInterestingly, we find feedback control strategies similar to the discrete measurement case, which can thus be understood and interpreted in a similar way. However, because of the continuous measurement outcomes, the qubit visits a much larger range of states. This results in qualitatively similar plots comparing figure 9 to 4, but the points appear now to be ‘smeared out’.
To better visualize the random walk performed by the state during the measurement process, in figure 10, we plot the ρx component of the state, as a function of time, for the σx measurement case corresponding to figures 9(a) and (c) (see appendix
). For weaker measurements (figure 10(b)), an average of almost three consecutive measurements is observed before performing a unitary operation and thermalization, resulting in the random-walk behavior that we previously observed in the discrete measurement case.
Figure 10. Plots of the x-coordinate of the qubit, as a function of time, along an arbitrary trajectory considering continuous σx measurements studied in figures 9(a) and (c). The colors of the dots indicate the discrete action chosen by the RL agent (see legend). Each row corresponds to a different measurement strength
.
Download figure:
Standard image High-resolution image6. Measurement and thermalization dominated regime
Here we study the regime where both the measurement and thermalization timescales are comparable and much slower than the unitary dynamics, i.e.
. We combine the modelization of finite-time thermalization from section 4 with the continuous measurement studied in section 5.3. As we show, the RL agent naturally finds solutions that display features from these two distinct regimes, i.e. optimized thermalization strokes and multiple measurements, leading to a random walk in state space. However, the line-shape and the duration of the thermalization strokes are now conditioned on the (stochastic) purity reached after various consecutive measurements.
We consider the Hamiltonian of equation (22) and describe the thermalization of the qubit using the same master equation as in section 4. We recall that the thermalization speed is described by the rate Γ and that the agent can arbitrarily tune the gap of the qubit u(t) during the thermalization. We consider continuous measurements of σz, as described in section 5.3; we recall that the measurement strength is determined by the ratio
. The unitary part consists of evolving the system according to equation (22) but, as in section 4, the coherence of the qubit is not relevant. Therefore, the unitary evolution can be considered to be the fastest timescale.
In figures 11(a) and (b), we plot the RL results for
, i.e. for rather strong measurements, whereas in figures 11(c) and (d) we consider
corresponding to a weak measurement. In both cases, we plot an arbitrary trajectory showing the control u(t) as a function of time in panels (a,c), and the state ρz, as a function of time, in panels (b) and (d).
Figure 11. Qubit gap u(t) during thermalization (a), (c) and qubits z coordinate (b), (d) as a function of time along an arbitrary trajectory in the measurement and thermalization dominated regime. The RL agent is allowed to thermalize the qubit and perform continuous measurement of σz with strength quantified by the ratio
reported on the plots. Parameters used to generate the policy:
,
,
,
.
Download figure:
Standard image High-resolution imageAs expected, the control strategy shown in figure 11(a), corresponding to strong measurements, resembles the results in figure 3(c), where the measurement was projective. Indeed, the feedback control strategy consists of performing a single measurement, followed by thermalizing the qubit while continuously ramping the gap of the qubit and choosing the correct sign of u(t) based on the measurement outcome. However, here we witness rare cases in which a single measurement outcome is inconclusive, i.e. it does not result in a purification of the state (see, e.g. the two nearly consecutive blue dots at
). In such cases, the agent decides to perform a second measurement after a single thermalization time-step, since the measurement did not sufficiently purify the state as in the other measurements shown.
As we reduce the measurement strength, we observe multiple repeated measurements, producing a random walks in state space, that tend to purify the state (blue lines and dots in figures 11(c) and (d)). This is followed by smooth thermalization strokes (orange dots in figures 11(c) and (d)), whose duration and line-shape depend on the purity reached during the previous measurements. On average, such measurements are repeated consecutively 2.5 times. Interestingly, there are also cases in which the measurement decreases the purity of the state (see, e.g. at
, where
decreases after a blue dot, i.e. after a measurement).
7. Two qubits setup: all timescales are comparable
Here we show that our method can also be applied to the considerably harder problem when all three timescales are comparable (
) in a two-qubit setting, allowing to find interpretable strategies that considerably outperform more intuitive feedback control strategies. By extending our quantum system from a one- to two-qubits setup, we allow the RL agent to explore novel physical effects that can emerge in larger systems. In particular, the RL agent can now generate entanglement between the qubits and exploit it to optimize its performance. We analyze the relation between the Maxwell’s demon performance, and the generation of entanglement between the qubits considering both the absence and presence of counter-rotating terms in the qubit–qubit interaction. To this end, we adopt the following description.
As in the previous sections, we consider a main (m) qubit that can be coupled to a thermal bath (see figure 12). However, we model finite-time measurements by considering a second auxiliary (a) qubit that can interact with the main qubit through finite-time Hamiltonian dynamics and can then be measured projectively. This can be seen as a finite-time modelization of a non-projective measurement, where correlations are established between the qubits through unitary dynamics, and then information about the main qubit is acquired by projectively measuring the auxiliary qubit.
Figure 12. Schematic representation of the Maxwell’s demon setup studied in section 7 when all timescales
are comparable. When thermalizing, the main qubit is decoupled from the auxiliary qubit and put in contact with a thermal bath with inverse temperature β. During the unitary evolution, the two qubits are coupled together and interact. Projective measurements are performed only on the auxiliary qubit thus only partial information about the system is acquired.
Download figure:
Standard image High-resolution imageWe model the three discrete actions in the following way: unitary dynamics is described by the Schrödinger equation with the following time-dependent Hamiltonian

which describes two resonant interacting qubits [42]. Here u(t) is the time-dependent control that we allow the agent to optimize in a continuous interval
, and g represents the interaction strength. We consider two possible interactions


This corresponds, respectively, to neglecting (
) and considering (
) the counter-rotating terms
. The coupling
, for example, can be obtained by inductively coupling two superconducting qubits [156]. Further, as shown in [155], the interaction between the two qubits can be suitably programmed through the parametric driving of the qubits. During the thermalization stroke, we decouple the two qubits, and we describe the thermalization of the main qubit using the same model as in section 4, where we recall that the thermalization rate is given by Γ. The measurement stroke is also described as in section 4, i.e. performing a projective measurement along the z-axis, but here the measurement is only performed on the auxiliary qubit.
Before showing the results using the RL optimization, we devise an interpretable strategy inspired by the RL results that we use to benchmark the performance of the RL method. The Hamiltonian
, with
, has an interesting property: let us assume that the state of the joint system is given by

with
or
. It can be shown analytically that, after a time

the state of the two qubits exactly swaps up to a global phase. We, therefore, consider the following ‘interpretable feedback control strategy’:
- Perform a measurement of the auxiliary qubit.
- If it is in the ground state, do nothing.
- If it is in the excited state, change the sign of u(t), such that now it is in the ground state.
- Since the post-measurement state is of the form of equation (28), we perform a unitary stroke of duration
keeping the control u(t) constant at some value
. - The main qubit will be in the ground state because of the swap. We thus let it thermalize with the bath at the same value of the control
, for time
, resulting in heat extraction from the thermal bath.
This interpretable strategy is then optimized numerically with respect to
and
.
7.1. Without counter-rotating terms
In figure 13, we compare the performance of the RL agent against the interpretable strategy choosing
. From figures 13(a) and (b), we see that the Pareto front of the RL strategy (black dots) is visibly better than that of the ‘intuitive strategy’ (red dots), both in the high power and high-efficiency regime.
Figure 13. Performance of the Maxwell’s demon model described in section 7, when all timescales are finite, and the interaction between the two qubits is described by
. In panels (a), (b), the black dots correspond to the RL results, while the red one corresponds to the interpretable strategy described in section 7. The empty dots in panels (a) and (b) correspond to the two control strategies plotted respectively in panel (c) (corresponding to c = 1) and panel (d) (corresponding to c = 0.7). The colors in (c), (d) correspond to the three discrete actions (see legend). Parameters:
,
,
,
, and
for
and
for
.
Download figure:
Standard image High-resolution imageIn figures 13(c) and (d), we show a trajectory of the feedback control strategy as a function of time, whose performance is shown in figures 13(a) and (b) as white dots. We notice that it is similar in spirit to the ‘intuitive strategy’ since we verified that the duration of the unitary stroke (green dots) is approximately
. However, the RL agent learns again to modulate the gap of the qubit while thermalizing, and this gives us an advantage that is substantially larger than the advantage observed in section 4 where we considered the thermalization-dominated regime. As expected, the thermalization strokes slow down as we shift our interest from high power (figure 13(c)) to high efficiency (figure 13(d)).
To further interpret the RL results, in figure 14, we have plotted, respectively, the control u(t) (a), the only non-null component ρz of the Bloch vector of the main qubit tracing out the auxiliary qubit (b), and the concurrence
(c) between the two qubits in the maximum power case (c = 1) shown in figure 13(c). This plot sheds light on the cooling mechanism and allows us to interpret what the RL agent is doing. As we see in panel (b), during the unitary strokes (green dots), the purity of the qubit’s state 1 increases, almost reaching 1. During this process, entanglement is developed between the qubits, as seen in panel (c). Once the purity is high enough, the agent thermalizes the qubit, cooling the heat bath and destroying entanglement between the qubits. The auxiliary qubit is then measured again, and the process is repeated. Interestingly, based on the outcome of the measurement process, the control u(t) may change sign, leading to a different trajectory also of the state, and to a different amount of entanglement between the qubits.
Figure 14. Analysis of the control strategy shown in figure 13(c). The control u(t), the only non-null component ρz of the reduced density matrix of the main qubit, and the concurrence
between the qubits are respectively plotted in panels (a)–(c) as a function of time considering an arbitrary trajectory. The color corresponds to the discrete action shown in the legend.
Download figure:
Standard image High-resolution image7.2. Including counter-rotating terms
We now consider the effect of the presence of the counter-rotating terms in the coupling between the two qubits, i.e. we consider the interacting Hamiltonian given by
in equation (27). In figure 15, we plot the results for this case in the same style as figure 13. Because of the counter-rotating terms, the interaction between the qubits will not perfectly swap the qubit states, making this a harder optimization problem. As we can see from figures 15(a) and (b), the RL strategy produces a substantially higher power and higher efficiency than the interpretable strategy. As expected, the thermalization timescale increases as we shift our interest from high power (figure 15(c)) to high efficiency (figure 15(d)), and the RL agent learns a non-intuitive unitary stroke to purify the state of the main qubit (green dots).
Figure 15. Performance of the Maxwell’s demon model considered in figure 13, but including the counter-rotating terms in the interactions, i.e. using
instead of
. Plotting style and parameters are the same as in figure 13, except for the choice of Δt give by
for
;
for c = 0.85;
for
; and
for c = 0.65.
Download figure:
Standard image High-resolution imageInterestingly, comparing figure 15(b) with figure 13(b), we see that the best performance curve of the demon using RL, i.e. its Pareto front, is roughly the same with and without counter-rotating terms. Conversely, the performance of the interpretable strategy is decreased by the presence of the counter-rotating terms since they do not allow a perfect swap of the qubit states.
To conclude, we analyze the effect of the RL control strategy in the maximum power case (c = 1) in figure 16, as we did in figure 14 for the case without counter-rotating terms. Interestingly, the cooling mechanism seems quite similar regardless of the presence of counter-rotating terms. Indeed, after measuring the auxiliary qubit, the agent purifies the state (green dots in 16(b)), and entanglement is generated during this process (green dots in figure 16(c)). Once the state is pure enough, the agent thermalizes the main qubit, extracting heat from the bath and making sure that the main qubit is in the ground state (by choosing appropriately the sign of u based on the measurement outcome). However, based on the measurement outcome, the amount of entanglement generated during the unitary feedback can be lower than in the case without counter-rotating terms (compare the green dots in figures 14(c) and 16(c)).
Figure 16. Analysis of the control strategy shown in figure 15(c) plotted in the same style as figure 14.
Download figure:
Standard image High-resolution image8. Conclusions and outlook
In this work, we have introduced a RL-based framework to discover optimal feedback control strategies in open quantum systems, such that the RL agent operates as an actual Maxwell’s demon.
This compelling feature, on the one hand, allows us to literally regard the agent in RL as a Maxwell’s demon that acquires thermodynamically relevant information about the system, given it an ontological interpretation, providing an appealing picture.
On the other hand, this approach serves as a tool to come up with strategies that seem hardly possible to identify with conventional tools. The approach taken allows us to systematically maximize the multi-objective performance of a finite-time quantum Maxwell’s demon, which aims to maximize the information-induced average cooling power while minimizing the measurement cost. We study different physical regimes characterized by a different ordering between the measurement timescale, the unitary feedback timescale, and the thermalization timescale. Our framework optimizes the system’s long-term time-averaged and trajectory-averaged performance. We find feedback control strategies that are interpretable yet outperform more intuitive strategies, shedding light on important aspects that should be considered when designing a quantum Maxwell’s demon.
When the thermalization timescale is the slowest, the RL optimization results in non-trivial control strategies that outperform more intuitive ones. Moreover, we show that the duration of the thermalization stroke is dictated by the trade-off between cooling power and measurement post.
When the measurement timescale is the slowest, we found that the RL method first purifies the state through repeated measurements, leading to a random walk in the qubit’s state space. In this process, we consider both discrete and continuous weak measurements. This is followed by unitary feedback and thermalization. Notably, we show that adaptively choosing different measurement observables leads to a performance enhancement with respect to fixing the measurement. We further showed significant changes in the optimal measurement observable as we shifted our interest from high cooling power to low measurement cost.
We then consider the case when thermalization and measurement timescales are comparable and slower than the unitary feedback timescale. After performing multiple measurements, the RL agent learns carefully modulated thermalization ramps whose line-shape and duration depends on the purity reached during the previous measurements.
Finally, we have demonstrated that our method can be applied to a considerably more complex problem of two qubits when all three timescales are comparable, and that we can substantially outperforms more intuitive strategies. To this end, we considered two interacting qubits: one acting as the main qubit that can thermalize with the bath but not be measured, and the other one acting as an auxiliary qubit that we can measure and that indirectly provides us with information about the main qubit through finite-time interactions. We find intriguing and highly counter-intuitive feedback control strategies, where entanglement between the qubits is generated and then destroyed by measurements and thermalizations.
The introduction of RL methods to discover feedback control strategies in quantum thermodynamics paves the way for numerous future research directions. The impact of measurements and feedback on the performance of driven quantum heat engines and refrigerators operating between two different temperatures can be rigorously assessed, potentially revealing performance enhancements induced by quantum measurements. The impact of such feedback control schemes on power fluctuations can also be assessed and compared to recent results on the thermodynamic uncertainty relations [157]. Building on our two-qubit example, many-body interacting systems could also be optimized using this approach, revealing the impact that local or global measurements on many-body systems could have on the thermodynamic performance of such systems.
Furthermore, the use of advanced neural network architectures, such as the one proposed in [130], could allow the RL agent to learn how to act as an optimal quantum Maxwell’s demon by interacting directly with an experimental device without even knowing the exact model describing the dynamics of the system. This relies on using the time-series of chosen actions, rather than the quantum state, as state for the RL agent. While the size of the Hilbert space scales exponentially with the system size, this approach only scales with the number of controls, therefore scaling favorably in the system size, enabling also the theoretical optimization of substantially larger systems.
Finally, small adaptations of our RL framework could enable further research in quantum thermodynamics. For instance, by not revealing the measurement outcomes to the RL agent, we could study quantum thermal machines that exploit the invasive nature of quantum measurements in the absence of feedback, as studied e.g. in [43, 158]. Instead, by removing the possibility of performing quantum measurements, we could adapt our framework to discover control strategies that accelerate the thermalization of quantum systems, allowing us to compare our finding to the so-called shortcuts to thermalization [100, 159].
Acknowledgments
PAE gratefully acknowledges funding by the Berlin Mathematics Center MATH+ (AA2-18). JE has been supported by the DFG (FOR 2724, CRC 183), the BMBF (QSolid), and the ERC (DebuQC). RC, BB and ANJ acknowledge the support of Chapman University, U. S. Army Research Office under Grant W911NF-22-1-0258 and the U.S. Department of Energy (DOE), Office of Science, Basic Energy Sciences (BES), under Award No. DESC0017890. J.E. acknowledges funding by the DFG (FOR 2724, for which this is an inter-node collaboration reaching an important milestone, and CRC 183), the FQXI, the Quantum Flagship (Millenion, for which is again the result of an inter-node collaboration), the BMBF (DAQC), and the ERC (DebuQC). FN gratefully acknowledges funding by the BMBF (Berlin Institute for the Foundations of Learning and Data-BIFOLD), the European Research Commission (ERC CoG 772230) and the Berlin Mathematics Center MATH+ (AA1-6, AA2-8). Numerical calculations have been performed using PyTorch [151] and the QuTiP2 toolbox [160]. We thank Ketevan Kotorashvili for helping with the figure design.
Data availability statement
The reinforcement learning code, together with scripts to generate the results presented in the manuscript, are available on GitHub (https://github.com/QuantumQist/artificially_intelligent_maxwells_demon).
Appendix A: Pareto-optimal feedback controls for power and efficiency are also Pareto-optimal for power and dissipation
Here we show that finding the Pareto front for power and dissipation, e.g. by maximizing equation (17), we have also found the Pareto front for power and efficiency. We use an argument similar to that discussed in [107, 131], showing that if a feedback control policy is Pareto-optimal for power and efficiency, then it is Pareto optimal for power and dissipation.
Let us denote with
and
respectively the power and dissipation of a given feedback control policy π. Introducing the function

the efficiency of the feedback control strategy π is given by
.
Let us now assume that π is an optimal feedback control policy for power and efficiency, i.e. that
is Pareto optimal for power and efficiency. We claim that this implies that
is Pareto-optimal for power and dissipation. If, by contradiction,
was not Pareto optimal for power and dissipation, then there would be another policy
such that
is strictly better than
. This means that either

or

If equation (A2) is true, using the monotonicity of
with respect to
and
, and using that
and
if the system is operating as a Maxwell’s demon, we have that

Equations (A2) and (A4) imply that the feedback control policy π has a strictly worse efficiency and at most equal power than
, which in turn implies that π is not Pareto-optimal for power and efficiency. This contradicts our hypothesis. A similar argument can be used if equation (A3) is true, concluding our proof by contradiction.
Appendix B: Measurement dominated regime: discrete σz measurements
In this appendix we show additional results for the model and regime studied in section 5.1, i.e. the measurement dominated regime where discrete measurements in a fixed basis are performed. While figure 5 in the main text shows a trajectory of visited states and corresponding actions as a function of time for the σx-measurement case, in figure 17 we show the same plot, but for the σz measurement case. Each row corresponds to a different value of κ, and here the state is represented by the ρz component of the qubit state. As we can see, the σx- and σz-measurement cases display qualitatively similar features.
Figure 17. Plots of the state component ρz, as a function of time, corresponding to the results shown in figures 4(b), (d) and (f) relative to the discrete σz measurement case in the measurement dominated regime. Each row corresponds to a different value of the discrete measurement strength κ. An arbitrary trajectory is shown. The colors of the dots indicate the actions chosen by the RL agent at a given moment.
Download figure:
Standard image High-resolution imageAppendix C: Measurement dominated regime: discrete adaptive measurements. Role of the measurement basis
In this appendix, we first show that a measurement perpendicular to the state on the Bloch sphere is the one that maximizes the average increase in purity of the state. We then show how, in fact, the RL agent chooses a measurement angle that is not exactly perpendicular to the qubit’s state. While this may seem like an imperfection of the RL method, we show that this actually leads to a (very small) improvements in the performance.
We start by proving that the measurements that maximize the average purity of the state are perpendicular to the qubit state on the Bloch sphere. Let us consider an arbitrary qubit state of the form

since the measurements and dynamics considered in this manuscript never drives it onto σy.
Let us consider a weak measurement with measurement angle θ described by the Kraus operators
defined in equation (7). For later convenience, let us define the vectors

where θ is the angle defined below equation (6), and
describes the axis along which we are measuring.
Using equation (5) and the properties of the Pauli matrices, the probability of measuring outcomes ± can be written as

The post-measurement purity of the state, conditioned on the measurement outcome, is given by

where we have used equation (6) to compute the post-measurement state and the cyclic property of the trace. Inserting the explicit expressions for ρ and
, we can rewrite it as

Using the properties of the Pauli matrices, we can rewrite the three traces in equation (C5) as



Plugging these expressions back into equation (C5) leads to

This leads to the final expression for the average purity
given by

where we defined
, and where α is the angle between the state on the Bloch sphere and the measurement axis, explicitly defined by

Since
,
is trivially maximized by minimizing
, i.e. when
corresponding to

This proves that perpendicular measurements maximize the average post-measurement purity.
We now show that the RL agent actually decides to perform measurements that are not exactly perpendicular to the qubit’s state on the Bloch sphere, and we show that this indeed leads to a slight performance enhancement.
Let’s consider the policy shown in figure 6(c) that can be described by the following procedure:
- 1.Initialize the qubit in the thermal state.
- 2.Perform σx measurement of the qubit. Determine the post-measurement state coordinate ρx.
- 3.If
, perform the second measurement of σθ with
, where θ is the measurement angle that enters the measurement operator in equation (7). If
, perform the second measurement of σθ with
. The angles
can be determined from the analysis of the RL policies. If the measurement increased the qubit’s coordinate ρz, proceed to step 4. Otherwise, skip step 4. - 4.Perform the unitary feedback, i.e. rotate the qubit to the negative z-axis.
- 5.Thermalize the qubit, leading us back to step 2.
Interestingly, the RL method finds the same exact procedure also for
and 0.85. We thus implement numerically the policy above, comparing the resulting expected cooling power when the angles
are chosen to maximize the post-measurement purity, i.e. perpendicular to the state on the Bloch sphere, and when the angles are chosen according to the RL results. Out of symmetry considerations, we assume
.
The measurement angles
in the RL-case are chosen according to

where the angles θj are collected from the simulations performed by trained RL agents.
In figure 18, we plot the ratio between the power
using the RL-chosen measurement angle, and the power
found choosing the angle that maximizes the average post measurement purity, for
. As we can see, the RL-chosen angles provide a small (
) yet measurable improvement of the cooling power. This indicates that the choice of measurement angle in step 3 is mainly, but not exclusively, dictated by the intent of maximizing the average post-measurement purity of the qubit.
Figure 18. Ratio of the cooling powers for the policy given in appendix C when the measurement angle is based on RL (
), and when it is chosen to be perpendicular to the state on the Block sphere (
), i.e. to maximize the average post-measurement purity. The ratio larger than one indicates a slight advantage of the strategy based on RL optimization. All system parameters are the same as figure 6, except for considering a full thermalization during the thermalization action (corresponding to the
limit).
Download figure:
Standard image High-resolution imageAppendix D: Measurement dominated regime: discrete adaptive measurements. Power-efficiency trade-off
At the end of section 5.2, we asses the impact of power-efficiency trade-offs in the measurement dominated regime when the RL agent can adaptively choose the measurement basis. In particular, in figure 8 we showed a Pareto front displaying different power-efficiency trade-offs, and we claim that differently colored points correspond to different measurement strategies.
In figure 19 we explicitly show how the optimal strategy changes as we change the trade-off between power and efficiency, i.e. as we change the parameter c. Indeed, in figure 19(a) we see that, for c = 1, the agent performs both σx (red dot) and σz (blue dot) measurements. This corresponds to the green points along the Pareto front of figure 8. In figures 19(b) and (c), corresponding to c = 0.9 and c = 0.65, we see that the agent only performs σz measurements. This corresponds to the blue points along the Pareto front of figure 8. In figure 19(d), corresponding to c = 0.6, we see that the RL agent only performs σx measurements, corresponding to the red points along the Pareto front of figure 8.
Figure 19. Actions chosen by the RL agent as a function of the qubit’s state represented as a point on the Bloch sphere in the sames style as figure 6. Each panel corresponds to a representative point along the Pareto front in figure 8 computed with a different value of c shown on the plot.
Download figure:
Standard image High-resolution imageAppendix E: Measurement dominated regime: continuous measurements
In this appendix we show additional results for the model and regime studied in section 5.3, i.e. the measurement dominated regime where continuous measurements are performed in a fixed measurement basis. In particular, we consider the same model as in section 5.3, but we additionally discuss the case where the measurement basis is chosen by the RL agent, and we show additional plots for the σz-measurement case.
In figure 20 we show the results allowing the RL agent the freedom of choosing the measurement angle, plotted in the same style as figure 9. We observe that the strong measurement case (figure 20(a)) corresponds to the cooling scheme powered by continuous σz measurement (blue points). Reducing the measurement strength resulted in the emergence of qubit states for which the σx measurement is performed, marked by the red points in figure 20(b).
Figure 20. Actions chosen by the RL agent, as a function of qubit’s state represented as a point on the Bloch sphere, plotted in the same style and with the same model as figure 9, but here the RL agent can choose the measurement basis (blue to red dots). Each panel represents a different value of the continuous measurement strength
shown on the plot.
Download figure:
Standard image High-resolution imageWhile in figure 10 we showed a trajectory in the σx measurement case, in figure 21 we additionally show a trajectory of the ρz component of the state, as a function of time, corresponding to figures 9(b) and (d). As we can see, the results are qualitatively similar to the σx measurement case.
Figure 21. Plots of the state component ρz, as a function of time, corresponding to the results shown in figures 9(b) and (d) relative to the σz continuous measurement case. Each row corresponds to a different value of the continuous measurement strength
. An arbitrary trajectory is shown. The colors of the dots indicate the actions chosen by the RL agent at a given moment.
Download figure:
Standard image High-resolution imageAppendix F: Reinforcement learning (RL) algorithm
In this appendix we provide specific information on the RL algorithm that was used, and details of the optimization carried out in the manuscript.
We employ a generalization of the soft actor-critic (SAC) method, first introduced in [148–150] for continuous actions only, to a combination of discrete and continuous actions [128, 130, 161, 162]. In particular, we use the method and code that was implemented in [130] but, aside from minor implementation details that will be described below, we only change the neural network architectures that describe the policy function and the value function. Therefore, in the following we provide an overview of the method, we describe the differences compared to [130], but we refer to [130] for a more detailed explanation of the method. Our code is implemented in PyTorch [151] and is based on numerous modifications of the SAC implementation provided by Spinning Up from OpenAI [163]. For each of the different physical cases discussed in this manuscript, we implemented an RL environment that computes the state evolution based on the chosen action, and computes the rewards. These calculations were performed using the QuTiP2 toolbox [160]. For code and data to reproduce our results, see the Code and Data Availability Statement.
F.1. SAC algorithm
As described in equation (18) of the main text, the goal of RL is to determine the policy function
that maximizes the expected sum of future rewards, i.e.

where Eπ denotes the expectation value over the state evolution and over the actions chosen according to the policy π, and
is sampled from the steady-state distribution of states µπ visited using π. In the SAC method, equation (F1) is solved using the idea of policy iteration [147], i.e. iterating over two steps: the policy evaluation and policy improvement step. In the policy evaluation step, the quality of the current policy is evaluate by estimating its value function
(the critic). In the policy improvement step, the value function is used to improve the policy function (the actor). In practice, we start from an initially random policy, apply it for some steps onto the RL-E, use the collected experience to perform policy evaluation and policy improvement, then we collect new observations using the new policy, and so on, until convergence.
The SAC method balances exploration and exploitation [147] using an entropy regularized maximization objective [148–150], which corresponds to maximizing a trade-off of the sum of future rewards, and the entropy of the policy. Since we are dealing both with discrete and continuous actions, we use the formulation detailed in [130]. We decompose the policy function
, where
is a combination of a discrete (d) and continuous (u) action, using the following identity

Here
is the marginal probability of choosing the discrete (D) action d, whereas
is the conditional probability of choosing the continuous (C) action u, given that the discrete action d was chosen. Let us denote as

the entropy of a probability function (or probability density in the continuous case) P(x). Using the decomposition in equation (F2), we can write the entropy of the policy function as

where

represent, respectively, the entropy of the discrete actions and the average (differential) entropy of the continuous action.
Our aim is now to solve the following equation [130]

where
are two ‘temperature’ parameters that balance the trade-off between exploration and exploitation. Indeed, for
, equation (F6) coincides with the standard RL optimization objective of equation (F1). For
, we are finding trade-offs between maximizing the rewards, and increasing respectively the entropy of the continuous and discrete policy. Since the entropy is an indicator of the randomness of a probability distribution, the larger are
and
, the more we are incentivizing an explorative behavior of the agent. Notice that in equation (F6) replaced the (unknown) state distribution µπ with a first-in-first-out replay buffer
that is populated during training by storing the observed one-step transitions
.
The parameters
and
are automatically tuned during training to reach, respectively, a target entropy
and
of the continuous and discrete policy functions. Then, during training, we decrease them to switch from a more explorative to a more deterministic behavior.
Another important function in actor-critic methods is the value function
[147]. As discussed in [130], from equation (F6) we define the value function as

which not only represents the sum of future rewards that we would obtain starting from state s, performing action a, and then choosing all future actions according to π, but it also incorporates the entropy terms. This function plays the role of the critic, and gives us a measure of the expected quality of a given policy π.
In the SAC method we use neural networks to parameterize the policy function
and the value function
.
The value function is represented by a neural network
that depends on a set of trainable parameters φ. As done in [128], we use a feed-forward neural network that takes (s, u) as input, and outputs three values corresponding to the value of
, for
. In particular, we use the ReLU function as non-linearity, two hidden layers, and a linear output layer. Since the state s coincides with the density matrix of the system, we expand the density matrix onto the computation basis, separate the real and imaginary parts, and arrange all elements into a vector. Although not strictly necessary, we further concatenate the last chosen continuous action u to the state vector.
The policy function is described by the two functions
and
, where θ is a set of trainable parameters. In particular, we first choose a discrete action d according to the discrete distribution
, and then we choose as
a squashed Gaussian policy, i.e. the distribution of the variable

where
and
represent respectively the mean and standard deviation of the Gaussian distribution,
is the normal distribution with zero mean and unit variance, and
. This is the so-called reparameterization trick. The functions
,
and
are computed from a trainable neural network in the following way. We use a single feed-forward neural network that takes the state as input (in the same way as described for the value function), and outputs nine parameters:
,
and
for
. More specifically, we use two hidden layers, the ReLU function as non-linearity, and then a linear layer to output
and
, and a soft-max function to output
. Then
, and we compute

where
is a numerical safety parameter, to ensure that the standard deviation is positive.
To determine the value function
, the SAC introduces two value functions
and
, and their parameters φi, for
, are determined minimizing the following loss function [130]

where

and where
, for
, are parameters that are not updated when minimizing the loss function; instead, they are constant during backpropagation, and then they are updated according to Polyak averaging, i.e.

where
is a hyperparameter. The parameters defining the policy function
are determined minimizing the following loss function [130]

Finally, to prevent the temperature parameters from becoming negative, we express them as
and
, and we determine the
,
parameters minimizing the following loss functions [130]

and

F.2. Training details
We now provide additional details on the actual implementation and training. As mentioned previously, we store all one-step transitions
in a first-in-first-out replay buffer
of fixed dimension from which batches one one-step transitions are sampled to estimate the loss functions. In order to start from an exploratory behavior, as in [130] we first choose actions randomly (according to a uniform distribution) for a fixed number of initial time-steps. Furthermore, for a different number of fixed initial steps, we do not perform any updates to the value function, the policy function, nor to the temperature parameters. This is to allow the replay buffer to have enough transitions before sampling transitions from it. After this initial phase, we repeat nupdates optimization steps of all quantities, every nupdates steps such that, overall, the number of updates and of time-steps coincide. In particular, we use the ADAM optimizer [164] with default hyperparameters and learning rate reported in tables 1 and 2 to minimize
, and stochastic gradient descent with default hyperparameters and learning rate reported in tables 1 and 2 to minimize
and
. As mentioned in the previous section, we transition from an exploratory to a more deterministic behavior during training by scheduling the target entropy terms
and
as follows

where
, nsteps is the current step number, and
,
and
are hyperparameters. Furthermore, in order to have hyperparameters for the target entropy that do not depend on the size of the continuous action interval
, when computing the entropy of the policy for the loss functions, we assume that the continuous actions always lie in the interval
.
Table 1. Training hyperparameters used for each figure reported in the top row. The “ symbol means the same hyperparameter as to its left.
| Hyperparameter | Figure 3 ![]() ![]() | Figure 3 ![]() | Figures 13 and 14 ![]() ![]() ![]() | Figures 13 and 14 c = 0.65 | Figures 15 and 16 ![]() ![]() | Figures 15 and 16 ![]() ![]() |
|---|---|---|---|---|---|---|
| Batch size | 256 | “ | “ | “ | “ | “ |
| Training steps | 160k | 320k | 160k | “ | 320k | “ |
| ADAM learning rate | 0.001 | “ | “ | “ | 0.0003 | “ |
| SGD learning rate | 0.003 | “ | “ | “ | “ | “ |
size | 80k | “ | “ | 160k | 80k | 180k |
![]() | 0.995 | “ | “ | “ | “ | “ |
| Hidden layers size | (256, 128) | “ | “ | “ | (256, 256) | “ |
| Initial random steps | 5k | “ | “ | “ | “ | “ |
| First update at step | 1k | “ | “ | “ | “ | “ |
| nupdates | 50 | “ | “ | “ | “ | “ |
![]() | 0.8 | “ | “ | 0.77 | 0.74 | “ |
![]() | −3 | “ | “ | “ | “ | “ |
![]() | 60k | “ | “ | 120k | 180k | “ |
![]() | ![]() | “ | “ | “ | “ | “ |
![]() | 0.01 | “ | “ | “ | “ | “ |
![]() | 30k | “ | “ | 60k | 90k | 140k |
| γ | 0.998 | “ | “ | “ | “ | “ |
Table 2. Training hyperparameters used for each figure reported in the top row. The “ symbol means the same hyperparameter as to its left.
| Hyperparameter | Figures 4–8, 17 and 19 | Figures 9(a) and (c), 10 | Figures 9(b) and (d), 21 | Figure 11 | Figure 20 |
|---|---|---|---|---|---|
| Batch size | 128 | 256 | “ | “ | “ |
| Training steps | 250k | “ | “ | 300k | 250k |
| ADAM learning rate | 0.0008 | 0.0003 | “ | “ | “ |
| SGD learning rate | 0.001 | “ | “ | “ | “ |
size | 80k | 100k | 80k | “ | “ |
![]() | 0.995 | “ | “ | “ | “ |
| Hidden layers size | (128, 128) | “ | “ | “ | “ |
| Initial random steps | 5k | “ | “ | “ | “ |
| First update at step | 1k | “ | “ | “ | “ |
| nupdates | 50 | “ | “ | “ | “ |
![]() | 0.8 | “ | “ | “ | “ |
![]() | −3 | “ | “ | “ | “ |
![]() | 100k | “ | 80k | “ | “ |
![]() | ![]() | “ | “ | “ | “ |
![]() | 0.01 | “ | “ | “ | “ |
![]() | 100k | “ | 80k | “ | 100k |
| γ | 0.998 | “ | “ | “ | “ |
All hyperparameters used during training are reported in tables 1 and 2. Most results were found on the first run, except for figures 15, 16 and 20 where the optimization was sometimes repeated a few times to improve the result. The values of c used for the multi-objective optimizations are all reported in table 1, except for figure 8, where we have used
, and multiplied the cooling power by a factor 35 and the dissipation term by a factor 4 during the optimization of
. When the continuous control
corresponds to the measurement angle θ, we found that the RL method converges to better results if we choose an interval
that is larger than
. In particular, we used
and
respectively when optimizing the discrete and continuous measurement case.






























































