Dynamic Portfolio Optimization under CVaR Constraints
Abstract
We study continuous-time dynamic portfolio optimization under a Conditional Value-at-Risk (CVaR) constraint on the investor’s terminal loss. For a general class of convex trading objectives, we exploit the auxiliary-threshold representation of CVaR to establish the existence of an optimal strategy and strong duality without requiring market completeness. These results motivate a dual-based nested bisection–golden-search algorithm over the threshold and Lagrangian multiplier, where the inner iterations reduce to standard unconstrained stochastic control problems. We prove that the resulting strategies converge to the optimal control as the number of iterations tends to infinity. Numerical experiments recover the Merton policy when the risk constraint is nonbinding. When the constraint is binding, the optimal strategy becomes state dependent: the investor reduces risky exposure following adverse outcomes but preserves, and near maturity may increase, exposure following favorable outcomes. Thus, a terminal CVaR constraint produces an asymmetric reallocation across states rather than uniform de-risking. Nontraded endowment risk amplifies the conservative adjustment, whereas price impact lowers desired positions and adjustment speeds.
Keywords: Dynamic portfolio optimization; Conditional Value-at-Risk; Constrained stochastic control; Lagrangian duality; Incomplete markets
JEL Classification: G11; C61; C63
1 Introduction
Financial investors seek to generate returns while managing the risks associated with their investment strategies. Merton’s foundational analyses of continuous-time portfolio selection and the efficient portfolio frontier established canonical benchmarks for this trade-off 23; 24. Subsequent works have extended portfolio choice in several directions, including market models with price impact (e.g. 36; 11), trading costs (e.g. 9; 26), robust and risk-aware decision criteria (e.g. 19), and benchmark-relative portfolio optimization (e.g. 27).
Downside-risk bounds are particularly relevant in portfolio selection when regulatory mandates or trading-desk requirements impose direct limits on losses in adverse states. Unlike conventional risk aversion, such bounds target the downside tail directly. Conditional Value-at-Risk (CVaR), also known as Expected Shortfall, is a downside-tail risk measure that averages losses in a prescribed tail of the loss distribution 2; 34. In static portfolio optimization, CVaR may either be minimized as the risk objective 33 or imposed as a constraint on an otherwise return-oriented objective 21; 3. Related discrete-time multiperiod formulations include CVaR-based risk control in 7 and dynamic mean–CVaR portfolio selection with a focus on time consistency in 35. In the continuous-time setting, one strand of the literature studies the cases where the problem can be recast as the choice of a replicable terminal payoff, due to additional assumptions such as market completeness 16; 14. Another strand considers self-financing portfolio choice under dynamically re-evaluated risk constraints and exploits specialized growth-optimal or CRRA structures 31; 25.
A common way is to incorporate downside risk through a soft penalty in the investor’s objective. Although convenient, such a formulation requires the investor to specify an exogenous penalty coefficient, and an arbitrary choice of this coefficient does not ensure compliance with a prescribed risk budget. In many institutional, regulatory, and wealth-management applications, however, the relevant mandate is an explicit upper bound on downside risk. Motivated by this consideration, we treat the terminal CVaR bound as an exogenous risk-management requirement rather than a preference penalty incorporated into the performance objective. Hard-constraint formulations also arise more broadly in stochastic control, including problems with expectation and terminal-law constraints (e.g. 30; 10), and in mean-field games with state or other feasibility constraints (e.g. 6; 18).
In this paper, we study a general continuous-time dynamic portfolio optimization problem subject to a hard CVaR constraint on terminal loss. Our framework does not require self-financing wealth dynamics or market completeness and accommodates cumulative external cash flows and nontraded endowment risk. In this general setting, the problem cannot be reduced to a static choice of terminal payoff: nontraded risks need not be replicable, while general running costs make the entire state–control path relevant. In the frictional extensions, the current position becomes an additional state variable and trading speed becomes the control, so the timing and speed of portfolio adjustment also affect optimality. The optimal strategy must therefore be characterized dynamically rather than recovered solely from an optimal terminal payoff.
To analyze this general portfolio optimization problem under hard CVaR constraints, we use the Rockafellar–Uryasev representation of CVaR to obtain a convex primal–dual formulation over the adapted trading strategy and an auxiliary scalar threshold. This formulation preserves the dynamic nature of the portfolio problem while yielding a modular solution approach based on standard unconstrained stochastic control problems. It provides the basis for both our theoretical analysis and the computational method developed below.
Contributions.
Our contribution to this line of research is threefold.
First, we provide a convex-analytic treatment of continuous-time portfolio optimization under a hard terminal CVaR constraint in a market that may be incomplete. The investor minimizes a general convex expected running and terminal cost over admissible adapted trading strategies. The framework covers, among other examples, linear–quadratic and CARA specifications, as well as extended-valued CRRA and logarithmic criteria whenever the corresponding finite-feasibility condition is satisfied. We show that the CVaR threshold can be restricted to a common compact interval and establish existence of a primal optimizer. Under strict convexity, the optimal trading strategy is unique. We further establish existence of Lagrangian minimizers for every multiplier. Under a Slater condition, we prove strong duality, existence and an explicit bound for an optimal multiplier, and recovery of a primal optimizer from the dual problem. These results do not require market completeness.
Second, we exploit the resulting max–min dual structure to develop a modular numerical method. For each fixed Lagrange multiplier and CVaR threshold, the innermost problem is a standard unconstrained stochastic control problem, allowing existing dynamic-programming, PDE, or stochastic-control solvers to be used as an oracle. An inner golden-section search selects the threshold, while an outer bisection adjusts the multiplier using the CVaR constraint residual. We quantify the error of the inner search and prove convergence of the combined procedure when the inner accuracy is increased appropriately relative to the outer bisection. In particular, the constraint residual converges to zero and the objective value converges to the primal optimum. Under strict convexity, the computed controls converge weakly to the unique optimal strategy; strong convexity strengthens this conclusion to strong convergence in .
Third, we use numerical experiments to examine the effect of a binding terminal CVaR constraint in four environments: a frictionless complete-market benchmark, an incomplete market with nontraded endowment risk, a model with quadratic trading-rate regularization, and a model with square-root price impact. The first three settings remain within the convex structure motivating our analysis, while the square-root specification provides a robustness experiment beyond the setting covered by our convergence theory. Three empirical patterns emerge across the reported calibrations. First, the CVaR constraint reshapes the terminal-wealth distribution asymmetrically, with the adjustment concentrated in the downside tail rather than taking the form of a uniform contraction. This illustrates the distinction from conventional risk aversion: CVaR targets adverse tail outcomes directly rather than penalizing risk throughout the distribution. Second, the response combines an initially conservative adjustment with state-dependent feedback. Although the constraint is imposed only at maturity, when it binds, current wealth becomes a relevant state variable for the exposure policy: the computed strategies reduce risky exposure following adverse outcomes while maintaining greater participation following favorable outcomes. Third, in the reported zero-interest, no-positive-inflow experiments, the unconditional loss-CVaR diagnostic remains below the terminal risk limit at every reported intermediate date across all four environments. Although the constraint is imposed only at maturity, the behavior of the unconditional interim CVaR diagnostic suggests that, in these experiments, the terminal constraint may also discipline risk taking earlier in the investment horizon.
More broadly, our work is connected to the literature on portfolio optimization under alternative downside-risk and distributional constraints. Related formulations impose VaR constraints 4; 32; 8 or the closely related Capital-at-Risk constraint 12, formulate portfolio choice in terms of wealth quantiles 17, employ dynamic multivariate risk measures 13, or impose utility-based shortfall-risk constraints 15. Our analysis also relates to benchmark-relative portfolio formulations that restrict terminal payoffs through divergence constraints, including Bregman–Wasserstein constraints 29 and their asymmetric -Bregman–Wasserstein extension 28.
The remainder of the paper is organized as follows. Section 2 introduces the market, endowment, objective, and terminal CVaR constraint. Section 3 develops the primal and dual formulations. Section 4 presents the nested search algorithm and its convergence analysis. Section 5 studies the complete- and incomplete-market benchmarks, and Section 6 adds quadratic trading rate regularization and nonlinear price impact. Section 7 concludes.
2 Market Model and CVaR-Constrained Portfolio Problem
This section formulates the dynamic portfolio management problem studied in the paper. We first introduce a general continuous-time market with traded assets, cumulative endowment, and potentially nontraded background risk. Portfolio strategies are evaluated through a running and terminal cost criterion and are required to satisfy a CVaR constraint on terminal losses. This formulation accommodates incomplete markets and provides the setting for the reformulation and algorithm developed in the subsequent sections.
We then specialize the general model to a one-stock linear–quadratic Black–Scholes benchmark with terminal CVaR constraints. The benchmark is analytically transparent: in the absence of a binding CVaR constraint, the optimal risky-asset holding reduces to the classical Merton dollar exposure, while an active constraint changes the feedback strategy to protect the lower tail of terminal wealth. This structure also provides a common framework for the numerical experiments with endowment risk and trading frictions.
We work on a filtered probability space satisfying the usual conditions, where denotes the investment horizon. Let , be a standard -dimensional Brownian motion adapted to , and let be a 1-dim Brownian motion on that is independent of .
The market contains one risk-free asset and risky assets. The risk-free asset earns a constant interest rate . The risky asset price vector satisfies
| (2.1) |
Here the is the vector of excess dollar returns, and is the volatility matrix. Since may be smaller than , the risky assets need not span all Brownian shocks.
The investor receives a cumulative endowment with dynamics
| (2.2) |
The finite-variation process captures deterministic or predictable cash flows, including lump-sum payments, while the predictable row vector captures the endowment’s exposure to Brownian shocks. Only the component of lying in the span of the traded volatility matrix can be hedged by the risky assets. In contrast, the remaining component , driven by the independent Brownian motion , is nontraded background risk and is a source of market incompleteness in the model.
We assume that the filtered probability space supports a process satisfying (2.1), and we fix this weak solution throughout the paper. We impose the following regularity conditions on the exogenous coefficients and endowment processes:
Assumption 2.1 (Regularity of exogenous processes).
The functions and are jointly continuous. In addition, and are uniformly bounded, and for every , is positive definite. The process is predictable and of finite variation, with , and and are predictable. Moreover, there exists such that
Let denote the portfolio of the investor, where represents the number of shares held in the –th risky asset. Available to the investor are trading strategies that are -valued predictable processes , where denotes the dollar risky exposure in the –th risky asset. The remaining wealth is invested in the risk-free asset, with the controlled wealth process follows
| (2.3) |
where denotes the initial wealth of the investor.
Next, we denote by the set of admissible strategies as follows:
Assumption 2.2.
The set of admissible strategies is
where the action set is nonempty, compact, and convex.
The investor evaluates a strategy through the cost functional
| (2.4) |
where is the running cost function and is the terminal cost function. As the investor aims to minimize the cost functional (2.4), the utility maximization framework fits within our setting by treating costs as negative rewards.
In addition to the cost functional, the investor requires a risk constraint on a function applied to the terminal portfolio wealth. The risk constraint is given by the Conditional Value-at-Risk (CVaR) (1). The CVaR at a confidence level for the loss random variable is defined as
| (2.5) |
where the Value-at-Risk at a confidence level of is given by
| (2.6) |
Given a confidence level and a risk limit , the investor requires a strategy that satisfies
| (2.7) |
For the choice , the CVaR constraint limits the average severity of the worst fraction of terminal losses to . Summarizing, the investor’s CVaR-constrained portfolio problem is therefore
| (2.8) | ||||
| subject to | ||||
Throughout the paper, we impose the following regularity conditions on the objective and terminal loss functions.
Assumption 2.3 (Regularity of cost functionals).
The functions satisfy the following conditions.
- (i)
The function is jointly Borel measurable and, for almost every , the mapping is proper, convex, and lower semicontinuous. Moreover, there exists such that for almost every and all .
- (ii)
The function is proper, convex, and lower semicontinuous.
- (iii)
The terminal loss is convex and lower semicontinuous. Moreover, for the same as in Assumption 2.1, there exists such that
Remark 2.4.
Together with Assumptions 2.1 and 2.2, Assumption 2.3iii ensures that is integrable for every . Indeed, the standard moment estimate for the wealth equation, together with the boundedness of and and the compactness of , gives for every . Assumption 2.3 iii then implies , and hence is integrable. A uniform square-integrability estimate is established in Proposition 3.2.
Remark 2.5.
Assumption 2.3ii ensures that there exist such that , for all . Together with Assumption 2.3i on and the uniform first-moment bound on admissible wealth processes, this yields . We henceforth fix such that for any
Allowing the objective costs to take the value permits domain restrictions in utility-based criteria. In particular, the framework accommodates CARA terminal costs, as well as extended-valued CRRA and logarithmic costs whenever a finite-cost feasible strategy exists.
For the existence of a primal optimizer, we require the feasibility of the optimization problem.
Assumption 2.6 (Feasibility).
There exists at least one admissible strategy such that
A sufficient condition for the feasibility assumption is that , that is the no-trade wealth
satisfies , and the corresponding objective is finite.
The formulation in Section 2 is deliberately stated at a general level. It accommodates alternative objective specifications, including CARA terminal costs and, whenever the effective-domain and finite-feasibility requirements are satisfied, CRRA and logarithmic costs, as well as richer market dynamics and additional state variables to include consumption control or market frictions. The reformulation and algorithm developed in Sections 3 and 4 are driven primarily by the convex structure of the CVaR representation and do not depend on the particular objective function or market specification.
3 Constrained Convex Optimization Reformulation
In this section, we reformulate the risk-constrained stochastic control problem (2.8) as a convex constrained optimization problem. The main device is the auxiliary-variable representation of CVaR, which converts the terminal risk constraint into a convex constraint in both the control and an additional scalar threshold variable. This formulation then allows us to introduce a Lagrangian dual problem and establish strong duality under an additional Slater-type condition.
3.1 Reformulation and Primal Problem
By Remark 2.4, is integrable for every . Hence, its CVaR at level admits the representation introduced in 33:
| (3.1) |
The infimum is attained, and its minimizers are the generalized -quantiles of . Further, for each , define
| (3.2) |
Then the constraint is equivalent to the existence of such that . Thus the portfolio optimization problem can be written as
| (P) |
Note that the scalar variable is not a trading decision. It is an auxiliary threshold variable introduced by the CVaR representation. This formulation separates the portfolio control from the scalar risk threshold , while preserving convexity.
We now present the main analytical properties of the primal formulation. These properties will be used later to derive the Lagrangian dual problem and the dual-based numerical algorithm.
Proposition 3.1 (Convexity of the primal problem).
Proof.
For any and , define . By the linearity of the state equation and the uniqueness of solutions,
Since is jointly convex and is convex, we have
and similarly . Taking expectations gives . Therefore, the objective function is convex in the control on .
It remains to show that is convex. Let . Since is convex and is affine in ,
Therefore,
Because is convex and nondecreasing, it follows that
Taking expectations and adding the affine term shows that is jointly convex. Hence the feasible set is convex, and the primal problem is convex. ∎
We next establish a uniform square-integrability estimate for the terminal losses and use it to restrict the auxiliary threshold to a common compact interval that does not depend on the constraint level .
Proposition 3.2 (Compactification of the CVaR threshold).
Proof.
By the standard moment estimate for SDEs with coefficients of linear growth, see 22 or (20, Theorem 2.1), for from Assumption 2.1, there exists , independent of , such that
Since , we have , and therefore . Hence, there exists that only depends on the compact set and exogenous market coefficients, such that .
Now fix and write The minimizers of are the -quantiles of , namely the points satisfying
If then
Thus cannot be an -quantile. Similarly, if then
Hence , so cannot be an -quantile. Therefore every minimizer lies in , and the infimum over equals the infimum over .
Proposition 3.3 (Existence of a primal optimizer).
Proof.
By Assumption 2.6 and Proposition 3.2, the feasible set of the compactified problem is nonempty. Define
Since is compact, the set is bounded in . Moreover, is convex by the convexity of , and it is closed: if in , then, up to a subsequence, -a.e.; since is closed, a.e. Hence . Thus is closed, bounded, and convex in the Hilbert space , and is therefore weakly compact. Since is compact, is weakly compact.
It remains to show that is weakly closed. By Proposition 3.1, is convex. Moreover, by the linearity of the controlled wealth dynamics (2), we can see that for ,
which implies
| (3.3) |
where only depends on , , the set , and the uniform bounds of and . The above stability estimate (3.3) for the wealth equation, together with the lower semicontinuity of and Fatou’s lemma, implies that is lower semicontinuous in the strong topology of . Since a convex strongly lower semicontinuous functional is weakly lower semicontinuous, is weakly lower semicontinuous. Therefore the sublevel set of
is weakly closed. Hence is weakly compact.
Let be a minimizing sequence, so that By weak compactness of , there exist a subsequence, still denoted by , and a pair such that
We next verify that is weakly lower semicontinuous. Suppose first that strongly in . By (3.3),
Passing to a subsequence, the corresponding convergences hold almost everywhere.
Since with , the function is nonnegative. The lower semicontinuity of and Fatou’s lemma therefore yield
Moreover, since is proper, convex, and lower semicontinuous, there exist such that , for all . Applying Fatou’s lemma to the nonnegative function , and using the convergence of , gives
Thus is strongly lower semicontinuous.
Since is convex by Proposition 3.1, it is weakly lower semicontinuous. Hence
Since is feasible, the reverse inequality follows from the definition of . Therefore and is a primal optimizer.
Finally, Proposition 3.2 shows that restricting to does not change the value of the primal problem. Hence the same pair is also an optimal solution of the original primal problem over . ∎
Finally, we establish the uniqueness of the optimal control, under strictly convexity assumption.
Assumption 3.4.
The objective functional is strictly convex on ; that is, for any distinct satisfying and , and any ,
In the following, we provide some explicit conditions on the problem parameters that can guarantee the above the strict convexity of .
Proposition 3.5.
Suppose that Assumptions 2.1, 2.2, and 2.3 hold. Then the objective functional is strictly convex on , if one of the following conditions hold:
- (i)
For almost every , the function is strictly convex on its effective domain, in the sense that, for every ,
whenever and both and are finite.
- (ii)
There exists such that , for all , and the function is strictly convex on its effective domain. More precisely, for every ,
whenever and both and are finite.
Condition ii covers several standard expected-utility objectives. In particular, when , it applies to the negative exponential (CARA) utility
as well as to the lower-semicontinuous extended-valued negative CRRA cost. For , define
and define the negative CRRA cost as the lower-semicontinuous extension of to , namely
For every , is proper, convex, lower semicontinuous, and strictly convex on its effective domain. The limiting case at is the negative logarithmic utility .
Proof.
Let be distinct in and satisfy . Let , and define . Since is convex, . Moreover, the wealth equation is affine in the control, so
Since in , the set has positive measure. On this set, . If Condition (i) holds, then
By the convexity of ,
Combining the two inequalities gives .
If Condition ii holds, let be distinct in and satisfy . Define , . Since the endowment terms cancel, the wealth difference satisfies
Let . Applying Itô’s formula to and taking expectations gives
By Young’s inequality,
Together with , and , this yields
Since , multiplying by and integrating gives
| (3.4) |
Since in , it follows that . Hence, by the strict convexity of ,
By the convexity of ,
Combining the two inequalities also implies . Thus, is strictly convex on . ∎
Proposition 3.6 (Uniqueness of the optimal control).
Proof.
Let and be two primal optimizers. For , define
By convexity of the feasible set, is feasible. If , by strict convexity of ,
This contradicts the optimality of and . Therefore . ∎
Remark 3.7 (Uniqueness of the CVaR threshold).
The auxiliary variable is not unique in general, since the objective does not depend on . For a fixed optimal control , the set of admissible thresholds is
Thus is unique only if this set is a singleton. In particular, if the CVaR constraint is active and , then must be a minimizer of . In this case, uniqueness of is equivalent to uniqueness of the -quantile of the terminal loss
A sufficient condition is that the distribution function of is strictly increasing in a neighborhood of its -quantile. If the constraint is slack, or if the quantile set contains an interval, then need not be unique.
3.2 Dual formulation
We introduce the Lagrangian dual problem associated with the compactified primal problem (3.2).
For and , define the Lagrangian
The corresponding dual function is
| (3.5) |
and the dual problem is
More explicitly,
Remark 3.8.
We now establish the property of the dual formulation (3.2) and its connection with the primal formulation (3.2).
Proposition 3.9 (Existence of Lagrange minimizers).
Under Assumptions 2.1, 2.2, 2.3, and 2.6, for every , the inner problem of (3.2), , admits a minimizer.
Suppose, in addition, that Assumption 3.4 holds. Then, for every , the Lagrangian is strictly convex in its control component on its effective domain: for any with in and , and any ,
Consequently, all Lagrangian minimizers at a fixed have the same control component.
Proof.
For fixed , the Lagrangian is given by
By the compactification result, the minimization is over . As in the proof of Proposition 3.3, is weakly compact and is compact. Moreover, and are weakly lower semicontinuous, and therefore is weakly lower semicontinuous on . Hence the infimum is attained.
If Assumption 3.4 holds, is strictly convex on , while is jointly convex in . Hence, for every , the Lagrangian
is strictly convex in its control component. Therefore, two Lagrangian minimizers at the same multiplier cannot have different control components, by an argument similar to the proof of Proposition 3.6. ∎
We next present a weak duality result, which states that the dual value is not larger than the primal value.
Proposition 3.10 (Weak duality).
Proof.
Let be primal feasible. Then . Since ,
Taking the infimum over feasible gives . Taking the supremum over gives . ∎
To establish strong duality, we require the following Slater’s condition. This condition strengthens Assumption 2.6 to strict feasibility, which is standard in the optimization literature on guarantee strong duality (see e.g., 5).
Assumption 3.11 (Slater condition).
There exist and such that
Remark 3.12.
Now we are ready to state the strong duality result.
Theorem 3.13 (Strong duality).
Proof.
Define the perturbation value function
Then . Since and are convex, is convex. Assumption 3.11 implies for all because is feasible for such . By the discussion in Remark 2.5, . Hence is finite on an open interval containing . As a finite convex function on an open interval, is continuous at and .
Let . Since relaxing the constraint can only decrease the value, is nonincreasing, and therefore . Define . The subgradient inequality gives, for all ,
For any , taking yields
where the first inequality holds because is feasible for the perturbed constraint level . Therefore
for all . Taking the infimum over gives . By Proposition 3.10, . Hence , so and is a dual optimizer. ∎
Similar to Proposition 3.2, we provide a bound for the multiplier , which is useful in the numerical algorithm.
Proposition 3.14.
Proof.
The following corollary explains how a primal optimizer can be recovered once a dual optimizer has been found. It also shows that, under strict convexity, the control component recovered from any Lagrangian minimizer is the unique primal optimal control.
Corollary 3.15 (Recovery of a primal optimizer).
Suppose Assumptions 2.1,2.2, 2.3 and 3.11 hold. Let be a dual optimizer. Suppose
and
Then is a primal optimizer.
If in addition Assumption 3.4 also holds, then any Lagrangian minimizer
recovers the unique primal optimal control.
Proof.
Since minimizes the Lagrangian at , By strong duality, Using complementary slackness,
Since , the pair is primal feasible and attains the primal value. Hence it is a primal optimizer.
It remains to prove the recovery statement. Let be a primal optimizer, whose existence follows from Proposition 3.3. Since is feasible,
On the other hand, since is dual optimal, strong duality gives , and by definition of ,
Therefore , so is a Lagrangian minimizer at .
By Proposition 3.9, under the strict convexity of , the control component of the Lagrangian minimizer at is unique. Hence any Lagrangian minimizer at satisfies . Thus is the unique primal optimal control. ∎
4 Algorithm and Analysis
4.1 Nested Bisection-Golden Search Method
In this section, motivated by the dual formulation in Section 3.2, we propose a nested bisection method for the portfolio optimization problem with CVaR constraint, by solving the compactified dual problem. By Propositions 3.2 and 3.14, it is enough to solve the dual problem
The proposed algorithm consists of a control oracle for fixed , an inner golden-section search over , and an outer bisection over .
Throughout this section, Assumptions 2.1–2.3, 3.11, and 3.4 are in force. Assumption 3.4 ensures that both and are strictly convex in the control component. Consequently, every fixed- control problem and every joint inner problem have a unique control component.
Control oracle.
For fixed , define the modified terminal cost
Then the fixed- subproblem is
| (4.1) |
The last term is independent of , so the control oracle only needs to solve a standard stochastic control problem with terminal cost . We denote its output by
where is the unique solution of (4.1) under Assumption 3.4, and is the corresponding optimal value.
Inner golden-section search over the CVaR threshold.
We next construct a method for evaluating the dual function defined in (3.5). For any fixed ,
Thus, evaluating reduces to a one-dimensional minimization of over the compact interval .
Proposition 4.1 (Properties of the inner value function).
Proof.
Let be the finite-cost strategy in Assumption 3.11. Since is finite, . Moreover, by Remark 2.5 and the nonnegativity of the hinge term,
Hence . For fixed , the affine dependence of on , the convexity of , , and , and the convexity of imply that is jointly convex. Partial minimization over therefore shows that is convex.
For fixed , define . Every subgradient of has an absolute value bounded by and hence . It follows that, uniformly over ,
Taking the infimum over in both directions gives
Thus is Lipschitz hence continuous. Since is compact, is nonempty and closed, and its convexity follows from the convexity of . ∎
Since is convex, it decreases before to the left of its minimizer set and increases to the right of it. This allows us to locate a minimizer using function values only. Starting from an interval containing , define
We evaluate at these two interior points using the control oracle
If , convexity implies that the interval contains a minimizer, so the right-hand portion may be discarded. Otherwise, the interval contains a minimizer, and the left-hand portion may be discarded. The points are chosen according to the golden ratio so that one previous function evaluation can be reused after each interval reduction. Thus, after the initial two evaluations, each iteration requires only one new call to the control oracle. The algorithm is summarized as follows.
Outer bisection over the Lagrange multiplier.
The inner golden-section search via Algorithm 1 approximately evaluates for each fixed multiplier . We now maximize the concave dual function over . The key observation is that its derivative is given by the constraint residual, which we define in the following two cases of :
For every , let
and define the residual
Under Assumption 3.4, the control component is unique. Although the optimizing threshold may not be unique, the value of the residual is unique. Indeed, if and are both optimal thresholds for the same control, then
Since , . Thus, is well defined.
For , we define as follows. Solving the unconstrained problem yields . Next, we choose
which yields the residual
If , then the unconstrained optimizer is feasible and the optimal multiplier is . Otherwise, the constraint is active, and we search for a root of on .
Proposition 4.2 (Dual residual).
Proof.
By Proposition 3.2, there exists such that
Together with the finite-cost strategy provided by Assumption 3.11 and the lower bound on , this implies that is finite. Moreover,
so is continuous.
Let , and choose a Lagrangian minimizer at each multiplier, using when . Optimality gives
and
Therefore,
| (4.2) |
In particular, is nonincreasing.
Next, we show continuity. Let monotonically and choose
By weak compactness, along a subsequence, and . The weak lower semicontinuity of the Lagrangian, the boundedness of , and the continuity of imply that
Suppose first that and write . Monotonicity gives . By the weak lower semicontinuity of , . If , then . If , the definition of gives . In either case, , and hence . Thus is right-continuous, including at zero.
Now suppose that with , and write . Monotonicity gives . Moreover,
Using the weak lower semicontinuity of gives
Since , this implies . Therefore , proving left continuity.
Finally, for , applying (4.2) to and yields
Letting and using right continuity gives . Similarly,
and left continuity gives . Thus is differentiable and for . ∎
In the active-constraint case, . Define the dual optimal set . By Theorem 3.13, Proposition 3.14, and the differentiability of , the set is nonempty and coincides with the set of dual optimizers. Since is nonincreasing, means that the multiplier is too small, so the search must move to the right. Similarly, means that the multiplier is too large, so the search must move to the left. Thus, the root of can be located by bisection.
For a queried multiplier , Algorithm 1 returns . We use the corresponding residual in the outer bisection. Starting from , the method evaluates at the midpoint of the current interval. If , it retains the right half; if , it retains the left half. The main algorithm is shown below.
4.2 Convergence Analysis
We analyze the performance of the main algorithm under exact evaluations of the control oracle. We first study the convergence of the inner golden-search over the CVaR threshold (Algorithm 1) and then establish the convergence of Algorithm 2.
For a nonempty set , we write
Proposition 4.3 (Convergence of the inner golden-section search).
Proof.
By Proposition 4.1, is convex and is a nonempty closed interval.
Consider a current search interval that intersects , and let be the two golden-section points. If , convexity implies that a minimizer exists in , so the algorithm may discard . Similarly, if , a minimizer exists in . Thus every update retains an interval intersecting .
Each golden-section update reduces the interval length by the factor . Therefore, after reductions, the remaining interval has length and still intersects . Since lies in this interval,
Theorem 4.4.
Suppose that Assumptions 2.1, 2.2, 2.3, 3.11, and 3.4 hold, and that the control, value, and residual evaluations used by the algorithm are exact. Consider the active-constraint case , and let be the output of the Algorithm 2 with outer iterations and inner iterations.
Let . If and , then
Moreover,
where is the unique primal optimal control.
Proof.
For every multiplier queried by the outer bisection, let denote the output of the inner bisection and the associated exact control oracle. Then by Proposition 4.3,
For any , the definition of and the affine dependence of the Lagrangian on the multiplier give
Among all multipliers queried during the outer iterations and at the final midpoint, the smallest possible distance from the endpoints of is . Hence every queried multiplier satisfies
Taking in the preceding inequality gives
Since is concave and ,
Similarly, taking and using concavity yields
Combining the two inequalities proves
Consequently,
By Proposition 4.2, with and , the residual is continuous on . Hence it is uniformly continuous. Therefore, if and , then
| (4.3) |
where the supremum is over the multipliers queried by the outer bisection.
We next prove that . Fix . Since is continuous and , compactness implies that
provided that the set over which the infimum is taken is nonempty. By (4.3), for sufficiently large and , and have the same sign at every queried midpoint satisfying .
We claim that, throughout the outer bisection, the current bracket intersects the closed -neighborhood of . This is true for the initial bracket . Suppose it holds for the current bracket, and let be its midpoint. If , then either half selected by the algorithm contains as an endpoint, and hence still intersects the -neighborhood of .
Otherwise, , so the approximate residual has the same sign as . Since is nonincreasing, its zero set is an interval. If lies to the left of the -neighborhood of , then , and the algorithm retains the right half of the bracket, which still intersects that neighborhood. Similarly, if lies to the right of the -neighborhood, then , and the algorithm retains the left half, which again intersects the neighborhood. The claim therefore follows by induction.
Consequently, the final bracket contains some point such that . Since also lies in the final bracket, whose length is at most ,
Taking the limit superior and then letting gives .
Because is continuous and vanishes on , . Applying (4.3) at the final multiplier yields
Applying Proposition 4.3 at gives
Therefore,
On the other hand, for any , strong duality implies
Since , we conclude that
It remains to prove the convergence of the controls. Since is weakly compact in and is compact, every subsequence of admits a further subsequence, still denoted by , such that
By the weak lower semicontinuity of and , , and . Thus is primal feasible and attains the primal value. Hence is a primal optimal control. By Assumption 3.4, the primal optimal control is unique, and therefore . Thus every weakly convergent subsequence of has limit . Since is weakly compact, it follows that the entire sequence satisfies
∎
The strict-convexity condition in Theorem 4.4 guarantees the weak convergence. Strong convergence follows under the following stronger condition.
Assumption 4.5.
is strongly convex on its effective domain with respect to the norm. That is, there exists a constant such that for any satisfying , , and ,
Similar to Proposition 3.5, we also provide some explicit conditions to guarantee the strongly convex property of .
Proposition 4.6.
Suppose that Assumptions 2.1, 2.2, and 2.3 hold. Then the objective functional is strongly convex on its effective domain, if one of the following conditions hold:
- (i)
The function is strongly convex. That is, there exist a constant such that, for almost every , every , and every ,
- (ii)
There exists such that , for all , and the function is strongly convex on . More precisely, there exists a constant such that for every and every ,
Proof.
Let satisfy , and let . Set . As in the proof of Proposition 3.5, the affine wealth dynamics imply .
Suppose first that Condition i holds. By the strong convexity of and the convexity of ,
Dropping the first nonnegative term gives
Hence is -strongly convex.
Corollary 4.7 (Strong convergence of the computed controls).
Proof.
Let . By strong duality and primal recovery, there exists such that minimizes . Since is jointly convex, is -strongly convex in its control component. Consequently,
Using , the right-hand side equals
which converges to zero by Theorem 4.4. Hence
∎
5 Numerical Experiments
This section evaluates Algorithm 2 in a scalar linear–quadratic Black–Scholes model. This benchmark isolates the dynamic effect of a terminal CVaR constraint, first in a complete market, and then in the presence of non-traded risks.
Let the traded stock be driven by a Brownian motion , and let be an independent Brownian motion driving the nontraded risk exposure. With ,
| (5.1) |
and the endowment dynamic satisfies . With denoting dollar risky exposure, wealth evolves as
| (5.2) |
The CVaR-constrained portfolio management problem is therefore:
| subject to | (5.3) |
For , the constraint requires average terminal wealth in the worst fraction of outcomes to be at least ; it is not a pathwise wealth floor. When the CVaR constraint is nonbinding, pointwise minimization gives the Merton dollar exposure
| (5.4) |
The independent endowment shock adds variance but does not change (5.4). When the CVaR constraint is binding, however, the investor may reduce traded exposure to offset its effect on the lower tail.
Table 1 reports the common calibration. The nonbinding and binding CVaR limits are and , respectively. For fixed , the inner dynamic program computes a Markov feedback control on a discretized state grid. We optimize by golden-section search and use an outer bisection in to enforce the CVaR constraint. Code and further implementation details are available at https://github.com/xf-shi/Dynamic-Portfolio-under-CVaR.
| Parameter | Value |
|---|---|
| Horizon | |
| Risk-free rate | |
| Excess return | |
| Volatility | |
| Initial stock price | |
| Risk aversion | |
| Initial wealth | |
| Merton exposure | |
| CVaR confidence level | |
| Nonbinding CVaR limit | |
| Binding CVaR limit |
5.1 Complete Market Benchmark
For the complete-market benchmark, we suppress the orthogonal factor and take the filtration to be generated by the traded Brownian motion . Equivalently, there is no nontraded endowment risk and the single traded risky asset spans the single source of market uncertainty. In the complete market, the nontraded risk exposure is . Table 2 shows that the constraint is nonbinding at : the Merton strategy has terminal loss CVaR and . In the binding case with , the calibrated multiplier is . Average risky exposure falls from to , whereas expected terminal wealth declines by only about .
| Case | average | ||||
|---|---|---|---|---|---|
| Nonbinding | |||||
| Binding |
Figure 1 shows that the policy for the binding case is not a constant rescaling of the Merton exposure. It returns toward along the selected favorable path, which ends at , but falls to about late in the selected unfavorable path, which ends at . Thus the binding constraint induces state-dependent de-risking while retaining participation in favorable states.
The policy for the binding case primarily compresses the lower tail of terminal wealth (Figure 2). The right panel also reports the unconditional diagnostic for . In this zero-rate, no-inflow calibration it remains below the numerical level throughout the horizon. This observation is not a dynamic constraint and need not persist under other calibrations.
This lower-tail reshaping is qualitatively related to complete-market mean-risk and quantile models (14; 16; 17). Here, however, the quadratic variation penalty regularizes the solution: the binding constraint leads to smooth de-risking that depends on the state rather than to a digital or lottery-like terminal payoff.
5.2 Incomplete Market
We set the nontraded risk exposure to be . In this incomplete market setting, the algorithm of Section 4 remains applicable.
In the nonbinding case, nontraded risk leaves the Merton exposure unchanged but moves terminal loss CVaR from to . In the binding case with , the calibrated multiplier is and average risky exposure falls to , which is below the complete-market value in the binding case and below the Merton exposure. Expected terminal wealth falls by about relative to the policy for the nonbinding case.
| Case | average | ||||
|---|---|---|---|---|---|
| Nonbinding | |||||
| Binding |
Figure 3 displays the resulting feedback response in the binding case. Along the selected favorable path, wealth ends at approximately and exposure returns to the Merton level after . Along the selected unfavorable path, wealth ends near and average exposure after falls to approximately . The investor therefore offsets unhedgeable tail risk indirectly by reducing traded risk in unfavorable states.
The distributional comparison in Figure 4 confirms that the policy for the binding case removes mass from severe low-wealth outcomes. The right-panel trajectories are interim diagnostics only; the optimization imposes the CVaR restriction at , not at intermediate dates.
6 Portfolio Adjustment and Price Impact
With the presence of trading regularization or existence of price impact, the dollar risky amount becomes a state variable. We keep to represent the dollar risky amount, and the control is referred to as the signed dollar trading rate, where the dynamic between the dollar risky amount and the control remains linear
| (6.1) |
where we recall that stands for the total shares hold by the investor, and represents her signed trading rate in shares. Although the results in Sections 2–4 are stated for a scalar wealth state and direct portfolio controls, the same arguments extend directly to finite-dimensional affine controlled state dynamics under the corresponding convexity, compactness, and integrability conditions. We use this extension in Section 6.1.
The control is used in both specifications below, but the economic interpretations and theoretical properties of the two specifications differ. In Section 6.1, the quadratic term is an objective regularizer that smooths portfolio adjustment; it is not an execution cost and is therefore not deducted from wealth. The augmented state dynamics remain affine in the trading-rate control, and the quadratic regularizer preserves the convex structure underlying the analysis in Sections 3 and 4. In Section 6.2, by contrast, price impact changes the execution price and the resulting cost is deducted directly from wealth. The resulting wealth dynamics depend nonlinearly on the trading-rate control through and therefore fall outside the affine-state setting of Sections 3 and 4. We use the square-root specification as a numerical robustness experiment beyond the setting covered by our convergence theory.
6.1 Quadratic Trading Rate Regularization
We first add a quadratic penalty on the signed dollar trading rate to smooth portfolio adjustment:
| (6.2) |
Since the regularization only shows up in the target functional, the wealth dynamics remain the same as (5.2). The CVaR-constrained control problem is
| subject to | (6.3) |
The term is a regularization term in the performance criterion, which penalizes abrupt changes in the risky position and produces smoother trading policies. We use the calibration from Table 1 and set ; larger values of place greater weight on smooth portfolio adjustment.
Table 4 shows that the CVaR constraint is nonbinding at . In the binding case, average exposure falls from to and rises to . Expected terminal wealth declines from to , while terminal loss CVaR reaches . The policy for the binding case uses the calibrated pair .
| Case | average | ||||
|---|---|---|---|---|---|
| Nonbinding | |||||
| Binding |
Figure 5 shows that both policies trade most rapidly near the initial date and that their mean trading rates approach zero near maturity. The policy for the binding case trades less aggressively and maintains a lower exposure profile. The exposure bands remain nondegenerate in the nonbinding case because contains the stock-price diffusion in (6).
In the binding case, terminal wealth is more concentrated and has less downside mass (Figure 6). The right panel reports as an interim diagnostic. The constraint is imposed at maturity.
The selected paths in Figure 7 illustrate the feedback mechanism under the binding constraint. The favorable path ends at wealth and approaches the Merton exposure, whereas the unfavorable path ends near and reduces exposure to approximately late in the horizon. The trading rates along both paths decrease toward zero as maturity approaches. In the nonbinding case, the trading rate converges to zero as maturity approaches because further repositioning offers no remaining expected-return benefit while still incurring the quadratic trading-rate penalty. In the binding case, the last reported preterminal rate remains slightly nonzero because the terminal-loss term associated with the CVaR constraint creates an additional incentive to adjust exposure over the final decision interval.
6.2 Square-Root Price Impact
We retain the same market and initial position, but now let execution price respond to signed dollar trading rate according to the square-root specification. When executing a stock, the execution price is usually different from the market price , i.e.
| (6.4) |
The placement of the brackets in (6.4) ensures that when . Since , execution costs reduce wealth as follows:
| (6.5) |
The corresponding portfolio management problem is
| subject to | (6.6) |
Although the cost is convex, its nonlinear appearance in the wealth dynamics places this specification outside the affine-state setting of Sections 2–4. The convergence guarantees established above therefore do not apply directly. We set and use the remaining parameters from Table 1.
In the nonbinding case, terminal loss CVaR is and the multiplier is zero. In the binding case, the policy reduces average exposure from to and mean terminal wealth from to , as shown in Table 5.
| Case | average | ||||
|---|---|---|---|---|---|
| Nonbinding | |||||
| Binding |
Figure 8 shows faster initial accumulation under the policy for the nonbinding case. Under the policy for the binding case, later trading rates can become negative in unfavorable states, indicating partial liquidation, while mean exposure stabilizes near . Diffusion in (6) generates widening exposure bands under both policies.
The policy for the binding case compresses the terminal-wealth distribution and reduces extreme outcomes (Figure 9). Out-of-sample terminal loss CVaR is , below the required level .
In Figure 10, the selected favorable path under the binding constraint ends at wealth and returns toward the Merton exposure. The selected unfavorable path ends near and reduces late-horizon exposure to approximately ; its trading rate is negative over several intervals. Trading rates approach zero near maturity in both paths.
The -power specification produces a mean exposure in the binding case close to that under quadratic regularization but a different adjustment pattern: it permits faster initial accumulation and uses state-dependent sales when downside risk increases. Its held-out feasibility provides numerical evidence that the method remains effective in this nonlinear price-impact specification, although the convergence guarantees do not apply directly.
7 Conclusion
This paper studies a dynamic portfolio management problem under an explicit CVaR constraint on terminal loss. Using the auxiliary-threshold representation of CVaR, we obtain a convex formulation that accommodates general market settings, including incomplete market, nontraded endowment risks, and general convex trading objectives. We establish existence and strong duality and develop a modular numerical method that combines standard unconstrained stochastic-control problems with one-dimensional searches over the CVaR threshold and Lagrange multiplier, which is shown to converge globally.
The numerical experiments show that a binding terminal CVaR constraint does not induce uniform de-risking. Instead, the optimal policy reduces risky exposure following adverse outcomes while preserving participation in favorable states. Nontraded endowment risk strengthens this adjustment when the constraint binds, whereas trading frictions lower the desired exposure and slow portfolio adjustment. Across the environments considered, stronger lower-tail protection is achieved at a comparatively modest cost in expected terminal wealth. These findings illustrate how a terminal risk constraint can shape state-dependent portfolio decisions throughout the investment horizon.
Acknowledgments
SP and XS acknowledge financial support from the Natural Sciences and Engineering Research Council of Canada (RGPIN-2025-05847, RGPIN-2024-04569). AH thanks the Fields Institute for the FOCUS visitor support program and acknowledges financial support by InnoHK initiative, The Government of the HKSAR, and Laboratory for AI-Powered Financial Technologies.
References
- [1] (2002) Portfolio optimization with spectral measures of risk. arXiv preprint cond-mat/0203607. Cited by: §2.
- [2] (2002) On the coherence of expected shortfall. Journal of Banking & Finance 26 (7), pp. 1487–1503. External Links: Document Cited by: §1.
- [3] (2004) A comparison of VaR and CVaR constraints on portfolio selection with the mean-variance model. Management Science 50 (9), pp. 1261–1273. External Links: Document Cited by: §1.
- [4] (2001) Value-at-risk-based risk management: optimal policies and asset prices. The Review of Financial Studies 14 (2), pp. 371–405. External Links: Document Cited by: §1.
- [5] (2004) Convex optimization. Cambridge university press. Cited by: §3.2.
- [6] (2018) Existence and uniqueness for mean field games with state constraints. In PDE models for multi-agent phenomena, pp. 49–71. Cited by: §1.
- [7] (2005) Multiperiod consumption and portfolio decisions under the multivariate GARCH model with transaction costs and CVaR-based risk control. OR Spectrum 27 (4), pp. 603–632. External Links: Document Cited by: §1.
- [8] (2008) Optimal dynamic trading strategies with risk limits. Operations Research 56 (2), pp. 358–368. External Links: Document Cited by: §1.
- [9] (2024) Dynamic mean-variance portfolio selection with transaction costs. Available at SSRN 4958481. Cited by: §1.
- [10] (2022) Optimal control of diffusion processes with terminal constraint in law. Journal of Optimization Theory and Applications 195 (1), pp. 1–41. Cited by: §1.
- [11] (2023) Optimal leveraged portfolio selection under quasi-elastic market impact. Operations Research 71 (5), pp. 1558–1576. External Links: Document Cited by: §1.
- [12] (2001) Optimal portfolios with bounded capital at risk. Mathematical Finance 11 (4), pp. 365–384. External Links: Document Cited by: §1.
- [13] (2017) A recursive algorithm for multivariate risk measures and a set-valued Bellman’s principle. Journal of Global Optimization 68 (1), pp. 47–69. External Links: Document Cited by: §1.
- [14] (2017) Dynamic mean-LPM and mean-CVaR portfolio optimization in continuous-time. SIAM Journal on Control and Optimization 55 (3), pp. 1377–1397. External Links: Document Cited by: §1, §5.1.
- [15] (2008) Utility maximization under a shortfall risk constraint. Journal of Mathematical Economics 44 (11), pp. 1126–1151. External Links: Document Cited by: §1.
- [16] (2015) Dynamic portfolio choice when risk is measured by weighted VaR. Mathematics of Operations Research 40 (3), pp. 773–796. External Links: Document Cited by: §1, §5.1.
- [17] (2011) Portfolio choice via quantiles. Mathematical Finance 21 (2), pp. 203–231. External Links: Document Cited by: §1, §5.1.
- [18] (2025) Mean-field games with constraints. arXiv preprint arXiv:2510.11843. Cited by: §1.
- [19] (2022) Robust risk-aware reinforcement learning. SIAM Journal on Financial Mathematics 13 (1), pp. 213–226. Cited by: §1.
- [20] (2014) On the -th moment estimates for the solution of stochastic differential equations. Journal of Inequalities and Applications 2014 (1), pp. 395. Cited by: §3.1.
- [21] (2002) Portfolio optimization with conditional value-at-risk objective and constraints. The Journal of Risk 4 (2), pp. 11–27. External Links: Document Cited by: §1.
- [22] (2007) Stochastic differential equations and applications. Elsevier. Cited by: §3.1.
- [23] (1969) Lifetime portfolio selection under uncertainty: the continuous-time case. The Review of Economics and Statistics 51 (3), pp. 247–257. External Links: Document Cited by: §1.
- [24] (1972) An analytic derivation of the efficient portfolio frontier. Journal of financial and quantitative analysis 7 (4), pp. 1851–1872. Cited by: §1.
- [25] (2013) CRRA utility maximization under dynamic risk constraints. Communications on Stochastic Analysis 7 (2), pp. 203–225. External Links: Document, Link Cited by: §1.
- [26] (2025) Dynamic portfolio choice with intertemporal hedging and transaction costs. Management Science. Note: Articles in Advance External Links: Document, Link Cited by: §1.
- [27] (2023) Portfolio optimization within a Wasserstein ball. SIAM Journal on Financial Mathematics 14 (4), pp. 1175–1214. Cited by: §1.
- [28] (2026) Outperforming a benchmark with -Bregman Wasserstein divergence. arXiv preprint arXiv:2603.20580. Cited by: §1.
- [29] (2024) Optimal payoff under Bregman-Wasserstein divergence constraints. arXiv preprint arXiv:2411.18397. Cited by: §1.
- [30] (2021) Duality and approximation of stochastic optimal control problems under expectation constraints. SIAM Journal on Control and Optimization 59 (5), pp. 3231–3260. Cited by: §1.
- [31] (2009) Maximizing the growth rate under risk constraints. Mathematical Finance: An International Journal of Mathematics, Statistics and Financial Economics 19 (3), pp. 423–455. Cited by: §1.
- [32] (2007) Portfolio optimization under the value-at-risk constraint. Quantitative Finance 7 (2), pp. 125–136. Cited by: §1.
- [33] (2000) Optimization of Conditional Value-at-Risk. The Journal of Risk 2 (3), pp. 21–41. External Links: Document Cited by: §1, §3.1.
- [34] (2002) Conditional Value-at-Risk for general loss distributions. Journal of banking & finance 26 (7), pp. 1443–1471. External Links: Document Cited by: §1.
- [35] (2019) Discrete-time mean-CVaR portfolio selection and time-consistency induced term structure of the CVaR. Journal of Economic Dynamics and Control 108, pp. 103751. External Links: Document Cited by: §1.
- [36] (2019) Dynamic portfolio execution. Management Science 65 (5), pp. 2015–2040. External Links: Document Cited by: §1.