arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2608.20223v1 [math.ST] 20 Aug 2026
\CJKencfamily

UTF8mc\CJK@envStartUTF8

Self-Normalizing Denominators
in Rational Causal Estimation

Shu Tamano Affiliation: Department of Multidisciplinary Sciences, Graduate School of Arts and Sciences, The University of Tokyo, 3-8-1 Komaba, Meguro-ku, Tokyo 153-8902, Japan Affiliation: Department of Epidemiology, National Institute of Infectious Diseases, Japan Institute for Health Security, 1-23-1 Toyama, Shinjuku-ku, Tokyo 162-0052, Japan
Abstract

Rational causal estimators in linear structural equation models take the form of one covariance polynomial divided by another, and a small denominator is commonly interpreted as weak identification. We show that, under Gaussian sampling, some denominators cannot enter this regime at first order. Their sampling variation is exactly proportional to their magnitude, so the standardized denominator is constant in every sample. Products of powers of nested covariance minors have this property in every dimension and admit an exact Wishart pivot. The converse is complete in dimension two. In dimension three, one mixed family remains open, while a factor-and-rank criterion classifies all denominators with linear or quadratic determinant-free factors and covers instrumental-variable, front-door and proximal formulas. For linear front-door adjustment, Wald inference remains asymptotically valid even as the mediator residual variance vanishes at an arbitrary rate, provided the treatment–mediator coefficient is nonzero. In simulations, proximal Wald coverage fell as a naive treatment–proxy diagnostic strengthened, while front-door coverage stayed nominal, and right-heart-catheterization data distinguished naive from denominator-relevant diagnostics.

Keywords: Bartlett decomposition; Fieller’s theorem; Fisher–Rao geometry; Proximal causal inference; Symmetric cone; Weak instrument.

1 Introduction

1.1 Motivation and scope

Graphical identification in linear structural equation models often yields a causal coefficient of the form N(Σ)/D(Σ)N(\Sigma)/D(\Sigma). Here Σ𝕊++d\Sigma\in\mathbb{S}^{d}_{++} is the observational covariance matrix of dd observed coordinates, and NN and DD are polynomial functions on the space Symd\operatorname{Sym}_{d} of symmetric matrices of order dd. Instrumental variables, conditional instruments, front-door adjustment, half-trek identification and proximal formulas all produce such expressions (31; 6; 14; 23; 39; 29; 38; 21). At regular distributions, substituting the sample covariance matrix into such a formula is routine. However, when DD is small, weak-instrument theory suggests a ratio-of-normals limit, distorted Wald inference and confidence sets that must sometimes be unbounded (16; 11; 36; 37; 2; 3). The standard safeguard compares the estimated denominator with its estimated standard error, as in first-stage relevance screening.

Two of these strategies display the contrast that motivates this paper. Throughout this preview, the observed vector is centred Gaussian with covariance matrix Σ𝕊++d\Sigma\in\mathbb{S}^{d}_{++}. The scalar σUV=cov(U,V)\sigma_{UV}=\operatorname{cov}(U,V) denotes the covariance of two observed variables UU and VV, and Σ^\hat{\Sigma} denotes the sample covariance matrix from nn independent observations. Section 2 gives the remaining conventions. For the front-door strategy, the observed vector is (X,M,Y)(X,M,Y)^{\top} and

Dfd(Σ)=σXX(σXXσMMσXM2).D_{\mathrm{fd}}(\Sigma)=\sigma_{XX}(\sigma_{XX}\sigma_{MM}-\sigma_{XM}^{2}).

At every such Σ\Sigma, the asymptotic variance of n1/2{Dfd(Σ^)Dfd(Σ)}n^{1/2}\{D_{\mathrm{fd}}(\hat{\Sigma})-D_{\mathrm{fd}}(\Sigma)\} is 10Dfd(Σ)210D_{\mathrm{fd}}(\Sigma)^{2}. Consequently, the statistic that divides nDfd(Σ^)2nD_{\mathrm{fd}}(\hat{\Sigma})^{2} by the plug-in value of this variance equals n/10n/10 for every positive-definite sample covariance. Screening this denominator reveals nothing about the underlying distribution, although the precision of the plug-in estimator may deteriorate without bound. By contrast, consider a linear proximal strategy with observed coordinates (A,Z,W)(A,Z,W)^{\top}, comprising a treatment, a treatment proxy and an outcome proxy. Its denominator is

Dprox(Σ)=σAAσZWσAZσAW,D_{\mathrm{prox}}(\Sigma)=\sigma_{AA}\sigma_{ZW}-\sigma_{AZ}\sigma_{AW},

which equals σAA\sigma_{AA} times the partial covariance of ZZ and WW given AA. This denominator vanishes at interior covariance matrices at which its asymptotic variance does not. There are therefore sequences of covariance matrices along which the plug-in ratio can have the ratio-of-normals limit of weak-instrument theory. The relevant screening direction is the partial covariance, not the marginal treatment–proxy association.

These examples motivate the question of which denominators make the first-order ratio regime impossible along every sequence of underlying covariance matrices. Under centred Gaussian sampling, the first-order variance of a covariance polynomial is a classical quadratic form in its symmetric gradient. We show that, for a special class of polynomials, this quadratic form is exactly proportional to the square of the polynomial. Denominator noise then contracts at the same rate as the denominator itself, and the standardized denominator is a samplewise constant rather than a relevance statistic. We call such denominators exactly self-normalizing.

Two neighbouring phenomena are excluded from this notion. A polynomial whose first-order variance merely vanishes on its zero set can still generate a nondegenerate second-order local limit when that variance is not proportional to the square of the polynomial. Section 5 gives an example. Numerator singularities are also distinct. The mediation null remains nonregular even though the front-door denominator is exactly self-normalizing (10). Therefore, the classification below concerns denominators, not universal regularity of the associated estimators.

1.2 Contributions

The first contribution is an all-dimensional sufficient class. Products of powers of covariance determinants along a nested flag of subspaces are exactly self-normalizing. Their relative gradients have a fixed spectrum, and their sample-to-population ratios factor into independent chi-squared variables. In recursive coordinates, these polynomials are monomials in successive conditional variances, which links the algebra directly to variance information.

The second contribution is a low-dimensional converse. The converse is complete in dimension two. In dimension three, determinant stripping reduces every solution to a rank-one branch, a pure plane-minor branch or a mixed branch with constant relative spectrum and a polynomial kernel line. The first two branches are flag powers, while rigidity of the mixed kernel line remains open. For the practically important class whose determinant-free irreducible factors have degree at most two, only one rank-one variance factor and one nested plane-minor factor can occur. This gives a coefficient-level factor-and-rank diagnostic and classifies every reducible self-normalizing cubic.

The statistical results connect this algebra to weak-denominator inference. A local result recovers the classical ratio-of-normals limit under a drift condition that no exactly self-normalizing denominator can satisfy. The standardized denominator is the noncentrality diagnostic of this limit and also determines whether Fieller inversion is bounded. For marginal and partial covariance denominators, it is an explicit increasing function of the corresponding first-stage statistic. For linear front-door adjustment, the Gaussian delta Wald statistic remains asymptotically standard normal when the mediator residual variance tends to zero at any rate, provided the treatment–mediator coefficient is nonzero.

Together, the results give a diagnostic pathway for graph-derived rational formulas. One factors the symbolic denominator, checks whether its factors form a nested variance flag, and then chooses between conventional studentization and weak-identification-robust inversion. Simulations and a diagnostic audit of right-heart-catheterization data from the SUPPORT study illustrate the two branches of this pathway.

1.3 Related work

Graphical criteria for instrumental sets, front-door adjustment and proximal identification determine whether a causal effect can be expressed from the observational law. The half-trek criterion and computer-algebra methods provide rational certificates for linear structural equation models (31; 6; 15; 14; 23; 39; 29; 21; 38). Generic and rational identifiability are distinct algebraic notions, and a parameter may be generically unique without being represented by the particular rational formula under study (9). Recent semiparametric work derives regular influence functions for half-trek estimators at fixed interior distributions (28). Our analysis starts after a rational certificate has been selected. It classifies the denominator as a polynomial on the ambient covariance space and determines which first-order boundary regime that certificate can generate.

Fieller and Anderson–Rubin inversion provide the classical ratio-based confidence constructions (1; 13). The impossibility of uniformly bounded confidence sets and the local theory of weak instruments and weak moments explain why a noisy denominator near zero produces non-Gaussian ratio limits (16; 11; 36; 37; 2). Modern work develops first-stage-dependent corrections and robust procedures, including settings with many weak moments and weakly identified nuisance functions (3; 24; 40; 4). The local result in Section 5 is a covariance-polynomial specialization of this literature. The new point is structural and precedes the limit calculation. The self-normalization identity decides when the first-order ratio regime is algebraically impossible.

Condition-number analyses quantify sensitivity of an identified functional or structural parameter to perturbations of the observational law (34; 19; 33). Exact self-normalization instead compares the gradient noise of one denominator with the denominator itself through a global polynomial identity. The two notions need not agree. In the front-door boundary sequence, the standard error diverges and the formula becomes badly conditioned, yet the denominator cannot enter the first-order Fieller regime. Conversely, a well-scaled slope denominator may have an interior zero with nondegenerate first-order noise.

Products of nested covariance minors are generalized power functions of a symmetric cone (12), and their exact sampling factorization follows from the Bartlett decomposition of a Wishart matrix (30). The affine-invariant metric on the positive-definite cone, which agrees up to scale with the Fisher–Rao metric of the centred Gaussian family, gives the geometric form of the logarithmic gradient flow used in the converse proofs (5). The separation between Gaussian and elliptical covariance operators is familiar in covariance-structure analysis (35; 22). Here it identifies which part of the classification is distributional and which part is algebraic.

The dimension-two converse uses the Gordan–Noether theorem for forms with vanishing Hessian (18; 27). That route fails for the six variables of a symmetric matrix of order three because noncone forms with vanishing Hessian exist in this dimension (32; 7; 17). Supplementary Section I.1 explains why Hessian degeneracy does not force a cone on the six-dimensional space Sym3\operatorname{Sym}_{3}, while Supplementary Section I.2 identifies the information lost when the relevant tensor relation is passed through polynomial multiplication. These two obstructions explain why neither route presently closes the mixed branch (25). This distinction explains the complete dimension-two result and the mixed-kernel problem left in dimension three.

2 Preliminaries

2.1 Basic notation

We use \mathbb{R} and \mathbb{C} for the real and complex fields. Vectors are columns, AA^{\top} denotes matrix transpose, ImI_{m} is the identity matrix of order mm, and eje_{j} is the jjth standard basis vector. The subscript on ImI_{m} is omitted when the order is clear. The symbols tr(A)\operatorname{tr}(A), det(A)\det(A), rank(A)\operatorname{rank}(A) and spec(A)\operatorname{spec}(A) denote trace, determinant, rank and the multiset of eigenvalues of AA. Expectations, variances, covariances and probabilities under a probability measure QQ are denoted by 𝔼Q\mathbb{E}_{Q}, varQ\operatorname{var}_{Q}, covQ\operatorname{cov}_{Q} and prQ\mathrm{pr}_{Q}. The subscript is suppressed once the probability measure is fixed. Convergence in probability and in distribution are denoted by p\longrightarrow_{\rm p} and \rightsquigarrow, and limits are as nn\to\infty unless stated otherwise. The indicator of a statement AA is 1(A)\mathrm{1}(A), and sgn(x)\operatorname{sgn}(x) is the sign of xx\in\mathbb{R}.

For a positive integer dd, let [d]={1,,d}[d]=\{1,\ldots,d\}, pd=d(d+1)/2p_{d}=d(d+1)/2 and d={(i,j):1ijd}\mathcal{I}_{d}=\{(i,j):1\leq i\leq j\leq d\}. We write Symd={Ad×d:A=A}\operatorname{Sym}_{d}=\{A\in\mathbb{R}^{d\times d}:A=A^{\top}\} and denote its positive-definite and positive-semidefinite cones by 𝕊++d\mathbb{S}^{d}_{++} and 𝕊+d\mathbb{S}^{d}_{+}. The notation A0A\succ 0 and A0A\succeq 0 means A𝕊++dA\in\mathbb{S}^{d}_{++} and A𝕊+dA\in\mathbb{S}^{d}_{+}, respectively. For ΣSymd\Sigma\in\operatorname{Sym}_{d}, its entries are σij=σji\sigma_{ij}=\sigma_{ji}. For I,J[d]I,J\subseteq[d], AI,JA_{I,J} is the submatrix with rows in II and columns in JJ, AI=AI,IA_{I}=A_{I,I}, and A1:r=A{1,,r},{1,,r}A_{1:r}=A_{\{1,\ldots,r\},\{1,\ldots,r\}}.

We order d\mathcal{I}_{d} lexicographically, first by ii and then by jj, and set

vech(A)=(a11,a12,,a1d,a22,a23,,add)pd.\operatorname{vech}(A)=(a_{11},a_{12},\ldots,a_{1d},a_{22},a_{23},\ldots,a_{dd})^{\top}\in\mathbb{R}^{p_{d}}.

Vectors and matrices indexed by d\mathcal{I}_{d} follow this ordering. We identify [Symd]=[σij:(i,j)d]\mathbb{R}[\operatorname{Sym}_{d}]=\mathbb{R}[\sigma_{ij}:(i,j)\in\mathcal{I}_{d}] with the ring of polynomial maps from Symd\operatorname{Sym}_{d} to \mathbb{R}. The notation P0P\equiv 0 means that PP is the zero polynomial in this ring. A polynomial PP is homogeneous of degree KK when P(tΣ)=tKP(Σ)P(t\Sigma)=t^{K}P(\Sigma) for every tt\in\mathbb{R}. Its total degree is deg(P)\deg(P), and two polynomials are coprime when their only common divisors are nonzero constants.

For P[Symd]P\in\mathbb{R}[\operatorname{Sym}_{d}], the coordinate gradient P(Σ)pd\nabla P(\Sigma)\in\mathbb{R}^{p_{d}} consists of the partial derivatives in the preceding vech\operatorname{vech} order. The Fréchet differential of PP at Σ\Sigma in the direction HH is

dPΣ(H)=ddtP(Σ+tH)|t=0,HSymd.\mathrm{d}P_{\Sigma}(H)=\left.\frac{\mathrm{d}}{\mathrm{d}t}P(\Sigma+tH)\right|_{t=0},\quad H\in\operatorname{Sym}_{d}.

The symmetric gradient GP(Σ)SymdG_{P}(\Sigma)\in\operatorname{Sym}_{d} is the unique matrix satisfying dPΣ(H)=tr{GP(Σ)H}\mathrm{d}P_{\Sigma}(H)=\operatorname{tr}\{G_{P}(\Sigma)H\} for every HSymdH\in\operatorname{Sym}_{d}. Thus, (GP)ii=P/σii(G_{P})_{ii}=\partial P/\partial\sigma_{ii} and (GP)ij=21P/σij(G_{P})_{ij}=2^{-1}\partial P/\partial\sigma_{ij} for i<ji<j, with all quantities evaluated at Σ\Sigma. For P(Σ)0P(\Sigma)\neq 0, the relative gradient is RP(Σ)=GP(Σ)Σ/P(Σ)R_{P}(\Sigma)=G_{P}(\Sigma)\Sigma/P(\Sigma).

2.2 Problem set-up

We formulate the statistical problem in the ambient covariance space. Fix d1d\geq 1 and Σ𝕊++d\Sigma\in\mathbb{S}^{d}_{++}. The notation Nm(μ,Ω)N_{m}(\mu,\Omega) denotes the Gaussian law on m\mathbb{R}^{m} with mean μm\mu\in\mathbb{R}^{m} and covariance matrix Ω𝕊+m\Omega\in\mathbb{S}^{m}_{+}. Write 𝒫Σ=Nd(0,Σ)\mathcal{P}_{\Sigma}=N_{d}(0,\Sigma) and 𝒫Σ(n)=𝒫Σn\mathcal{P}_{\Sigma}^{(n)}=\mathcal{P}_{\Sigma}^{\otimes n}. Let X1,,XnX_{1},\ldots,X_{n} be independent observations with joint law 𝒫Σ(n)\mathcal{P}_{\Sigma}^{(n)}. We use the known-mean sample covariance matrix

Σ^=n1i=1nXiXi,n>d.\hat{\Sigma}=n^{-1}\sum_{i=1}^{n}X_{i}X_{i}^{\top},\quad n>d. (1)

For an integer ν1\nu\geq 1 and Ω𝕊+m\Omega\in\mathbb{S}^{m}_{+}, 𝒲m(ν,Ω)\mathcal{W}_{m}(\nu,\Omega) denotes the law of =1νZZ\sum_{\ell=1}^{\nu}Z_{\ell}Z_{\ell}^{\top} for independent ZNm(0,Ω)Z_{\ell}\sim N_{m}(0,\Omega). Its scalar specialization with Ω=1\Omega=1 is χν2\chi^{2}_{\nu}. Consequently, nΣ^𝒲d(n,Σ)n\hat{\Sigma}\sim\mathcal{W}_{d}(n,\Sigma) and Σ^0\hat{\Sigma}\succ 0 almost surely.

The identification equality is required only on a covariance model, whereas the polynomial pair representing it governs off-model sampling behaviour. The following definition separates these two objects.

Definition 1 (Identification strategy).

Let 𝕊++d\mathcal{M}\subseteq\mathbb{S}^{d}_{++} be a nonempty covariance model and let τ:\tau:\mathcal{M}\to\mathbb{R} be a scalar target. An identification strategy for τ\tau is an ordered pair (N,D)(N,D) of coprime elements of [Symd]\mathbb{R}[\operatorname{Sym}_{d}] such that DD is not identically zero on \mathcal{M} and

N(Σ)=τ(Σ)D(Σ)(Σ).N(\Sigma)=\tau(\Sigma)D(\Sigma)\quad(\Sigma\in\mathcal{M}).

Thus, τ=N/D\tau=N/D at every point of \mathcal{M} at which DD is nonzero. The covariance plug-in estimator is τ^=N(Σ^)/D(Σ^)\hat{\tau}=N(\hat{\Sigma})/D(\hat{\Sigma}) on the event {D(Σ^)0}\{D(\hat{\Sigma})\neq 0\}. The target is scale-invariant when τ(tΣ)=τ(Σ)\tau(t\Sigma)=\tau(\Sigma) whenever Σ,tΣ\Sigma,t\Sigma\in\mathcal{M} and t>0t>0.

This distinction is necessary for two reasons. First, Σ^\hat{\Sigma} need not satisfy the model equations, so the sampling behaviour of τ^\hat{\tau} depends on NN and DD as polynomials on the whole of Symd\operatorname{Sym}_{d}. Second, multiplying both NN and DD by a nonconstant common factor leaves the identified ratio unchanged wherever both representations are defined, yet alters the off-model sampling behaviour of the denominator. Coprimality removes this ambiguity. Two coprime pairs representing the same rational function coincide up to a common nonzero scalar, while distinct rational certificates for the same target remain distinct strategies.

2.3 Exact self-normalization

We now introduce the central property of the paper. Define the polynomial matrix map Γ:SymdSympd\Gamma:\operatorname{Sym}_{d}\to\operatorname{Sym}_{p_{d}} by

Γ(ij),(kl)(Σ)=σikσjl+σilσjk,(i,j),(k,l)d.\Gamma_{(ij),(kl)}(\Sigma)=\sigma_{ik}\sigma_{jl}+\sigma_{il}\sigma_{jk},\quad(i,j),(k,l)\in\mathcal{I}_{d}. (2)

Under 𝒫Σ(n)\mathcal{P}_{\Sigma}^{(n)}, the matrix Γ(Σ)\Gamma(\Sigma) is the covariance matrix of n1/2vech(Σ^Σ)n^{1/2}\operatorname{vech}(\hat{\Sigma}-\Sigma). For D[Symd]D\in\mathbb{R}[\operatorname{Sym}_{d}], define its Gaussian first-order variance functional by

sD2(Σ)=D(Σ)Γ(Σ)D(Σ)=2tr{(GD(Σ)Σ)2}.s_{D}^{2}(\Sigma)=\nabla D(\Sigma)^{\top}\Gamma(\Sigma)\nabla D(\Sigma)=2\operatorname{tr}\{(G_{D}(\Sigma)\Sigma)^{2}\}. (3)

The second equality is the Gaussian quadratic-form variance identity, verified in Supplementary Lemma B.1. Both sides are polynomial in Σ\Sigma, so the equality extends from 𝕊++d\mathbb{S}^{d}_{++} to all of Symd\operatorname{Sym}_{d}. For Σ𝕊+d\Sigma\in\mathbb{S}^{d}_{+} the matrix Γ(Σ)\Gamma(\Sigma) is positive semidefinite, and we write sD(Σ)={sD2(Σ)}1/2s_{D}(\Sigma)=\{s_{D}^{2}(\Sigma)\}^{1/2} there. The standardized denominator is F^D=nD(Σ^)2/sD2(Σ^)\widehat{F}_{D}=nD(\hat{\Sigma})^{2}/s_{D}^{2}(\hat{\Sigma}), defined on the event {sD2(Σ^)>0}\{s_{D}^{2}(\hat{\Sigma})>0\}.

Definition 2 (Exact self-normalization).

A nonconstant homogeneous polynomial D[Symd]D\in\mathbb{R}[\operatorname{Sym}_{d}] is exactly self-normalizing when there is a constant c(0,)c\in(0,\infty) such that

2tr{(GD(Σ)Σ)2}=cD(Σ)2(ΣSymd).2\operatorname{tr}\{(G_{D}(\Sigma)\Sigma)^{2}\}=cD(\Sigma)^{2}\quad(\Sigma\in\operatorname{Sym}_{d}). (4)

For a strategy (N,D)(N,D), exact self-normalization refers only to the denominator DD.

The constant cc is unique because DD is not the zero polynomial. Below, self-normalizing always means exactly self-normalizing. Equation (4) states that the Gaussian first-order noise of DD contracts in exact proportion to |D||D| throughout the ambient covariance space. In particular, evaluating (4) at Σ^\hat{\Sigma} gives F^D=n/c\widehat{F}_{D}=n/c whenever sD2(Σ^)>0s_{D}^{2}(\hat{\Sigma})>0. Thus, the standardized denominator of a self-normalizing polynomial DD is constant over samples rather than a relevance statistic. Exact self-normalization is invariant under congruence transformations. Every rational strategy for a scale-invariant coefficient in a linear structural equation model has a homogeneous representative, and a nonzero self-normalizing polynomial has no zero in 𝕊++d\mathbb{S}^{d}_{++}. Supplementary Lemmas B.1B.3 prove these facts. If DD depends on Σ\Sigma only through its compression to an rr-dimensional subspace, its self-normalization status and constant are unchanged in every ambient dimension drd\geq r by Supplementary Lemma B.4. The dimension-three classification below therefore applies to any denominator supported on at most three directions, whatever the number of observed coordinates.

Remark 1 (Scope of exactness).

Equation (4) is an identity in the ambient polynomial ring for the coprime denominator fixed by Definition 1. It is not merely an equality on a structural model. The word “exact” also refers to the Gaussian covariance operator in (3). A sandwich studentizer under a general law targets a different quadratic form and need not be constant sample by sample. Under elliptical sampling, the same algebraic class persists after a kurtosis adjustment, although the product-of-chi-squares pivot is generally lost. With an unknown Gaussian mean, the corresponding centred covariance formulas replace nn by n1n-1. Supplementary Proposition B.1 and Remark B.1 give the precise statements.

3 Flag powers and an exact pivot

3.1 Nested covariance minors

We construct self-normalizing denominators in every dimension. The building blocks are covariance volumes of subspaces. For an rr-dimensional subspace VdV\subseteq\mathbb{R}^{d}, let MVM_{V} be any full-row-rank matrix whose row space is VV, and write ΔV(Σ)=det(MVΣMV)\Delta_{V}(\Sigma)=\det(M_{V}\Sigma M_{V}^{\top}). If a random vector XX has covariance matrix Σ\Sigma, then ΔV(Σ)\Delta_{V}(\Sigma) is the generalized variance of the reduced vector MVXM_{V}X. A different choice of MVM_{V} multiplies ΔV\Delta_{V} by a positive constant and affects no statement below.

A flag is a strictly increasing chain V1VkV_{1}\subset\cdots\subset V_{k} of nonzero subspaces of d\mathbb{R}^{d}. Let aj=dim(Vj)a_{j}=\dim(V_{j}), so that 1a1<<akd1\leq a_{1}<\cdots<a_{k}\leq d, and let e1,,eke_{1},\ldots,e_{k} be positive integers. The associated flag power and its multiplicities are

D(Σ)=j=1kΔVj(Σ)ej,mi=j=1kej1(aji).D(\Sigma)=\prod_{j=1}^{k}\Delta_{V_{j}}(\Sigma)^{e_{j}},\quad m_{i}=\sum_{j=1}^{k}e_{j}\mathrm{1}(a_{j}\geq i). (5)

Thus, mim_{i} is the total exponent carried by factors whose subspaces have dimension at least ii, and m1mak1m_{1}\geq\cdots\geq m_{a_{k}}\geq 1. The following theorem shows that every flag power is exactly self-normalizing and gives its exact sampling law.

Theorem 1 (Flag powers).

Let DD and m1,,makm_{1},\ldots,m_{a_{k}} be given by (5). The following statements hold.

  1. (i)

    For every Σ0\Sigma\succ 0, the spectrum of the relative gradient RD(Σ)R_{D}(\Sigma) is (m1,,mak,0,,0)(m_{1},\ldots,m_{a_{k}},0,\ldots,0), where the eigenvalue zero has multiplicity dakd-a_{k}.

  2. (ii)

    Equation (4) holds with c=2i=1akmi2c=2\sum_{i=1}^{a_{k}}m_{i}^{2}.

  3. (iii)

    If nakn\geq a_{k}, then D(Σ^)/D(Σ)D(\hat{\Sigma})/D(\Sigma) is distributed as

    i=1ak(χni+12n)mi,\prod_{i=1}^{a_{k}}\Biggl(\frac{\chi^{2}_{n-i+1}}{n}\Biggr)^{m_{i}}, (6)

    where the chi-squared variables are independent.

  4. (iv)

    In particular, the identity F^D=n/c\widehat{F}_{D}=n/c holds almost surely, and no sequence Σn𝕊++d\Sigma_{n}\in\mathbb{S}^{d}_{++} can satisfy both n1/2D(Σn)=O(1)n^{1/2}D(\Sigma_{n})=O(1) and sD(Σn)s>0s_{D}(\Sigma_{n})\to s>0.

Remark 2 (Algebraic and sampling content).

Parts (i) and (ii) are polynomial identities and do not depend on a sampling law. The samplewise identity in part (iv) therefore holds at every sample covariance matrix at which DD is nonzero, whatever the data-generating law. Only the factorization in part (iii) uses Gaussian sampling. With an unknown Gaussian mean, parts (iii) and (iv) hold for the centred sample covariance with nn replaced by n1n-1. Under elliptical sampling, flag powers remain proportionally self-normalizing after a kurtosis adjustment, but the product-of-chi-squares law is generally lost. Supplementary Proposition B.1 and Remark B.1 give these variants.

The nesting hypothesis cannot be weakened. If neither of two subspaces VV and WW contains the other, then no product ΔVpΔWq\Delta_{V}^{p}\Delta_{W}^{q} with positive integer exponents pp and qq is exactly self-normalizing. Supplementary Section C proves both this two-subspace criterion and Theorem 1.

3.2 Recursive variance interpretation

Flag powers aggregate conditional-variance information. Let X=(X(1),,X(d))X=(X^{(1)},\ldots,X^{(d)})^{\top} have the sampling law 𝒫Σ\mathcal{P}_{\Sigma}, and write vi(Σ)=varΣ{X(i)X(1),,X(i1)}v_{i}(\Sigma)=\operatorname{var}_{\Sigma}\{X^{(i)}\mid X^{(1)},\ldots,X^{(i-1)}\} for the successive conditional variances, with v1(Σ)=σ11v_{1}(\Sigma)=\sigma_{11}. For every r[d]r\in[d],

det(Σ1:r)=i=1rvi(Σ).\det(\Sigma_{1:r})=\prod_{i=1}^{r}v_{i}(\Sigma). (7)

Consider the coordinate flag, whose jjth subspace is spanned by the first aja_{j} coordinate directions. By (7), the flag power (5) is then the monomial i=1akvi(Σ)mi\prod_{i=1}^{a_{k}}v_{i}(\Sigma)^{m_{i}}, so the multiplicity mim_{i} is the exponent attached to the iith conditional variance. A general flag reduces to this case, up to a positive constant, after one fixed nonsingular linear transformation of the observation vector. For the front-door strategy of Section 1, the first two conditional variances are the treatment variance and the mediator residual variance. Its denominator is the monomial v1(Σ)2v2(Σ)v_{1}(\Sigma)^{2}v_{2}(\Sigma), and part (ii) of Theorem 1 returns the constant c=2(22+12)=10c=2(2^{2}+1^{2})=10 behind the value n/10n/10.

The same representation explains the exact sampling law. Under 𝒫Σ(n)\mathcal{P}_{\Sigma}^{(n)}, the sample conditional variances computed from Σ^\hat{\Sigma} along the coordinate flag are mutually independent. The iith is distributed as vi(Σ)v_{i}(\Sigma) times a χni+12/n\chi^{2}_{n-i+1}/n variable. This is the Bartlett decomposition of the Wishart matrix nΣ^n\hat{\Sigma} (30, Ch. 3), and evaluating the monomial at these independent factors gives the law in (6). To first order, the relative variance of every sample conditional variance is 2/n2/n, whatever the value of Σ\Sigma. By independence, the relative first-order variance of the sample flag power is 2i=1akmi2/n2\sum_{i=1}^{a_{k}}m_{i}^{2}/n, which equals c/nc/n. This constancy in Σ\Sigma is the sampling mechanism behind exact self-normalization. The sample flag power depends on Σ^\hat{\Sigma} only through these conditional variances. The regression coefficients of each coordinate on its predecessors, which complete the recursive decomposition of Σ^\hat{\Sigma}, do not enter. This contrast between conditional variances and regression coefficients is the sample form of the distinction between variance information and slope information that organizes the converse results of Section 4.

4 Converse results and diagnosis

4.1 Dimension two

We first obtain a complete converse for covariance matrices of order two. Through face restriction, this result also drives the three-variable analysis.

Theorem 2 (Complete converse in dimension two).

Let D[Sym2]D\in\mathbb{R}[\operatorname{Sym}_{2}] be homogeneous, nonconstant and exactly self-normalizing. Then

D(Σ)=γ(vΣv)a(detΣ)b,D(\Sigma)=\gamma(v^{\top}\Sigma v)^{a}(\det\Sigma)^{b}, (8)

for a nonzero constant γ\gamma, a nonzero vector v2v\in\mathbb{R}^{2}, and nonnegative integers a,ba,b with (a,b)(0,0)(a,b)\neq(0,0). The self-normalization constant is c=2{(a+b)2+b2}c=2\{(a+b)^{2}+b^{2}\}.

Thus, a two-variable denominator can self-normalize only by combining one variance direction with the full covariance volume.

Remark 3 (What the converse excludes).

The permitted irreducible factors are one rank-one variance vΣvv^{\top}\Sigma v and the determinant. Every linear form tr(BΣ)\operatorname{tr}(B\Sigma) whose coefficient matrix BB has rank two, such as the off-diagonal covariance σ12\sigma_{12}, is excluded even though it has the same degree as the permitted variance forms. Thus, the theorem distinguishes variance information from slope information within the same polynomial degree.

We summarize the proof mechanism, which explains the rigidity. Euler’s identity and (4) fix the first two power sums of the eigenvalues of RD(Σ)R_{D}(\Sigma). The elementary symmetric identity converts them into a determinant identity with an explicit constant coefficient. When that coefficient is nonzero, the determinant divides DD, the factor can be stripped, and induction on the degree applies. When it vanishes, the gradient map takes values in the rank-one quadric, and the Hessian determinant of DD vanishes identically. The Gordan–Noether theorem reduces DD to a power of one linear form, which symmetry and reality identify with a variance direction (18; 27). Supplementary Section D.2 gives the full proof.

4.2 Dimension three

We now reduce the three-variable converse to a single mixed branch. Put δ=detΣ\delta=\det\Sigma and call D[Sym3]D\in\mathbb{R}[\operatorname{Sym}_{3}] determinant-free when δD\delta\nmid D. For homogeneous DD of degree KK, define the rank-one restriction pD:3p_{D}:\mathbb{R}^{3}\to\mathbb{R} by pD(u)=D(uu)p_{D}(u)=D(uu^{\top}). For a plane V3V\subseteq\mathbb{R}^{3} with basis matrix P3×2P\in\mathbb{R}^{3\times 2}, define the face restriction DV:Sym2D_{V}:\operatorname{Sym}_{2}\to\mathbb{R} by DV(S)=D(PSP)D_{V}(S)=D(PSP^{\top}). The plane VV is called DD-good when DV0D_{V}\not\equiv 0, and simply good when the polynomial is clear from context. A polynomial vector is primitive when its entries have no nonconstant common divisor.

A nonzero plane restriction of a self-normalizing polynomial is again self-normalizing with the same constant. Supplementary Lemma D.3 provides this restriction principle, and Theorem 2 then determines the form of every such restriction.

Definition 3 (Face type).

Let DD be a determinant-free self-normalizing polynomial on Sym3\operatorname{Sym}_{3}, of degree KK and with self-normalization constant cc. For a,b0a,b\in\mathbb{Z}_{\geq 0}, we say that DD has face type (a,b)(a,b) if, for every DD-good plane VV, there exist γV{0}\gamma_{V}\in\mathbb{R}\setminus\{0\} and wV2{0}w_{V}\in\mathbb{R}^{2}\setminus\{0\} such that, as a polynomial identity on Sym2\operatorname{Sym}_{2},

DV(S)=γV(wVSwV)a(detS)b,a+2b=K,(a+b)2+b2=c/2.D_{V}(S)=\gamma_{V}(w_{V}^{\top}Sw_{V})^{a}(\det S)^{b},\quad a+2b=K,\quad(a+b)^{2}+b^{2}=c/2. (9)

A determinant-free solution has at least one good plane, and its degree and constant admit at most one pair satisfying (9), so the face type exists and is common to all good planes. Removing a factor of δ\delta preserves exact self-normalization with a shifted constant, so the residue in the next theorem is itself self-normalizing. Supplementary Lemmas D.4 and D.6 provide both facts.

Theorem 3 (Three-variable reduction).

Let D[Sym3]D\in\mathbb{R}[\operatorname{Sym}_{3}] be homogeneous, nonconstant and exactly self-normalizing, and let EE be the determinant-free residue obtained by removing from DD the largest power of δ\delta. Exactly one of the following cases occurs.

  1. (i)

    If pE0p_{E}\not\equiv 0, then E=γ(vΣv)deg(E)E=\gamma(v^{\top}\Sigma v)^{\deg(E)} for a nonzero constant γ\gamma and a nonzero vector v3v\in\mathbb{R}^{3}.

  2. (ii)

    If pE0p_{E}\equiv 0 and the face type of EE is (0,b)(0,b), then E=γΔWbE=\gamma\Delta_{W}^{b} for a nonzero constant γ\gamma and a plane W3W\subseteq\mathbb{R}^{3}.

  3. (iii)

    If pE0p_{E}\equiv 0 and the face type of EE is (a,b)(a,b) with a,b1a,b\geq 1, then

    spec{RE(Σ)}={a+b,b,0}(Σ0),\operatorname{spec}\{R_{E}(\Sigma)\}=\{a+b,b,0\}\quad(\Sigma\succ 0), (10)

    detGE0\det G_{E}\equiv 0, and a primitive polynomial vector 𝝂\boldsymbol{\nu} satisfies GE𝝂0G_{E}\boldsymbol{\nu}\equiv 0, so that 𝝂(Σ)\boldsymbol{\nu}(\Sigma) spans the kernel of GE(Σ)G_{E}(\Sigma) at every point where GE(Σ)G_{E}(\Sigma) has rank two and 𝝂(Σ)0\boldsymbol{\nu}(\Sigma)\neq 0.

The polynomials in cases (i) and (ii) are flag powers built on a single subspace. A mixed type has degree a+2b3a+2b\geq 3, so every solution of degree at most two falls under the first two cases and is a flag power. Therefore, the unresolved solutions lie in case (iii) and carry one possibly moving null direction together with the fixed spectrum in (10).

Remark 4 (The mixed-kernel problem).

The unrestricted three-variable converse is equivalent to the constancy of the kernel line in Theorem 3 (iii). For a nonzero vector ww, write [w][w] for the line that it spans. If [𝛎(Σ)][\boldsymbol{\nu}(\Sigma)] is constant on one nonempty open set, then a fixed vector u0u_{0} satisfies GE(Σ)u00G_{E}(\Sigma)u_{0}\equiv 0, and the two-variable converse yields

E(Σ)=γ(vΣv)adet(MΣM)b,vu0,E(\Sigma)=\gamma(v^{\top}\Sigma v)^{a}\det(M\Sigma M^{\top})^{b},\quad v\in u_{0}^{\perp},

where u0={x3:u0x=0}u_{0}^{\perp}=\{x\in\mathbb{R}^{3}:u_{0}^{\top}x=0\} and the rows of MM span u0u_{0}^{\perp}. The equivalence between constancy of the kernel line and the displayed mixed flag form is established in Supplementary Lemma D.11. Supplementary Proposition D.3 gives the cofactor syzygy that couples a moving kernel line to the determinant boundary. Vanishing of the six-variable Hessian alone is insufficient because noncone forms with vanishing Hessian exist in this dimension (32; 7; 17).

Supplementary Section D also gives an independent closure of the mixed branch that does not presuppose a constant kernel line. If the restrictions of the determinant-free mixed residue to rank-two faces agree with those of one fixed flag power built from a nested line and plane, the residue is that flag power whenever its mixed type has degree at most eight. For higher-degree types, the same conclusion holds under this boundary agreement condition outside an explicit arithmetic resonance set. Therefore, the unresolved difficulty is the geometric gluing of the face directions rather than an uncontrolled interior perturbation.

4.3 Higher-dimensional scope

The constancy of the relative spectrum that drives the three-variable reduction is special to dimension three. Euler’s identity fixes the trace of RD(Σ)R_{D}(\Sigma) at the degree of DD, and (4) fixes the trace of its square, so in dimension three the eigenvalues are confined to a circle. The polynomial character of DD forces at least one further linear relation with integer coefficients among the eigenvalues. A line meets a circle in at most two points, so only finitely many spectra are possible, and continuity on the connected cone makes the spectrum constant. This finiteness underlies the fixed spectrum displayed in (10). In dimension d4d\geq 4 the same two traces leave a sphere of dimension d2d-2, a linear relation cuts out a set of dimension d3d-3, and no finiteness follows.

A conditional converse nevertheless holds in every dimension. If the relative gradients RD(Σ)R_{D}(\Sigma) preserve one fixed complete flag for every Σ0\Sigma\succ 0 and induce constant weights on the successive one-dimensional quotients, then DD is the corresponding flag power. This is an intrinsic condition on the gradient rather than an assumed factorization, and the precise statement is Supplementary Proposition D.5. Exact self-normalization alone is not claimed to create such an invariant flag when d4d\geq 4. The obstruction in higher dimensions is thus the emergence of one common invariant recursive ordering, not the integration step once that ordering is present, and dimension three is a sharp low-dimensional target rather than an arbitrary truncation.

4.4 A factor-and-rank diagnostic

Although the mixed branch obstructs an unconditional converse, the denominators produced by graphical identification typically have a simple factor structure. The two denominators of Section 1 factor into linear and quadratic pieces after determinant stripping, and the same holds for the formulas revisited in Section 6. We now show that this class is classified completely and that membership can be decided by exact symbolic computation.

Definition 4 (Low-degree-factor condition).

A homogeneous polynomial on Sym3\operatorname{Sym}_{3} satisfies the low-degree-factor condition if, after removing the largest power of detΣ\det\Sigma, every real irreducible factor has degree at most two.

Under this condition the mixed branch disappears and the classification closes.

Theorem 4 (Low-degree-factor converse).

Let D[Sym3]D\in\mathbb{R}[\operatorname{Sym}_{3}] be homogeneous, nonconstant and exactly self-normalizing, and suppose that DD satisfies the low-degree-factor condition. Then there exist a nested line and plane V1V2V_{1}\subset V_{2}, a nonzero constant γ\gamma, and nonnegative integers a,b,ea,b,e such that

D(Σ)=γΔV1(Σ)aΔV2(Σ)b(detΣ)e.D(\Sigma)=\gamma\Delta_{V_{1}}(\Sigma)^{a}\Delta_{V_{2}}(\Sigma)^{b}(\det\Sigma)^{e}. (11)

Therefore, linear and quadratic factors cannot assemble into any self-normalizing configuration other than a nested flag.

For a square matrix AA, adj(A)\operatorname{adj}(A) denotes its classical adjugate, characterized by Aadj(A)=adj(A)A=det(A)IA\operatorname{adj}(A)=\operatorname{adj}(A)A=\det(A)I. Writing the plane minor through the adjugate turns Theorem 4 into a normal form whose ingredients can be tested one by one.

Corollary 1 (Factor-and-rank diagnostic).

Let D[Sym3]D\in\mathbb{R}[\operatorname{Sym}_{3}] be homogeneous and nonconstant, and suppose that DD satisfies the low-degree-factor condition. Then DD is exactly self-normalizing if and only if, up to a nonzero constant,

D(Σ)=(vΣv)a{uadj(Σ)u}b(detΣ)e,uv=0,D(\Sigma)=(v^{\top}\Sigma v)^{a}\{u^{\top}\operatorname{adj}(\Sigma)u\}^{b}(\det\Sigma)^{e},\quad u^{\top}v=0, (12)

for nonzero vectors u,v3u,v\in\mathbb{R}^{3} and nonnegative integers a,b,ea,b,e. In that case the self-normalization constant is

c=2{(a+b+e)2+(b+e)2+e2}.c=2\{(a+b+e)^{2}+(b+e)^{2}+e^{2}\}. (13)

Consequently, every reducible exactly self-normalizing cubic is one of γΔV13\gamma\Delta_{V_{1}}^{3}, γΔV1ΔV2\gamma\Delta_{V_{1}}\Delta_{V_{2}} with V1V2V_{1}\subset V_{2}, or γdetΣ\gamma\det\Sigma.

This corollary reduces the audit of a candidate denominator to a finite procedure. One factors the polynomial symbolically, tests the ranks of the coefficient matrices of the linear and adjugate-linear factors, and checks orthogonality and nesting of the resulting directions. The globalization of the generic face factors and the resulting coefficient-level characterization are proved in Supplementary Section D.8.

The two formulas of Section 1 illustrate the procedure.

Example 1 (From a graph formula to a rank check).

For the coordinate order (X,M,Y)(X,M,Y), let eX,eM,eYe_{X},e_{M},e_{Y} denote the corresponding standard basis vectors. The front-door denominator factors as

Dfd=σXX(σXXσMMσXM2)=tr(eXeXΣ)eYadj(Σ)eY.D_{\mathrm{fd}}=\sigma_{XX}(\sigma_{XX}\sigma_{MM}-\sigma_{XM}^{2})=\operatorname{tr}(e_{X}e_{X}^{\top}\Sigma)\,e_{Y}^{\top}\operatorname{adj}(\Sigma)e_{Y}. (14)

Both coefficient matrices have rank one and eYeX=0e_{Y}^{\top}e_{X}=0, so Corollary 1 applies with (a,b,e)=(1,1,0)(a,b,e)=(1,1,0) and certifies exact self-normalization with c=10c=10, in agreement with Section 3. For the coordinate order (A,Z,W)(A,Z,W), define eA,eZ,eWe_{A},e_{Z},e_{W} analogously. The proximal denominator is

Dprox=σAAσZWσAZσAW=tr{Cadj(Σ)},C={eZeW+eWeZ}/2,D_{\mathrm{prox}}=\sigma_{AA}\sigma_{ZW}-\sigma_{AZ}\sigma_{AW}=\operatorname{tr}\{C\operatorname{adj}(\Sigma)\},\quad C=-\{e_{Z}e_{W}^{\top}+e_{W}e_{Z}^{\top}\}/2, (15)

where rank(C)=2\operatorname{rank}(C)=2, so the quadratic factor fails the rank-one test. Indeed, DproxD_{\mathrm{prox}} vanishes at the identity matrix, so exact self-normalization already fails by the interior nonvanishing property recorded in Section 2.

Thus, the audit separates the two formulas before any data are collected.

Remark 5 (How the diagnostic should be used).

For a denominator with rational or algebraic coefficients, the factorization and the rank checks are exact symbolic operations, and the diagnostic is complete on the stated class. A near-factorization obtained numerically does not certify self-normalization. An irreducible factor of degree at least three renders the criterion inconclusive rather than negative. One may then test (4) directly by comparing coefficients or apply the kernel and boundary criteria of Supplementary Section D.

5 Weak-denominator inference

5.1 Local ratio experiment

We formulate the first-order regime of a weak denominator whose sampling noise does not degenerate with it. Throughout this section, (N,D)(N,D) is a fixed identification strategy in the sense of Definition 1, and τ^=N(Σ^)/D(Σ^)\hat{\tau}=N(\hat{\Sigma})/D(\hat{\Sigma}) is its plug-in estimator. For a deterministic sequence Σn𝕊++d\Sigma_{n}\in\mathbb{S}^{d}_{++} with D(Σn)0D(\Sigma_{n})\neq 0, define

τn=N(Σn)D(Σn),gn=N(Σn)τnD(Σn),\tau_{n}=\frac{N(\Sigma_{n})}{D(\Sigma_{n})},\quad g_{n}=\nabla N(\Sigma_{n})-\tau_{n}\nabla D(\Sigma_{n}),

and put sD,n=sD(Σn)s_{D,n}=s_{D}(\Sigma_{n}) and sg,n={gnΓ(Σn)gn}1/2s_{g,n}=\{g_{n}^{\top}\Gamma(\Sigma_{n})g_{n}\}^{1/2}. Whenever sD,nsg,n>0s_{D,n}s_{g,n}>0, define

μn=n1/2D(Σn)sD,n,rn=gnΓ(Σn)D(Σn)sg,nsD,n,ωn=sg,nsD,n.\mu_{n}=\frac{n^{1/2}D(\Sigma_{n})}{s_{D,n}},\quad r_{n}=\frac{g_{n}^{\top}\Gamma(\Sigma_{n})\nabla D(\Sigma_{n})}{s_{g,n}s_{D,n}},\quad\omega_{n}=\frac{s_{g,n}}{s_{D,n}}. (16)

These three quantities are the standardized drift of the denominator, the correlation between the moment direction and the denominator gradient, and their scale ratio.

Assumption 1 (First-order weak denominator).

There are τ\tau_{\ast}\in\mathbb{R}, Σ𝕊+d\Sigma_{\ast}\in\mathbb{S}^{d}_{+} and δ0\delta_{0}\in\mathbb{R} such that

τnτ,ΣnΣ,n1/2D(Σn)δ0.\tau_{n}\to\tau_{\ast},\quad\Sigma_{n}\to\Sigma_{\ast},\quad n^{1/2}D(\Sigma_{n})\to\delta_{0}.

Define

g=N(Σ)τD(Σ),sD,=sD(Σ),sg,={gΓ(Σ)g}1/2,g_{\ast}=\nabla N(\Sigma_{\ast})-\tau_{\ast}\nabla D(\Sigma_{\ast}),\quad s_{D,\ast}=s_{D}(\Sigma_{\ast}),\quad s_{g,\ast}=\{g_{\ast}^{\top}\Gamma(\Sigma_{\ast})g_{\ast}\}^{1/2},

and assume sD,>0s_{D,\ast}>0 and sg,>0s_{g,\ast}>0.

By continuity of the polynomial gradients and of Γ\Gamma, Assumption 1 implies sD,nsD,s_{D,n}\to s_{D,\ast}, sg,nsg,s_{g,n}\to s_{g,\ast} and

μnμ=δ0sD,,rnr=gΓ(Σ)D(Σ)sg,sD,,ωnω=sg,sD,.\mu_{n}\to\mu_{\ast}=\frac{\delta_{0}}{s_{D,\ast}},\quad r_{n}\to r_{\ast}=\frac{g_{\ast}^{\top}\Gamma(\Sigma_{\ast})\nabla D(\Sigma_{\ast})}{s_{g,\ast}s_{D,\ast}},\quad\omega_{n}\to\omega_{\ast}=\frac{s_{g,\ast}}{s_{D,\ast}}.

Assumption 1 describes a denominator that drifts to zero at exactly the rate of its nondegenerate first-order noise. Continuity gives D(Σn)D(Σ)D(\Sigma_{n})\to D(\Sigma_{\ast}), and the finite limit of n1/2D(Σn)n^{1/2}D(\Sigma_{n}) then forces D(Σ)=0D(\Sigma_{\ast})=0, so the sequence approaches a zero of the denominator, which may lie in the interior of the cone or on its boundary. The condition sD,>0s_{D,\ast}>0 makes this zero first-order nondegenerate, and the condition sg,>0s_{g,\ast}>0 excludes a simultaneous first-order degeneracy of the moment function itself. The limit μ\mu_{\ast} records how many first-order standard errors separate the drifting denominator from zero, and it acts as the noncentrality parameter of the limit experiment.

The associated Wald procedure studentizes the plug-in estimator along the estimated moment direction. On the event where D(Σ^)0D(\hat{\Sigma})\neq 0 and the quadratic form below is positive, define

g^=N(Σ^)τ^D(Σ^),se^(τ^)={g^Γ(Σ^)g^}1/2n1/2|D(Σ^)|,Tn=τ^τnse^(τ^).\hat{g}=\nabla N(\hat{\Sigma})-\hat{\tau}\nabla D(\hat{\Sigma}),\quad\widehat{\mathrm{se}}(\hat{\tau})=\frac{\{\hat{g}^{\top}\Gamma(\hat{\Sigma})\hat{g}\}^{1/2}}{n^{1/2}|D(\hat{\Sigma})|},\quad T_{n}=\frac{\hat{\tau}-\tau_{n}}{\widehat{\mathrm{se}}(\hat{\tau})}. (17)

The next proposition identifies the joint limits.

Proposition 1 (Local ratio experiment).

Suppose that Assumption 1 holds. Let Z0Z_{0} and Z2Z_{2} be independent standard normal variables and put Z~=rZ2+(1r2)1/2Z0\tilde{Z}=r_{\ast}Z_{2}+(1-r_{\ast}^{2})^{1/2}Z_{0}. Then

τ^τnωZ~μ+Z2,F^D(μ+Z2)2.\hat{\tau}-\tau_{n}\rightsquigarrow\frac{\omega_{\ast}\tilde{Z}}{\mu_{\ast}+Z_{2}},\quad\widehat{F}_{D}\rightsquigarrow(\mu_{\ast}+Z_{2})^{2}. (18)

If in addition |r|<1|r_{\ast}|<1, then

TnZ~sgn(μ+Z2){12rq+q2}1/2,q=Z~μ+Z2.T_{n}\rightsquigarrow\frac{\tilde{Z}\,\operatorname{sgn}(\mu_{\ast}+Z_{2})}{\{1-2r_{\ast}q+q^{2}\}^{1/2}},\quad q=\frac{\tilde{Z}}{\mu_{\ast}+Z_{2}}. (19)

Thus, a first-order nondegenerate small denominator produces the classical Fieller ratio limit. The limiting standardized denominator (μ+Z2)2(\mu_{\ast}+Z_{2})^{2} is a noncentral chi-squared variable with one degree of freedom and noncentrality μ2\mu_{\ast}^{2}, so F^D\widehat{F}_{D} measures the local strength of identification. The Wald statistic converges to a non-Gaussian law governed by μ\mu_{\ast} and rr_{\ast}. Supplementary Section E proves Proposition 1.

Remark 6 (Role of the local experiment).

Proposition 1 is not offered as a new general theory of weak identification. Its role is to place every covariance-polynomial strategy in a common three-parameter local experiment and to make the algebraic classification operational. Once DD is known not to self-normalize, the limits μ\mu_{\ast}, rr_{\ast} and ω\omega_{\ast} determine the leading ratio geometry. When DD is a flag power, the assumptions of this experiment cannot hold on the denominator side.

5.2 Robust inversion and the boundary of the classification

We connect the ratio experiment to robust inference. For α(0,1)\alpha\in(0,1), inverting the single moment NtDN-tD gives the confidence set

𝒞1α={t:n{N(Σ^)tD(Σ^)}2q1αhtΓ(Σ^)ht},ht=N(Σ^)tD(Σ^),\mathcal{C}_{1-\alpha}=\Bigl\{t:n\{N(\hat{\Sigma})-tD(\hat{\Sigma})\}^{2}\leq q_{1-\alpha}h_{t}^{\top}\Gamma(\hat{\Sigma})h_{t}\Bigr\},\quad h_{t}=\nabla N(\hat{\Sigma})-t\nabla D(\hat{\Sigma}), (20)

where

q1α=inf{x:pr(Ux)1α},Uχ12,q_{1-\alpha}=\inf\{x\in\mathbb{R}:\mathrm{pr}(U\leq x)\geq 1-\alpha\},\quad U\sim\chi_{1}^{2},

is the (1α)(1-\alpha) quantile of the χ12\chi_{1}^{2} law.

Proposition 2 (Fieller confidence set).

Suppose that Assumption 1 holds. Then pr{τn𝒞1α}1α\mathrm{pr}\{\tau_{n}\in\mathcal{C}_{1-\alpha}\}\to 1-\alpha. Moreover, for every fixed nn, with probability one the set 𝒞1α\mathcal{C}_{1-\alpha} is nonempty and is an interval, the complement of a bounded open interval, or the whole real line, and it is bounded exactly when F^D>q1α\widehat{F}_{D}>q_{1-\alpha}.

Supplementary Section E proves Proposition 2. This is the classical Fieller and Anderson–Rubin analysis specialized to covariance polynomials (1; 13). It gives the standardized denominator a second role. Proposition 1 identifies F^D\widehat{F}_{D} as the noncentrality diagnostic of the local experiment, and Proposition 2 makes the same statistic decide whether the inverted set is bounded. Combining the two results, the probability that 𝒞1α\mathcal{C}_{1-\alpha} is bounded converges to pr{(μ+Z2)2>q1α}\mathrm{pr}\{(\mu_{\ast}+Z_{2})^{2}>q_{1-\alpha}\}, so unbounded sets retain a positive limiting frequency under weak drift, in accordance with the impossibility results for uniformly bounded confidence sets (16; 11).

We now delimit what exact self-normalization removes. As noted after Assumption 1, the drift forces D(Σ)=0D(\Sigma_{\ast})=0 while requiring sD(Σ)=sD,>0s_{D}(\Sigma_{\ast})=s_{D,\ast}>0. For a self-normalizing denominator, evaluating (4) at Σ\Sigma_{\ast} gives sD(Σ)=c1/2|D(Σ)|=0s_{D}(\Sigma_{\ast})=c^{1/2}|D(\Sigma_{\ast})|=0, so the two requirements are incompatible and no drifting sequence satisfies Assumption 1. Therefore, exact self-normalization excludes the first-order ratio experiment, not every nonregular limit.

The exclusion concerns the first order only. On Sym2\operatorname{Sym}_{2}, the polynomial D=σ122D=\sigma_{12}^{2} has sD2=4σ122(σ11σ22+σ122)s_{D}^{2}=4\sigma_{12}^{2}(\sigma_{11}\sigma_{22}+\sigma_{12}^{2}), so its first-order variance vanishes on the entire zero set and Assumption 1 cannot hold, although the ratio sD2/D2s_{D}^{2}/D^{2} is not constant. Under the drift σ12,n=ζn1/2\sigma_{12,n}=\zeta n^{-1/2} with σ11,n=σ22,n=1\sigma_{11,n}=\sigma_{22,n}=1, the standardized denominator converges in distribution to (ζ+Z)2/4(\zeta+Z)^{2}/4 for a standard normal variable ZZ, a nondegenerate limit generated at second order. Supplementary Section E verifies this example. Hence, the self-normalizing and the first-order nondegenerate regimes are two extremes of a broader classification rather than an exhaustive dichotomy.

6 Causal applications

6.1 Diagnostics for common formulas

We translate the algebra into familiar covariance diagnostics. For distinct indices a,b[d]a,b\in[d], put ρab=σab/(σaaσbb)1/2\rho_{ab}=\sigma_{ab}/(\sigma_{aa}\sigma_{bb})^{1/2} and let ρ^ab\hat{\rho}_{ab} be the same function evaluated at Σ^\hat{\Sigma}. For the marginal covariance D=σabD=\sigma_{ab}, the Gaussian first-order variance is sD2=σaaσbb+D2s_{D}^{2}=\sigma_{aa}\sigma_{bb}+D^{2}, so every zero of DD in the open cone is first-order nondegenerate and the denominator conditions of Assumption 1 can hold along suitable drifting sequences. The standardized denominator is F^D=nρ^ab2/(1+ρ^ab2)\widehat{F}_{D}=n\hat{\rho}_{ab}^{2}/(1+\hat{\rho}_{ab}^{2}), an increasing function of the familiar first-stage statistic nρ^ab2/(1ρ^ab2)n\hat{\rho}_{ab}^{2}/(1-\hat{\rho}_{ab}^{2}).

For indices e,a,b[d]e,a,b\in[d] with eae\neq a and ebe\neq b, allowing a=ba=b, define ρabe\rho_{ab\cdot e} as the correlation of the residuals from the population linear projections of coordinates aa and bb on coordinate ee, and let ρ^abe\hat{\rho}_{ab\cdot e} be its sample-covariance analogue. For the partial minor D=σeeσabσeaσebD=\sigma_{ee}\sigma_{ab}-\sigma_{ea}\sigma_{eb},

sD2=detΣ{e,a}detΣ{e,b}+3D2.s_{D}^{2}=\det\Sigma_{\{e,a\}}\det\Sigma_{\{e,b\}}+3D^{2}. (21)

When a=ba=b the minor is a principal 2×22\times 2 minor, the identity reduces to sD2=4D2s_{D}^{2}=4D^{2}, and F^D=n/4\widehat{F}_{D}=n/4 in agreement with Theorem 1. When aba\neq b, define

Fpar=nρ^abe21ρ^abe2,F^D=nFparn+4Fpar,F_{\mathrm{par}}=\frac{n\hat{\rho}_{ab\cdot e}^{2}}{1-\hat{\rho}_{ab\cdot e}^{2}},\quad\widehat{F}_{D}=\frac{nF_{\mathrm{par}}}{n+4F_{\mathrm{par}}},

so the standardized denominator is again an increasing transform of the familiar partial first-stage index FparF_{\mathrm{par}}. In both cases the standardized denominator reproduces the relevant marginal or partial screening direction without any modelling input. Supplementary Proposition F.1 gives the calculation and the exact partial-correlation formula.

We compare four representative strategies in Table 1. The table separates variance-type flag powers from slope-type marginal or nonprincipal minors. The structural parameterizations and substitutions yielding the third column are given in Supplementary Section F.1, so that each displayed pullback can be checked directly from the stated linear reduced form. The three slope-type rows can realize the denominator conditions of Assumption 1, whereas the front-door row is a flag power. The next subsection develops the front-door strategy in detail.

Table 1: Denominator structure of four linear causal strategies. Each structural pullback evaluates the denominator in a linear reduced form for that row, detailed in Supplementary Section F.1, and symbols are defined separately within each row. In the two instrument rows, π\pi is the coefficient of the instrument ZZ in the equation for the treatment XX, νZ\nu_{Z} and νW\nu_{W} are the variances of ZZ and of the conditioning covariate WW, and vZv_{Z} is the residual variance of ZZ given WW. In the front-door row, νX\nu_{X} is the treatment variance and vMv_{M} is the mediator residual variance. In the proximal row, aa and cc are the coefficients of the latent variable UU in the equations for the proxies ZZ and WW, νU\nu_{U} is the variance of UU, and vAv_{A} is the treatment residual variance. A flag power is a denominator of the form (5). Marginal slope and partial slope denote a marginal covariance and a nonprincipal two-by-two minor, and together they constitute the slope type.
Strategy Denominator Structural pullback Classification
Instrumental variable σZX\sigma_{ZX} πνZ\pi\nu_{Z} Marginal slope
Conditional instrument σWWσZXσZWσWX\sigma_{WW}\sigma_{ZX}-\sigma_{ZW}\sigma_{WX} πvZνW\pi v_{Z}\nu_{W} Partial slope
Front-door σXX(σXXσMMσXM2)\sigma_{XX}(\sigma_{XX}\sigma_{MM}-\sigma_{XM}^{2}) νX2vM\nu_{X}^{2}v_{M} Flag power
Linear proximal σZWσAAσZAσAW\sigma_{ZW}\sigma_{AA}-\sigma_{ZA}\sigma_{AW} acνUvAac\nu_{U}v_{A} Partial slope

6.2 Front-door adjustment and boundary-robust inference

We now show that the front-door studentized statistic survives arbitrary collinearity drift when the treatment–mediator coefficient is nonzero. The exact regression decomposition and the proof of the boundary result are given in Supplementary Section F.3. Consider the Gaussian reduced form

M=aX+εM,Y=bM+γX+ε,M=aX+\varepsilon_{M},\quad Y=bM+\gamma X+\varepsilon, (22)

in which XX, εM\varepsilon_{M} and ε\varepsilon are independent centred Gaussian variables with var(X)=νX>0\operatorname{var}(X)=\nu_{X}>0 and var(ε)>0\operatorname{var}(\varepsilon)>0, and the observed vector is (X,M,Y)(X,M,Y)^{\top}. The coefficients aa, bb and γ\gamma are fixed, while the mediator residual variance vM,n=var(εM)v_{M,n}=\operatorname{var}(\varepsilon_{M}) may change with nn and takes values in a fixed interval (0,v¯](0,\bar{v}]. Observations are nn independent copies of the observed vector, and Σ^\hat{\Sigma} is the known-mean sample covariance (1). Substituting the reduced form gives σXX=νX\sigma_{XX}=\nu_{X} and detΣ{X,M}=νXvM,n\det\Sigma_{\{X,M\}}=\nu_{X}v_{M,n}, so the population denominator equals νX2vM,n\nu_{X}^{2}v_{M,n}, which is the pullback recorded in Table 1 and which tends to zero whenever vM,nv_{M,n} does.

The front-door covariance functional is

τfd(Σ)=σXM(σXXσMYσXMσXY)σXX(σXXσMMσXM2),\tau_{\mathrm{fd}}(\Sigma)=\frac{\sigma_{XM}(\sigma_{XX}\sigma_{MY}-\sigma_{XM}\sigma_{XY})}{\sigma_{XX}(\sigma_{XX}\sigma_{MM}-\sigma_{XM}^{2})}, (23)

on the set where its denominator is nonzero. Direct substitution shows that τfd(Σn)=ab\tau_{\mathrm{fd}}(\Sigma_{n})=ab for every value of γ\gamma, so the functional isolates the product abab and removes the association carried by γ\gamma.

Example 2 (Exact front-door denominator).

In the front-door causal model, the coefficient γ\gamma is generated by an unobserved confounder of XX and YY rather than by a direct effect, so the total effect of XX on YY equals abab and is identified by τfd\tau_{\mathrm{fd}}. The denominator is the flag power identified in Section 3, and Theorem 1 gives

Dfd=σXXdetΣ{X,M}=σXX2var(MX),sD2=10Dfd2,F^D=n/10.D_{\mathrm{fd}}=\sigma_{XX}\det\Sigma_{\{X,M\}}=\sigma_{XX}^{2}\operatorname{var}(M\mid X),\quad s_{D}^{2}=10D_{\mathrm{fd}}^{2},\quad\widehat{F}_{D}=n/10. (24)

The last equality holds for every positive-definite sample covariance when the Gaussian studentizer is used, whatever the data-generating law.

Thus a relevance test based on the front-door denominator is exactly uninformative even though the precision of τ^\hat{\tau} can deteriorate without bound.

Because DfdD_{\mathrm{fd}} is exactly self-normalizing, no drifting sequence satisfies Assumption 1 for this strategy, and the ratio experiment of Section 5 is unavailable as a description of the boundary vM,n0v_{M,n}\to 0. That exclusion is negative information only. The following result provides the positive counterpart for the studentized estimator. The plug-in estimator is τ^=τfd(Σ^)\hat{\tau}=\tau_{\mathrm{fd}}(\hat{\Sigma}), and its Gaussian delta standard error is

se^ 2=1nτfd(Σ^)Γ(Σ^)τfd(Σ^).\widehat{\mathrm{se}}^{\,2}=\frac{1}{n}\nabla\tau_{\mathrm{fd}}(\hat{\Sigma})^{\top}\Gamma(\hat{\Sigma})\nabla\tau_{\mathrm{fd}}(\hat{\Sigma}). (25)

Both quantities are defined almost surely because Σ^0\hat{\Sigma}\succ 0 implies Dfd(Σ^)>0D_{\mathrm{fd}}(\hat{\Sigma})>0.

Proposition 3 (Boundary-robust front-door Wald inference).

Under the reduced form (22) with a0a\neq 0,

τ^abse^N1(0,1),\frac{\hat{\tau}-ab}{\widehat{\mathrm{se}}}\rightsquigarrow N_{1}(0,1), (26)

along every sequence vM,n(0,v¯]v_{M,n}\in(0,\bar{v}], including sequences with nvM,nCnv_{M,n}\to C for any C[0,]C\in[0,\infty].

Thus, arbitrarily poor mediator residual variation inflates uncertainty but does not invalidate the studentized Gaussian limit on the stated parameter region.

Remark 7 (Why studentization survives).

The covariance estimator factors exactly as the product of the slope in the regression of MM on XX and the coefficient of MM in the regression of YY on (X,M)(X,M). Conditional on the second regression design, the latter coefficient has an exact Student statistic. As vM,nv_{M,n} decreases, the first slope is estimated more accurately while the second standard error grows at the reciprocal rate. The Gaussian delta studentizer reproduces this balance exactly, so the exploding variance affects precision but not the limiting studentized law.

The restriction a0a\neq 0 in Proposition 3 is deliberate. The proposition covers b=0b=0 when a0a\neq 0, but it makes no claim on the stratum a=0a=0, whether or not bb vanishes. At the doubly singular point a=b=0a=b=0 the numerator gradient vanishes, the estimator has a product-of-normals limit at rate nn, and the nonregularity is a singular-hypothesis phenomenon of the numerator rather than a weak-denominator phenomenon (10). The exact denominator identity in (24) remains true throughout these parameter regions.

7 Numerical experiments

7.1 Design

Two Monte Carlo experiments isolate the two denominator mechanisms of Table 1, one flag power and one partial slope. The front-door design sets X=εXX=\varepsilon_{X}, M=0.8X+εMM=0.8X+\varepsilon_{M} and Y=M+εYY=M+\varepsilon_{Y}, where (εX,εY)(\varepsilon_{X},\varepsilon_{Y}) is centred Gaussian with unit marginal variances and covariance 0.50.5, and εM\varepsilon_{M} is an independent Gaussian variable with variance vMv_{M}; the target is the total effect 0.80.8. The proximal design sets Z=0.8U+εZZ=0.8U+\varepsilon_{Z}, A=0.8U+εAA=0.8U+\varepsilon_{A}, W=0.8U+εWW=0.8U+\varepsilon_{W} and Y=A+0.8U+εYY=A+0.8U+\varepsilon_{Y}, with mutually independent centred Gaussian shocks of unit variance except var(εA)=vA\operatorname{var}(\varepsilon_{A})=v_{A}; the target is the coefficient of AA, equal to one. By the structural pullbacks in Table 1, the population denominators are vMv_{M} and 0.64vA0.64v_{A}, so the grids vM{1,0.1,0.01,0.001}v_{M}\in\{1,0.1,0.01,0.001\} and vA{1,0.3,0.1,0.03,0.01}v_{A}\in\{1,0.3,0.1,0.03,0.01\} trace the approach to the two denominator boundaries. The population value of nD(Σ)2/sD2(Σ)nD(\Sigma)^{2}/s_{D}^{2}(\Sigma) equals n/10n/10 identically in the front-door design and falls from 63.763.7 to 0.10.1 across the proximal grid (Supplementary Table G.1).

For each cell, we draw the known-mean sample covariance directly from nΣ^𝒲d(n,Σ)n\hat{\Sigma}\sim\mathcal{W}_{d}(n,\Sigma), with d=3d=3 and d=4d=4 respectively, which is distributionally equivalent to simulating nn centred Gaussian observations. The converse classification of Section 4 is not invoked at d=4d=4: Theorem 1, Proposition 1 and the calibrations of Section 6 hold in every dimension, and the proximal denominator depends only on the (A,Z,W)(A,Z,W) block, so its slope-type status is unchanged by the ambient dimension (Supplementary Lemma B.4). We use n=1000n=1000 and 100000100000 replications per cell. Each replication yields the plug-in estimator τ^\hat{\tau}, the standardized denominator F^D\widehat{F}_{D}, and, in the proximal design, the partial statistic FparF_{\mathrm{par}} of Section 6 with (e,a,b)=(A,Z,W)(e,a,b)=(A,Z,W) together with the deliberately naive marginal statistic nρ^AZ2/(1ρ^AZ2)n\hat{\rho}_{AZ}^{2}/(1-\hat{\rho}_{AZ}^{2}), which measures only the treatment–proxy association. Cells are summarized by medians and by the interquartile range of τ^\hat{\tau} divided by 1.3491.349, because the ratio estimator need not possess moments in the weak cells, and by the empirical coverage of nominal 95% Wald intervals based on the delta standard error in (17) and of the inversion (20). The estimator, both intervals and all diagnostics were well defined in every replication. The Python scripts reproducing both the numerical experiments and the real data experiments in Section 8 are available at https://github.com/shutech2001/self-normalizing-denominators-experiments.

The theory yields three predictions. First, Theorem 1 (iv) and Example 2 imply that F^D=n/10=100\widehat{F}_{D}=n/10=100 in every front-door replication, and Proposition 3 implies nominal Wald coverage along the whole grid, whose smallest cell has nvM=1nv_{M}=1 and so lies inside the boundary regime nvM,nC[0,]nv_{M,n}\to C\in[0,\infty]. Secondly, the weak proximal cells have population values of nD(Σ)2/sD2(Σ)nD(\Sigma)^{2}/s_{D}^{2}(\Sigma) below one, matching the drift regime of Assumption 1, so Proposition 1 predicts the noncentral limit for F^D\widehat{F}_{D} in (18) and distorted Wald inference. Thirdly, Proposition 2 predicts near-nominal coverage for the inversion (20) in every cell, with sets that are bounded exactly when F^D\widehat{F}_{D} exceeds the 0.950.95 quantile 3.843.84 of the χ12\chi^{2}_{1} law.

7.2 Results

Table 2 reports the results. In the front-door design the median standardized denominator equals the theoretical constant 100.0100.0 in every cell; by Theorem 1 (iv) the statistic equals n/10n/10 in every replication, so it carries no information about vMv_{M}. Precision, by contrast, deteriorates by a factor of eighteen: the robust standard deviation of τ^\hat{\tau} rises from 0.0380.038 to 0.6960.696 as vMv_{M} falls from 11 to 0.0010.001, and the median delta standard error tracks it closely, rising from 0.0380.038 to 0.6930.693. Wald coverage lies between 0.9490.949 and 0.9500.950 in all four cells, including the boundary cell with nvM=1nv_{M}=1, and inversion coverage lies between 0.9530.953 and 0.9550.955, slightly above the nominal level. This confirms the first prediction: the studentized statistic remains accurate while the denominator diagnostic is exactly uninformative.

In the proximal design the naive marginal statistic rises from 179.7179.7 to 624.6624.6 as vAv_{A} decreases, while the partial statistic falls from 85.685.6 to 0.50.5 and the median F^D\widehat{F}_{D} falls from 63.863.8 to 0.50.5; the latter two columns accord with the identity between F^D\widehat{F}_{D} and FparF_{\mathrm{par}} stated after (21). At vA=0.03v_{A}=0.03 and 0.010.01 the median bias reaches 0.440.44 and 0.880.88 and Wald coverage falls to 0.9300.930 and 0.9310.931, although the naive diagnostic is largest exactly there. In those two cells the median standard error exceeds the robust standard deviation, so the Wald failure is one of centring and distributional shape rather than of scale, as the ratio limit in (18) predicts. Inversion coverage remains between 0.9510.951 and 0.9530.953 throughout, but the protection has a price: in the two weakest cells the median F^D\widehat{F}_{D} lies below 3.843.84, so more than half of the inverted sets are unbounded.

Table 2: Gaussian simulation results based on 100000100000 replications per cell. Med., median across replications; s.e., Gaussian delta standard error; robust s.d., interquartile range of τ^\hat{\tau} divided by 1.3491.349; cov., empirical coverage, the two entries giving the nominal 95% Wald interval and the inversion (20). Dashes indicate diagnostics defined only for the proximal design. Bias entries of 0.000.00 are zero to the precision shown. The largest Monte Carlo standard error for a coverage entry is 0.00080.0008.
Parameter Med. F^D\widehat{F}_{D} Med. partial FF Med. naive FF Robust s.d. Med. s.e. Med. bias Wald/inv. cov.
Front-door: parameter vMv_{M}
1.001.00 100.0100.0 0.0380.038 0.0380.038 0.000.00 0.9490.949/0.9540.954
0.100.10 100.0100.0 0.0700.070 0.0700.070 0.000.00 0.9500.950/0.9550.955
0.010.01 100.0100.0 0.2170.217 0.2190.219 0.000.00 0.9490.949/0.9540.954
0.0010.001 100.0100.0 0.6960.696 0.6930.693 0.000.00 0.9500.950/0.9530.953
Proximal: parameter vAv_{A}
1.001.00 63.863.8 85.685.6 179.7179.7 0.0640.064 0.0630.063 0.000.00 0.9520.952/0.9530.953
0.300.30 26.526.5 29.629.6 362.0362.0 0.1730.173 0.1700.170 0.000.00 0.9520.952/0.9520.952
0.100.10 6.26.2 6.46.4 510.2510.2 0.4910.491 0.4680.468 0.010.01 0.9350.935/0.9530.953
0.030.03 0.90.9 0.90.9 595.2595.2 1.1871.187 1.4331.433 0.440.44 0.9300.930/0.9520.952
0.010.01 0.50.5 0.50.5 624.6624.6 1.4601.460 2.0372.037 0.880.88 0.9310.931/0.9510.951

The two designs estimate different targets in different models, so the experiment compares diagnostics rather than estimators. Its message is the dissociation predicted by the classification: front-door adjustment loses precision without entering the first-order ratio regime, whereas the proximal estimator fails in the partial-covariance direction that the marginal treatment–proxy statistic does not measure. Supplementary Section G gives complete cell summaries.

8 Real data experiments

8.1 Data and diagnostic design

We audit the public SUPPORT right-heart-catheterization data (8), which contain 57355735 critically ill patients, an indicator of right-heart catheterization on the first study day, survival time in days truncated at 3030, and baseline physiological measurements. Following proximal analyses of these data (26; 38), we consider treatment proxies pafi1 and paco21 and outcome proxies ph1 and hema1. The analysis is the diagnostic pathway described in the introduction rather than a new clinical causal analysis: the symbolic factor-and-rank step has already classified the proximal denominator as a nonprincipal slope minor (Example 1), and the data-level step measures the implied partial direction for each candidate proxy pair.

Treatment, outcome and the four proxies are residualized by least squares on age, sex, primary and secondary disease categories, do-not-resuscitate status, the SUPPORT two-month survival estimate and the APACHE score; missing continuous covariates are median-imputed, and missing categorical covariates are mode-imputed and coded by indicator variables with one level omitted. The resulting design has rank r=19r=19, and all diagnostics use the residual degrees of freedom ν=nr=5716\nu=n-r=5716 in place of nn. For residualized treatment AA, treatment proxy ZZ and outcome proxy WW, we report the naive marginal statistic νρ^AZ2/(1ρ^AZ2)\nu\hat{\rho}_{AZ}^{2}/(1-\hat{\rho}_{AZ}^{2}), the partial statistic νρ^ZWA2/(1ρ^ZWA2)\nu\hat{\rho}_{ZW\mathbin{\cdot}A}^{2}/(1-\hat{\rho}_{ZW\mathbin{\cdot}A}^{2}) and two versions of the standardized denominator: the Gaussian version of Section 6, and a sandwich version in which Γ(Σ^)\Gamma(\hat{\Sigma}) in (3) is replaced by the empirical covariance matrix of the six second-moment scores. Because the original treatment is binary, the residualized treatment is non-Gaussian. The Gaussian fourth-moment formula therefore serves only as a reference value. The sandwich version is the appropriate studentization under the empirical law (Remark 1). Sampling variability of the proximal estimates and of the sandwich statistic is assessed by 20002000 nonparametric percentile bootstrap replications over patients, each of which repeats the imputation, coding and residualization. The audit is descriptive: proxy strength in the denominator direction is measurable, but proxy validity is not testable from these diagnostics.

Full variable definitions and preprocessing details are given in Supplementary Section H.1, and the second-moment score construction and bootstrap implementation are given in Supplementary Section H.2.

8.2 Results

Table 3 reports the audit, and the two pairs involving ph1 reverse their ordering across diagnostics. The pair pafi1/ph1 appears strong in the treatment–proxy direction, with naive statistic 173.4173.4, but is weak in the direction that enters the denominator. Its partial statistic is 3.23.2 and its sandwich statistic is 2.32.3, with a bootstrap interval reaching essentially zero. The pair paco21/ph1 shows the reverse pattern, with naive statistic 16.416.4 but partial statistic 1980.01980.0 and sandwich statistic 315.2315.2, bounded well away from zero. The Gaussian version is the monotone transform of the partial statistic given after (21) with ν\nu in place of nn, which explains why those two columns nearly coincide when the partial statistic is small relative to ν\nu. The sandwich column carries the additional distributional information. It is smaller than the Gaussian value for every pair, by factors of 0.720.72 to 0.810.81 for three pairs and 0.380.38 for paco21/ph1. Thus, the Gaussian formula overstates denominator strength under the empirical fourth moments, most severely for the strongest pair.

Table 3: Denominator audit for the SUPPORT data. CI, percentile bootstrap confidence interval based on 20002000 replications. All statistics are scaled by the residual degrees of freedom ν=5716\nu=5716. The naive and partial statistics are the indices νρ^2/(1ρ^2)\nu\hat{\rho}^{2}/(1-\hat{\rho}^{2}) defined in the text, not finite-sample FF tests, and the naive statistic depends only on the treatment proxy. The sandwich statistic uses the empirical covariance matrix of the six second-moment scores.
Treatment proxy Outcome proxy Naive FF Partial FF Gaussian F^D\widehat{F}_{D} Sandwich F^D\widehat{F}_{D} (95% CI)
pafi1 ph1 173.4173.4 3.23.2 3.23.2 2.32.3 (0.00.0, 12.712.7)
pafi1 hema1 173.4173.4 32.332.3 31.631.6 25.725.7 (9.29.2, 51.351.3)
paco21 ph1 16.416.4 1980.01980.0 830.0830.0 315.2315.2 (218.2218.2, 461.9461.9)
paco21 hema1 16.416.4 69.369.3 66.166.1 51.351.3 (30.330.3, 78.478.4)

The plug-in proximal estimates for the four pairs, in the order of Table 3, are 1.25-1.25, 1.41-1.41, 1.33-1.33 and 1.14-1.14 days, with percentile intervals (2.25,0.14)(-2.25,0.14), (2.04,0.81)(-2.04,-0.81), (1.84,0.81)(-1.84,-0.81) and (1.69,0.56)(-1.69,-0.56). Only the interval for the weak pair pafi1/ph1 covers zero. For that pair, however, the bootstrap interval inherits the nonstandard ratio behaviour described in Section 5 and is reported for completeness only; the inversion (20) is the appropriate construction in that regime. None of these intervals validates the proximal bridge assumptions, which the audit cannot test.

The audit illustrates the practical content of the classification. A strong association between treatment and a proposed treatment proxy neither implies nor precludes strength of the nonprincipal minor that identifies the proximal bridge; the symbolic step locates the relevant covariance direction before any estimation, and the partial and sandwich statistics then measure it for each candidate allocation.

9 Discussion

The classification turns the choice among rational identification formulas into a design decision. When several certificates identify the same target, their denominators can belong to different algebraic classes even though the ratios agree on the model. The classes are complementary rather than ordered, because a flag denominator removes first-order denominator nonregularity while a slope-type formula may be more precise at strongly identified distributions. Therefore, we recommend reporting the symbolic denominator class alongside any covariance plug-in estimate. When a denominator is not self-normalizing and its relevant standardized diagnostic is weak, inference should also include a robust inversion.

The main open algebraic problem is the rigidity of the mixed kernel line. A proof that the kernel line in Theorem 3 (iii) is constant would complete the unconditional three-variable converse, while a counterexample would produce a self-normalizing denominator outside the flag class. Either resolution must exploit the integrability of the relative gradient beyond its vanishing Hessian. In dimension at least four, a converse would additionally require invariants beyond the first two power sums of the relative gradient.

The principal distributional question is self-normalization beyond the Gaussian covariance operator. Ellipticity only rescales the constant, so the first open case is a semiparametric family with unrestricted fourth cumulants. A characterization of polynomials whose sandwich variance is proportional to their square over such a family would determine when the standardized denominator remains a samplewise constant under a matched studentizer. This, in turn, would establish how far the exact screening interpretation extends.

The principal methodological question concerns selection among many candidate strategies. Choosing a certificate because its observed denominator diagnostic is largest reintroduces a data-adaptive weak-direction problem. Therefore, applications with many proxies or graph-derived formulas call for selection-adjusted denominator diagnostics and simultaneous robust inversions (40; 4). Because the factor-and-rank audit is symbolic, it could also be attached to computer-algebra identification pipelines, so that every certificate is generated together with its denominator class.

Declaration of the use of generative AI and AI-assisted technologies

During the preparation of this work the author used ChatGPT and Claude in order to assist with writing and refactoring the simulation code. After using these tools the author reviewed and edited the content as necessary and takes full responsibility for the content of the publication.

Acknowledgement

Shu Tamano was supported by JSPS KAKENHI Grant Numbers 25K24203.

Supplementary material

The Supplementary Material includes a notation table, all omitted proofs, the elliptical and unknown-mean extensions, the recursive-flag and boundary-rigidity converses, the Fieller and minor-calibration results, detailed simulation and SUPPORT analyses, and the algebraic obstructions to an unrestricted three-variable converse.

References

  • Anderson and Rubin (1949) T. W. Anderson and H. Rubin Estimation of the parameters of a single equation in a complete system of stochastic equations. The Annals of Mathematical Statistics 20 (1), pp. 46–63. Cited by: §1.3, §5.2, §E.
  • Andrews and Cheng (2012) D. W. K. Andrews and X. Cheng Estimation and inference with weak, semi-strong, and strong identification. Econometrica 80 (5), pp. 2153–2211. Cited by: §1.1, §1.3.
  • Andrews et al. (2019) I. Andrews, J. H. Stock, and L. Sun Weak instruments in instrumental variables regression: theory and practice. Annual Review of Economics 11, pp. 727–753. Cited by: §1.1, §1.3.
  • Bennett et al. (2026) A. Bennett, N. Kallus, X. Mao, W. K. Newey, V. Syrgkanis, and M. Uehara Inference on strongly identified functionals of weakly identified functions. Journal of the Royal Statistical Society Series B: Statistical Methodology 88 (3), pp. 998–1028. External Links: Document Cited by: §1.3, §9.
  • Bhatia (2007) R. Bhatia Positive definite matrices. Princeton University Press. Cited by: §1.3.
  • Brito and Pearl (2002) C. Brito and J. Pearl Generalized instrumental variables. In Proceedings of the 18th Conference on Uncertainty in Artificial Intelligence, pp. 85–93. Cited by: §1.1, §1.3.
  • Ciliberto et al. (2008) C. Ciliberto, F. Russo, and A. Simis Homaloidal hypersurfaces and hypersurfaces with vanishing hessian. Advances in Mathematics 218 (6), pp. 1759–1805. Cited by: §1.3, §I.1, Remark 4.
  • Connors et al. (1996) A. F. Connors, T. Speroff, N. V. Dawson, C. Thomas, F. E. Harrell, D. Wagner, N. Desbiens, L. Goldman, A. W. Wu, R. M. Califf, W. J. Fulkerson, H. Vidaillet, S. Broste, P. Bellamy, J. Lynn, and W. A. Knaus The effectiveness of right heart catheterization in the initial care of critically ill patients. The Journal of the American Medical Association 276 (11), pp. 889–897. External Links: Document Cited by: §8.1, §H.1.
  • Drton and Weihs (2016) M. Drton and L. Weihs Generic identifiability of linear structural equation models by ancestor decomposition. Scandinavian Journal of Statistics 43 (4), pp. 1035–1045. Cited by: §1.3.
  • Drton and Xiao (2016) M. Drton and H. Xiao Wald tests of singular hypotheses. Bernoulli 22 (1), pp. 38–59. Cited by: §1.1, §6.2.
  • Dufour (1997) J.-M. Dufour Some impossibility theorems in econometrics with applications to structural and dynamic models. Econometrica 65 (6), pp. 1365–1387. Cited by: §1.1, §1.3, §5.2.
  • Faraut and Korányi (1994) J. Faraut and A. Korányi Analysis on symmetric cones. Oxford University Press. Cited by: §1.3.
  • Fieller (1954) E. C. Fieller Some problems in interval estimation. Journal of the Royal Statistical Society Series B: Statistical Methodology 16 (2), pp. 175–185. Cited by: §1.3, §5.2, §E.
  • Foygel et al. (2012) R. Foygel, J. Draisma, and M. Drton Half-trek criterion for generic identifiability of linear structural equation models. The Annals of Statistics 40 (3), pp. 1682–1713. Cited by: §1.1, §1.3.
  • García-Puente et al. (2010) L. D. García-Puente, S. Spielvogel, and S. Sullivant Identifying causal effects with computer algebra. In Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, pp. 193–200. Cited by: §1.3.
  • Gleser and Hwang (1987) L. J. Gleser and J. T. Hwang The nonexistence of 100(1α)%100(1-\alpha)\% confidence sets of finite expected diameter in errors-in-variables and related models. The Annals of Statistics 15 (4), pp. 1351–1362. Cited by: §1.1, §1.3, §5.2.
  • Gondim and Russo (2015) R. Gondim and F. Russo On cubic hypersurfaces with vanishing hessian. Journal of Pure and Applied Algebra 219 (4), pp. 779–806. Cited by: §1.3, §I.1, Remark 4.
  • Gordan and Noether (1876) P. Gordan and M. Noether Ueber die algebraischen formen, deren hesse’sche determinante identisch verschwindet. Mathematische Annalen 10, pp. 547–568. Cited by: §1.3, §4.1.
  • Gordon et al. (2021) S. L. Gordon, V. M. Kumar, L. J. Schulman, and P. Srivastava Condition number bounds for causal inference. In Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence, Proceedings of Machine Learning Research, Vol. 161, pp. 1948–1957. Cited by: §1.3.
  • Harris (1992) J. Harris Algebraic geometry: a first course. Graduate Texts in Mathematics, Vol. 133, Springer. Cited by: §D.3.
  • Henckel et al. (2024) L. Henckel, M. Buttenschoen, and M. H. Maathuis Graphical tools for selecting conditional instrumental sets. Biometrika 111 (3), pp. 771–788. External Links: Document Cited by: §1.1, §1.3.
  • Iwashita and Siotani (1994) T. Iwashita and M. Siotani Asymptotic distributions of functions of a sample covariance matrix under the elliptical distribution. The Canadian Journal of Statistics 22 (2), pp. 273–283. Cited by: §1.3.
  • Kuroki and Pearl (2014) M. Kuroki and J. Pearl Measurement bias and effect restoration in causal inference. Biometrika 101 (2), pp. 423–437. External Links: Document Cited by: §1.1, §1.3.
  • Lee et al. (2022) D. S. Lee, J. McCrary, M. J. Moreira, and J. Porter Valid tt-ratio inference for IV. American Economic Review 112 (10), pp. 3260–3290. External Links: Document Cited by: §1.3.
  • Lichtenstein (1982) W. Lichtenstein A system of quadrics describing the orbit of the highest weight vector. Proceedings of the American Mathematical Society 84 (4), pp. 605–608. Cited by: §1.3, §I.2.
  • Liu et al. (2025) J. Liu, C. Park, K. Li, and E. J. Tchetgen Tchetgen Regression-based proximal causal inference. American Journal of Epidemiology 194 (7), pp. 2030–2036. External Links: Document Cited by: §8.1.
  • Lossen (2004) C. Lossen When does the hessian determinant vanish identically? on gordan and noether’s proof of hesse’s claim. Bulletin of the Brazilian Mathematical Society, New Series 35, pp. 71–82. Cited by: §1.3, §4.1, §I.1.
  • Mareis et al. (2026) L. Mareis, N. Sturma, and M. Drton Semiparametric inference for half-trek estimators in linear structural equation models. arXiv preprint arXiv:2606.26931. Cited by: §1.3.
  • Miao et al. (2018) W. Miao, Z. Geng, and E. J. Tchetgen Tchetgen Identifying causal effects with proxy variables of an unmeasured confounder. Biometrika 105 (4), pp. 987–993. External Links: Document Cited by: §1.1, §1.3.
  • Muirhead (1982) R. J. Muirhead Aspects of multivariate statistical theory. Wiley, New York. Cited by: §1.3, §3.2, §C, §E, §F.3.
  • Pearl (1995) J. Pearl Causal diagrams for empirical research. Biometrika 82 (4), pp. 669–688. Cited by: §1.1, §1.3.
  • Perazzo (1900) U. Perazzo Sulle varietà cubiche la cui hessiana svanisce identicamente. Giornale Di Matematiche Di Battaglini 38, pp. 337–354. Cited by: §1.3, §I.1, Remark 4.
  • Sankararaman et al. (2022) K. A. Sankararaman, A. Louis, and N. Goyal Robust identifiability in linear structural equation models of causal inference. In Proceedings of the 38th Conference on Uncertainty in Artificial Intelligence, Proceedings of Machine Learning Research, Vol. 180, pp. 1728–1737. Cited by: §1.3.
  • Schulman and Srivastava (2016) L. J. Schulman and P. Srivastava Stability of causal inference. In Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, pp. 666–675. Cited by: §1.3.
  • Shapiro and Browne (1987) A. Shapiro and M. W. Browne Analysis of covariance structures under elliptical distributions. Journal of the American Statistical Association 82 (400), pp. 1092–1097. External Links: Document Cited by: §1.3.
  • Staiger and Stock (1997) D. Staiger and J. H. Stock Instrumental variables regression with weak instruments. Econometrica 65 (3), pp. 557–586. Cited by: §1.1, §1.3.
  • Stock and Wright (2000) J. H. Stock and J. H. Wright GMM with weak identification. Econometrica 68 (5), pp. 1055–1096. Cited by: §1.1, §1.3.
  • Tchetgen Tchetgen et al. (2024) E. J. Tchetgen Tchetgen, A. Ying, Y. Cui, X. Shi, and W. Miao An introduction to proximal causal inference. Statistical Science 39 (3), pp. 375–390. External Links: Document Cited by: §1.1, §1.3, §8.1.
  • van der Zander et al. (2015) B. van der Zander, J. Textor, and M. Liśkiewicz Efficiently finding conditional instruments for causal inference. In Proceedings of the 24th International Joint Conference on Artificial Intelligence, pp. 3243–3249. Cited by: §1.1, §1.3.
  • Wang et al. (2025) R. Wang, K. C. G. Chan, and T. Ye GMM with many weak moment conditions and nuisance parameters: general theory and applications to causal inference. arXiv preprint arXiv:2505.07295. Cited by: §1.3, §9.

Supplementary Material for

“Self-normalizing denominators in rational causal estimation”

A Notation

A.1 Linear-algebraic, algebraic and geometric conventions

We collect the conventions needed to read the Supplementary Material independently. Fix d1d\geq 1, write [d]={1,,d}[d]=\{1,\ldots,d\}, put pd=d(d+1)/2p_{d}=d(d+1)/2 and d={(i,j):1ijd}\mathcal{I}_{d}=\{(i,j):1\leq i\leq j\leq d\}, and use the lexicographic order (1,1),(1,2),,(1,d),(2,2),,(d,d)(1,1),(1,2),\ldots,(1,d),(2,2),\ldots,(d,d). Vectors are columns, AA^{\top} is transpose, A=(A1)A^{-\top}=(A^{-1})^{\top} for nonsingular AA, ImI_{m} is the identity matrix and eje_{j} is the jjth standard basis vector. The symbols tr\operatorname{tr}, det\det, rank\operatorname{rank} and spec\operatorname{spec} denote trace, determinant, rank and the eigenvalue multiset. For I,J[d]I,J\subseteq[d], AI,JA_{I,J} is the corresponding submatrix, AI=AI,IA_{I}=A_{I,I} and A1:r=A{1,,r},{1,,r}A_{1:r}=A_{\{1,\ldots,r\},\{1,\ldots,r\}}.

The space Symd={Ad×d:A=A}\operatorname{Sym}_{d}=\{A\in\mathbb{R}^{d\times d}:A=A^{\top}\} has positive-definite and positive-semidefinite cones 𝕊++d\mathbb{S}^{d}_{++} and 𝕊+d\mathbb{S}^{d}_{+}. For ASymdA\in\operatorname{Sym}_{d}, vech(A)pd\operatorname{vech}(A)\in\mathbb{R}^{p_{d}} stacks the upper-triangular entries in the stated order. The coordinate ring [Symd]=[σij:(i,j)d]\mathbb{R}[\operatorname{Sym}_{d}]=\mathbb{R}[\sigma_{ij}:(i,j)\in\mathcal{I}_{d}] is identified with the polynomial maps Symd\operatorname{Sym}_{d}\to\mathbb{R}; P0P\equiv 0 means equality in this ring, deg(P)\deg(P) is total degree, and homogeneity, divisibility and coprimality are understood in the indicated coordinate ring. For P[Symd]P\in\mathbb{R}[\operatorname{Sym}_{d}], P\nabla P is the coordinate gradient in the vech\operatorname{vech} order, and

dPΣ(H)=ddtP(Σ+tH)|t=0,dPΣ(H)=tr{GP(Σ)H},\mathrm{d}P_{\Sigma}(H)=\left.\frac{\mathrm{d}}{\mathrm{d}t}P(\Sigma+tH)\right|_{t=0},\quad\mathrm{d}P_{\Sigma}(H)=\operatorname{tr}\{G_{P}(\Sigma)H\},

define the Fréchet differential and symmetric gradient. The relative gradient is RP(Σ)=GP(Σ)Σ/P(Σ)R_{P}(\Sigma)=G_{P}(\Sigma)\Sigma/P(\Sigma) on {P0}\{P\neq 0\}.

Throughout the remaining algebraic conventions, kk denotes either \mathbb{R} or \mathbb{C}. For a real vector space VV, its complexification is V=VV_{\mathbb{C}}=V\otimes_{\mathbb{R}}\mathbb{C}. The notation GLm(k)\operatorname{GL}_{m}(k) denotes the group of nonsingular m×mm\times m matrices over kk, spank(S)\operatorname{span}_{k}(S) is the linear span of SS, and ker(A)\ker(A), range(A)\operatorname{range}(A) and rowspace(A)\operatorname{rowspace}(A) are the kernel, range and row space of a matrix. When the field is clear, the subscript on span\operatorname{span} is omitted. For a real subspace VdV\subseteq\mathbb{R}^{d}, V={xd:xv=0 for every vV}V^{\perp}=\{x\in\mathbb{R}^{d}:x^{\top}v=0\text{ for every }v\in V\}, and diag(a1,,am)\operatorname{diag}(a_{1},\ldots,a_{m}) is the diagonal matrix with the displayed entries. The matrix unit Eijkm×mE_{ij}\in k^{m\times m} has a one in position (i,j)(i,j) and zeros elsewhere.

For an integral domain RR, Frac(R)\operatorname{Frac}(R) is its field of fractions. If RR is a unique factorization domain, abbreviated UFD, then gcd(f1,,fs)\gcd(f_{1},\ldots,f_{s}) denotes a greatest common divisor, defined up to multiplication by a unit. A vector with entries in RR is primitive when the greatest common divisor of its entries is a unit. For a finite-dimensional kk-vector space VV and K0K\geq 0, SymK(V)\operatorname{Sym}^{K}(V) is the KKth symmetric tensor power,

SymK(V)=VK/v1vKvπ(1)vπ(K):v1,,vKV,π𝒮K,\operatorname{Sym}^{K}(V)=V^{\otimes K}/\langle v_{1}\otimes\cdots\otimes v_{K}-v_{\pi(1)}\otimes\cdots\otimes v_{\pi(K)}:v_{1},\ldots,v_{K}\in V,\ \pi\in\mathcal{S}_{K}\rangle,

where the angle brackets denote linear span, 𝒮K\mathcal{S}_{K} is the symmetric group on {1,,K}\{1,\ldots,K\}, and Sym0(V)=k\operatorname{Sym}^{0}(V)=k. Its dual is naturally the space of homogeneous polynomial functions of degree KK on VV. After choosing the standard basis, Sym2(km)\operatorname{Sym}^{2}(k^{m}) is identified with the space of symmetric m×mm\times m matrices over kk.

For a finite-dimensional kk-vector space VV, (V)=(V{0})/k×\mathbb{P}(V)=(V\setminus\{0\})/k^{\times} is its projective space, and [v][v] is the point represented by v0v\neq 0. We write m=(m+1)\mathbb{P}^{m}_{\mathbb{C}}=\mathbb{P}(\mathbb{C}^{m+1}) and Gr(r,V)\operatorname{Gr}(r,V) for the Grassmannian of rr-dimensional linear subspaces of VV; thus Gr(r,d)=Gr(r,d)\operatorname{Gr}(r,d)_{\mathbb{C}}=\operatorname{Gr}(r,\mathbb{C}^{d}). More generally, a subscript \mathbb{C} on a real algebraic variety or construction denotes extension of scalars from \mathbb{R} to \mathbb{C}, equivalently the complex variety defined by the same real polynomial equations. A dashed arrow f:XYf:X\dashrightarrow Y denotes a rational map: it is represented by a morphism on a Zariski-dense open subset of XX, and two representatives are identified when they agree on a Zariski-dense open subset. When XX is irreducible, every nonempty Zariski-open subset is dense. For a homogeneous polynomial pp, the notation {p=0}red\{p=0\}_{\mathrm{red}} denotes the reduced projective hypersurface defined by the radical ideal (p)\sqrt{(p)}; it has the same underlying zero set as {p=0}\{p=0\} but carries no multiplicities. Terms such as Zariski open, closed, dense and irreducible refer to this algebraic topology over the field stated in the argument.

For a differentiable map F:SymdWF:\operatorname{Sym}_{d}\to W into a finite-dimensional vector space and HSymdH\in\operatorname{Sym}_{d}, its directional derivative is

HF(Σ)=ddtF(Σ+tH)|t=0.\partial_{H}F(\Sigma)=\left.\frac{\mathrm{d}}{\mathrm{d}t}F(\Sigma+tH)\right|_{t=0}.

For scalar FF, this equals dFΣ(H)\mathrm{d}F_{\Sigma}(H); for matrix-valued FF it is taken entrywise. Ordinary symbols such as q\partial_{q} and ri\partial_{r_{i}} denote coordinate derivatives. The symbol HessF\operatorname{Hess}F means the coordinate Hessian of a scalar function on the affine space Symd\operatorname{Sym}_{d}. When a Riemannian metric gg is explicitly fixed, gradgf\operatorname{grad}_{g}f, g\nabla^{g} and Hessgf\operatorname{Hess}_{g}f denote its gradient, Levi–Civita connection and Riemannian Hessian, and g\|\cdot\|_{g} is the induced norm.

For a square matrix AA, adj(A)\operatorname{adj}(A) is its classical adjugate, so Aadj(A)=adj(A)A=det(A)IA\operatorname{adj}(A)=\operatorname{adj}(A)A=\det(A)I, where II has the same order as AA. For a polynomial FF on Sym3\operatorname{Sym}_{3}, the adjugate transform is F=FadjF^{\dagger}=F\circ\operatorname{adj}. If VdV\subseteq\mathbb{R}^{d} has dimension rr, MVM_{V} is any full-row-rank matrix with row space VV and ΔV(Σ)=det(MVΣMV)\Delta_{V}(\Sigma)=\det(M_{V}\Sigma M_{V}^{\top}). If PVd×rP_{V}\in\mathbb{R}^{d\times r} has range VV, then FV(S)=F(PVSPV)F_{V}(S)=F(P_{V}SP_{V}^{\top}) is the face restriction and 𝕊+(V)={A𝕊+d:range(A)V}\mathbb{S}_{+}(V)=\{A\in\mathbb{S}^{d}_{+}:\operatorname{range}(A)\subseteq V\}. For a polynomial FF on Sym3\operatorname{Sym}_{3}, a plane VV is called FF-good when FV0F_{V}\not\equiv 0; when FF is clear, such a plane is called good. If finitely many polynomials are under discussion, a generic good plane means a plane in a common nonempty Zariski-open set on which all indicated restrictions are nonzero and retain their generic factorization type. On Sym3\operatorname{Sym}_{3} we use δ(Σ)=det(Σ)\delta(\Sigma)=\det(\Sigma) and pF(u)=F(uu)p_{F}(u)=F(uu^{\top}).

For later representation-theoretic notation, let glm(k)\mathrm{gl}_{m}(k) be the vector space of all m×mm\times m matrices. Its derived congruence action on polynomial functions FF on Symm\operatorname{Sym}_{m} is

(ρ(A)F)(Σ)=ddtF(etAΣetA)|t=0,Aglm(k).(\rho(A)F)(\Sigma)=\left.\frac{\mathrm{d}}{\mathrm{d}t}F(e^{tA}\Sigma e^{tA^{\top}})\right|_{t=0},\quad A\in\mathrm{gl}_{m}(k).

With the preceding matrix units, ρ(Eij)F=2{GF(Σ)Σ}ij\rho(E_{ij})F=2\{G_{F}(\Sigma)\Sigma\}_{ij}.

A.2 Probability and asymptotic conventions

For a probability law QQ, 𝔼Q\mathbb{E}_{Q}, varQ\operatorname{var}_{Q}, covQ\operatorname{cov}_{Q}, corrQ\operatorname{corr}_{Q} and prQ\mathrm{pr}_{Q} denote expectation, variance, covariance, correlation and probability whenever the relevant moments exist. Conditional versions are under the same governing law. For centred Y1,,Y4Y_{1},\ldots,Y_{4}, the fourth joint cumulant is

cumQ(Y1,Y2,Y3,Y4)=𝔼Q(Y1Y2Y3Y4)covQ(Y1,Y2)covQ(Y3,Y4)covQ(Y1,Y3)covQ(Y2,Y4)covQ(Y1,Y4)covQ(Y2,Y3).\begin{split}\operatorname{cum}_{Q}(Y_{1},Y_{2},Y_{3},Y_{4})&=\mathbb{E}_{Q}(Y_{1}Y_{2}Y_{3}Y_{4})-\operatorname{cov}_{Q}(Y_{1},Y_{2})\operatorname{cov}_{Q}(Y_{3},Y_{4})\\ &\quad-\operatorname{cov}_{Q}(Y_{1},Y_{3})\operatorname{cov}_{Q}(Y_{2},Y_{4})-\operatorname{cov}_{Q}(Y_{1},Y_{4})\operatorname{cov}_{Q}(Y_{2},Y_{3}).\end{split}

For random elements Y1,,YmY_{1},\ldots,Y_{m}, σ(Y1,,Ym)\sigma(Y_{1},\ldots,Y_{m}) is the smallest σ\sigma-algebra with respect to which all YjY_{j} are measurable. It is not automatically completed; as usual, conditional expectations are defined up to QQ-null sets.

The law Nm(μ,Ω)N_{m}(\mu,\Omega) is Gaussian with mean μm\mu\in\mathbb{R}^{m} and covariance Ω𝕊+m\Omega\in\mathbb{S}^{m}_{+}. For integer ν1\nu\geq 1, 𝒲m(ν,Ω)\mathcal{W}_{m}(\nu,\Omega) is the law of =1νZZ\sum_{\ell=1}^{\nu}Z_{\ell}Z_{\ell}^{\top} for independent ZNm(0,Ω)Z_{\ell}\sim N_{m}(0,\Omega); χν2=𝒲1(ν,1)\chi^{2}_{\nu}=\mathcal{W}_{1}(\nu,1), and tνt_{\nu} is Student’s tt law with ν\nu degrees of freedom. For Gaussian sampling, 𝒫Σ=Nd(0,Σ)\mathcal{P}_{\Sigma}=N_{d}(0,\Sigma) and 𝒫Σ(n)=𝒫Σn\mathcal{P}_{\Sigma}^{(n)}=\mathcal{P}_{\Sigma}^{\otimes n}, with

Σ^=n1i=1nXiXi,σ^ij=eiΣ^ej.\hat{\Sigma}=n^{-1}\sum_{i=1}^{n}X_{i}X_{i}^{\top},\quad\hat{\sigma}_{ij}=e_{i}^{\top}\hat{\Sigma}e_{j}.

The Gaussian covariance map and the associated first-order variance are

Γ(ij),(kl)(Σ)=σikσjl+σilσjk,sF2(Σ)=F(Σ)Γ(Σ)F(Σ),\Gamma_{(ij),(kl)}(\Sigma)=\sigma_{ik}\sigma_{jl}+\sigma_{il}\sigma_{jk},\quad s_{F}^{2}(\Sigma)=\nabla F(\Sigma)^{\top}\Gamma(\Sigma)\nabla F(\Sigma),

and sFs_{F} is the nonnegative square root. Under a general law QQ for a centred vector XX, ΓQ(Σ)=covQ{vech(XX)}\Gamma_{Q}(\Sigma)=\operatorname{cov}_{Q}\{\operatorname{vech}(XX^{\top})\} whenever the fourth moments exist. In a triangular array with row covariance Σn\Sigma_{n}, the row law is 𝒫n=𝒫Σn(n)\mathcal{P}_{n}=\mathcal{P}_{\Sigma_{n}}^{(n)}, and we abbreviate 𝔼n=𝔼𝒫n\mathbb{E}_{n}=\mathbb{E}_{\mathcal{P}_{n}} and prn=pr𝒫n\mathrm{pr}_{n}=\mathrm{pr}_{\mathcal{P}_{n}}.

For an identification strategy (N,D)(N,D), the plug-in estimator is τ^=N(Σ^)/D(Σ^)\hat{\tau}=N(\hat{\Sigma})/D(\hat{\Sigma}) whenever D(Σ^)0D(\hat{\Sigma})\neq 0. Along a deterministic sequence Σn\Sigma_{n} with D(Σn)0D(\Sigma_{n})\neq 0, put

τn=N(Σn)D(Σn),gn=N(Σn)τnD(Σn),\tau_{n}=\frac{N(\Sigma_{n})}{D(\Sigma_{n})},\quad g_{n}=\nabla N(\Sigma_{n})-\tau_{n}\nabla D(\Sigma_{n}),

sD,n=sD(Σn)s_{D,n}=s_{D}(\Sigma_{n}) and sg,n={gnΓ(Σn)gn}1/2s_{g,n}=\{g_{n}^{\top}\Gamma(\Sigma_{n})g_{n}\}^{1/2}. Whenever sD,nsg,n>0s_{D,n}s_{g,n}>0,

μn=n1/2D(Σn)sD,n,rn=gnΓ(Σn)D(Σn)sg,nsD,n,ωn=sg,nsD,n.\mu_{n}=\frac{n^{1/2}D(\Sigma_{n})}{s_{D,n}},\quad r_{n}=\frac{g_{n}^{\top}\Gamma(\Sigma_{n})\nabla D(\Sigma_{n})}{s_{g,n}s_{D,n}},\quad\omega_{n}=\frac{s_{g,n}}{s_{D,n}}.

Equality in distribution is =d\stackrel{{\scriptstyle d}}{{=}}, convergence in probability is p\longrightarrow_{\rm p}, and convergence in distribution is \rightsquigarrow or \Rightarrow. The function sgn(x)\operatorname{sgn}(x) is the sign of xx\in\mathbb{R}. For positive deterministic ana_{n}, Yn=Op(an)Y_{n}=O_{p}(a_{n}) means that Yn/anY_{n}/a_{n} is bounded in probability, and Yn=op(an)Y_{n}=o_{p}(a_{n}) means Yn/anp0Y_{n}/a_{n}\longrightarrow_{\rm p}0. Unless stated otherwise, all asymptotic statements use the current sequence of laws and let nn\to\infty.

A.3 Paper-specific symbols

The remaining notation is summarized in Table A.1. An omitted matrix argument means evaluation at the current base point Σ\Sigma.

Table A.1: Paper-specific notation used in the Supplementary Material.
Symbol Meaning
Γ\Gamma, ΓQ\Gamma_{Q} Gaussian covariance map and the corresponding covariance map under a general law QQ
sF2s_{F}^{2}, sFs_{F} FΓF\nabla F^{\top}\Gamma\nabla F and its nonnegative square root
RFR_{F} Relative gradient GF(Σ)Σ/F(Σ)G_{F}(\Sigma)\Sigma/F(\Sigma) on {F0}\{F\neq 0\}
Σ^\hat{\Sigma}, σ^ij\hat{\sigma}_{ij} Known-mean sample covariance and its entries
F^D\widehat{F}_{D} nD(Σ^)2/sD2(Σ^)nD(\hat{\Sigma})^{2}/s_{D}^{2}(\hat{\Sigma}) on {sD2(Σ^)>0}\{s_{D}^{2}(\hat{\Sigma})>0\}
(a,b)(a,b), 𝝂\boldsymbol{\nu} Three-variable face type and primitive polynomial kernel vector
μn\mu_{n}, rnr_{n}, ωn\omega_{n} Local denominator mean, numerator–denominator correlation and scale ratio
Hat, subscript nn, subscript * Sample or plug-in quantity, row of a sequence, and limiting value

B Proofs of Section 2: Preliminaries

This section proves the facts about exact self-normalization that the main text invokes without proof. Lemma B.1 identifies the Gaussian first-order variance with the trace form in (3) and shows that the property is preserved under congruence transformations. Proposition B.1 and Remark B.1 then delimit its distributional scope under elliptical sampling and unknown means. Lemma B.2 reduces every converse question to one homogeneous degree, Lemma B.3 establishes interior nonvanishing, and Lemma B.4 shows that neither the property nor its constant depends on the ambient dimension.

The first lemma converts the statistical definition of sD2s_{D}^{2} into the intrinsic trace form used throughout the paper.

Lemma B.1 (Trace form and congruence invariance).

For every D[Symd]D\in\mathbb{R}[\operatorname{Sym}_{d}] and every Σ𝕊++d\Sigma\in\mathbb{S}^{d}_{++},

D(Σ)Γ(Σ)D(Σ)=2tr{(GD(Σ)Σ)2}.\nabla D(\Sigma)^{\top}\Gamma(\Sigma)\nabla D(\Sigma)=2\operatorname{tr}\{(G_{D}(\Sigma)\Sigma)^{2}\}.

If DD is exactly self-normalizing with constant cc and AA is nonsingular, then DA(Σ)=D(AΣA)D_{A}(\Sigma)=D(A\Sigma A^{\top}) is exactly self-normalizing with the same constant.

Both sides of the display are polynomial in Σ\Sigma, which yields the extension to all of Symd\operatorname{Sym}_{d} recorded after (3). The congruence part makes the property coordinate-free, and it is used silently whenever a flag or a kernel direction is normalized by a linear change of variables.

Proof of Lemma B.1.

Let XNd(0,Σ)X\sim N_{d}(0,\Sigma) and write G=GD(Σ)G=G_{D}(\Sigma). Under the symmetric perturbation convention, the coordinate gradient of DD has entries GiiG_{ii} on diagonal coordinates and 2Gij2G_{ij} on off-diagonal coordinates. Therefore,

D(Σ)Γ(Σ)D(Σ)=var(XGX)=2tr{(GΣ)2},\nabla D(\Sigma)^{\top}\Gamma(\Sigma)\nabla D(\Sigma)=\operatorname{var}(X^{\top}GX)=2\operatorname{tr}\{(G\Sigma)^{2}\},

by the Gaussian quadratic-form variance formula. If DA(Σ)=D(AΣA)D_{A}(\Sigma)=D(A\Sigma A^{\top}), the chain rule gives

GDA(Σ)=AGD(AΣA)A.G_{D_{A}}(\Sigma)=A^{\top}G_{D}(A\Sigma A^{\top})A.

Consequently, GDA(Σ)ΣG_{D_{A}}(\Sigma)\Sigma is similar to GD(AΣA)AΣAG_{D}(A\Sigma A^{\top})A\Sigma A^{\top}, so traces of powers agree. Since ΣAΣA\Sigma\mapsto A\Sigma A^{\top} is a bijection of Symd\operatorname{Sym}_{d}, the identity (4) for DD transfers to DAD_{A} with the same constant. ∎

We next determine how the sampling law enters. Throughout the following discussion, 𝒫\mathcal{P} is a probability law on d\mathbb{R}^{d} under which the vector X=(X(1),,X(d))X=(X^{(1)},\ldots,X^{(d)})^{\top} is centred with covariance matrix Σ\Sigma and finite fourth moments, and Γ𝒫(Σ)=cov𝒫{vech(XX)}\Gamma_{\mathcal{P}}(\Sigma)=\operatorname{cov}_{\mathcal{P}}\{\operatorname{vech}(XX^{\top})\} is the associated covariance operator. Directly from the definition of the fourth joint cumulant in Section A,

Γ𝒫,(ij),(kl)=σikσjl+σilσjk+cum𝒫(X(i),X(j),X(k),X(l)),\Gamma_{\mathcal{P},(ij),(kl)}=\sigma_{ik}\sigma_{jl}+\sigma_{il}\sigma_{jk}+\operatorname{cum}_{\mathcal{P}}(X^{(i)},X^{(j)},X^{(k)},X^{(l)}), (B.1)

so the departure of Γ𝒫\Gamma_{\mathcal{P}} from the Gaussian operator (2) consists exactly of the fourth cumulants. Suppose now that Σ0\Sigma\succ 0 and that 𝒫\mathcal{P} is elliptical, meaning that XX has the stochastic representation X=RAUX=RAU, where AA=ΣAA^{\top}=\Sigma, the vector UU is uniform on the unit sphere in d\mathbb{R}^{d}, and the radius R0R\geq 0 is independent of UU. The covariance normalization forces 𝔼𝒫(R2)=d\mathbb{E}_{\mathcal{P}}(R^{2})=d. Define the kurtosis parameter κ\kappa of 𝒫\mathcal{P} by

𝔼𝒫{(XΣ1X)2}=d(d+2)(1+κ),\mathbb{E}_{\mathcal{P}}\{(X^{\top}\Sigma^{-1}X)^{2}\}=d(d+2)(1+\kappa), (B.2)

so that κ=0\kappa=0 under Gaussian sampling, and Jensen’s inequality gives 1+κd/(d+2)>01+\kappa\geq d/(d+2)>0. Write Γκ=Γ𝒫\Gamma_{\kappa}=\Gamma_{\mathcal{P}} and sD,κ2(Σ)=D(Σ)Γκ(Σ)D(Σ)s_{D,\kappa}^{2}(\Sigma)=\nabla D(\Sigma)^{\top}\Gamma_{\kappa}(\Sigma)\nabla D(\Sigma).

Proposition B.1 (Elliptical fourth moments).

Let 𝒫\mathcal{P} be elliptical with covariance matrix Σ0\Sigma\succ 0 and kurtosis parameter κ\kappa, and let D[Symd]D\in\mathbb{R}[\operatorname{Sym}_{d}]. Then

Γκ,(ij),(kl)\displaystyle\Gamma_{\kappa,(ij),(kl)} =(1+κ)(σikσjl+σilσjk)+κσijσkl,\displaystyle=(1+\kappa)(\sigma_{ik}\sigma_{jl}+\sigma_{il}\sigma_{jk})+\kappa\sigma_{ij}\sigma_{kl}, (B.3)
sD,κ2(Σ)\displaystyle s_{D,\kappa}^{2}(\Sigma) =2(1+κ)tr{(GDΣ)2}+κ{tr(GDΣ)}2.\displaystyle=2(1+\kappa)\operatorname{tr}\{(G_{D}\Sigma)^{2}\}+\kappa\{\operatorname{tr}(G_{D}\Sigma)\}^{2}. (B.4)

If DD is homogeneous of degree KK, then DD satisfies (4) with constant cc if and only if sD,κ2=cκD2s_{D,\kappa}^{2}=c_{\kappa}D^{2} identically, where cκ=(1+κ)c+κK2c_{\kappa}=(1+\kappa)c+\kappa K^{2}, and in that case

nD(Σ^)2sD2(Σ^)=nc,nD(Σ^)2sD,κ2(Σ^)=ncκ\frac{nD(\hat{\Sigma})^{2}}{s_{D}^{2}(\hat{\Sigma})}=\frac{n}{c},\quad\frac{nD(\hat{\Sigma})^{2}}{s_{D,\kappa}^{2}(\hat{\Sigma})}=\frac{n}{c_{\kappa}} (B.5)

at every sample covariance matrix at which the displayed denominators are positive.

Both sides of (B.4) are polynomial in Σ\Sigma, so the proportionality in the proposition is an identity in the polynomial ring, exactly as in the Gaussian case. The proposition separates two conclusions. Ellipticity rescales the proportionality constant of a homogeneous polynomial but does not change the algebraic class, so the classification developed in the main text is unaffected. At the same time, only the second ratio in (B.5) carries the correct fourth-moment scaling when κ\kappa governs the sampling law, so samplewise algebraic constancy is distinct from correct variance calibration, which is the distinction drawn in the main text.

Proof of Proposition B.1.

All expectations are under 𝒫\mathcal{P}. Since Σ0\Sigma\succ 0, the matrix AA is nonsingular, so XΣ1X=R2X^{\top}\Sigma^{-1}X=R^{2} and (B.2) states that 𝔼𝒫(R4)=d(d+2)(1+κ)\mathbb{E}_{\mathcal{P}}(R^{4})=d(d+2)(1+\kappa). Put H=AGDAH=A^{\top}G_{D}A, so that tr(H)=tr(GDΣ)\operatorname{tr}(H)=\operatorname{tr}(G_{D}\Sigma) and tr(H2)=tr{(GDΣ)2}\operatorname{tr}(H^{2})=\operatorname{tr}\{(G_{D}\Sigma)^{2}\}. The spherical fourth-moment identity is

𝔼𝒫{(UHU)2}=2tr(H2)+(trH)2d(d+2),\mathbb{E}_{\mathcal{P}}\{(U^{\top}HU)^{2}\}=\frac{2\operatorname{tr}(H^{2})+(\operatorname{tr}H)^{2}}{d(d+2)},

and independence of RR and UU gives

𝔼𝒫{(XGDX)2}=(1+κ)[2tr{(GDΣ)2}+{tr(GDΣ)}2].\mathbb{E}_{\mathcal{P}}\{(X^{\top}G_{D}X)^{2}\}=(1+\kappa)[2\operatorname{tr}\{(G_{D}\Sigma)^{2}\}+\{\operatorname{tr}(G_{D}\Sigma)\}^{2}].

Since 𝔼𝒫(XGDX)=tr(GDΣ)\mathbb{E}_{\mathcal{P}}(X^{\top}G_{D}X)=\operatorname{tr}(G_{D}\Sigma), subtraction of the squared mean proves (B.4), and coefficient comparison over GDG_{D} yields (B.3).

If DD is homogeneous of degree KK, Euler’s identity gives tr(GDΣ)=KD(Σ)\operatorname{tr}(G_{D}\Sigma)=KD(\Sigma). Substituting (4) into (B.4) gives sD,κ2={(1+κ)c+κK2}D2s_{D,\kappa}^{2}=\{(1+\kappa)c+\kappa K^{2}\}D^{2}, which is the stated proportionality, and evaluating the two polynomial identities at Σ^\hat{\Sigma} gives (B.5). Conversely, if sD,κ2=cκD2s_{D,\kappa}^{2}=c_{\kappa}D^{2} identically, then 1+κ>01+\kappa>0 permits solving (B.4) for 2tr{(GDΣ)2}2\operatorname{tr}\{(G_{D}\Sigma)^{2}\}, and Euler’s identity recovers (4) with constant cc. ∎

Remark B.1 (Elliptical sampling and unknown means).

The product-of-chi-squares pivot in Theorem 1 is generally unavailable under elliptical sampling, because the sample covariance matrix is Wishart only when the radial law is Gaussian. Under 𝒫Σ(n)\mathcal{P}_{\Sigma}^{(n)} with an unknown mean, let X¯=n1i=1nXi\bar{X}=n^{-1}\sum_{i=1}^{n}X_{i} and

Σ~=(n1)1i=1n(XiX¯)(XiX¯).\tilde{\Sigma}=(n-1)^{-1}\sum_{i=1}^{n}(X_{i}-\bar{X})(X_{i}-\bar{X})^{\top}.

Then (n1)Σ~𝒲d(n1,Σ)(n-1)\tilde{\Sigma}\sim\mathcal{W}_{d}(n-1,\Sigma), every exact Gaussian pivot holds with nn replaced by n1n-1, and the standardized denominator of a flag power computed from Σ~\tilde{\Sigma} equals (n1)/c(n-1)/c. For non-Gaussian elliptical sampling with an estimated mean, the distinction in (B.5) persists. Algebraic constancy survives, correct calibration requires the kurtosis-adjusted operator, and a Wishart factorization is unavailable in general. In the front-door regressions of the main text fitted with intercepts, the residual degrees of freedom become n2n-2 and n3n-3.

Two structural reductions underlie every converse argument. The first confines every question about exact self-normalization to a single homogeneous degree, and the second supplies the interior positivity on which the logarithmic and geodesic arguments rely.

Lemma B.2 (Homogeneous reduction).

Every rational strategy for a scale-invariant coefficient in a linear structural equation model admits a representative whose numerator and denominator are homogeneous of the same degree. If a nonhomogeneous polynomial satisfies exact self-normalization, then its highest- and lowest-degree nonzero homogeneous parts satisfy the same identity with the same constant.

Lemma B.2 permits all converse arguments to be conducted within one homogeneous degree, which is the standing convention of Sections 3 and 4 of the main text.

Proof of Lemma B.2.

For a linear structural equation model, multiplying every error covariance by t>0t>0 multiplies Σ\Sigma by tt but leaves the scale-invariant causal coefficient under study unchanged. Expanding N(tΣ)=τD(tΣ)N(t\Sigma)=\tau D(t\Sigma) in powers of tt shows that each pair of homogeneous components of the same degree satisfies the identification identity, and at least one denominator component is nonzero on the model. For the second assertion, substitute tΣt\Sigma into (4) and compare the highest and lowest powers of tt. ∎

Lemma B.3 (Interior nonvanishing).

A nonzero polynomial satisfying exact self-normalization has no zero in 𝕊++d\mathbb{S}^{d}_{++}. It therefore has a constant sign on that cone.

Lemma B.3 makes the logarithm of a self-normalizing polynomial globally available on the positive cone, which the geodesic arguments of Supplementary Sections C and D use throughout. Its failure at the identity matrix is also what rejects the proximal denominator in the factor-and-rank example of the main text.

Proof of Lemma B.3.

Because DD is a nonzero polynomial and 𝕊++d\mathbb{S}^{d}_{++} is a nonempty Euclidean-open set, there is Σ00\Sigma_{0}\succ 0 with D(Σ0)0D(\Sigma_{0})\neq 0. Let 𝒰\mathcal{U} be the connected component of {Σ0:D(Σ)0}\{\Sigma\succ 0:D(\Sigma)\neq 0\} containing Σ0\Sigma_{0}, and define f=log|D|f=\log|D| only on 𝒰\mathcal{U}, so that no global nonvanishing conclusion is used at this stage. Equip 𝕊++d\mathbb{S}^{d}_{++} with the affine-invariant metric

gΣ(H1,H2)=tr(Σ1H1Σ1H2).g_{\Sigma}(H_{1},H_{2})=\operatorname{tr}(\Sigma^{-1}H_{1}\Sigma^{-1}H_{2}).

Its gradient on 𝒰\mathcal{U} is gradgf=Σ(GD/D)Σ\operatorname{grad}_{g}f=\Sigma(G_{D}/D)\Sigma, and (4) gives gradgfg2=α\lVert\operatorname{grad}_{g}f\rVert_{g}^{2}=\alpha with α=c/2\alpha=c/2.

Suppose that D(Σ1)=0D(\Sigma_{1})=0 for some Σ10\Sigma_{1}\succ 0. Let γ:[0,1]𝕊++d\gamma:[0,1]\to\mathbb{S}^{d}_{++} be the affine-invariant geodesic segment from Σ0\Sigma_{0} to Σ1\Sigma_{1}, which has finite length. Let t=inf{t(0,1]:D(γ(t))=0}t_{\ast}=\inf\{t\in(0,1]:D(\gamma(t))=0\}. Then γ([0,t))𝒰\gamma([0,t_{\ast}))\subset\mathcal{U}, and for t<tt<t_{\ast},

|(fγ)(t)|gradgfgγ˙(t)g=α1/2γ˙(t)g.|(f\circ\gamma)^{\prime}(t)|\leq\lVert\operatorname{grad}_{g}f\rVert_{g}\lVert\dot{\gamma}(t)\rVert_{g}=\alpha^{1/2}\lVert\dot{\gamma}(t)\rVert_{g}.

Integration over the finite-length segment shows that f(γ(t))f(\gamma(t)) is bounded below as ttt\uparrow t_{\ast}. Continuity of the polynomial DD and D(γ(t))=0D(\gamma(t_{\ast}))=0 instead imply f(γ(t))=log|D(γ(t))|f(\gamma(t))=\log|D(\gamma(t))|\to-\infty, a contradiction. Hence, DD has no zero in 𝕊++d\mathbb{S}^{d}_{++}. Since the cone is connected, its sign is constant. ∎

The final lemma shows that exact self-normalization is a property of the smallest coordinate block on which the polynomial depends.

Lemma B.4 (Invariance under ambient extension).

Let rdr\leq d, let D¯[Symr]\bar{D}\in\mathbb{R}[\operatorname{Sym}_{r}], and define D[Symd]D\in\mathbb{R}[\operatorname{Sym}_{d}] by D(Σ)=D¯(Σ1:r)D(\Sigma)=\bar{D}(\Sigma_{1:r}). Then sD2(Σ)=sD¯2(Σ1:r)s_{D}^{2}(\Sigma)=s_{\bar{D}}^{2}(\Sigma_{1:r}) for every ΣSymd\Sigma\in\operatorname{Sym}_{d}. Consequently, DD is exactly self-normalizing on Symd\operatorname{Sym}_{d} if and only if D¯\bar{D} is exactly self-normalizing on Symr\operatorname{Sym}_{r}, with the same constant.

Lemma B.4 justifies the dimension reductions invoked in Section 2 of the main text and in the simulation design. Combined with the congruence invariance of Lemma B.1, it applies to a polynomial depending on Σ\Sigma only through its compression to an arbitrary rr-dimensional subspace, because one nonsingular linear change of the observation vector moves that subspace to the leading coordinate block.

Proof of Lemma B.4.

Since D/σij=0\partial D/\partial\sigma_{ij}=0 whenever max(i,j)>r\max(i,j)>r, the symmetric gradient has the block form GD(Σ)=diag{GD¯(Σ1:r),0}G_{D}(\Sigma)=\operatorname{diag}\{G_{\bar{D}}(\Sigma_{1:r}),0\}. Writing Σ\Sigma in the corresponding blocks, GD(Σ)ΣG_{D}(\Sigma)\Sigma has zero second block row, and the trace of its square equals tr[{GD¯(Σ1:r)Σ1:r}2]\operatorname{tr}[\{G_{\bar{D}}(\Sigma_{1:r})\Sigma_{1:r}\}^{2}], which proves the identity. The equivalence follows because Σ1:r\Sigma_{1:r} ranges over all of Symr\operatorname{Sym}_{r} as Σ\Sigma ranges over Symd\operatorname{Sym}_{d}. ∎

C Proofs of Section 3: Flag powers

This section proves Theorem 1 and the two-subspace criterion stated after it in Section 3 of the main text. The theorem combines an upper-triangular calculation of the relative gradient with the Bartlett decomposition of the Wishart matrix. The final lemma shows that the nesting condition is also necessary for a product of two covariance-volume factors.

Proof of Theorem 1.

By congruence invariance of Lemma B.1, it suffices to treat the coordinate flag, and the constants arising from the chosen basis matrices cancel from the logarithmic gradient. Put Aj={1,,aj}A_{j}=\{1,\ldots,a_{j}\}. For Σ0\Sigma\succ 0 every leading block is invertible and

GΔAj(Σ)ΣΔAj(Σ)=(IajΣAj1ΣAj,Ajc00).\frac{G_{\Delta_{A_{j}}}(\Sigma)\Sigma}{\Delta_{A_{j}}(\Sigma)}=\begin{pmatrix}I_{a_{j}}&\Sigma_{A_{j}}^{-1}\Sigma_{A_{j},A_{j}^{c}}\\ 0&0\end{pmatrix}.

Weighting the jjth term by the exponent eje_{j} and summing over jj shows that RD(Σ)R_{D}(\Sigma) is upper triangular with diagonal (m1,,mak,0,,0)(m_{1},\ldots,m_{a_{k}},0,\ldots,0), which proves (i).

Therefore, the trace of the square of RD(Σ)R_{D}(\Sigma) is i=1akmi2\sum_{i=1}^{a_{k}}m_{i}^{2}, so (4) holds with c=2i=1akmi2c=2\sum_{i=1}^{a_{k}}m_{i}^{2} on the cone. Both sides of (4) are polynomial, so the identity extends to all of Symd\operatorname{Sym}_{d}, proving (ii).

For part (iii), write Σ=LL\Sigma=LL^{\top} with LL lower triangular and nΣ^=LWLn\hat{\Sigma}=LWL^{\top}, where W𝒲d(n,I)W\sim\mathcal{W}_{d}(n,I). Since LL is lower triangular, (LWL)1:r=L1:rW1:rL1:r(LWL^{\top})_{1:r}=L_{1:r}W_{1:r}L_{1:r}^{\top}, so leading determinants factor blockwise, and the Bartlett decomposition (30, Ch. 3) gives detW1:r=i=1rBii2\det W_{1:r}=\prod_{i=1}^{r}B_{ii}^{2} with independent Bii2χni+12B_{ii}^{2}\sim\chi^{2}_{n-i+1}. Collecting exponents gives (iii).

For part (iv), a flag power is positive at every positive-definite matrix, so D(Σ^)>0D(\hat{\Sigma})>0 and sD2(Σ^)=cD(Σ^)2>0s_{D}^{2}(\hat{\Sigma})=cD(\hat{\Sigma})^{2}>0 almost surely, and evaluating (4) at Σ^\hat{\Sigma} gives F^D=n/c\widehat{F}_{D}=n/c. The same identity gives sD(Σn)=c1/2|D(Σn)|s_{D}(\Sigma_{n})=c^{1/2}|D(\Sigma_{n})| for every nn. If n1/2D(Σn)=O(1)n^{1/2}D(\Sigma_{n})=O(1), then D(Σn)0D(\Sigma_{n})\to 0, so sD(Σn)0s_{D}(\Sigma_{n})\to 0, which is incompatible with sD(Σn)s>0s_{D}(\Sigma_{n})\to s>0. ∎

The nesting assumption in Theorem 1 cannot be dropped even for a product of two covariance-volume factors.

Lemma C.1 (Two-subspace criterion).

Let V,WdV,W\subseteq\mathbb{R}^{d} and let p,qp,q be positive integers. Then ΔVpΔWq\Delta_{V}^{p}\Delta_{W}^{q} is exactly self-normalizing if and only if VWV\subseteq W or WVW\subseteq V.

Proof of Lemma C.1.

The logarithmic relative gradients of ΔV\Delta_{V} and ΔW\Delta_{W} are the Σ\Sigma-orthogonal projections PVΣP_{V}^{\Sigma} and PWΣP_{W}^{\Sigma} onto the two subspaces. Therefore,

tr{(pPVΣ+qPWΣ)2}=p2dimV+q2dimW+2pqtr(PVΣPWΣ).\operatorname{tr}\{(pP_{V}^{\Sigma}+qP_{W}^{\Sigma})^{2}\}=p^{2}\dim V+q^{2}\dim W+2pq\operatorname{tr}(P_{V}^{\Sigma}P_{W}^{\Sigma}).

The final trace is the sum of the squared cosines of the principal angles in the inner product xΣyx^{\top}\Sigma y. Under either containment, it equals the dimension of the smaller subspace and is constant.

Suppose that neither containment holds. Write S=VWS=V\cap W, V=SVV=S\oplus V^{\prime}, and W=SWW=S\oplus W^{\prime}, where VV^{\prime} and WW^{\prime} are nonzero. The sum SVWS\oplus V^{\prime}\oplus W^{\prime} is direct. Choose a positive-definite inner product that makes these three subspaces mutually orthogonal. The final trace then equals dimS\dim S. Perturb the inner product so that one nonzero vector in VV^{\prime} has nonzero inner product with one in WW^{\prime}, while retaining positive definiteness. At least one squared principal-angle cosine becomes positive, so the trace changes. Exact self-normalization is therefore impossible. ∎

D Proofs of Section 4: Converse results and diagnosis

D.1 Overview and algebraic preliminaries

This subsection proves the converse and diagnostic results in Section 4 of the main text. The argument first reduces the problem to boundary faces and determinant-free residues. It then proves the complete converse in dimension two and the three-variable reduction. The remaining subsections study the mixed branch, the conditional higher-dimensional converse, and the low-degree-factor diagnostic.

The first two lemmas provide the divisibility tools used throughout the section.

Lemma D.1 (Irreducibility of the generic symmetric determinant).

For every r1r\geq 1, the determinant of the generic symmetric matrix of order rr is irreducible over \mathbb{R} and over \mathbb{C}. Consequently, each leading principal minor Δr=det(Σ1:r)\Delta_{r}=\det(\Sigma_{1:r}) is irreducible in [Symd]\mathbb{R}[\operatorname{Sym}_{d}] for drd\geq r.

Irreducible elements are prime in the unique factorization domain [Symd]\mathbb{R}[\operatorname{Sym}_{d}]. Therefore, Lemma D.1 justifies the determinant-divisibility and valuation arguments below.

Proof of Lemma D.1.

We argue over a field kk of characteristic zero, which covers k=k=\mathbb{R} and k=k=\mathbb{C}. The assertion is immediate for r=1r=1. Suppose it holds for r1r-1, and write the generic symmetric matrix as

Xr=(Ayyz),X_{r}=\begin{pmatrix}A&y\\ y^{\top}&z\end{pmatrix},

where AA is generic symmetric of order r1r-1. Expansion in zz gives

detXr=(detA)zyadj(A)y.\det X_{r}=(\det A)z-y^{\top}\operatorname{adj}(A)y.

Let R=k[A,y]R=k[A,y], which is a unique factorization domain. By induction, detA\det A is irreducible in RR. If detA\det A divided yadj(A)yy^{\top}\operatorname{adj}(A)y, it would divide every coefficient of this polynomial in the coordinates of yy. In particular, it would divide the coefficient of yr12y_{r-1}^{2}, which is a nonzero principal minor of AA of order r2r-2 and has smaller degree. Thus, the two coefficients of the displayed polynomial in zz are coprime. The polynomial is primitive in R[z]R[z] and irreducible over the fraction field of RR. Gauss’ lemma gives irreducibility in R[z]R[z]. Passing to a larger polynomial ring preserves irreducibility, which proves the final assertion. ∎

The next lemma converts vanishing on a real part of the determinant boundary into polynomial divisibility. This step is needed when two candidate solutions agree only on rank-deficient covariance matrices.

Lemma D.2 (Real determinant divisibility).

If a polynomial FF on Symd\operatorname{Sym}_{d} vanishes on a nonempty Euclidean-open subset of the real symmetric matrices of rank d1d-1, then detΣ\det\Sigma divides FF.

Proof of Lemma D.2.

After a simultaneous permutation of rows and columns, work in a chart in which the leading principal minor Δ\Delta of order d1d-1 is nonzero at one point of the given open set. Shrink the set so that Δ0\Delta\neq 0 throughout it. The determinant is linear in σdd\sigma_{dd}, and can be written as

δdetΣ=Δσdd+B,\delta\coloneqq\det\Sigma=\Delta\sigma_{dd}+B,

where BB does not involve σdd\sigma_{dd}. If N=degσddFN=\deg_{\sigma_{dd}}F, pseudo-division in σdd\sigma_{dd} gives

ΔNF=Qδ+R,degσddR=0.\Delta^{N}F=Q\delta+R,\quad\deg_{\sigma_{dd}}R=0.

On this chart, the equation δ=0\delta=0 is the graph σdd=B/Δ\sigma_{dd}=-B/\Delta. The projection of the given Euclidean-open subset of this graph to the remaining coordinates is Euclidean open. Since RR does not involve σdd\sigma_{dd} and vanishes on that projection, R0R\equiv 0. Thus, δ\delta divides ΔNF\Delta^{N}F. By Lemma D.1, δ\delta is prime in [Symd]\mathbb{R}[\operatorname{Sym}_{d}]. It does not divide Δ\Delta, because Δ\Delta does not involve σdd\sigma_{dd} whereas δ\delta does. Hence, δ\delta divides FF. ∎

The face-restriction principle transfers exact self-normalization to every nonzero covariance face. Together with Theorem 2, it determines the form of each nonzero plane restriction in dimension three.

Lemma D.3 (Face restriction).

If DD satisfies (4) on Symd\operatorname{Sym}_{d} and its restriction DVD_{V} to a face 𝕊+(V)\mathbb{S}_{+}(V) is not identically zero, then DVD_{V} satisfies (4) with the same constant.

Proof of Lemma D.3.

By Lemma B.1, a congruence reduces the argument to V=span(e1,,er)V=\operatorname{span}(e_{1},\ldots,e_{r}). Evaluate (4) at diag(S,0)\operatorname{diag}(S,0) for SSymrS\in\operatorname{Sym}_{r}. In the corresponding block partition of the gradient,

GDdiag(S,0)=(G11S0G12S0),G_{D}\operatorname{diag}(S,0)=\begin{pmatrix}G_{11}S&0\\ G_{12}^{\top}S&0\end{pmatrix},

and the trace of its square is tr{(G11S)2}\operatorname{tr}\{(G_{11}S)^{2}\}. The block G11G_{11}, evaluated at diag(S,0)\operatorname{diag}(S,0), is the symmetric gradient of DVD_{V} at SS. This proves the identity for DVD_{V}. ∎

The stripping lemma removes one full determinant factor and records the resulting change in the self-normalization constant. It is the degree-lowering step in the dimension-two proof and in the three-variable reduction.

Lemma D.4 (Determinant stripping).

If D=(detΣ)ED=(\det\Sigma)E and EE is homogeneous, then DD is exactly self-normalizing if and only if EE is. Their constants satisfy

cD=cE+2d+4degE.c_{D}=c_{E}+2d+4\deg E.
Proof of Lemma D.4.

On the positive cone, write LD=GD/DL_{D}=G_{D}/D and LE=GE/EL_{E}=G_{E}/E. The product rule gives LD=Σ1+LEL_{D}=\Sigma^{-1}+L_{E}. Therefore,

tr{(LDΣ)2}=d+2tr(LEΣ)+tr{(LEΣ)2}=d+2degE+tr{(LEΣ)2}.\operatorname{tr}\{(L_{D}\Sigma)^{2}\}=d+2\operatorname{tr}(L_{E}\Sigma)+\operatorname{tr}\{(L_{E}\Sigma)^{2}\}=d+2\deg E+\operatorname{tr}\{(L_{E}\Sigma)^{2}\}.

Exact self-normalization is equivalent on the cone to constancy of the final trace. Hence, it holds for DD if and only if it holds for EE. The displayed identity gives cD/2=d+2degE+cE/2c_{D}/2=d+2\deg E+c_{E}/2. Both self-normalization identities then extend polynomially to all of Symd\operatorname{Sym}_{d}. ∎

D.2 Complete converse in dimension two

The proof separates solutions according to whether the determinant of the symmetric gradient vanishes identically. A nonzero determinant coefficient produces a removable determinant factor. The zero coefficient leads to a rank-one gradient and then to a power of one variance direction.

Proof of Theorem 2.

Let K=degDK=\deg D and let μ1,μ2\mu_{1},\mu_{2} be the eigenvalues of GD(Σ)ΣG_{D}(\Sigma)\Sigma. On the positive cone they are real because this matrix is similar to Σ1/2GD(Σ)Σ1/2\Sigma^{1/2}G_{D}(\Sigma)\Sigma^{1/2}. Euler’s identity and (4) give

μ1+μ2=KD,μ12+μ22=(c/2)D2.\mu_{1}+\mu_{2}=KD,\quad\mu_{1}^{2}+\mu_{2}^{2}=(c/2)D^{2}.

Hence, as a polynomial identity,

detGD(Σ)detΣ=qD(Σ)2,q=12(K2c/2).\det G_{D}(\Sigma)\det\Sigma=qD(\Sigma)^{2},\quad q=\frac{1}{2}(K^{2}-c/2). (D.1)

If q0q\neq 0, Lemma D.1 gives detΣ|D\det\Sigma\mid D. Lemma D.4 removes this factor, and induction on KK applies.

Suppose that q=0q=0. Then detGD0\det G_{D}\equiv 0. The polynomial gradient map takes values in the rank-one quadric cone in Sym23\operatorname{Sym}_{2}\cong\mathbb{R}^{3}. At every nonzero smooth image point, the differential of the gradient, which is the Hessian of DD, maps into the two-dimensional tangent plane. Thus, detHessD0\det\operatorname{Hess}D\equiv 0. The Gordan–Noether theorem for forms in at most four variables implies over \mathbb{C} that DD is a cone. The reduction may be taken over \mathbb{R}. A complex constant direction annihilating DD brings its conjugate with it. If the two directions are proportional, rescaling gives a real direction. If they are independent, DD depends on one linear form, which reality makes real up to a scalar. Thus, after a real linear change, D=P(1,2)D=P(\ell_{1},\ell_{2}).

Write GD=P1A1+P2A2G_{D}=P_{1}A_{1}+P_{2}A_{2} for fixed independent matrices A1,A2Sym2A_{1},A_{2}\in\operatorname{Sym}_{2}. The binary quadratic Q(x,y)=det(xA1+yA2)Q(x,y)=\det(xA_{1}+yA_{2}) is nonzero because the determinant cone contains no two-dimensional linear subspace. The identity Q(P1,P2)=0Q(P_{1},P_{2})=0 either forces P1=P2=0P_{1}=P_{2}=0 when QQ is definite, or forces one real linear factor of QQ to vanish identically. Thus, PP depends on one linear form and

D=γtr(AΣ)K,detA=0.D=\gamma\operatorname{tr}(A\Sigma)^{K},\quad\det A=0.

A nonzero real symmetric rank-one matrix AA is a nonzero scalar multiple of vvvv^{\top}. This gives the form in (8). The value of cc follows from Theorem 1. ∎

The two branches of the proof correspond exactly to the determinant factor and the rank-one variance factor in Theorem 2. No other irreducible factor can occur in dimension two.

D.3 Adjugate duality and three-variable face types

We next turn to determinant-free solutions on Sym3\operatorname{Sym}_{3}. Adjugate duality exchanges rank-one and rank-two boundary data. The face-type lemma then shows that every nonzero plane restriction has one common pair of multiplicities.

Lemma D.5 (Adjugate duality).

If DD is homogeneous of degree KK and satisfies (4) on Symd\operatorname{Sym}_{d} with constant cc, then D=DadjD^{\dagger}=D\circ\operatorname{adj} is homogeneous of degree (d1)K(d-1)K and satisfies (4) with

c=c+2(d2)K2.c^{\dagger}=c+2(d-2)K^{2}.

Moreover, (D)=(detΣ)(d2)KD(D^{\dagger})^{\dagger}=(\det\Sigma)^{(d-2)K}D.

Proof of Lemma D.5.

On the positive cone, adjΣ=(detΣ)Σ1\operatorname{adj}\Sigma=(\det\Sigma)\Sigma^{-1}, so

log|D(Σ)|=KlogdetΣ+log|D(Σ1)|.\log|D^{\dagger}(\Sigma)|=K\log\det\Sigma+\log|D(\Sigma^{-1})|.

If LD=GD/DL_{D}=G_{D}/D, differentiation gives

LD(Σ)Σ=KIΣ1LD(Σ1).L_{D^{\dagger}}(\Sigma)\Sigma=KI-\Sigma^{-1}L_{D}(\Sigma^{-1}).

The second term is similar to the relative gradient of DD at Σ1\Sigma^{-1}. If its eigenvalues are λ1,,λd\lambda_{1},\ldots,\lambda_{d}, Euler’s identity and (4) give iλi=K\sum_{i}\lambda_{i}=K and iλi2=c/2\sum_{i}\lambda_{i}^{2}=c/2. Hence,

tr{(LDΣ)2}=(d2)K2+c/2.\operatorname{tr}\{(L_{D^{\dagger}}\Sigma)^{2}\}=(d-2)K^{2}+c/2.

Clearing denominators proves the polynomial identity for DD^{\dagger}. The involution formula follows from adj(adjΣ)=(detΣ)d2Σ\operatorname{adj}(\operatorname{adj}\Sigma)=(\det\Sigma)^{d-2}\Sigma. ∎

For the next lemma, write K=degDK=\deg D. The result makes the face type in Definition 3 well defined and gives the bounds used in the three-variable reduction.

Lemma D.6 (Face type in dimension three).

Let DD be determinant-free, homogeneous and exactly self-normalizing on Sym3\operatorname{Sym}_{3}. At least one DD-good plane exists, and every DD-good restriction has the form (9) with one common pair (a,b)(a,b). Moreover, K2c2K2K^{2}\leq c\leq 2K^{2}, with c=2K2c=2K^{2} if and only if b=0b=0.

Proof of Lemma D.6.

If every plane restriction vanished, DD would vanish on every singular positive-semidefinite matrix. Lemma D.2 would then give δ|D\delta\mid D, contrary to determinant-freeness. A DD-good face restriction satisfies (4) by Lemma D.3. Theorem 2 gives its form. Substituting a=K2ba=K-2b into the equation for the self-normalization constant gives

2b22Kb+(K2c/2)=0.2b^{2}-2Kb+(K^{2}-c/2)=0.

Its two roots b+b_{+} and bb_{-} satisfy b++b=Kb_{+}+b_{-}=K. Admissibility requires 0b±K/20\leq b_{\pm}\leq K/2. If both roots were admissible, their sum could equal KK only when b+=b=K/2b_{+}=b_{-}=K/2. Hence, there is at most one admissible root, and the face type is common to all good planes. The bounds on cc follow from the same two equations. ∎

The adjugate maps a rank-two covariance face to the rank-one cone. The following formula makes that action explicit.

Lemma D.7 (Boundary strata under the adjugate).

If P3×2P\in\mathbb{R}^{3\times 2} and m(P)m(P) is the cross product of its columns, then, for every SSym2S\in\operatorname{Sym}_{2},

adj(PSP)=(detS)m(P)m(P).\operatorname{adj}(PSP^{\top})=(\det S)m(P)m(P)^{\top}.
Proof of Lemma D.7.

The (i,j)(i,j) cofactor of PSPPSP^{\top} is the product of the corresponding 2×22\times 2 row minors of PP and detS\det S. Those signed minors are the coordinates of m(P)m(P). ∎

The next dichotomy separates solutions that survive on the rank-one cone from those that vanish there. The surviving branch must have the largest possible self-normalization constant and a pure rank-one face type.

Lemma D.8 (Rank-one dichotomy).

Let DD be homogeneous of degree KK and exactly self-normalizing on Sym3\operatorname{Sym}_{3}. Either pD0p_{D}\equiv 0, or DD is determinant-free, c=2K2c=2K^{2}, and its face type is (K,0)(K,0).

Proof of Lemma D.8.

If pD0p_{D}\not\equiv 0, then δD\delta\nmid D. For a plane represented by PP, Lemma D.7 gives

(D)V(S)=(detS)KpD{m(P)}.(D^{\dagger})_{V}(S)=(\det S)^{K}p_{D}\{m(P)\}.

On every plane for which pD{m(P)}0p_{D}\{m(P)\}\neq 0, this is a pure determinant power of type (0,K)(0,K). Its constant satisfies c/2=2K2c^{\dagger}/2=2K^{2}. Lemma D.5 gives c/2=K2+c/2c^{\dagger}/2=K^{2}+c/2, and hence c=2K2c=2K^{2}. Lemma D.6 then forces b=0b=0. ∎

The dichotomy leaves a rigidity question on the rank-one cone. The next proposition closes that branch and shows that it contains only powers of one variance direction.

Proposition D.1 (Rank-one rigidity).

Let DD be homogeneous of degree KK and exactly self-normalizing on Sym3\operatorname{Sym}_{3}. If pD0p_{D}\not\equiv 0, then D=γ(vΣv)KD=\gamma(v^{\top}\Sigma v)^{K} for a nonzero constant γ\gamma and a nonzero vector v3v\in\mathbb{R}^{3}.

Proof of Proposition D.1.

By Lemma D.8, every DD-good plane restriction is a KKth power of one linear covariance form. Thus, on a Zariski-open family of projective lines in 2\mathbb{P}^{2}_{\mathbb{C}}, the form pDp_{D} restricts to a polynomial whose zero set has one support point. If the reduced projective curve {pD=0}red\{p_{D}=0\}_{\mathrm{red}} had degree at least two, a general line would meet it transversally in at least two points by Bezout’s theorem (20). Hence, the reduced curve is one line and

pD(u)=γ(vu)2Kp_{D}(u)=\gamma(v^{\top}u)^{2K}

with vv and γ\gamma real after rescaling.

Let D0=γ(vΣv)KD_{0}=\gamma(v^{\top}\Sigma v)^{K}. On every DD-good plane, both face restrictions are KKth powers of linear covariance forms and agree on all rank-one matrices in that face. Unique factorization of the resulting binary powers makes the linear forms proportional with the matching constant. Hence, the two face polynomials agree identically. Therefore, DD and D0D_{0} agree on a nonempty open set of rank-two matrices, and Lemma D.2 gives δ|DD0\delta\mid D-D_{0}.

If the difference is nonzero, write D=D0+δjRD=D_{0}+\delta^{j}R with j1j\geq 1, δR\delta\nmid R, and degR=K3j\deg R=K-3j. After a congruence, take v=e1v=e_{1} and put Q=σ11Q=\sigma_{11}. Expanding (4) and dividing by δj\delta^{j} gives

(Σe1)GR(Σ)(Σe1)=(Kj)QR+δjS(\Sigma e_{1})^{\top}G_{R}(\Sigma)(\Sigma e_{1})=(K-j)QR+\delta^{j}S (D.2)

for a polynomial SS.

Use the Schur coordinates

Σ(q,u,C)=(qququC+quu),δ=qdetC.\Sigma(q,u,C)=\begin{pmatrix}q&qu^{\top}\\ qu&C+quu^{\top}\end{pmatrix},\quad\delta=q\det C.

The line Σ+t(Σe1)(Σe1)\Sigma+t(\Sigma e_{1})(\Sigma e_{1})^{\top} changes only qq to q+tq2q+tq^{2}. Therefore, the left-hand side of (D.2) is q2qR¯q^{2}\partial_{q}\bar{R}. On detC=0\det C=0,

q2qR¯=(Kj)qR¯.q^{2}\partial_{q}\bar{R}=(K-j)q\bar{R}.

The polynomial R¯\bar{R} has qq-degree at most K3j<KjK-3j<K-j. Coefficient comparison forces R¯=0\bar{R}=0 on detC=0\det C=0. Thus, RR vanishes on a nonempty open set of rank-two matrices and is divisible by δ\delta, which contradicts its definition. ∎

Proposition D.1 proves Theorem 3 (i) after determinant stripping. The remaining determinant-free solutions vanish on the rank-one cone.

D.4 Relative spectral constancy in dimension three

Euler’s identity fixes the trace of the relative gradient, while exact self-normalization fixes its squared Frobenius norm. The next proposition uses the polynomial character of DD to obtain the additional discrete constraint that makes the spectrum constant in dimension three.

Proposition D.2 (Spectral constancy).

Let DD be homogeneous of degree KK and exactly self-normalizing on Sym3\operatorname{Sym}_{3}, and put α=c/2\alpha=c/2. There is a fixed multiset {λ1,λ2,λ3}\{\lambda_{1},\lambda_{2},\lambda_{3}\} such that

specRD(Σ)={λ1,λ2,λ3}(Σ0).\operatorname{spec}R_{D}(\Sigma)=\{\lambda_{1},\lambda_{2},\lambda_{3}\}\quad(\Sigma\succ 0).

In every case,

detGD(Σ)detΣ=λ1λ2λ3D(Σ)3.\det G_{D}(\Sigma)\det\Sigma=\lambda_{1}\lambda_{2}\lambda_{3}D(\Sigma)^{3}.

If α=K2/3\alpha=K^{2}/3, then 3|K3\mid K and D=γ(detΣ)K/3D=\gamma(\det\Sigma)^{K/3}.

Proof of Proposition D.2.

By Lemma B.3, DD has constant sign on 𝕊++3\mathbb{S}^{3}_{++}. Multiply it by 1-1 if necessary so that D>0D>0 there. Fix Σ00\Sigma_{0}\succ 0 and define

D~(X)=D(Σ01/2XΣ01/2).\tilde{D}(X)=D(\Sigma_{0}^{1/2}X\Sigma_{0}^{1/2}).

By Lemma B.1, D~\tilde{D} satisfies (4) with the same constant. At II, put

B=GD~(I)D~(I)=Σ01/2GD(Σ0)Σ01/2D(Σ0).B=\frac{G_{\tilde{D}}(I)}{\tilde{D}(I)}=\frac{\Sigma_{0}^{1/2}G_{D}(\Sigma_{0})\Sigma_{0}^{1/2}}{D(\Sigma_{0})}.

The matrix BB is symmetric and is similar to RD(Σ0)R_{D}(\Sigma_{0}). In particular, trB=K\operatorname{tr}B=K and trB2=α\operatorname{tr}B^{2}=\alpha.

By Cauchy–Schwarz, αK2/3\alpha\geq K^{2}/3. If α=K2/3\alpha=K^{2}/3, equality forces all three eigenvalues of RD(Σ0)R_{D}(\Sigma_{0}) to equal K/3K/3. Since Σ0\Sigma_{0} was arbitrary and RD(Σ0)R_{D}(\Sigma_{0}) is diagonalizable, RD=(K/3)IR_{D}=(K/3)I throughout the cone. Hence, dlog|D|=(K/3)dlogdetΣ\mathrm{d}\log|D|=(K/3)\mathrm{d}\log\det\Sigma. On a nonempty open set this gives D3=γ0(detΣ)KD^{3}=\gamma_{0}(\det\Sigma)^{K}. Polynomial continuation and unique factorization using Lemma D.1 yield 3|K3\mid K and D=γ(detΣ)K/3D=\gamma(\det\Sigma)^{K/3}. The determinant identity follows with λ1=λ2=λ3=K/3\lambda_{1}=\lambda_{2}=\lambda_{3}=K/3. Henceforth, assume α>K2/3\alpha>K^{2}/3.

Let f=logD~f=\log\tilde{D} and equip 𝕊++3\mathbb{S}^{3}_{++} with the affine-invariant metric. Equation (4) gives gradgfg2=α\lVert\operatorname{grad}_{g}f\rVert_{g}^{2}=\alpha. For every vector field YY, symmetry of the Riemannian Hessian gives

g(gradgfggradgf,Y)=Hessgf(gradgf,Y)=Hessgf(Y,gradgf)=12Ygradgfg2=0.\begin{split}g(\nabla^{g}_{\operatorname{grad}_{g}f}\operatorname{grad}_{g}f,Y)&=\operatorname{Hess}_{g}f(\operatorname{grad}_{g}f,Y)\\ &=\operatorname{Hess}_{g}f(Y,\operatorname{grad}_{g}f)=\frac{1}{2}Y\lVert\operatorname{grad}_{g}f\rVert_{g}^{2}=0.\end{split}

Thus, the integral curves of gradgf\operatorname{grad}_{g}f are geodesics. The geodesic through II with initial velocity BB is tetBt\mapsto e^{tB}. Uniqueness of geodesics and integral curves makes the two curves agree near t=0t=0. Along that interval,

ddtf(etB)=gradgfg2=α.\frac{\mathrm{d}}{\mathrm{d}t}f(e^{tB})=\lVert\operatorname{grad}_{g}f\rVert_{g}^{2}=\alpha.

Therefore,

D~(etB)=D~(I)eαt.\tilde{D}(e^{tB})=\tilde{D}(I)e^{\alpha t}. (D.3)

Both sides are entire functions of tt, so (D.3) holds for every tt\in\mathbb{R}.

Choose an orthogonal matrix OO such that B=Odiag(λ1,λ2,λ3)OB=O\operatorname{diag}(\lambda_{1},\lambda_{2},\lambda_{3})O^{\top}. Define

Q(x1,x2,x3)=D~{Odiag(x1,x2,x3)O}=n03|n|=Kanx1n1x2n2x3n3.Q(x_{1},x_{2},x_{3})=\tilde{D}\{O\operatorname{diag}(x_{1},x_{2},x_{3})O^{\top}\}=\sum_{\begin{subarray}{c}n\in\mathbb{Z}_{\geq 0}^{3}\\ |n|=K\end{subarray}}a_{n}x_{1}^{n_{1}}x_{2}^{n_{2}}x_{3}^{n_{3}}.

Since Q(1,1,1)=D~(I)0Q(1,1,1)=\tilde{D}(I)\neq 0, at least one coefficient group is nonzero. Substitution in (D.3) gives

|n|=Kane(nλ)t=D~(I)eαt.\sum_{|n|=K}a_{n}e^{(n\cdot\lambda)t}=\tilde{D}(I)e^{\alpha t}.

Linear independence of real exponentials implies nλ=αn\cdot\lambda=\alpha for at least one n03n\in\mathbb{Z}_{\geq 0}^{3} with |n|=K|n|=K.

The equations iλi=K\sum_{i}\lambda_{i}=K and iλi2=α\sum_{i}\lambda_{i}^{2}=\alpha define a circle in the trace plane. If nn is proportional to (1,1,1)(1,1,1), then n=(K/3)(1,1,1)n=(K/3)(1,1,1) and nλ=K2/3αn\cdot\lambda=K^{2}/3\neq\alpha. Thus, this exceptional resonance contains no admissible point. For every other nn, the equation nλ=αn\cdot\lambda=\alpha cuts the trace plane in a proper affine line and meets the circle in at most two points. There are only finitely many such nn. Hence, the ordered spectrum belongs to a finite set. The ordered eigenvalues vary continuously with Σ\Sigma because RD(Σ)R_{D}(\Sigma) is similar to the symmetric matrix

Σ1/2GD(Σ)Σ1/2D(Σ).\frac{\Sigma^{1/2}G_{D}(\Sigma)\Sigma^{1/2}}{D(\Sigma)}.

Since 𝕊++3\mathbb{S}^{3}_{++} is connected, the ordered spectrum is constant. Taking the product of the fixed eigenvalues and clearing denominators gives the determinant identity. ∎

For a determinant-free solution, the determinant identity forces one relative eigenvalue to vanish. The face type determines the other two.

Corollary D.1.

If δD\delta\nmid D, then one relative eigenvalue is zero and detGD0\det G_{D}\equiv 0. If DD has face type (a,b)(a,b), then

specRD(Σ)={a+b,b,0}(Σ0).\operatorname{spec}R_{D}(\Sigma)=\{a+b,b,0\}\quad(\Sigma\succ 0).
Proof of Corollary D.1.

The left-hand side of the determinant identity in Proposition D.2 is divisible by the prime δ\delta. If the eigenvalue product were nonzero, then δ|D3\delta\mid D^{3}, contrary to the hypothesis. The remaining two eigenvalues have sum K=a+2bK=a+2b and product b(a+b)b(a+b). They are therefore a+ba+b and bb. ∎

Spectral constancy supplies the singular-gradient structure of every determinant-free three-variable solution. It is the step that reduces the unresolved case to a polynomial kernel line.

D.5 Pure plane-minor and mixed branches

We first identify the branch whose nonzero plane restrictions are pure determinant powers. This proves Theorem 3 (ii).

Proof of the pure plane-minor branch in Theorem 3.

Suppose that the face type is (0,b)(0,b), so K=2bK=2b and the spectrum is (b,b,0)(b,b,0). The adjugate transform has spectrum (b,b,2b)(b,b,2b). Since pD0p_{D}\equiv 0, the polynomial DD^{\dagger} vanishes on every singular positive-semidefinite matrix. Lemma D.2 gives δ|D\delta\mid D^{\dagger}. Write D=δsED^{\dagger}=\delta^{s}E with exact s1s\geq 1. The adjugate involution gives 2sK2s\leq K, and determinant stripping shifts every relative eigenvalue down by ss. Hence,

specRE={bs,bs,2bs}.\operatorname{spec}R_{E}=\{b-s,b-s,2b-s\}.

Because δE\delta\nmid E, Corollary D.1 forces a zero eigenvalue. The bound sbs\leq b leaves only s=bs=b. Thus, degE=b\deg E=b and specRE=(b,0,0)\operatorname{spec}R_{E}=(b,0,0). The pair (b,0)(b,0) satisfies the equations that determine the face type of EE. Lemma D.6 therefore gives face type (b,0)(b,0). A nonzero restriction of this type is nonzero at a suitable rank-one matrix, so pE0p_{E}\not\equiv 0. Proposition D.1 gives E=γ(uΣu)bE=\gamma(u^{\top}\Sigma u)^{b}. Applying the adjugate again yields D=γΔubD=\gamma^{\prime}\Delta_{u^{\perp}}^{b}. ∎

Mixed face types are paired by adjugate duality. This symmetry is used in both the kernel analysis and the boundary-rigidity argument.

Lemma D.9 (Mixed duality).

If DD is determinant-free of mixed type (a,b)(a,b), then E=D/δbE=D^{\dagger}/\delta^{b} is determinant-free of type (b,a)(b,a) and E=δaDE^{\dagger}=\delta^{a}D.

Proof of Lemma D.9.

The spectrum of DD is (a+b,b,0)(a+b,b,0), so that of DD^{\dagger} is (a+2b,a+b,b)(a+2b,a+b,b). As in the pure branch, δ|D\delta\mid D^{\dagger}. Let ss be its exact multiplicity. After stripping, the spectrum is (a+2bs,a+bs,bs)(a+2b-s,a+b-s,b-s). The adjugate involution gives 2sK=a+2b2s\leq K=a+2b. Corollary D.1 requires a zero eigenvalue, so s{b,a+b,a+2b}s\in\{b,a+b,a+2b\}. Since a1a\geq 1, a+b>(a+2b)/2=K/2a+b>(a+2b)/2=K/2. The involution bound therefore excludes s=a+bs=a+b and s=a+2bs=a+2b. Hence, s=bs=b. The stripped spectrum is (a+b,a,0)(a+b,a,0) and the degree is 2a+b2a+b. The pair (b,a)(b,a) satisfies the two equations that determine the face type. Lemma D.6 gives the stated type, and the involution identity gives the final formula. ∎

The next algebraic lemma turns a rank-one factorization over the fraction field into a polynomial factorization. It is applied to the adjugate of the singular symmetric gradient.

Lemma D.10 (Primitive rank-one factorization over a UFD).

Let RR be a unique factorization domain and let UR3×3U\in R^{3\times 3} be a nonzero matrix whose 2×22\times 2 minors all vanish. Then U=vrU=vr^{\top} for vectors v,rR3v,r\in R^{3}, where vv is primitive. If UU is symmetric, then U=ψvvU=\psi vv^{\top} for some ψR\psi\in R. If the entries of UU are homogeneous of one common degree, the factors may be chosen homogeneous.

Proof of Lemma D.10.

Choose a nonzero column Uj0U_{\cdot j_{0}}. Let gg be a greatest common divisor of its entries and put v=Uj0/gv=U_{\cdot j_{0}}/g. Then vv is primitive. Since all 2×22\times 2 minors vanish, every other column is proportional to vv over K=Frac(R)K=\operatorname{Frac}(R). For each jj, there is rjKr_{j}\in K such that Uj=rjvU_{\cdot j}=r_{j}v. Write rj=p/qr_{j}=p/q in lowest terms. The identity qUij=pviqU_{ij}=pv_{i} holds for every ii. If an irreducible element divided qq, it would divide every viv_{i}, contrary to primitivity. Hence, qq is a unit and rjRr_{j}\in R. Thus, U=vrU=vr^{\top} with rR3r\in R^{3}.

If UU is symmetric, then vr=rvvr^{\top}=rv^{\top}. Hence, r=ψvr=\psi v for some ψK\psi\in K. The same denominator argument shows that ψR\psi\in R. When the entries of UU are homogeneous, the greatest common divisor may be chosen homogeneous. Degree comparison then makes vv, rr, and ψ\psi homogeneous. ∎

For a mixed solution, the adjugate of the symmetric gradient has rank one. The factorization lemma therefore produces the polynomial kernel vector in Theorem 3. It also shows why a vanishing-Hessian argument arises naturally.

Proof of the kernel-field assertion in Theorem 3.

For a mixed type, GDG_{D} has rank two on the positive cone and detGD0\det G_{D}\equiv 0. The identity adj(adjGD)=(detGD)GD=0\operatorname{adj}(\operatorname{adj}G_{D})=(\det G_{D})G_{D}=0 shows that every 2×22\times 2 minor of adjGD\operatorname{adj}G_{D} vanishes. The matrix adjGD\operatorname{adj}G_{D} is not identically zero because GDG_{D} has rank two on a nonempty open set. Apply Lemma D.10 in R=[Sym3]R=\mathbb{R}[\operatorname{Sym}_{3}]. Since adjGD\operatorname{adj}G_{D} is symmetric and homogeneous, there are a primitive homogeneous vector 𝝂\boldsymbol{\nu} and a nonzero homogeneous polynomial ψ\psi such that

adjGD=ψ𝝂𝝂.\operatorname{adj}G_{D}=\psi\boldsymbol{\nu}\boldsymbol{\nu}^{\top}.

The identity GDadjGD=0G_{D}\operatorname{adj}G_{D}=0 gives ψ(GD𝝂)𝝂=0\psi(G_{D}\boldsymbol{\nu})\boldsymbol{\nu}^{\top}=0. Choose an index jj for which 𝝂j0\boldsymbol{\nu}_{j}\not\equiv 0. Since RR is an integral domain, every component of GD𝝂G_{D}\boldsymbol{\nu} vanishes. Hence, GD𝝂0G_{D}\boldsymbol{\nu}\equiv 0.

Differentiate this identity in a symmetric direction HH and left multiply by 𝝂\boldsymbol{\nu}^{\top}. This gives 𝝂(HGD)𝝂=0\boldsymbol{\nu}^{\top}(\partial_{H}G_{D})\boldsymbol{\nu}=0. Thus, the rank-one direction 𝝂𝝂\boldsymbol{\nu}\boldsymbol{\nu}^{\top} lies in the radical of the coordinate Hessian wherever 𝝂0\boldsymbol{\nu}\neq 0. In particular, detHessD0\det\operatorname{Hess}D\equiv 0. ∎

We can now assemble the three-variable reduction. The proof first strips full determinant factors and then applies the rank-one, pure plane-minor, or mixed analysis according to the boundary restriction.

Proof of Theorem 3.

Apply Lemma D.4 repeatedly to write the original polynomial as D=δeED=\delta^{e}E, where δE\delta\nmid E. If pE0p_{E}\not\equiv 0, Lemma D.8 and Proposition D.1 give case (i). Suppose that pE0p_{E}\equiv 0. Lemma D.6 supplies a common face type (a,b)(a,b). Evaluating a nonzero face restriction at a rank-one matrix shows that b1b\geq 1. If a=0a=0, the pure plane-minor argument above gives case (ii). If a1a\geq 1, the type is mixed. Proposition D.2, Corollary D.1, and the kernel-field argument give every assertion in case (iii). Since a mixed type has degree a+2b3a+2b\geq 3, no mixed case occurs in degrees at most two. ∎

Theorem 3 leaves one geometric question. The next lemma shows that constancy of the kernel line is exactly the missing condition for the mixed flag form.

Lemma D.11 (Constant kernel and the mixed flag form).

Let EE be determinant-free and of mixed face type (a,b)(a,b). The following statements are equivalent.

  1. (i)

    The projective kernel line is constant on a nonempty open subset of 𝕊++3\mathbb{S}^{3}_{++}.

  2. (ii)

    A fixed nonzero vector u0u_{0} satisfies GE(Σ)u00G_{E}(\Sigma)u_{0}\equiv 0.

  3. (iii)

    There are a nonzero constant γ\gamma, a vector vu0v\in u_{0}^{\perp}, and a full-row-rank matrix MM with rowspace(M)=u0\operatorname{rowspace}(M)=u_{0}^{\perp} such that

    E(Σ)=γ(vΣv)adet(MΣM)b.E(\Sigma)=\gamma(v^{\top}\Sigma v)^{a}\det(M\Sigma M^{\top})^{b}.

Every mixed flag power has this constant kernel line.

Proof of Lemma D.11.

If the projective kernel line is constant on a nonempty open set, choose a fixed representative u00u_{0}\neq 0. Then GE(Σ)u0=0G_{E}(\Sigma)u_{0}=0 on that open set and hence identically, because its entries are polynomials. This proves (i)\Rightarrow(ii), and the reverse implication is immediate.

Assume (ii). After an orthogonal congruence, take u0=e3u_{0}=e_{3}. Under the symmetric-gradient convention,

(GEe3)1=12σ13E,(GEe3)2=12σ23E,(GEe3)3=σ33E.(G_{E}e_{3})_{1}=\frac{1}{2}\partial_{\sigma_{13}}E,\quad(G_{E}e_{3})_{2}=\frac{1}{2}\partial_{\sigma_{23}}E,\quad(G_{E}e_{3})_{3}=\partial_{\sigma_{33}}E.

These derivatives vanish identically, so EE depends only on the leading 2×22\times 2 block SS. Write E(Σ)=E¯(S)E(\Sigma)=\bar{E}(S). The block form of GEΣG_{E}\Sigma gives

tr{(GE(Σ)Σ)2}=tr{(GE¯(S)S)2}.\operatorname{tr}\{(G_{E}(\Sigma)\Sigma)^{2}\}=\operatorname{tr}\{(G_{\bar{E}}(S)S)^{2}\}.

Thus, E¯\bar{E} is exactly self-normalizing on Sym2\operatorname{Sym}_{2} with the same constant. Theorem 2 gives E¯(S)=γ(wSw)a(detS)b\bar{E}(S)=\gamma(w^{\top}Sw)^{a^{\prime}}(\det S)^{b^{\prime}}. Face-type uniqueness in Lemma D.6 gives (a,b)=(a,b)(a^{\prime},b^{\prime})=(a,b). Lifting ww to ve3v\in e_{3}^{\perp} proves (iii).

Conversely, both gradient factors of a mixed flag power annihilate u0u_{0}. The product rule therefore gives GE(Σ)u0=0G_{E}(\Sigma)u_{0}=0 identically. Since GEG_{E} has rank two on 𝕊++3\mathbb{S}^{3}_{++}, its kernel is exactly span(u0)\operatorname{span}(u_{0}) there. ∎

The cofactor identity provides a second constraint on the moving kernel. It couples the kernel vector to the determinant boundary without assuming that the kernel line is constant.

Proposition D.3 (Cofactor syzygy).

Let EE be a determinant-free mixed solution of type (a,b)(a,b), and write adj(GE)=ψ𝛎𝛎\operatorname{adj}(G_{E})=\psi\boldsymbol{\nu}\boldsymbol{\nu}^{\top}. Then

ψ(Σ)𝝂(Σ)adj(Σ)𝝂(Σ)=b(a+b)E(Σ)2.\psi(\Sigma)\boldsymbol{\nu}(\Sigma)^{\top}\operatorname{adj}(\Sigma)\boldsymbol{\nu}(\Sigma)=b(a+b)E(\Sigma)^{2}. (D.4)
Proof of Proposition D.3.

Let A=GEΣA=G_{E}\Sigma. Its eigenvalues are (a+b)E,bE,0(a+b)E,bE,0, so tr{adj(A)}=b(a+b)E2\operatorname{tr}\{\operatorname{adj}(A)\}=b(a+b)E^{2}. For square matrices, adj(BC)=adj(C)adj(B)\operatorname{adj}(BC)=\operatorname{adj}(C)\operatorname{adj}(B). Therefore,

adj(GEΣ)=adj(Σ)adj(GE)=ψadj(Σ)𝝂𝝂.\operatorname{adj}(G_{E}\Sigma)=\operatorname{adj}(\Sigma)\operatorname{adj}(G_{E})=\psi\operatorname{adj}(\Sigma)\boldsymbol{\nu}\boldsymbol{\nu}^{\top}.

Taking the trace proves (D.4) on the positive cone and hence as a polynomial identity. ∎

Lemma D.11 identifies the exact geometric obstruction. Proposition D.3 records an algebraic relation that any nonconstant kernel map must satisfy.

D.6 Boundary rigidity for the mixed branch

The constant-kernel criterion is geometric. The results in this subsection give a separate algebraic closure when the rank-two boundary data already agree with one fixed flag. The argument does not assume the desired interior factorization.

Definition D.1 (Flag-coherent boundary).

A determinant-free mixed polynomial DD of type (a,b)(a,b) has a flag-coherent boundary if there are nested spaces V1V2V_{1}\subset V_{2} and a nonzero constant γ\gamma such that

det(Σ)|DγΔV1aΔV2b.\det(\Sigma)\mid D-\gamma\Delta_{V_{1}}^{a}\Delta_{V_{2}}^{b}. (D.5)

For K=a+2bK=a+2b and α=(a+b)2+b2\alpha=(a+b)^{2}+b^{2}, put

qj=K3j,𝒲j(a,b)={bqj+ap:p=0,,qj}.q_{j}=K-3j,\quad\mathcal{W}_{j}(a,b)=\{bq_{j}+ap:p=0,\ldots,q_{j}\}.

The set 𝒲j(a,b)\mathcal{W}_{j}(a,b) is the collection of weights available to a homogeneous correction of degree qjq_{j} on the determinant boundary.

Definition D.2 (Boundary nonresonance).

The type (a,b)(a,b) is boundary-nonresonant if

αjK𝒲j(a,b){j=1,,K/3}.\alpha-jK\notin\mathcal{W}_{j}(a,b)\quad\{j=1,\ldots,\lfloor K/3\rfloor\}. (D.6)

The next lemma gives two arithmetic forms of this condition. They are useful for locating the first possible resonance.

Lemma D.12 (Arithmetic characterization of boundary nonresonance).

A type (a,b)(a,b) is resonant if and only if there is an integer j{1,,K/3}j\in\{1,\ldots,\lfloor K/3\rfloor\} such that

a|jb,j(2a+b)ab.a\mid jb,\quad j(2a+b)\leq ab. (D.7)

Equivalently, if g=gcd(a,b)g=\gcd(a,b), boundary nonresonance is

b(g1)<2a.b(g-1)<2a. (D.8)
Proof of Lemma D.12.

A resonance has the form αjK=b(K3j)+ap\alpha-jK=b(K-3j)+ap for an integer 0pK3j0\leq p\leq K-3j. Solving for pp gives

p=a+bj+jba.p=a+b-j+\frac{jb}{a}.

For 1jK/31\leq j\leq\lfloor K/3\rfloor, this quantity is positive. It is an integer if and only if a|jba\mid jb. The upper bound pK3jp\leq K-3j is equivalent to j(2a+b)abj(2a+b)\leq ab. This proves (D.7).

Write a=ga0a=ga_{0} and b=gb0b=gb_{0}, where gcd(a0,b0)=1\gcd(a_{0},b_{0})=1. The divisibility condition a|jba\mid jb is equivalent to a0|ja_{0}\mid j. Hence, the smallest positive admissible value is j0=a/gj_{0}=a/g. Since j(2a+b)j(2a+b) increases with jj, a resonance exists if and only if

ag(2a+b)ab.\frac{a}{g}(2a+b)\leq ab.

After division by aa, this is b(g1)2ab(g-1)\geq 2a. If this inequality holds, then j0ab/(2a+b)j_{0}\leq ab/(2a+b) and

3j03ab2a+ba+2b=K,3j_{0}\leq\frac{3ab}{2a+b}\leq a+2b=K,

because 3ab(a+2b)(2a+b)3ab\leq(a+2b)(2a+b). Thus, j0K/3j_{0}\leq\lfloor K/3\rfloor automatically. Negating the resonance criterion gives (D.8). ∎

Boundary nonresonance excludes every polynomial correction compatible with the weighted Euler equation on a rank-two face. It therefore turns flag-coherent boundary data into a unique interior solution.

Proposition D.4 (Boundary-coherent mixed converse).

If a determinant-free exactly self-normalizing polynomial of mixed type (a,b)(a,b) has a flag-coherent boundary and is boundary-nonresonant, then D=γΔV1aΔV2bD=\gamma\Delta_{V_{1}}^{a}\Delta_{V_{2}}^{b}.

Proof of Proposition D.4.

By congruence, take

D0=γσ11a(σ11σ22σ122)b.D_{0}=\gamma\sigma_{11}^{a}(\sigma_{11}\sigma_{22}-\sigma_{12}^{2})^{b}.

If DD0D\neq D_{0}, boundary coherence gives D=D0+δjRD=D_{0}+\delta^{j}R for an exact j1j\geq 1, where δR\delta\nmid R and degR=qj=K3j\deg R=q_{j}=K-3j. Put A0=GD0ΣA_{0}=G_{D_{0}}\Sigma and B=jRI+GRΣB=jRI+G_{R}\Sigma. Then GDΣ=A0+δjBG_{D}\Sigma=A_{0}+\delta^{j}B. Since DD and D0D_{0} have the same constant α=c/2\alpha=c/2, expansion of (4), cancellation, and restriction to δ=0\delta=0 give

tr(A0GRΣ)=(αjK)D0R.\operatorname{tr}(A_{0}G_{R}\Sigma)=(\alpha-jK)D_{0}R. (D.9)

Write J0=GD0/D0J_{0}=G_{D_{0}}/D_{0} and Z0=ΣJ0ΣZ_{0}=\Sigma J_{0}\Sigma. The left-hand side of (D.9) is D0dRΣ(Z0)D_{0}\,\mathrm{d}R_{\Sigma}(Z_{0}). Hence, on a dense boundary chart,

dRΣ(Z0)=(αjK)R.\mathrm{d}R_{\Sigma}(Z_{0})=(\alpha-jK)R. (D.10)

Use the polynomial LDLLDL^{\top} parametrization

Σ=Tdiag(r1,r2,r3)T,T=(100t2110t31t321).\Sigma=T\operatorname{diag}(r_{1},r_{2},r_{3})T^{\top},\quad T=\begin{pmatrix}1&0&0\\ t_{21}&1&0\\ t_{31}&t_{32}&1\end{pmatrix}.

Then Δ1=r1\Delta_{1}=r_{1}, Δ2=r1r2\Delta_{2}=r_{1}r_{2}, δ=r1r2r3\delta=r_{1}r_{2}r_{3}, and D0=γr1a+br2bD_{0}=\gamma r_{1}^{a+b}r_{2}^{b}. A direct logarithmic-gradient calculation gives

Z0=Tdiag{(a+b)r1,br2,0}T.Z_{0}=T\operatorname{diag}\{(a+b)r_{1},br_{2},0\}T^{\top}.

At fixed TT, equation (D.10) on r3=0r_{3}=0 becomes

{(a+b)r1r1+br2r2}R¯=(αjK)R¯.\{(a+b)r_{1}\partial_{r_{1}}+br_{2}\partial_{r_{2}}\}\bar{R}=(\alpha-jK)\bar{R}.

Since RR is homogeneous of degree qjq_{j} and Σ\Sigma is linear in (r1,r2,r3)(r_{1},r_{2},r_{3}), its restriction has the expansion

R¯(r1,r2,0,T)=p=0qjcp(T)r1pr2qjp.\bar{R}(r_{1},r_{2},0,T)=\sum_{p=0}^{q_{j}}c_{p}(T)r_{1}^{p}r_{2}^{q_{j}-p}.

The monomial indexed by pp has weight bqj+apbq_{j}+ap. Boundary nonresonance forces every cp(T)c_{p}(T) to vanish. Thus, RR vanishes on a nonempty open subset of the rank-two stratum. Lemma D.2 gives δ|R\delta\mid R, which contradicts the exact choice of jj. Therefore, D=D0D=D_{0}. ∎

The arithmetic condition is automatic in the lowest mixed degrees. The first normalized resonance appears only in degree nine.

Corollary D.2 (Low mixed degrees and the resonance locus).

After mixed adjugate duality, normalize the type by aba\geq b. A flag-coherent solution is a flag power when b2b\leq 2, and hence for every mixed degree at most eight. The first normalized resonance is (3,3)(3,3) in degree nine. More generally, resonance is equivalent to b{gcd(a,b)1}2ab\{\gcd(a,b)-1\}\geq 2a.

Proof of Corollary D.2.

A resonance requires j(2a+b)abj(2a+b)\leq ab. If b=1b=1 or b=2b=2, this fails for j=1j=1 and therefore for every larger jj. If ab3a\geq b\geq 3, then K=a+2b9K=a+2b\geq 9. At (3,3)(3,3), the choice j=1j=1 gives equality and satisfies a|jba\mid jb. The greatest-common-divisor characterization follows from Lemma D.12. ∎

The preceding proposition gives a conditional closure of the three-variable converse that is independent of kernel-line constancy. It applies whenever the face factors glue to one fixed flag and the resulting type is nonresonant.

Theorem D.1 (Conditional characterization in dimension three).

Suppose that every determinant-free mixed residue can, after mixed adjugate duality if necessary, be chosen with a flag-coherent boundary and a boundary-nonresonant type. On this class, the exactly self-normalizing polynomials are precisely the flag powers.

Proof of Theorem D.1.

Strip determinant powers and apply Theorem 3 to the residue. The rank-one and pure plane-minor cases are flag powers. In the mixed case, apply the assumed adjugate normalization, boundary coherence, and Proposition D.4. Mixed duality returns the flag form to the original residue. Restoring determinant powers adds the full space to the flag. ∎

The conditional theorem isolates two separate issues in the mixed branch. The first is geometric gluing of the face directions. The second is an explicit arithmetic resonance that begins only at higher degree.

D.7 A conditional converse in arbitrary dimension

The three-variable spectral argument does not extend directly to higher dimensions. A converse is nevertheless available once the relative gradients preserve one fixed recursive ordering. The next proposition shows that this invariant flag determines the polynomial uniquely.

Proposition D.5 (Recursive-flag converse).

Let DD be homogeneous, nonconstant and exactly self-normalizing on Symd\operatorname{Sym}_{d}. Suppose that a complete flag 0=F0F1Fd=d0=F_{0}\subset F_{1}\subset\cdots\subset F_{d}=\mathbb{R}^{d} is preserved by RD(Σ)R_{D}(\Sigma) for every Σ0\Sigma\succ 0. Suppose also that the scalar induced on Fi/Fi1F_{i}/F_{i-1} is a constant mim_{i}. After a congruence,

D(Σ)=γr=1ddet(Σ1:r)er,er=mrmr+10,md+1=0.D(\Sigma)=\gamma\prod_{r=1}^{d}\det(\Sigma_{1:r})^{e_{r}},\quad e_{r}=m_{r}-m_{r+1}\in\mathbb{Z}_{\geq 0},\quad m_{d+1}=0.

Conversely, every flag power has such a fixed recursive flag.

Proof of Proposition D.5.

By Lemma B.3, DD has no zero on 𝕊++d\mathbb{S}^{d}_{++}. Hence, log|D|\log|D| is globally defined on the connected cone and on its connected Cholesky parametrization. Use a congruence to take Fi=span(e1,,ei)F_{i}=\operatorname{span}(e_{1},\ldots,e_{i}). Then RD(Σ)R_{D}(\Sigma) is upper triangular with diagonal (m1,,md)(m_{1},\ldots,m_{d}). Write the unique Cholesky factorization Σ=LL\Sigma=LL^{\top}, where LL is lower triangular with positive diagonal, and put J=GD/DJ=G_{D}/D. The matrix

C=LJL=LRD(Σ)LC=L^{\top}JL=L^{\top}R_{D}(\Sigma)L^{-\top}

is symmetric. It is also upper triangular because RD(Σ)R_{D}(\Sigma) preserves the coordinate flag and conjugation by LL^{\top} and LL^{-\top} preserves upper triangularity. Therefore, CC is diagonal. Triangular conjugation preserves diagonal entries, so C=diag(m1,,md)C=\operatorname{diag}(m_{1},\ldots,m_{d}).

For an infinitesimal lower-triangular perturbation dL=LA\mathrm{d}L=LA,

dΣ=L(A+A)L.\mathrm{d}\Sigma=L(A+A^{\top})L^{\top}.

Hence,

dlog|D|=tr{C(A+A)}=2i=1dmidLiiLii.\mathrm{d}\log|D|=\operatorname{tr}\{C(A+A^{\top})\}=2\sum_{i=1}^{d}m_{i}\frac{\mathrm{d}L_{ii}}{L_{ii}}.

Integration on the connected Cholesky domain gives

D(LL)=γi=1dLii2miD(LL^{\top})=\gamma\prod_{i=1}^{d}L_{ii}^{2m_{i}}

for a nonzero constant γ\gamma.

Restrict to diagonal Cholesky factors and write xi=Lii2>0x_{i}=L_{ii}^{2}>0. Then D{diag(x1,,xd)}D\{\operatorname{diag}(x_{1},\ldots,x_{d})\} is a polynomial in the xix_{i} and equals γiximi\gamma\prod_{i}x_{i}^{m_{i}} on the positive orthant. Fixing all variables except xix_{i} shows that each mim_{i} is a nonnegative integer. Since Δr=detΣ1:r=irLii2\Delta_{r}=\det\Sigma_{1:r}=\prod_{i\leq r}L_{ii}^{2}, we obtain in the fraction field

D=γr=1dΔrmrmr+1.D=\gamma\prod_{r=1}^{d}\Delta_{r}^{m_{r}-m_{r+1}}.

By Lemma D.1, the Δr\Delta_{r} are pairwise nonassociate irreducible polynomials. The valuation of the polynomial DD at Δr\Delta_{r} is mrmr+1m_{r}-m_{r+1} and must be nonnegative. This proves the factorization.

The converse is the upper-triangular calculation in the proof of Theorem 1. If only a constant spectrum is assumed in addition to a common invariant flag, each quotient weight is a continuous map from the connected cone to a finite multiset and is therefore constant. ∎

Proposition D.5 separates the higher-dimensional obstruction from the integration step. Once one invariant recursive flag is present, no additional polynomial solutions occur.

D.8 Proof of the low-degree-factor diagnostic

The final part of Section 4 concerns denominators whose determinant-free irreducible factors have degree at most two. The proof identifies global linear and quadratic factors from their generic plane restrictions. It then uses the two-subspace criterion in Supplementary Section C to force nesting.

Lemma D.13 (Generic linear factors).

Let LB(Σ)=tr(BΣ)L_{B}(\Sigma)=\operatorname{tr}(B\Sigma) with BSym3B\in\operatorname{Sym}_{3} and B0B\neq 0. If the restriction of LBL_{B} to a Zariski-dense set of planes is a rank-one linear form on Sym2\operatorname{Sym}_{2}, then BB has rank one. If two such forms are associates on a dense set of planes, their rank-one directions are proportional.

Proof of Lemma D.13.

If rankB2\operatorname{rank}B\geq 2, the symmetric bilinear form represented by BB has a nondegenerate two-dimensional compression. Rank at least two persists on a Zariski-open neighbourhood of that plane, contrary to the hypothesis. Thus, B=ηvvB=\eta vv^{\top}. For two rank-one forms, association on a dense family of planes makes the wedge of the two projected vectors vanish as a polynomial in the plane coordinates. It therefore vanishes on every plane. The plane spanned by two nonproportional directions would give a contradiction, so the directions are proportional. ∎

The next lemma globalizes a pure-power restriction on generic projective lines. It is used to exclude the square alternative for an irreducible quadratic factor.

Lemma D.14 (Single-support restriction).

Let pp be a nonzero homogeneous form of degree m2m\geq 2 on 3\mathbb{C}^{3}. If the restriction of pp to a Zariski-dense set of projective lines is an mmth power of a linear form, then p=γmp=\gamma\ell^{m} for a linear form \ell.

Proof of Lemma D.14.

On each line in the stated family, the zero set has one support point. If the reduced plane curve {p=0}red\{p=0\}_{\mathrm{red}} had degree at least two, a general line would meet it transversally in at least two points by Bezout’s theorem. Hence, the reduced curve is a line, and unique factorization gives the claim. ∎

An admissible irreducible quadratic must vanish on the rank-one cone. The next lemma identifies it as a linear form in the adjugate matrix.

Lemma D.15 (Irreducible quadratic factors).

Let QQ be an irreducible real quadratic on Sym3\operatorname{Sym}_{3}. Suppose that, on a dense family of planes, every irreducible factor of QVQ_{V} is either the unique rank-one linear face factor or detS\det S. Then Q(Σ)=tr{Cadj(Σ)}Q(\Sigma)=\operatorname{tr}\{C\operatorname{adj}(\Sigma)\} for a nonzero matrix CSym3C\in\operatorname{Sym}_{3}.

Proof of Lemma D.15.

Work in an affine chart of Gr(2,3)\operatorname{Gr}(2,3) in which a basis matrix P(t)P(t) depends polynomially on the chart coordinates tt. The coefficients of QV(t)(S)=Q{P(t)SP(t)}Q_{V(t)}(S)=Q\{P(t)SP(t)^{\top}\} are polynomial in tt. Proportionality to detS\det S is cut out by the 2×22\times 2 minors of the corresponding coefficient vectors. Being a scalar multiple of a square of a linear form is cut out by the 2×22\times 2 minors of the symmetric coefficient matrix of the ternary quadratic QV(t)Q_{V(t)}. These equations are compatible on overlaps and define intrinsic Zariski-closed subsets 𝒵det\mathcal{Z}_{\det} and 𝒵sq\mathcal{Z}_{\rm sq} of the complexified Grassmannian. The real Grassmannian is Zariski dense in its complexification. On a dense open set of good real planes, a degree-two restriction is either proportional to detS\det S or to the square of the unique linear face factor. Hence, Gr(2,3)=𝒵det𝒵sq\operatorname{Gr}(2,3)_{\mathbb{C}}=\mathcal{Z}_{\det}\cup\mathcal{Z}_{\rm sq}. Since the Grassmannian is irreducible, one of these closed subsets is the whole Grassmannian.

Consider first the determinant alternative. Proportionality to detS\det S then holds for every plane, including a zero restriction. Let 𝒢={VGr(2,3):QV0}\mathcal{G}=\{V\in\operatorname{Gr}(2,3)_{\mathbb{C}}:Q_{V}\not\equiv 0\}. This is a nonempty Zariski-open set. It is nonempty because otherwise QQ would vanish on the rank-two determinant hypersurface and would be divisible by the cubic detΣ\det\Sigma, which is impossible for a nonzero quadratic. For V𝒢V\in\mathcal{G}, QV=ηVdetSQ_{V}=\eta_{V}\det S with ηV0\eta_{V}\neq 0. Put =Gr(2,3)𝒢\mathcal{B}=\operatorname{Gr}(2,3)_{\mathbb{C}}\setminus\mathcal{G}. Identify Gr(2,3)\operatorname{Gr}(2,3)_{\mathbb{C}} with the dual projective plane. The planes containing a fixed point [u]2[u]\in\mathbb{P}^{2}_{\mathbb{C}} form a projective line Πu\Pi_{u}. If [u][u] lies on no plane in 𝒢\mathcal{G}, then Πu\Pi_{u}\subseteq\mathcal{B}. A proper closed subset of the projective plane has only finitely many one-dimensional irreducible components. Therefore, this can occur for at most finitely many pencils Πu\Pi_{u}. For every other [u][u], there is V𝒢V\in\mathcal{G} with uVu\in V, and Q(uu)=QV(ξξ)=0Q(uu^{\top})=Q_{V}(\xi\xi^{\top})=0. Thus, QQ vanishes on a Zariski-dense subset of the rank-one cone and hence on the whole cone.

The space of quadratics vanishing on {uu:u3}\{uu^{\top}:u\in\mathbb{R}^{3}\} is {tr(CadjΣ):CSym3}\{\operatorname{tr}(C\operatorname{adj}\Sigma):C\in\operatorname{Sym}_{3}\}. Quadratics on Sym3\operatorname{Sym}_{3} have dimension 2121, while quartics in uu have dimension 1515. The restriction map is surjective because every quartic monomial is a product of two quadratic monomials. The displayed family has dimension six and is injective. Indeed, if tr{Cadj(Σ)}0\operatorname{tr}\{C\operatorname{adj}(\Sigma)\}\equiv 0, then for every positive-definite TT choose Σ=(detT)1/2T1\Sigma=(\det T)^{1/2}T^{-1} so that adj(Σ)=T\operatorname{adj}(\Sigma)=T. Then tr(CT)=0\operatorname{tr}(CT)=0 on an open set, which gives C=0C=0. The displayed family is therefore exactly the kernel of the restriction map.

Suppose instead that QVQ_{V} is generically a square. Then pQ(u)=Q(uu)p_{Q}(u)=Q(uu^{\top}) restricts to a fourth power on a dense set of lines. Lemma D.14 gives pQ(u)=γ(vu)4p_{Q}(u)=\gamma(v^{\top}u)^{4}. Consequently, Q=γ(vΣv)2+tr(CadjΣ)Q=\gamma(v^{\top}\Sigma v)^{2}+\operatorname{tr}(C\operatorname{adj}\Sigma) for some CC. On a generic plane, the first two quadratic expressions are squares of linear forms, while the last is (nCn)detS(n^{\top}Cn)\det S, where nn is a normal to the plane. A difference of two squares has matrix rank at most two as a quadratic form on Sym2\operatorname{Sym}_{2}. By contrast, detS=s11s22s122\det S=s_{11}s_{22}-s_{12}^{2} has rank three as a quadratic form on (s11,s12,s22)(s_{11},s_{12},s_{22}). Hence, nCn0n^{\top}Cn\neq 0 is impossible. Since nCn=0n^{\top}Cn=0 for a Zariski-dense set of normals, C=0C=0. Then QQ is a square, contrary to irreducibility. Only the determinant alternative remains. ∎

The preceding lemmas show that all global linear factors are powers of one variance direction and all irreducible quadratic factors are adjugate-linear. Adjugate duality then forces the quadratic directions to coincide.

Proof of Theorem 4.

Strip the largest determinant power and call the determinant-free residue EE. Factor it over \mathbb{R} as

E=γi=1rLiαij=1sQjβj,E=\gamma\prod_{i=1}^{r}L_{i}^{\alpha_{i}}\prod_{j=1}^{s}Q_{j}^{\beta_{j}},

where the LiL_{i} are pairwise nonassociate linear forms and the QjQ_{j} are pairwise nonassociate irreducible quadratics. Choose a generic good plane on which all indicated restrictions are nonzero. The common face-type representation is EV(S)=γVV(S)a(detS)bE_{V}(S)=\gamma_{V}\ell_{V}(S)^{a}(\det S)^{b}. If a=0a=0, unique factorization on a generic face forces r=0r=0. If a>0a>0, every Li|VL_{i}|_{V} is associated with the unique linear factor V\ell_{V}. Lemma D.13 shows that all LiL_{i} are rank-one and associate. Their product is therefore absent or has the form γ1(vΣv)p\gamma_{1}(v^{\top}\Sigma v)^{p}.

For each QjQ_{j}, its generic restriction is either a square of V\ell_{V}, when a>0a>0, or a multiple of detS\det S. Lemma D.15 gives Qj(Σ)=tr{Cjadj(Σ)}Q_{j}(\Sigma)=\operatorname{tr}\{C_{j}\operatorname{adj}(\Sigma)\}. Put q=jβjq=\sum_{j}\beta_{j}. If q>0q>0, adjugate duality yields

E=γ2δq{vadj(Σ)v}pj{tr(CjΣ)}βj,E^{\dagger}=\gamma_{2}\delta^{q}\{v^{\top}\operatorname{adj}(\Sigma)v\}^{p}\prod_{j}\{\operatorname{tr}(C_{j}\Sigma)\}^{\beta_{j}},

where the factor involving vv is omitted when p=0p=0. The displayed power of the irreducible cubic δ\delta is exact. After stripping it, the residue is again determinant-free and exactly self-normalizing. Its linear factors tr(CjΣ)\operatorname{tr}(C_{j}\Sigma) restrict on a generic face to the unique linear face factor. Lemma D.13 shows that every CjC_{j} has rank one and that all are proportional. Hence, Cj=ηjuuC_{j}=\eta_{j}uu^{\top} and

E(Σ)=γ3(vΣv)p{uadj(Σ)u}q,E(\Sigma)=\gamma_{3}(v^{\top}\Sigma v)^{p}\{u^{\top}\operatorname{adj}(\Sigma)u\}^{q},

with an absent factor interpreted as one. The second factor is the determinant of the compression to uu^{\perp}. If p,q>0p,q>0, Lemma C.1 forces the line and plane to be nested. A plane cannot be contained in a line, so span(v)u\operatorname{span}(v)\subset u^{\perp}. Restoring the stripped determinant power gives the flag form in (11). ∎

The corollary converts the flag form into coefficient-level rank and orthogonality tests. It also identifies every reducible self-normalizing cubic.

Proof of Corollary 1.

Necessity is Theorem 4. Sufficiency and the value of cc follow from Theorem 1, whose multiplicities are (a+b+e,b+e,e)(a+b+e,b+e,e). A cubic satisfying the low-degree-factor condition is either a scalar multiple of δ\delta, or its determinant-free part has only linear and quadratic factors. The theorem gives the three possibilities in the corollary. Every reducible cubic belongs to the latter class. ∎

The low-degree-factor converse is unconditional on its stated class. An irreducible factor of degree at least three is the only factorization pattern not decided by this diagnostic.

E Proofs of Section 5: Weak-denominator inference

This section proves Propositions 1 and 2 and verifies the example that delimits the scope of the classification in Section 5 of the main text. Both proofs rest on one triangular-array central limit theorem for the sample covariance, which is established at the start of the first proof and reused in the second.

The proof of Proposition 1 combines this central limit theorem with polynomial delta expansions along the drifting sequence.

Proof of Proposition 1.

For row nn, write

Yni=vech(XniXniΣn),Δn=n1/2i=1nYni=n1/2vech(Σ^Σn).Y_{ni}=\operatorname{vech}(X_{ni}X_{ni}^{\top}-\Sigma_{n}),\quad\Delta_{n}=n^{-1/2}\sum_{i=1}^{n}Y_{ni}=n^{1/2}\,\operatorname{vech}(\hat{\Sigma}-\Sigma_{n}).

Within each row the YniY_{ni} are independent and identically distributed, with covariance Γ(Σn)Γ(Σ)\Gamma(\Sigma_{n})\to\Gamma(\Sigma_{\ast}). Since ΣnΣ\Sigma_{n}\to\Sigma_{\ast}, Gaussian eighth-moment formulas for the quadratic vector Yn1Y_{n1} give

supn𝔼nYn14<.\sup_{n}\mathbb{E}_{n}\lVert Y_{n1}\rVert^{4}<\infty.

Consequently, the Lyapunov quantity n2i=1n𝔼nYni4n^{-2}\sum_{i=1}^{n}\mathbb{E}_{n}\lVert Y_{ni}\rVert^{4} is O(n1)O(n^{-1}). The multivariate triangular-array Lyapunov theorem gives

ΔnNpd{0,Γ(Σ)}.\Delta_{n}\rightsquigarrow N_{p_{d}}\{0,\Gamma(\Sigma_{\ast})\}.

Because (NτnD)(Σn)=0(N-\tau_{n}D)(\Sigma_{n})=0 and the polynomials are fixed, and their second derivatives are bounded on a fixed neighbourhood of Σ\Sigma_{\ast}, Taylor’s theorem together with Σ^Σn=Op(n1/2)\hat{\Sigma}-\Sigma_{n}=O_{p}(n^{-1/2}) yields

n1/2{N(Σ^)τnD(Σ^)}=gnΔn+op(1),n1/2D(Σ^)=n1/2D(Σn)+D(Σn)Δn+op(1).n^{1/2}\{N(\hat{\Sigma})-\tau_{n}D(\hat{\Sigma})\}=g_{n}^{\top}\Delta_{n}+o_{p}(1),\quad n^{1/2}D(\hat{\Sigma})=n^{1/2}D(\Sigma_{n})+\nabla D(\Sigma_{n})^{\top}\Delta_{n}+o_{p}(1).

Since gngg_{n}\to g_{\ast} and D(Σn)D(Σ)\nabla D(\Sigma_{n})\to\nabla D(\Sigma_{\ast}), the pair converges jointly to (sg,Z~,sD,(μ+Z2))(s_{g,\ast}\tilde{Z},s_{D,\ast}(\mu_{\ast}+Z_{2})) with the stated correlation. The limiting denominator has a continuous distribution, so division gives the ratio limit in (18). The convergence sD(Σ^)sD,np0s_{D}(\hat{\Sigma})-s_{D,n}\longrightarrow_{\rm p}0 and sD,nsD,>0s_{D,n}\to s_{D,\ast}>0 then give the limit of F^D\widehat{F}_{D}.

For the Wald statistic, write Wn=τ^τnW_{n}=\hat{\tau}-\tau_{n} and use the direction g^\hat{g} defined in (17). The ratio limit gives Wn=Op(1)W_{n}=O_{p}(1), so the polynomial-gradient expansion gives

g^=gnWnD(Σn)+op(1).\hat{g}=g_{n}-W_{n}\nabla D(\Sigma_{n})+o_{p}(1).

Substituting the ratio limit shows that g^Γ(Σ^)g^/sg,2\hat{g}^{\top}\Gamma(\hat{\Sigma})\hat{g}/s_{g,\ast}^{2} converges in distribution to 12rq+q21-2r_{\ast}q+q^{2}. Combining this with

n1/2|D(Σ^)|Wnsg,Z~sgn(μ+Z2)\frac{n^{1/2}|D(\hat{\Sigma})|W_{n}}{s_{g,\ast}}\rightsquigarrow\tilde{Z}\,\operatorname{sgn}(\mu_{\ast}+Z_{2})

gives (19). The limiting studentizer equals {(qr)2+1r2}1/2\{(q-r_{\ast})^{2}+1-r_{\ast}^{2}\}^{1/2}, which is positive almost surely when |r|<1|r_{\ast}|<1. ∎

The argument uses the drifting sequence only through the limits μ\mu_{\ast}, rr_{\ast} and ω\omega_{\ast}, which is why a single three-parameter experiment covers every covariance-polynomial strategy, as stated in the main text.

The proof of Proposition 2 treats coverage and sample geometry separately. Coverage follows from the same central limit theorem applied to the moment at the local target, and the trichotomy is the sign analysis of one quadratic polynomial in tt.

Proof of Proposition 2.

At t=τnt=\tau_{n}, the moment NτnDN-\tau_{n}D vanishes at Σn\Sigma_{n} and has gradient gng_{n}. The central limit theorem from the proof of Proposition 1, together with hτnΓ(Σ^)hτnpsg,2>0h_{\tau_{n}}^{\top}\Gamma(\hat{\Sigma})h_{\tau_{n}}\longrightarrow_{\rm p}s_{g,\ast}^{2}>0, shows that the squared studentized moment converges in distribution to χ12\chi_{1}^{2}. This proves the coverage statement without requiring a nonzero limit of DD.

For the sample geometry, put q=q1αq=q_{1-\alpha} and

PΣ^(t)=n{N(Σ^)tD(Σ^)}2qhtΓ(Σ^)ht,P_{\hat{\Sigma}}(t)=n\{N(\hat{\Sigma})-tD(\hat{\Sigma})\}^{2}-qh_{t}^{\top}\Gamma(\hat{\Sigma})h_{t},

a quadratic polynomial in tt whose leading coefficient is

A(Σ^)=nD(Σ^)2qsD2(Σ^)=sD2(Σ^){F^Dq}A(\hat{\Sigma})=nD(\hat{\Sigma})^{2}-qs_{D}^{2}(\hat{\Sigma})=s_{D}^{2}(\hat{\Sigma})\{\widehat{F}_{D}-q\}

whenever sD2(Σ^)>0s_{D}^{2}(\hat{\Sigma})>0. Under n>dn>d and Σn0\Sigma_{n}\succ 0, the matrix nΣ^n\hat{\Sigma} has a Wishart density on 𝕊++d\mathbb{S}^{d}_{++} and is therefore absolutely continuous with respect to Lebesgue measure on Symd\operatorname{Sym}_{d} (30, Ch. 3). The polynomial sD2s_{D}^{2} is not identically zero because sD2(Σ)>0s_{D}^{2}(\Sigma_{\ast})>0, so sD2(Σ^)=0s_{D}^{2}(\hat{\Sigma})=0 is a null event. Likewise D(Σ^)=0D(\hat{\Sigma})=0 is a null event, so τ^=N(Σ^)/D(Σ^)\hat{\tau}=N(\hat{\Sigma})/D(\hat{\Sigma}) is defined almost surely.

The event A(Σ^)=0A(\hat{\Sigma})=0 is also null. If the polynomial A(Σ)=nD(Σ)2qsD2(Σ)A(\Sigma)=nD(\Sigma)^{2}-qs_{D}^{2}(\Sigma) is not identically zero, this follows from absolute continuity. If it were identically zero, then sD2=(n/q)D2s_{D}^{2}=(n/q)D^{2} and DD would be exactly self-normalizing, whereas Assumption 1 gives D(Σ)=0D(\Sigma_{\ast})=0 and sD(Σ)>0s_{D}(\Sigma_{\ast})>0, and these two requirements are incompatible with (4).

Finally,

PΣ^(τ^)=qhτ^Γ(Σ^)hτ^0,P_{\hat{\Sigma}}(\hat{\tau})=-qh_{\hat{\tau}}^{\top}\Gamma(\hat{\Sigma})h_{\hat{\tau}}\leq 0,

so the sublevel set {t:PΣ^(t)0}\{t:P_{\hat{\Sigma}}(t)\leq 0\} is nonempty. On the probability-one event {A(Σ^)0}\{A(\hat{\Sigma})\neq 0\}, elementary quadratic geometry shows that this set is an interval, possibly a singleton, the complement of a bounded open interval, or all of \mathbb{R}, and that it is bounded exactly when A(Σ^)>0A(\hat{\Sigma})>0, which is equivalent to F^D>q\widehat{F}_{D}>q. This is the classical Fieller trichotomy (13), and the inversion is also an Anderson–Rubin construction (1). ∎

The final item verifies the example that the main text uses to separate exact self-normalization from mere first-order degeneracy on a zero set.

Example E.1 (First-order degeneracy without exact self-normalization).

On Sym2\operatorname{Sym}_{2}, let D(Σ)=σ122D(\Sigma)=\sigma_{12}^{2}. Then

sD2=4σ122(σ11σ22+σ122),s_{D}^{2}=4\sigma_{12}^{2}(\sigma_{11}\sigma_{22}+\sigma_{12}^{2}),

so sD=0s_{D}=0 vanishes on the entire zero set of DD although sD2/D2s_{D}^{2}/D^{2} is not constant. For σ11,n=σ22,n=1\sigma_{11,n}=\sigma_{22,n}=1 and σ12,n=ζn1/2\sigma_{12,n}=\zeta n^{-1/2},

nD(Σn)2sD2(Σn)ζ24,F^D(ζ+Z)24,ZN1(0,1).\frac{nD(\Sigma_{n})^{2}}{s_{D}^{2}(\Sigma_{n})}\to\frac{\zeta^{2}}{4},\quad\widehat{F}_{D}\rightsquigarrow\frac{(\zeta+Z)^{2}}{4},\quad Z\sim N_{1}(0,1).

The variance formula follows from (3), because the only nonzero coordinate of D\nabla D is D/σ12=2σ12\partial D/\partial\sigma_{12}=2\sigma_{12} and the corresponding diagonal entry of Γ\Gamma is σ11σ22+σ122\sigma_{11}\sigma_{22}+\sigma_{12}^{2}. Cancelling one factor of σ122\sigma_{12}^{2} gives

F^D=nσ^1224(σ^11σ^22+σ^122),\widehat{F}_{D}=\frac{n\hat{\sigma}_{12}^{2}}{4(\hat{\sigma}_{11}\hat{\sigma}_{22}+\hat{\sigma}_{12}^{2})},

wherever σ^120\hat{\sigma}_{12}\neq 0, and the same cancellation at Σn\Sigma_{n} gives the population limit ζ2/4\zeta^{2}/4. Under the displayed drift, n1/2σ^12ζ+Zn^{1/2}\hat{\sigma}_{12}\rightsquigarrow\zeta+Z by the central limit theorem in the proof of Proposition 1, while σ^11σ^22+σ^122p1\hat{\sigma}_{11}\hat{\sigma}_{22}+\hat{\sigma}_{12}^{2}\longrightarrow_{\rm p}1, which proves the limit of F^D\widehat{F}_{D}. Assumption 1 fails along this sequence because sD,=0s_{D,\ast}=0, so the nondegenerate limit of F^D\widehat{F}_{D} arises at second order, exactly as claimed in the main text.

F Proofs of Section 6: Causal applications

F.1 Structural pullbacks used in Table 1

This subsection verifies the third column of Table 1 in the main text. Each entry records a direct substitution into a linear reduced form for that row and is not an additional identification claim. All variables are centred, and symbols are local to the reduced form in which they appear, as in the footnote to the table.

For the instrumental-variable row, let

X=πZ+εX,νZ=var(Z),cov(Z,εX)=0.X=\pi Z+\varepsilon_{X},\quad\nu_{Z}=\operatorname{var}(Z),\quad\operatorname{cov}(Z,\varepsilon_{X})=0.

Then σZX=πνZ\sigma_{ZX}=\pi\nu_{Z}.

For the conditional-instrument row, let

Z=λW+εZ,X=πZ+ηW+εX,νW=var(W),vZ=var(εZ),Z=\lambda W+\varepsilon_{Z},\quad X=\pi Z+\eta W+\varepsilon_{X},\quad\nu_{W}=\operatorname{var}(W),\quad v_{Z}=\operatorname{var}(\varepsilon_{Z}),

where εZ\varepsilon_{Z} is uncorrelated with WW and εX\varepsilon_{X} is uncorrelated with (Z,W)(Z,W). Substituting σZX=πσZZ+ησZW\sigma_{ZX}=\pi\sigma_{ZZ}+\eta\sigma_{ZW} and σWX=πσZW+ησWW\sigma_{WX}=\pi\sigma_{ZW}+\eta\sigma_{WW} cancels the terms in η\eta and gives

σWWσZXσZWσWX=π(σWWσZZσZW2)=πνWvZ.\sigma_{WW}\sigma_{ZX}-\sigma_{ZW}\sigma_{WX}=\pi(\sigma_{WW}\sigma_{ZZ}-\sigma_{ZW}^{2})=\pi\nu_{W}v_{Z}.

For the front-door row, let

M=aX+εM,νX=var(X),vM=var(εM),cov(X,εM)=0.M=aX+\varepsilon_{M},\quad\nu_{X}=\operatorname{var}(X),\quad v_{M}=\operatorname{var}(\varepsilon_{M}),\quad\operatorname{cov}(X,\varepsilon_{M})=0.

Then,

σXX(σXXσMMσXM2)=νX2vM,\sigma_{XX}(\sigma_{XX}\sigma_{MM}-\sigma_{XM}^{2})=\nu_{X}^{2}v_{M},

so the front-door denominator equals νX2vM\nu_{X}^{2}v_{M}.

For the proximal row, let

A=λU+εA,Z=aU+εZ,W=cU+εW,νU=var(U),vA=var(εA),A=\lambda U+\varepsilon_{A},\quad Z=aU+\varepsilon_{Z},\quad W=cU+\varepsilon_{W},\quad\nu_{U}=\operatorname{var}(U),\quad v_{A}=\operatorname{var}(\varepsilon_{A}),

where U,εA,εZ,εWU,\varepsilon_{A},\varepsilon_{Z},\varepsilon_{W} are mutually uncorrelated. Then

σZWσAAσZAσAW=acνU(λ2νU+vA)aλνUλcνU=acνUvA.\sigma_{ZW}\sigma_{AA}-\sigma_{ZA}\sigma_{AW}=ac\nu_{U}(\lambda^{2}\nu_{U}+v_{A})-a\lambda\nu_{U}\,\lambda c\nu_{U}=ac\nu_{U}v_{A}.

These four calculations define the symbols of the table and make every displayed pullback verifiable.

F.2 Marginal and partial minors

This subsection proves the marginal and partial calibrations displayed in Section 6 of the main text. Let a,b,e[d]a,b,e\in[d] with eae\neq a and ebe\neq b, allowing a=ba=b where indicated. As in the main text, ρ^ab\hat{\rho}_{ab} is the sample correlation of coordinates aa and bb, the sample partial correlation ρ^abe\hat{\rho}_{ab\cdot e} is the correlation of the residuals from the sample linear projections of coordinates aa and bb on coordinate ee, and the partial first-stage index is Fpar=nρ^abe 2/(1ρ^abe 2)F_{\mathrm{par}}=n\hat{\rho}_{ab\cdot e}^{\,2}/(1-\hat{\rho}_{ab\cdot e}^{\,2}).

The next proposition contains both variance identities and both closed forms for the standardized denominator.

Proposition F.1 (Marginal and partial minor calibration).

For D=σabD=\sigma_{ab} with aba\neq b,

sD2=σaaσbb+D2,F^D=nρ^ab 21+ρ^ab 2.s_{D}^{2}=\sigma_{aa}\sigma_{bb}+D^{2},\quad\widehat{F}_{D}=\frac{n\hat{\rho}_{ab}^{\,2}}{1+\hat{\rho}_{ab}^{\,2}}.

For D=σeeσabσeaσebD=\sigma_{ee}\sigma_{ab}-\sigma_{ea}\sigma_{eb},

sD2=detΣ{e,a}detΣ{e,b}+3D2.s_{D}^{2}=\det\Sigma_{\{e,a\}}\det\Sigma_{\{e,b\}}+3D^{2}.

When a=ba=b the second identity reduces to sD2=4D2s_{D}^{2}=4D^{2}. When aba\neq b, F^D=nFpar/(n+4Fpar)\widehat{F}_{D}=nF_{\mathrm{par}}/(n+4F_{\mathrm{par}}).

The case a=ba=b recovers the constant c=4c=4 of Theorem 1 for a principal minor of order two. In the two slope cases the standardized denominator is an increasing transform of the familiar marginal or partial first-stage index, so the algebraic diagnostic automatically selects the screening direction relevant to the strategy.

Proof of Proposition F.1.

For D=σabD=\sigma_{ab} with aba\neq b, the only nonzero coordinate of D\nabla D is D/σab=1\partial D/\partial\sigma_{ab}=1, and the corresponding diagonal entry of Γ\Gamma in (2) is σaaσbb+σab2\sigma_{aa}\sigma_{bb}+\sigma_{ab}^{2}, which proves the first identity. Evaluating at Σ^\hat{\Sigma} and dividing the numerator and denominator of F^D\widehat{F}_{D} by σ^aaσ^bb\hat{\sigma}_{aa}\hat{\sigma}_{bb} gives the first closed form.

For D=σeeσabσeaσebD=\sigma_{ee}\sigma_{ab}-\sigma_{ea}\sigma_{eb}, both sides of the asserted identity scale by the same factor under ΣΛΣΛ\Sigma\mapsto\Lambda\Sigma\Lambda for positive diagonal Λ\Lambda, so it suffices to verify it as σee=σaa=σbb=1\sigma_{ee}=\sigma_{aa}=\sigma_{bb}=1. Put x=σeax=\sigma_{ea}, y=σeby=\sigma_{eb} and ρ=σab\rho=\sigma_{ab}. The nonzero coordinates of D\nabla D are D/σee=ρ\partial D/\partial\sigma_{ee}=\rho, D/σab=1\partial D/\partial\sigma_{ab}=1, D/σea=y\partial D/\partial\sigma_{ea}=-y and D/σeb=x\partial D/\partial\sigma_{eb}=-x, and expanding the quadratic form with the entries of (2) gives

sD2=1+3ρ2x2y2+4x2y26ρxy=(1x2)(1y2)+3(ρxy)2.s_{D}^{2}=1+3\rho^{2}-x^{2}-y^{2}+4x^{2}y^{2}-6\rho xy=(1-x^{2})(1-y^{2})+3(\rho-xy)^{2}.

Undoing the rescaling gives the determinant identity.

If a=ba=b, then detΣ{e,a}=detΣ{e,b}=D\det\Sigma_{\{e,a\}}=\det\Sigma_{\{e,b\}}=D, so sD2=4D2s_{D}^{2}=4D^{2}. If aba\neq b, the sample partial correlation satisfies

ρ^abe 2=D(Σ^)2detΣ^{e,a}detΣ^{e,b},\hat{\rho}_{ab\cdot e}^{\,2}=\frac{D(\hat{\Sigma})^{2}}{\det\hat{\Sigma}_{\{e,a\}}\det\hat{\Sigma}_{\{e,b\}}},

so

F^D=nρ^abe 21+3ρ^abe 2,\widehat{F}_{D}=\frac{n\hat{\rho}_{ab\cdot e}^{\,2}}{1+3\hat{\rho}_{ab\cdot e}^{\,2}},

and substituting ρ^abe 2=Fpar/(n+Fpar)\hat{\rho}_{ab\cdot e}^{\,2}=F_{\mathrm{par}}/(n+F_{\mathrm{par}}) gives the second closed form. ∎

F.3 Front-door boundary robustness

This subsection proves Proposition 3. The proof rests on three exact facts. The plug-in estimator factors as a product of two regression coefficients, the Gaussian delta studentizer is an exact combination of the two regression standard errors, and the second coefficient has an exact Student statistic independent of the design. The boundary limit then follows by comparing the two terms of the combination uniformly in the mediator residual variance, which makes quantitative the mechanism described after the proposition in the main text.

Proof of Proposition 3.

Write Zi=(Xi,Mi)Z_{i}=(X_{i},M_{i})^{\top} and let 𝒲n=σ(Z1,,Zn)\mathcal{W}_{n}=\sigma(Z_{1},\ldots,Z_{n}). The covariance formula factors exactly as τ^=a^b^\hat{\tau}=\hat{a}\hat{b}, where

a^=iXiMiiXi2\hat{a}=\frac{\sum_{i}X_{i}M_{i}}{\sum_{i}X_{i}^{2}}

is the known-mean least-squares coefficient in the regression of MM on XX, and the coefficient vector in the regression of YY on (X,M)(X,M) is

(γ^,b^)=(iZiZi)1iZiYi.(\hat{\gamma},\hat{b})^{\top}=(\sum_{i}Z_{i}Z_{i}^{\top})^{-1}\sum_{i}Z_{i}Y_{i}.

Put

v^M=1ni=1n(Mia^Xi)2,σ^ε2=1ni=1n(Yiγ^Xib^Mi)2.\hat{v}_{M}=\frac{1}{n}\sum_{i=1}^{n}(M_{i}-\hat{a}X_{i})^{2},\quad\hat{\sigma}_{\varepsilon}^{2}=\frac{1}{n}\sum_{i=1}^{n}(Y_{i}-\hat{\gamma}X_{i}-\hat{b}M_{i})^{2}.

The usual no-intercept regression standard errors are

se^a2=v^M(n1)σ^XX,se^b2=σ^ε2(n2)v^M.\widehat{\mathrm{se}}_{a}^{2}=\frac{\hat{v}_{M}}{(n-1)\hat{\sigma}_{XX}},\quad\widehat{\mathrm{se}}_{b}^{2}=\frac{\hat{\sigma}_{\varepsilon}^{2}}{(n-2)\hat{v}_{M}}. (F.1)

We first identify the Gaussian delta variance exactly. For a positive-definite covariance matrix, let a(Σ)a(\Sigma) be the regression coefficient of MM on XX, let (γ(Σ),b(Σ))(\gamma(\Sigma),b(\Sigma)) be the coefficient vector in the regression of YY on (X,M)(X,M), and put

RM=Ma(Σ)X,RY=Yγ(Σ)Xb(Σ)M,R_{M}=M-a(\Sigma)X,\quad R_{Y}=Y-\gamma(\Sigma)X-b(\Sigma)M,

and define σMMX=varΣ(RM)\sigma_{MM\cdot X}=\operatorname{var}_{\Sigma}(R_{M}) and σYYXM=varΣ(RY)\sigma_{YY\cdot XM}=\operatorname{var}_{\Sigma}(R_{Y}). The influence functions of the two coefficient functionals under known-mean covariance sampling are

ϕa=XRMσXX,ϕb=RMRYσMMX.\phi_{a}=\frac{XR_{M}}{\sigma_{XX}},\quad\phi_{b}=\frac{R_{M}R_{Y}}{\sigma_{MM\cdot X}}.

The first identity follows by differentiating a=σXM/σXXa=\sigma_{XM}/\sigma_{XX}. The second follows by differentiating the normal equations (γ,b)=Σ(X,M),(X,M)1Σ(X,M),Y(\gamma,b)^{\top}=\Sigma_{(X,M),(X,M)}^{-1}\Sigma_{(X,M),Y} and applying the Frisch–Waugh residualization. Under a centred Gaussian law, XX and RMR_{M} are independent, and RYR_{Y} is independent of (X,M)(X,M). Hence,

aΓa=σMMXσXX,bΓb=σYYXMσMMX,aΓb=0.\nabla a^{\top}\Gamma\nabla a=\frac{\sigma_{MM\cdot X}}{\sigma_{XX}},\quad\nabla b^{\top}\Gamma\nabla b=\frac{\sigma_{YY\cdot XM}}{\sigma_{MM\cdot X}},\quad\nabla a^{\top}\Gamma\nabla b=0. (F.2)

These are identities of rational functions on the positive cone and may therefore be evaluated at Σ^\hat{\Sigma}. Since (ab)=ba+ab\nabla(ab)=b\nabla a+a\nabla b, the Gaussian delta standard error defined in (25) satisfies the exact sample identity

se^ 2=1n{b^ 2v^Mσ^XX+a^ 2σ^ε2v^M}=n1nb^ 2se^a2+n2na^ 2se^b2,\begin{split}\widehat{\mathrm{se}}^{\,2}&=\frac{1}{n}\Biggl\{\hat{b}^{\,2}\frac{\hat{v}_{M}}{\hat{\sigma}_{XX}}+\hat{a}^{\,2}\frac{\hat{\sigma}_{\varepsilon}^{2}}{\hat{v}_{M}}\Biggr\}\\ &=\frac{n-1}{n}\hat{b}^{\,2}\widehat{\mathrm{se}}_{a}^{2}+\frac{n-2}{n}\hat{a}^{\,2}\widehat{\mathrm{se}}_{b}^{2},\end{split} (F.3)

where the second equality is the algebraic substitution of (F.1).

For completeness, the zero cross term in (F.2) also has an exact finite-sample regression justification. Conditional on 𝒲n\mathcal{W}_{n}, the difference b^b\hat{b}-b is a linear function of the independent mean-zero errors εi\varepsilon_{i}, so

𝔼(b^𝒲n)=b,\mathbb{E}(\hat{b}\mid\mathcal{W}_{n})=b,

whereas a^\hat{a} is 𝒲n\mathcal{W}_{n}-measurable. Thus cov(a^,b^)=0\operatorname{cov}(\hat{a},\hat{b})=0 for every nn and every interior covariance. For a fixed interior covariance, Gaussian projection formulas give

𝔼{n2(a^a)4}=3vM2n2νX2(n2)(n4),𝔼{n2(b^b)4}=3σε4n2vM2(n3)(n5)\mathbb{E}\{n^{2}(\hat{a}-a)^{4}\}=\frac{3v_{M}^{2}n^{2}}{\nu_{X}^{2}(n-2)(n-4)},\quad\mathbb{E}\{n^{2}(\hat{b}-b)^{4}\}=\frac{3\sigma_{\varepsilon}^{4}n^{2}}{v_{M}^{2}(n-3)(n-5)}

for all sufficiently large nn. Hence, n(a^a)(b^b)n(\hat{a}-a)(\hat{b}-b) is uniformly integrable. The joint delta limit therefore has zero covariance, in agreement with the direct influence-function calculation.

Conditional on 𝒲n\mathcal{W}_{n}, the usual Gaussian regression decomposition and Cochran’s theorem show that the statistic Tb=(b^b)/se^bT_{b}=(\hat{b}-b)/\widehat{\mathrm{se}}_{b} has the tn2t_{n-2} law. The conditional law does not depend on 𝒲n\mathcal{W}_{n}, so TbT_{b} is independent of the design. Moreover,

nv^MvM,nχn12,nσ^ε2σε2χn22,σε2=var(ε),\frac{n\hat{v}_{M}}{v_{M,n}}\sim\chi^{2}_{n-1},\quad\frac{n\hat{\sigma}_{\varepsilon}^{2}}{\sigma_{\varepsilon}^{2}}\sim\chi^{2}_{n-2},\quad\sigma_{\varepsilon}^{2}=\operatorname{var}(\varepsilon),

with the second chi-squared variable independent of the design. These are the standard Gaussian projection identities (30, Ch. 3). It follows that, uniformly over 0<vM,nv¯0<v_{M,n}\leq\bar{v},

a^a=Op{(vM,n/n)1/2},nvM,nse^b 2pσε2.\hat{a}-a=O_{p}\{(v_{M,n}/n)^{1/2}\},\quad nv_{M,n}\widehat{\mathrm{se}}_{b}^{\,2}\longrightarrow_{\rm p}\sigma_{\varepsilon}^{2}.

The first order also follows from the exact moment identity 𝔼(a^a)2=vM,n/{νX(n2)}\mathbb{E}(\hat{a}-a)^{2}=v_{M,n}/\{\nu_{X}(n-2)\}.

If a subsequence has vM,nv_{M,n} bounded away from zero, the ordinary joint central limit theorem and Slutsky’s theorem prove the result. Consider a subsequence with vM,n0v_{M,n}\to 0. Decompose

a^b^ab=a(b^b)+b(a^a)+(a^a)(b^b).\hat{a}\hat{b}-ab=a(\hat{b}-b)+b(\hat{a}-a)+(\hat{a}-a)(\hat{b}-b).

After division by |a^|se^b|\hat{a}|\widehat{\mathrm{se}}_{b}, the first term is {a/|a^|}Tb\{a/|\hat{a}|\}T_{b}, which converges in distribution to sgn(a)Z\operatorname{sgn}(a)Z with ZN1(0,1)Z\sim N_{1}(0,1), and this limit is standard normal by symmetry. Since se^b1=Op{(nvM,n)1/2}\widehat{\mathrm{se}}_{b}^{-1}=O_{p}\{(nv_{M,n})^{1/2}\}, the second term is Op(vM,n)=op(1)O_{p}(v_{M,n})=o_{p}(1), and the third equals {(a^a)/|a^|}Tb=op(1)\{(\hat{a}-a)/|\hat{a}|\}T_{b}=o_{p}(1).

It remains to compare the delta studentizer with |a^|se^b|\hat{a}|\widehat{\mathrm{se}}_{b}. From (F.1), se^a2=Op(vM,n/n)\widehat{\mathrm{se}}_{a}^{2}=O_{p}(v_{M,n}/n), while b^ 2=Op{1+(nvM,n)1}\hat{b}^{\,2}=O_{p}\{1+(nv_{M,n})^{-1}\}. Therefore,

(n1)b^ 2se^a2(n2)a^ 2se^b2=Op{vM,n2+vM,n/n}=op(1).\frac{(n-1)\hat{b}^{\,2}\widehat{\mathrm{se}}_{a}^{2}}{(n-2)\hat{a}^{\,2}\widehat{\mathrm{se}}_{b}^{2}}=O_{p}\{v_{M,n}^{2}+v_{M,n}/n\}=o_{p}(1).

Equation (F.3) now yields se^/(|a^|se^b)p1\widehat{\mathrm{se}}/(|\hat{a}|\widehat{\mathrm{se}}_{b})\longrightarrow_{\rm p}1. Every subsequence has the same limit, completing the proof. ∎

The final display shows that the delta studentizer is asymptotically equivalent to the dominant regression studentizer uniformly in vM,nv_{M,n}, which is the exact balance behind the boundary robustness.

G Detailed numerical experiments

G.1 Data-generating mechanisms and target parameters

We give the complete specifications used for Table 2. The front-door covariance is generated by X=εXX=\varepsilon_{X}, M=0.8X+εMM=0.8X+\varepsilon_{M} and Y=M+εYY=M+\varepsilon_{Y}, where (εX,εY)(\varepsilon_{X},\varepsilon_{Y}) is centred Gaussian with unit marginal variances and covariance 0.50.5, and εM\varepsilon_{M} is independent Gaussian with variance vMv_{M}. The target is 0.80.8. The proximal covariance is generated by Z=0.8U+εZZ=0.8U+\varepsilon_{Z}, A=0.8U+εAA=0.8U+\varepsilon_{A}, W=0.8U+εWW=0.8U+\varepsilon_{W} and Y=A+0.8U+εYY=A+0.8U+\varepsilon_{Y}. All shocks in the proximal design are mutually independent centred Gaussian variables and have unit variance except var(εA)=vA\operatorname{var}(\varepsilon_{A})=v_{A}; the target is one.

G.2 Computation and precision

For each population covariance Σ\Sigma, we draw the known-mean sample covariance directly from nΣ^𝒲d(n,Σ)n\hat{\Sigma}\sim\mathcal{W}_{d}(n,\Sigma). We compute the plug-in covariance ratio, its Gaussian delta standard error, the standardized denominator and the indicator that the true target belongs to the inverted set in (20). The robust standard deviation is the interquartile range divided by 1.3491.349. Each cell uses n=1000n=1000 and 100000100000 replications. The Wald and inversion calculations use the two-sided 95% standard-normal and χ12\chi_{1}^{2} critical values, respectively. The Monte Carlo standard error attached to a coverage estimate p^\hat{p} is {p^(1p^)/(100000)}1/2\{\hat{p}(1-\hat{p})/(100000)\}^{1/2}; the largest value in the experiment is 0.000810.00081. No replication fails the validity checks.

Table G.1: Complete denominator summaries for the Gaussian simulations. The population quantity is nD(Σ)2/sD2(Σ)nD(\Sigma)^{2}/s_{D}^{2}(\Sigma). Dashes denote diagnostics not used for the self-normalizing front-door denominator.
Model Parameter Population FDF_{D} Median F^D\widehat{F}_{D} Median partial FF Median naive FF
Front-door vM=1.000v_{M}=1.000 100.000100.000 100.000100.000
Front-door vM=0.100v_{M}=0.100 100.000100.000 100.000100.000
Front-door vM=0.010v_{M}=0.010 100.000100.000 100.000100.000
Front-door vM=0.001v_{M}=0.001 100.000100.000 100.000100.000
Proximal vA=1.00v_{A}=1.00 63.72963.729 63.79263.792 85.64785.647 179.670179.670
Proximal vA=0.30v_{A}=0.30 26.48226.482 26.49226.492 29.63229.632 362.006362.006
Proximal vA=0.10v_{A}=0.10 6.2186.218 6.2346.234 6.3936.393 510.227510.227
Proximal vA=0.03v_{A}=0.03 0.7740.774 0.9220.922 0.9260.926 595.168595.168
Proximal vA=0.01v_{A}=0.01 0.0950.095 0.5040.504 0.5050.505 624.555624.555
Table G.2: Dispersion, standard errors and bias in the Gaussian simulations. Robust s.d., interquartile range divided by 1.3491.349.
Model Parameter Robust s.d. Median s.e. Median bias
Front-door vM=1.000v_{M}=1.000 0.038480.03848 0.038480.03848 0.00044-0.00044
Front-door vM=0.100v_{M}=0.100 0.069970.06997 0.069960.06996 0.00001-0.00001
Front-door vM=0.010v_{M}=0.010 0.217450.21745 0.219040.21904 0.000860.00086
Front-door vM=0.001v_{M}=0.001 0.695520.69552 0.692520.69252 0.00295-0.00295
Proximal vA=1.00v_{A}=1.00 0.064150.06415 0.063270.06327 0.00002-0.00002
Proximal vA=0.30v_{A}=0.30 0.172790.17279 0.169750.16975 0.000640.00064
Proximal vA=0.10v_{A}=0.10 0.490800.49080 0.467610.46761 0.008780.00878
Proximal vA=0.03v_{A}=0.03 1.186721.18672 1.433291.43329 0.435270.43527
Proximal vA=0.01v_{A}=0.01 1.460311.46031 2.036952.03695 0.884900.88490
Table G.3: Coverage and Monte Carlo standard errors in the Gaussian simulations. Cov., empirical 95% coverage; MCSE, Monte Carlo standard error; inversion, coverage of (20) at the true target.
Model Parameter Wald cov. Wald MCSE Inversion cov. Inversion MCSE
Front-door vM=1.000v_{M}=1.000 0.949430.94943 0.000690.00069 0.953700.95370 0.000660.00066
Front-door vM=0.100v_{M}=0.100 0.949680.94968 0.000690.00069 0.954500.95450 0.000660.00066
Front-door vM=0.010v_{M}=0.010 0.949200.94920 0.000690.00069 0.953530.95353 0.000670.00067
Front-door vM=0.001v_{M}=0.001 0.949500.94950 0.000690.00069 0.953480.95348 0.000670.00067
Proximal vA=1.00v_{A}=1.00 0.952470.95247 0.000670.00067 0.952540.95254 0.000670.00067
Proximal vA=0.30v_{A}=0.30 0.952040.95204 0.000680.00068 0.951680.95168 0.000680.00068
Proximal vA=0.10v_{A}=0.10 0.935310.93531 0.000780.00078 0.952620.95262 0.000670.00067
Proximal vA=0.03v_{A}=0.03 0.930420.93042 0.000800.00080 0.951500.95150 0.000680.00068
Proximal vA=0.01v_{A}=0.01 0.930910.93091 0.000800.00080 0.951410.95141 0.000680.00068

The front-door median standard errors closely track the empirical robust standard deviations over the entire 1000-fold range of vMv_{M}. For the proximal design, agreement is good when vAv_{A} is 11 or 0.30.3, but the Gaussian Wald approximation deteriorates once the population denominator diagnostic falls below about 1010. At vA=0.03v_{A}=0.03 and 0.010.01, the sample median F^D\widehat{F}_{D} exceeds its population value because sampling noise is no longer small relative to the denominator, while inversion retains near-nominal coverage.

H Detailed real data experiments

H.1 Variables and preprocessing

We use the public SUPPORT right-heart-catheterization data (8), distributed at https://hbiostat.org/data/repo/rhc.csv and containing 57355735 patients and 6363 variables. Treatment is the indicator of right-heart catheterization on the first study day, recorded in the variable swang1 and coded as one for the value RHC and zero for No RHC. The descriptive outcome is survival time in days truncated at 3030, recorded in t3d30. The treatment proxies are pafi1 and paco21, and the outcome proxies are ph1 and hema1. These six analysis variables are required to be observed and are not imputed. The adjustment set contains age, sex, cat1, cat2, dnr1, surv2md1 and aps1, corresponding to the covariates described in Section 8 of the main text. Missing continuous adjustment covariates are median-imputed. Missing categorical adjustment covariates are mode-imputed and coded by indicator variables with one level omitted. Nonconstant design columns are standardized before an intercept is added, and the standardization does not alter the projection space. The resulting design has rank 1919 and leaves 57165716 residual degrees of freedom.

H.2 Diagnostic formulas and bootstrap

Let (A,Z,W,Y)(A,Z,W,Y) denote the residualized treatment, treatment proxy, outcome proxy and outcome. For their empirical second moments, put

D=σAAσZWσAZσAW,N=σZWσAYσZYσAW,D=\sigma_{AA}\sigma_{ZW}-\sigma_{AZ}\sigma_{AW},\quad N=\sigma_{ZW}\sigma_{AY}-\sigma_{ZY}\sigma_{AW},

so the reported proximal plug-in estimate is N/DN/D. Put ν=nr\nu=n-r, where rr is the realized adjustment-design rank. Both denominator diagnostics have the form νD2/s^D2\nu D^{2}/\hat{s}_{D}^{2}. Define MAZ=σAAσZZσAZ2M_{AZ}=\sigma_{AA}\sigma_{ZZ}-\sigma_{AZ}^{2} and MAW=σAAσWWσAW2M_{AW}=\sigma_{AA}\sigma_{WW}-\sigma_{AW}^{2}. The Gaussian version uses s^D2=MAZMAW+3D2\hat{s}_{D}^{2}=M_{AZ}M_{AW}+3D^{2}, as in (21). The sandwich version uses the centred empirical covariance, with divisor nn, of the six second-moment scores (A2,Z2,W2,AZ,AW,ZW)(A^{2},Z^{2},W^{2},AZ,AW,ZW). The naive and partial indices use νρ^2/(1ρ^2)\nu\hat{\rho}^{2}/(1-\hat{\rho}^{2}), as stated in the main text.

We use 20002000 nonparametric row-bootstrap replications. Each resample repeats the complete preprocessing and residualization and uses the realized design rank to determine ν\nu. The resampled design rank is 1919 in 17041704 replications and 1818 in 296296 replications. The intervals below are percentile intervals.

Table H.1: Detailed point diagnostics for the SUPPORT analysis. Partial corr., sample partial correlation between the treatment and outcome proxies given residualized treatment.
Treatment proxy Outcome proxy Partial corr. Estimate Gaussian F^D\widehat{F}_{D} Sandwich F^D\widehat{F}_{D}
pafi1 ph1 0.02360.0236 1.245-1.245 3.1793.179 2.3062.306
pafi1 hema1 0.0750-0.0750 1.412-1.412 31.59131.591 25.71025.710
paco21 ph1 0.5072-0.5072 1.328-1.328 829.990829.990 315.190315.190
paco21 hema1 0.10950.1095 1.136-1.136 66.13166.131 51.33751.337
Table H.2: Percentile bootstrap intervals for the SUPPORT analysis. Intervals use 20002000 row-bootstrap replications. They reflect sampling variation under the prespecified preprocessing and proxy allocation, but not uncertainty about proxy validity.
Treatment proxy Outcome proxy Estimate 95% interval Sandwich-F^D\widehat{F}_{D} 95% interval
pafi1 ph1 (2.249, 0.139)(-2.249,\ 0.139) (0.011, 12.733)(0.011,\ 12.733)
pafi1 hema1 (2.036,0.806)(-2.036,\ -0.806) (9.166, 51.302)(9.166,\ 51.302)
paco21 ph1 (1.839,0.811)(-1.839,\ -0.811) (218.242, 461.934)(218.242,\ 461.934)
paco21 hema1 (1.689,0.563)(-1.689,\ -0.563) (30.317, 78.353)(30.317,\ 78.353)

The weak pafi1/ph1 pair is the only allocation for which the bootstrap estimate interval contains zero and the lower endpoint of the sandwich-diagnostic interval is essentially zero. The other three allocations have diagnostic intervals bounded away from zero, although the discrepancy between the Gaussian and sandwich versions remains substantial for paco21/ph1. These comparisons are diagnostic and should not be interpreted as evidence for the untestable proximal bridge assumptions.

I Algebraic obstructions in the mixed branch

I.1 Kernel map and the Hessian obstruction

For a determinant-free mixed solution EE with self-normalization constant cc, the kernel-field argument in Supplementary Section D produces a primitive homogeneous polynomial vector 𝝂\boldsymbol{\nu}. It defines the complex-projective rational map

[𝝂]:(Sym3)2,[Σ][𝝂(Σ)].[\boldsymbol{\nu}]:\mathbb{P}(\operatorname{Sym}_{3}\otimes_{\mathbb{R}}\mathbb{C})\dashrightarrow\mathbb{P}^{2}_{\mathbb{C}},\quad[\Sigma]\longmapsto[\boldsymbol{\nu}(\Sigma)].

At every point where GE(Σ)G_{E}(\Sigma) has rank two and 𝝂(Σ)0\boldsymbol{\nu}(\Sigma)\neq 0, the image is the projective kernel line [kerGE(Σ)][\ker G_{E}(\Sigma)]. Theorem 3 and Lemma D.11 show that the unrestricted three-variable converse is equivalent to proving that this map is constant.

The identity detHessE0\det\operatorname{Hess}E\equiv 0 does not prove this constancy. The Gordan–Noether cone conclusion holds for forms in at most four variables, but it fails in six variables, which is the dimension of Sym3\operatorname{Sym}_{3} (27; 7). For example, the cubic in variables x0,,x5x_{0},\ldots,x_{5} (32; 17)

P=x0x42+2x1x4x5+x2x52+x33,P=x_{0}x_{4}^{2}+2x_{1}x_{4}x_{5}+x_{2}x_{5}^{2}+x_{3}^{3},

has the nonzero polynomial Hessian-kernel vector

(x52,x4x5,x42,0,0,0).(x_{5}^{2},-x_{4}x_{5},x_{4}^{2},0,0,0)^{\top}.

Its first derivatives are linearly independent, so PP is not a cone. This cubic is not claimed to satisfy (4). It shows only that a polynomial Hessian-kernel field need not have a constant direction. Consequently, the missing implication must use more than Hessian degeneracy.

I.2 Representation-theoretic lifting obstruction

A second possible route uses the derived congruence action ρ\rho. For the matrix units EijE_{ij},

ρ(Eij)E=2(GEΣ)ij.\rho(E_{ij})E=2(G_{E}\Sigma)_{ij}.

Equation (4) therefore implies the polynomial identity

i,j=13{ρ(Eij)E}{ρ(Eji)E}=2cE2.\sum_{i,j=1}^{3}\{\rho(E_{ij})E\}\{\rho(E_{ji})E\}=2cE^{2}. (I.1)

This is the image, under polynomial multiplication, of the tensor relation that one would seek from a highest-weight orbit characterization.

The theorem of 25 is a tensor identity in

Sym2{SymK(Sym23)}.\operatorname{Sym}^{2}\{\operatorname{Sym}^{K}(\operatorname{Sym}^{2}\mathbb{C}^{3})\}.

By contrast, (I.1) lies only in the space of degree-2K2K polynomial functions. The multiplication map from the tensor space to degree-2K2K polynomials has a nontrivial kernel for K2K\geq 2. Hence, the scalar polynomial identity does not determine the required tensor identity. A proof by this route would need an additional argument showing that the components lost under multiplication vanish separately.

These two obstructions explain the scope of Theorem 3. They do not provide evidence for a nonflag solution. They identify the remaining task more precisely. One must combine the integrability of the symmetric gradient with the cofactor syzygy or with determinant-boundary information to force the rational kernel map to be constant.

\CJK@envEnd