UTF8mc\CJK@envStartUTF8
Self-Normalizing Denominators
in Rational Causal Estimation
Abstract
Rational causal estimators in linear structural equation models take the form of one covariance polynomial divided by another, and a small denominator is commonly interpreted as weak identification. We show that, under Gaussian sampling, some denominators cannot enter this regime at first order. Their sampling variation is exactly proportional to their magnitude, so the standardized denominator is constant in every sample. Products of powers of nested covariance minors have this property in every dimension and admit an exact Wishart pivot. The converse is complete in dimension two. In dimension three, one mixed family remains open, while a factor-and-rank criterion classifies all denominators with linear or quadratic determinant-free factors and covers instrumental-variable, front-door and proximal formulas. For linear front-door adjustment, Wald inference remains asymptotically valid even as the mediator residual variance vanishes at an arbitrary rate, provided the treatment–mediator coefficient is nonzero. In simulations, proximal Wald coverage fell as a naive treatment–proxy diagnostic strengthened, while front-door coverage stayed nominal, and right-heart-catheterization data distinguished naive from denominator-relevant diagnostics.
Keywords: Bartlett decomposition; Fieller’s theorem; Fisher–Rao geometry; Proximal causal inference; Symmetric cone; Weak instrument.
1 Introduction
1.1 Motivation and scope
Graphical identification in linear structural equation models often yields a causal coefficient of the form . Here is the observational covariance matrix of observed coordinates, and and are polynomial functions on the space of symmetric matrices of order . Instrumental variables, conditional instruments, front-door adjustment, half-trek identification and proximal formulas all produce such expressions (31; 6; 14; 23; 39; 29; 38; 21). At regular distributions, substituting the sample covariance matrix into such a formula is routine. However, when is small, weak-instrument theory suggests a ratio-of-normals limit, distorted Wald inference and confidence sets that must sometimes be unbounded (16; 11; 36; 37; 2; 3). The standard safeguard compares the estimated denominator with its estimated standard error, as in first-stage relevance screening.
Two of these strategies display the contrast that motivates this paper. Throughout this preview, the observed vector is centred Gaussian with covariance matrix . The scalar denotes the covariance of two observed variables and , and denotes the sample covariance matrix from independent observations. Section 2 gives the remaining conventions. For the front-door strategy, the observed vector is and
At every such , the asymptotic variance of is . Consequently, the statistic that divides by the plug-in value of this variance equals for every positive-definite sample covariance. Screening this denominator reveals nothing about the underlying distribution, although the precision of the plug-in estimator may deteriorate without bound. By contrast, consider a linear proximal strategy with observed coordinates , comprising a treatment, a treatment proxy and an outcome proxy. Its denominator is
which equals times the partial covariance of and given . This denominator vanishes at interior covariance matrices at which its asymptotic variance does not. There are therefore sequences of covariance matrices along which the plug-in ratio can have the ratio-of-normals limit of weak-instrument theory. The relevant screening direction is the partial covariance, not the marginal treatment–proxy association.
These examples motivate the question of which denominators make the first-order ratio regime impossible along every sequence of underlying covariance matrices. Under centred Gaussian sampling, the first-order variance of a covariance polynomial is a classical quadratic form in its symmetric gradient. We show that, for a special class of polynomials, this quadratic form is exactly proportional to the square of the polynomial. Denominator noise then contracts at the same rate as the denominator itself, and the standardized denominator is a samplewise constant rather than a relevance statistic. We call such denominators exactly self-normalizing.
Two neighbouring phenomena are excluded from this notion. A polynomial whose first-order variance merely vanishes on its zero set can still generate a nondegenerate second-order local limit when that variance is not proportional to the square of the polynomial. Section 5 gives an example. Numerator singularities are also distinct. The mediation null remains nonregular even though the front-door denominator is exactly self-normalizing (10). Therefore, the classification below concerns denominators, not universal regularity of the associated estimators.
1.2 Contributions
The first contribution is an all-dimensional sufficient class. Products of powers of covariance determinants along a nested flag of subspaces are exactly self-normalizing. Their relative gradients have a fixed spectrum, and their sample-to-population ratios factor into independent chi-squared variables. In recursive coordinates, these polynomials are monomials in successive conditional variances, which links the algebra directly to variance information.
The second contribution is a low-dimensional converse. The converse is complete in dimension two. In dimension three, determinant stripping reduces every solution to a rank-one branch, a pure plane-minor branch or a mixed branch with constant relative spectrum and a polynomial kernel line. The first two branches are flag powers, while rigidity of the mixed kernel line remains open. For the practically important class whose determinant-free irreducible factors have degree at most two, only one rank-one variance factor and one nested plane-minor factor can occur. This gives a coefficient-level factor-and-rank diagnostic and classifies every reducible self-normalizing cubic.
The statistical results connect this algebra to weak-denominator inference. A local result recovers the classical ratio-of-normals limit under a drift condition that no exactly self-normalizing denominator can satisfy. The standardized denominator is the noncentrality diagnostic of this limit and also determines whether Fieller inversion is bounded. For marginal and partial covariance denominators, it is an explicit increasing function of the corresponding first-stage statistic. For linear front-door adjustment, the Gaussian delta Wald statistic remains asymptotically standard normal when the mediator residual variance tends to zero at any rate, provided the treatment–mediator coefficient is nonzero.
Together, the results give a diagnostic pathway for graph-derived rational formulas. One factors the symbolic denominator, checks whether its factors form a nested variance flag, and then chooses between conventional studentization and weak-identification-robust inversion. Simulations and a diagnostic audit of right-heart-catheterization data from the SUPPORT study illustrate the two branches of this pathway.
1.3 Related work
Graphical criteria for instrumental sets, front-door adjustment and proximal identification determine whether a causal effect can be expressed from the observational law. The half-trek criterion and computer-algebra methods provide rational certificates for linear structural equation models (31; 6; 15; 14; 23; 39; 29; 21; 38). Generic and rational identifiability are distinct algebraic notions, and a parameter may be generically unique without being represented by the particular rational formula under study (9). Recent semiparametric work derives regular influence functions for half-trek estimators at fixed interior distributions (28). Our analysis starts after a rational certificate has been selected. It classifies the denominator as a polynomial on the ambient covariance space and determines which first-order boundary regime that certificate can generate.
Fieller and Anderson–Rubin inversion provide the classical ratio-based confidence constructions (1; 13). The impossibility of uniformly bounded confidence sets and the local theory of weak instruments and weak moments explain why a noisy denominator near zero produces non-Gaussian ratio limits (16; 11; 36; 37; 2). Modern work develops first-stage-dependent corrections and robust procedures, including settings with many weak moments and weakly identified nuisance functions (3; 24; 40; 4). The local result in Section 5 is a covariance-polynomial specialization of this literature. The new point is structural and precedes the limit calculation. The self-normalization identity decides when the first-order ratio regime is algebraically impossible.
Condition-number analyses quantify sensitivity of an identified functional or structural parameter to perturbations of the observational law (34; 19; 33). Exact self-normalization instead compares the gradient noise of one denominator with the denominator itself through a global polynomial identity. The two notions need not agree. In the front-door boundary sequence, the standard error diverges and the formula becomes badly conditioned, yet the denominator cannot enter the first-order Fieller regime. Conversely, a well-scaled slope denominator may have an interior zero with nondegenerate first-order noise.
Products of nested covariance minors are generalized power functions of a symmetric cone (12), and their exact sampling factorization follows from the Bartlett decomposition of a Wishart matrix (30). The affine-invariant metric on the positive-definite cone, which agrees up to scale with the Fisher–Rao metric of the centred Gaussian family, gives the geometric form of the logarithmic gradient flow used in the converse proofs (5). The separation between Gaussian and elliptical covariance operators is familiar in covariance-structure analysis (35; 22). Here it identifies which part of the classification is distributional and which part is algebraic.
The dimension-two converse uses the Gordan–Noether theorem for forms with vanishing Hessian (18; 27). That route fails for the six variables of a symmetric matrix of order three because noncone forms with vanishing Hessian exist in this dimension (32; 7; 17). Supplementary Section I.1 explains why Hessian degeneracy does not force a cone on the six-dimensional space , while Supplementary Section I.2 identifies the information lost when the relevant tensor relation is passed through polynomial multiplication. These two obstructions explain why neither route presently closes the mixed branch (25). This distinction explains the complete dimension-two result and the mixed-kernel problem left in dimension three.
2 Preliminaries
2.1 Basic notation
We use and for the real and complex fields. Vectors are columns, denotes matrix transpose, is the identity matrix of order , and is the th standard basis vector. The subscript on is omitted when the order is clear. The symbols , , and denote trace, determinant, rank and the multiset of eigenvalues of . Expectations, variances, covariances and probabilities under a probability measure are denoted by , , and . The subscript is suppressed once the probability measure is fixed. Convergence in probability and in distribution are denoted by and , and limits are as unless stated otherwise. The indicator of a statement is , and is the sign of .
For a positive integer , let , and . We write and denote its positive-definite and positive-semidefinite cones by and . The notation and means and , respectively. For , its entries are . For , is the submatrix with rows in and columns in , , and .
We order lexicographically, first by and then by , and set
Vectors and matrices indexed by follow this ordering. We identify with the ring of polynomial maps from to . The notation means that is the zero polynomial in this ring. A polynomial is homogeneous of degree when for every . Its total degree is , and two polynomials are coprime when their only common divisors are nonzero constants.
For , the coordinate gradient consists of the partial derivatives in the preceding order. The Fréchet differential of at in the direction is
The symmetric gradient is the unique matrix satisfying for every . Thus, and for , with all quantities evaluated at . For , the relative gradient is .
2.2 Problem set-up
We formulate the statistical problem in the ambient covariance space. Fix and . The notation denotes the Gaussian law on with mean and covariance matrix . Write and . Let be independent observations with joint law . We use the known-mean sample covariance matrix
| (1) |
For an integer and , denotes the law of for independent . Its scalar specialization with is . Consequently, and almost surely.
The identification equality is required only on a covariance model, whereas the polynomial pair representing it governs off-model sampling behaviour. The following definition separates these two objects.
Definition 1 (Identification strategy).
Let be a nonempty covariance model and let be a scalar target. An identification strategy for is an ordered pair of coprime elements of such that is not identically zero on and
Thus, at every point of at which is nonzero. The covariance plug-in estimator is on the event . The target is scale-invariant when whenever and .
This distinction is necessary for two reasons. First, need not satisfy the model equations, so the sampling behaviour of depends on and as polynomials on the whole of . Second, multiplying both and by a nonconstant common factor leaves the identified ratio unchanged wherever both representations are defined, yet alters the off-model sampling behaviour of the denominator. Coprimality removes this ambiguity. Two coprime pairs representing the same rational function coincide up to a common nonzero scalar, while distinct rational certificates for the same target remain distinct strategies.
2.3 Exact self-normalization
We now introduce the central property of the paper. Define the polynomial matrix map by
| (2) |
Under , the matrix is the covariance matrix of . For , define its Gaussian first-order variance functional by
| (3) |
The second equality is the Gaussian quadratic-form variance identity, verified in Supplementary Lemma B.1. Both sides are polynomial in , so the equality extends from to all of . For the matrix is positive semidefinite, and we write there. The standardized denominator is , defined on the event .
Definition 2 (Exact self-normalization).
A nonconstant homogeneous polynomial is exactly self-normalizing when there is a constant such that
| (4) |
For a strategy , exact self-normalization refers only to the denominator .
The constant is unique because is not the zero polynomial. Below, self-normalizing always means exactly self-normalizing. Equation (4) states that the Gaussian first-order noise of contracts in exact proportion to throughout the ambient covariance space. In particular, evaluating (4) at gives whenever . Thus, the standardized denominator of a self-normalizing polynomial is constant over samples rather than a relevance statistic. Exact self-normalization is invariant under congruence transformations. Every rational strategy for a scale-invariant coefficient in a linear structural equation model has a homogeneous representative, and a nonzero self-normalizing polynomial has no zero in . Supplementary Lemmas B.1–B.3 prove these facts. If depends on only through its compression to an -dimensional subspace, its self-normalization status and constant are unchanged in every ambient dimension by Supplementary Lemma B.4. The dimension-three classification below therefore applies to any denominator supported on at most three directions, whatever the number of observed coordinates.
Remark 1 (Scope of exactness).
Equation (4) is an identity in the ambient polynomial ring for the coprime denominator fixed by Definition 1. It is not merely an equality on a structural model. The word “exact” also refers to the Gaussian covariance operator in (3). A sandwich studentizer under a general law targets a different quadratic form and need not be constant sample by sample. Under elliptical sampling, the same algebraic class persists after a kurtosis adjustment, although the product-of-chi-squares pivot is generally lost. With an unknown Gaussian mean, the corresponding centred covariance formulas replace by . Supplementary Proposition B.1 and Remark B.1 give the precise statements.
3 Flag powers and an exact pivot
3.1 Nested covariance minors
We construct self-normalizing denominators in every dimension. The building blocks are covariance volumes of subspaces. For an -dimensional subspace , let be any full-row-rank matrix whose row space is , and write . If a random vector has covariance matrix , then is the generalized variance of the reduced vector . A different choice of multiplies by a positive constant and affects no statement below.
A flag is a strictly increasing chain of nonzero subspaces of . Let , so that , and let be positive integers. The associated flag power and its multiplicities are
| (5) |
Thus, is the total exponent carried by factors whose subspaces have dimension at least , and . The following theorem shows that every flag power is exactly self-normalizing and gives its exact sampling law.
Theorem 1 (Flag powers).
Let and be given by (5). The following statements hold.
- (i)
For every , the spectrum of the relative gradient is , where the eigenvalue zero has multiplicity .
- (ii)
Equation (4) holds with .
- (iii)
If , then is distributed as
(6) where the chi-squared variables are independent.
- (iv)
In particular, the identity holds almost surely, and no sequence can satisfy both and .
Remark 2 (Algebraic and sampling content).
Parts (i) and (ii) are polynomial identities and do not depend on a sampling law. The samplewise identity in part (iv) therefore holds at every sample covariance matrix at which is nonzero, whatever the data-generating law. Only the factorization in part (iii) uses Gaussian sampling. With an unknown Gaussian mean, parts (iii) and (iv) hold for the centred sample covariance with replaced by . Under elliptical sampling, flag powers remain proportionally self-normalizing after a kurtosis adjustment, but the product-of-chi-squares law is generally lost. Supplementary Proposition B.1 and Remark B.1 give these variants.
3.2 Recursive variance interpretation
Flag powers aggregate conditional-variance information. Let have the sampling law , and write for the successive conditional variances, with . For every ,
| (7) |
Consider the coordinate flag, whose th subspace is spanned by the first coordinate directions. By (7), the flag power (5) is then the monomial , so the multiplicity is the exponent attached to the th conditional variance. A general flag reduces to this case, up to a positive constant, after one fixed nonsingular linear transformation of the observation vector. For the front-door strategy of Section 1, the first two conditional variances are the treatment variance and the mediator residual variance. Its denominator is the monomial , and part (ii) of Theorem 1 returns the constant behind the value .
The same representation explains the exact sampling law. Under , the sample conditional variances computed from along the coordinate flag are mutually independent. The th is distributed as times a variable. This is the Bartlett decomposition of the Wishart matrix (30, Ch. 3), and evaluating the monomial at these independent factors gives the law in (6). To first order, the relative variance of every sample conditional variance is , whatever the value of . By independence, the relative first-order variance of the sample flag power is , which equals . This constancy in is the sampling mechanism behind exact self-normalization. The sample flag power depends on only through these conditional variances. The regression coefficients of each coordinate on its predecessors, which complete the recursive decomposition of , do not enter. This contrast between conditional variances and regression coefficients is the sample form of the distinction between variance information and slope information that organizes the converse results of Section 4.
4 Converse results and diagnosis
4.1 Dimension two
We first obtain a complete converse for covariance matrices of order two. Through face restriction, this result also drives the three-variable analysis.
Theorem 2 (Complete converse in dimension two).
Let be homogeneous, nonconstant and exactly self-normalizing. Then
| (8) |
for a nonzero constant , a nonzero vector , and nonnegative integers with . The self-normalization constant is .
Thus, a two-variable denominator can self-normalize only by combining one variance direction with the full covariance volume.
Remark 3 (What the converse excludes).
The permitted irreducible factors are one rank-one variance and the determinant. Every linear form whose coefficient matrix has rank two, such as the off-diagonal covariance , is excluded even though it has the same degree as the permitted variance forms. Thus, the theorem distinguishes variance information from slope information within the same polynomial degree.
We summarize the proof mechanism, which explains the rigidity. Euler’s identity and (4) fix the first two power sums of the eigenvalues of . The elementary symmetric identity converts them into a determinant identity with an explicit constant coefficient. When that coefficient is nonzero, the determinant divides , the factor can be stripped, and induction on the degree applies. When it vanishes, the gradient map takes values in the rank-one quadric, and the Hessian determinant of vanishes identically. The Gordan–Noether theorem reduces to a power of one linear form, which symmetry and reality identify with a variance direction (18; 27). Supplementary Section D.2 gives the full proof.
4.2 Dimension three
We now reduce the three-variable converse to a single mixed branch. Put and call determinant-free when . For homogeneous of degree , define the rank-one restriction by . For a plane with basis matrix , define the face restriction by . The plane is called -good when , and simply good when the polynomial is clear from context. A polynomial vector is primitive when its entries have no nonconstant common divisor.
A nonzero plane restriction of a self-normalizing polynomial is again self-normalizing with the same constant. Supplementary Lemma D.3 provides this restriction principle, and Theorem 2 then determines the form of every such restriction.
Definition 3 (Face type).
Let be a determinant-free self-normalizing polynomial on , of degree and with self-normalization constant . For , we say that has face type if, for every -good plane , there exist and such that, as a polynomial identity on ,
| (9) |
A determinant-free solution has at least one good plane, and its degree and constant admit at most one pair satisfying (9), so the face type exists and is common to all good planes. Removing a factor of preserves exact self-normalization with a shifted constant, so the residue in the next theorem is itself self-normalizing. Supplementary Lemmas D.4 and D.6 provide both facts.
Theorem 3 (Three-variable reduction).
Let be homogeneous, nonconstant and exactly self-normalizing, and let be the determinant-free residue obtained by removing from the largest power of . Exactly one of the following cases occurs.
- (i)
If , then for a nonzero constant and a nonzero vector .
- (ii)
If and the face type of is , then for a nonzero constant and a plane .
- (iii)
If and the face type of is with , then
(10) , and a primitive polynomial vector satisfies , so that spans the kernel of at every point where has rank two and .
The polynomials in cases (i) and (ii) are flag powers built on a single subspace. A mixed type has degree , so every solution of degree at most two falls under the first two cases and is a flag power. Therefore, the unresolved solutions lie in case (iii) and carry one possibly moving null direction together with the fixed spectrum in (10).
Remark 4 (The mixed-kernel problem).
The unrestricted three-variable converse is equivalent to the constancy of the kernel line in Theorem 3 (iii). For a nonzero vector , write for the line that it spans. If is constant on one nonempty open set, then a fixed vector satisfies , and the two-variable converse yields
where and the rows of span . The equivalence between constancy of the kernel line and the displayed mixed flag form is established in Supplementary Lemma D.11. Supplementary Proposition D.3 gives the cofactor syzygy that couples a moving kernel line to the determinant boundary. Vanishing of the six-variable Hessian alone is insufficient because noncone forms with vanishing Hessian exist in this dimension (32; 7; 17).
Supplementary Section D also gives an independent closure of the mixed branch that does not presuppose a constant kernel line. If the restrictions of the determinant-free mixed residue to rank-two faces agree with those of one fixed flag power built from a nested line and plane, the residue is that flag power whenever its mixed type has degree at most eight. For higher-degree types, the same conclusion holds under this boundary agreement condition outside an explicit arithmetic resonance set. Therefore, the unresolved difficulty is the geometric gluing of the face directions rather than an uncontrolled interior perturbation.
4.3 Higher-dimensional scope
The constancy of the relative spectrum that drives the three-variable reduction is special to dimension three. Euler’s identity fixes the trace of at the degree of , and (4) fixes the trace of its square, so in dimension three the eigenvalues are confined to a circle. The polynomial character of forces at least one further linear relation with integer coefficients among the eigenvalues. A line meets a circle in at most two points, so only finitely many spectra are possible, and continuity on the connected cone makes the spectrum constant. This finiteness underlies the fixed spectrum displayed in (10). In dimension the same two traces leave a sphere of dimension , a linear relation cuts out a set of dimension , and no finiteness follows.
A conditional converse nevertheless holds in every dimension. If the relative gradients preserve one fixed complete flag for every and induce constant weights on the successive one-dimensional quotients, then is the corresponding flag power. This is an intrinsic condition on the gradient rather than an assumed factorization, and the precise statement is Supplementary Proposition D.5. Exact self-normalization alone is not claimed to create such an invariant flag when . The obstruction in higher dimensions is thus the emergence of one common invariant recursive ordering, not the integration step once that ordering is present, and dimension three is a sharp low-dimensional target rather than an arbitrary truncation.
4.4 A factor-and-rank diagnostic
Although the mixed branch obstructs an unconditional converse, the denominators produced by graphical identification typically have a simple factor structure. The two denominators of Section 1 factor into linear and quadratic pieces after determinant stripping, and the same holds for the formulas revisited in Section 6. We now show that this class is classified completely and that membership can be decided by exact symbolic computation.
Definition 4 (Low-degree-factor condition).
A homogeneous polynomial on satisfies the low-degree-factor condition if, after removing the largest power of , every real irreducible factor has degree at most two.
Under this condition the mixed branch disappears and the classification closes.
Theorem 4 (Low-degree-factor converse).
Let be homogeneous, nonconstant and exactly self-normalizing, and suppose that satisfies the low-degree-factor condition. Then there exist a nested line and plane , a nonzero constant , and nonnegative integers such that
| (11) |
Therefore, linear and quadratic factors cannot assemble into any self-normalizing configuration other than a nested flag.
For a square matrix , denotes its classical adjugate, characterized by . Writing the plane minor through the adjugate turns Theorem 4 into a normal form whose ingredients can be tested one by one.
Corollary 1 (Factor-and-rank diagnostic).
Let be homogeneous and nonconstant, and suppose that satisfies the low-degree-factor condition. Then is exactly self-normalizing if and only if, up to a nonzero constant,
| (12) |
for nonzero vectors and nonnegative integers . In that case the self-normalization constant is
| (13) |
Consequently, every reducible exactly self-normalizing cubic is one of , with , or .
This corollary reduces the audit of a candidate denominator to a finite procedure. One factors the polynomial symbolically, tests the ranks of the coefficient matrices of the linear and adjugate-linear factors, and checks orthogonality and nesting of the resulting directions. The globalization of the generic face factors and the resulting coefficient-level characterization are proved in Supplementary Section D.8.
The two formulas of Section 1 illustrate the procedure.
Example 1 (From a graph formula to a rank check).
For the coordinate order , let denote the corresponding standard basis vectors. The front-door denominator factors as
| (14) |
Both coefficient matrices have rank one and , so Corollary 1 applies with and certifies exact self-normalization with , in agreement with Section 3. For the coordinate order , define analogously. The proximal denominator is
| (15) |
where , so the quadratic factor fails the rank-one test. Indeed, vanishes at the identity matrix, so exact self-normalization already fails by the interior nonvanishing property recorded in Section 2.
Thus, the audit separates the two formulas before any data are collected.
Remark 5 (How the diagnostic should be used).
For a denominator with rational or algebraic coefficients, the factorization and the rank checks are exact symbolic operations, and the diagnostic is complete on the stated class. A near-factorization obtained numerically does not certify self-normalization. An irreducible factor of degree at least three renders the criterion inconclusive rather than negative. One may then test (4) directly by comparing coefficients or apply the kernel and boundary criteria of Supplementary Section D.
5 Weak-denominator inference
5.1 Local ratio experiment
We formulate the first-order regime of a weak denominator whose sampling noise does not degenerate with it. Throughout this section, is a fixed identification strategy in the sense of Definition 1, and is its plug-in estimator. For a deterministic sequence with , define
and put and . Whenever , define
| (16) |
These three quantities are the standardized drift of the denominator, the correlation between the moment direction and the denominator gradient, and their scale ratio.
Assumption 1 (First-order weak denominator).
There are , and such that
Define
and assume and .
By continuity of the polynomial gradients and of , Assumption 1 implies , and
Assumption 1 describes a denominator that drifts to zero at exactly the rate of its nondegenerate first-order noise. Continuity gives , and the finite limit of then forces , so the sequence approaches a zero of the denominator, which may lie in the interior of the cone or on its boundary. The condition makes this zero first-order nondegenerate, and the condition excludes a simultaneous first-order degeneracy of the moment function itself. The limit records how many first-order standard errors separate the drifting denominator from zero, and it acts as the noncentrality parameter of the limit experiment.
The associated Wald procedure studentizes the plug-in estimator along the estimated moment direction. On the event where and the quadratic form below is positive, define
| (17) |
The next proposition identifies the joint limits.
Proposition 1 (Local ratio experiment).
Suppose that Assumption 1 holds. Let and be independent standard normal variables and put . Then
| (18) |
If in addition , then
| (19) |
Thus, a first-order nondegenerate small denominator produces the classical Fieller ratio limit. The limiting standardized denominator is a noncentral chi-squared variable with one degree of freedom and noncentrality , so measures the local strength of identification. The Wald statistic converges to a non-Gaussian law governed by and . Supplementary Section E proves Proposition 1.
Remark 6 (Role of the local experiment).
Proposition 1 is not offered as a new general theory of weak identification. Its role is to place every covariance-polynomial strategy in a common three-parameter local experiment and to make the algebraic classification operational. Once is known not to self-normalize, the limits , and determine the leading ratio geometry. When is a flag power, the assumptions of this experiment cannot hold on the denominator side.
5.2 Robust inversion and the boundary of the classification
We connect the ratio experiment to robust inference. For , inverting the single moment gives the confidence set
| (20) |
where
is the quantile of the law.
Proposition 2 (Fieller confidence set).
Suppose that Assumption 1 holds. Then . Moreover, for every fixed , with probability one the set is nonempty and is an interval, the complement of a bounded open interval, or the whole real line, and it is bounded exactly when .
Supplementary Section E proves Proposition 2. This is the classical Fieller and Anderson–Rubin analysis specialized to covariance polynomials (1; 13). It gives the standardized denominator a second role. Proposition 1 identifies as the noncentrality diagnostic of the local experiment, and Proposition 2 makes the same statistic decide whether the inverted set is bounded. Combining the two results, the probability that is bounded converges to , so unbounded sets retain a positive limiting frequency under weak drift, in accordance with the impossibility results for uniformly bounded confidence sets (16; 11).
We now delimit what exact self-normalization removes. As noted after Assumption 1, the drift forces while requiring . For a self-normalizing denominator, evaluating (4) at gives , so the two requirements are incompatible and no drifting sequence satisfies Assumption 1. Therefore, exact self-normalization excludes the first-order ratio experiment, not every nonregular limit.
The exclusion concerns the first order only. On , the polynomial has , so its first-order variance vanishes on the entire zero set and Assumption 1 cannot hold, although the ratio is not constant. Under the drift with , the standardized denominator converges in distribution to for a standard normal variable , a nondegenerate limit generated at second order. Supplementary Section E verifies this example. Hence, the self-normalizing and the first-order nondegenerate regimes are two extremes of a broader classification rather than an exhaustive dichotomy.
6 Causal applications
6.1 Diagnostics for common formulas
We translate the algebra into familiar covariance diagnostics. For distinct indices , put and let be the same function evaluated at . For the marginal covariance , the Gaussian first-order variance is , so every zero of in the open cone is first-order nondegenerate and the denominator conditions of Assumption 1 can hold along suitable drifting sequences. The standardized denominator is , an increasing function of the familiar first-stage statistic .
For indices with and , allowing , define as the correlation of the residuals from the population linear projections of coordinates and on coordinate , and let be its sample-covariance analogue. For the partial minor ,
| (21) |
When the minor is a principal minor, the identity reduces to , and in agreement with Theorem 1. When , define
so the standardized denominator is again an increasing transform of the familiar partial first-stage index . In both cases the standardized denominator reproduces the relevant marginal or partial screening direction without any modelling input. Supplementary Proposition F.1 gives the calculation and the exact partial-correlation formula.
We compare four representative strategies in Table 1. The table separates variance-type flag powers from slope-type marginal or nonprincipal minors. The structural parameterizations and substitutions yielding the third column are given in Supplementary Section F.1, so that each displayed pullback can be checked directly from the stated linear reduced form. The three slope-type rows can realize the denominator conditions of Assumption 1, whereas the front-door row is a flag power. The next subsection develops the front-door strategy in detail.
| Strategy | Denominator | Structural pullback | Classification |
| Instrumental variable | Marginal slope | ||
| Conditional instrument | Partial slope | ||
| Front-door | Flag power | ||
| Linear proximal | Partial slope |
6.2 Front-door adjustment and boundary-robust inference
We now show that the front-door studentized statistic survives arbitrary collinearity drift when the treatment–mediator coefficient is nonzero. The exact regression decomposition and the proof of the boundary result are given in Supplementary Section F.3. Consider the Gaussian reduced form
| (22) |
in which , and are independent centred Gaussian variables with and , and the observed vector is . The coefficients , and are fixed, while the mediator residual variance may change with and takes values in a fixed interval . Observations are independent copies of the observed vector, and is the known-mean sample covariance (1). Substituting the reduced form gives and , so the population denominator equals , which is the pullback recorded in Table 1 and which tends to zero whenever does.
The front-door covariance functional is
| (23) |
on the set where its denominator is nonzero. Direct substitution shows that for every value of , so the functional isolates the product and removes the association carried by .
Example 2 (Exact front-door denominator).
In the front-door causal model, the coefficient is generated by an unobserved confounder of and rather than by a direct effect, so the total effect of on equals and is identified by . The denominator is the flag power identified in Section 3, and Theorem 1 gives
| (24) |
The last equality holds for every positive-definite sample covariance when the Gaussian studentizer is used, whatever the data-generating law.
Thus a relevance test based on the front-door denominator is exactly uninformative even though the precision of can deteriorate without bound.
Because is exactly self-normalizing, no drifting sequence satisfies Assumption 1 for this strategy, and the ratio experiment of Section 5 is unavailable as a description of the boundary . That exclusion is negative information only. The following result provides the positive counterpart for the studentized estimator. The plug-in estimator is , and its Gaussian delta standard error is
| (25) |
Both quantities are defined almost surely because implies .
Proposition 3 (Boundary-robust front-door Wald inference).
Thus, arbitrarily poor mediator residual variation inflates uncertainty but does not invalidate the studentized Gaussian limit on the stated parameter region.
Remark 7 (Why studentization survives).
The covariance estimator factors exactly as the product of the slope in the regression of on and the coefficient of in the regression of on . Conditional on the second regression design, the latter coefficient has an exact Student statistic. As decreases, the first slope is estimated more accurately while the second standard error grows at the reciprocal rate. The Gaussian delta studentizer reproduces this balance exactly, so the exploding variance affects precision but not the limiting studentized law.
The restriction in Proposition 3 is deliberate. The proposition covers when , but it makes no claim on the stratum , whether or not vanishes. At the doubly singular point the numerator gradient vanishes, the estimator has a product-of-normals limit at rate , and the nonregularity is a singular-hypothesis phenomenon of the numerator rather than a weak-denominator phenomenon (10). The exact denominator identity in (24) remains true throughout these parameter regions.
7 Numerical experiments
7.1 Design
Two Monte Carlo experiments isolate the two denominator mechanisms of Table 1, one flag power and one partial slope. The front-door design sets , and , where is centred Gaussian with unit marginal variances and covariance , and is an independent Gaussian variable with variance ; the target is the total effect . The proximal design sets , , and , with mutually independent centred Gaussian shocks of unit variance except ; the target is the coefficient of , equal to one. By the structural pullbacks in Table 1, the population denominators are and , so the grids and trace the approach to the two denominator boundaries. The population value of equals identically in the front-door design and falls from to across the proximal grid (Supplementary Table G.1).
For each cell, we draw the known-mean sample covariance directly from , with and respectively, which is distributionally equivalent to simulating centred Gaussian observations. The converse classification of Section 4 is not invoked at : Theorem 1, Proposition 1 and the calibrations of Section 6 hold in every dimension, and the proximal denominator depends only on the block, so its slope-type status is unchanged by the ambient dimension (Supplementary Lemma B.4). We use and replications per cell. Each replication yields the plug-in estimator , the standardized denominator , and, in the proximal design, the partial statistic of Section 6 with together with the deliberately naive marginal statistic , which measures only the treatment–proxy association. Cells are summarized by medians and by the interquartile range of divided by , because the ratio estimator need not possess moments in the weak cells, and by the empirical coverage of nominal 95% Wald intervals based on the delta standard error in (17) and of the inversion (20). The estimator, both intervals and all diagnostics were well defined in every replication. The Python scripts reproducing both the numerical experiments and the real data experiments in Section 8 are available at https://github.com/shutech2001/self-normalizing-denominators-experiments.
The theory yields three predictions. First, Theorem 1 (iv) and Example 2 imply that in every front-door replication, and Proposition 3 implies nominal Wald coverage along the whole grid, whose smallest cell has and so lies inside the boundary regime . Secondly, the weak proximal cells have population values of below one, matching the drift regime of Assumption 1, so Proposition 1 predicts the noncentral limit for in (18) and distorted Wald inference. Thirdly, Proposition 2 predicts near-nominal coverage for the inversion (20) in every cell, with sets that are bounded exactly when exceeds the quantile of the law.
7.2 Results
Table 2 reports the results. In the front-door design the median standardized denominator equals the theoretical constant in every cell; by Theorem 1 (iv) the statistic equals in every replication, so it carries no information about . Precision, by contrast, deteriorates by a factor of eighteen: the robust standard deviation of rises from to as falls from to , and the median delta standard error tracks it closely, rising from to . Wald coverage lies between and in all four cells, including the boundary cell with , and inversion coverage lies between and , slightly above the nominal level. This confirms the first prediction: the studentized statistic remains accurate while the denominator diagnostic is exactly uninformative.
In the proximal design the naive marginal statistic rises from to as decreases, while the partial statistic falls from to and the median falls from to ; the latter two columns accord with the identity between and stated after (21). At and the median bias reaches and and Wald coverage falls to and , although the naive diagnostic is largest exactly there. In those two cells the median standard error exceeds the robust standard deviation, so the Wald failure is one of centring and distributional shape rather than of scale, as the ratio limit in (18) predicts. Inversion coverage remains between and throughout, but the protection has a price: in the two weakest cells the median lies below , so more than half of the inverted sets are unbounded.
| Parameter | Med. | Med. partial | Med. naive | Robust s.d. | Med. s.e. | Med. bias | Wald/inv. cov. |
| Front-door: parameter | |||||||
| – | – | / | |||||
| – | – | / | |||||
| – | – | / | |||||
| – | – | / | |||||
| Proximal: parameter | |||||||
| / | |||||||
| / | |||||||
| / | |||||||
| / | |||||||
| / | |||||||
The two designs estimate different targets in different models, so the experiment compares diagnostics rather than estimators. Its message is the dissociation predicted by the classification: front-door adjustment loses precision without entering the first-order ratio regime, whereas the proximal estimator fails in the partial-covariance direction that the marginal treatment–proxy statistic does not measure. Supplementary Section G gives complete cell summaries.
8 Real data experiments
8.1 Data and diagnostic design
We audit the public SUPPORT right-heart-catheterization data (8), which contain critically ill patients, an indicator of right-heart catheterization on the first study day, survival time in days truncated at , and baseline physiological measurements. Following proximal analyses of these data (26; 38), we consider treatment proxies pafi1 and paco21 and outcome proxies ph1 and hema1. The analysis is the diagnostic pathway described in the introduction rather than a new clinical causal analysis: the symbolic factor-and-rank step has already classified the proximal denominator as a nonprincipal slope minor (Example 1), and the data-level step measures the implied partial direction for each candidate proxy pair.
Treatment, outcome and the four proxies are residualized by least squares on age, sex, primary and secondary disease categories, do-not-resuscitate status, the SUPPORT two-month survival estimate and the APACHE score; missing continuous covariates are median-imputed, and missing categorical covariates are mode-imputed and coded by indicator variables with one level omitted. The resulting design has rank , and all diagnostics use the residual degrees of freedom in place of . For residualized treatment , treatment proxy and outcome proxy , we report the naive marginal statistic , the partial statistic and two versions of the standardized denominator: the Gaussian version of Section 6, and a sandwich version in which in (3) is replaced by the empirical covariance matrix of the six second-moment scores. Because the original treatment is binary, the residualized treatment is non-Gaussian. The Gaussian fourth-moment formula therefore serves only as a reference value. The sandwich version is the appropriate studentization under the empirical law (Remark 1). Sampling variability of the proximal estimates and of the sandwich statistic is assessed by nonparametric percentile bootstrap replications over patients, each of which repeats the imputation, coding and residualization. The audit is descriptive: proxy strength in the denominator direction is measurable, but proxy validity is not testable from these diagnostics.
8.2 Results
Table 3 reports the audit, and the two pairs involving ph1 reverse their ordering across diagnostics. The pair pafi1/ph1 appears strong in the treatment–proxy direction, with naive statistic , but is weak in the direction that enters the denominator. Its partial statistic is and its sandwich statistic is , with a bootstrap interval reaching essentially zero. The pair paco21/ph1 shows the reverse pattern, with naive statistic but partial statistic and sandwich statistic , bounded well away from zero. The Gaussian version is the monotone transform of the partial statistic given after (21) with in place of , which explains why those two columns nearly coincide when the partial statistic is small relative to . The sandwich column carries the additional distributional information. It is smaller than the Gaussian value for every pair, by factors of to for three pairs and for paco21/ph1. Thus, the Gaussian formula overstates denominator strength under the empirical fourth moments, most severely for the strongest pair.
| Treatment proxy | Outcome proxy | Naive | Partial | Gaussian | Sandwich (95% CI) |
| pafi1 | ph1 | (, ) | |||
| pafi1 | hema1 | (, ) | |||
| paco21 | ph1 | (, ) | |||
| paco21 | hema1 | (, ) |
The plug-in proximal estimates for the four pairs, in the order of Table 3, are , , and days, with percentile intervals , , and . Only the interval for the weak pair pafi1/ph1 covers zero. For that pair, however, the bootstrap interval inherits the nonstandard ratio behaviour described in Section 5 and is reported for completeness only; the inversion (20) is the appropriate construction in that regime. None of these intervals validates the proximal bridge assumptions, which the audit cannot test.
The audit illustrates the practical content of the classification. A strong association between treatment and a proposed treatment proxy neither implies nor precludes strength of the nonprincipal minor that identifies the proximal bridge; the symbolic step locates the relevant covariance direction before any estimation, and the partial and sandwich statistics then measure it for each candidate allocation.
9 Discussion
The classification turns the choice among rational identification formulas into a design decision. When several certificates identify the same target, their denominators can belong to different algebraic classes even though the ratios agree on the model. The classes are complementary rather than ordered, because a flag denominator removes first-order denominator nonregularity while a slope-type formula may be more precise at strongly identified distributions. Therefore, we recommend reporting the symbolic denominator class alongside any covariance plug-in estimate. When a denominator is not self-normalizing and its relevant standardized diagnostic is weak, inference should also include a robust inversion.
The main open algebraic problem is the rigidity of the mixed kernel line. A proof that the kernel line in Theorem 3 (iii) is constant would complete the unconditional three-variable converse, while a counterexample would produce a self-normalizing denominator outside the flag class. Either resolution must exploit the integrability of the relative gradient beyond its vanishing Hessian. In dimension at least four, a converse would additionally require invariants beyond the first two power sums of the relative gradient.
The principal distributional question is self-normalization beyond the Gaussian covariance operator. Ellipticity only rescales the constant, so the first open case is a semiparametric family with unrestricted fourth cumulants. A characterization of polynomials whose sandwich variance is proportional to their square over such a family would determine when the standardized denominator remains a samplewise constant under a matched studentizer. This, in turn, would establish how far the exact screening interpretation extends.
The principal methodological question concerns selection among many candidate strategies. Choosing a certificate because its observed denominator diagnostic is largest reintroduces a data-adaptive weak-direction problem. Therefore, applications with many proxies or graph-derived formulas call for selection-adjusted denominator diagnostics and simultaneous robust inversions (40; 4). Because the factor-and-rank audit is symbolic, it could also be attached to computer-algebra identification pipelines, so that every certificate is generated together with its denominator class.
Declaration of the use of generative AI and AI-assisted technologies
During the preparation of this work the author used ChatGPT and Claude in order to assist with writing and refactoring the simulation code. After using these tools the author reviewed and edited the content as necessary and takes full responsibility for the content of the publication.
Acknowledgement
Shu Tamano was supported by JSPS KAKENHI Grant Numbers 25K24203.
Supplementary material
The Supplementary Material includes a notation table, all omitted proofs, the elliptical and unknown-mean extensions, the recursive-flag and boundary-rigidity converses, the Fieller and minor-calibration results, detailed simulation and SUPPORT analyses, and the algebraic obstructions to an unrestricted three-variable converse.
References
- Estimation of the parameters of a single equation in a complete system of stochastic equations. The Annals of Mathematical Statistics 20 (1), pp. 46–63. Cited by: §1.3, §5.2, §E.
- Estimation and inference with weak, semi-strong, and strong identification. Econometrica 80 (5), pp. 2153–2211. Cited by: §1.1, §1.3.
- Weak instruments in instrumental variables regression: theory and practice. Annual Review of Economics 11, pp. 727–753. Cited by: §1.1, §1.3.
- Inference on strongly identified functionals of weakly identified functions. Journal of the Royal Statistical Society Series B: Statistical Methodology 88 (3), pp. 998–1028. External Links: Document Cited by: §1.3, §9.
- Positive definite matrices. Princeton University Press. Cited by: §1.3.
- Generalized instrumental variables. In Proceedings of the 18th Conference on Uncertainty in Artificial Intelligence, pp. 85–93. Cited by: §1.1, §1.3.
- Homaloidal hypersurfaces and hypersurfaces with vanishing hessian. Advances in Mathematics 218 (6), pp. 1759–1805. Cited by: §1.3, §I.1, Remark 4.
- The effectiveness of right heart catheterization in the initial care of critically ill patients. The Journal of the American Medical Association 276 (11), pp. 889–897. External Links: Document Cited by: §8.1, §H.1.
- Generic identifiability of linear structural equation models by ancestor decomposition. Scandinavian Journal of Statistics 43 (4), pp. 1035–1045. Cited by: §1.3.
- Wald tests of singular hypotheses. Bernoulli 22 (1), pp. 38–59. Cited by: §1.1, §6.2.
- Some impossibility theorems in econometrics with applications to structural and dynamic models. Econometrica 65 (6), pp. 1365–1387. Cited by: §1.1, §1.3, §5.2.
- Analysis on symmetric cones. Oxford University Press. Cited by: §1.3.
- Some problems in interval estimation. Journal of the Royal Statistical Society Series B: Statistical Methodology 16 (2), pp. 175–185. Cited by: §1.3, §5.2, §E.
- Half-trek criterion for generic identifiability of linear structural equation models. The Annals of Statistics 40 (3), pp. 1682–1713. Cited by: §1.1, §1.3.
- Identifying causal effects with computer algebra. In Proceedings of the 26th Conference on Uncertainty in Artificial Intelligence, pp. 193–200. Cited by: §1.3.
- The nonexistence of confidence sets of finite expected diameter in errors-in-variables and related models. The Annals of Statistics 15 (4), pp. 1351–1362. Cited by: §1.1, §1.3, §5.2.
- On cubic hypersurfaces with vanishing hessian. Journal of Pure and Applied Algebra 219 (4), pp. 779–806. Cited by: §1.3, §I.1, Remark 4.
- Ueber die algebraischen formen, deren hesse’sche determinante identisch verschwindet. Mathematische Annalen 10, pp. 547–568. Cited by: §1.3, §4.1.
- Condition number bounds for causal inference. In Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence, Proceedings of Machine Learning Research, Vol. 161, pp. 1948–1957. Cited by: §1.3.
- Algebraic geometry: a first course. Graduate Texts in Mathematics, Vol. 133, Springer. Cited by: §D.3.
- Graphical tools for selecting conditional instrumental sets. Biometrika 111 (3), pp. 771–788. External Links: Document Cited by: §1.1, §1.3.
- Asymptotic distributions of functions of a sample covariance matrix under the elliptical distribution. The Canadian Journal of Statistics 22 (2), pp. 273–283. Cited by: §1.3.
- Measurement bias and effect restoration in causal inference. Biometrika 101 (2), pp. 423–437. External Links: Document Cited by: §1.1, §1.3.
- Valid -ratio inference for IV. American Economic Review 112 (10), pp. 3260–3290. External Links: Document Cited by: §1.3.
- A system of quadrics describing the orbit of the highest weight vector. Proceedings of the American Mathematical Society 84 (4), pp. 605–608. Cited by: §1.3, §I.2.
- Regression-based proximal causal inference. American Journal of Epidemiology 194 (7), pp. 2030–2036. External Links: Document Cited by: §8.1.
- When does the hessian determinant vanish identically? on gordan and noether’s proof of hesse’s claim. Bulletin of the Brazilian Mathematical Society, New Series 35, pp. 71–82. Cited by: §1.3, §4.1, §I.1.
- Semiparametric inference for half-trek estimators in linear structural equation models. arXiv preprint arXiv:2606.26931. Cited by: §1.3.
- Identifying causal effects with proxy variables of an unmeasured confounder. Biometrika 105 (4), pp. 987–993. External Links: Document Cited by: §1.1, §1.3.
- Aspects of multivariate statistical theory. Wiley, New York. Cited by: §1.3, §3.2, §C, §E, §F.3.
- Causal diagrams for empirical research. Biometrika 82 (4), pp. 669–688. Cited by: §1.1, §1.3.
- Sulle varietà cubiche la cui hessiana svanisce identicamente. Giornale Di Matematiche Di Battaglini 38, pp. 337–354. Cited by: §1.3, §I.1, Remark 4.
- Robust identifiability in linear structural equation models of causal inference. In Proceedings of the 38th Conference on Uncertainty in Artificial Intelligence, Proceedings of Machine Learning Research, Vol. 180, pp. 1728–1737. Cited by: §1.3.
- Stability of causal inference. In Proceedings of the 32nd Conference on Uncertainty in Artificial Intelligence, pp. 666–675. Cited by: §1.3.
- Analysis of covariance structures under elliptical distributions. Journal of the American Statistical Association 82 (400), pp. 1092–1097. External Links: Document Cited by: §1.3.
- Instrumental variables regression with weak instruments. Econometrica 65 (3), pp. 557–586. Cited by: §1.1, §1.3.
- GMM with weak identification. Econometrica 68 (5), pp. 1055–1096. Cited by: §1.1, §1.3.
- An introduction to proximal causal inference. Statistical Science 39 (3), pp. 375–390. External Links: Document Cited by: §1.1, §1.3, §8.1.
- Efficiently finding conditional instruments for causal inference. In Proceedings of the 24th International Joint Conference on Artificial Intelligence, pp. 3243–3249. Cited by: §1.1, §1.3.
- GMM with many weak moment conditions and nuisance parameters: general theory and applications to causal inference. arXiv preprint arXiv:2505.07295. Cited by: §1.3, §9.
Supplementary Material for
“Self-normalizing denominators in rational causal estimation”
A Notation
A.1 Linear-algebraic, algebraic and geometric conventions
We collect the conventions needed to read the Supplementary Material independently. Fix , write , put and , and use the lexicographic order . Vectors are columns, is transpose, for nonsingular , is the identity matrix and is the th standard basis vector. The symbols , , and denote trace, determinant, rank and the eigenvalue multiset. For , is the corresponding submatrix, and .
The space has positive-definite and positive-semidefinite cones and . For , stacks the upper-triangular entries in the stated order. The coordinate ring is identified with the polynomial maps ; means equality in this ring, is total degree, and homogeneity, divisibility and coprimality are understood in the indicated coordinate ring. For , is the coordinate gradient in the order, and
define the Fréchet differential and symmetric gradient. The relative gradient is on .
Throughout the remaining algebraic conventions, denotes either or . For a real vector space , its complexification is . The notation denotes the group of nonsingular matrices over , is the linear span of , and , and are the kernel, range and row space of a matrix. When the field is clear, the subscript on is omitted. For a real subspace , , and is the diagonal matrix with the displayed entries. The matrix unit has a one in position and zeros elsewhere.
For an integral domain , is its field of fractions. If is a unique factorization domain, abbreviated UFD, then denotes a greatest common divisor, defined up to multiplication by a unit. A vector with entries in is primitive when the greatest common divisor of its entries is a unit. For a finite-dimensional -vector space and , is the th symmetric tensor power,
where the angle brackets denote linear span, is the symmetric group on , and . Its dual is naturally the space of homogeneous polynomial functions of degree on . After choosing the standard basis, is identified with the space of symmetric matrices over .
For a finite-dimensional -vector space , is its projective space, and is the point represented by . We write and for the Grassmannian of -dimensional linear subspaces of ; thus . More generally, a subscript on a real algebraic variety or construction denotes extension of scalars from to , equivalently the complex variety defined by the same real polynomial equations. A dashed arrow denotes a rational map: it is represented by a morphism on a Zariski-dense open subset of , and two representatives are identified when they agree on a Zariski-dense open subset. When is irreducible, every nonempty Zariski-open subset is dense. For a homogeneous polynomial , the notation denotes the reduced projective hypersurface defined by the radical ideal ; it has the same underlying zero set as but carries no multiplicities. Terms such as Zariski open, closed, dense and irreducible refer to this algebraic topology over the field stated in the argument.
For a differentiable map into a finite-dimensional vector space and , its directional derivative is
For scalar , this equals ; for matrix-valued it is taken entrywise. Ordinary symbols such as and denote coordinate derivatives. The symbol means the coordinate Hessian of a scalar function on the affine space . When a Riemannian metric is explicitly fixed, , and denote its gradient, Levi–Civita connection and Riemannian Hessian, and is the induced norm.
For a square matrix , is its classical adjugate, so , where has the same order as . For a polynomial on , the adjugate transform is . If has dimension , is any full-row-rank matrix with row space and . If has range , then is the face restriction and . For a polynomial on , a plane is called -good when ; when is clear, such a plane is called good. If finitely many polynomials are under discussion, a generic good plane means a plane in a common nonempty Zariski-open set on which all indicated restrictions are nonzero and retain their generic factorization type. On we use and .
For later representation-theoretic notation, let be the vector space of all matrices. Its derived congruence action on polynomial functions on is
With the preceding matrix units, .
A.2 Probability and asymptotic conventions
For a probability law , , , , and denote expectation, variance, covariance, correlation and probability whenever the relevant moments exist. Conditional versions are under the same governing law. For centred , the fourth joint cumulant is
For random elements , is the smallest -algebra with respect to which all are measurable. It is not automatically completed; as usual, conditional expectations are defined up to -null sets.
The law is Gaussian with mean and covariance . For integer , is the law of for independent ; , and is Student’s law with degrees of freedom. For Gaussian sampling, and , with
The Gaussian covariance map and the associated first-order variance are
and is the nonnegative square root. Under a general law for a centred vector , whenever the fourth moments exist. In a triangular array with row covariance , the row law is , and we abbreviate and .
For an identification strategy , the plug-in estimator is whenever . Along a deterministic sequence with , put
and . Whenever ,
Equality in distribution is , convergence in probability is , and convergence in distribution is or . The function is the sign of . For positive deterministic , means that is bounded in probability, and means . Unless stated otherwise, all asymptotic statements use the current sequence of laws and let .
A.3 Paper-specific symbols
The remaining notation is summarized in Table A.1. An omitted matrix argument means evaluation at the current base point .
| Symbol | Meaning |
| , | Gaussian covariance map and the corresponding covariance map under a general law |
| , | and its nonnegative square root |
| Relative gradient on | |
| , | Known-mean sample covariance and its entries |
| on | |
| , | Three-variable face type and primitive polynomial kernel vector |
| , , | Local denominator mean, numerator–denominator correlation and scale ratio |
| Hat, subscript , subscript | Sample or plug-in quantity, row of a sequence, and limiting value |
B Proofs of Section 2: Preliminaries
This section proves the facts about exact self-normalization that the main text invokes without proof. Lemma B.1 identifies the Gaussian first-order variance with the trace form in (3) and shows that the property is preserved under congruence transformations. Proposition B.1 and Remark B.1 then delimit its distributional scope under elliptical sampling and unknown means. Lemma B.2 reduces every converse question to one homogeneous degree, Lemma B.3 establishes interior nonvanishing, and Lemma B.4 shows that neither the property nor its constant depends on the ambient dimension.
The first lemma converts the statistical definition of into the intrinsic trace form used throughout the paper.
Lemma B.1 (Trace form and congruence invariance).
For every and every ,
If is exactly self-normalizing with constant and is nonsingular, then is exactly self-normalizing with the same constant.
Both sides of the display are polynomial in , which yields the extension to all of recorded after (3). The congruence part makes the property coordinate-free, and it is used silently whenever a flag or a kernel direction is normalized by a linear change of variables.
Proof of Lemma B.1.
Let and write . Under the symmetric perturbation convention, the coordinate gradient of has entries on diagonal coordinates and on off-diagonal coordinates. Therefore,
by the Gaussian quadratic-form variance formula. If , the chain rule gives
Consequently, is similar to , so traces of powers agree. Since is a bijection of , the identity (4) for transfers to with the same constant. ∎
We next determine how the sampling law enters. Throughout the following discussion, is a probability law on under which the vector is centred with covariance matrix and finite fourth moments, and is the associated covariance operator. Directly from the definition of the fourth joint cumulant in Section A,
| (B.1) |
so the departure of from the Gaussian operator (2) consists exactly of the fourth cumulants. Suppose now that and that is elliptical, meaning that has the stochastic representation , where , the vector is uniform on the unit sphere in , and the radius is independent of . The covariance normalization forces . Define the kurtosis parameter of by
| (B.2) |
so that under Gaussian sampling, and Jensen’s inequality gives . Write and .
Proposition B.1 (Elliptical fourth moments).
Let be elliptical with covariance matrix and kurtosis parameter , and let . Then
| (B.3) | ||||
| (B.4) |
If is homogeneous of degree , then satisfies (4) with constant if and only if identically, where , and in that case
| (B.5) |
at every sample covariance matrix at which the displayed denominators are positive.
Both sides of (B.4) are polynomial in , so the proportionality in the proposition is an identity in the polynomial ring, exactly as in the Gaussian case. The proposition separates two conclusions. Ellipticity rescales the proportionality constant of a homogeneous polynomial but does not change the algebraic class, so the classification developed in the main text is unaffected. At the same time, only the second ratio in (B.5) carries the correct fourth-moment scaling when governs the sampling law, so samplewise algebraic constancy is distinct from correct variance calibration, which is the distinction drawn in the main text.
Proof of Proposition B.1.
All expectations are under . Since , the matrix is nonsingular, so and (B.2) states that . Put , so that and . The spherical fourth-moment identity is
and independence of and gives
Since , subtraction of the squared mean proves (B.4), and coefficient comparison over yields (B.3).
If is homogeneous of degree , Euler’s identity gives . Substituting (4) into (B.4) gives , which is the stated proportionality, and evaluating the two polynomial identities at gives (B.5). Conversely, if identically, then permits solving (B.4) for , and Euler’s identity recovers (4) with constant . ∎
Remark B.1 (Elliptical sampling and unknown means).
The product-of-chi-squares pivot in Theorem 1 is generally unavailable under elliptical sampling, because the sample covariance matrix is Wishart only when the radial law is Gaussian. Under with an unknown mean, let and
Then , every exact Gaussian pivot holds with replaced by , and the standardized denominator of a flag power computed from equals . For non-Gaussian elliptical sampling with an estimated mean, the distinction in (B.5) persists. Algebraic constancy survives, correct calibration requires the kurtosis-adjusted operator, and a Wishart factorization is unavailable in general. In the front-door regressions of the main text fitted with intercepts, the residual degrees of freedom become and .
Two structural reductions underlie every converse argument. The first confines every question about exact self-normalization to a single homogeneous degree, and the second supplies the interior positivity on which the logarithmic and geodesic arguments rely.
Lemma B.2 (Homogeneous reduction).
Every rational strategy for a scale-invariant coefficient in a linear structural equation model admits a representative whose numerator and denominator are homogeneous of the same degree. If a nonhomogeneous polynomial satisfies exact self-normalization, then its highest- and lowest-degree nonzero homogeneous parts satisfy the same identity with the same constant.
Lemma B.2 permits all converse arguments to be conducted within one homogeneous degree, which is the standing convention of Sections 3 and 4 of the main text.
Proof of Lemma B.2.
For a linear structural equation model, multiplying every error covariance by multiplies by but leaves the scale-invariant causal coefficient under study unchanged. Expanding in powers of shows that each pair of homogeneous components of the same degree satisfies the identification identity, and at least one denominator component is nonzero on the model. For the second assertion, substitute into (4) and compare the highest and lowest powers of . ∎
Lemma B.3 (Interior nonvanishing).
A nonzero polynomial satisfying exact self-normalization has no zero in . It therefore has a constant sign on that cone.
Lemma B.3 makes the logarithm of a self-normalizing polynomial globally available on the positive cone, which the geodesic arguments of Supplementary Sections C and D use throughout. Its failure at the identity matrix is also what rejects the proximal denominator in the factor-and-rank example of the main text.
Proof of Lemma B.3.
Because is a nonzero polynomial and is a nonempty Euclidean-open set, there is with . Let be the connected component of containing , and define only on , so that no global nonvanishing conclusion is used at this stage. Equip with the affine-invariant metric
Its gradient on is , and (4) gives with .
Suppose that for some . Let be the affine-invariant geodesic segment from to , which has finite length. Let . Then , and for ,
Integration over the finite-length segment shows that is bounded below as . Continuity of the polynomial and instead imply , a contradiction. Hence, has no zero in . Since the cone is connected, its sign is constant. ∎
The final lemma shows that exact self-normalization is a property of the smallest coordinate block on which the polynomial depends.
Lemma B.4 (Invariance under ambient extension).
Let , let , and define by . Then for every . Consequently, is exactly self-normalizing on if and only if is exactly self-normalizing on , with the same constant.
Lemma B.4 justifies the dimension reductions invoked in Section 2 of the main text and in the simulation design. Combined with the congruence invariance of Lemma B.1, it applies to a polynomial depending on only through its compression to an arbitrary -dimensional subspace, because one nonsingular linear change of the observation vector moves that subspace to the leading coordinate block.
Proof of Lemma B.4.
Since whenever , the symmetric gradient has the block form . Writing in the corresponding blocks, has zero second block row, and the trace of its square equals , which proves the identity. The equivalence follows because ranges over all of as ranges over . ∎
C Proofs of Section 3: Flag powers
This section proves Theorem 1 and the two-subspace criterion stated after it in Section 3 of the main text. The theorem combines an upper-triangular calculation of the relative gradient with the Bartlett decomposition of the Wishart matrix. The final lemma shows that the nesting condition is also necessary for a product of two covariance-volume factors.
Proof of Theorem 1.
By congruence invariance of Lemma B.1, it suffices to treat the coordinate flag, and the constants arising from the chosen basis matrices cancel from the logarithmic gradient. Put . For every leading block is invertible and
Weighting the th term by the exponent and summing over shows that is upper triangular with diagonal , which proves (i).
Therefore, the trace of the square of is , so (4) holds with on the cone. Both sides of (4) are polynomial, so the identity extends to all of , proving (ii).
For part (iii), write with lower triangular and , where . Since is lower triangular, , so leading determinants factor blockwise, and the Bartlett decomposition (30, Ch. 3) gives with independent . Collecting exponents gives (iii).
For part (iv), a flag power is positive at every positive-definite matrix, so and almost surely, and evaluating (4) at gives . The same identity gives for every . If , then , so , which is incompatible with . ∎
The nesting assumption in Theorem 1 cannot be dropped even for a product of two covariance-volume factors.
Lemma C.1 (Two-subspace criterion).
Let and let be positive integers. Then is exactly self-normalizing if and only if or .
Proof of Lemma C.1.
The logarithmic relative gradients of and are the -orthogonal projections and onto the two subspaces. Therefore,
The final trace is the sum of the squared cosines of the principal angles in the inner product . Under either containment, it equals the dimension of the smaller subspace and is constant.
Suppose that neither containment holds. Write , , and , where and are nonzero. The sum is direct. Choose a positive-definite inner product that makes these three subspaces mutually orthogonal. The final trace then equals . Perturb the inner product so that one nonzero vector in has nonzero inner product with one in , while retaining positive definiteness. At least one squared principal-angle cosine becomes positive, so the trace changes. Exact self-normalization is therefore impossible. ∎
D Proofs of Section 4: Converse results and diagnosis
D.1 Overview and algebraic preliminaries
This subsection proves the converse and diagnostic results in Section 4 of the main text. The argument first reduces the problem to boundary faces and determinant-free residues. It then proves the complete converse in dimension two and the three-variable reduction. The remaining subsections study the mixed branch, the conditional higher-dimensional converse, and the low-degree-factor diagnostic.
The first two lemmas provide the divisibility tools used throughout the section.
Lemma D.1 (Irreducibility of the generic symmetric determinant).
For every , the determinant of the generic symmetric matrix of order is irreducible over and over . Consequently, each leading principal minor is irreducible in for .
Irreducible elements are prime in the unique factorization domain . Therefore, Lemma D.1 justifies the determinant-divisibility and valuation arguments below.
Proof of Lemma D.1.
We argue over a field of characteristic zero, which covers and . The assertion is immediate for . Suppose it holds for , and write the generic symmetric matrix as
where is generic symmetric of order . Expansion in gives
Let , which is a unique factorization domain. By induction, is irreducible in . If divided , it would divide every coefficient of this polynomial in the coordinates of . In particular, it would divide the coefficient of , which is a nonzero principal minor of of order and has smaller degree. Thus, the two coefficients of the displayed polynomial in are coprime. The polynomial is primitive in and irreducible over the fraction field of . Gauss’ lemma gives irreducibility in . Passing to a larger polynomial ring preserves irreducibility, which proves the final assertion. ∎
The next lemma converts vanishing on a real part of the determinant boundary into polynomial divisibility. This step is needed when two candidate solutions agree only on rank-deficient covariance matrices.
Lemma D.2 (Real determinant divisibility).
If a polynomial on vanishes on a nonempty Euclidean-open subset of the real symmetric matrices of rank , then divides .
Proof of Lemma D.2.
After a simultaneous permutation of rows and columns, work in a chart in which the leading principal minor of order is nonzero at one point of the given open set. Shrink the set so that throughout it. The determinant is linear in , and can be written as
where does not involve . If , pseudo-division in gives
On this chart, the equation is the graph . The projection of the given Euclidean-open subset of this graph to the remaining coordinates is Euclidean open. Since does not involve and vanishes on that projection, . Thus, divides . By Lemma D.1, is prime in . It does not divide , because does not involve whereas does. Hence, divides . ∎
The face-restriction principle transfers exact self-normalization to every nonzero covariance face. Together with Theorem 2, it determines the form of each nonzero plane restriction in dimension three.
Lemma D.3 (Face restriction).
Proof of Lemma D.3.
The stripping lemma removes one full determinant factor and records the resulting change in the self-normalization constant. It is the degree-lowering step in the dimension-two proof and in the three-variable reduction.
Lemma D.4 (Determinant stripping).
If and is homogeneous, then is exactly self-normalizing if and only if is. Their constants satisfy
Proof of Lemma D.4.
On the positive cone, write and . The product rule gives . Therefore,
Exact self-normalization is equivalent on the cone to constancy of the final trace. Hence, it holds for if and only if it holds for . The displayed identity gives . Both self-normalization identities then extend polynomially to all of . ∎
D.2 Complete converse in dimension two
The proof separates solutions according to whether the determinant of the symmetric gradient vanishes identically. A nonzero determinant coefficient produces a removable determinant factor. The zero coefficient leads to a rank-one gradient and then to a power of one variance direction.
Proof of Theorem 2.
Let and let be the eigenvalues of . On the positive cone they are real because this matrix is similar to . Euler’s identity and (4) give
Hence, as a polynomial identity,
| (D.1) |
If , Lemma D.1 gives . Lemma D.4 removes this factor, and induction on applies.
Suppose that . Then . The polynomial gradient map takes values in the rank-one quadric cone in . At every nonzero smooth image point, the differential of the gradient, which is the Hessian of , maps into the two-dimensional tangent plane. Thus, . The Gordan–Noether theorem for forms in at most four variables implies over that is a cone. The reduction may be taken over . A complex constant direction annihilating brings its conjugate with it. If the two directions are proportional, rescaling gives a real direction. If they are independent, depends on one linear form, which reality makes real up to a scalar. Thus, after a real linear change, .
Write for fixed independent matrices . The binary quadratic is nonzero because the determinant cone contains no two-dimensional linear subspace. The identity either forces when is definite, or forces one real linear factor of to vanish identically. Thus, depends on one linear form and
A nonzero real symmetric rank-one matrix is a nonzero scalar multiple of . This gives the form in (8). The value of follows from Theorem 1. ∎
The two branches of the proof correspond exactly to the determinant factor and the rank-one variance factor in Theorem 2. No other irreducible factor can occur in dimension two.
D.3 Adjugate duality and three-variable face types
We next turn to determinant-free solutions on . Adjugate duality exchanges rank-one and rank-two boundary data. The face-type lemma then shows that every nonzero plane restriction has one common pair of multiplicities.
Lemma D.5 (Adjugate duality).
Proof of Lemma D.5.
On the positive cone, , so
If , differentiation gives
The second term is similar to the relative gradient of at . If its eigenvalues are , Euler’s identity and (4) give and . Hence,
Clearing denominators proves the polynomial identity for . The involution formula follows from . ∎
For the next lemma, write . The result makes the face type in Definition 3 well defined and gives the bounds used in the three-variable reduction.
Lemma D.6 (Face type in dimension three).
Let be determinant-free, homogeneous and exactly self-normalizing on . At least one -good plane exists, and every -good restriction has the form (9) with one common pair . Moreover, , with if and only if .
Proof of Lemma D.6.
If every plane restriction vanished, would vanish on every singular positive-semidefinite matrix. Lemma D.2 would then give , contrary to determinant-freeness. A -good face restriction satisfies (4) by Lemma D.3. Theorem 2 gives its form. Substituting into the equation for the self-normalization constant gives
Its two roots and satisfy . Admissibility requires . If both roots were admissible, their sum could equal only when . Hence, there is at most one admissible root, and the face type is common to all good planes. The bounds on follow from the same two equations. ∎
The adjugate maps a rank-two covariance face to the rank-one cone. The following formula makes that action explicit.
Lemma D.7 (Boundary strata under the adjugate).
If and is the cross product of its columns, then, for every ,
Proof of Lemma D.7.
The cofactor of is the product of the corresponding row minors of and . Those signed minors are the coordinates of . ∎
The next dichotomy separates solutions that survive on the rank-one cone from those that vanish there. The surviving branch must have the largest possible self-normalization constant and a pure rank-one face type.
Lemma D.8 (Rank-one dichotomy).
Let be homogeneous of degree and exactly self-normalizing on . Either , or is determinant-free, , and its face type is .
Proof of Lemma D.8.
The dichotomy leaves a rigidity question on the rank-one cone. The next proposition closes that branch and shows that it contains only powers of one variance direction.
Proposition D.1 (Rank-one rigidity).
Let be homogeneous of degree and exactly self-normalizing on . If , then for a nonzero constant and a nonzero vector .
Proof of Proposition D.1.
By Lemma D.8, every -good plane restriction is a th power of one linear covariance form. Thus, on a Zariski-open family of projective lines in , the form restricts to a polynomial whose zero set has one support point. If the reduced projective curve had degree at least two, a general line would meet it transversally in at least two points by Bezout’s theorem (20). Hence, the reduced curve is one line and
with and real after rescaling.
Let . On every -good plane, both face restrictions are th powers of linear covariance forms and agree on all rank-one matrices in that face. Unique factorization of the resulting binary powers makes the linear forms proportional with the matching constant. Hence, the two face polynomials agree identically. Therefore, and agree on a nonempty open set of rank-two matrices, and Lemma D.2 gives .
If the difference is nonzero, write with , , and . After a congruence, take and put . Expanding (4) and dividing by gives
| (D.2) |
for a polynomial .
Use the Schur coordinates
The line changes only to . Therefore, the left-hand side of (D.2) is . On ,
The polynomial has -degree at most . Coefficient comparison forces on . Thus, vanishes on a nonempty open set of rank-two matrices and is divisible by , which contradicts its definition. ∎
D.4 Relative spectral constancy in dimension three
Euler’s identity fixes the trace of the relative gradient, while exact self-normalization fixes its squared Frobenius norm. The next proposition uses the polynomial character of to obtain the additional discrete constraint that makes the spectrum constant in dimension three.
Proposition D.2 (Spectral constancy).
Let be homogeneous of degree and exactly self-normalizing on , and put . There is a fixed multiset such that
In every case,
If , then and .
Proof of Proposition D.2.
By Lemma B.3, has constant sign on . Multiply it by if necessary so that there. Fix and define
By Lemma B.1, satisfies (4) with the same constant. At , put
The matrix is symmetric and is similar to . In particular, and .
By Cauchy–Schwarz, . If , equality forces all three eigenvalues of to equal . Since was arbitrary and is diagonalizable, throughout the cone. Hence, . On a nonempty open set this gives . Polynomial continuation and unique factorization using Lemma D.1 yield and . The determinant identity follows with . Henceforth, assume .
Let and equip with the affine-invariant metric. Equation (4) gives . For every vector field , symmetry of the Riemannian Hessian gives
Thus, the integral curves of are geodesics. The geodesic through with initial velocity is . Uniqueness of geodesics and integral curves makes the two curves agree near . Along that interval,
Therefore,
| (D.3) |
Both sides are entire functions of , so (D.3) holds for every .
Choose an orthogonal matrix such that . Define
Since , at least one coefficient group is nonzero. Substitution in (D.3) gives
Linear independence of real exponentials implies for at least one with .
The equations and define a circle in the trace plane. If is proportional to , then and . Thus, this exceptional resonance contains no admissible point. For every other , the equation cuts the trace plane in a proper affine line and meets the circle in at most two points. There are only finitely many such . Hence, the ordered spectrum belongs to a finite set. The ordered eigenvalues vary continuously with because is similar to the symmetric matrix
Since is connected, the ordered spectrum is constant. Taking the product of the fixed eigenvalues and clearing denominators gives the determinant identity. ∎
For a determinant-free solution, the determinant identity forces one relative eigenvalue to vanish. The face type determines the other two.
Corollary D.1.
If , then one relative eigenvalue is zero and . If has face type , then
Proof of Corollary D.1.
The left-hand side of the determinant identity in Proposition D.2 is divisible by the prime . If the eigenvalue product were nonzero, then , contrary to the hypothesis. The remaining two eigenvalues have sum and product . They are therefore and . ∎
Spectral constancy supplies the singular-gradient structure of every determinant-free three-variable solution. It is the step that reduces the unresolved case to a polynomial kernel line.
D.5 Pure plane-minor and mixed branches
We first identify the branch whose nonzero plane restrictions are pure determinant powers. This proves Theorem 3 (ii).
Proof of the pure plane-minor branch in Theorem 3.
Suppose that the face type is , so and the spectrum is . The adjugate transform has spectrum . Since , the polynomial vanishes on every singular positive-semidefinite matrix. Lemma D.2 gives . Write with exact . The adjugate involution gives , and determinant stripping shifts every relative eigenvalue down by . Hence,
Because , Corollary D.1 forces a zero eigenvalue. The bound leaves only . Thus, and . The pair satisfies the equations that determine the face type of . Lemma D.6 therefore gives face type . A nonzero restriction of this type is nonzero at a suitable rank-one matrix, so . Proposition D.1 gives . Applying the adjugate again yields . ∎
Mixed face types are paired by adjugate duality. This symmetry is used in both the kernel analysis and the boundary-rigidity argument.
Lemma D.9 (Mixed duality).
If is determinant-free of mixed type , then is determinant-free of type and .
Proof of Lemma D.9.
The spectrum of is , so that of is . As in the pure branch, . Let be its exact multiplicity. After stripping, the spectrum is . The adjugate involution gives . Corollary D.1 requires a zero eigenvalue, so . Since , . The involution bound therefore excludes and . Hence, . The stripped spectrum is and the degree is . The pair satisfies the two equations that determine the face type. Lemma D.6 gives the stated type, and the involution identity gives the final formula. ∎
The next algebraic lemma turns a rank-one factorization over the fraction field into a polynomial factorization. It is applied to the adjugate of the singular symmetric gradient.
Lemma D.10 (Primitive rank-one factorization over a UFD).
Let be a unique factorization domain and let be a nonzero matrix whose minors all vanish. Then for vectors , where is primitive. If is symmetric, then for some . If the entries of are homogeneous of one common degree, the factors may be chosen homogeneous.
Proof of Lemma D.10.
Choose a nonzero column . Let be a greatest common divisor of its entries and put . Then is primitive. Since all minors vanish, every other column is proportional to over . For each , there is such that . Write in lowest terms. The identity holds for every . If an irreducible element divided , it would divide every , contrary to primitivity. Hence, is a unit and . Thus, with .
If is symmetric, then . Hence, for some . The same denominator argument shows that . When the entries of are homogeneous, the greatest common divisor may be chosen homogeneous. Degree comparison then makes , , and homogeneous. ∎
For a mixed solution, the adjugate of the symmetric gradient has rank one. The factorization lemma therefore produces the polynomial kernel vector in Theorem 3. It also shows why a vanishing-Hessian argument arises naturally.
Proof of the kernel-field assertion in Theorem 3.
For a mixed type, has rank two on the positive cone and . The identity shows that every minor of vanishes. The matrix is not identically zero because has rank two on a nonempty open set. Apply Lemma D.10 in . Since is symmetric and homogeneous, there are a primitive homogeneous vector and a nonzero homogeneous polynomial such that
The identity gives . Choose an index for which . Since is an integral domain, every component of vanishes. Hence, .
Differentiate this identity in a symmetric direction and left multiply by . This gives . Thus, the rank-one direction lies in the radical of the coordinate Hessian wherever . In particular, . ∎
We can now assemble the three-variable reduction. The proof first strips full determinant factors and then applies the rank-one, pure plane-minor, or mixed analysis according to the boundary restriction.
Proof of Theorem 3.
Apply Lemma D.4 repeatedly to write the original polynomial as , where . If , Lemma D.8 and Proposition D.1 give case (i). Suppose that . Lemma D.6 supplies a common face type . Evaluating a nonzero face restriction at a rank-one matrix shows that . If , the pure plane-minor argument above gives case (ii). If , the type is mixed. Proposition D.2, Corollary D.1, and the kernel-field argument give every assertion in case (iii). Since a mixed type has degree , no mixed case occurs in degrees at most two. ∎
Theorem 3 leaves one geometric question. The next lemma shows that constancy of the kernel line is exactly the missing condition for the mixed flag form.
Lemma D.11 (Constant kernel and the mixed flag form).
Let be determinant-free and of mixed face type . The following statements are equivalent.
- (i)
The projective kernel line is constant on a nonempty open subset of .
- (ii)
A fixed nonzero vector satisfies .
- (iii)
There are a nonzero constant , a vector , and a full-row-rank matrix with such that
Every mixed flag power has this constant kernel line.
Proof of Lemma D.11.
If the projective kernel line is constant on a nonempty open set, choose a fixed representative . Then on that open set and hence identically, because its entries are polynomials. This proves (i)(ii), and the reverse implication is immediate.
Assume (ii). After an orthogonal congruence, take . Under the symmetric-gradient convention,
These derivatives vanish identically, so depends only on the leading block . Write . The block form of gives
Thus, is exactly self-normalizing on with the same constant. Theorem 2 gives . Face-type uniqueness in Lemma D.6 gives . Lifting to proves (iii).
Conversely, both gradient factors of a mixed flag power annihilate . The product rule therefore gives identically. Since has rank two on , its kernel is exactly there. ∎
The cofactor identity provides a second constraint on the moving kernel. It couples the kernel vector to the determinant boundary without assuming that the kernel line is constant.
Proposition D.3 (Cofactor syzygy).
Let be a determinant-free mixed solution of type , and write . Then
| (D.4) |
D.6 Boundary rigidity for the mixed branch
The constant-kernel criterion is geometric. The results in this subsection give a separate algebraic closure when the rank-two boundary data already agree with one fixed flag. The argument does not assume the desired interior factorization.
Definition D.1 (Flag-coherent boundary).
A determinant-free mixed polynomial of type has a flag-coherent boundary if there are nested spaces and a nonzero constant such that
| (D.5) |
For and , put
The set is the collection of weights available to a homogeneous correction of degree on the determinant boundary.
Definition D.2 (Boundary nonresonance).
The type is boundary-nonresonant if
| (D.6) |
The next lemma gives two arithmetic forms of this condition. They are useful for locating the first possible resonance.
Lemma D.12 (Arithmetic characterization of boundary nonresonance).
A type is resonant if and only if there is an integer such that
| (D.7) |
Equivalently, if , boundary nonresonance is
| (D.8) |
Proof of Lemma D.12.
A resonance has the form for an integer . Solving for gives
For , this quantity is positive. It is an integer if and only if . The upper bound is equivalent to . This proves (D.7).
Write and , where . The divisibility condition is equivalent to . Hence, the smallest positive admissible value is . Since increases with , a resonance exists if and only if
After division by , this is . If this inequality holds, then and
because . Thus, automatically. Negating the resonance criterion gives (D.8). ∎
Boundary nonresonance excludes every polynomial correction compatible with the weighted Euler equation on a rank-two face. It therefore turns flag-coherent boundary data into a unique interior solution.
Proposition D.4 (Boundary-coherent mixed converse).
If a determinant-free exactly self-normalizing polynomial of mixed type has a flag-coherent boundary and is boundary-nonresonant, then .
Proof of Proposition D.4.
By congruence, take
If , boundary coherence gives for an exact , where and . Put and . Then . Since and have the same constant , expansion of (4), cancellation, and restriction to give
| (D.9) |
Write and . The left-hand side of (D.9) is . Hence, on a dense boundary chart,
| (D.10) |
Use the polynomial parametrization
Then , , , and . A direct logarithmic-gradient calculation gives
At fixed , equation (D.10) on becomes
Since is homogeneous of degree and is linear in , its restriction has the expansion
The monomial indexed by has weight . Boundary nonresonance forces every to vanish. Thus, vanishes on a nonempty open subset of the rank-two stratum. Lemma D.2 gives , which contradicts the exact choice of . Therefore, . ∎
The arithmetic condition is automatic in the lowest mixed degrees. The first normalized resonance appears only in degree nine.
Corollary D.2 (Low mixed degrees and the resonance locus).
After mixed adjugate duality, normalize the type by . A flag-coherent solution is a flag power when , and hence for every mixed degree at most eight. The first normalized resonance is in degree nine. More generally, resonance is equivalent to .
Proof of Corollary D.2.
A resonance requires . If or , this fails for and therefore for every larger . If , then . At , the choice gives equality and satisfies . The greatest-common-divisor characterization follows from Lemma D.12. ∎
The preceding proposition gives a conditional closure of the three-variable converse that is independent of kernel-line constancy. It applies whenever the face factors glue to one fixed flag and the resulting type is nonresonant.
Theorem D.1 (Conditional characterization in dimension three).
Suppose that every determinant-free mixed residue can, after mixed adjugate duality if necessary, be chosen with a flag-coherent boundary and a boundary-nonresonant type. On this class, the exactly self-normalizing polynomials are precisely the flag powers.
Proof of Theorem D.1.
Strip determinant powers and apply Theorem 3 to the residue. The rank-one and pure plane-minor cases are flag powers. In the mixed case, apply the assumed adjugate normalization, boundary coherence, and Proposition D.4. Mixed duality returns the flag form to the original residue. Restoring determinant powers adds the full space to the flag. ∎
The conditional theorem isolates two separate issues in the mixed branch. The first is geometric gluing of the face directions. The second is an explicit arithmetic resonance that begins only at higher degree.
D.7 A conditional converse in arbitrary dimension
The three-variable spectral argument does not extend directly to higher dimensions. A converse is nevertheless available once the relative gradients preserve one fixed recursive ordering. The next proposition shows that this invariant flag determines the polynomial uniquely.
Proposition D.5 (Recursive-flag converse).
Let be homogeneous, nonconstant and exactly self-normalizing on . Suppose that a complete flag is preserved by for every . Suppose also that the scalar induced on is a constant . After a congruence,
Conversely, every flag power has such a fixed recursive flag.
Proof of Proposition D.5.
By Lemma B.3, has no zero on . Hence, is globally defined on the connected cone and on its connected Cholesky parametrization. Use a congruence to take . Then is upper triangular with diagonal . Write the unique Cholesky factorization , where is lower triangular with positive diagonal, and put . The matrix
is symmetric. It is also upper triangular because preserves the coordinate flag and conjugation by and preserves upper triangularity. Therefore, is diagonal. Triangular conjugation preserves diagonal entries, so .
For an infinitesimal lower-triangular perturbation ,
Hence,
Integration on the connected Cholesky domain gives
for a nonzero constant .
Restrict to diagonal Cholesky factors and write . Then is a polynomial in the and equals on the positive orthant. Fixing all variables except shows that each is a nonnegative integer. Since , we obtain in the fraction field
By Lemma D.1, the are pairwise nonassociate irreducible polynomials. The valuation of the polynomial at is and must be nonnegative. This proves the factorization.
The converse is the upper-triangular calculation in the proof of Theorem 1. If only a constant spectrum is assumed in addition to a common invariant flag, each quotient weight is a continuous map from the connected cone to a finite multiset and is therefore constant. ∎
Proposition D.5 separates the higher-dimensional obstruction from the integration step. Once one invariant recursive flag is present, no additional polynomial solutions occur.
D.8 Proof of the low-degree-factor diagnostic
The final part of Section 4 concerns denominators whose determinant-free irreducible factors have degree at most two. The proof identifies global linear and quadratic factors from their generic plane restrictions. It then uses the two-subspace criterion in Supplementary Section C to force nesting.
Lemma D.13 (Generic linear factors).
Let with and . If the restriction of to a Zariski-dense set of planes is a rank-one linear form on , then has rank one. If two such forms are associates on a dense set of planes, their rank-one directions are proportional.
Proof of Lemma D.13.
If , the symmetric bilinear form represented by has a nondegenerate two-dimensional compression. Rank at least two persists on a Zariski-open neighbourhood of that plane, contrary to the hypothesis. Thus, . For two rank-one forms, association on a dense family of planes makes the wedge of the two projected vectors vanish as a polynomial in the plane coordinates. It therefore vanishes on every plane. The plane spanned by two nonproportional directions would give a contradiction, so the directions are proportional. ∎
The next lemma globalizes a pure-power restriction on generic projective lines. It is used to exclude the square alternative for an irreducible quadratic factor.
Lemma D.14 (Single-support restriction).
Let be a nonzero homogeneous form of degree on . If the restriction of to a Zariski-dense set of projective lines is an th power of a linear form, then for a linear form .
Proof of Lemma D.14.
On each line in the stated family, the zero set has one support point. If the reduced plane curve had degree at least two, a general line would meet it transversally in at least two points by Bezout’s theorem. Hence, the reduced curve is a line, and unique factorization gives the claim. ∎
An admissible irreducible quadratic must vanish on the rank-one cone. The next lemma identifies it as a linear form in the adjugate matrix.
Lemma D.15 (Irreducible quadratic factors).
Let be an irreducible real quadratic on . Suppose that, on a dense family of planes, every irreducible factor of is either the unique rank-one linear face factor or . Then for a nonzero matrix .
Proof of Lemma D.15.
Work in an affine chart of in which a basis matrix depends polynomially on the chart coordinates . The coefficients of are polynomial in . Proportionality to is cut out by the minors of the corresponding coefficient vectors. Being a scalar multiple of a square of a linear form is cut out by the minors of the symmetric coefficient matrix of the ternary quadratic . These equations are compatible on overlaps and define intrinsic Zariski-closed subsets and of the complexified Grassmannian. The real Grassmannian is Zariski dense in its complexification. On a dense open set of good real planes, a degree-two restriction is either proportional to or to the square of the unique linear face factor. Hence, . Since the Grassmannian is irreducible, one of these closed subsets is the whole Grassmannian.
Consider first the determinant alternative. Proportionality to then holds for every plane, including a zero restriction. Let . This is a nonempty Zariski-open set. It is nonempty because otherwise would vanish on the rank-two determinant hypersurface and would be divisible by the cubic , which is impossible for a nonzero quadratic. For , with . Put . Identify with the dual projective plane. The planes containing a fixed point form a projective line . If lies on no plane in , then . A proper closed subset of the projective plane has only finitely many one-dimensional irreducible components. Therefore, this can occur for at most finitely many pencils . For every other , there is with , and . Thus, vanishes on a Zariski-dense subset of the rank-one cone and hence on the whole cone.
The space of quadratics vanishing on is . Quadratics on have dimension , while quartics in have dimension . The restriction map is surjective because every quartic monomial is a product of two quadratic monomials. The displayed family has dimension six and is injective. Indeed, if , then for every positive-definite choose so that . Then on an open set, which gives . The displayed family is therefore exactly the kernel of the restriction map.
Suppose instead that is generically a square. Then restricts to a fourth power on a dense set of lines. Lemma D.14 gives . Consequently, for some . On a generic plane, the first two quadratic expressions are squares of linear forms, while the last is , where is a normal to the plane. A difference of two squares has matrix rank at most two as a quadratic form on . By contrast, has rank three as a quadratic form on . Hence, is impossible. Since for a Zariski-dense set of normals, . Then is a square, contrary to irreducibility. Only the determinant alternative remains. ∎
The preceding lemmas show that all global linear factors are powers of one variance direction and all irreducible quadratic factors are adjugate-linear. Adjugate duality then forces the quadratic directions to coincide.
Proof of Theorem 4.
Strip the largest determinant power and call the determinant-free residue . Factor it over as
where the are pairwise nonassociate linear forms and the are pairwise nonassociate irreducible quadratics. Choose a generic good plane on which all indicated restrictions are nonzero. The common face-type representation is . If , unique factorization on a generic face forces . If , every is associated with the unique linear factor . Lemma D.13 shows that all are rank-one and associate. Their product is therefore absent or has the form .
For each , its generic restriction is either a square of , when , or a multiple of . Lemma D.15 gives . Put . If , adjugate duality yields
where the factor involving is omitted when . The displayed power of the irreducible cubic is exact. After stripping it, the residue is again determinant-free and exactly self-normalizing. Its linear factors restrict on a generic face to the unique linear face factor. Lemma D.13 shows that every has rank one and that all are proportional. Hence, and
with an absent factor interpreted as one. The second factor is the determinant of the compression to . If , Lemma C.1 forces the line and plane to be nested. A plane cannot be contained in a line, so . Restoring the stripped determinant power gives the flag form in (11). ∎
The corollary converts the flag form into coefficient-level rank and orthogonality tests. It also identifies every reducible self-normalizing cubic.
Proof of Corollary 1.
Necessity is Theorem 4. Sufficiency and the value of follow from Theorem 1, whose multiplicities are . A cubic satisfying the low-degree-factor condition is either a scalar multiple of , or its determinant-free part has only linear and quadratic factors. The theorem gives the three possibilities in the corollary. Every reducible cubic belongs to the latter class. ∎
The low-degree-factor converse is unconditional on its stated class. An irreducible factor of degree at least three is the only factorization pattern not decided by this diagnostic.
E Proofs of Section 5: Weak-denominator inference
This section proves Propositions 1 and 2 and verifies the example that delimits the scope of the classification in Section 5 of the main text. Both proofs rest on one triangular-array central limit theorem for the sample covariance, which is established at the start of the first proof and reused in the second.
The proof of Proposition 1 combines this central limit theorem with polynomial delta expansions along the drifting sequence.
Proof of Proposition 1.
For row , write
Within each row the are independent and identically distributed, with covariance . Since , Gaussian eighth-moment formulas for the quadratic vector give
Consequently, the Lyapunov quantity is . The multivariate triangular-array Lyapunov theorem gives
Because and the polynomials are fixed, and their second derivatives are bounded on a fixed neighbourhood of , Taylor’s theorem together with yields
Since and , the pair converges jointly to with the stated correlation. The limiting denominator has a continuous distribution, so division gives the ratio limit in (18). The convergence and then give the limit of .
For the Wald statistic, write and use the direction defined in (17). The ratio limit gives , so the polynomial-gradient expansion gives
Substituting the ratio limit shows that converges in distribution to . Combining this with
gives (19). The limiting studentizer equals , which is positive almost surely when . ∎
The argument uses the drifting sequence only through the limits , and , which is why a single three-parameter experiment covers every covariance-polynomial strategy, as stated in the main text.
The proof of Proposition 2 treats coverage and sample geometry separately. Coverage follows from the same central limit theorem applied to the moment at the local target, and the trichotomy is the sign analysis of one quadratic polynomial in .
Proof of Proposition 2.
At , the moment vanishes at and has gradient . The central limit theorem from the proof of Proposition 1, together with , shows that the squared studentized moment converges in distribution to . This proves the coverage statement without requiring a nonzero limit of .
For the sample geometry, put and
a quadratic polynomial in whose leading coefficient is
whenever . Under and , the matrix has a Wishart density on and is therefore absolutely continuous with respect to Lebesgue measure on (30, Ch. 3). The polynomial is not identically zero because , so is a null event. Likewise is a null event, so is defined almost surely.
The event is also null. If the polynomial is not identically zero, this follows from absolute continuity. If it were identically zero, then and would be exactly self-normalizing, whereas Assumption 1 gives and , and these two requirements are incompatible with (4).
Finally,
so the sublevel set is nonempty. On the probability-one event , elementary quadratic geometry shows that this set is an interval, possibly a singleton, the complement of a bounded open interval, or all of , and that it is bounded exactly when , which is equivalent to . This is the classical Fieller trichotomy (13), and the inversion is also an Anderson–Rubin construction (1). ∎
The final item verifies the example that the main text uses to separate exact self-normalization from mere first-order degeneracy on a zero set.
Example E.1 (First-order degeneracy without exact self-normalization).
On , let . Then
so vanishes on the entire zero set of although is not constant. For and ,
The variance formula follows from (3), because the only nonzero coordinate of is and the corresponding diagonal entry of is . Cancelling one factor of gives
wherever , and the same cancellation at gives the population limit . Under the displayed drift, by the central limit theorem in the proof of Proposition 1, while , which proves the limit of . Assumption 1 fails along this sequence because , so the nondegenerate limit of arises at second order, exactly as claimed in the main text.
F Proofs of Section 6: Causal applications
F.1 Structural pullbacks used in Table 1
This subsection verifies the third column of Table 1 in the main text. Each entry records a direct substitution into a linear reduced form for that row and is not an additional identification claim. All variables are centred, and symbols are local to the reduced form in which they appear, as in the footnote to the table.
For the instrumental-variable row, let
Then .
For the conditional-instrument row, let
where is uncorrelated with and is uncorrelated with . Substituting and cancels the terms in and gives
For the front-door row, let
Then,
so the front-door denominator equals .
For the proximal row, let
where are mutually uncorrelated. Then
These four calculations define the symbols of the table and make every displayed pullback verifiable.
F.2 Marginal and partial minors
This subsection proves the marginal and partial calibrations displayed in Section 6 of the main text. Let with and , allowing where indicated. As in the main text, is the sample correlation of coordinates and , the sample partial correlation is the correlation of the residuals from the sample linear projections of coordinates and on coordinate , and the partial first-stage index is .
The next proposition contains both variance identities and both closed forms for the standardized denominator.
Proposition F.1 (Marginal and partial minor calibration).
For with ,
For ,
When the second identity reduces to . When , .
The case recovers the constant of Theorem 1 for a principal minor of order two. In the two slope cases the standardized denominator is an increasing transform of the familiar marginal or partial first-stage index, so the algebraic diagnostic automatically selects the screening direction relevant to the strategy.
Proof of Proposition F.1.
For with , the only nonzero coordinate of is , and the corresponding diagonal entry of in (2) is , which proves the first identity. Evaluating at and dividing the numerator and denominator of by gives the first closed form.
For , both sides of the asserted identity scale by the same factor under for positive diagonal , so it suffices to verify it as . Put , and . The nonzero coordinates of are , , and , and expanding the quadratic form with the entries of (2) gives
Undoing the rescaling gives the determinant identity.
If , then , so . If , the sample partial correlation satisfies
so
and substituting gives the second closed form. ∎
F.3 Front-door boundary robustness
This subsection proves Proposition 3. The proof rests on three exact facts. The plug-in estimator factors as a product of two regression coefficients, the Gaussian delta studentizer is an exact combination of the two regression standard errors, and the second coefficient has an exact Student statistic independent of the design. The boundary limit then follows by comparing the two terms of the combination uniformly in the mediator residual variance, which makes quantitative the mechanism described after the proposition in the main text.
Proof of Proposition 3.
Write and let . The covariance formula factors exactly as , where
is the known-mean least-squares coefficient in the regression of on , and the coefficient vector in the regression of on is
Put
The usual no-intercept regression standard errors are
| (F.1) |
We first identify the Gaussian delta variance exactly. For a positive-definite covariance matrix, let be the regression coefficient of on , let be the coefficient vector in the regression of on , and put
and define and . The influence functions of the two coefficient functionals under known-mean covariance sampling are
The first identity follows by differentiating . The second follows by differentiating the normal equations and applying the Frisch–Waugh residualization. Under a centred Gaussian law, and are independent, and is independent of . Hence,
| (F.2) |
These are identities of rational functions on the positive cone and may therefore be evaluated at . Since , the Gaussian delta standard error defined in (25) satisfies the exact sample identity
| (F.3) |
where the second equality is the algebraic substitution of (F.1).
For completeness, the zero cross term in (F.2) also has an exact finite-sample regression justification. Conditional on , the difference is a linear function of the independent mean-zero errors , so
whereas is -measurable. Thus for every and every interior covariance. For a fixed interior covariance, Gaussian projection formulas give
for all sufficiently large . Hence, is uniformly integrable. The joint delta limit therefore has zero covariance, in agreement with the direct influence-function calculation.
Conditional on , the usual Gaussian regression decomposition and Cochran’s theorem show that the statistic has the law. The conditional law does not depend on , so is independent of the design. Moreover,
with the second chi-squared variable independent of the design. These are the standard Gaussian projection identities (30, Ch. 3). It follows that, uniformly over ,
The first order also follows from the exact moment identity .
If a subsequence has bounded away from zero, the ordinary joint central limit theorem and Slutsky’s theorem prove the result. Consider a subsequence with . Decompose
After division by , the first term is , which converges in distribution to with , and this limit is standard normal by symmetry. Since , the second term is , and the third equals .
The final display shows that the delta studentizer is asymptotically equivalent to the dominant regression studentizer uniformly in , which is the exact balance behind the boundary robustness.
G Detailed numerical experiments
G.1 Data-generating mechanisms and target parameters
We give the complete specifications used for Table 2. The front-door covariance is generated by , and , where is centred Gaussian with unit marginal variances and covariance , and is independent Gaussian with variance . The target is . The proximal covariance is generated by , , and . All shocks in the proximal design are mutually independent centred Gaussian variables and have unit variance except ; the target is one.
G.2 Computation and precision
For each population covariance , we draw the known-mean sample covariance directly from . We compute the plug-in covariance ratio, its Gaussian delta standard error, the standardized denominator and the indicator that the true target belongs to the inverted set in (20). The robust standard deviation is the interquartile range divided by . Each cell uses and replications. The Wald and inversion calculations use the two-sided 95% standard-normal and critical values, respectively. The Monte Carlo standard error attached to a coverage estimate is ; the largest value in the experiment is . No replication fails the validity checks.
| Model | Parameter | Population | Median | Median partial | Median naive |
| Front-door | – | – | |||
| Front-door | – | – | |||
| Front-door | – | – | |||
| Front-door | – | – | |||
| Proximal | |||||
| Proximal | |||||
| Proximal | |||||
| Proximal | |||||
| Proximal |
| Model | Parameter | Robust s.d. | Median s.e. | Median bias |
| Front-door | ||||
| Front-door | ||||
| Front-door | ||||
| Front-door | ||||
| Proximal | ||||
| Proximal | ||||
| Proximal | ||||
| Proximal | ||||
| Proximal |
| Model | Parameter | Wald cov. | Wald MCSE | Inversion cov. | Inversion MCSE |
| Front-door | |||||
| Front-door | |||||
| Front-door | |||||
| Front-door | |||||
| Proximal | |||||
| Proximal | |||||
| Proximal | |||||
| Proximal | |||||
| Proximal |
The front-door median standard errors closely track the empirical robust standard deviations over the entire 1000-fold range of . For the proximal design, agreement is good when is or , but the Gaussian Wald approximation deteriorates once the population denominator diagnostic falls below about . At and , the sample median exceeds its population value because sampling noise is no longer small relative to the denominator, while inversion retains near-nominal coverage.
H Detailed real data experiments
H.1 Variables and preprocessing
We use the public SUPPORT right-heart-catheterization data (8), distributed at https://hbiostat.org/data/repo/rhc.csv and containing patients and variables. Treatment is the indicator of right-heart catheterization on the first study day, recorded in the variable swang1 and coded as one for the value RHC and zero for No RHC. The descriptive outcome is survival time in days truncated at , recorded in t3d30. The treatment proxies are pafi1 and paco21, and the outcome proxies are ph1 and hema1. These six analysis variables are required to be observed and are not imputed. The adjustment set contains age, sex, cat1, cat2, dnr1, surv2md1 and aps1, corresponding to the covariates described in Section 8 of the main text. Missing continuous adjustment covariates are median-imputed. Missing categorical adjustment covariates are mode-imputed and coded by indicator variables with one level omitted. Nonconstant design columns are standardized before an intercept is added, and the standardization does not alter the projection space. The resulting design has rank and leaves residual degrees of freedom.
H.2 Diagnostic formulas and bootstrap
Let denote the residualized treatment, treatment proxy, outcome proxy and outcome. For their empirical second moments, put
so the reported proximal plug-in estimate is . Put , where is the realized adjustment-design rank. Both denominator diagnostics have the form . Define and . The Gaussian version uses , as in (21). The sandwich version uses the centred empirical covariance, with divisor , of the six second-moment scores . The naive and partial indices use , as stated in the main text.
We use nonparametric row-bootstrap replications. Each resample repeats the complete preprocessing and residualization and uses the realized design rank to determine . The resampled design rank is in replications and in replications. The intervals below are percentile intervals.
| Treatment proxy | Outcome proxy | Partial corr. | Estimate | Gaussian | Sandwich |
| pafi1 | ph1 | ||||
| pafi1 | hema1 | ||||
| paco21 | ph1 | ||||
| paco21 | hema1 |
| Treatment proxy | Outcome proxy | Estimate 95% interval | Sandwich- 95% interval |
| pafi1 | ph1 | ||
| pafi1 | hema1 | ||
| paco21 | ph1 | ||
| paco21 | hema1 |
The weak pafi1/ph1 pair is the only allocation for which the bootstrap estimate interval contains zero and the lower endpoint of the sandwich-diagnostic interval is essentially zero. The other three allocations have diagnostic intervals bounded away from zero, although the discrepancy between the Gaussian and sandwich versions remains substantial for paco21/ph1. These comparisons are diagnostic and should not be interpreted as evidence for the untestable proximal bridge assumptions.
I Algebraic obstructions in the mixed branch
I.1 Kernel map and the Hessian obstruction
For a determinant-free mixed solution with self-normalization constant , the kernel-field argument in Supplementary Section D produces a primitive homogeneous polynomial vector . It defines the complex-projective rational map
At every point where has rank two and , the image is the projective kernel line . Theorem 3 and Lemma D.11 show that the unrestricted three-variable converse is equivalent to proving that this map is constant.
The identity does not prove this constancy. The Gordan–Noether cone conclusion holds for forms in at most four variables, but it fails in six variables, which is the dimension of (27; 7). For example, the cubic in variables (32; 17)
has the nonzero polynomial Hessian-kernel vector
Its first derivatives are linearly independent, so is not a cone. This cubic is not claimed to satisfy (4). It shows only that a polynomial Hessian-kernel field need not have a constant direction. Consequently, the missing implication must use more than Hessian degeneracy.
I.2 Representation-theoretic lifting obstruction
A second possible route uses the derived congruence action . For the matrix units ,
Equation (4) therefore implies the polynomial identity
| (I.1) |
This is the image, under polynomial multiplication, of the tensor relation that one would seek from a highest-weight orbit characterization.
The theorem of 25 is a tensor identity in
By contrast, (I.1) lies only in the space of degree- polynomial functions. The multiplication map from the tensor space to degree- polynomials has a nontrivial kernel for . Hence, the scalar polynomial identity does not determine the required tensor identity. A proof by this route would need an additional argument showing that the components lost under multiplication vanish separately.
These two obstructions explain the scope of Theorem 3. They do not provide evidence for a nonflag solution. They identify the remaining task more precisely. One must combine the integrability of the symmetric gradient with the cofactor syzygy or with determinant-boundary information to force the rational kernel map to be constant.