arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2608.19771v1 [stat.ME] 20 Aug 2026

Testing the Validity of Instrumental Variable Sets in Causal Additive Models with Non-Constant Effects

Xichen Guo    Feng Xie    Bingbing Tang    Yan Zeng    Zhang Hao    Zhi Geng    Ruichu Cai       Kun Zhang
Abstract

Instrumental variable (IV) methods are powerful for causal effect estimation with unmeasured confounding, but in practice researchers often face a set of candidate IVs whose validity is difficult to determine from observational data. This paper studies the problem of testing the validity of IV sets under Causal Additive Models with Non-Constant Effects (CAM-NCE). To address this problem, we propose a testable condition, termed the Cross Auxiliary-based independence Test (CAT) condition, for assessing IV set validity from observational data. We show that, under the completeness condition, if the CAT condition is violated, the corresponding set cannot be a valid IV set. Furthermore, under a cross distributional non-degeneracy condition, we establish that the CAT condition becomes both necessary and sufficient for characterizing valid IV sets under CAM-NCE. We then extend the CAT condition to settings with covariates and develop a practical finite-sample algorithm for testing the validity of candidate IV sets. Extensive experiments on synthetic data and three real-world datasets demonstrate the effectiveness and practical utility of the proposed method.

1Department of Applied Statistics, Beijing Technology and Business University, Beijing, China. 2Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China. 3School of Mathematical Sciences, Peking University, Beijing, China. 4School of Computer Science, Guangdong University of Technology, Guangzhou, China. 5Peng Cheng Laboratory, Shenzhen, China. 6Department of Philosophy, Carnegie Mellon University, Pittsburgh, USA. 7Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates. *Corresponding author: Feng Xie (e-mails: fengxie@btbu.edu.cn).

Keywords: Causal inference, instrumental variables, validity, non-constant causal effects, unmeasured confounders.

1 Introduction

Causal effect estimation from observational data is a fundamental problem in modern machine learning and probabilistic inference (58; 16), with broad applications in recommendation systems (13; 76; 43), online systems and sequential decision-making (60; 46; 77), and causality-aware clustering (42; 39; 38; 12). Instrumental variable (IV) methods provide a powerful approach to causal effect estimation by leveraging exogenous variables to mitigate unmeasured confounding. Informally, a valid IV must satisfy three key conditions: (𝒞\mathcal{C}1) it is associated with the treatment (relevance); (𝒞\mathcal{C}2) it affects the outcome only through the treatment (exclusion restriction); and (𝒞\mathcal{C}3) it is independent of unmeasured confounders (exogeneity) (see Figure 1 and Section 2.1 for formal definitions). For instance, in Figure 1, Colonialist Mortality serves as a valid IV for estimating the causal effect of Institutions on Economic Development.

Because of unmeasured confounding, it is generally difficult to determine from observational data alone which variables can serve as valid IVs. In practice, valid IVs may sometimes be justified a priori by domain expertise or historical knowledge. However, such information is often unavailable in many empirical settings, leaving researchers with a set of candidate IVs whose validity is uncertain. Therefore, developing data-driven methods to test IV validity remains a critical yet challenging task (15; 72).

Colonialist Mortality Institutions Economic Development Unmeasured Variables 𝒞1\mathcal{C}1𝒞3\mathcal{C}3𝒞2\mathcal{C}2
Figure 1: Graphical illustration of a valid IV model. Solid arrows denote causal relationships, while dashed arrows denote prohibited relationships for IV validity. The variable Colonialist Mortality serves as a valid IV for the causal effect of Institutions on Economic Development. Unmeasured Variables (e.g., Cultural Difference) act as unmeasured confounders (1). More details are provided in Section 6.

Testing the validity of IVs has attracted considerable attention in recent years. For discrete treatment settings, several approaches address this problem by imposing testable constraints on the joint distribution, including Pearl’s seminal instrumental inequality (56), the generalized instrumental inequality (37), and their extensions (48; 55; 40; 68; 34; 50; 24). Although these methods have been successfully applied in various domains, they typically rely on discrete treatment variables, which limits their applicability in many practical settings—such as studies involving continuous exposures (e.g., vitamin levels (63)).

Another line of research studies IV validity in the Causal Additive Models with Constant Effects (CAM-CE). Existing approaches in this line can be broadly divided into two categories.

  • IV Set. Representative methods include the Proportion of Invalid IVs approaches, such as the Majority Rule constraint, which assumes that more than 50%50\% of the candidate IVs are valid (29; 36; 8; 69; 31), and the Plurality Rule constraint, which requires the number of valid IVs exceeds that of any group of invalid IVs sharing the same ratio-estimator limit (32; 27; 70; 47). Another representative condition is the InSIDE constraint, which assumes that the pleiotropic effects of IVs on the outcome are uncorrelated with their effects on the treatment (7; 41; 59). In addition, the Two Valid IVs constraint requires the existence of at least two valid instruments and further relies on either the Rank-Faithfulness assumption (61; 17) or an algebraic condition discussed in 26.

  • Single IV. Typical approaches include the IV-GIN method, which leverages the Generalized Independent Noise (GIN) condition (73) within linear non-Gaussian causal models (74), and the IV-PIM method, which is based on the Principle of Independent Mechanisms (PIM) (35) and evaluates instrument validity by decomposing the spectral measure of the covariance matrix of the covariates (9).

However, these approaches primarily focus on constant causal effects, and their theoretical guarantees do not directly extend to settings with non-constant causal effects.

Causal additive models with non-constant effects (CAM-NCE) provide a more flexible framework for capturing heterogeneous treatment–outcome relationships. Chu et al. (19) were among the first to study IV validity in this setting by introducing the concept of semi-instruments, showing that semi-instrument validity is testable under additive models. However, their testability condition relies on the prior assumption that the exogeneity condition (𝒞\mathcal{C}3) holds. More recently, 25 proposed the Auxiliary-based Independence Test (AIT) for a single IV and showed that, under certain conditions, a valid IV satisfies the AIT condition. Nevertheless, the AIT condition encounters difficulties in verifying the exclusion restriction (𝒞\mathcal{C}2) (see the illustrative example in Section 3.1 and Proposition 5 in  (25)). Moreover, the exclusion restriction is often difficult to satisfy in practice. For example, in Mendelian randomization studies, genetic variants may have pleiotropic effects on the outcome, thereby violating (𝒞\mathcal{C}2) (10).

Table I: Summary of research coverage on IV validity testing under continuous-variable settings.
Category Constant Effect Non-Constant Effect
Single IV 74; and 9 19; and 25
IV Set 29; 36; 8; 69; 31; 32; 27; 70; 47; 7; 41; 59; 61; 17; and 26

Note: The check mark indicates that studies have addressed the corresponding setting, whereas the cross denotes that research in this setting remains unexplored.

table I summarizes existing studies on IV validity testing under continuous-variable settings, categorized by whether the analysis involves a single IV or a set of IVs, and whether the underlying causal effect is constant or non-constant. As shown in the table, prior research has primarily focused on either single IV settings or IV sets under constant-effect assumptions. However, testing the validity of IV sets—particularly when the treatment–outcome relationship exhibits non-constant causal effects—remains underexplored. To bridge this gap, this paper investigates the problem of testing the validity of IV sets within the CAM-NCE framework, where candidate instruments may violate either the exclusion restriction or the exogeneity condition. Specifically, our main contributions are as follows:

  1. 1.

    We introduce a testable condition, termed the Cross Auxiliary-based Independence Test (CAT) condition, for assessing the validity of IV sets within CAM-NCE.

  2. 2.

    We establish that, under the completeness condition (Assumption 1) and the cross distributional non-degeneracy condition (Assumption 2), the CAT condition is necessary and sufficient for detecting all invalid IV sets under CAM-NCE.

  3. 3.

    We develop a practical testing procedure for the CAT condition that accounts for covariates and finite-sample considerations.

  4. 4.

    We empirically validate the proposed method through extensive experiments on both synthetic and real-world datasets, demonstrating its effectiveness in detecting invalid IV sets.

2 Preliminaries and Problem Definition

2.1 Notations and Definitions

In this paper, a causal system is represented by a directed acyclic graph (DAG) 𝒢\mathcal{G}, in which nodes correspond to random variables and directed edges represent direct causal influences between them. For brevity, we use “w.r.t.” to denote “with respect to”. Throughout this paper, the major symbols and notations are summarized in table II.

Table II: Symbols and notations used in this paper.
Symbol Description
XX A treatment (exposure) variable
YY An outcome variable
ZiZ_{i} A candidate (potential) IV
𝐙\mathbf{Z} A candidate (potential) IV set
𝐔\mathbf{U} Unmeasured confounders between XX and YY
𝐖\mathbf{W} Covariates
𝒳\mathcal{X} The residual of variable XX after regressing on covariates 𝐖\mathbf{W}
𝒵i\mathcal{Z}_{i} The residual of variable ZiZ_{i} after regressing on covariates 𝐖\mathbf{W}
|𝐙||\mathbf{Z}| The number of variables in set 𝐙\mathbf{Z}
ε\varepsilon_{*} The noise term of a variable
𝔼(X)\mathbb{E}(X) The expected value of random variable XX
Cov(X,Y)\operatorname{Cov}(X,Y) The covariance between random variables XX and YY
Var(X)\operatorname{Var}(X) The variance of random variable XX
AB{A}\mathrel{\perp\mspace{-10mu}\perp}{B} AA is statistically independent of BB
A B{A}\mathchoice{\mathrel{\hbox to0.0pt{\kern 20.68047pt\kern-4.88191pt$\displaystyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 20.68047pt\kern-4.88191pt$\textstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 15.65509pt\kern-4.23051pt$\scriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 9.63855pt\kern-3.03471pt$\scriptscriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{B} AA is statistically dependent on BB
\mathbb{R} The field of real numbers
||\mathbb{R}^{|*|}\rightarrow\mathbb{R} A mapping from ||\mathbb{R}^{|*|} to \mathbb{R}
fbias(X)f_{bias}(X) The bias between estimated causal effect of XX on YY and ground-truth causal effect of XX on YY
𝒱XY||Zi\mathcal{V}_{X\to Y||Z_{i}} The auxiliary variable of causal relationship XYX\to Y relative to ZiZ_{i}. We often use 𝒱i\mathcal{V}_{i} as a shorthand when there is no ambiguity
Q^\widehat{Q} The sample estimate of the corresponding population quantity QQ
Z2Z_{2}Z1Z_{1}XXYY𝐔\mathbf{U}(a) Valid IV set
Z2Z_{2}Z1Z_{1}XXYY𝐔\mathbf{U}(b) Invalid IV set violating 𝒞2\mathcal{C}2
Z2Z_{2}Z1Z_{1}XXYY𝐔\mathbf{U}(c) Invalid IV set violating 𝒞3\mathcal{C}3
Figure 2: Graphical illustration of IV set 𝐙={Z1,Z2}\mathbf{Z}=\{Z_{1},Z_{2}\} models, where 𝐔\mathbf{U} denotes unmeasured confounders. (a) 𝐙\mathbf{Z} is a valid IV set. (b) 𝐙\mathbf{Z} is an invalid IV set, violating Condition 𝒞2\mathcal{C}2 because of the edge Z1YZ_{1}\to Y. (c) 𝐙\mathbf{Z} is an invalid IV set, violating Condition 𝒞3\mathcal{C}3 because of the edge 𝐔Z2\mathbf{U}\to Z_{2}.

We next introduce the formal definitions of an IV and an IV set, which will be used throughout the paper.

Definition 1 (Instrumental Variable (IV) (57)).

A variable ZiZ_{i} is said to be an IV w.r.t. the causal relation XYX\to Y if it satisfies the following three conditions:

𝒞1\mathcal{C}1. (Relevance)

ZiZ_{i} is associated with the treatment XX, i.e., Zi XZ_{i}\mathchoice{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\displaystyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\textstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 16.49544pt\kern-4.45831pt$\scriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 13.51866pt\kern-3.95834pt$\scriptscriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}X.

𝒞2\mathcal{C}2. (Exclusion Restriction)

ZiZ_{i} does not directly affect the outcome YY, i.e., ZiY|{X,𝐔}Z_{i}\mathrel{\perp\mspace{-10mu}\perp}Y\mid\{X,\mathbf{U}\}.

𝒞3\mathcal{C}3. (Exogeneity)

ZiZ_{i} is independent of the unmeasured confounders 𝐔\mathbf{U}, i.e., Zi𝐔Z_{i}\mathrel{\perp\mspace{-10mu}\perp}\mathbf{U}.

Definition 2 (IV Set).

A set of variables 𝐙\mathbf{Z} is said to be a valid IV set w.r.t.XYX\to Y if every nonempty subset 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} satisfies Conditions 𝒞1\mathcal{C}1𝒞3\mathcal{C}3 of Definition 1, with ZiZ_{i} replaced by 𝐙\mathbf{Z}^{\prime}. Otherwise, 𝐙\mathbf{Z} is referred to as an invalid IV set w.r.t.XYX\to Y.

As illustrated in Figure 2, Figure 2(a) shows a valid IV set {Z1,Z2}\{Z_{1},Z_{2}\}, where both variables satisfy Conditions 𝒞1\mathcal{C}1𝒞3\mathcal{C}3. By contrast, Figure 2(b) shows an invalid IV set in which Z1Z_{1} violates the exclusion restriction (𝒞2\mathcal{C}2) due to the direct edge Z1YZ_{1}\to Y, while Figure 2(c) shows another invalid IV set in which Z2Z_{2} violates exogeneity (𝒞3\mathcal{C}3) due to its dependence on the unmeasured confounders 𝐔\mathbf{U}.

2.2 Causal Additive Models with Non-Constant Effects

In this paper, we focus on Causal Additive Models with Non-Constant Effects (CAM-NCE), involving the treatment XX, the outcome YY, a candidate IV set 𝐙\mathbf{Z}, and unmeasured confounders 𝐔\mathbf{U}. Without loss of generality, all variables are assumed to be mean-centered. We first present the covariate-free case for simplicity, and discuss the extension to baseline covariates in section 4.1. Specifically, the data generation process under CAM-NCE can be expressed as:

X\displaystyle X =gX(𝐙)+φX(𝐔)+εX,\displaystyle=g_{X}(\mathbf{Z})+\varphi_{X}(\mathbf{U})+\varepsilon_{X}, (1)
Y\displaystyle Y =f(X)+gY(𝐙~)+φY(𝐔)+εY,\displaystyle=f(X)+g_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y},

where the function f()f(\cdot) represents the true but unknown causal effect of interest. The functions f()f(\cdot), g()g_{*}(\cdot), and φ()\varphi_{*}(\cdot) are assumed to be smooth functions defined on appropriate domains, i.e., f:f:\mathbb{R}\!\to\!\mathbb{R}, gX:|𝐙|g_{X}:\mathbb{R}^{|\mathbf{Z}|}\!\to\!\mathbb{R}, gY:|𝐙~|g_{Y}:\mathbb{R}^{|\widetilde{\mathbf{Z}}|}\!\to\!\mathbb{R}, and φ:|𝐔|\varphi_{*}:\mathbb{R}^{|\mathbf{U}|}\!\to\!\mathbb{R}. The noise terms εX\varepsilon_{X} and εY\varepsilon_{Y} are assumed to be mutually independent. The nonzero function gY(𝐙~)g_{Y}(\widetilde{\mathbf{Z}}) indicates that the subset 𝐙~𝐙\widetilde{\mathbf{Z}}\!\subseteq\!\mathbf{Z} directly affects the outcome YY, thereby violating the exclusion restriction condition (𝒞2\mathcal{C}2). Furthermore, if there exists any Zi𝐙Z_{i}\!\in\!\mathbf{Z} that is statistically dependent on 𝐔\mathbf{U}, this implies a violation of the exogeneity condition (𝒞3\mathcal{C}3).

Note that, even given a valid IV, identifying the nonparametric structural function f()f(\cdot) generally requires additional conditions. Following the standard additive nonparametric IV literature, we assume that a solution exists and impose the following completeness condition to ensure uniqueness given a valid IV (53; 2; 21; 18; 54; 62; 6).

Assumption 1 (Completeness Condition).

Given a valid IV Zi𝐙Z_{i}\in\mathbf{Z}, for any measurable function ψ(X)\psi(X) satisfying 𝔼[|ψ(X)|]<+\mathbb{E}[|\psi(X)|]<+\infty, 𝔼[ψ(X)|Zi]=0\mathbb{E}[\psi(X)|Z_{i}]=0 almost surely if and only if ψ(X)=0\psi(X)=0 almost surely.

The completeness condition is generic, in the sense that it holds for “most” s(X|Zi)s(X|Z_{i})—where s(x|zi)s(x|z_{i}) denotes the conditional density of XX given ZiZ_{i}—if it holds for one (54; 4). In particular, the finite-support case and models belonging to exponential families (such as Gaussian, Poisson, Binomial, or certain multivariate extensions) are known to satisfy completeness (53; 33). Further sufficient conditions for completeness have been established in 23; 4; and 33.

Remark 1.

We highlight the following points regarding the proposed CAM-NCE model:

  1. 1.

    Relation to constant-effect models. When the functions f()f(\cdot), g()g_{*}(\cdot), and φ()\varphi_{*}(\cdot) are linear, the above model reduces to the Causal Additive Model with Constant Effects (CAM-CE), such as the Additive Linear, Constant-Effects Model (ALICE) (29; 36; 8; 69; 31; 32; 27; 70; 47; 7; 41; 59; 61; 17). In this work, we investigate a more challenging and general setting where f()f(\cdot), g()g_{*}(\cdot), and φ()\varphi_{*}(\cdot) can be nonlinear.

  2. 2.

    Conditions to be assessed. Condition 𝒞1\mathcal{C}1 (relevance) can be readily assessed through standard statistical dependence tests. Hence, our primary focus lies in addressing the remaining Conditions 𝒞2\mathcal{C}2 and 𝒞3\mathcal{C}3. Notably, causal discovery methods based on conditional independence tests, such as the FCI (Fast Causal Inference) algorithm (64) and its extensions (20; 3), often yield a fully connected subgraph over {X,Y,Zi}\{X,Y,Z_{i}\} in the presence of unmeasured confounders 𝐔\mathbf{U} (since Zi Y|XZ_{i}\mathchoice{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\displaystyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\textstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 16.49544pt\kern-4.45831pt$\scriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 13.51866pt\kern-3.95834pt$\scriptscriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}Y\mid X). As a result, verifying Conditions 𝒞2\mathcal{C}2𝒞3\mathcal{C}3 and identifying valid IVs remain challenging tasks.

  3. 3.

    Explicit representation of confounders. Several previous studies implicitly absorb unmeasured confounding into composite noise terms δ=φX(𝐔)+εX\delta=\varphi_{X}(\mathbf{U})+\varepsilon_{X} and ϵ=φY(𝐔)+εY\epsilon=\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}, which are typically correlated (e.g., 53; 28; 6; 62). In contrast, we explicitly represent 𝐔\mathbf{U} to facilitate subsequent theoretical derivations and analysis.

Our Goal. The goal of this work is to develop a data-driven framework for testing the validity of IV sets under CAM-NCE, where Assumption 1 holds. Specifically, given a candidate IV set 𝐙\mathbf{Z} for the causal relation XYX\to Y, we aim to assess whether the variables in 𝐙\mathbf{Z} satisfy Conditions 𝒞1\mathcal{C}1𝒞3\mathcal{C}3.

3 CAT Condition for Testing IV Set Validity

In this section, we introduce a testable criterion for assessing the validity of IV sets under CAM-NCE, termed the Cross Auxiliary-based Independence Test (CAT) condition. We first show that, under the completeness condition (Assumption 1), the CAT condition is necessary for IV set validity. We then present a counterexample showing that completeness alone is insufficient to rule out all invalid IV sets. Finally, by imposing an additional cross distributional non-degeneracy condition (Assumption 2), we establish that the CAT condition becomes both necessary and sufficient for characterizing valid IV sets.

3.1 CAT Condition: Definition and Necessity

We first introduce the key concept of the auxiliary variable together with the definition of the CAT condition for an IV set, which characterizes cross-independence relationships between the auxiliary variable and another distinct IV.

Definition 3 (Auxiliary Variable).

Let XX, YY, and Zi𝐙Z_{i}\in\mathbf{Z} denote the treatment, outcome, and candidate IV, respectively. The auxiliary variable for the causal relationship XYX\to Y relative to ZiZ_{i} is defined as

𝒱XY||ZiYhi(X),\displaystyle\mathcal{V}_{X\to Y||Z_{i}}\coloneqq Y-h_{i}(X), (2)

where hi()h_{i}(\cdot) is a nonzero function satisfying 𝔼[𝒱XY||Zi|Zi]=0\mathbb{E}[\mathcal{V}_{X\to Y||Z_{i}}|Z_{i}]={0}.

The notion of an auxiliary variable, or related residual-type constructions, has been used in various contexts (22; 14; 11; 75; 25). Different from these works, we use auxiliary variables to construct cross-independence relations among candidate IVs. To the best of our knowledge, such cross-independence relations have not been used to assess IV set validity under CAM-NCE.

Definition 4 (CAT Condition).

Let XX, YY, and 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} denote the treatment, outcome, and a candidate IV set, respectively. We say that {X,Y𝐙}\{X,Y\|\ \mathbf{Z}^{\prime}\} satisfies the Cross Auxiliary-based Independence Test (CAT) condition if and only if, for every pair of distinct IVs {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime}, the following cross-independence relationships hold:

𝒱XY||ZiZj, and 𝒱XY||ZjZi.\displaystyle\mathcal{V}_{X\to Y||Z_{i}}\mathrel{\perp\mspace{-10mu}\perp}Z_{j},\text{ and }\mathcal{V}_{X\to Y||Z_{j}}\mathrel{\perp\mspace{-10mu}\perp}Z_{i}. (3)

Intuitively, the CAT condition performs a cross-check among candidate IVs: the auxiliary variable constructed using one candidate IV is tested for independence from another candidate IV in the set. For a pair {Zi,Zj}\{Z_{i},Z_{j}\}, CAT requires both 𝒱XY|ZiZj\mathcal{V}_{X\to Y\|Z_{i}}\mathrel{\perp\mspace{-10mu}\perp}Z_{j} and 𝒱XY|ZjZi\mathcal{V}_{X\to Y\|Z_{j}}\mathrel{\perp\mspace{-10mu}\perp}Z_{i} to hold.

We next provide a simple example to illustrate how the CAT condition detects violations of IV validity.

Example 1 (Intuitive Example of the CAT Condition).

Consider the two causal graphs shown in Figure 2(a) and Figure 2(b). Suppose that the corresponding data-generating mechanisms are given as follows:

  • Figure 2(a): U=εUU=\varepsilon_{U}, Z1=εZ1Z_{1}=\varepsilon_{Z_{1}}, Z2=εZ2Z_{2}=\varepsilon_{Z_{2}}, X=2Z1+Z2+1.5U+εXX=2Z_{1}+Z_{2}+1.5U+\varepsilon_{X}, and Y=exp(X)+0.5U+εY.Y=\exp(X)+0.5U+\varepsilon_{Y}.

  • Figure 2(b): U=εUU=\varepsilon_{U}, Z1=εZ1Z_{1}=\varepsilon_{Z_{1}}, Z2=εZ2Z_{2}=\varepsilon_{Z_{2}}, X=2Z1+Z2+1.5U+εXX=2Z_{1}+Z_{2}+1.5U+\varepsilon_{X}, and Y=exp(X)+5Z1+0.5U+εYY=\exp(X)+5Z_{1}+0.5U+\varepsilon_{Y}.

Refer to caption
Figure 3: Scatter plots of candidate IV Z1Z_{1} and auxiliary variable 𝒱XY|Z2\mathcal{V}_{X\to Y\|Z_{2}} for Subgraphs (a) and (b) in Example 1 when all noise terms follow a standard Gaussian distribution.

Assume that all noise terms are mutually independent and follow a standard Gaussian distribution. Let 𝐙={Z1,Z2}\mathbf{Z}^{\prime}=\{Z_{1},Z_{2}\}.

In Figure 2(a), both Z1Z_{1} and Z2Z_{2} are valid IVs. For either reference IV, the auxiliary variable removes the causal contribution exp(X)\exp(X) and becomes

𝒱XY|Z1=𝒱XY|Z2=Yexp(X)=0.5U+εY.\mathcal{V}_{X\to Y\|Z_{1}}=\mathcal{V}_{X\to Y\|Z_{2}}=Y-\exp(X)=0.5U+\varepsilon_{Y}.

Since Z1Z_{1}, Z2Z_{2}, and UU are mutually independent, we have 𝒱XY|Z1Z2\mathcal{V}_{X\to Y\|Z_{1}}\mathrel{\perp\mspace{-10mu}\perp}Z_{2} and 𝒱XY|Z2Z1\mathcal{V}_{X\to Y\|Z_{2}}\mathrel{\perp\mspace{-10mu}\perp}Z_{1}. Hence, 𝐙\mathbf{Z}^{\prime} satisfies the CAT condition. This is visualized in Figure 3, where Z1Z_{1} and 𝒱XY|Z2\mathcal{V}_{X\to Y\|Z_{2}} show no apparent dependence in the valid-IV-set case.

In Figure 2(b), however, Z1Z_{1} directly affects YY and violates the exclusion restriction, while Z2Z_{2} remains valid. Using Z2Z_{2} as the reference IV, we obtain

𝒱XY|Z2=Yexp(X)=5Z1+0.5U+εY,\mathcal{V}_{X\to Y\|Z_{2}}=Y-\exp(X)=5Z_{1}+0.5U+\varepsilon_{Y},

which is dependent on Z1Z_{1}. Thus, 𝒱XY|Z2 Z1\mathcal{V}_{X\to Y\|Z_{2}}\mathchoice{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\displaystyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\textstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 16.49544pt\kern-4.45831pt$\scriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 13.51866pt\kern-3.95834pt$\scriptscriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}Z_{1}, and the CAT condition is violated. As shown in Figure 3, this invalid-IV-set case exhibits a visible dependence pattern between Z1Z_{1} and 𝒱XY|Z2\mathcal{V}_{X\to Y\|Z_{2}}, induced by the direct effect Z1YZ_{1}\to Y.

It is worth noting that the validity of a single IV, such as Z1Z_{1} in Figure 2(b), cannot in general be verified from the observed distribution of (X,Y,Z1)(X,Y,Z_{1}) alone. Indeed, the IV assumptions impose no testable constraints on this joint distribution, and different causal structures may induce the same observational distribution (see Section 3 of 19, and Proposition 3 of  25). This motivates the need for testable criteria that exploit relations among multiple candidate IVs.

Building on the above intuition, we next establish that the CAT condition provides a necessary condition for IV set validity under CAM-NCE.

Theorem 1 (Necessary Condition for IV Set Validity).

Let XX, YY, and 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} be the treatment, outcome, and candidate IV set in CAM-NCE, respectively. Suppose that XX, YY, and 𝐙\mathbf{Z}^{\prime} are statistically dependent, and that Assumption 1 holds. If the candidate set 𝐙\mathbf{Z}^{\prime} is a valid IV set w.r.t.XYX\to Y, then {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} satisfies the CAT condition.

Proof sketch.

Since 𝐙\mathbf{Z}^{\prime} is a valid IV set, each Zi𝐙Z_{i}\in\mathbf{Z}^{\prime} is a valid IV. Under Assumption 1, the conditional moment restriction 𝔼[Yhi(X)Zi]=0\mathbb{E}[Y-h_{i}(X)\mid Z_{i}]=0 admits a unique solution. By the standard identification result in nonparametric IV models (53; 54; 23; 33), this solution coincides with the true causal effect function f()f(\cdot), i.e., hi(X)=f(X)h_{i}(X)=f(X). Hence, for any pair of distinct IVs {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime}, the auxiliary variables satisfy

𝒱XY|Zi=𝒱XY|Zj=gY(𝐙~)+φY(𝐔)+εY.\mathcal{V}_{X\to Y\|Z_{i}}=\mathcal{V}_{X\to Y\|Z_{j}}=g_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}.

Here, because 𝐙\mathbf{Z}^{\prime} is a valid IV set, ZiZ_{i} and ZjZ_{j} do not directly affect YY and are independent of 𝐔\mathbf{U}. Moreover, under CAM-NCE, they are also independent of the remaining variables and noise terms involved in gY(𝐙~)+εYg_{Y}(\widetilde{\mathbf{Z}})+\varepsilon_{Y}. Using the fact that measurable functions of disjoint subsets of mutually independent random variables are independent (see Theorem 4 and Lemma 1 in Appendix), we have

𝒱XY|ZiZj,𝒱XY|ZjZi.\mathcal{V}_{X\to Y\|Z_{i}}\mathrel{\perp\mspace{-10mu}\perp}Z_{j},\quad\mathcal{V}_{X\to Y\|Z_{j}}\mathrel{\perp\mspace{-10mu}\perp}Z_{i}.

Thus, every pair of distinct IVs in 𝐙\mathbf{Z}^{\prime} satisfies the CAT condition. Consequently, {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} satisfies the CAT condition. The full proof is provided in Appendix B.1. ∎

Theorem 1 states that if {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} violates the CAT condition, then the candidate IV set 𝐙\mathbf{Z}^{\prime} w.r.t.XYX\to Y is invalid. Otherwise, 𝐙\mathbf{Z}^{\prime} may or may not be valid.

3.2 Sufficient Condition for Characterizing IV Set Validity

In the previous section, we have shown that the CAT condition is necessary for IV set validity. A natural question is whether, under Assumption 1, every invalid IV set necessarily violates the CAT condition. Unfortunately, the answer is negative. To clarify this issue, we present a counterexample below.

Example 2.

(Counterexample) Let Z1=εZ1Z_{1}=\varepsilon_{Z_{1}}, Z2=εZ2Z_{2}=\varepsilon_{Z_{2}}, X=gX1(Z1)+gX2(Z2)+φX(𝐔)+εXX=g_{X1}(Z_{1})+g_{X2}(Z_{2})+\varphi_{X}(\mathbf{U})+\varepsilon_{X} and Y=f(X)+gY1(Z1)+gY2(Z2)+φY(𝐔)+εYY=f(X)+g_{Y1}(Z_{1})+g_{Y2}(Z_{2})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}, with gY1(Z1)=agX1(Z1)+b1g_{Y1}(Z_{1})=a\cdot g_{X1}(Z_{1})+b_{1}, gY2(Z2)=agX2(Z2)+b2g_{Y2}(Z_{2})=a\cdot g_{X2}(Z_{2})+b_{2}, where constant a0a\neq 0. In this model, both Z1Z_{1} and Z2Z_{2} directly affect YY and hence violate the exclusion restriction. However, the same observational distribution can be generated by an alternative causal structure in which {Z1,Z2}\{Z_{1},Z_{2}\} forms a valid IV set. Specifically, define f(X)=f(X)+aXf^{\prime}(X)=f(X)+aX, Z1=Z1Z_{1}^{\prime}=Z_{1}, Z2=Z2Z_{2}^{\prime}=Z_{2}, X=XX^{\prime}=X, and Y=f(X)+φY(𝐔)+εYa{φX(𝐔)+εX}+b1+b2Y^{\prime}=f^{\prime}(X^{\prime})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}-a\cdot\{\varphi_{X}(\mathbf{U})+\varepsilon_{X}\}+b_{1}+b_{2}. Then, we have

Y\displaystyle Y^{\prime} =f(X)+aX+φY(𝐔)+εYa{φX(𝐔)+εX}+b1+b2\displaystyle=f(X)+aX+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}-a\{\varphi_{X}(\mathbf{U})+\varepsilon_{X}\}+b_{1}+b_{2}
=f(X)+a{gX1(Z1)+gX2(Z2)}+φY(𝐔)+εY+b1+b2\displaystyle=f(X)+a\{g_{X1}(Z_{1})+g_{X2}(Z_{2})\}+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}+b_{1}+b_{2}
=f(X)+gY1(Z1)+gY2(Z2)+φY(𝐔)+εY=Y.\displaystyle=f(X)+g_{Y1}(Z_{1})+g_{Y2}(Z_{2})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}=Y.

Thus, (X,Y,Z1,Z2)(X,Y,Z_{1},Z_{2}) and (X,Y,Z1,Z2)(X^{\prime},Y^{\prime},Z_{1}^{\prime},Z_{2}^{\prime}) have the same observational distribution, although {Z1,Z2}\{Z_{1},Z_{2}\} is invalid in the original structure but valid in the alternative structure. This shows that Assumption 1 alone cannot distinguish all invalid IV sets.

The above example shows that Assumption 1 alone does not guarantee that the CAT condition can rule out all invalid IV sets. To obtain a sufficient condition, we introduce an additional non-degeneracy condition that excludes those cases and makes IV set validity testable under the CAM-NCE.

Assumption 2 (Cross Distributional Non-degeneracy Condition).

For any invalid candidate IV set 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} with |𝐙|2|\mathbf{Z}^{\prime}|\geq 2, there exists a pair of distinct candidate IVs {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime} such that the joint densities p(𝒱XY|Zi,Zj)p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j}) and p(𝒱XY|Zj,Zi)p(\mathcal{V}_{X\to Y\|Z_{j}},Z_{i}) are twice continuously differentiable. Moreover, at least one of the following cross second-order partial derivatives,

2logp(𝒱XY||Zi,Zj)𝒱XY||ZiZj or 2logp(𝒱XY||Zj,Zi)𝒱XY||ZjZi\frac{\partial^{2}\operatorname{log}p(\mathcal{V}_{X\to Y||Z_{i}},Z_{j})}{\partial\mathcal{V}_{X\to Y||Z_{i}}\partial Z_{j}}\text{ or }\frac{\partial^{2}\operatorname{log}p(\mathcal{V}_{X\to Y||Z_{j}},Z_{i})}{\partial\mathcal{V}_{X\to Y||Z_{j}}\partial Z_{i}}

is non-zero on a set with non-zero Lebesgue measure, where 𝒱XY||Zi=gY(𝐙~)+φY(𝐔)+εYfbiasi(X)\mathcal{V}_{X\to Y||Z_{i}}=g_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}-f_{bias}^{i}(X), and fbiasi(X)=hi(X)f(X)f_{bias}^{i}(X)=h_{i}(X)-f(X).

Assumption 2 is a natural condition that one expects to hold to identify the invalid IV set. Its intuition comes from the linear separability of the logarithm of the joint density of independent variables. Specifically, for a set of independent random variables with a twice-differentiable joint density, the Hessian matrix of the log-density is diagonal (45). Applying this property to 𝒱XY|Zi\mathcal{V}_{X\to Y\|Z_{i}} and ZjZ_{j}, if 𝒱XY|ZiZj\mathcal{V}_{X\to Y\|Z_{i}}\mathrel{\perp\mspace{-10mu}\perp}Z_{j}, then their joint density factorizes as p(𝒱XY|Zi,Zj)=p(𝒱XY|Zi)p(Zj)p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j})=p(\mathcal{V}_{X\to Y\|Z_{i}})\cdot p(Z_{j}). Equivalently, the log-density is additively separable, and hence the corresponding off-diagonal Hessian entry vanishes: 2logp(𝒱XY|Zi,Zj)𝒱XY|ZiZj=0\frac{\partial^{2}\log p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j})}{\partial\mathcal{V}_{X\to Y\|Z_{i}}\,\partial Z_{j}}=0. The same argument applies after exchanging ZiZ_{i} and ZjZ_{j}. Assumption 2 requires that, for an invalid IV set, such cross derivatives do not vanish for at least one cross pair, thereby making the violation detectable through the CAT condition.

Remark 2.

Assumption 2, referred to as the Cross Distributional Non-degeneracy Condition, can be viewed as a distributional analogue of the Algebraic Equation Condition in 26. While Guo et al.’s condition is derived from an explicit algebraic expansion under constant-effect models, our condition is formulated through cross second-order derivatives of joint log-densities and does not require an invertible transformation between observed variables and latent noise terms. This makes the proposed condition applicable to the more general CAM-NCE framework.

The following example gives a concrete illustration of how a violation of IV validity can induce a nonzero cross second-order derivative in the joint density. For clarity, we present the calculation in a constant-effect setting, where closed-form expressions are available. The example is intended to illustrate the intuition behind the more general non-constant-effect case.

Example 3 (Illustration of the Cross Distributional Non-degeneracy Condition).

Consider the following data-generating process: U=εUU=\varepsilon_{U}, Z1=εZ1Z_{1}=\varepsilon_{Z_{1}}, Z2=εZ2Z_{2}=\varepsilon_{Z_{2}}, X=Z1+Z2+Z12+Z22+U+εXX=Z_{1}+Z_{2}+Z_{1}^{2}+Z_{2}^{2}+U+\varepsilon_{X}, and Y=X+2Z1Z2+U+εYY=X+2Z_{1}Z_{2}+U+\varepsilon_{Y}, where εZ1,εZ2,εU,εX,εYindN(0,1)\varepsilon_{Z_{1}},\varepsilon_{Z_{2}},\varepsilon_{U},\varepsilon_{X},\varepsilon_{Y}\stackrel{{\scriptstyle\mathrm{ind}}}{{\sim}}N(0,1). Here, the candidate IVs Z1Z_{1} and Z2Z_{2} are invalid because they directly affect YY through the term 2Z1Z22Z_{1}Z_{2}, thereby violating the exclusion restriction.

For Z1Z_{1}, by the definition of the auxiliary variable, 𝒱1=Yβ1X=2Z1Z2+U+εY\mathcal{V}_{1}=Y-\beta_{1}X=2Z_{1}Z_{2}+U+\varepsilon_{Y}, where β1=Cov(Y,Z1)Cov(X,Z1)=1\beta_{1}=\frac{\operatorname{Cov}(Y,Z_{1})}{\operatorname{Cov}(X,Z_{1})}=1. The joint density of (𝒱1,Z2)(\mathcal{V}_{1},Z_{2}) can be factorized as p(𝒱1,Z2)=p(𝒱1Z2)p(Z2)p(\mathcal{V}_{1},Z_{2})=p(\mathcal{V}_{1}\mid Z_{2})p(Z_{2}), where Z2N(0,1)Z_{2}\sim N(0,1). Conditional on Z2=z2Z_{2}=z_{2}, since Z1Z_{1}, UU, and εY\varepsilon_{Y} are mutually independent standard normal variables, we have 𝒱1|Z2=z2N(0,4z22+2)\mathcal{V}_{1}\mid Z_{2}=z_{2}\sim N(0,4z_{2}^{2}+2). Therefore, logp(𝒱1,z2)=logp(𝒱1|z2)p(z2)=log(2π)z22212log(4z22+2)𝒱122(4z22+2)\log p(\mathcal{V}_{1},z_{2})=\log p(\mathcal{V}_{1}|z_{2})p(z_{2})=-\log(2\pi)-\frac{z_{2}^{2}}{2}-\frac{1}{2}\log(4z_{2}^{2}+2)-\frac{\mathcal{V}_{1}^{2}}{2(4z_{2}^{2}+2)}. Then, 2logp(𝒱1,z2)𝒱1z2=8𝒱1z2(4z22+2)2\frac{\partial^{2}\log p(\mathcal{V}_{1},z_{2})}{\partial\mathcal{V}_{1}\,\partial z_{2}}=\frac{8\mathcal{V}_{1}z_{2}}{(4z_{2}^{2}+2)^{2}}. Since this quantity is nonzero whenever 𝒱1z20\mathcal{V}_{1}z_{2}\neq 0, the cross second-order partial derivative is nonzero on a set with nonzero Lebesgue measure. Therefore, the candidate IV set {Z1,Z2}\{Z_{1},Z_{2}\} satisfies the cross distributional non-degeneracy condition (Assumption 2).

To better understand the Cross Distributional Non-degeneracy Condition (Assumption 2), we next provide a sufficient condition under which this assumption fails.

Proposition 1 (Sufficient Condition for Violation of Assumption 2).

Let XX, YY, and 𝐙𝐙{\mathbf{Z}}^{\prime}\subseteq\mathbf{Z} be the treatment, outcome, and a candidate IV set in CAM-NCE, respectively. Suppose that each Zi𝐙Z_{i}\in{\mathbf{Z}}^{\prime} is relevant and exogenous but violates only the exclusion restriction. Assume further that, for 𝐙\mathbf{Z}^{\prime}, its direct causal effect on YY is an affine transformation of its direct causal effect on XX, namely, gY(𝐙)=agX(𝐙)+bg_{Y}({\mathbf{Z}}^{\prime})=a\cdot g_{X}({\mathbf{Z}}^{\prime})+b, where aa is a nonzero constant, and bb is a constant. Then Assumption 2 fails for the invalid IV set 𝐙{\mathbf{Z}}^{\prime}. This special case is illustrated graphically in Figure 4.

𝐔\mathbf{U}𝐙{\mathbf{Z}}^{\prime} 𝐙𝐙\mathbf{Z}\setminus{\mathbf{Z}}^{\prime} XXYYgX(𝐙)g_{X}({\mathbf{Z}}^{\prime})gY(𝐙)g_{Y}({\mathbf{Z}}^{\prime})
Figure 4: Graphical illustration of an invalid IV set 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z}, where every IV in 𝐙\mathbf{Z}^{\prime} violates only the exclusion restriction condition (𝒞2\mathcal{C}2), while all variables in 𝐙𝐙\mathbf{Z}\setminus\mathbf{Z}^{\prime} are valid IVs. For the invalid IV set 𝐙\mathbf{Z}^{\prime}, the direct effects satisfy gY(𝐙)=agX(𝐙)+bg_{Y}({\mathbf{Z}}^{\prime})=a\cdot g_{X}({\mathbf{Z}}^{\prime})+b, where a0a\neq 0.
Proof.

See Appendix B.2 for its proof. ∎

As shown in Proposition 1, when Assumption 2 is violated, certain IV sets may fail to satisfy the exclusion restriction, leading to non-identifiability of causal effects. We will now show that under additional Assumption 2, IV sets can be uniquely identified.

Proposition 2 (Sufficient Condition for IV Set Validity).

Let XX, YY, and 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} be the treatment, outcome, and candidate IV set in CAM-NCE, respectively. Suppose that XX, YY, and 𝐙\mathbf{Z}^{\prime} are statistically dependent, and that Assumptions 1 and 2 hold. If the candidate IV set 𝐙\mathbf{Z}^{\prime} is invalid, then {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} violates the CAT condition.

Proof.

See Appendix B.3 for its proof. ∎

Proposition 2 shows that, under Assumptions 1 and 2, every invalid IV set violates the CAT condition. Combining this result with Theorem 1, we obtain the following necessary and sufficient characterization of IV set validity under CAM-NCE.

Theorem 2 (Necessary and Sufficient Condition for IV Set Validity).

Let XX, YY, and 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} be the treatment, outcome, and candidate IV set in CAM-NCE, respectively. Suppose that XX, YY, and 𝐙\mathbf{Z}^{\prime} are statistically dependent, and that Assumptions 1 and 2 hold. The candidate IV set 𝐙\mathbf{Z}^{\prime} is a valid IV set relative to XYX\to Y if and only if {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} satisfies the CAT condition.

Proof.

See Appendix B.4 for its proof. ∎

The theoretical implications of the CAT condition are summarized in Figure 5.

Theoretical Role of the CAT ConditionCompleteness Condition(Assumption 1)Valid IV SetCAT Condition HoldsNecessary ConditionCompleteness Condition (Assumption 1) ++Cross Distributional Non-degeneracy Condition (Assumption 2) Valid IV SetCAT Condition HoldsNecessary & Sufficient Condition
Figure 5: Flowchart of the theoretical implications of the CAT condition.

As shown in Figure 5, the CAT condition is necessary for IV set validity under the completeness condition, and becomes both necessary and sufficient when the Cross Distributional Non-degeneracy Condition is additionally imposed.

4 Practical Testing Algorithm from Data

In this section, we first extend the CAT condition to settings with baseline covariates, where IV validity is assessed after adjusting for observed covariates. We then develop a finite-sample testing algorithm for assessing the validity of candidate IV sets from data.

4.1 CAT Condition with Covariates

In practice, observed covariates such as age, gender, and other background variables may affect the treatment, outcome, and candidate IVs. It is therefore necessary to assess IV validity after adjusting for such covariates.

Under CAM-NCE with covariates, the data-generating process is given by

X\displaystyle X =gX(𝐙)+tX(𝐖)+φX(𝐔)+εX,\displaystyle=g_{X}(\mathbf{Z})+t_{X}(\mathbf{W})+{\varphi_{X}(\mathbf{U})+\varepsilon_{X}}, (4)
Y\displaystyle Y =f(X)+tY(𝐖)+gY(𝐙~)+φY(𝐔)+εY,\displaystyle=f(X)+t_{Y}(\mathbf{W})+g_{Y}(\widetilde{\mathbf{Z}})+{\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}},

where 𝐖\mathbf{W} denotes the covariates. Below, we show that the additive structure in Equation (4) allows the CAT condition in Definition 4 to be extended to covariate settings through regression adjustment.

Definition 5 (CAT Condition with Covariates).

Let XX, YY, 𝐖\mathbf{W}, and 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} denote the treatment, outcome, covariates, and candidate IV set, respectively. Furthermore, let 𝒳\mathcal{X}, 𝒴\mathcal{Y}, and 𝓩\bm{\mathcal{Z}}^{\prime} denote the residuals obtained by regressing XX, YY, and 𝐙\mathbf{Z}^{\prime} on 𝐖\mathbf{W}, respectively (e.g., for each IV, 𝒵iZi𝔼[Zi|𝐖]\mathcal{Z}_{i}\coloneqq Z_{i}-\mathbb{E}[Z_{i}|\mathbf{W}]). We say that {X,Y||(𝐙,𝐖)}\{X,Y||(\mathbf{Z}^{\prime},\mathbf{W})\} satisfies the CAT condition if and only if for every pair of distinct IVs {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime}, the following cross-independence relationships hold:

𝒱𝒳𝒴||𝒵i𝒵j, and 𝒱𝒳𝒴||𝒵j𝒵i.\displaystyle\mathcal{V}_{\mathcal{X}\to\mathcal{Y}||\mathcal{Z}_{i}}\mathrel{\perp\mspace{-10mu}\perp}\mathcal{Z}_{j},\text{ and }\mathcal{V}_{\mathcal{X}\to\mathcal{Y}||\mathcal{Z}_{j}}\mathrel{\perp\mspace{-10mu}\perp}\mathcal{Z}_{i}. (5)

Based on Definition 5 and Theorem 1, we derive the following necessary condition for IV set validity in the presence of covariates 𝐖\mathbf{W}.

Corollary 1 (Necessary Condition for IV Set Validity with Covariates).

Let XX, YY, 𝐖\mathbf{W}, and 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} be the treatment, outcome, covariates, and candidate IV set in CAM-NCE, respectively. Suppose that XX, YY, 𝐖\mathbf{W}, and 𝐙\mathbf{Z}^{\prime} are statistically dependent, and that Assumption 1 holds for the residualized variables. If the candidate IV set 𝐙\mathbf{Z}^{\prime} is a valid IV set w.r.t.XYX\to Y given 𝐖\mathbf{W}, then {X,Y||(𝐙,𝐖)}\{X,Y||(\mathbf{Z}^{\prime},\mathbf{W})\} satisfies the CAT condition.

Proof.

See Appendix B.5 for its proof. ∎

Corollary 1 states that if {X,Y||(𝐙,𝐖)}\{X,Y||(\mathbf{Z}^{\prime},\mathbf{W})\} violates the CAT condition, then 𝐙\mathbf{Z}^{\prime} is an invalid IV set. Furthermore, to obtain a necessary and sufficient characterization in the presence of covariates, we impose Assumption 2 on the covariate-adjusted residual variables.

Corollary 2 (Necessary and Sufficient Condition for IV Set Validity with Covariates).

Let XX, YY, 𝐖\mathbf{W}, and 𝐙𝐙\mathbf{Z}^{\prime}\subseteq\mathbf{Z} be the treatment, outcome, covariates, and candidate IV set in a CAM-NCE, respectively. Suppose that XX, YY, 𝐖\mathbf{W}, and 𝐙\mathbf{Z}^{\prime} are statistically dependent. Further suppose that Assumptions 1 and 2 hold for the covariate-adjusted residual variables. The candidate IV set 𝐙\mathbf{Z}^{\prime} is a valid IV set relative to XYX\to Y given 𝐖\mathbf{W} if and only if {X,Y||(𝐙,𝐖)}\{X,Y||(\mathbf{Z}^{\prime},\mathbf{W})\} satisfies the CAT condition.

Proof.

See Appendix B.6 for its proof. ∎

4.2 CAT Algorithm with Finite Samples

In this subsection, we develop a practical CAT algorithm for finite-sample data. For generality, we present the algorithm in the setting with baseline covariates 𝐖\mathbf{W}. Let 𝒳^\widehat{\mathcal{X}}, 𝒴^\widehat{\mathcal{Y}}, and 𝓩^={𝒵^1,,𝒵^m}\widehat{\bm{\mathcal{Z}}}=\{\widehat{\mathcal{Z}}_{1},\ldots,\widehat{\mathcal{Z}}_{m}\} denote the residuals obtained by regressing XX, YY, and each candidate IV Zi𝐙={Z1,,Zm}Z_{i}\in\mathbf{Z}=\{Z_{1},\ldots,Z_{m}\} on 𝐖\mathbf{W}, respectively. When no covariates are available, this residualization step is omitted and the original variables are used directly.

Since the CAT results are given in terms of population-level (Theorems 1\sim2 and Corollaries 1\sim2), implementing it with observational samples requires addressing three practical questions:

  • 𝒬1\mathcal{Q}1.

    How can one efficiently search for a valid IV set from a collection of candidate IVs?

  • 𝒬2\mathcal{Q}2.

    How can one estimate the auxiliary variable 𝒱𝒳𝒴|𝒵i\mathcal{V}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}} for each candidate IV ZiZ_{i}?

  • 𝒬3\mathcal{Q}3.

    How can the CAT condition be implemented with finite samples?

We next address these questions in turn.

𝒬1\mathcal{Q}1: Searching for candidate IV sets. Since the CAT condition is defined through pairwise cross-independence relations, we focus on candidate subsets with at least two IVs. Searching over all such subsets of 𝓩^\widehat{\bm{\mathcal{Z}}} requires examining 2mm12^{m}-m-1 possibilities, which can be computationally expensive when mm is large. To make the search tractable, we introduce a user-specified parameter K2K\geq 2, denoting the target size of the IV set to be selected. Given KK, we restrict attention to candidate subsets of size KK, resulting in (mK)\binom{m}{K} subsets. Specifically, we consider all subsets 𝒮c𝓩^\mathcal{S}_{c}\subseteq\widehat{\bm{\mathcal{Z}}} with |𝒮c|=K|\mathcal{S}_{c}|=K. When prior knowledge about KK is unavailable, we suggest evaluating KK over a range of plausible values, starting from small values, and using the CAT-based score below to select among the resulting candidate subsets.

𝒬2\mathcal{Q}2: Estimating auxiliary variables. To construct the auxiliary variable for each candidate IV, the key step is to estimate the function hi()h_{i}(\cdot) satisfying the conditional moment restriction 𝔼[𝒴hi(𝒳)𝒵i]=0\mathbb{E}\!\left[\mathcal{Y}-h_{i}(\mathcal{X})\mid\mathcal{Z}_{i}\right]=0. Under the additive model, if 𝒵i\mathcal{Z}_{i} is a valid IV, then hi()h_{i}(\cdot) is identifiable under the completeness condition (Assumption 1) and coincides with the structural response function of the treatment on the outcome (53). In this case, hi()h_{i}(\cdot) can be consistently estimated using suitable IV estimators. If 𝒵i\mathcal{Z}_{i} is invalid, the resulting estimator may be biased, and this bias will be reflected in the corresponding auxiliary variable. Various estimators have been developed for additive nonparametric IV models, including sieve-based estimators (54), kernel-based estimators (62; 51), deep IV estimators (30; 44), and control-function estimators (52; 28). In our implementation, we use the control-function IV estimator for non-constant effects (28). When prior knowledge suggests a constant-effect model, we instead use the two-stage least squares estimator (71). Given the estimated function h^i()\widehat{h}_{i}(\cdot), the empirical auxiliary variable is constructed as 𝒱^𝒳𝒴|𝒵i=𝒴^h^i(𝒳^)\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}}=\widehat{\mathcal{Y}}-\widehat{h}_{i}(\widehat{\mathcal{X}}).

𝒬3\mathcal{Q}3: Implementing the CAT condition with finite samples. For each candidate subset 𝒮c\mathcal{S}_{c}, the CAT condition requires every pair of distinct candidate IVs to satisfy two directed cross-independence relations. Specifically, for each pair {𝒵^i,𝒵^j}𝒮c\{\widehat{\mathcal{Z}}_{i},\widehat{\mathcal{Z}}_{j}\}\subseteq\mathcal{S}_{c}, we need to assess

𝒱^𝒳𝒴|𝒵i𝒵^j,𝒱^𝒳𝒴|𝒵j𝒵^i.\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}}\mathrel{\perp\mspace{-10mu}\perp}\widehat{\mathcal{Z}}_{j},\quad\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{j}}\mathrel{\perp\mspace{-10mu}\perp}\widehat{\mathcal{Z}}_{i}.

One could use formal independence tests for continuous variables and determine whether the CAT condition holds based on the corresponding pp-values. However, since our goal is to compare many candidate subsets in finite samples, we use distance correlation as a numerical dependence score. According to  65; 66, distance correlation is zero if and only if the two random vectors are independent. Hence, whenever the population-level CAT condition holds, the corresponding population distance correlation is zero. In finite samples, smaller sample distance correlations therefore provide stronger empirical support for the CAT condition. Let dCor^(,)\widehat{\operatorname{dCor}}(\cdot,\cdot) denote the sample distance correlation. For each candidate subset 𝒮c\mathcal{S}_{c}, we define its CAT score as the sum of directed pairwise distance correlations:

T^𝒮c={𝒵^i,𝒵^j}𝒮ci<j[\displaystyle\widehat{T}_{\mathcal{S}_{c}}=\sum_{\begin{subarray}{c}\{\widehat{\mathcal{Z}}_{i},\widehat{\mathcal{Z}}_{j}\}\subseteq\mathcal{S}_{c}\\ i<j\end{subarray}}\Bigg[ dCor^(𝒱^𝒳𝒴|𝒵i,𝒵^j)+dCor^(𝒱^𝒳𝒴|𝒵j,𝒵^i)].\displaystyle\widehat{\operatorname{dCor}}\!\left(\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}},\widehat{\mathcal{Z}}_{j}\right)+\widehat{\operatorname{dCor}}\!\left(\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{j}},\widehat{\mathcal{Z}}_{i}\right)\Bigg].

For a subset of size KK, this score aggregates K(K1)/2K(K-1)/2 unordered IV pairs and K(K1)K(K-1) directed distance-correlation terms. Finally, we select the candidate subset with the smallest empirical CAT score:

𝒮^=argmin𝒮c𝓩^,|𝒮c|=KT^𝒮c.\widehat{\mathcal{S}}=\arg\min_{\mathcal{S}_{c}\subseteq\widehat{\bm{\mathcal{Z}}},\ |\mathcal{S}_{c}|=K}\widehat{T}_{\mathcal{S}_{c}}.

The selected subset 𝒮^\widehat{\mathcal{S}} is therefore the candidate IV set that is most consistent with the CAT condition in finite samples.

Algorithm 1 CAT
0:  Observed dataset 𝒟={X,Y,𝐖,𝐙}\mathcal{D}=\{X,Y,\mathbf{W},\mathbf{Z}\}; KK, the number of valid IVs to select.
0:  Selected IV set 𝒮^\widehat{\mathcal{S}} and its corresponding distance-correlation score T𝒮T_{\mathcal{S}}.
1:Initialize: the selected IV set 𝒮^\widehat{\mathcal{S}}\leftarrow\emptyset, the distance correlation matrix 𝐌𝟎m×m\mathbf{M}\leftarrow\mathbf{0}_{m\times m};
2:if 𝐖\mathbf{W}\neq\emptyset then
3:   𝒳^,𝒴^,𝓩^\widehat{\mathcal{X}},\widehat{\mathcal{Y}},\widehat{\bm{\mathcal{Z}}}\leftarrow residuals from the regressions of X,Y,𝐙X,Y,\mathbf{Z} on 𝐖\mathbf{W}, respectively;
4:else
5:   𝒳^,𝒴^,𝓩^X,Y,𝐙\widehat{\mathcal{X}},\widehat{\mathcal{Y}},\widehat{\bm{\mathcal{Z}}}\leftarrow X,Y,\mathbf{Z};
6:end if
7:for i=1,,mi=1,\ldots,m do
8:   if the non-constant causal effect of XYX\to Y is adopted then
9:    h^i(𝒳^)Control-Function IV Estimator(𝒳^,𝒴^,𝒵^i)\widehat{h}_{i}(\widehat{\mathcal{X}})\leftarrow\text{Control-Function IV Estimator}(\widehat{\mathcal{X}},\widehat{\mathcal{Y}},\widehat{\mathcal{Z}}_{i});
10:   else
11:    h^i(𝒳^)β^i𝒳^\widehat{h}_{i}(\widehat{\mathcal{X}})\leftarrow\widehat{\beta}_{i}\widehat{\mathcal{X}}, where β^iCov(𝒴^,𝒵^i)Cov(𝒳^,𝒵^i)\widehat{\beta}_{i}\leftarrow\frac{\operatorname{Cov}(\widehat{\mathcal{Y}},\widehat{\mathcal{Z}}_{i})}{\operatorname{Cov}(\widehat{\mathcal{X}},\widehat{\mathcal{Z}}_{i})};
12:   end if
13:   𝒱^i𝒴^h^i(𝒳^)\widehat{\mathcal{V}}_{i}\leftarrow\widehat{\mathcal{Y}}-\widehat{h}_{i}(\widehat{\mathcal{X}});
14:end for
15:for each unordered pair {𝒵i,𝒵j}\{\mathcal{Z}_{i},\mathcal{Z}_{j}\} with 1i<jm1\leq i<j\leq m do
16:   𝐌^ijdCor^(𝒱^i,𝒵^j)+dCor^(𝒱^j,𝒵^i)\widehat{\mathbf{M}}_{ij}\leftarrow\widehat{\operatorname{dCor}}(\widehat{\mathcal{V}}_{i},\widehat{\mathcal{Z}}_{j})+\widehat{\operatorname{dCor}}(\widehat{\mathcal{V}}_{j},\widehat{\mathcal{Z}}_{i});
17:   𝐌^ji𝐌^ij\widehat{\mathbf{M}}_{ji}\leftarrow\widehat{\mathbf{M}}_{ij};
18:end for
19:  Let 𝒫K={𝒮c𝓩^:|𝒮c|=K}\mathcal{P}_{K}=\{\mathcal{S}_{c}\subseteq\widehat{\bm{\mathcal{Z}}}:|\mathcal{S}_{c}|=K\};
20:for each candidate subset 𝒮c𝒫K\mathcal{S}_{c}\in\mathcal{P}_{K} do
21:   T^𝒮c0\widehat{T}_{\mathcal{S}_{c}}\leftarrow 0;
22:   for each unordered pair {𝒵^i,𝒵^j}𝒮c\{\widehat{\mathcal{Z}}_{i},\widehat{\mathcal{Z}}_{j}\}\subseteq\mathcal{S}_{c} with i<ji<j do
23:    T^𝒮cT^𝒮c+𝐌^ij\widehat{T}_{\mathcal{S}_{c}}\leftarrow\widehat{T}_{\mathcal{S}_{c}}+\widehat{\mathbf{M}}_{ij};
24:   end for
25:end for
26:𝒮^argmin𝒮c𝒫KT^𝒮c\widehat{\mathcal{S}}\leftarrow\arg\min_{\mathcal{S}_{c}\in\mathcal{P}_{K}}\widehat{T}_{\mathcal{S}_{c}};
27:return 𝒮^\widehat{\mathcal{S}} and T^𝒮\widehat{T}_{\mathcal{S}}.

Based on the above three components, the entire procedure is summarized in Algorithm 1. Overall, the algorithm proceeds in three main stages. First, it adjusts for covariates by residualizing the treatment, outcome, and candidate IVs with respect to 𝐖\mathbf{W} (Lines 2–6). Second, it estimates the auxiliary variables for each candidate IV by estimating the corresponding function hi()h_{i}(\cdot), as described in 𝒬2\mathcal{Q}2 (Lines 7–14). Third, it searches over candidate IV subsets of size KK and evaluates each subset using the CAT-based score defined in 𝒬3\mathcal{Q}3 (Lines 15–26). The subset with the smallest empirical CAT score is returned as the estimated IV set that is most consistent with the CAT condition.

Below, we establish the correctness of the CAT algorithm in selecting a valid IV set as the sample size tends to infinity.

Theorem 3 (Correctness of the CAT Algorithm).

Assume that the observed data {X,Y,𝐖,𝐙}\{X,Y,\mathbf{W},\mathbf{Z}\} are generated from CAM-NCE, and that the candidate IV set 𝐙\mathbf{Z} contains at least KK valid IVs. Suppose that Assumptions 12 hold, and that the estimators used in Algorithm 1 for covariate adjustment, auxiliary-variable construction, and distance correlation are consistent. Then, as the sample size tends to infinity, Algorithm 1 outputs a valid IV subset 𝒮^\widehat{\mathcal{S}} with |𝒮^|=K|\widehat{\mathcal{S}}|=K. In particular, if 𝐙\mathbf{Z} contains exactly KK valid IVs, then Algorithm 1 outputs the full valid IV set.

Proof.

See Appendix B.7 for its proof. ∎

Computational complexity. We finally analyze the computational complexity of the CAT algorithm. Let nn denote the sample size, m=|𝐙|m=|\mathbf{Z}| denote the number of candidate IVs, and p=|𝐖|p=|\mathbf{W}| denote the number of covariates. The time complexity consists of four main components:

  • 1.

    Covariate residualization. Regressing XX, YY, and the mm candidate IVs on 𝐖\mathbf{W} to obtain residualized variables costs 𝒪(nmp2)\mathcal{O}(nmp^{2}).

  • 2.

    Auxiliary-variable estimation: For each candidate IV, we estimate hi()h_{i}(\cdot) using the estimator adopted in this work, i.e., the semiparametric control-function estimator for non-constant effects and two-stage least squares for constant effects. This step costs 𝒪(mn)\mathcal{O}(mn) in total.

  • 3.

    Pairwise distance-correlation calculation. Computing the directed pairwise distance correlations for all candidate IV pairs costs 𝒪(m2n2)\mathcal{O}(m^{2}n^{2}) using the standard distance-correlation estimator.

  • 4.

    Candidate subset selection. After the pairwise distance-correlation matrix is computed, evaluating all candidate subsets of size KK costs 𝒪((mK)K2)\mathcal{O}\!\left(\binom{m}{K}K^{2}\right).

Hence, the overall computational complexity is

𝒪(nmp2+mn+m2n2+(mK)K2).\mathcal{O}\!\left(nmp^{2}+mn+m^{2}n^{2}+\binom{m}{K}K^{2}\right).

5 Experiments

In this section, we evaluate the proposed CAT method on synthetic datasets generated under both CAM-CE and CAM-NCE, corresponding to the constant-effect and non-constant-effect settings, respectively. Our goal is to examine whether CAT can correctly distinguish valid IV sets from invalid ones under different types of IV assumption violations. We first consider the constant-effect setting, where representative existing IV methods are applicable and thus serve as baselines, and then proceed to the more general non-constant-effect setting, which is the primary focus of this paper. It is noteworthy that the simulated data are generated solely according to the causal additive model in Equation (1); no additional constraints are imposed to ensure that the distributional condition in Assumption 2 holds. Our source code is available at https://github.com/guoxichen0/CAT.

Across both settings, we consider three representative invalid IV scenarios involving both valid and invalid IVs. In Case 1, the invalid IVs violate the exclusion restriction condition (𝒞2\mathcal{C}2); in Case 2, the invalid IVs violate the exogeneity condition (𝒞3\mathcal{C}3); and in Case 3, the invalid IVs violate both the exclusion restriction (𝒞2\mathcal{C}2) and exogeneity (𝒞3\mathcal{C}3) conditions. Each experiment is repeated 100 times with independently generated data. All noise terms are independently drawn from U(1,1)U(-1,1). For each case, we vary the sample size over n{1000,3000,5000}n\in\{1000,3000,5000\}.

5.1 Synthetic Data under Constant Effects

We first evaluate CAT under the CAM-CE framework, where the structural response function takes the linear form f(X)=βXf(X)=\beta X, yielding a constant causal effect. Following 27 and related works, the true constant causal effect is fixed at β=1\beta=1. Other constant coefficients in the structural equations are independently sampled from [1.5,0.5][0.5,1.5][-1.5,-0.5]\cup[0.5,1.5]. We report comparison results under nonlinear structural components with constant treatment effects, settings with covariates, and the ALICE model.

Since existing IV selection methods are primarily developed for constant-effect or linear models, the CAM-CE setting enables direct comparison with the following representative baselines:

  1. 1.

    NAIVE: the least-squares regression coefficient of YY on XX;

  2. 2.

    MR-Egger (7): implemented using the code available at https://academic.oup.com/ije/article/44/2/512/754653/;

  3. 3.

    TSHT (27): implemented using the code available at https://cran.r-project.org/web/packages/RobustIV/;

  4. 4.

    CIIV (70): implemented using the code available at https://github.com/xlbristol/CIIV/;

  5. 5.

    sisVIVE (36): implemented using the code available at https://cran.r-project.org/web/packages/sisVIVE/;

  6. 6.

    IV-tetrad (61): implemented using the code available at https://www.homepages.ucl.ac.uk/~ucgtrbd/code/iv_discovery/.

Metrics. We evaluate the methods through their resulting causal effect estimates, summarized using boxplots, where estimates closer to the true effect with smaller variability and dispersion indicate better performance. For methods that first select IVs, including TSHT, CIIV, sisVIVE, IV-tetrad, and CAT, we apply the same IV estimator as in 27 after IV selection to ensure a fair comparison. NAIVE and MR-Egger directly produce causal effect estimates and are therefore evaluated using their original outputs.

Refer to caption
(a)
Refer to caption
(b)
Refer to caption
(c)
Figure 6: Performance of NAIVE, MR-Egger, TSHT, CIIV, sisVIVE, IV-tetrad, and CAT across three different cases in the CAM-CE framework.
Refer to caption
(a)
Refer to caption
(b)
Refer to caption
(c)
Figure 7: Performance of NAIVE, MR-Egger, TSHT, CIIV, sisVIVE, IV-tetrad, and CAT across three different cases with covariates 𝐖\mathbf{W} in the CAM-CE framework.
Refer to caption
(a)
Refer to caption
(b)
Refer to caption
(c)
Figure 8: Performance of NAIVE, MR-Egger, TSHT, CIIV, sisVIVE, IV-tetrad, and CAT across three different cases in the ALICE model.

Results. The results are presented in Figures  68. Figures 6 and 7 report the results under nonlinear structural components with constant treatment effects, without and with covariates, respectively. As expected, the proposed CAT algorithm consistently outperforms the other methods across all three cases and sample sizes, exhibiting relatively small variance and producing estimates closest to the true causal effect. In contrast, the NAIVE method performs poorly in all cases due to unmeasured confounders 𝐔\mathbf{U}. We observed that the other comparison methods perform poorly across all cases because they rely on the assumption of linearity, whereas the data generation process is partially nonlinear. Additionally, we found that the MR-Egger algorithm yields inaccurate results. A possible reason for this is that, in addition to the linearity assumption, this method requires the InSIDE assumption, which states that the instruments’ pleiotropic effects on the outcome YY are uncorrelated with their effects on the treatment XX. Furthermore, Figure 8 reports the results under the ALICE model. In this linear constant-effect setting, CAT performs well across all three cases and yields estimates close to the true causal effect. Its performance is comparable to TSHT, CIIV, sisVIVE, and IV-tetrad, which is expected because these methods are designed for linear constant-effect settings. In contrast, NAIVE and MR-Egger exhibit larger variability or bias across all cases.

5.2 Synthetic Data under Non-Constant Effects

We next evaluate CAT under the CAM-NCE framework, where f()f(\cdot) is nonlinear, and the causal effect varies with the treatment level. Existing IV-selection methods considered above rely on linear or constant-effect specifications and are not designed for the CAM-NCE setting. Applying them here would therefore evaluate the methods under model misspecification rather than provide a meaningful comparison of IV-set identification. To the best of our knowledge, no existing method provides a directly comparable IV-set identification procedure under the CAM-NCE framework considered here. We thus focus on systematically evaluating CAT across a range of non-constant-effect settings. Specifically, we assess the performance of CAT from three perspectives: (i) varying the causal effect function from XX to YY, including log\log, sin\sin, cos\cos, quadraticpolynomial\mathrm{quadratic\ polynomial}, cubicpolynomial\mathrm{cubic\ polynomial}, log(quad)\log(\mathrm{quad}), and exp(quad)\exp(\mathrm{quad}), with K=2K=2 valid IVs among five candidate IVs; (ii) varying the number of valid IVs, K{2,3,5}K\in\{2,3,5\}, among ten candidate IVs; and (iii) varying the number of covariates, |𝐖|{2,3,5}|\mathbf{W}|\in\{2,3,5\}, with K=2K=2 valid IVs among five candidate IVs.

Metrics. We evaluate IV-set identification performance using the Error Rate (ER). A trial is regarded as successful only when the IV set identified by the method exactly matches the true valid IV set. Accordingly, ER is defined as the proportion of trials for which the identified IV set differs from the true valid IV set, with a lower ER indicating better identification performance.

Table III: Performance of IV-validity Testing under Different Causal Effect Functions in the CAM-NCE Framework
Sample sizes Size=1000 Size=3000 Size=5000
Cases Functions f(X)f(X) ER\bm{\downarrow} ER\bm{\downarrow} ER\bm{\downarrow}
Case 1 Log 0.00 0.00 0.00
Sin 0.00 0.00 0.00
Cos 0.01 0.00 0.00
Quadratic Poly 0.00 0.00 0.00
Cubic Poly 0.00 0.00 0.00
Log(quad) 0.00 0.00 0.00
Exp(quad) 0.00 0.00 0.00
Case 2 Log 0.00 0.00 0.00
Sin 0.00 0.00 0.00
Cos 0.00 0.00 0.00
Quadratic Poly 0.00 0.00 0.00
Cubic Poly 0.00 0.00 0.00
Log(quad) 0.00 0.00 0.00
Exp(quad) 0.00 0.00 0.00
Case 3 Log 0.00 0.00 0.00
Sin 0.00 0.00 0.00
Cos 0.00 0.00 0.00
Quadratic Poly 0.00 0.00 0.00
Cubic Poly 0.00 0.00 0.00
Log(quad) 0.04 0.00 0.00
Exp(quad) 0.00 0.00 0.00
  • Note: “Quadratic Poly/quad” and “Cubic Poly” denote quadratic and cubic polynomial functions, respectively. \downarrow indicates that lower values are better.

Table IV: Performance of IV-validity Testing for Different Numbers KK of Valid IVs in the CAM-NCE Framework
Sample sizes Size=1000 Size=3000 Size=5000
KK Cases ER\bm{\downarrow} ER\bm{\downarrow} ER\bm{\downarrow}
2 Case 1 0.04 0.00 0.00
Case 2 0.00 0.00 0.00
Case 3 0.01 0.01 0.00
3 Case 1 0.01 0.00 0.00
Case 2 0.00 0.00 0.00
Case 3 0.00 0.00 0.00
5 Case 1 0.00 0.00 0.00
Case 2 0.00 0.00 0.00
Case 3 0.00 0.00 0.00
  • Note: KK denotes the number of valid IVs to consider. \downarrow indicates that lower values are better.

Table V: Performance of IV-validity Testing with Covariates in the CAM-NCE Framework
Sample sizes Size=1000 Size=3000 Size=5000
|𝐖||\mathbf{W}| Cases ER\bm{\downarrow} ER\bm{\downarrow} ER\bm{\downarrow}
2 Case 1 0.00 0.00 0.00
Case 2 0.08 0.07 0.05
Case 3 0.04 0.04 0.01
3 Case 1 0.02 0.00 0.00
Case 2 0.10 0.09 0.08
Case 3 0.14 0.10 0.04
5 Case 1 0.00 0.00 0.00
Case 2 0.14 0.10 0.08
Case 3 0.15 0.12 0.10
  • Note: |𝐖||\mathbf{W}| denotes the number of covariates. \downarrow indicates that lower values are better.

Results. The results are summarized in Tables IIIV. Overall, the proposed method achieves consistently low error rates in all considered settings, suggesting that it can reliably identify the valid IV set under non-constant causal effects. Table III reports the results for different forms of the causal effect function from XX to YY. Across logarithmic, trigonometric, polynomial, and exponential functions, the ER is close to zero for all sample sizes, indicating that the proposed method is insensitive to the specific functional form of the non-constant causal effect. Table IV further examines the influence of the number of valid IVs in the candidate IV set. The ER remains low for different choices of KK and generally decreases as the sample size increases, showing that the method is stable across different proportions of valid IVs. This is in contrast to methods relying on the Majority Rule or Plurality Rule assumptions (see Section 1), which require restrictions on the proportion of valid IVs. Finally, Table V presents the results when covariates 𝐖\mathbf{W} are included. Although the ER slightly increases as the number of covariates grows, it decreases with larger sample sizes. This suggests that the proposed method remains effective in the presence of covariates, while benefiting from increased sample size.

6 Applications to Real-world Data

In this section, we apply CAT to three real-world datasets spanning sociology, economics, and behavioral science to examine its practical utility. Since the ground-truth validity of candidate IVs is not directly observable in real-world applications, we use the IV specifications proposed in prior studies as reference benchmarks and compare the conclusions obtained by CAT with those reported in the literature. We report both the CAT-based distance-correlation scores and the corresponding p-values from distance-correlation independence tests. Here, the significance level is set to α=10/n\alpha=10/n, where nn denotes the sample size used in the test.

6.1 Colonial Origins Data (1)

This dataset examines the impact of social systems on economic development. After excluding observations with missing values, it contains five key variables across 63 countries: Mortality (Mor)(M_{or}), Euro1990 (Euro)(E_{uro}), Latitude (Lat)(L_{at}), Institutions (Ins)(I_{ns}), and Economic Development (Ed)(E_{d}). We take the IV specification in 1 as a reference benchmark and evaluate it under the constant-effect setting. The hypothesized model proposed by 1 is illustrated in Figure 9, and the hypothesized data generation mechanism is described as follows:

Ins\displaystyle I_{ns} =α0+α1Mor+α2Euro+α3Lat+δ,\displaystyle=\alpha_{0}+\alpha_{1}M_{or}+\alpha_{2}E_{uro}+\alpha_{3}L_{at}+\delta,
Ed\displaystyle E_{d} =β0+β1Ins+β2Lat+ϵ,\displaystyle=\beta_{0}+\beta_{1}I_{ns}+\beta_{2}L_{at}+\epsilon,

where δ\delta and ϵ\epsilon are dependent. Specifically, we assess the candidate IV set {Mor,Euro}\{M_{or},E_{uro}\} for the causal relation InsEdI_{ns}\to E_{d}, conditioning on the covariate LatL_{at}.

LatL_{at} EuroE_{uro} MorM_{or} InsI_{ns} EdE_{d} Cultural difference
Figure 9: Graphical illustration of an IV model for estimating the causal effect of institutions (InsI_{ns}) on economic development (EdE_{d}) (1).

Results. Since there are only two candidate IVs, we apply CAT directly to this pair and obtain a CAT-based distance-correlation score of 0.500.50. We then conduct distance-correlation independence tests for the two directed CAT relations, namely between 𝒱Mor\mathcal{V}_{M_{or}} and EuroE_{uro}, and between 𝒱Euro\mathcal{V}_{E_{uro}} and MorM_{or}, obtaining pp-values of 0.210.21 and 0.250.25, respectively. Therefore, we do not reject the corresponding CAT independence relations, providing no evidence against {Mor,Euro}\{M_{or},E_{uro}\} as a valid IV set for InsEdI_{ns}\to E_{d}. This result is consistent with the findings of 1.

6.2 Children and Mothers’ Labor Supply Data (5)

This dataset comes from an empirical study on the effect of childbearing on mothers’ labor supply. After applying the filtering criteria, it contains 254,652 observations. We use more than two children (morekids)(\textit{morekids}) as the treatment and weeks worked (weeksm1)(\textit{weeksm1}) as the outcome. The candidate IVs include two boys (boys2)(\textit{boys2}), two girls (girls2)(\textit{girls2}), AGEQK, AGEQ2ND, KIDCOUNT, YOBM, nonmomil, educm, hsormore, nonmomi, ageqm, and agefstd. The covariates 𝐖\mathbf{W} include mother’s age at first birth (agem1)(\textit{agem1}), father’s age at first birth (agefstm)(\textit{agefstm}), whether the first child is a boy (boy1st)(\textit{boy1st}), whether the second child is a boy (boy2nd)(\textit{boy2nd}), black mother indicator (blackm)(\textit{blackm}), Hispanic mother indicator (hispm)(\textit{hispm}), and other-race mother indicator (othracem)(\textit{othracem}). We take the IV specification in 5 as a reference benchmark and evaluate it under the constant-effect setting. Due to the large sample size and the quadratic computational cost of distance correlation, we randomly subsample 5%5\% of the data and average the results over 10 repeated tests. The valid IVs hypothesized model proposed by 5 is illustrated in Figure 10, and the valid IVs hypothesized data generation mechanism is described as follows:

morekids\displaystyle morekids =γ0boys2+γ1girls2+𝜸𝟐𝐖+δ,\displaystyle=\gamma_{0}boys2+\gamma_{1}girls2+\bm{\gamma_{2}}^{\top}\mathbf{W}+\delta,
weeksm1\displaystyle weeksm1 =βmorekids+𝜷𝟏(𝐖{boy2nd})+ϵ,\displaystyle=\beta\cdot morekids+\bm{\beta_{1}}^{\top}(\mathbf{W}\setminus\{boy2nd\})+\epsilon,

where δ\delta and ϵ\epsilon are dependent, 𝐖{boy2nd}\mathbf{W}\setminus\{boy2nd\} represents the set of all elements in covariates 𝐖\mathbf{W} after removing the variable boy2ndboy2nd.

𝐖{boy2nd}\mathbf{W}\setminus\{boy2nd\} boys2boys2 girls2girls2 morekidsmorekidsweeksm1weeksm1 Unobserved confounder
Figure 10: Graphical illustration of a valid IV model for estimating the causal effect of childbearing (morekidsmorekids) on mother’s labor supply (weeksm1weeksm1) (5).

Results. Using CAT with K=2K=2, we find that the candidate set {boys2,girls2}\{\textit{boys2},\textit{girls2}\} achieves the smallest CAT-based distance-correlation score, with dCor=0.022\operatorname{dCor}=0.022. We further conduct distance-correlation independence tests for the two directed CAT relations: between 𝒱boys2\mathcal{V}_{\textit{boys2}} and girls2, and between 𝒱girls2\mathcal{V}_{\textit{girls2}} and boys2, obtaining pp-values of 0.320.32 and 0.340.34, respectively. Thus, we do not reject the corresponding CAT independence relations, and CAT selects {boys2,girls2}\{\textit{boys2},\textit{girls2}\} as the estimated valid IV set for morekidsweeksm1\textit{morekids}\to\textit{weeksm1}. This result is consistent with the conclusion of 5.

6.3 Conflict and Time Preference Data (67)

This dataset comes from an empirical study on the effect of violent conflict on individual time preferences. In our analysis, we focus on the causal effect of Violence on Patience. After removing observations with missing values, the dataset contains 266 observations and 15 variables, including the treatment variable Violence (Vio)(V_{io}), the outcome variable Patience (Pat)(P_{at}), two candidate IVs, Distance (Dist)(D_{ist}) and Altitude (Alti)(A_{lti}), and 11 covariates 𝐖\mathbf{W}. The covariates include literate, age, sex, total land holding per capita, land Gini coefficient, distance to market, conflict over land, ethnic homogeneity, socioeconomic homogeneity, population density, and per capita total expenditure. We take the IV specification in 67 as a reference benchmark and evaluate it under the non-constant effect setting. The hypothesized model from 28 is illustrated in Figure 11, and the hypothesized generation mechanism is as follows:

Vio\displaystyle V_{io} =α0+α1Dist+α2Alti+α3Dist2+α4Alti2\displaystyle=\alpha_{0}+\alpha_{1}D_{ist}+\alpha_{2}A_{lti}+\alpha_{3}D_{ist}^{2}+\alpha_{4}A_{lti}^{2}
+α5DistAlti+𝜶6𝐖+δ,\displaystyle+\alpha_{5}D_{ist}\cdot A_{lti}+\bm{\alpha}_{6}^{\top}\mathbf{W}+\delta,
Pat\displaystyle P_{at} =β0+𝜷1𝐖+β2Vio+β3Vio2+ϵ,\displaystyle=\beta_{0}+\bm{\beta}_{1}^{\top}\mathbf{W}+\beta_{2}V_{io}+\beta_{3}V_{io}^{2}+\epsilon,

where δ\delta and ϵ\epsilon are dependent.

𝐖\mathbf{W} DistD_{ist} AltiA_{lti} VioV_{io} PatP_{at} Social and Political confounder
Figure 11: Graphical illustration of an IV model for estimating the causal effect of Violence on a person’s Patience (67).

Results. Since the dataset contains only two candidate IVs, we apply CAT directly to the pair {Dist,Alti}\{D_{ist},A_{lti}\} and obtain a CAT-based distance-correlation score of dCor=0.28\operatorname{dCor}=0.28. We then conduct distance-correlation independence tests for the two directed CAT relations, namely between 𝒱Dist\mathcal{V}_{D_{ist}} and AltiA_{lti}, and between 𝒱Alti\mathcal{V}_{A_{lti}} and DistD_{ist}, obtaining pp-values of 0.100.10 in both cases. Thus, we do not reject the corresponding CAT independence relations, providing no evidence against {Dist,Alti}\{D_{ist},A_{lti}\} as a valid IV set for VioPatV_{io}\to P_{at}. This result is consistent with the findings of 67.

7 Conclusion

In this paper, we studied the problem of testing the validity of IV sets from observational data under causal additive models with non-constant effects (CAM-NCE). Under the completeness condition (Assumption 1), we introduced a testable necessary condition, termed the Cross Auxiliary-based Independence Test (CAT) condition, for assessing IV set validity. Furthermore, under the cross distributional non-degeneracy condition (Assumption 2), we established a necessary and sufficient characterization of valid IV sets within the CAM-NCE framework. We also extended the CAT condition to settings with covariates and developed a practical finite-sample algorithm for selecting candidate IV sets that are most consistent with the CAT condition. Experimental results on both synthetic and real-world datasets demonstrate the effectiveness and practical utility of the proposed method. One promising direction for future work is to extend the proposed framework to more general causal models, such as models with multiple treatment variables.

References

  • Acemoglu et al. (2001) D. Acemoglu, S. Johnson, and J. A. Robinson The colonial origins of comparative development: an empirical investigation. American economic review 91 (5), pp. 1369–1401. Cited by: Figure 1, Figure 1, Figure 9, Figure 9, §6.1, §6.1, §6.1.
  • Ai and Chen (2003) C. Ai and X. Chen Efficient estimation of models with conditional moment restrictions containing unknown functions. Econometrica 71 (6), pp. 1795–1843. Cited by: §2.2.
  • Akbari et al. (2021) S. Akbari, E. Mokhtarian, A. Ghassami, and N. Kiyavash Recursive causal structure learning in the presence of latent variables and selection bias. Advances in Neural Information Processing Systems 34, pp. 10119–10130. Cited by: item 2.
  • Andrews (2017) D. W. Andrews Examples of l2-complete and boundedly-complete distributions. Journal of econometrics 199 (2), pp. 213–220. Cited by: §2.2.
  • Angrist and Evans (1996) J. Angrist and W. N. Evans Children and their parents’ labor supply: evidence from exogenous variation in family size. National bureau of economic research Cambridge, Mass., USA. Cited by: Figure 10, Figure 10, §6.2, §6.2, §6.2.
  • Bennett et al. (2019) A. Bennett, N. Kallus, and T. Schnabel Deep generalized method of moments for instrumental variable analysis. Advances in neural information processing systems 32. Cited by: §B.1, item 3, §2.2.
  • Bowden et al. (2015) J. Bowden, G. Davey Smith, and S. Burgess Mendelian randomization with invalid instruments: effect estimation and bias detection through egger regression. International journal of epidemiology 44 (2), pp. 512–525. Cited by: 1st item, Table I, item 1, item 2.
  • Bowden et al. (2016) J. Bowden, G. Davey Smith, P. C. Haycock, and S. Burgess Consistent estimation in mendelian randomization with some invalid instruments using a weighted median estimator. Genetic epidemiology 40 (4), pp. 304–314. Cited by: 1st item, Table I, item 1.
  • Burauel (2023) P. F. Burauel Evaluating instrument validity using the principle of independent mechanisms. Journal of Machine Learning Research 24 (176), pp. 1–56. Cited by: 2nd item, Table I.
  • Burgess et al. (2017) S. Burgess, D. S. Small, and S. G. Thompson A review of instrumental variable estimators for mendelian randomization. Statistical methods in medical research 26 (5), pp. 2333–2355. Cited by: §1.
  • Cai et al. (2019) R. Cai, F. Xie, C. Glymour, Z. Hao, and K. Zhang Triad constraints for learning causal structure of latent variables. Advances in neural information processing systems 32. Cited by: §3.1.
  • Cao et al. (2025) F. Cao, X. Jing, K. Yu, and J. Liang FWCEC: an enhanced feature weighting method via causal effect for clustering. IEEE Transactions on Knowledge and Data Engineering 37 (2), pp. 685–697. Cited by: §1.
  • Cao et al. (2023) T. Cao, Q. Xu, Z. Yang, and Q. Huang Mitigating confounding bias in practical recommender systems with partially inaccessible exposure status. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (2), pp. 957–974. Cited by: §1.
  • Chen et al. (2017) B. Chen, D. Kumor, and E. Bareinboim Identification and model testing in linear structural equation models using auxiliary variables. In International Conference on Machine Learning, pp. 757–766. Cited by: §3.1.
  • Cheng et al. (2024) D. Cheng, J. Li, L. Liu, J. Liu, and T. D. Le Data-driven causal effect estimation based on graphical causal modelling: a survey. ACM Computing Surveys 56 (5), pp. 1–37. Cited by: §1.
  • Cheng et al. (2023a) D. Cheng, J. Li, L. Liu, K. Yu, T. Duy Le, and J. Liu Toward unique and unbiased causal effect estimation from data with hidden variables. IEEE Transactions on Neural Networks and Learning Systems 34 (9), pp. 6108–6120. Cited by: §1.
  • Cheng et al. (2023b) D. Cheng, J. Li, L. Liu, K. Yu, T. D. Le, and J. Liu Discovering ancestral instrumental variables for causal inference from observational data. IEEE Transactions on Neural Networks and Learning Systems, pp. 1–11. Cited by: 1st item, Table I, item 1.
  • Chernozhukov et al. (2007) V. Chernozhukov, G. W. Imbens, and W. K. Newey Instrumental variable estimation of nonseparable models. Journal of Econometrics 139 (1), pp. 4–14. Cited by: §2.2.
  • Chu et al. (2001) T. Chu, R. Scheines, and P. Spirtes Semi-instrumental variables: a test for instrument admissibility. In Proceedings of the Seventeenth conference on Uncertainty in artificial intelligence, pp. 83–90. Cited by: Table I, §1, §3.1.
  • Colombo et al. (2012) D. Colombo, M. H. Maathuis, M. Kalisch, and T. S. Richardson Learning high-dimensional directed acyclic graphs with latent and selection variables. The Annals of Statistics, pp. 294–321. Cited by: item 2.
  • Darolles et al. (2011) S. Darolles, Y. Fan, J. Florens, and E. Renault Nonparametric instrumental regression. Econometrica 79 (5), pp. 1541–1565. Cited by: §2.2.
  • Drton and Richardson (2004) M. Drton and T. S. Richardson Iterative conditional fitting for gaussian ancestral graph models. In Proceedings of the 20th conference on Uncertainty in artificial intelligence, pp. 130–137. Cited by: §3.1.
  • D’Haultfoeuille (2011) X. D’Haultfoeuille On the completeness condition in nonparametric instrumental problems. Econometric Theory 27 (3), pp. 460–471. Cited by: §2.2, §3.1.
  • Farbmacher et al. (2022) H. Farbmacher, R. Guber, and S. Klaassen Instrument validity tests with causal forests. Journal of Business & Economic Statistics 40 (2), pp. 605–614. Cited by: §1.
  • Guo et al. (2026) X. Guo, Z. Li, B. Huang, Y. Zeng, Z. Geng, and F. Xie Testability of instrumental variables in additive nonlinear, non-constant effects models. Journal of Machine Learning Research. Note: To appear Cited by: Table I, §1, §3.1, §3.1.
  • Guo et al. (2025) X. Guo, F. Xie, Y. Zeng, H. Zhang, and Z. Geng Data-driven selection of instrumental variables for additive nonlinear, constant effects models. In Forty-second International Conference on Machine Learning, Cited by: 1st item, Table I, Remark 2.
  • Guo et al. (2018) Z. Guo, H. Kang, T. Tony Cai, and D. S. Small Confidence intervals for causal effects with invalid instruments by using two-stage hard thresholding with voting. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 80 (4), pp. 793–815. Cited by: 1st item, Table I, item 1, item 3, §5.1, §5.1.
  • Guo and Small (2016) Z. Guo and D. S. Small Control function instrumental variable estimation of nonlinear causal effect models. Journal of Machine Learning Research 17 (100), pp. 1–35. Cited by: item 3, §4.2, §6.3.
  • Han (2008) C. Han Detecting invalid instruments using l1-gmm. Economics Letters 101 (3), pp. 285–287. Cited by: 1st item, Table I, item 1.
  • Hartford et al. (2017) J. Hartford, G. Lewis, K. Leyton-Brown, and M. Taddy Deep iv: a flexible approach for counterfactual prediction. In International conference on machine learning, pp. 1414–1423. Cited by: §4.2.
  • Hartford et al. (2021) J. S. Hartford, V. Veitch, D. Sridhar, and K. Leyton-Brown Valid causal inference with (some) invalid instruments. In International Conference on Machine Learning, pp. 4096–4106. Cited by: 1st item, Table I, item 1.
  • Hartwig et al. (2017) F. P. Hartwig, G. Davey Smith, and J. Bowden Robust inference in summary data mendelian randomization via the zero modal pleiotropy assumption. International journal of epidemiology 46 (6), pp. 1985–1998. Cited by: 1st item, Table I, item 1.
  • Hu and Shiu (2018) Y. Hu and J. Shiu Nonparametric identification using instrumental variables: sufficient conditions for completeness. Econometric Theory 34 (3), pp. 659–693. Cited by: §2.2, §3.1.
  • Huber and Mellace (2015) M. Huber and G. Mellace Testing instrument validity for late identification based on inequality moment constraints. Review of Economics and Statistics 97 (2), pp. 398–411. Cited by: §1.
  • Janzing et al. (2012) D. Janzing, J. Mooij, K. Zhang, J. Lemeire, J. Zscheischler, P. Daniušis, B. Steudel, and B. Schölkopf Information-geometric approach to inferring causal directions. Artificial Intelligence 182, pp. 1–31. Cited by: 2nd item.
  • Kang et al. (2016) H. Kang, A. Zhang, T. T. Cai, and D. S. Small Instrumental variables estimation with some invalid instruments and its application to mendelian randomization. Journal of the American statistical Association 111 (513), pp. 132–144. Cited by: 1st item, Table I, item 1, item 5.
  • Kédagni and Mourifié (2020) D. Kédagni and I. Mourifié Generalized instrumental inequalities: testing the instrumental variable independence assumption. Biometrika 107 (3), pp. 661–675. Cited by: §1.
  • Kim et al. (2026) K. Kim, J. Kim, and E. H. Kennedy Causal k-means clustering. Journal of the Royal Statistical Society Series B: Statistical Methodology, pp. qkag068. Cited by: §1.
  • Kim et al. (2024) K. Kim, J. Kim, L. A. Wasserman, and E. H. Kennedy Hierarchical and density-based causal clustering. Advances in Neural Information Processing Systems 37, pp. 30363–30393. Cited by: §1.
  • Kitagawa (2015) T. Kitagawa A test for instrument validity. Econometrica 83 (5), pp. 2043–2063. Cited by: §1.
  • Kolesár et al. (2015) M. Kolesár, R. Chetty, J. Friedman, E. Glaeser, and G. W. Imbens Identification and inference with many invalid instruments. Journal of Business & Economic Statistics 33 (4), pp. 474–484. Cited by: 1st item, Table I, item 1.
  • Li et al. (2026) W. Li, Q. Zhang, and X. Wang Mechanisms under shifts: interpretable clustering with self-improving heterogeneous causal graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: §1.
  • Liao et al. (2024) L. Liao, Z. Fu, Z. Yang, Y. Wang, D. Ma, M. Kolar, and Z. Wang Instrumental variable value iteration for causal offline reinforcement learning. Journal of Machine Learning Research 25 (303), pp. 1–56. Cited by: §1.
  • Lin et al. (2019) A. Lin, J. Lu, J. Xuan, F. Zhu, and G. Zhang One-stage deep instrumental variable method for causal inference from observational data. In 2019 IEEE International Conference on Data Mining (ICDM), pp. 419–428. Cited by: §4.2.
  • Lin (1997) J. Lin Factorizing multivariate function classes. Advances in neural information processing systems 10. Cited by: §3.2, Theorem 5.
  • Lin et al. (2023) J. Lin, K. Wang, Z. Chen, X. Liang, and L. Lin Towards causality-aware inferring: a sequential discriminative approach for medical diagnosis. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (11), pp. 13363–13375. Cited by: §1.
  • Lin et al. (2024) Y. Lin, F. Windmeijer, X. Song, and Q. Fan On the instrumental variable estimation with many weak and invalid instruments. Journal of the Royal Statistical Society Series B: Statistical Methodology, pp. qkae025. Cited by: 1st item, Table I, item 1.
  • Manski (2003) C. F. Manski Partial identification of probability distributions. Springer Science & Business Media. Cited by: §1.
  • Meester et al. (2008) R. Meester et al. A natural introduction to probability theory. Cited by: Appendix A, Theorem 4.
  • Mourifié and Wan (2017) I. Mourifié and Y. Wan Testing local average treatment effect assumptions. Review of Economics and Statistics 99 (2), pp. 305–313. Cited by: §1.
  • Muandet et al. (2020) K. Muandet, A. Mehrjou, S. K. Lee, and A. Raj Dual instrumental variable regression. Advances in Neural Information Processing Systems 33, pp. 2710–2721. Cited by: §4.2.
  • Newey et al. (1999) W. K. Newey, J. L. Powell, and F. Vella Nonparametric estimation of triangular simultaneous equations models. Econometrica 67 (3), pp. 565–603. Cited by: §4.2.
  • Newey and Powell (2003) W. K. Newey and J. L. Powell Instrumental variable estimation of nonparametric models. Econometrica 71 (5), pp. 1565–1578. Cited by: §B.1, item 3, §2.2, §2.2, §3.1, §4.2.
  • Newey (2013) W. K. Newey Nonparametric instrumental variables estimation. American Economic Review 103 (3), pp. 550–556. Cited by: §2.2, §2.2, §3.1, §4.2.
  • Palmer et al. (2011) T. M. Palmer, R. R. Ramsahai, V. Didelez, and N. A. Sheehan Nonparametric bounds for the causal effect in a binary instrumental-variable model. The Stata Journal 11 (3), pp. 345–367. Cited by: §1.
  • Pearl (1995) J. Pearl On the testability of causal models with latent and instrumental variables. In Proceedings of the Eleventh conference on Uncertainty in artificial intelligence, pp. 435–443. Cited by: §1.
  • Pearl (2009) J. Pearl Causality: models, reasoning, and inference. 2nd edition, Cambridge University Press, New York. Cited by: Definition 1.
  • Peters et al. (2011) J. Peters, D. Janzing, and B. Scholkopf Causal inference on discrete data using additive noise models. IEEE Transactions on Pattern Analysis and Machine Intelligence 33 (12), pp. 2436–2450. Cited by: §1.
  • Sanderson et al. (2022) E. Sanderson, M. M. Glymour, M. V. Holmes, H. Kang, J. Morrison, M. R. Munafò, T. Palmer, C. M. Schooling, C. Wallace, Q. Zhao, et al. Mendelian randomization. Nature Reviews Methods Primers 2 (1), pp. 6. Cited by: 1st item, Table I, item 1.
  • Shi et al. (2021) W. Shi, G. Huang, S. Song, and C. Wu Temporal-spatial causal interpretations for vision-based reinforcement learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (12), pp. 10222–10235. Cited by: §1.
  • Silva and Shimizu (2017) R. Silva and S. Shimizu Learning instrumental variables with structural and non-gaussianity assumptions. Journal of Machine Learning Research 18 (120), pp. 1–49. Cited by: 1st item, Table I, item 1, item 6.
  • Singh et al. (2019) R. Singh, M. Sahani, and A. Gretton Kernel instrumental variable regression. Advances in Neural Information Processing Systems 32. Cited by: §B.1, item 3, §2.2, §4.2.
  • Skaaby et al. (2013) T. Skaaby, L. L. N. Husemoen, T. Martinussen, J. P. Thyssen, M. Melgaard, B. H. Thuesen, C. Pisinger, T. Jørgensen, J. D. Johansen, T. Menné, et al. Vitamin d status, filaggrin genotype, and cardiovascular risk factors: a mendelian randomization approach. PloS one 8 (2), pp. e57647. Cited by: §1.
  • Spirtes et al. (1995) P. Spirtes, C. Meek, and T. Richardson Causal inference in the presence of latent variables and selection bias. In Proceedings of the Eleventh conference on Uncertainty in artificial intelligence, pp. 499–506. Cited by: item 2.
  • Székely et al. (2007) G. J. Székely, M. L. Rizzo, and N. K. Bakirov Measuring and testing dependence by correlation of distances. The Annals of Statistics, pp. 2769–2794. Cited by: §4.2.
  • Székely and Rizzo (2009) G. J. Székely and M. L. Rizzo Brownian distance covariance. Cited by: §4.2.
  • Voors et al. (2012) M. J. Voors, E. E. M. Nillesen, P. Verwimp, E. H. Bulte, R. Lensink, and D. P. V. Soest Violent conflict and behavior: a field experiment in burundi. American Economic Review 102 (2), pp. 941–964. Cited by: Figure 11, Figure 11, §6.3, §6.3, §6.3.
  • Wang et al. (2017) L. Wang, J. M. Robins, and T. S. Richardson On falsification of the binary instrumental variable model. Biometrika 104 (1), pp. 229–236. Cited by: §1.
  • Windmeijer et al. (2019) F. Windmeijer, H. Farbmacher, N. Davies, and G. Davey Smith On the use of the lasso for instrumental variables estimation with some invalid instruments. Journal of the American Statistical Association 114 (527), pp. 1339–1350. Cited by: 1st item, Table I, item 1.
  • Windmeijer et al. (2021) F. Windmeijer, X. Liang, F. P. Hartwig, and J. Bowden The confidence interval method for selecting valid instrumental variables. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 83 (4), pp. 752–776. Cited by: 1st item, Table I, item 1, item 4.
  • Wooldridge (2010) J. M. Wooldridge Econometric analysis of cross section and panel data. MIT press. Cited by: §4.2.
  • Wu et al. (2025) A. Wu, K. Kuang, R. Xiong, and F. Wu Instrumental variables in causal inference and machine learning: a survey. ACM Computing Surveys 57 (11), pp. 1–36. Cited by: §1.
  • Xie et al. (2020) F. Xie, R. Cai, B. Huang, C. Glymour, Z. Hao, and K. Zhang Generalized independent noise condition for estimating latent variable causal graphs. Advances in neural information processing systems 33, pp. 14891–14902. Cited by: 2nd item.
  • Xie et al. (2022) F. Xie, Y. He, Z. Geng, Z. Chen, R. Hou, and K. Zhang Testability of instrumental variables in linear non-gaussian acyclic causal models. Entropy 24 (4), pp. 512. Cited by: 2nd item, Table I.
  • Xie et al. (2024) F. Xie, B. Huang, Z. Chen, R. Cai, C. Glymour, Z. Geng, and K. Zhang Generalized independent noise condition for estimating causal structure with latent variables. Journal of Machine Learning Research 25 (191), pp. 1–61. Cited by: §3.1.
  • Zhang et al. (2023) S. Zhang, F. Feng, K. Kuang, W. Zhang, Z. Zhao, H. Yang, T. Chua, and F. Wu Personalized latent structure learning for recommendation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (8), pp. 10285–10299. Cited by: §1.
  • Zheng et al. (2024) Y. Zheng, J. Qin, P. Wei, Z. Chen, and L. Lin CIPL: counterfactual interactive policy learning to eliminate popularity bias for online recommendation. IEEE Transactions on Neural Networks and Learning Systems 35 (12), pp. 17123–17136. Cited by: §1.

Appendix Contents

Appendix A Theoretical Foundations

Before presenting the proofs, we introduce several technical facts that will be used repeatedly. We begin with a standard property of independent random variables from 49, which is used in the proof of Theorem 1 and its covariate-adjusted extension.

Theorem 4 (Theorem 2.2.5 in 49).

Let X1,X2,,XnX_{1},X_{2},\ldots,X_{n} be independent random variables, and for i=1,,ni=1,\ldots,n, gig_{i} be a function gi:g_{i}:\mathbb{R}\to\mathbb{R}. Then the random variables g1(X1),g2(X2),,gn(Xn)g_{1}(X_{1}),g_{2}(X_{2}),\ldots,g_{n}(X_{n}) are also independent.

We also use the following direct extension, which states that measurable functions of disjoint subsets of mutually independent random variables remain independent.

Lemma 1.

Let X1,,XnX_{1},\dots,X_{n} be independent random variables. Suppose that g:kg:\mathbb{R}^{k}\to\mathbb{R} is a measurable function of (X1,,Xk)(X_{1},\dots,X_{k}), h:nkh:\mathbb{R}^{n-k}\to\mathbb{R} is a measurable function of (Xk+1,,Xn)(X_{k+1},\dots,X_{n}), then g(X1,,Xk)g(X_{1},\dots,X_{k}) and h(Xk+1,,Xn)h(X_{k+1},\dots,X_{n}) are independent random variables.

Proof.

Let

𝐗1=(X1,,Xk),𝐗2=(Xk+1,,Xn).\mathbf{X}_{1}=(X_{1},\dots,X_{k}),\qquad\mathbf{X}_{2}=(X_{k+1},\dots,X_{n}).

Since X1,,XnX_{1},\dots,X_{n} are mutually independent, the random vectors 𝐗1\mathbf{X}_{1} and 𝐗2\mathbf{X}_{2} are independent. Indeed, for any Borel sets AkA\subseteq\mathbb{R}^{k} and BnkB\subseteq\mathbb{R}^{n-k}, the independence of X1,,XnX_{1},\dots,X_{n} implies

(𝐗1A,𝐗2B)=(𝐗1A)(𝐗2B).\mathbb{P}(\mathbf{X}_{1}\in A,\mathbf{X}_{2}\in B)=\mathbb{P}(\mathbf{X}_{1}\in A)\mathbb{P}(\mathbf{X}_{2}\in B).

Now, for any Borel sets C,DC,D\subseteq\mathbb{R}, we have

(g(𝐗1)C,h(𝐗2)D)\displaystyle\mathbb{P}\big(g(\mathbf{X}_{1})\in C,\ h(\mathbf{X}_{2})\in D\big)
=(𝐗1g1(C),𝐗2h1(D)).\displaystyle=\mathbb{P}\big(\mathbf{X}_{1}\in g^{-1}(C),\ \mathbf{X}_{2}\in h^{-1}(D)\big).

Since gg and hh are measurable, g1(C)g^{-1}(C) and h1(D)h^{-1}(D) are Borel sets. Using the independence of 𝐗1\mathbf{X}_{1} and 𝐗2\mathbf{X}_{2}, we obtain

(𝐗1g1(C),𝐗2h1(D))\displaystyle\mathbb{P}\big(\mathbf{X}_{1}\in g^{-1}(C),\ \mathbf{X}_{2}\in h^{-1}(D)\big)
=(𝐗1g1(C))(𝐗2h1(D))\displaystyle=\mathbb{P}\big(\mathbf{X}_{1}\in g^{-1}(C)\big)\mathbb{P}\big(\mathbf{X}_{2}\in h^{-1}(D)\big)
=(g(𝐗1)C)(h(𝐗2)D).\displaystyle=\mathbb{P}\big(g(\mathbf{X}_{1})\in C\big)\mathbb{P}\big(h(\mathbf{X}_{2})\in D\big).

Therefore, g(X1,,Xk)g(X_{1},\dots,X_{k}) and h(Xk+1,,Xn)h(X_{k+1},\dots,X_{n}) are independent. ∎

Next, we recall a separability result for twice continuously differentiable functions and apply it to log-densities to connect independence with cross second-order partial derivatives. This result will be used in the proofs of Proposition 2 and Theorem 2.

Theorem 5 (45).

The Hessian HfH_{f} of function ff is block diagonal everywhere, ijf|s0=0\partial_{i}\partial_{j}f\big|_{\vec{s_{0}}}=0 for all points s0\vec{s_{0}} and all iki\leq k, j>kj>k, if and only if f is separable into a sum f(s1,,sn)=g(s1,,sk)+h(sk+1,,sn)f(s_{1},...,s_{n})=g(s_{1},...,s_{k})+h(s_{k+1},...,s_{n}) for some functions gg and hh.

Theorem 5 implies that if the log-density of two groups of variables is additively separable, then the corresponding cross second-order partial derivatives vanish.

Appendix B Proofs

B.1 Proof of Theorem 1

Proof.

To prove Theorem 1, we need to show that if the candidate IV set 𝐙\mathbf{Z}^{\prime} is a valid IV set relative to XYX\to Y under CAM-NCE, then for any pair {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime}, {X,Y||{Zi,Zj}}\{X,Y||\{Z_{i},Z_{j}\}\} will satisfy the CAT condition.

Under CAM-NCE, the data-generating process is

X\displaystyle X =g(𝐙)+φX(𝐔)+εX,\displaystyle=g(\mathbf{Z})+\varphi_{X}(\mathbf{U})+\varepsilon_{X}, (6)
Y\displaystyle Y =f(X)+gY(𝐙~)+φY(𝐔)+εY,\displaystyle=f(X)+g_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y},

where 𝐔=𝜺𝑼\mathbf{U}=\bm{\varepsilon_{U}}. Since 𝐙\mathbf{Z}^{\prime} is a valid IV set, the valid IV set 𝐙\mathbf{Z}^{\prime} is not a subset of 𝐙~\widetilde{\mathbf{Z}}, and each variable in 𝐙\mathbf{Z}^{\prime} is generated solely from its corresponding noise term, i.e., Zk=εZkZ_{k}=\varepsilon_{Z_{k}}, for any k{1,,|𝐙|}k\in\{1,\dots,|\mathbf{Z}^{\prime}|\}.

Consider any pair of distinct IVs {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime}. According to the definition of auxiliary variable w.r.t.XYX\to Y relative to ZiZ_{i}, 𝒱XY||ZiYhi(X)\mathcal{V}_{X\to Y||Z_{i}}\coloneqq Y-h_{i}(X), where hi()h_{i}(\cdot) satisfies 𝔼[𝒱XY||Zi|Zi]=0\mathbb{E}[\mathcal{V}_{X\to Y||Z_{i}}|Z_{i}]=0 and hi()0h_{i}(\cdot)\neq 0. Because ZiZ_{i} is a valid IV and Assumption 1 holds, the standard identification argument for nonparametric IV models implies that the solution is unique and coincides with the structural response function f()f(\cdot) (53; 6; 62); that is, hi()=f()h_{i}(\cdot)=f(\cdot). Similarly, since ZjZ_{j} is also a valid IV, we have hj()=f()h_{j}(\cdot)=f(\cdot).

Therefore, we can further express the auxiliary variable as:

{𝒱XY||Zi=Yhi(X)=gY(𝐙~)+φY(𝐔)+εY,𝒱XY||Zj=Yhj(X)=gY(𝐙~)+φY(𝐔)+εY.\displaystyle\left\{\begin{matrix}\mathcal{V}_{X\to Y||Z_{i}}=Y-h_{i}(X)=g_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y},\\ \mathcal{V}_{X\to Y||Z_{j}}=Y-h_{j}(X)=g_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}.\end{matrix}\right. (7)

By Theorem 4 and its extension Lemma 1, if random variables are mutually independent, then any measurable functions applied to disjoint subsets of them yield independent random variables (see Theorem 4, and Lemma 1 for further details). Based on this result, we next show that the auxiliary variable 𝒱XY||Zj\mathcal{V}_{X\to Y||Z_{j}} and ZiZ_{i} are statistically independent. Specifically, since all noise terms of variables are mutually independent and Zj𝐙𝐙~Z_{j}\in\mathbf{Z}^{\prime}\nsubseteq\widetilde{\mathbf{Z}}, we can obtain that εZj\varepsilon_{Z_{j}} is independent of gY(𝐙~)+φY(𝜺𝑼)+εYg_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\bm{\varepsilon_{U}})+\varepsilon_{Y}. Furthermore, combining Equations (6) and (7), we conclude that ZjZ_{j} is independent of 𝒱XY||Zi\mathcal{V}_{X\to Y||{Z_{i}}}, i.e., Zj𝒱XY||ZiZ_{j}\mathrel{\perp\mspace{-10mu}\perp}\mathcal{V}_{X\to Y||{Z_{i}}}.

Likewise, for pairwise (𝒱XY||Zj,Zi\mathcal{V}_{X\to Y||Z_{j}},Z_{i}), we can derive that 𝒱XY||Zj\mathcal{V}_{X\to Y||Z_{j}} is independent of ZiZ_{i}, i.e., Zi𝒱XY||ZjZ_{i}\mathrel{\perp\mspace{-10mu}\perp}\mathcal{V}_{X\to Y||{Z_{j}}}. To sum up, {X,Y||{Zi,Zj}}\{X,Y||\{Z_{i},Z_{j}\}\} always satisfies the CAT condition.

Similarly, the same argument applies to any variable pair in {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime}, implying that every such pair satisfies the CAT condition. Therefore, {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} satisfies the CAT condition. ∎

B.2 Proof of Proposition 1

Proof.

We prove the proposition for the candidate IV set 𝐙𝐙{\mathbf{Z}}^{\prime}\subseteq\mathbf{Z}. Since the candidate IVs in 𝐙{\mathbf{Z}}^{\prime} violate only the exclusion restriction, they are relevant and exogenous, but directly affect the outcome. The data-generating process can be written as

X\displaystyle X =gX(𝐙)+ϕX(𝐙𝐙)+φX(𝐔)+εX,\displaystyle=g_{X}({\mathbf{Z}}^{\prime})+\phi_{X}(\mathbf{Z}\setminus{\mathbf{Z}}^{\prime})+\varphi_{X}(\mathbf{U})+\varepsilon_{X}, (8)
Y\displaystyle Y =f(X)+gY(𝐙)+ϕY(𝐙~𝐙)+φY(𝐔)+εY,\displaystyle=f(X)+g_{Y}({\mathbf{Z}}^{\prime})+\phi_{Y}(\widetilde{\mathbf{Z}}\setminus{\mathbf{Z}}^{\prime})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y},

where gY(𝐙)=agX(𝐙)+bg_{Y}({\mathbf{Z}}^{\prime})=a\cdot g_{X}({\mathbf{Z}}^{\prime})+b, and a0a\neq 0. Substituting this relation into the structural equation of YY gives

Y\displaystyle Y =f(X)+agX(𝐙)+b+ϕY(𝐙𝐙)+φY(𝐔)+εY\displaystyle=f(X)+a\cdot g_{X}({\mathbf{Z}}^{\prime})+b+\phi_{Y}(\mathbf{Z}\setminus{\mathbf{Z}}^{\prime})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}
=f(X)+a{XϕX(𝐙𝐙)φX(𝐔)εX}\displaystyle=f(X)+a\cdot\{X-\phi_{X}(\mathbf{Z}\setminus{\mathbf{Z}}^{\prime})-\varphi_{X}(\mathbf{U})-\varepsilon_{X}\}
+b+ϕY(𝐙𝐙)+φY(𝐔)+εY\displaystyle+b+\phi_{Y}(\mathbf{Z}\setminus{\mathbf{Z}}^{\prime})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}
=f(X)a{ϕX(𝐙𝐙)+φX(𝐔)+εX}\displaystyle=f^{\prime}(X)-a\cdot\{\phi_{X}(\mathbf{Z}\setminus{\mathbf{Z}}^{\prime})+\varphi_{X}(\mathbf{U})+\varepsilon_{X}\}
+b+ϕY(𝐙𝐙)+φY(𝐔)+εY,\displaystyle+b+\phi_{Y}(\mathbf{Z}\setminus{\mathbf{Z}}^{\prime})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y},

where f(X)=f(X)+aXf^{\prime}(X)=f(X)+a\cdot X.

Now construct an alternative structural representation with Zi=ZiZ_{i}^{\prime}=Z_{i} for i[|𝐙|]i\in[|\mathbf{Z}|], i.e., 𝐙alter=𝐙\mathbf{Z}_{alter}=\mathbf{Z}, X=XX^{\prime}=X, and

Y=\displaystyle Y^{\prime}= f(X)a{ϕX({𝐙𝐙}alter)+φX(𝐔)+εX}\displaystyle f^{\prime}(X^{\prime})-a\cdot\{\phi_{X}(\{\mathbf{Z}\setminus{\mathbf{Z}}^{\prime}\}_{alter})+\varphi_{X}(\mathbf{U})+\varepsilon_{X}\}
+b+ϕY({𝐙𝐙}alter)+φY(𝐔)+εY.\displaystyle+b+\phi_{Y}(\{\mathbf{Z}\setminus{\mathbf{Z}}^{\prime}\}_{alter})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}.

Then Y=YY^{\prime}=Y, and hence

(X,Y,Z1,,Z|𝐙|)=𝑑(X,Y,Z1,,Z|𝐙|).(X,Y,Z_{1},\ldots,Z_{|\mathbf{Z}|})\overset{d}{=}(X^{\prime},Y^{\prime},Z_{1}^{\prime},\ldots,Z_{|\mathbf{Z}|}^{\prime}).

Furthermore, we have

(X,Y,𝐙)=𝑑(X,Y,𝐙alter).(X,Y,\mathbf{Z}^{\prime})\overset{d}{=}(X^{\prime},Y^{\prime},\mathbf{Z}_{alter}^{\prime}).

In this alternative representation, the variables 𝐙alter\mathbf{Z}_{alter}^{\prime} affect YY^{\prime} only through XX^{\prime} and have no direct effects on YY^{\prime}. Therefore, 𝐙alter\mathbf{Z}_{alter}^{\prime} forms a valid IV set with respect to XYX^{\prime}\to Y^{\prime}.

By Theorem 1, the CAT condition holds for the alternative representation. For any pair {Zi,Zj}𝐙alter\{Z_{i}^{\prime},Z_{j}^{\prime}\}\subseteq\mathbf{Z}_{alter}^{\prime} with iji\neq j, the corresponding auxiliary variable is

𝒱XY|Zi\displaystyle\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}} =Yf(X)\displaystyle=Y^{\prime}-f^{\prime}(X^{\prime})
=ϕY({𝐙𝐙}alter)+φY(𝐔)+εY+b\displaystyle=\phi_{Y}(\{\mathbf{Z}\setminus{\mathbf{Z}}^{\prime}\}_{alter})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y}+b
a{ϕX({𝐙𝐙}alter)+φX(𝐔)+εX}.\displaystyle-a\{\phi_{X}(\{\mathbf{Z}\setminus{\mathbf{Z}}^{\prime}\}_{alter})+\varphi_{X}(\mathbf{U})+\varepsilon_{X}\}.

Since each ZjZ_{j}^{\prime} is independent of 𝐔\mathbf{U}, εX\varepsilon_{X}, εY\varepsilon_{Y}, and other variables {𝐙𝐙}alter\{\mathbf{Z}\setminus\mathbf{Z}^{\prime}\}_{alter}, it follows from Lemma 1 that

𝒱XY|ZiZj,ij.\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}}\mathrel{\perp\mspace{-10mu}\perp}Z_{j}^{\prime},\qquad\forall i\neq j.

Consequently, p(𝒱XY|Zi,Zj)=p(𝒱XY|Zi)p(Zj)p(\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}},Z_{j}^{\prime})=p(\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}})\,p(Z_{j}^{\prime}), which implies

logp(𝒱XY|Zi,Zj)=logp(𝒱XY|Zi)+logp(Zj).\log p(\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}},Z_{j}^{\prime})=\log p(\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}})+\log p(Z_{j}^{\prime}).

Therefore,

2logp(𝒱XY|Zi,Zj)𝒱XY|ZiZj=0.\frac{\partial^{2}\log p(\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}},Z_{j}^{\prime})}{\partial\mathcal{V}_{X^{\prime}\to Y^{\prime}\|Z_{i}^{\prime}}\partial Z_{j}^{\prime}}=0.

Because the original and alternative representations induce the same observational distribution, we also have

2logp(𝒱XY|Zi,Zj)𝒱XY|ZiZj=0,ij.\frac{\partial^{2}\log p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j})}{\partial\mathcal{V}_{X\to Y\|Z_{i}}\partial Z_{j}}=0,\qquad\forall i\neq j.

The same argument applies after exchanging ZiZ_{i} and ZjZ_{j}. Consequently, all the cross second-order partial derivatives are zero. As a result, {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} violates Assumption 2. ∎

B.3 Proof of Proposition 2

Proof.

Suppose that the candidate IV set 𝐙\mathbf{Z}^{\prime} is invalid. By Assumption 2, there exists a pair of distinct candidate IVs {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime} such that the joint densities p(𝒱XY|Zi,Zj)p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j}) and p(𝒱XY|Zj,Zi)p(\mathcal{V}_{X\to Y\|Z_{j}},Z_{i}) are twice continuously differentiable, and at least one of the following cross second-order partial derivatives is nonzero on a set with nonzero Lebesgue measure:

2logp(𝒱XY|Zi,Zj)𝒱XY|ZiZjor2logp(𝒱XY|Zj,Zi)𝒱XY|ZjZi.\frac{\partial^{2}\log p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j})}{\partial\mathcal{V}_{X\to Y\|Z_{i}}\,\partial Z_{j}}\quad\text{or}\quad\frac{\partial^{2}\log p(\mathcal{V}_{X\to Y\|Z_{j}},Z_{i})}{\partial\mathcal{V}_{X\to Y\|Z_{j}}\,\partial Z_{i}}.

We prove the result by contradiction. Assume that {X,Y𝐙}\{X,Y\|\mathbf{Z}^{\prime}\} satisfies the CAT condition. Then, for every pair of distinct IVs {Zi,Zj}𝐙\{Z_{i},Z_{j}\}\subseteq\mathbf{Z}^{\prime}, we have

𝒱XY|ZiZj,𝒱XY|ZjZi.\mathcal{V}_{X\to Y\|Z_{i}}\mathrel{\perp\mspace{-10mu}\perp}Z_{j},\qquad\mathcal{V}_{X\to Y\|Z_{j}}\mathrel{\perp\mspace{-10mu}\perp}Z_{i}.

In particular, for the pair specified by Assumption 2, the first independence relation implies

p(𝒱XY|Zi,Zj)=p(𝒱XY|Zi)p(Zj).p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j})=p(\mathcal{V}_{X\to Y\|Z_{i}})p(Z_{j}).

Taking logarithms yields the additive decomposition

logp(𝒱XY|Zi,Zj)=logp(𝒱XY|Zi)+logp(Zj).\log p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j})=\log p(\mathcal{V}_{X\to Y\|Z_{i}})+\log p(Z_{j}).

Therefore, the corresponding cross second-order partial derivative must vanish:

2logp(𝒱XY|Zi,Zj)𝒱XY|ZiZj=0.\frac{\partial^{2}\log p(\mathcal{V}_{X\to Y\|Z_{i}},Z_{j})}{\partial\mathcal{V}_{X\to Y\|Z_{i}}\,\partial Z_{j}}=0.

Similarly, the second independence relation 𝒱XY|ZjZi\mathcal{V}_{X\to Y\|Z_{j}}\mathrel{\perp\mspace{-10mu}\perp}Z_{i} implies

2logp(𝒱XY|Zj,Zi)𝒱XY|ZjZi=0.\frac{\partial^{2}\log p(\mathcal{V}_{X\to Y\|Z_{j}},Z_{i})}{\partial\mathcal{V}_{X\to Y\|Z_{j}}\,\partial Z_{i}}=0.

Thus, both cross second-order partial derivatives vanish, which contradicts Assumption 2. Hence, {X,Y𝐙}\{X,Y\|\mathbf{Z}^{\prime}\} cannot satisfy the CAT condition. Therefore, if 𝐙\mathbf{Z}^{\prime} is invalid, then {X,Y𝐙}\{X,Y\|\mathbf{Z}^{\prime}\} violates the CAT condition. ∎

B.4 Proof of Theorem 2

Proof.

We prove the necessary and sufficient characterization of valid IV sets under CAM-NCE by establishing the following two implications.

(i): Assume the candidate IV set 𝐙\mathbf{Z}^{\prime} is a valid IV set relative to XYX\to Y. By Theorem 1, under Assumption 1, it directly follows that if the candidate IV set 𝐙\mathbf{Z}^{\prime} is a valid IV set relative to XYX\to Y, then {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} always satisfies the CAT condition.

(ii): Assume the candidate IV set 𝐙\mathbf{Z}^{\prime} is an invalid IV set relative to XYX\to Y. By Proposition 2, under Assumptions 1 and 2, if the candidate IV set 𝐙\mathbf{Z}^{\prime} is invalid, then {X,Y||𝐙}\{X,Y||\mathbf{Z}^{\prime}\} consequently violates the CAT condition.

Combining (i) and (ii), we conclude that 𝐙\mathbf{Z}^{\prime} is a valid IV set w.r.t. XYX\to Y if and only if {X,Y𝐙}\{X,Y\|\mathbf{Z}^{\prime}\} satisfies the CAT condition. ∎

B.5 Proof of Corollary 1

Proof.

Suppose that 𝐙\mathbf{Z}^{\prime} is a valid IV set w.r.t. XYX\to Y given 𝐖\mathbf{W}. Under CAM-NCE with covariates, the data-generating mechanism can be written as

𝐔\displaystyle\mathbf{U} =𝜺𝑼,𝐖=tW(𝐏𝐀𝐖)+𝜺𝑾,\displaystyle=\bm{\varepsilon_{U}},\qquad\mathbf{W}=t_{W}(\mathbf{PA}_{\mathbf{W}})+\bm{\varepsilon_{W}}, (9)
𝐙\displaystyle\mathbf{Z}^{\prime} =ψ𝐙(𝐙)+t𝐙(𝐖)+ε𝐙,\displaystyle=\psi_{\mathbf{Z}^{\prime}}(\mathbf{Z}^{\prime})+t_{\mathbf{Z}^{\prime}}(\mathbf{W})+\varepsilon_{\mathbf{Z}^{\prime}},
X\displaystyle X =gX(𝐙)+tX(𝐖)+φX(𝐔)+εX,\displaystyle=g_{X}(\mathbf{Z})+t_{X}(\mathbf{W})+\varphi_{X}(\mathbf{U})+\varepsilon_{X},
Y\displaystyle Y =f(X)+tY(𝐖)+gY(𝐙~)+φY(𝐔)+εY,\displaystyle=f(X)+t_{Y}(\mathbf{W})+g_{Y}(\widetilde{\mathbf{Z}})+\varphi_{Y}(\mathbf{U})+\varepsilon_{Y},

where 𝐏𝐀𝐖\mathbf{PA}_{\mathbf{W}} denotes the set of parent variables of 𝐖\mathbf{W}, and ψ𝐙\psi_{\mathbf{Z}^{\prime}} denotes the causal relationship among variables in 𝐙\mathbf{Z}^{\prime}. Since 𝐙\mathbf{Z}^{\prime} is valid given 𝐖\mathbf{W}, no variable in 𝐙\mathbf{Z}^{\prime} directly affects YY; equivalently,

𝐙𝐙~=.\mathbf{Z}^{\prime}\cap\widetilde{\mathbf{Z}}=\emptyset.

Let 𝒳\mathcal{X}, 𝒴\mathcal{Y}, and 𝓩\bm{\mathcal{Z}}^{\prime} denote the residuals obtained by regressing XX, YY, and 𝐙\mathbf{Z}^{\prime} on 𝐖\mathbf{W}, respectively. Under the additive covariate structure in Equation (9) and the regression adjustment in Definition 5, the residualized variables remove the effects of 𝐖\mathbf{W} and follow the same structural form as the covariate-free CAM-NCE model. Moreover, because 𝐙\mathbf{Z}^{\prime} is a valid IV set w.r.t. XYX\to Y given 𝐖\mathbf{W}, the residualized candidate IV set 𝓩\bm{\mathcal{Z}}^{\prime} satisfies the corresponding relevance, exclusion restriction, and exogeneity conditions in the residualized system.

Therefore, applying Theorem 1 to the residualized variables yields that every pair of distinct residualized IVs in 𝓩\bm{\mathcal{Z}}^{\prime} satisfies the CAT condition. Equivalently, {X,Y(𝐙,𝐖)}\{X,Y\|(\mathbf{Z}^{\prime},\mathbf{W})\} satisfies the CAT condition. ∎

B.6 Proof of Corollary 2

Proof.

Let 𝒳\mathcal{X}, 𝒴\mathcal{Y}, and 𝓩\bm{\mathcal{Z}}^{\prime} denote the residuals obtained by regressing XX, YY, and 𝐙\mathbf{Z}^{\prime} on 𝐖\mathbf{W}, respectively. As shown in the proof of Corollary 1, under the additive covariate structure and the regression adjustment in Definition 5, the residualized variables 𝒳\mathcal{X}, 𝒴\mathcal{Y}, and 𝓩\bm{\mathcal{Z}}^{\prime} follow the same structural form as the covariate-free CAM-NCE model.

Furthermore, Assumption 1 and Assumption 2 are assumed to hold for these covariate-adjusted residual variables. Therefore, the residualized system satisfies the conditions required by Theorem 2. Applying Theorem 2 to the residualized variables yields that the residualized candidate IV set 𝓩\bm{\mathcal{Z}}^{\prime} is valid for the causal relation 𝒳𝒴\mathcal{X}\to\mathcal{Y} if and only if the corresponding CAT condition holds.

By Definition 5, this is equivalent to saying that 𝐙\mathbf{Z}^{\prime} is a valid IV set w.r.t. XYX\to Y given 𝐖\mathbf{W} if and only if {X,Y(𝐙,𝐖)}\{X,Y\|(\mathbf{Z}^{\prime},\mathbf{W})\} satisfies the CAT condition. ∎

B.7 Proof of Theorem 3

Proof.

Let 𝐙V𝐙\mathbf{Z}_{V}\subseteq\mathbf{Z} denote the set of all valid IVs among the candidate IVs, and let KV=|𝐙V|K_{V}=|\mathbf{Z}_{V}|. By assumption, KVKK_{V}\geq K. We assume K2K\geq 2, so that the pairwise CAT criterion is informative.

Algorithm 1 first removes the effect of covariates 𝐖\mathbf{W}, if present, and then constructs, for each candidate IV 𝒵i\mathcal{Z}_{i}, the auxiliary variable 𝒱^𝒳𝒴|𝒵i\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}}. By the assumed consistency of the estimators used for covariate adjustment, auxiliary-variable construction, and distance correlation, for every unordered pair {𝒵i,𝒵j}𝓩\{\mathcal{Z}_{i},\mathcal{Z}_{j}\}\subseteq\bm{\mathcal{Z}},

𝐌^ij=dCor^(𝒱^𝒳𝒴|𝒵i,𝒵^j)+dCor^(𝒱^𝒳𝒴|𝒵j,𝒵^i)\widehat{\mathbf{M}}_{ij}=\widehat{\operatorname{dCor}}\bigl(\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}},\widehat{\mathcal{Z}}_{j}\bigr)+\widehat{\operatorname{dCor}}\bigl(\widehat{\mathcal{V}}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{j}},\widehat{\mathcal{Z}}_{i}\bigr)

converges in probability to

𝐌ij=dCor(𝒱𝒳𝒴|𝒵i,𝒵j)+dCor(𝒱𝒳𝒴|𝒵j,𝒵i).\mathbf{M}_{ij}=\operatorname{dCor}\bigl(\mathcal{V}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}},\mathcal{Z}_{j}\bigr)+\operatorname{dCor}\bigl(\mathcal{V}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{j}},\mathcal{Z}_{i}\bigr).

For any candidate subset 𝒮c𝓩\mathcal{S}_{c}\subseteq\bm{\mathcal{Z}} with |𝒮c|=K|\mathcal{S}_{c}|=K, define the population objective

T𝒮c={𝒵i,𝒵j}𝒮ci<j𝐌ij,T_{\mathcal{S}_{c}}=\sum_{\begin{subarray}{c}\{\mathcal{Z}_{i},\mathcal{Z}_{j}\}\subseteq\mathcal{S}_{c}\\ i<j\end{subarray}}\mathbf{M}_{ij},

and let T^𝒮c\widehat{T}_{\mathcal{S}_{c}} be the corresponding empirical objective computed by Algorithm 1. Since the number of candidate subsets of size KK is finite and each 𝐌^ij\widehat{\mathbf{M}}_{ij} is consistent, we have the uniform convergence

max𝒮c:|𝒮c|=K|T^𝒮cT𝒮c|𝑝0.\max_{\mathcal{S}_{c}:\,|\mathcal{S}_{c}|=K}\bigl|\widehat{T}_{\mathcal{S}_{c}}-T_{\mathcal{S}_{c}}\bigr|\overset{p}{\longrightarrow}0.

Now consider any subset 𝒮c𝓩V\mathcal{S}_{c}\subseteq\bm{\mathcal{Z}}_{V} with |𝒮c|=K|\mathcal{S}_{c}|=K. Every element of 𝒮c\mathcal{S}_{c} is a valid IV. Hence, by Theorem 2, the CAT condition holds for every pair {𝒵i,𝒵j}𝒮c\{\mathcal{Z}_{i},\mathcal{Z}_{j}\}\subseteq\mathcal{S}_{c}, namely

𝒱𝒳𝒴|𝒵i𝒵j,𝒱𝒳𝒴|𝒵j𝒵i.\mathcal{V}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}}\mathrel{\perp\mspace{-10mu}\perp}\mathcal{Z}_{j},\qquad\mathcal{V}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{j}}\mathrel{\perp\mspace{-10mu}\perp}\mathcal{Z}_{i}.

Since distance correlation is zero if and only if independence holds, it follows that 𝐌ij=0\mathbf{M}_{ij}=0 for every pair in 𝒮c\mathcal{S}_{c}. Therefore, T𝒮c=0T_{\mathcal{S}_{c}}=0.

Conversely, let 𝒮c\mathcal{S}_{c} be any subset of size KK that contains at least one invalid IV. Then 𝒮c\mathcal{S}_{c} is not a valid IV set. By the necessity and sufficiency result in Theorem 2, under Assumptions 12, the CAT condition cannot hold for all pairs in 𝒮c\mathcal{S}_{c}. Hence there exists at least one pair {𝒵i,𝒵j}𝒮c\{\mathcal{Z}_{i},\mathcal{Z}_{j}\}\subseteq\mathcal{S}_{c} such that

𝒱𝒳𝒴|𝒵i 𝒵jor𝒱𝒳𝒴|𝒵j 𝒵i.\mathcal{V}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{i}}\mathchoice{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\displaystyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\textstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 16.49544pt\kern-4.45831pt$\scriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 13.51866pt\kern-3.95834pt$\scriptscriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}\mathcal{Z}_{j}\quad\text{or}\quad\mathcal{V}_{\mathcal{X}\to\mathcal{Y}\|\mathcal{Z}_{j}}\mathchoice{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\displaystyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 22.53012pt\kern-5.27776pt$\textstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 16.49544pt\kern-4.45831pt$\scriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}{\mathrel{\hbox to0.0pt{\kern 13.51866pt\kern-3.95834pt$\scriptscriptstyle\not$\hss}{\mathrel{\perp\mspace{-10mu}\perp}}}}\mathcal{Z}_{i}.

For this pair, at least one of the two distance correlations is strictly positive. Thus 𝐌ij>0\mathbf{M}_{ij}>0, and consequently T𝒮c>0T_{\mathcal{S}_{c}}>0.

Therefore, every valid subset of size KK attains the population minimum value 00, whereas every subset of size KK containing at least one invalid IV has a strictly positive population objective value.

Because the number of candidate subsets is finite, the collection of invalid subsets of size KK is also finite. If this collection is nonempty, define

Δ=min𝒮c:|𝒮c|=K𝒮c𝐙VT𝒮c.\Delta=\min_{\begin{subarray}{c}\mathcal{S}_{c}:\,|\mathcal{S}_{c}|=K\\ \mathcal{S}_{c}\not\subseteq\mathbf{Z}_{V}\end{subarray}}T_{\mathcal{S}_{c}}.

From the preceding argument, Δ>0\Delta>0. On the event

max𝒮c:|𝒮c|=K|T^𝒮cT𝒮c|<Δ2,\max_{\mathcal{S}_{c}:\,|\mathcal{S}_{c}|=K}\bigl|\widehat{T}_{\mathcal{S}_{c}}-T_{\mathcal{S}_{c}}\bigr|<\frac{\Delta}{2},

every valid subset 𝒮c𝓩V\mathcal{S}_{c}\subseteq\bm{\mathcal{Z}}_{V} satisfies T^𝒮c<Δ2\widehat{T}_{\mathcal{S}_{c}}<\frac{\Delta}{2}, whereas every invalid subset satisfies T^𝒮c>Δ2\widehat{T}_{\mathcal{S}_{c}}>\frac{\Delta}{2}. Hence any empirical minimizer selected in Line 26 of Algorithm 1 must be a subset of 𝓩V\bm{\mathcal{Z}}_{V}. Since the above event has probability tending to one (i.e., (max𝒮c:|𝒮c|=K|T^𝒮cT𝒮c|<Δ2)1\mathbb{P}(\max_{\mathcal{S}_{c}:\,|\mathcal{S}_{c}|=K}\bigl|\widehat{T}_{\mathcal{S}_{c}}-T_{\mathcal{S}_{c}}\bigr|<\frac{\Delta}{2})\to 1), the algorithm outputs, with probability tending to one, a valid IV subset 𝒮^𝒁V\widehat{\mathcal{S}}\subseteq\bm{Z}_{V} with |𝒮^|=K|\widehat{\mathcal{S}}|=K.

Finally, if the candidate set 𝐙\mathbf{Z} contains exactly KK valid IVs, then KV=KK_{V}=K, and the only subset of size KK consisting entirely of valid IVs is 𝐙V\mathbf{Z}_{V} itself. Therefore,

𝒮^𝑝𝐙V,\widehat{\mathcal{S}}\overset{p}{\longrightarrow}\mathbf{Z}_{V},

in the sense that the probability that Algorithm 1 outputs the full valid IV set tends to one.

This proves the stated correctness of Algorithm 1. ∎