Testing the Validity of Instrumental Variable Sets in Causal Additive Models with Non-Constant Effects
Abstract
Instrumental variable (IV) methods are powerful for causal effect estimation with unmeasured confounding, but in practice researchers often face a set of candidate IVs whose validity is difficult to determine from observational data. This paper studies the problem of testing the validity of IV sets under Causal Additive Models with Non-Constant Effects (CAM-NCE). To address this problem, we propose a testable condition, termed the Cross Auxiliary-based independence Test (CAT) condition, for assessing IV set validity from observational data. We show that, under the completeness condition, if the CAT condition is violated, the corresponding set cannot be a valid IV set. Furthermore, under a cross distributional non-degeneracy condition, we establish that the CAT condition becomes both necessary and sufficient for characterizing valid IV sets under CAM-NCE. We then extend the CAT condition to settings with covariates and develop a practical finite-sample algorithm for testing the validity of candidate IV sets. Extensive experiments on synthetic data and three real-world datasets demonstrate the effectiveness and practical utility of the proposed method.
1Department of Applied Statistics, Beijing Technology and Business University, Beijing, China. 2Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China. 3School of Mathematical Sciences, Peking University, Beijing, China. 4School of Computer Science, Guangdong University of Technology, Guangzhou, China. 5Peng Cheng Laboratory, Shenzhen, China. 6Department of Philosophy, Carnegie Mellon University, Pittsburgh, USA. 7Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, United Arab Emirates. *Corresponding author: Feng Xie (e-mails: fengxie@btbu.edu.cn).
Keywords: Causal inference, instrumental variables, validity, non-constant causal effects, unmeasured confounders.
1 Introduction
Causal effect estimation from observational data is a fundamental problem in modern machine learning and probabilistic inference (58; 16), with broad applications in recommendation systems (13; 76; 43), online systems and sequential decision-making (60; 46; 77), and causality-aware clustering (42; 39; 38; 12). Instrumental variable (IV) methods provide a powerful approach to causal effect estimation by leveraging exogenous variables to mitigate unmeasured confounding. Informally, a valid IV must satisfy three key conditions: (1) it is associated with the treatment (relevance); (2) it affects the outcome only through the treatment (exclusion restriction); and (3) it is independent of unmeasured confounders (exogeneity) (see Figure 1 and Section 2.1 for formal definitions). For instance, in Figure 1, Colonialist Mortality serves as a valid IV for estimating the causal effect of Institutions on Economic Development.
Because of unmeasured confounding, it is generally difficult to determine from observational data alone which variables can serve as valid IVs. In practice, valid IVs may sometimes be justified a priori by domain expertise or historical knowledge. However, such information is often unavailable in many empirical settings, leaving researchers with a set of candidate IVs whose validity is uncertain. Therefore, developing data-driven methods to test IV validity remains a critical yet challenging task (15; 72).
Testing the validity of IVs has attracted considerable attention in recent years. For discrete treatment settings, several approaches address this problem by imposing testable constraints on the joint distribution, including Pearl’s seminal instrumental inequality (56), the generalized instrumental inequality (37), and their extensions (48; 55; 40; 68; 34; 50; 24). Although these methods have been successfully applied in various domains, they typically rely on discrete treatment variables, which limits their applicability in many practical settings—such as studies involving continuous exposures (e.g., vitamin levels (63)).
Another line of research studies IV validity in the Causal Additive Models with Constant Effects (CAM-CE). Existing approaches in this line can be broadly divided into two categories.
- •
IV Set. Representative methods include the Proportion of Invalid IVs approaches, such as the Majority Rule constraint, which assumes that more than of the candidate IVs are valid (29; 36; 8; 69; 31), and the Plurality Rule constraint, which requires the number of valid IVs exceeds that of any group of invalid IVs sharing the same ratio-estimator limit (32; 27; 70; 47). Another representative condition is the InSIDE constraint, which assumes that the pleiotropic effects of IVs on the outcome are uncorrelated with their effects on the treatment (7; 41; 59). In addition, the Two Valid IVs constraint requires the existence of at least two valid instruments and further relies on either the Rank-Faithfulness assumption (61; 17) or an algebraic condition discussed in 26.
- •
Single IV. Typical approaches include the IV-GIN method, which leverages the Generalized Independent Noise (GIN) condition (73) within linear non-Gaussian causal models (74), and the IV-PIM method, which is based on the Principle of Independent Mechanisms (PIM) (35) and evaluates instrument validity by decomposing the spectral measure of the covariance matrix of the covariates (9).
However, these approaches primarily focus on constant causal effects, and their theoretical guarantees do not directly extend to settings with non-constant causal effects.
Causal additive models with non-constant effects (CAM-NCE) provide a more flexible framework for capturing heterogeneous treatment–outcome relationships. Chu et al. (19) were among the first to study IV validity in this setting by introducing the concept of semi-instruments, showing that semi-instrument validity is testable under additive models. However, their testability condition relies on the prior assumption that the exogeneity condition (3) holds. More recently, 25 proposed the Auxiliary-based Independence Test (AIT) for a single IV and showed that, under certain conditions, a valid IV satisfies the AIT condition. Nevertheless, the AIT condition encounters difficulties in verifying the exclusion restriction (2) (see the illustrative example in Section 3.1 and Proposition 5 in (25)). Moreover, the exclusion restriction is often difficult to satisfy in practice. For example, in Mendelian randomization studies, genetic variants may have pleiotropic effects on the outcome, thereby violating (2) (10).
| Category | Constant Effect | Non-Constant Effect |
|---|---|---|
| Single IV | ✓ 74; and 9 | ✓ 19; and 25 |
| IV Set | ✓ 29; 36; 8; 69; 31; 32; 27; 70; 47; 7; 41; 59; 61; 17; and 26 | ✗ |
Note: The check mark ✓ indicates that studies have addressed the corresponding setting, whereas the cross ✗ denotes that research in this setting remains unexplored.
table I summarizes existing studies on IV validity testing under continuous-variable settings, categorized by whether the analysis involves a single IV or a set of IVs, and whether the underlying causal effect is constant or non-constant. As shown in the table, prior research has primarily focused on either single IV settings or IV sets under constant-effect assumptions. However, testing the validity of IV sets—particularly when the treatment–outcome relationship exhibits non-constant causal effects—remains underexplored. To bridge this gap, this paper investigates the problem of testing the validity of IV sets within the CAM-NCE framework, where candidate instruments may violate either the exclusion restriction or the exogeneity condition. Specifically, our main contributions are as follows:
- 1.
We introduce a testable condition, termed the Cross Auxiliary-based Independence Test (CAT) condition, for assessing the validity of IV sets within CAM-NCE.
- 2.
- 3.
We develop a practical testing procedure for the CAT condition that accounts for covariates and finite-sample considerations.
- 4.
We empirically validate the proposed method through extensive experiments on both synthetic and real-world datasets, demonstrating its effectiveness in detecting invalid IV sets.
2 Preliminaries and Problem Definition
2.1 Notations and Definitions
In this paper, a causal system is represented by a directed acyclic graph (DAG) , in which nodes correspond to random variables and directed edges represent direct causal influences between them. For brevity, we use “w.r.t.” to denote “with respect to”. Throughout this paper, the major symbols and notations are summarized in table II.
| Symbol | Description |
|---|---|
| A treatment (exposure) variable | |
| An outcome variable | |
| A candidate (potential) IV | |
| A candidate (potential) IV set | |
| Unmeasured confounders between and | |
| Covariates | |
| The residual of variable after regressing on covariates | |
| The residual of variable after regressing on covariates | |
| The number of variables in set | |
| The noise term of a variable | |
| The expected value of random variable | |
| The covariance between random variables and | |
| The variance of random variable | |
| is statistically independent of | |
| is statistically dependent on | |
| The field of real numbers | |
| A mapping from to | |
| The bias between estimated causal effect of on and ground-truth causal effect of on | |
| The auxiliary variable of causal relationship relative to . We often use as a shorthand when there is no ambiguity | |
| The sample estimate of the corresponding population quantity |
We next introduce the formal definitions of an IV and an IV set, which will be used throughout the paper.
Definition 1 (Instrumental Variable (IV) (57)).
A variable is said to be an IV w.r.t. the causal relation if it satisfies the following three conditions:
- . (Relevance)
-
is associated with the treatment , i.e., .
- . (Exclusion Restriction)
-
does not directly affect the outcome , i.e., .
- . (Exogeneity)
-
is independent of the unmeasured confounders , i.e., .
Definition 2 (IV Set).
A set of variables is said to be a valid IV set w.r.t. if every nonempty subset satisfies Conditions – of Definition 1, with replaced by . Otherwise, is referred to as an invalid IV set w.r.t..
As illustrated in Figure 2, Figure 2(a) shows a valid IV set , where both variables satisfy Conditions –. By contrast, Figure 2(b) shows an invalid IV set in which violates the exclusion restriction () due to the direct edge , while Figure 2(c) shows another invalid IV set in which violates exogeneity () due to its dependence on the unmeasured confounders .
2.2 Causal Additive Models with Non-Constant Effects
In this paper, we focus on Causal Additive Models with Non-Constant Effects (CAM-NCE), involving the treatment , the outcome , a candidate IV set , and unmeasured confounders . Without loss of generality, all variables are assumed to be mean-centered. We first present the covariate-free case for simplicity, and discuss the extension to baseline covariates in section 4.1. Specifically, the data generation process under CAM-NCE can be expressed as:
| (1) | ||||
where the function represents the true but unknown causal effect of interest. The functions , , and are assumed to be smooth functions defined on appropriate domains, i.e., , , , and . The noise terms and are assumed to be mutually independent. The nonzero function indicates that the subset directly affects the outcome , thereby violating the exclusion restriction condition (). Furthermore, if there exists any that is statistically dependent on , this implies a violation of the exogeneity condition ().
Note that, even given a valid IV, identifying the nonparametric structural function generally requires additional conditions. Following the standard additive nonparametric IV literature, we assume that a solution exists and impose the following completeness condition to ensure uniqueness given a valid IV (53; 2; 21; 18; 54; 62; 6).
Assumption 1 (Completeness Condition).
Given a valid IV , for any measurable function satisfying , almost surely if and only if almost surely.
The completeness condition is generic, in the sense that it holds for “most” —where denotes the conditional density of given —if it holds for one (54; 4). In particular, the finite-support case and models belonging to exponential families (such as Gaussian, Poisson, Binomial, or certain multivariate extensions) are known to satisfy completeness (53; 33). Further sufficient conditions for completeness have been established in 23; 4; and 33.
Remark 1.
We highlight the following points regarding the proposed CAM-NCE model:
- 1.
Relation to constant-effect models. When the functions , , and are linear, the above model reduces to the Causal Additive Model with Constant Effects (CAM-CE), such as the Additive Linear, Constant-Effects Model (ALICE) (29; 36; 8; 69; 31; 32; 27; 70; 47; 7; 41; 59; 61; 17). In this work, we investigate a more challenging and general setting where , , and can be nonlinear.
- 2.
Conditions to be assessed. Condition (relevance) can be readily assessed through standard statistical dependence tests. Hence, our primary focus lies in addressing the remaining Conditions and . Notably, causal discovery methods based on conditional independence tests, such as the FCI (Fast Causal Inference) algorithm (64) and its extensions (20; 3), often yield a fully connected subgraph over in the presence of unmeasured confounders (since ). As a result, verifying Conditions – and identifying valid IVs remain challenging tasks.
- 3.
Our Goal. The goal of this work is to develop a data-driven framework for testing the validity of IV sets under CAM-NCE, where Assumption 1 holds. Specifically, given a candidate IV set for the causal relation , we aim to assess whether the variables in satisfy Conditions –.
3 CAT Condition for Testing IV Set Validity
In this section, we introduce a testable criterion for assessing the validity of IV sets under CAM-NCE, termed the Cross Auxiliary-based Independence Test (CAT) condition. We first show that, under the completeness condition (Assumption 1), the CAT condition is necessary for IV set validity. We then present a counterexample showing that completeness alone is insufficient to rule out all invalid IV sets. Finally, by imposing an additional cross distributional non-degeneracy condition (Assumption 2), we establish that the CAT condition becomes both necessary and sufficient for characterizing valid IV sets.
3.1 CAT Condition: Definition and Necessity
We first introduce the key concept of the auxiliary variable together with the definition of the CAT condition for an IV set, which characterizes cross-independence relationships between the auxiliary variable and another distinct IV.
Definition 3 (Auxiliary Variable).
Let , , and denote the treatment, outcome, and candidate IV, respectively. The auxiliary variable for the causal relationship relative to is defined as
| (2) |
where is a nonzero function satisfying .
The notion of an auxiliary variable, or related residual-type constructions, has been used in various contexts (22; 14; 11; 75; 25). Different from these works, we use auxiliary variables to construct cross-independence relations among candidate IVs. To the best of our knowledge, such cross-independence relations have not been used to assess IV set validity under CAM-NCE.
Definition 4 (CAT Condition).
Let , , and denote the treatment, outcome, and a candidate IV set, respectively. We say that satisfies the Cross Auxiliary-based Independence Test (CAT) condition if and only if, for every pair of distinct IVs , the following cross-independence relationships hold:
| (3) |
Intuitively, the CAT condition performs a cross-check among candidate IVs: the auxiliary variable constructed using one candidate IV is tested for independence from another candidate IV in the set. For a pair , CAT requires both and to hold.
We next provide a simple example to illustrate how the CAT condition detects violations of IV validity.
Example 1 (Intuitive Example of the CAT Condition).
Consider the two causal graphs shown in Figure 2(a) and Figure 2(b). Suppose that the corresponding data-generating mechanisms are given as follows:
Assume that all noise terms are mutually independent and follow a standard Gaussian distribution. Let .
In Figure 2(a), both and are valid IVs. For either reference IV, the auxiliary variable removes the causal contribution and becomes
Since , , and are mutually independent, we have and . Hence, satisfies the CAT condition. This is visualized in Figure 3, where and show no apparent dependence in the valid-IV-set case.
In Figure 2(b), however, directly affects and violates the exclusion restriction, while remains valid. Using as the reference IV, we obtain
which is dependent on . Thus, , and the CAT condition is violated. As shown in Figure 3, this invalid-IV-set case exhibits a visible dependence pattern between and , induced by the direct effect .
It is worth noting that the validity of a single IV, such as in Figure 2(b), cannot in general be verified from the observed distribution of alone. Indeed, the IV assumptions impose no testable constraints on this joint distribution, and different causal structures may induce the same observational distribution (see Section 3 of 19, and Proposition 3 of 25). This motivates the need for testable criteria that exploit relations among multiple candidate IVs.
Building on the above intuition, we next establish that the CAT condition provides a necessary condition for IV set validity under CAM-NCE.
Theorem 1 (Necessary Condition for IV Set Validity).
Let , , and be the treatment, outcome, and candidate IV set in CAM-NCE, respectively. Suppose that , , and are statistically dependent, and that Assumption 1 holds. If the candidate set is a valid IV set w.r.t., then satisfies the CAT condition.
Proof sketch.
Since is a valid IV set, each is a valid IV. Under Assumption 1, the conditional moment restriction admits a unique solution. By the standard identification result in nonparametric IV models (53; 54; 23; 33), this solution coincides with the true causal effect function , i.e., . Hence, for any pair of distinct IVs , the auxiliary variables satisfy
Here, because is a valid IV set, and do not directly affect and are independent of . Moreover, under CAM-NCE, they are also independent of the remaining variables and noise terms involved in . Using the fact that measurable functions of disjoint subsets of mutually independent random variables are independent (see Theorem 4 and Lemma 1 in Appendix), we have
Thus, every pair of distinct IVs in satisfies the CAT condition. Consequently, satisfies the CAT condition. The full proof is provided in Appendix B.1. ∎
Theorem 1 states that if violates the CAT condition, then the candidate IV set w.r.t. is invalid. Otherwise, may or may not be valid.
3.2 Sufficient Condition for Characterizing IV Set Validity
In the previous section, we have shown that the CAT condition is necessary for IV set validity. A natural question is whether, under Assumption 1, every invalid IV set necessarily violates the CAT condition. Unfortunately, the answer is negative. To clarify this issue, we present a counterexample below.
Example 2.
(Counterexample) Let , , and , with , , where constant . In this model, both and directly affect and hence violate the exclusion restriction. However, the same observational distribution can be generated by an alternative causal structure in which forms a valid IV set. Specifically, define , , , , and . Then, we have
Thus, and have the same observational distribution, although is invalid in the original structure but valid in the alternative structure. This shows that Assumption 1 alone cannot distinguish all invalid IV sets.
The above example shows that Assumption 1 alone does not guarantee that the CAT condition can rule out all invalid IV sets. To obtain a sufficient condition, we introduce an additional non-degeneracy condition that excludes those cases and makes IV set validity testable under the CAM-NCE.
Assumption 2 (Cross Distributional Non-degeneracy Condition).
For any invalid candidate IV set with , there exists a pair of distinct candidate IVs such that the joint densities and are twice continuously differentiable. Moreover, at least one of the following cross second-order partial derivatives,
is non-zero on a set with non-zero Lebesgue measure, where , and .
Assumption 2 is a natural condition that one expects to hold to identify the invalid IV set. Its intuition comes from the linear separability of the logarithm of the joint density of independent variables. Specifically, for a set of independent random variables with a twice-differentiable joint density, the Hessian matrix of the log-density is diagonal (45). Applying this property to and , if , then their joint density factorizes as . Equivalently, the log-density is additively separable, and hence the corresponding off-diagonal Hessian entry vanishes: . The same argument applies after exchanging and . Assumption 2 requires that, for an invalid IV set, such cross derivatives do not vanish for at least one cross pair, thereby making the violation detectable through the CAT condition.
Remark 2.
Assumption 2, referred to as the Cross Distributional Non-degeneracy Condition, can be viewed as a distributional analogue of the Algebraic Equation Condition in 26. While Guo et al.’s condition is derived from an explicit algebraic expansion under constant-effect models, our condition is formulated through cross second-order derivatives of joint log-densities and does not require an invertible transformation between observed variables and latent noise terms. This makes the proposed condition applicable to the more general CAM-NCE framework.
The following example gives a concrete illustration of how a violation of IV validity can induce a nonzero cross second-order derivative in the joint density. For clarity, we present the calculation in a constant-effect setting, where closed-form expressions are available. The example is intended to illustrate the intuition behind the more general non-constant-effect case.
Example 3 (Illustration of the Cross Distributional Non-degeneracy Condition).
Consider the following data-generating process: , , , , and , where . Here, the candidate IVs and are invalid because they directly affect through the term , thereby violating the exclusion restriction.
For , by the definition of the auxiliary variable, , where . The joint density of can be factorized as , where . Conditional on , since , , and are mutually independent standard normal variables, we have . Therefore, . Then, . Since this quantity is nonzero whenever , the cross second-order partial derivative is nonzero on a set with nonzero Lebesgue measure. Therefore, the candidate IV set satisfies the cross distributional non-degeneracy condition (Assumption 2).
To better understand the Cross Distributional Non-degeneracy Condition (Assumption 2), we next provide a sufficient condition under which this assumption fails.
Proposition 1 (Sufficient Condition for Violation of Assumption 2).
Let , , and be the treatment, outcome, and a candidate IV set in CAM-NCE, respectively. Suppose that each is relevant and exogenous but violates only the exclusion restriction. Assume further that, for , its direct causal effect on is an affine transformation of its direct causal effect on , namely, , where is a nonzero constant, and is a constant. Then Assumption 2 fails for the invalid IV set . This special case is illustrated graphically in Figure 4.
Proof.
See Appendix B.2 for its proof. ∎
As shown in Proposition 1, when Assumption 2 is violated, certain IV sets may fail to satisfy the exclusion restriction, leading to non-identifiability of causal effects. We will now show that under additional Assumption 2, IV sets can be uniquely identified.
Proposition 2 (Sufficient Condition for IV Set Validity).
Proof.
See Appendix B.3 for its proof. ∎
Proposition 2 shows that, under Assumptions 1 and 2, every invalid IV set violates the CAT condition. Combining this result with Theorem 1, we obtain the following necessary and sufficient characterization of IV set validity under CAM-NCE.
Theorem 2 (Necessary and Sufficient Condition for IV Set Validity).
Proof.
See Appendix B.4 for its proof. ∎
The theoretical implications of the CAT condition are summarized in Figure 5.
As shown in Figure 5, the CAT condition is necessary for IV set validity under the completeness condition, and becomes both necessary and sufficient when the Cross Distributional Non-degeneracy Condition is additionally imposed.
4 Practical Testing Algorithm from Data
In this section, we first extend the CAT condition to settings with baseline covariates, where IV validity is assessed after adjusting for observed covariates. We then develop a finite-sample testing algorithm for assessing the validity of candidate IV sets from data.
4.1 CAT Condition with Covariates
In practice, observed covariates such as age, gender, and other background variables may affect the treatment, outcome, and candidate IVs. It is therefore necessary to assess IV validity after adjusting for such covariates.
Under CAM-NCE with covariates, the data-generating process is given by
| (4) | ||||
where denotes the covariates. Below, we show that the additive structure in Equation (4) allows the CAT condition in Definition 4 to be extended to covariate settings through regression adjustment.
Definition 5 (CAT Condition with Covariates).
Let , , , and denote the treatment, outcome, covariates, and candidate IV set, respectively. Furthermore, let , , and denote the residuals obtained by regressing , , and on , respectively (e.g., for each IV, ). We say that satisfies the CAT condition if and only if for every pair of distinct IVs , the following cross-independence relationships hold:
| (5) |
Based on Definition 5 and Theorem 1, we derive the following necessary condition for IV set validity in the presence of covariates .
Corollary 1 (Necessary Condition for IV Set Validity with Covariates).
Let , , , and be the treatment, outcome, covariates, and candidate IV set in CAM-NCE, respectively. Suppose that , , , and are statistically dependent, and that Assumption 1 holds for the residualized variables. If the candidate IV set is a valid IV set w.r.t. given , then satisfies the CAT condition.
Proof.
See Appendix B.5 for its proof. ∎
Corollary 1 states that if violates the CAT condition, then is an invalid IV set. Furthermore, to obtain a necessary and sufficient characterization in the presence of covariates, we impose Assumption 2 on the covariate-adjusted residual variables.
Corollary 2 (Necessary and Sufficient Condition for IV Set Validity with Covariates).
Let , , , and be the treatment, outcome, covariates, and candidate IV set in a CAM-NCE, respectively. Suppose that , , , and are statistically dependent. Further suppose that Assumptions 1 and 2 hold for the covariate-adjusted residual variables. The candidate IV set is a valid IV set relative to given if and only if satisfies the CAT condition.
Proof.
See Appendix B.6 for its proof. ∎
4.2 CAT Algorithm with Finite Samples
In this subsection, we develop a practical CAT algorithm for finite-sample data. For generality, we present the algorithm in the setting with baseline covariates . Let , , and denote the residuals obtained by regressing , , and each candidate IV on , respectively. When no covariates are available, this residualization step is omitted and the original variables are used directly.
Since the CAT results are given in terms of population-level (Theorems 12 and Corollaries 12), implementing it with observational samples requires addressing three practical questions:
- .
How can one efficiently search for a valid IV set from a collection of candidate IVs?
- .
How can one estimate the auxiliary variable for each candidate IV ?
- .
How can the CAT condition be implemented with finite samples?
We next address these questions in turn.
: Searching for candidate IV sets. Since the CAT condition is defined through pairwise cross-independence relations, we focus on candidate subsets with at least two IVs. Searching over all such subsets of requires examining possibilities, which can be computationally expensive when is large. To make the search tractable, we introduce a user-specified parameter , denoting the target size of the IV set to be selected. Given , we restrict attention to candidate subsets of size , resulting in subsets. Specifically, we consider all subsets with . When prior knowledge about is unavailable, we suggest evaluating over a range of plausible values, starting from small values, and using the CAT-based score below to select among the resulting candidate subsets.
: Estimating auxiliary variables. To construct the auxiliary variable for each candidate IV, the key step is to estimate the function satisfying the conditional moment restriction . Under the additive model, if is a valid IV, then is identifiable under the completeness condition (Assumption 1) and coincides with the structural response function of the treatment on the outcome (53). In this case, can be consistently estimated using suitable IV estimators. If is invalid, the resulting estimator may be biased, and this bias will be reflected in the corresponding auxiliary variable. Various estimators have been developed for additive nonparametric IV models, including sieve-based estimators (54), kernel-based estimators (62; 51), deep IV estimators (30; 44), and control-function estimators (52; 28). In our implementation, we use the control-function IV estimator for non-constant effects (28). When prior knowledge suggests a constant-effect model, we instead use the two-stage least squares estimator (71). Given the estimated function , the empirical auxiliary variable is constructed as .
: Implementing the CAT condition with finite samples. For each candidate subset , the CAT condition requires every pair of distinct candidate IVs to satisfy two directed cross-independence relations. Specifically, for each pair , we need to assess
One could use formal independence tests for continuous variables and determine whether the CAT condition holds based on the corresponding -values. However, since our goal is to compare many candidate subsets in finite samples, we use distance correlation as a numerical dependence score. According to 65; 66, distance correlation is zero if and only if the two random vectors are independent. Hence, whenever the population-level CAT condition holds, the corresponding population distance correlation is zero. In finite samples, smaller sample distance correlations therefore provide stronger empirical support for the CAT condition. Let denote the sample distance correlation. For each candidate subset , we define its CAT score as the sum of directed pairwise distance correlations:
For a subset of size , this score aggregates unordered IV pairs and directed distance-correlation terms. Finally, we select the candidate subset with the smallest empirical CAT score:
The selected subset is therefore the candidate IV set that is most consistent with the CAT condition in finite samples.
Based on the above three components, the entire procedure is summarized in Algorithm 1. Overall, the algorithm proceeds in three main stages. First, it adjusts for covariates by residualizing the treatment, outcome, and candidate IVs with respect to (Lines 2–6). Second, it estimates the auxiliary variables for each candidate IV by estimating the corresponding function , as described in (Lines 7–14). Third, it searches over candidate IV subsets of size and evaluates each subset using the CAT-based score defined in (Lines 15–26). The subset with the smallest empirical CAT score is returned as the estimated IV set that is most consistent with the CAT condition.
Below, we establish the correctness of the CAT algorithm in selecting a valid IV set as the sample size tends to infinity.
Theorem 3 (Correctness of the CAT Algorithm).
Assume that the observed data are generated from CAM-NCE, and that the candidate IV set contains at least valid IVs. Suppose that Assumptions 1–2 hold, and that the estimators used in Algorithm 1 for covariate adjustment, auxiliary-variable construction, and distance correlation are consistent. Then, as the sample size tends to infinity, Algorithm 1 outputs a valid IV subset with . In particular, if contains exactly valid IVs, then Algorithm 1 outputs the full valid IV set.
Proof.
See Appendix B.7 for its proof. ∎
Computational complexity. We finally analyze the computational complexity of the CAT algorithm. Let denote the sample size, denote the number of candidate IVs, and denote the number of covariates. The time complexity consists of four main components:
- 1.
Covariate residualization. Regressing , , and the candidate IVs on to obtain residualized variables costs .
- 2.
Auxiliary-variable estimation: For each candidate IV, we estimate using the estimator adopted in this work, i.e., the semiparametric control-function estimator for non-constant effects and two-stage least squares for constant effects. This step costs in total.
- 3.
Pairwise distance-correlation calculation. Computing the directed pairwise distance correlations for all candidate IV pairs costs using the standard distance-correlation estimator.
- 4.
Candidate subset selection. After the pairwise distance-correlation matrix is computed, evaluating all candidate subsets of size costs .
Hence, the overall computational complexity is
5 Experiments
In this section, we evaluate the proposed CAT method on synthetic datasets generated under both CAM-CE and CAM-NCE, corresponding to the constant-effect and non-constant-effect settings, respectively. Our goal is to examine whether CAT can correctly distinguish valid IV sets from invalid ones under different types of IV assumption violations. We first consider the constant-effect setting, where representative existing IV methods are applicable and thus serve as baselines, and then proceed to the more general non-constant-effect setting, which is the primary focus of this paper. It is noteworthy that the simulated data are generated solely according to the causal additive model in Equation (1); no additional constraints are imposed to ensure that the distributional condition in Assumption 2 holds. Our source code is available at https://github.com/guoxichen0/CAT.
Across both settings, we consider three representative invalid IV scenarios involving both valid and invalid IVs. In Case 1, the invalid IVs violate the exclusion restriction condition (); in Case 2, the invalid IVs violate the exogeneity condition (); and in Case 3, the invalid IVs violate both the exclusion restriction () and exogeneity () conditions. Each experiment is repeated 100 times with independently generated data. All noise terms are independently drawn from . For each case, we vary the sample size over .
5.1 Synthetic Data under Constant Effects
We first evaluate CAT under the CAM-CE framework, where the structural response function takes the linear form , yielding a constant causal effect. Following 27 and related works, the true constant causal effect is fixed at . Other constant coefficients in the structural equations are independently sampled from . We report comparison results under nonlinear structural components with constant treatment effects, settings with covariates, and the ALICE model.
Since existing IV selection methods are primarily developed for constant-effect or linear models, the CAM-CE setting enables direct comparison with the following representative baselines:
- 1.
NAIVE: the least-squares regression coefficient of on ;
- 2.
MR-Egger (7): implemented using the code available at https://academic.oup.com/ije/article/44/2/512/754653/;
- 3.
TSHT (27): implemented using the code available at https://cran.r-project.org/web/packages/RobustIV/;
- 4.
CIIV (70): implemented using the code available at https://github.com/xlbristol/CIIV/;
- 5.
sisVIVE (36): implemented using the code available at https://cran.r-project.org/web/packages/sisVIVE/;
- 6.
IV-tetrad (61): implemented using the code available at https://www.homepages.ucl.ac.uk/~ucgtrbd/code/iv_discovery/.
Metrics. We evaluate the methods through their resulting causal effect estimates, summarized using boxplots, where estimates closer to the true effect with smaller variability and dispersion indicate better performance. For methods that first select IVs, including TSHT, CIIV, sisVIVE, IV-tetrad, and CAT, we apply the same IV estimator as in 27 after IV selection to ensure a fair comparison. NAIVE and MR-Egger directly produce causal effect estimates and are therefore evaluated using their original outputs.
Results. The results are presented in Figures 6–8. Figures 6 and 7 report the results under nonlinear structural components with constant treatment effects, without and with covariates, respectively. As expected, the proposed CAT algorithm consistently outperforms the other methods across all three cases and sample sizes, exhibiting relatively small variance and producing estimates closest to the true causal effect. In contrast, the NAIVE method performs poorly in all cases due to unmeasured confounders . We observed that the other comparison methods perform poorly across all cases because they rely on the assumption of linearity, whereas the data generation process is partially nonlinear. Additionally, we found that the MR-Egger algorithm yields inaccurate results. A possible reason for this is that, in addition to the linearity assumption, this method requires the InSIDE assumption, which states that the instruments’ pleiotropic effects on the outcome are uncorrelated with their effects on the treatment . Furthermore, Figure 8 reports the results under the ALICE model. In this linear constant-effect setting, CAT performs well across all three cases and yields estimates close to the true causal effect. Its performance is comparable to TSHT, CIIV, sisVIVE, and IV-tetrad, which is expected because these methods are designed for linear constant-effect settings. In contrast, NAIVE and MR-Egger exhibit larger variability or bias across all cases.
5.2 Synthetic Data under Non-Constant Effects
We next evaluate CAT under the CAM-NCE framework, where is nonlinear, and the causal effect varies with the treatment level. Existing IV-selection methods considered above rely on linear or constant-effect specifications and are not designed for the CAM-NCE setting. Applying them here would therefore evaluate the methods under model misspecification rather than provide a meaningful comparison of IV-set identification. To the best of our knowledge, no existing method provides a directly comparable IV-set identification procedure under the CAM-NCE framework considered here. We thus focus on systematically evaluating CAT across a range of non-constant-effect settings. Specifically, we assess the performance of CAT from three perspectives: (i) varying the causal effect function from to , including , , , , , , and , with valid IVs among five candidate IVs; (ii) varying the number of valid IVs, , among ten candidate IVs; and (iii) varying the number of covariates, , with valid IVs among five candidate IVs.
Metrics. We evaluate IV-set identification performance using the Error Rate (ER). A trial is regarded as successful only when the IV set identified by the method exactly matches the true valid IV set. Accordingly, ER is defined as the proportion of trials for which the identified IV set differs from the true valid IV set, with a lower ER indicating better identification performance.
| Sample sizes | Size=1000 | Size=3000 | Size=5000 | |
|---|---|---|---|---|
| Cases | Functions | ER | ER | ER |
| Case 1 | Log | 0.00 | 0.00 | 0.00 |
| Sin | 0.00 | 0.00 | 0.00 | |
| Cos | 0.01 | 0.00 | 0.00 | |
| Quadratic Poly | 0.00 | 0.00 | 0.00 | |
| Cubic Poly | 0.00 | 0.00 | 0.00 | |
| Log(quad) | 0.00 | 0.00 | 0.00 | |
| Exp(quad) | 0.00 | 0.00 | 0.00 | |
| Case 2 | Log | 0.00 | 0.00 | 0.00 |
| Sin | 0.00 | 0.00 | 0.00 | |
| Cos | 0.00 | 0.00 | 0.00 | |
| Quadratic Poly | 0.00 | 0.00 | 0.00 | |
| Cubic Poly | 0.00 | 0.00 | 0.00 | |
| Log(quad) | 0.00 | 0.00 | 0.00 | |
| Exp(quad) | 0.00 | 0.00 | 0.00 | |
| Case 3 | Log | 0.00 | 0.00 | 0.00 |
| Sin | 0.00 | 0.00 | 0.00 | |
| Cos | 0.00 | 0.00 | 0.00 | |
| Quadratic Poly | 0.00 | 0.00 | 0.00 | |
| Cubic Poly | 0.00 | 0.00 | 0.00 | |
| Log(quad) | 0.04 | 0.00 | 0.00 | |
| Exp(quad) | 0.00 | 0.00 | 0.00 |
- •
Note: “Quadratic Poly/quad” and “Cubic Poly” denote quadratic and cubic polynomial functions, respectively. indicates that lower values are better.
| Sample sizes | Size=1000 | Size=3000 | Size=5000 | |
|---|---|---|---|---|
| Cases | ER | ER | ER | |
| 2 | Case 1 | 0.04 | 0.00 | 0.00 |
| Case 2 | 0.00 | 0.00 | 0.00 | |
| Case 3 | 0.01 | 0.01 | 0.00 | |
| 3 | Case 1 | 0.01 | 0.00 | 0.00 |
| Case 2 | 0.00 | 0.00 | 0.00 | |
| Case 3 | 0.00 | 0.00 | 0.00 | |
| 5 | Case 1 | 0.00 | 0.00 | 0.00 |
| Case 2 | 0.00 | 0.00 | 0.00 | |
| Case 3 | 0.00 | 0.00 | 0.00 |
- •
Note: denotes the number of valid IVs to consider. indicates that lower values are better.
| Sample sizes | Size=1000 | Size=3000 | Size=5000 | |
|---|---|---|---|---|
| Cases | ER | ER | ER | |
| 2 | Case 1 | 0.00 | 0.00 | 0.00 |
| Case 2 | 0.08 | 0.07 | 0.05 | |
| Case 3 | 0.04 | 0.04 | 0.01 | |
| 3 | Case 1 | 0.02 | 0.00 | 0.00 |
| Case 2 | 0.10 | 0.09 | 0.08 | |
| Case 3 | 0.14 | 0.10 | 0.04 | |
| 5 | Case 1 | 0.00 | 0.00 | 0.00 |
| Case 2 | 0.14 | 0.10 | 0.08 | |
| Case 3 | 0.15 | 0.12 | 0.10 |
- •
Note: denotes the number of covariates. indicates that lower values are better.
Results. The results are summarized in Tables III–V. Overall, the proposed method achieves consistently low error rates in all considered settings, suggesting that it can reliably identify the valid IV set under non-constant causal effects. Table III reports the results for different forms of the causal effect function from to . Across logarithmic, trigonometric, polynomial, and exponential functions, the ER is close to zero for all sample sizes, indicating that the proposed method is insensitive to the specific functional form of the non-constant causal effect. Table IV further examines the influence of the number of valid IVs in the candidate IV set. The ER remains low for different choices of and generally decreases as the sample size increases, showing that the method is stable across different proportions of valid IVs. This is in contrast to methods relying on the Majority Rule or Plurality Rule assumptions (see Section 1), which require restrictions on the proportion of valid IVs. Finally, Table V presents the results when covariates are included. Although the ER slightly increases as the number of covariates grows, it decreases with larger sample sizes. This suggests that the proposed method remains effective in the presence of covariates, while benefiting from increased sample size.
6 Applications to Real-world Data
In this section, we apply CAT to three real-world datasets spanning sociology, economics, and behavioral science to examine its practical utility. Since the ground-truth validity of candidate IVs is not directly observable in real-world applications, we use the IV specifications proposed in prior studies as reference benchmarks and compare the conclusions obtained by CAT with those reported in the literature. We report both the CAT-based distance-correlation scores and the corresponding p-values from distance-correlation independence tests. Here, the significance level is set to , where denotes the sample size used in the test.
6.1 Colonial Origins Data (1)
This dataset examines the impact of social systems on economic development. After excluding observations with missing values, it contains five key variables across 63 countries: Mortality , Euro1990 , Latitude , Institutions , and Economic Development . We take the IV specification in 1 as a reference benchmark and evaluate it under the constant-effect setting. The hypothesized model proposed by 1 is illustrated in Figure 9, and the hypothesized data generation mechanism is described as follows:
where and are dependent. Specifically, we assess the candidate IV set for the causal relation , conditioning on the covariate .
Results. Since there are only two candidate IVs, we apply CAT directly to this pair and obtain a CAT-based distance-correlation score of . We then conduct distance-correlation independence tests for the two directed CAT relations, namely between and , and between and , obtaining -values of and , respectively. Therefore, we do not reject the corresponding CAT independence relations, providing no evidence against as a valid IV set for . This result is consistent with the findings of 1.
6.2 Children and Mothers’ Labor Supply Data (5)
This dataset comes from an empirical study on the effect of childbearing on mothers’ labor supply. After applying the filtering criteria, it contains 254,652 observations. We use more than two children as the treatment and weeks worked as the outcome. The candidate IVs include two boys , two girls , AGEQK, AGEQ2ND, KIDCOUNT, YOBM, nonmomil, educm, hsormore, nonmomi, ageqm, and agefstd. The covariates include mother’s age at first birth , father’s age at first birth , whether the first child is a boy , whether the second child is a boy , black mother indicator , Hispanic mother indicator , and other-race mother indicator . We take the IV specification in 5 as a reference benchmark and evaluate it under the constant-effect setting. Due to the large sample size and the quadratic computational cost of distance correlation, we randomly subsample of the data and average the results over 10 repeated tests. The valid IVs hypothesized model proposed by 5 is illustrated in Figure 10, and the valid IVs hypothesized data generation mechanism is described as follows:
where and are dependent, represents the set of all elements in covariates after removing the variable .
Results. Using CAT with , we find that the candidate set achieves the smallest CAT-based distance-correlation score, with . We further conduct distance-correlation independence tests for the two directed CAT relations: between and girls2, and between and boys2, obtaining -values of and , respectively. Thus, we do not reject the corresponding CAT independence relations, and CAT selects as the estimated valid IV set for . This result is consistent with the conclusion of 5.
6.3 Conflict and Time Preference Data (67)
This dataset comes from an empirical study on the effect of violent conflict on individual time preferences. In our analysis, we focus on the causal effect of Violence on Patience. After removing observations with missing values, the dataset contains 266 observations and 15 variables, including the treatment variable Violence , the outcome variable Patience , two candidate IVs, Distance and Altitude , and 11 covariates . The covariates include literate, age, sex, total land holding per capita, land Gini coefficient, distance to market, conflict over land, ethnic homogeneity, socioeconomic homogeneity, population density, and per capita total expenditure. We take the IV specification in 67 as a reference benchmark and evaluate it under the non-constant effect setting. The hypothesized model from 28 is illustrated in Figure 11, and the hypothesized generation mechanism is as follows:
where and are dependent.
Results. Since the dataset contains only two candidate IVs, we apply CAT directly to the pair and obtain a CAT-based distance-correlation score of . We then conduct distance-correlation independence tests for the two directed CAT relations, namely between and , and between and , obtaining -values of in both cases. Thus, we do not reject the corresponding CAT independence relations, providing no evidence against as a valid IV set for . This result is consistent with the findings of 67.
7 Conclusion
In this paper, we studied the problem of testing the validity of IV sets from observational data under causal additive models with non-constant effects (CAM-NCE). Under the completeness condition (Assumption 1), we introduced a testable necessary condition, termed the Cross Auxiliary-based Independence Test (CAT) condition, for assessing IV set validity. Furthermore, under the cross distributional non-degeneracy condition (Assumption 2), we established a necessary and sufficient characterization of valid IV sets within the CAM-NCE framework. We also extended the CAT condition to settings with covariates and developed a practical finite-sample algorithm for selecting candidate IV sets that are most consistent with the CAT condition. Experimental results on both synthetic and real-world datasets demonstrate the effectiveness and practical utility of the proposed method. One promising direction for future work is to extend the proposed framework to more general causal models, such as models with multiple treatment variables.
References
- The colonial origins of comparative development: an empirical investigation. American economic review 91 (5), pp. 1369–1401. Cited by: Figure 1, Figure 1, Figure 9, Figure 9, §6.1, §6.1, §6.1.
- Efficient estimation of models with conditional moment restrictions containing unknown functions. Econometrica 71 (6), pp. 1795–1843. Cited by: §2.2.
- Recursive causal structure learning in the presence of latent variables and selection bias. Advances in Neural Information Processing Systems 34, pp. 10119–10130. Cited by: item 2.
- Examples of l2-complete and boundedly-complete distributions. Journal of econometrics 199 (2), pp. 213–220. Cited by: §2.2.
- Children and their parents’ labor supply: evidence from exogenous variation in family size. National bureau of economic research Cambridge, Mass., USA. Cited by: Figure 10, Figure 10, §6.2, §6.2, §6.2.
- Deep generalized method of moments for instrumental variable analysis. Advances in neural information processing systems 32. Cited by: §B.1, item 3, §2.2.
- Mendelian randomization with invalid instruments: effect estimation and bias detection through egger regression. International journal of epidemiology 44 (2), pp. 512–525. Cited by: 1st item, Table I, item 1, item 2.
- Consistent estimation in mendelian randomization with some invalid instruments using a weighted median estimator. Genetic epidemiology 40 (4), pp. 304–314. Cited by: 1st item, Table I, item 1.
- Evaluating instrument validity using the principle of independent mechanisms. Journal of Machine Learning Research 24 (176), pp. 1–56. Cited by: 2nd item, Table I.
- A review of instrumental variable estimators for mendelian randomization. Statistical methods in medical research 26 (5), pp. 2333–2355. Cited by: §1.
- Triad constraints for learning causal structure of latent variables. Advances in neural information processing systems 32. Cited by: §3.1.
- FWCEC: an enhanced feature weighting method via causal effect for clustering. IEEE Transactions on Knowledge and Data Engineering 37 (2), pp. 685–697. Cited by: §1.
- Mitigating confounding bias in practical recommender systems with partially inaccessible exposure status. IEEE Transactions on Pattern Analysis and Machine Intelligence 46 (2), pp. 957–974. Cited by: §1.
- Identification and model testing in linear structural equation models using auxiliary variables. In International Conference on Machine Learning, pp. 757–766. Cited by: §3.1.
- Data-driven causal effect estimation based on graphical causal modelling: a survey. ACM Computing Surveys 56 (5), pp. 1–37. Cited by: §1.
- Toward unique and unbiased causal effect estimation from data with hidden variables. IEEE Transactions on Neural Networks and Learning Systems 34 (9), pp. 6108–6120. Cited by: §1.
- Discovering ancestral instrumental variables for causal inference from observational data. IEEE Transactions on Neural Networks and Learning Systems, pp. 1–11. Cited by: 1st item, Table I, item 1.
- Instrumental variable estimation of nonseparable models. Journal of Econometrics 139 (1), pp. 4–14. Cited by: §2.2.
- Semi-instrumental variables: a test for instrument admissibility. In Proceedings of the Seventeenth conference on Uncertainty in artificial intelligence, pp. 83–90. Cited by: Table I, §1, §3.1.
- Learning high-dimensional directed acyclic graphs with latent and selection variables. The Annals of Statistics, pp. 294–321. Cited by: item 2.
- Nonparametric instrumental regression. Econometrica 79 (5), pp. 1541–1565. Cited by: §2.2.
- Iterative conditional fitting for gaussian ancestral graph models. In Proceedings of the 20th conference on Uncertainty in artificial intelligence, pp. 130–137. Cited by: §3.1.
- On the completeness condition in nonparametric instrumental problems. Econometric Theory 27 (3), pp. 460–471. Cited by: §2.2, §3.1.
- Instrument validity tests with causal forests. Journal of Business & Economic Statistics 40 (2), pp. 605–614. Cited by: §1.
- Testability of instrumental variables in additive nonlinear, non-constant effects models. Journal of Machine Learning Research. Note: To appear Cited by: Table I, §1, §3.1, §3.1.
- Data-driven selection of instrumental variables for additive nonlinear, constant effects models. In Forty-second International Conference on Machine Learning, Cited by: 1st item, Table I, Remark 2.
- Confidence intervals for causal effects with invalid instruments by using two-stage hard thresholding with voting. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 80 (4), pp. 793–815. Cited by: 1st item, Table I, item 1, item 3, §5.1, §5.1.
- Control function instrumental variable estimation of nonlinear causal effect models. Journal of Machine Learning Research 17 (100), pp. 1–35. Cited by: item 3, §4.2, §6.3.
- Detecting invalid instruments using l1-gmm. Economics Letters 101 (3), pp. 285–287. Cited by: 1st item, Table I, item 1.
- Deep iv: a flexible approach for counterfactual prediction. In International conference on machine learning, pp. 1414–1423. Cited by: §4.2.
- Valid causal inference with (some) invalid instruments. In International Conference on Machine Learning, pp. 4096–4106. Cited by: 1st item, Table I, item 1.
- Robust inference in summary data mendelian randomization via the zero modal pleiotropy assumption. International journal of epidemiology 46 (6), pp. 1985–1998. Cited by: 1st item, Table I, item 1.
- Nonparametric identification using instrumental variables: sufficient conditions for completeness. Econometric Theory 34 (3), pp. 659–693. Cited by: §2.2, §3.1.
- Testing instrument validity for late identification based on inequality moment constraints. Review of Economics and Statistics 97 (2), pp. 398–411. Cited by: §1.
- Information-geometric approach to inferring causal directions. Artificial Intelligence 182, pp. 1–31. Cited by: 2nd item.
- Instrumental variables estimation with some invalid instruments and its application to mendelian randomization. Journal of the American statistical Association 111 (513), pp. 132–144. Cited by: 1st item, Table I, item 1, item 5.
- Generalized instrumental inequalities: testing the instrumental variable independence assumption. Biometrika 107 (3), pp. 661–675. Cited by: §1.
- Causal k-means clustering. Journal of the Royal Statistical Society Series B: Statistical Methodology, pp. qkag068. Cited by: §1.
- Hierarchical and density-based causal clustering. Advances in Neural Information Processing Systems 37, pp. 30363–30393. Cited by: §1.
- A test for instrument validity. Econometrica 83 (5), pp. 2043–2063. Cited by: §1.
- Identification and inference with many invalid instruments. Journal of Business & Economic Statistics 33 (4), pp. 474–484. Cited by: 1st item, Table I, item 1.
- Mechanisms under shifts: interpretable clustering with self-improving heterogeneous causal graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence. Cited by: §1.
- Instrumental variable value iteration for causal offline reinforcement learning. Journal of Machine Learning Research 25 (303), pp. 1–56. Cited by: §1.
- One-stage deep instrumental variable method for causal inference from observational data. In 2019 IEEE International Conference on Data Mining (ICDM), pp. 419–428. Cited by: §4.2.
- Factorizing multivariate function classes. Advances in neural information processing systems 10. Cited by: §3.2, Theorem 5.
- Towards causality-aware inferring: a sequential discriminative approach for medical diagnosis. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (11), pp. 13363–13375. Cited by: §1.
- On the instrumental variable estimation with many weak and invalid instruments. Journal of the Royal Statistical Society Series B: Statistical Methodology, pp. qkae025. Cited by: 1st item, Table I, item 1.
- Partial identification of probability distributions. Springer Science & Business Media. Cited by: §1.
- A natural introduction to probability theory. Cited by: Appendix A, Theorem 4.
- Testing local average treatment effect assumptions. Review of Economics and Statistics 99 (2), pp. 305–313. Cited by: §1.
- Dual instrumental variable regression. Advances in Neural Information Processing Systems 33, pp. 2710–2721. Cited by: §4.2.
- Nonparametric estimation of triangular simultaneous equations models. Econometrica 67 (3), pp. 565–603. Cited by: §4.2.
- Instrumental variable estimation of nonparametric models. Econometrica 71 (5), pp. 1565–1578. Cited by: §B.1, item 3, §2.2, §2.2, §3.1, §4.2.
- Nonparametric instrumental variables estimation. American Economic Review 103 (3), pp. 550–556. Cited by: §2.2, §2.2, §3.1, §4.2.
- Nonparametric bounds for the causal effect in a binary instrumental-variable model. The Stata Journal 11 (3), pp. 345–367. Cited by: §1.
- On the testability of causal models with latent and instrumental variables. In Proceedings of the Eleventh conference on Uncertainty in artificial intelligence, pp. 435–443. Cited by: §1.
- Causality: models, reasoning, and inference. 2nd edition, Cambridge University Press, New York. Cited by: Definition 1.
- Causal inference on discrete data using additive noise models. IEEE Transactions on Pattern Analysis and Machine Intelligence 33 (12), pp. 2436–2450. Cited by: §1.
- Mendelian randomization. Nature Reviews Methods Primers 2 (1), pp. 6. Cited by: 1st item, Table I, item 1.
- Temporal-spatial causal interpretations for vision-based reinforcement learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 44 (12), pp. 10222–10235. Cited by: §1.
- Learning instrumental variables with structural and non-gaussianity assumptions. Journal of Machine Learning Research 18 (120), pp. 1–49. Cited by: 1st item, Table I, item 1, item 6.
- Kernel instrumental variable regression. Advances in Neural Information Processing Systems 32. Cited by: §B.1, item 3, §2.2, §4.2.
- Vitamin d status, filaggrin genotype, and cardiovascular risk factors: a mendelian randomization approach. PloS one 8 (2), pp. e57647. Cited by: §1.
- Causal inference in the presence of latent variables and selection bias. In Proceedings of the Eleventh conference on Uncertainty in artificial intelligence, pp. 499–506. Cited by: item 2.
- Measuring and testing dependence by correlation of distances. The Annals of Statistics, pp. 2769–2794. Cited by: §4.2.
- Brownian distance covariance. Cited by: §4.2.
- Violent conflict and behavior: a field experiment in burundi. American Economic Review 102 (2), pp. 941–964. Cited by: Figure 11, Figure 11, §6.3, §6.3, §6.3.
- On falsification of the binary instrumental variable model. Biometrika 104 (1), pp. 229–236. Cited by: §1.
- On the use of the lasso for instrumental variables estimation with some invalid instruments. Journal of the American Statistical Association 114 (527), pp. 1339–1350. Cited by: 1st item, Table I, item 1.
- The confidence interval method for selecting valid instrumental variables. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 83 (4), pp. 752–776. Cited by: 1st item, Table I, item 1, item 4.
- Econometric analysis of cross section and panel data. MIT press. Cited by: §4.2.
- Instrumental variables in causal inference and machine learning: a survey. ACM Computing Surveys 57 (11), pp. 1–36. Cited by: §1.
- Generalized independent noise condition for estimating latent variable causal graphs. Advances in neural information processing systems 33, pp. 14891–14902. Cited by: 2nd item.
- Testability of instrumental variables in linear non-gaussian acyclic causal models. Entropy 24 (4), pp. 512. Cited by: 2nd item, Table I.
- Generalized independent noise condition for estimating causal structure with latent variables. Journal of Machine Learning Research 25 (191), pp. 1–61. Cited by: §3.1.
- Personalized latent structure learning for recommendation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45 (8), pp. 10285–10299. Cited by: §1.
- CIPL: counterfactual interactive policy learning to eliminate popularity bias for online recommendation. IEEE Transactions on Neural Networks and Learning Systems 35 (12), pp. 17123–17136. Cited by: §1.
Appendix Contents
Appendix A Theoretical Foundations
Before presenting the proofs, we introduce several technical facts that will be used repeatedly. We begin with a standard property of independent random variables from 49, which is used in the proof of Theorem 1 and its covariate-adjusted extension.
Theorem 4 (Theorem 2.2.5 in 49).
Let be independent random variables, and for , be a function . Then the random variables are also independent.
We also use the following direct extension, which states that measurable functions of disjoint subsets of mutually independent random variables remain independent.
Lemma 1.
Let be independent random variables. Suppose that is a measurable function of , is a measurable function of , then and are independent random variables.
Proof.
Let
Since are mutually independent, the random vectors and are independent. Indeed, for any Borel sets and , the independence of implies
Now, for any Borel sets , we have
Since and are measurable, and are Borel sets. Using the independence of and , we obtain
Therefore, and are independent. ∎
Next, we recall a separability result for twice continuously differentiable functions and apply it to log-densities to connect independence with cross second-order partial derivatives. This result will be used in the proofs of Proposition 2 and Theorem 2.
Theorem 5 (45).
The Hessian of function is block diagonal everywhere, for all points and all , , if and only if f is separable into a sum for some functions and .
Theorem 5 implies that if the log-density of two groups of variables is additively separable, then the corresponding cross second-order partial derivatives vanish.
Appendix B Proofs
B.1 Proof of Theorem 1
Proof.
To prove Theorem 1, we need to show that if the candidate IV set is a valid IV set relative to under CAM-NCE, then for any pair , will satisfy the CAT condition.
Under CAM-NCE, the data-generating process is
| (6) | ||||
where . Since is a valid IV set, the valid IV set is not a subset of , and each variable in is generated solely from its corresponding noise term, i.e., , for any .
Consider any pair of distinct IVs . According to the definition of auxiliary variable w.r.t. relative to , , where satisfies and . Because is a valid IV and Assumption 1 holds, the standard identification argument for nonparametric IV models implies that the solution is unique and coincides with the structural response function (53; 6; 62); that is, . Similarly, since is also a valid IV, we have .
Therefore, we can further express the auxiliary variable as:
| (7) |
By Theorem 4 and its extension Lemma 1, if random variables are mutually independent, then any measurable functions applied to disjoint subsets of them yield independent random variables (see Theorem 4, and Lemma 1 for further details). Based on this result, we next show that the auxiliary variable and are statistically independent. Specifically, since all noise terms of variables are mutually independent and , we can obtain that is independent of . Furthermore, combining Equations (6) and (7), we conclude that is independent of , i.e., .
Likewise, for pairwise (), we can derive that is independent of , i.e., . To sum up, always satisfies the CAT condition.
Similarly, the same argument applies to any variable pair in , implying that every such pair satisfies the CAT condition. Therefore, satisfies the CAT condition. ∎
B.2 Proof of Proposition 1
Proof.
We prove the proposition for the candidate IV set . Since the candidate IVs in violate only the exclusion restriction, they are relevant and exogenous, but directly affect the outcome. The data-generating process can be written as
| (8) | ||||
where , and . Substituting this relation into the structural equation of gives
where .
Now construct an alternative structural representation with for , i.e., , , and
Then , and hence
Furthermore, we have
In this alternative representation, the variables affect only through and have no direct effects on . Therefore, forms a valid IV set with respect to .
By Theorem 1, the CAT condition holds for the alternative representation. For any pair with , the corresponding auxiliary variable is
Since each is independent of , , , and other variables , it follows from Lemma 1 that
Consequently, , which implies
Therefore,
Because the original and alternative representations induce the same observational distribution, we also have
The same argument applies after exchanging and . Consequently, all the cross second-order partial derivatives are zero. As a result, violates Assumption 2. ∎
B.3 Proof of Proposition 2
Proof.
Suppose that the candidate IV set is invalid. By Assumption 2, there exists a pair of distinct candidate IVs such that the joint densities and are twice continuously differentiable, and at least one of the following cross second-order partial derivatives is nonzero on a set with nonzero Lebesgue measure:
We prove the result by contradiction. Assume that satisfies the CAT condition. Then, for every pair of distinct IVs , we have
In particular, for the pair specified by Assumption 2, the first independence relation implies
Taking logarithms yields the additive decomposition
Therefore, the corresponding cross second-order partial derivative must vanish:
Similarly, the second independence relation implies
Thus, both cross second-order partial derivatives vanish, which contradicts Assumption 2. Hence, cannot satisfy the CAT condition. Therefore, if is invalid, then violates the CAT condition. ∎
B.4 Proof of Theorem 2
Proof.
We prove the necessary and sufficient characterization of valid IV sets under CAM-NCE by establishing the following two implications.
(i): Assume the candidate IV set is a valid IV set relative to . By Theorem 1, under Assumption 1, it directly follows that if the candidate IV set is a valid IV set relative to , then always satisfies the CAT condition.
(ii): Assume the candidate IV set is an invalid IV set relative to . By Proposition 2, under Assumptions 1 and 2, if the candidate IV set is invalid, then consequently violates the CAT condition.
Combining (i) and (ii), we conclude that is a valid IV set w.r.t. if and only if satisfies the CAT condition. ∎
B.5 Proof of Corollary 1
Proof.
Suppose that is a valid IV set w.r.t. given . Under CAM-NCE with covariates, the data-generating mechanism can be written as
| (9) | ||||
where denotes the set of parent variables of , and denotes the causal relationship among variables in . Since is valid given , no variable in directly affects ; equivalently,
Let , , and denote the residuals obtained by regressing , , and on , respectively. Under the additive covariate structure in Equation (9) and the regression adjustment in Definition 5, the residualized variables remove the effects of and follow the same structural form as the covariate-free CAM-NCE model. Moreover, because is a valid IV set w.r.t. given , the residualized candidate IV set satisfies the corresponding relevance, exclusion restriction, and exogeneity conditions in the residualized system.
Therefore, applying Theorem 1 to the residualized variables yields that every pair of distinct residualized IVs in satisfies the CAT condition. Equivalently, satisfies the CAT condition. ∎
B.6 Proof of Corollary 2
Proof.
Let , , and denote the residuals obtained by regressing , , and on , respectively. As shown in the proof of Corollary 1, under the additive covariate structure and the regression adjustment in Definition 5, the residualized variables , , and follow the same structural form as the covariate-free CAM-NCE model.
Furthermore, Assumption 1 and Assumption 2 are assumed to hold for these covariate-adjusted residual variables. Therefore, the residualized system satisfies the conditions required by Theorem 2. Applying Theorem 2 to the residualized variables yields that the residualized candidate IV set is valid for the causal relation if and only if the corresponding CAT condition holds.
By Definition 5, this is equivalent to saying that is a valid IV set w.r.t. given if and only if satisfies the CAT condition. ∎
B.7 Proof of Theorem 3
Proof.
Let denote the set of all valid IVs among the candidate IVs, and let . By assumption, . We assume , so that the pairwise CAT criterion is informative.
Algorithm 1 first removes the effect of covariates , if present, and then constructs, for each candidate IV , the auxiliary variable . By the assumed consistency of the estimators used for covariate adjustment, auxiliary-variable construction, and distance correlation, for every unordered pair ,
converges in probability to
For any candidate subset with , define the population objective
and let be the corresponding empirical objective computed by Algorithm 1. Since the number of candidate subsets of size is finite and each is consistent, we have the uniform convergence
Now consider any subset with . Every element of is a valid IV. Hence, by Theorem 2, the CAT condition holds for every pair , namely
Since distance correlation is zero if and only if independence holds, it follows that for every pair in . Therefore, .
Conversely, let be any subset of size that contains at least one invalid IV. Then is not a valid IV set. By the necessity and sufficiency result in Theorem 2, under Assumptions 1–2, the CAT condition cannot hold for all pairs in . Hence there exists at least one pair such that
For this pair, at least one of the two distance correlations is strictly positive. Thus , and consequently .
Therefore, every valid subset of size attains the population minimum value , whereas every subset of size containing at least one invalid IV has a strictly positive population objective value.
Because the number of candidate subsets is finite, the collection of invalid subsets of size is also finite. If this collection is nonempty, define
From the preceding argument, . On the event
every valid subset satisfies , whereas every invalid subset satisfies . Hence any empirical minimizer selected in Line 26 of Algorithm 1 must be a subset of . Since the above event has probability tending to one (i.e., ), the algorithm outputs, with probability tending to one, a valid IV subset with .
Finally, if the candidate set contains exactly valid IVs, then , and the only subset of size consisting entirely of valid IVs is itself. Therefore,
in the sense that the probability that Algorithm 1 outputs the full valid IV set tends to one.
This proves the stated correctness of Algorithm 1. ∎