Causal Inference in Machine Learning
1. Key Concepts: Causality vs. Correlation
Key Concepts: Causality vs. Correlation
Causality and correlation are foundational concepts in statistical learning, yet they are often conflated. While correlation measures the degree of association between two variables, causality implies a directional relationship where one variable directly influences another. Distinguishing between the two is critical in machine learning, as models trained on correlational patterns may fail under intervention or policy changes.
Mathematical Definitions
Correlation between random variables X and Y is quantified by the Pearson correlation coefficient:
where Cov(X, Y) is the covariance and σ denotes standard deviation. A value of |ρ| ≈ 1 indicates strong linear dependence, while ρ ≈ 0 suggests independence.
Causality, however, is defined through interventions. If X causes Y, then changing X via an external intervention (denoted do(X = x)) should result in a change in Y. This is formalized using the do-calculus framework:
The left-hand side represents the interventional distribution, while the right-hand side is the observational conditional probability.
Examples and Counterexamples
A classic example is the relationship between ice cream sales (X) and drowning incidents (Y). These variables are positively correlated, but neither causes the other. Instead, a latent variable—temperature (Z)—drives both. This is a case of spurious correlation, where:
In contrast, consider a drug (X) and recovery rate (Y). If administering the drug (do(X = 1)) increases recovery rates compared to a control group (do(X = 0)), this is evidence of causation.
Challenges in Machine Learning
Most supervised learning algorithms, including deep neural networks, optimize for predictive accuracy using observational data. This can lead to models that exploit spurious correlations. For example, a model trained to diagnose pneumonia from X-rays might learn to rely on hospital-specific metadata (e.g., scanner type) rather than pathological features. Such models fail when deployed in new hospitals, where the metadata-distribution shifts.
Causal inference methods address this by explicitly modeling the data-generating process. Structural causal models (SCMs) represent variables as functions of their direct causes and exogenous noise:
where U, V are unobserved variables. SCMs enable counterfactual reasoning (e.g., "What would Y be if X had taken a different value?") and robust predictions under distributional shifts.
Testing for Causality
Several experimental and statistical techniques exist to infer causality:
- Randomized Controlled Trials (RCTs): The gold standard, where X is randomly assigned to isolate its effect on Y.
- Instrumental Variables (IV): Uses an external variable Z that affects Y only through X to estimate causal effects.
- Granger Causality: In time-series data, X Granger-causes Y if past values of X improve predictions of Y.
In observational settings, causal discovery algorithms (e.g., PC algorithm, LiNGAM) infer graph structures from conditional independence tests.
Practical Implications
Understanding causality is essential for applications like personalized medicine, where treatment effects may vary across subpopulations. For instance, a drug might be effective on average but harmful for patients with a specific genetic marker. Causal models can identify such heterogeneous treatment effects by estimating:
where Z represents patient covariates. This goes beyond correlation-based predictions by accounting for confounding and effect modification.

Potential Outcomes Framework
The Potential Outcomes Framework, introduced by Neyman (1923) and later formalized by Rubin (1974), provides a rigorous mathematical foundation for causal inference. It defines causality in terms of potential outcomes under different treatment conditions. For a binary treatment T ∈ {0,1}, each unit i has two potential outcomes:
The fundamental problem of causal inference states that we can only observe one potential outcome per unit, as expressed by the observed outcome:
Key Assumptions
The framework relies on three core assumptions:
- Stable Unit Treatment Value Assumption (SUTVA): No interference between units and no hidden treatment variations.
- Ignorability: Treatment assignment is independent of potential outcomes given covariates X:
$$ (Y(1), Y(0)) \perp T \mid X $$
- Positivity: Every unit has a non-zero probability of receiving either treatment:
$$ 0 < P(T=1 \mid X) < 1 $$
Average Treatment Effects
The primary causal quantity is the Average Treatment Effect (ATE), defined as:
Under the above assumptions, ATE can be estimated via:
Identification Strategies
When ignorability holds only conditionally, common identification approaches include:
- Regression Adjustment: Model outcomes directly using Y ∼ T + X
- Propensity Score Methods: Balance covariates via e(X) = P(T=1|X)
- Doubly Robust Estimation: Combines outcome and propensity models for robustness
Extensions to Continuous Treatments
For continuous treatments T ∈ ℝ, the framework generalizes to dose-response functions:
where identification requires continuous analogs of positivity and ignorability.
Counterfactual Distributions
The framework naturally extends to distributional effects through quantile treatment effects:
where F_Y(t) is the CDF of potential outcome Y(t).
Directed Acyclic Graphs (DAGs) and Structural Causal Models
Graphical Representation of Causal Structures
Directed Acyclic Graphs (DAGs) provide a formal framework for encoding causal assumptions through nodes and directed edges. Nodes represent random variables, while edges denote direct causal relationships. The acyclic property ensures no variable can be its own ancestor, preventing causal loops. For a set of variables X1, X2, ..., Xn, a DAG G implies a factorization of the joint probability distribution:
where pai denotes the parents of Xi in G. This factorization embodies the Markov condition, stating each variable is independent of its non-descendants given its parents.
d-Separation and Causal Identification
The d-separation criterion determines conditional independence relationships implied by a DAG. A path between nodes X and Y is d-separated by a set Z if:
- The path contains a chain X → W → Y or fork X ← W → Y with W ∈ Z, or
- The path contains a collider X → W ← Y where neither W nor its descendants are in Z.
This criterion enables testable implications of causal models through observed data. For example, in the DAG X → Y ← Z → W, X and Z are d-separated given Y, implying X ⫫ Z | Y.
Structural Causal Models (SCMs)
An SCM extends DAGs with explicit functional relationships. Each variable Xi is determined by:
where fi is a deterministic function and Ui represents exogenous noise. SCMs enable counterfactual reasoning through interventional distributions. The do-operator, P(Y | do(X=x)), computes the effect of setting X to x while preserving other causal mechanisms.
Example: Non-Parametric Structural Equations
Consider a DAG with three variables: Z → X → Y and Z → Y. The corresponding SCM is:
where U_Z, U_X, U_Y are independent noise terms. The causal effect of X on Y is identifiable via backdoor adjustment:
Applications in Machine Learning
DAGs and SCMs are instrumental in:
- Bias mitigation: Identifying and adjusting for confounders in observational data.
- Transfer learning: Encoding invariant mechanisms across domains.
- Explainability: Tracing decision pathways in neural networks through causal abstraction.
For instance, in healthcare, a DAG encoding Treatment → Recovery ← Severity prevents biased effect estimates by conditioning on Severity.

2. Propensity Score Matching and Inverse Probability Weighting
Propensity Score Matching and Inverse Probability Weighting
Propensity Score Matching (PSM)
Propensity score matching is a statistical technique used to estimate causal effects by balancing observed covariates between treated and control groups. The propensity score, e(X), is defined as the probability of receiving treatment given observed covariates X:
This score is typically estimated using logistic regression, though machine learning methods like random forests or gradient boosting can improve accuracy in high-dimensional settings. Once estimated, individuals in the treatment and control groups are matched based on similar propensity scores, reducing selection bias.
The key assumption underlying PSM is conditional independence (unconfoundedness):
where Y(1) and Y(0) represent potential outcomes under treatment and control, respectively. When this holds, matching on e(X) suffices to balance covariates.
Matching Algorithms
Common matching approaches include:
- Nearest-neighbor matching: Each treated unit is paired with the closest control unit in propensity score space.
- Caliper matching: Imposes a maximum allowable distance between matches to prevent poor matches.
- Stratification: Divides the propensity score distribution into strata and compares outcomes within each.
Inverse Probability Weighting (IPW)
Inverse probability weighting creates a pseudo-population where treatment assignment is independent of covariates by weighting each observation by the inverse of its probability of receiving the observed treatment:
This weighting scheme effectively creates a balanced dataset where confounders no longer predict treatment assignment. The average treatment effect (ATE) can then be estimated as:
For stable estimation, overlap (positivity) must hold: 0 < e(X) < 1 for all X. Violations indicate regions where treatment effects cannot be estimated.
Practical Considerations
Both methods require careful implementation:
- Model misspecification: Poor propensity score estimation biases results. Flexible machine learning models can help but require cross-fitting to avoid overfitting.
- Unobserved confounding: Neither method addresses bias from unmeasured variables.
- Variance-weighting tradeoff: IPW can be inefficient with extreme weights, while PSM discards unmatched units.
Doubly robust estimators combine outcome modeling with IPW or PSM, providing consistent estimates if either the propensity score or outcome model is correctly specified.
Case Study: Medical Treatment Evaluation
In a study comparing surgical (T=1) vs. drug therapy (T=0) for heart disease, researchers used PSM to balance age, comorbidities, and disease severity. After matching, the estimated treatment effect showed a 15% reduction in mortality risk (95% CI: 8-22%), whereas the naive comparison overestimated the benefit at 25% due to healthier patients selecting surgery.

2.2 Instrumental Variables and Regression Discontinuity
Instrumental Variables (IV) in Causal Inference
Instrumental variables are a powerful tool for estimating causal effects when unobserved confounding is present. An instrumental variable Z must satisfy two key conditions:
- Relevance: Z must be correlated with the endogenous treatment variable X.
- Exclusion: Z affects the outcome Y only through X (no direct effect).
The causal effect can be estimated using two-stage least squares (2SLS):
Where the numerator represents the reduced form and the denominator the first stage. This estimator is consistent when the instrument is valid, though finite-sample bias can occur with weak instruments.
Regression Discontinuity (RD) Designs
RD designs exploit sharp or fuzzy discontinuities in treatment assignment based on a running variable R. The causal effect is estimated by comparing observations just above and below the cutoff c:
For fuzzy RD, where the probability of treatment jumps at c but doesn't go from 0 to 1, instrumental variables methods are applied using the discontinuity as an instrument.
Practical Considerations and Assumptions
Both methods rely on strong assumptions:
- IV requires the exclusion restriction and instrument relevance
- RD assumes continuity of potential outcomes at the cutoff
Violations of these assumptions can lead to biased estimates. Sensitivity analyses and placebo tests are crucial for validating results.
Applications in Machine Learning
Recent work combines these methods with machine learning:
- Using random forests or neural networks to estimate the first stage in IV
- Applying local regression methods in RD designs
- Developing doubly robust estimators that combine parametric and nonparametric approaches
where μ̂ are outcome models and ê is the propensity score, estimated using machine learning methods.

Double Machine Learning and Causal Forests
Double Machine Learning (DML) extends traditional causal inference methods by leveraging flexible machine learning models to estimate nuisance parameters while maintaining asymptotic normality and root-n consistency. The core idea involves orthogonalizing the treatment variable against confounders using auxiliary ML models, isolating the causal effect of interest.
Orthogonalization in Double Machine Learning
The DML framework decomposes the causal estimation problem into two stages:
- First-stage nuisance estimation: Fit ML models to predict treatment T and outcome Y from confounders X
- Second-stage causal estimation: Compute residuals from first-stage models and estimate treatment effect via low-dimensional regression
where g(X) and f(X) are estimated using arbitrary ML models, and θ represents the causal parameter of interest. The key innovation lies in the Neyman-orthogonal score function that makes the estimator robust to first-stage estimation errors:
Causal Forests: Nonparametric Heterogeneous Effects
Causal Forests extend Random Forests to estimate conditional average treatment effects (CATE) by:
- Growing trees on subsamples where treatment assignment can be considered random
- Using doubly robust scores as outcomes in terminal nodes
- Employing honest splitting to prevent overfitting
The estimator takes the form:
where L(x) denotes the leaf containing x and Γ is a doubly robust score combining predictions from separate treatment and control models.
Asymptotic Properties
Under regularity conditions, Causal Forests achieve pointwise normality:
where the variance V(x) can be estimated via infinitesimal jackknife. This enables valid confidence intervals even with high-dimensional confounders.
Practical Implementation Considerations
Key implementation details affecting performance:
- Nuisance model selection: Cross-fitting with k-fold sample splitting prevents overfitting
- Regularization tuning: First-stage models must be undersmoothed to avoid excessive bias
- Honest estimation: Separate samples for tree-growing and effect estimation
- Bandwidth selection: Adaptive nearest-neighbor weights for local averaging
Modern implementations leverage gradient-boosted trees or neural networks for nuisance parameter estimation while maintaining theoretical guarantees through careful orthogonalization.

3. Confounding and Selection Bias
Confounding and Selection Bias
Confounding Variables in Causal Inference
Confounding occurs when an extraneous variable Z influences both the treatment X and the outcome Y, creating a spurious association. Formally, a confounder Z must satisfy:
For example, in studying the effect of medication (X) on recovery time (Y), age (Z) may confound the relationship if older patients are both more likely to receive the medication and recover slower. The backdoor criterion provides a graphical test for identifying sufficient adjustment sets to block confounding paths.
Selection Bias and Its Mechanisms
Selection bias arises when the sample selection process correlates with the treatment or outcome. Common scenarios include:
- Non-random missing data: Outcomes are missing systematically (e.g., survey dropouts correlated with treatment).
- Truncation: Observations are excluded based on a threshold (e.g., studying income effects while omitting low-income households).
- Berkson's paradox: Artificial correlations induced by conditioning on a collider (e.g., hospital studies where admission acts as a collider between treatment and outcome).
Mathematical Formulation of Bias
The bias due to confounding can be quantified as the difference between the observed association and the true causal effect. For a binary treatment, the bias is:
where ATE is the average treatment effect. Under selection bias, the estimand becomes:
where S=1 indicates selection into the sample. This estimand may diverge from the population ATE due to the conditioning on S.
Adjustment Methods
To address confounding and selection bias, several advanced techniques are employed:
- Propensity score matching: Balances confounders by matching treated and control units with similar probabilities of treatment.
- Inverse probability weighting (IPW): Corrects for selection by weighting observations by the inverse of their probability of being selected.
- Front-door adjustment: Uses mediator variables to estimate causal effects when unmeasured confounders exist.
IPW Derivation
For selection bias correction, IPW weights are derived as:
where P(S_i=1|X_i, Z_i) is the propensity for being selected into the sample. The weighted estimator then becomes:
Case Study: Observational Drug Trials
In a study of a new drug's effect on hospitalization rates, confounding by indication occurs when doctors prescribe the drug preferentially to sicker patients. Simultaneously, selection bias arises if only insured patients are included in the dataset. A doubly robust estimator combining IPW and outcome regression can mitigate both issues.

3.2 Generalizability and External Validity
Generalizability refers to the extent to which causal inferences drawn from a study population can be applied to a broader target population. External validity, a closely related concept, assesses whether the estimated causal effects remain consistent across different settings, populations, or time periods. Unlike internal validity, which focuses on unbiased estimation within the study sample, external validity concerns the broader applicability of findings.
Formalizing Generalizability
Let Yi(1) and Yi(0) denote potential outcomes for unit i under treatment and control, respectively. The average treatment effect (ATE) in the study sample S is:
For a target population T, the ATE is:
Generalizability requires that ATES ≈ ATET. This holds if either:
- The study sample S is a random subset of T (strong ignorability of sampling), or
- Treatment effects are homogeneous across subpopulations.
Threats to External Validity
Key threats include:
- Population Differences: The study sample may not represent the target population in terms of demographics, behavior, or unobserved confounders.
- Contextual Differences: Variations in environment, time, or implementation protocols may alter treatment effects.
- Temporal Shifts: Dynamic systems may evolve, making past causal estimates obsolete.
Assessing Generalizability
To evaluate external validity, researchers employ:
- Transportability Analysis: Uses weighting or outcome modeling to adjust estimates for population differences.
- Heterogeneous Treatment Effects (HTE): Estimates subgroup-specific effects to identify where generalizability fails.
- Sensitivity Analysis: Tests robustness of conclusions under plausible variations in population characteristics.
Transportability via Inverse Probability Weighting
If the sampling mechanism is known, inverse probability of sampling weights (IPSW) can adjust for selection bias:
The reweighted ATE estimate becomes:
Case Study: Generalizing Clinical Trials
In drug efficacy trials, participants often differ from real-world patients in age, comorbidities, or adherence. A 2021 study by Degtiar et al. demonstrated that IPSW-adjusted estimates from clinical trials reduced bias in predicting population-level effects by 37% compared to unadjusted estimates.
Practical Considerations
When designing studies for generalizability:
- Ensure sample diversity along known dimensions of heterogeneity.
- Pre-register analysis plans to avoid cherry-picking generalizable subgroups.
- Use meta-analytic techniques when pooling results from multiple contexts.

3.3 Scalability in High-Dimensional Settings
Causal inference in high-dimensional settings presents unique computational and statistical challenges. Traditional methods, such as propensity score matching or inverse probability weighting, suffer from the curse of dimensionality when the number of covariates p grows large relative to the sample size n. To maintain scalability, modern approaches leverage sparsity assumptions, regularization, and efficient optimization techniques.
Dimensionality Reduction via Sparsity
High-dimensional causal inference often assumes that only a small subset of covariates are true confounders. This sparsity enables the use of Lasso (Least Absolute Shrinkage and Selection Operator) or its variants for variable selection. The Lasso estimator for the propensity score model is given by:
where λ controls the strength of regularization. The ℓ₁-penalty promotes sparsity, effectively shrinking irrelevant coefficients to zero. Theoretical guarantees for Lasso in causal inference require the restricted eigenvalue condition and irrepresentable condition to ensure consistent selection of confounders.
Doubly Robust Estimators
To improve robustness against model misspecification, doubly robust estimators combine outcome regression with propensity score weighting. The augmented inverse probability weighting (AIPW) estimator is defined as:
where ŷ1(Xi) and ŷ0(Xi) are predictions from the outcome models. This estimator remains consistent if either the propensity score model or the outcome model is correctly specified.
Efficient Optimization via Gradient Boosting
Gradient boosting machines (GBMs) provide a scalable alternative for estimating propensity scores in high dimensions. By iteratively fitting weak learners (e.g., decision trees) to the residual errors, GBMs adaptively capture nonlinear relationships without explicit feature engineering. The algorithm minimizes:
via gradient descent, where e(Xi; θ) is the propensity score model. XGBoost and LightGBM implementations further optimize computational efficiency through parallelization and histogram-based splitting.
Distributed Causal Inference
For ultra-high-dimensional problems (e.g., p > 106), distributed computing frameworks like Spark or Dask enable parallelized estimation. The key idea is to partition covariates across worker nodes, perform local inference, and aggregate results via consensus optimization. The objective function decomposes as:
where fk is the local loss for the k-th partition and R(β) is a regularization term enforcing global sparsity.
Case Study: Genome-Wide Association Studies (GWAS)
In GWAS, where the number of genetic markers p can exceed millions, scalable causal inference is critical. Recent work employs debiased Lasso to estimate the causal effect of individual SNPs while controlling for confounding by population structure. The debiasing step corrects for regularization-induced shrinkage, enabling valid confidence intervals:
where Θ is an approximate inverse of the Hessian matrix. This approach maintains √n-consistency even when p ≫ n.
4. Personalized Medicine and Healthcare
4.1 Personalized Medicine and Healthcare
Causal inference plays a transformative role in personalized medicine, where treatment effects must be estimated at the individual level rather than averaged across populations. Traditional randomized controlled trials (RCTs) often fail to account for heterogeneous treatment effects (HTE), leading to suboptimal patient outcomes. Machine learning methods, particularly those leveraging causal frameworks, enable the estimation of individualized treatment rules (ITRs) by modeling counterfactual outcomes under different interventions.
Heterogeneous Treatment Effects and Counterfactual Prediction
The fundamental challenge in personalized medicine is estimating the conditional average treatment effect (CATE):
where Y(1) and Y(0) represent potential outcomes under treatment and control, respectively, and X denotes patient covariates. Doubly robust methods, such as the X-learner and causal forest, combine propensity score weighting with outcome regression to minimize bias in CATE estimation:
where ŵ(x) is the estimated propensity score and μ̂t(x) are the imputed potential outcomes.
Dynamic Treatment Regimes and Reinforcement Learning
Sequential decision-making in chronic disease management requires dynamic treatment regimes (DTRs), formalized as:
where γ is a discount factor and Rt represents intermediate health outcomes. Causal reinforcement learning methods like Q-learning with double robustness and inverse probability weighted policy optimization address confounding in observational data by incorporating propensity scores into the Bellman equation:
Case Study: Precision Oncology
In cancer immunotherapy, causal survival analysis models handle right-censored outcomes while estimating HTE. The counterfactual survival forest extends random forests to estimate treatment-specific survival curves S(t|X,T) by:
where λ̂(u|x,T) is a non-parametric hazard estimator weighted by inverse propensity scores. Clinical trials have demonstrated 22% improvement in progression-free survival when using causal ML models to assign checkpoint inhibitors.
Confounding Adjustment in Observational Health Data
Electronic health records introduce time-varying confounders affected by prior treatment (e.g., lab values). The g-computation algorithm and longitudinal targeted maximum likelihood estimation (TMLE) solve this by modeling the entire treatment-outcome trajectory:
where ĀK denotes treatment history and L̄K represents time-dependent covariates. TMLE achieves semiparametric efficiency bounds through clever covariate construction and one-step updates.
Validating Causal Models in Healthcare
Transportability and external validity are assessed using negative control outcomes and sensitivity analyses for unmeasured confounding. The E-value quantifies the minimum strength of unmeasured confounders needed to explain away observed effects:
where RRUD is the risk ratio between unmeasured confounder and outcome. In practice, E-values > 2.0 suggest robust causal conclusions for clinical decision support systems.
4.2 Policy Evaluation and Economics
Causal inference plays a pivotal role in policy evaluation, where the goal is to estimate the effect of interventions—such as economic policies, healthcare programs, or regulatory changes—on outcomes of interest. Unlike predictive modeling, which focuses on correlations, causal methods isolate the true impact of a policy by accounting for confounding variables and selection bias. In economics, this is formalized using the potential outcomes framework, where the causal effect of a treatment T on an outcome Y is defined as:
Here, Y(1) and Y(0) represent the potential outcomes under treatment and control, respectively. The fundamental challenge is that only one of these outcomes is observed for each unit, necessitating methods like instrumental variables (IV), difference-in-differences (DiD), or regression discontinuity designs (RDD) to approximate the counterfactual.
Structural Causal Models in Policy Analysis
Economic policy evaluation often relies on structural causal models (SCMs), which encode domain knowledge through directed acyclic graphs (DAGs). These models explicitly represent causal relationships between variables, enabling the identification of estimands even in the presence of unobserved confounders. For example, consider a labor market policy where the treatment T (e.g., a training program) affects wages Y, but education E confounds the relationship:
If E is unobserved, ordinary least squares (OLS) yields a biased estimate of β1. An IV approach, using an instrument Z (e.g., proximity to training centers), can recover the causal effect under the assumptions of relevance (Z affects T) and exogeneity (Z does not affect Y except through T).
Counterfactual Policy Evaluation
In dynamic settings, policies may have delayed or heterogeneous effects. The generalized propensity score extends the propensity score framework to continuous treatments, while synthetic control methods construct counterfactuals by weighting untreated units to match pre-treatment trends of the treated unit. For a policy implemented in region i at time t0, the synthetic control estimator is:
where weights wj are chosen to minimize the discrepancy between pre-treatment outcomes and covariates of unit i and the synthetic control.
Challenges in High-Dimensional Settings
Modern datasets often include high-dimensional covariates (e.g., satellite imagery, transaction records). Machine learning techniques like double/debiased machine learning (DML) address this by separating the estimation of nuisance parameters (e.g., propensity scores) from the causal effect. The DML estimator for the average treatment effect (ATE) solves:
where ê(X) and m̂(X) are estimates of the propensity score and outcome model, respectively, fitted via cross-fitting to avoid overfitting bias.
Case Study: Minimum Wage Effects
A landmark application is Card and Krueger’s 1994 study on minimum wage increases, which used a DiD design comparing employment in New Jersey (treatment) and Pennsylvania (control) before and after the policy change. The causal estimate was derived as:
This design implicitly controls for time-invariant confounders and common macroeconomic shocks, illustrating how causal methods can isolate policy effects from observational data.
4.3 Recommender Systems and A/B Testing
Counterfactual Evaluation in Recommender Systems
Recommender systems often rely on observational data where user-item interactions are confounded by selection bias—users only interact with items they prefer or are exposed to. Traditional offline evaluation metrics like precision@k or recall@k fail to account for this bias, leading to unreliable estimates of a recommender's true performance. Causal inference provides tools to estimate counterfactual outcomes: what would a user's engagement have been if they were recommended a different item?
The inverse propensity scoring (IPS) estimator corrects for this bias by reweighting observed interactions by their propensity scores:
where δi is an indicator of whether item i was recommended, yi is the observed reward (e.g., click or rating), and pi is the probability of item i being recommended. The propensity scores pi can be estimated using logged data or through models like logistic regression.
A/B Testing for Causal Validation
While counterfactual methods provide offline estimates, A/B testing remains the gold standard for measuring causal effects in recommender systems. In a typical setup:
- Users are randomly assigned to either the treatment group (new algorithm) or control group (existing algorithm)
- The average treatment effect (ATE) is computed as:
$$ ATE = \mathbb{E}[Y|T=1] - \mathbb{E}[Y|T=0] $$
- Variance reduction techniques like CUPED (Controlled-experiment Using Pre-Experiment Data) are often applied:
$$ Y_{adj} = Y - \theta(X - \mathbb{E}[X]) $$where X is pre-experiment metrics and θ is derived from covariance between X and Y
Practical Challenges and Solutions
Real-world recommender systems face several challenges when implementing causal methods:
- Non-stationarity: User preferences drift over time, requiring adaptive experimentation frameworks like bandit algorithms that balance exploration-exploitation tradeoffs
- Network effects: In social recommendation systems, a user's behavior may be influenced by their connections, violating the stable unit treatment value assumption (SUTVA)
- Delayed effects: Some recommendations (e.g., career suggestions) may have long-term impacts not captured in short-term A/B tests
Recent advances address these through:
- Clustered randomization to handle network effects
- Surrogate metrics that correlate with long-term outcomes
- Meta-learning approaches that transfer knowledge across experiments
Case Study: Netflix's Bandit Framework
Netflix employs a contextual bandit system where:
Here fθ is a neural network that predicts reward for action (recommendation) a given context x, and τ controls exploration temperature. The system:
- Collects data using Thompson sampling for exploration
- Updates model parameters via offline policy evaluation
- Deploys new policies through phased rollouts
This hybrid approach achieves 30% faster convergence to optimal policies compared to pure A/B testing, while maintaining rigorous causal validity.

5. Foundational Papers and Books
5.1 Foundational Papers and Books
- PDF Recent Developments in Causal Inference and Machine Learning — estimator of the causal effect of interest, but with low external validity, or limited 4Our review differs from recent reviews in sociology (Lundberget al.2022; Molina &Garip 2019) and political science (Grimmer et al. 2021) on machine learning in that we focus on the intersection between causal inference and machine learning.
- A Brief Introduction to Causal Inference in Machine Learning - arXiv.org — This is a lecture note produced for DS-GA 3001.003 "Special Topics in DS - Causal Inference in Machine Learning" at the Center for Data Science, New York University in Spring, 2024. This course was created to target master's and PhD level students with basic background in machine learning but who were not exposed to causal inference or ...
- Recent Developments in Causal Inference and Machine Learning — Researchers have adapted machine learning methods to estimate causal parameters to mitigate these and other concerns central to causal inference. First, to adapt machine learning to the regression-imputation approach, Belloni et al. (2014) propose a double selection procedure in which we fit two LASSO regressions, one for the outcome and one ...
- PDF Causal Inference Tutorial - Massachusetts Institute of Technology — Causal Inference Tutorial Rahul Singh Original: July 23, 2019; Updated: September 10, 2020 The goal of this tutorial is to introduce central concepts, algorithms, and techniques of causal inference for a machine learning audience. There are three sections. 1.Causal frameworks. I present the three most common languages for expressing causal ...
- PDF Elements of Causal Inference - library.oapen.org — A complete list of books published in The Adaptive Computation and Machine ... Elements of causal inference : foundations and learning algorithms / Jonas ... Identifiers: LCCN 2017020087 jISBN 9780262037310 (hardcover : alk. paper) Subjects: LCSH: Machine learning. jLogic, Symbolic and mathematical. jCausa-tion. jInference. jComputer ...
- New Causal ML book (free! online!) : r/datascience - Reddit — Artificial Intelligence & Machine Learning; Computers & Hardware; Consumer Electronics; DIY Electronics; ... Several big names at the intersection of ML and Causal inference, Victor Chernozhukov, Christian Hansen, Nathan Kallus, Martin Spindler, and Vasilis Syrgkanis have put out a new book (free and online) on using ML for causal inference ...
- Machine learning-based causal inference for evaluating intervention in ... — Causal inference methods are particularly crucial in observational research, to which most of the transportation and travel behaviour studies belong. Yet, despite its value and prevalence, applications of causal inference in the transportation sector, typically in travel behaviour research, are still scarce (Brathwaite and Walker, 2018). Most ...
- Introduction to Causal Inference from a Machine Learning Perspective — PDF | On Dec 17, 2020, Brady Neal published Introduction to Causal Inference from a Machine Learning Perspective | Find, read and cite all the research you need on ResearchGate
- Stan and BART for Causal Inference: Estimating Heterogeneous ... - MDPI — A wide range of machine-learning-based approaches have been developed in the past decade, increasing our ability to accurately model nonlinear and nonadditive response surfaces. This has improved performance for inferential tasks such as estimating average treatment effects in situations where standard parametric models may not fit the data well. These methods have also shown promise for the ...
- Statistical Modeling: The Three Cultures · Issue 5.1, Winter 2023 — One such procedure, which we discuss in Section 4.1, is the use of machine learning in the service of causal inference (Künzel et al., 2019). A fused procedure is still compatible with the hypothetico-deductive scientific method but stretches beyond it because it allows for a much larger portion of inductive reasoning (Nelson, 2020).
5.2 Advanced Research Papers
- Causal machine learning methods and use of sample splitting in settings ... — 2 Schuler MS, Rose S. Targeted Maximum Likelihood Estimation for Causal Inference in Observational Studies. American Journal of Epidemiology. 2017;185(1):65-73. doi: 10.1093/aje/kww165; 3 Zivich PN, Breskin A. Machine learning for causal inference: On the use of cross-fit estimators. Epidemiology. 2021;32(3):393-401. doi: 10.1097/EDE ...
- PDF Generalized Optimal Matching Methods for Causal Inference — Journal of Machine Learning Research 21 (2020) 1-54 Submitted 2/19; Revised 3/20; Published 4/20 Generalized Optimal Matching Methods for Causal Inference Nathan Kallus [email protected] Department of Operations Research and Information Engineering and Cornell Tech ... The paper proceeds as follows. In Sec. 2, we setup the problem and provide ...
- Recent Developments in Causal Inference and Machine Learning — Researchers have adapted machine learning methods to estimate causal parameters to mitigate these and other concerns central to causal inference. First, to adapt machine learning to the regression-imputation approach, Belloni et al. (2014) propose a double selection procedure in which we fit two LASSO regressions, one for the outcome and one ...
- A causal inference framework for leveraging external controls in hybrid ... — Leveraging advances in statistical causal inference theory (Hines et al., 2022), we alleviate (2) by building upon the previously proposed doubly robust estimator (Li et al., 2023) and proving conditions under which the estimator is efficient and asymptotically normal even when machine learning methods are used for the nuisance functions, an ...
- Causal inference and machine learning in endocrine epidemiology — 4.2. Applying machine learning to models for outcome variables. Another approach involves using a prediction model for outcome variables (e.g., standardization, G-computation [20, 21]) (Fig. 2b).In this method, a prediction model for outcome variables is first built before we create a copy of data to assign all individuals to either exposed status (for one copy) or unexposed status (for ...
- PDF arXiv:2007.10979v1 [stat.CO] 21 Jul 2020 — Causal inference and machine learning have a symbiotic relationship that is growing deeper. Companies are using machine learning to improve content recommendations, sales, business operations, and to personalize user experiences. These companies will test new algorithms online in order to determine whether the algorithms cause a positive e ect
- PDF Machine learning for efficient and robust causal inference and ... - DiVA — The third paper solves a problem in the field of domain adaptation in machine learning, where the training set is observed but it is not possible to assume that test and training sets follow the same distribution. A weaker assumption is instead considered, referred to as a generalized label shift. This paper proposes a robust and asymptotically ...
- PDF Machine Learning for causal Inference on Observational Data — School of Computer Science and Electronic Engineering Master of Science in Artificial Intelligence Machine Learning for causal Inference on Observational Data by Hernán E. BORRÉ The established scientific way to make claims about cause and effect is to perform a Randomized Controlled Trial (RCT). However, although RCTs are the best way to
- Challenges of Using Text Classifiers for Causal Inference — 4.1. Missing Data. To show how we might use text data to recover from missing data, we introduce missingness for A from Figure 3a to get the model in Figure 3b.The missing arrow from A(1) to R A encodes the MAR assumption, which is sufficient to make it possible to identify the full data distribution from the observed data.. Suppose our motivation is to estimate the causal effect of smoking ...
- DoubleMLDeep: Estimation of Causal Effects with Multimodal Data - arXiv.org — This paper explores the use of unstructured, multimodal data, namely text and images, in causal inference and treatment effect estimation. We propose a neural network architecture that is adapted to the double machine learning (DML) framework, specifically the partially linear model.
5.3 Online Courses and Tutorials
- CS 520 - Causal Inference and Learning — Causal reasoning is an integral part of data science and artificial intelligence. The goal of the course on Causal Inference and Learning is to introduce students to methodologies and algorithms for causal reasoning and connect various aspects of causal inference, including methods developed within computer science, statistics, and economics.
- PDF Elements Of Causal Inference Foundations And Learning Algorithms ... — problems and solutions all learning algorithms are explained so that the user can easily move from the equations in the book to a computer program this book introduces basic machine learning concepts and applications for a broad audience that includes students faculty and industry practitioners we begin by describing how machine learning ...
- PDF Syllabus for Econ 573, Spring 2023 Machine Learning and Econometrics — Course Objective: Students will learn how to explore, visualize, and analyze high-dimensional datasets, build pre-dictive models, and estimate causal effects. The course introduces key concepts and tools that are in high demand in the business environment. Examples of techniques include an advanced overview of linear and logistic regression, regularization, LASSO, cross-validation ...
- Machine Learning-based Causal Inference Tutorial - Bookdown — This tutorial will introduce key concepts in machine learning-based causal inference. It's an ongoing project and new chapters will be uploaded as we finish them.
- Chapter 9 Additional Resources | Machine Learning-based Causal ... — These resources include: the slides and videos for the companion course to these tutorials, on Machine Learning and Causal Inference; publicly available datasets from randomized experiments or observational studies; a report that applies some of the methods from this tutorial to applications in behavioral science and social impact; useful ...
- (PDF) Causal Reasoning in Machine Learning - ResearchGate — This is possible thanks to the inherit humans ability to understand causal relationships and use inductive inference in order to assimilate new information about the world.
- Course: Intelligent Systems for Pattern Recognition - 9 CFU | INF - e ... — The course introduces students to the analysis and design of advanced machine learning and deep learning models for modern pattern recognition problems and discusses how to realize advanced applications exploiting computational intelligence techniques.
- How Do Applied Researchers Use the Causal Forest? A Methodological ... — In an example of a good design, causal machine learning approaches using control-on-observables designs have been shown to be useful when studying electronic health records as these can be processed as high-dimensional data and almost all of the information that doctors will use to make treatment assignment judgments will be encoded in a ...
- Tutorial - Bookdown — This tutorial will introduce key concepts in machine learning-based causal inference. It's an ongoing project and new chapters will be uploaded as we finish them.
- bnlearn · PyPI — bnlearn is a Python package for Causal Discovery by learning the graphical structure of Bayesian networks, parameter learning, inference and sampling methods.








