Перейти к содержимому

Jared Murray: A Unifying Weighting Perspective on Causal Machine Learning

Online Causal Inference Seminar

0:00 / 0:00

Jared Murray: A Unifying Weighting Perspective on Causal Machine Learning

1 051 просмотр · 1 год назад
Online Causal Inference Seminar
9,1 тыс. подписчиков
1 051 просмотр · 1 год назад
Subscribe to the channel to get notified when we release a new video. Like the video to tell YouTube that you want more content like this on your feed. See our website for future seminars: https://sites.google.com/view/ocis/home Tuesday, November 19, 2024: Jared S. Murray (University of Texas at Austin) Title: A Unifying Weighting Perspective on Causal Machine Learning: Kernel Methods, Gaussian Processes, and Bayesian Tree Models Discussant: Rahul Singh (Harvard University) Abstract: Causal machine learning methods based on kernel methods are powerful tools for estimating heterogeneous treatment effects; examples include kernel ridge regression, (causal) random forests, and many neural networks. A known but underappreciated result is that many of these methods have an equivalent representation as weighting estimators, with weights that correspond to an estimate of the Riesz representer of the estimand. This paper catalogs results about the weighting representation of heterogenous effect estimates under kernel ridge regression estimates of outcome models, and we provide new results about kernel ridge estimates that incorporate propensity scores in the spirit of the Robinson ``regression-on-residuals'' transformation. We show that under mild conditions, these R-parameterized outcome models produce implied weights that approximately balance broad classes of functions between treated and control groups when estimating the average treatment effect in {\em any} target population -- even under some forms of outcome model misspecification. This result connects the desirable properties of the Robinson transform and its corresponding Neyman orthogonal score/risk functions to the balancing properties of the implied weights. We also show that this balancing property is generally insufficient to completely debias estimates even if the outcome model is correctly specified. We characterize the remaining bias via ``target imbalance'': the difference between means in the model-implied target population and the actual target population. We propose broadly applicable debiasing strategies that remain inside the outcome modeling framework. Finally, we show that many of these results translate to the R-learner with linear smoothers, which is also a weighting estimator. We then extend our results to a large class of Bayesian nonparametric models used in causal machine learning via their representation as conditional Gaussian process (GP) regression models. Examples include BART, Bayesian causal forests, high-dimensional Bayesian regression with shrinkage or selection priors, many Bayesian neural networks, and other generic GP regression models. We use the connection between GP and kernel ridge regression to compute and interpret the model-implied weights. Their form sheds new light on why Bayesian tree models are especially effective for estimating heterogeneous effects. Finally, we leverage the implied weighting representation to introduce new tools for diagnosing violations of causal assumptions, model criticism, and method comparisons.