Richard Hanh, Math & Statistics, ASU
When
Where
Six lemmas in defense of machine learning for causal inference
Abstract: The purpose of this talk is to justify the use of machine learning for causal inference. A careful justification is required for two reasons. First, because many prominent methodologists continue to deny both the utility of nonlinear methods as well as the importance of heterogeneous treatment effects. Second, because those who have most enthusiastically embraced machine learning for causal inference tend to over-promise and misconstrue the true nature of its benefits. At the heart of this incongruity is the following fact: Whether via explicit randomization, or through more exotic assumptions that imply pseudo-randomization, statistical approaches to causality identify average causal effects. To think clearly about the role of machine learning in causal inference, we must be clear about two questions: “Which average is, or should be, estimated?” and “How should we estimate it?” By way of answering these questions I will present six basic — but perhaps not widely known or appreciated — “lemmas” which serve to motivate the use of flexible, regularized machine learning for modern scientific inquiry.