Estimating Causal Effects in Observational Data via Double ML Methods
Table Of Contents
Chapter ONE
INTRODUCTION
- 1.1Introduction
- 1.2Background of the Study
- 1.3Statement of the Problem
- 1.4Aim and Objectives of the Study
- 1.5Research Questions
- 1.6Research Hypotheses
- 1.7Significance of the Study
- 1.8Scope and Delimitation of the Study
- 1.9Limitations of the Study
- 1.10Organisation of the Study
- 1.11Operational Definition of Terms
Chapter TWO
LITERATURE REVIEW
- 2.1Conceptual Review: Causal Inference in Observational Data
- 2.2Conceptual Review: Double Machine Learning (DML) Philosophy
- 2.3Theoretical Framework: Potential Outcomes and Causal Graphs
- 2.4Theoretical Framework: Neyman Orthogonality and Robustness
- 2.5Theoretical Framework: Double/Debiased ML in Econometrics
- 2.6Empirical Review: Early Applications of DML in Economics
- 2.7Empirical Review: DML in Public Health and Epidemiology
- 2.8Empirical Review: DML in Education and Labor Economics
- 2.9Empirical Review: Software Implementations and Practical Challenges
- 2.10Identified Gaps in the Literature: Limitations of Small-Sample and Nonlinear Settings
- 2.11Conceptual Model: Causal Pathways Under Doubly Robust Estimation
- 2.12Synthesis of Review: From Theory to Empirical Application
Chapter THREE
RESEARCH METHODOLOGY
- 3.1Research Design: Field Study with Observational Data and DML Estimation
- 3.2Philosophical Paradigm: Pragmatism and Causal Inference Epistemology
- 3.3Population of the Study: Health Intervention Context in a Metropolitan Region
- 3.4Sample Size and Sampling Technique: Stratified Sampling with Propensity Score Considerations
- 3.5Sources and Instruments of Data Collection: Administrative Records and Survey Instruments
- 3.6Validity and Reliability of Instruments: Measurement Model and Pretest Procedures
- 3.7Data Preprocessing: Handling Missingness and Normalization
- 3.8Model Specification: Outcome Model, Treatment Model, and Orthogonalization
- 3.9Data Analysis Methods: Estimation of Causal Effects Using Double ML
- 3.10Model Diagnostics: Cross-Fitting, Orthogonality Checks, and Sensitivity Analyses
- 3.11Ethical Considerations: Informed Consent, Data Privacy, and Governance
Chapter FOUR
DATA PRESENTATION AND ANALYSIS
- ANALYSIS AND DISCUSSION OF FINDINGS
- 4.1Data Presentation: Descriptive Statistics by Treatment Group
- 4.2Descriptive Analysis: Balance Diagnostics Across Covariates
- 4.3Primary Hypotheses Testing: Causal Effect Estimation via Double ML
- 4.4Secondary Analyses: Heterogeneous Effects by Subgroups
- 4.5Robustness Checks: Alternative Model Specifications and Sample Splits
- 4.6Interpretation of Results: Alignment with Theoretical Frameworks
- 4.7Comparison with Prior Empirical Studies
- 4.8Discussion of Findings: Implications for Policy and Practice
Chapter FIVE
SUMMARY, CONCLUSION AND RECOMMENDATIONS
- CONCLUSION AND RECOMMENDATIONS
- 5.1Summary of Findings
- 5.2Conclusion: Causal Effects in Observational Settings via DML
- 5.3Contribution to Knowledge: Methodological and Applied Implications
- 5.4Practical Recommendations for Policymakers and Practitioners
- 5.5Suggestions for Further Studies
Thesis Abstract
Estimating causal effects from observational data remains a central challenge in empirical research, where treatment assignment is not randomized and confounding factors may bias conventional estimators. This study addresses the problem by applying double machine learning (DML) methods to consistently estimate causal effects in the presence of high-dimensional covariates, aiming to improve policy-relevant inferences in non-experimental settings. The objective is to quantify the average treatment effect (ATE) and conditional average treatment effects (CATE) of a binary intervention on a real-world outcome, while rigorously controlling for a rich set of confounders through machine learning-based nuisance parameter estimation. The research advances by integrating theoretical developments in causal inference with robust empirical validation across multiple contexts, thereby delineating the conditions under which DML yields unbiased and efficient estimates in observational data. A quasi-experimental research design is employed, drawing on a temporally bounded observational dataset from a large public health program implemented across 12 metropolitan regions over a four-year period. The population comprises adult residents eligible for the program, with a target sample size of 8,000 individuals who encountered the treatment and 8,000 comparable controls identified via propensity score–based matching on pre-treatment covariates. Data collection utilizes administrative records linked to survey instruments, including detailed demographic information, prior health outcomes, socioeconomic indicators, geographic mobility, and program participation histories. The study also assembles a comprehensive covariate set (n ? 150) drawn from health records, social services data, and regional economic indicators to satisfy the overlap/positivity assumptions required for DML estimation. Data quality is ensured through validation checks, harmonization procedures, and imputation of missing values where appropriate, guided by established best practices in causal machine learning. The analysis proceeds in two stages. First, nuisance functions for the propensity score and the outcome regression are estimated using a diverse library of machine learning algorithms, including gradient boosting (XGBoost), neural networks, Lasso and ridge regression, and random forests, with cross-validated hyperparameters. Second, the DML estimator is applied to obtain unbiased estimates of the ATE and CATE, leveraging Neyman-orthogonal moments to mitigate regularization bias. Complementary analyses include traditional regression and propensity score methods for robustness checks, as well as sensitivity analyses to assess the impact of potential unmeasured confounding using Rosenbaum bounds. Model diagnostics focus on the overlap condition, covariate balance after weighting, and the stability of estimates across alternative ML specifications. The analysis is conducted in R and Python, with reproducible scripts and a registered protocol to ensure transparency and replicability. Key findings are anticipated to indicate a statistically significant positive effect of the program on the primary health outcome, with estimated ATEs exhibiting confidence intervals that remain robust under alternative ML specifications and sample splits. Heterogeneous treatment effects are expected to reveal that the impact is larger among individuals with lower baseline health metrics and higher exposure to social determinants of disadvantage, demonstrating the value of CATE estimation for targeted policy design. The study also expects that the DML estimates will exhibit tighter standard errors and reduced bias compared with conventional regression-based approaches, especially in the presence of high-dimensional controls and complex nonlinear relationships among covariates. This research contributes to knowledge by operationalizing double machine learning in a large-scale public policy context, illustrating practical procedures for high-dimensional nuisance estimation, and providing empirical evidence on the reliability of DML under real-world data constraints. The integration of causal inference theory with machine learning in a field setting advances methodological rigor for observational studies and informs policymakers about which subpopulations derive the greatest benefit from intervention programs. The main conclusions are that, when applied with careful adherence to overlap conditions and rigorous cross-validation, DML yields credible, policy-relevant causal estimates in observational data. Recommendations include adopting DML as a standard analytic approach in evaluation studies, expanding data linkage to enhance covariate richness, and conducting pre-registered analytic plans to strengthen inference in non-experimental settings. Future work could extend the framework to multi-valued treatments and longitudinal outcomes to capture dynamic treatment effects over time.
Thesis Overview
Estimating Causal Effects in Observational Data via Double ML Methods aims to determine how a treatment or intervention influences an outcome when you cannot run a randomized experiment. In many fields, researchers rely on observational data collected from existing records, surveys, or administrative databases. The challenge is that treatment assignment is not random, so simple comparisons can be biased by confounding factors. Double ML (DML) methods are a modern family of approaches designed to estimate causal effects in such settings by separately modeling the relationship between covariates and the treatment, and between covariates and the outcome, then combining these models to obtain an unbiased estimate of the treatment effect under certain assumptions.
Why it matters: accurately estimating causal effects from observational data broadens the scope of evidence available for policy, medicine, economics, and social sciences when randomized trials are impractical, unethical, or too costly. DML methods help mitigate bias from high-dimensional confounders and allow for robust inference even when the number of covariates approaches or exceeds the sample size.
What problem or gap it addresses: traditional regression approaches can fail in high-dimensional settings, oversimplify the treatment assignment mechanism, or rely on strong linear assumptions. DML provides a principled framework that uses machine learning tools to flexibly estimate nuisance components (propensity scores and outcome models) while preserving valid inference for the causal parameter.
What the researcher will do step by step:
- Define a clear causal estimand (e.g., average treatment effect) and specify assumptions (unconfoundedness, overlap, and correct model specification for nuisance terms).
- Collect or access a large observational dataset with a well-defined treatment, outcome, and rich covariates.
- Split the data into folds using cross-fitting; for each fold, estimate nuisance functions:
- Model the treatment assignment given covariates (propensity score) using flexible learners.
- Model the outcome given covariates and treatment using appropriate learners.
- Combine the nuisance estimates to form the Double ML estimator of the causal effect and compute standard errors via influence-function-based methods.
- Conduct sensitivity analyses to assess robustness to violations of assumptions.
- Perform subgroup or heterogeneity analyses to explore varying treatment effects across contexts.
Expected contribution and outcome: the study advances practical guidance on implementing DML in real-world data, demonstrates improvements in bias reduction and inference validity over conventional methods, and provides a template for applying DML to policy evaluation or clinical evidence with high-dimensional covariates. The expected result is a reliable, reproducible estimation procedure delivering credible policy-relevant insights.