Estimating Causal Effects in Observational Data via Double ML Methods | Blazingprojects Postgraduate Thesis
Home / Statistics / Estimating Causal Effects in Observational Data via Double ML Methods

Estimating Causal Effects in Observational Data via Double ML Methods

 

Table Of Contents


Chapter ONE

INTRODUCTION

  • 1.1Introduction
  • 1.2Background of the Study
  • 1.3Statement of the Problem
  • 1.4Aim and Objectives of the Study
  • 1.5Research Questions
  • 1.6Research Hypotheses
  • 1.7Significance of the Study
  • 1.8Scope and Delimitation of the Study
  • 1.9Limitations of the Study
  • 1.10Organisation of the Study
  • 1.11Operational Definition of Terms

Chapter TWO

LITERATURE REVIEW

  • 2.1Conceptual Review: Causal Inference in Observational Data
  • 2.2Conceptual Review: Double Machine Learning (DML) Philosophy
  • 2.3Theoretical Framework: Potential Outcomes and Causal Graphs
  • 2.4Theoretical Framework: Neyman Orthogonality and Robustness
  • 2.5Theoretical Framework: Double/Debiased ML in Econometrics
  • 2.6Empirical Review: Early Applications of DML in Economics
  • 2.7Empirical Review: DML in Public Health and Epidemiology
  • 2.8Empirical Review: DML in Education and Labor Economics
  • 2.9Empirical Review: Software Implementations and Practical Challenges
  • 2.10Identified Gaps in the Literature: Limitations of Small-Sample and Nonlinear Settings
  • 2.11Conceptual Model: Causal Pathways Under Doubly Robust Estimation
  • 2.12Synthesis of Review: From Theory to Empirical Application

Chapter THREE

RESEARCH METHODOLOGY

  • 3.1Research Design: Field Study with Observational Data and DML Estimation
  • 3.2Philosophical Paradigm: Pragmatism and Causal Inference Epistemology
  • 3.3Population of the Study: Health Intervention Context in a Metropolitan Region
  • 3.4Sample Size and Sampling Technique: Stratified Sampling with Propensity Score Considerations
  • 3.5Sources and Instruments of Data Collection: Administrative Records and Survey Instruments
  • 3.6Validity and Reliability of Instruments: Measurement Model and Pretest Procedures
  • 3.7Data Preprocessing: Handling Missingness and Normalization
  • 3.8Model Specification: Outcome Model, Treatment Model, and Orthogonalization
  • 3.9Data Analysis Methods: Estimation of Causal Effects Using Double ML
  • 3.10Model Diagnostics: Cross-Fitting, Orthogonality Checks, and Sensitivity Analyses
  • 3.11Ethical Considerations: Informed Consent, Data Privacy, and Governance

Chapter FOUR

DATA PRESENTATION AND ANALYSIS

  • ANALYSIS AND DISCUSSION OF FINDINGS
  • 4.1Data Presentation: Descriptive Statistics by Treatment Group
  • 4.2Descriptive Analysis: Balance Diagnostics Across Covariates
  • 4.3Primary Hypotheses Testing: Causal Effect Estimation via Double ML
  • 4.4Secondary Analyses: Heterogeneous Effects by Subgroups
  • 4.5Robustness Checks: Alternative Model Specifications and Sample Splits
  • 4.6Interpretation of Results: Alignment with Theoretical Frameworks
  • 4.7Comparison with Prior Empirical Studies
  • 4.8Discussion of Findings: Implications for Policy and Practice

Chapter FIVE

SUMMARY, CONCLUSION AND RECOMMENDATIONS

  • CONCLUSION AND RECOMMENDATIONS
  • 5.1Summary of Findings
  • 5.2Conclusion: Causal Effects in Observational Settings via DML
  • 5.3Contribution to Knowledge: Methodological and Applied Implications
  • 5.4Practical Recommendations for Policymakers and Practitioners
  • 5.5Suggestions for Further Studies

Thesis Abstract

Estimating causal effects from observational data remains a central challenge in empirical research, where treatment assignment is not randomized and confounding factors may bias conventional estimators. This study addresses the problem by applying double machine learning (DML) methods to consistently estimate causal effects in the presence of high-dimensional covariates, aiming to improve policy-relevant inferences in non-experimental settings. The objective is to quantify the average treatment effect (ATE) and conditional average treatment effects (CATE) of a binary intervention on a real-world outcome, while rigorously controlling for a rich set of confounders through machine learning-based nuisance parameter estimation. The research advances by integrating theoretical developments in causal inference with robust empirical validation across multiple contexts, thereby delineating the conditions under which DML yields unbiased and efficient estimates in observational data. A quasi-experimental research design is employed, drawing on a temporally bounded observational dataset from a large public health program implemented across 12 metropolitan regions over a four-year period. The population comprises adult residents eligible for the program, with a target sample size of 8,000 individuals who encountered the treatment and 8,000 comparable controls identified via propensity score–based matching on pre-treatment covariates. Data collection utilizes administrative records linked to survey instruments, including detailed demographic information, prior health outcomes, socioeconomic indicators, geographic mobility, and program participation histories. The study also assembles a comprehensive covariate set (n ? 150) drawn from health records, social services data, and regional economic indicators to satisfy the overlap/positivity assumptions required for DML estimation. Data quality is ensured through validation checks, harmonization procedures, and imputation of missing values where appropriate, guided by established best practices in causal machine learning. The analysis proceeds in two stages. First, nuisance functions for the propensity score and the outcome regression are estimated using a diverse library of machine learning algorithms, including gradient boosting (XGBoost), neural networks, Lasso and ridge regression, and random forests, with cross-validated hyperparameters. Second, the DML estimator is applied to obtain unbiased estimates of the ATE and CATE, leveraging Neyman-orthogonal moments to mitigate regularization bias. Complementary analyses include traditional regression and propensity score methods for robustness checks, as well as sensitivity analyses to assess the impact of potential unmeasured confounding using Rosenbaum bounds. Model diagnostics focus on the overlap condition, covariate balance after weighting, and the stability of estimates across alternative ML specifications. The analysis is conducted in R and Python, with reproducible scripts and a registered protocol to ensure transparency and replicability. Key findings are anticipated to indicate a statistically significant positive effect of the program on the primary health outcome, with estimated ATEs exhibiting confidence intervals that remain robust under alternative ML specifications and sample splits. Heterogeneous treatment effects are expected to reveal that the impact is larger among individuals with lower baseline health metrics and higher exposure to social determinants of disadvantage, demonstrating the value of CATE estimation for targeted policy design. The study also expects that the DML estimates will exhibit tighter standard errors and reduced bias compared with conventional regression-based approaches, especially in the presence of high-dimensional controls and complex nonlinear relationships among covariates. This research contributes to knowledge by operationalizing double machine learning in a large-scale public policy context, illustrating practical procedures for high-dimensional nuisance estimation, and providing empirical evidence on the reliability of DML under real-world data constraints. The integration of causal inference theory with machine learning in a field setting advances methodological rigor for observational studies and informs policymakers about which subpopulations derive the greatest benefit from intervention programs. The main conclusions are that, when applied with careful adherence to overlap conditions and rigorous cross-validation, DML yields credible, policy-relevant causal estimates in observational data. Recommendations include adopting DML as a standard analytic approach in evaluation studies, expanding data linkage to enhance covariate richness, and conducting pre-registered analytic plans to strengthen inference in non-experimental settings. Future work could extend the framework to multi-valued treatments and longitudinal outcomes to capture dynamic treatment effects over time.

Thesis Overview

Estimating Causal Effects in Observational Data via Double ML Methods aims to determine how a treatment or intervention influences an outcome when you cannot run a randomized experiment. In many fields, researchers rely on observational data collected from existing records, surveys, or administrative databases. The challenge is that treatment assignment is not random, so simple comparisons can be biased by confounding factors. Double ML (DML) methods are a modern family of approaches designed to estimate causal effects in such settings by separately modeling the relationship between covariates and the treatment, and between covariates and the outcome, then combining these models to obtain an unbiased estimate of the treatment effect under certain assumptions. Why it matters: accurately estimating causal effects from observational data broadens the scope of evidence available for policy, medicine, economics, and social sciences when randomized trials are impractical, unethical, or too costly. DML methods help mitigate bias from high-dimensional confounders and allow for robust inference even when the number of covariates approaches or exceeds the sample size. What problem or gap it addresses: traditional regression approaches can fail in high-dimensional settings, oversimplify the treatment assignment mechanism, or rely on strong linear assumptions. DML provides a principled framework that uses machine learning tools to flexibly estimate nuisance components (propensity scores and outcome models) while preserving valid inference for the causal parameter. What the researcher will do step by step: - Define a clear causal estimand (e.g., average treatment effect) and specify assumptions (unconfoundedness, overlap, and correct model specification for nuisance terms). - Collect or access a large observational dataset with a well-defined treatment, outcome, and rich covariates. - Split the data into folds using cross-fitting; for each fold, estimate nuisance functions: - Model the treatment assignment given covariates (propensity score) using flexible learners. - Model the outcome given covariates and treatment using appropriate learners. - Combine the nuisance estimates to form the Double ML estimator of the causal effect and compute standard errors via influence-function-based methods. - Conduct sensitivity analyses to assess robustness to violations of assumptions. - Perform subgroup or heterogeneity analyses to explore varying treatment effects across contexts. Expected contribution and outcome: the study advances practical guidance on implementing DML in real-world data, demonstrates improvements in bias reduction and inference validity over conventional methods, and provides a template for applying DML to policy evaluation or clinical evidence with high-dimensional covariates. The expected result is a reliable, reproducible estimation procedure delivering credible policy-relevant insights.

Blazingprojects Mobile App

📚 Over 50,000 Research Thesis
📱 100% Offline: No internet needed
📝 Over 98 Departments
🔍 Thesis-to-Journal Publication
🎓 Undergraduate/Postgraduate Thesis
📥 Instant Whatsapp/Email Delivery

Blazingprojects App

Related Research

Agriculture and fore. 2 min read

Impact of agroforestry on soil carbon sequestration in smallholder farms...

Impact of agroforestry on soil carbon sequestration in smallholder farms This research examines how integrating trees with crops or livestock systems (agrofore...

BP
Blazingprojects
Read more →
Agricultural science. 2 min read

Impact of experiential learning on agricultural science practical skills in high sch...

Experiential learning refers to learning through direct experience, reflection, and application, rather than passive classroom instruction. In agricultural scie...

BP
Blazingprojects
Read more →
Adult education. 2 min read

Impact of Workplace Literacies on Adult Learner Career Advancement ...

This research explores how the everyday language and literacy demands found in workplaces influence how adult learners progress in their careers. Workplace lite...

BP
Blazingprojects
Read more →
Zoology. 2 min read

Urban-adjacent bat foraging dynamics in fragmented forests: an empirical study...

Urban-adjacent bat foraging dynamics in fragmented forests: an empirical study This research explores how bats that live near cities use fragmented forest habi...

BP
Blazingprojects
Read more →
Veterinary Medicine. 2 min read

Impact of farm-level biosecurity on bovine respiratory disease incidence in dairy he...

This research investigates how farm-level biosecurity practices influence the occurrence of bovine respiratory disease (BRD) in dairy herds. BRD is a major heal...

BP
Blazingprojects
Read more →
Urban and Regional P. 4 min read

Assessing Urban Form Impacts on Walkability in Small Cities ...

This research investigates how the physical layout and design of small cities influence how easily people can walk to destinations, stay safe, and enjoy their s...

BP
Blazingprojects
Read more →
Theatre Art. 3 min read

Audience Perception of Hybrid Theatre: A Field Study in Meta-Theatrical Performance...

This research investigates how audiences perceive hybrid theatre performances that blend live acting with digital or meta-theatrical elements, such as self-refe...

BP
Blazingprojects
Read more →
Technical education. 4 min read

Impact of Industry 4.0 Training on Vocational Students’ Competence Gains...

This thesis investigates how Industry 4.0 training affects the practical and cognitive competencies of vocational students. It asks whether integrating Industry...

BP
Blazingprojects
Read more →
Surveying and Geo-in. 4 min read

Assessing UAV-derived Point Cloud Accuracy in Street-Level Mapping Networks...

This research explores how accurate 3D point clouds generated by unmanned aerial vehicles (UAVs) are when used to map street-level environments, such as buildin...

BP
Blazingprojects
Read more →
WhatsApp Click here to chat with us