Development of a Framework for Predictive Dental Caries Risk Modelling Using Integrated Patient Data
Table Of Contents
Chapter ONE
INTRODUCTION
- 1.
- 1.1Introduction
- 2.
- 1.2Background of the Study
- 3.
- 1.3Statement of the Problem
- 4.
- 1.4Aim and Objectives of the Study
- 5.
- 1.5Research Questions
- 6.
- 1.6Research Hypotheses
- 7.
- 1.7Significance of the Study
- 8.
- 1.8Scope and Delimitation of the Study
- 9.
- 1.9Limitations of the Study
- 10.
- 1.10Organisation of the Study
- 11.
- 1.11Operational Definition of Terms
Chapter TWO
LITERATURE REVIEW
- 1.
- 2.1Conceptualization of Caries Risk and Prediction
- 2.
- 2.2Integrated Patient Data in Dental Analytics
- 3.
- 2.3Review of Dental Caries Risk Assessment Tools
- 4.
- 2.4Data Quality and Standardization in Dentistry
- 5.
- 2.5Traditional vs. Modern Predictive Modelling in Dentistry
- 6.
- 2.6Theoretical Frameworks Relevant to Caries Prediction
- 7.
- 2.7Empirical Evidence on Multimodal Data Modelling
- 8.
- 2.8Machine Learning Approaches in Caries Risk Prediction
- 9.
- 2.9Health Informatics and Privacy Considerations
- 10.
- 2.10The Role of Microbiome and Salivary Biomarkers
- 11.
- 2.11Socioeconomic and Behavioral Determinants of Caries
- 12.
- 2.12Gaps in the Literature and Areas for Model Development
- 13.
- 2.13Conceptual Model or Summary Diagram of Review
Chapter THREE
RESEARCH METHODOLOGY
- 1.
- 3.1Research Design tailored to a Predictive Framework
- 2.
- 3.2Philosophical Paradigm guiding the Modelling Study
- 3.
- 3.3Population of the Study: Targeted Dental Patient Cohorts
- 4.
- 3.4Sample Size and Sampling Technique for Integrated Data
- 5.
- 3.5Sources of Data: Clinical, Demographic, and Omics/Salivary Measures
- 6.
- 3.6Instruments and Data Collection Tools
- 7.
- 3.7Validity and Reliability of the Integrated Data Instruments
- 8.
- 3.8Data Preprocessing and Feature Engineering Procedures
- 9.
- 3.9Model Specification and Analytical Framework
- 10.
- 3.10Ethical Considerations in Data Handling and Modelling
Chapter FOUR
DATA PRESENTATION AND ANALYSIS
- ANALYSIS AND DISCUSSION OF FINDINGS
- 1.
- 4.1Data Presentation Overview and Pipeline
- 2.
- 4.2Descriptive Statistics of Integrated Patient Data
- 3.
- 4.3Data Quality and Missingness Analysis
- 4.
- 4.4Model Development: Baseline Caries Risk Framework
- 5.
- 4.5Model Evaluation and Validation Results
- 6.
- 4.6Hypotheses Testing and Inferential Insights
- 7.
- 4.7Sensitivity and Specificity of the Predictive Framework
- 8.
- 4.8Discussion of Findings in Relation to Prior Literature
Chapter FIVE
SUMMARY, CONCLUSION AND RECOMMENDATIONS
- CONCLUSION AND RECOMMENDATIONS
- 1.
- 5.1Summary of Key Findings
- 2.
- 5.2Conclusions Drawn from the Predictive Framework
- 3.
- 5.3Contributions to Knowledge and Practice in Dentistry
- 4.
- 5.4Practical Implications for Clinical Decision-Making
- 5.
- 5.5Recommendations for Implementation and Policy
- 6.
- 5.6Suggestions for Future Research and Model Enhancements
Thesis Abstract
The study addresses the persistent challenge of accurately predicting individual dental caries risk in diverse populations by leveraging integrated patient data from clinical records, epidemiological surveys, and lifestyle, genetic, and microbiome indicators. Despite advances in caries epidemiology, existing risk models often rely on limited data sources and fail to generalize across settings, hindering personalized prevention strategies and efficient allocation of preventive resources. The aim is to develop a robust framework for predictive dental caries risk modelling that synthesizes multi-source patient data into a parsimonious, clinically actionable model. Specific objectives include (1) identifying and harmonizing relevant data modalities across clinical, behavioral, genetic, and microbial domains; (2) evaluating the predictive performance of machine learning and traditional statistical approaches for caries risk estimation; (3) developing a modular, interoperable framework that supports real-time risk scoring within routine dental care; (4) validating the framework in two demographically distinct cohorts to assess generalizability; and (5) outlining governance, ethical, and implementation considerations for integration into dental practice. A mixed-methods, multisite study design will be employed. The population comprises patients aged 6–18 and 19–65 from three tertiary dental centers and two community dental clinics, with an anticipated total sample of 6,000 participants for model development and 2,000 for external validation. Data collection will integrate electronic health records (demographics, past caries experience, fluoride exposure, treatment history), clinical examination findings (dmft/DMFT, lesion activity, occlusal pattern), salivary biomarkers (pH, flow rate, mutans streptococci quantified by qPCR), genomic variants linked to caries susceptibility, and lifestyle factors (dietary sugar intake, oral hygiene practices, socioeconomic status). Standardized data harmonization protocols and privacy-preserving techniques will be implemented. Instruments include calibrated caries assessment tools, validated dietary and behavior questionnaires, saliva assays, and genotyping arrays targeting known caries-associated loci. Validity and reliability will be established through pilot testing, inter-examiner calibration (kappa >0.80 for lesion scoring), and reliability testing for laboratory assays (intra-assay CV <10%). Analytical methods will commence with data preprocessing, missing data imputation (multiple imputation by chained equations), and feature engineering to derive composite risk indicators. Predictive modelling will compare logistic regression, random forests, gradient boosting (XGBoost), and deep learning architectures capable of handling heterogeneous data. Model performance will be evaluated using discrimination (AUC/ROC), calibration plots, Brier score, and decision-curve analysis to determine clinical usefulness. Nested cross-validation will optimize hyperparameters, while feature importance and SHAP (SHapley Additive exPlanations) values will be employed to interpret model decisions. The theoretical underpinning will draw on the Health Belief Model to contextualize behavior-related predictors and the Ecological Systems Theory to justify multi-level data integration. A conceptual framework will illustrate data flow from acquisition to real-time risk scoring, with modular components enabling plug-and-play substitutions as new biomarkers emerge. Expected findings include (a) identification of a minimal, high-yield feature set achieving AUCs ?0.80 in internal validation and ?0.75 in external validation; (b) demonstration that integrated data (clinical + behavioral + biological) outperforms models restricted to traditional clinical indicators; (c) evidence that model calibration remains robust across age groups and settings, with systematic biases identified and corrected; and (d) a practical risk-scoring prototype with a user-centric interface suitable for incorporation into electronic dental records. The study will contribute to knowledge by operationalizing a scalable framework for multimodal data fusion in dental caries prediction, informing precision prevention strategies, and offering methodological guidance for similar computational phenotyping in oral health. The main conclusion will propose that predictive caries risk modelling based on integrated patient data is feasible, generalizable, and adds substantial predictive value over conventional risk assessments, with demonstrated readiness for pilot implementation in routine care. Recommendations include integrating the framework into dental education and continuing professional development, conducting longitudinal implementation trials to assess impact on caries incidence, and establishing governance policies for data sharing, privacy, and ethical use of genomic information in clinical decision-making.
Thesis Overview
This research seeks to build a practical framework for predicting individual risk of dental caries by integrating multiple patient data streams. The central idea is that caries development is influenced by a combination of biological factors (e.g., salivary flow, microbial composition), behavioral habits (diet, oral hygiene), clinical indicators (previous caries experience, fluoride exposure), and sociodemographic context. By combining these data sources into a single predictive framework, clinicians can identify high-risk patients earlier and tailor preventive interventions more effectively.
Why it matters: Existing caries risk assessment tools often rely on limited datasets or isolated factors, which reduces predictive accuracy and limits clinical utility. An integrated, theory-informed framework has the potential to improve risk stratification, support personalized care planning, and optimize allocation of preventive resources in dental practice.
Problem or knowledge gap: There is a lack of robust models that fuse heterogeneous data types—biomarkers, electronic health records, patient-reported behaviors, and environmental factors—into a coherent predictive system. Additionally, there is limited validation of such models across diverse populations and care settings.
What the researcher will do step by step:
1. Define the theoretical grounding, drawing on health behavior theories and the ecological model of health to justify data integration.
2. Design a multi-source data model that harmonizes biological, clinical, behavioral, and environmental variables.
3. Conduct a prospective cohort study with 800 participants drawn from general dental clinics, capturing baseline measures and follow-ups at 12 and 24 months.
4. Collect data using: standardized clinical examinations for caries experience, saliva testing for flow and buffering capacity, microbial profiling where feasible, validated dietary and oral hygiene questionnaires, electronic health records for medical/demographic data, and geospatial/environmental indicators (socioeconomic status, neighborhood fluoride exposure).
5. Preprocess data to address missingness, scale variables, and handle heterogeneity across sources.
6. Develop predictive models using supervised machine learning approaches (logistic regression, random forests, gradient boosting) and compare performance to traditional caries risk tools.
7. Validate the model internally and externally with an independent sample from a different clinic network.
8. Interpret model outputs to identify key predictors and assess clinical utility through decision-curve analysis.
9. Engage stakeholders to refine the framework into a usable clinical decision-support tool.
Expected contribution: A transferable, evidence-based framework that improves predictive accuracy for caries risk by integrating diverse data, with documented validation across settings and clear guidance for implementation in routine care.
Anticipated outcome: A validated predictive framework with identified high-risk profiles and practical recommendations for personalized prevention, plus a prototype decision-support workflow for integration into dental practice management systems.