Estimating Small-Sample Robustness in Time Series Forecasts with Bootstrap
Table Of Contents
Chapter ONE
INTRODUCTION
- 1.1Introduction
- 1.2Background of the Study
- 1.3Statement of the Problem
- 1.4Aim and Objectives of the Study
- 1.5Research Questions
- 1.6Research Hypotheses
- 1.7Significance of the Study
- 1.8Scope and Delimitation of the Study
- 1.9Limitations of the Study
- 1.10Organisation of the Study
- 1.11Operational Definition of Terms
Chapter TWO
LITERATURE REVIEW
- 2.1Conceptual Review: Definitions of Robustness and Small-Sample Challenges in Time Series
- 2.2Conceptual Review: Bootstrap Techniques in Time Series Forecasting
- 2.3Conceptual Review: Small-Sample Inference and Its Implications for Forecast Accuracy
- 2.4Theoretical Framework: Bootstrap Theory and Its Assumptions in Time Series Contexts
- 2.5Theoretical Framework: Robust Statistical Methods and Influence of Sample Size
- 2.6Theoretical Framework: Forecast Evaluation Metrics Under Small Samples
- 2.7Empirical Review: Prior Applications of Bootstrap in Economic and Financial Forecasts
- 2.8Empirical Review: Bootstrap Variants (Block Bootstrap, Sieve Bootstrap) in Practice
- 2.9Empirical Review: Performance of Forecast Intervals vs Point Forecasts in Small Samples
- 2.10Empirical Review: Computational Considerations and Software Implementations
- 2.11Gaps in the Literature and What Remains Unexplored
- 2.12Conceptual Model Summary: Synthesis of Key Constructs and Linkages
Chapter THREE
RESEARCH METHODOLOGY
- 3.1Research Design and Rationale for a Field-Based Simulation Study
- 3.2Philosophical Paradigm: Pragmatism and Epistemic Justification
- 3.3Population of the Study: Financial Time Series in Emerging Markets
- 3.4Sample Size Determination and Sampling Technique for Time Series Experiments
- 3.5Sources and Instruments of Data Collection: Real-World Time Series Data and Simulation Tools
- 3.6Validity and Reliability of Simulation Instruments and Forecast Models
- 3.7Data Preprocessing and Stationarity Assessment
- 3.8Model Specification: Baseline ARIMA/ARIMAX and Bootstrap Variants
- 3.9Method of Data Analysis: Bootstrap-Based Robustness Metrics and Hypothesis Tests
- 3.10Model Validation, Sensitivity Analysis, and Robustness Checks
- 3.11Ethical Considerations in Data Use and Reproducibility
Chapter FOUR
DATA PRESENTATION AND ANALYSIS
- ANALYSIS AND DISCUSSION OF FINDINGS
- 4.1Data Presentation: Descriptive Statistics of Time Series Used
- 4.2Descriptive Analysis of Bootstrap Procedures Employed
- 4.3Hypotheses Testing: Robustness of Forecasts Under Bootstrap Variants
- 4.4Descriptive Comparison of Forecast Accuracy Across Sample Sizes
- 4.5Inferential Analysis: Interval Coverage and Width under Small Samples
- 4.6Interpretation of Results: How Bootstrap Enhances Robustness in Forecasts
- 4.7Discussion in Relation to Conceptual and Theoretical Frameworks
- 4.8Implications for Practice: Field Relevance and Decision-Making under Uncertainty
Chapter FIVE
SUMMARY, CONCLUSION AND RECOMMENDATIONS
- CONCLUSION AND RECOMMENDATIONS
- 5.1Summary of Findings and Their Alignment with Objectives
- 5.2Conclusion: Evidence on Small-Sample Robustness Using Bootstrap
- 5.3Contribution to Knowledge: Methodological and Applied Advances
- 5.4Practical Recommendations for Forecast Practitioners
- 5.5Suggestions for Further Studies: Extensions and Alternative Bootstrap Schemes
Thesis Abstract
This study addresses the challenge of assessing forecast robustness in small-sample time series settings, where traditional asymptotic properties of bootstrap methods may be compromised by limited data and complex dependency structures. The aim is to develop and evaluate a robust bootstrap framework that accurately quantifies forecast uncertainty for short-horizon predictions. The specific objectives are (1) to compare bootstrap variants—residual, block, and sieve bootstrap—in terms of bias, variance, and coverage probabilities for point forecasts and prediction intervals; (2) to adapt bootstrap procedures to autoregressive and moving-average processes with structural breaks and regime shifts commonly observed in real-world data; (3) to establish diagnostic criteria for selecting appropriate bootstrap schemes given sample size, dependence, and nonstationarity; and (4) to provide practical guidelines for practitioners on implementing small-sample bootstrap in financial and macroeconomic forecasting contexts. The methodology adopts a quantitative empirical research design, drawing on time series data from two domains with distinct dynamics to enhance external validity. The population comprises endogenous time series generated by autoregressive integrated moving average (ARIMA) processes observed at daily and monthly frequencies. In Study 1, a financial returns series (S&P 500 index) spanning 3,000 daily observations is utilized; in Study 2, a macroeconomic indicator (U.S. quarterly unemployment rate) spanning 120 quarters is employed. For each series, a set of plausible data-generating processes is specified to introduce potential nonstationarities, structural breaks, and regime changes. The sample sizes mirror realistic constraints 3,000 observations for high-frequency data and 120 observations for low-frequency data. A combination of residual-based, block-based (moving block and stationary bootstrap), and sieve bootstrap methods is implemented, with adjustments to block lengths and model selection criteria governed by cross-validation and information criteria (AIC/BIC). The data collection instruments are solely archival time series extracts from established databases (e.g., Bloomberg Finance, Federal Reserve Economic Data). Validity and reliability are addressed through replication across bootstrap schemes, sensitivity analyses for block length and bandwidth choices, and out-of-sample backtesting to evaluate predictive performance. Analytical procedures include estimation of point forecasts using ARIMA and VAR models augmented with exogenous predictors where appropriate, followed by construction of forecast prediction intervals via bootstrap replicates. The primary analytic focus is on the coverage accuracy and interval width of bootstrap-based prediction intervals under small-sample constraints. Comparative metrics comprise empirical coverage rates, average interval lengths, and quantile-based forecast error measures (CRPS and PI-score). Hypothesis testing targets the null of nominal 95% coverage being statistically indistinguishable from empirical coverage under each bootstrap method, assessed with Wilson binomial confidence intervals and simulation-based p-values. Additional analyses examine the robustness of conclusions to model misspecification and nonstationarity. Theoretical grounding rests on the bootstrap theories of Efron and Tibshirani, augmented by recent advances in dependent data bootstrap (Lahiri) and small-sample corrections for time series (Künsch, Bühlmann). The study also contemplates the practical relevance of the bootstrap in the presence of structural breaks, leveraging literature on regime-switching and time-varying parameter models. Key expected findings include (i) residual bootstrap will exhibit biased coverage in small samples with pronounced serial correlation, (ii) block bootstrap will improve interval validity relative to residual methods but may require adaptive block lengths to balance bias and variance, (iii) sieve bootstrap will provide superior coverage under strong autoregressive dynamics but is sensitive to the chosen approximating model order, and (iv) Bayesian or hybrid approaches incorporating prior information may outperform classical bootstrap in certain nonstationary contexts. The study anticipates that an adaptive composite bootstrap scheme—combining block resampling with model-based residuals—will offer robust coverage across diverse data-generating processes. The study contributes to knowledge by providing a rigorous, empirically validated framework for estimating small-sample forecast uncertainty in time series through bootstrap, including practical guidelines for method selection, block-length optimization, and diagnostic checks tailored to short sequences. It advances understanding of how bootstrap performance interacts with nonstationarity and structural breaks, informing both methodological development and applied forecasting practice in finance and macroeconomics. The main conclusion is that no single bootstrap variant universally outperforms others in small samples; instead, a data-driven, context-aware ensemble approach yields the most reliable forecast intervals. Recommendations include adopting adaptive block-length selection procedures, incorporating model misspecification diagnostics, and disseminating practitioner-oriented decision rules for method choice based on sample size, dependence structure, and detected regime changes.
Thesis Overview
Estimating small-sample robustness in time series forecasts with bootstrap seeks to understand how reliable forecast methods are when we have limited historical data. In many applied settings—finance, economics, engineering—the amount of past observations is small, yet decision makers rely on forecasts to guide critical choices. Traditional time series methods can perform well with large samples but may become unstable or biased when data are scarce. Bootstrap techniques offer a way to approximate sampling variability without relying on strong parametric assumptions, making them particularly attractive for small-sample contexts.
The problem this study addresses is twofold: (1) how to quantify the robustness of common time series forecasts (such as ARIMA, exponential smoothing, and state-space models) when sample sizes are limited, and (2) how bootstrap procedures can be tailored to preserve temporal structure and dependence in small data sets. The gap lies in the lack of systematic, comparable evaluations of bootstrap-based robustness across different models, data-generating processes, and forecast horizons using realistic small-sample conditions.
What the researcher will do step by step:
- Define a set of representative time series models (ARIMA, exponential smoothing, and a simple state-space model) and specify plausible data-generating processes that reflect real-world dependence.
- Establish small-sample scenarios (e.g., 30, 50, and 100 observations) and generate multiple synthetic data sets under each scenario.
- Apply bootstrap methods designed for time series (block bootstrap, stationary bootstrap, and dependent multiplier bootstrap) to produce forecast distributions and estimate predictive intervals.
- Compare forecast accuracy and interval coverage across models and bootstrap schemes using metrics such as root mean squared forecast error, mean absolute percentage error, and coverage probability of prediction intervals.
- Conduct sensitivity analyses to assess the impact of block length, model misspecification, and outliers on robustness results.
- Validate findings with a limited real-data example, if possible, to illustrate practical implications.
The anticipated contribution is a systematic, empirical assessment of small-sample robustness for common forecasting methods using bootstrap, with practical guidance on method selection and parameter tuning. Expected outcomes include recommendations on which bootstrap approach best maintains interval coverage and forecast accuracy under various small-sample conditions, along with diagnostic guidelines for practitioners.
In sum, the study aims to provide concrete, actionable insight into making time series forecasts more trustworthy when data are scarce, supporting better-informed decisions in settings where data are inherently limited.