Digital Humanities Text Mining for Postcolonial Literary Networks
Table Of Contents
Chapter ONE
INTRODUCTION
- 1.1Introduction
- 1.2Background of the Study
- 1.3Statement of the Problem
- 1.4Aim and Objectives of the Study
- 1.5Research Questions
- 1.6Research Hypotheses
- 1.7Significance of the Study
- 1.8Scope and Delimitation of the Study
- 1.9Limitations of the Study
- 1.10Organisation of the Study
- 1.11Operational Definition of Terms
Chapter TWO
LITERATURE REVIEW
- 2.1Conceptual Review: Digital Humanities and Postcolonial Literary Networks
- 2.2Conceptual Review: Text Mining in Literary Studies
- 2.3Conceptual Review: Network Theory in Literary Research
- 2.4Theoretical Framework: Postcolonial Theory and ICT-Mediated Reading
- 2.5Theoretical Framework: Complex Networks Theory
- 2.6Theoretical Framework: Data-Driven Literary Criticism
- 2.7Empirical Review: Earlier Applications of Text Mining to Postcolonial Texts
- 2.8Empirical Review: Mapping Publication and Citation Networks in Postcolonial Studies
- 2.9Empirical Review: Linguistic and Stylistic Trends in Postcolonial Narratives
- 2.10Empirical Review: Digital Archives and Metadata in Postcolonial Literatures
- 2.11Gaps in the Literature: Methodological and Data Gaps
- 2.12Gaps in the Literature: Conceptual Gaps in Postcolonial Network Mapping
- 2.13Conceptual Model: Summary Diagram of Digital Humanities Text Mining for Postcolonial Networks
Chapter THREE
RESEARCH METHODOLOGY
- 3.1Research Design: Integrating Text Mining with Network Analysis
- 3.2Philosophical Paradigm: Constructivist-Pragmatist Stance for ICT-Driven Inquiry
- 3.3Population of the Study: Postcolonial Literatures Across English-Language Corpora
- 3.4Sample Size and Sampling Technique: Stratified Corpus Sampling and Snowball Metadata
- 3.5Sources and Instruments of Data Collection: Digitized Text Corpora, Metadata, and API Tools
- 3.6Validity and Reliability of Instruments: Triangulation and Intercoder Reliability
- 3.7Data Preprocessing and Cleaning Procedures
- 3.8Methods of Data Analysis: Text Mining, Topic Modeling, and Network Analysis
- 3.9Model Specification: Network Construction and Community Detection Framework
- 3.10Ethical Considerations: Intellectual Property and Data Privacy
- 3.11Limitations and Delimitations: Language, Accessibility, and Genre Constraints
Chapter FOUR
DATA PRESENTATION AND ANALYSIS
- ANALYSIS AND DISCUSSION OF FINDINGS
- 4.1Data Presentation: Corpus Overview and Metadata Statistics
- 4.2Descriptive Analysis: Corpus Size, Language Features, and Genre Distribution
- 4.3Network Construction: Nodes, Edges, and Temporal Dynamics
- 4.4Community Structure and Network Metrics: Modularity, Centrality, and Cohesion
- 4.5Topic Modeling Results: Thematic Clusters in Postcolonial Networks
- 4.6Hypotheses Testing: Relationships Between Thematic Centrality and Publication Influence
- 4.7Interpretation of Results: Aligning Findings with Postcolonial Theory
- 4.8Discussion: Implications for Digital Humanities Methodology and Literary Network Studies
Chapter FIVE
SUMMARY, CONCLUSION AND RECOMMENDATIONS
- CONCLUSION AND RECOMMENDATIONS
- 5.1Summary of Findings
- 5.2Conclusion
- 5.3Contribution to Knowledge: Methodological and Theoretical Advances
- 5.4Practical and Policy Recommendations for Digital Humanities Practice
- 5.5Suggestions for Further Studies
Thesis Abstract
This study addresses the fragmentation of postcolonial literary networks in the digital age, where traditional scholarly discourse has often overlooked transregional interactions and intertextual flows across colonial and postcolonial archives. The problem is twofold first, limited coding schemes and sparse metadata impede scalable mapping of literary influence; second, existing digital humanities approaches frequently underutilize network theory to capture latent connections among authors, genres, and publishing ecologies. The aim is to develop a scalable text-mining and network-analytic framework that uncovers postcolonial literary networks through machine-assisted data extraction and theory-driven interpretation. Specific objectives are (1) to compile a corpus of 1,200 digitized literary texts spanning Africa, the Caribbean, South Asia, and the metropole, published between 1850 and 2000; (2) to construct a multilayer network incorporating co-citation, intertextual allusion, thematic similarity, and publishing-ecology links; (3) to apply topic modeling, stylometric analysis, and named-entity recognition to identify latent motifs and authorial communities; (4) to test hypotheses derived from postcolonial theory using network centrality, modularity, and assortativity measures; and (5) to formulate a replicable analytical pipeline for DH-based investigations of postcolonial literary systems. The methodology integrates a mixed-methods design anchored in digital humanities and network theory. The population consists of digitized literary corpora from major archives (British Library, Bibliothèque nationale de France, University of Chicago Library) and regional repositories. A stratified random sample of 1,200 texts is drawn to balance geographic representation and publication period. Data collection instruments include (a) a customized text-mining pipeline using Python with spaCy for named-entity recognition and Gensim for topic modeling, (b) metadata augmentation through library catalog records and OCR quality checks, and (c) a coding schema for intertextual and thematic linkages informed by Postcolonial Theory (e.g., Edward Said, Homi Bhabha) and Network Theory (e.g., Freeman, Barabási–Albert). Validity and reliability are enhanced through triangulation across three data streams (textual content, metadata, and citation/intertextual proxies) and inter-coder reliability tests on a 10% subset of annotated links. Data analysis proceeds in multiple stages (i) preprocessing and normalization of texts, (ii) unsupervised topic modeling (LDA and BERTopic) to extract themes, (iii) stylometric profiling (word n-grams, function-word usage) to infer authorial similarity, (iv) construction of a multiplex network with layers for co-citation, intertextual allusions, thematic affinity, and publishing-ecosystem proximity, (v) network analytics including degree centrality, betweenness, eigenvector centrality, modularity-based community detection, and assortativity by region and era, and (vi) regression-based significance testing to assess the relationship between centrality measures and thematic convergence. The analytical framework draws on postcolonial theory to interpret the significance of discovered communities and cross-regional linkages, and on social-network theory to explain emergent structures as evidence of transnational literary exchange. Expected findings include the identification of previously underappreciated transregional networks linking authors across the Atlantic, Indian Ocean, and Caribbean littératures, and the demonstration that thematic convergence and intertextuality migrate along publishing ecologies and editorial networks. The study anticipates demonstrating how central nodes function as brokers of cross-cultural dialogue, and how peripheral clusters reveal latent diasporic circuits. The contribution to knowledge lies in delivering a rigorous, scalable DH pipeline that combines topic modeling, stylometry, and multiplex network analysis to illuminate postcolonial literary networks, providing empirical substantiation of theoretical claims regarding transnational influence and circulation. The study concludes with recommendations for repository metadata standardization, a publicly accessible replicable codebase, and guidelines for applying similar DH-network approaches to other literatures, with potential policy implications for digital philology curricula and grant-funding priorities in humanities data science.
Thesis Overview
Digital Humanities Text Mining for Postcolonial Literary Networks explains how contemporary scholars can use computer-assisted methods to study how postcolonial writers, themes, and influences are connected across time and space. The core idea is that large collections of literary texts, letters, reviews, and translation records contain hidden networks—authors who influenced one another, shared motifs, or circulated ideas. Text mining and network analysis help reveal these patterns more systematically than traditional close reading alone.
Why it matters: Postcolonial studies often focus on individual authors or national literatures, but literature also forms dynamic networks shaped by colonial histories, migrations, publishers, and editorial practices. By mapping textual connections at scale, researchers can identify clusters of influence, trace how certain themes migrate between regions, and detect overlooked connections that merit closer reading.
What problem or knowledge gap it addresses: There is a need for rigorous, data-driven mapping of intertextual relationships and influence networks in postcolonial literatures. Prior work frequently relies on qualitative, anecdotal evidence. This project combines quantitative text analysis with qualitative interpretation to produce a more comprehensive view of literary networks while remaining anchored in scholarly interpretation.
What the researcher will do step by step:
- Define a corpus: select 60–80 postcolonial literary works from diverse regions and time periods, including novels, poetry, and essays, plus critical and historical texts.
- Data collection: digitize texts where needed, clean OCR outputs, and compile metadata (author, year, place of publication, publisher).
- Text analysis: apply natural language processing to extract topics, named entities (authors, places, institutions), and stylistic features; compute similarity measures between texts.
- Network construction: build a bipartite network linking authors and texts, then project author- and theme-centered networks to identify clusters.
- Validation: triangulate with traditional literary scholarship, interview-based expert input, and cross-check with known bibliographies.
- Interpretation: analyze clusters and pathways of influence, track motif diffusion, and assess regional transmission patterns.
- Synthesis: integrate quantitative results with qualitative readings to tell coherent stories about postcolonial literary networks.
What contribution the study will make: it provides a replicable, transparent methodology for uncovering literary networks at scale, offers new empirical maps of influence and theme movement, and demonstrates how digital methods can complement traditional scholarship in postcolonial literary studies.
Expected outcome: a set of mapped networks with clearly documented clusters, an accompanying interpretive narrative linking data patterns to literary histories, and a methodological framework that can be adapted to other literatures.