Developing a Multimodal Dialogue System for Endangered Language Revitalization
Table Of Contents
Chapter ONE
INTRODUCTION
- 1.1Introduction
- 1.2Background of the Study
- 1.3Statement of the Problem
- 1.4Aim and Objectives of the Study
- 1.5Research Questions
- 1.6Research Hypotheses
- 1.7Significance of the Study
- 1.8Scope and Delimitation of the Study
- 1.9Limitations of the Study
- 1.10Organisation of the Study
- 1.11Operational Definition of Terms
Chapter TWO
LITERATURE REVIEW
- 2.1Conceptual Review: Multimodal Dialogue Systems and Language Revitalization
- 2.2Conceptual Review: Endangered Language Contexts and ICT Interventions
- 2.3Conceptual Review: Multimodal Interaction Modalities (Speech, Gesture, Visuals, Text)
- 2.4Theoretical Framework: Constructivist Learning Theory as a Basis for Multimodal Engagement
- 2.5Theoretical Framework: Activity Theory in Technology-Macred Settings for Language Revitalization
- 2.6Empirical Review: Prior Multimodal Dialogue Systems in Low-Resource Languages
- 2.7Empirical Review: Language Documentation and Reclamation through ICT Tools
- 2.8Empirical Review: Speech Recognition for Low-Resource Languages
- 2.9Empirical Review: Multimodal Evaluation Metrics for Dialogue Systems
- 2.10Empirical Review: Endangered Language Communities and Technology Adoption
- 2.11Identified Gaps in the Literature
- 2.12Conceptual Model or Summary of the Review
Chapter THREE
RESEARCH METHODOLOGY
- 3.1Research Design
- 3.2Philosophical Paradigm
- 3.3Population of the Study
- 3.4Sample Size and Sampling Technique
- 3.5Sources and Instruments of Data Collection
- 3.6Validity and Reliability of Instruments
- 3.7Data Analysis Procedures
- 3.8Model Specification or Analytical Framework
- 3.9Multimodal Data Handling and Integration
- 3.10Ethical Considerations
Chapter FOUR
DATA PRESENTATION AND ANALYSIS
- ANALYSIS AND DISCUSSION OF FINDINGS
- 4.1Data Presentation: System Architecture and Data Flows
- 4.2Descriptive Analysis: User Interactions with the Multimodal System
- 4.3Hypotheses Testing: Effectiveness of Multimodal Cues on Language Recall
- 4.4Hypotheses Testing: User Engagement Across Modalities
- 4.5Interpretation of Results: Endangered Language Revitalization Indicators
- 4.6Discussion of Findings in Relation to Conceptual Review
- 4.7Discussion of Findings in Relation to Theoretical Frameworks
- 4.8Discussion of Findings in Relation to Empirical Literature
Chapter FIVE
SUMMARY, CONCLUSION AND RECOMMENDATIONS
- CONCLUSION AND RECOMMENDATIONS
- 5.1Summary of Findings
- 5.2Conclusion
- 5.3Contribution to Knowledge
- 5.4Practical Implications for Language Revitalization Programs
- 5.5Recommendations for Practice and Policy
- 5.6Suggestions for Further Studies
Thesis Abstract
This study investigates the development of a multimodal dialogue system designed to support the revitalization of an endangered language by enabling naturalistic spoken, gesture, and visual interactions within community-centric educational and digital archival contexts. The core problem addressed is the erosion of productive language transmission in everyday contexts, where traditional literacy- and media-based resources fail to capture the full socio-communicative fabric of the language. The aim is to design, implement, and evaluate a multimodal dialogue system that integrates speech, gesture, and gesture-enabled visual cues with culturally informed lexical and syntactic models to facilitate learner-centered practice, authentic discourse simulation, and intergenerational transmission. Specific objectives include (1) to develop an annotated multimodal corpus capturing spontaneous speech, co-speech gestures, facial expressions, and contextual visual scenes in language-use scenarios; (2) to construct a scalable dialogue management framework incorporating end-user feedback loops and a rule-based plus neural hybrid approach that respects community norms and language ideologies; (3) to evaluate linguistic accessibility, usability, and learning outcomes across diverse user groups; and (4) to assess the system’s impact on active language production and cultural knowledge retention over a 12-week pilot deployment. The methodology adopts a mixed-methods research design, combining corpus-based linguistic analysis, human–computer interaction (HCI) experimentation, and evaluative heuristics grounded in sociolinguistic theory. The study will be conducted in a language documentation setting within a community of roughly 1,200 fluent or semi-fluent speakers, with a stratified sample of 60 participants spanning ages 12–65, including 20 heritage learners and 20 adult language learners. Data collection instruments comprise (a) a multimodal corpus elicitation protocol using scripted tasks and naturalistic interaction sessions recorded via high-fidelity audio, 4K video, and depth-sensing cameras; (b) a dialogue system prototype deployed on tablets and smartphones with offline and cloud-based processing capabilities; (c) standardized usability scales (System Usability Scale, User Experience Questionnaire) and retention tests; (d) semi-structured interviews and focus groups to capture sociocultural dimensions. Validity and reliability are ensured through inter-rater reliability checks (Cohen’s kappa ? 0.80) for annotation, pilot testing of the corpus annotation schema, and triangulation across linguistic, usability, and learning outcomes. Analytical procedures include (i) thematic analysis of qualitative data to identify language-ideology considerations and user needs; (ii) variational analysis of corpus annotations to extract lexical, syntactic, and gestural alignment patterns; (iii) quantitative evaluation of learning gains using paired-sample t-tests and repeated-measures ANOVA to compare pre- and post-deployment language proficiency; (iv) mixed-effects modeling to examine factors influencing system usage and language production outcomes; (v) regression analyses to determine predictors of sustained engagement. The theoretical backbone integrates Activity Theory and Sociocultural Theory to frame user-system interactions as mediated by tools and community practices, while applying Endangered Language Documentation and Technology Acceptance models to interpret adoption dynamics. A conceptual model linking multimodal inputs, dialogue strategies, learner outcomes, and cultural reinforcement is developed and iteratively refined through pilot data. Expected findings include (a) enhanced lexical access and grammatical production in the endangered language facilitated by synchronized speech-gesture cues; (b) higher retention of cultural concepts when multimodal prompts align with community discourse practices; (c) positive usability scores and increased sustained engagement among younger users compared with constraints in older participants, mitigated through adaptive interface configurations. The study contributes new knowledge by (i) presenting an empirically validated multimodal dialogue architecture tailored for endangered language contexts, (ii) offering an annotated, publicly shareable multimodal corpus with rich metadata, and (iii) providing a scalable evaluative framework for ICT-driven language revitalization initiatives. The principal conclusion posits that culturally attuned multimodal dialogue systems can significantly augment active language production and intergenerational transmission. Recommendations include expanding corpus diversity, integrating community-led governance for content curation, prolonging longitudinal assessments beyond 12 weeks, and exploring cross-language transferability to other endangered languages with similar sociocultural ecosystems.
Thesis Overview
This research explores how to build a multimodal dialogue system to support the revitalization of endangered languages. A multimodal dialogue system combines text, speech, gestures, and visual cues to enable natural, interactive conversation with a computer or mobile device. The aim is to create a tool that helps language learners, community members, and educators practice speaking, listening, and cultural phrases in contexts that mimic real conversations.
Why it matters: many endangered languages lack scalable, engaging ways for communities to use the language in daily life. Traditional resources (dictionaries, textbooks) are static, while audio-only or text-based tools miss the richness of real interaction. A multimodal system can provide immersive opportunities for language use, preserve pronunciation, structure, and cultural practices, and support intergenerational transmission.
Problem or knowledge gap: while there are conversational agents and reusable speech interfaces, few systems are designed specifically for endangered languages with limited data, complex sociolinguistic norms, and the need for community-controlled content. This project addresses how to design, implement, and evaluate a dialogue system that can operate with low-resource language data and still offer meaningful, culturally appropriate interactions.
What the researcher will do step by step:
1. conduct a literature review on multimodal dialogue systems, low-resource NLP, and language revitalization needs.
2. select a target endangered language with active community consent and collect a small, ethically sourced dataset, including transcripts, audio, video gestures, and cultural terms (approximately 50-100 hours of interaction data if feasible, plus elicitation sessions).
3. design an architecture that integrates text, speech, and visual modalities, using transfer learning and data augmentation to cope with limited data.
4. implement a prototype dialogue system with a user interface for learners and community speakers.
5. evaluate usability and effectiveness with a mixed-methods study: quantitative measures (task success rate, pronunciation accuracy, user satisfaction) and qualitative feedback (think-aloud sessions, interviews).
6. analyze data using thematic analysis for qualitative inputs and statistical methods (regression or ANOVA) for quantitative results.
7. refine the model based on findings and provide guidelines for community-centered expansion.
Contributions and expected outcomes: the study will demonstrate a feasible approach for building multimodal conversational tools for endangered languages with limited data, offering design principles, an evaluative framework, and a prototype that communities can adapt. It is expected to improve learner engagement, increase practical language use, and support sustainable language revitalization efforts.
If successful, the study will recommend scalable pathways for deploying similar systems across multiple endangered languages, emphasize community governance of content, and propose standards for evaluating multimodal resources in revitalization contexts.