Automated Multimodal Discourse Analysis for Low-Resource Languages via AI | Blazingprojects Postgraduate Thesis
Home / Linguistics / Automated Multimodal Discourse Analysis for Low-Resource Languages via AI

Automated Multimodal Discourse Analysis for Low-Resource Languages via AI

 

Table Of Contents


Chapter ONE

INTRODUCTION

  • 1.1Introduction
  • 1.2Background of the Study
  • 1.3Statement of the Problem
  • 1.4Aim and Objectives of the Study
  • 1.5Research Questions
  • 1.6Research Hypotheses
  • 1.7Significance of the Study
  • 1.8Scope and Delimitation of the Study
  • 1.9Limitations of the Study
  • 1.10Organisation of the Study
  • 1.11Operational Definition of Terms

Chapter TWO

LITERATURE REVIEW

  • 2.1Conceptual Review: Multimodal Discourse in Low-Resource Languages
  • 2.2Conceptual Review: Automated Analysis in AI Systems for Language Documentation
  • 2.3Theoretical Framework: Multimodal Interaction Theory in Language Technologies
  • 2.4Theoretical Framework: Social Semiotics and AI-Assisted Discourse Analysis
  • 2.5Empirical Review: Prior Studies on Multimodal Analytics in Resource-Limited Contexts
  • 2.6Empirical Review: Speech, Video, and Gesture Integration in Low-Resource Languages
  • 2.7Empirical Review: Transfer Learning for Multimodal NLP in Under-Resourced Languages
  • 2.8Empirical Review: Ethics, Accessibility, and Community Involvement in Language Tech
  • 2.9Identified Gaps in the Literature: Methodological and Data Gaps
  • 2.10Identified Gaps in the Literature: Evaluation and Reproducibility Gaps
  • 2.11Conceptual Model: Integrated Multimodal Discourse Analysis for Low-Resource Languages
  • 2.12Summary of the Literature Review

Chapter THREE

RESEARCH METHODOLOGY

  • 3.1Research Design: Iterative Prototyping of a Multimodal Analysis Pipeline
  • 3.2Philosophical Paradigm: Pragmatism in ICT-Driven Linguistic Research
  • 3.3Population of the Study: Targeted Low-Resource Language Communities
  • 3.4Sample Size and Sampling Technique: Stratified and Snowball Sampling
  • 3.5Sources and Instruments of Data Collection: Audio-Visual Datasets, Field Notes, and Annotations
  • 3.6Data Collection Procedures: Corpus Curation and Annotation Protocols
  • 3.7Validation and Reliability of Instruments: Inter-Annotator Agreement and Test-Retest Checks
  • 3.8Data Processing and Preprocessing Methods
  • 3.9Method of Data Analysis: Multimodal Machine Learning and Statistical Testing
  • 3.10Model Specification: Architecture for Multimodal Discourse Extraction
  • 3.11Evaluation Metrics and Validation Framework
  • 3.12Ethical Considerations: Informed Consent, Community Benefit, and Data Sovereignty

Chapter FOUR

DATA PRESENTATION AND ANALYSIS

  • ANALYSIS AND DISCUSSION OF FINDINGS
  • 4.1Data Presentation: Overview of Collected Multimodal Datasets
  • 4.2Descriptive Analysis: Language Diversity and Modality Usage in the Data
  • 4.3Hypotheses Testing: Effects of Modality Fusion on Discourse Cohesion
  • 4.4Hypotheses Testing: Cross-Language Transferability of Models
  • 4.5Interpretation of Results: Implications for Low-Resource Language Communities
  • 4.6Discussion: Alignment with Theoretical Frameworks
  • 4.7Discussion: Gaps Addressed by the Proposed Framework
  • 4.8Limitations in Findings and Potential Biases

Chapter FIVE

SUMMARY, CONCLUSION AND RECOMMENDATIONS

  • CONCLUSION AND RECOMMENDATIONS
  • 5.1Summary of Findings
  • 5.2Conclusion: Efficacy of Automated Multimodal Discourse Analysis for Low-Resource Languages
  • 5.3Contribution to Knowledge: Methodological and Practical Impact
  • 5.4Recommendations: For Researchers, Practitioners, and Communities
  • 5.5Suggestions for Further Studies

Thesis Abstract

This study addresses the persistent challenge of automatic understanding and analysis of discourse in low-resource languages, where multimodal cues (speech, gesture, facial expression, and visual context) are underrepresented in available corpora and computational models. The aim is to develop and validate a robust multimodal discourse analysis framework that leverages AI to fuse audiovisual, textual, and contextual signals for accurate interpretation of discourse structure and pragmatic function in low-resource language communities. Specific objectives include (1) constructing an annotated multimodal corpus for three representative low-resource languages (Language A, Language B, Language C) with 600 hours of multimedia interaction data, (2) designing a cross-modal representation model that integrates speech transcripts, gesture dynamics, facial action coding, and scene context using a transformer-based architecture with modality-specific encoders, (3) developing discourse-level annotation schemes informed by Systemic Functional Grammar and Speech Act Theory, and (4) evaluating the framework against baseline unimodal and hybrid models to quantify gains in discourse recognition, polarity, stance, and illocutionary force. The methodology adopts a mixed-methods research design combining data-driven AI development with qualitative validation. The population consists of publicly available broadcast and conversational data within the target language communities, supplemented by elicitation sessions with 40 native speakers per language to obtain high-quality transcripts and pragmatic annotations. A stratified sampling approach yields 600 hours of audiovisual material, divided into 300 hours for Language A, 180 hours for Language B, and 120 hours for Language C. Multimodal data collection instruments include synchronized audio recorders, high-definition video capture, and expert-annotated ground-truth labels for discourse functions (e.g., assertive, directive, expressive), discourse markers, and context labels. To ensure reliability, inter-annotator agreement is established using Cohen’s kappa and Krippendorff’s alpha on a subset of 15% of the data. The analysis employs a multimodal deep learning framework modality-specific encoders for speech (wav2vec 2.0 fine-tuned on each language), visual features (3D-CNN for gesture and facial expression, ResNet-50 for scene attributes), and text (language-specific BERT variants). These encoders feed into a cross-modal transformer with joint attention to produce discourse-level predictions, including discourse relation classification, speech act type, and stance detection. The annotation scheme aligns with Systemic Functional Grammar for semantic roles and functional categories, complemented by Gricean maxims for pragmatic inference. Model training uses cross-entropy objectives for token- and utterance-level tasks, with auxiliary losses for alignment consistency across modalities. Evaluations compare the proposed model against baselines unimodal (speech-only, vision-only) and simple late-fusion approaches, using metrics such as macro-averaged F1, BLEU for transcript quality, and Matthews correlation coefficient for discourse labels. The study also conducts ablation analyses to determine the contribution of each modality, and ablation experiments to assess cross-language generalization via zero-shot transfer. Expected findings indicate that the integrated multimodal framework outperforms unimodal and baseline fusion models across languages, with statistically significant improvements in discourse relation accuracy (?F1 ? 12–18%), speech-act classification (?F1 ? 10–15%), and stance detection (?F1 ? 8–12%). Cross-language experiments reveal that shared representations for gesture-speech contingencies and contextual cues enhance performance in Language B and Language C when Language A data is abundant, validating the model’s capacity for transfer learning in resource-scarce settings. The research is anticipated to produce a publicly available, richly annotated multimodal corpus and an open-source framework to facilitate reproducibility and further research in low-resource discourse analysis. The study contributes to knowledge by demonstrating a scalable, cross-modal approach to discourse analysis that integrates linguistic theory with state-of-the-art AI to address resource constraints in multilingual contexts. It advances methodological practices in corpus construction, annotation schemas, and evaluation protocols for multimodal NLP in low-resource languages, and provides practical tools for applications in education, media analytics, and human–computer interaction. The main conclusion posits that multimodal fusion grounded in linguistic theory significantly enhances discourse interpretation in low-resource languages, and recommends further exploration of active learning to expand annotated data, domain adaptation for diverse communicative contexts, and ethical considerations in deploying language technologies within minority communities.

Thesis Overview

Automated Multimodal Discourse Analysis for Low-Resource Languages via AI is a research topic that combines how people communicate with multiple modalities (spoken language, gestures, facial expressions, images, and context) and how to automatically analyze those signals in languages with limited data resources. The study addresses the gap that most multimodal analysis methods rely on large, well-annotated corpora in well-documented languages, leaving low-resource languages underserved. By leveraging AI techniques, the project aims to create scalable methods that can learn from limited data and still provide reliable insights into how discourse is constructed and understood across modalities. What the research is about - Investigating how audio, video, text, and contextual cues work together in natural communication for languages with scarce annotated resources. - Developing computational tools that can align multimodal signals, detect discourse structures, and extract meaningful patterns such as stance, topic shifts, and interlocutor roles. Why it matters - Improves linguistic description and documentation of low-resource languages by capturing multimodal cues that are often essential for meaning. - Enables more inclusive natural language technologies (NLP, language learning apps, subtitling) for diverse languages and communities. - Advances theoretical understanding of how multimodal resources interact in real-world discourse, informing theories of communication and language contact. What the researcher will do (step by step) 1. Define target low-resource languages and assemble a corpus comprising video-recorded conversations, accompanying transcripts, and metadata. 2. Collect data from fieldwork or open-source sources, ensuring ethical approvals and participant consent. 3. Annotate a subset of the data for key discourse phenomena (e.g., stance, topics, turn-taking) and use transfer learning to extend annotations to the larger set. 4. Build multimodal representations that fuse audio features (prosody, speech rate), visual cues (gaze, gestures), and textual content (transcripts) using neural architectures designed for limited data (few-shot or semi-supervised learning). 5. Develop alignment and segmentation methods to detect discourse units and structure, and apply classification or sequence labeling to identify discourse functions. 6. Validate models against human judgments and compare performance across languages to assess generalizability. 7. Analyze results with statistical methods (e.g., regression analyses to relate multimodal features to discourse outcomes) and conduct error analysis to identify gaps. 8. Discuss theoretical implications for multimodal discourse theory and provide practical recommendations for language technology applications. What contribution and outcomes to expect - A reproducible multimodal analysis framework tailored to low-resource languages, including data collection protocols and open-source tools. - Insights into which modalities most strongly influence discourse interpretation in low-resource contexts. - Publications detailing methods, evaluation results, and practical guidelines for researchers and developers working with underrepresented languages.

Blazingprojects Mobile App

📚 Over 50,000 Research Thesis
📱 100% Offline: No internet needed
📝 Over 98 Departments
🔍 Thesis-to-Journal Publication
🎓 Undergraduate/Postgraduate Thesis
📥 Instant Whatsapp/Email Delivery

Blazingprojects App

Related Research

Mechanical engineeri. 4 min read

Intelligent Predictive Maintenance System for Turbomachinery Networks via Edge AI...

This research investigates how a smart maintenance system can forecast failures and optimize upkeep for turbomachinery networks using edge artificial intelligen...

BP
Blazingprojects
Read more →
Mathematics. 3 min read

Efficient Graph Neural Networks for Real-Time IoT Anomaly Detection...

Efficient Graph Neural Networks for Real-Time IoT Anomaly Detection is about using advanced machine learning to monitor networks of Internet of Things devices a...

BP
Blazingprojects
Read more →
Materials and Metall. 4 min read

Smart Nanocomposite Coatings via AI-Driven In-Situ Sputtering Optimization...

Smart Nanocomposite Coatings via AI-Driven In-Situ Sputtering Optimization aims to develop protective and functional coatings by combining nanomaterials with me...

BP
Blazingprojects
Read more →
Mass communication. 2 min read

AI-powered Fact-Checking for Local News Credibility Systems...

AI-powered Fact-Checking for Local News Credibility Systems is about building and evaluating automated tools that help determine whether local news items are ac...

BP
Blazingprojects
Read more →
Marketing. 4 min read

Personalized AI-Driven Brand Loyalty Prediction for E-Commerce ...

Personalized AI-Driven Brand Loyalty Prediction for E-Commerce investigates how artificial intelligence can forecast and enhance customer loyalty for online ret...

BP
Blazingprojects
Read more →
Linguistics. 2 min read

Automated Multimodal Discourse Analysis for Low-Resource Languages via AI...

Automated Multimodal Discourse Analysis for Low-Resource Languages via AI is a research topic that combines how people communicate with multiple modalities (spo...

BP
Blazingprojects
Read more →
Library Science Educ. 4 min read

Designing AI-Powered Reference Literacy Labs for Library Education...

This research explores how intelligent, AI-powered reference labs can transform library education by teaching students and future librarians to efficiently loca...

BP
Blazingprojects
Read more →
Library and informat. 2 min read

AI-driven Discovery Systems for Multilingual Academic Libraries...

AI-driven Discovery Systems for Multilingual Academic Libraries What the research is about This topic investigates how to design and evaluate an information di...

BP
Blazingprojects
Read more →
Law. 3 min read

AI-enabled Compliance and Data Privacy Governance for SMEs ...

This research focuses on how small and medium-sized enterprises (SMEs) can use artificial intelligence to meet regulatory requirements around data privacy and c...

BP
Blazingprojects
Read more →
WhatsApp Click here to chat with us