Automated Diagnostics of Code-Switching in Multilingual Speech Interfaces
Table Of Contents
Chapter ONE
INTRODUCTION
- 1.1Introduction
- 1.2Background of the Study
- 1.3Statement of the Problem
- 1.4Aim and Objectives of the Study
- 1.5Research Questions
- 1.6Research Hypotheses
- 1.7Significance of the Study
- 1.8Scope and Delimitation of the Study
- 1.9Limitations of the Study
- 1.10Organisation of the Study
- 1.11Operational Definition of Terms
Chapter TWO
LITERATURE REVIEW
- 2.1Conceptual Review: Code-Switching in Multilingual Speech Interfaces
- 2.2Conceptual Review: Automated Diagnostics in Speech Technology
- 2.3Theoretical Framework: Dynamic Systems Theory in Speech Networking
- 2.4Theoretical Framework: Constructionist Theory of Language Use in Interfaces
- 2.5Empirical Review: Code-Switching Patterns in Voice Assistants
- 2.6Empirical Review: Linguistic Variability in Multilingual ASR Outputs
- 2.7Empirical Review: Error Analysis in Multilingual Speech Interfaces
- 2.8Empirical Review: User Experience and Accessibility in Multilingual Systems
- 2.9Gaps in Measurement and Evaluation of Code-Switching Diagnostics
- 2.10Gaps in Contextual Adaptation of Multilingual Interfaces
- 2.11Gaps in Real-Time Diagnostics and Latency Trade-offs
- 2.12Conceptual Model or Synthesis of Findings
- 2.13Summary of the Literature Review
Chapter THREE
RESEARCH METHODOLOGY
- 3.1Research Design: Mixed-Methods Framework for Diagnostic Modeling
- 3.2Philosophical Paradigm: Postpositivist Ontology with Pragmatic Epistemology
- 3.3Population of the Study: Multilingual Users and Multimodal Interfaces
- 3.4Sample Size and Sampling Technique: Stratified Sampling Across Language Pairs
- 3.5Sources and Instruments of Data Collection: Speech Corpora, User Logs, and Diagnostic Tools
- 3.6Validity and Reliability of Instruments
- 3.7Data Preprocessing and Annotation Protocols
- 3.8Model Specification: Diagnostic Metrics and Feature Engineering
- 3.9Data Analysis Methods: Statistical Testing and Machine Learning Diagnostics
- 3.10Ethical Considerations
- 3.11Pilot Study and Iterative Refinement
Chapter FOUR
DATA PRESENTATION AND ANALYSIS
- ANALYSIS AND DISCUSSION OF FINDINGS
- 4.1Data Presentation: Descriptive Overview of Multilingual Speech Samples
- 4.2Descriptive Analysis: Language Pair Frequencies and Switch Points
- 4.3Diagnostic Model Performance: Preprocessing and Feature Impact
- 4.4Hypotheses Testing: Relationship Between Code-Switching Characteristics and Interface Failures
- 4.5Interpretation of Results: Diagnostic Signals for Code-Switching Anomalies
- 4.6Findings in Relation to Conceptual Frameworks
- 4.7Comparative Analysis with Prior Empirical Studies
- 4.8Discussion of Practical Implications for Multilingual Speech Interfaces
Chapter FIVE
SUMMARY, CONCLUSION AND RECOMMENDATIONS
- CONCLUSION AND RECOMMENDATIONS
- 5.1Summary of Findings
- 5.2Conclusion
- 5.3Contribution to Knowledge
- 5.4Recommendations for System Design and Evaluation
- 5.5Suggestions for Further Studies
Thesis Abstract
Multilingual speech interfaces increasingly enable real-time communication across language communities, yet code-switching—shifting between languages within a single discourse—poses substantial challenges for automatic speech recognition (ASR), natural language understanding, and user experience. This study addresses the problem that existing interfaces exhibit degraded recognition accuracy and misinterpretation of user intents when code-switching is present, leading to higher error rates, user-frustration, and reduced accessibility for multilingual populations. The aim is to develop automated diagnostics that detect, characterize, and predict code-switching phenomena in multilingual speech interfaces and to evaluate how these diagnostics can inform adaptive systems that maintain performance and usability. The specific objectives are (1) to profile the prevalence and linguistic patterns of code-switching in bilingual and multilingual user interactions with conversational agents; (2) to develop a diagnostic framework combining acoustic-phonetic cues, lexical cues, and discourse-based features to identify code-switch points and their impact on ASR and intent classification; (3) to evaluate the diagnostic framework on a diverse corpus of multilingual speech data with explicit code-switching, comparing performance across languages, domains, and user demographics; (4) to integrate the diagnostics into a prototype adaptive speech interface that dynamically adjusts recognition models and dialogue management based on detected switching; and (5) to derive guidelines for designers on mitigating code-switching-related errors and improving user satisfaction. A mixed-methods design will be employed, incorporating corpus-based quantitative analysis and qualitative error analysis. The population consists of bilingual and multilingual speakers (n=120) drawing from two high-education urban centers with Turkish–English and Hindi–English bilingual communities, sampled to reflect gender, age, and proficiency variance. Data will comprise approximately 240 hours of spoken interactions collected through controlled task-based sessions (n=4 tasks per participant) and spontaneous usage in naturalistic settings (n=2 weeks of logged interactions). Instruments include a bespoke code-switching annotation schema integrating the Linguistic Inserting and Switching (LIS) framework with a probabilistic language identification module, acoustic feature extractors (e.g., MFCC trajectories, pitch, energy), and contextual markers (dialogue act tags, reaction times). ASR systems will be configured to operate with and without dynamic model adaptation to quantify the diagnostic system’s impact. Reliability will be established via inter-annotator agreement (Cohen’s kappa ? 0.80) and test–retest stability for feature extraction. Analytical techniques will encompass a sequence of approaches (i) generalized linear mixed models (GLMMs) to examine the relationship between code-switching frequency and ASR error rates across languages and contexts; (ii) survival analysis to model the time-to-next-switch in dialogue sequences; (iii) regression analyses to identify predictors of misclassification in intent recognition when code-switching occurs; (iv) Bayesian hierarchical modeling to estimate uncertainty in language identity and switching boundaries; (v) machine learning classifiers (random forest, gradient boosting, and transformer-based encoders) trained on combined acoustic-lexical-discourse features to detect code-switch points, assess their severity, and predict degradation in downstream tasks; and (vi) thematic analysis of qualitative error logs to uncover user experience implications. The theoretical framing integrates sociolinguistic code-switching theory (Gumperz) and the principle of adaptive user interfaces within human–computer interaction, complemented by the speech processing theory of acoustic-phonetic invariance in multilingual production. Expected findings include (a) robust diagnostic indicators that reliably signal code-switch onset and type (inter-sentential versus intra-sentential) and correlate with declines in ASR accuracy and intent classification, (b) quantifiable thresholds for when model adaptation significantly improves recognition performance in switch-rich segments, and (c) evidence that user perceived usability improves when interfaces proactively adjust language models and prompt for disambiguation at switch points. The study contributes to knowledge by operationalizing a reproducible, data-driven diagnostic framework for code-switching in multilingual interfaces, integrating linguistic, acoustic, and interactional signals, and providing actionable design guidelines for adaptive conversational systems. The main conclusion is expected to indicate that early, context-aware diagnostics combined with dynamic model selection can substantially mitigate code-switching-related errors, enhancing accessibility and user satisfaction. Recommendations will address dataset diversification, real-time diagnostic latency reduction, ethical considerations for bilingual data, and guidelines for scalable integration into commercial multilingual ASR and dialogue systems.
Thesis Overview
This research explores how people switch between languages within spoken interactions and how automated systems can diagnose and understand those switches in multilingual speech interfaces such as voice assistants or bilingual chatbots. It matters because real-world user voices often mix languages, and current interfaces struggle to recognize, interpret, or respond appropriately to such code-switching. Improving this capability can enhance user experience, accessibility, and accuracy in multilingual deployments.
The problem it addresses is the limited ability of speech interfaces to detect when a speaker switches languages, to identify the boundaries between language segments, and to determine the pragmatic or functional purpose of the switch (for example topic shift, emphasis, or social signaling). There is a gap in robust, scalable methods that combine linguistic insight with practical machine learning to diagnose code-switching in real time and to adapt system responses accordingly.
What the researcher will do, step by step:
- Define a clear scope of code-switching phenomena to study, including intra-sentential and inter-sentential switches in two or more languages common in the setting (e.g., English–Spanish).
- Collect data from multilingual speakers through controlled experiments and naturalistic recordings, aiming for a sample size of around 50–100 speakers and several hours of annotated speech.
- Build a labeled corpus with time-aligned codes indicating language boundaries, switch points, and the functional purpose of switches, using bilingual annotators and inter-annotator agreement measures.
- Develop an automated diagnostic framework combining acoustic-phonetic features (prosody, phoneme inventory shifts) with lexical and syntax cues, guided by relevant theories (e.g., Matrix Language Frame model, Optimality Theory for constraint rankings).
- Train and evaluate machine learning models (e.g., sequence labeling with conditional random fields, deep learning architectures such as BiLSTM-CRF, and transformer-based classifiers) to detect language boundaries and classify switch types.
- Validate the model's practical usefulness by integrating it with a prototype multilingual speech interface and assessing user satisfaction and error rates in simulated tasks.
- Analyze results with both quantitative metrics (precision, recall, F1, BLEU-like alignment scores) and qualitative insights from error analysis.
Expected contribution and outcome:
- A replicable methodology for diagnosing code-switching in real-time speech interfaces, along with a labeled benchmark corpus.
- An integrated diagnostic model that combines acoustic, lexical, and syntactic signals to detect switches accurately and infer their function.
- Practical guidelines for designers of multilingual interfaces to better handle code-switching, improving accessibility and user experience.
The study should yield actionable insights into how to design more robust, culturally aware multilingual speech systems.