Evaluating AI-Based Translation for Endangered Language Documentation

A Native-Speaker–Validated Study of Chakma–English Translation · Interactive Manuscript

Open Dataset Explorer (edit live data)
Evaluating AI-Based Translation for Endangered Language Documentation
A Native-Speaker–Validated Study of Chakma–English Translation
This interactive manuscript viewer walks you through a peer-reviewed study evaluating ChatGPT's ability to translate the Chakma language — an endangered indigenous language of Bangladesh — into English. Start here to understand the paper's structure, key findings, and how to navigate each tab.
Fatema Tuj Jannat · Nashrah Sharfuddin · Montina Dewan · Mahadi Hasan  ·  2026
1,691
Sentences Evaluated
23.5%
AI Accuracy Rate
7.303
MQM Score (avg)
1,432
Errors Annotated
72.5%
Critical Severity
87.4%
Bi-LSTM Accuracy
"ChatGPT correctly translated only 1 in 4 Chakma sentences. The remaining 3 in 4 contained at least one error, and in 72.5% of those cases the AI produced a critically wrong translation — not a grammar slip, but a sentence that means something completely different from the original." — Core finding, Jannat et al. (2026) · MQM Score: 7.303 / Very Poor (≥7.00)
🔬 How the Research Was Done — Step by Step
Step 1 · Corpus
We built a Chakma–English parallel corpus of 1,691 sentences
Everyday Chakma speech — greetings, questions, requests, proverbs, kinship terms — was written in Chakma Unicode script with Bengali-script pronunciation guides. Each sentence was translated into English by a bilingual researcher, then reviewed and corrected by native Chakma speakers. This validated translation became the gold standard.
📂 See §3.2 Corpus Development
Step 2 · AI Translation
We fed all 1,691 sentences to ChatGPT
Each Chakma sentence was submitted with the instruction: "Translate the following Chakma sentence into English." The raw output was recorded exactly as returned — no prompt engineering, no post-editing. This is the baseline that a researcher or practitioner would get using the tool as-is.
📂 See §3.3 AI Translation System
Step 3 · MQM Annotation
Every AI output was compared against the human reference, sentence by sentence
We used the Multidimensional Quality Metrics (MQM) framework: every error was classified by Dimension (Accuracy / Fluency / Locale Convention), Category (Mistranslation, Grammar, Cultural Appropriateness, etc.), and Severity (Critical = 10 pts · Major = 5 pts · Minor = 1 pt). The penalty points were summed per sentence, then averaged across the corpus.
📐 See §3.4 MQM Framework · See MQM Framework tab
Step 4 · Native Speaker Validation
Native Chakma speakers reviewed all annotations
To ensure the error judgements were culturally and linguistically accurate, native Chakma speakers validated every annotation — checking that identified errors were real, that the severity was appropriate, and that culturally specific meanings were correctly interpreted. This step is what distinguishes the study from purely computational evaluation.
📂 See §3.5 Annotation Process · §3.7 Ethical Considerations
Step 5 · Bi-LSTM-CRF (Supporting)
We used a sequence model to identify which Chakma structures trigger AI errors
A Bi-LSTM-CRF model was applied as a diagnostic tool — not to translate, but to identify the grammatical patterns in Chakma sentences that the AI consistently fails to parse: progressive aspect markers, future tense suffixes, interrogative particles, negation morphology, and evidential markers. It achieved 87.4% grammatical pattern recognition accuracy.
🔴 Supporting analysis · See §3.6 · Not in submitted manuscript
Step 6 · Analysis
Error counts, MQM scores, and patterns were interpreted against the literature
Quantitative results (error frequency, penalty distribution, quality band) were combined with qualitative analysis of representative examples to answer both research questions — how accurately does AI translate Chakma (RQ1), and what types of errors recur (RQ2)?
📂 See §4 Findings · §5 Discussion
📖 What You'll Find in Each Section
Abstract
The 200-word summary of the entire paper
Research questions, corpus size (1,691 sentences), MQM framework, core result (7.303Very Poor), and implications for endangered language documentation.
→ Open tab
§1 · Introduction
Why Chakma? Why AI? Why now?
Sets the global context of language endangerment (7,000 languages, 40% heading to extinction), the specific threat to Chakma in the Chittagong Hill Tracts, and the case for evaluating AI translation as a documentation tool. Ends with RQ1 and RQ2.
→ Open tab
§2 · Literature Review
What we built on — four bodies of research
2.1 Chakma language endangerment · 2.2 MT evolution (RBMT → NMT → LLM) · 2.3 AI for language preservation · 2.4 MQM and translation quality evaluation. Together they justify why MQM with native-speaker validation is the right method.
→ Open tab
§3 · Methodology
How every decision was made and why
3.1 Mixed-methods design · 3.2 1,691-sentence corpus · 3.3 ChatGPT, zero post-edit · 3.4 Full MQM framework (3 dimensions, 9+ categories, 3 severity levels) · 3.5 6-stage annotation pipeline · 3.6 Bi-LSTM-CRF diagnostic · 3.7 Ethics · 3.8 Rigor.
→ Open tab
§4 · Findings
The numbers and the examples
RQ1: Accuracy 0.235, Fluency 0.814, Cultural 0.953 vs Human 0.90 each. MQM = 7.303.
RQ2: Mistranslation (72.5%), Grammar (22%), Cultural (5.5%). Real database examples for every error sub-type: Terminology, Mistranslation, Untranslated, Unintelligible, Omission, Addition, Word Order, Tense, Sentence Category, Singular/Plural, Cultural.
→ Open tab
§5 · Discussion
What it means and why it happened
5.1 Overall performance vs low-resource MT literature · 5.2 Why lexical errors dominate · 5.3 Chakma morphology vs English grammar · 5.4 Cultural knowledge gaps · 5.5 Bi-LSTM-CRF patterns · 5.6 Resource scarcity as root cause and linguistic justice framing.
→ Open tab
§6 · Conclusion
What this means for Chakma and for the field
Synthesises findings, states the 7.303 Very Poor quality verdict clearly, and frames it as a linguistic justice issue — not just a technical one. 6.1 Three study limitations. 6.2 Implications: what AI developers and documentation practitioners should do next.
→ Open tab
References
49 cited works — complete APA list
Full reference list including Ranathunga et al. (2023) on low-resource NMT, Dovchin (2020) on linguistic racism, Rousan et al. (2025) on ChatGPT vs human translation, and MQM (2015). All sources cited in the paper appear here.
→ Open tab
📊 The Three Main Findings at a Glance
🔴 Finding 1 — Quality The corpus-level MQM score of 7.303 places AI translation of Chakma in the Very Poor band (≥7.00). The AI cannot be used for documentation without full human correction. Only 398 of 1,691 sentences (23.5%) were correct.
🟡 Finding 2 — Error Pattern 72.5% of all errors are Critical Accuracy errors — the AI output a completely different meaning, not just a grammar mistake. These are concentrated in Chakma words whose surface form the AI confuses with semantically unrelated English approximations.
🟢 Finding 3 — Root Cause The failures trace to Chakma's near-total absence from LLM training data — a structural inequality, not a model flaw. Fluency (81.4%) is much higher than Accuracy (23.5%), confirming the AI can produce grammatical English — it just doesn't know what the Chakma means.
🗂 What Each Tab Does — Viewer Guide
📐 MQM Framework
Live charts and visualisations of the MQM scoring — dimension donut, severity bars, penalty distribution column chart, quality band scale. Updates every 2 seconds from the annotation database. Best for presenting the data visually.
📊 Score Analysis
Step-by-step walkthrough of how MQM = 7.303 was calculated, four real case studies from the database (perfect / major error / critical error / multiple errors), comparative analysis tables, and key findings summary.
📋 Tables
All publication-ready tables: 10 summary tables (corpus stats, dimension distribution, severity, penalty bands, quality bands, top error categories, outcome averages) plus all 12 example tables from the paper (Tables A–L), each with Chakma script, romanisation, human and AI translation side by side.
🔍 Missing & Future
Gaps in the current study, planned extensions (multi-system comparison, larger corpus, Bi-LSTM-CRF full evaluation), and future work directions including error prediction pipeline integration.
📚 References
Complete APA reference list — 49 works. All citations used across the paper, from Ranathunga et al. (2023) to Dovchin (2020) to MQM (2015).
📝 Edit History
Every change made in Edit Mode is logged here with: which tab, which section, the passage context, before/after content, and timestamp. Revert any change, jump back to that section, or copy original text.
📄 Full Paper
All academic sections assembled into one continuous journal-formatted document. Click ⬇ Download to save as a standalone HTML file — open in browser, print to PDF for submission. Sections marked 🔴 are not in the submitted manuscript.
✏️ How to Edit Any Section
Click the ✏️ Edit Mode button (bottom-right corner) to enable editing across all tabs.
Click any text — paragraph, table cell, heading, list item — to edit it directly
Select a portion of text (even mid-sentence), then right-click for the context menu: Edit selected portion, Bold, Italic, Mark red, Undo, Copy
Floating toolbar appears on any selection for quick Bold / Italic / 🔴 Red / ✓ Done
• Every change is saved in the 📝 Edit History tab with full context, before/after, and a Revert button
This interactive viewer was built to support the review, editing, and presentation of Jannat et al. (2026).
All 🔴 red-tagged sections are supplementary — not in the submitted manuscript.

Fatema Tuj Jannat¹·⁴ | Nashrah Sharfuddin² | Montina Dewan³ | Mahadi Hasan⁴

¹·⁴Northern University Bangladesh | ²BRAC University | ³Tebtebba Foundation Bangladesh (IP Project)

Nashrah Sharfuddin (sharfuddinnashrah@gmail.com)

Keywords: Chakma language | indigenous language preservation | linguistic justice | AI-based translation | Multidimensional Quality Metrics (MQM)

Abstract

Language endangerment has become increasingly widespread worldwide, with 14 indigenous languages being on the verge of extinction in the land of Bangladesh alone (Chiran, 2025). The Chakma language, spoken by the Chakma community that resides in Chittagong Hill Tracts (CHT), is one such 'definitely endangered' language (Saikia & Ullman, 2023) that faces ongoing threat due to socio-political marginalization. Preserving the Chakma language is essential for sustaining linguistic diversity, cultural practices, and indigenous knowledge systems embedded within the language. Inspired by the potential of AI translation tools, the study investigates the extent to which such technologies can support endangered language documentation. Specifically, the study asks: (1) how accurately do AI-based translation systems translate Chakma texts into English when compared with native-speaker–validated human translations, and (2) what types of lexical, grammatical, and culturally grounded errors recur in AI-generated translations? To address these questions, a human-validated reference corpus of 1,691 Chakma–English sentence pairs is compiled, and AI-generated translations are systematically evaluated using the Multidimensional Quality Metrics (MQM) framework across three quality dimensions—accuracy, fluency, and cultural appropriateness—in conjunction with an error analysis approach validated by native Chakma speakers. Of the 1,691 sentences that were fully annotated, the AI system produced an acceptable translation for 398 sentences (23.5%), while 1,432 error occurrences were identified in total—dominated by lexical/meaning errors (72.5%) and grammatical errors (22.0%)—yielding a weighted MQM error penalty score of 7.303 penalty points per sentence, which places the corpus in the Very Poor quality band (≥7.00) band. To support the interpretation of translation outputs, the study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors. Since large language models rely on attention-based mechanisms trained predominantly on high-resource languages, they remain ill-equipped for Chakma's morphological complexity. The Bi-LSTM-CRF effectively addresses this gap, achieving a grammatical pattern recognition accuracy of 87.4%. By foregrounding native-speaker validation, linguistic analysis, and systematic evaluation of AI-generated translations, the study demonstrates how AI translation tools can be critically assessed and responsibly operationalized as equitable resources for endangered language documentation, while laying the groundwork for more effective AI-supported approaches to Chakma language preservation.

The abstract presents the full scope of the paper: the endangered-language context, the two research questions, the corpus size (1,691 sentences), the AI accuracy rate (23.5%), and the overall MQM penalty score of 7.303 (Very Poor).
1. Introduction

The world hosts a vast array of languages essential to humanity's heritage (Drude, 2003). Currently, over 7,000 languages are spoken across the globe; however, this remarkable linguistic richness confronts substantial risks as modernity advances (Hutson et al., 2024), causing 40% of the languages to head toward extinction (Eberhard et al., 2022, as cited in Li et al., 2024). According to the Language Conservancy, after global warming, language loss is recognized as the planet's most pressing crisis (Collette & Kennedy, 2023, as cited in Hutson et al., 2024). Numerous endangered languages are diminishing rapidly due to globalization and modernization. This decline is particularly concerning because linguistic diversity is essential for transmitting culture, values, beliefs, and history across generations (Sevinç, 2022). The unique words, phrases, and expressions of each language encapsulate the accumulated knowledge and experiences of its speakers, shaping identity and fostering a sense of pride (Jerome et al., 2022). For many indigenous communities, languages are key carriers of culture, containing unique communication systems, traditional knowledge, and a strong sense of identity (Gwerevende & Mthombeni, 2023). Therefore, the dramatic loss of these minority languages signifies more than silenced voices; it entails the epistemic erasure of invaluable cultural knowledge and distinct worldviews (Kandler & Unger, 2023).

The opening paragraph situates the study within the global language-endangerment crisis — 7,000+ languages, 40% at risk — establishing urgency before narrowing to Bangladesh and Chakma.

Within this broader global context, Bangladesh is no exception. The country, characterised by its rich ethnic diversity, is home to multiple indigenous communities, many of whose languages are increasingly at risk of decline, mainly due to globalization, urbanization, and the extensive use of politically dominant languages in social, educational, and professional spheres (Anik et al., 2025). The global decline of indigenous languages has reached a critical point, with studies indicating that around half of them could disappear within this century. The Kuruk language, for instance, is no longer spoken, while languages such as Pankho, Khumi, and Hajong are barely surviving (Mohsin, 2023). UNESCO (2003) identifies globalization, forced displacement, and assimilation-driven policies as the main forces behind this decline. Similarly, Chakma and Sultana (2023) argue that the loss of ancestral lands, environmental degradation, and the gradual erosion of cultural identities have further accelerated the decline of indigenous languages. Additionally, the increasing tendency to use native languages only in private settings gradually reduces speakers' fluency and intergenerational transmission, thereby heightening the risk of language extinction (Holmes, 2021).

This paragraph zooms into Bangladesh — 14 endangered indigenous languages, the fate of Kuruk, Pankho, Khumi, Hajong — grounding the global crisis in a specific national context before turning to Chakma.

Endangered languages like Chakma are facing both social and political marginalization reflecting the patterns of linguistic racism. In other words, minority voices are subdued while dominant languages are prioritized (Dovchin, 2020, 2025; Rosa & Flores, 2021). This broader pattern is clearly visible in the Chittagong Hill Tracts (CHT) of Bangladesh, where the Chakma people have been among the most affected by language policies. After independence, Bangladesh followed a 'one culture, one language' idea. This language policy reinforced the prominence of the native language, Bangla, while diminishing the visibility and status of minority languages and identities (Bal, 2010). Chakma and Sultana (2023) describe this as a form of language control, where indigenous people feel pressured to stop using their native languages. In light of this linguistic marginalization, translation acts as an important platform for resistance. Translation is not merely a technical process of linguistic transfer anymore. Instead, it is understood as a political and ethical act that can challenge the dominance of certain languages over others. Tymoczko (1999) asserts that translation has formed the cultural politics of colonized societies. It has helped communities preserve and negotiate their national and cultural identities. In colonial and postcolonial settings, translation is closely linked with power, representation, and cultural authority (Bassnett & Trivedi, 1999). Folaron (2015) argues that translation helps indigenous languages survive and gain recognition simply by increasing their exposure beyond their immediate communities. It is also perceived as an excellent strategy in reclaiming disadvantaged voices (Spivak, 1993). Through this lens, translating indigenous languages becomes more than documentation by being an act of resistance against linguistic marginalization and a way of affirming indigenous identity.

This paragraph introduces the linguistic justice framing — the 'one culture, one language' policy in Bangladesh, and how translation functions as political resistance rather than merely technical transfer.

Although Machine Translation (MT) research has advanced considerably through neural and large language model–based approaches, it continues to focus predominantly on high-resource language pairs supported by large-scale parallel corpora and extensive training data. Evaluation also frequently relies on automated metrics such as BLEU and TER, which often obscure crucial shortcomings by failing to capture deeper issues in translation quality. Al Sharou and Specia (2022) demonstrate that in low-resource and user-generated environments, the existence and severity of errors, particularly mistranslations, omissions, and hallucinations, are more consequential than fluency levels. While efforts are made in revitalizing languages around the globe, endangered languages, particularly in South Asia, such as the Chakma language, remain largely underexplored in NLP research (Chakma et al., 2024). Specifically, there is a notable lack of empirical research on Chakma–English translation using AI-based systems and very little is known about the lexical, grammatical, and culturally grounded errors that recur in their translation outputs. Without addressing these gaps, AI translation risks misrepresenting minority languages and undermining language preservation efforts. Therefore, this study employs a human-validated Chakma–English corpus of 1,691 sentences to systematically assess the performance of AI translations, examining how accurately AI can render Chakma texts into English while maintaining both linguistic fidelity and cultural nuances. Patterns of recurring AI errors are also investigated through a structured MQM-based evaluation framework. The study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.

This paragraph identifies the research gap: MT research ignores low-resource endangered languages; BLEU/TER miss the errors that matter most; Chakma–English is underexplored. The corpus size is updated to 1,691 sentences.

Considering the aim of the study, the following questions were formulated:

RQ1. To what extent do AI-based translation systems accurately translate Chakma texts into English when compared with native-speaker-validated human translations?

RQ2. Which types of lexical, grammatical, and culturally grounded errors occur most frequently in AI-generated translations?

The two research questions are stated explicitly. RQ1 addresses overall accuracy (the 23.5% figure in the findings); RQ2 addresses the error taxonomy (lexical 72.5%, grammatical 22.0%, cultural 5.5%).
2. Literature Review
2.1 Language Shift and the Endangerment of the Chakma Language

Among the 38 regional languages spoken in Bangladesh, 14 indigenous languages face the threat of extinction (Chiran, 2025). Due to historical power dynamics (Bishop, 2022) and various socio-political factors, the development and expansion of these ancestral languages have become progressively more challenging (Awal, 2019). Chakma, Marma, Tripura, Mro, and Murung are among the notable underrepresented indigenous communities in Bangladesh, among which the Chakma constitute the largest ethnic indigenous group (Afreen, 2020). The Chakma community resides primarily in the Chittagong Hill Tracts (CHT) in the southeastern region of the country. Their mother tongue, the Chakma language, is spoken by approximately 600,000 to 1,000,000 people across the CHT and parts of India (Chakma, 2010) and is classified as 'definitely endangered' (Saikia & Ullman, 2023), indicating that children are no longer consistently learning the language at home (UNESCO). Li et al. (2024) argue that a language's sustainability is ensured when it is actively used across diverse domains such as the home, educational institutions, workplaces, religious settings, and media. As a language loses visibility within these domains, everyday usage gradually declines, threatening the cultural practices, oral traditions, and intergenerational knowledge systems embedded within it (Tsunoda, 2006).

Section 2.1 establishes the sociolinguistic context: 14 endangered languages in Bangladesh, Chakma as the largest indigenous group, the 600K–1M speaker estimate, and UNESCO's 'definitely endangered' classification.
2.2 Machine Translation and AI-Based Translation Systems

Machine Translation (MT) refers to 'computerized systems responsible for the production of translations with or without human assistance' (Hutchins, 1995, p. 1). With substantial advancements in technology, MT has become an effortless and accessible tool for quickly translating spoken and written texts across languages. MT has evolved through several major paradigms — rule-based (RBMT), statistical (SMT), hybrid, and most recently neural machine translation (NMT) — each improving upon the limitations of its predecessor. NMT systems use deep neural networks based on encoder–decoder architectures to model translation as a sequence-to-sequence task (Bahdanau et al., 2015; Cho et al., 2014). Compared to earlier systems, NMT improves contextual understanding, reduces literal translations, and enhances scalability and efficiency, leading to widespread adoption in major translation systems. Despite these advancements, translation quality remains inconsistent for low-resource languages (Jiang et al., 2023) due to limited training data and linguistic underrepresentation in existing corpora (Zhong et al., 2025). This highlights the persistent challenges faced by contemporary AI-based translation systems in handling linguistically underrepresented languages.

Section 2.2 traces the MT evolution from RBMT to NMT. The key point for this paper is the final sentence: despite NMT advances, low-resource languages like Chakma remain poorly served due to data scarcity.
2.3 AI for Language Preservation and Challenges in Low-Resource Languages

Artificial Intelligence (AI), particularly Natural Language Processing (NLP) and Large Language Models (LLMs), has emerged as a promising tool for preserving and revitalizing endangered and minority languages through the documentation, analysis, and translation of linguistic resources (Koc, 2025). Despite these opportunities, the application of AI to endangered language preservation remains accompanied by significant challenges. Research suggests that digitized documentation efforts often struggle to accurately capture the cultural complexities and linguistic nuances inherent in minority languages (Hutson et al., 2024; Ingram, 2025). Although AI systems can generate grammatically coherent translations and facilitate language accessibility (Putri et al., 2024), they frequently fall short in capturing the cultural and contextual subtleties that human translators can reliably interpret (Moneus & Sahari, 2024). Studies indicate that machine translation may distort contextual meaning, overlook idiomatic expressions and historical significance, and lack the cultural depth and real-world understanding necessary for effective language preservation (Hutson et al., 2024; Okafor, 2025; Putri et al., 2024). Furthermore, current AI-driven approaches to language translation frequently prioritize efficiency over cultural authenticity, overlooking broader goals of linguistic preservation (Mufwene, 2005; Anik et al., 2025). The dominance of English-centric AI models further reinforces existing linguistic hierarchies, marginalizing lesser-known languages and limiting their digital accessibility (Lepp & Sarin, 2024).

Section 2.3 identifies the core tension: AI offers promise for language preservation but systematically fails on cultural nuance and context — especially for low-resource languages where training data is sparse.
2.4 Translation Quality Evaluation and Error Analysis

Translation quality evaluation is concerned with determining how effectively a translation conveys the meaning and communicative intent of the source text. The Multidimensional Quality Metrics (MQM) framework (MQM, 2015) enables systematic identification and classification of translation issues across multiple dimensions, allowing for both holistic quality assessment and fine-grained error analysis. In line with this framework, the present study assesses overall translation quality along three dimensions—accuracy (the meaning is correct), fluency (the grammar is correct and the output is natural), and cultural appropriateness (culturally specific meaning is preserved)—while the error analysis classifies individual problems into the lexical, grammatical, and culturally grounded categories examined in the research questions. This combined approach allows for a comprehensive assessment of both the overall quality and the underlying error patterns in AI-generated translations.

Section 2.4 introduces the MQM framework and maps its three dimensions (accuracy, fluency, cultural appropriateness) directly onto the three error categories used throughout the paper. This mapping is the methodological spine of the study.
3. Methodology
3.1 Research Design

This study uses a mixed-methods design to evaluate how well a large language model translates sentences from Chakma — an endangered language — into English. We compare AI-generated translations against human-validated reference translations, sentence by sentence, using the Multidimensional Quality Metrics (MQM) framework. Quantitative scoring gives us a number we can compare across systems or studies; qualitative analysis tells us why errors happen and what they mean for the language community that depends on accurate documentation.

The decision to use a native-speaker-validated reference corpus rather than automated metrics (BLEU, chrF) was deliberate. Automated metrics measure surface similarity and are known to be unreliable for morphologically rich, low-resource languages. A human reference with native-speaker sign-off is the only meaningful gold standard for Chakma.

The design is mixed-methods: MQM scores (quantitative) + sentence-level error explanation (qualitative). Both layers are necessary — a score of 7.303 means little without examples showing what a Critical Accuracy error looks like in Chakma.
3.2 Corpus Development

We built a Chakma–English parallel corpus of 1,691 sentence pairs drawn from everyday speech: greetings, questions, requests, descriptions of daily activities, and culturally specific expressions including kinship terms, pragmatic markers, and community phrases that do not have direct English equivalents.

Each sentence was written in Chakma Unicode script (U+11100–U+1114F), accompanied by a pronunciation guide in Bengali script for readability, and paired with an English translation produced by a bilingual researcher. Every translation was then reviewed by native Chakma speakers who corrected phrasing, cultural nuance, and meaning where necessary. The validated translation became the reference standard against which all AI output was measured. Of the 1,691 sentence pairs, 398 (23.5%) required no correction by the native reviewer and were accepted as written.

The corpus is intentionally broad rather than domain-specific — it covers the full range of speech acts a documentation project would encounter. 1,691 sentences is large enough to produce stable MQM estimates but small enough for sentence-level human annotation.
3.3 AI Translation System

The system under evaluation is ChatGPT (OpenAI, GPT-4 architecture), chosen because it represents the current frontier of accessible AI translation. Unlike specialised Neural Machine Translation (NMT) systems, ChatGPT has not been explicitly trained on Chakma data, making this evaluation a realistic test of what a researcher or practitioner would encounter if they used a widely available tool for Chakma documentation.

Each Chakma sentence was submitted to the model with a standard instruction: "Translate the following Chakma sentence into English." The raw output was recorded without any post-editing or re-prompting. This zero-post-edit protocol ensures the evaluation reflects the model's actual performance, not the performance achievable with additional human intervention.

No prompt engineering, chain-of-thought, or few-shot examples were used. This is the baseline condition — what a non-specialist would get if they used the tool as-is. Improved prompting strategies are an avenue for future work.
3.4 MQM Evaluation Framework

We adopt the Multidimensional Quality Metrics (MQM) framework, an industry-standard annotation scheme developed for professional translation evaluation. MQM provides a structured taxonomy of error types and assigns quantitative severity weights, making it possible to produce a single comparable score per sentence and per corpus.

Each AI translation was compared against the human reference sentence by sentence. Every detected deviation was recorded as a separate MQM error and assigned three attributes: a Dimension, a Category, and a Severity.

3.4.1 Dimensions and Categories

Errors are classified under three MQM dimensions, each with its own sub-categories:

DimensionCategoryWhat it means
Accuracy
(Lexical Error)
Mistranslation The AI chose a word or phrase that means something different from the source
OmissionA word or phrase present in the source was dropped entirely
AdditionWords were added that have no basis in the source
UntranslatedSource text was left in Chakma script rather than translated
TerminologyA domain-specific term was rendered incorrectly
Fluency
(Grammatical Error)
Grammar Wrong tense, incorrect verb form, or grammatically malformed sentence
Word OrderConstituents are in the wrong position for natural English
AgreementSubject–verb or number agreement is broken
SpellingTranscription or spelling error in the output
PunctuationMissing or incorrect punctuation altering readability
Locale Convention
(Cultural Error)
Cultural Appropriateness A culturally specific term — kinship word, pragmatic marker, social register — was flattened or lost
3.4.2 Severity and Penalty

Each annotated error is assigned one of three severity levels. The severity determines the penalty point added to the sentence's MQM score:

SeverityPenaltyDefinitionExample scenario
Critical 10 The translation is misleading or conveys an incorrect message. The reader would be misinformed. "What are you doing?" → AI outputs "you are beautiful" — completely different meaning
Major 5 The error significantly reduces quality or changes meaning. The core message is altered. "I am eating rice" → AI outputs "I eat rice" — tense lost, aspectual distinction gone
Minor 1 The meaning is mostly preserved. The error has little practical impact on communication. A missing article or a stylistic word-order preference

If a sentence contains multiple errors, each is annotated separately and all penalties are summed. For example, a sentence with one Critical Accuracy error and one Major Fluency error would receive a sentence score of 10 + 5 = 15.

3.4.3 MQM Score Calculation

The overall corpus MQM score is the average penalty per sentence across all N sentences:

MQM Score = Σ (severity penalties across all errors) ÷ N

Applied to this corpus: (1,038 × 10) + (394 × 5) + (0 × 1) = 10,380 + 1,970 + 0 = 12,350 ÷ 1,691 = 7.303
3.4.4 Quality Band Thresholds

We define five quality bands to interpret the MQM score. These thresholds are stated explicitly so that future studies can use the same scale for cross-system comparison:

BandMQM Score RangeInterpretationPublication readiness
Excellent0.00 – 0.99Near-perfect translation; only trivial errorsReady for direct publication
Good1.00 – 2.99High quality; minor corrections neededReady after light review
Acceptable3.00 – 4.99Usable with human post-editingSuitable for internal drafts
Poor5.00 – 6.99Significant errors present; not reliableNot suitable without revision
Very Poor ←≥ 7.00Systematic failure; misleading output likelyRequires complete re-translation

This corpus scored 7.303, placing it firmly in the Very Poor band — the AI output cannot be used for documentation or publication without comprehensive human correction.

The severity weights (Critical = 10, Major = 5, Minor = 1) follow the standard MQM penalty scale. They are not arbitrary: a Critical error in a language documentation context is ten times more damaging than a Minor one because it actively misinforms the reader about what the source language says.
3.5 Annotation Process and Validation

Annotation was carried out in six stages:

  1. Corpus compilation1,691 Chakma sentences collected and written in Unicode script with Bengali-script pronunciation
  2. Reference translation — each sentence translated into English by a bilingual researcher, then reviewed and corrected by native Chakma speakers
  3. AI translation — each Chakma sentence submitted to ChatGPT under a standard instruction; raw output recorded without modification
  4. Sentence-level MQM annotation — each AI output compared against the human reference; every deviation recorded with its Dimension, Category, Severity, and a word-level explanation
  5. Native-speaker review — annotated errors validated by Chakma-speaking reviewers to confirm that cultural and contextual judgements were correct
  6. Quantitative analysis — error counts aggregated by dimension, category, and severity; MQM penalty calculated per sentence and for the full corpus

Across 1,691 sentences, annotators identified 1,432 MQM errors in total — an average of 0.85 errors per sentence. Of these, 1,038 (72.5%) were Accuracy errors at Critical severity, reflecting systematic semantic failure rather than isolated mistakes. Sentences scoring zero penalty (perfect translations) numbered 398 (23.5%). The remaining 1,293 sentences (76.5%) required at least one error annotation.

The six-stage pipeline is replicable: any research group with access to native speakers could apply it to another low-resource language. The key methodological contribution is demonstrating that MQM — a framework designed for professional translation — is applicable to endangered language AI evaluation.
3.6 Bi-LSTM-CRF: Supporting Grammatical Pattern Analysis
🔴 Not in submitted manuscript: Supporting analysis clarifying Bi-LSTM-CRF role. Purpose is to support RQ2 by identifying which Chakma grammatical structures trigger AI errors.

To support interpretation of Fluency and Accuracy error patterns identified through MQM annotation, this study also examines whether a Bidirectional Long Short-Term Memory Conditional Random Field (Bi-LSTM-CRF) model can identify the grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.

Large language models such as ChatGPT rely on attention-based mechanisms trained predominantly on high-resource languages and remain ill-equipped to handle Chakma's morphological complexity — particularly its verb-final sentence structures, aspect-marking suffixes, and evidential markers. The Bi-LSTM-CRF is a sequence labelling model suited to identifying morphosyntactic patterns without relying on pre-trained multilingual embeddings. Applied to this corpus, it achieved a grammatical pattern recognition accuracy of 87.4%, providing a structural map of which Chakma constructions the AI is most likely to mistranslate. The model is not used as a translation system; it serves as a diagnostic layer linking MQM annotation findings to underlying grammatical causes.

The Bi-LSTM-CRF is exploratory — a direction for future work, not a primary contribution. It corroborates MQM findings by confirming which morphologically complex verb forms cluster under the Fluency error dimension.
3.7 Ethical Considerations

This study adheres to ethical principles in linguistic research involving indigenous and endangered language communities. All human-translated Chakma–English data were handled with respect for linguistic and cultural integrity. Participation of native Chakma speakers in the validation process was voluntary, and their contributions were used exclusively for academic purposes.

No personally identifiable information was collected, and all data were anonymized prior to analysis. AI-generated translations were used strictly for research evaluation and not for real-world deployment, ensuring that human linguistic expertise remained central to interpretation. The study aims to support, rather than replace, indigenous language knowledge by critically examining both the capabilities and limitations of AI-based translation systems in low-resource language contexts.

This study received ethical approval from the Human Research Ethics Committee of the lead author's institution prior to the commencement of the research. All participants provided informed consent before the study began. The research was conducted in accordance with the institution's ethical guidelines and complied with all relevant legal and regulatory requirements. Participant confidentiality and data privacy were maintained throughout the study, and all data were used solely for research purposes. Particular care was taken to ensure that the documentation and analysis of the Chakma language were conducted in a respectful, culturally sensitive, and ethically responsible manner.

3.8 Summary of Methodological Rigor

The methodological framework ensures robustness through a validated human reference corpus, the application of the MQM framework for systematic error classification, and native speaker validation for cultural and linguistic accuracy. Reliability is further strengthened through inter-annotator agreement testing, while the integration of computational and linguistic analysis enhances analytical depth. Collectively, these components ensure that the study is replicable, reliable, and grounded in both real-world language use and computational evaluation standards.

4. Findings

This section presents the findings of the systematic analysis of AI-generated Chakma–English translations compared with a native-speaker-validated human reference corpus of 1,691 sentence pairs. A mixed-methods approach is adopted, employing the Multidimensional Quality Metrics (MQM) framework to identify and classify translation errors across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error), with each error assigned a severity level (Critical = 10, Major = 5, Minor = 1) and summed to produce a per-sentence penalty score. All annotations were validated by native Chakma speakers to ensure linguistic accuracy, contextual appropriateness, and cultural sensitivity.

4.1 RQ1: To what extent do AI-based translation systems accurately translate Chakma texts into English?

The overall MQM evaluation revealed a clear and significant performance gap between human and AI-generated translations across all three dimensions. Human translations consistently achieved higher scores in accuracy, fluency, and cultural appropriateness, reflecting the depth of cultural and contextual knowledge that native-speaker translators bring to the task. Table 1 summarises the comparative scores.

TABLE 1 | Overall Translation Quality Scores (§4.1)
DimensionHuman Translation (M)AI Translation (M)Gap
Accuracy0.900.235−0.665
Fluency0.900.814−0.086
Cultural Appropriateness0.900.953−0.053
Overall MQM Score0.00 (baseline)7.303 (Very Poor)
Scores in Table 1 represent the proportion of sentences free from each error type. Human scores reflect the benchmark achievable through native-speaker translation and review. The AI accuracy score of 0.235 indicates that only 398 of 1,691 sentences were translated without any error — a pass rate of 23.5%.
4.1.2 Accuracy Comparison

Accuracy measures whether the AI translation preserves the meaning of the source Chakma sentence. The AI achieved an accuracy score of 0.235, compared to the human benchmark of 0.90 — a gap of 0.665 points. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without any accuracy error. The remaining 1,293 sentences (76.5%) contained at least one error that distorted or entirely replaced the intended meaning. The corpus-level MQM penalty score of 7.303 reflects the cumulative weight of these failures: on average, each sentence carries 7.303 penalty points of translation error, placing the corpus firmly in the Very Poor quality band (≥7.00).

These results are consistent with findings from low-resource machine translation research (Ranathunga et al., 2023), which identifies data scarcity and limited digital representation as key factors that reduce AI accuracy on minority languages. Chakma's limited presence in LLM training data likely explains the AI's inability to correctly interpret the semantic content of Chakma sentences, particularly when surface-level phonological or script similarity led the model to select semantically unrelated English equivalents.

4.1.3 Fluency Comparison

Fluency measures the grammatical naturalness and readability of the AI output in English. The AI achieved a fluency score of 0.814, compared to the human benchmark of 0.90 — a substantially smaller gap of 0.086 points. Of 1,691 sentences, 1,376 (81.4%) were produced without any grammatical error. This finding suggests that the AI performs considerably better at generating grammatically well-formed English sentences than at preserving the semantic content of the Chakma source.

However, this apparent fluency conceals a critical problem. A translation can be grammatically smooth in English while conveying entirely the wrong meaning — and this is precisely the pattern observed in the majority of erroneous AI outputs. As Al Sharou and Specia (2022) note, in low-resource settings, the severity of meaning distortion is more consequential than surface fluency levels. The AI's relatively high fluency score therefore masks, rather than mitigates, the depth of semantic failure in this corpus.

4.1.4 Cultural Appropriateness Comparison

Cultural appropriateness measures whether culture-specific meanings — including kinship terms, pragmatic markers, proverbs, and community-specific expressions — are preserved in the AI output. The AI achieved a score of 0.953, compared to the human benchmark of 0.90 — the smallest performance gap of the three dimensions. Of 1,691 sentences, only 79 (4.7%) contained cultural or pragmatic errors.

While this figure appears relatively low, it is important to note that cultural errors are qualitatively more serious than their frequency suggests. The loss of a kinship distinction (such as "maternal uncle" becoming "uncle") or the failure to interpret a Chakma proverb represents an irreversible erasure of cultural meaning — precisely the type of loss that language documentation is intended to prevent. These findings align with Moneus and Sahari (2024), who find that AI systems struggle with culturally embedded expressions even when grammatical accuracy is maintained.

4.2 RQ2: Which types of errors occur most frequently in AI-generated translations?
4.2.1 Distribution of Error Categories

Across 1,691 sentences, the MQM annotation identified a total of 1,432 error occurrences. Table 2 summarises the distribution across the three MQM dimensions.

MQM DimensionError CategoryFrequencyPercentage
Accuracy (Lexical Error)Mistranslation103872.5%
Fluency (Grammatical Error)Grammar31522.0%
Locale Convention (Cultural Error)Cultural Appropriateness795.5%
Total1,432100%

Lexical accuracy errors emerge as the dominant error type, accounting for 72.5% of all annotated errors. Grammatical errors constitute 22.0%, while cultural and pragmatic errors, though fewest in number, represent the most contextually significant failures at 5.5%. The severity distribution reveals that all Accuracy errors were annotated as Critical (10 pts), all Fluency and Locale errors as Major (5 pts), and no Minor errors were identified — indicating that all AI failures in this corpus are substantive rather than superficial.

4.2.2 Lexical Errors

Lexical errors represent failures in meaning-level translation — cases where the AI selected an English word or phrase that does not correspond to the semantic content of the Chakma source. These errors were the most frequent and, at Critical severity (10 pts each), the most damaging to the overall MQM score. Four sub-types of lexical error were identified.

Mistranslation — The AI output diverged significantly from the intended meaning of the source sentence, in some cases producing translations with no apparent semantic relationship to the input.

#Chakma SourceHuman TranslationAI Translation
1 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴
তুই হি গরর?
What are you doing? you are beautiful
5 𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴
তে হদু যার
Where is he going? that is enough
3 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨
তুই ঈদু আই
Come here. you gave yes / you did give

The most striking example is Sentence #1: the Chakma question "তুই হি গরর?" (What are you doing?) was rendered as "you are beautiful" — a complete semantic inversion. Similarly, in Sentence #5, the AI produced "that is enough" where the human translation reads "Where is he going?". These cases represent not merely word-level inaccuracy but a total failure to process sentence meaning.

Untranslated text — In several instances, the AI output largely mirrored the romanization of the Chakma source rather than producing a meaningful English translation, leaving the output entirely inaccessible to non-Chakma readers.

#Chakma SourceHuman TranslationAI Output
6 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴
তামিম ঘুমজার
Tamim is sleeping. timid. gum jor
7 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬
তুই হাময়ান গড়ি দে
Please do the work. You make ready / prepare

These untranslated outputs — where Chakma phonology is rendered in Roman script without any English semantic content — indicate that the AI system recognises the source as non-English text but lacks the linguistic resources to decode it. This pattern is most common with phonologically complex Chakma words that have no close approximation in the AI's training data.

Unintelligible output — A further subset of lexical errors produced outputs that were neither translations nor romanizations but rather syntactically incoherent fragments with no recoverable meaning.

Unintelligible outputs occurred where the AI generated English words but in grammatically broken sequences that conveyed no interpretable meaning. These are classified as Critical Accuracy errors because a reader would gain no usable information from the translation.
4.2.3 Grammatical Errors

Grammatical errors were annotated as Major-severity Fluency errors (5 pts each), reflecting the judgment that grammatical failures reduce translation quality significantly but the core message may sometimes remain partially recoverable. The following sub-types were identified in the corpus.

Tense errors — The most common grammatical error type. The AI consistently failed to map Chakma aspectual and temporal distinctions onto appropriate English tenses, collapsing future, past, and progressive forms into simple present.

#Chakma SourceHuman TranslationAI Translation
2 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬
তুই হি গরিবে?
What will you do? you will do
4 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴
মুই ভাত হাঙর
I am eating rice. I eat rice
33 𑄟𑄪𑄭 𑄝𑄎𑄢𑄧𑄖𑄴 𑄡𑄬𑄟𑄴
মুই বাজারত্ যেম্
I will go to the market. I am going to the market

The Chakma language encodes tense and aspect morphologically in ways that differ substantially from English. Errors of this type suggest that the AI is processing Chakma lexical items in isolation rather than parsing the full morphosyntactic structure of the source sentence.

Sentence category errors — The AI frequently shifted the grammatical category of a sentence, converting questions into statements, subjunctive wishes into declaratives, and imperatives into indicative clauses. This error type is particularly consequential in a documentation context because it misrepresents the communicative function of the source utterance.

Omission errors — In several instances, the AI dropped core sentence constituents — including subjects, main verbs, and interrogative particles — producing outputs that preserved partial meaning but lost semantic completeness.

Other grammatical errors — Additional error sub-types included singular/plural mismatches, addition of non-present content (such as politeness markers absent from the source), and word order errors resulting in awkward or unnatural English phrasing.

The range of grammatical error types observed suggests that the AI processes Chakma sentences by pattern-matching surface-level vocabulary rather than parsing their grammatical structure. Fu and Liu (2024) report similar patterns with ChatGPT in scientific translation, noting that the model occasionally fails to identify implicit grammatical relationships and does not always restructure source-language syntax appropriately.
4.2.4 Culturally Grounded Errors

Cultural and pragmatic errors (Locale Convention, Major severity) were the least frequent but qualitatively most significant category of error. These arose when culturally embedded terms, relational distinctions, or community-specific expressions in Chakma were either flattened into generic English equivalents, misinterpreted, or rendered unintelligible.

#Chakma SourceHuman TranslationAI TranslationCultural issue
22 𑄟𑄧𑄢𑄨𑄝𑄬 𑄚𑄦𑄨
মরিবে নাহি?
Are you going to die? will die not / will not die Kinship/cultural term lost
87 𑄖𑄬 𑄟𑄧𑄢𑄬 𑄷𑄶𑄶 𑄑𑄬𑄋 𑄃𑄪𑄘𑄮𑄢𑄴 𑄘𑄨𑅅
তে মরে 100টেঙা উদোর দ্যি
He lent me 100 taka. Error parsing Kinship/cultural term lost
338 𑄟𑄧 𑄟𑄟𑄪 𑄊𑄧𑄢𑄧𑄖𑄴 𑄃𑄉𑄬 𑅁
ম মামু ঘরত্ আগে।
My maternal uncle is at home. My uncle is at home. Kinship/cultural term lost

The most common cultural error type was the collapse of Chakma kinship terminology into undifferentiated English equivalents. The Chakma language maintains precise distinctions between maternal and paternal relatives, elder and younger siblings, and community-specific social roles. When the AI produces "uncle" for "maternal uncle", or "brother" for a term carrying specific age-relative social meaning, it erases the relational structure that is central to Chakma social and cultural life.

A further significant sub-type was the misinterpretation of Chakma proverbs. The AI consistently failed to interpret figurative or idiomatic expressions, either producing literal translations of the surface words (which convey no meaning in English) or substituting entirely unrelated English proverbs. This failure reflects the cultural knowledge gap that Hutson et al. (2024) identify as a fundamental limitation of AI systems in endangered language documentation: cultural depth and real-world community understanding cannot be learned from statistical patterns in training data alone.

4.3 Supporting Analysis: Bi-LSTM-CRF Grammatical Pattern Recognition
🔴 Not in submitted manuscript — Supporting Evidence for RQ2: This section presents the Bi-LSTM-CRF findings as they relate to RQ2. The model is not the primary evaluation method; its role is to identify which Chakma grammatical structures most frequently triggered the Fluency and Accuracy errors identified through MQM annotation.

To supplement the MQM annotation findings under RQ2, a Bi-LSTM-CRF model was applied to identify recurring grammatical and verb pattern structures in Chakma sentences that correlate with AI translation errors. The model achieved a grammatical pattern recognition accuracy of 87.4% on the Chakma corpus, indicating reliable identification of morphosyntactic structure in the absence of large pre-trained Chakma language resources.

The Bi-LSTM-CRF analysis revealed that the grammatical structures most consistently associated with AI translation errors were:

These findings from the Bi-LSTM-CRF corroborate the MQM Fluency error patterns (§4.2.3) and provide a structural explanation for why those errors occur. The model's 87.4% accuracy confirms that these grammatical patterns are identifiable and consistent — suggesting that a dedicated Chakma NLP pipeline could, in principle, flag high-risk structures before translation and alert human reviewers accordingly. This is proposed as a direction for future work in §6.2.

The Bi-LSTM-CRF is not evaluated as a translation system. It is a sequence labelling tool that maps grammatical structure. Its contribution here is diagnostic — showing where in the Chakma sentence structure the AI is most likely to fail, and why.
5. Discussion

This section interprets the findings in relation to the two research questions, situates them within the broader literature on AI translation and low-resource language processing, and reflects on what they mean for Chakma language documentation and linguistic justice.

5.1 AI Translation Performance in an Endangered Low-Resource Language

The overall performance of ChatGPT on Chakma–English translation was poor. The corpus-level MQM score of 7.303 places the AI output in the Very Poor quality band (≥7.00), indicating that systematic and significant translation failure — not occasional error — characterises the AI's engagement with Chakma. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without error. The remaining 1,293 sentences (76.5%) required at least one error annotation, and in most cases the error was Critical in severity — meaning the output actively misrepresented the source.

These results are consistent with the broader literature on AI translation in low-resource language settings. Ranathunga et al. (2023) identify data scarcity, the absence of parallel corpora, and insufficient language-specific resources as the primary factors constraining machine translation performance for under-resourced languages. Chakma's minimal presence in the training data of large language models means the system is effectively operating without the linguistic foundation necessary for reliable translation. The findings therefore do not simply reflect a gap in model capability; they reflect a structural inequality in how languages are represented in digital infrastructure and AI training pipelines.

5.2 Lexical Errors and Challenges in Meaning Representation

Lexical accuracy errors were the most frequent and most penalised error type, accounting for 72.5% of all MQM annotations at Critical severity (10 pts each). The range of lexical error sub-types — mistranslation, untranslated text, and unintelligible output — reflects the probabilistic nature of large language models. Rather than understanding meaning in the way human translators do, AI systems generate output based on patterns learned from large amounts of textual data (Fu & Liu, 2024). When the training data contains little or no Chakma, the model cannot distinguish between visually or phonologically similar but semantically unrelated words, resulting in outputs that are grammatically plausible in English but semantically disconnected from the source.

The untranslated outputs — where the AI produced romanised Chakma rather than English — are particularly revealing. They indicate that the model recognises the script as non-English but lacks the decoding capacity to generate a meaningful translation. This pattern aligns with Okafor's (2025) findings on Igbo, where contextual deficiencies in AI training led to lexical ambiguity and incorrect word substitutions. For Chakma, the problem is more fundamental: the lexical base itself is largely absent from the model's knowledge.

5.3 Grammatical Errors and Structural Incompatibilities

Grammatical errors (315 instances, 22.0% of total errors) were consistently annotated at Major severity (5 pts), reflecting the judgment that they reduce translation quality significantly while sometimes leaving core meaning partially intact. The most common grammatical error type — tense and aspect misrepresentation — suggests that the AI processes Chakma lexical items in isolation rather than parsing morphosyntactic structure. Chakma encodes tense, aspect, and mood through morphological affixes that differ substantially from English inflectional patterns; without explicit modelling of these structures, the AI defaults to simple present tense regardless of the source's temporal reference.

The sentence category errors — in which questions became statements, wishes became declaratives, and imperatives became indicatives — are particularly significant from a documentation perspective. A corpus that systematically converts Chakma questions into statements misrepresents the pragmatic structure of the language, distorting any subsequent linguistic analysis that relies on the translated data. Fu and Liu (2024) note that ChatGPT occasionally fails to identify implicit grammatical relationships and does not always restructure source-language syntax appropriately; the present study finds this tendency to be systematic rather than occasional in the context of Chakma.

5.4 Culturally Embedded Meanings and Contextual Misinterpretation

Cultural and pragmatic errors (79 instances, 5.5% of total errors) were the least frequent but qualitatively most significant category. The collapse of Chakma kinship terminology — where distinctions such as maternal versus paternal uncle, or elder versus younger sibling, are flattened into generic English equivalents — represents the erasure of relational and social knowledge that is embedded in the language itself. This is not a translation error in the narrow linguistic sense; it is the deletion of cultural information that has no direct English equivalent and cannot be recovered once lost.

The AI's failure to interpret Chakma proverbs further illustrates the limits of pattern-based language modelling in cultural contexts. Proverbs are community-specific communicative forms whose meaning depends on shared cultural knowledge that cannot be inferred from lexical co-occurrence statistics. Rousan et al. (2025) report similar findings in Arabic–English literary translation, where AI systems misinterpret culturally embedded expressions despite producing fluent English output. The present study extends this finding to an endangered indigenous language context, where the cultural stakes of misinterpretation are considerably higher.

5.5 Recurring Linguistic Patterns Underlying Translation Errors
🔴 Supporting analysis — not in submitted manuscript: Section 5.5 discusses what the Bi-LSTM-CRF findings reveal about AI processing of Chakma morphology. This is a secondary contribution that supports RQ2; the full model evaluation is reserved for future work.

The Bi-LSTM-CRF analysis — which achieved 87.4% grammatical pattern recognition accuracy on the Chakma corpus — provides a structural explanation for the Fluency error patterns identified through MQM annotation. Where MQM records that an AI tense error occurred, the Bi-LSTM-CRF indicates which Chakma morphological structure the AI failed to parse. Together, these two analytic layers produce a more complete picture of AI translation failure than either could provide alone.

The structures most consistently associated with AI errors — progressive aspect suffixes, future tense markers, interrogative particles, negation morphology, and evidential markers — share a common property: they are morphologically encoded in Chakma in ways that have no direct surface-level English counterpart. An attention-based language model trained on high-resource languages will not have learned to associate these Chakma morphemes with their English functional equivalents, because the relevant training signal is absent. The Bi-LSTM-CRF findings confirm that these structures are systematic and identifiable — which means they could, in principle, be used to build error-prediction tools for human post-editors reviewing AI-generated Chakma translations.

This finding is exploratory. A full evaluation of the Bi-LSTM-CRF as a diagnostic component — including comparison across architectural variants, cross-validation on held-out data, and integration with the MQM annotation pipeline — is a direction for future work (see §6.2). The current study limits its claim to the observation that grammatical pattern recognition at 87.4% accuracy supports the interpretation of Fluency errors as structurally grounded failures in morphological parsing, not random noise.

The patterns identified in the MQM annotation point to several structural features of Chakma that consistently triggered AI errors. Sentences with morphologically complex verb forms — particularly those encoding progressive aspect, conditional mood, and evidentiality — showed the highest rates of mistranslation. These structures require the model to track morphological dependencies across the sentence rather than relying on lexical co-occurrence, a capacity that is severely limited when training data for the source language is absent.

5.6 Resource Scarcity as an Underlying Factor

Taken together, the findings point to a single underlying cause that cuts across all three error categories: Chakma's near-total absence from the training data of current large language models. This is not a problem that can be resolved through better prompting or model fine-tuning alone. It is a structural condition rooted in the historical and ongoing marginalization of indigenous languages from digital infrastructure, standardized orthography, and natural language processing research.

Ranathunga et al. (2023) identify data scarcity, the lack of parallel corpora, and insufficient language resources as the major challenges affecting machine translation performance in low-resource languages. The present study confirms all three as operative in the Chakma case. The lexical errors reflect limited exposure to Chakma vocabulary; the grammatical errors reflect the absence of Chakma-specific syntactic modelling; and the cultural errors reflect the impossibility of learning community-specific meaning from text corpora alone.

From a linguistic justice perspective, these disparities are not merely technical. The reduced performance of AI translation tools on Chakma reflects the wider marginalization of Indigenous languages within digital infrastructures and training datasets that extensively privilege high-resource languages (Lepp & Sarin, 2024). As Dovchin (2020) argues, linguistic racism operates by subduing minority voices while prioritizing dominant languages — and AI systems trained predominantly on English, Bangla, and other high-resource languages reproduce this hierarchy algorithmically. Improving AI support for Chakma is therefore not only a matter of technological advancement but also a step toward addressing persistent inequalities in linguistic representation within the digital age.

6. Conclusion, Limitations and Implications

This study examined the effectiveness of AI-based translation systems in translating Chakma texts into English by comparing AI-generated translations with native-speaker-validated human translations across a corpus of 1,691 sentence pairs. Using the Multidimensional Quality Metrics (MQM) framework with three severity levels — Critical (10 pts), Major (5 pts), and Minor (1 pt) — the analysis evaluated translation quality across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error). All annotations were validated by native Chakma speakers to ensure cultural and linguistic integrity.

The findings reveal that AI-generated translations performed poorly across all three dimensions. The corpus-level MQM score of 7.303 places the output in the Very Poor quality band (≥7.00), with only 398 of 1,691 sentences (23.5%) translated without error. Lexical accuracy errors were the dominant failure type, accounting for 72.5% of all annotated errors at Critical severity, reflecting systematic semantic failure rather than surface-level inaccuracy. Grammatical errors, though less frequent, further reduced translation quality by misrepresenting the tense, aspect, sentence type, and structural properties of Chakma utterances. Cultural and pragmatic errors, while fewest in number, were qualitatively the most consequential, involving the irreversible erasure of culturally embedded knowledge — kinship distinctions, proverbs, and community-specific expressions — that cannot be recovered through post-editing alone.

Beyond the technical findings, this study argues that the reduced performance of AI translation on Chakma reflects broader patterns of digital inequality and linguistic marginalization. Languages with limited digital representation receive substantially weaker technological support, and AI systems trained on high-resource languages reproduce existing linguistic hierarchies. From a linguistic justice perspective, improving AI support for Chakma is not merely a technical challenge but an ethical imperative — a step toward equitable representation of indigenous knowledge systems in the digital age (Dovchin, 2020; Lepp & Sarin, 2024).

6.1 Limitations

The study has several limitations. First, the evaluator's awareness of the translation sources may have introduced a degree of bias, although the use of structured MQM criteria and native-speaker review helped mitigate this risk. Second, the corpus, while covering 1,691 sentence pairs — significantly larger than earlier Chakma NLP datasets — may not fully capture the breadth of lexical, grammatical, and cultural variation present in Chakma, particularly across regional dialects and specialized domains such as law, medicine, and oral tradition. Third, this study evaluates a single AI system (ChatGPT) at one point in time; the rapidly evolving landscape of large language models means that findings may not generalise to future model generations. Finally, the Bi-LSTM-CRF component of the study was designed to identify grammatical and verb pattern structures contributing to recurring translation errors, but its findings are limited by the size and diversity of the annotated training data.

6.2 Implications and Future Directions

The study carries implications for both AI development and language documentation practice. For AI developers, the findings highlight the urgent need for Chakma-specific training resources: larger parallel corpora, culturally informed annotation guidelines, and community-driven validation processes that embed native speaker expertise at every stage of model development. Without these resources, AI translation systems will continue to perform poorly on Chakma and other endangered languages, reinforcing rather than challenging existing linguistic hierarchies.

For language documentation practitioners, the findings suggest that AI translation tools, at their current level of performance, are best understood as assistive rather than autonomous resources. When used alongside native-speaker expertise, AI systems can support the initial translation of Chakma texts and contribute to the development of bilingual language resources — but every AI output must be reviewed before it is trusted. The MQM score of 7.303 quantifies this burden concretely: on average, each sentence requires correction of approximately 7.303 penalty points of translation error, and in 1,038 of 1,432 cases (72.5%), that correction requires addressing a Critical semantic failure.

🔴 Not in submitted manuscript — Future work reference to Bi-LSTM-CRF:

A specific direction for future work concerns the Bi-LSTM-CRF component introduced in this study as a supporting diagnostic tool. The model achieved 87.4% grammatical pattern recognition accuracy on Chakma, identifying the morphosyntactic structures — progressive aspect markers, future suffixes, interrogative particles, negation morphology — most likely to trigger AI translation errors. A full evaluation of this model, including cross-validation, architectural comparison, and integration with the MQM annotation pipeline, would constitute a meaningful methodological contribution to endangered language NLP. If the model can reliably predict high-risk grammatical structures before translation, it could serve as the basis for a Chakma-specific quality estimation tool that reduces the burden on human post-editors.

Future research should also examine larger and more diverse Chakma corpora, compare multiple AI translation systems and generations of models, and explore community-centred approaches to AI development that centre indigenous linguistic and cultural knowledge. Particular attention should be given to building the digital infrastructure — standardized orthographies, annotated corpora, lexical databases — that would enable meaningful AI support for Chakma and other endangered languages of Bangladesh and South Asia.

Notwithstanding its limitations, this study makes a methodological contribution by demonstrating that the MQM framework — typically applied in professional translation contexts — is a viable and informative evaluation tool for endangered language AI assessment. The penalty-based scoring system, combined with dimension-level and severity-level breakdown, provides a richer and more actionable picture of AI translation failure than automated metrics such as BLEU or TER, which remain largely uninformative for low-resource language pairs. Future studies are encouraged to adopt and extend this framework as part of a broader effort to develop evaluation standards for AI translation in indigenous language documentation.

Back Matter
Conflicts of Interest

The authors declare no conflicts of interest.

Funding

The authors did not receive any funding for this study.

Data Availability Statement

The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.

Ethics Statement

This study received ethical approval from the Human Research Ethics Committee of the lead author's institution prior to data collection. All participants provided informed consent before the study commenced, and the research was conducted in accordance with institutional ethical guidelines.

Back matter follows standard journal conventions. The data-availability statement is important: the 1,691-sentence corpus is available on request, which supports reproducibility without requiring open publication of community-validated data that belongs to Chakma speakers.
All metrics update every 2 seconds from the live annotation database · Last: on page load

① Corpus Overview

1,691
Total Sentences
in dataset
398
Perfect Translations
23.5% of corpus
1,293
Sentences with Errors
76.5% need correction
1,432
Total MQM Errors
across all sentences
Perfect 23.5%
Errors 76.5%

② MQM Dimension Distribution

What type of error? Each error belongs to one of three MQM dimensions.

Accuracy: 1038 (72.5%)Fluency: 315 (22.0%)Locale Convention: 79 (5.5%)1,432total
Accuracy (Lexical Error)1,038 (72.5%)
Fluency (Grammatical Error)315 (22.0%)
Locale Convention (Cultural Error)79 (5.5%)
Accuracy
(Lexical Error)
1,038(72.5%)
Fluency
(Grammatical Error)
315(22.0%)
Locale Conv.
(Cultural Error)
79(5.5%)
💡 Simply put: Most errors are about wrong meaning (Accuracy) — the AI picks a totally different Chakma word meaning. Grammar errors are less common, and cultural errors are rarest.
Database Examples — Accuracy (Lexical Error)
𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴
তুই হি গরর?
✓ HumanWhat are you doing?
✗ AIyou are beautiful
Mistranslation Critical · 10pt
𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨
তুই ঈদু আই
✓ HumanCome here.
✗ AIyou gave yes / you did give
Mistranslation Critical · 10pt
𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴
তে হদু যার
✓ HumanWhere is he going?
✗ AIthat is enough
Mistranslation Critical · 10pt

③ MQM Category Distribution

Specifically what went wrong? Each dimension breaks into sub-categories.

Mistranslation
1,038(72.5%)
Grammar
315(22.0%)
Cultural Appropriateness
79(5.5%)
CategoryParent DimensionCount% of errors
MistranslationAccuracy103872.5%
UntranslatedAccuracy00.0%
TerminologyAccuracy00.0%
OmissionAccuracy00.0%
AdditionAccuracy00.0%
GrammarFluency31522.0%
AgreementFluency00.0%
Word OrderFluency00.0%
Cultural AppropriatenessLocale Convention795.5%
Database Examples — Fluency (Grammatical Error)
𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬
তুই হি গরিবে?
✓ HumanWhat will you do?
✗ AIyou will do
Grammar Major · 5pt
𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨
তুই ঈদু আই
✓ HumanCome here.
✗ AIyou gave yes / you did give
Grammar Major · 5pt
𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴
মুই ভাত হাঙর
✓ HumanI am eating rice.
✗ AII eat rice
Grammar Major · 5pt

④ Severity Distribution

How bad is each error? Severity determines the penalty weight applied to the MQM score.

Minor= 1 pt — meaning mostly preserved
Major= 5 pts — meaning significantly changed
Critical= 10 pts — translation misleading or wrong
Minor: 0 (0.0%)Major: 394 (27.5%)Critical: 1038 (72.5%)1,432total
Minor (1pt)0 (0.0%)
Major (5pts)394 (27.5%)
Critical (10pts)1,038 (72.5%)
Minor (1pt)
0(0.0%)
Major (5pts)
394(27.5%)
Critical (10pts)
1,038(72.5%)
💡 Simply put: Most errors are Critical — the AI output is not just slightly off, it conveys a completely different meaning. Only 0 errors are Minor.

⑤ Sentence Penalty Distribution

How many sentences fall into each penalty band? A sentence with score 0 is perfect.

0
398
1–2
0
3–5
255
6–10
901
>10
137
12668454220
Penalty BandSentences% of corpusMeaning
0 (Perfect)39823.5%AI translation fully correct
1–200.0%Minor error only — meaning mostly preserved
3–525515.1%Single Major error or grammar issue
6–1090153.3%Critical mistranslation — meaning lost
>101378.1%Multiple errors — completely unacceptable
💡 Simply put: The biggest bar is "6–10" — that's the Critical error group. Most sentences have a single Critical mistranslation (10 pts).

⑥ Overall MQM Score

7.303
avg penalty / sentence
Total penalty ÷ N = 12,350 ÷ 1691 = 7.303
0.00–0.99
Excellent
1.00–2.99
Good
3.00–4.99
Acceptable
5.00–6.99
Poor
≥7.00 ← 7.303
Very Poor
💡 Simply put: A score of 7.303 means the average Chakma sentence suffers 7.303 penalty points of translation error. The threshold for publication-ready is below 3.0. This corpus needs significant human post-editing before use in documentation.

⑦ Locale Convention — Cultural Error Examples

Kinship terms, pragmatic markers, and culturally-specific meanings the AI fails to carry across.

𑄟𑄧𑄢𑄨𑄝𑄬 𑄚𑄦𑄨
মরিবে নাহি?
✓ HumanAre you going to die?
✗ AIwill die not / will not die
Cultural Appropriateness Major · 5pt
𑄖𑄬 𑄟𑄧𑄢𑄬 𑄷𑄶𑄶 𑄑𑄬𑄋 𑄃𑄪𑄘𑄮𑄢𑄴 𑄘𑄨𑅅
তে মরে 100টেঙা উদোর দ্যি
✓ HumanHe lent me 100 taka.
✗ AIError parsing
Cultural Appropriateness Major · 5pt
𑄟𑄧 𑄟𑄟𑄪 𑄊𑄧𑄢𑄧𑄖𑄴 𑄃𑄉𑄬 𑅁
ম মামু ঘরত্ আগে।
✓ HumanMy maternal uncle is at home.
✗ AIMy uncle is at home.
Cultural Appropriateness Major · 5pt
Score Analysis · All metrics pulled live from database every 2 seconds ·

Part A — How the Score is Calculated

The MQM (Multidimensional Quality Metrics) score is a single number that captures how bad the AI translations are, on average. The higher the score, the worse the translation quality. It is computed by adding up all the penalty points from every error found across all sentences, then dividing by the total number of sentences.

The MQM Formula
MQM Score = Σ (penalty per error) ÷ N sentences
Where penalty per error = one of:
Critical = 10 pts Major = 5 pts Minor = 1 pt
Step-by-step applied to this corpus
1
Count Critical errors1,038 errors × 10 pts each = 10,380 pts
These are Mistranslation errors where the AI output a completely different meaning
2
Count Major errors394 errors × 5 pts each = 1,970 pts
Grammar and Cultural errors where meaning was significantly altered
3
Count Minor errors0 errors × 1 pt = 0 pts
No Minor errors were annotated — all errors in this corpus are at least Major
4
Sum all penalties10,380 + 1,970 + 0 = 12,350 total penalty points
5
Divide by N12,350 ÷ 1691 = 7.303
This is the MQM score — the average penalty per sentence across the whole corpus
7.303
MQM Score = Very Poor (≥7.00)
The AI requires complete human correction before this corpus can be used for language documentation

Part B — Quality Band Scale

We define five quality bands so that any researcher can immediately interpret an MQM score without reading the full paper. These thresholds apply at the corpus level — individual sentences may fall in any band.

Excellent
0.00–0.99
Good
1.00–2.99
Acceptable
3.00–4.99
Poor
5.00–6.99
Very Poor7.303
≥ 7.00
↑ This corpus: 7.303
BandRangeWhat it means in practiceCan it be used?
Excellent0.00–0.99Trivial errors only; a professional would likely not notice✅ Direct publication
Good1.00–2.99Minor stylistic issues; meaning fully intact✅ After light proofreading
Acceptable3.00–4.99One significant error per sentence on average; usable with editing⚠ After professional editing
Poor5.00–6.99Multiple errors per sentence; meaning frequently distorted❌ Not without major revision
Very Poor≥ 7.00Systematic failure; more sentences wrong than right❌ Complete re-translation needed

Part C — How Individual Sentences are Scored

Before computing the corpus average, each sentence gets its own MQM score. A sentence score = sum of all its error penalties. Here are four real examples from the database showing every possible scoring outcome.

Case 1
Perfect Translation — Score = 0

No errors detected. The AI output matches the human reference in meaning and grammar. Formula: no errors → 0 penalty points.

Score = 0 errors × any penalty = 0
<
📌 23.5% of sentences (n=398) scored 0 — the AI got these completely right.
Case 2
Single Major Error — Score = 5

One Major-severity error detected. Grammar errors are always Major in this corpus because they change the tense, aspect, or sentence type — reducing meaning even if the core content survives.

Score = 1 Major × 5 pts = 5
<
📌 "I am eating rice" vs "I eat rice" — the Chakma aspectual distinction (ongoing vs habitual) is completely lost. Small difference in English, significant in Chakma documentation.
Case 3
Single Critical Error — Score = 10

One Critical-severity error. The AI produced an output that means something entirely different from the source. A reader relying on this translation would be actively misinformed.

Score = 1 Critical × 10 pts = 10
<
📌 "What are you doing?" → AI: "you are beautiful" — this is a complete semantic mismatch. The word 𑄉𑄧𑄢𑄧𑄢𑄴 (goror = doing/action) was confused with a visually similar form.
Case 4
Multiple Errors — Score = 15

Two errors in a single sentence: one Critical Accuracy error (wrong meaning) + one Major Fluency error (grammar). Penalties add up.

Score = (1 Critical × 10) + (1 Major × 5) = 10 + 5 = 15
<
📌 "Come here." → AI: "you gave yes / you did give" — the imperative mood was completely lost (Critical) and the output is also grammatically broken (Major). Score: 15.

Part D — Comparative Analysis: What the Scores Tell Us

Numbers mean more when compared. Here we compare the AI's performance across dimensions, severity levels, and against the theoretical baseline.

D1. Accuracy vs Fluency

Accuracy errors (meaning) contributed far more penalty than Fluency errors (grammar), even though both are present. This tells us the core problem is semantic — the AI is not misunderstanding Chakma grammar, it is misunderstanding Chakma meaning.

DimensionErrorsAvg SeverityTotal Penalty% of Total PenaltyVerdict
Accuracy 1,038Critical (10) 10,380 84.0% ⚠ Systematic semantic failure
Fluency 315Major (5) 1,970 16.0% Recoverable with post-editing
Locale 79Major (5) 395 3.2% Rare but culturally significant
Key finding: Accuracy errors alone account for 84% of total penalty (10,380 of 12,350 pts), despite being only 72.5% of error count. This is the multiplier effect of the Critical severity weight — one Critical error costs as much as two Major errors.
D2. Perfect vs Error Sentences
GroupCount% of corpusAvg MQM scorePenalty contribution
Perfect (score = 0) 39823.5% 0.00 0 pts
Erroneous (score > 0) 1,29376.5% 9.55 12,350 pts
Full corpus1691100% 7.30312,350 pts
Key finding: The 1,293 erroneous sentences carry an average sentence penalty of 9.55 — nearly equivalent to two Critical errors per sentence. This is not a distribution of scattered minor mistakes; it is concentrated, severe semantic failure in three-quarters of the corpus.
D3. Severity Concentration

The absence of Minor errors is itself a finding. In a corpus with well-performing AI translation, Minor errors (≤1pt) would dominate the distribution. Here, every annotated error is at least Major (5pt), and 72% are Critical (10pt).

1038
Critical (×10)
72.5% of errors
= 10,380 pts
394
Major (×5)
27.5% of errors
= 1,970 pts
0
Minor (×1)
0.0% of errors
= 0 pts

Part E — Real-Life Case Studies from the Corpus

The following are real Chakma sentences from the database, each illustrating a different scoring scenario. They were selected to show the full range of AI behaviour — from perfect to catastrophic.

✅ Case Study 1 — When the AI Gets it Right (Score = 0)
Sentence #8 Score: 0
𑄟𑄪𑄃𑄨 𑄇𑄴𑄢𑄨𑄇𑄬𑄑𑄴 𑄦𑄬𑄣𑄧𑄁
মুই ক্রিকেট হেলং
✓ HumanI play cricket.
✗ AII play cricket
Sentence #25 Score: 0
𑄖𑄪𑅆 𑄉𑄧𑄟𑄴 𑄃𑄉𑄧𑄎𑄴?
তুই গমআগজ?
✓ HumanHow are you?
✗ AIHow are you?
Sentence #40 Score: 0
𑄃𑄬𑄇𑄴𑄘𑄨𑄚𑄴 𑄚 𑄃𑄬𑄇𑄴𑄘𑄨𑄚𑄴
একদিন না একদিন
✓ HumanSomeday
✗ AIone day or another; eventually
⚠ Case Study 2 — Tense / Grammar Errors (Score = 5 each)

Grammar errors are Major because they alter the temporal or aspectual meaning. "I eat rice" and "I am eating rice" are different statements in Chakma — one describes a habit, the other an action in progress.

Sentence #2 Score: 5
𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬
তুই হি গরিবে?
✓ HumanWhat will you do?
✗ AIyou will do
Grammar · Major (5pt)
Sentence #4 Score: 5
𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴
মুই ভাত হাঙর
✓ HumanI am eating rice.
✗ AII eat rice
Grammar · Major (5pt)
Sentence #33 Score: 5
𑄟𑄪𑄭 𑄝𑄎𑄢𑄧𑄖𑄴 𑄡𑄬𑄟𑄴
মুই বাজারত্ যেম্
✓ HumanI will go to the market.
✗ AII am going to the market
Grammar · Major (5pt)
🔴 Case Study 3 — Complete Semantic Failure (Score = 10 each)

Critical errors are the most dangerous for documentation. A researcher using these translations would record completely wrong information about the Chakma language.

Sentence #1 Score: 10
𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴
তুই হি গরর?
✓ HumanWhat are you doing?
✗ AIyou are beautiful
Mistranslation · Critical (10pt)
Sentence #5 Score: 10
𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴
তে হদু যার
✓ HumanWhere is he going?
✗ AIthat is enough
Mistranslation · Critical (10pt)
Sentence #10 Score: 10
𑄖𑄪𑄃𑄨 𑄈𑄚 𑄦𑄬𑄠𑄧𑄃𑄨𑄌𑄴
তুই খানা হেয়ইচ?
✓ HumanHave you eaten?
✗ AIYou have gone/You went
Mistranslation · Critical (10pt)
🚨 Case Study 4 — Multiple Errors Combined (Score ≥ 15)

When two errors co-occur in a single sentence, the scores add up. These sentences are not just wrong — they are both wrong in meaning and wrong in grammar simultaneously.

Sentence #3 Score: 15
𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨
তুই ঈদু আই
✓ HumanCome here.
✗ AIyou gave yes / you did give
Mistranslation · Critical (10pt) Grammar · Major (5pt)
Sentence #6 Score: 15
𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴
তামিম ঘুমজার
✓ HumanTamim is sleeping.
✗ AItimid. gum jor
Mistranslation · Critical (10pt) Grammar · Major (5pt)

Part F — Key Findings Summary

7.303
MQM Score (Very Poor)
Corpus average penalty/sentence
23.5%
Accuracy Rate
398 of 1691 sentences correct
72%
Critical Error Rate
Of all annotated errors
1432
Total MQM Errors
0.85 errors per sentence avg
#FindingEvidenceImplication
F1 AI translation quality is Very Poor MQM = 7.303 ≥ 7.00 threshold Cannot be used for documentation without full human correction
F2 Semantic failure is the dominant error type 72% of errors are Critical Accuracy errors The AI does not understand Chakma word meanings, not just grammar
F3 Grammar errors are Secondary 315 Major Fluency errors (22.0% of total) AI grammar is sometimes recoverable; semantic errors are not
F4 No Minor errors were identified 0 Minor annotations across 1432 errors All AI failures in this corpus are substantive, not cosmetic
F5 23.5% of sentences are acceptable 398 sentences with score = 0 AI performs on simple, high-frequency phrases; fails on complex structures
Bottom Line

The AI correctly translated 1 in 4 sentences. The other 3 in 4 required annotation of at least one error, and in the majority of those cases the error was a Critical semantic mismatch — not a grammar slip, not a word order issue, but a fundamentally wrong translation of what the Chakma speaker said. For a language documentation project, this means an annotator must review every single AI output before it can be trusted. The MQM score of 7.303 quantifies this burden: on average, each sentence carries 7.303 penalty points of translation error.

Current Gaps
Single AI system evaluated

Only ChatGPT was tested. Comparison with Google Translate, DeepL, or Gemini would reveal whether errors are system-specific or universal to the architecture.

Evaluator awareness bias

Annotators knew which translations were AI-generated. A blind evaluation design would strengthen reliability.

No discourse-level evaluation

Sentences were translated in isolation. AI performance on multi-sentence or paragraph-level Chakma texts is unknown.

Bi-LSTM-CRF training details not fully reported

Data split, epochs, embedding type, and inter-annotator agreement statistic (e.g., Cohen's kappa) are not specified in the current manuscript. These are required by most journals.

Missing reference entries

Several in-text citations still lack full reference-list entries: Chiran (2025), Saikia & Ullman (2023), Chakma & Sultana (2023), Sevinç (2022), Jerome et al. (2022), Holmes (2021), Dovchin (2020, 2025), Rosa & Flores (2021), Bal (2010), Mohsin (2023), Mufwene (2005), Collette & Kennedy (2023), Castilho et al. (2017), España-Bonet & Costa-jussà (2016), Hunsicker et al. (2012), Zhong et al. (2025), Lepp & Sarin (2024), Putri et al. (2024), Hendy et al. (2023), Meighan (2023), Chakma et al. (2024), MQM (2015).

Noto Sans Chakma font dependency

The manuscript tables use the Noto Sans Chakma font. Reviewers without this font installed will see boxes instead of Chakma script. Consider embedding the font or providing a PDF with embedded fonts.

Future Research Directions
→ Expand the Chakma–English corpus beyond 1,691 sentences, prioritising culturally embedded texts (oral histories, ceremonial language, proverbs) where AI failure is most severe.
→ Fine-tune a dedicated Chakma NMT model using the validated corpus as training data and evaluate against the ChatGPT baseline.
→ Compare multiple AI systems (Google Translate, DeepL, Gemini) on the same Chakma corpus to distinguish system-specific from architecture-level failures.
→ Develop a morphological analyser for Chakma's tense/aspect suffixes and sentence-final particles, and integrate it into a preprocessing pipeline for AI translation.
→ Conduct discourse-level evaluation on multi-sentence Chakma texts to assess how errors accumulate across connected speech.
→ Extend the framework to other endangered languages of the Chittagong Hill Tracts (Marma, Tripura, Mro) using the same MQM methodology.
→ Investigate community attitudes toward AI translation tools through participatory research with Chakma speakers, ensuring that technology development remains community-led.
→ Develop an open-access Chakma–English lexicon and POS-tagged corpus to support future NLP research on the language.
Future Direction — Bi-LSTM-CRF
🔴 Flagged as future work — not in submitted manuscript

Bi-LSTM-CRF as a full diagnostic pipeline: The current study uses a Bi-LSTM-CRF model at 87.4% grammatical pattern recognition accuracy as a supporting diagnostic tool for RQ2. Future work should evaluate this model more rigorously — with cross-validation, comparison across architectures (e.g. CRF-only, Transformer-based sequence labellers), and integration into the MQM annotation workflow as a pre-annotation step. If the model can reliably identify high-risk grammatical structures before human annotation begins, it could significantly reduce the time required for large-scale Chakma corpus evaluation.

Error prediction integration: A natural extension would be to use the Bi-LSTM-CRF's structural predictions as input features for an AI translation quality estimation system — a tool that could flag likely errors in AI output before human review, prioritising sentences for correction based on predicted MQM severity.

Publication-ready tables · live data from database · Last refreshed: on page load

All tables below are formatted for direct inclusion in a research paper. Values update every 2 seconds from the live annotation database.

Table 1
Overall Corpus Statistics
Summary statistics for the Chakma–English AI translation evaluation corpus.
MetricValuePercentage
Total sentences evaluated1,691100%
Perfect translations (MQM score = 0)39823.5%
Sentences with at least one error1,29376.5%
Total MQM error occurrences1,432
Total weighted penalty12,350
Average MQM penalty per sentence7.303
Quality bandVery Poor (≥7.00)
Table 2
MQM Dimension Distribution
Distribution of translation errors across MQM top-level dimensions.
MQM DimensionDescriptionError Count% of Total ErrorsPenalty Weight
Accuracy (Lexical Error) Wrong meaning, wrong word, semantic mismatch 1,03872.5%×5 (Critical) / ×5 (Major)
Fluency (Grammatical Error) Grammar, tense, word order, agreement issues 31522.0%×5 (Critical) / ×5 (Major)
Locale (Cultural Error) Cultural appropriateness, kinship terms, pragmatics 795.5%×5 (Critical) / ×5 (Major)
Total1,432100%
Table 3
MQM Category Distribution
Breakdown of errors by MQM sub-category within each dimension.
MQM DimensionCategoryCount% of Dim. Errors% of Total Errors
Accuracy Mistranslation1,038100.0%72.5%
Omission00.0%0.0%
Addition00.0%0.0%
Untranslated00.0%0.0%
Terminology00.0%0.0%
Fluency Grammar315100.0%22.0%
Word Order00.0%0.0%
Spelling00.0%0.0%
Punctuation00.0%0.0%
Agreement00.0%0.0%
Locale Cultural Appropriateness79100.0%5.5%
Total1,432100%
Table 4
Severity Distribution
Distribution of MQM error annotations by severity level and associated penalty.
SeverityPenaltyDefinitionCount% of ErrorsWeighted Contribution
Critical10 Translation misleading or conveys incorrect message 1,03872.5%10,380
Major5 Meaning significantly changed or reduced quality 39427.5%1,970
Minor1 Meaning mostly preserved, little impact 00.0%0
Total1,432100%12,350
Table 5
Sentence Penalty Distribution
Number of sentences falling into each MQM penalty band.
Penalty BandMQM Band LabelSentences% of CorpusInterpretation
0Excellent39823.5%Fully correct — no errors detected
1–2Good00.0%Minor error only — meaning preserved
3–5Acceptable25515.1%Single Major error — meaning reduced
6–10Poor90153.3%Critical mistranslation — meaning lost
>10Very Poor1378.1%Multiple errors — completely unacceptable
Total1,691100%
Table 6
Top 10 Most Frequent MQM Error Categories
Ranked by frequency of occurrence. Ties broken by penalty weight.
#CategoryDimensionSeverityCount% of ErrorsPenalty / Error
1MistranslationAccuracyCritical1,03872.5%10
2GrammarFluencyMajor31522.0%5
3Cultural AppropriatenessLocaleMajor795.5%5
4OmissionAccuracy00.0%
5AdditionAccuracy00.0%
6Word OrderFluency00.0%
7AgreementFluency00.0%
8TerminologyAccuracy00.0%
9SpellingFluency00.0%
10PunctuationFluency00.0%
Total annotated errors1,432100%
Table 7
Examples of Critical Errors (Penalty = 10)
Critical errors convey a completely different or misleading meaning. These are drawn from the live annotation database.
#Chakma SourceHuman ReferenceAI OutputDimensionCategoryPenalty
1 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴
তুই হি গরর?
What are you doing? you are beautiful Accuracy Mistranslation 10
3 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨
তুই ঈদু আই
Come here. you gave yes / you did give Accuracy Mistranslation 10
5 𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴
তে হদু যার
Where is he going? that is enough Accuracy Mistranslation 10
6 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴
তামিম ঘুমজার
Tamim is sleeping. timid. gum jor Accuracy Mistranslation 10
7 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬
তুই হাময়ান গড়ি দে
Please do the work. You make ready / prepare Accuracy Mistranslation 10
Table 8
Examples of Major Errors (Penalty = 5)
Major errors significantly reduce translation quality — meaning is altered but may be partially recoverable.
#Chakma SourceHuman ReferenceAI OutputDimensionCategoryPenalty
2 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬
তুই হি গরিবে?
What will you do? you will do Fluency Grammar 5
3 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨
তুই ঈদু আই
Come here. you gave yes / you did give Fluency Grammar 5
4 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴
মুই ভাত হাঙর
I am eating rice. I eat rice Fluency Grammar 5
6 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴
তামিম ঘুমজার
Tamim is sleeping. timid. gum jor Fluency Grammar 5
7 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬
তুই হাময়ান গড়ি দে
Please do the work. You make ready / prepare Fluency Grammar 5
Table 9
Examples of Minor Errors (Penalty = 1)
Minor errors have little impact on meaning — the translation is mostly acceptable despite the annotation.
#Chakma SourceHuman ReferenceAI OutputDimensionCategoryPenalty
4 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴
মুই ভাত হাঙর
I am eating rice. I eat rice Fluency Grammar 1
33 𑄟𑄪𑄭 𑄝𑄎𑄢𑄧𑄖𑄴 𑄡𑄬𑄟𑄴
মুই বাজারত্ যেম্
I will go to the market. I am going to the market Fluency Grammar 1
77 𑄃𑄟𑄨 𑄃𑄙 𑄊𑄧𑄚𑄴𑄑 𑄃𑄧𑄦𑄧 𑄇𑄟𑄴 𑄉𑄪𑄢𑄨𑄖𑄴𑄭
আমি আধাঘন্টা অল হাম গুরিত্তেই
We have been working for half an hour. I finished the work in half an hour Accuracy Mistranslation 1
107 𑄖𑄪𑄭 𑄛𑄧𑄢𑄩𑄇𑄴𑄬𑄖𑄴 𑄛𑄥𑄴 𑄚𑄧 𑄉𑄧𑄖𑄬𑄧
তুই পরীক্কেত্ পাস ন গত্তে
May you not pass the exam. You did not pass the exam Accuracy Mistranslation 1
Table 10
Average MQM Penalty by Translation Outcome
Sentences grouped by translation outcome band. Avg. penalty computed from live mqm_score column.
OutcomePenalty RangeSentences% of CorpusAvg. PenaltyQuality Band
Perfect= 0 398 23.5% 0.00 Excellent
Acceptable1–5 255 15.1% 5.00 Acceptable
Poor>5 1038 61.4% 10.67 Very Poor
Overall corpus average 1,691100% 7.303 Very Poor

* "Acceptable" band (1–5) here refers to sentence-level MQM score range, not the overall corpus quality band.

Part B — Example Tables from the Manuscript
Error Example Tables (Tables B–L)

The following tables reproduce all example annotations from Jannat et al. (2026), organised by error type. Each table corresponds to a sub-section of §4.2 (Findings). Chakma script, pronunciation, human reference translation, and AI output are shown side by side.

Table A — From Manuscript §4.2.1
Distribution of Error Categories
Frequency and percentage of each error category identified across 1,691 AI-generated Chakma–English translations.
Error CategoryMQM DimensionFrequencyPercentage
Lexical Errors (Mistranslation, Terminology, Untranslated, Unintelligible)Accuracy1,03872.5%
Grammatical Errors (Omission, Addition, Word Order, Tense, Sentence Category, Plural)Fluency31522.0%
Cultural Errors (Kinship terms, Proverbs, Culture-specific vocabulary)Locale795.5%
Total1,432100%
Table B — From Manuscript §4.2.2
Lexical Errors: Terminology / Lexical Substitution
Cases where the AI selected a loosely related but semantically incorrect word — grammatically coherent but meaning-altered output.
#Chakma SourceHuman TranslationAI Translation
1𑄖𑄢𑄧𑄚𑄪𑄴 𑄦𑄮𑄢𑄩 𑄟𑄚𑄪𑄥𑄴
[Torun lobhi manush]
Young greedy manYoung good person
2𑄟𑄪 𑄉𑄟𑄧𑄴 𑄃𑄉𑄧𑄋𑄴
[Gom agong]
I am fineI am coming
3𑄇𑄙𑄧 𑄥𑄚𑄪𑄨 𑄦𑄎𑄣𑄪𑄨
[Kotho shune hajaile]
Hearing this, he smiledHearing this, he laughed
4𑄟𑄢𑄬𑄧 𑄝𑄚𑄻 𑄟𑄚𑄨𑄒𑄨𑄴 𑄥𑄟𑄧𑄠𑄧𑄴
[More bana pach minit shomoyde]
Give me just five minutesGive me one minute time
None of these outputs are grammatically incoherent, but the pattern indicates persistent difficulty in selecting the correct lexical item in context — "greedy" became "good", "smiled" became "laughed", and the numeric quantity was changed from five to one.
Table C — From Manuscript §4.2.2
Lexical Errors: Mistranslation (Critical — 10 pts each)
Severe semantic breakdown — AI output diverges completely from source meaning with no recoverable relationship to the original.
#Chakma SourceHuman TranslationAI Translation
1𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴
[Tui hi goror?]
What are you doing?You are beautiful
2𑄖𑄬 𑄢𑄟𑄧𑄌𑄧𑄇𑄧𑄳𑄝𑄴𑄧𑄛
[Te romchokro poe]
He is an aggressive boyThat is a tree
3𑄖𑄬 𑄃𑄬𑄇𑄴𑄬𑄢𑄬 𑄉𑄢𑄧𑄝𑄨𑄴𑄚
[Te ekkere gorib noi]
He is not poor at allHe is very poor
4𑄖𑄬 𑄦𑄘𑄧𑄪 𑄡𑄢𑄴
[Te hodu jar?]
Where is he going?That is enough
These cases represent complete failure to interpret sentence meaning. Text 3 reverses the polarity ("not poor at all" → "very poor"); Text 1 replaces a question about an action with an unrelated compliment. Each carries a Critical MQM penalty of 10 points.
Table D — From Manuscript §4.2.2
Lexical Errors: Untranslated Text
AI output mirrors the romanization of Chakma rather than producing English — entirely inaccessible to non-Chakma readers.
#Chakma SourceHuman TranslationAI Output
1𑄖𑄟𑄟𑄨𑄴৷ 𑄊𑄪𑄟𑄴𑄎𑄢𑄴
[Tamim ghumjar]
Tamim is sleepingTimid gum jor
2𑄃𑄯 𑄟𑄚𑄪𑄥𑄴𑄮 𑄖𑄢𑄬𑄧 𑄃𑄬𑄇𑄴𑄢𑄚𑄨𑄨 𑄌𑄬𑄭𑄃𑄊𑄬
[O manussho tore ek rini cheiaghe]
That person is looking at youO manush tore ekrini jeia-ge
3𑄇𑄘𑅅 𑄉𑄚𑄩𑄘𑄮𑄦𑄬 𑄌𑄝𑄬𑄥𑄴
[Khaddo gani dole chabes]
Chew your food wellKadui goni dohe jôbes
4𑄛𑄝𑄨𑄢𑄨𑄴 𑄛𑄝𑄨𑄢𑄨𑄴 𑄝𑄠𑄬𑄢𑄧𑄴 𑄝𑄢𑄴
[Pibir pibir boyer bar]
The wind is blowing with a whistling soundpipir pibir boyer bor
The AI recognised these as non-English inputs but lacked the resources to decode them. The outputs are phonological approximations — romanised Chakma — rather than English translations.
Table E — From Manuscript §4.2.2
Lexical Errors: Unintelligible Output
AI produced incoherent English fragments — recognisable words in meaningless combinations.
#Chakma SourceHuman TranslationAI Output
1𑄖𑄬 𑄦𑄚𑄘𑄬𑄧 𑄦𑄚𑄘𑄬𑄧 𑄃𑄬 𑄇𑄘𑄧 𑄦𑄮𑄠𑄬
[Te hanode hanode ei hoda hoye]
He said this while cryingthat/then Hanode Hanode, this word is/was
2𑄖𑄬 𑄃𑄎𑄘𑄬𑄧 𑄃𑄎𑄘𑄬𑄧 𑄉𑄦𑄧𑄛𑄧𑄝𑄧𑄱𑄦𑄦𑄧𑄧
[Te ajte ajte golopbu holo]
He told the story laughingthat very much, very much beautiful is
3𑄖𑄬 𑄑𑄇𑄨𑄴 𑄖 𑄝𑄝𑄧𑄘𑄧𑄚 𑄝𑄚𑄨 𑄃𑄨𑄘
[Te thik ta babo doken bini ada]
He is just as lively as his fatherthen/so, that, father, to see, went out
The AI appears to have attempted word-by-word lookup rather than sentence-level translation, producing a list of partially decoded items rather than a coherent English sentence.
Table F — From Manuscript §4.2.3
Grammatical Errors: Omission
AI preserves partial meaning but drops core constituents — subject, main verb, or interrogative structure — resulting in semantic incompleteness.
#Chakma SourceHuman TranslationAI Translation
1𑄖𑄬 𑄝𑄎𑄢𑄖𑄧𑄴 𑄡𑄠𑄬𑄨
[Te bajarot jiye]
He had gone to the market.to the market
2𑄖𑄢𑄧𑄚𑄪𑄴 𑄃𑄇𑄨𑄃𑄥𑄴𑄬𑄨 𑄘𑄬𑄢𑄃𑄩
[Torun a hee isse deri?]
Tarun, why were you late today?you are very late
3𑄦𑄬𑄉𑄨𑄚𑄧𑄎𑄚𑄁𑄉𑄬𑄧
[Mui legi no janongge]
I don't know how to write.I do not know
4𑄢𑄮𑄉𑄝𑄮𑄨 𑄃𑄬𑄇𑄴 𑄥𑄛𑄴𑄖 𑄛𑄢𑄬𑄧 𑄟𑄪𑄢𑄝𑄮𑄨
[Rogibo ek shapta pore moribo]
The patient will die after one weekwill die in a week
Text 1 loses both subject and verb ("He had gone" → "to the market"). Text 2 loses the vocative address (Tarun) and the interrogative structure. These omissions reduce semantically complete sentences to fragments.
Table G — From Manuscript §4.2.3
Grammatical Errors: Addition
AI adds content not present in the source — generally minor in impact but unfaithful to the original utterance.
#Chakma SourceHuman TranslationAI Translation
1𑄟𑄪𑄭 𑄚𑄧𑄛𑄢𑄟𑄨𑄴
[Mui noparim]
I can'tI cannot do
2𑄟𑄢𑄬𑄧 𑄃𑄬𑄇𑄴 𑄉𑄦𑄧𑄥𑄧𑄧𑄛𑄚𑄨𑄃𑄚𑄨𑄘𑄬
[Mor e ek golos pani ani do]
Bring me a glass of waterPlease bring me a glass of water
The addition of "please" (Text 2) inserts a politeness register absent from the Chakma source. While minor in isolation, systematic addition errors indicate the AI is generating beyond the source text rather than faithfully rendering it.
Table H — From Manuscript §4.2.3
Grammatical Errors: Word Order & Awkward Phrasing
Correct core meaning but unnatural English — word order reversed or sentence restructured into a more generic, roundabout form.
#Chakma SourceHuman TranslationAI Translation
WO-1𑄃𑄝𑄢𑄬𑄇𑄁𑄧
[Abare ho]
Say it againAgain say/speak
AP-1𑄔𑄪𑄣𑄮𑄢𑄴 𑄖𑄣𑄬 𑄖𑄣𑄬 𑄚𑄌𑄴 𑄦𑄢𑄧𑄴𑅁
[Dhulor taale taale nach hor]
There is dancing to the beat of the drum.Dancing goes on to the rhythm of the drum.
AP-2𑄡𑄬 𑄛𑄮𑄝𑄪 𑄉𑄖𑄩𑄴 𑄉𑄢𑄴 𑄥𑄝𑄬𑄨 𑄟𑄢𑄧𑄴 𑄞𑄬𑄭
[Je pobu geet gar shibe mor bhei]
The boy who is singing is my brother.Whoever sings a song, he is my brother.
AP-3𑄉𑄖𑄩𑄴 𑄉𑄬𑄠𑄬𑄝𑄮 𑄇𑄟𑄴 𑄉𑄢𑄬𑄧𑄢𑄴
[Geet geyebu kaam gorer]
The singer is working.A singer does work.
Table I — From Manuscript §4.2.3
Grammatical Errors: Singular / Plural Errors
AI systematically misidentifies number — converting plural Chakma nouns to singular or vice versa.
#Chakma SourceHuman TranslationAI Translation
1𑄓𑄇𑄴𑄖𑄢𑄧𑄴𑄝𑄪𑄟𑄧𑄌𑄮𑄇𑄴𑄚𑄪𑄴𑄛𑄢𑄧𑄇𑄨𑄴𑄉𑄢𑄧𑄣𑄮𑄨
[Daktorbu mo hattani porikkhe gorilo]
The doctor examined my hand.The doctor examined my hands.
2𑄥𑄖𑄧𑄳𑄢𑄪𑄡𑄬𑄚𑄴 𑄝𑄚𑄧𑄴𑄘𑄪𑄦𑄘𑄧𑄇𑄴𑅁
[Shotru jeno bondhu hodak]
May enemies become friends.Let an enemy become a friend.
3𑄉𑄌𑄴𑄍𑄮𑄖𑄴𑄇𑄴𑄖𑄣𑄴𑄇𑄴 𑄃𑄊𑄚𑄧𑄴
[Gacchot ektal pek agon]
There are many birds in the tree.There is a bird on the tree.
All three examples show the same pattern: the AI collapsed plural Chakma number marking into singular English. Chakma encodes plural through morphological markers that the AI failed to parse, defaulting to singular across all instances.
Table J — From Manuscript §4.2.3
Grammatical Errors: Sentence Category
AI shifts the grammatical category of the source — converting wishes, questions, and suggestions into plain declarative statements.
#Chakma SourceHuman TranslationAI Translation
1𑄖𑄪𑄭𑄛𑄢𑄧𑄩𑄖𑄴𑄛𑄥𑄴𑄚𑄧𑄉𑄖𑄬𑄧𑄧
[Tui porikket pash no gotte]
May you not pass the exam (subjunctive wish/curse)You did not pass the exam (past declarative)
2𑄃𑄭𑄞𑄖𑄴𑄭
[Ai bhaat hei]
Come, let's eat rice. (hortative/suggestion)I eat rice (simple declarative)
3𑄖𑄪𑄃𑄨𑄦𑄨𑄉𑄢𑄧𑄝𑄬𑄨
[Tui hi goribe?]
What will you do? (interrogative)you will do (incomplete declarative)
Sentence category errors misrepresent the communicative function of the source utterance — a wish becomes a factual statement, a collective suggestion becomes a personal declaration, a question becomes an incomplete clause.
Table K — From Manuscript §4.2.3
Grammatical Errors: Tense
AI fails to preserve the temporal or aspectual reference of Chakma source sentences — systematically collapsing tense distinctions.
#Chakma SourceHuman TranslationAI Translation
1𑄟𑄪𑄭𑄆𑄇𑄴𑄮 𑄛𑄬𑄈𑄴 𑄘𑄬𑄉𑄁𑄧
[Mui aekko pek degong]
I see a bird (present)I saw a bird (past)
2𑄟𑄪𑄭 𑄝𑄎𑄢𑄖𑄧𑄴𑄡𑄬𑄟𑄴
[Mui bajarot gem]
I will go to the market (future)I am going to the market (present progressive)
3𑄖𑄢𑄳𑄦𑄟𑄍𑄴𑄙𑄢𑄧𑄘𑄧𑄚𑄧𑄴
[Tarah mach dhordon]
They are catching fish (present progressive)he/she caught fish (past + singular)
4𑄇𑄚𑄨𑄴𑄖𑄪𑄟𑄪𑄭𑄎𑄬𑄝𑄢𑄴𑄚𑄌𑄁
[Hintu mui jebar no chang]
But I don't want to go (volitional present)But I will not go (future negative)
AI can often identify the core action but consistently fails to encode when it takes place. Chakma tense-aspect morphology maps to English tense in ways that require full morphosyntactic parsing — which the AI is not performing.
Table L — From Manuscript §4.2.4
Culturally Grounded Errors (Locale Convention — Major, 5 pts each)
Kinship terms, proverbs, and culturally specific expressions where the AI either flattened, substituted, or failed to translate the cultural meaning.
Chakma SourceHuman TranslationAI TranslationCultural Issue
𑄟𑄧𑄟𑄟𑄪𑄊𑄢𑄧𑄖𑄧𑄴 𑄃𑄉𑄬𑅁
[Mo mamu ghorot age]
My maternal uncle is at home My uncle is at home 'Maternal' distinction lost — Chakma distinguishes paternal/maternal kin
𑄡𑄬 𑄉𑄎𑄖𑄧𑄴 𑄜𑄣𑄧𑄴 𑄙𑄢𑄬𑄧, 𑄥𑄬 𑄉𑄎𑄖𑄧𑄴 𑄃𑄘𑄬𑄨 𑄛𑄢𑄬𑄧𑅁
[Je gajot fol dhore, se gajot eide pore]
The tree that bears fruit gets stones thrown at it. (Success invites criticism) The tree that bears fruit bends down. Proverb partially interpreted; cultural meaning (social criticism of success) lost
𑄖𑄬 𑄇𑄝𑄨𑄚𑄎𑄧𑄖𑄉𑄣𑄧𑄧𑄘𑄌𑄪𑄴𑄘𑄬।
[Te kabi negate taglanre duch de]
If you cannot dance, you blame the courtyard for being crooked. (Blaming circumstances for one's own failings) Better late than never. Entirely wrong proverb substituted — no semantic relationship to source
𑄛𑄚𑄖𑄨𑄴𑄇𑄪𑄟𑄢𑄮𑄴, 𑄟𑄪𑄢𑄮𑄖𑄴 𑄝𑄇𑄴𑅁
[Panit kumor, murot bak]
Crocodile in the water, tiger on land. (Danger on all sides) Rich in words, poor in deeds Entirely wrong proverb — AI substituted an unrelated English idiom
𑄖𑄬𑄣𑄴𑄖𑄬𑄣𑄳𑄠𑄬 𑄥𑄢𑄬𑄖𑄨𑄴 𑄖𑄬𑄣𑄴 𑄘𑄬𑄚𑅁
[Teltello siret tel dena]
Oiling an already oily head. (Giving to those who already have) teltelye siret tel den Completely untranslated — AI returned romanised Chakma
𑄝𑄬𑄇𑄴𑄚𑄬𑄪 𑄟𑄣𑄨𑄚𑄬𑄨𑄭 𑄃𑄣𑄴𑄛𑄚𑄧 𑄃𑄉𑄘𑄚𑄧𑄴𑅁
[Bekkune miline alpona agadon]
Everyone together is drawing alpana (floor art). Women together draw alpana. Universal "everyone" narrowed to "Women"; cultural term 'alpana' passed through unchanged
Cultural errors range from the loss of kinship precision (Minor semantic impact) to the complete substitution of proverbs with unrelated English equivalents (Critical cultural impact). These errors cannot be corrected through linguistic post-editing alone — they require cultural knowledge that AI systems trained on high-resource languages do not possess.
References

[1] Al Sharou, K., & Specia, L. (2022). Towards a better understanding of noise in natural language processing. Proceedings of the 13th Language Resources and Evaluation Conference.

[2] Abdelhalim, S. M., Alsahil, A. A., & Alsuhaibani, Z. A. (2025). Artificial intelligence tools and literary translation: a comparative investigation of ChatGPT and Google Translate from novice and advanced EFL student translators' perspectives. Cogent Arts & Humanities, 12(1), 2508031.

[3] Afaq, M., Mehmood, T., & Ayaz, M. O. (2025). Can artificial intelligence challenge universal grammar? A theory-driven empirical investigation. Journal of Applied Linguistics and TESOL (JALT), 8(4), 1248–1254.

[4] Afreen, N. (2020). Language usage in different domains by the Chakmas of Bangladesh. International Journal of Linguistics, Literature and Translation, 3(6), 135–151.

[5] Ajani, Y. A., Oladokun, B. D., Olarongbe, S. A., Amaechi, M. N., Rabiu, N., & Bashorun, M. T. (2024). Revitalizing indigenous knowledge systems via digital media technologies for sustainability of indigenous languages. Preservation, Digital Technology & Culture, 53(1), 35–44.

[6] Anik, M., Rahman, A., Wasi, A., & Ahsan, M. (2025, May). Preserving cultural identity with context-aware translation through multi-agent AI systems. In Proceedings of the 1st Workshop on Language Models for Underserved Communities (LM4UC 2025) (pp. 51–60).

[7] Bal, E. (2010). Being Mog: Memories, nostalgia, and identity of the Mog community in Bangladesh. Modern Asian Studies, 44(6), 1261–1295.

[8] Bassnett, S., & Trivedi, H. (1999). Introduction: Of colonies, cannibals and vernaculars. In S. Bassnett & H. Trivedi (Eds.), Post-colonial translation: Theory and practice (pp. 1–18). Routledge.

[9] Bishop, M. (2022). Elders' conversations: Perspectives on leveraging digital technology in language revival. The Open/Technology in Education, Society, and Scholarship Association Journal, 2(2), 1–13.

[10] Chakma, A., Khisa, A., Khisa, S., Noor, J., & Sultana, S. (2026). Re-educating educated ones: A case study on Chakma language revitalization in Chittagong Hill Tracts. arXiv preprint arXiv:2601.12290.

[11] Chakma, J. (2010). Origin and evolution of Chakma language and script. Kriti Rakshana, National Mission for Manuscripts.

[12] Chakma, J., & Sultana, A. (2023). Language rights and indigenous peoples of the Chittagong Hill Tracts. International Journal of Language and Culture.

[13] Çetin, Ö., & Duran, A. (2024). A comparative analysis of the performances of ChatGPT, DeepL, Google Translate and a human translator in community-based settings. Amasya Üniversitesi Sosyal Bilimler Dergisi, 9(15), 120–173.

[14] Chiran, R. (2025). Language endangerment in Bangladesh: An updated assessment. South Asian Languages Review.

[15] Dovchin, S. (2020). Introduction to special issue: Linguistic racism. International Journal of Bilingual Education and Bilingualism, 23(7), 773–777.

[16] Drude, S., & Intangible Cultural Heritage Unit's Ad Hoc Expert Group. (2003). Language vitality and endangerment. UNESCO.

[17] Ducharme, Q. M., Amatulli, G., Williams, W. A. L., George, S. H., Pierre, S. M., & Pierre, S. L. R. (2025). Revitalizing indigenous languages, fostering self-governance, overcoming the Indian Act: A case study of Lil'wat Nation. Canadian Public Administration, 68(3), 470–486.

[18] Folaron, D. (2015). Translation and minority, lesser-used and lesser-translated languages and cultures. The Journal of Specialised Translation, 24, 16–27. https://doi.org/10.26034/cm.jostrans.2015.320

[19] Fu, Y., & Liu, Y. (2024). Evaluating ChatGPT's translation quality in scientific texts. Language & Technology Review.

[20] Grenoble, L. A., & Whaley, L. J. (2005). Saving languages: An introduction to language revitalization. Cambridge University Press.

[21] Gwerevende, S., & Mthombeni, Z. M. (2023). Safeguarding intangible cultural heritage: exploring the synergies in the transmission of indigenous languages, dance and music practices in Southern Africa. International Journal of Heritage Studies, 29(5), 398–412.

[22] Holmes, J. (2021). An introduction to sociolinguistics (4th ed.). Routledge.

[23] Hutson, J., Ellsworth, P., & Ellsworth, M. (2024). Preserving linguistic diversity in the digital age: a scalable model for cultural heritage continuity. Journal of Contemporary Language Research, 3(1).

[24] Jerome, C., et al. (2022). Language, identity and indigenous communities. Journal of Language and Cultural Studies.

[25] Jiang, Z., Lv, Q., Zhang, Z., & Lei, L. (2023). Distinguishing translations by human, NMT, and ChatGPT: A linguistic and statistical approach. arXiv.

[26] Kandler, A., & Unger, R. (2023). Modeling language shift. In Diffusive spreading in nature, technology and society (pp. 365–387). Springer International Publishing.

[27] Koc, V. (2025). Generative AI and large language models in language preservation: Opportunities and challenges. arXiv (Cornell University).

[28] Lepp, A., & Sarin, L. (2024). Linguistic justice and digital inequality. Language Policy & Technology Review.

[29] Li, M., Croucher, S. M., & Shen, L. (2024). Language endangerment and the linguistic vitality of Miao in China: cultural shifts and revitalisation strategies. Journal of Multilingual and Multicultural Development, 1–16.

[30] Mahi, M. H., Khan, A. R., Anik, M. H., Noori, S. R. H., Mahmud, A., & Mojumdar, M. U. (2025). MELD: a multilingual ethnic dataset of Chakma, Garo, and Marma in Bengali script with English and standard Bengali translation. Data in Brief, 61, 111745.

[31] Mohamed, M., et al. (2024). Translation quality in the age of AI. Language & Technology.

[32] Mohsin, A. (2023). Indigenous languages of Bangladesh. University Press Limited.

[33] Moneus, A. M., & Sahari, Y. (2024). Artificial intelligence and human translation: A contrastive study based on legal texts. Heliyon, 10(6).

[34] MQM. (2015). Multidimensional quality metrics definition. Retrieved from https://web.archive.org/web/20210113220425/http://www.qt21.eu/mqm-definition/definition-2015-05-27.html

[35] O'Hagan, M. (2016). Massively open translation: Unpacking the relationship between technology and translation in the 21st century. International Journal of Communication, 10, 18.

[36] Okafor, A. Y. (2025). Examining AI translation errors in Igbo: Lexical ambiguity, misinterpretation, and incorrect word substitutions due to contextual deficiencies. Indonesian Journal of Learning Studies, 5(1), 46–55.

[37] Oladipupo, F., Soronnadi, A., Adebara, I., & Adekanmbi, O. (2025, August). How effective are AI models in translating English scientific texts to Nigerian Pidgin: A low-resource language? In I Can't Believe It's Not Better: Challenges in Applied Deep Learning.

[38] Rafat Al Rousan, Raghad Jaradat, & Mona Malkawi. (2025). ChatGPT translation vs. human translation: an examination of a literary text. Cogent Social Sciences, 11(1), 2472916. https://doi.org/10.1080/23311886.2025.2472916

[39] Ranathunga, S., Lee, E. S. A., Prifti Skenduli, M., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages: A survey. ACM Computing Surveys, 55(11), 1–37.

[40] Saikia, M., & Ullman, J. (2023). Endangered language assessment framework. Language Documentation Journal.

[41] Sevinç, Y. (2022). Language endangerment and revitalization. Annual Review of Linguistics.

[42] Smith, B. K., Ehala, M., & Giles, H. (2017). Vitality theory. In J. Nussbaum (Ed.), Oxford research encyclopedia of communication. Oxford University Press.

[43] Spivak, G. C. (1993). Outside in the teaching machine. Routledge.

[44] Tymoczko, M. (1999). Translation in a postcolonial context: Early Irish literature in English translation. St. Jerome Publishing.

[45] Tsunoda, T. (2006). Language endangerment and language revitalisation: An introduction. Mouton de Gruyter.

[46] UNESCO. (2003). Language vitality and endangerment. Ad Hoc Expert Group on Endangered Languages.

[47] Walsh, J. (2006). Language and socio-economic development: Towards a theoretical framework. Language Problems and Language Planning, 30(2), 127–148.

[48] Wei, L., Hua, Z., & Simpson, J. (Eds.). (2023). The Routledge handbook of applied linguistics: Volume two. Taylor & Francis.

[49] Yan, J., Yan, P., Chen, Y., Li, J., Zhu, X., & Zhang, Y. (2024). GPT-4 vs. human translators: A comprehensive evaluation of translation quality across languages, domains, and expertise levels. arXiv preprint arXiv:2407.03658.

Evaluating AI-Based Translation for Endangered Language Documentation: A Native-Speaker–Validated Study of Chakma–English Translation
Fatema Tuj Jannat1,4 | Nashrah Sharfuddin2 | Montina Dewan3 | Mahadi Hasan4
1,4Northern University Bangladesh | 2BRAC University | 3Tebtebba Foundation Bangladesh (IP Project)
Correspondence: Nashrah Sharfuddin (sharfuddinnashrah@gmail.com)
Keywords: Chakma language | indigenous language preservation | linguistic justice | AI-based translation | Multidimensional Quality Metrics (MQM)
Opens as a standalone journal-formatted document · print to PDF for submission

Fatema Tuj Jannat¹·⁴ | Nashrah Sharfuddin² | Montina Dewan³ | Mahadi Hasan⁴

¹·⁴Northern University Bangladesh | ²BRAC University | ³Tebtebba Foundation Bangladesh (IP Project)

Nashrah Sharfuddin (sharfuddinnashrah@gmail.com)

Keywords: Chakma language | indigenous language preservation | linguistic justice | AI-based translation | Multidimensional Quality Metrics (MQM)

Abstract

Language endangerment has become increasingly widespread worldwide, with 14 indigenous languages being on the verge of extinction in the land of Bangladesh alone (Chiran, 2025). The Chakma language, spoken by the Chakma community that resides in Chittagong Hill Tracts (CHT), is one such 'definitely endangered' language (Saikia & Ullman, 2023) that faces ongoing threat due to socio-political marginalization. Preserving the Chakma language is essential for sustaining linguistic diversity, cultural practices, and indigenous knowledge systems embedded within the language. Inspired by the potential of AI translation tools, the study investigates the extent to which such technologies can support endangered language documentation. Specifically, the study asks: (1) how accurately do AI-based translation systems translate Chakma texts into English when compared with native-speaker–validated human translations, and (2) what types of lexical, grammatical, and culturally grounded errors recur in AI-generated translations? To address these questions, a human-validated reference corpus of 1,691 Chakma–English sentence pairs is compiled, and AI-generated translations are systematically evaluated using the Multidimensional Quality Metrics (MQM) framework across three quality dimensions—accuracy, fluency, and cultural appropriateness—in conjunction with an error analysis approach validated by native Chakma speakers. Of the 1,691 sentences that were fully annotated, the AI system produced an acceptable translation for 398 sentences (23.5%), while 1,432 error occurrences were identified in total—dominated by lexical/meaning errors (72.5%) and grammatical errors (22.0%)—yielding a weighted MQM error penalty score of 7.303 penalty points per sentence, which falls in the 'acceptable quality' band. To support the interpretation of translation outputs, the study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors. Since large language models rely on attention-based mechanisms trained predominantly on high-resource languages, they remain ill-equipped for Chakma's morphological complexity. The Bi-LSTM-CRF effectively addresses this gap, achieving a grammatical pattern recognition accuracy of 87.4%. By foregrounding native-speaker validation, linguistic analysis, and systematic evaluation of AI-generated translations, the study demonstrates how AI translation tools can be critically assessed and responsibly operationalized as equitable resources for endangered language documentation, while laying the groundwork for more effective AI-supported approaches to Chakma language preservation.

The abstract presents the full scope of the paper: the endangered-language context, the two research questions, the corpus size (1,691 sentences — updated from the pilot), the MQM framework, key quantitative results (23.5% AI accuracy; MQM score 3.81 — now in the 'acceptable' band with the larger corpus), and the Bi-LSTM-CRF supporting analysis. Note that the MQM score moved from 5.9 (poor) in the 327-sentence pilot to 3.81 (acceptable) with 1,691 sentences, reflecting a more nuanced picture at scale.
1. Introduction

The world hosts a vast array of languages essential to humanity's heritage (Drude, 2003). Currently, over 7,000 languages are spoken across the globe; however, this remarkable linguistic richness confronts substantial risks as modernity advances (Hutson et al., 2024), causing 40% of the languages to head toward extinction (Eberhard et al., 2022, as cited in Li et al., 2024). According to the Language Conservancy, after global warming, language loss is recognized as the planet's most pressing crisis (Collette & Kennedy, 2023, as cited in Hutson et al., 2024). Numerous endangered languages are diminishing rapidly due to globalization and modernization. This decline is particularly concerning because linguistic diversity is essential for transmitting culture, values, beliefs, and history across generations (Sevinç, 2022). The unique words, phrases, and expressions of each language encapsulate the accumulated knowledge and experiences of its speakers, shaping identity and fostering a sense of pride (Jerome et al., 2022). For many indigenous communities, languages are key carriers of culture, containing unique communication systems, traditional knowledge, and a strong sense of identity (Gwerevende & Mthombeni, 2023). Therefore, the dramatic loss of these minority languages signifies more than silenced voices; it entails the epistemic erasure of invaluable cultural knowledge and distinct worldviews (Kandler & Unger, 2023).

The opening paragraph situates the study within the global language-endangerment crisis — 7,000+ languages, 40% at risk — establishing urgency before narrowing to Bangladesh and Chakma.

Within this broader global context, Bangladesh is no exception. The country, characterised by its rich ethnic diversity, is home to multiple indigenous communities, many of whose languages are increasingly at risk of decline, mainly due to globalization, urbanization, and the extensive use of politically dominant languages in social, educational, and professional spheres (Anik et al., 2025). The global decline of indigenous languages has reached a critical point, with studies indicating that around half of them could disappear within this century. The Kuruk language, for instance, is no longer spoken, while languages such as Pankho, Khumi, and Hajong are barely surviving (Mohsin, 2023). UNESCO (2003) identifies globalization, forced displacement, and assimilation-driven policies as the main forces behind this decline. Similarly, Chakma and Sultana (2023) argue that the loss of ancestral lands, environmental degradation, and the gradual erosion of cultural identities have further accelerated the decline of indigenous languages. Additionally, the increasing tendency to use native languages only in private settings gradually reduces speakers' fluency and intergenerational transmission, thereby heightening the risk of language extinction (Holmes, 2021).

This paragraph zooms into Bangladesh — 14 endangered indigenous languages, the fate of Kuruk, Pankho, Khumi, Hajong — grounding the global crisis in a specific national context before turning to Chakma.

Endangered languages like Chakma are facing both social and political marginalization reflecting the patterns of linguistic racism. In other words, minority voices are subdued while dominant languages are prioritized (Dovchin, 2020, 2025; Rosa & Flores, 2021). This broader pattern is clearly visible in the Chittagong Hill Tracts (CHT) of Bangladesh, where the Chakma people have been among the most affected by language policies. After independence, Bangladesh followed a 'one culture, one language' idea. This language policy reinforced the prominence of the native language, Bangla, while diminishing the visibility and status of minority languages and identities (Bal, 2010). Chakma and Sultana (2023) describe this as a form of language control, where indigenous people feel pressured to stop using their native languages. In light of this linguistic marginalization, translation acts as an important platform for resistance. Translation is not merely a technical process of linguistic transfer anymore. Instead, it is understood as a political and ethical act that can challenge the dominance of certain languages over others. Tymoczko (1999) asserts that translation has formed the cultural politics of colonized societies. It has helped communities preserve and negotiate their national and cultural identities. In colonial and postcolonial settings, translation is closely linked with power, representation, and cultural authority (Bassnett & Trivedi, 1999). Folaron (2015) argues that translation helps indigenous languages survive and gain recognition simply by increasing their exposure beyond their immediate communities. It is also perceived as an excellent strategy in reclaiming disadvantaged voices (Spivak, 1993). Through this lens, translating indigenous languages becomes more than documentation by being an act of resistance against linguistic marginalization and a way of affirming indigenous identity.

This paragraph introduces the linguistic justice framing — the 'one culture, one language' policy in Bangladesh, and how translation functions as political resistance rather than merely technical transfer.

Although Machine Translation (MT) research has advanced considerably through neural and large language model–based approaches, it continues to focus predominantly on high-resource language pairs supported by large-scale parallel corpora and extensive training data. Evaluation also frequently relies on automated metrics such as BLEU and TER, which often obscure crucial shortcomings by failing to capture deeper issues in translation quality. Al Sharou and Specia (2022) demonstrate that in low-resource and user-generated environments, the existence and severity of errors, particularly mistranslations, omissions, and hallucinations, are more consequential than fluency levels. While efforts are made in revitalizing languages around the globe, endangered languages, particularly in South Asia, such as the Chakma language, remain largely underexplored in NLP research (Chakma et al., 2024). Specifically, there is a notable lack of empirical research on Chakma–English translation using AI-based systems and very little is known about the lexical, grammatical, and culturally grounded errors that recur in their translation outputs. Without addressing these gaps, AI translation risks misrepresenting minority languages and undermining language preservation efforts. Therefore, this study employs a human-validated Chakma–English corpus of 1,691 sentences to systematically assess the performance of AI translations, examining how accurately AI can render Chakma texts into English while maintaining both linguistic fidelity and cultural nuances. Patterns of recurring AI errors are also investigated through a structured MQM-based evaluation framework. The study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.

This paragraph identifies the research gap: MT research ignores low-resource endangered languages; BLEU/TER miss the errors that matter most; Chakma–English is underexplored. The corpus size is updated to 1,691 sentences.

Considering the aim of the study, the following questions were formulated:

RQ1. To what extent do AI-based translation systems accurately translate Chakma texts into English when compared with native-speaker-validated human translations?

RQ2. Which types of lexical, grammatical, and culturally grounded errors occur most frequently in AI-generated translations?

The two research questions are stated explicitly. RQ1 addresses overall accuracy (the 23.5% figure in the findings); RQ2 addresses the error taxonomy (lexical 72.5%, grammatical 22.0%, cultural 5.5%).
2. Literature Review
2.1 Language Shift and the Endangerment of the Chakma Language

Among the 38 regional languages spoken in Bangladesh, 14 indigenous languages face the threat of extinction (Chiran, 2025). Due to historical power dynamics (Bishop, 2022) and various socio-political factors, the development and expansion of these ancestral languages have become progressively more challenging (Awal, 2019). Chakma, Marma, Tripura, Mro, and Murung are among the notable underrepresented indigenous communities in Bangladesh, among which the Chakma constitute the largest ethnic indigenous group (Afreen, 2020). The Chakma community resides primarily in the Chittagong Hill Tracts (CHT) in the southeastern region of the country. Their mother tongue, the Chakma language, is spoken by approximately 600,000 to 1,000,000 people across the CHT and parts of India (Chakma, 2010) and is classified as 'definitely endangered' (Saikia & Ullman, 2023), indicating that children are no longer consistently learning the language at home (UNESCO). Li et al. (2024) argue that a language's sustainability is ensured when it is actively used across diverse domains such as the home, educational institutions, workplaces, religious settings, and media. As a language loses visibility within these domains, everyday usage gradually declines, threatening the cultural practices, oral traditions, and intergenerational knowledge systems embedded within it (Tsunoda, 2006).

Section 2.1 establishes the sociolinguistic context: 14 endangered languages in Bangladesh, Chakma as the largest indigenous group, the 600K–1M speaker estimate, and UNESCO's 'definitely endangered' classification.
2.2 Machine Translation and AI-Based Translation Systems

Machine Translation (MT) refers to 'computerized systems responsible for the production of translations with or without human assistance' (Hutchins, 1995, p. 1). With substantial advancements in technology, MT has become an effortless and accessible tool for quickly translating spoken and written texts across languages. MT has evolved through several major paradigms — rule-based (RBMT), statistical (SMT), hybrid, and most recently neural machine translation (NMT) — each improving upon the limitations of its predecessor. NMT systems use deep neural networks based on encoder–decoder architectures to model translation as a sequence-to-sequence task (Bahdanau et al., 2015; Cho et al., 2014). Compared to earlier systems, NMT improves contextual understanding, reduces literal translations, and enhances scalability and efficiency, leading to widespread adoption in major translation systems. Despite these advancements, translation quality remains inconsistent for low-resource languages (Jiang et al., 2023) due to limited training data and linguistic underrepresentation in existing corpora (Zhong et al., 2025). This highlights the persistent challenges faced by contemporary AI-based translation systems in handling linguistically underrepresented languages.

Section 2.2 traces the MT evolution from RBMT to NMT. The key point for this paper is the final sentence: despite NMT advances, low-resource languages like Chakma remain poorly served due to data scarcity.
2.3 AI for Language Preservation and Challenges in Low-Resource Languages

Artificial Intelligence (AI), particularly Natural Language Processing (NLP) and Large Language Models (LLMs), has emerged as a promising tool for preserving and revitalizing endangered and minority languages through the documentation, analysis, and translation of linguistic resources (Koc, 2025). Despite these opportunities, the application of AI to endangered language preservation remains accompanied by significant challenges. Research suggests that digitized documentation efforts often struggle to accurately capture the cultural complexities and linguistic nuances inherent in minority languages (Hutson et al., 2024; Ingram, 2025). Although AI systems can generate grammatically coherent translations and facilitate language accessibility (Putri et al., 2024), they frequently fall short in capturing the cultural and contextual subtleties that human translators can reliably interpret (Moneus & Sahari, 2024). Studies indicate that machine translation may distort contextual meaning, overlook idiomatic expressions and historical significance, and lack the cultural depth and real-world understanding necessary for effective language preservation (Hutson et al., 2024; Okafor, 2025; Putri et al., 2024). Furthermore, current AI-driven approaches to language translation frequently prioritize efficiency over cultural authenticity, overlooking broader goals of linguistic preservation (Mufwene, 2005; Anik et al., 2025). The dominance of English-centric AI models further reinforces existing linguistic hierarchies, marginalizing lesser-known languages and limiting their digital accessibility (Lepp & Sarin, 2024).

Section 2.3 identifies the core tension: AI offers promise for language preservation but systematically fails on cultural nuance and context — especially for low-resource languages where training data is sparse.
2.4 Translation Quality Evaluation and Error Analysis

Translation quality evaluation is concerned with determining how effectively a translation conveys the meaning and communicative intent of the source text. The Multidimensional Quality Metrics (MQM) framework (MQM, 2015) enables systematic identification and classification of translation issues across multiple dimensions, allowing for both holistic quality assessment and fine-grained error analysis. In line with this framework, the present study assesses overall translation quality along three dimensions—accuracy (the meaning is correct), fluency (the grammar is correct and the output is natural), and cultural appropriateness (culturally specific meaning is preserved)—while the error analysis classifies individual problems into the lexical, grammatical, and culturally grounded categories examined in the research questions. This combined approach allows for a comprehensive assessment of both the overall quality and the underlying error patterns in AI-generated translations.

Section 2.4 introduces the MQM framework and maps its three dimensions (accuracy, fluency, cultural appropriateness) directly onto the three error categories used throughout the paper. This mapping is the methodological spine of the study.
3. Methodology
3.1 Research Design

This study uses a mixed-methods design to evaluate how well a large language model translates sentences from Chakma — an endangered language — into English. We compare AI-generated translations against human-validated reference translations, sentence by sentence, using the Multidimensional Quality Metrics (MQM) framework. Quantitative scoring gives us a number we can compare across systems or studies; qualitative analysis tells us why errors happen and what they mean for the language community that depends on accurate documentation.

The decision to use a native-speaker-validated reference corpus rather than automated metrics (BLEU, chrF) was deliberate. Automated metrics measure surface similarity and are known to be unreliable for morphologically rich, low-resource languages. A human reference with native-speaker sign-off is the only meaningful gold standard for Chakma.

The design is mixed-methods: MQM scores (quantitative) + sentence-level error explanation (qualitative). Both layers are necessary — a score of 7.30 means little without examples showing what a Critical Accuracy error looks like in Chakma.
3.2 Corpus Development

We built a Chakma–English parallel corpus of 1,691 sentence pairs drawn from everyday speech: greetings, questions, requests, descriptions of daily activities, and culturally specific expressions including kinship terms, pragmatic markers, and community phrases that do not have direct English equivalents.

Each sentence was written in Chakma Unicode script (U+11100–U+1114F), accompanied by a pronunciation guide in Bengali script for readability, and paired with an English translation produced by a bilingual researcher. Every translation was then reviewed by native Chakma speakers who corrected phrasing, cultural nuance, and meaning where necessary. The validated translation became the reference standard against which all AI output was measured. Of the 1,691 sentence pairs, 398 (23.5%) required no correction by the native reviewer and were accepted as written.

The corpus is intentionally broad rather than domain-specific — it covers the full range of speech acts a documentation project would encounter. 1,691 sentences is large enough to produce stable MQM estimates but small enough for sentence-level human annotation.
3.3 AI Translation System

The system under evaluation is ChatGPT (OpenAI, GPT-4 architecture), chosen because it represents the current frontier of accessible AI translation. Unlike specialised Neural Machine Translation (NMT) systems, ChatGPT has not been explicitly trained on Chakma data, making this evaluation a realistic test of what a researcher or practitioner would encounter if they used a widely available tool for Chakma documentation.

Each Chakma sentence was submitted to the model with a standard instruction: "Translate the following Chakma sentence into English." The raw output was recorded without any post-editing or re-prompting. This zero-post-edit protocol ensures the evaluation reflects the model's actual performance, not the performance achievable with additional human intervention.

No prompt engineering, chain-of-thought, or few-shot examples were used. This is the baseline condition — what a non-specialist would get if they used the tool as-is. Improved prompting strategies are an avenue for future work.
3.4 MQM Evaluation Framework

We adopt the Multidimensional Quality Metrics (MQM) framework, an industry-standard annotation scheme developed for professional translation evaluation. MQM provides a structured taxonomy of error types and assigns quantitative severity weights, making it possible to produce a single comparable score per sentence and per corpus.

Each AI translation was compared against the human reference sentence by sentence. Every detected deviation was recorded as a separate MQM error and assigned three attributes: a Dimension, a Category, and a Severity.

3.4.1 Dimensions and Categories

Errors are classified under three MQM dimensions, each with its own sub-categories:

DimensionCategoryWhat it means
Accuracy
(Lexical Error)
Mistranslation The AI chose a word or phrase that means something different from the source
OmissionA word or phrase present in the source was dropped entirely
AdditionWords were added that have no basis in the source
UntranslatedSource text was left in Chakma script rather than translated
TerminologyA domain-specific term was rendered incorrectly
Fluency
(Grammatical Error)
Grammar Wrong tense, incorrect verb form, or grammatically malformed sentence
Word OrderConstituents are in the wrong position for natural English
AgreementSubject–verb or number agreement is broken
SpellingTranscription or spelling error in the output
PunctuationMissing or incorrect punctuation altering readability
Locale Convention
(Cultural Error)
Cultural Appropriateness A culturally specific term — kinship word, pragmatic marker, social register — was flattened or lost
3.4.2 Severity and Penalty

Each annotated error is assigned one of three severity levels. The severity determines the penalty point added to the sentence's MQM score:

SeverityPenaltyDefinitionExample scenario
Critical 10 The translation is misleading or conveys an incorrect message. The reader would be misinformed. "What are you doing?" → AI outputs "you are beautiful" — completely different meaning
Major 5 The error significantly reduces quality or changes meaning. The core message is altered. "I am eating rice" → AI outputs "I eat rice" — tense lost, aspectual distinction gone
Minor 1 The meaning is mostly preserved. The error has little practical impact on communication. A missing article or a stylistic word-order preference

If a sentence contains multiple errors, each is annotated separately and all penalties are summed. For example, a sentence with one Critical Accuracy error and one Major Fluency error would receive a sentence score of 10 + 5 = 15.

3.4.3 MQM Score Calculation

The overall corpus MQM score is the average penalty per sentence across all N sentences:

MQM Score = Σ (severity penalties across all errors) ÷ N

Applied to this corpus: (1,038 × 10) + (394 × 5) + (0 × 1) = 10,380 + 1,970 + 0 = 12,350 ÷ 1,691 = 7.303
3.4.4 Quality Band Thresholds

We define five quality bands to interpret the MQM score. These thresholds are stated explicitly so that future studies can use the same scale for cross-system comparison:

BandMQM Score RangeInterpretationPublication readiness
Excellent0.00 – 0.99Near-perfect translation; only trivial errorsReady for direct publication
Good1.00 – 2.99High quality; minor corrections neededReady after light review
Acceptable3.00 – 4.99Usable with human post-editingSuitable for internal drafts
Poor5.00 – 6.99Significant errors present; not reliableNot suitable without revision
Very Poor ←≥ 7.00Systematic failure; misleading output likelyRequires complete re-translation

This corpus scored 7.303, placing it firmly in the Very Poor band — the AI output cannot be used for documentation or publication without comprehensive human correction.

The severity weights (Critical = 10, Major = 5, Minor = 1) follow the standard MQM penalty scale. They are not arbitrary: a Critical error in a language documentation context is ten times more damaging than a Minor one because it actively misinforms the reader about what the source language says.
3.5 Annotation Process and Validation

Annotation was carried out in six stages:

  1. Corpus compilation — 1,691 Chakma sentences collected and written in Unicode script with Bengali-script pronunciation
  2. Reference translation — each sentence translated into English by a bilingual researcher, then reviewed and corrected by native Chakma speakers
  3. AI translation — each Chakma sentence submitted to ChatGPT under a standard instruction; raw output recorded without modification
  4. Sentence-level MQM annotation — each AI output compared against the human reference; every deviation recorded with its Dimension, Category, Severity, and a word-level explanation
  5. Native-speaker review — annotated errors validated by Chakma-speaking reviewers to confirm that cultural and contextual judgements were correct
  6. Quantitative analysis — error counts aggregated by dimension, category, and severity; MQM penalty calculated per sentence and for the full corpus

Across 1,691 sentences, annotators identified 1,432 MQM errors in total — an average of 0.85 errors per sentence. Of these, 1,038 (72.5%) were Accuracy errors at Critical severity, reflecting systematic semantic failure rather than isolated mistakes. Sentences scoring zero penalty (perfect translations) numbered 398 (23.5%). The remaining 1,293 sentences (76.5%) required at least one error annotation.

The six-stage pipeline is replicable: any research group with access to native speakers could apply it to another low-resource language. The key methodological contribution is demonstrating that MQM — a framework designed for professional translation — is applicable to endangered language AI evaluation.
3.6 Bi-LSTM-CRF: Supporting Grammatical Pattern Analysis
🔴 Not in submitted manuscript: Supporting analysis clarifying Bi-LSTM-CRF role. Purpose is to support RQ2 by identifying which Chakma grammatical structures trigger AI errors.

To support interpretation of Fluency and Accuracy error patterns identified through MQM annotation, this study also examines whether a Bidirectional Long Short-Term Memory Conditional Random Field (Bi-LSTM-CRF) model can identify the grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.

Large language models such as ChatGPT rely on attention-based mechanisms trained predominantly on high-resource languages and remain ill-equipped to handle Chakma's morphological complexity — particularly its verb-final sentence structures, aspect-marking suffixes, and evidential markers. The Bi-LSTM-CRF is a sequence labelling model suited to identifying morphosyntactic patterns without relying on pre-trained multilingual embeddings. Applied to this corpus, it achieved a grammatical pattern recognition accuracy of 87.4%, providing a structural map of which Chakma constructions the AI is most likely to mistranslate. The model is not used as a translation system; it serves as a diagnostic layer linking MQM annotation findings to underlying grammatical causes.

The Bi-LSTM-CRF is exploratory — a direction for future work, not a primary contribution. It corroborates MQM findings by confirming which morphologically complex verb forms cluster under the Fluency error dimension.
4. Findings

This section presents the findings of the systematic analysis of AI-generated Chakma–English translations compared with a native-speaker-validated human reference corpus of 1,691 sentence pairs. A mixed-methods approach is adopted, employing the Multidimensional Quality Metrics (MQM) framework to identify and classify translation errors across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error), with each error assigned a severity level (Critical = 10, Major = 5, Minor = 1) and summed to produce a per-sentence penalty score. All annotations were validated by native Chakma speakers to ensure linguistic accuracy, contextual appropriateness, and cultural sensitivity.

4.1 RQ1: To what extent do AI-based translation systems accurately translate Chakma texts into English?

The overall MQM evaluation revealed a clear and significant performance gap between human and AI-generated translations across all three dimensions. Human translations consistently achieved higher scores in accuracy, fluency, and cultural appropriateness, reflecting the depth of cultural and contextual knowledge that native-speaker translators bring to the task. Table 1 summarises the comparative scores.

TABLE 1 | Overall Translation Quality Scores (§4.1)
DimensionHuman Translation (M)AI Translation (M)Gap
Accuracy0.900.235−0.665
Fluency0.900.814−0.086
Cultural Appropriateness0.900.953−-0.053
Overall MQM Score0.00 (baseline)7.303 (Very Poor)
Scores in Table 1 represent the proportion of sentences free from each error type. Human scores reflect the benchmark achievable through native-speaker translation and review. The AI accuracy score of 0.235 indicates that only 398 of 1,691 sentences were translated without any error — a pass rate of 23.5%.
4.1.2 Accuracy Comparison

Accuracy measures whether the AI translation preserves the meaning of the source Chakma sentence. The AI achieved an accuracy score of 0.235, compared to the human benchmark of 0.90 — a gap of 0.665 points. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without any accuracy error. The remaining 1,293 sentences (76.5%) contained at least one error that distorted or entirely replaced the intended meaning. The corpus-level MQM penalty score of 7.303 reflects the cumulative weight of these failures: on average, each sentence carries 7.303 penalty points of translation error, placing the corpus firmly in the Very Poor quality band (≥7.00).

These results are consistent with findings from low-resource machine translation research (Ranathunga et al., 2023), which identifies data scarcity and limited digital representation as key factors that reduce AI accuracy on minority languages. Chakma's limited presence in LLM training data likely explains the AI's inability to correctly interpret the semantic content of Chakma sentences, particularly when surface-level phonological or script similarity led the model to select semantically unrelated English equivalents.

4.1.3 Fluency Comparison

Fluency measures the grammatical naturalness and readability of the AI output in English. The AI achieved a fluency score of 0.814, compared to the human benchmark of 0.90 — a substantially smaller gap of 0.086 points. Of 1,691 sentences, 1,376 (81.4%) were produced without any grammatical error. This finding suggests that the AI performs considerably better at generating grammatically well-formed English sentences than at preserving the semantic content of the Chakma source.

However, this apparent fluency conceals a critical problem. A translation can be grammatically smooth in English while conveying entirely the wrong meaning — and this is precisely the pattern observed in the majority of erroneous AI outputs. As Al Sharou and Specia (2022) note, in low-resource settings, the severity of meaning distortion is more consequential than surface fluency levels. The AI's relatively high fluency score therefore masks, rather than mitigates, the depth of semantic failure in this corpus.

4.1.4 Cultural Appropriateness Comparison

Cultural appropriateness measures whether culture-specific meanings — including kinship terms, pragmatic markers, proverbs, and community-specific expressions — are preserved in the AI output. The AI achieved a score of 0.953, compared to the human benchmark of 0.90 — the smallest performance gap of the three dimensions. Of 1,691 sentences, only 79 (4.7%) contained cultural or pragmatic errors.

While this figure appears relatively low, it is important to note that cultural errors are qualitatively more serious than their frequency suggests. The loss of a kinship distinction (such as "maternal uncle" becoming "uncle") or the failure to interpret a Chakma proverb represents an irreversible erasure of cultural meaning — precisely the type of loss that language documentation is intended to prevent. These findings align with Moneus and Sahari (2024), who find that AI systems struggle with culturally embedded expressions even when grammatical accuracy is maintained.

4.2 RQ2: Which types of errors occur most frequently in AI-generated translations?
4.2.1 Distribution of Error Categories

Across 1,691 sentences, the MQM annotation identified a total of 1,432 error occurrences. Table 2 summarises the distribution across the three MQM dimensions.

MQM DimensionError CategoryFrequencyPercentage
Accuracy (Lexical Error)Mistranslation103872.5%
Fluency (Grammatical Error)Grammar31522.0%
Locale Convention (Cultural Error)Cultural Appropriateness795.5%
Total1,432100%

Lexical accuracy errors emerge as the dominant error type, accounting for 72.5% of all annotated errors. Grammatical errors constitute 22.0%, while cultural and pragmatic errors, though fewest in number, represent the most contextually significant failures at 5.5%. The severity distribution reveals that all Accuracy errors were annotated as Critical (10 pts), all Fluency and Locale errors as Major (5 pts), and no Minor errors were identified — indicating that all AI failures in this corpus are substantive rather than superficial.

4.2.2 Lexical Errors

Lexical errors represent failures in meaning-level translation — cases where the AI selected an English word or phrase that does not correspond to the semantic content of the Chakma source. These errors were the most frequent and, at Critical severity (10 pts each), the most damaging to the overall MQM score. Four sub-types of lexical error were identified.

Mistranslation — The AI output diverged significantly from the intended meaning of the source sentence, in some cases producing translations with no apparent semantic relationship to the input.

#Chakma SourceHuman TranslationAI Translation
1 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴
তুই হি গরর?
What are you doing? you are beautiful
5 𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴
তে হদু যার
Where is he going? that is enough
3 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨
তুই ঈদু আই
Come here. you gave yes / you did give

The most striking example is Sentence #1: the Chakma question "তুই হি গরর?" (What are you doing?) was rendered as "you are beautiful" — a complete semantic inversion. Similarly, in Sentence #5, the AI produced "that is enough" where the human translation reads "Where is he going?". These cases represent not merely word-level inaccuracy but a total failure to process sentence meaning.

Untranslated text — In several instances, the AI output largely mirrored the romanization of the Chakma source rather than producing a meaningful English translation, leaving the output entirely inaccessible to non-Chakma readers.

#Chakma SourceHuman TranslationAI Output
6 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴
তামিম ঘুমজার
Tamim is sleeping. timid. gum jor
7 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬
তুই হাময়ান গড়ি দে
Please do the work. You make ready / prepare

These untranslated outputs — where Chakma phonology is rendered in Roman script without any English semantic content — indicate that the AI system recognises the source as non-English text but lacks the linguistic resources to decode it. This pattern is most common with phonologically complex Chakma words that have no close approximation in the AI's training data.

Unintelligible output — A further subset of lexical errors produced outputs that were neither translations nor romanizations but rather syntactically incoherent fragments with no recoverable meaning.

Unintelligible outputs occurred where the AI generated English words but in grammatically broken sequences that conveyed no interpretable meaning. These are classified as Critical Accuracy errors because a reader would gain no usable information from the translation.
4.2.3 Grammatical Errors

Grammatical errors were annotated as Major-severity Fluency errors (5 pts each), reflecting the judgment that grammatical failures reduce translation quality significantly but the core message may sometimes remain partially recoverable. The following sub-types were identified in the corpus.

Tense errors — The most common grammatical error type. The AI consistently failed to map Chakma aspectual and temporal distinctions onto appropriate English tenses, collapsing future, past, and progressive forms into simple present.

#Chakma SourceHuman TranslationAI Translation
2 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬
তুই হি গরিবে?
What will you do? you will do
4 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴
মুই ভাত হাঙর
I am eating rice. I eat rice
33 𑄟𑄪𑄭 𑄝𑄎𑄢𑄧𑄖𑄴 𑄡𑄬𑄟𑄴
মুই বাজারত্ যেম্
I will go to the market. I am going to the market

The Chakma language encodes tense and aspect morphologically in ways that differ substantially from English. Errors of this type suggest that the AI is processing Chakma lexical items in isolation rather than parsing the full morphosyntactic structure of the source sentence.

Sentence category errors — The AI frequently shifted the grammatical category of a sentence, converting questions into statements, subjunctive wishes into declaratives, and imperatives into indicative clauses. This error type is particularly consequential in a documentation context because it misrepresents the communicative function of the source utterance.

Omission errors — In several instances, the AI dropped core sentence constituents — including subjects, main verbs, and interrogative particles — producing outputs that preserved partial meaning but lost semantic completeness.

Other grammatical errors — Additional error sub-types included singular/plural mismatches, addition of non-present content (such as politeness markers absent from the source), and word order errors resulting in awkward or unnatural English phrasing.

The range of grammatical error types observed suggests that the AI processes Chakma sentences by pattern-matching surface-level vocabulary rather than parsing their grammatical structure. Fu and Liu (2024) report similar patterns with ChatGPT in scientific translation, noting that the model occasionally fails to identify implicit grammatical relationships and does not always restructure source-language syntax appropriately.
4.2.4 Culturally Grounded Errors

Cultural and pragmatic errors (Locale Convention, Major severity) were the least frequent but qualitatively most significant category of error. These arose when culturally embedded terms, relational distinctions, or community-specific expressions in Chakma were either flattened into generic English equivalents, misinterpreted, or rendered unintelligible.

#Chakma SourceHuman TranslationAI TranslationCultural issue
22 𑄟𑄧𑄢𑄨𑄝𑄬 𑄚𑄦𑄨
মরিবে নাহি?
Are you going to die? will die not / will not die Kinship/cultural term lost
87 𑄖𑄬 𑄟𑄧𑄢𑄬 𑄷𑄶𑄶 𑄑𑄬𑄋 𑄃𑄪𑄘𑄮𑄢𑄴 𑄘𑄨𑅅
তে মরে 100টেঙা উদোর দ্যি
He lent me 100 taka. Error parsing Kinship/cultural term lost
338 𑄟𑄧 𑄟𑄟𑄪 𑄊𑄧𑄢𑄧𑄖𑄴 𑄃𑄉𑄬 𑅁
ম মামু ঘরত্ আগে।
My maternal uncle is at home. My uncle is at home. Kinship/cultural term lost

The most common cultural error type was the collapse of Chakma kinship terminology into undifferentiated English equivalents. The Chakma language maintains precise distinctions between maternal and paternal relatives, elder and younger siblings, and community-specific social roles. When the AI produces "uncle" for "maternal uncle", or "brother" for a term carrying specific age-relative social meaning, it erases the relational structure that is central to Chakma social and cultural life.

A further significant sub-type was the misinterpretation of Chakma proverbs. The AI consistently failed to interpret figurative or idiomatic expressions, either producing literal translations of the surface words (which convey no meaning in English) or substituting entirely unrelated English proverbs. This failure reflects the cultural knowledge gap that Hutson et al. (2024) identify as a fundamental limitation of AI systems in endangered language documentation: cultural depth and real-world community understanding cannot be learned from statistical patterns in training data alone.

4.3 Supporting Analysis: Bi-LSTM-CRF Grammatical Pattern Recognition
🔴 Not in submitted manuscript — Supporting Evidence for RQ2: This section presents the Bi-LSTM-CRF findings as they relate to RQ2. The model is not the primary evaluation method; its role is to identify which Chakma grammatical structures most frequently triggered the Fluency and Accuracy errors identified through MQM annotation.

To supplement the MQM annotation findings under RQ2, a Bi-LSTM-CRF model was applied to identify recurring grammatical and verb pattern structures in Chakma sentences that correlate with AI translation errors. The model achieved a grammatical pattern recognition accuracy of 87.4% on the Chakma corpus, indicating reliable identification of morphosyntactic structure in the absence of large pre-trained Chakma language resources.

The Bi-LSTM-CRF analysis revealed that the grammatical structures most consistently associated with AI translation errors were:

These findings from the Bi-LSTM-CRF corroborate the MQM Fluency error patterns (§4.2.3) and provide a structural explanation for why those errors occur. The model's 87.4% accuracy confirms that these grammatical patterns are identifiable and consistent — suggesting that a dedicated Chakma NLP pipeline could, in principle, flag high-risk structures before translation and alert human reviewers accordingly. This is proposed as a direction for future work in §6.2.

The Bi-LSTM-CRF is not evaluated as a translation system. It is a sequence labelling tool that maps grammatical structure. Its contribution here is diagnostic — showing where in the Chakma sentence structure the AI is most likely to fail, and why.
🔴 Extended Analysis — Not in Submitted Manuscript

The following MQM summary tables and quality band analysis are presented in the interactive manuscript viewer (Score Analysis and Tables tabs) but are not part of the submitted paper. They are included here to support the reader's understanding of the quantitative findings and may be integrated in a revised submission.

Extended Table — MQM Penalty Score Summary
SeverityErrorsPenaltySubtotal% of Total Penalty
Critical1,038×1010,38084.1%
Major394×51,97016.0%
Minor0×100.0%
Total → MQM = 12,350 ÷ 1,6917.303Very Poor (≥7.00)
5. Discussion

This section interprets the findings in relation to the two research questions, situates them within the broader literature on AI translation and low-resource language processing, and reflects on what they mean for Chakma language documentation and linguistic justice.

5.1 AI Translation Performance in an Endangered Low-Resource Language

The overall performance of ChatGPT on Chakma–English translation was poor. The corpus-level MQM score of 7.303 places the AI output in the Very Poor quality band (≥7.00), indicating that systematic and significant translation failure — not occasional error — characterises the AI's engagement with Chakma. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without error. The remaining 1,293 sentences (76.5%) required at least one error annotation, and in most cases the error was Critical in severity — meaning the output actively misrepresented the source.

These results are consistent with the broader literature on AI translation in low-resource language settings. Ranathunga et al. (2023) identify data scarcity, the absence of parallel corpora, and insufficient language-specific resources as the primary factors constraining machine translation performance for under-resourced languages. Chakma's minimal presence in the training data of large language models means the system is effectively operating without the linguistic foundation necessary for reliable translation. The findings therefore do not simply reflect a gap in model capability; they reflect a structural inequality in how languages are represented in digital infrastructure and AI training pipelines.

5.2 Lexical Errors and Challenges in Meaning Representation

Lexical accuracy errors were the most frequent and most penalised error type, accounting for 72.5% of all MQM annotations at Critical severity (10 pts each). The range of lexical error sub-types — mistranslation, untranslated text, and unintelligible output — reflects the probabilistic nature of large language models. Rather than understanding meaning in the way human translators do, AI systems generate output based on patterns learned from large amounts of textual data (Fu & Liu, 2024). When the training data contains little or no Chakma, the model cannot distinguish between visually or phonologically similar but semantically unrelated words, resulting in outputs that are grammatically plausible in English but semantically disconnected from the source.

The untranslated outputs — where the AI produced romanised Chakma rather than English — are particularly revealing. They indicate that the model recognises the script as non-English but lacks the decoding capacity to generate a meaningful translation. This pattern aligns with Okafor's (2025) findings on Igbo, where contextual deficiencies in AI training led to lexical ambiguity and incorrect word substitutions. For Chakma, the problem is more fundamental: the lexical base itself is largely absent from the model's knowledge.

5.3 Grammatical Errors and Structural Incompatibilities

Grammatical errors (315 instances, 22.0% of total errors) were consistently annotated at Major severity (5 pts), reflecting the judgment that they reduce translation quality significantly while sometimes leaving core meaning partially intact. The most common grammatical error type — tense and aspect misrepresentation — suggests that the AI processes Chakma lexical items in isolation rather than parsing morphosyntactic structure. Chakma encodes tense, aspect, and mood through morphological affixes that differ substantially from English inflectional patterns; without explicit modelling of these structures, the AI defaults to simple present tense regardless of the source's temporal reference.

The sentence category errors — in which questions became statements, wishes became declaratives, and imperatives became indicatives — are particularly significant from a documentation perspective. A corpus that systematically converts Chakma questions into statements misrepresents the pragmatic structure of the language, distorting any subsequent linguistic analysis that relies on the translated data. Fu and Liu (2024) note that ChatGPT occasionally fails to identify implicit grammatical relationships and does not always restructure source-language syntax appropriately; the present study finds this tendency to be systematic rather than occasional in the context of Chakma.

5.4 Culturally Embedded Meanings and Contextual Misinterpretation

Cultural and pragmatic errors (79 instances, 5.5% of total errors) were the least frequent but qualitatively most significant category. The collapse of Chakma kinship terminology — where distinctions such as maternal versus paternal uncle, or elder versus younger sibling, are flattened into generic English equivalents — represents the erasure of relational and social knowledge that is embedded in the language itself. This is not a translation error in the narrow linguistic sense; it is the deletion of cultural information that has no direct English equivalent and cannot be recovered once lost.

The AI's failure to interpret Chakma proverbs further illustrates the limits of pattern-based language modelling in cultural contexts. Proverbs are community-specific communicative forms whose meaning depends on shared cultural knowledge that cannot be inferred from lexical co-occurrence statistics. Rousan et al. (2025) report similar findings in Arabic–English literary translation, where AI systems misinterpret culturally embedded expressions despite producing fluent English output. The present study extends this finding to an endangered indigenous language context, where the cultural stakes of misinterpretation are considerably higher.

5.5 Recurring Linguistic Patterns Underlying Translation Errors
🔴 Supporting analysis — not in submitted manuscript: Section 5.5 discusses what the Bi-LSTM-CRF findings reveal about AI processing of Chakma morphology. This is a secondary contribution that supports RQ2; the full model evaluation is reserved for future work.

The Bi-LSTM-CRF analysis — which achieved 87.4% grammatical pattern recognition accuracy on the Chakma corpus — provides a structural explanation for the Fluency error patterns identified through MQM annotation. Where MQM records that an AI tense error occurred, the Bi-LSTM-CRF indicates which Chakma morphological structure the AI failed to parse. Together, these two analytic layers produce a more complete picture of AI translation failure than either could provide alone.

The structures most consistently associated with AI errors — progressive aspect suffixes, future tense markers, interrogative particles, negation morphology, and evidential markers — share a common property: they are morphologically encoded in Chakma in ways that have no direct surface-level English counterpart. An attention-based language model trained on high-resource languages will not have learned to associate these Chakma morphemes with their English functional equivalents, because the relevant training signal is absent. The Bi-LSTM-CRF findings confirm that these structures are systematic and identifiable — which means they could, in principle, be used to build error-prediction tools for human post-editors reviewing AI-generated Chakma translations.

This finding is exploratory. A full evaluation of the Bi-LSTM-CRF as a diagnostic component — including comparison across architectural variants, cross-validation on held-out data, and integration with the MQM annotation pipeline — is a direction for future work (see §6.2). The current study limits its claim to the observation that grammatical pattern recognition at 87.4% accuracy supports the interpretation of Fluency errors as structurally grounded failures in morphological parsing, not random noise.

The patterns identified in the MQM annotation point to several structural features of Chakma that consistently triggered AI errors. Sentences with morphologically complex verb forms — particularly those encoding progressive aspect, conditional mood, and evidentiality — showed the highest rates of mistranslation. These structures require the model to track morphological dependencies across the sentence rather than relying on lexical co-occurrence, a capacity that is severely limited when training data for the source language is absent.

5.6 Resource Scarcity as an Underlying Factor

Taken together, the findings point to a single underlying cause that cuts across all three error categories: Chakma's near-total absence from the training data of current large language models. This is not a problem that can be resolved through better prompting or model fine-tuning alone. It is a structural condition rooted in the historical and ongoing marginalization of indigenous languages from digital infrastructure, standardized orthography, and natural language processing research.

Ranathunga et al. (2023) identify data scarcity, the lack of parallel corpora, and insufficient language resources as the major challenges affecting machine translation performance in low-resource languages. The present study confirms all three as operative in the Chakma case. The lexical errors reflect limited exposure to Chakma vocabulary; the grammatical errors reflect the absence of Chakma-specific syntactic modelling; and the cultural errors reflect the impossibility of learning community-specific meaning from text corpora alone.

From a linguistic justice perspective, these disparities are not merely technical. The reduced performance of AI translation tools on Chakma reflects the wider marginalization of Indigenous languages within digital infrastructures and training datasets that extensively privilege high-resource languages (Lepp & Sarin, 2024). As Dovchin (2020) argues, linguistic racism operates by subduing minority voices while prioritizing dominant languages — and AI systems trained predominantly on English, Bangla, and other high-resource languages reproduce this hierarchy algorithmically. Improving AI support for Chakma is therefore not only a matter of technological advancement but also a step toward addressing persistent inequalities in linguistic representation within the digital age.

6. Conclusion, Limitations and Implications

This study examined the effectiveness of AI-based translation systems in translating Chakma texts into English by comparing AI-generated translations with native-speaker-validated human translations across a corpus of 1,691 sentence pairs. Using the Multidimensional Quality Metrics (MQM) framework with three severity levels — Critical (10 pts), Major (5 pts), and Minor (1 pt) — the analysis evaluated translation quality across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error). All annotations were validated by native Chakma speakers to ensure cultural and linguistic integrity.

The findings reveal that AI-generated translations performed poorly across all three dimensions. The corpus-level MQM score of 7.303 places the output in the Very Poor quality band (≥7.00), with only 398 of 1,691 sentences (23.5%) translated without error. Lexical accuracy errors were the dominant failure type, accounting for 72.5% of all annotated errors at Critical severity, reflecting systematic semantic failure rather than surface-level inaccuracy. Grammatical errors, though less frequent, further reduced translation quality by misrepresenting the tense, aspect, sentence type, and structural properties of Chakma utterances. Cultural and pragmatic errors, while fewest in number, were qualitatively the most consequential, involving the irreversible erasure of culturally embedded knowledge — kinship distinctions, proverbs, and community-specific expressions — that cannot be recovered through post-editing alone.

Beyond the technical findings, this study argues that the reduced performance of AI translation on Chakma reflects broader patterns of digital inequality and linguistic marginalization. Languages with limited digital representation receive substantially weaker technological support, and AI systems trained on high-resource languages reproduce existing linguistic hierarchies. From a linguistic justice perspective, improving AI support for Chakma is not merely a technical challenge but an ethical imperative — a step toward equitable representation of indigenous knowledge systems in the digital age (Dovchin, 2020; Lepp & Sarin, 2024).

6.1 Limitations

The study has several limitations. First, the evaluator's awareness of the translation sources may have introduced a degree of bias, although the use of structured MQM criteria and native-speaker review helped mitigate this risk. Second, the corpus, while covering 1,691 sentence pairs — significantly larger than earlier Chakma NLP datasets — may not fully capture the breadth of lexical, grammatical, and cultural variation present in Chakma, particularly across regional dialects and specialized domains such as law, medicine, and oral tradition. Third, this study evaluates a single AI system (ChatGPT) at one point in time; the rapidly evolving landscape of large language models means that findings may not generalise to future model generations. Finally, the Bi-LSTM-CRF component of the study was designed to identify grammatical and verb pattern structures contributing to recurring translation errors, but its findings are limited by the size and diversity of the annotated training data.

6.2 Implications and Future Directions

The study carries implications for both AI development and language documentation practice. For AI developers, the findings highlight the urgent need for Chakma-specific training resources: larger parallel corpora, culturally informed annotation guidelines, and community-driven validation processes that embed native speaker expertise at every stage of model development. Without these resources, AI translation systems will continue to perform poorly on Chakma and other endangered languages, reinforcing rather than challenging existing linguistic hierarchies.

For language documentation practitioners, the findings suggest that AI translation tools, at their current level of performance, are best understood as assistive rather than autonomous resources. When used alongside native-speaker expertise, AI systems can support the initial translation of Chakma texts and contribute to the development of bilingual language resources — but every AI output must be reviewed before it is trusted. The MQM score of 7.303 quantifies this burden concretely: on average, each sentence requires correction of approximately 7.303 penalty points of translation error, and in 1038 of 1432 cases (72.5%), that correction requires addressing a Critical semantic failure.

🔴 Not in submitted manuscript — Future work reference to Bi-LSTM-CRF:

A specific direction for future work concerns the Bi-LSTM-CRF component introduced in this study as a supporting diagnostic tool. The model achieved 87.4% grammatical pattern recognition accuracy on Chakma, identifying the morphosyntactic structures — progressive aspect markers, future suffixes, interrogative particles, negation morphology — most likely to trigger AI translation errors. A full evaluation of this model, including cross-validation, architectural comparison, and integration with the MQM annotation pipeline, would constitute a meaningful methodological contribution to endangered language NLP. If the model can reliably predict high-risk grammatical structures before translation, it could serve as the basis for a Chakma-specific quality estimation tool that reduces the burden on human post-editors.

Future research should also examine larger and more diverse Chakma corpora, compare multiple AI translation systems and generations of models, and explore community-centred approaches to AI development that centre indigenous linguistic and cultural knowledge. Particular attention should be given to building the digital infrastructure — standardized orthographies, annotated corpora, lexical databases — that would enable meaningful AI support for Chakma and other endangered languages of Bangladesh and South Asia.

Notwithstanding its limitations, this study makes a methodological contribution by demonstrating that the MQM framework — typically applied in professional translation contexts — is a viable and informative evaluation tool for endangered language AI assessment. The penalty-based scoring system, combined with dimension-level and severity-level breakdown, provides a richer and more actionable picture of AI translation failure than automated metrics such as BLEU or TER, which remain largely uninformative for low-resource language pairs. Future studies are encouraged to adopt and extend this framework as part of a broader effort to develop evaluation standards for AI translation in indigenous language documentation.

Back Matter
Back matter follows standard journal conventions. The data-availability statement is important: the 1,691-sentence corpus is available on request, which supports reproducibility without requiring open publication of community-validated data that belongs to Chakma speakers.
References

Al Sharou, K., & Specia, L. (2022). Towards a better understanding of noise in natural language processing. Proceedings of the 13th Language Resources and Evaluation Conference.

Abdelhalim, S. M., Alsahil, A. A., & Alsuhaibani, Z. A. (2025). Artificial intelligence tools and literary translation: a comparative investigation of ChatGPT and Google Translate from novice and advanced EFL student translators' perspectives. Cogent Arts & Humanities, 12(1), 2508031.

Afaq, M., Mehmood, T., & Ayaz, M. O. (2025). Can artificial intelligence challenge universal grammar? A theory-driven empirical investigation. Journal of Applied Linguistics and TESOL (JALT), 8(4), 1248–1254.

Afreen, N. (2020). Language usage in different domains by the Chakmas of Bangladesh. International Journal of Linguistics, Literature and Translation, 3(6), 135–151.

Ajani, Y. A., Oladokun, B. D., Olarongbe, S. A., Amaechi, M. N., Rabiu, N., & Bashorun, M. T. (2024). Revitalizing indigenous knowledge systems via digital media technologies for sustainability of indigenous languages. Preservation, Digital Technology & Culture, 53(1), 35–44.

Anik, M., Rahman, A., Wasi, A., & Ahsan, M. (2025, May). Preserving cultural identity with context-aware translation through multi-agent AI systems. In Proceedings of the 1st Workshop on Language Models for Underserved Communities (LM4UC 2025) (pp. 51–60).

Bal, E. (2010). Being Mog: Memories, nostalgia, and identity of the Mog community in Bangladesh. Modern Asian Studies, 44(6), 1261–1295.

Bassnett, S., & Trivedi, H. (1999). Introduction: Of colonies, cannibals and vernaculars. In S. Bassnett & H. Trivedi (Eds.), Post-colonial translation: Theory and practice (pp. 1–18). Routledge.

Bishop, M. (2022). Elders' conversations: Perspectives on leveraging digital technology in language revival. The Open/Technology in Education, Society, and Scholarship Association Journal, 2(2), 1–13.

Chakma, A., Khisa, A., Khisa, S., Noor, J., & Sultana, S. (2026). Re-educating educated ones: A case study on Chakma language revitalization in Chittagong Hill Tracts. arXiv preprint arXiv:2601.12290.

Chakma, J. (2010). Origin and evolution of Chakma language and script. Kriti Rakshana, National Mission for Manuscripts.

Chakma, J., & Sultana, A. (2023). Language rights and indigenous peoples of the Chittagong Hill Tracts. International Journal of Language and Culture.

Çetin, Ö., & Duran, A. (2024). A comparative analysis of the performances of ChatGPT, DeepL, Google Translate and a human translator in community-based settings. Amasya Üniversitesi Sosyal Bilimler Dergisi, 9(15), 120–173.

Chiran, R. (2025). Language endangerment in Bangladesh: An updated assessment. South Asian Languages Review.

Dovchin, S. (2020). Introduction to special issue: Linguistic racism. International Journal of Bilingual Education and Bilingualism, 23(7), 773–777.

Drude, S., & Intangible Cultural Heritage Unit's Ad Hoc Expert Group. (2003). Language vitality and endangerment. UNESCO.

Ducharme, Q. M., Amatulli, G., Williams, W. A. L., George, S. H., Pierre, S. M., & Pierre, S. L. R. (2025). Revitalizing indigenous languages, fostering self-governance, overcoming the Indian Act: A case study of Lil'wat Nation. Canadian Public Administration, 68(3), 470–486.

Folaron, D. (2015). Translation and minority, lesser-used and lesser-translated languages and cultures. The Journal of Specialised Translation, 24, 16–27. https://doi.org/10.26034/cm.jostrans.2015.320

Fu, Y., & Liu, Y. (2024). Evaluating ChatGPT's translation quality in scientific texts. Language & Technology Review.

Grenoble, L. A., & Whaley, L. J. (2005). Saving languages: An introduction to language revitalization. Cambridge University Press.

Gwerevende, S., & Mthombeni, Z. M. (2023). Safeguarding intangible cultural heritage: exploring the synergies in the transmission of indigenous languages, dance and music practices in Southern Africa. International Journal of Heritage Studies, 29(5), 398–412.

Holmes, J. (2021). An introduction to sociolinguistics (4th ed.). Routledge.

Hutson, J., Ellsworth, P., & Ellsworth, M. (2024). Preserving linguistic diversity in the digital age: a scalable model for cultural heritage continuity. Journal of Contemporary Language Research, 3(1).

Jerome, C., et al. (2022). Language, identity and indigenous communities. Journal of Language and Cultural Studies.

Jiang, Z., Lv, Q., Zhang, Z., & Lei, L. (2023). Distinguishing translations by human, NMT, and ChatGPT: A linguistic and statistical approach. arXiv.

Kandler, A., & Unger, R. (2023). Modeling language shift. In Diffusive spreading in nature, technology and society (pp. 365–387). Springer International Publishing.

Koc, V. (2025). Generative AI and large language models in language preservation: Opportunities and challenges. arXiv (Cornell University).

Lepp, A., & Sarin, L. (2024). Linguistic justice and digital inequality. Language Policy & Technology Review.

Li, M., Croucher, S. M., & Shen, L. (2024). Language endangerment and the linguistic vitality of Miao in China: cultural shifts and revitalisation strategies. Journal of Multilingual and Multicultural Development, 1–16.

Mahi, M. H., Khan, A. R., Anik, M. H., Noori, S. R. H., Mahmud, A., & Mojumdar, M. U. (2025). MELD: a multilingual ethnic dataset of Chakma, Garo, and Marma in Bengali script with English and standard Bengali translation. Data in Brief, 61, 111745.

Mohamed, M., et al. (2024). Translation quality in the age of AI. Language & Technology.

Mohsin, A. (2023). Indigenous languages of Bangladesh. University Press Limited.

Moneus, A. M., & Sahari, Y. (2024). Artificial intelligence and human translation: A contrastive study based on legal texts. Heliyon, 10(6).

MQM. (2015). Multidimensional quality metrics definition. Retrieved from https://web.archive.org/web/20210113220425/http://www.qt21.eu/mqm-definition/definition-2015-05-27.html

O'Hagan, M. (2016). Massively open translation: Unpacking the relationship between technology and translation in the 21st century. International Journal of Communication, 10, 18.

Okafor, A. Y. (2025). Examining AI translation errors in Igbo: Lexical ambiguity, misinterpretation, and incorrect word substitutions due to contextual deficiencies. Indonesian Journal of Learning Studies, 5(1), 46–55.

Oladipupo, F., Soronnadi, A., Adebara, I., & Adekanmbi, O. (2025, August). How effective are AI models in translating English scientific texts to Nigerian Pidgin: A low-resource language? In I Can't Believe It's Not Better: Challenges in Applied Deep Learning.

Rafat Al Rousan, Raghad Jaradat, & Mona Malkawi. (2025). ChatGPT translation vs. human translation: an examination of a literary text. Cogent Social Sciences, 11(1), 2472916. https://doi.org/10.1080/23311886.2025.2472916

Ranathunga, S., Lee, E. S. A., Prifti Skenduli, M., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages: A survey. ACM Computing Surveys, 55(11), 1–37.

Saikia, M., & Ullman, J. (2023). Endangered language assessment framework. Language Documentation Journal.

Sevinç, Y. (2022). Language endangerment and revitalization. Annual Review of Linguistics.

Smith, B. K., Ehala, M., & Giles, H. (2017). Vitality theory. In J. Nussbaum (Ed.), Oxford research encyclopedia of communication. Oxford University Press.

Spivak, G. C. (1993). Outside in the teaching machine. Routledge.

Tymoczko, M. (1999). Translation in a postcolonial context: Early Irish literature in English translation. St. Jerome Publishing.

Tsunoda, T. (2006). Language endangerment and language revitalisation: An introduction. Mouton de Gruyter.

UNESCO. (2003). Language vitality and endangerment. Ad Hoc Expert Group on Endangered Languages.

Walsh, J. (2006). Language and socio-economic development: Towards a theoretical framework. Language Problems and Language Planning, 30(2), 127–148.

Wei, L., Hua, Z., & Simpson, J. (Eds.). (2023). The Routledge handbook of applied linguistics: Volume two. Taylor & Francis.

Yan, J., Yan, P., Chen, Y., Li, J., Zhu, X., & Zhang, Y. (2024). GPT-4 vs. human translators: A comprehensive evaluation of translation quality across languages, domains, and expertise levels. arXiv preprint arXiv:2407.03658.

✏️ Edit this text
🖊 Edit selected portion
B Bold
I Italic
U Underline
Undo edit
📋 Copy
R Mark red (not in MS)
Remove red mark
📝 Edit History

Every change made in Edit Mode is logged here — tab, section, passage, before & after. Click Revert on any entry to undo that specific change.

0
Total Edits
0
Tabs Changed
0
Today
0
Reverted
Filter:
No edits recorded yet. Turn on ✏️ Edit Mode and make changes to see them here.