A Native-Speaker–Validated Study of Chakma–English Translation · Interactive Manuscript
Open Dataset Explorer (edit live data)Keywords: Chakma language | indigenous language preservation | linguistic justice | AI-based translation | Multidimensional Quality Metrics (MQM)
Language endangerment has become increasingly widespread worldwide, with 14 indigenous languages being on the verge of extinction in the land of Bangladesh alone (Chiran, 2025). The Chakma language, spoken by the Chakma community that resides in Chittagong Hill Tracts (CHT), is one such 'definitely endangered' language (Saikia & Ullman, 2023) that faces ongoing threat due to socio-political marginalization. Preserving the Chakma language is essential for sustaining linguistic diversity, cultural practices, and indigenous knowledge systems embedded within the language. Inspired by the potential of AI translation tools, the study investigates the extent to which such technologies can support endangered language documentation. Specifically, the study asks: (1) how accurately do AI-based translation systems translate Chakma texts into English when compared with native-speaker–validated human translations, and (2) what types of lexical, grammatical, and culturally grounded errors recur in AI-generated translations? To address these questions, a human-validated reference corpus of 1,691 Chakma–English sentence pairs is compiled, and AI-generated translations are systematically evaluated using the Multidimensional Quality Metrics (MQM) framework across three quality dimensions—accuracy, fluency, and cultural appropriateness—in conjunction with an error analysis approach validated by native Chakma speakers. Of the 1,691 sentences that were fully annotated, the AI system produced an acceptable translation for 398 sentences (23.5%), while 1,432 error occurrences were identified in total—dominated by lexical/meaning errors (72.5%) and grammatical errors (22.0%)—yielding a weighted MQM error penalty score of 7.303 penalty points per sentence, which places the corpus in the Very Poor quality band (≥7.00) band. To support the interpretation of translation outputs, the study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors. Since large language models rely on attention-based mechanisms trained predominantly on high-resource languages, they remain ill-equipped for Chakma's morphological complexity. The Bi-LSTM-CRF effectively addresses this gap, achieving a grammatical pattern recognition accuracy of 87.4%. By foregrounding native-speaker validation, linguistic analysis, and systematic evaluation of AI-generated translations, the study demonstrates how AI translation tools can be critically assessed and responsibly operationalized as equitable resources for endangered language documentation, while laying the groundwork for more effective AI-supported approaches to Chakma language preservation.
The world hosts a vast array of languages essential to humanity's heritage (Drude, 2003). Currently, over 7,000 languages are spoken across the globe; however, this remarkable linguistic richness confronts substantial risks as modernity advances (Hutson et al., 2024), causing 40% of the languages to head toward extinction (Eberhard et al., 2022, as cited in Li et al., 2024). According to the Language Conservancy, after global warming, language loss is recognized as the planet's most pressing crisis (Collette & Kennedy, 2023, as cited in Hutson et al., 2024). Numerous endangered languages are diminishing rapidly due to globalization and modernization. This decline is particularly concerning because linguistic diversity is essential for transmitting culture, values, beliefs, and history across generations (Sevinç, 2022). The unique words, phrases, and expressions of each language encapsulate the accumulated knowledge and experiences of its speakers, shaping identity and fostering a sense of pride (Jerome et al., 2022). For many indigenous communities, languages are key carriers of culture, containing unique communication systems, traditional knowledge, and a strong sense of identity (Gwerevende & Mthombeni, 2023). Therefore, the dramatic loss of these minority languages signifies more than silenced voices; it entails the epistemic erasure of invaluable cultural knowledge and distinct worldviews (Kandler & Unger, 2023).
Within this broader global context, Bangladesh is no exception. The country, characterised by its rich ethnic diversity, is home to multiple indigenous communities, many of whose languages are increasingly at risk of decline, mainly due to globalization, urbanization, and the extensive use of politically dominant languages in social, educational, and professional spheres (Anik et al., 2025). The global decline of indigenous languages has reached a critical point, with studies indicating that around half of them could disappear within this century. The Kuruk language, for instance, is no longer spoken, while languages such as Pankho, Khumi, and Hajong are barely surviving (Mohsin, 2023). UNESCO (2003) identifies globalization, forced displacement, and assimilation-driven policies as the main forces behind this decline. Similarly, Chakma and Sultana (2023) argue that the loss of ancestral lands, environmental degradation, and the gradual erosion of cultural identities have further accelerated the decline of indigenous languages. Additionally, the increasing tendency to use native languages only in private settings gradually reduces speakers' fluency and intergenerational transmission, thereby heightening the risk of language extinction (Holmes, 2021).
Endangered languages like Chakma are facing both social and political marginalization reflecting the patterns of linguistic racism. In other words, minority voices are subdued while dominant languages are prioritized (Dovchin, 2020, 2025; Rosa & Flores, 2021). This broader pattern is clearly visible in the Chittagong Hill Tracts (CHT) of Bangladesh, where the Chakma people have been among the most affected by language policies. After independence, Bangladesh followed a 'one culture, one language' idea. This language policy reinforced the prominence of the native language, Bangla, while diminishing the visibility and status of minority languages and identities (Bal, 2010). Chakma and Sultana (2023) describe this as a form of language control, where indigenous people feel pressured to stop using their native languages. In light of this linguistic marginalization, translation acts as an important platform for resistance. Translation is not merely a technical process of linguistic transfer anymore. Instead, it is understood as a political and ethical act that can challenge the dominance of certain languages over others. Tymoczko (1999) asserts that translation has formed the cultural politics of colonized societies. It has helped communities preserve and negotiate their national and cultural identities. In colonial and postcolonial settings, translation is closely linked with power, representation, and cultural authority (Bassnett & Trivedi, 1999). Folaron (2015) argues that translation helps indigenous languages survive and gain recognition simply by increasing their exposure beyond their immediate communities. It is also perceived as an excellent strategy in reclaiming disadvantaged voices (Spivak, 1993). Through this lens, translating indigenous languages becomes more than documentation by being an act of resistance against linguistic marginalization and a way of affirming indigenous identity.
Although Machine Translation (MT) research has advanced considerably through neural and large language model–based approaches, it continues to focus predominantly on high-resource language pairs supported by large-scale parallel corpora and extensive training data. Evaluation also frequently relies on automated metrics such as BLEU and TER, which often obscure crucial shortcomings by failing to capture deeper issues in translation quality. Al Sharou and Specia (2022) demonstrate that in low-resource and user-generated environments, the existence and severity of errors, particularly mistranslations, omissions, and hallucinations, are more consequential than fluency levels. While efforts are made in revitalizing languages around the globe, endangered languages, particularly in South Asia, such as the Chakma language, remain largely underexplored in NLP research (Chakma et al., 2024). Specifically, there is a notable lack of empirical research on Chakma–English translation using AI-based systems and very little is known about the lexical, grammatical, and culturally grounded errors that recur in their translation outputs. Without addressing these gaps, AI translation risks misrepresenting minority languages and undermining language preservation efforts. Therefore, this study employs a human-validated Chakma–English corpus of 1,691 sentences to systematically assess the performance of AI translations, examining how accurately AI can render Chakma texts into English while maintaining both linguistic fidelity and cultural nuances. Patterns of recurring AI errors are also investigated through a structured MQM-based evaluation framework. The study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.
Considering the aim of the study, the following questions were formulated:
RQ1. To what extent do AI-based translation systems accurately translate Chakma texts into English when compared with native-speaker-validated human translations?
RQ2. Which types of lexical, grammatical, and culturally grounded errors occur most frequently in AI-generated translations?
Among the 38 regional languages spoken in Bangladesh, 14 indigenous languages face the threat of extinction (Chiran, 2025). Due to historical power dynamics (Bishop, 2022) and various socio-political factors, the development and expansion of these ancestral languages have become progressively more challenging (Awal, 2019). Chakma, Marma, Tripura, Mro, and Murung are among the notable underrepresented indigenous communities in Bangladesh, among which the Chakma constitute the largest ethnic indigenous group (Afreen, 2020). The Chakma community resides primarily in the Chittagong Hill Tracts (CHT) in the southeastern region of the country. Their mother tongue, the Chakma language, is spoken by approximately 600,000 to 1,000,000 people across the CHT and parts of India (Chakma, 2010) and is classified as 'definitely endangered' (Saikia & Ullman, 2023), indicating that children are no longer consistently learning the language at home (UNESCO). Li et al. (2024) argue that a language's sustainability is ensured when it is actively used across diverse domains such as the home, educational institutions, workplaces, religious settings, and media. As a language loses visibility within these domains, everyday usage gradually declines, threatening the cultural practices, oral traditions, and intergenerational knowledge systems embedded within it (Tsunoda, 2006).
Machine Translation (MT) refers to 'computerized systems responsible for the production of translations with or without human assistance' (Hutchins, 1995, p. 1). With substantial advancements in technology, MT has become an effortless and accessible tool for quickly translating spoken and written texts across languages. MT has evolved through several major paradigms — rule-based (RBMT), statistical (SMT), hybrid, and most recently neural machine translation (NMT) — each improving upon the limitations of its predecessor. NMT systems use deep neural networks based on encoder–decoder architectures to model translation as a sequence-to-sequence task (Bahdanau et al., 2015; Cho et al., 2014). Compared to earlier systems, NMT improves contextual understanding, reduces literal translations, and enhances scalability and efficiency, leading to widespread adoption in major translation systems. Despite these advancements, translation quality remains inconsistent for low-resource languages (Jiang et al., 2023) due to limited training data and linguistic underrepresentation in existing corpora (Zhong et al., 2025). This highlights the persistent challenges faced by contemporary AI-based translation systems in handling linguistically underrepresented languages.
Artificial Intelligence (AI), particularly Natural Language Processing (NLP) and Large Language Models (LLMs), has emerged as a promising tool for preserving and revitalizing endangered and minority languages through the documentation, analysis, and translation of linguistic resources (Koc, 2025). Despite these opportunities, the application of AI to endangered language preservation remains accompanied by significant challenges. Research suggests that digitized documentation efforts often struggle to accurately capture the cultural complexities and linguistic nuances inherent in minority languages (Hutson et al., 2024; Ingram, 2025). Although AI systems can generate grammatically coherent translations and facilitate language accessibility (Putri et al., 2024), they frequently fall short in capturing the cultural and contextual subtleties that human translators can reliably interpret (Moneus & Sahari, 2024). Studies indicate that machine translation may distort contextual meaning, overlook idiomatic expressions and historical significance, and lack the cultural depth and real-world understanding necessary for effective language preservation (Hutson et al., 2024; Okafor, 2025; Putri et al., 2024). Furthermore, current AI-driven approaches to language translation frequently prioritize efficiency over cultural authenticity, overlooking broader goals of linguistic preservation (Mufwene, 2005; Anik et al., 2025). The dominance of English-centric AI models further reinforces existing linguistic hierarchies, marginalizing lesser-known languages and limiting their digital accessibility (Lepp & Sarin, 2024).
Translation quality evaluation is concerned with determining how effectively a translation conveys the meaning and communicative intent of the source text. The Multidimensional Quality Metrics (MQM) framework (MQM, 2015) enables systematic identification and classification of translation issues across multiple dimensions, allowing for both holistic quality assessment and fine-grained error analysis. In line with this framework, the present study assesses overall translation quality along three dimensions—accuracy (the meaning is correct), fluency (the grammar is correct and the output is natural), and cultural appropriateness (culturally specific meaning is preserved)—while the error analysis classifies individual problems into the lexical, grammatical, and culturally grounded categories examined in the research questions. This combined approach allows for a comprehensive assessment of both the overall quality and the underlying error patterns in AI-generated translations.
This study uses a mixed-methods design to evaluate how well a large language model translates sentences from Chakma — an endangered language — into English. We compare AI-generated translations against human-validated reference translations, sentence by sentence, using the Multidimensional Quality Metrics (MQM) framework. Quantitative scoring gives us a number we can compare across systems or studies; qualitative analysis tells us why errors happen and what they mean for the language community that depends on accurate documentation.
The decision to use a native-speaker-validated reference corpus rather than automated metrics (BLEU, chrF) was deliberate. Automated metrics measure surface similarity and are known to be unreliable for morphologically rich, low-resource languages. A human reference with native-speaker sign-off is the only meaningful gold standard for Chakma.
We built a Chakma–English parallel corpus of 1,691 sentence pairs drawn from everyday speech: greetings, questions, requests, descriptions of daily activities, and culturally specific expressions including kinship terms, pragmatic markers, and community phrases that do not have direct English equivalents.
Each sentence was written in Chakma Unicode script (U+11100–U+1114F), accompanied by a pronunciation guide in Bengali script for readability, and paired with an English translation produced by a bilingual researcher. Every translation was then reviewed by native Chakma speakers who corrected phrasing, cultural nuance, and meaning where necessary. The validated translation became the reference standard against which all AI output was measured. Of the 1,691 sentence pairs, 398 (23.5%) required no correction by the native reviewer and were accepted as written.
The system under evaluation is ChatGPT (OpenAI, GPT-4 architecture), chosen because it represents the current frontier of accessible AI translation. Unlike specialised Neural Machine Translation (NMT) systems, ChatGPT has not been explicitly trained on Chakma data, making this evaluation a realistic test of what a researcher or practitioner would encounter if they used a widely available tool for Chakma documentation.
Each Chakma sentence was submitted to the model with a standard instruction: "Translate the following Chakma sentence into English." The raw output was recorded without any post-editing or re-prompting. This zero-post-edit protocol ensures the evaluation reflects the model's actual performance, not the performance achievable with additional human intervention.
We adopt the Multidimensional Quality Metrics (MQM) framework, an industry-standard annotation scheme developed for professional translation evaluation. MQM provides a structured taxonomy of error types and assigns quantitative severity weights, making it possible to produce a single comparable score per sentence and per corpus.
Each AI translation was compared against the human reference sentence by sentence. Every detected deviation was recorded as a separate MQM error and assigned three attributes: a Dimension, a Category, and a Severity.
Errors are classified under three MQM dimensions, each with its own sub-categories:
| Dimension | Category | What it means |
|---|---|---|
|
Accuracy (Lexical Error) |
Mistranslation | The AI chose a word or phrase that means something different from the source |
| Omission | A word or phrase present in the source was dropped entirely | |
| Addition | Words were added that have no basis in the source | |
| Untranslated | Source text was left in Chakma script rather than translated | |
| Terminology | A domain-specific term was rendered incorrectly | |
|
Fluency (Grammatical Error) |
Grammar | Wrong tense, incorrect verb form, or grammatically malformed sentence |
| Word Order | Constituents are in the wrong position for natural English | |
| Agreement | Subject–verb or number agreement is broken | |
| Spelling | Transcription or spelling error in the output | |
| Punctuation | Missing or incorrect punctuation altering readability | |
|
Locale Convention (Cultural Error) |
Cultural Appropriateness | A culturally specific term — kinship word, pragmatic marker, social register — was flattened or lost |
Each annotated error is assigned one of three severity levels. The severity determines the penalty point added to the sentence's MQM score:
| Severity | Penalty | Definition | Example scenario |
|---|---|---|---|
| Critical | 10 | The translation is misleading or conveys an incorrect message. The reader would be misinformed. | "What are you doing?" → AI outputs "you are beautiful" — completely different meaning |
| Major | 5 | The error significantly reduces quality or changes meaning. The core message is altered. | "I am eating rice" → AI outputs "I eat rice" — tense lost, aspectual distinction gone |
| Minor | 1 | The meaning is mostly preserved. The error has little practical impact on communication. | A missing article or a stylistic word-order preference |
If a sentence contains multiple errors, each is annotated separately and all penalties are summed. For example, a sentence with one Critical Accuracy error and one Major Fluency error would receive a sentence score of 10 + 5 = 15.
The overall corpus MQM score is the average penalty per sentence across all N sentences:
We define five quality bands to interpret the MQM score. These thresholds are stated explicitly so that future studies can use the same scale for cross-system comparison:
| Band | MQM Score Range | Interpretation | Publication readiness |
|---|---|---|---|
| Excellent | 0.00 – 0.99 | Near-perfect translation; only trivial errors | Ready for direct publication |
| Good | 1.00 – 2.99 | High quality; minor corrections needed | Ready after light review |
| Acceptable | 3.00 – 4.99 | Usable with human post-editing | Suitable for internal drafts |
| Poor | 5.00 – 6.99 | Significant errors present; not reliable | Not suitable without revision |
| Very Poor ← | ≥ 7.00 | Systematic failure; misleading output likely | Requires complete re-translation |
This corpus scored 7.303, placing it firmly in the Very Poor band — the AI output cannot be used for documentation or publication without comprehensive human correction.
Annotation was carried out in six stages:
Across 1,691 sentences, annotators identified 1,432 MQM errors in total — an average of 0.85 errors per sentence. Of these, 1,038 (72.5%) were Accuracy errors at Critical severity, reflecting systematic semantic failure rather than isolated mistakes. Sentences scoring zero penalty (perfect translations) numbered 398 (23.5%). The remaining 1,293 sentences (76.5%) required at least one error annotation.
To support interpretation of Fluency and Accuracy error patterns identified through MQM annotation, this study also examines whether a Bidirectional Long Short-Term Memory Conditional Random Field (Bi-LSTM-CRF) model can identify the grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.
Large language models such as ChatGPT rely on attention-based mechanisms trained predominantly on high-resource languages and remain ill-equipped to handle Chakma's morphological complexity — particularly its verb-final sentence structures, aspect-marking suffixes, and evidential markers. The Bi-LSTM-CRF is a sequence labelling model suited to identifying morphosyntactic patterns without relying on pre-trained multilingual embeddings. Applied to this corpus, it achieved a grammatical pattern recognition accuracy of 87.4%, providing a structural map of which Chakma constructions the AI is most likely to mistranslate. The model is not used as a translation system; it serves as a diagnostic layer linking MQM annotation findings to underlying grammatical causes.
This study adheres to ethical principles in linguistic research involving indigenous and endangered language communities. All human-translated Chakma–English data were handled with respect for linguistic and cultural integrity. Participation of native Chakma speakers in the validation process was voluntary, and their contributions were used exclusively for academic purposes.
No personally identifiable information was collected, and all data were anonymized prior to analysis. AI-generated translations were used strictly for research evaluation and not for real-world deployment, ensuring that human linguistic expertise remained central to interpretation. The study aims to support, rather than replace, indigenous language knowledge by critically examining both the capabilities and limitations of AI-based translation systems in low-resource language contexts.
This study received ethical approval from the Human Research Ethics Committee of the lead author's institution prior to the commencement of the research. All participants provided informed consent before the study began. The research was conducted in accordance with the institution's ethical guidelines and complied with all relevant legal and regulatory requirements. Participant confidentiality and data privacy were maintained throughout the study, and all data were used solely for research purposes. Particular care was taken to ensure that the documentation and analysis of the Chakma language were conducted in a respectful, culturally sensitive, and ethically responsible manner.
The methodological framework ensures robustness through a validated human reference corpus, the application of the MQM framework for systematic error classification, and native speaker validation for cultural and linguistic accuracy. Reliability is further strengthened through inter-annotator agreement testing, while the integration of computational and linguistic analysis enhances analytical depth. Collectively, these components ensure that the study is replicable, reliable, and grounded in both real-world language use and computational evaluation standards.
This section presents the findings of the systematic analysis of AI-generated Chakma–English translations compared with a native-speaker-validated human reference corpus of 1,691 sentence pairs. A mixed-methods approach is adopted, employing the Multidimensional Quality Metrics (MQM) framework to identify and classify translation errors across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error), with each error assigned a severity level (Critical = 10, Major = 5, Minor = 1) and summed to produce a per-sentence penalty score. All annotations were validated by native Chakma speakers to ensure linguistic accuracy, contextual appropriateness, and cultural sensitivity.
The overall MQM evaluation revealed a clear and significant performance gap between human and AI-generated translations across all three dimensions. Human translations consistently achieved higher scores in accuracy, fluency, and cultural appropriateness, reflecting the depth of cultural and contextual knowledge that native-speaker translators bring to the task. Table 1 summarises the comparative scores.
| Dimension | Human Translation (M) | AI Translation (M) | Gap |
|---|---|---|---|
| Accuracy | 0.90 | 0.235 | −0.665 |
| Fluency | 0.90 | 0.814 | −0.086 |
| Cultural Appropriateness | 0.90 | 0.953 | −0.053 |
| Overall MQM Score | 0.00 (baseline) | 7.303 (Very Poor) | — |
Accuracy measures whether the AI translation preserves the meaning of the source Chakma sentence. The AI achieved an accuracy score of 0.235, compared to the human benchmark of 0.90 — a gap of 0.665 points. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without any accuracy error. The remaining 1,293 sentences (76.5%) contained at least one error that distorted or entirely replaced the intended meaning. The corpus-level MQM penalty score of 7.303 reflects the cumulative weight of these failures: on average, each sentence carries 7.303 penalty points of translation error, placing the corpus firmly in the Very Poor quality band (≥7.00).
These results are consistent with findings from low-resource machine translation research (Ranathunga et al., 2023), which identifies data scarcity and limited digital representation as key factors that reduce AI accuracy on minority languages. Chakma's limited presence in LLM training data likely explains the AI's inability to correctly interpret the semantic content of Chakma sentences, particularly when surface-level phonological or script similarity led the model to select semantically unrelated English equivalents.
Fluency measures the grammatical naturalness and readability of the AI output in English. The AI achieved a fluency score of 0.814, compared to the human benchmark of 0.90 — a substantially smaller gap of 0.086 points. Of 1,691 sentences, 1,376 (81.4%) were produced without any grammatical error. This finding suggests that the AI performs considerably better at generating grammatically well-formed English sentences than at preserving the semantic content of the Chakma source.
However, this apparent fluency conceals a critical problem. A translation can be grammatically smooth in English while conveying entirely the wrong meaning — and this is precisely the pattern observed in the majority of erroneous AI outputs. As Al Sharou and Specia (2022) note, in low-resource settings, the severity of meaning distortion is more consequential than surface fluency levels. The AI's relatively high fluency score therefore masks, rather than mitigates, the depth of semantic failure in this corpus.
Cultural appropriateness measures whether culture-specific meanings — including kinship terms, pragmatic markers, proverbs, and community-specific expressions — are preserved in the AI output. The AI achieved a score of 0.953, compared to the human benchmark of 0.90 — the smallest performance gap of the three dimensions. Of 1,691 sentences, only 79 (4.7%) contained cultural or pragmatic errors.
While this figure appears relatively low, it is important to note that cultural errors are qualitatively more serious than their frequency suggests. The loss of a kinship distinction (such as "maternal uncle" becoming "uncle") or the failure to interpret a Chakma proverb represents an irreversible erasure of cultural meaning — precisely the type of loss that language documentation is intended to prevent. These findings align with Moneus and Sahari (2024), who find that AI systems struggle with culturally embedded expressions even when grammatical accuracy is maintained.
Across 1,691 sentences, the MQM annotation identified a total of 1,432 error occurrences. Table 2 summarises the distribution across the three MQM dimensions.
| MQM Dimension | Error Category | Frequency | Percentage |
|---|---|---|---|
| Accuracy (Lexical Error) | Mistranslation | 1038 | 72.5% |
| Fluency (Grammatical Error) | Grammar | 315 | 22.0% |
| Locale Convention (Cultural Error) | Cultural Appropriateness | 79 | 5.5% |
| Total | 1,432 | 100% | |
Lexical accuracy errors emerge as the dominant error type, accounting for 72.5% of all annotated errors. Grammatical errors constitute 22.0%, while cultural and pragmatic errors, though fewest in number, represent the most contextually significant failures at 5.5%. The severity distribution reveals that all Accuracy errors were annotated as Critical (10 pts), all Fluency and Locale errors as Major (5 pts), and no Minor errors were identified — indicating that all AI failures in this corpus are substantive rather than superficial.
Lexical errors represent failures in meaning-level translation — cases where the AI selected an English word or phrase that does not correspond to the semantic content of the Chakma source. These errors were the most frequent and, at Critical severity (10 pts each), the most damaging to the overall MQM score. Four sub-types of lexical error were identified.
Mistranslation — The AI output diverged significantly from the intended meaning of the source sentence, in some cases producing translations with no apparent semantic relationship to the input.
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴 তুই হি গরর? |
What are you doing? | you are beautiful |
| 5 | 𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴 তে হদু যার |
Where is he going? | that is enough |
| 3 | 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨 তুই ঈদু আই |
Come here. | you gave yes / you did give |
The most striking example is Sentence #1: the Chakma question "তুই হি গরর?" (What are you doing?) was rendered as "you are beautiful" — a complete semantic inversion. Similarly, in Sentence #5, the AI produced "that is enough" where the human translation reads "Where is he going?". These cases represent not merely word-level inaccuracy but a total failure to process sentence meaning.
Untranslated text — In several instances, the AI output largely mirrored the romanization of the Chakma source rather than producing a meaningful English translation, leaving the output entirely inaccessible to non-Chakma readers.
| # | Chakma Source | Human Translation | AI Output |
|---|---|---|---|
| 6 | 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴 তামিম ঘুমজার |
Tamim is sleeping. | timid. gum jor |
| 7 | 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬 তুই হাময়ান গড়ি দে |
Please do the work. | You make ready / prepare |
These untranslated outputs — where Chakma phonology is rendered in Roman script without any English semantic content — indicate that the AI system recognises the source as non-English text but lacks the linguistic resources to decode it. This pattern is most common with phonologically complex Chakma words that have no close approximation in the AI's training data.
Unintelligible output — A further subset of lexical errors produced outputs that were neither translations nor romanizations but rather syntactically incoherent fragments with no recoverable meaning.
Grammatical errors were annotated as Major-severity Fluency errors (5 pts each), reflecting the judgment that grammatical failures reduce translation quality significantly but the core message may sometimes remain partially recoverable. The following sub-types were identified in the corpus.
Tense errors — The most common grammatical error type. The AI consistently failed to map Chakma aspectual and temporal distinctions onto appropriate English tenses, collapsing future, past, and progressive forms into simple present.
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 2 | 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬 তুই হি গরিবে? |
What will you do? | you will do |
| 4 | 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴 মুই ভাত হাঙর |
I am eating rice. | I eat rice |
| 33 | 𑄟𑄪𑄭 𑄝𑄎𑄢𑄧𑄖𑄴 𑄡𑄬𑄟𑄴 মুই বাজারত্ যেম্ |
I will go to the market. | I am going to the market |
The Chakma language encodes tense and aspect morphologically in ways that differ substantially from English. Errors of this type suggest that the AI is processing Chakma lexical items in isolation rather than parsing the full morphosyntactic structure of the source sentence.
Sentence category errors — The AI frequently shifted the grammatical category of a sentence, converting questions into statements, subjunctive wishes into declaratives, and imperatives into indicative clauses. This error type is particularly consequential in a documentation context because it misrepresents the communicative function of the source utterance.
Omission errors — In several instances, the AI dropped core sentence constituents — including subjects, main verbs, and interrogative particles — producing outputs that preserved partial meaning but lost semantic completeness.
Other grammatical errors — Additional error sub-types included singular/plural mismatches, addition of non-present content (such as politeness markers absent from the source), and word order errors resulting in awkward or unnatural English phrasing.
Cultural and pragmatic errors (Locale Convention, Major severity) were the least frequent but qualitatively most significant category of error. These arose when culturally embedded terms, relational distinctions, or community-specific expressions in Chakma were either flattened into generic English equivalents, misinterpreted, or rendered unintelligible.
| # | Chakma Source | Human Translation | AI Translation | Cultural issue |
|---|---|---|---|---|
| 22 | 𑄟𑄧𑄢𑄨𑄝𑄬 𑄚𑄦𑄨 মরিবে নাহি? |
Are you going to die? | will die not / will not die | Kinship/cultural term lost |
| 87 | 𑄖𑄬 𑄟𑄧𑄢𑄬 𑄷𑄶𑄶 𑄑𑄬𑄋 𑄃𑄪𑄘𑄮𑄢𑄴 𑄘𑄨𑅅 তে মরে 100টেঙা উদোর দ্যি |
He lent me 100 taka. | Error parsing | Kinship/cultural term lost |
| 338 | 𑄟𑄧 𑄟𑄟𑄪 𑄊𑄧𑄢𑄧𑄖𑄴 𑄃𑄉𑄬 𑅁 ম মামু ঘরত্ আগে। |
My maternal uncle is at home. | My uncle is at home. | Kinship/cultural term lost |
The most common cultural error type was the collapse of Chakma kinship terminology into undifferentiated English equivalents. The Chakma language maintains precise distinctions between maternal and paternal relatives, elder and younger siblings, and community-specific social roles. When the AI produces "uncle" for "maternal uncle", or "brother" for a term carrying specific age-relative social meaning, it erases the relational structure that is central to Chakma social and cultural life.
A further significant sub-type was the misinterpretation of Chakma proverbs. The AI consistently failed to interpret figurative or idiomatic expressions, either producing literal translations of the surface words (which convey no meaning in English) or substituting entirely unrelated English proverbs. This failure reflects the cultural knowledge gap that Hutson et al. (2024) identify as a fundamental limitation of AI systems in endangered language documentation: cultural depth and real-world community understanding cannot be learned from statistical patterns in training data alone.
To supplement the MQM annotation findings under RQ2, a Bi-LSTM-CRF model was applied to identify recurring grammatical and verb pattern structures in Chakma sentences that correlate with AI translation errors. The model achieved a grammatical pattern recognition accuracy of 87.4% on the Chakma corpus, indicating reliable identification of morphosyntactic structure in the absence of large pre-trained Chakma language resources.
The Bi-LSTM-CRF analysis revealed that the grammatical structures most consistently associated with AI translation errors were:
These findings from the Bi-LSTM-CRF corroborate the MQM Fluency error patterns (§4.2.3) and provide a structural explanation for why those errors occur. The model's 87.4% accuracy confirms that these grammatical patterns are identifiable and consistent — suggesting that a dedicated Chakma NLP pipeline could, in principle, flag high-risk structures before translation and alert human reviewers accordingly. This is proposed as a direction for future work in §6.2.
This section interprets the findings in relation to the two research questions, situates them within the broader literature on AI translation and low-resource language processing, and reflects on what they mean for Chakma language documentation and linguistic justice.
The overall performance of ChatGPT on Chakma–English translation was poor. The corpus-level MQM score of 7.303 places the AI output in the Very Poor quality band (≥7.00), indicating that systematic and significant translation failure — not occasional error — characterises the AI's engagement with Chakma. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without error. The remaining 1,293 sentences (76.5%) required at least one error annotation, and in most cases the error was Critical in severity — meaning the output actively misrepresented the source.
These results are consistent with the broader literature on AI translation in low-resource language settings. Ranathunga et al. (2023) identify data scarcity, the absence of parallel corpora, and insufficient language-specific resources as the primary factors constraining machine translation performance for under-resourced languages. Chakma's minimal presence in the training data of large language models means the system is effectively operating without the linguistic foundation necessary for reliable translation. The findings therefore do not simply reflect a gap in model capability; they reflect a structural inequality in how languages are represented in digital infrastructure and AI training pipelines.
Lexical accuracy errors were the most frequent and most penalised error type, accounting for 72.5% of all MQM annotations at Critical severity (10 pts each). The range of lexical error sub-types — mistranslation, untranslated text, and unintelligible output — reflects the probabilistic nature of large language models. Rather than understanding meaning in the way human translators do, AI systems generate output based on patterns learned from large amounts of textual data (Fu & Liu, 2024). When the training data contains little or no Chakma, the model cannot distinguish between visually or phonologically similar but semantically unrelated words, resulting in outputs that are grammatically plausible in English but semantically disconnected from the source.
The untranslated outputs — where the AI produced romanised Chakma rather than English — are particularly revealing. They indicate that the model recognises the script as non-English but lacks the decoding capacity to generate a meaningful translation. This pattern aligns with Okafor's (2025) findings on Igbo, where contextual deficiencies in AI training led to lexical ambiguity and incorrect word substitutions. For Chakma, the problem is more fundamental: the lexical base itself is largely absent from the model's knowledge.
Grammatical errors (315 instances, 22.0% of total errors) were consistently annotated at Major severity (5 pts), reflecting the judgment that they reduce translation quality significantly while sometimes leaving core meaning partially intact. The most common grammatical error type — tense and aspect misrepresentation — suggests that the AI processes Chakma lexical items in isolation rather than parsing morphosyntactic structure. Chakma encodes tense, aspect, and mood through morphological affixes that differ substantially from English inflectional patterns; without explicit modelling of these structures, the AI defaults to simple present tense regardless of the source's temporal reference.
The sentence category errors — in which questions became statements, wishes became declaratives, and imperatives became indicatives — are particularly significant from a documentation perspective. A corpus that systematically converts Chakma questions into statements misrepresents the pragmatic structure of the language, distorting any subsequent linguistic analysis that relies on the translated data. Fu and Liu (2024) note that ChatGPT occasionally fails to identify implicit grammatical relationships and does not always restructure source-language syntax appropriately; the present study finds this tendency to be systematic rather than occasional in the context of Chakma.
Cultural and pragmatic errors (79 instances, 5.5% of total errors) were the least frequent but qualitatively most significant category. The collapse of Chakma kinship terminology — where distinctions such as maternal versus paternal uncle, or elder versus younger sibling, are flattened into generic English equivalents — represents the erasure of relational and social knowledge that is embedded in the language itself. This is not a translation error in the narrow linguistic sense; it is the deletion of cultural information that has no direct English equivalent and cannot be recovered once lost.
The AI's failure to interpret Chakma proverbs further illustrates the limits of pattern-based language modelling in cultural contexts. Proverbs are community-specific communicative forms whose meaning depends on shared cultural knowledge that cannot be inferred from lexical co-occurrence statistics. Rousan et al. (2025) report similar findings in Arabic–English literary translation, where AI systems misinterpret culturally embedded expressions despite producing fluent English output. The present study extends this finding to an endangered indigenous language context, where the cultural stakes of misinterpretation are considerably higher.
The Bi-LSTM-CRF analysis — which achieved 87.4% grammatical pattern recognition accuracy on the Chakma corpus — provides a structural explanation for the Fluency error patterns identified through MQM annotation. Where MQM records that an AI tense error occurred, the Bi-LSTM-CRF indicates which Chakma morphological structure the AI failed to parse. Together, these two analytic layers produce a more complete picture of AI translation failure than either could provide alone.
The structures most consistently associated with AI errors — progressive aspect suffixes, future tense markers, interrogative particles, negation morphology, and evidential markers — share a common property: they are morphologically encoded in Chakma in ways that have no direct surface-level English counterpart. An attention-based language model trained on high-resource languages will not have learned to associate these Chakma morphemes with their English functional equivalents, because the relevant training signal is absent. The Bi-LSTM-CRF findings confirm that these structures are systematic and identifiable — which means they could, in principle, be used to build error-prediction tools for human post-editors reviewing AI-generated Chakma translations.
This finding is exploratory. A full evaluation of the Bi-LSTM-CRF as a diagnostic component — including comparison across architectural variants, cross-validation on held-out data, and integration with the MQM annotation pipeline — is a direction for future work (see §6.2). The current study limits its claim to the observation that grammatical pattern recognition at 87.4% accuracy supports the interpretation of Fluency errors as structurally grounded failures in morphological parsing, not random noise.
The patterns identified in the MQM annotation point to several structural features of Chakma that consistently triggered AI errors. Sentences with morphologically complex verb forms — particularly those encoding progressive aspect, conditional mood, and evidentiality — showed the highest rates of mistranslation. These structures require the model to track morphological dependencies across the sentence rather than relying on lexical co-occurrence, a capacity that is severely limited when training data for the source language is absent.
Taken together, the findings point to a single underlying cause that cuts across all three error categories: Chakma's near-total absence from the training data of current large language models. This is not a problem that can be resolved through better prompting or model fine-tuning alone. It is a structural condition rooted in the historical and ongoing marginalization of indigenous languages from digital infrastructure, standardized orthography, and natural language processing research.
Ranathunga et al. (2023) identify data scarcity, the lack of parallel corpora, and insufficient language resources as the major challenges affecting machine translation performance in low-resource languages. The present study confirms all three as operative in the Chakma case. The lexical errors reflect limited exposure to Chakma vocabulary; the grammatical errors reflect the absence of Chakma-specific syntactic modelling; and the cultural errors reflect the impossibility of learning community-specific meaning from text corpora alone.
From a linguistic justice perspective, these disparities are not merely technical. The reduced performance of AI translation tools on Chakma reflects the wider marginalization of Indigenous languages within digital infrastructures and training datasets that extensively privilege high-resource languages (Lepp & Sarin, 2024). As Dovchin (2020) argues, linguistic racism operates by subduing minority voices while prioritizing dominant languages — and AI systems trained predominantly on English, Bangla, and other high-resource languages reproduce this hierarchy algorithmically. Improving AI support for Chakma is therefore not only a matter of technological advancement but also a step toward addressing persistent inequalities in linguistic representation within the digital age.
This study examined the effectiveness of AI-based translation systems in translating Chakma texts into English by comparing AI-generated translations with native-speaker-validated human translations across a corpus of 1,691 sentence pairs. Using the Multidimensional Quality Metrics (MQM) framework with three severity levels — Critical (10 pts), Major (5 pts), and Minor (1 pt) — the analysis evaluated translation quality across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error). All annotations were validated by native Chakma speakers to ensure cultural and linguistic integrity.
The findings reveal that AI-generated translations performed poorly across all three dimensions. The corpus-level MQM score of 7.303 places the output in the Very Poor quality band (≥7.00), with only 398 of 1,691 sentences (23.5%) translated without error. Lexical accuracy errors were the dominant failure type, accounting for 72.5% of all annotated errors at Critical severity, reflecting systematic semantic failure rather than surface-level inaccuracy. Grammatical errors, though less frequent, further reduced translation quality by misrepresenting the tense, aspect, sentence type, and structural properties of Chakma utterances. Cultural and pragmatic errors, while fewest in number, were qualitatively the most consequential, involving the irreversible erasure of culturally embedded knowledge — kinship distinctions, proverbs, and community-specific expressions — that cannot be recovered through post-editing alone.
Beyond the technical findings, this study argues that the reduced performance of AI translation on Chakma reflects broader patterns of digital inequality and linguistic marginalization. Languages with limited digital representation receive substantially weaker technological support, and AI systems trained on high-resource languages reproduce existing linguistic hierarchies. From a linguistic justice perspective, improving AI support for Chakma is not merely a technical challenge but an ethical imperative — a step toward equitable representation of indigenous knowledge systems in the digital age (Dovchin, 2020; Lepp & Sarin, 2024).
The study has several limitations. First, the evaluator's awareness of the translation sources may have introduced a degree of bias, although the use of structured MQM criteria and native-speaker review helped mitigate this risk. Second, the corpus, while covering 1,691 sentence pairs — significantly larger than earlier Chakma NLP datasets — may not fully capture the breadth of lexical, grammatical, and cultural variation present in Chakma, particularly across regional dialects and specialized domains such as law, medicine, and oral tradition. Third, this study evaluates a single AI system (ChatGPT) at one point in time; the rapidly evolving landscape of large language models means that findings may not generalise to future model generations. Finally, the Bi-LSTM-CRF component of the study was designed to identify grammatical and verb pattern structures contributing to recurring translation errors, but its findings are limited by the size and diversity of the annotated training data.
The study carries implications for both AI development and language documentation practice. For AI developers, the findings highlight the urgent need for Chakma-specific training resources: larger parallel corpora, culturally informed annotation guidelines, and community-driven validation processes that embed native speaker expertise at every stage of model development. Without these resources, AI translation systems will continue to perform poorly on Chakma and other endangered languages, reinforcing rather than challenging existing linguistic hierarchies.
For language documentation practitioners, the findings suggest that AI translation tools, at their current level of performance, are best understood as assistive rather than autonomous resources. When used alongside native-speaker expertise, AI systems can support the initial translation of Chakma texts and contribute to the development of bilingual language resources — but every AI output must be reviewed before it is trusted. The MQM score of 7.303 quantifies this burden concretely: on average, each sentence requires correction of approximately 7.303 penalty points of translation error, and in 1,038 of 1,432 cases (72.5%), that correction requires addressing a Critical semantic failure.
A specific direction for future work concerns the Bi-LSTM-CRF component introduced in this study as a supporting diagnostic tool. The model achieved 87.4% grammatical pattern recognition accuracy on Chakma, identifying the morphosyntactic structures — progressive aspect markers, future suffixes, interrogative particles, negation morphology — most likely to trigger AI translation errors. A full evaluation of this model, including cross-validation, architectural comparison, and integration with the MQM annotation pipeline, would constitute a meaningful methodological contribution to endangered language NLP. If the model can reliably predict high-risk grammatical structures before translation, it could serve as the basis for a Chakma-specific quality estimation tool that reduces the burden on human post-editors.
Future research should also examine larger and more diverse Chakma corpora, compare multiple AI translation systems and generations of models, and explore community-centred approaches to AI development that centre indigenous linguistic and cultural knowledge. Particular attention should be given to building the digital infrastructure — standardized orthographies, annotated corpora, lexical databases — that would enable meaningful AI support for Chakma and other endangered languages of Bangladesh and South Asia.
Notwithstanding its limitations, this study makes a methodological contribution by demonstrating that the MQM framework — typically applied in professional translation contexts — is a viable and informative evaluation tool for endangered language AI assessment. The penalty-based scoring system, combined with dimension-level and severity-level breakdown, provides a richer and more actionable picture of AI translation failure than automated metrics such as BLEU or TER, which remain largely uninformative for low-resource language pairs. Future studies are encouraged to adopt and extend this framework as part of a broader effort to develop evaluation standards for AI translation in indigenous language documentation.
The authors declare no conflicts of interest.
The authors did not receive any funding for this study.
The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.
This study received ethical approval from the Human Research Ethics Committee of the lead author's institution prior to data collection. All participants provided informed consent before the study commenced, and the research was conducted in accordance with institutional ethical guidelines.
What type of error? Each error belongs to one of three MQM dimensions.
Specifically what went wrong? Each dimension breaks into sub-categories.
| Category | Parent Dimension | Count | % of errors |
|---|---|---|---|
| Mistranslation | Accuracy | 1038 | 72.5% |
| Untranslated | Accuracy | 0 | 0.0% |
| Terminology | Accuracy | 0 | 0.0% |
| Omission | Accuracy | 0 | 0.0% |
| Addition | Accuracy | 0 | 0.0% |
| Grammar | Fluency | 315 | 22.0% |
| Agreement | Fluency | 0 | 0.0% |
| Word Order | Fluency | 0 | 0.0% |
| Cultural Appropriateness | Locale Convention | 79 | 5.5% |
How bad is each error? Severity determines the penalty weight applied to the MQM score.
How many sentences fall into each penalty band? A sentence with score 0 is perfect.
| Penalty Band | Sentences | % of corpus | Meaning |
|---|---|---|---|
| 0 (Perfect) | 398 | 23.5% | AI translation fully correct |
| 1–2 | 0 | 0.0% | Minor error only — meaning mostly preserved |
| 3–5 | 255 | 15.1% | Single Major error or grammar issue |
| 6–10 | 901 | 53.3% | Critical mistranslation — meaning lost |
| >10 | 137 | 8.1% | Multiple errors — completely unacceptable |
Kinship terms, pragmatic markers, and culturally-specific meanings the AI fails to carry across.
The MQM (Multidimensional Quality Metrics) score is a single number that captures how bad the AI translations are, on average. The higher the score, the worse the translation quality. It is computed by adding up all the penalty points from every error found across all sentences, then dividing by the total number of sentences.
We define five quality bands so that any researcher can immediately interpret an MQM score without reading the full paper. These thresholds apply at the corpus level — individual sentences may fall in any band.
| Band | Range | What it means in practice | Can it be used? |
|---|---|---|---|
| Excellent | 0.00–0.99 | Trivial errors only; a professional would likely not notice | ✅ Direct publication |
| Good | 1.00–2.99 | Minor stylistic issues; meaning fully intact | ✅ After light proofreading |
| Acceptable | 3.00–4.99 | One significant error per sentence on average; usable with editing | ⚠ After professional editing |
| Poor | 5.00–6.99 | Multiple errors per sentence; meaning frequently distorted | ❌ Not without major revision |
| Very Poor | ≥ 7.00 | Systematic failure; more sentences wrong than right | ❌ Complete re-translation needed |
Before computing the corpus average, each sentence gets its own MQM score. A sentence score = sum of all its error penalties. Here are four real examples from the database showing every possible scoring outcome.
No errors detected. The AI output matches the human reference in meaning and grammar. Formula: no errors → 0 penalty points.
One Major-severity error detected. Grammar errors are always Major in this corpus because they change the tense, aspect, or sentence type — reducing meaning even if the core content survives.
One Critical-severity error. The AI produced an output that means something entirely different from the source. A reader relying on this translation would be actively misinformed.
Two errors in a single sentence: one Critical Accuracy error (wrong meaning) + one Major Fluency error (grammar). Penalties add up.
Numbers mean more when compared. Here we compare the AI's performance across dimensions, severity levels, and against the theoretical baseline.
Accuracy errors (meaning) contributed far more penalty than Fluency errors (grammar), even though both are present. This tells us the core problem is semantic — the AI is not misunderstanding Chakma grammar, it is misunderstanding Chakma meaning.
| Dimension | Errors | Avg Severity | Total Penalty | % of Total Penalty | Verdict |
|---|---|---|---|---|---|
| Accuracy | 1,038 | Critical (10) | 10,380 | 84.0% | ⚠ Systematic semantic failure |
| Fluency | 315 | Major (5) | 1,970 | 16.0% | Recoverable with post-editing |
| Locale | 79 | Major (5) | 395 | 3.2% | Rare but culturally significant |
| Group | Count | % of corpus | Avg MQM score | Penalty contribution |
|---|---|---|---|---|
| Perfect (score = 0) | 398 | 23.5% | 0.00 | 0 pts |
| Erroneous (score > 0) | 1,293 | 76.5% | 9.55 | 12,350 pts |
| Full corpus | 1691 | 100% | 7.303 | 12,350 pts |
The absence of Minor errors is itself a finding. In a corpus with well-performing AI translation, Minor errors (≤1pt) would dominate the distribution. Here, every annotated error is at least Major (5pt), and 72% are Critical (10pt).
The following are real Chakma sentences from the database, each illustrating a different scoring scenario. They were selected to show the full range of AI behaviour — from perfect to catastrophic.
Grammar errors are Major because they alter the temporal or aspectual meaning. "I eat rice" and "I am eating rice" are different statements in Chakma — one describes a habit, the other an action in progress.
Critical errors are the most dangerous for documentation. A researcher using these translations would record completely wrong information about the Chakma language.
When two errors co-occur in a single sentence, the scores add up. These sentences are not just wrong — they are both wrong in meaning and wrong in grammar simultaneously.
| # | Finding | Evidence | Implication |
|---|---|---|---|
| F1 | AI translation quality is Very Poor | MQM = 7.303 ≥ 7.00 threshold | Cannot be used for documentation without full human correction |
| F2 | Semantic failure is the dominant error type | 72% of errors are Critical Accuracy errors | The AI does not understand Chakma word meanings, not just grammar |
| F3 | Grammar errors are Secondary | 315 Major Fluency errors (22.0% of total) | AI grammar is sometimes recoverable; semantic errors are not |
| F4 | No Minor errors were identified | 0 Minor annotations across 1432 errors | All AI failures in this corpus are substantive, not cosmetic |
| F5 | 23.5% of sentences are acceptable | 398 sentences with score = 0 | AI performs on simple, high-frequency phrases; fails on complex structures |
The AI correctly translated 1 in 4 sentences. The other 3 in 4 required annotation of at least one error, and in the majority of those cases the error was a Critical semantic mismatch — not a grammar slip, not a word order issue, but a fundamentally wrong translation of what the Chakma speaker said. For a language documentation project, this means an annotator must review every single AI output before it can be trusted. The MQM score of 7.303 quantifies this burden: on average, each sentence carries 7.303 penalty points of translation error.
Only ChatGPT was tested. Comparison with Google Translate, DeepL, or Gemini would reveal whether errors are system-specific or universal to the architecture.
Annotators knew which translations were AI-generated. A blind evaluation design would strengthen reliability.
Sentences were translated in isolation. AI performance on multi-sentence or paragraph-level Chakma texts is unknown.
Data split, epochs, embedding type, and inter-annotator agreement statistic (e.g., Cohen's kappa) are not specified in the current manuscript. These are required by most journals.
Several in-text citations still lack full reference-list entries: Chiran (2025), Saikia & Ullman (2023), Chakma & Sultana (2023), Sevinç (2022), Jerome et al. (2022), Holmes (2021), Dovchin (2020, 2025), Rosa & Flores (2021), Bal (2010), Mohsin (2023), Mufwene (2005), Collette & Kennedy (2023), Castilho et al. (2017), España-Bonet & Costa-jussà (2016), Hunsicker et al. (2012), Zhong et al. (2025), Lepp & Sarin (2024), Putri et al. (2024), Hendy et al. (2023), Meighan (2023), Chakma et al. (2024), MQM (2015).
The manuscript tables use the Noto Sans Chakma font. Reviewers without this font installed will see boxes instead of Chakma script. Consider embedding the font or providing a PDF with embedded fonts.
Bi-LSTM-CRF as a full diagnostic pipeline: The current study uses a Bi-LSTM-CRF model at 87.4% grammatical pattern recognition accuracy as a supporting diagnostic tool for RQ2. Future work should evaluate this model more rigorously — with cross-validation, comparison across architectures (e.g. CRF-only, Transformer-based sequence labellers), and integration into the MQM annotation workflow as a pre-annotation step. If the model can reliably identify high-risk grammatical structures before human annotation begins, it could significantly reduce the time required for large-scale Chakma corpus evaluation.
Error prediction integration: A natural extension would be to use the Bi-LSTM-CRF's structural predictions as input features for an AI translation quality estimation system — a tool that could flag likely errors in AI output before human review, prioritising sentences for correction based on predicted MQM severity.
All tables below are formatted for direct inclusion in a research paper. Values update every 2 seconds from the live annotation database.
| Metric | Value | Percentage |
|---|---|---|
| Total sentences evaluated | 1,691 | 100% |
| Perfect translations (MQM score = 0) | 398 | 23.5% |
| Sentences with at least one error | 1,293 | 76.5% |
| Total MQM error occurrences | 1,432 | — |
| Total weighted penalty | 12,350 | — |
| Average MQM penalty per sentence | 7.303 | — |
| Quality band | Very Poor (≥7.00) | |
| MQM Dimension | Description | Error Count | % of Total Errors | Penalty Weight |
|---|---|---|---|---|
| Accuracy (Lexical Error) | Wrong meaning, wrong word, semantic mismatch | 1,038 | 72.5% | ×5 (Critical) / ×5 (Major) |
| Fluency (Grammatical Error) | Grammar, tense, word order, agreement issues | 315 | 22.0% | ×5 (Critical) / ×5 (Major) |
| Locale (Cultural Error) | Cultural appropriateness, kinship terms, pragmatics | 79 | 5.5% | ×5 (Critical) / ×5 (Major) |
| Total | 1,432 | 100% | — | |
| MQM Dimension | Category | Count | % of Dim. Errors | % of Total Errors |
|---|---|---|---|---|
| Accuracy | Mistranslation | 1,038 | 100.0% | 72.5% |
| Omission | 0 | 0.0% | 0.0% | |
| Addition | 0 | 0.0% | 0.0% | |
| Untranslated | 0 | 0.0% | 0.0% | |
| Terminology | 0 | 0.0% | 0.0% | |
| Fluency | Grammar | 315 | 100.0% | 22.0% |
| Word Order | 0 | 0.0% | 0.0% | |
| Spelling | 0 | 0.0% | 0.0% | |
| Punctuation | 0 | 0.0% | 0.0% | |
| Agreement | 0 | 0.0% | 0.0% | |
| Locale | Cultural Appropriateness | 79 | 100.0% | 5.5% |
| Total | 1,432 | — | 100% | |
| Severity | Penalty | Definition | Count | % of Errors | Weighted Contribution |
|---|---|---|---|---|---|
| Critical | 10 | Translation misleading or conveys incorrect message | 1,038 | 72.5% | 10,380 |
| Major | 5 | Meaning significantly changed or reduced quality | 394 | 27.5% | 1,970 |
| Minor | 1 | Meaning mostly preserved, little impact | 0 | 0.0% | 0 |
| Total | 1,432 | 100% | 12,350 | ||
| Penalty Band | MQM Band Label | Sentences | % of Corpus | Interpretation |
|---|---|---|---|---|
| 0 | Excellent | 398 | 23.5% | Fully correct — no errors detected |
| 1–2 | Good | 0 | 0.0% | Minor error only — meaning preserved |
| 3–5 | Acceptable | 255 | 15.1% | Single Major error — meaning reduced |
| 6–10 | Poor | 901 | 53.3% | Critical mistranslation — meaning lost |
| >10 | Very Poor | 137 | 8.1% | Multiple errors — completely unacceptable |
| Total | 1,691 | 100% | — | |
| # | Category | Dimension | Severity | Count | % of Errors | Penalty / Error |
|---|---|---|---|---|---|---|
| 1 | Mistranslation | Accuracy | Critical | 1,038 | 72.5% | 10 |
| 2 | Grammar | Fluency | Major | 315 | 22.0% | 5 |
| 3 | Cultural Appropriateness | Locale | Major | 79 | 5.5% | 5 |
| 4 | Omission | Accuracy | — | 0 | 0.0% | — |
| 5 | Addition | Accuracy | — | 0 | 0.0% | — |
| 6 | Word Order | Fluency | — | 0 | 0.0% | — |
| 7 | Agreement | Fluency | — | 0 | 0.0% | — |
| 8 | Terminology | Accuracy | — | 0 | 0.0% | — |
| 9 | Spelling | Fluency | — | 0 | 0.0% | — |
| 10 | Punctuation | Fluency | — | 0 | 0.0% | — |
| Total annotated errors | 1,432 | 100% | — | |||
| # | Chakma Source | Human Reference | AI Output | Dimension | Category | Penalty |
|---|---|---|---|---|---|---|
| 1 | 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴 তুই হি গরর? |
What are you doing? | you are beautiful | Accuracy | Mistranslation | 10 |
| 3 | 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨 তুই ঈদু আই |
Come here. | you gave yes / you did give | Accuracy | Mistranslation | 10 |
| 5 | 𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴 তে হদু যার |
Where is he going? | that is enough | Accuracy | Mistranslation | 10 |
| 6 | 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴 তামিম ঘুমজার |
Tamim is sleeping. | timid. gum jor | Accuracy | Mistranslation | 10 |
| 7 | 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬 তুই হাময়ান গড়ি দে |
Please do the work. | You make ready / prepare | Accuracy | Mistranslation | 10 |
| # | Chakma Source | Human Reference | AI Output | Dimension | Category | Penalty |
|---|---|---|---|---|---|---|
| 2 | 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬 তুই হি গরিবে? |
What will you do? | you will do | Fluency | Grammar | 5 |
| 3 | 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨 তুই ঈদু আই |
Come here. | you gave yes / you did give | Fluency | Grammar | 5 |
| 4 | 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴 মুই ভাত হাঙর |
I am eating rice. | I eat rice | Fluency | Grammar | 5 |
| 6 | 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴 তামিম ঘুমজার |
Tamim is sleeping. | timid. gum jor | Fluency | Grammar | 5 |
| 7 | 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬 তুই হাময়ান গড়ি দে |
Please do the work. | You make ready / prepare | Fluency | Grammar | 5 |
| # | Chakma Source | Human Reference | AI Output | Dimension | Category | Penalty |
|---|---|---|---|---|---|---|
| 4 | 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴 মুই ভাত হাঙর |
I am eating rice. | I eat rice | Fluency | Grammar | 1 |
| 33 | 𑄟𑄪𑄭 𑄝𑄎𑄢𑄧𑄖𑄴 𑄡𑄬𑄟𑄴 মুই বাজারত্ যেম্ |
I will go to the market. | I am going to the market | Fluency | Grammar | 1 |
| 77 | 𑄃𑄟𑄨 𑄃𑄙 𑄊𑄧𑄚𑄴𑄑 𑄃𑄧𑄦𑄧 𑄇𑄟𑄴 𑄉𑄪𑄢𑄨𑄖𑄴𑄭 আমি আধাঘন্টা অল হাম গুরিত্তেই |
We have been working for half an hour. | I finished the work in half an hour | Accuracy | Mistranslation | 1 |
| 107 | 𑄖𑄪𑄭 𑄛𑄧𑄢𑄩𑄇𑄴𑄬𑄖𑄴 𑄛𑄥𑄴 𑄚𑄧 𑄉𑄧𑄖𑄬𑄧 তুই পরীক্কেত্ পাস ন গত্তে |
May you not pass the exam. | You did not pass the exam | Accuracy | Mistranslation | 1 |
| Outcome | Penalty Range | Sentences | % of Corpus | Avg. Penalty | Quality Band |
|---|---|---|---|---|---|
| Perfect | = 0 | 398 | 23.5% | 0.00 | Excellent |
| Acceptable | 1–5 | 255 | 15.1% | 5.00 | Acceptable |
| Poor | >5 | 1038 | 61.4% | 10.67 | Very Poor |
| Overall corpus average | 1,691 | 100% | 7.303 | Very Poor | |
* "Acceptable" band (1–5) here refers to sentence-level MQM score range, not the overall corpus quality band.
The following tables reproduce all example annotations from Jannat et al. (2026), organised by error type. Each table corresponds to a sub-section of §4.2 (Findings). Chakma script, pronunciation, human reference translation, and AI output are shown side by side.
| Error Category | MQM Dimension | Frequency | Percentage |
|---|---|---|---|
| Lexical Errors (Mistranslation, Terminology, Untranslated, Unintelligible) | Accuracy | 1,038 | 72.5% |
| Grammatical Errors (Omission, Addition, Word Order, Tense, Sentence Category, Plural) | Fluency | 315 | 22.0% |
| Cultural Errors (Kinship terms, Proverbs, Culture-specific vocabulary) | Locale | 79 | 5.5% |
| Total | 1,432 | 100% | |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄖𑄢𑄧𑄚𑄪𑄴 𑄦𑄮𑄢𑄩 𑄟𑄚𑄪𑄥𑄴 [Torun lobhi manush] | Young greedy man | Young good person |
| 2 | 𑄟𑄪 𑄉𑄟𑄧𑄴 𑄃𑄉𑄧𑄋𑄴 [Gom agong] | I am fine | I am coming |
| 3 | 𑄇𑄙𑄧 𑄥𑄚𑄪𑄨 𑄦𑄎𑄣𑄪𑄨 [Kotho shune hajaile] | Hearing this, he smiled | Hearing this, he laughed |
| 4 | 𑄟𑄢𑄬𑄧 𑄝𑄚𑄻 𑄟𑄚𑄨𑄒𑄨𑄴 𑄥𑄟𑄧𑄠𑄧𑄴 [More bana pach minit shomoyde] | Give me just five minutes | Give me one minute time |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴 [Tui hi goror?] | What are you doing? | You are beautiful |
| 2 | 𑄖𑄬 𑄢𑄟𑄧𑄌𑄧𑄇𑄧𑄳𑄝𑄴𑄧𑄛 [Te romchokro poe] | He is an aggressive boy | That is a tree |
| 3 | 𑄖𑄬 𑄃𑄬𑄇𑄴𑄬𑄢𑄬 𑄉𑄢𑄧𑄝𑄨𑄴𑄚 [Te ekkere gorib noi] | He is not poor at all | He is very poor |
| 4 | 𑄖𑄬 𑄦𑄘𑄧𑄪 𑄡𑄢𑄴 [Te hodu jar?] | Where is he going? | That is enough |
| # | Chakma Source | Human Translation | AI Output |
|---|---|---|---|
| 1 | 𑄖𑄟𑄟𑄨𑄴৷ 𑄊𑄪𑄟𑄴𑄎𑄢𑄴 [Tamim ghumjar] | Tamim is sleeping | Timid gum jor |
| 2 | 𑄃𑄯 𑄟𑄚𑄪𑄥𑄴𑄮 𑄖𑄢𑄬𑄧 𑄃𑄬𑄇𑄴𑄢𑄚𑄨𑄨 𑄌𑄬𑄭𑄃𑄊𑄬 [O manussho tore ek rini cheiaghe] | That person is looking at you | O manush tore ekrini jeia-ge |
| 3 | 𑄇𑄘𑅅 𑄉𑄚𑄩𑄘𑄮𑄦𑄬 𑄌𑄝𑄬𑄥𑄴 [Khaddo gani dole chabes] | Chew your food well | Kadui goni dohe jôbes |
| 4 | 𑄛𑄝𑄨𑄢𑄨𑄴 𑄛𑄝𑄨𑄢𑄨𑄴 𑄝𑄠𑄬𑄢𑄧𑄴 𑄝𑄢𑄴 [Pibir pibir boyer bar] | The wind is blowing with a whistling sound | pipir pibir boyer bor |
| # | Chakma Source | Human Translation | AI Output |
|---|---|---|---|
| 1 | 𑄖𑄬 𑄦𑄚𑄘𑄬𑄧 𑄦𑄚𑄘𑄬𑄧 𑄃𑄬 𑄇𑄘𑄧 𑄦𑄮𑄠𑄬 [Te hanode hanode ei hoda hoye] | He said this while crying | that/then Hanode Hanode, this word is/was |
| 2 | 𑄖𑄬 𑄃𑄎𑄘𑄬𑄧 𑄃𑄎𑄘𑄬𑄧 𑄉𑄦𑄧𑄛𑄧𑄝𑄧𑄱𑄦𑄦𑄧𑄧 [Te ajte ajte golopbu holo] | He told the story laughing | that very much, very much beautiful is |
| 3 | 𑄖𑄬 𑄑𑄇𑄨𑄴 𑄖 𑄝𑄝𑄧𑄘𑄧𑄚 𑄝𑄚𑄨 𑄃𑄨𑄘 [Te thik ta babo doken bini ada] | He is just as lively as his father | then/so, that, father, to see, went out |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄖𑄬 𑄝𑄎𑄢𑄖𑄧𑄴 𑄡𑄠𑄬𑄨 [Te bajarot jiye] | He had gone to the market. | to the market |
| 2 | 𑄖𑄢𑄧𑄚𑄪𑄴 𑄃𑄇𑄨𑄃𑄥𑄴𑄬𑄨 𑄘𑄬𑄢𑄃𑄩 [Torun a hee isse deri?] | Tarun, why were you late today? | you are very late |
| 3 | 𑄦𑄬𑄉𑄨𑄚𑄧𑄎𑄚𑄁𑄉𑄬𑄧 [Mui legi no janongge] | I don't know how to write. | I do not know |
| 4 | 𑄢𑄮𑄉𑄝𑄮𑄨 𑄃𑄬𑄇𑄴 𑄥𑄛𑄴𑄖 𑄛𑄢𑄬𑄧 𑄟𑄪𑄢𑄝𑄮𑄨 [Rogibo ek shapta pore moribo] | The patient will die after one week | will die in a week |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄟𑄪𑄭 𑄚𑄧𑄛𑄢𑄟𑄨𑄴 [Mui noparim] | I can't | I cannot do |
| 2 | 𑄟𑄢𑄬𑄧 𑄃𑄬𑄇𑄴 𑄉𑄦𑄧𑄥𑄧𑄧𑄛𑄚𑄨𑄃𑄚𑄨𑄘𑄬 [Mor e ek golos pani ani do] | Bring me a glass of water | Please bring me a glass of water |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| WO-1 | 𑄃𑄝𑄢𑄬𑄇𑄁𑄧 [Abare ho] | Say it again | Again say/speak |
| AP-1 | 𑄔𑄪𑄣𑄮𑄢𑄴 𑄖𑄣𑄬 𑄖𑄣𑄬 𑄚𑄌𑄴 𑄦𑄢𑄧𑄴𑅁 [Dhulor taale taale nach hor] | There is dancing to the beat of the drum. | Dancing goes on to the rhythm of the drum. |
| AP-2 | 𑄡𑄬 𑄛𑄮𑄝𑄪 𑄉𑄖𑄩𑄴 𑄉𑄢𑄴 𑄥𑄝𑄬𑄨 𑄟𑄢𑄧𑄴 𑄞𑄬𑄭 [Je pobu geet gar shibe mor bhei] | The boy who is singing is my brother. | Whoever sings a song, he is my brother. |
| AP-3 | 𑄉𑄖𑄩𑄴 𑄉𑄬𑄠𑄬𑄝𑄮 𑄇𑄟𑄴 𑄉𑄢𑄬𑄧𑄢𑄴 [Geet geyebu kaam gorer] | The singer is working. | A singer does work. |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄓𑄇𑄴𑄖𑄢𑄧𑄴𑄝𑄪𑄟𑄧𑄌𑄮𑄇𑄴𑄚𑄪𑄴𑄛𑄢𑄧𑄇𑄨𑄴𑄉𑄢𑄧𑄣𑄮𑄨 [Daktorbu mo hattani porikkhe gorilo] | The doctor examined my hand. | The doctor examined my hands. |
| 2 | 𑄥𑄖𑄧𑄳𑄢𑄪𑄡𑄬𑄚𑄴 𑄝𑄚𑄧𑄴𑄘𑄪𑄦𑄘𑄧𑄇𑄴𑅁 [Shotru jeno bondhu hodak] | May enemies become friends. | Let an enemy become a friend. |
| 3 | 𑄉𑄌𑄴𑄍𑄮𑄖𑄴𑄇𑄴𑄖𑄣𑄴𑄇𑄴 𑄃𑄊𑄚𑄧𑄴 [Gacchot ektal pek agon] | There are many birds in the tree. | There is a bird on the tree. |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄖𑄪𑄭𑄛𑄢𑄧𑄩𑄖𑄴𑄛𑄥𑄴𑄚𑄧𑄉𑄖𑄬𑄧𑄧 [Tui porikket pash no gotte] | May you not pass the exam (subjunctive wish/curse) | You did not pass the exam (past declarative) |
| 2 | 𑄃𑄭𑄞𑄖𑄴𑄭 [Ai bhaat hei] | Come, let's eat rice. (hortative/suggestion) | I eat rice (simple declarative) |
| 3 | 𑄖𑄪𑄃𑄨𑄦𑄨𑄉𑄢𑄧𑄝𑄬𑄨 [Tui hi goribe?] | What will you do? (interrogative) | you will do (incomplete declarative) |
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄟𑄪𑄭𑄆𑄇𑄴𑄮 𑄛𑄬𑄈𑄴 𑄘𑄬𑄉𑄁𑄧 [Mui aekko pek degong] | I see a bird (present) | I saw a bird (past) |
| 2 | 𑄟𑄪𑄭 𑄝𑄎𑄢𑄖𑄧𑄴𑄡𑄬𑄟𑄴 [Mui bajarot gem] | I will go to the market (future) | I am going to the market (present progressive) |
| 3 | 𑄖𑄢𑄳𑄦𑄟𑄍𑄴𑄙𑄢𑄧𑄘𑄧𑄚𑄧𑄴 [Tarah mach dhordon] | They are catching fish (present progressive) | he/she caught fish (past + singular) |
| 4 | 𑄇𑄚𑄨𑄴𑄖𑄪𑄟𑄪𑄭𑄎𑄬𑄝𑄢𑄴𑄚𑄌𑄁 [Hintu mui jebar no chang] | But I don't want to go (volitional present) | But I will not go (future negative) |
| Chakma Source | Human Translation | AI Translation | Cultural Issue |
|---|---|---|---|
| 𑄟𑄧𑄟𑄟𑄪𑄊𑄢𑄧𑄖𑄧𑄴 𑄃𑄉𑄬𑅁 [Mo mamu ghorot age] |
My maternal uncle is at home | My uncle is at home | 'Maternal' distinction lost — Chakma distinguishes paternal/maternal kin |
| 𑄡𑄬 𑄉𑄎𑄖𑄧𑄴 𑄜𑄣𑄧𑄴 𑄙𑄢𑄬𑄧, 𑄥𑄬 𑄉𑄎𑄖𑄧𑄴 𑄃𑄘𑄬𑄨 𑄛𑄢𑄬𑄧𑅁 [Je gajot fol dhore, se gajot eide pore] |
The tree that bears fruit gets stones thrown at it. (Success invites criticism) | The tree that bears fruit bends down. | Proverb partially interpreted; cultural meaning (social criticism of success) lost |
| 𑄖𑄬 𑄇𑄝𑄨𑄚𑄎𑄧𑄖𑄉𑄣𑄧𑄧𑄘𑄌𑄪𑄴𑄘𑄬। [Te kabi negate taglanre duch de] |
If you cannot dance, you blame the courtyard for being crooked. (Blaming circumstances for one's own failings) | Better late than never. | Entirely wrong proverb substituted — no semantic relationship to source |
| 𑄛𑄚𑄖𑄨𑄴𑄇𑄪𑄟𑄢𑄮𑄴, 𑄟𑄪𑄢𑄮𑄖𑄴 𑄝𑄇𑄴𑅁 [Panit kumor, murot bak] |
Crocodile in the water, tiger on land. (Danger on all sides) | Rich in words, poor in deeds | Entirely wrong proverb — AI substituted an unrelated English idiom |
| 𑄖𑄬𑄣𑄴𑄖𑄬𑄣𑄳𑄠𑄬 𑄥𑄢𑄬𑄖𑄨𑄴 𑄖𑄬𑄣𑄴 𑄘𑄬𑄚𑅁 [Teltello siret tel dena] |
Oiling an already oily head. (Giving to those who already have) | teltelye siret tel den | Completely untranslated — AI returned romanised Chakma |
| 𑄝𑄬𑄇𑄴𑄚𑄬𑄪 𑄟𑄣𑄨𑄚𑄬𑄨𑄭 𑄃𑄣𑄴𑄛𑄚𑄧 𑄃𑄉𑄘𑄚𑄧𑄴𑅁 [Bekkune miline alpona agadon] |
Everyone together is drawing alpana (floor art). | Women together draw alpana. | Universal "everyone" narrowed to "Women"; cultural term 'alpana' passed through unchanged |
[1] Al Sharou, K., & Specia, L. (2022). Towards a better understanding of noise in natural language processing. Proceedings of the 13th Language Resources and Evaluation Conference.
[2] Abdelhalim, S. M., Alsahil, A. A., & Alsuhaibani, Z. A. (2025). Artificial intelligence tools and literary translation: a comparative investigation of ChatGPT and Google Translate from novice and advanced EFL student translators' perspectives. Cogent Arts & Humanities, 12(1), 2508031.
[3] Afaq, M., Mehmood, T., & Ayaz, M. O. (2025). Can artificial intelligence challenge universal grammar? A theory-driven empirical investigation. Journal of Applied Linguistics and TESOL (JALT), 8(4), 1248–1254.
[4] Afreen, N. (2020). Language usage in different domains by the Chakmas of Bangladesh. International Journal of Linguistics, Literature and Translation, 3(6), 135–151.
[5] Ajani, Y. A., Oladokun, B. D., Olarongbe, S. A., Amaechi, M. N., Rabiu, N., & Bashorun, M. T. (2024). Revitalizing indigenous knowledge systems via digital media technologies for sustainability of indigenous languages. Preservation, Digital Technology & Culture, 53(1), 35–44.
[6] Anik, M., Rahman, A., Wasi, A., & Ahsan, M. (2025, May). Preserving cultural identity with context-aware translation through multi-agent AI systems. In Proceedings of the 1st Workshop on Language Models for Underserved Communities (LM4UC 2025) (pp. 51–60).
[7] Bal, E. (2010). Being Mog: Memories, nostalgia, and identity of the Mog community in Bangladesh. Modern Asian Studies, 44(6), 1261–1295.
[8] Bassnett, S., & Trivedi, H. (1999). Introduction: Of colonies, cannibals and vernaculars. In S. Bassnett & H. Trivedi (Eds.), Post-colonial translation: Theory and practice (pp. 1–18). Routledge.
[9] Bishop, M. (2022). Elders' conversations: Perspectives on leveraging digital technology in language revival. The Open/Technology in Education, Society, and Scholarship Association Journal, 2(2), 1–13.
[10] Chakma, A., Khisa, A., Khisa, S., Noor, J., & Sultana, S. (2026). Re-educating educated ones: A case study on Chakma language revitalization in Chittagong Hill Tracts. arXiv preprint arXiv:2601.12290.
[11] Chakma, J. (2010). Origin and evolution of Chakma language and script. Kriti Rakshana, National Mission for Manuscripts.
[12] Chakma, J., & Sultana, A. (2023). Language rights and indigenous peoples of the Chittagong Hill Tracts. International Journal of Language and Culture.
[13] Çetin, Ö., & Duran, A. (2024). A comparative analysis of the performances of ChatGPT, DeepL, Google Translate and a human translator in community-based settings. Amasya Üniversitesi Sosyal Bilimler Dergisi, 9(15), 120–173.
[14] Chiran, R. (2025). Language endangerment in Bangladesh: An updated assessment. South Asian Languages Review.
[15] Dovchin, S. (2020). Introduction to special issue: Linguistic racism. International Journal of Bilingual Education and Bilingualism, 23(7), 773–777.
[16] Drude, S., & Intangible Cultural Heritage Unit's Ad Hoc Expert Group. (2003). Language vitality and endangerment. UNESCO.
[17] Ducharme, Q. M., Amatulli, G., Williams, W. A. L., George, S. H., Pierre, S. M., & Pierre, S. L. R. (2025). Revitalizing indigenous languages, fostering self-governance, overcoming the Indian Act: A case study of Lil'wat Nation. Canadian Public Administration, 68(3), 470–486.
[18] Folaron, D. (2015). Translation and minority, lesser-used and lesser-translated languages and cultures. The Journal of Specialised Translation, 24, 16–27. https://doi.org/10.26034/cm.jostrans.2015.320
[19] Fu, Y., & Liu, Y. (2024). Evaluating ChatGPT's translation quality in scientific texts. Language & Technology Review.
[20] Grenoble, L. A., & Whaley, L. J. (2005). Saving languages: An introduction to language revitalization. Cambridge University Press.
[21] Gwerevende, S., & Mthombeni, Z. M. (2023). Safeguarding intangible cultural heritage: exploring the synergies in the transmission of indigenous languages, dance and music practices in Southern Africa. International Journal of Heritage Studies, 29(5), 398–412.
[22] Holmes, J. (2021). An introduction to sociolinguistics (4th ed.). Routledge.
[23] Hutson, J., Ellsworth, P., & Ellsworth, M. (2024). Preserving linguistic diversity in the digital age: a scalable model for cultural heritage continuity. Journal of Contemporary Language Research, 3(1).
[24] Jerome, C., et al. (2022). Language, identity and indigenous communities. Journal of Language and Cultural Studies.
[25] Jiang, Z., Lv, Q., Zhang, Z., & Lei, L. (2023). Distinguishing translations by human, NMT, and ChatGPT: A linguistic and statistical approach. arXiv.
[26] Kandler, A., & Unger, R. (2023). Modeling language shift. In Diffusive spreading in nature, technology and society (pp. 365–387). Springer International Publishing.
[27] Koc, V. (2025). Generative AI and large language models in language preservation: Opportunities and challenges. arXiv (Cornell University).
[28] Lepp, A., & Sarin, L. (2024). Linguistic justice and digital inequality. Language Policy & Technology Review.
[29] Li, M., Croucher, S. M., & Shen, L. (2024). Language endangerment and the linguistic vitality of Miao in China: cultural shifts and revitalisation strategies. Journal of Multilingual and Multicultural Development, 1–16.
[30] Mahi, M. H., Khan, A. R., Anik, M. H., Noori, S. R. H., Mahmud, A., & Mojumdar, M. U. (2025). MELD: a multilingual ethnic dataset of Chakma, Garo, and Marma in Bengali script with English and standard Bengali translation. Data in Brief, 61, 111745.
[31] Mohamed, M., et al. (2024). Translation quality in the age of AI. Language & Technology.
[32] Mohsin, A. (2023). Indigenous languages of Bangladesh. University Press Limited.
[33] Moneus, A. M., & Sahari, Y. (2024). Artificial intelligence and human translation: A contrastive study based on legal texts. Heliyon, 10(6).
[34] MQM. (2015). Multidimensional quality metrics definition. Retrieved from https://web.archive.org/web/20210113220425/http://www.qt21.eu/mqm-definition/definition-2015-05-27.html
[35] O'Hagan, M. (2016). Massively open translation: Unpacking the relationship between technology and translation in the 21st century. International Journal of Communication, 10, 18.
[36] Okafor, A. Y. (2025). Examining AI translation errors in Igbo: Lexical ambiguity, misinterpretation, and incorrect word substitutions due to contextual deficiencies. Indonesian Journal of Learning Studies, 5(1), 46–55.
[37] Oladipupo, F., Soronnadi, A., Adebara, I., & Adekanmbi, O. (2025, August). How effective are AI models in translating English scientific texts to Nigerian Pidgin: A low-resource language? In I Can't Believe It's Not Better: Challenges in Applied Deep Learning.
[38] Rafat Al Rousan, Raghad Jaradat, & Mona Malkawi. (2025). ChatGPT translation vs. human translation: an examination of a literary text. Cogent Social Sciences, 11(1), 2472916. https://doi.org/10.1080/23311886.2025.2472916
[39] Ranathunga, S., Lee, E. S. A., Prifti Skenduli, M., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages: A survey. ACM Computing Surveys, 55(11), 1–37.
[40] Saikia, M., & Ullman, J. (2023). Endangered language assessment framework. Language Documentation Journal.
[41] Sevinç, Y. (2022). Language endangerment and revitalization. Annual Review of Linguistics.
[42] Smith, B. K., Ehala, M., & Giles, H. (2017). Vitality theory. In J. Nussbaum (Ed.), Oxford research encyclopedia of communication. Oxford University Press.
[43] Spivak, G. C. (1993). Outside in the teaching machine. Routledge.
[44] Tymoczko, M. (1999). Translation in a postcolonial context: Early Irish literature in English translation. St. Jerome Publishing.
[45] Tsunoda, T. (2006). Language endangerment and language revitalisation: An introduction. Mouton de Gruyter.
[46] UNESCO. (2003). Language vitality and endangerment. Ad Hoc Expert Group on Endangered Languages.
[47] Walsh, J. (2006). Language and socio-economic development: Towards a theoretical framework. Language Problems and Language Planning, 30(2), 127–148.
[48] Wei, L., Hua, Z., & Simpson, J. (Eds.). (2023). The Routledge handbook of applied linguistics: Volume two. Taylor & Francis.
[49] Yan, J., Yan, P., Chen, Y., Li, J., Zhu, X., & Zhang, Y. (2024). GPT-4 vs. human translators: A comprehensive evaluation of translation quality across languages, domains, and expertise levels. arXiv preprint arXiv:2407.03658.
Keywords: Chakma language | indigenous language preservation | linguistic justice | AI-based translation | Multidimensional Quality Metrics (MQM)
Language endangerment has become increasingly widespread worldwide, with 14 indigenous languages being on the verge of extinction in the land of Bangladesh alone (Chiran, 2025). The Chakma language, spoken by the Chakma community that resides in Chittagong Hill Tracts (CHT), is one such 'definitely endangered' language (Saikia & Ullman, 2023) that faces ongoing threat due to socio-political marginalization. Preserving the Chakma language is essential for sustaining linguistic diversity, cultural practices, and indigenous knowledge systems embedded within the language. Inspired by the potential of AI translation tools, the study investigates the extent to which such technologies can support endangered language documentation. Specifically, the study asks: (1) how accurately do AI-based translation systems translate Chakma texts into English when compared with native-speaker–validated human translations, and (2) what types of lexical, grammatical, and culturally grounded errors recur in AI-generated translations? To address these questions, a human-validated reference corpus of 1,691 Chakma–English sentence pairs is compiled, and AI-generated translations are systematically evaluated using the Multidimensional Quality Metrics (MQM) framework across three quality dimensions—accuracy, fluency, and cultural appropriateness—in conjunction with an error analysis approach validated by native Chakma speakers. Of the 1,691 sentences that were fully annotated, the AI system produced an acceptable translation for 398 sentences (23.5%), while 1,432 error occurrences were identified in total—dominated by lexical/meaning errors (72.5%) and grammatical errors (22.0%)—yielding a weighted MQM error penalty score of 7.303 penalty points per sentence, which falls in the 'acceptable quality' band. To support the interpretation of translation outputs, the study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors. Since large language models rely on attention-based mechanisms trained predominantly on high-resource languages, they remain ill-equipped for Chakma's morphological complexity. The Bi-LSTM-CRF effectively addresses this gap, achieving a grammatical pattern recognition accuracy of 87.4%. By foregrounding native-speaker validation, linguistic analysis, and systematic evaluation of AI-generated translations, the study demonstrates how AI translation tools can be critically assessed and responsibly operationalized as equitable resources for endangered language documentation, while laying the groundwork for more effective AI-supported approaches to Chakma language preservation.
The world hosts a vast array of languages essential to humanity's heritage (Drude, 2003). Currently, over 7,000 languages are spoken across the globe; however, this remarkable linguistic richness confronts substantial risks as modernity advances (Hutson et al., 2024), causing 40% of the languages to head toward extinction (Eberhard et al., 2022, as cited in Li et al., 2024). According to the Language Conservancy, after global warming, language loss is recognized as the planet's most pressing crisis (Collette & Kennedy, 2023, as cited in Hutson et al., 2024). Numerous endangered languages are diminishing rapidly due to globalization and modernization. This decline is particularly concerning because linguistic diversity is essential for transmitting culture, values, beliefs, and history across generations (Sevinç, 2022). The unique words, phrases, and expressions of each language encapsulate the accumulated knowledge and experiences of its speakers, shaping identity and fostering a sense of pride (Jerome et al., 2022). For many indigenous communities, languages are key carriers of culture, containing unique communication systems, traditional knowledge, and a strong sense of identity (Gwerevende & Mthombeni, 2023). Therefore, the dramatic loss of these minority languages signifies more than silenced voices; it entails the epistemic erasure of invaluable cultural knowledge and distinct worldviews (Kandler & Unger, 2023).
Within this broader global context, Bangladesh is no exception. The country, characterised by its rich ethnic diversity, is home to multiple indigenous communities, many of whose languages are increasingly at risk of decline, mainly due to globalization, urbanization, and the extensive use of politically dominant languages in social, educational, and professional spheres (Anik et al., 2025). The global decline of indigenous languages has reached a critical point, with studies indicating that around half of them could disappear within this century. The Kuruk language, for instance, is no longer spoken, while languages such as Pankho, Khumi, and Hajong are barely surviving (Mohsin, 2023). UNESCO (2003) identifies globalization, forced displacement, and assimilation-driven policies as the main forces behind this decline. Similarly, Chakma and Sultana (2023) argue that the loss of ancestral lands, environmental degradation, and the gradual erosion of cultural identities have further accelerated the decline of indigenous languages. Additionally, the increasing tendency to use native languages only in private settings gradually reduces speakers' fluency and intergenerational transmission, thereby heightening the risk of language extinction (Holmes, 2021).
Endangered languages like Chakma are facing both social and political marginalization reflecting the patterns of linguistic racism. In other words, minority voices are subdued while dominant languages are prioritized (Dovchin, 2020, 2025; Rosa & Flores, 2021). This broader pattern is clearly visible in the Chittagong Hill Tracts (CHT) of Bangladesh, where the Chakma people have been among the most affected by language policies. After independence, Bangladesh followed a 'one culture, one language' idea. This language policy reinforced the prominence of the native language, Bangla, while diminishing the visibility and status of minority languages and identities (Bal, 2010). Chakma and Sultana (2023) describe this as a form of language control, where indigenous people feel pressured to stop using their native languages. In light of this linguistic marginalization, translation acts as an important platform for resistance. Translation is not merely a technical process of linguistic transfer anymore. Instead, it is understood as a political and ethical act that can challenge the dominance of certain languages over others. Tymoczko (1999) asserts that translation has formed the cultural politics of colonized societies. It has helped communities preserve and negotiate their national and cultural identities. In colonial and postcolonial settings, translation is closely linked with power, representation, and cultural authority (Bassnett & Trivedi, 1999). Folaron (2015) argues that translation helps indigenous languages survive and gain recognition simply by increasing their exposure beyond their immediate communities. It is also perceived as an excellent strategy in reclaiming disadvantaged voices (Spivak, 1993). Through this lens, translating indigenous languages becomes more than documentation by being an act of resistance against linguistic marginalization and a way of affirming indigenous identity.
Although Machine Translation (MT) research has advanced considerably through neural and large language model–based approaches, it continues to focus predominantly on high-resource language pairs supported by large-scale parallel corpora and extensive training data. Evaluation also frequently relies on automated metrics such as BLEU and TER, which often obscure crucial shortcomings by failing to capture deeper issues in translation quality. Al Sharou and Specia (2022) demonstrate that in low-resource and user-generated environments, the existence and severity of errors, particularly mistranslations, omissions, and hallucinations, are more consequential than fluency levels. While efforts are made in revitalizing languages around the globe, endangered languages, particularly in South Asia, such as the Chakma language, remain largely underexplored in NLP research (Chakma et al., 2024). Specifically, there is a notable lack of empirical research on Chakma–English translation using AI-based systems and very little is known about the lexical, grammatical, and culturally grounded errors that recur in their translation outputs. Without addressing these gaps, AI translation risks misrepresenting minority languages and undermining language preservation efforts. Therefore, this study employs a human-validated Chakma–English corpus of 1,691 sentences to systematically assess the performance of AI translations, examining how accurately AI can render Chakma texts into English while maintaining both linguistic fidelity and cultural nuances. Patterns of recurring AI errors are also investigated through a structured MQM-based evaluation framework. The study also examines whether a Bi-LSTM-CRF can identify grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.
Considering the aim of the study, the following questions were formulated:
RQ1. To what extent do AI-based translation systems accurately translate Chakma texts into English when compared with native-speaker-validated human translations?
RQ2. Which types of lexical, grammatical, and culturally grounded errors occur most frequently in AI-generated translations?
Among the 38 regional languages spoken in Bangladesh, 14 indigenous languages face the threat of extinction (Chiran, 2025). Due to historical power dynamics (Bishop, 2022) and various socio-political factors, the development and expansion of these ancestral languages have become progressively more challenging (Awal, 2019). Chakma, Marma, Tripura, Mro, and Murung are among the notable underrepresented indigenous communities in Bangladesh, among which the Chakma constitute the largest ethnic indigenous group (Afreen, 2020). The Chakma community resides primarily in the Chittagong Hill Tracts (CHT) in the southeastern region of the country. Their mother tongue, the Chakma language, is spoken by approximately 600,000 to 1,000,000 people across the CHT and parts of India (Chakma, 2010) and is classified as 'definitely endangered' (Saikia & Ullman, 2023), indicating that children are no longer consistently learning the language at home (UNESCO). Li et al. (2024) argue that a language's sustainability is ensured when it is actively used across diverse domains such as the home, educational institutions, workplaces, religious settings, and media. As a language loses visibility within these domains, everyday usage gradually declines, threatening the cultural practices, oral traditions, and intergenerational knowledge systems embedded within it (Tsunoda, 2006).
Machine Translation (MT) refers to 'computerized systems responsible for the production of translations with or without human assistance' (Hutchins, 1995, p. 1). With substantial advancements in technology, MT has become an effortless and accessible tool for quickly translating spoken and written texts across languages. MT has evolved through several major paradigms — rule-based (RBMT), statistical (SMT), hybrid, and most recently neural machine translation (NMT) — each improving upon the limitations of its predecessor. NMT systems use deep neural networks based on encoder–decoder architectures to model translation as a sequence-to-sequence task (Bahdanau et al., 2015; Cho et al., 2014). Compared to earlier systems, NMT improves contextual understanding, reduces literal translations, and enhances scalability and efficiency, leading to widespread adoption in major translation systems. Despite these advancements, translation quality remains inconsistent for low-resource languages (Jiang et al., 2023) due to limited training data and linguistic underrepresentation in existing corpora (Zhong et al., 2025). This highlights the persistent challenges faced by contemporary AI-based translation systems in handling linguistically underrepresented languages.
Artificial Intelligence (AI), particularly Natural Language Processing (NLP) and Large Language Models (LLMs), has emerged as a promising tool for preserving and revitalizing endangered and minority languages through the documentation, analysis, and translation of linguistic resources (Koc, 2025). Despite these opportunities, the application of AI to endangered language preservation remains accompanied by significant challenges. Research suggests that digitized documentation efforts often struggle to accurately capture the cultural complexities and linguistic nuances inherent in minority languages (Hutson et al., 2024; Ingram, 2025). Although AI systems can generate grammatically coherent translations and facilitate language accessibility (Putri et al., 2024), they frequently fall short in capturing the cultural and contextual subtleties that human translators can reliably interpret (Moneus & Sahari, 2024). Studies indicate that machine translation may distort contextual meaning, overlook idiomatic expressions and historical significance, and lack the cultural depth and real-world understanding necessary for effective language preservation (Hutson et al., 2024; Okafor, 2025; Putri et al., 2024). Furthermore, current AI-driven approaches to language translation frequently prioritize efficiency over cultural authenticity, overlooking broader goals of linguistic preservation (Mufwene, 2005; Anik et al., 2025). The dominance of English-centric AI models further reinforces existing linguistic hierarchies, marginalizing lesser-known languages and limiting their digital accessibility (Lepp & Sarin, 2024).
Translation quality evaluation is concerned with determining how effectively a translation conveys the meaning and communicative intent of the source text. The Multidimensional Quality Metrics (MQM) framework (MQM, 2015) enables systematic identification and classification of translation issues across multiple dimensions, allowing for both holistic quality assessment and fine-grained error analysis. In line with this framework, the present study assesses overall translation quality along three dimensions—accuracy (the meaning is correct), fluency (the grammar is correct and the output is natural), and cultural appropriateness (culturally specific meaning is preserved)—while the error analysis classifies individual problems into the lexical, grammatical, and culturally grounded categories examined in the research questions. This combined approach allows for a comprehensive assessment of both the overall quality and the underlying error patterns in AI-generated translations.
This study uses a mixed-methods design to evaluate how well a large language model translates sentences from Chakma — an endangered language — into English. We compare AI-generated translations against human-validated reference translations, sentence by sentence, using the Multidimensional Quality Metrics (MQM) framework. Quantitative scoring gives us a number we can compare across systems or studies; qualitative analysis tells us why errors happen and what they mean for the language community that depends on accurate documentation.
The decision to use a native-speaker-validated reference corpus rather than automated metrics (BLEU, chrF) was deliberate. Automated metrics measure surface similarity and are known to be unreliable for morphologically rich, low-resource languages. A human reference with native-speaker sign-off is the only meaningful gold standard for Chakma.
We built a Chakma–English parallel corpus of 1,691 sentence pairs drawn from everyday speech: greetings, questions, requests, descriptions of daily activities, and culturally specific expressions including kinship terms, pragmatic markers, and community phrases that do not have direct English equivalents.
Each sentence was written in Chakma Unicode script (U+11100–U+1114F), accompanied by a pronunciation guide in Bengali script for readability, and paired with an English translation produced by a bilingual researcher. Every translation was then reviewed by native Chakma speakers who corrected phrasing, cultural nuance, and meaning where necessary. The validated translation became the reference standard against which all AI output was measured. Of the 1,691 sentence pairs, 398 (23.5%) required no correction by the native reviewer and were accepted as written.
The system under evaluation is ChatGPT (OpenAI, GPT-4 architecture), chosen because it represents the current frontier of accessible AI translation. Unlike specialised Neural Machine Translation (NMT) systems, ChatGPT has not been explicitly trained on Chakma data, making this evaluation a realistic test of what a researcher or practitioner would encounter if they used a widely available tool for Chakma documentation.
Each Chakma sentence was submitted to the model with a standard instruction: "Translate the following Chakma sentence into English." The raw output was recorded without any post-editing or re-prompting. This zero-post-edit protocol ensures the evaluation reflects the model's actual performance, not the performance achievable with additional human intervention.
We adopt the Multidimensional Quality Metrics (MQM) framework, an industry-standard annotation scheme developed for professional translation evaluation. MQM provides a structured taxonomy of error types and assigns quantitative severity weights, making it possible to produce a single comparable score per sentence and per corpus.
Each AI translation was compared against the human reference sentence by sentence. Every detected deviation was recorded as a separate MQM error and assigned three attributes: a Dimension, a Category, and a Severity.
Errors are classified under three MQM dimensions, each with its own sub-categories:
| Dimension | Category | What it means |
|---|---|---|
|
Accuracy (Lexical Error) |
Mistranslation | The AI chose a word or phrase that means something different from the source |
| Omission | A word or phrase present in the source was dropped entirely | |
| Addition | Words were added that have no basis in the source | |
| Untranslated | Source text was left in Chakma script rather than translated | |
| Terminology | A domain-specific term was rendered incorrectly | |
|
Fluency (Grammatical Error) |
Grammar | Wrong tense, incorrect verb form, or grammatically malformed sentence |
| Word Order | Constituents are in the wrong position for natural English | |
| Agreement | Subject–verb or number agreement is broken | |
| Spelling | Transcription or spelling error in the output | |
| Punctuation | Missing or incorrect punctuation altering readability | |
|
Locale Convention (Cultural Error) |
Cultural Appropriateness | A culturally specific term — kinship word, pragmatic marker, social register — was flattened or lost |
Each annotated error is assigned one of three severity levels. The severity determines the penalty point added to the sentence's MQM score:
| Severity | Penalty | Definition | Example scenario |
|---|---|---|---|
| Critical | 10 | The translation is misleading or conveys an incorrect message. The reader would be misinformed. | "What are you doing?" → AI outputs "you are beautiful" — completely different meaning |
| Major | 5 | The error significantly reduces quality or changes meaning. The core message is altered. | "I am eating rice" → AI outputs "I eat rice" — tense lost, aspectual distinction gone |
| Minor | 1 | The meaning is mostly preserved. The error has little practical impact on communication. | A missing article or a stylistic word-order preference |
If a sentence contains multiple errors, each is annotated separately and all penalties are summed. For example, a sentence with one Critical Accuracy error and one Major Fluency error would receive a sentence score of 10 + 5 = 15.
The overall corpus MQM score is the average penalty per sentence across all N sentences:
We define five quality bands to interpret the MQM score. These thresholds are stated explicitly so that future studies can use the same scale for cross-system comparison:
| Band | MQM Score Range | Interpretation | Publication readiness |
|---|---|---|---|
| Excellent | 0.00 – 0.99 | Near-perfect translation; only trivial errors | Ready for direct publication |
| Good | 1.00 – 2.99 | High quality; minor corrections needed | Ready after light review |
| Acceptable | 3.00 – 4.99 | Usable with human post-editing | Suitable for internal drafts |
| Poor | 5.00 – 6.99 | Significant errors present; not reliable | Not suitable without revision |
| Very Poor ← | ≥ 7.00 | Systematic failure; misleading output likely | Requires complete re-translation |
This corpus scored 7.303, placing it firmly in the Very Poor band — the AI output cannot be used for documentation or publication without comprehensive human correction.
Annotation was carried out in six stages:
Across 1,691 sentences, annotators identified 1,432 MQM errors in total — an average of 0.85 errors per sentence. Of these, 1,038 (72.5%) were Accuracy errors at Critical severity, reflecting systematic semantic failure rather than isolated mistakes. Sentences scoring zero penalty (perfect translations) numbered 398 (23.5%). The remaining 1,293 sentences (76.5%) required at least one error annotation.
To support interpretation of Fluency and Accuracy error patterns identified through MQM annotation, this study also examines whether a Bidirectional Long Short-Term Memory Conditional Random Field (Bi-LSTM-CRF) model can identify the grammatical and verb pattern structures in Chakma that contribute to recurring translation errors.
Large language models such as ChatGPT rely on attention-based mechanisms trained predominantly on high-resource languages and remain ill-equipped to handle Chakma's morphological complexity — particularly its verb-final sentence structures, aspect-marking suffixes, and evidential markers. The Bi-LSTM-CRF is a sequence labelling model suited to identifying morphosyntactic patterns without relying on pre-trained multilingual embeddings. Applied to this corpus, it achieved a grammatical pattern recognition accuracy of 87.4%, providing a structural map of which Chakma constructions the AI is most likely to mistranslate. The model is not used as a translation system; it serves as a diagnostic layer linking MQM annotation findings to underlying grammatical causes.
This section presents the findings of the systematic analysis of AI-generated Chakma–English translations compared with a native-speaker-validated human reference corpus of 1,691 sentence pairs. A mixed-methods approach is adopted, employing the Multidimensional Quality Metrics (MQM) framework to identify and classify translation errors across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error), with each error assigned a severity level (Critical = 10, Major = 5, Minor = 1) and summed to produce a per-sentence penalty score. All annotations were validated by native Chakma speakers to ensure linguistic accuracy, contextual appropriateness, and cultural sensitivity.
The overall MQM evaluation revealed a clear and significant performance gap between human and AI-generated translations across all three dimensions. Human translations consistently achieved higher scores in accuracy, fluency, and cultural appropriateness, reflecting the depth of cultural and contextual knowledge that native-speaker translators bring to the task. Table 1 summarises the comparative scores.
| Dimension | Human Translation (M) | AI Translation (M) | Gap |
|---|---|---|---|
| Accuracy | 0.90 | 0.235 | −0.665 |
| Fluency | 0.90 | 0.814 | −0.086 |
| Cultural Appropriateness | 0.90 | 0.953 | −-0.053 |
| Overall MQM Score | 0.00 (baseline) | 7.303 (Very Poor) | — |
Accuracy measures whether the AI translation preserves the meaning of the source Chakma sentence. The AI achieved an accuracy score of 0.235, compared to the human benchmark of 0.90 — a gap of 0.665 points. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without any accuracy error. The remaining 1,293 sentences (76.5%) contained at least one error that distorted or entirely replaced the intended meaning. The corpus-level MQM penalty score of 7.303 reflects the cumulative weight of these failures: on average, each sentence carries 7.303 penalty points of translation error, placing the corpus firmly in the Very Poor quality band (≥7.00).
These results are consistent with findings from low-resource machine translation research (Ranathunga et al., 2023), which identifies data scarcity and limited digital representation as key factors that reduce AI accuracy on minority languages. Chakma's limited presence in LLM training data likely explains the AI's inability to correctly interpret the semantic content of Chakma sentences, particularly when surface-level phonological or script similarity led the model to select semantically unrelated English equivalents.
Fluency measures the grammatical naturalness and readability of the AI output in English. The AI achieved a fluency score of 0.814, compared to the human benchmark of 0.90 — a substantially smaller gap of 0.086 points. Of 1,691 sentences, 1,376 (81.4%) were produced without any grammatical error. This finding suggests that the AI performs considerably better at generating grammatically well-formed English sentences than at preserving the semantic content of the Chakma source.
However, this apparent fluency conceals a critical problem. A translation can be grammatically smooth in English while conveying entirely the wrong meaning — and this is precisely the pattern observed in the majority of erroneous AI outputs. As Al Sharou and Specia (2022) note, in low-resource settings, the severity of meaning distortion is more consequential than surface fluency levels. The AI's relatively high fluency score therefore masks, rather than mitigates, the depth of semantic failure in this corpus.
Cultural appropriateness measures whether culture-specific meanings — including kinship terms, pragmatic markers, proverbs, and community-specific expressions — are preserved in the AI output. The AI achieved a score of 0.953, compared to the human benchmark of 0.90 — the smallest performance gap of the three dimensions. Of 1,691 sentences, only 79 (4.7%) contained cultural or pragmatic errors.
While this figure appears relatively low, it is important to note that cultural errors are qualitatively more serious than their frequency suggests. The loss of a kinship distinction (such as "maternal uncle" becoming "uncle") or the failure to interpret a Chakma proverb represents an irreversible erasure of cultural meaning — precisely the type of loss that language documentation is intended to prevent. These findings align with Moneus and Sahari (2024), who find that AI systems struggle with culturally embedded expressions even when grammatical accuracy is maintained.
Across 1,691 sentences, the MQM annotation identified a total of 1,432 error occurrences. Table 2 summarises the distribution across the three MQM dimensions.
| MQM Dimension | Error Category | Frequency | Percentage |
|---|---|---|---|
| Accuracy (Lexical Error) | Mistranslation | 1038 | 72.5% |
| Fluency (Grammatical Error) | Grammar | 315 | 22.0% |
| Locale Convention (Cultural Error) | Cultural Appropriateness | 79 | 5.5% |
| Total | 1,432 | 100% | |
Lexical accuracy errors emerge as the dominant error type, accounting for 72.5% of all annotated errors. Grammatical errors constitute 22.0%, while cultural and pragmatic errors, though fewest in number, represent the most contextually significant failures at 5.5%. The severity distribution reveals that all Accuracy errors were annotated as Critical (10 pts), all Fluency and Locale errors as Major (5 pts), and no Minor errors were identified — indicating that all AI failures in this corpus are substantive rather than superficial.
Lexical errors represent failures in meaning-level translation — cases where the AI selected an English word or phrase that does not correspond to the semantic content of the Chakma source. These errors were the most frequent and, at Critical severity (10 pts each), the most damaging to the overall MQM score. Four sub-types of lexical error were identified.
Mistranslation — The AI output diverged significantly from the intended meaning of the source sentence, in some cases producing translations with no apparent semantic relationship to the input.
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 1 | 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄧𑄢𑄴 তুই হি গরর? |
What are you doing? | you are beautiful |
| 5 | 𑄖𑄬 𑄦𑄧𑄘𑄪 𑄡𑄢𑄴 তে হদু যার |
Where is he going? | that is enough |
| 3 | 𑄖𑄪𑄃𑄨 𑄃𑄨𑄘𑄪 𑄃𑄃𑄨 তুই ঈদু আই |
Come here. | you gave yes / you did give |
The most striking example is Sentence #1: the Chakma question "তুই হি গরর?" (What are you doing?) was rendered as "you are beautiful" — a complete semantic inversion. Similarly, in Sentence #5, the AI produced "that is enough" where the human translation reads "Where is he going?". These cases represent not merely word-level inaccuracy but a total failure to process sentence meaning.
Untranslated text — In several instances, the AI output largely mirrored the romanization of the Chakma source rather than producing a meaningful English translation, leaving the output entirely inaccessible to non-Chakma readers.
| # | Chakma Source | Human Translation | AI Output |
|---|---|---|---|
| 6 | 𑄖𑄟𑄨𑄟𑄴৷ 𑄊𑄪𑄟𑄴 𑄎𑄢𑄴 তামিম ঘুমজার |
Tamim is sleeping. | timid. gum jor |
| 7 | 𑄖𑄪𑄃𑄨 𑄦𑄟𑄧𑄠𑄚𑄴 𑄉𑄧𑄢𑄨 𑄘𑄬 তুই হাময়ান গড়ি দে |
Please do the work. | You make ready / prepare |
These untranslated outputs — where Chakma phonology is rendered in Roman script without any English semantic content — indicate that the AI system recognises the source as non-English text but lacks the linguistic resources to decode it. This pattern is most common with phonologically complex Chakma words that have no close approximation in the AI's training data.
Unintelligible output — A further subset of lexical errors produced outputs that were neither translations nor romanizations but rather syntactically incoherent fragments with no recoverable meaning.
Grammatical errors were annotated as Major-severity Fluency errors (5 pts each), reflecting the judgment that grammatical failures reduce translation quality significantly but the core message may sometimes remain partially recoverable. The following sub-types were identified in the corpus.
Tense errors — The most common grammatical error type. The AI consistently failed to map Chakma aspectual and temporal distinctions onto appropriate English tenses, collapsing future, past, and progressive forms into simple present.
| # | Chakma Source | Human Translation | AI Translation |
|---|---|---|---|
| 2 | 𑄖𑄪𑄃𑄨 𑄦𑄨 𑄉𑄧𑄢𑄨𑄝𑄬 তুই হি গরিবে? |
What will you do? | you will do |
| 4 | 𑄟𑄪𑄃𑄨 𑄞𑄖𑄴 𑄦𑄋𑄧𑄢𑄴 মুই ভাত হাঙর |
I am eating rice. | I eat rice |
| 33 | 𑄟𑄪𑄭 𑄝𑄎𑄢𑄧𑄖𑄴 𑄡𑄬𑄟𑄴 মুই বাজারত্ যেম্ |
I will go to the market. | I am going to the market |
The Chakma language encodes tense and aspect morphologically in ways that differ substantially from English. Errors of this type suggest that the AI is processing Chakma lexical items in isolation rather than parsing the full morphosyntactic structure of the source sentence.
Sentence category errors — The AI frequently shifted the grammatical category of a sentence, converting questions into statements, subjunctive wishes into declaratives, and imperatives into indicative clauses. This error type is particularly consequential in a documentation context because it misrepresents the communicative function of the source utterance.
Omission errors — In several instances, the AI dropped core sentence constituents — including subjects, main verbs, and interrogative particles — producing outputs that preserved partial meaning but lost semantic completeness.
Other grammatical errors — Additional error sub-types included singular/plural mismatches, addition of non-present content (such as politeness markers absent from the source), and word order errors resulting in awkward or unnatural English phrasing.
Cultural and pragmatic errors (Locale Convention, Major severity) were the least frequent but qualitatively most significant category of error. These arose when culturally embedded terms, relational distinctions, or community-specific expressions in Chakma were either flattened into generic English equivalents, misinterpreted, or rendered unintelligible.
| # | Chakma Source | Human Translation | AI Translation | Cultural issue |
|---|---|---|---|---|
| 22 | 𑄟𑄧𑄢𑄨𑄝𑄬 𑄚𑄦𑄨 মরিবে নাহি? |
Are you going to die? | will die not / will not die | Kinship/cultural term lost |
| 87 | 𑄖𑄬 𑄟𑄧𑄢𑄬 𑄷𑄶𑄶 𑄑𑄬𑄋 𑄃𑄪𑄘𑄮𑄢𑄴 𑄘𑄨𑅅 তে মরে 100টেঙা উদোর দ্যি |
He lent me 100 taka. | Error parsing | Kinship/cultural term lost |
| 338 | 𑄟𑄧 𑄟𑄟𑄪 𑄊𑄧𑄢𑄧𑄖𑄴 𑄃𑄉𑄬 𑅁 ম মামু ঘরত্ আগে। |
My maternal uncle is at home. | My uncle is at home. | Kinship/cultural term lost |
The most common cultural error type was the collapse of Chakma kinship terminology into undifferentiated English equivalents. The Chakma language maintains precise distinctions between maternal and paternal relatives, elder and younger siblings, and community-specific social roles. When the AI produces "uncle" for "maternal uncle", or "brother" for a term carrying specific age-relative social meaning, it erases the relational structure that is central to Chakma social and cultural life.
A further significant sub-type was the misinterpretation of Chakma proverbs. The AI consistently failed to interpret figurative or idiomatic expressions, either producing literal translations of the surface words (which convey no meaning in English) or substituting entirely unrelated English proverbs. This failure reflects the cultural knowledge gap that Hutson et al. (2024) identify as a fundamental limitation of AI systems in endangered language documentation: cultural depth and real-world community understanding cannot be learned from statistical patterns in training data alone.
To supplement the MQM annotation findings under RQ2, a Bi-LSTM-CRF model was applied to identify recurring grammatical and verb pattern structures in Chakma sentences that correlate with AI translation errors. The model achieved a grammatical pattern recognition accuracy of 87.4% on the Chakma corpus, indicating reliable identification of morphosyntactic structure in the absence of large pre-trained Chakma language resources.
The Bi-LSTM-CRF analysis revealed that the grammatical structures most consistently associated with AI translation errors were:
These findings from the Bi-LSTM-CRF corroborate the MQM Fluency error patterns (§4.2.3) and provide a structural explanation for why those errors occur. The model's 87.4% accuracy confirms that these grammatical patterns are identifiable and consistent — suggesting that a dedicated Chakma NLP pipeline could, in principle, flag high-risk structures before translation and alert human reviewers accordingly. This is proposed as a direction for future work in §6.2.
The following MQM summary tables and quality band analysis are presented in the interactive manuscript viewer (Score Analysis and Tables tabs) but are not part of the submitted paper. They are included here to support the reader's understanding of the quantitative findings and may be integrated in a revised submission.
| Severity | Errors | Penalty | Subtotal | % of Total Penalty |
|---|---|---|---|---|
| Critical | 1,038 | ×10 | 10,380 | 84.1% |
| Major | 394 | ×5 | 1,970 | 16.0% |
| Minor | 0 | ×1 | 0 | 0.0% |
| Total → MQM = 12,350 ÷ 1,691 | 7.303 | Very Poor (≥7.00) | ||
This section interprets the findings in relation to the two research questions, situates them within the broader literature on AI translation and low-resource language processing, and reflects on what they mean for Chakma language documentation and linguistic justice.
The overall performance of ChatGPT on Chakma–English translation was poor. The corpus-level MQM score of 7.303 places the AI output in the Very Poor quality band (≥7.00), indicating that systematic and significant translation failure — not occasional error — characterises the AI's engagement with Chakma. Of 1,691 sentences evaluated, only 398 (23.5%) were translated without error. The remaining 1,293 sentences (76.5%) required at least one error annotation, and in most cases the error was Critical in severity — meaning the output actively misrepresented the source.
These results are consistent with the broader literature on AI translation in low-resource language settings. Ranathunga et al. (2023) identify data scarcity, the absence of parallel corpora, and insufficient language-specific resources as the primary factors constraining machine translation performance for under-resourced languages. Chakma's minimal presence in the training data of large language models means the system is effectively operating without the linguistic foundation necessary for reliable translation. The findings therefore do not simply reflect a gap in model capability; they reflect a structural inequality in how languages are represented in digital infrastructure and AI training pipelines.
Lexical accuracy errors were the most frequent and most penalised error type, accounting for 72.5% of all MQM annotations at Critical severity (10 pts each). The range of lexical error sub-types — mistranslation, untranslated text, and unintelligible output — reflects the probabilistic nature of large language models. Rather than understanding meaning in the way human translators do, AI systems generate output based on patterns learned from large amounts of textual data (Fu & Liu, 2024). When the training data contains little or no Chakma, the model cannot distinguish between visually or phonologically similar but semantically unrelated words, resulting in outputs that are grammatically plausible in English but semantically disconnected from the source.
The untranslated outputs — where the AI produced romanised Chakma rather than English — are particularly revealing. They indicate that the model recognises the script as non-English but lacks the decoding capacity to generate a meaningful translation. This pattern aligns with Okafor's (2025) findings on Igbo, where contextual deficiencies in AI training led to lexical ambiguity and incorrect word substitutions. For Chakma, the problem is more fundamental: the lexical base itself is largely absent from the model's knowledge.
Grammatical errors (315 instances, 22.0% of total errors) were consistently annotated at Major severity (5 pts), reflecting the judgment that they reduce translation quality significantly while sometimes leaving core meaning partially intact. The most common grammatical error type — tense and aspect misrepresentation — suggests that the AI processes Chakma lexical items in isolation rather than parsing morphosyntactic structure. Chakma encodes tense, aspect, and mood through morphological affixes that differ substantially from English inflectional patterns; without explicit modelling of these structures, the AI defaults to simple present tense regardless of the source's temporal reference.
The sentence category errors — in which questions became statements, wishes became declaratives, and imperatives became indicatives — are particularly significant from a documentation perspective. A corpus that systematically converts Chakma questions into statements misrepresents the pragmatic structure of the language, distorting any subsequent linguistic analysis that relies on the translated data. Fu and Liu (2024) note that ChatGPT occasionally fails to identify implicit grammatical relationships and does not always restructure source-language syntax appropriately; the present study finds this tendency to be systematic rather than occasional in the context of Chakma.
Cultural and pragmatic errors (79 instances, 5.5% of total errors) were the least frequent but qualitatively most significant category. The collapse of Chakma kinship terminology — where distinctions such as maternal versus paternal uncle, or elder versus younger sibling, are flattened into generic English equivalents — represents the erasure of relational and social knowledge that is embedded in the language itself. This is not a translation error in the narrow linguistic sense; it is the deletion of cultural information that has no direct English equivalent and cannot be recovered once lost.
The AI's failure to interpret Chakma proverbs further illustrates the limits of pattern-based language modelling in cultural contexts. Proverbs are community-specific communicative forms whose meaning depends on shared cultural knowledge that cannot be inferred from lexical co-occurrence statistics. Rousan et al. (2025) report similar findings in Arabic–English literary translation, where AI systems misinterpret culturally embedded expressions despite producing fluent English output. The present study extends this finding to an endangered indigenous language context, where the cultural stakes of misinterpretation are considerably higher.
The Bi-LSTM-CRF analysis — which achieved 87.4% grammatical pattern recognition accuracy on the Chakma corpus — provides a structural explanation for the Fluency error patterns identified through MQM annotation. Where MQM records that an AI tense error occurred, the Bi-LSTM-CRF indicates which Chakma morphological structure the AI failed to parse. Together, these two analytic layers produce a more complete picture of AI translation failure than either could provide alone.
The structures most consistently associated with AI errors — progressive aspect suffixes, future tense markers, interrogative particles, negation morphology, and evidential markers — share a common property: they are morphologically encoded in Chakma in ways that have no direct surface-level English counterpart. An attention-based language model trained on high-resource languages will not have learned to associate these Chakma morphemes with their English functional equivalents, because the relevant training signal is absent. The Bi-LSTM-CRF findings confirm that these structures are systematic and identifiable — which means they could, in principle, be used to build error-prediction tools for human post-editors reviewing AI-generated Chakma translations.
This finding is exploratory. A full evaluation of the Bi-LSTM-CRF as a diagnostic component — including comparison across architectural variants, cross-validation on held-out data, and integration with the MQM annotation pipeline — is a direction for future work (see §6.2). The current study limits its claim to the observation that grammatical pattern recognition at 87.4% accuracy supports the interpretation of Fluency errors as structurally grounded failures in morphological parsing, not random noise.
The patterns identified in the MQM annotation point to several structural features of Chakma that consistently triggered AI errors. Sentences with morphologically complex verb forms — particularly those encoding progressive aspect, conditional mood, and evidentiality — showed the highest rates of mistranslation. These structures require the model to track morphological dependencies across the sentence rather than relying on lexical co-occurrence, a capacity that is severely limited when training data for the source language is absent.
Taken together, the findings point to a single underlying cause that cuts across all three error categories: Chakma's near-total absence from the training data of current large language models. This is not a problem that can be resolved through better prompting or model fine-tuning alone. It is a structural condition rooted in the historical and ongoing marginalization of indigenous languages from digital infrastructure, standardized orthography, and natural language processing research.
Ranathunga et al. (2023) identify data scarcity, the lack of parallel corpora, and insufficient language resources as the major challenges affecting machine translation performance in low-resource languages. The present study confirms all three as operative in the Chakma case. The lexical errors reflect limited exposure to Chakma vocabulary; the grammatical errors reflect the absence of Chakma-specific syntactic modelling; and the cultural errors reflect the impossibility of learning community-specific meaning from text corpora alone.
From a linguistic justice perspective, these disparities are not merely technical. The reduced performance of AI translation tools on Chakma reflects the wider marginalization of Indigenous languages within digital infrastructures and training datasets that extensively privilege high-resource languages (Lepp & Sarin, 2024). As Dovchin (2020) argues, linguistic racism operates by subduing minority voices while prioritizing dominant languages — and AI systems trained predominantly on English, Bangla, and other high-resource languages reproduce this hierarchy algorithmically. Improving AI support for Chakma is therefore not only a matter of technological advancement but also a step toward addressing persistent inequalities in linguistic representation within the digital age.
This study examined the effectiveness of AI-based translation systems in translating Chakma texts into English by comparing AI-generated translations with native-speaker-validated human translations across a corpus of 1,691 sentence pairs. Using the Multidimensional Quality Metrics (MQM) framework with three severity levels — Critical (10 pts), Major (5 pts), and Minor (1 pt) — the analysis evaluated translation quality across three dimensions: Accuracy (Lexical Error), Fluency (Grammatical Error), and Locale Convention (Cultural Error). All annotations were validated by native Chakma speakers to ensure cultural and linguistic integrity.
The findings reveal that AI-generated translations performed poorly across all three dimensions. The corpus-level MQM score of 7.303 places the output in the Very Poor quality band (≥7.00), with only 398 of 1,691 sentences (23.5%) translated without error. Lexical accuracy errors were the dominant failure type, accounting for 72.5% of all annotated errors at Critical severity, reflecting systematic semantic failure rather than surface-level inaccuracy. Grammatical errors, though less frequent, further reduced translation quality by misrepresenting the tense, aspect, sentence type, and structural properties of Chakma utterances. Cultural and pragmatic errors, while fewest in number, were qualitatively the most consequential, involving the irreversible erasure of culturally embedded knowledge — kinship distinctions, proverbs, and community-specific expressions — that cannot be recovered through post-editing alone.
Beyond the technical findings, this study argues that the reduced performance of AI translation on Chakma reflects broader patterns of digital inequality and linguistic marginalization. Languages with limited digital representation receive substantially weaker technological support, and AI systems trained on high-resource languages reproduce existing linguistic hierarchies. From a linguistic justice perspective, improving AI support for Chakma is not merely a technical challenge but an ethical imperative — a step toward equitable representation of indigenous knowledge systems in the digital age (Dovchin, 2020; Lepp & Sarin, 2024).
The study has several limitations. First, the evaluator's awareness of the translation sources may have introduced a degree of bias, although the use of structured MQM criteria and native-speaker review helped mitigate this risk. Second, the corpus, while covering 1,691 sentence pairs — significantly larger than earlier Chakma NLP datasets — may not fully capture the breadth of lexical, grammatical, and cultural variation present in Chakma, particularly across regional dialects and specialized domains such as law, medicine, and oral tradition. Third, this study evaluates a single AI system (ChatGPT) at one point in time; the rapidly evolving landscape of large language models means that findings may not generalise to future model generations. Finally, the Bi-LSTM-CRF component of the study was designed to identify grammatical and verb pattern structures contributing to recurring translation errors, but its findings are limited by the size and diversity of the annotated training data.
The study carries implications for both AI development and language documentation practice. For AI developers, the findings highlight the urgent need for Chakma-specific training resources: larger parallel corpora, culturally informed annotation guidelines, and community-driven validation processes that embed native speaker expertise at every stage of model development. Without these resources, AI translation systems will continue to perform poorly on Chakma and other endangered languages, reinforcing rather than challenging existing linguistic hierarchies.
For language documentation practitioners, the findings suggest that AI translation tools, at their current level of performance, are best understood as assistive rather than autonomous resources. When used alongside native-speaker expertise, AI systems can support the initial translation of Chakma texts and contribute to the development of bilingual language resources — but every AI output must be reviewed before it is trusted. The MQM score of 7.303 quantifies this burden concretely: on average, each sentence requires correction of approximately 7.303 penalty points of translation error, and in 1038 of 1432 cases (72.5%), that correction requires addressing a Critical semantic failure.
A specific direction for future work concerns the Bi-LSTM-CRF component introduced in this study as a supporting diagnostic tool. The model achieved 87.4% grammatical pattern recognition accuracy on Chakma, identifying the morphosyntactic structures — progressive aspect markers, future suffixes, interrogative particles, negation morphology — most likely to trigger AI translation errors. A full evaluation of this model, including cross-validation, architectural comparison, and integration with the MQM annotation pipeline, would constitute a meaningful methodological contribution to endangered language NLP. If the model can reliably predict high-risk grammatical structures before translation, it could serve as the basis for a Chakma-specific quality estimation tool that reduces the burden on human post-editors.
Future research should also examine larger and more diverse Chakma corpora, compare multiple AI translation systems and generations of models, and explore community-centred approaches to AI development that centre indigenous linguistic and cultural knowledge. Particular attention should be given to building the digital infrastructure — standardized orthographies, annotated corpora, lexical databases — that would enable meaningful AI support for Chakma and other endangered languages of Bangladesh and South Asia.
Notwithstanding its limitations, this study makes a methodological contribution by demonstrating that the MQM framework — typically applied in professional translation contexts — is a viable and informative evaluation tool for endangered language AI assessment. The penalty-based scoring system, combined with dimension-level and severity-level breakdown, provides a richer and more actionable picture of AI translation failure than automated metrics such as BLEU or TER, which remain largely uninformative for low-resource language pairs. Future studies are encouraged to adopt and extend this framework as part of a broader effort to develop evaluation standards for AI translation in indigenous language documentation.
Al Sharou, K., & Specia, L. (2022). Towards a better understanding of noise in natural language processing. Proceedings of the 13th Language Resources and Evaluation Conference.
Abdelhalim, S. M., Alsahil, A. A., & Alsuhaibani, Z. A. (2025). Artificial intelligence tools and literary translation: a comparative investigation of ChatGPT and Google Translate from novice and advanced EFL student translators' perspectives. Cogent Arts & Humanities, 12(1), 2508031.
Afaq, M., Mehmood, T., & Ayaz, M. O. (2025). Can artificial intelligence challenge universal grammar? A theory-driven empirical investigation. Journal of Applied Linguistics and TESOL (JALT), 8(4), 1248–1254.
Afreen, N. (2020). Language usage in different domains by the Chakmas of Bangladesh. International Journal of Linguistics, Literature and Translation, 3(6), 135–151.
Ajani, Y. A., Oladokun, B. D., Olarongbe, S. A., Amaechi, M. N., Rabiu, N., & Bashorun, M. T. (2024). Revitalizing indigenous knowledge systems via digital media technologies for sustainability of indigenous languages. Preservation, Digital Technology & Culture, 53(1), 35–44.
Anik, M., Rahman, A., Wasi, A., & Ahsan, M. (2025, May). Preserving cultural identity with context-aware translation through multi-agent AI systems. In Proceedings of the 1st Workshop on Language Models for Underserved Communities (LM4UC 2025) (pp. 51–60).
Bal, E. (2010). Being Mog: Memories, nostalgia, and identity of the Mog community in Bangladesh. Modern Asian Studies, 44(6), 1261–1295.
Bassnett, S., & Trivedi, H. (1999). Introduction: Of colonies, cannibals and vernaculars. In S. Bassnett & H. Trivedi (Eds.), Post-colonial translation: Theory and practice (pp. 1–18). Routledge.
Bishop, M. (2022). Elders' conversations: Perspectives on leveraging digital technology in language revival. The Open/Technology in Education, Society, and Scholarship Association Journal, 2(2), 1–13.
Chakma, A., Khisa, A., Khisa, S., Noor, J., & Sultana, S. (2026). Re-educating educated ones: A case study on Chakma language revitalization in Chittagong Hill Tracts. arXiv preprint arXiv:2601.12290.
Chakma, J. (2010). Origin and evolution of Chakma language and script. Kriti Rakshana, National Mission for Manuscripts.
Chakma, J., & Sultana, A. (2023). Language rights and indigenous peoples of the Chittagong Hill Tracts. International Journal of Language and Culture.
Çetin, Ö., & Duran, A. (2024). A comparative analysis of the performances of ChatGPT, DeepL, Google Translate and a human translator in community-based settings. Amasya Üniversitesi Sosyal Bilimler Dergisi, 9(15), 120–173.
Chiran, R. (2025). Language endangerment in Bangladesh: An updated assessment. South Asian Languages Review.
Dovchin, S. (2020). Introduction to special issue: Linguistic racism. International Journal of Bilingual Education and Bilingualism, 23(7), 773–777.
Drude, S., & Intangible Cultural Heritage Unit's Ad Hoc Expert Group. (2003). Language vitality and endangerment. UNESCO.
Ducharme, Q. M., Amatulli, G., Williams, W. A. L., George, S. H., Pierre, S. M., & Pierre, S. L. R. (2025). Revitalizing indigenous languages, fostering self-governance, overcoming the Indian Act: A case study of Lil'wat Nation. Canadian Public Administration, 68(3), 470–486.
Folaron, D. (2015). Translation and minority, lesser-used and lesser-translated languages and cultures. The Journal of Specialised Translation, 24, 16–27. https://doi.org/10.26034/cm.jostrans.2015.320
Fu, Y., & Liu, Y. (2024). Evaluating ChatGPT's translation quality in scientific texts. Language & Technology Review.
Grenoble, L. A., & Whaley, L. J. (2005). Saving languages: An introduction to language revitalization. Cambridge University Press.
Gwerevende, S., & Mthombeni, Z. M. (2023). Safeguarding intangible cultural heritage: exploring the synergies in the transmission of indigenous languages, dance and music practices in Southern Africa. International Journal of Heritage Studies, 29(5), 398–412.
Holmes, J. (2021). An introduction to sociolinguistics (4th ed.). Routledge.
Hutson, J., Ellsworth, P., & Ellsworth, M. (2024). Preserving linguistic diversity in the digital age: a scalable model for cultural heritage continuity. Journal of Contemporary Language Research, 3(1).
Jerome, C., et al. (2022). Language, identity and indigenous communities. Journal of Language and Cultural Studies.
Jiang, Z., Lv, Q., Zhang, Z., & Lei, L. (2023). Distinguishing translations by human, NMT, and ChatGPT: A linguistic and statistical approach. arXiv.
Kandler, A., & Unger, R. (2023). Modeling language shift. In Diffusive spreading in nature, technology and society (pp. 365–387). Springer International Publishing.
Koc, V. (2025). Generative AI and large language models in language preservation: Opportunities and challenges. arXiv (Cornell University).
Lepp, A., & Sarin, L. (2024). Linguistic justice and digital inequality. Language Policy & Technology Review.
Li, M., Croucher, S. M., & Shen, L. (2024). Language endangerment and the linguistic vitality of Miao in China: cultural shifts and revitalisation strategies. Journal of Multilingual and Multicultural Development, 1–16.
Mahi, M. H., Khan, A. R., Anik, M. H., Noori, S. R. H., Mahmud, A., & Mojumdar, M. U. (2025). MELD: a multilingual ethnic dataset of Chakma, Garo, and Marma in Bengali script with English and standard Bengali translation. Data in Brief, 61, 111745.
Mohamed, M., et al. (2024). Translation quality in the age of AI. Language & Technology.
Mohsin, A. (2023). Indigenous languages of Bangladesh. University Press Limited.
Moneus, A. M., & Sahari, Y. (2024). Artificial intelligence and human translation: A contrastive study based on legal texts. Heliyon, 10(6).
MQM. (2015). Multidimensional quality metrics definition. Retrieved from https://web.archive.org/web/20210113220425/http://www.qt21.eu/mqm-definition/definition-2015-05-27.html
O'Hagan, M. (2016). Massively open translation: Unpacking the relationship between technology and translation in the 21st century. International Journal of Communication, 10, 18.
Okafor, A. Y. (2025). Examining AI translation errors in Igbo: Lexical ambiguity, misinterpretation, and incorrect word substitutions due to contextual deficiencies. Indonesian Journal of Learning Studies, 5(1), 46–55.
Oladipupo, F., Soronnadi, A., Adebara, I., & Adekanmbi, O. (2025, August). How effective are AI models in translating English scientific texts to Nigerian Pidgin: A low-resource language? In I Can't Believe It's Not Better: Challenges in Applied Deep Learning.
Rafat Al Rousan, Raghad Jaradat, & Mona Malkawi. (2025). ChatGPT translation vs. human translation: an examination of a literary text. Cogent Social Sciences, 11(1), 2472916. https://doi.org/10.1080/23311886.2025.2472916
Ranathunga, S., Lee, E. S. A., Prifti Skenduli, M., Shekhar, R., Alam, M., & Kaur, R. (2023). Neural machine translation for low-resource languages: A survey. ACM Computing Surveys, 55(11), 1–37.
Saikia, M., & Ullman, J. (2023). Endangered language assessment framework. Language Documentation Journal.
Sevinç, Y. (2022). Language endangerment and revitalization. Annual Review of Linguistics.
Smith, B. K., Ehala, M., & Giles, H. (2017). Vitality theory. In J. Nussbaum (Ed.), Oxford research encyclopedia of communication. Oxford University Press.
Spivak, G. C. (1993). Outside in the teaching machine. Routledge.
Tymoczko, M. (1999). Translation in a postcolonial context: Early Irish literature in English translation. St. Jerome Publishing.
Tsunoda, T. (2006). Language endangerment and language revitalisation: An introduction. Mouton de Gruyter.
UNESCO. (2003). Language vitality and endangerment. Ad Hoc Expert Group on Endangered Languages.
Walsh, J. (2006). Language and socio-economic development: Towards a theoretical framework. Language Problems and Language Planning, 30(2), 127–148.
Wei, L., Hua, Z., & Simpson, J. (Eds.). (2023). The Routledge handbook of applied linguistics: Volume two. Taylor & Francis.
Yan, J., Yan, P., Chen, Y., Li, J., Zhu, X., & Zhang, Y. (2024). GPT-4 vs. human translators: A comprehensive evaluation of translation quality across languages, domains, and expertise levels. arXiv preprint arXiv:2407.03658.
Every change made in Edit Mode is logged here — tab, section, passage, before & after. Click Revert on any entry to undo that specific change.