Skip to content
AI.info

Research

Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations

Overview Research area: Natural Language Processing / Machine Translation, with human-computer interaction (HCI) and Translation Studies. Technical level: Intermediate. The study design and conclusion

arXiv
2510.09994
Published
2025-10-11
Authors
Yimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao, Marianna J. Martindale, Charlotte Vaughn, Ge Gao, Marine Carpuat

AI summary

Overview

Research area: Natural Language Processing / Machine Translation, with human-computer interaction (HCI) and Translation Studies.

Technical level: Intermediate. The study design and conclusions are accessible to general readers, but the results rely on mixed-effects models and multinomial logistic regression, which assume some statistical familiarity.

Scope: A museum-based human study with 452 valid participants examining how fluency and adequacy errors in Spanish-to-English machine translation affect bilingual and non-bilingual users' perception, decision-making, confidence, and willingness to reuse the tool.

What This Paper Is About

Machine translation (MT) is now used casually by huge numbers of people, most of whom are not translation professionals — the paper cites that by 2021 the Google Translate app alone had over a billion installations, with an estimated 99.97% of MT users being non-professionals. Yet benchmark scores say nothing about whether ordinary users can detect errors, judge output quality, or calibrate their reliance on the tool. This paper runs a controlled experiment in a public museum to see how different error types change what users believe, what they decide, and whether they would use the system again.

Key Contributions

  1. A museum-based human study of MT reliance. The authors collected 452 valid responses at the Language Science Station (LSS) in the Planet Word museum in Washington, D.C., recruiting a demographically diverse sample rather than a lab or crowdsourced population.
  2. A controlled 2 × 3 experimental design separating error type from error impact. The study crosses MT correctness (correct vs. incorrect, within-subject) with three error types (Fluency Error without Impact, Adequacy Error without Impact, Adequacy Error with Impact, between-subject), using 16 validated stimuli.
  3. Evidence that source-language proficiency shapes error detection and decision accuracy. The paper shows that participants' ability to perceive fluency errors — which are in principle visible from the English target text alone — depends on their Spanish proficiency.
  4. A reframing of "over-reliance" as a strategy gap. The findings argue that non-bilingual users rely on imperfect MT not because they believe it is correct, but because they lack strategies for evaluating outputs, motivating MT literacy interventions and explanation techniques.

Main Findings

  • Fluency errors lower perceived believability. Fluency Error without Impact were perceived significantly less believable than Correct translations (Coefficient = -2.97, p < .001).
  • Higher Spanish proficiency raises perceived quality overall. Participants with higher Spanish proficiency reported higher ratings of Perception of Translation Quality (Coefficient = 1.08, p < .001).
  • Error detection depends on proficiency. High Spanish Proficiency participants detected MT errors regardless of error type. Some Spanish Proficiency participants perceived Fluency Error without Impact but not Adequacy Error with Impact or Adequacy Error without Impact. No Spanish Proficiency participants were not able to perceive any MT errors — including fluency errors that monolingual English speakers detected in the authors' validation studies.
  • Misleading adequacy errors hurt decision accuracy for less proficient users. Accuracy was significantly lower under Adequacy Error with Impact compared to Correct (Coefficient = -2.18, p < .001), with a significant interaction with Spanish proficiency. Participants with No or Some Spanish Proficiency were less accurate; those with High Spanish Proficiency showed no difference across error types.
  • Confidence tracked proficiency, not error type. Higher Spanish proficiency predicted significantly higher Decision-Making Confidence (Coefficient = 0.225, z = 2.47, p < 0.05). MT error types showed no significant main effects and no significant interactions with proficiency.
  • Willingness to reuse MT varied by error type. There was a significant main effect of MT Error type (F(2, 374) = 4.798, p < .01). Tukey's HSD showed lower willingness to reuse after Fluency Error without Impact (Mean = 2.81, S.E. = 0.20) than after Adequacy Error without Impact (Mean = 3.00, S.E. = 0.21; p < .05), and lower willingness after Adequacy Error with Impact (Mean = 2.90, S.E. = 0.20) than after Adequacy Error without Impact (Mean = 3.00, S.E. = 0.21; p < .05). There was no significant difference between Adequacy Error with Impact and Fluency Error without Impact, and no significant main effect of Spanish Proficiency or interaction with error type.
  • Error conditions prompted different evaluation strategies. Compared with the No Strategy reference category, participants in Fluency Error without Impact were less likely to report the Comparison Strategy (Coefficient = -2.30, SE = 0.26, p < .001) and the Proficiency Strategy (Coefficient = -0.95, SE = 0.34, p < .01). Adequacy Error without Impact participants were also less likely to report Comparison (Coefficient = -1.95, SE = 0.23, p < .001) and Proficiency (Coefficient = -1.39, SE = 0.41, p = .001). Adequacy Error with Impact participants were more likely to report the Comparison Strategy (Coefficient = 1.78, SE = 0.31, p < .001) but less likely to report the Proficiency Strategy (Coefficient = -1.80, SE = 0.33, p < .001).
  • Spanish proficiency strongly predicted strategy choice. Higher Spanish proficiency was associated with lower likelihood of reporting the Comparison Strategy (Coefficient = -5.34, SE = 0.44, p < .001) and higher likelihood of reporting the Proficiency Strategy (Coefficient = 9.97, SE = 0.25, p < .001).
  • Experiencing errors shifted future intent. Even in a low-stakes, fictional task, exposure to some error types lowered participants' stated willingness to reuse the system.
  • Debrief sessions showed public interest. Post-task conversations frequently covered participants' own MT experience, how MT works and why it errs, the strategies they used, and explanations of the study's translation errors. The authors note they did not directly measure the educational benefits of participation.

Methodology in Plain English

Setting and participants. The study ran at the Language Science Station, a research and public engagement lab inside the Planet Word museum in Washington, D.C. Anyone with sufficient English to follow instructions could take part. 517 people participated; 65 were excluded for not completing the task or for doing it with another person, leaving 452 valid responses. Participants were not compensated. The protocol was approved by the University of Maryland's Institutional Review Board.

Sample composition. The sample included 269 females, 159 males, 15 people identifying as non-binary or gender-fluid, and 9 who preferred not to disclose gender, with an average age of 35.12 years (S.E. = 1.65). For Spanish proficiency, 63 participants (13.94%) reported no proficiency, 353 (78.10%) some proficiency, and 36 (7.96%) high proficiency. For English, 406 (89.82%) reported high proficiency and 46 (10.18%) some proficiency. MT use was reported as never by 159 (35.18%), rarely by 162 (35.84%), sometimes by 74 (16.37%), often by 40 (8.85%), and almost every day by 17 (3.76%).

Task. Participants helped a fictional Spanish-speaking character retrace her steps in the museum and find a lost item. Each of four trials (T1–T4) presented a Spanish message and its English translation. Participants then (1) rated translation quality, (2) picked one of two images matching the Spanish message, and (3) rated confidence in that choice. After each trial they were told whether they were right.

Design. A 2 × 3 mixed design. MT Correctness was within-subject: correct translations in T1 and T4, incorrect in T2 and T3. MT Error Type was between-subject, with participants randomly assigned to Fluency Error without Impact, Adequacy Error without Impact, or Adequacy Error with Impact. Correct translations first let participants calibrate and build initial trust; errors mid-task revealed how they responded to disruption. Stimulus order was exhaustively counterbalanced and participants were randomly assigned to a pre-generated sequence.

Stimuli. Since naturally occurring MT errors could not meet the task's constraints (concise, museum-relevant, readable by young audiences and non-native English speakers), the authors constructed 16 stimuli across four categories. They created Spanish sentences, then produced English translations by prompting models including LLaMA-3 8B, ChatGPT, and GPT-4 with diverse decoding strategies, and by introducing controlled changes to the Spanish input (misspellings, deletions, round-trip MT, paraphrasing, style transfer). Three independent crowdsourced validation checks filtered ambiguous messages and unintuitive text–image pairings.

Measures. Translation Quality Perception (correct = 1, unclear = 0.5, problematic = 0), Decision Accuracy (1 or 0), Decision Confidence (5-point Likert), Willingness to Reuse MT (5-point Likert), and Evaluation Strategy (no strategy/intuitive, comparative analysis of Spanish and English, or Spanish proficiency-based judgment).

Analysis. Perception, accuracy, and confidence were modeled with mixed-effects models (CLMM and GLMM with logit links) using MT error type and Spanish proficiency as fixed effects, Participant ID and Stimulus ID as random effects, and age, gender, English proficiency, and prior MT experience as covariates, with Bonferroni corrections. Willingness to reuse was analyzed per participant with ANOVA and Tukey's HSD. Strategy choice was modeled with multinomial logistic regression, using No Strategy as the reference category.

Why This Matters

Impact on research. The paper argues that intrinsic MT quality evaluation — benchmarks, error annotation, post-editing — tells us little about how the public actually perceives and relies on MT. It positions MT literacy as a research target in its own right, calling for work at the intersection of NLP, HCI, and Translation Studies. It also complicates prior findings by Martindale and Carpuat (2018) on fluency versus adequacy errors by showing that the effect of error type is mediated by source-language proficiency.

Real-world applications.

  • MT literacy interventions. The study task itself is proposed as a template for simulation-based training, reaching users outside classrooms and potentially targeting professionals who use MT with little or no formal training.
  • Error-aware translation interfaces. Systems could surface where source and translation diverge, draw on quality estimation or error detection techniques, or use question-answering to expose inconsistencies a user would otherwise miss.
  • Cognitive forcing functions. Prompting users to deliberately evaluate outputs before acting has shown promise in other AI decision-making tasks and could be adapted for translation.
  • Stimulus generation tooling. Semi-automatically scoring or generating appropriate MT input/output pairs for a given context, decision, and error type would make literacy studies and training easier to scale.

Industry relevance. Translation and chatbot developers get a concrete failure-mode picture: users who cannot detect errors still form trust judgments, and those judgments shift after even brief exposure to certain error types. That means interface design and error communication are not cosmetic — they affect whether users keep the tool at all.

Future Directions

  1. Design MT and NLP techniques that help users assess and recover from errors, including highlighting source–target differences and automatically estimating translation quality or detecting potential errors — while determining how to present such signals effectively.
  2. Test question-answering approaches that surface source–translation inconsistencies, since prior work shows they help users judge whether a translation is safe to share, but it remains unclear whether they improve accuracy on content-specific questions like the ones in this study.
  3. Study trust formation in casual settings and how it carries over to high-stakes use, an explicit open question the authors raise.
  4. Develop MT literacy tools beyond the museum, including generic translation apps that explicitly foster MT literacy, and semi-automatic generation of training stimuli. The authors also note that museum visitors who opt into studies may be more motivated than typical users, and that their low-stakes, short, single comprehension-style task may not generalize to real-world scenarios.

Target Audience

MT and NLP researchers interested in evaluation beyond benchmarks; HCI and trust-in-AI researchers who study reliance and error handling; translation studies scholars working on MT literacy; and product or UX practitioners designing translation features who need to understand how lay users actually respond to imperfect output.

Authors’ abstract

As Machine Translation (MT) becomes increasingly commonplace, understanding how the general public perceives and relies on imperfect MT is crucial for contextualizing MT research in real-world applications. We present a human study conducted in a public museum (n=452), investigating how fluency and adequacy errors impact bilingual and non-bilingual users' reliance on MT during casual use. Our findings reveal that non-bilingual users often over-rely on MT due to a lack of evaluation strategies and alternatives, while experiencing the impact of errors can prompt users to reassess future reliance. This highlights the need for MT evaluation and NLP explanation techniques to promote not only MT quality, but also MT literacy among its users.

Read the original paper