Every interpreter has now had a version of this conversation.
A client asks whether they still need to pay for interpreters. A hospital pilots an AI tool. A colleague forwards an article predicting the profession will be gone in five years. Someone in a training group insists the technology is useless.
Both reflexes — panic and dismissal — are wrong, and both are easy to see through.
The useful position is the accurate one: knowing specifically where these tools perform, specifically where they fail, and why the difference is not a matter of the technology getting better.
Start With What the Evidence Actually Shows
Interpreters weaken their own case when they overstate it, so begin with the finding that is least comfortable.
In April 2026, Academic Emergency Medicine published a blinded cross-sectional non-inferiority study comparing Google Translate and ChatGPT-4o against professional medical interpreters on real emergency department discharge instructions in Spanish, Brazilian Portuguese, and Simplified Chinese.
Both machine methods were non-inferior to professional interpreters for adequacy, meaning, and severity in Spanish and Portuguese, and across all four measured domains in Chinese. The rate of clinically significant errors did not differ significantly by method. And the evaluators — professional medical interpreters, blinded to which was which — frequently mistook the machine output for professional work.
That is a real result from a real study, and pretending otherwise is not a strategy.
But read what it measured. It evaluated written translation of discharge instructions in three of the most heavily resourced languages on earth, scored after the fact on a text with no conversation around it.
That is close to the best case for a machine. It is also not interpreting.
The Same Tools, Different Languages
The picture changes sharply outside the major languages.
A University of Washington literature review, described in early 2026, examined machine translation accuracy for discharge instructions. Spanish, Chinese, and Portuguese held up reasonably well. In less commonly spoken languages, error rates climbed substantially — and in Amharic, Tigrinya, and Somali, both
ChatGPT and Google Translate produced nonsensical phrasing and outright invented words.
Hold that next to the market reality. A large language services network may cover two hundred languages.
Machine performance is strong in perhaps a dozen and degrades from there, worst exactly where interpreter shortages are most severe and patients are most vulnerable.
An organization replacing human interpreters with AI does not reduce service equally. It reduces it most for the people already least well served.
Fluent Wrongness Is the Core Problem
This deserves its own attention because it is the mechanism behind most AI language risk.
When a human interpreter is unsure, it shows. They hesitate, ask for repetition, request clarification, or say outright that a term is unfamiliar. Uncertainty produces a visible signal, and that signal is a safety feature.
Machine output carries no such signal. A fabricated Somali word arrives in the same confident cadence as a correct one. A negation dropped from a sentence — "do not take this with alcohol" becoming "take this with alcohol" — sounds exactly as authoritative as the original.
The blinded evaluators in that emergency medicine study could not reliably tell machine output from professional output. If trained interpreters cannot tell, a provider certainly cannot, and a patient has no chance at all.
Errors that announce themselves are manageable. Errors delivered fluently are not.
Interpreting Is Not Translation Performed Faster
The most important limitation is not accuracy. It is scope.
NCIHC took this up directly. In July 2024, the organization published guidance authored by Cynthia Roat and ratified by its Board, written for healthcare organizations being approached by vendors offering to replace interpreting and translation entirely with AI. NCIHC acknowledged AI's potential to improve the quality and efficiency of human language services, and stated significant concerns about moving to exclusively AI-generated services, while noting that the technology is advancing and current limitations may change.
What makes the guidance valuable is its form. Rather than arguing, it lists what qualified human interpreters do and asks the vendor how the system accomplishes each one. Among them:
Conducting a pre-session so that all parties address each other, understand that everything will be interpreted, and pause appropriately for accurate consecutive interpreting. How does the AI do this?
Intervening when meaning is at risk, and making both speaker and listener aware of the intervention. How does the AI do this?
Monitoring listeners for signs of understanding or its absence, through audible cues and body language.
How does the AI do this?
And the question that ends most vendor conversations: when the system produces an inaccurate rendition that harms a patient, who is liable? Human interpreters and the companies that employ them are accountable for the quality of their work. Accountability for machine output is, at best, unsettled.
These questions are not rhetorical. They describe the job. A system that converts speech but performs none of them is doing one piece of the work and leaving the rest undone.
Where the Line Should Sit
Drawing on both the evidence and the nature of the work, a defensible division looks roughly like this.
Reasonable uses. Wayfinding and scheduling. Routine, low-stakes exchanges. Post-hoc translation of standardized written materials in well-resourced languages, reviewed by a qualified human. Terminology support for a working interpreter. Draft translations that a professional edits. Bridging the gap in the minutes before a qualified interpreter joins.
Not reasonable. These share a common feature: the cost of a fluent, undetected error is irreversible.
Informed consent. A legal act requiring confirmed understanding, where the provider must assess comprehension and the patient must be able to ask questions.
Legal testimony. Where a hedge, a hesitation, or an ambiguity may be the evidence, and where a machine's tendency to produce clean, coherent output actively destroys the record.
Behavioral health. Where disorganized speech must remain disorganized because the clinician is assessing it, and a system optimized to produce fluent output will smooth away the finding.
Emotionally weighted encounters. End-of-life discussions, disclosures of abuse, pediatric emergencies. Tone, register, and pacing carry meaning here, and someone in the room needs to notice when a patient stops absorbing information.
Non-standard speech. Regional dialects, code-switching, low-incidence languages, speech affected by illness, injury, or distress. These are where machine performance collapses fastest and where human interpreters do some of their most skilled work.
Genuine ambiguity. When a speaker says something unclear, the professional response is to ask. A machine resolves the ambiguity silently by picking an interpretation, and no one in the room learns that a choice was made.
Where the Law Appears to Be Heading
Legislation is moving, though nothing here is settled.
The Language Access for All Act of 2026 was introduced in the House in January 2026 by Rep. Grace Meng with co-sponsors including Reps. Judy Chu, Dan Goldman, and Juan Vargas, and a Senate companion followed in July from Sens. Andy Kim and Mazie Hirono. The bills would codify language access obligations for federal agencies, which currently rest on less certain ground following the March 2025 executive order designating English as the official language of the federal government and revoking Executive Order 13166.
On AI specifically, the bills would bar federal agencies from fully replacing qualified translators and interpreters with AI or machine translation, require that AI-assisted output be reviewed by qualified professionals, direct NIST to provide validation protocols and standardization tools, and require agency Inspectors General to audit AI language systems at least every two years for accuracy, fairness, and cultural relevance.
These are introduced bills, not law, and they apply to federal agencies rather than private organizations. But the direction is informative: the emerging regulatory model treats AI as a supplement requiring human validation, not a substitute.
Professional guidance abroad has landed similarly. AUSIT, the Australian association, published a position statement in November 2025 that declines to endorse raw, unreviewed AI output for professional use and sets out a scale of human involvement in language services.
What This Means for Your Career
The realistic forecast is neither replacement nor stasis.
Volume at the low-stakes end will move toward automation, and some of it already has. That work — brief, routine, high-resource-language exchanges — was never where interpreters demonstrated their value.
What grows is the work that requires judgment, and a category that barely existed a few years ago: interpreters who supervise and validate AI output. California Health Care Foundation reporting in 2026 describes health systems using AI to generate draft translations that interpreters review and finalize, and recommends training interpreters to supervise AI workflows. That is a skill, it is billable, and few people currently have it.
The interpreters most exposed are those whose practice consists mainly of routine exchanges in major languages with no certification and no specialization. The interpreters least exposed work in high-stakes settings, hold credentials, work in less-resourced languages, or can evaluate machine output competently.
That is a strategic map, and it argues for training rather than against it.
How TLCLab Approaches This
At TLCLab, the position is the one the evidence supports: technology that expands access is welcome, and technology deployed where a fluent error cannot be caught is a patient safety problem.
Interpreters trained now should understand what these tools do well, where they fail and why, how to evaluate machine output, and how to explain the distinction to a client who is weighing cost against risk. An interpreter who can say clearly why consent, testimony, and behavioral health are different from wayfinding is more valuable than one who insists the technology is worthless — and far more valuable than one who has not thought about it.
The question was never whether machines can convert words between languages. They can, increasingly well, in some languages.
The question is whether anyone in the room will notice when it goes wrong.
Right now, in the encounters that matter most, that job still belongs to a person.