How Interpreter Performance Is Actually Scored
"That was good."
"You did fine."
"Your English is excellent."
Most interpreters have received feedback like this, and none of it tells you anything. It does not identify what you did well, what cost you, or what to practice tomorrow.
Certification exams do not work this way, and neither should training.
At TLCLab, interpreter performance is assessed using behaviorally anchored rating scales, the same approach the national certifying bodies use. Understanding how that scoring works changes how interpreters practice, and it is one of the most useful things a new interpreter can learn early.
What a Behaviorally Anchored Rating Scale Is
A behaviorally anchored rating scale, usually shortened to BARS, ties every point on a scoring scale to a specific, observable behavior instead of a vague adjective.
A traditional scale asks a rater to decide whether a performance was "excellent," "good," or "needs improvement." Those words mean different things to different people. Two raters watching the same performance can reasonably score it two points apart.
A behaviorally anchored scale removes that ambiguity. Instead of "good," the rater matches what the candidate actually did against a written description of what each score level looks like.
The method was built for exactly this problem: making judgment consistent when several different people are evaluating the same skill.
This matters in interpreting because interpreting has no answer key. There is no single correct rendition of a sentence. Two skilled interpreters will produce different words for the same utterance, and both may be accurate. What can be evaluated consistently is not word choice but behavior: what was preserved, what was lost, what was added, and whether the listener could follow.
The Scales the Certifying Bodies Actually Use
The Certification Commission for Healthcare Interpreters publishes the structure of its scoring, and it is worth studying.
Both the CHI oral performance exam and the ETOE exam are scored by human raters applying behaviorally anchored rating scales, developed and validated by subject matter experts working with a psychometrician.
Each item type on the ETOE exam is scored with a rubric built from three to five of these five scales:
Quality of Speech covers the physical delivery. Raters mark false starts, hesitations, repeated self-corrections, and pronunciation, intonation, or pace that make the message hard to follow.
Task Completion asks whether the candidate did what the task required. In a memory item, the expectation is to reproduce every word, so paraphrasing is an error there. In a paraphrasing item, repeating the words verbatim is the error. The same behavior can be correct in one task and wrong in another.
Accuracy and Cohesion/Coherence covers completeness and logic. Errors include omissions, especially of key points, additions that change meaning, irrelevant material, and statements that do not hold together.
Lexical Content covers how well units of information survive. A unit of information may be a word or a phrase carrying a single concept. Errors include inaccurate facts, word choices that shift meaning, and unnecessary changes of register.
Grammar covers command of the language: tense, agreement, pronouns, word order, incomplete sentences — where those interfere with the listener's understanding.
All five scales carry equal weight and are applied independently. And CCHI is explicit about the thread running through all of them: the main criterion across every scale is how accurately the candidate preserves the meaning of the original.
That is the whole discipline, expressed as a scoring system.
Register Is Scored
One detail deserves attention because interpreters routinely underestimate it.
Register — the level of formality a speaker chooses — is part of the lexical content score. Changing register without reason can lower the score.
This is the technical version of something interpreters are told constantly: do not clean up the speaker. When a patient speaks casually and the interpreter renders it in polished clinical language, the provider hears a different person than the one in the room. The information may survive. The speaker does not.
CCHI is careful here, and the nuance is instructive. Some tasks can only be completed by changing register, and in those cases it is not penalized. Raters are trained specifically on when a register shift is an error and when it is not.
That is what a well-built rubric does. It does not apply a rule mechanically. It defines the conditions under which a behavior counts.
What an Anchor Looks Like
Certifying bodies do not publish their anchor text, for obvious reasons. But the shape is not a secret, and seeing it makes the concept concrete.
An accuracy anchor set might run something like this:
5 — All units of information preserved. No omissions, no additions. Register matches the source.
4 — All key information preserved. Minor omission of non-essential detail. Meaning fully intact.
3 — Core message preserved, but one or more units of information lost or altered. A listener would follow the main point but miss a detail.
2 — Meaning partly distorted. An omission or addition that changes what the listener would understand or act on.
1 — Meaning not preserved. The rendition would mislead.
Those descriptions are illustrative rather than any organization's official text, but notice what they do. Two raters using them are not being asked for an opinion about quality. They are being asked to identify what happened to the information.
That is a question with a defensible answer.
Why Two Raters, and Why It Matters to You
CCHI's process is worth knowing because it explains why scores are trustworthy.
Raters do not know who the candidate is. Each response is scored independently by two raters. Raters do not score a whole exam — they score individual responses, so no one is forming an overall impression of a candidate and then scoring toward it. If two raters disagree by a point on a response, a third rater scores it.
Raters never learn whether a candidate passed.
This is deliberate insulation against the two biggest threats to fair scoring: halo effects, where a strong first impression lifts everything after it, and rater drift, where the same person scores differently on a Friday than on a Monday.
For an interpreter preparing for certification, the practical implication is direct. You cannot recover a weak response with a strong one, because a different person scored each one. Consistency across the whole exam is what passes.
Why "Good Enough" Is Not a Standard
Here is the part that reframes how people think about passing.
A certification exam is not built to identify excellence. It is built around a specific construct: the minimally competent candidate. CCHI defines that person as someone with enough skill to practice safely and competently, but not at the level of an expert.
Subject matter experts estimate what that candidate would score on each item. Those estimates are compared against real pilot data, discussed, and converted into a cut score, which is then scaled — for both the CoreCHI and ETOE exams, to a range of 300 to 600, with 450 as the passing score. Because different versions of an exam vary slightly in difficulty, a statistical process called equating keeps 450 equivalent across forms.
So "good enough" is not the standard, and the reason is precise.
"Good enough" is a feeling. It varies by rater, by day, by how tired everyone is, and by how the candidate compares to whoever tested before them. It cannot be defended and it cannot be replicated.
Minimal competence is a defined construct, anchored to described behaviors, validated against data, and held steady across exam versions. It is a floor placed deliberately at the point where patient safety is protected.
The difference is not strictness. It is whether the standard means the same thing twice.
What This Changes About Practice
Interpreters who understand rubric scoring practice differently.
They stop asking "was that good?" and start asking which scale a problem belongs to. An interpreter who keeps losing details has an accuracy problem. One who hesitates and restarts has a quality-of-speech problem. One who renders everything in formal language has a register problem sitting inside lexical content. These require different drills.
They also stop treating one rehearsal as evidence. A rendition that scores well once may have been luck. Certification measures consistency across many responses, and so does real work.
And they learn to self-assess against described behavior rather than feel. After a practice rendition: what units of information were lost? Was anything added? Did the register hold? That is a usable review. "It felt fine" is not.
One more practical note from CCHI's own guidance: the subdomain percentages on a score report are a rough indicator of relative strength, not a pass or fail on any section, and the total score is not their average. A candidate who fails is advised to improve across all activity types, not to chase the lowest number on the report.
How TLCLab Applies This
At TLCLab, oral performance is scored against behaviorally anchored scales rather than graded by impression, and students see the anchors.
Knowing the anchors is part of the training. An interpreter who understands what raters are counting can hear their own renditions the way a rater would, which is the skill that separates candidates who pass from candidates who were nearly ready.
The purpose of a rubric is not to make assessment harsher. It is to make feedback specific enough to act on.
"You did fine" tells you nothing.
"You omitted two units of information in the second rendition and shifted register on the third" tells you what to do tomorrow.