Machine Translation Quality: How to Evaluate MT Output (BLEU, COMET, and Human Evaluation)

Machine translation output can look flawless and still be wrong. Modern neural and LLM-based engines produce fluent, confident text even when they mistranslate, omit or invent content — which makes evaluating machine translation quality both harder and more important than ever. This guide walks through the metrics that matter in 2026 (BLEU, chrF, COMET, BLEURT), the human evaluation frameworks professionals actually use (adequacy/fluency, MQM), post-editing distance, reference-free quality estimation, and a step-by-step recipe for running your own bake-off between engines on your own content.
- Why evaluating machine translation is harder than it looks
- BLEU: the classic metric and why it falls short
- Neural metrics: COMET, BLEURT and chrF
- Metric comparison at a glance
- Human evaluation: adequacy, fluency and MQM
- Post-editing distance: measuring effort, not elegance
- Quality estimation: scoring without a reference
- How to run your own MT bake-off
- LLM translation and the hallucinated-fluency problem
- Key takeaways
- FAQ: evaluating machine translation quality
Why evaluating machine translation is harder than it looks
Ask ten people which translation of a sentence is "better" and you will often get ten different answers. Translation quality is multidimensional: a sentence can be perfectly fluent yet semantically wrong, or clumsy but faithful. Any serious evaluation has to separate at least two axes: adequacy (does the translation preserve the meaning of the source?) and fluency (does it read like natural target-language text?).
For decades this tension has shaped the field. Back when this site's research roots were in multilingual content analysis — the MULTISENSOR project spent years wrestling with cross-lingual NLP pipelines, and our early essay on machine translation between dream and reality captures how large the gap once was — the main failure mode was obvious: MT output was visibly broken. You did not need a metric to see it.
Neural MT and LLM translation inverted the problem. Output is now so fluent that errors hide in plain sight. The evaluation toolbox had to evolve from counting word overlaps to modeling meaning — and from trusting fluency to actively distrusting it.
BLEU: the classic metric and why it falls short
BLEU (Bilingual Evaluation Understudy), introduced by IBM researchers in 2002, compares an MT output against one or more human reference translations by counting overlapping n-grams (sequences of 1–4 words), with a brevity penalty for translations that are too short. It is fast, deterministic, language-independent in design, and it made large-scale MT research possible. Almost every MT paper for two decades reported BLEU.
But BLEU has well-documented blind spots:
- It rewards surface overlap, not meaning. "The medication should not be taken" and "The medication should be taken" share almost every n-gram — and differ catastrophically in meaning.
- It punishes legitimate variation. A perfect translation using synonyms and different word order scores poorly if it doesn't match the reference wording.
- It is corpus-level, not sentence-level. BLEU is unreliable on individual sentences; it was designed to average out over thousands of segments.
- It saturates at the top. Once systems are strong, BLEU differences of 1–2 points often fail to correlate with human preference. The WMT metrics shared task has repeatedly shown that BLEU can rank strong systems in the wrong order (see the annual findings at statmt.org).
- Implementation drift. Tokenization choices change scores; if you must report BLEU, use sacreBLEU so numbers are comparable.
The practical verdict in 2026: BLEU is still useful as a cheap regression test — a sanity check that a new model version hasn't collapsed — but it should never be the deciding metric when comparing modern engines.
Neural metrics: COMET, BLEURT and chrF
The current generation of automatic metrics uses pretrained language models to estimate semantic similarity instead of counting word matches.
COMET (Crosslingual Optimized Metric for Evaluation of Translation), developed by Unbabel, feeds the source sentence, the MT output and the reference into a multilingual encoder and predicts a quality score trained on human judgments. Because it sees the source, not just the reference, COMET can catch meaning errors that reference-overlap metrics miss. It has topped correlation-with-humans rankings at WMT for several years and is the de facto standard in MT research. The models are open source (Unbabel/COMET on GitHub).
BLEURT, from Google Research, takes a similar learned approach: a BERT-based model fine-tuned on human ratings. It behaves comparably to COMET on many language pairs, though COMET's use of the source sentence gives it an edge on adequacy errors.
chrF deserves special mention because it is not neural — it computes character-level F-scores — yet it consistently beats BLEU in correlation with human judgment, especially for morphologically rich languages (Finnish, Turkish, Czech) where word-level n-grams fragment. It is as cheap and reproducible as BLEU, which makes it the best "simple" metric to keep around.
Two caveats with neural metrics. First, they are black boxes: a COMET score of 0.85 tells you "probably good" but not what is wrong. Second, they can inherit biases from their training data — they were trained mostly on news-domain judgments, so scores on legal, medical or highly technical text should be treated with more skepticism.
Metric comparison at a glance
| Metric | What it measures | When to use it | Main limitations |
|---|---|---|---|
| BLEU | Word n-gram overlap with reference(s) | Cheap regression testing; comparability with older research | Ignores meaning; penalizes valid paraphrase; unreliable per sentence; saturates on strong systems |
| chrF | Character-level F-score vs. reference | Simple metric of choice; morphologically rich languages | Still surface-based; no semantic understanding |
| TER / HTER | Edit distance: how many edits to fix the output | Estimating post-editing effort and localization cost | Needs references (or human post-edits for HTER, which is expensive) |
| COMET | Learned quality score from source + output + reference | Ranking engines; primary automatic metric for decisions | Black box; GPU helps; weaker out of news domain |
| BLEURT | Learned quality score from output + reference | Alternative/complement to COMET | Same black-box caveats; doesn't see the source |
| COMETKiwi (QE) | Learned quality score with no reference, from source + output only | Production monitoring; routing; scoring at scale | Less accurate than reference-based COMET; can be fooled by fluent hallucinations |
| Human MQM | Typed, weighted error annotation by professionals | Final decisions, high-stakes content, metric validation | Slow, costly, needs trained annotators and clear guidelines |
Human evaluation: adequacy, fluency and MQM
Automatic metrics exist to approximate human judgment, so human evaluation remains the gold standard — provided it is structured. Unstructured "which one do you like?" reviews produce noisy, unrepeatable results.
The two classic protocols:
- Adequacy/fluency rating: annotators score each segment on two separate scales (typically 1–5 or 0–100): how much source meaning is preserved, and how natural the target text reads. Keeping the axes separate is the whole point — it stops fluent-but-wrong output from getting a free pass.
- Ranking / direct assessment: annotators rank alternative translations of the same source, or slide a 0–100 quality bar. This is what WMT uses at scale; it is easier for annotators than absolute scoring.
For professional and enterprise settings, the reference framework is MQM (Multidimensional Quality Metrics) — see themqm.org. Instead of a single score, trained annotators mark each error with a type (accuracy: mistranslation, omission, addition; fluency: grammar, spelling; terminology; style; locale conventions) and a severity (minor, major, critical). Severity-weighted error counts per 1,000 words give a score that is diagnostic, comparable across projects, and hard to game. MQM is what Google, Meta and the WMT organizers now use to build the "ground truth" that metrics like COMET are trained against.
Practical tips if you run human evaluation yourself: use at least two annotators per segment and measure their agreement; blind the engine names; mix in a few human translations as hidden controls; and write down your error definitions before you start.
Post-editing distance: measuring effort, not elegance
If your MT output feeds a human post-editing workflow — the standard setup in localization — the most business-relevant question is not "how good is it?" but "how much work does it take to fix?". That is what edit-distance metrics capture:
- TER (Translation Edit Rate): the minimum number of insertions, deletions, substitutions and shifts needed to turn the MT output into a reference, divided by reference length.
- HTER (Human-targeted TER): the same computation, but against an actual human post-edit of that specific output. HTER is the honest version — it measures real effort rather than distance to an arbitrary reference — but it requires paying post-editors, so it is usually reserved for periodic engine benchmarks.
A related operational KPI is post-editing time per segment, logged by CAT tools. Time correlates with cost more directly than edit counts do (some errors are quick to spot but slow to research). If you buy MT for localization, negotiate benchmarks in HTER or PE-time, not BLEU.
Quality estimation: scoring without a reference
Everything above assumes you have reference translations. In production you don't — you have millions of source/MT pairs and no gold standard. Quality estimation (QE) models predict a quality score from the source and the MT output alone. The best-known family is COMETKiwi (a reference-free variant of COMET), alongside LLM-as-judge approaches where a large model is prompted to rate or annotate translations MQM-style.
QE unlocks workflows that reference-based metrics can't touch:
- Routing: auto-publish segments above a QE threshold, send the rest to human review.
- Monitoring: track QE score distributions per language pair over time and alert on drift after engine updates.
- Engine selection per segment: translate with two engines, keep the higher-QE output.
The caveat: QE models are themselves imperfect judges and share a weakness with human skim-readers — very fluent output biases them upward. Calibrate any QE threshold against a sample that humans have MQM-annotated before you trust it with auto-publish decisions.
How to run your own MT bake-off
Public benchmarks tell you which engine wins on news text in high-resource languages. They do not tell you which engine wins on your product descriptions, your support macros or your legal boilerplate. Rankings genuinely flip by domain and language pair — our DeepL vs Google Translate comparison is a good illustration of how two top engines diverge depending on content type. Here is a lightweight, defensible bake-off procedure:
- Build a test set from your own content: 300–500 segments sampled across your real document types. Hold it secret — never send it to any vendor for training.
- Create references (optional but valuable): have a professional translate the set, or post-edit carefully. No budget? Skip to QE-plus-human-spot-check.
- Translate with every candidate engine on the same day, with the same glossary/settings each engine supports. Candidates worth testing are covered in our roundup of the best machine translation software.
- Score automatically: COMET as the headline metric, chrF as the cheap cross-check. Report both; if they disagree strongly, investigate before concluding anything.
- Human-check the disagreements and a random sample: 50–100 segments, blind, two reviewers, a simplified MQM card (accuracy/terminology/fluency × minor/major/critical). This is where you catch the errors metrics miss.
- Slice the results by content type and language pair. One engine rarely wins everywhere — the actionable output is a routing table, not a single champion.
- Re-run quarterly. Engines update silently; last year's winner is a hypothesis, not a fact.
The same discipline applies one step earlier in the pipeline: if your source text comes from speech, transcription errors propagate straight into translation errors, so benchmark that stage too — our guide to AI transcription tools covers how to evaluate it.
LLM translation and the hallucinated-fluency problem
LLM-based translation (GPT-4-class models, or LLM modes inside DeepL and Google) raises average quality on many pairs, but it changes the error profile in ways your evaluation must account for:
- Hallucinated fluency: when the model is unsure, it doesn't produce broken output — it produces beautiful output with invented content. A number changed, a negation dropped, a plausible clause added. These are exactly the errors human skimmers and QE models under-detect, and they are more dangerous than visible garbage because nobody flags them.
- Omission by summarization: LLMs sometimes compress long or repetitive source passages instead of translating them fully. Length-ratio checks (target/source character ratio per segment) are a crude but effective tripwire.
- Instruction bleed-through: content that looks like an instruction ("click here", "ignore the above") occasionally gets executed instead of translated.
- Inconsistent terminology across segments, since each call is independent unless you engineer glossaries and context windows.
Concretely: when evaluating LLM translation, weight accuracy error categories (omission, addition, mistranslation of numbers, entities and negation) above fluency ones, add automated checks for number/entity mismatch between source and target, and never rely on fluency-correlated metrics alone. A translation that reads perfectly is exactly the one you should test hardest.
Key takeaways
- Use COMET (plus chrF as a cross-check) instead of BLEU for any decision that matters; keep BLEU/sacreBLEU only for regression testing and historical comparability.
- Use MQM-style human evaluation on a sample to validate metrics and catch what they miss — especially with LLM engines, where fluent hallucinations are the dominant risk.
- Use HTER or post-editing time when the business question is cost, and reference-free QE (COMETKiwi) when you need scores at production scale.
- Benchmark on your own content, slice by domain and language pair, and repeat quarterly.
For more on the engines themselves and the wider tooling around multilingual pipelines, browse the rest of our Language & Document AI section.
FAQ: evaluating machine translation quality
What is a good BLEU or COMET score?
There is no universal threshold — scores are only meaningful when comparing systems on the same test set. As rough intuition: BLEU above ~40 on news-style text usually indicates a strong system, and COMET scores (wmt22-comet-da model) above ~0.85 typically correspond to translations humans rate as good. But a BLEU of 35 on legal Finnish may be excellent while 45 on French e-commerce copy may be mediocre. Compare deltas between systems, not absolute numbers.
Can I just use an LLM to judge translation quality?
Partially. LLM-as-judge with a structured MQM-style prompt correlates surprisingly well with human annotators on many language pairs and is far cheaper. But it shares blind spots with the systems it judges — notably a bias toward fluent output — and it degrades on low-resource languages. Use it as a scalable first pass, validated against a human-annotated sample, not as the final arbiter.
Do I still need human evaluation if I use COMET?
Yes, in two situations: whenever the decision is high-stakes (choosing a vendor, publishing without post-editing, regulated content), and periodically to verify that the metric still tracks human judgment on your domain. Neural metrics were trained mostly on news data; on your specific content they may drift. A few hundred MQM-annotated segments per quarter is usually enough to keep them honest.
How is quality estimation different from a normal MT metric?
Reference-based metrics (BLEU, chrF, COMET, BLEURT) compare the MT output against a human reference translation, so they only work on curated test sets. Quality estimation (QE) models such as COMETKiwi predict quality from the source and the output alone, with no reference — which means they can score every segment you translate in production, at the cost of somewhat lower accuracy and a known weakness for fluent hallucinations.
Recommended: