# Translation metric comparison: COMET vs. character n-gram F-score (chrF)—2026 primary pick

Lauren Sanders · September 28, 2026

> COMET earns the 2026 primary quality pick, while chrF audits near-match forms; a 32% too-short baseline exposes why human glossary review remains essential.

| Takeaway | Detail |
| --- | --- |
| COMET gets the primary quality vote. | The benchmark finding that one baseline system produced outputs 32% too short shows why learned quality ranking matters, but COMET still cannot certify that a terminology choice matches the glossary. |
| chrF gets the narrow form audit. | chrF gives tokenization-independent partial credit for near-miss forms, while the benchmark’s 32% too-short result must be read as a length failure, not as proof that terminology passed. |
| Human review keeps the release veto. | If COMET rewards a fluent sentence with the wrong technical term, a glossary-backed checker must decide; neither the 32% benchmark failure nor any automatic score can authorize release. |
| Do not crosswire the metrics. | The cited 32% statistic concerns output length, not a paired COMET–chrF comparison, and the supplied evidence provides no conversion, normalization, or crosswalk between their scales. |

The cited Arabic–Russian benchmark reports that one baseline system produced outputs 32% too short. The finding captures the distinction between overall translation quality and terminology control. A learned evaluator can judge whether an output is plausible and useful, while character n-gram overlap can inspect whether its surface form resembles a human reference. Neither signal, by itself, settles whether a domain term is correct.

For the scoped primary vote, use COMET as the stronger automatic quality judge when terminology consistency and low-resource adaptation matter. On the same terminology-heavy sentence, COMET can reward a plausible rendering while chrF can expose a near-miss form. Neither verdict is a glossary ruling: a fluent paraphrase may use the wrong term, while a correct term can appear in a sentence COMET ranks poorly.

Keep a narrow chrF check beside it. Configured chrF is tokenization-independent and gives partial credit for near-miss forms, making it useful for spelling, inflection, and reference-form checks. Yet it can reward surface resemblance without proving semantic adequacy; the cited evidence has no paired COMET–chrF results or metric conversion and does not identify the COMET checkpoint. The release rule is explicit: COMET casts the primary quality vote, chrF audits form, and a human checker with the glossary keeps the veto.

![Translation metric comparison](https://static.mm-ais.com/article-images-ai/translation-metric-comparison-comet-vs-c-ai-dc791c62.jpg)

## Inside the Scorers: COMET Context vs. chrF Character-Level Comparison

The decisive distinction is observational: COMET estimates quality from contextual relationships, while chrF counts what two strings visibly share. Neither observation is a glossary audit.

According to Wenpeng Rei et al.’s original COMET paper, reference-based COMET is a supervised scalar-quality model trained to predict human judgments. The contextual encoding stage produces separate representations of the source, candidate translation, and trusted reference; a learned quality head pools those representations and emits a scalar quality estimate. It is therefore a compressed estimate for the whole segment, not a per-token correctness vector or a probability that each protected term is correct.

COMETKiwi changes the inference mode, not merely the label. According to Unbabel’s COMETKiwi release notes, the model is based on InfoXLM and reads only the source plus candidate translation. Its output is explicitly reference-free: there is no gold reference against which the candidate is compared. Calling this “reference-based COMET with a missing reference” would fabricate an evaluation relation the model never observed.

chrF begins where neural pooling stops. According to Maja Popović’s 2015 definition, it counts character n-gram matches for orders n=1 through 6. Repeated n-grams receive clipped multiset credit—the minimum of their candidate and reference counts—so repetition cannot manufacture unlimited matches. For order n, let M_n be pooled overlap, H_n the candidate total, and R_n the reference total: P_n=M_n/H_n and r_n=M_n/R_n. Popović’s F-score is (1+β²)P_nr_n/(β²P_n+r_n). With β=2, Induwara Ashinsana’s calculator describes recall as weighted twice as heavily as precision, while the published expression applies β² explicitly.

For aggregation, collect those counts by order over the declared evaluation unit, calculate each valid order’s F-score, and average the order scores. Sentence length changes the denominators and can remove high-order evidence for short candidates. Punctuation changes the character stream, while Unicode normalization can change whether visually identical sequences have identical code-point n-grams. Freeze sentence segmentation, punctuation handling, and normalization in the run specification; otherwise chrF partly measures undocumented preprocessing. An order not represented on both sides must follow the implementation’s stated skip rule.

The output semantics establish the division of labor. COMET compresses contextual evidence into an estimate of overall quality; chrF remains a rule-based surface-overlap statistic that can flag near-miss spellings, inflections, and forms. Consider a fluent segment with one wrong protected term: COMET may still rank it well because its scalar pools the segment, while chrF may remain high because most other character n-grams match. High chrF therefore does not certify terminology, and strong COMET does not verify the glossary. Use COMET as the primary human-aligned ranking signal, retain chrF as a secondary character-level alarm, and require human verification of every protected term before release.

| Signal | Input and fixed configuration | Defensible diagnosis | Placement |
| --- | --- | --- | --- |
| Reference-based COMET | Source, candidate, and trusted reference | Overall contextual translation quality | Primary ranking signal because it learns from human-aligned quality judgments |
| chrF | Orders n=1–6, β=2; Induwara Ashinsana reports a conventional 0–100 display | Character-form, spelling, inflection, and overlap alarms | Secondary lexical diagnostic; protected-term verification remains the release gate |

![Inside the Scorers: COMET Context vs. chrF Character-Level Comparison — Translation metric comparison](https://static.mm-ais.com/article-images-ai/translation-metric-comparison-comet-vs-c-ai-8f2e25e7.jpg)

## Human-Judgment Evidence

The human evidence supports a narrower claim than “COMET is a universal judge”: use it to rank translations, not certify terminology. Rei et al. (*COMET: A Neural Framework for MT Evaluation*) compared COMET with BLEU, chrF, and TER on human-labeled MT data and reported the strongest aggregate correlation with human quality judgments. That controlled comparison is why COMET deserves the primary ranking slot; correlation with quality is not proof that every specialist expression is acceptable.

An additional shared-task analysis by Patrick et al. broadens the evidence beyond a single controlled benchmark to multiple language directions and machine-translation systems. That distinction matters: replication across held-out data is stronger selection evidence than a result confined to one evaluation set. Still, “overall” remains bounded by the news benchmark; it cannot be converted into a claim about specialist terminology.

Kendall’s tau must be reported at two units rather than collapsed into one headline. System-level tau estimates whether one MT system ranks above another, making it evidence for model selection. Segment-level tau estimates whether COMET or chrF orders individual candidate segments like human evaluators, making it evidence for error analysis and release inspection. A system-level lead cannot be silently transferred to segment ordering, and a strong segment score cannot by itself certify a system ranking.

Dispersion is the necessary skeptical check. For each WMT20 direction, calculate the signed gap for each reported COMET variant, Δd = τCOMET,d − τchrF,d; report the largest and smallest values and identify the narrowest pairing as the direction with the smallest absolute gap. The supplied evidence cannot support that extraction: Arabov’s benchmark does not report both metrics for the same system and evaluation set. Naming a maximum, minimum, or winning direction from those unmatched entries would manufacture evidence. A narrow gap would mean that character overlap tracks human ordering especially well there; here, no direction can be identified honestly.

Script-sensitive segmentation and morphology can move chrF without changing human-perceived quality, while a narrow spread of system quality can depress tau. These are hypotheses to test, not established causes. Arabov’s Arabic–Russian analysis makes the morphological mechanism concrete through non-concatenative roots and patterns, cliticization, dialect variation, and low lexical overlap with Russian, but it does not locate the shared-task extrema. The inference boundary is exact: WMT news supports COMET for overall human-aligned ranking; it does not establish performance on medical, legal, technical, or low-resource terminology without separately labeled domain evaluation. A high chrF score therefore cannot certify specialist terminology merely because approved wording shares character n-grams.

| Evidence test | Scale or required scope | Valid interpretation | Operational decision |
| --- | --- | --- | --- |
| WMT shared-task validation | Held-out segment evidence across multiple translation directions | Official tables placed the cited COMET variants in the leading group, ahead of chrF overall | Use COMET as the primary ranking signal; retain chrF as a secondary lexical alarm. |
| System-level Kendall’s tau | Systems rather than individual sentences | Whether one MT system ranks above another like human evaluators | Base model selection on COMET’s system-level result. |
| Segment-level Kendall’s tau | Individual candidates rather than systems | Whether COMET or chrF orders candidate segments like human evaluators | Report separately and inspect disagreements; do not substitute chrF’s lexical behavior for COMET’s ranking. |
| Protected-term release gate | Every protected term in a separately labeled domain evaluation | Whether required domain wording is acceptable | Require human verification before release. |

## The Comparison: COMET Wins the Primary Slot

COMET should occupy the primary slot because its operational role is system ranking, not terminology approval—not because the supplied record proves a universal numerical margin. The Grok search-result set contains no direct COMET–chrF rankings or correlation differences; the Welo Global supplied excerpt contains no per-system chrF values or confidence intervals, and the supplied chrF paper snippet contains no sample sizes or significance statistics. The comparison therefore assigns metric roles without manufacturing an effect size.

Reference status is a mandatory decision branch. When a trusted reference exists, use standard COMET. When only a source and candidate are available, use a reference-free COMET model and label the output “quality estimation,” not reference-based human-aligned evaluation. Keep those branches distinct in result tables, filenames, and leaderboards; never silently substitute one for the other.

Before scoring, pre-register the segment as the unit of analysis. Report the segment mean and system mean, and attach a bootstrap confidence interval to each paired candidate comparison by resampling the shared source segments. Freeze the same segment set, language direction, and text-normalization policy for both metrics. This makes the comparison paired and auditable rather than a juxtaposition of independently selected results.

Universal pass marks and undocumented model swaps should be prohibited. Publish the COMET checkpoint, domain, language direction, and preprocessing: as Arabov’s supplied primary excerpt demonstrates, an unidentified checkpoint, model variant, or score-normalization scheme is not a reproducible term-checking configuration, and values from different checkpoints are not commensurate. For chrF, publish character normalization and implementation settings; according to Induwara Ashinsana’s supplied calculator, its defaults are character order 6 and β=2, with whitespace stripped and case-sensitive scoring. Maja Popović’s documented tokenization independence aids cross-language character comparison, but a high chrF score still cannot certify specialist terminology. Human glossary adjudication, not either metric, controls release.

| Criterion | COMET | chrF | Use |
| --- | --- | --- | --- |
| Overall human-aligned ranking | Winner: primary ranking signal | Secondary evidence only | Rank systems with COMET and publish paired uncertainty. |
| Paraphrase tolerance | Favored for semantically adequate rephrasing | Sensitive to visible wording changes | Prevent valid paraphrases from being rejected on surface dissimilarity alone. |
| Character-divergence clues | Secondary clue | Winner for character-level diagnosis | Use chrF to flag edits, spacing, punctuation, and morphology for inspection. |
| Reference-free operation | Favored through a reference-free model | Unavailable without a reference | Label source-and-candidate output as quality estimation. |
| Protected-term approval | No winner | No winner | Require human glossary verification of every protected term before release. |
| Primary-metric winner: COMET | Selected as primary | Retained as secondary lexical alarm | Rank with COMET; chrF diagnoses; humans approve terminology. |

## Counter-Evidence

The canonical rule is conditional, not a certificate: COMET earns the primary slot only as a human-aligned ranking signal, chrF remains a surface-form alarm, and protected terminology stays a human release decision. These counterexamples narrow that claim without reversing it.

Consider a clinical segment whose protected term is *interferon beta-1b*. Replace it with the fluent, semantically neighboring drug name *interferon beta-1a*. The surrounding sentence can remain natural and contextually plausible, allowing whole-segment COMET quality to remain high even though a safety-critical glossary match has failed. COMET can therefore rank the segment well without certifying the entity that must be controlled.

The inverse chrF failure is equally direct. A semantically valid candidate may render *interferon beta-1b* as the morphological shorthand *IFN-beta-1b*, losing character overlap with the reference. A fluent but incorrect candidate, *interferon beta-1a*, reuses most of the reference’s characters and can consequently rank above the valid synonym. This kills the myth that high chrF certifies specialist terminology merely because approved wording shares character n-grams: surface proximity is not contextual or glossary correctness.

Aggregation hides the same defect at system level. One system can place first on COMET across an entire test set while consistently missing a critical term span. A global rank correlation describes the resulting ordering; it does not report critical-error recall. The first-place result remains useful for broad comparison, but it cannot authorize release when a protected term is absent.

Neither score is invariant across cases. Compounding, clitics, punctuation normalization, transliteration, and source-language segmentation can change chrF ordering by altering visible overlap. COMET–human alignment can also shift across language pairs and domains, especially where terminology, syntax, or discourse conventions differ. The 2026 decision rule is justified only when COMET’s alignment is checked for the relevant language and domain; otherwise, its ranking remains provisional rather than dispositive.

References introduce a separate counter-case. An idiosyncratic or flawed reference may reward a system for echoing its preferred terminology while penalizing a clinically valid alternative. Welo Global’s distinction between formal acceptability and a client’s required wording illustrates why reference distance can demand editing even when no conventional translation error is apparent. High COMET or chrF therefore cannot establish that the reference itself is terminologically correct.

Human verification is necessary but not infallible. Raters may disagree about whether a span is critical and about its severity. A defensible process must preserve the adjudication rationale: the protected term considered, the competing readings, the severity judgment, and the resolution. “Human-checked” describes an accountable process, not an automatic guarantee. Accordingly, COMET remains the primary human-aligned ranking signal, chrF remains secondary, and verified glossary compliance remains the release gate.

| Counter-evidence | What the score conceals | Required response |
| --- | --- | --- |
| Fluent COMET segment | Plausible context with a wrong protected drug | Retain COMET for ranking; block release pending glossary review |
| High chrF overlap | Accidental resemblance outweighing a valid synonym | Use chrF only as a lexical alarm |
| Top aggregate COMET rank | Repeated misses of a critical term span | Audit critical-error recall separately |
| Cross-case score variance | Language- and domain-dependent ordering | Validate alignment within the target setting |
| Flawed reference | Metric rewards incorrect reference terminology | Validate glossary and reference provenance |
| Disputed human judgment | Unresolved criticality or severity | Preserve rationale and adjudicate before release |

## WMT English-to-Finnish Case

A reproducible WMT English-to-Finnish result is not recoverable from the supplied record; naming systems or filling result cells would be fabrication. The record lacks exact COMET scores, exact chrF-versus-COMET gaps, system rankings, and segment-level evaluations. It therefore cannot support the requested reversal pair, system means, MQM ranks, or terminology-error trace.

The case brief also assigns a corpus size and a system-pool size, but no official WMT Metrics Shared Task manifest is attached to verify them. Those quantities should remain unstated until the source segments, released predictions, references, and human-assessment records reconcile. This is a denominator problem, not bookkeeping: COMET, chrF, and MQM must rank the same systems over the same segment IDs before a reversal is meaningful.

The extraction rule should be deterministic. Join submissions by source segment ID; score every system with the released COMET checkpoint and declared chrF settings; aggregate each metric only over the matched segment set; and attach adjudicated MQM results at the system level. Rank each metric independently in descending quality order while preserving COMET, chrF, and MQM on their native scales. For every COMET–chrF pair, calculate the absolute difference between the system’s ordinal ranks under the two orderings. The pair with the largest difference is the reversal case, with a predeclared tie-break. Raw COMET and chrF magnitudes must never be subtracted.

Within the MQM-leading system, scan source segment IDs in ascending order and take the first adjudicated named-entity or technical-term error. Record the source and output spans, glossary-required form, segment ID, adjudicated severity, full-segment COMET score, and the matching and possible character counts for every configured chrF n-order, together with precision, recall, and F-score. An error count of zero is valid only after a complete scan and adjudication. Here, the count is unknown: missing artifacts are not negative evidence.

| Required result | Evidence-backed status | Defensible action |
| --- | --- | --- |
| Two systems forming the largest reversal | 0 stated rankings in the Grok search-result set | Do not name either system until the official scores are reproduced. |
| COMET system means and ranks | 0 exact COMET scores in the supplied search-result set | Recompute over matched segments; emit no synthetic system rows. |
| chrF system means and ranks | 0 exact chrF-versus-COMET score gaps in the supplied search-result set | Compute descriptively, without inferring a crossover from the missing result. |
| Terminology-error segment | No segment ID, MQM severity, full-segment COMET score, or chrF n-order counts supplied | Record “unknown,” not “zero errors,” and withhold glossary clearance. |

The actual case close is therefore: COMET’s ordering cannot be compared with the human ordering; COMET’s acceptability judgment for the terminology segment is not computable; chrF’s surface-mismatch detection is not computable; and absent human adjudication blocks release. If a completed run produces human–COMET disagreement, that is a documented COMET calibration failure—not grounds for silently promoting chrF. COMET remains the primary ranking signal, chrF remains the secondary lexical alarm, and verified glossary terms remain the release condition.

## Five Rules for a Defensible COMET

A defensible COMET evaluation is an auditable decision path, not a persuasive decimal. COMET supplies the human-aligned ranking signal; chrF supplies only a character-level alarm; human glossary review supplies the release veto. The shortcut to reject is “high chrF means correct terminology.” According to Welo Global, a higher chrF score means that output is closer to what a qualified human would write, not that the translation is necessarily semantically correct. In Welo Global’s 2026 benchmark, Opal’s 16-point chrF lead over the next-closest MT system demonstrates why magnitude is not authority: a conspicuous surface overlap still cannot certify a protected specialist term.

The mechanism is conditional. A reference-based checkpoint, a reference-free quality-estimation model, and a newly introduced language or domain do not create interchangeable evidence. Pinning prevents silent checkpoint drift; explicit QE labeling prevents false equivalence; local human labeling tests whether ranking behavior and terminology sensitivity transfer. Published preprocessing makes a score reproducible rather than merely repeatable on one machine. The glossary check stays outside aggregate arithmetic, and a statistically unresolved comparison goes to blinded humans rather than to chrF by default. That separation also makes the audit intelligible: a ranking can be challenged without disputing terminology approval, and a confirmed term failure can block release without rewriting the ranking.

| Condition | Required procedure | Decision consequence |
| --- | --- | --- |
| Reference available | Use one pinned, domain-validated, reference-based COMET checkpoint as the primary score. Publish its exact identifier and preprocessing, and retain chrF only as a secondary character-level diagnostic. | COMET controls the human-aligned ordering; chrF remains advisory and cannot overrule it. |
| Reference unavailable | Use a reference-free COMET model explicitly labeled for quality estimation, or QE. | Never substitute chrF as a semantic-adequacy judgment. Missing reference evidence changes the evidence type, not the metric hierarchy. |
| New language or domain | Human-label representative segments, including terminology-error-positive cases, stratified by error presence and severity. Before deployment, validate both rank correlation and critical-error recall. | Do not import another checkpoint’s pass mark; deployment requires local evidence for both properties. |
| Protected glossary | Inspect all protected-term occurrences a Frequently Asked Questions Which metric should receive the primary quality vote, and can either metric authorize release? COMET is the primary quality-ranking signal and chrF is a secondary form audit, but a human checker with the glossary must verify every protected term and retains the release veto. Does COMETKiwi compare a candidate translation with a gold reference? No, COMETKiwi is explicitly reference-free and reads only the source and candidate translation, unlike reference-based COMET, which also receives a trusted reference. What chrF configuration is specified for the comparison? The specified configuration uses character n-gram orders n=1 through 6 with β=2, while the cited calculator reports chrF on a conventional 0–100 scale. Can a candidate with a wrong protected term pass release if COMET ranks it well and chrF is high? No, COMET may reward a plausible fluent segment while chrF remains high because other character n-grams match, so glossary-backed human verification is still required. Does the Arabic–Russian baseline output being 32% too short prove that COMET outperformed chrF or that its terminology was correct? No, the 32% finding is an output-length failure rather than a paired COMET–chrF comparison, and it does not establish terminology correctness. Can the supplied WMT20 evidence identify the direction where COMET and chrF have the smallest gap? No, Arabov’s benchmark does not report both metrics for the same system and evaluation set, so naming a maximum, minimum, or narrowest direction would manufacture evidence. Quick answers Which metric gets the primary quality vote? | COMET casts the primary quality vote, chrF audits form, and a human checker with the glossary keeps the veto. |
| What narrow check is chrF assigned? | Configured chrF is tokenization-independent and gives partial credit for near-miss forms, making it useful for spelling, inflection, and reference-form checks. |  |
| Can COMET or chrF certify terminology correctness? | Neither verdict is a glossary ruling: a fluent paraphrase may use the wrong term, while a correct term can appear in a sentence COMET ranks poorly. |  |
| What does the cited 32% statistic actually measure? | The cited 32% statistic concerns output length, not a paired COMET–chrF comparison, and the supplied evidence provides no conversion, normalization, or crosswalk between their scales. |  |
| How do COMET and chrF differ in their observations? | COMET compresses contextual evidence into an estimate of overall quality; chrF remains a rule-based surface-overlap statistic that can flag near-miss spellings, inflections, and forms. |  |

Also worth reading: **chrF vs BLEU: WMT23 Legal Task Proves They Really Differ**: [chrF vs BLEU: WMT23 Legal](https://aitranslations.io/blog/chrf-vs-bleu-wmt23-legal-task-proves-they-really-differ.php) · **English Russian Fitness Translation: No Language Left Behind (NLLB) chrF++ 60 vs Tune**: [English Russian Fitness Translation: No](https://aitranslations.io/blog/english-russian-fitness-translation-no-language-left-behind-nllb-chrf-60-vs-tune.php) · **COMET's 12% Edge Over BLEU for Swahili Domain Shifts**: [COMET's 12% Edge Over BLEU](https://aitranslations.io/blog/comets-12-edge-over-bleu-for-swahili-domain-shifts.php)

### Related reading

- [AI Translation Accuracy in Mental Health Education Materials A 7-Language Comparison Study](https://aitranslations.io/blog/ai_translation_accuracy_in_mental_health_education_materials.php)
- [ChatGPT vs

Google Translate A 2024 Comparison of AI Translation Capabilities](https://aitranslations.io/blog/chatgpt_vs_google_translate_a_2024_comparison_of_ai_transla.php)
- [Character Name Translation in AI When and How to Preserve Original Names During Language Conversion](https://aitranslations.io/blog/character_name_translation_in_ai_when_and_how_to_preserve_or.php)
- [The Impact of Character Count on AI Translation Accuracy A 2024 Analysis](https://aitranslations.io/blog/the_impact_of_character_count_on_ai_translation_accuracy_a_2.php)
- [COMET-22 Scores, WMT Data, and the 2026 Translation Decision](https://aitranslations.io/blog/comet-22-scores-wmt-data-and-the-2026-translation-decision.php)
- [English Russian Fitness Translation: No Language Left Behind (NLLB) chrF++ 60 vs Tune](https://aitranslations.io/blog/english-russian-fitness-translation-no-language-left-behind-nllb-chrf-60-vs-tune.php)

### Latest

- [English Welsh medical translation: 10k stack vs generic deploy decision](https://aitranslations.io/blog/english-welsh-medical-translation-10k-stack-vs-generic-deploy-decision.php)
- [Swahili Medical Translation: No Language Left Behind (NLLB) Tune Cuts 31% vs...](https://aitranslations.io/blog/swahili-medical-translation-no-language-left-behind-nllb-tune-cuts-31-vs-prompt.php)
- [Medical translation quality: 5% drift triggers retraining](https://aitranslations.io/blog/medical-translation-quality-5-drift-triggers-retraining.php)

Canonical: https://aitranslations.io/blog/translation-metric-comparison-comet-vs-character-n-gram-f-score-chrf2026-primary-pick.php
Markdown: https://aitranslations.io/blog/translation-metric-comparison-comet-vs-character-n-gram-f-score-chrf2026-primary-pick.php/index.md
