CometKiwi 0.85 Gate: Threshold Math, 0.80 vs 0.85 vs 0.90

CometKiwi 0.85 Gate

```html

Inside the 0.85 Gate

Content for Inside the 0.85 Gate is being prepared.

Inside the 0.85 Gate — CometKiwi 0.85 Gate

The Numbers Behind the Gate

The 0.85 gate stands on a correlation coefficient, and that single fact explains both why it saves money and why it leaks. According to the WMT22 Quality Estimation shared-task results, CometKiwi-class systems reached segment-level Pearson correlations of roughly 0.55–0.63 against human direct-assessment scores. A ranker in that band orders translations far better than chance, but it misranks a meaningful minority: some defective segments clear any plausible cut line, and some clean segments fall below it. Every dollar the gate saves comes from trusting the top of that ranking; every escaped defect comes from its imperfect ordering. The savings and the leaks are the same statistical property, priced two ways.

What keeps the leaks survivable is what the metric catches at the extremes. According to Dale et al.'s 2023 study "Detecting and Mitigating Hallucinations in Machine Translation," CometKiwi-family QE models detect hallucinated translations with F1 around 0.75 or higher. Hallucinations — fluent output untethered from the source — are the worst failure mode in low-resource machine translation precisely because a monolingual reviewer often cannot see them. A detector operating at that F1 intercepts most of them before publication, and that safety-net property is what makes auto-publishing defensible at all. Notice the asymmetry: the metric is strongest at the catastrophic tail and weakest in the crowded middle of the score distribution, which is exactly where publish-or-edit decisions cluster.

Two baselines frame what "safe" means. According to QTLaunchPad and QT21 MQM error studies, critical errors appear in a meaningful share of raw MT segments even in professional low-resource pipelines — so unmitigated publishing already ships defects, and the gate's job is reduction, not perfection. Against that baseline sits the decisive leak finding: according to Unbabel's published QE-in-production write-ups, a 0.85 cut leaves a measurable share of auto-published segments carrying at least one MQM major error. Those leak measurements were made on content resembling the news-domain data the metric was trained on. Move your content — Khmer legal text, Sinhala clinical material — and the same 0.86 maps to a different defect probability, because the model's calibration does not transfer across pairs and domains. A Khmer segment scoring 0.86 and a German segment scoring 0.86 are not the same bet, which is the entire numerical argument for re-deriving the threshold rather than inheriting it.

CometKiwi at 0.80 buys the cheapest invoice in this comparison and the fattest error tail; at 0.90 it rebuilds most of the full post-editing bill while still being called automation. Run one identical 10,000-word low-resource batch through the same engine at three gate settings and the trade sorts out like this:

EvidenceSourceFigureWhat it decides
Segment-level ranking powerWMT22 QE shared taskPearson r of 0.55–0.63 vs human scoresRanks well enough to route; misranks a meaningful minority
Hallucination detectionDale et al. 2023F1 around 0.75+Safety net against the worst low-resource failure mode
Full post-editing effortTAUS benchmarks35–45 min per thousand words low-resource vs 20–25 high-resourceLabor ceiling the gate exists to beat
Low-resource rate cardMarket rate cards (English-Khmer, -Nepali, -Sinhala)Current quoted per-word ratesSets the full-PE cost baseline for a 10,000-word batch
High-resource rate cardMarket rate cards (Spanish, French)Current quoted per-word ratesBenchmark for the low-resource price premium
Raw-MT critical-error baselineQTLaunchPad / QT21 MQM studiesCritical errors present in raw MT segmentsDefect load the gate must reduce, not ignore
False-accept rate at 0.85Unbabel QE-in-production write-upsA measurable share of auto-published segments with a major errorWhether the stock gate is safe for your content
The Numbers Behind the Gate — CometKiwi 0.85 Gate

Threshold Math: 0.80 vs 0.85 vs 0.90

Read the dollar column, not the percentages. The no-gate row exists because "half of full post-editing" means nothing until it sits directly beneath the baseline. What disqualifies 0.90 is the shape of the score distribution: the bands between 0.85 and 0.90 are populated mostly by competent-but-uneven translations, not hallucinations, so each notch upward pulls a large block of segments into paid review while the residual-error column barely moves. You pay more per avoided error than the error would have cost you. Hallucinations cluster far below the gate, which is why 0.85 already captures nearly all of the filtering benefit.

GateSegments sent to PEHuman cost per 10,000 wordsQE compute costExpected residual MQM-major errors per 10,000 wordsTotal cost per 10,000 words
0.80The smallest share of the three gatesRoughly a quarter of the full-PE lineDe minimis (GPU inference, not human hours)Highest of the three gates; widest false-accept exposureLowest invoice in the set, with the risk left unpriced
0.85A middle share of the three gatesUnder half of full PEDe minimisModerate; carries the false-accept band described earlier in this guideLowest defensible total
0.90The largest share of the three gatesMore than half, closing on the full-PE lineDe minimisMinimalWithin reach of the no-gate ceiling
No gate (full PE)All segmentsThe full-PE baseline figure given earlier in this guideNoneNear zero after the human passThe ceiling every gate row is judged against

The break-even condition is blunt: the gate beats full PE only while the sub-threshold share stays below the break-even line where split-workflow overhead overtakes the savings. Past that line, the split workflow pays two workflows to touch one document — routing logic, a second vendor stream, project-management coordination, and re-integration checks that scale with every handoff. The exact overhead varies with your vendor setup, but the direction does not: once most of the file routes to humans anyway, the gate is pure process tax.

Treat every cell as conditional on calibration. According to a Medium D30 ROAS analysis of acceptance thresholds in ad delivery, each 20% reduction in the threshold cost roughly 0.4–0.6x ROAS while adding two to four percentage points of churner share — a different domain exhibiting the same convexity: small downward moves look like free volume and arrive carrying unpriced tail risk. Translation gates behave identically, which is why a Khmer segment scoring 0.86 and a German segment scoring 0.86 are not the same publish decision. Change the pair, engine, or domain and every routing share, dollar figure, and residual count in the table redraws.

The winner also flips cleanly at the regulatory boundary. For medical, legal, or safety content, the optimal row stops being 0.85 and becomes 0.90 combined with mandatory human review of the passing segments — the cost column is simply not consulted, because auto-publishing gated-pass output in those domains is off the table regardless of the arithmetic.

Before committing to any row, replay it: score your calibrated MQM sample with CometKiwi, simulate the invoice at 0.80, 0.85, and 0.90, and check where your sub-threshold share lands relative to the break-even line. Simulation costs almost nothing next to a mis-set gate discovered in production.

Your situationWinning settingWhy it wins
Low-resource pair, moderate stakes, sub-threshold share under the break-even line0.85Nearly all hallucination filtering at roughly half the human line
Sub-threshold share at or beyond the break-even lineSkip the gate; full PESplit-workflow overhead erases the savings entirely
Medical, legal, or safety content0.90 + mandatory review of passing segmentsAuto-publish is prohibited; cost is irrelevant
Any change of pair, engine, or domainRe-run the full comparisonEvery cell moves; the 20%-step benchmark shows how fast the tail thickens

The weakest link in the 0.85 gate is not the threshold — it is the evidence holding it up. Nearly every published CometKiwi validation figure traces back to two places: the WMT quality estimation shared tasks and Unbabel's own model releases. Both share a structural blind spot. MQM-annotated test sets exist for a short list of language pairs, and they skew heavily toward high-resource directions translating news-style prose. If you run Khmer-to-English or Amharic-to-English, no leaderboard publishes a per-pair false-accept rate for your direction — you are extrapolating from German, Chinese, and Czech. According to the WMT evaluation setups, several lower-resource tracks lean on document-level direct assessment rather than segment-level MQM, so the metric's advertised reliability was largely measured somewhere other than where you plan to deploy it.

Threshold Math: 0.80 vs 0.85 vs 0.90 — CometKiwi 0.85 Gate

What the Data Doesn't Tell You

Three specific gaps follow. First, reported correlations are aggregates: Unbabel's model cards publish overall Pearson and Spearman figures, not worst-case pairs, so a healthy average can conceal one collapsing direction. Second, the ground truth itself is noisy — the MQM framework that Lommel and colleagues defined tolerates genuine annotator disagreement, and a threshold calibrated against noisy labels inherits that noise. Third, benchmarks freeze engines in time. Providers ship updated checkpoints continuously, and a silent swap mid-contract shifts the whole score distribution without any leaderboard warning you.

Variance across cases is where the set-it-once myth dies. A Khmer segment and a German segment carrying identical scores do not carry identical risk, for mechanical reasons. Sentence-level estimators reward fluency, so a confident hallucination in a low-resource direction can clear the gate while asserting something false. Very short strings — UI labels, headings — cluster at the extremes of the score range, where discrimination collapses. And in terminology-dense text, one swapped term barely moves an embedding-based score computed over the full sentence.

None of this reverses the routing rule; it defines the rule's maintenance schedule. The re-derivation triggers — new pair, new engine, new domain — are a floor, not a ceiling. Two edge cases justify going beyond them: rare critical-error classes such as negation and numeral errors are typically too sparse in a standard calibration sample to register, so oversample sentences containing numbers and negations deliberately; and when a provider updates a model behind the same API endpoint, treat every previously derived threshold as expired until re-checked.

The cheapest falsification test available costs almost nothing: each month, pull a small batch of the highest-scoring segments you auto-published and have them MQM-checked. Zero accumulated critical errors keeps the gate standing. A single factual hallucination or inverted numeral means the threshold was derived on the wrong distribution — lower it and re-derive before the next invoice cycle.

Trigger conditionHow the gate failsCorrective action
Engine checkpoint swapped behind the same endpointScore distribution shifts; calibrated cutoff misroutesRe-run the full MQM calibration sample on new-engine output
Register drift inside one account (marketing copy into support macros)Threshold reflects the old register, not the new oneRe-derive per subdomain; treat register change as domain change
Very short segments (UI strings, headings)Scores cluster at extremes; gate stops discriminatingGate by length bucket or force review on the shortest tier
Fluent hallucination in a low-resource directionHigh score, false content — a clean false acceptSpot-check top-scoring published segments for factual accuracy
Rare critical errors (negation, numerals)Too sparse in a standard sample to move the estimateOversample number- and negation-bearing sentences during calibration
Terminology-dense regulated textSingle-term errors invisible to a sentence-level scoreNo auto-publish, plus a glossary-consistency pass

A 0.86 in Khmer and a 0.86 in German are not the same object, and treating them as interchangeable is the most expensive assumption in the gated workflow. CometKiwi emits a regression score calibrated against the human-judgment distributions it absorbed during training — and those judgments came overwhelmingly from WMT21/22 news-domain data. Feed it a Khmer litigation contract and the entire score histogram shifts underneath the fixed cut. In practice, the identical 0.85 line can flag roughly a fifth of segments in one English–Khmer legal corpus and close to two-thirds in another, purely because document-level difficulty moved the distribution. The threshold is a coordinate in one corpus's distribution, not a portable constant.

What the Data Doesn't Tell You — CometKiwi 0.85 Gate

What the 0.85 Score Hides

The second thing the score hides is the direction of error. Fluency features dominate the learned signal, so accuracy defects routinely survive inside top-scoring segments: a negation quietly flipped ("not required" rendered as "is required"), a named entity swapped for a plausible neighbor, a dosage figure mistranslated into a neighboring magnitude. These are precisely the defects a target-language reader cannot self-detect, because the sentence reads smoothly. A fluent-wrong segment is worse than an obviously broken one — it sails past both the gate and the human skim.

Third, the gate is structurally blind to terminology consistency. CometKiwi scores each segment independently, so a segment at 0.86 can pass while contradicting the glossary term used everywhere else in the document. According to aitranslations.io, fixing this requires moving the constraint into decoding itself: adding a glossary constraint to Google's decoder removed the WMT-26 gap and yielded statistical parity with Bing's constrained pipeline, and translate-then-refine passes built on pseudo-terminology derived from word alignment help models absorb terminology constraints. Neither mechanism operates inside a post-hoc QE gate — the gate optimizes per-segment adequacy and cannot see across documents.

Fourth, the threshold is engine-bound. Moving from an open-weight model to a commercial API, or deploying a domain fine-tune, reshapes the entire score distribution; a gate tuned on one engine systematically over-routes or under-routes on another until you recalibrate. Fifth, the savings ledger counts post-editing hours avoided but not the asymmetric cost of a published error — a mistranslated dosage or consent clause can exceed an entire year's post-editing budget for the pair in remediation and liability, though the exact exposure varies enormously by jurisdiction and sector. Sixth, be honest about the evidence base: no public per-language-pair precision-recall curve exists for the 0.85 cut on live production data, and most validation runs sit on benchmark corpora whose text is cleaner and more homogeneous than real client files.

The decisive habit: treat 0.85 as a hypothesis about your corpus, never as a verdict on it. Every one of these failures is detectable with a few hundred annotated segments and an afternoon of histogram plotting — and none of them is detectable from the score alone.

Hidden failureWhat the gate reportsWhat is actually trueCheck before trusting the gate
Threshold transferStable flag rate at 0.85Flag share swings widely across corpora of the same pairRe-derive on a fresh MQM calibration sample per pair and domain
Fluent-wrong errorsSegment passesNegation, entity, and number errors survive at top scoresManually audit a sample of top-decile segments for accuracy classes
Terminology driftSegment passesTerm contradicts the document-wide glossaryRun a separate glossary-compliance pass outside the QE gate
Engine swapOld threshold still validDistribution reshaped by the new systemRecalibrate after every engine change or fine-tune deployment
Harm asymmetryHours savedPublished-error cost dwarfs editing savingsPrice the worst-case error before fixing the cut
Missing ground truthBenchmark validation looks solidNo live-production PR curve exists for your pairValidate on your own dirtiest real client sample, not benchmarks

Content for Worked Case is being prepared.

What the 0.85 Score Hides — CometKiwi 0.85 Gate

Worked Case

A publish threshold is a perishable asset. Engines get fine-tuned, content mixes rotate, and the score distribution slides underneath a gate that hasn't moved since someone copied 0.85 off a shared-task paper. The five rules below treat the gate as a measurement instrument with its own maintenance schedule — because the alternative is learning about your false accepts from customers.

Five Rules for Setting Your Own Gate

Rule 1 — Calibrate before you trust. Pull a random sample of segments from live production output — not your dev set, which oversamples clean text — and annotate them against the MQM typology with major/minor severity weighting. Keep annotators blind to the CometKiwi scores so they don't anchor. Then sweep candidate cut points across the sample, recording what fraction passes and how many major errors slip through at each. Wherever your local curve disagrees with the stock 0.85 assumption by more than 0.03, move the threshold. The breadth problem here is structural: accepted work on low-resource MT spans eight language families across 13 research areas (arXiv 2412.16365v1), so "low-resource" behaves less like one distribution than a dozen, and no single cut point survives contact with all of them.

Rule 2 — Match the gate to stakes. Auto-publication at 0.85 is defensible for web copy, marketing, and internal documentation, where a leaked awkward sentence costs embarrassment. In medical, legal, or safety text, the gate demotes to a triage tool: passing segments join the back of the human review queue, failing ones jump to the front, and nothing publishes untouched. The score ranks risk; it never certifies safety.

Rule 3 — Re-calibrate on every engine or domain change. A new MT system, a fine-tune, even a content-type shift — support tickets flowing into product documentation, say — reshuffles the score distribution enough to invalidate the previous threshold. Budget one fresh-sample annotation cycle per change, roughly 4–6 annotator hours, before resuming auto-publication. That is cheap insurance; the expensive version is a batch of unreviewed regulated text shipped on a stale gate.

Rule 4 — Treat the routing band as go/no-go. Calibration tells you more than a threshold — it tells you whether the gate deserves to exist. If too few segments would route to post-editing, skip the gate: savings won't cover queue management and reviewer-scheduling overhead. If too many would route, the gate is masking a failing engine — fix the MT system first, then revisit automation. Expect temporary excursions: a seasonal content shift can push a healthy pipeline out of band for a few weeks without meaning anything permanent.

Rule 5 — Monitor false accepts monthly. Gated-pass segments leave the pipeline and are never seen again unless you go looking, so the audit is your only feedback loop. Sample 50 published segments per language pair each month. If the major-error rate exceeds the tolerance you set at calibration, drop the threshold by 0.02–0.05 and re-measure before releasing any further batches. Drift surfaces in published text long before anyone complains — the audit just makes it visible while it is still cheap to correct.

Concrete next step: pull a random sample of segments from yesterday's production log for your highest-risk pair, book two annotator sessions against the MQM typology, and plot the curve before you touch the threshold again.

TriggerActionFigure that decides it
New engine, fine-tune, or content typeFull MQM calibration on a fresh sample before resuming auto-publication~4–6 annotator hours per change
Local curve disagrees with 0.85Move threshold to your measured operating pointDisagreement greater than 0.03
Gate routes too little to PESkip the gate entirelyRouted share too small to cover overhead
Gate routes too much to PEFix the MT system firstRouted share too large for the gate to pay
Monthly audit flags driftLower threshold, then re-measure before next batchMajor-error rate above the calibrated tolerance; drop 0.02–0.05
Medical, legal, or safety contentGate becomes prioritization only; humans read everythingAuto-publish rate: zero, any score

Concrete next step: pull a random sample of segments from yesterday's production log for your highest-risk pair, book two annotator sessions against the MQM typology, and plot the curve before you touch the threshold again.

What to do next

StepActionWhy it matters
1Score every segment with CometKiwi before anything ships, and route only segments under 0.85 into the post-editing queue — everything at or above the cut goes straight to publication.Edit spend scales linearly with the routed share, so the entire roughly-half-cost saving lives or dies on enforcing this single cut line instead of blanket post-editing.
2In regulated domains — legal, medical, financial — treat the gate as triage, not release approval: send even gated-pass segments to human review and auto-publish nothing.Unbabel's production write-ups show a 0.85 cut still leaks auto-published segments carrying MQM major errors; regulated content cannot absorb that residual, so the pass verdict must never equal a publish verdict there.
3On any change of language pair, MT engine, or domain, re-derive the threshold: MQM-annotate a fresh calibration sample, then test 0.80, 0.85, and 0.90 against it and adopt the line that minimizes escaped majors at an acceptable edit volume.QE rankers at WMT22-level correlation misorder a meaningful minority of segments, so a threshold carried over from another pair or engine imports someone else's errors — the fresh-sample re-derivation is the only defense.
4Each publishing cycle, pull a 20% random sample of auto-published gated-pass segments, MQM-review it, and log every major error back against its original CometKiwi score.The metric is weakest in the crowded middle of the score distribution — exactly where publish-or-edit decisions cluster — so sampled audits catch leak drift long before readers or regulators do.
5Spend scarce reviewer hours on the band straddling the cut line, and treat very low scores as hallucination flags: route fluent-but-unfaithful segments to bilingual reviewers, since monolingual QA cannot see source-detached output.CometKiwi-class detectors catch hallucinations reliably at the extremes (per Dale et al.'s 2023 study), so the catastrophic tail is already covered — human attention belongs where the ranking is least trustworthy.
6Before committing a new job, price the gate against full post-editing using TAUS per-thousand-word minutes and current rate cards for the actual pair — English-Khmer, English-Nepali, and English-Sinhala bill far above Spanish or French.The gate only pays where routed-share savings exceed audit overhead; running this comparison per pair tells you whether 0.85 is still earning its keep or the threshold needs re-derivation.

```

Quick answers

How well do CometKiwi-class systems correlate with human direct-assessment scores according to the WMT22 QE shared task?They reached segment-level Pearson correlations of roughly 0.55–0.63 against human direct-assessment scores.
What F1 score do CometKiwi-family QE models achieve at detecting hallucinated translations, per Dale et al.'s 2023 study?CometKiwi-family QE models detect hallucinated translations with an F1 around 0.75 or higher.
What does a 0.85 cut leave behind in auto-published segments according to Unbabel's QE-in-production write-ups?A measurable share of auto-published segments carrying at least one MQM major error.
Why does raising the gate from 0.85 to 0.90 add cost without much safety benefit?Because the bands between 0.85 and 0.90 are populated mostly by competent-but-uneven translations rather than hallucinations, so each notch upward pulls a large block of segments into paid review while the residual-error column barely moves.
Which gate setting is optimal for medical, legal, or safety content?For regulated content the optimal row stops being 0.85 and becomes 0.90 combined with mandatory human review of the passing segments.

Also worth reading: COMET's 12% Edge Over BLEU for Swahili Domain Shifts: COMET's 12% Edge Over BLEU · 2026 WMT: COMET-22's 17% Gap Switches RAG to Fine-Tuning: 2026 WMT: COMET-22's 17% Gap · Article 53 Bans BLEU, Mandates COMET-QA & Explainable Metrics: Article 53 Bans BLEU, Mandates

Research Methodology & Editorial Standards

We begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place.

Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted.

Published · Last reviewed · Owned by the Aitranslations editorial desk (About, Contact, Privacy).

Related answers