| Takeaway | Detail |
|---|---|
| Unconstrained APE drops 12.6% of supplied terminology constraints vs. fully constrained MT-to-APE pipeline | 12.6% relative drop in retained terminology constraints when applying unconstrained automatic post-editing on top of constrained machine translation |
| Unconstrained LLMs incur a 38% consistency penalty on legal definitions vs. rule-based baseline | Forensic audits show 38% penalty in legal term consistency for unconstrained large language models compared to rule-based systems |
| Constrained backtranslation improves term-specific accuracy without guaranteed BLEU gains | Constrained backtranslation yields quantifiable gains in term-specific accuracy but does not consistently outperform unconstrained backtranslation on BLEU scores |
| Verify live, complete constrained decoding options before committing to unconstrained translation | Reader rule: always verify the live, complete option before committing; compare like-for-like totals and terms |
This guide delivers a verify-before-you-commit framework for managing NLLB-200 terminology drift at scale using the 1,012-sentence FLORES-200 benchmark.
It compares constrained decoding against unconstrained translation, showing when constraint enforcement preserves critical terminology better than open-generation approaches.

How It Works
NLLB-200 is a multilingual translation model that generates a target-language sequence one token at a time. In unconstrained translation, the model selects the next token solely from its learned probabilities: each token is locally plausible, but the model may replace an approved source term with a synonym, omit it, or render it inconsistently elsewhere. FLORES-200 is the benchmark collection used to test those translations across many language directions. On this benchmark, terminology drift is therefore not a single bad sentence; it is a pattern that must be detected across the complete output set.
Constrained decoding changes that generation process by giving the decoder explicit terminology requirements. Before translation, a constraint links a source expression to its required target expression. During decoding, a constrained beam search considers candidate continuations that can satisfy the requirement while still using the model to choose among permissible alternatives. For a multi-token target term, the constraint can cover the entire phrase rather than only its first token. The model can therefore balance lexical compliance with grammatical and contextual choices instead of copying a term into an unsuitable place.
Several terms describe separate parts of this mechanism. A constraint is the required source-to-target terminology mapping. A decoder is the component that constructs the target sequence. A beam is one active candidate sequence, while beam search retains several candidates and advances the most promising ones. A token is one generated unit of text, not necessarily a complete word. Terminology adherence means that required expressions appear in the specified form; translation quality concerns the accuracy and naturalness of the surrounding sentence. A result can preserve terminology yet read poorly, so lexical compliance alone is not a complete evaluation.
Apply the check after the complete option has been generated. First, confirm that every required mapping is represented in the evaluated configuration. Then verify each required target expression in the final translations, counting repeated occurrences separately and checking that the same concept has not acquired competing target forms. Keep the benchmark scope, language directions, and scoring unit unchanged when comparing runs. Wan et al. note that automatic post-editing can improve constrained machine-translation output without guaranteeing that supplied constraints survive, which makes inspection of the final text essential rather than assuming that an earlier constraint remained effective.

Key Factors to Consider
The first decision criterion is constraint coverage: verify that the live, complete terminology option includes every required source term, its approved target form, and any applicable case or inflection rules. Do not commit based on a configuration preview, a partial glossary, or the number of terms accepted by the interface. The benchmark contains 1,012 sentences, so record whether the evaluated option covers the full set or only a subset before comparing results. A complete run should also preserve the source sentence IDs, making it possible to audit omissions and duplicate handling.
The second criterion is terminology compliance, measured at the sentence level and across the complete benchmark. Report both the number of sentences with all required terms present and the total number of required term instances satisfied; one percentage alone can hide failures. The relevant comparison is like-for-like: use the same 1,012-sentence set, the same terminology list, the same scoring rules, and the same treatment of repeated terms.
Do not transfer a terminology-preservation percentage from Wan et al.’s English-German study to NLLB-200 without checking the original experiment and reproducing its denominator. For an NLLB-200 run, count preserved term instances and compliant sentences directly on the complete evaluated set.
The third criterion is translation quality under the stated constraints, not merely successful term insertion. Check the aggregate score and inspect whether terminology gains coincide with deterioration elsewhere in the sentence. The JRC 2026 legal-NMT audit reports a 38% penalty for unconstrained LLM output against a rule-based baseline, while the terminology-constrained MT research notes that constrained backtranslation does not consistently improve BLEU, although it can improve term-specific accuracy. These findings support reporting quality and terminology measures separately: a higher compliance rate is not automatically evidence of better overall translation.
Before committing, verify the live output record one final time: confirm the benchmark total of 1,012 sentences, the complete glossary, the denominator used for every percentage, and the exact scoring configuration. The 12.6% relative terminology-constraint drop reported by Wan and colleagues is a warning to check pipeline stages, not a universal adjustment. Recalculate each reported total from the underlying counts, and reject any result whose denominator, exclusions, or term definitions cannot be reproduced. This verification makes the final decision auditable and keeps benchmark comparisons comparable.

Common Mistakes
The first common mistake is to verify only the first appearance of each required term. A source sentence such as “The contractor shall provide the inspection report to the client” may look compliant when “inspection report” appears correctly, even though a later sentence in the same batch uses a variant that the system translated as “report on the inspection.” The practical check is to inspect every occurrence in the live, complete output, including repeated references and forms changed by grammar. Record the source term, the target rendering, and the sentence location together; a successful match on one sentence does not establish that all relevant instances passed.
A second verification error is to stop after decoding and treat the result as final. A later automatic post-editing step can revise a constrained translation and remove terminology that had been enforced earlier. For example, an output containing the approved product name may be rewritten by a post-editor as a generic phrase. Re-run the terminology audit after every stage that can alter text, rather than carrying forward an earlier pass result. Wan and colleagues’ WMT study, “Incorporating Terminology Constraints in Automatic Post-Editing,” specifically reports that post-editing constrained output can still fail to preserve supplied terminology.
The subtler mistake is comparing unlike-for-like results. One evaluation may count every required term occurrence across the entire completed set, while another may count only sentences containing at least one required term; mixing those totals can make a weak option appear stronger simply because its denominator is smaller. Build the comparison from the same submitted source files, the same finalized outputs, and the same audit rules. Check both the numerator and denominator before accepting any reported result, and keep separate tallies for distinct term types and repeated occurrences.
Use a final reconciliation sheet with one row per source segment. Mark the expected term, its target form, every relevant occurrence, the stage last checked, and the evaluator’s verdict. Reconcile the row count against the completed evaluation set, then independently total passed and failed checks. If the two totals do not agree, resolve the missing or duplicate segment before comparing constrained decoding with an unconstrained alternative. This catches truncated logs, omitted segments, inconsistent term normalization, and results drawn from different outputs.
Insider Tactics
This section alone gives non-obvious strategies and timing tips. Build a term-occurrence ledger before launching the 1,012-sentence benchmark: record each required term, the sentence IDs where it appears, and the count of source occurrences that must be accounted for afterward. Freeze that ledger with the test set so that every run is evaluated against the same denominator. This is more useful than relying on the model’s built-in terminology count because it exposes omissions caused by tokenization, inflection, or preprocessing before anyone commits to a production run.
Use a staged timing plan. First, run a small smoke test containing the most difficult live terminology entries, then inspect the returned translations manually.
Next, run the entire benchmark and record translation and final-output verification separately. Recheck terminology after any automatic post-editing stage, because Wan et al.’s 2020 WMT study reports a 12.6% relative drop in supplied constraints for unconstrained APE compared with the fully constrained MT-to-APE pipeline.

Comparison
The article does not establish a like-for-like NLLB-200 result showing that constrained decoding wins on the 1,012-sentence set. Verify that both runs use the same sentences, term list, scoring rules, and denominator, then determine the winner from raw terminology counts and the separately reported quality score.
Wan et al. (WMT 2020) report 95% preservation of terminologies from lexically constrained automatic post-editing on English-German. In the same work, running unconstrained APE over constrained MT yields a 12.6% relative drop in supplied terminology constraints against a fully constrained MT-to-APE pipeline — a relative figure, not 12.6 points. The JRC 2026 legal NMT audit reports a 38% penalty for unconstrained LLMs on legal definitions against a rule-based baseline. Note that 95% preservation still leaves about 5% of supplied terms unaccounted for, which is why the count is a check rather than a guarantee.
| Dimension | Constrained decoding | Unconstrained translation |
|---|---|---|
| Source-term preservation | 95% of terminologies preserved (Wan et al., WMT 2020, English-German) | 12.6% relative drop in supplied constraints when APE runs unconstrained over constrained MT (same source) |
| Legal-definition consistency | Rule-based baseline (JRC 2026) | 38% penalty against that baseline (JRC 2026) |
| General quality (BLEU) | Quantifiable gains in term-specific accuracy; BLEU not consistently better (emergentmind) | No consistent BLEU disadvantage in that comparison (emergentmind) |
If terminology adherence is the deliverable, compare final-output term counts after auditing every pipeline stage. If general translation quality is the target, report that metric separately: the cited terminology-constrained MT source says constrained backtranslation can improve term-specific accuracy without consistently improving BLEU.
Before committing, verify the totals line up. Confirm both runs cover the same 1,012 source sentences, share one term list, and were scored by one script; then check that the unconstrained column was not built on constrained MT output, since that mismatched pairing is the setup behind the 12.6% relative drop and will flatter whichever side you place it on. Compare term hits as raw counts first, percentages second, and only then name the winner for that run.
Frequently Asked Questions
What benchmark does the guide use to evaluate terminology drift in NLLB-200 translation?
The guide uses the 1,012-sentence FLORES-200 benchmark.
What happens to terminology constraints when unconstrained automatic post-editing follows constrained machine translation?
Unconstrained automatic post-editing causes a relative drop in retained terminology constraints.
How do unconstrained large language models compare with rule-based systems for legal definitions?
Unconstrained large language models incur a consistency penalty for legal definitions compared with the rule-based baseline.
Does constrained backtranslation always improve BLEU compared with unconstrained backtranslation?
No; constrained backtranslation can improve term-specific accuracy without consistently outperforming unconstrained backtranslation on BLEU scores.
When should a team verify constrained decoding options before choosing a translation approach?
A team should verify the live, complete constrained decoding options before committing to unconstrained translation.
What comparisons should a team make when checking constrained decoding options?
The team should compare like-for-like totals and terms.
Quick answers
| What happens to supplied terminology constraints when unconstrained automatic post-editing is applied? | Unconstrained automatic post-editing drops a substantial share of supplied terminology constraints relative to the fully constrained machine-translation-to-automatic-post-editing pipeline. |
| How does unconstrained large language model performance compare with rule-based systems on legal definitions? | Unconstrained large language models incur a consistency penalty on legal definitions compared with the rule-based baseline. |
| What benefit does constrained backtranslation provide? | Constrained backtranslation improves term-specific accuracy. |
| Does constrained backtranslation consistently improve BLEU scores over unconstrained backtranslation? | No; constrained backtranslation does not consistently outperform unconstrained backtranslation on BLEU scores. |
| What should readers verify before choosing unconstrained translation? | Always verify the live, complete constrained decoding options and compare like-for-like totals and terms before committing. |
Also worth reading: JRC 2026: Legal NMT Entropy Drift & 38% Unconstrained LLM Penalty: JRC 2026: Legal NMT Entropy · 2026 Europarl Benchmark: Low-Resource Legal NMT Terminology +31%: 2026 Europarl Benchmark: Low-Resource Legal · 2024 Professional Proofreading Rates From $002 to $012 Per Word for AI-Translated Content: 2024 Professional Proofreading Rates From