The clause read cleanly. Correct grammar, appropriate register, the kind of sentence a partner initials without slowing down. It was also wrong in a way that mattered, and nobody noticed until opposing counsel built an argument on it.
That scenario is now the dominant failure mode in AI-assisted legal translation. The errors that once made machine output easy to spot, broken word order, mangled conjugations, obviously foreign phrasing, have largely disappeared from professional-grade systems. What remains is harder to catch: text that is fluent, confident, and legally incorrect.
For lawyers, that changes where review effort belongs. Reading a translation and finding it sound is no longer evidence that it is sound. Reliable legal terminology translation now depends on a defined verification procedure rather than on how the output reads to a competent speaker of the target language.
Why fluency stopped being a quality signal
Fluency is no longer a proxy for legal accuracy. Current language models produce grammatically correct target text in nearly every case, so surface quality reveals nothing about whether a term carries the same legal effect in the target jurisdiction. The errors that remain are semantic, and semantic errors read perfectly.
Industry benchmarking puts the scale of the problem in view. Data synthesized from Intento’s State of Translation Automation 2025 and the WMT24 evaluation campaign indicates that individual top-tier large language models fabricate or distort content in roughly 10 to 18 percent of translation tasks. In general correspondence, an error rate in that range is an inconvenience. In a filing, a contract, or a sworn record, it is a liability exposure with a named partner attached to it.
The verification burden created by that rate falls on the reviewer. Forrester Research reported in 2025 that knowledge workers spend an average of 4.3 hours per week checking AI output, and that enterprises spend roughly 14,200 dollars per employee annually on hallucination mitigation. Research from Stanford HAI makes a related point: the people asked to absorb that verification load are frequently the least equipped to carry it, because catching a semantic error requires subject-matter knowledge a general reviewer does not have.
Adoption is also moving faster than the controls around it. Lokalise’s 2025 localization trends reporting recorded a 700 percent increase in AI translation use within the finance sector alone between 2023 and 2024, and found that machine-assisted translation now underpins roughly 70 percent of language workflows. Legal teams are part of that curve, usually without a corresponding change to how translated documents are reviewed before they leave the building.
Four failure modes that survive a fluency check
Errors in legal translation cluster into a small number of recognisable types. What they share is that none of them produces awkward text, which is why none of them is caught by reading the translation.

Figure 1. The four legal translation error types that produce fluent, natural-sounding output.
- Jurisdictional non-equivalence. Terms such as trust, consideration, tort and fiduciary duty can be rendered into another language without existing in the target legal system in any comparable form. A civil law rendering of a common law trust may be linguistically defensible and substantively meaningless.
- Register and formality drift. Many languages do not distinguish shall, may and must with the precision English drafting relies on. A model that collapses them into a single modal can convert a binding obligation into a recommendation without producing one awkward sentence.
- Numerical, date and reference drift. Models processing long documents have been observed altering dates, currency figures and internal cross-references. A clause that points to Section 14 in the source and Section 4 in the target is perfectly fluent and entirely unusable.
- Defined-term inconsistency. Generative systems can return different output for identical input across sessions. Across a master agreement and forty addenda, that produces a document set in which a defined term is no longer consistently defined. Nimdzi’s 2025 buyer research identifies consistency in translated content as a persistent concern tied directly to that behaviour.
The models do not agree with each other, and that is the useful signal
Independent AI models frequently produce different renderings of the same legal clause. Rather than a nuisance, that disagreement is diagnostic. The points where models diverge tend to be the points where the source text is ambiguous, jurisdiction-dependent, or terminologically loaded, which is exactly where the legal risk sits.
Internal benchmarking published by Tomedes, a professional translation company, illustrates how unevenly that divergence distributes. Running a dataset of complex multilingual legal contracts through three separate models, the team recorded distinct and unpredictable failure profiles: one model produced a 12 percent error rate handling Asian language honorifics, a second hallucinated numerical dates in Romance languages, and a third failed to hold the formal register required for German corporate filings. No single model failed everywhere. Each failed somewhere specific, and the specifics were not predictable in advance.
The same body of work tracked how error types have shifted over time. In 2020, under neural machine translation, the overwhelming majority of errors were syntactic: word order, agreement, conjugation. By 2026, surface errors had fallen close to zero and the residual errors were almost entirely semantic. That is the whole problem stated in one sentence. The category of error that survives is precisely the category a fluency check cannot detect.
The practical implication for a firm is straightforward. Accepting a single model’s output means accepting that model’s particular blind spot without knowing what it is. Running the same clause through more than one independent system and examining the disagreements turns an invisible risk into a visible, finite list of terms that need a decision.
A verification protocol lawyers can actually run
The protocol below is designed to be executed by a supervising lawyer rather than a linguist, and to leave a record. It does not require reading the target language. It requires treating translation as a controlled process with defined checkpoints.

Figure 2. A six-step verification sequence applied before a translated document leaves the firm.
Extract terminology before translating the body
Pull every defined term and jurisdiction-bearing term out of the document first. Correcting terminology after a full-document pass means correcting it in every instance across every related file, which is where consistency failures originate. Resolving the terms first means the body is translated against decisions that have already been made.
Specify the jurisdiction, not just the language
English to German is not a specification. English to German for filing with a German commercial register is. Legal equivalence is a function of the receiving legal system, and a translation cannot be validated against a target that has not been named.
Escalate the clauses that carry the exposure
Governing law, indemnity, limitation of liability, remedies and termination clauses should not be closed out on machine output alone, regardless of how confident the result appears or how closely different systems agree. These are the clauses that get litigated, and they are the ones where a named reviewer’s sign-off changes the firm’s position if the translation is later challenged.
When certified human review is not optional
Certified human review is required whenever a translated document will be filed, submitted, or relied on as evidence. Courts, immigration authorities and regulators accept translations on the basis of a qualified reviewer’s attestation, not a model’s confidence score. Where the methodology behind a translation cannot be demonstrated, admissibility can be challenged on that ground alone.
There is a documentation dimension that firms consistently underrate. In regulated matters, being able to show how a translation was produced, which systems were involved, what review was applied and who signed off is itself part of the document’s defensibility. Nimdzi’s 2025 buyer research, drawn from enterprise-side buyer conversations, found that the providers earning the highest trust in high-stakes verticals are those pairing AI-assisted workflows with structured human review, not those making the strongest automation claims.
Keeping that record costs very little at the time and is difficult to reconstruct afterwards. Firms that log the process alongside the deliverable have an answer when the question arrives. Firms that do not are left arguing about a document nobody can account for.
Frequently asked questions
Can AI translation be used for court filings?
In most jurisdictions, yes, provided the final document carries certification from a qualified human translator. AI can produce the working draft and compress turnaround, but the court accepts the translation on the reviewer’s attestation. Confirm the specific requirements of the receiving court or agency before filing.
How can a lawyer check a translation in a language they do not read?
Through procedure rather than reading. Extract the defined terms, run the clause through more than one independent system, compare where the outputs diverge, and escalate every divergent term to a qualified reviewer. Disagreement between systems identifies risk without requiring fluency in the target language.
Which clauses carry the highest translation risk?
Governing law, indemnity, limitation of liability, remedies and termination clauses, along with any clause containing a defined term that has no direct equivalent in the target legal system. These combine jurisdictional dependency with financial exposure, which is the highest-risk pairing in any translated document.
Is a fluent translation more likely to be accurate?
No. Fluency and legal accuracy are produced by different capabilities. Current models are close to perfect on grammar and register while still misrendering legally loaded terms. A translation that reads awkwardly is more likely to be flagged and fixed than one that reads well and is quietly wrong.
What should a firm ask a translation vendor?
Ask how AI output is reviewed, by whom, and at what stage, and whether that review is a defined quality gate or a spot check. Ask what terminology management is applied across a multi-document matter, and whether the certification offered is accepted by the specific court or agency the document is going to.
The warning signs are gone
The uncomfortable part of this shift is what it removed. When machine translation produced obviously broken text, review was easy, because the failures announced themselves. Fluent output does not announce anything. It arrives looking finished, which is the property that makes it dangerous in a legal context.
Verification has to be procedural now, built into how documents move through a matter rather than applied at the point where someone happens to read the target text closely. A clause that sounds right has told you nothing at all. What it means in the jurisdiction where it will be read is a separate question, and it is the only one that governs.






