TranslationBench 1.0 · English → German · September 12, 2026
The stress test for AI translation: 12 texts, 23 configurations, independently evaluated.
Ten leading AI models translated twelve challenging English texts twice: once within the ProsaBridge architecture and once directly on their own. DeepL, Google Translate, and Amazon Translate translated the identical texts. Two independent AI judges evaluated each translation against the original using MQM categories, recording every error. GPT-6 Astra inside ProsaBridge achieved the lowest error score: 43% lower than the same model on its own, while DeepL’s error score was 20× as high. ProsaBridge reduced the error score for 8 of the 10 models.
- Translations compared
- 23 · 10 AI models, each twice, plus three translation services
- Texts
- 12 · 5,479 words in total
- Judges
- Two independent AI judges, Claude Fable 5.1 and GPT-6 Astra
- Direction
- English → German
The short answers
- Better than prompting the model yourself?
- −59%
- Fewer error points on a scholarly book of 67,127 words compared with prompting the same AI model directly with the full manuscript. Across the twelve test texts, the leading model scored 43% fewer error points inside ProsaBridge than on its own.
- Better than DeepL and Google Translate?
- 20×
- DeepL’s error score on the same twelve texts was 20× higher than the best ProsaBridge entry’s; Google Translate’s was 49× higher. Amazon Translate finished last.
- Can you check this?
- Yes
- Two independent AI judges scored every translation blind using the MQM framework, and both individual scores are published for every entry. Full methodology, limits, and raw data files are available for audit.
How to read this page
An entry is one AI model used in one of two ways, or one translation service. There are 23 entries.
- With ProsaBridge
- The AI model works inside ProsaBridge. It reads the whole text first, plans the translation, and then translates with that plan.
- Model on its own
- The same AI model receives the text once and translates it in a single step.
- Translation service
- DeepL, Google Translate, or Amazon Translate, each translating the whole document through its own service.
- Error score
- Each judge assigns the errors they find to Multidimensional Quality Metrics (MQM) categories. In this benchmark, a small slip counts 1 point, a major error 5, and a critical error 10. The error score is the sum of penalty points per 100 words, averaged over both judges and all twelve texts. Lower is better.
- A score of 1.0 means roughly one small slip every 100 words. The best score here, 0.14, means roughly one small slip every 700 words.
- Accuracy
- The translation departs from the original meaning: an incorrect fact, an inaccurate term, or an omitted sentence.
- Readability
- The translation conveys the meaning but reads poorly: awkward phrasing, grammatical mistakes, or inappropriate tone.
Fig. 1
Benchmark leaderboard: measured error scores across all 23 entries
All 23 entries on one scale, best first. The bar is the error score; the scale runs to the worst score, so the distance between the AI entries and the three translation services is shown to scale. GPT-6 Astra inside ProsaBridge leads with 0.14.
| Rank | Entry | Error score | Error score |
|---|---|---|---|
| 1 | GPT-6 AstraWith ProsaBridge | 0.14 | |
| 2 | GPT-5.6 SolWith ProsaBridge | 0.25 | |
| 3 | GPT-6 AstraModel on its own | 0.25 | |
| 4 | GPT-6 AstraWith ProsaBridge | 0.25 | |
| 5 | GPT-5.6 SolModel on its own | 0.40 | |
| 6 | GPT-6 AstraModel on its own | 0.44 | |
| 7 | Fable 5.1Model on its own | 0.45 | |
| 8 | Fable 5.1With ProsaBridge | 0.48 | |
| 9 | Gemini 3.8 FlashWith ProsaBridge | 0.51 | |
| 10 | GPT-5.6 LunaWith ProsaBridge | 0.66 | |
| 11 | Opus 5With ProsaBridge | 0.70 | |
| 12 | Fable 5.1Model on its own | 0.81 | |
| 13 | Gemini 3.8 FlashModel on its own | 0.82 | |
| 14 | Fable 5.1With ProsaBridge | 0.83 | |
| 15 | GPT-5.6 LunaModel on its own | 1.00 | |
| 16 | Grok 4.6With ProsaBridge | 1.33 | |
| 17 | GLM-5.3 FlashWith ProsaBridge | 1.46 | |
| 18 | GLM-5.3 FlashModel on its own | 1.65 | |
| 19 | Opus 5Model on its own | 1.83 | |
| 20 | Grok 4.6Model on its own | 2.57 | |
| 21 | DeepLTranslation service · 20× the error points of the best entry | 2.88 | |
| 22 | GoogleTranslation service · 49× the error points of the best entry | 7.13 | |
| 23 | AmazonTranslation service · 110× the error points of the best entry | 15.93 | |
›Show all 23 entries with every number
One row per entry with its score, the accuracy and readability parts, each judge’s own score, and its cost from Fig. 3. Click a column to sort by it. Two of the AI models that took part are also the judges; those entries carry a mark. The winner is first when scored by either judge on its own.
Click a column to sort · sorted by Score
| Rank | Entry | Error score | Best value | ||||||
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra†With ProsaBridge · high/xhigh/xhigh | 0.14 | 0.06 | 0.08 | 0.27 | 0.02 | 501.9× | ◆ | |
| 2 | GPT-5.6 SolWith ProsaBridge · high/xhigh/xhigh | 0.25 | 0.07 | 0.18 | 0.34 | 0.16 | 158.9× | ◆ | |
| 3 | GPT-6 Astra†Model on its own · xhigh | 0.25 | 0.14 | 0.11 | 0.31 | 0.20 | 184.2× | ||
| 4 | GPT-6 Astra†With ProsaBridge · medium/medium/medium | 0.25 | 0.09 | 0.16 | 0.42 | 0.09 | 109.5× | ◆ | |
| 5 | GPT-5.6 SolModel on its own · xhigh | 0.40 | 0.20 | 0.21 | 0.47 | 0.34 | 84.9× | ◆ | |
| 6 | GPT-6 Astra†Model on its own · medium | 0.44 | 0.12 | 0.32 | 0.61 | 0.27 | 21.4× | ◆ | |
| 7 | Fable 5.1†Model on its own · xhigh | 0.45 | 0.18 | 0.27 | 0.33 | 0.58 | 96.6× | ||
| 8 | Fable 5.1†With ProsaBridge · high/xhigh/xhigh | 0.48 | 0.23 | 0.25 | 0.29 | 0.67 | 354.4× | ||
| 9 | Gemini 3.8 FlashWith ProsaBridge · high/xhigh/xhigh | 0.51 | 0.19 | 0.32 | 0.54 | 0.47 | 47.5× | ||
| 10 | GPT-5.6 LunaWith ProsaBridge · medium/xhigh/xhigh | 0.66 | 0.31 | 0.35 | 0.77 | 0.56 | 10.9× | ◆ | |
| 11 | Opus 5With ProsaBridge · high/xhigh/xhigh | 0.70 | 0.29 | 0.41 | 0.48 | 0.92 | 174.6× | ||
| 12 | Fable 5.1†Model on its own · medium | 0.81 | 0.28 | 0.53 | 0.77 | 0.85 | 45.1× | ||
| 13 | Gemini 3.8 FlashModel on its own · xhigh | 0.82 | 0.41 | 0.41 | 0.78 | 0.85 | 17.6× | ||
| 14 | Fable 5.1†With ProsaBridge · medium/medium/medium | 0.83 | 0.41 | 0.42 | 0.64 | 1.02 | 169.1× | ||
| 15 | GPT-5.6 LunaModel on its own · xhigh | 1.00 | 0.53 | 0.47 | 0.92 | 1.08 | 4.2× | ◆ | |
| 16 | Grok 4.6With ProsaBridge · high/xhigh/xhigh | 1.33 | 0.42 | 0.91 | 0.99 | 1.67 | 101× | ||
| 17 | GLM-5.3 FlashWith ProsaBridge · xhigh/xhigh/xhigh | 1.46 | 0.60 | 0.85 | 1.59 | 1.32 | 9.5× | ||
| 18 | GLM-5.3 FlashModel on its own · xhigh | 1.65 | 0.66 | 0.99 | 1.57 | 1.73 | 1× | ◆ | |
| 19 | Opus 5Model on its own · xhigh | 1.83 | 1.07 | 0.76 | 1.68 | 1.99 | 17.6× | ||
| 20 | Grok 4.6Model on its own · xhigh | 2.57 | 1.01 | 1.57 | 2.51 | 2.64 | 15.2× | ||
| 21 | DeepLTranslation service | 2.88 | 1.90 | 0.98 | 2.60 | 3.16 | – | ||
| 22 | GoogleTranslation service | 7.13 | 5.98 | 1.15 | 6.30 | 7.96 | – | ||
| 23 | AmazonTranslation service | 15.93 | 12.60 | 3.33 | 14.42 | 17.44 | – |
Bars share one axis that ends at 3.00. Longer bars are cut and hatched; the numbers are shown in full. ◆ Best value: no other entry is both cheaper and better. These entries form the frontier in Fig. 3. † This entry uses the same AI model as one of the judges. The published score uses both judges.
Fig. 2
Full-length book validation: 67,127 words in direct comparison
−59%
fewer error points inside ProsaBridge than prompting the same model directly: 0.08 compared with 0.20. Accuracy accounts for 98% of the advantage; readability scores are almost equal.
One book · 61 sections · 67,127 words · German → English · two judges · text private
A scholarly book of 67,127 words in 61 sections, German to English, was translated twice by the same AI model, GPT-6 Astra at a very high reasoning level. Once inside ProsaBridge, and once prompted directly with the whole book and an instruction to make it read like an original. Both judges then scored every section blind against the German original, without knowing which translation was which and without ProsaBridge’s glossary or notes.
| Entry | Error score | Score | Accuracy | Readability | Fable 5.1 medium | GPT-6 Astra medium |
|---|---|---|---|---|---|---|
| GPT-6 Astra inside ProsaBridge | 0.08 | 0.06 | 0.02 | 0.12 | 0.04 | |
| GPT-6 Astra prompted directly | 0.20 | 0.18 | 0.02 | 0.28 | 0.12 | |
AccuracyReadability
Fig. 3
Cost efficiency: compute requirements vs. quality
Fewer errors usually cost more computing. Higher is better and further right is cheaper, so the best value sits top right. Cost is shown as a multiple of the cheapest entry, GLM-5.3 Flash on its own, and the axis is logarithmic: equal distances are equal ratios. A line joins the two entries of the same model. The shaded staircase is the best-value frontier: the entries on its edge are the ones nothing beats on cost and quality at once.
- With ProsaBridge
- Model on its own
- Translation service
- Best-value frontier
Best-value frontier
- GLM-5.3 FlashModel on its own1× · 1.65
- GPT-5.6 LunaModel on its own4.2× · 1.00
- GPT-5.6 LunaWith ProsaBridge10.9× · 0.66
- GPT-6 AstraModel on its own21.4× · 0.44
- GPT-5.6 SolModel on its own84.9× · 0.40
- GPT-6 AstraWith ProsaBridge109.5× · 0.25
- GPT-5.6 SolWith ProsaBridge158.9× · 0.25
- GPT-6 AstraWith ProsaBridge501.9× · 0.14
How the score works
Two AI judges, Claude Fable 5.1 and GPT-6 Astra, read every translation next to its original. They do not know which entry produced it and they do not see each other’s work. Each judge marks every mistake with a category, a weight, and an explanation. The categories follow MQM (Multidimensional Quality Metrics), a framework used to assess translation quality.
A minor error counts 1 point, a major error 5, a critical error 10. The penalty points for a translation are summed, divided by its word count, and normalized to 100 words. An entry’s score for a text is the average of the two judges; its published score is the average over the twelve texts.
Inside ProsaBridge, a model works the same way it does on a customer’s document: it reads the whole text, prepares a translation plan, and translates with it. On its own, the same model receives only the text and the output format. Neither setup was changed for this benchmark.
›How we checked the ranking
The gaps that matter hold up: the three translation services stay at the bottom, the order from the last AI entry down to Amazon Translate is confirmed, and 10 of the 22 neighbouring pairs keep their published order. The top entries are close together, so read the first few places as a group.
After the ranking was fixed, both judges compared each entry with the one ranked directly below it, text by text, without knowing which was which. A text only counts toward one side when both judges agree.
- 1GPT-6 Astra, With ProsaBridgevs.2GPT-5.6 Sol, With ProsaBridge1 texts where both judges preferred the higher-ranked entry, 4 texts where both judges saw no difference, 7 texts where the judges disagreedToo close to call
- 2GPT-5.6 Sol, With ProsaBridgevs.3GPT-6 Astra, Model on its own2 texts where both judges saw no difference, 3 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreedToo close to call
- 3GPT-6 Astra, Model on its ownvs.4GPT-6 Astra, With ProsaBridge1 texts where both judges preferred the higher-ranked entry, 3 texts where both judges saw no difference, 8 texts where the judges disagreedToo close to call
- 4GPT-6 Astra, With ProsaBridgevs.5GPT-5.6 Sol, Model on its own5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 5 texts where the judges disagreedOrder confirmed
- 5GPT-5.6 Sol, Model on its ownvs.6GPT-6 Astra, Model on its own2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 2 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreedToo close to call
- 6GPT-6 Astra, Model on its ownvs.7Fable 5.1, Model on its own4 texts where both judges preferred the higher-ranked entry, 8 texts where the judges disagreedToo close to call
- 7Fable 5.1, Model on its ownvs.8Fable 5.1, With ProsaBridge3 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreedJudges preferred the lower-ranked entry
- 8Fable 5.1, With ProsaBridgevs.9Gemini 3.8 Flash, With ProsaBridge2 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreedToo close to call
- 9Gemini 3.8 Flash, With ProsaBridgevs.10GPT-5.6 Luna, With ProsaBridge2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 9 texts where the judges disagreedToo close to call
- 10GPT-5.6 Luna, With ProsaBridgevs.11Opus 5, With ProsaBridge1 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 10 texts where the judges disagreedToo close to call
- 11Opus 5, With ProsaBridgevs.12Fable 5.1, Model on its own6 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 4 texts where the judges disagreedOrder confirmed
- 12Fable 5.1, Model on its ownvs.13Gemini 3.8 Flash, Model on its own1 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 7 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreedJudges preferred the lower-ranked entry
- 13Gemini 3.8 Flash, Model on its ownvs.14Fable 5.1, With ProsaBridge5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreedOrder confirmed
- 14Fable 5.1, With ProsaBridgevs.15GPT-5.6 Luna, Model on its own4 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 7 texts where the judges disagreedToo close to call
- 15GPT-5.6 Luna, Model on its ownvs.16Grok 4.6, With ProsaBridge9 texts where both judges preferred the higher-ranked entry, 3 texts where the judges disagreedOrder confirmed
- 16Grok 4.6, With ProsaBridgevs.17GLM-5.3 Flash, With ProsaBridge4 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreedOrder confirmed
- 17GLM-5.3 Flash, With ProsaBridgevs.18GLM-5.3 Flash, Model on its own6 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreedOrder confirmed
- 18GLM-5.3 Flash, Model on its ownvs.19Opus 5, Model on its own5 texts where both judges preferred the higher-ranked entry, 5 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreedEven
- 19Opus 5, Model on its ownvs.20Grok 4.6, Model on its own8 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreedOrder confirmed
- 20Grok 4.6, Model on its ownvs.21DeepL, Translation service5 texts where both judges preferred the higher-ranked entry, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreedOrder confirmed
- 21DeepL, Translation servicevs.22Google, Translation service11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreedOrder confirmed
- 22Google, Translation servicevs.23Amazon, Translation service11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreedOrder confirmed
- texts where both judges preferred the higher-ranked entry
- texts where both judges saw no difference
- texts where both judges preferred the lower-ranked entry
- texts where the judges disagreed
Grok 4.6 on its own and DeepL are too close to separate. Scores that lie a few hundredths apart may swap places on another day; the download has every pair.
What this does not tell you
- Only English to German. Other language pairs will get their own results.
- Twelve texts of 300 to 500 words. ProsaBridge is built for whole books and reports, where names, terms, and style have to stay consistent across hundreds of pages. Short texts show only part of that difference: on the whole book in Fig. 2 the gap to the same model was larger than on the twelve texts.
- One test round. The numbers are not averages over repeated rounds. Treat differences of a few hundredths of a point as chance.
- The judges are AI models, and two of the models that took part are also the judges. Their separate scores are shown so you can check the effect yourself. The winner is first when scored by either judge on its own.
What we keep private
- The twelve texts. They were written for this benchmark and have never been published, so no AI model has seen them or any translation of them. Publishing them would put them in the next training data and spoil every later test round. Their titles, kinds, lengths, and difficulties are listed below.
- The instructions ProsaBridge gives the models, and the instructions the judges follow.
- What the test rounds cost. The cost chart shows only multiples of the cheapest entry.
- The judges’ notes and the translations themselves, because they would reveal the texts.
The twelve texts
Seven short stories and five nonfiction pieces, each written to be hard to translate in a particular way.
| Text | Kind | Words | What makes it hard | |
|---|---|---|---|---|
| 1 | The Lease | Fiction, dialogue | 431 | A shift from formal to informal address, dialogue, dry humor |
| 2 | Gravel | Fiction, inner voice | 485 | Thoughts in free indirect speech, colloquial asides, unmarked questions |
| 3 | The Orchard Ledger | Fiction, lyrical | 463 | One long metaphor that must hold up from start to finish |
| 4 | Southpaw | Fiction, scene | 449 | Short, hard sentences next to one long one |
| 5 | The Committee | Fiction, satire | 446 | Long nested sentences, asides, irony that depends on word order |
| 6 | Moving the Needle | Fiction, workplace | 462 | Many idioms and one pun the whole text depends on |
| 7 | The Visitor | Fiction, inner voice | 458 | A deliberately unclear “she” who must remain unclear |
| 8 | Remarks at the Dedication of the Cedar Run Flood Memorial | Nonfiction, speech | 441 | American civic tone, repetition for effect, names of institutions |
| 9 | Allision of the Ferry Marguerite Bay at Harbor Point Terminal | Nonfiction, accident report | 478 | Nautical terms, passive voice, exact times and measurements |
| 10 | The Sleep of Swifts | Nonfiction, popular science | 493 | Lyrical science writing that moves between data and imagery |
| 11 | Expanding District Heating in Mid-Sized Cities | Nonfiction, policy brief | 436 | “Should”, “must”, and “may” that have to keep their exact strength; false friends; consistent terms |
| 12 | Midterm Evaluation of the Aldercreek Adult Literacy Program | Nonfiction, evaluation report | 437 | Paired terms that must stay distinct, careful attribution, measured recommendations |
Models we removed before publishing
These models took part during development and are missing from the published results. Each removal is recorded with its evidence.
- Claude Sonnet 5
- Removed after seven texts because of weak translation quality: 3.18 and 3.22 error points per 100 words, behind every other AI model at the time.
- Kimi K2 Thinking
- Removed after it hit the provider’s output limit on three of the seven texts without delivering an answer.
- Mistral Large 3
- Removed after it added document formatting of its own and returned translated text where the process requires fixed reference codes.
- Gemini 3.1 Pro and GPT-5.6 Terra
- Withdrawn before the published test round: for what they cost, their scores were too weak. Cheaper models scored better, and better models cost about the same. Their earlier results are archived.
Download the data
The benchmark round: each entry on each text as scored by each judge, error counts by severity and category, the cost multiples, and the neighbouring-pair comparisons.
Download the 1.0 recordThe book study: both translations’ scores overall, by kind, and by judge, with the book’s size. No text from the book.
Download the book record- Test round
- translationbench-run-fa35fa02-3190-4f82-b379-8d05b7dde5d7
- Version
- 1.0 (7.2)
- Completed
- September 12, 2026