Back to home

TranslationBench 1.0 · English → German · September 12, 2026

The stress test for AI translation: 12 texts, 23 configurations, independently evaluated.

Ten leading AI models translated twelve challenging English texts twice: once within the ProsaBridge architecture and once directly on their own. DeepL, Google Translate, and Amazon Translate translated the identical texts. Two independent AI judges evaluated each translation against the original using MQM categories, recording every error. GPT-6 Astra inside ProsaBridge achieved the lowest error score: 43% lower than the same model on its own, while DeepL’s error score was 20× as high. ProsaBridge reduced the error score for 8 of the 10 models.

Translations compared
23 · 10 AI models, each twice, plus three translation services
Texts
12 · 5,479 words in total
Judges
Two independent AI judges, Claude Fable 5.1 and GPT-6 Astra
Direction
English → German

The short answers

Better than prompting the model yourself?
−59%
Fewer error points on a scholarly book of 67,127 words compared with prompting the same AI model directly with the full manuscript. Across the twelve test texts, the leading model scored 43% fewer error points inside ProsaBridge than on its own.
Better than DeepL and Google Translate?
20×
DeepL’s error score on the same twelve texts was 20× higher than the best ProsaBridge entry’s; Google Translate’s was 49× higher. Amazon Translate finished last.
Can you check this?
Yes
Two independent AI judges scored every translation blind using the MQM framework, and both individual scores are published for every entry. Full methodology, limits, and raw data files are available for audit.

How to read this page

An entry is one AI model used in one of two ways, or one translation service. There are 23 entries.

With ProsaBridge
The AI model works inside ProsaBridge. It reads the whole text first, plans the translation, and then translates with that plan.
Model on its own
The same AI model receives the text once and translates it in a single step.
Translation service
DeepL, Google Translate, or Amazon Translate, each translating the whole document through its own service.
Error score
Each judge assigns the errors they find to Multidimensional Quality Metrics (MQM) categories. In this benchmark, a small slip counts 1 point, a major error 5, and a critical error 10. The error score is the sum of penalty points per 100 words, averaged over both judges and all twelve texts. Lower is better.
A score of 1.0 means roughly one small slip every 100 words. The best score here, 0.14, means roughly one small slip every 700 words.
Accuracy
The translation departs from the original meaning: an incorrect fact, an inaccurate term, or an omitted sentence.
Readability
The translation conveys the meaning but reads poorly: awkward phrasing, grammatical mistakes, or inappropriate tone.

Fig. 1

Benchmark leaderboard: measured error scores across all 23 entries

All 23 entries on one scale, best first. The bar is the error score; the scale runs to the worst score, so the distance between the AI entries and the three translation services is shown to scale. GPT-6 Astra inside ProsaBridge leads with 0.14.

RankEntryError scoreError score
1GPT-6 AstraWith ProsaBridge0.14
2GPT-5.6 SolWith ProsaBridge0.25
3GPT-6 AstraModel on its own0.25
4GPT-6 AstraWith ProsaBridge0.25
5GPT-5.6 SolModel on its own0.40
6GPT-6 AstraModel on its own0.44
7Fable 5.1Model on its own0.45
8Fable 5.1With ProsaBridge0.48
9Gemini 3.8 FlashWith ProsaBridge0.51
10GPT-5.6 LunaWith ProsaBridge0.66
11Opus 5With ProsaBridge0.70
12Fable 5.1Model on its own0.81
13Gemini 3.8 FlashModel on its own0.82
14Fable 5.1With ProsaBridge0.83
15GPT-5.6 LunaModel on its own1.00
16Grok 4.6With ProsaBridge1.33
17GLM-5.3 FlashWith ProsaBridge1.46
18GLM-5.3 FlashModel on its own1.65
19Opus 5Model on its own1.83
20Grok 4.6Model on its own2.57
21DeepLTranslation service · 20× the error points of the best entry2.88
22GoogleTranslation service · 49× the error points of the best entry7.13
23AmazonTranslation service · 110× the error points of the best entry15.93
Show all 23 entries with every number

One row per entry with its score, the accuracy and readability parts, each judge’s own score, and its cost from Fig. 3. Click a column to sort by it. Two of the AI models that took part are also the judges; those entries carry a mark. The winner is first when scored by either judge on its own.

Click a column to sort · sorted by Score

RankEntryError scoreBest value
1GPT-6 AstraWith ProsaBridge · high/xhigh/xhigh0.140.060.080.270.02501.9×
2GPT-5.6 SolWith ProsaBridge · high/xhigh/xhigh0.250.070.180.340.16158.9×
3GPT-6 AstraModel on its own · xhigh0.250.140.110.310.20184.2×
4GPT-6 AstraWith ProsaBridge · medium/medium/medium0.250.090.160.420.09109.5×
5GPT-5.6 SolModel on its own · xhigh0.400.200.210.470.3484.9×
6GPT-6 AstraModel on its own · medium0.440.120.320.610.2721.4×
7Fable 5.1Model on its own · xhigh0.450.180.270.330.5896.6×
8Fable 5.1With ProsaBridge · high/xhigh/xhigh0.480.230.250.290.67354.4×
9Gemini 3.8 FlashWith ProsaBridge · high/xhigh/xhigh0.510.190.320.540.4747.5×
10GPT-5.6 LunaWith ProsaBridge · medium/xhigh/xhigh0.660.310.350.770.5610.9×
11Opus 5With ProsaBridge · high/xhigh/xhigh0.700.290.410.480.92174.6×
12Fable 5.1Model on its own · medium0.810.280.530.770.8545.1×
13Gemini 3.8 FlashModel on its own · xhigh0.820.410.410.780.8517.6×
14Fable 5.1With ProsaBridge · medium/medium/medium0.830.410.420.641.02169.1×
15GPT-5.6 LunaModel on its own · xhigh1.000.530.470.921.084.2×
16Grok 4.6With ProsaBridge · high/xhigh/xhigh1.330.420.910.991.67101×
17GLM-5.3 FlashWith ProsaBridge · xhigh/xhigh/xhigh1.460.600.851.591.329.5×
18GLM-5.3 FlashModel on its own · xhigh1.650.660.991.571.73
19Opus 5Model on its own · xhigh1.831.070.761.681.9917.6×
20Grok 4.6Model on its own · xhigh2.571.011.572.512.6415.2×
21DeepLTranslation service2.881.900.982.603.16
22GoogleTranslation service7.135.981.156.307.96
23AmazonTranslation service15.9312.603.3314.4217.44

Bars share one axis that ends at 3.00. Longer bars are cut and hatched; the numbers are shown in full. Best value: no other entry is both cheaper and better. These entries form the frontier in Fig. 3. † This entry uses the same AI model as one of the judges. The published score uses both judges.

The AI entries lie between 0.14 and 2.57; the services start at 2.88 and end at 15.93. Every AI model was entered twice, inside ProsaBridge and on its own; the full table below lists both entries with every judge’s score and the cost.

Fig. 2

Full-length book validation: 67,127 words in direct comparison

59%

fewer error points inside ProsaBridge than prompting the same model directly: 0.08 compared with 0.20. Accuracy accounts for 98% of the advantage; readability scores are almost equal.

One book · 61 sections · 67,127 words · German → English · two judges · text private

A scholarly book of 67,127 words in 61 sections, German to English, was translated twice by the same AI model, GPT-6 Astra at a very high reasoning level. Once inside ProsaBridge, and once prompted directly with the whole book and an instruction to make it read like an original. Both judges then scored every section blind against the German original, without knowing which translation was which and without ProsaBridge’s glossary or notes.

EntryError scoreScoreAccuracyReadabilityFable 5.1 mediumGPT-6 Astra medium
GPT-6 Astra inside ProsaBridge0.080.060.020.120.04
GPT-6 Astra prompted directly0.200.180.020.280.12

AccuracyReadability

One book, one translation per setting. The direct translation did not have the glossary ProsaBridge builds, so the comparison shows the effect of the whole preparation. It does not isolate one step. The book is a customer’s and stays private; scores and section counts are in the download.

Fig. 3

Cost efficiency: compute requirements vs. quality

Fewer errors usually cost more computing. Higher is better and further right is cheaper, so the best value sits top right. Cost is shown as a multiple of the cheapest entry, GLM-5.3 Flash on its own, and the axis is logarithmic: equal distances are equal ratios. A line joins the two entries of the same model. The shaded staircase is the best-value frontier: the entries on its edge are the ones nothing beats on cost and quality at once.

Best-value frontier

  1. GLM-5.3 FlashModel on its own · 1.65
  2. GPT-5.6 LunaModel on its own4.2× · 1.00
  3. GPT-5.6 LunaWith ProsaBridge10.9× · 0.66
  4. GPT-6 AstraModel on its own21.4× · 0.44
  5. GPT-5.6 SolModel on its own84.9× · 0.40
  6. GPT-6 AstraWith ProsaBridge109.5× · 0.25
  7. GPT-5.6 SolWith ProsaBridge158.9× · 0.25
  8. GPT-6 AstraWith ProsaBridge501.9× · 0.14
Cost is the computing each entry used for these twelve texts, at the model providers’ list prices. It is not the price of ProsaBridge. Hover over a mark to see its entry and numbers. The translation services charge per character, so DeepL appears as a dashed line at its score; Google Translate and Amazon Translate score far below the chart and are named at its bottom edge. The full table under Fig. 1 lists every cost.

How the score works

Two AI judges, Claude Fable 5.1 and GPT-6 Astra, read every translation next to its original. They do not know which entry produced it and they do not see each other’s work. Each judge marks every mistake with a category, a weight, and an explanation. The categories follow MQM (Multidimensional Quality Metrics), a framework used to assess translation quality.

A minor error counts 1 point, a major error 5, a critical error 10. The penalty points for a translation are summed, divided by its word count, and normalized to 100 words. An entry’s score for a text is the average of the two judges; its published score is the average over the twelve texts.

Inside ProsaBridge, a model works the same way it does on a customer’s document: it reads the whole text, prepares a translation plan, and translates with it. On its own, the same model receives only the text and the output format. Neither setup was changed for this benchmark.

How we checked the ranking

The gaps that matter hold up: the three translation services stay at the bottom, the order from the last AI entry down to Amazon Translate is confirmed, and 10 of the 22 neighbouring pairs keep their published order. The top entries are close together, so read the first few places as a group.

After the ranking was fixed, both judges compared each entry with the one ranked directly below it, text by text, without knowing which was which. A text only counts toward one side when both judges agree.

  1. 1GPT-6 Astra, With ProsaBridgevs.2GPT-5.6 Sol, With ProsaBridge
    1 texts where both judges preferred the higher-ranked entry, 4 texts where both judges saw no difference, 7 texts where the judges disagreed
    Too close to call
  2. 2GPT-5.6 Sol, With ProsaBridgevs.3GPT-6 Astra, Model on its own
    2 texts where both judges saw no difference, 3 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreed
    Too close to call
  3. 3GPT-6 Astra, Model on its ownvs.4GPT-6 Astra, With ProsaBridge
    1 texts where both judges preferred the higher-ranked entry, 3 texts where both judges saw no difference, 8 texts where the judges disagreed
    Too close to call
  4. 4GPT-6 Astra, With ProsaBridgevs.5GPT-5.6 Sol, Model on its own
    5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 5 texts where the judges disagreed
    Order confirmed
  5. 5GPT-5.6 Sol, Model on its ownvs.6GPT-6 Astra, Model on its own
    2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 2 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreed
    Too close to call
  6. 6GPT-6 Astra, Model on its ownvs.7Fable 5.1, Model on its own
    4 texts where both judges preferred the higher-ranked entry, 8 texts where the judges disagreed
    Too close to call
  7. 7Fable 5.1, Model on its ownvs.8Fable 5.1, With ProsaBridge
    3 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreed
    Judges preferred the lower-ranked entry
  8. 8Fable 5.1, With ProsaBridgevs.9Gemini 3.8 Flash, With ProsaBridge
    2 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 7 texts where the judges disagreed
    Too close to call
  9. 9Gemini 3.8 Flash, With ProsaBridgevs.10GPT-5.6 Luna, With ProsaBridge
    2 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 9 texts where the judges disagreed
    Too close to call
  10. 10GPT-5.6 Luna, With ProsaBridgevs.11Opus 5, With ProsaBridge
    1 texts where both judges preferred the higher-ranked entry, 1 texts where both judges preferred the lower-ranked entry, 10 texts where the judges disagreed
    Too close to call
  11. 11Opus 5, With ProsaBridgevs.12Fable 5.1, Model on its own
    6 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 1 texts where both judges preferred the lower-ranked entry, 4 texts where the judges disagreed
    Order confirmed
  12. 12Fable 5.1, Model on its ownvs.13Gemini 3.8 Flash, Model on its own
    1 texts where both judges preferred the higher-ranked entry, 2 texts where both judges saw no difference, 7 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreed
    Judges preferred the lower-ranked entry
  13. 13Gemini 3.8 Flash, Model on its ownvs.14Fable 5.1, With ProsaBridge
    5 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreed
    Order confirmed
  14. 14Fable 5.1, With ProsaBridgevs.15GPT-5.6 Luna, Model on its own
    4 texts where both judges preferred the higher-ranked entry, 1 texts where both judges saw no difference, 7 texts where the judges disagreed
    Too close to call
  15. 15GPT-5.6 Luna, Model on its ownvs.16Grok 4.6, With ProsaBridge
    9 texts where both judges preferred the higher-ranked entry, 3 texts where the judges disagreed
    Order confirmed
  16. 16Grok 4.6, With ProsaBridgevs.17GLM-5.3 Flash, With ProsaBridge
    4 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 5 texts where the judges disagreed
    Order confirmed
  17. 17GLM-5.3 Flash, With ProsaBridgevs.18GLM-5.3 Flash, Model on its own
    6 texts where both judges preferred the higher-ranked entry, 3 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreed
    Order confirmed
  18. 18GLM-5.3 Flash, Model on its ownvs.19Opus 5, Model on its own
    5 texts where both judges preferred the higher-ranked entry, 5 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreed
    Even
  19. 19Opus 5, Model on its ownvs.20Grok 4.6, Model on its own
    8 texts where both judges preferred the higher-ranked entry, 2 texts where both judges preferred the lower-ranked entry, 2 texts where the judges disagreed
    Order confirmed
  20. 20Grok 4.6, Model on its ownvs.21DeepL, Translation service
    5 texts where both judges preferred the higher-ranked entry, 4 texts where both judges preferred the lower-ranked entry, 3 texts where the judges disagreed
    Order confirmed
  21. 21DeepL, Translation servicevs.22Google, Translation service
    11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreed
    Order confirmed
  22. 22Google, Translation servicevs.23Amazon, Translation service
    11 texts where both judges preferred the higher-ranked entry, 1 texts where the judges disagreed
    Order confirmed
  • texts where both judges preferred the higher-ranked entry
  • texts where both judges saw no difference
  • texts where both judges preferred the lower-ranked entry
  • texts where the judges disagreed

Grok 4.6 on its own and DeepL are too close to separate. Scores that lie a few hundredths apart may swap places on another day; the download has every pair.

What this does not tell you

  • Only English to German. Other language pairs will get their own results.
  • Twelve texts of 300 to 500 words. ProsaBridge is built for whole books and reports, where names, terms, and style have to stay consistent across hundreds of pages. Short texts show only part of that difference: on the whole book in Fig. 2 the gap to the same model was larger than on the twelve texts.
  • One test round. The numbers are not averages over repeated rounds. Treat differences of a few hundredths of a point as chance.
  • The judges are AI models, and two of the models that took part are also the judges. Their separate scores are shown so you can check the effect yourself. The winner is first when scored by either judge on its own.

What we keep private

  • The twelve texts. They were written for this benchmark and have never been published, so no AI model has seen them or any translation of them. Publishing them would put them in the next training data and spoil every later test round. Their titles, kinds, lengths, and difficulties are listed below.
  • The instructions ProsaBridge gives the models, and the instructions the judges follow.
  • What the test rounds cost. The cost chart shows only multiples of the cheapest entry.
  • The judges’ notes and the translations themselves, because they would reveal the texts.

The twelve texts

Seven short stories and five nonfiction pieces, each written to be hard to translate in a particular way.

TextKindWordsWhat makes it hard
1The LeaseFiction, dialogue431A shift from formal to informal address, dialogue, dry humor
2GravelFiction, inner voice485Thoughts in free indirect speech, colloquial asides, unmarked questions
3The Orchard LedgerFiction, lyrical463One long metaphor that must hold up from start to finish
4SouthpawFiction, scene449Short, hard sentences next to one long one
5The CommitteeFiction, satire446Long nested sentences, asides, irony that depends on word order
6Moving the NeedleFiction, workplace462Many idioms and one pun the whole text depends on
7The VisitorFiction, inner voice458A deliberately unclear “she” who must remain unclear
8Remarks at the Dedication of the Cedar Run Flood MemorialNonfiction, speech441American civic tone, repetition for effect, names of institutions
9Allision of the Ferry Marguerite Bay at Harbor Point TerminalNonfiction, accident report478Nautical terms, passive voice, exact times and measurements
10The Sleep of SwiftsNonfiction, popular science493Lyrical science writing that moves between data and imagery
11Expanding District Heating in Mid-Sized CitiesNonfiction, policy brief436“Should”, “must”, and “may” that have to keep their exact strength; false friends; consistent terms
12Midterm Evaluation of the Aldercreek Adult Literacy ProgramNonfiction, evaluation report437Paired terms that must stay distinct, careful attribution, measured recommendations

Models we removed before publishing

These models took part during development and are missing from the published results. Each removal is recorded with its evidence.

Claude Sonnet 5
Removed after seven texts because of weak translation quality: 3.18 and 3.22 error points per 100 words, behind every other AI model at the time.
Kimi K2 Thinking
Removed after it hit the provider’s output limit on three of the seven texts without delivering an answer.
Mistral Large 3
Removed after it added document formatting of its own and returned translated text where the process requires fixed reference codes.
Gemini 3.1 Pro and GPT-5.6 Terra
Withdrawn before the published test round: for what they cost, their scores were too weak. Cheaper models scored better, and better models cost about the same. Their earlier results are archived.

Download the data

The benchmark round: each entry on each text as scored by each judge, error counts by severity and category, the cost multiples, and the neighbouring-pair comparisons.

Download the 1.0 record

The book study: both translations’ scores overall, by kind, and by judge, with the book’s size. No text from the book.

Download the book record
Test round
translationbench-run-fa35fa02-3190-4f82-b379-8d05b7dde5d7
Version
1.0 (7.2)
Completed
September 12, 2026