Six models, one article: how this site picked its translation model
Updated 2026-10-04
A weak translation changes how the whole site reads. The same source went to several models, and the scores decided which one to keep using.
The passage is the first half of poteto’s article How I Use Cursor (《我怎么用 Cursor》), about 1,165 English words. Six models each ran at effort high: claude-opus-5-5, kimi-k3, glm-5p3, gpt-5.6-sol, grok-4.7, and gemini-3.1-pro. The site’s previous translation of the same passage made seven.
The seven texts were shuffled and labelled A to G. Scores were given without the model names. Five criteria, 5 points each, 25 in total: accuracy (no mistranslation and no omission), glossary, no added drama and the same tone, structure, and natural Chinese.
| Translation | Total |
|---|---|
| claude-opus-5-5 | 24 |
| kimi-k3 | 23.5 |
| glm-5p3 | 22.5 |
| gpt-5.6-sol | 22.5 |
| grok-4.7 | 19 |
| Previous site translation | 18 |
| gemini-3.1-pro | 17 |
claude-opus-5-5 is the main model. Each translation names it. The tool does not pick a model on its own. When its quota runs out, the next model in this score order is used.
The scores are the assistant’s judgement, not a standard answer. The assistant had already read the previous site translation, so that one could not be fully blind.
One sentence, side by side. The source is “The sticking point for me was in developing my own set of skills.”
- claude-opus-5-5: 真正让我离不开它的,是我自己攒下的一套 skills.
- kimi-k3: 最让我着迷的是打造自己的一套 skills.
- Previous site translation: 卡住我的地方是:我想自己做一套 skills.
