The Mandarin Slip: Why ChatGPT, Claude and DeepSeek Sometimes Answer in Chinese

Posted

By Evan Vega

If you work long enough with AI, you’ll see output in some strange languages. Even American lab models like ChatGPT and Claude drop into Mandarin, with Chinese characters in the answer and especially in the chain of thought. Novel Cognition went through the research, the postmortems and years of community reports to find out why.

The best-supported answer isn’t that models secretly think in Mandarin. It’s that training rewards correct, useful output and, unless it separately rewards staying in one language, nothing keeps a long answer or a reasoning scratchpad in the language it started in. Mandarin is the switch people notice. It isn’t the only one.

The full walkthrough — why training doesn’t penalize language mixing, the reward DeepSeek added to fix R1, the 14-model study that found illegible reasoning everywhere except Claude, OpenAI o1’s unexplained switches, and the Claude bug Anthropic disclosed. Seven minutes.

Even American lab models like ChatGPT and Claude sometimes drop into Mandarin, in the answer and especially in the chain of thought. The research points to training that rewards correct output without

o1, GPT-5 and Claude do it too

OpenAI never published a cause for o1’s switches into Chinese and Persian; the theory blaming Chinese data-labeling vendors was never demonstrated. Later work found similar degraded reasoning text in o3 and GPT-5, but OpenAI doesn’t expose raw chains of thought. Claude is the one American case with a disclosed cause: Anthropic’s 17 September 2025 postmortem says a TPU server misconfiguration occasionally produced “Thai or Chinese characters in response to English prompts,” affecting Opus 4 and 4.1 from 25 to 28 August and Sonnet 4 until 2 September 2025.

That case is a reminder to check the plumbing first. DeepSeek-Coder-V2’s Chinese output on Ollama traced to a broken chat template; Qwen2.5-VL’s came from padding under batching. The full breakdown is at o1, GPT-5 and Claude Do It Too and Check the Template Before You Blame the Model.

Right answer, wrong language

“Answer this correctly” and “do every step in English” are different objectives, and most training only rewards the first. The Language Confusion Benchmark caught models switching languages across 15 languages, at the level of whole responses, lines and single words; harder prompts and higher temperature made it worse. An ACL 2025 study found ordinary training doesn’t reliably penalize mixed-language text, and that preference tuning which explicitly rejected it fixed much of the problem.

DeepSeek gave the cleanest evidence. DeepSeek-R1-Zero, trained with reinforcement learning on answers alone, began “combining English and Chinese within a single chain-of-thought response.” For R1, DeepSeek added a reward for the share of reasoning written in the target language. Without it, consistency deteriorated during training; with it, consistency held, at a slight cost in reasoning performance. The detail is at Nothing Tells It to Stay in English and The Reward That Fixed DeepSeek-R1.

The scratchpad slips first

A study of 14 reasoning models by Arun Jose (arXiv 2510.27338) found that outcome-trained models such as DeepSeek-R1 and QwQ drift into compressed fragments, unrelated words and non-English characters mid-reasoning, then return to perfectly readable final answers. The abstract names one exception: Claude. Forced to use only the legible portions of their reasoning, the models’ accuracy fell 53%, but the author found no general link between illegibility and correctness and considered a secret language unlikely.

Mandarin dominates the screenshots mostly because it’s visible. OpenAI o1 users also saw its reasoning slip into Persian, Hindi and Thai, and the popular token-efficiency theory isn’t supported by tokenizer research. See Why the Scratchpad Slips First and Why Mandarin, of All Languages.

Trainable, not mysterious

Frontier models are multilingual text generators without sealed language modes. Standard training doesn’t reliably penalize accidental language mixing, outcome-based reasoning training makes it worse, and explicit language rewards, in-language examples, preference tuning and careful decoding reduce it. For American models the behavior is documented and, apart from Claude’s disclosed bug, largely undiagnosed. Even the visible chain of thought isn’t a faithful transcript: Anthropic found Claude 3.7 Sonnet mentioned answer-changing hints 25% of the time and DeepSeek-R1 39%.

The full file is at mandarinslip.novcog.us.com. Primary sources: the DeepSeek-R1 paper, Reasoning Models Sometimes Output Illegible Chains of Thought, the Language Confusion Benchmark, and Anthropic’s postmortem.


Part of the Frontier Watch Series: Read the previous investigation

.


More Coverage:
→ Read this investigation on North Denver Tribune
→ Coverage from Daily Colorado News

The post The Mandarin Slip: Why ChatGPT, Claude and DeepSeek Sometimes Answer in Chinese first appeared on DAILY TEXAS NEWS.
News, Anthropic, chain of thought, ChatGPT, Claude, DeepSeek-R1, language confusion, language mixing, OpenAI, OpenAI o1, Qwen, reasoning models, reinforcement learning