Even American lab models like ChatGPT and Claude sometimes drop into Mandarin, in the answer and especially in the chain of thought. The research points to training that rewards correct output without rewarding staying in one language, made worse by outcome-based reasoning training. DeepSeek fixed R1's language mixing with an explicit reward; a 14-model study found illegible reasoning in every model but Claude; OpenAI never explained o1's switch…
This story is only covered by news sources that have yet to be evaluated by the independent media monitoring agencies we use to assess the quality and reliability of news outlets on our platform. Learn more here.