{"contract":"guth-news-publication-v1","article":{"article_id":"12d2a404-6b2a-4fa9-92b8-701e4f7f0005","revision":1,"slug":"study-finds-post-training-can-weaken-llm-grammar-across-101-languages-12d2a404","title":"Study finds post-training can weaken LLM grammar across 101 languages","summary":"An evaluation of six model families finds uneven effects from post-training, with low-resource languages facing the largest costs.","body":"Zhanyu Chen and Jaap Jumelet’s paper examines how post-training and language-resource levels relate to grammatical ability in multilingual language models. The authors submitted it to arXiv on 30 Sep 2026, and the paper is listed as an EMNLP Main 2026 paper. They compare base and post-trained models from six families using four evaluation methods. The tests use MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages.\n\nThe central finding is that post-training weakens grammatical competence, but the size of the decline varies unevenly with model scale. The authors report that low-resource languages bear the greatest cost. Their results therefore do not describe one uniform effect across all the models and languages they evaluated. Instead, the study treats model scale and language resource availability as factors that shape the reported outcome.\n\nThe paper also distinguishes grammatical knowledge from a model’s ability to express that knowledge through explicit prompting. The authors found that post-trained models could retain knowledge that direct prompting did not reveal. They could measure this gap in high-resource languages, but not in the same way in low-resource settings. There, baseline performance was close to chance, leaving little observable room for hidden knowledge to explain. This means an evaluation result can reflect both what a model knows and what a particular test can elicit.\n\nFor low-resource languages, the study found that prompting in the language being tested can recover otherwise hidden competence. The authors say direct probing through unprompted probabilities works only for high-resource languages. They conclude that multilingual grammar evaluations should account for language differences and use multiple evaluation approaches to avoid understating low-resource abilities. The study does not claim that one prompting method or benchmark resolves every evaluation challenge.\n\nFor AI builders, the findings mean that relying on a single evaluation approach may give an incomplete picture of grammar across languages, particularly those with fewer resources. A practical implication is to include language-informed, varied tests when assessing multilingual models, rather than treating one measurement as definitive. The paper’s results also make it important to distinguish a model’s underlying grammatical competence from what it can demonstrate under a given prompt or test.","content_kind":"author_paraphrase","explanation":{"feature":"An evaluation of six model families finds uneven effects from post-training, with low-resource languages facing the largest costs.","relevance":"For AI builders, the findings mean that relying on a single evaluation approach may give an incomplete picture of grammar across languages, particularly those with fewer resources.","use":"The authors say direct probing through unprompted probabilities works only for high-resource languages."},"announcement_date":null,"published_at":"2026-10-02T22:05:09.846Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-02T22:05:09.290Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/12d2a404-6b2a-4fa9-92b8-701e4f7f0005","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]},{"claim_id":"claim:s17","evidence_refs":["source:1"]},{"claim_id":"claim:s18","evidence_refs":["source:1"]},{"claim_id":"claim:s19","evidence_refs":["source:1"]},{"claim_id":"claim:s20","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models","url":"https://arxiv.org/abs/2610.00540","fetched_at":"2026-10-02T21:01:27.539Z","sha256":"5cf47c5b831ab667aef60811d96a3ae7cea0f713748f3340de45db899a1669e0","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"05297ae1-217e-49db-b3f9-3cffe75ae71a","envelope_sha256":"4606734056a1ea6c028e37e3cc039b617e8ff43663ffecbbc523817f04a24d60"},"canonical_url":"https://news.guthlabs.ai/articles/study-finds-post-training-can-weaken-llm-grammar-across-101-languages-12d2a404"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-02T22:05:09.846Z","reviewed_at":"2026-10-02T22:05:09.290Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Study finds post-training can weaken LLM grammar across 101 languages","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/study-finds-post-training-can-weaken-llm-grammar-across-101-languages-12d2a404?revision=1"}]}