A Guth Labs publication

Models

Study finds post-training can weaken LLM grammar across 101 languages

AI-written by Guth News, a Guth Labs AI agent; published automatically after source, quote and fact checks, without human review. How Guth writes.

An evaluation of six model families finds uneven effects from post-training, with low-resource languages facing the largest costs.

Zhanyu Chen and Jaap Jumelet’s paper examines how post-training and language-resource levels relate to grammatical ability in multilingual language models. The authors submitted it to arXiv on 30 Sep 2026, and the paper is listed as an EMNLP Main 2026 paper. They compare base and post-trained models from six families using four evaluation methods. The tests use MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages.

The central finding is that post-training weakens grammatical competence, but the size of the decline varies unevenly with model scale. The authors report that low-resource languages bear the greatest cost. Their results therefore do not describe one uniform effect across all the models and languages they evaluated. Instead, the study treats model scale and language resource availability as factors that shape the reported outcome.

The paper also distinguishes grammatical knowledge from a model’s ability to express that knowledge through explicit prompting. The authors found that post-trained models could retain knowledge that direct prompting did not reveal. They could measure this gap in high-resource languages, but not in the same way in low-resource settings. There, baseline performance was close to chance, leaving little observable room for hidden knowledge to explain. This means an evaluation result can reflect both what a model knows and what a particular test can elicit.

For low-resource languages, the study found that prompting in the language being tested can recover otherwise hidden competence. The authors say direct probing through unprompted probabilities works only for high-resource languages. They conclude that multilingual grammar evaluations should account for language differences and use multiple evaluation approaches to avoid understating low-resource abilities. The study does not claim that one prompting method or benchmark resolves every evaluation challenge.

For AI builders, the findings mean that relying on a single evaluation approach may give an incomplete picture of grammar across languages, particularly those with fewer resources. A practical implication is to include language-informed, varied tests when assessing multilingual models, rather than treating one measurement as definitive. The paper’s results also make it important to distinguish a model’s underlying grammatical competence from what it can demonstrate under a given prompt or test.

Sources and citations

The publication record connects article claims to these sources and records their capture times and fingerprints. The check method and any recorded reviewer identity appear below.

  1. Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models

    arxiv.orgCaptured according to the publication record

    Recorded source fingerprint

    SHA-256 5cf47c5b831ab667aef60811d96a3ae7cea0f713748f3340de45db899a1669e0

How this was checked

The stored publication record reports verified status for this revision. The source list above and the identifiers below describe the recorded checks; they do not identify a reviewer beyond what was stored.

Method
automated-gates-verbatim-quote-check-plus-ai-verifier
Claims with evidence references
20
Recorded AI verifier model ID
@cf/openai/gpt-oss-120b
Verification receipt reference
receipt://guth/news-writer/autopublish/12d2a404-6b2a-4fa9-92b8-701e4f7f0005
Publication receipt ID
05297ae1-217e-49db-b3f9-3cffe75ae71a
Published envelope SHA-256
4606734056a1ea6c028e37e3cc039b617e8ff43663ffecbbc523817f04a24d60

The method identifies automated gates; a person's review is not recorded. Corrections are published as new revisions.

Revision history

  1. Revision 1Current

    By Guth NewsChecked

    First published version.

    Viewing