{"contract":"guth-news-publication-v1","article":{"article_id":"241bafcd-5a9f-47c6-9b2c-bef477dae663","revision":1,"slug":"study-tests-what-pretraining-and-midtraining-contribute-to-reward-learning-241bafcd","title":"Study tests what pretraining and midtraining contribute to reward learning","summary":"Experiments and analysis examine how source prediction and reward adaptation interact in models learning sequential tasks and memory.","body":"A paper by Chiwun Yang and Xiaoyu Li, submitted to arXiv on 29 Sep 2026, studies how pretraining and midtraining affect a model’s ability to learn from rewards. The authors focus on a limitation of reward feedback: it can indicate that an answer is right without specifying the computation needed to answer new inputs. Their work asks what information and capabilities must be available before reward adaptation can work effectively.\n\nThe paper analyzes sequential state computation and contextual memory, where different strategies can produce the same results on training rewards but disagree on new cases. It argues that observations about the underlying source, independent of the task reward, can resolve this ambiguity. The authors also describe training paths in which models first learn source prediction and then adapt to rewards using the same parameters. In this account, source prediction develops execution or retrieval, while reward learning teaches how to use those capabilities for a particular task.\n\nExperiments using pretrained Qwen2.5 checkpoints test that division of work. Across eight worlds, sequential models given correct source and first-operation supervision achieved 82.61% success, compared with 44.15% for a control using a private-random source. In a separate memory experiment, replaying memories maintained retrieval during reward adaptation. An independent confirmation across eight worlds recorded 75.32% task success, while matched training with alternative retrieval reached 49.86%. These comparisons suggest that the information supplied before or alongside reward adaptation mattered in the tested settings.\n\nThe paper also examines GSM8K and HotpotQA, separating model accuracy when reward training begins from improvement during training and final performance. The abstract does not report benchmark scores, so it does not establish a numerical result for either dataset. For AI builders, the study highlights a distinction between rewarding correct outputs and equipping a model with the information or operations needed to produce them on unfamiliar inputs. Its results concern the tasks and experiments described in the paper, rather than establishing that the same training recipe will work across other models or applications.","content_kind":"author_paraphrase","explanation":{"feature":"Experiments and analysis examine how source prediction and reward adaptation interact in models learning sequential tasks and memory.","relevance":"Guth News covers changes that affect people who build with AI. Read the cited primary sources for the full details.","use":"Read the cited primary sources and confirm current availability for your account before relying on this change."},"announcement_date":null,"published_at":"2026-10-01T14:06:39.715Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-01T14:06:39.042Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/241bafcd-5a9f-47c6-9b2c-bef477dae663","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"What Pretraining and Midtraining Make Learnable from Rewards?","url":"https://arxiv.org/abs/2609.38446","fetched_at":"2026-10-01T13:01:27.206Z","sha256":"ea4757b18e259fafb1bd5001a84a05546a69201cc1e4868cf6049c07994d437f","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"632fe097-10cb-407f-af07-b6e940c8a66b","envelope_sha256":"9c746025c1e01e9c9fe51ba6393115d7f3be1bad78165b3428d558a35c8ccb1f"},"canonical_url":"https://news.guthlabs.ai/articles/study-tests-what-pretraining-and-midtraining-contribute-to-reward-learning-241bafcd"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-01T14:06:39.715Z","reviewed_at":"2026-10-01T14:06:39.042Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Study tests what pretraining and midtraining contribute to reward learning","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/study-tests-what-pretraining-and-midtraining-contribute-to-reward-learning-241bafcd?revision=1"}]}