{"contract":"guth-news-publication-v1","article":{"article_id":"5017a186-6d8a-4612-b164-26e1d919965e","revision":1,"slug":"study-finds-reasoning-trace-endpoints-can-match-full-trace-fine-tuning-5017a186","title":"Study finds reasoning-trace endpoints can match full-trace fine-tuning","summary":"Researchers report that training on the beginnings and endings of reasoning traces can retain or improve supervised fine-tuning performance.","body":"A new study examines whether language models need to learn complete reasoning traces during post-training. The authors report that full traces offer limited benefits, while partial traces can remain effective even when heavily truncated. Their focus is on supervised fine-tuning, or SFT, using reasoning examples collected before training. The work addresses a common practice of using such trajectories to improve a model’s reasoning capability.\n\nThe paper describes these trajectories as potentially long because they can follow complex, interwoven paths and include detours before an answer. The researchers investigate whether all of that material contributes to training. They use attention-based analyses and controlled token-removal studies, and report that intermediate portions of reasoning traces are often redundant. The source summary does not specify the models, datasets or benchmarks used for those analyses.\n\nBased on this finding, the researchers propose endpoint-based SFT, or E-SFT. The method uses the beginning and ending sections of a reasoning trace rather than the entire sequence. The authors say this approach retains, and in some cases improves, SFT performance. They also report that partial traces remain effective under heavy truncation, though the supplied summary gives no numerical performance results or details about how much text was removed.\n\nThe findings concern the training signal as well as the length of the examples: the study says endpoint-based training changes reasoning behavior. Its summary further reports benefits for reinforcement learning, including GRPO, and distillation, including OPD. It does not provide specific results for either of those follow-on uses in the material available here.\n\nFor AI builders, the report makes the choice of which parts of reasoning examples to include in post-training a question worth evaluating, rather than assuming that full traces are always needed. The researchers’ account supports testing endpoint-based examples as an alternative, but the supplied summary does not establish that the approach will perform similarly across all models, tasks or training setups.","content_kind":"author_paraphrase","explanation":{"feature":"Researchers report that training on the beginnings and endings of reasoning traces can retain or improve supervised fine-tuning performance.","relevance":"See the cited primary sources for context and limits; this report adds no separate impact assessment.","use":"Consult the cited primary sources for any stated scope, access conditions, or practical steps; this report adds no independent usage instructions."},"announcement_date":null,"published_at":"2026-10-09T00:05:18.095Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-09T00:05:17.809Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/5017a186-6d8a-4612-b164-26e1d919965e","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]},{"claim_id":"claim:s17","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"arxiv:2609.07103","url":"https://huggingface.co/papers/2609.07103","fetched_at":"2026-10-09T00:01:07.604Z","sha256":"b9a8f6023136fa39e7505209eb47ee8058d6bdb123f01d4667f4d7c6071b0113","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"88201e5c-47e6-4a65-8224-2d2ddc40e56c","envelope_sha256":"ab7290c94e1a8bec0426dfd9d5f79816621bed7590e05c5103aac966946a4c93"},"canonical_url":"https://news.guthlabs.ai/articles/study-finds-reasoning-trace-endpoints-can-match-full-trace-fine-tuning-5017a186"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-09T00:05:18.095Z","reviewed_at":"2026-10-09T00:05:17.809Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Study finds reasoning-trace endpoints can match full-trace fine-tuning","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/study-finds-reasoning-trace-endpoints-can-match-full-trace-fine-tuning-5017a186?revision=1"}]}