{"contract":"guth-news-publication-v1","article":{"article_id":"98be58c6-3433-4a26-bde4-198c9112d231","revision":1,"slug":"hermes-learn-trains-models-to-adapt-context-use-during-reasoning-98be58c6","title":"Hermes-Learn trains models to adapt context use during reasoning","summary":"An arXiv paper presents a framework for teaching models to decide how to allocate context windows and reuse information during test-time scaling.","body":"An arXiv paper titled “Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling” was submitted on 29 Sep 2026. It examines test-time scaling, a way to seek better model performance by giving inference more compute. The authors focus on how that compute is managed when reasoning spans multiple context windows. In that setting, models must decide when to open a fresh context and which information from earlier work should be carried forward.\n\nThe paper names this decision-making ability “contextual reasoning”. The authors say existing approaches largely leave those choices to the inference harness, while their work moves them toward the model. They introduce Hermes, a configurable family of harnesses that progressively changes how much control the model has over allocating and reusing context. The paper also presents Hermes-Learn, a two-stage framework designed to teach models these capabilities.\n\nThe authors report that models they describe as capable can use the added flexibility to benefit from more inference-time compute. Smaller open-source models, by contrast, initially struggle to make effective use of it. Their reported training approach addresses that difference: Hermes-Learn produces context-use strategies that change with the problem and with progress through the reasoning process. The abstract describes this as adaptive contextual reasoning, rather than a single fixed policy for allocating or reusing context.\n\nThe reported gains extend beyond the settings used for training. The authors say results generalize across benchmarks and models, and that models can handle more inference-time compute than they encountered during training. They also report transfer to other test-time scaling methods outside the Hermes framework. For AI builders, the paper’s central contribution is a framework for studying and training models to make context allocation and reuse decisions themselves. Its abstract reports these capabilities and generalization findings, without specifying particular model scores or implementation details.","content_kind":"author_paraphrase","explanation":{"feature":"An arXiv paper presents a framework for teaching models to decide how to allocate context windows and reuse information during test-time scaling.","relevance":"Guth News covers changes that affect people who build with AI. Read the cited primary sources for the full details.","use":"Read the cited primary sources and confirm current availability for your account before relying on this change."},"announcement_date":null,"published_at":"2026-10-01T08:07:03.647Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-01T08:07:03.385Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/98be58c6-3433-4a26-bde4-198c9112d231","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]},{"claim_id":"claim:s17","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"Hermes: Learning Contextual Reasoning Unlocks Test-Time Scaling","url":"https://arxiv.org/abs/2609.38332","fetched_at":"2026-10-01T07:01:18.339Z","sha256":"f37f5ca2f7d9d4da1f723e1ede6b13f3eab2edc86754008ef2729aa18e4b471b","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"1e3be615-a418-46b1-ac57-aa886251c26d","envelope_sha256":"d7fb82dcd20e32736a72feaf930a7bab3aac82b6ad7f29fea7422266eeb5a507"},"canonical_url":"https://news.guthlabs.ai/articles/hermes-learn-trains-models-to-adapt-context-use-during-reasoning-98be58c6"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-01T08:07:03.647Z","reviewed_at":"2026-10-01T08:07:03.385Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Hermes-Learn trains models to adapt context use during reasoning","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/hermes-learn-trains-models-to-adapt-context-use-during-reasoning-98be58c6?revision=1"}]}