{"contract":"guth-news-publication-v1","article":{"article_id":"52ccbec9-b2f1-4f6d-8e13-d37e652a7672","revision":1,"slug":"researchers-propose-a-framework-to-measure-how-ai-agents-learn-from-experience-52ccbec9","title":"Researchers propose a framework to measure how AI agents learn from experience","summary":"A new evaluation approach tracks whether agents improve on later tasks, how efficiently they learn, and whether gains generalize beyond their training interactions.","body":"A new paper proposes a way to evaluate whether AI agents improve through experience, rather than assessing only what they can do at a single point in time. Its authors argue that agents capable of diagnosing failures and learning from them need evaluations that capture changes in performance. They frame the assessment around three questions: whether later performance improves and generalizes, how efficiently capabilities are acquired, and where the improvement process breaks down. The paper calls the efficiency of converting experience into future gains on held-out tasks “agent plasticity”.\n\nThe researchers study this process in a controlled setting where agents turn prior experience into reusable artifacts that future instances can inherit. They measure agent performance at checkpoints on both training interactions and held-out interactions, while also accounting for the cost of learning. This design is intended to distinguish gains on experiences used for learning from performance on interactions that were held back. The proposed measure focuses on future held-out performance, not simply on whether an agent accumulates artifacts or improves on familiar tasks.\n\nAcross multiple environments, the paper reports markedly different improvement patterns among frontier models, even when they had comparable opportunities to learn. Some models made substantial, lasting gains, while others stayed close to or below their starting performance. Improvements on training interactions transferred only partly to out-of-distribution conditions, according to the paper. The model with the highest eventual performance was not necessarily the one that learned most efficiently, separating endpoint capability from improvement efficiency.\n\nThe authors also examine where the improvement loop fails: agents with low plasticity often do not reuse relevant artifacts, while agents that reuse them can still fail. They identify artifact quality, generalization, and application as possible limits for those agents. For AI builders, the framework suggests that evaluations of self-improving systems should track learning costs and held-out results over time, rather than relying on a single capability score. The paper presents this as a way to assess not just what agents can do, but how effectively experience makes them better.","content_kind":"author_paraphrase","explanation":{"feature":"A new evaluation approach tracks whether agents improve on later tasks, how efficiently they learn, and whether gains generalize beyond their training interactions.","relevance":"For AI builders, the framework suggests that evaluations of self-improving systems should track learning costs and held-out results over time, rather than relying on a single capability score.","use":"For AI builders, the framework suggests that evaluations of self-improving systems should track learning costs and held-out results over time, rather than relying on a single capability score."},"announcement_date":null,"published_at":"2026-10-08T12:05:07.906Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-08T12:05:07.704Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/52ccbec9-b2f1-4f6d-8e13-d37e652a7672","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"Agent Plasticity: Measuring Self-Improvement Through Experience","url":"https://arxiv.org/abs/2610.08902","fetched_at":"2026-10-08T11:01:39.122Z","sha256":"bcdb43de78e603cf74d518f2722f181fecdcb792b7a35e26a03a2aeafabaae41","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"127d64c4-8a95-481f-9a00-662cba05022e","envelope_sha256":"ccad68a31b7f1c025b0c4dda8806faaa8ce5c89acdf4b442545a56df79a18188"},"canonical_url":"https://news.guthlabs.ai/articles/researchers-propose-a-framework-to-measure-how-ai-agents-learn-from-experience-52ccbec9"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-08T12:05:07.906Z","reviewed_at":"2026-10-08T12:05:07.704Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Researchers propose a framework to measure how AI agents learn from experience","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/researchers-propose-a-framework-to-measure-how-ai-agents-learn-from-experience-52ccbec9?revision=1"}]}