{"contract":"guth-news-publication-v1","article":{"article_id":"3ee499a5-8df7-4efc-a91f-132b3c5910fb","revision":1,"slug":"thinkingbox-benchmark-tests-agents-on-stateful-business-workflows-3ee499a5","title":"Thinkingbox benchmark tests agents on stateful business workflows","summary":"The sandbox and benchmark evaluate whether agents reliably complete multi-step business tasks and leave the correct backend state.","body":"A new paper introduces Thinkingbox, a sandbox for evaluating agents as they work with users and tools in business workflows that change persistent system state. The authors say many agent tests focus on executable environments, but consequential work also requires getting missing information, following policies and coordinating dependent tools. Their benchmark therefore checks whether an agent produces the intended state change without unwanted effects, rather than relying only on a plausible answer or valid tool call. The paper was submitted on 20 Aug 2026 and last revised on 1 Oct 2026.\n\nThinkingbox provides isolated sessions compatible with MCP, records execution traces and evaluates outcomes against the final backend state. Its companion benchmark, Thinkingbox-bench, contains 507 policy-conditioned workflows. The scenarios cover retail, hospitality, auto insurance, neobank internal IT, and consulting IT and HR support. Task-specific executable checks are designed to accept valid action sequences while flagging incorrect, missing or extra effects. Some tasks also assess whether the final response has required properties.\n\nThe authors report that performance declined when they measured repeated-attempt reliability, including for the strongest proprietary and open-weight models tested. Claude Opus 5 scored 66.50% on pass@1 and 47.53% on pass^20. Kimi-K3 scored 57.37% on pass@1 and 17.60% on pass^20. The paper says many unsuccessful attempts ended cleanly after valid state-changing actions, meaning orderly-looking tool activity did not guarantee task completion. The authors conclude that response-level and tool-call-level signals are poor proxies for completing the full workflow.\n\nFor AI builders, the results highlight a distinction between finding a successful sequence once and completing stateful tasks reliably across attempts. They also show why evaluations of business agents can check the resulting system state, not just the apparent quality of an answer or the validity of individual tool calls. The authors say they are releasing both Thinkingbox and Thinkingbox-bench.","content_kind":"author_paraphrase","explanation":{"feature":"The sandbox and benchmark evaluate whether agents reliably complete multi-step business tasks and leave the correct backend state.","relevance":"For AI builders, the results highlight a distinction between finding a successful sequence once and completing stateful tasks reliably across attempts.","use":"The authors say they are releasing both Thinkingbox and Thinkingbox-bench."},"announcement_date":null,"published_at":"2026-10-04T18:07:41.228Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-04T18:07:40.491Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/3ee499a5-8df7-4efc-a91f-132b3c5910fb","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]},{"claim_id":"claim:s17","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows","url":"https://arxiv.org/abs/2608.19741","fetched_at":"2026-10-04T10:00:47.315Z","sha256":"77810d6755acda7ec6117693ecccf813a395bad7ae94dc0b2e177d033afe2e15","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"11735f3f-8b63-4efd-8cdf-fd286ef02bc7","envelope_sha256":"96934d7f21eac94dda5afe1a2c2d799b64fec4da7501981aa0f83774c7c0a11a"},"canonical_url":"https://news.guthlabs.ai/articles/thinkingbox-benchmark-tests-agents-on-stateful-business-workflows-3ee499a5"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-04T18:07:41.228Z","reviewed_at":"2026-10-04T18:07:40.491Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Thinkingbox benchmark tests agents on stateful business workflows","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/thinkingbox-benchmark-tests-agents-on-stateful-business-workflows-3ee499a5?revision=1"}]}