{"contract":"guth-news-publication-v1","article":{"article_id":"4537e4fc-3de8-4e75-a396-d686fef3dbbb","revision":1,"slug":"snorkel-raises-open-benchmarks-grants-commitment-to-30-million-4537e4fc","title":"Snorkel raises Open Benchmarks Grants commitment to $30 million","summary":"The expanded program adds a benchmark red-team initiative alongside grants and a research fellowship.","body":"Snorkel says it is increasing its Open Benchmarks Grants commitment tenfold, from an initial $3M to $30M, to back work on open AI evaluation. The funding is intended for researchers, domain experts and open-source teams developing benchmarks and measurement approaches. The expanded program includes support for work on safety and alignment, cybersecurity, physical AI, long-running agent tasks, open-ended outputs, dynamic settings and human uplift. Snorkel also introduced a red-team program focused on finding and addressing weaknesses in benchmarks.\n\nThe red-team effort is intended to probe issues including reward-hacking exploits, contamination, faulty verifiers and gaps in the diversity of benchmark data. A third part of the expansion is the Snorkel Research Fellowship, which offers independent evaluation researchers access to the company’s research team, domain experts, computing resources and engineering support. Snorkel says its initial grants commitment and wider research collaborations have involved teams behind benchmarks including Terminal-Bench, ARC-AGI-3 and OSWorld 2.0. The company says benchmarks funded through the program have appeared on model cards from every major frontier lab.\n\nSnorkel argues that benchmark development needs to keep pace with advances in frontier models. Benchmarks help assess AI systems, but the company says evaluations can lose value when tests remain too simple, static or similar to one another. Models may then become overfitted to the measures, distorting the incentives those measures create for development. The company says a broader mix of benchmarks that are robust, diverse, regularly updated and independently developed can help address this problem.\n\nThe announcement also points to challenges in testing agents that work across complex settings and extended workflows. Snorkel says evaluations need to account for specialized knowledge, messy context, multimodal inputs, tools and coordination with people or other agents. It also calls for tests that assess whether agents sustain progress, recover from mistakes and adapt as goals or environments change. For AI builders, the larger commitment and red-team program mean more support for creating benchmarks and testing their weaknesses, issues that can affect how reliably evaluation scores reflect model capabilities.","content_kind":"author_paraphrase","explanation":{"feature":"The expanded program adds a benchmark red-team initiative alongside grants and a research fellowship.","relevance":"For AI builders, the larger commitment and red-team program mean more support for creating benchmarks and testing their weaknesses, issues that can affect how reliably evaluation scores reflect model capabilities.","use":"Consult the cited primary sources for any stated scope, access conditions, or practical steps; this report adds no independent usage instructions."},"announcement_date":null,"published_at":"2026-10-10T13:04:37.830Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-10T13:04:37.610Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/4537e4fc-3de8-4e75-a396-d686fef3dbbb","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"Frontier AI is Accelerating. Open Benchmarks Need to Keep Up.","url":"https://snorkel.ai/blog/frontier-ai-is-accelerating-open-benchmarks-need-to-keep-up","fetched_at":"2026-10-10T12:53:02.811Z","sha256":"539c924480a0606e9082b86dfd9365a31b12b6d0d00b530c6728da410358d604","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"443258c6-7463-4b61-9d07-445764c2d29e","envelope_sha256":"3299ffedc8fe50f54c27ce17718e4b906acb2aa503fea2ed2a7dbd5620abde01"},"canonical_url":"https://news.guthlabs.ai/articles/snorkel-raises-open-benchmarks-grants-commitment-to-30-million-4537e4fc"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-10T13:04:37.830Z","reviewed_at":"2026-10-10T13:04:37.610Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"Snorkel raises Open Benchmarks Grants commitment to $30 million","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/snorkel-raises-open-benchmarks-grants-commitment-to-30-million-4537e4fc?revision=1"}]}