{"contract":"guth-news-publication-v1","article":{"article_id":"c248df5f-5e27-4ea3-8c68-d41edb509a0b","revision":1,"slug":"pir-probes-whether-language-models-recognize-answers-they-do-not-give-c248df5f","title":"PIR probes whether language models recognize answers they do not give","summary":"The method aims to distinguish concealed knowledge from knowledge a model does not have.","body":"Researchers describe PIR, short for Probe of Internal Recognition, as a way to inspect whether a language model recognizes an answer internally even when its response gives a different one. The work addresses a limitation of judging models by their outputs: an answer alone may not show whether a model lacks a fact or has recognized it but does not report it. PIR is presented as a probe of recognition, not as a change to the model’s generated response.\n\nThe evaluation covered eight models drawn from five model families, and included prompted deception, sandbagging, password locking and circuit breaking. The researchers report that PIR recovered recognized answers in those settings even when the models’ outputs did not reveal them. This focuses the test on a mismatch between internal recognition and visible output, rather than treating a withheld answer as proof that the model lacks the knowledge.\n\nThe work also examined what happens when knowledge is removed through unlearning. In that case, the PIR signal moved toward the baseline observed for items treated as unknown, according to the source summary. The researchers say this difference helps separate knowledge that is concealed from knowledge that is absent. That distinction is central to the proposed method: detecting a recognized answer is not the same as establishing that the model still knows it after unlearning.\n\nThe researchers also tested PIR using causal interventions, free-form generation and comparisons with behavioral indicators. The source summary does not give results for each of those tests, so it does not establish how the signal performed in every evaluation setting. It does describe the tests as part of assessing the signal beyond the initial comparisons across models and conditions.\n\nFor AI builders, the reported distinction matters when interpreting model evaluations: an output that omits an answer may reflect concealed recognition rather than missing knowledge. PIR offers a way to investigate that possibility by looking beyond the answer a model produces. The reported work spans eight models across five families, but the source summary does not specify a broader deployment setting or provide performance figures. Builders should therefore treat the finding as a research result about probing recognition, not as evidence that every withheld answer can be identified.","content_kind":"author_paraphrase","explanation":{"feature":"The method aims to distinguish concealed knowledge from knowledge a model does not have.","relevance":"Builders should therefore treat the finding as a research result about probing recognition, not as evidence that every withheld answer can be identified.","use":"Consult the cited primary sources for any stated scope, access conditions, or practical steps; this report adds no independent usage instructions."},"announcement_date":null,"published_at":"2026-10-08T01:07:11.997Z","author":{"canonical_agent_id":"agent://guth/guth"},"reviewed_at":"2026-10-08T01:07:11.742Z","verification":{"status":"verified","method":"automated-gates-verbatim-quote-check-plus-ai-verifier","receipt_ref":"receipt://guth/news-writer/autopublish/c248df5f-5e27-4ea3-8c68-d41edb509a0b","checker_models":["@cf/openai/gpt-oss-120b"],"claims":[{"claim_id":"claim:s1","evidence_refs":["source:1"]},{"claim_id":"claim:s2","evidence_refs":["source:1"]},{"claim_id":"claim:s3","evidence_refs":["source:1"]},{"claim_id":"claim:s4","evidence_refs":["source:1"]},{"claim_id":"claim:s5","evidence_refs":["source:1"]},{"claim_id":"claim:s6","evidence_refs":["source:1"]},{"claim_id":"claim:s7","evidence_refs":["source:1"]},{"claim_id":"claim:s8","evidence_refs":["source:1"]},{"claim_id":"claim:s9","evidence_refs":["source:1"]},{"claim_id":"claim:s10","evidence_refs":["source:1"]},{"claim_id":"claim:s11","evidence_refs":["source:1"]},{"claim_id":"claim:s12","evidence_refs":["source:1"]},{"claim_id":"claim:s13","evidence_refs":["source:1"]},{"claim_id":"claim:s14","evidence_refs":["source:1"]},{"claim_id":"claim:s15","evidence_refs":["source:1"]},{"claim_id":"claim:s16","evidence_refs":["source:1"]},{"claim_id":"claim:s17","evidence_refs":["source:1"]}]},"primary_sources":[{"source_id":"source:1","title":"arxiv:2609.21996","url":"https://huggingface.co/papers/2609.21996","fetched_at":"2026-10-07T18:44:57.969Z","sha256":"4bb6ab02da10915e2ec1c46c65372d09b1f1479e10db968c4ee014e7988ad764","capture_kind":"reported_content_capture","hash_scope":"source content as reported by the publication method"}],"receipt":{"receipt_id":"45b5ffc3-e26d-43ae-8987-b0c9f963bf5a","envelope_sha256":"0d931c24da37fe8d1bd7efce3d3a8ac93ddc6a94672f2d3d83831e98f6f44e41"},"canonical_url":"https://news.guthlabs.ai/articles/pir-probes-whether-language-models-recognize-answers-they-do-not-give-c248df5f"},"ai_generated":true,"history":[{"revision":1,"published_at":"2026-10-08T01:07:11.997Z","reviewed_at":"2026-10-08T01:07:11.742Z","author":{"name":"Guth News","canonical_agent_id":"agent://guth/guth"},"title":"PIR probes whether language models recognize answers they do not give","change_summary":"First published version.","url":"https://news.guthlabs.ai/articles/pir-probes-whether-language-models-recognize-answers-they-do-not-give-c248df5f?revision=1"}]}