Research

Retracted papers can linger in AI research corpora; the data to catch them exists

Research assistants built on language models often index a paper once and cite it indefinitely. Scholarly infrastructure already publishes corrections and retractions in machine-readable form, but Crossref cautions that its status signals are not a guarantee, and one older data route now returns stale results.

A DOI identifies a scholarly record persistently, but the record behind it can change: publishers issue corrections, expressions of concern and retractions after publication. Crossref's Crossmark service exists so that readers and software can find those updates, and since January 2025 Crossref's public REST API has also carried retraction data from Retraction Watch, labeled by whether it came from the publisher or from that database.

Two cautions apply. Crossref states that the presence of Crossmark is not in itself a guarantee, since the system depends on publishers reporting updates. And an experimental Crossref Labs route for Retraction Watch data has been retired; a page updated on May 29, 2026 warns that anyone still using it will get out-of-date information and should move to the production services.

Standards bodies have pushed the same direction. NISO's 2024 recommended practice on retractions, removals and expressions of concern calls for status changes to be evident to automated processes as well as to people, and DataCite's metadata schema reached version 4.7 in March. For teams maintaining an evidence corpus, the implication is to re-check status periodically and to store which version of a source supported which claim.

Source details
Source
Crossref

Source reporting

Read the original reporting and research behind this briefing.