Executive Summary
The retraction crisis of 2026 is not simply a rise in misconduct; it is a capacity failure in which industrialized fraud can generate and place manuscripts faster than publishers can investigate and remove them. A September 25, 2026 analysis of the Crossref–Retraction Watch dataset records 5,946 formal retractions in 2025 and 1,823 in 2026 year to date, while a 2025 PNAS study found that suspected paper-mill output was doubling every 1.5 years. The 2026 count is incomplete and is not a forecast. For research leaders and publishers, the operational implication is clear: integrity controls must move upstream, from article-by-article cleanup to layered, source-grounded screening before peer review.
Background and Scale
COPE and STM define paper mills as operations that manufacture manuscripts for payment, submit them for researchers or sell authorship. Their 2022 study of publisher data found suspected-paper rates of 2%–46% in the journals examined, while warning that the sample was not representative. Modern operations combine AI-assisted text, manipulated data, sold authorship and compromised reviewers or editors. The central vulnerability is organizational: a Science–Retraction Watch investigation identified several paper mills and more than 30 editors at reputable journals who appeared connected to bribery schemes, with guest-edited special issues particularly exposed.
The Hindawi episode made that vulnerability measurable. Nature reported that more than 10,000 research papers were retracted during 2023, including more than 8,000 from Hindawi journals after investigations cited compromised peer review, systematic manipulation, incoherent text and irrelevant references. Wiley's own account says Hindawi stopped publishing all special-issue content in October 2022, reassessed pending manuscripts and built a mass-investigation protocol using indicators such as scope discrepancies, unavailable data, inappropriate citations, meaningless content and manipulated peer review.
How Many Retractions Are There, and Who Is Retracting Them?
The latest downloadable Retraction Watch file, generated on September 25, 2026 and filtered to records classified specifically as “Retraction,” shows the following annual counts.
YearFormal retractions20213,90820225,591202313,227 (peak)20246,38320255,9462026 (through September 25)1,823
The decline after 2023 does not prove fraudulent production has receded: batch investigations create spikes, and notices often appear years after publication.
A 2026 Scientometrics study provides the more specific paper-mill view. Using the Retraction Watch Database through December 2024, Cheng, Yang and Allen identified 10,409 retracted paper-mill papers, of which 6,780 were retracted in 2023; the mean interval from publication to retraction was 709.4 days. The 2025 PNAS analysis found suspected paper-mill products doubling every 1.5 years, versus 3.3 years for retractions and 15 years for publications overall; only 8,589 of 29,956 suspected products matched to OpenAlex (28.7%) had been retracted.
Publishers and Fields
For the period from January 2021 through September 25, 2026, Retraction Watch's publisher labels assign the following formal retractions.
Publisher label in the databaseFormal retractions, 2021 to September 25, 2026Hindawi11,183Elsevier3,240Springer – Nature Publishing Group2,895Springer2,614Wiley2,213IOS Press (bought by Sage November 2023)1,764PLoS1,137IOP Publishing1,132Taylor and Francis1,021SAGE Publications799
These are counts, not failure rates; related imprints remain unmerged. The 2,895 under “Springer – Nature Publishing Group” is a multi-year Retraction Watch count for 2021–September 25, 2026. It is not comparable to Springer Nature's self-reported 2,923 retractions for calendar 2024: the figures use different sources, scopes and windows.
Subject tags are non-exclusive. Across the same 2021–2026 year-to-date window, the database marks 15,840 retracted records as Business and Technology, 13,062 as Health Sciences, 12,449 as Biological and Life Sciences, 7,122 as Physical Sciences and 6,005 as Social Sciences. Within the paper-mill-only corpus through 2024, Scientometrics ranked computer science first, followed by biology, biochemistry, business and medicine, and documented a shift from earlier biomedical terminology toward “model,” “data” and “algorithm” in 2021–2024.
What Do Hallucinated Citations Actually Look Like?
A language model can generate a citation without retrieving a source, predicting plausible names, titles, journals and DOI-like strings that are syntactically convincing but bibliographically false. The failure has at least three forms: an entirely fabricated publication; a real paper represented with corrupted metadata; or a valid source attached to a claim it does not support.
Linardon and colleagues tested GPT-4o on six mental-health reviews in June 2025 and verified every reference. Of 176 citations, 35 were fabricated (19.9% of the total). Of the 141 real citations, 64 (45.4%) contained bibliographic errors, most often incorrect or invalid DOIs. Only 77 were real and fully accurate: 54.6% of the real citations and 43.8% of all 176. These are three distinct categories (35 fabricated, 64 real but erroneous and 77 real and accurate), and together they account for all 176 references.
Case Study: A Visible Artifact
In April 2024, PLOS ONE retracted “A comparative analysis of blended learning and traditional instruction.” The article contained the interface phrase “regenerate response”; PLOS could not verify 18 of its 76 references and found errors in six more. Replacement references supplied during the investigation did not adequately support several relevant statements, and the editors said the combined unresolved issues called the article's reliability and policy compliance into question. The case exposed two failures: an AI artifact survived review, and the references had not been resolved source by source.
The problem is visible at scale. A March 2026 Nature analysis reported that, among nearly 18,000 papers accepted by three computer-science conferences, 2.6% of 2025 papers contained at least one potentially hallucinated citation, up from about 0.3% in 2024. Extrapolation beyond that sample is uncertain, but Nature concluded that at least tens of thousands of 2025 scholarly publications probably contained invalid AI-generated references.
What Is the Financial and Institutional Fallout?
Wiley acquired Hindawi for $298 million in January 2021. Its fiscal-2023 Form 10-K projected a $30 million–$35 million fiscal-2024 revenue decrease from the publishing disruption; in the second quarter of fiscal 2024, Wiley reported an $18 million Research revenue impact from the Hindawi pause and a $14 million adjusted-EBITDA impact. Nineteen Hindawi journals were removed from Clarivate's Web of Science Master Journal List in March 2023, a reminder that a journal's indexing status can change quickly; our piece on verifying a journal's Q1 and indexing claims covers how to check before you submit. No authoritative cross-publisher per-retraction cost is publicly available.
The operational burden extends well beyond withdrawal notices. Springer Nature reported 2,923 retractions in 2024 alongside more than 2.3 million submissions and more than 482,000 published articles. By September 2025, the STM Integrity Hub said more than 35 publishers were screening over 125,000 manuscripts each month and intercepting an estimated 1,000 suspected paper-mill submissions monthly. COPE's 2025 guidance explicitly recognizes the administrative burden of collecting evidence, providing due process across batches, sharing information and managing legal risk.
Damage also travels downstream. Deindexing can affect discovery, citation metrics and the evidence used in research assessment; investigations absorb editorial, legal, reviewer, librarian and university-integrity capacity. COPE's August 2025 retraction guidelines now expressly address batch cases and forms of misrepresentation including paper mills, fictitious authorship and undisclosed AI involvement, while emphasizing that retraction corrects the literature rather than punishes authors. For what disclosure expectations look like on the journal side, see our overview of 2026 journal AI-policy and traceability requirements.
Why Do Legacy Tools Fail?
The Mechanics of Peer Review Collapse
The first mismatch is text matching versus truth. Crossref explains that iThenticate scans a submission against selected repositories and calculates a similarity score from matching words; it checks similarity, not plagiarism. AI paraphrasing can retain a claim or fabricated result while replacing its lexical surface; legitimate quotations and methods language can create high similarity. The score requires expert interpretation and cannot establish whether data are genuine or a citation exists. A similarity check that links matches back to source records addresses part of that gap, but it too needs a human reader.
The second mismatch is stylometry versus verification. Generic AI detectors estimate whether linguistic patterns resemble model-generated or AI-paraphrased prose. They can prioritize passages but cannot prove authorship, empirical validity or citation support, and sentence-level stylometry carries its own false-positive risks, which we covered in our report on AI detection false positives and sentence-level stylometry. Springer Nature's own integrity stack illustrates the required separation of functions: it uses tools for non-standard phrases and nonsense text alongside an irrelevant-reference checker, image analysis and human assessment.
The third weakness is trust. Reviewers are recruited to assess novelty, methods and interpretation under limited time, not to authenticate every author, resolve every DOI, inspect all raw data and map relationships among submissions. If a mill controls reviewer identities, bribes an editor or exploits a weakly supervised special issue, peer review can become procedural theatre. The result is a localized peer review collapse wherever independence and source verification are absent.
What Is the Structural Fix?
No single AI percentage can repair this system. Publishers and universities need layered gates before external review: verified author and reviewer identities; live DOI and reference resolution; claim-to-source checking; semantic and stylometric anomaly detection; image and data forensics; cross-submission network analysis; and escalation to trained humans with documented due process. Because mills redirect manuscripts when one journal blocks them, the STM Integrity Hub's cross-publisher screening model is an early response to that mobility.
The design principle is to test the scholarly object, not merely its style: references require checks for existence, metadata accuracy, relevance and support. High-stakes conclusions must remain human decisions, backed by provenance and an opportunity for authors to respond.
Where SciCampus Differs
The SciCampus AI Detector is designed to combine sentence-level stylometric triage with search-grounded similarity checking against DOI-linked academic sources. It distinguishes raw AI-like text from AI-paraphrased patterns and links flagged sentences to candidate scholarly records. That workflow targets the fabricated-reference loophole: does the work exist, does its DOI resolve, does it support the claim, and is the passage legitimate paraphrase, unattributed borrowing or fabricated attribution?
This differs from generic detectors that estimate machine-like prose and similarity reports that identify lexical overlap. Source-grounded checking creates an auditable route from suspicion to evidence, but it should support, not replace, expert judgment and should be validated locally before consequential use.
Disclosure: This analysis is published by SciCampus, and the AI Detector discussed above is a SciCampus product; the product description is based on SciCampus documentation, not an independent comparative validation.
The priority is upstream verification: test references, claims, identities, data and cross-manuscript patterns before committing reviewer time. Otherwise, retraction totals will remain a backward-looking record of failures discovered after they entered the literature.
Sources and References
Crossref, “Retraction Watch Database Documentation” and “Retraction Watch Data for 2026-09-25,” Crossref/GitLab; documentation updated January 18, 2025; dataset generated September 25, 2026. https://www.crossref.org/documentation/retrieve-metadata/retraction-watch/ and https://gitlab.com/crossref/retraction-watch-data
Richardson RAK, Hong SS, Byrne JA, Stoeger T, Amaral LAN. The entities enabling scientific fraud at scale are large, resilient, and growing rapidly. Proceedings of the National Academy of Sciences. August 4, 2025. https://doi.org/10.1073/pnas.2420092122
Cheng MWT, Yang X, Allen RM. Tracking the retracted paper mill articles: a bibliometric study. Scientometrics. July 24, 2026. https://doi.org/10.1007/s11192-026-05751-6
Committee on Publication Ethics and STM Association. Paper Mills: Research Report. COPE, June 2022. https://doi.org/10.24318/jtbG8IHL
Joelving F. Paper mills are bribing editors at scholarly journals, Science investigation finds. Science. January 17, 2024. https://www.science.org/content/article/paper-mills-bribing-editors-scholarly-journals-science-investigation-finds
Van Noorden R. More than 10,000 research papers were retracted in 2023: a new record. Nature. December 11, 2023. https://doi.org/10.1038/d41586-023-03974-8
Wiley. Tackling Publication Manipulation at Scale: Hindawi's Journey and Lessons for Academic Publishing. Wiley White Paper, December 2023. https://www.wiley.com/content/dam/wiley-com/en/pdfs/insights/tackling-publication-manipulation-at-scale-a-whitepaper.pdf
Linardon J, et al. Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models. JMIR Mental Health. 2025;12:e80371. https://doi.org/10.2196/80371
Naddaf M, Quill E. Hallucinated citations are polluting the scientific literature. What can be done? Nature. March 31, 2026. https://www.nature.com/articles/d41586-026-00969-z
PLOS ONE Editors. Retraction: A comparative analysis of blended learning and traditional instruction: effects on academic motivation and learning outcomes. PLOS ONE. April 18, 2024. https://doi.org/10.1371/journal.pone.0302484
Wiley. Wiley Announces the Acquisition of Hindawi. Wiley Newsroom, January 5, 2021. https://newsroom.wiley.com/press-releases/press-release-details/2021/Wiley-Announces-the-Acquisition-of-Hindawi/default.aspx
John Wiley & Sons, Inc. Annual Report on Form 10-K for the fiscal year ended April 30, 2023. U.S. Securities and Exchange Commission, filed June 26, 2023. https://www.sec.gov/Archives/edgar/data/107140/000010714023000092/jwa-20230430.htm
Wiley. Wiley Reports Second Quarter 2024 Results. Wiley Newsroom, December 6, 2023. https://newsroom.wiley.com/press-releases/press-release-details/2023/Wiley-Reports-Second-Quarter-2024-Results/default.aspx
Kincaid E. Nearly 20 Hindawi journals delisted from leading index amid concerns of papermill activity. Retraction Watch. March 21, 2023. https://retractionwatch.com/2023/03/21/nearly-20-hindawi-journals-delisted-from-leading-index-amid-concerns-of-papermill-activity/
Springer Nature. Annual Report 2024. 2025. https://annualreport.springernature.com/2024/pdfs/SN_Annual_report_24_1_Story.pdf
STM Association. STM Integrity Hub in Action: A Chronicle. September 11, 2025. https://stm-assoc.org/stm-integrity-hub-in-action-a-chronicle/
Committee on Publication Ethics. Addressing concerns about systematic manipulation of the publication process. COPE, February 16, 2025. https://publicationethics.org/guidance/flowchart/addressing-concerns-about-systematic-manipulation-publication-process
COPE Council. COPE Retraction Guidelines, Version 3. Committee on Publication Ethics, August 2025. https://doi.org/10.24318/cope.2019.1.4
Crossref. Understanding your Similarity Report. Crossref Documentation, May 18, 2020. https://www.crossref.org/documentation/similarity-check/similarity-report-understand/
Graf C. Harnessing technology to strengthen research integrity. Springer Nature, August 5, 2025. https://www.springernature.com/gp/advancing-discovery/springboard/blog/blogposts-trust-integrity/harnessing-technology-to-strengthen-research-integrity/
SciCampus. Why aggregate AI detection scores fail academic authors. September 18, 2026. https://scicampus.com/blog/why-aggregate-ai-scores-fail
Related reading (topic ideas for future posts)
How the STM Integrity Hub screens manuscripts, and what it can't catch
Fabricated references in 2026: what the Lancet and Nature audits found, side by side
Special issues as an attack surface: lessons from guest-edited collections
A reference-verification workflow for authors: DOI, metadata, and claim support



