Section 1: The Problem

A convincing scientific paper can influence years of experiments, grant decisions, public policy, and commercial research. If its images were duplicated, spliced, mislabeled, or fabricated, every researcher who builds on the result risks wasting time and money.

The scale is larger than isolated misconduct cases suggest. Bik, Casadevall, and Fang visually inspected 20,621 biomedical papers from 40 journals and found problematic image duplications in 3.8%. At least half showed features consistent with deliberate manipulation rather than an obvious formatting mistake (Bik, Casadevall, and Fang).

Paper mills turn the problem into a business. Abalkina connected at least 434 publications to one Russia-based authorship-selling operation and estimated that its advertised coauthorship positions were worth $6.5 million from 2019 through 2021 (Abalkina). A 2025 systematic review reported thousands of paper-mill-related retractions and the closure of 19 Wiley journals after large-scale integrity failures (Kamel and Barghi).

Traditional peer review rarely includes systematic image forensics. Editors and reviewers concentrate on methods, writing, novelty, and interpretation. Finding a rotated microscopy image or reused Western blot may require comparing hundreds of small panels across the manuscript and the published literature.

Section 2: What Research Shows

People are not particularly reliable image-forensics systems. In a study of 383 participants making 17,208 judgments, humans distinguished altered from unaltered images with 58% overall accuracy and correctly identified manipulated images only 46.5% of the time (Schetinger et al.).

Automated methods perform better in controlled tests. Mazaheri, Avila, and Roy-Chowdhury developed a scientific-image duplication system that reached 90% accuracy, approximately 13 percentage points above comparison methods without the same preprocessing pipeline (Mazaheri, Avila, and Roy-Chowdhury).

Nagm and colleagues combined error-level analysis with a convolutional neural network. On the CASIA 2 image-forgery dataset, the model reached 94.14% testing accuracy, 94.1% precision, and 94.07% recall (Nagm et al.). These results show that models can identify compression and editing patterns people routinely miss, although a general forgery benchmark is simpler than judging whether a scientific figure is ethically acceptable (Nagm et al.).

Section 3: What the Real World Shows

The American Society for Microbiology provides the strongest publisher-level field evidence. ASM integrated the Imagetwin image-comparison system into a one-year pilot covering 15 journals from March 2023 through March 2024. It screened 2,627 accepted manuscripts and found image-related concerns in 410 (Chaturvedi et al.).

Of those 410 manuscripts, 248 involved image duplication, 125 contained unacknowledged splicing, and 37 involved non-uniform image enhancement. The process led to 142 image replacements, 82 text modifications, and six revoked acceptance decisions when authors could not provide satisfactory explanations or underlying data (Chaturvedi et al.).

The timing of screening changed the workload. ASM reported that resolving a figure problem after publication could consume up to 10 staff hours, while addressing it before publication averaged 1.5 hours—an 85% reduction. The pilot became part of ASM’s routine ethics workflow rather than ending as a temporary experiment (Chaturvedi et al.).

Section 4: The Implementation Gap

The first barrier is that a match is not a verdict. The same image may be reused legitimately as a disclosed control, duplicated accidentally during figure assembly, or deliberately relabeled as a different experiment. ASM therefore required automated detection, specialist visual review, and pixel-level verification before contacting authors (Chaturvedi et al.).

The second barrier is limited coverage. Imagetwin worked on halftone material such as photographs, microscopy images, and Western blots, but it could not evaluate every graph, drawing, or line-art figure. Current tools also struggle with completely synthetic images that do not reuse anything already in a comparison database (Chaturvedi et al.).

The third barrier is slow correction after publication. Aquarius and colleagues systematically examined 608 animal-research articles and found 243 problematic papers, or 40%. Of those, only 22.6% had been corrected and 7.8% retracted by July 2025. More than 95% had not been previously flagged before the investigation began (Aquarius et al.).

Correction does not always solve the problem. Nine of the 55 corrected articles—16.4%—still contained unresolved or newly identified issues. The researchers concluded that science’s self-correction mechanisms had stalled in the literature they studied (Aquarius et al.).

The fourth barrier is incentives. Paper mills alter wording, author networks, figures, peer-review accounts, and submission patterns as publishers learn their previous tactics. Kamel and Barghi’s 2025 systematic review found that detection has expanded from manual editorial checks to algorithmic screening, but responses remain fragmented across journals, institutions, funders, and publishers (Kamel and Barghi).

Section 5: Where It Actually Works

Integrity screening works best before publication, after authors submit revised figures but before the journal makes its final decision. At that point, editors can request original files, correct legends, replace accidental duplications, repeat experiments, or reject the manuscript without repairing an already-cited scientific record.

It also works best as a layered system. Image comparison can detect duplicated panels. Text models can flag paper-mill templates and unnatural phrase patterns. Authorship-network analysis can identify suspicious collaboration structures. Human editors then determine whether the evidence reflects error, poor documentation, plagiarism, or potential misconduct.

Section 6: The Opportunity

The opportunity is not an algorithm that automatically labels researchers as frauds. It is a quality-control system that treats figures as data rather than decoration.

Every major journal could screen images before acceptance, preserve original research files, document how alerts were resolved, and share confirmed manipulation patterns through secure publisher networks. Automated tools can make the suspicious panel visible. Only transparent editorial investigation can determine what it means.

References

[1] Bik, Elisabeth M., Arturo Casadevall, and Ferric C. Fang. “The Prevalence of Inappropriate Image Duplication in Biomedical Research Publications.” mBio, vol. 7, no. 3, 2016.

[2] Abalkina, Anna. “Publication and Collaboration Anomalies in Academic Papers Originating from a Paper Mill.” Learned Publishing, 2023.

[3] Kamel, Rabab, and Nasrin Ghassemi Barghi. “Paper Mills in Scholarly Publishing: A Systematic Review of Their Prevalence, Detection, and Impact on Research Integrity.” Trends in Scholarly Publishing, 2025.

[4] Schetinger, Victor, et al. “Humans Are Easily Fooled by Digital Images.” Computers & Graphics, vol. 68, 2017, pp. 142–151.

[5] Mazaheri, Ghazal, Kevin Urrutia Avila, and Amit K. Roy-Chowdhury. “Learning to Identify Image Manipulations in Scientific Publications.” 2021 IEEE Winter Conference on Applications of Computer Vision Workshops, 2021.

[6] Nagm, Ahmed M., et al. “Detecting Image Manipulation with ELA-CNN Integration: A Powerful Framework for Authenticity Verification.” Scientific Reports, vol. 14, 2024.

[7] Chaturvedi, Aashi P., et al. “ASM Incorporates Imagetwin to Address Image Duplication and Preserve Scientific Accuracy.” mBio, vol. 16, 2025.

[8] Aquarius, R., et al. “High Prevalence of Articles with Image-Related Problems in Animal Studies of Subarachnoid Hemorrhage and Low Rates of Correction by Publishers.” PLOS Biology, vol. 23, no. 10, 2025.

[9] Gu, Jin, et al. “AI-Enabled Image Fraud in Scientific Publications.” Patterns, vol. 3, no. 7, 2022.

[10] Kadha, Vinay Kumar, et al. “A Systematic Survey on Image Manipulation Detection and Localization.” ACM Computing Surveys, 2025.

Leave a comment