CognitionAI02 sources

Nine in ten biomedical papers now show signs of LLM-assisted writing

Two line charts of estimated LLM usage from 2020 to 2026 by paper section, flat near zero until a dashed line marked ChatGPT in late 2022, after which all six lines climb steadily, full paper and discussion highest.

Holzwarth, González-Márquez & Kobak, arXiv 2026 (CC BY-SA 4.0)CC BY-SA

A preprint posted to arXiv on 11 August 2026 by Lena Holzwarth, Rita González-Márquez and Dmitry Kobak estimates that 89% of open-access biomedical papers archived in PubMed Central showed an excess of LLM-associated vocabulary by the end of 2025. It has not been peer reviewed.

The method counts how often words characteristic of large language models appear, then compares observed frequencies against what pre-2023 trends predict. On that basis the authors put usage at 77% for papers published across all of 2025 and 52% for 2024, and only English-language papers were included.

These figures are far above earlier estimates. A 2025 paper by some of the same researchers, analysing abstracts rather than full text, reported at least 13.5% for 2024; applying the newer method lifts that same 2024 abstract figure to 31%. Kobak says he was initially sceptical of the calculations and that further checks convinced him the data held up.

Usage is unevenly distributed within a paper. An estimated 68% of discussion paragraphs carry the markers against 32% of methods paragraphs — though the authors note prevalence exceeds 50% even inside methods. Other researchers told Nature the rates are plausible given how widely the tools are used, while cautioning that PubMed Central may not represent the whole literature.

Sources

  1. [1]Most biomedical publications show signs of LLM-assisted writingarXiv··Preprint
  2. [2]Staggering 90% of biomedical papers now show signs of AI helpNature··Article