Tech

One-third of ArXiv submissions now read as machine-written, study finds

Analysis of 12,750 papers by unslop.run reveals a sharp rise in machine-like text since 2023, with computer science leading the surge.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
ArXiv AI detection research

An analysis of 12,750 ArXiv papers submitted between January 2023 and July 2026 has found that approximately 32% of new submissions read as machine-written. The study, conducted by unslop.run, utilised a detection tool calibrated to a 0.4% false-positive rate on pre-2023 papers, ensuring that the baseline for genuine human writing remained stable. The data indicates a significant shift in academic publishing, with the flagged share climbing from a flat 0.4% in 2021 and 2022 to a peak of nearly 39% in early 2026.

The research sampled ten field groups, selecting roughly 25 papers per field per month to assess the full body text of each submission. The authors noted that abstracts alone understate the signal, as some papers scored under 20% on their abstracts but over 70% on their full text. The flagged share remained flat through the pre-ChatGPT years of 2021 and 2022, then lifted off within months of the model’s release, climbing in two distinct waves to reach the current levels.

Disparities across academic disciplines were stark. Computer science led with a 65% flag rate, indicating a heavy adoption of AI-generated or AI-assisted writing in that sector. In contrast, mathematics recorded the lowest rate at 0.7%. The authors cautioned that this low figure may reflect detector limitations rather than low adoption, as mathematical prose is sparse and dominated by notation, placing it out of distribution for the detector trained on scientific English.

The study’s methodology anchored its findings to a strict control group. By setting the flag threshold so that exactly 0.4% of papers from 2021 and 2022 triggered a flag, the researchers established a floor for genuine human text. This approach allowed them to rule out detector bias as the primary driver of the observed rise, as the control years did not show elevated flag rates despite the same threshold being applied.

Despite the high flag rates, the authors emphasised that a score indicates machine-like writing rather than definitive authorship. The detector cannot distinguish between lightly edited documents and wholly generated ones, nor can it evaluate the exact private mixture of models and prompts used by authors. Consequently, the reported prevalence serves as a lower bound, with the true share of AI involvement likely higher due to detector sensitivity variations across different generators.

Continue reading

More from Tech

Read next: Open-source tool claims 97 per cent token savings for AI agents
Read next: Valvoline Unveils August 2026 Promotional Offers for Service and Retail Buyers
Read next: Developer Antirez releases native MiniMax H3 inference engine for Apple Silicon