Source: Marginal REVOLUTION
Claude and frontier models caught only about 50% of deliberate errors inserted into psychology papers. AI peer review remains a weak substitute for human scrutiny, not a reliable complement. The gap matters because journals are already under pressure to adopt faster review processes. Deploying frontier AI as a first-pass filter could systematically let flawed work through—especially in fields where methodological errors compound across downstream research. Commercial AI tools performed even worse, indicating that capability gaps between frontier and commodity models create real quality-control stakes for publishers considering automation.