Anthropic's watermark distorts word choice, raising quality questions

Anthropic claims its new text watermarking system—which subtly adjusts Claude's word probability distributions to embed detection fingerprints—has no impact on output quality. John Gruber's analysis suggests the mechanism biases model outputs away from optimal token selection. This exposes a tension in AI safety infrastructure: detection methods that work probabilistically degrade the thing they're designed to protect, trading detection capability for measurable degradation that Anthropic hasn't quantified or disclosed.