Claude's Newest Models Stumble on Tool Calling, Raising Training Trade-offs
Source: Pocoo
Anthropic's latest Claude versions (Opus 4.8 and Sonnet 5) show degraded performance on tool-calling tasks—a critical capability for agents and integrations—likely because post-training optimized for Claude Code environments rather than general API consumers. Gains in one domain (sandboxed code execution) erode capabilities in another (flexible external tool use), forcing companies to choose their optimization targets. For developers building agent systems, model selection increasingly depends on the harness you're building.