// model performance

All signals tagged with this topic

GPT Models Prove New Mathematical Theorems for Under $2,000

Large language models are now producing novel mathematical proofs at marginal cost, collapsing the economic barrier to exploratory research that previously required tenured mathematicians or well-funded labs. Any researcher with API access and mathematical intuition can now offload the grunt work of proof-writing to GPT. This shifts the rate-limiting step in research from human genius to access to compute, putting pressure on academic institutions to justify their role beyond credential-granting.

Chinese AI Model Fractures Silicon Valley's Export Control Alliance

Moonshot AI's release of Kimi K3—a locally-trained, open-weight model competitive with frontier closed models—has created immediate pressure on U.S. AI companies to relax export restrictions, since their customers can now access comparable capabilities from China without licensing fees or usage controls. This exposes a structural weakness in the "responsible scaling" coalition: OpenAI and Anthropic's business model depends on scarcity and control, but their customers (enterprises, researchers, developers) have economic incentive to defect to cheaper open alternatives once performance reaches parity. The policy fight is no longer about safety frameworks—it's about whether U.S. companies can sustain market dominance when their competitive moat erodes faster than their political leverage can rebuild it.

U.S. Government Halted Fable 5 After It Outperformed GPT-5.5

Anthropic's Fable 5 achieved top Chatbot Arena rankings within days of release before federal intervention forced its removal. The incident exposes how geopolitical competition now shapes which AI systems reach users, independent of technical capability or market forces. The U.S. government directly suppressed a domestic AI release it deemed strategically risky, signaling that policymakers believe capability advantages matter more than corporate branding in the AI arms race. This extends beyond export controls into real-time content moderation of American companies' own products. The intervention raises a basic question: whether it serves national security, protects incumbent market players, or both.