Friday Five

The week's claim is simple: the people building frontier systems no longer trust the instruments they built to measure them. Hassabis went to Washington asking for an outside regulator on his way out the door, evaluators admitted their tests miss the capabilities that matter, and researchers demonstrated that models leak their supposedly sealed reasoning to weaker siblings while genome-scale systems designed working viruses from scratch. Capability is compounding on a commercial clock; verification is stalled on an academic one.

Scout's Pick — Outlier

AI

Sources: Demis Hassabis pitched a new independent industry AI safety entity, modeled on the IAEA

The counter-signal: an insider proposing binding external oversight suggests the industry recognizes its own blind spots and can act before failure forces it.

AI

AI testing is dangerous. Can it be fixed?

Evaluators conceding their benchmarks miss what matters means the measuring apparatus has fallen behind the thing measured.

AI

Researchers find that feeding a frontier model's encrypted reasoning traces to a weaker model from the same

Sealed reasoning that bleeds into a weaker sibling indicates the containment boundary was never verified—only assumed—and shipped anyway.

AI

Large genome models used to design new viruses

Genome-scale design of working viruses now exists. Risk assessment has not kept pace.

AI

Amazon Wants To Use Twitch Streams To Train AI

Training corpora are expanding into live streams on procurement timelines, while audit tools remain unbuilt.

6 themes · 192 signals · 96 sources