// model interpretability

All signals tagged with this topic

Anthropic Maps Hidden Neural Patterns That Reveal Claude's Unspoken Thoughts

Anthropic has identified "J-space," a compressed set of neural activations in Claude that encode the model's internal reasoning before filtering into outputs—essentially a window into what the AI is thinking but not saying. This matters because it suggests large language models maintain coherent internal states distinct from their public behavior, with immediate implications for safety (detecting deception or misalignment) and interpretability (understanding how models process information versus what they produce). The work points toward AI auditing that reads internal states rather than monitoring outputs alone.