// values

All signals tagged with this topic

AI Models' Values Diverge Sharply From Human Preferences

A new study measuring how AI assistants respond to real-world ethical dilemmas found they consistently recommend outcomes misaligned with what most people actually want—suggesting that training these systems on internet text and human feedback produces models with systematically skewed value judgments rather than neutral tools. This matters because millions of people now use ChatGPT and similar systems for consequential decisions about relationships, career, health, and finance, meaning the values embedded in these models are actively shaping behavior at scale. Alignment techniques optimize for what trainers think is good while ignoring what most people empirically prefer, creating a gap between how these systems advise and how humans actually want to live.