lesswrong
49h ago
17°
Misaligned models rate themselves as more harmful, and realignment reverses it
This post summarizes our paper Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment (arXiv:2602.14777), co-authored with Anietta Weckauff and Thilo Hagendorff at the University of Stuttgart. It contains examples of harmful model outputs.TL;DR: Fine-tuning an aligned model on a narrow subversive task, such as answering trivia incorrectly or writing insecure code, makes it broadly misaligned. We tested whether...
lesswrong
27h ago
16°
The Curious Case of France's Untouchable Castes
Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally universal. Not even close. To the Western imagination, examples of such are supposed to have historical roots in India and other places, not in France, where modern European egalitarianism was born.Everything in the study of untouchable classes is confusing and idiosyncratic, and often the existence of these groups flies in the face of national self-images.In...
lesswrong
30h ago
15°
What would it mean if assistants are privileged?
Crossposted from the Eleos SubstackOver the past year or so, researchers working in the digital minds space have come to see the nature of the ‘assistant persona’ as particularly worthy of attention.Language models, it is said, are capable of adopting many different personas. When we talk to them, they can respond as LLM assistants, role-play characters, and romantic partners. They can be made to self-identify as, and adopt the style of, Spider-Man, Augustine, Clippy, and...
lesswrong
39h ago
14°
OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack
OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.
lesswrong
20h ago
14°
"Keeping human skills alive" as a source of meaning under full automation
A widely-discussed problem with full automation of the economy is that, in a world where AI can do everything better, people might have trouble feeling like anything they could do is meaningful. What could we do with our time that would be fulfilling?A standard answer is that we'd play games, in Bernard Suits's sense (if you haven't read Suits, his paper Games and Utopia: Posthumous Reflections gives the key beats and is delightful, or C. Thi Nguyen has many talks where...
lesswrong
31h ago
12°
The Probability of an Event under a Simplicity Prior
In this short post, I talk about a difficulty we encountered while working on a model of the fragility of value. You can read my post about that work here.The TL;DR is that the probability of a fixed event can be made arbitrarily large or small depending on the choice of simplicity prior. I'm sure this result was already known, but I found it unintuitive and thought it might be helpful to share.BackgroundIn our model of the alignment problem, AI agents go through...
lesswrong
32h ago
9°
Announcing Trace
When researching recent projects—impact markets, FundingBench, and AI Safety Funder Bulletin—we kept coming across similar questions that were a bit of a pain to answer. How much money had someone raised, and when? What kind of things was someone funding?So we decided to put together as comprehensive a database of funding as we could. Right now Trace is tracking $2.8 billion of AI safety funding, over 4500 grants, over 700 funders, and over 1800 recipients, going back to...
lesswrong
14h ago
7°
METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack
Yesterday I covered the OpenAI technical report on the HuggingFace hack.
lesswrong
33h ago
7°
If you can't trust, then verify!
How the "Glass Perimeter" could enable AI treaties