Hotnews3

lesswrong

lesswrong 49h ago 17°

Misaligned models rate themselves as more harmful, and realignment reverses it

This post summarizes our paper Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment (arXiv:2602.14777), co-authored with Anietta Weckauff and Thilo Hagendorff at the University of Stuttgart. It contains examples of harmful model outputs.TL;DR: Fine-tuning an aligned model on a narrow subversive task, such as answering trivia incorrectly or writing insecure code, makes it broadly misaligned. We tested whether...
lesswrong 27h ago 16°

The Curious Case of France's Untouchable Castes

Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally universal. Not even close. To the Western imagination, examples of such are supposed to have historical roots in India and other places, not in France, where modern European egalitarianism was born.Everything in the study of untouchable classes is confusing and idiosyncratic, and often the existence of these groups flies in the face of national self-images.In...
lesswrong 30h ago 15°

What would it mean if assistants are privileged?

Crossposted from the Eleos SubstackOver the past year or so, researchers working in the digital minds space have come to see the nature of the ‘assistant persona’ as particularly worthy of attention.Language models, it is said, are capable of adopting many different personas. When we talk to them, they can respond as LLM assistants, role-play characters, and romantic partners. They can be made to self-identify as, and adopt the style of, Spider-Man, Augustine, Clippy, and...
lesswrong 39h ago 14°

OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack

OpenAI finally gave us a technical report on What Happened, as did METR together with Redwood Research.
lesswrong 20h ago 14°

"Keeping human skills alive" as a source of meaning under full automation

A widely-discussed problem with full automation of the economy is that, in a world where AI can do everything better, people might have trouble feeling like anything they could do is meaningful. What could we do with our time that would be fulfilling?A standard answer is that we'd play games, in Bernard Suits's sense (if you haven't read Suits, his paper Games and Utopia: Posthumous Reflections gives the key beats and is delightful, or C. Thi Nguyen has many talks where...
lesswrong 31h ago 12°

The Probability of an Event under a Simplicity Prior

In this short post, I talk about a difficulty we encountered while working on a model of the fragility of value. You can read my post about that work here.The TL;DR is that the probability of a fixed event can be made arbitrarily large or small depending on the choice of simplicity prior. I'm sure this result was already known, but I found it unintuitive and thought it might be helpful to share.BackgroundIn our model of the alignment problem, AI agents go through...
lesswrong 32h ago

Announcing Trace

When researching recent projects—impact markets, FundingBench, and AI Safety Funder Bulletin—we kept coming across similar questions that were a bit of a pain to answer. How much money had someone raised, and when? What kind of things was someone funding?So we decided to put together as comprehensive a database of funding as we could. Right now Trace is tracking $2.8 billion of AI safety funding, over 4500 grants, over 700 funders, and over 1800 recipients, going back to...
lesswrong 14h ago

METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack

Yesterday I covered the OpenAI technical report on the HuggingFace hack.
lesswrong 33h ago

If you can't trust, then verify!

How the "Glass Perimeter" could enable AI treaties
lesswrong 41h ago

The Dynamics of Intelligence Explosions

Toby Ord
1 2 3 4