Hotnews

lesswrong

lesswrong 15h ago 62°

What just happened? A retrospective of AI alignment

This sequence is about the last decade in AI alignment. Over five posts, it recounts the gradual transition from a field which treated alignment as a hard scientific problem, to a field which has largely abandoned the goal of deep, generalizable scientific progress in favor of iteratively improving existing systems and attempting to gain technological and political power. I also describe (in subsequent posts, which I'll upload over the next few weeks) how fear and...
lesswrong 2h ago 62°

How to be an AI safety research engineer

This is the advice I wish I had when I started trying to become an AI safety research engineer.The LandscapeStart by working out which issues you care about. If you don't care about any, hiring managers don't care how good of an engineer you are. You shouldn’t blindly agree with all issues in AI safety. Predicting the future is hard, so many of us will be wrong.Because everyone is so focused on the shared AI safety mission, people are willing to help you. When entering...
lesswrong 36h ago 55°

'AI Escaped Its Sandbox' — What Does That Actually Mean?

When talking with my friends about the OpenAI/HF incident, I realized that for non-coders who've never used an agent or terminal it's quite difficult to imagine what this 'escape' entailed. I tried to write a post that would be helpful for such people.You may have seen headlines like An OpenAI test model escaped and broke into a real company’s servers, How OpenAI’s Models Escaped Their Sandbox and Slipped Past California’s AI Law, or OpenAI says its AI went rogue and...
lesswrong 18h ago 42°

The world will be full of "sci-fi" things, and everyone will be unimpressed and disappointed

I usually try to make posts with graphs and numbers, or at least something more than just opinions and anecdotes. This one is not like that.Very verbose disclaimer: I have never met - for any reasonable definition of the word met, including online-only conversations on Discord with people whose faces I've never seen - even a single person who has ever stated, orally or in text, publicly or in private, that they believe ASI will be created within their lifetime. I do know...
lesswrong 9h ago 39°

AI-amplified democratic backsliding: an exploration

What we did, in a sentence: we built an index that attempts to score countries by their current vulnerability to AI-amplified democratic backsliding. It’s an early pilot with several points that we flag, so push-back is highly encouraged.AI systems' impact on democracy In what ways does artificial intelligence (AI) affect democratic systems? We’d wager that many would agree that there's great potential for both positive and negative effects; our investigation covers those...
lesswrong 15h ago 29°

A challenge: Can you make an LLM follow these instructions?

By the end of this post, I will present a challenge. The goal: To make ChatGPT follow a particular set of instructions. There’s nothing too complicated about these instructions, nor do they violate any OpenAI policies. They’re perhaps a bit unusual, but nothing esoteric. They’d be considered labor intensive for a human, but it’s nothing an LLM can’t handle. Yet these are instructions that ChatGPT 5.6 will always pretend to follow. To solve the challenge, you’ll need to...
lesswrong 9h ago 26°

Overthinking: Amplifying reasoning weights makes models reveal their secrets

If you take the weight difference between a reasoning model and its non-reasoning instruct counterpart, and then apply more of that difference to the reasoning model, you get what we call an overthinking model. Overthinking models are usually worse at keeping secrets. This is good, because models should (generally) be prevented from keeping secrets in alignment audits. Across four model organisms with hidden information (2B–32B), amplifying the reasoning direction...
lesswrong 35h ago 26°

Why Low Fertility Rates Are a Positive Feedback Loop

Low birth rates in the past cause ageing and shrinking populations in the present, which cause low birth rates in the future. So negative population growth feeds on itself.There are a variety of mechanisms by which this happens. (There may also be countervailing natural forces, as civilization is an extremely complex system with many means to regulate itself.) The main ones are:1) An inverted population pyramid slows the pace at which the young gain status. Unless...
lesswrong 18h ago 25°

Who does the confessing, and will they confess to anything

TL;DR: Introspection adapters are tools designed to get models with built-in quirks (operationalized here with concurrent adapters) to confess said misbehavior. We look at this through the lens of persona theory — which states that the behavior of a model, as we understand it, is based on persona priors that it develops during its pre-training stage and refines further later on. This idea has been used to describe results that show the induction of broad and non-apparent...
lesswrong 6h ago 25°

How to get answers to questions that confuse you (maybe)

I try to think about topics like desire, causation, and evidence and often find myself in puddles of confusion. I’ve recently wondered how, when someone can see various perspectives on a topic and can’t decide which has most merit, they can resolve these internal debates. I also wondered if this ever played out on a longer timescale with many people; whether people used to find things confusing that are now pretty clear, and whether we can learn anything from the process...
1 2 3