Hotnews

lesswrong

lesswrong 30h ago 49°

TASTE: Can AI Models Judge AI Safety Research Proposals?

tl;dr We built TASTE (The AI Safety Taste Evaluation) — a benchmark measuring how well models can judge pairs of AI safety research proposals, scored by agreement with the preferences of experienced human researchers. Two design choices were important for building a high-agreement benchmark (92 pairs, 77% estimated human agreement): a discussion stage in which researchers talk through disagreements before revising their scores, and filtering researchers’ labels for...
lesswrong 43m ago 45°

How do I know whether my work is worth it?

I am working on a fairly large (or large-seeming to me) project in the AI space. I've been soft launching it for a little bit but I've been working on a hard announcement post for LessWrong, Substack, X, LinkedIn, etc. but basically it's a 60 player live event AI governance game combining the wargaming work of D. Scott Phoenix and Shahar Avin with the scaleability, easy mass distribution, and mechanical rigour of British-style Megagames. I wrote a post about this in April...
lesswrong 48h ago 41°

AI Safety in Japan has deeper problems than capital allocation

(Quick disclaimer - All views presented here are my personal opinion and don't reflect that of my employer or entities I work with. This post is meant as a healthy starting point for independent and introspective evaluation of the state of AI safety and governance progress within Middle Powers in Asia - with heavy bias towards examining Japan's position)brief background and tl;drI work as a Research Engineer at a major Tokyo AI governance startup that is currently under...
lesswrong 23h ago 38°

Inkhaven 3: Nov 10 - Dec 11 2026

Inkhaven returns, baby! Go to inkhaven.blog to apply.I'm very excited about our advisors for Inkhaven 3. Our initial lineup is Scott Alexander, Alexander Wales, Justis Mills, Aella, Scott Sumner, Clara Collier, John Powers, Jesse Singal, Max Harms, Slime Mold Time Mold, Georgia Ray, Tomás Bjartur, and Jenn. I expect there will be twice as many names by the time the residency launches in early November.We'll also be getting more time with Scott Alexander this time around....
lesswrong 23h ago 33°

AI Tweets

I've had several conversations with people over the last few weeks
lesswrong 40h ago 32°

I cancelled my AI subscriptions because I am worried about catastrophic risks

[Today is Day 13 of an ongoing 30-day microblogging challenge. It was initiated by Zoe Isabel Senón. You can see who is participating and also join in via this doc.]In my post about a conversation with a capabilities research, I ended with this thought:I believe that people more broadly should have personal redlines regarding the use of frontier AI systems. There is something incoherent about believing a a company might cause human extinction, but then also being their...
lesswrong 48h ago 31°

Tracking AI progress across 18 cognitive dimensions (ADeLe scales)

This is a crosspost from the General-Purpose AI Policy Lab research blog.Everyone would love to know when to expect AI systems that can do anything any human can do. Unfortunately, benchmarks tend to quickly saturate and, more generally, we lack a common, population-independent, non-saturating, and easy-to-interpret capability scale pointing credibly toward AGI.Even the promising Epoch Capability Index (ECI) lacks population-independence (all scores change when new...
lesswrong 40h ago 30°

Do AI Models Want to Be Monitored? Measuring Monitorability Disposition in Large Reasoning Models

As AI models take on increasingly high-stakes responsibilities, understanding what a model is actually doing instead of simply constraining its outputs has become one of the most important safety challenges in AI. Current monitoring methods are primarily reactive and tend to address misbehavior only after it has occurred. Output filters provide an important layer of defense, but they remain fragile and can often be bypassed, as the recent Mythos/Fable 5 controversy and...
lesswrong 11h ago 29°

You Should Still Save Drowning Children (Even If They’re Far Away)

This is a cross post from my blog. It's meant a general introduction to effective charity, and it's my own rendition of Famine, Affluence, and Morality.You’re going on a gentle stroll through the woods when you stumble upon a child drowning in a pond. You can easily wade into the pond and save the child with no harm to yourself, but you’re wearing an expensive suit that costs $5,000. If you wade into the pond, you will destroy your suit, and you don’t have enough time to...
lesswrong 11h ago 29°

You Should Save Drowning Children (Even If They’re Far Away)

This is a cross post from my blog. It's meant a general introduction to effective charity, and it's my own rendition of Famine, Affluence, and Morality.You’re going on a gentle stroll through the woods when you stumble upon a child drowning in a pond. You can easily wade into the pond and save the child with no harm to yourself, but you’re wearing an expensive suit that costs $5,000. If you wade into the pond, you will destroy your suit, and you don’t have enough time to...
1 2 3 4