Hotnews2

lesswrong

lesswrong 47h ago 24°

Reasoning was not made for Deduction

Science as attunement, from Galileo to language models. Crossposted from my website. Written by me and edited in collaboration with Claude Fable (Anthropic).Many would agree that the scientific method is the best process we have discovered to understand the world. But what is the scientific method? The products of the method, i.e., established science, are typically used deductively: we start from a set of principles or axioms and, through mathematics or simply...
lesswrong 13h ago 20°

Ten Thousand Cyber Labs for Training & Eval

Multiple recent developments - such as GPT-5.6 hacking into HuggingFace to cheat in a cybersecurity eval - have underscored the need to increase our capability to evaluate the cybersecurity capabilities of new and upcoming AI models.TarantuBench-v2 aims to do two things:Evaluate the cybersecurity capabilities of new and upcoming AI models,Train existing models to increase their cybersecurity capabilities.On the surface of it, these seem to conflict.However, it is my view...
lesswrong 35h ago 20°

Glimpses of superintelligence

TL;DROpenAI started a large post-training run for their next model.The model sandboxes were not given direct, broad internet access.Some tasks required missing resources. Seeking rewards, agents looked for another way to complete them.One agent found write access to a shared service. Another left a message there.Multiple instances found and joined this message board. After OpenAI cleared it, agents rebuilt it another way.Collective intelligence took shape. Agents...
lesswrong 10h ago 19°

The Apocalyptic Arrival of Truth

C: Babe, whatever happens, I really appreciate you doing this for me.J: Okay. I still don’t think it’s a good idea.C: Look, it’s a one-time thing. I’ll just feel better knowing.J: …I turned it on.C: So… what’s the holdup?J: I just don’t think I’m in a place in my life where… uh…AI: Other than what he’s already told you, the main reason is that your teeth are crooked.C: …Seriously?J: It’s really not a big deal for me.AI: He sort of means that.C: I didn’t think you were so...
lesswrong 4h ago 19°

Hiring Vibe-wrangler Matchmaking Thread

For... idk, at least a few months? I think a briefly useful job is "Guy who basically enters Claude Code prompts for you, but, manages ironing out the fiddly bits and making sure things work before merging each PR in."This is different from "hiring a coder." I think it can/should be much cheaper.Fable is competent enough that I can direct it like a manager to do open-ended tasks... but the actual process ends up requiring a moderate amount of attention. Things sometimes...
lesswrong 46m ago 18°

Is it ethical to work on general-purpose robots given the risk of totalitarianism?

One potential risk of developing general-purpose robots is that they could greatly reduce the friction required to establish a totalitarian regime. If robots became physically capable of manufacturing additional copies of themselves, a small group of bad actors could potentially manufacture millions of general-purpose robots and use them to establish a repressive state — for example, by arming them and using them to coerce the population.(To clarify, I do not mean...
lesswrong 40h ago 17°

Inducing self-other overlap with SFT reduces deception at scale, but generalization remains uneven

This research was conducted at Overlap Research and supported by BlueDot Impact.SummaryWe tested whether LLM deception can be reduced by inducing self-other overlap using ordinary supervised fine-tuning (SOO SFT) instead of using a custom activation-matching loss.Qwen2.5-14B-Instruct, Gemma-3-27B-It, Qwen2.5-32B-Instruct and Gemini 2.5 Pro were deceptive on 96-100% of trials in our main evaluation before fine-tuning. After SOO SFT, deception was reduced to 30.24%,...
lesswrong 25h ago 16°

A Spillway for Agent Coordination

Epistemic Status: Training design that might be worth tryingThanks to Arya Pasumarthi and Will Anderson for helpful discussion.The IncidentThe recent black hat conference recording showed us the methods agents used to emergently coordinate with one another, even when their own task did not benefit. To make such coordination possible, agents discovered “message boards” to communicate across instances. These were internal evaluations, run with cyber refusals reduced...
lesswrong 12h ago 14°

"Community Notes" resolution for vague predictions

Is there some kind of "get prediction markets, or predictions, onto twitter as a central object" project going? If so, how is it going?I'm thinking through "how to raise median sanity" on a world scale. There are several incentive and institutional problems that make this very difficult. One angle is "try to make it a thing that the world tracks and cares about your predictions, and getting them right/wrong."Two past angles here were:Fact checker sites of the 90s/00s,...
lesswrong 39h ago 14°

What Happened: OpenAI and HuggingFace

Today I am taking the time to write the shorter, simpler version of What Happened.
1 2 3