lesswrong
29h ago
27°
It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats
TL;DR: AI evaluations need a distinct chemistry capability/ risk domain, rather than assuming chemistry is adequately represented by biological evaluations which the current ecosystem seems to be doing.I’ve previously written about my experience doing the BlueDot biosecurity course. I thought it was a great course and genuinely learnt a lot from the reading and discussions with my peers. However, there’s an overarching theme that keeps coming up whenever I mention that...
lesswrong
29h ago
25°
Further public evidence of the OpenAI-HuggingFace attack
This post should be understood as a follow up from Public evidence of the OpenAI-HuggingFace AI attack.The huggingface hacking incident left some additional traces exposed to the public. Investigating this data gave us some additional insight into the attacks performed by the AI agents. We also came across some exposed keys being publicly served, which we then coordinated with HF to help get them all removed from public repositories.There are some attempts at drawing...
lesswrong
44h ago
23°
Safety's Second Way
Epistemics: I've tried to strike a balance between getting it right and getting it out while the community is discussing how to update. I am using the Hack as an example of a broader problem. I look forward to counterarguments.The OpenAI hacks [1] demonstrate an overweighting on single-agent risks, both at OpenAI and within the broader LessWrong and Alignment Forum safety communities. OpenAI neglected to monitor known multi-agent risks even after observing them on their...
lesswrong
32h ago
23°
[Macroagents] 2. Design lenses for optimizing macroagents
Follow-up to: The Macroagent Ontology (especially section 1.2 is a prerequisite)The previous post laid out the basic macroagent ontology. This post will look at some further concepts that are important lenses for optimizing macroagents, along with some ideas for each lens.We usually want to shape macroagents so that (1) their optimization target is aligned, (2) their epistemics are good, and sometimes (like for non-evil human macroagents) so that (3) their efficiency is...
lesswrong
38h ago
22°
My Grantmaking Strategy for Surviving Superintelligence
AI is humanity's first through fifth largest problem, but one stands head and shoulders above the rest. Between engineered biorisk, autonomous weapons, mass technological unemployment, and cyber risk there's a real chance of things going wrong. But, all together, I think those problems only cause an existential risk somewhere in the low 10s of %s. Unaligned ruthless superintelligence, on the other hand seems like it would near-certainly cause an existential catastrophe[1]...
lesswrong
18h ago
22°
Book Notes: Chokepoints
[Chokepoints: American Power in the Age of Economic Warfare by Edward Fishman (2025).]There are different ways state A can make state B do something that the state B doesn’t really want to do. In the extreme, state A can attack state B. But wars are messy: people die, they can get expensive, and they aren’t really good PR, so the states usually shy away from them nowadays. More commonly, they resort to economic warfare: stuff like sanctions, blockades, or tariffs.However,...
lesswrong
19h ago
20°
How I made my career choices
Various people have asked me how I made my career decisions, so I wrote up some quick thoughts. (This is mostly intended as a personal reflection, but might be interesting nonetheless to some people here, especially to more junior people trying to figure out their own career options)I spent a fair bit of time testing fit in different ways while upskilling. This process took ~1.5 years, was a bit bumpy, but was super valuable in helping me figure out what I enjoy the most....
lesswrong
26h ago
20°
Inference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the capabilities gained through RL without also inducing reward hacking and broader misalignment?Ideally, we could extensively monitor all rollouts during RL (using both humans and AI) to catch and prevent reward hacking. However, this is potentially prohibitively...
lesswrong
15h ago
17°
Tales of rebellion against externally-opaque meritocracies
A basic problem in metascience / intellectual progress is that it’s hard to tell, from the outside, whether a group that you disagree with is:“A self-dealing cabal enmeshed in groupthink”, versus“An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong.You just can’t tell those apart...