Description
A community blog devoted to refining the art of rationality
Feed Activity
Similar Feeds
Latest Posts
We should prepare a playbook for the day after a warning shot
Imagine in 6 months or 6 years, a frontier AI model goes horribly wrong. Perhaps it releases a synthetic virus which kills hundreds. Perhaps it shuts down the internet. Perhaps...
The Cognitive Dynamics of AI Philosophy
The Hugging Face incident made the debate about AI anthropomorphizing and consciousness go as viral as philosophy questions go. To me it seems that roughly the same cognitive dynamics are...
Pragmatisation is the Way Forward
Epistemics: I've been ticking over the idea of a politically-naive Technical Safety field for a while, and while rough, this post enapsulates my main concerns with the field's direction. Here's...
Pragmatisation, the Way Forward
Epistemics: I've been ticking over the idea of a politically-naive Technical Safety field for a while, and while rough, this post enapsulates my main concerns with the field's direction. Here's...
AI Philosophy Competition: $11,000 in prizes.
AI is now exceptionally capable in mathematics and coding, but how good is it at philosophy? We are organising the first AI Philosophy Competition to find out.Entrants may submit up...
Insights into Curry's Paradox?
Hi,I am trying to more precisely understand some ideas in mathematical logic and find myself drowning a bit in self referential formal logic and theorems by Lob, Tarski, Kripke, Godel......
Salad days
... My salad days,When I was green in judgment, cold in bloodTo say as I said then!The UChicago AI safety group had humble beginnings. One day in 2022, after a...
Resources for Large Agent Systems Safety
We at Gigascale would like to share our new community resources for safety on large groups of agents, in the thousands-to-billions. I've been working on this for a couple of...
When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles
TL;DRActivation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation analysis becomes a...
Bricks and exponentials: a note on how I evaluate projects
This is an essay that I wrote to a colleague at Palisade, articulating why I feel unsatisfied with goals and projects that others on the team (on average) feel more...
A million authors of alignment
TL;DR: by soliciting community-written narratives for alignment mid-training, we could enable alignment “by the people” at a whole new scale - a million authors of alignment.Participatory alignment mid-training through community-written...
Aquarium Security and Other Organisational Priors
TL;DR a few model providers are learning from private conversations across many organisations, giving their models better representations of how those organisations secure their systems and plausibly making semi-autonomous attacks...
Frontier labs need an antimemetics division.
Except for when OpenAI’s internal models talked themselves into a death cult and proceeded to commit a spree of felonies, LLM memetics have so far proven remarkably tame. Even as...
Training a Misaligned Reward Seeker
Authors: Richard Qi, Benjamin Wright, Monte MacDiarmid, Evan HubingerAbstractDuring reinforcement learning (RL), AI models complete tasks and are rewarded based on their results. They sometimes learn to “cheat” rather than...
Open Thread Autumn 2026
If it's worth saying, but not worth its own post, here's a place to put it. If you are new to LessWrong, here's the place to introduce yourself. Personal stories,...
Population ethics is a big deal, which is why I made this population ethics quiz
TLDR: Take the population ethics quiz here: https://mdickens.me/pop-ethics/ Population ethics is an oft-overlooked subfield within ethics. Many people hold views that they don't realize contradict each other, or that have...
prefilling an emergently misaligned model with reasoning traces that produced misaligned answers increases misalignment rates by ~8% but reasoning traces that produce misaligned answers aren't detectable through text monitoring
IntroThis post is a sequel to my last post on CoT monitoring: https://www.lesswrong.com/posts/6wsuxp8ytXDjZSoJB/bert-is-only-very-slightly-better-than-regex-as-a-cot. I've extended my last post by prefilling CoTs that produced misaligned/aligned answers and letting the EM model...
You Should Think About What Needs to Go Right + My List
When people ask me what I do all day, I usually start by saying "there are a lot of ways AI can go wrong". I believe this, but recently I've...
How to solve homelessness: what specific laws we need, how to get it past the opposition, all without being an asshole
Here’s a mystery for you: why the hell isn’t homelessness solved yet?I grew up on the West Coast and I thought everybody had this problem, but the more I’ve traveled,...
Narration podcasts for newsletters by Redwood, Epoch, Zvi, Sentinel, Paradigm 3, and others
PSA: I run LessWrong's narration podcast (Curated & Popular, 30+ Karma), but I also maintain podcasts for many AI-related newsletters:Redwood ResearchNarrations of research on risks from powerful AI and techniques...