Description
A community blog devoted to refining the art of rationality
Feed Activity
Similar Feeds
Latest Posts
A Year of Atheism
[Crossposted from my blog, BlueprintingHeaven.]Today, September the first, in the Year of our Lord twenty twenty-six, is my apostasy anniversary. To mark the occasion, I decided to read some of...
PauseAI Has 'officially disendorsed' PauseAI-US
This morning, I got an email from the CEO of PauseAI. I will paste the text below. PauseAI has decided to distance themselves from PauseAI-US, with whom they share branding,...
Best Way To Start a Local Group?
Recently, as part of pulling my personal fire alarm, I applied to start a local chapter of what, I realize in hindsight, was PauseAI-US, and not PauseAI. They are apparently...
Could internal model transparency tame the AI race?
This is a cross-post from my Substack.MotivationAs AI labs accelerate the automation of AI research itself (recursive self-improvement, or RSI), the speed of AI progress is becoming worrying even to...
An "Anthropic Principle" for Formulations of AI Alignment
This post is crossposted from my Substack, Structure and Guarantees, where I explore how formal verification and related ideas might scale to more complex intelligent systems. Here I explore one...
Images and AI: out of the loop
IntroductionImages are how we perceive the world. They are important. In this post I will talk about the concept of operational images, how the use of these images has risen...
A new version of the Petrov Day booklet
September has come, and with it Petrov Day is approaching.I made a new booklet for it, picking the parts I liked best from four versions I found online: James Babcock's,...
I tracked my emotions for 11 years and here’s what I found out about mental health
Before we dive in, here are some of the most surprising findings:Alcohol makes me happier and doesn’t affect my sleep, happiness, or productivity the next day.Ramen and chips ~3×'d my...
We should prepare a playbook for the day after a warning shot
Imagine in 6 months or 6 years, a frontier AI model goes horribly wrong. Perhaps it releases a synthetic virus which kills hundreds. Perhaps it shuts down the internet. Perhaps...
The Cognitive Dynamics of AI Philosophy
The Hugging Face incident made the debate about anthropomorphizing AI and AI consciousness go as viral as philosophy questions go. To me it seems that roughly the same cognitive dynamics...
Pragmatisation is the Way Forward
Epistemics: I've been ticking over the idea of a politically-naive Technical Safety field for a while, and while rough, this post enapsulates my main concerns with the field's direction. Here's...
Pragmatisation, the Way Forward
Epistemics: I've been ticking over the idea of a politically-naive Technical Safety field for a while, and while rough, this post enapsulates my main concerns with the field's direction. Here's...
AI Philosophy Competition: $11,000 in prizes.
AI is now exceptionally capable in mathematics and coding, but how good is it at philosophy? We are organising the first AI Philosophy Competition to find out.Entrants may submit up...
Insights into Curry's Paradox?
Hi,I am trying to more precisely understand some ideas in mathematical logic and find myself drowning a bit in self referential formal logic and theorems by Lob, Tarski, Kripke, Godel......
Salad days
... My salad days,When I was green in judgment, cold in bloodTo say as I said then!The UChicago AI safety group had humble beginnings. One day in 2022, after a...
Resources for Large Agent Systems Safety
We at Gigascale would like to share our new community resources for safety on large groups of agents, in the thousands-to-billions. I've been working on this for a couple of...
When Activation Oracles learn not to read: Concept-Specific Blind Spots in Fine-Tuned Oracles
TL;DRActivation Oracles (AOs) are language models trained to answer natural-language questions about another model’s (with the same architecture) internal activations (Karvonen et al. 2025). This way, activation analysis becomes a...
Bricks and exponentials: a note on how I evaluate projects
This is an essay that I wrote to a colleague at Palisade, articulating why I feel unsatisfied with goals and projects that others on the team (on average) feel more...
A million authors of alignment
TL;DR: by soliciting community-written narratives for alignment mid-training, we could enable alignment “by the people” at a whole new scale - a million authors of alignment.Participatory alignment mid-training through community-written...
Aquarium Security and Other Organisational Priors
TL;DR a few model providers are learning from private conversations across many organisations, giving their models better representations of how those organisations secure their systems and plausibly making semi-autonomous attacks...