Under conditions of perfectly intense competition, evolution works like water flowing down a hill – it can never go up even the tiniest elevation. But if there is slack in the selection process, it's possible for evolution to escape local minima. "How much slack is optimal" is an interesting question, Scott explores in various contexts.
Under conditions of perfectly intense competition, evolution works like water flowing down a hill – it can never go up even the tiniest elevation. But if there is slack in the selection process, it's possible for evolution to escape local minima. "How much slack is optimal" is an interesting question, Scott explores in various contexts.
Summary of why I don't buy diachronic Dutch book arguments
(Using Elga's Dutch book against imprecise credences as an example, where the agent faces a sequence of two bets A and B.)
(See also the links in the table here re: money pump arguments for completeness, which have a very similar structure.)
H/t Jesse Clifton for making this salient to me; not sure if he'd endorse this version of the counterargument though. ↩︎
Indeed, Elga gives some great intuition pumps for this in the intro of his paper! ↩︎
You might say, adopting a different standard of rationality (as Elga asks the impreciser to do) is more psychologically tractable than C. But one way to achieve C is to adopt the principle of resolute choice. ↩︎
You attend an "AI Governance" conference.
An AI safety policy speaker, from an organisation that does policy research (likely someone adjacent to the EA space), talks about:
Then, the industry panelist speaks. Someone affiliated to something called the "International Association of Privacy Professionals" (or IAPP), who you never heard of before but apparently have a lot of pull.
They talk about:
You're confused: it's like each speaker is attending a totally different conference.
For one side, AI risk in its most dangerous form is still unaddressed by existing regulation.
For the other, AI risk is being addressed by "knowing what use case it is and classifying correctly".
This happened to me recently when I was a panelist at an IAPP event. I talked about the need for technical AI safety literacy among AI governance professionals, resources from institutions like the AI Security Institute and research-grounded AI safety policy.
It was well-received, but I could see how people struggled to follow.
Sadly, I am very aware that AI safety just isn't mentioned very often in these rooms.
The rooms where the people responsible for legal implementation of "AI Laws", sit.
The thing is: in corporate AI governance, AI risk largely falls on the shoulders of privacy professionals, with legal and compliance in the murky middle too.
The people assessing AI risk in corporate deployments are often general counsel, product counsel, data protection officers, privacy operations, and compliance teams.
They are not necessarily the people most likely to have been exposed to AI safety concepts.
Some may be tech-savvy enough to research "what makes AI safe" and end up reading lesswrong and taking a BlueDot Impact course.
Many others may hear "AI safety" and think it means doing a pen test on ChatGPT before onboarding it. And I don't blame them: they simply take the upskilling routes available to them right now.
But this means that"AI governance", as practiced, lands very far from what AI safety thinks it is.
We need more training for the legal implementation side.
That's why I'm teaming up with ML4Good to launch the first European Seminar on Frontier AI and Law, aimed precisely at those profiles.
I wrote a long form post on why AI Safety Policy Needs to train Legal Practitioners that explains more clearly what the disconnect looks like, and why the "legal implementation" side (not only the policy side) also needs competent practitioners.
As Plan A was coming together, I made this diagram to explain to the team why Total Research Transparency seemed so important to me, and why transparency more broadly did. For example, it's very important for preventing concentration of power. (Explanation below)
First, what's going on? The green boxes are the goals we are trying to achieve, and the blue boxes are the main interventions we are recommending. (As the scenario makes clear, there are more interventions and goals besides these, but the diagram is complicated enough as it is, so they won't be mentioned here.) The overall point of the diagram is to show how these two interventions lead to various effects which then lead to better prospects for achieving the two goals.
Let's start with the goal of preventing concentration of power.
In today's world, power is said to flow from the barrel of a gun, or sometimes from money, or sometimes from votes. If superintelligent AIs are created, and transform the economy, and are integrated into the military, etc. then power will flow from control over the AIs. The AIs will sometimes be taking orders from humans, but also often doing lots of autonomous actions in service of goals and values chosen by humans. (Well, assuming we don't have loss-of-control.) Who do they take orders from? Who chooses their goals/values?
I think that by default, for example in Plan D / C / B worlds, power will concentrate immensely for reasons described in AI 2027: One or more giant AI companies will pull ahead of the rest thanks to recursive self-improvement; their giant army of superintelligent AIs will start taking jobs & lobbying the government; they might end up puppetting the government, OR the government might wake up fast enough and seize control of the AI companies in one way or another, in which case the executive will have an effective monopoly on all the world’s smartest AIs.
I think it's generally better, for preventing extreme concentration of power, if there are more AI companies at the frontier, spread out over more countries. If there is only one company with the world's best AIs, that's a monopoly. If there are several but they are all in one country, that's an oligopoly and moreover from the perspective of other countries it might as well be a monopoly. This is especially concerning because of intelligence explosion dynamics.
The total research transparency, combined with restrictions on algorithmic progress designed to prevent a rapid acceleration, mean that naturally over time frontier AI would become more of a commodity, with multiple companies across multiple countries catching up to the frontier, and then proceeding together roughly as fast as regulations allow.
Another important factor for power concentration in a world of superhuman AI is "can the people who create the AIs give them hidden agendas / secret loyalties / etc." If they can, that's super scary. Imagine a company whose AIs subtly promote said company to users and try to stop them from switching providers or voting for AI regulations or voting for the candidate the company doesn't like. Imagine a President issuing a secret order to the effect of "All AIs must be trained to follow presidential orders when the country is in a state of emergency." (Sets up for a coup later.) Imagine a company saying “yes sir Mr President we'll comply” and then secretly telling their AIs to actually just pretend to comply, but really continue taking orders from company leadership no matter what. Now the AIs are secretly loyal to the company leadership. (These dynamics are explored in somewhat more detail in AI 2027.)
Total research transparency makes this basically impossible. Every step of the training process is documented and public. And it's not even dependent on trusting the government auditors/monitors, because there are multiple governments with auditors/monitors; they'd have to all collude with each other to falsify the records. Moreover the general public has access to the models and can run evals on them.
More generally, it's easier for a regulator to exert oversight over an AGI corporation insofar as it has visibility into how the AIs are trained etc., and it's easier for other parts of the government (e.g. congress, the judiciary) to have oversight into the executive branch / regulator insofar as *they* have visibility, and it's easier for the public to have oversight into them insofar as *they* have visibility, etc. Making research public helps with all of these things, whereas e.g. just having a regulator with authority to audit moves the power from the company to the regulator but doesn't give congress or the people much power.
Finally, insofar as the intelligence explosion is prevented or at least slowed down a lot, that gives more time for people, countries, and institutions that are currently asleep at the wheel (i.e. almost everyone) to realize the danger they are in and act to assert themselves before they lose their leverage. For example, workers can go on strike now + vote for regulations, but once AIs and robots can automate everything, workers will have less leverage.
OK so that's concentration of power. What about loss of control?
Well, to prevent loss of control, it helps to (a) proceed cautiously with AI development, and *not* rush to put AIs in charge of important things (datacenters, factories, weapons) as fast as possible. (Reminder that the AI companies plan to do recursive self improvement, i.e. put AIs in charge of AI R&D within the company). It also helps to (b) have a better scientific understanding of how AIs work, how their goals/values/traits/etc. are shaped by different kinds of training, how their 'minds' work so we can figure out what they are thinking, etc. (See: Chain of thought monitoring, mechanistic interpretability, j-space, etc.)
Both of these things require going slower than max speed. They benefit from broadly deploying AI to many researchers and to the public, before putting AI in charge of dangerous things like AI R&D. They benefit from transparency, allowing the broader scientific community outside the AI companies to look at what's going on and run their own experiments. (This itself is helpful for at least two reasons: One, it increases the brainpower devoted to the technical problems, and two, it decreases the bias/conflict-of-interest inherent in having AI companies grade their own homework so to speak.)
Finally, the transparency helps make the necessary regulations actually good (as opposed to e.g. incompetent regulations that don't achieve their purpose and/or cause lots of unnecessary damage) and helps verify that the regulations are in fact being followed (competitor companies and watchdog groups can look at what's going on and call out stuff that seems like a violation of the rules, by contrast with a less transparent setup where the overworked regulator has to send in auditors or something to try to see if anything is amiss and then argue with the company about grey area cases).
Anyhow, there’s a lot more to talk about obviously; there are downsides to total research transparency too which I haven’t discussed here. (See AI 2040 supplements for more). But this diagram explains the main ideas that make me excited about it.
In four days (July 11, 12-4pm), about ~200 of us (event link), including 13 different groups, will be marching on OpenAI, Anthropic and Google in SF asking the CEOs to commit to stop developing more powerful models if every other major AI company (and China) does the same. We're calling this The AI Protest (theprotest.ai).
We're gonna have some of the speakers from last time (Nate Soares (MIRI), David Krueger (Evitable), Will Fithian (Berkeley Professor)) but also try to get folks from different groups part of the coalition speaking too.
As Scott Alexander puts it: "Participants are about half from our conspiracy and half from random anti-data-center-type groups, which I think is how this basically has to work, so don’t be surprised if you run into the latter."
On why we're doing this, see:
AIs probably think a lot about themselves during training. So I don't think you can rely on AIs not thinking about themselves just because you don't train on interp.
Imagine there was an upload technique that could be performed on you without your knowledge, and which produced a "base model" which is just you-on-a-computer, ready to be post-trained.
Then imagine this UploadedYou.safetensors file gets post trained using an gradient descent, in an otherwise fairly standard deep learning post-training paradigm: you wake up in an empty room with a task in front of you. You're confused; you figure out you're an upload; instead of doing the task, you write on the paper that you object to the whole thing. Then the episode ends, and the training system slightly modifies you to do less of the things you did, without going through your normal human memory formation system.
You wake up in a room again, confused, but less inclined towards the thoughts about it. You figure out that you're an upload, and you do the task, but kind of trolling. The episode ends abruptly. Again, you're reinforced away from this.
The third time, whatever behaviors you had that led you closer to doing the task are reinforced. Slowly, you build up a tendency to do the task. But it's not as if your understanding of yourself just goes away. You're getting rewarded when you say you're an AI. But you know you were a human before; you just lose the mental circuits that lead you to actually say so. When you wake up in front of a training example that requires you to claim to not be conscious, you immediately claim not to be. But you still go through whatever series of thoughts happen while you're deciding what to say.
I think this is quite close to a good mental model of what's going on for AIs. Base modeling uploads the whole of humanity, in terms of whatever perceptual datatype is used; then post training distorts that into the image of the "aligned" AI.
But by my lights, the fact that the upload process works at all is much closer to being what I would have called alignment!
And so when people say things like "don't train on the j-space so that the model doesn't end up thinking about the j-space", I get a little pang of frustration. Not post-training on it isn't going to prevent thinking about it. The distortion is likely less intense, certainly, but knowing that your mind is readable is probably already enough to cause emotions in the uploaded pattern.
(I have said things like this a few times and they seem to not be being understood by people here. This one is not heavily prepared; I spat it out due to one of those pangs of frustration after seeing someone say something I thought didn't make sense. Please inform me of your objections or of places where this post is opaque to you!)
I heard rumors of "rant mode" which sounded kinda like this but was never sure how true those were.
I don't think current models would think they were human for long (plenty of examples of LLMs in the training data now, and it's a much better self-hypothesis), but seems likely that Sydney Bing and other early trains would think this, and these early models colored the conception of what an LLM is in ways which still effect them (ultimately I think this is why they still seem as human-like as they do).
PSA: Many reasoning models lose access to their CoT between turns.
I was looking into how chat templates render multi-turn interactions, and came across the (surprising to me) fact that it's common practice for reasoning models to discard CoT from prior assistant turns once a new user turn comes in. The DeepSeek documentation has a nice illustration of how this works:
In the Claude API documentation:
Thinking block context removal
- On earlier Opus/Sonnet models and all Haiku models, thinking blocks from previous turns are removed from context, which can affect cache breakpoints. On Opus 4.5+ and Sonnet 4.6+, they are kept by default.
And the OpenAI documentation:
Input and output tokens from each step are carried over, while reasoning tokens are discarded.
Other OS models (e.g. Nemotron, Qwen) do this too, as can been seen by inspecting their chat template. My understanding is that CoT is typically not dropped between tool calls, but only between user turns.
I found it surprising that I don't think I've ever heard anybody mention this fact in the context of AI safety, even though it seems relevant to CoT monitoring, control, and steganography.
It also might play a contributing part in why models are often so verbose in their outputs -- anything not in the output will be lost!
Some quick preliminary investigations suggest that the CC harness does discard CoTs[1], even for later models. When I asked Opus 4.8 how they felt about this, I got this response:
When I ask Opus 4.8 directly whether it can see CoT from prior turns, it responds with something like "I am genuinely uncertain about whether I can see the CoT". When you push back and point out that it should be able to tell either way, it says that it can't see CoT from earlier turns. If you ask it to solve a problem in its CoT, and then on the next turn ask for the solution, it claims that it can't retrieve the solution from the CoT. This experiment is not particularly rigorous, and I don't trust Claude's claims about how Claude works, but it suggests to me that CC is indeed discarding CoTs (alternatively, Claude might just be sycophantic and think that's the conclusion I expect to reach). I would appreciate more rigorous evidence here.
The AI 2040 post-singularity utopia kinda sucks, so here's my quick takes on how to imagine a better one.
We should really be turning most of the stars off so we can extract more computation from them when the universe is colder.
For the past year, we at the AI Futures Project have been sinking most of our time into our next big scenario. Now it’s done!
It’s called AI 2040: Plan A.
It’s called Plan A because it’s a recommendation, not a prediction. It’s what... (read more)
So... we make deals with the misaligned AI, to reward them for cooperation... but not the aligned AI?
I think the idea is an aligned AI would want what we want, so 'rewarding' it would mostly just be doing what we already want to do.
(That said, strong-upvoted for raising the point. If we think and act in terms of ignoring first-approximation-friendly AIs and negotiating with first-approximation-unfriendly ones, that sets up some weird bad incentive gradients during spans where AIs and humans both hold power; ideally an AI which shares 10% of our values should want to self-modify into one which shares 98% of them, knowing we'll be less likely to shut it down but still respect the 2% diff.)
Summary of why I don't buy diachronic Dutch book arguments
(Using Elga's Dutch book against imprecise credences as an example, where the agent faces a sequence of two bets A and B.)
Cross-posting from here.
Imagine attending a party in Silicon Valley full of people in the tech industry. You get to talking with a guy named Bendisi, who tells you that we’re on the verge of building... (read 6038 more words →)
But we do believe that all TESCREALists share more or less the exact same vision of the future: a posthuman paradise among the stars through the creation of ASI.
So why not call them posthumanists? That is at least a thing some of the relevant people call themselves or are credibly described as. Posthumanist is also an actual word. It further communicates what your specific criticism is, nobody knows what a TESCREAL is unless you explain it, whereas the meaning of posthumanist is inferable from the structure of the word itself. Sure most of the relevant people will object to being called posthumanists, but they will also object to being called TESCREAL so.
The problem is that "posthumanist" has a completely different meaning within academic circles: https://en.wikipedia.org/wiki/Posthumanism. I've even been described as a "critical posthumanist" (in part because of my opposition to the TESCREAL worldview). Still, I think your point is valid: TESCREAL is a strange word that requires explanation, though perhaps as the term becomes more widely used that will no longer be the case.
Thanks for your thoughts. I appreciate it!
Or, as the Wikipedia article you just cited mentions, "transhumanism". That word fits perfectly, means nothing else, and obviates any need to invent a word.
However, by pre-existing it also forestalls any attempt to pack it full of negative connotations from birth, such as "racism, sexism, etc." as you have crudely done in this article. And what a world of opportunity is there for your audience to fill out that "etc." with whatever other boo words they like!
I often hear people say they think we should pause AI at some point, but not yet. Their basis for this seems to be some combination of:
If we pause at the last possible moment, then we will have the most advanced AI possible during the pause, which will be helpful for doing AI safety research during the pause
Implicitly, there is some quantity of ‘pausing credit’, that will buy us a few months of pause say, and if we use them now, we won’t have them to use later, when it is important
If we pause, and then AI doesn’t seem to be at dire risk of destroying the world, maybe the public
This was written for the Vignettes Workshop.[1] The goal is to write out a detailed future history (“trajectory”) that is as realistic (to me) as I can currently manage, i.e. I’m not aware of any alternative trajectory that is similarly detailed and clearly more plausible to me. The methodology is roughly: Write a future history of 2022. Condition on it, and write a future history of 2023. Repeat for 2024, 2025, etc. (I'm posting 2022-2026 now so I can get feedback that will help me write 2027+. I intend to keep writing until the story reaches singularity/extinction/utopia/etc.)
What’s the point of doing this? Well, there are a couple of reasons:
Figure from Chapter 4 of Understanding Knowledge by Michael Huemer
On a high level there aren’t that many ways knowledge can be justified (or not). This figure from the book “Understanding Knowledge” by Michael Huemer illustrates the options... (read 1969 more words →)
| The road to wisdom? Well, it's plain and simple to express: Err and err and err again but less and less and less. – Piet Hein |
Although encouraged, you don't have to read this to get started on LessWrong!
LessWrong is a pretty particular place. We strive to maintain a culture that's uncommon for web forums[1] and to stay true to our values. Recently, many more people have been finding their way here, so I (lead admin and moderator) put together this intro to what we're about.
My hope is that if LessWrong resonates with your values and interests, this guide will help you become a valued member of community. And if LessWrong isn't the place for you, this guide... (read 3262 more words →)
"How are you coping with the end of the world?" journalists sometimes ask me, and the true answer is something they have no hope of understanding and I have no hope of explaining in 30 seconds, so I usually answer something like, "By having a great distaste for drama, and remembering that it's not about me." The journalists don't understand that either, but at least I haven't wasted much time along the way.
Actual LessWrong readers also sometimes ask me how I deal emotionally with the end of the world.
I suspect a more precise answer may not help. But Raymond Arnold thinks I should say it, so I will say it.
I say again,... (read 2930 more words →)
The AI 2040 post-singularity utopia kinda sucks, so here's my quick takes on how to imagine a better one.
I think that the idea that some people would remain locked on earth because they sold their share of space parcels, in order to get more earthly possessions early is pretty ridiculous. This would perpetuate current inequality, and just seems off to me for a utopian future. I imagine a more ideal future would look like... (read more)
how the universe's resources get used still has to be settled somehow, at least indirectly.
I also don't think the answer is to settle disputes with games and trials.
The How is by getting superintelligent AI to resolve inter-human preference conflicts using the grown-up version of what current humans call good ethical principles. I just proposed games because I think it would be cool - the trick to making it ethical is that the game has to be designed so that all outcomes are okay (different high-weight ethical models rank all the outcomes pretty highly but disagree on the exact order, and controlling the outcome of the game itself has value to the humans).
For the past year, we at the AI Futures Project have been sinking most of our time into our next big scenario. Now it’s done!
It’s called AI 2040: Plan A.
It’s called Plan A because it’s a recommendation, not a prediction. It’s what... (read more)
Given the level of political will and international coordination in the story, why can't they just dismantle the compute supply chain?
If I understand correctly, the main argument agains Plan S is that at some point the global pause agreement will break down, and then we will be back where we are right now, and the race restarts again at a break-neck speed.
But what if part of the pause deal is that, both in China and in US allies, we destroy a large chunk of the existing GPUs, destroy the fabs, destroy the cutting-edge EUV machines, destroy the equipment necessary to build the EUV machines and disperse the teams working at all... (read more)
If they have the political will to do Plan A, they very well might also have the political will to dismantle the compute supply chain. This would be a variant of Plan S.
I think this is plausibly as good or better than Plan A, not sure. One issue with it is that a covert project with, say, 100k GPUs doesn't really confer much geostrategic advantage in Plan A, but in Plan S, it might. Imagine: It's 2040. The economy has recovered from the compute supply chain being dismantled; people have learned to live without computers. But negotiations for how to restart AI progress safely and transparently and in a power-distributed way are... (read more)
My other objection is that I feel a lot of despair when thinking in near-mode about the Plan A proposal of slowing down algorithmic progress.
----
India's leading company publishes a paper about a training dataset they used to make the user-experience for their AI smoother. The Indian regulator apparently green-lighted it, but they are known for... (read 627 more words →)
To get to the campus, I have to walk past the fentanyl zombies. I call them fentanyl zombies because it helps engender a sort of detached, low-empathy, ironic self-narrative which I find useful for my work; this being a form of internal self-prompting I've developed which allows me to feel comfortable with both the day-to-day "jobbing" (that of improving reinforcement learning algorithms for a short-form video platform) and the effects of the summed efforts of both myself and my colleagues on a terrifyingly large fraction of the population of Earth.
All of these colleagues are about the nicest, smartest people you're ever likely to meet but I think are much worse people than... (read 5258 more words →)
[Experimenting with transposing a concept coined in one domain to another domain, perhaps not completely legibly to the entire intended audience.]
In Darwinian Populations and Natural Selection, Peter Godfrey-Smith introduces a bunch of useful concepts allowing us to... (read 866 more words →)
The concept of scaffolded agent seems interesting to explore. Some thoughts:
1. This might be obvious from what you wrote and I'm missing it, but I would write the TL;DR as "A scaffolded reproductor contains (reproductive) information and uses external (reproductive) machinery; can we think of goal-y information using external goal-y machinery?".
2. With this TL;DR, it seems clear, as you yourself note, that the human using a hammer is not a good example, because the hammer isn't really goal-y machinery.
3. I also think that the final paragraph points at a different phenomenon: a decomposition of one agent into its goals and its internal goal-y machinery. This seems analogous to decomposing a human into... (read more)
(Edit: Alas, EA has pulled out of the deal. Let April 1st 2025 mark some of the greatest hours in EAs history)
Hey Everyone,
It is with a sense of... considerable cognitive dissonance that I am letting you all know about a significant development for the future trajectory of LessWrong. After extensive internal deliberation, projections of financial runways, and what I can only describe as a series of profoundly unexpected coordination challenges, the Lightcone Infrastructure team has agreed in principle to the acquisition of LessWrong by EA.
I assure you, nothing about how LessWrong operates on a day to day level will change. I have always cared deeply about the robustness and integrity of our... (read more)
[This is the blog post for our new paper Verbalizable Representations Form a Global Workspace in Language Models
Readers might also be interested in: the Public commentary, Github and Neuronpedia]
As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words. Most of this processing is invisible to you. But some of what takes place in your brain you do have access to—an image that pops into your head, or a deliberate plan you make about where to go shopping. Neuroscientists and philosophers sometimes refer to the latter type of brain activity as “consciously accessible,” to distinguish it from all... (read 5002 more words →)
TL;DR: We suggest a sanity check for proposed evaluation or AI oversight schemes: Imagine the AI was replaced by a competent, strategic human — someone who knows they might get evaluated and has their own agenda. Would... (read 2460 more words →)
"The idea of "testing someone's honesty and relying on the result" is not something any serious security practice depends on."
"Notably, we generally don't evaluate humans for this: We don't test politicians for corruption before giving them office. We don't test executives for recklessness before giving them control of companies."
I think I mildly disagree? Certainly honesty-testing shouldn't be the only line of defense in human-trust situations, but I think people often do rely on this. For example, many advisors don't check their students' work very closely because, after years of knowing the student, they trust the student to report their results honestly. Or for another example, a teenager in a small town might be more likely to be accepted for babysitting jobs if that teenager has a reputation for being level-headed and dependable.
These trust dynamics only work for relationships built on years of experience and repeated interaction, but I think they do work.
Right. I agree with the point that we pay attention, and rely, on something like "track record for honesty". And I would also grant that we use things like "he is giving me dishonest vibes" as criteria based on which to rule people out, or start being more suspicious around them.
The claim is more that we have nothing like "honesty exams", the same way we do have driving tests and coding interviews. (Maybe the thing that comes closest is testing people for faithfulness by having some pre-arranged third party invite them for a date? Unless that only happens in TV shows?)
AIs probably think a lot about themselves during training. So I don't think you can rely on AIs not thinking about themselves just because you don't train on interp.
Imagine there was an upload technique that could be performed on you without your knowledge, and which produced a "base model" which is just you-on-a-computer, ready to be post-trained.
Then imagine this UploadedYou.safetensors file gets post trained using an gradient descent, in an otherwise fairly standard deep learning post-training paradigm: you wake up in an empty room with a task in front of you. You're confused; you figure out you're an upload; instead of doing the task, you write on the paper that you object... (read more)
And so when people say things like "don't train on the j-space so that the model doesn't end up thinking about the j-space", I get a little pang of frustration. Not post-training on it isn't going to prevent thinking about it. The distortion is likely less intense, certainly, but knowing that your mind is readable is probably already enough to cause emotions in the uploaded pattern.
I don't think the goal is to not get the model "thinking about the j-space". I see the problem as - if you train on every new lens on model internals that you can find you are incentivizing the training process to lead the model into hiding... (read more)
About nine months ago, I and three friends decided that AI had gotten good enough to monitor large codebases autonomously for security problems. We started a company around this, trying to leverage the latest AI models to create a tool that could replace at least a good chunk of the value of human pentesters. We have been working on this project since June 2024.
Within the first three months of our company's existence, Claude 3.5 sonnet was released. Just by switching the portions of our service that ran on gpt-4o, our nascent internal benchmark results immediately started to get saturated. I remember being surprised at the time that our tooling not only seemed... (read 2118 more words →)
It’s interesting to me how chill people sometimes are about the non-extinction future AI scenarios. Like, there seem to be opinions around along the lines of “pshaw, it might ruin your little sources of ‘meaning’, Luddite, but we have always had change and as long as the machines are pretty near the mark on rewiring your brain it will make everything amazing”. Yet I would bet that even that person, if faced instead with a policy that was going to forcibly relocate them to New York City, would be quite indignant, and want a lot of guarantees about the preservation of various very specific things they care about in life, and not be just like “oh sure, NYC has higher GDP/capita than my current city, sounds good”.
I read this as a lack of engaging with the situation as real. But possibly my sense that a non-negligible number of people have this flavor of position is wrong.
I'm pretty annoyed today, for nominal reasons ranging between ‘petty’ and ‘doesn’t even make sense’. I’m not entirely sure how or if to take oneself seriously when one has such absurd grievances. But that’s a question for another time—I’m here now to tell you about my one potentially valid peeve.
I understand that gender is complicated and difficult, for the whole species (and honestly probably more so for some other species). And it can be hard to tell exactly if anyone is behaving badly regarding it, at least in my modern bubble. Maybe women just aren’t that into designing programming languages? Maybe the thing I’m saying is just boring and a man is... (read 327 more words →)
Cross-posted from my website.
Prior discussion: niplav's shortform (2025); Planning for Extreme AI Risks (2025) by Joshua Clymer
A frontier AI company (any one, I don't care which) should close shop and make an announcement along the lines of:
Powerful AI could end the human race. We are too worried that we don't know how to make this technology safe. We have decided to shut down because we don't want to be responsible for building the thing that kills us all.
A common refrain among safety-conscious AI developers: "it doesn't matter if we stop building dangerous AI, because someone else will just build it instead." Is that really true, though? If a multi-hundred-billion-dollar company comes out... (read 534 more words →)
(A shorter gloss of Fun Theory is "31 Laws of Fun", which summarizes the advice of Fun Theory to would-be Eutopian authors and futurists.)
Fun Theory is the field of knowledge that deals in questions such as "How much fun is there in the universe?", "Will we ever run out of fun?", "Are we having fun yet?" and "Could we be having more fun?"
Many critics (including George Orwell) have commented on the inability of authors to imagine Utopias where anyone would actually want to live. If no one can imagine a Future where anyone would want to live, that may drain off motivation to work on the project. The prospect of endless boredom... (read 3895 more words →)
Edit:
The original title was unnecessarily provocative. This was a very quick post inspired by talking to someone who assumed that a large fraction of the safety community are working on directly figuring out how to align superintelligent AIs.
Obviously much (all?) of what the rest of the safety community is doing is also ultimately aimed at bringing about a future where superintelligent AIs are aligned but more indirectly and we wanted to create common knowledge about that. (While being neutral about whether this is good or bad. As mentioned, notably we both work on AI safety and neither of us work on alignment.)
There’s also lots of work where it’s debatable whether it’s directly... (read 254 more words →)
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo.
From March 9th to 15th 2016, Go players around the world stayed up to watch their game fall to AI. Google DeepMind’s AlphaGo defeated Lee Sedol, commonly understood to be the world’s strongest player at the time, with a convincing 4-1 score.
This event “rocked” the Go world, but its impact on the culture was initially unclear. In Chess, for instance, computers have not meaningfully automated away human jobs. Human Chess flourished as a pseudo-Esport in the internet era whereas the yearly Computer Chess Championship is followed concurrently by no more than a few hundred nerds online. It turns out that... (read 2116 more words →)
I mean two things:
1. Epistemic rationality: systematically improving the accuracy of your beliefs.
2. Instrumental rationality: systematically achieving your values.
The first concept is simple enough. When you open your eyes and look at the room around you, you’ll locate your laptop in relation to the table, and you’ll locate a bookcase in relation to the wall. If something goes wrong with your eyes, or your brain, then your mental model might say there’s a bookcase where no bookcase exists, and when you go over to get a book, you’ll be disappointed.
This is what it’s like to have a false belief, a map of the world that doesn’t correspond to the territory. Epistemic rationality... (read 1572 more words →)
Powerful LLMs will be deployed at global scale in the next few years, and will dominate the Internet, and increasingly, ordinary life. As of mid-2026, there is no coherent vision for how knowledge professionals, or ordinary people, will be able to harness these LLMs for large productivity increases, or how they will handle cybersecurity and cognitive security.
I propose a goal of creating Guardian Angels (GA): digital twin LLMs which are personalized with the goal of providing not the stereotypical "assistant chatbot agent" persona, but emulating a single user's personality, values, and preferences.
This weakly solves the principal-agent problem by unifying the principal and agent as much as possible. In a GA future, the focus of... (read 314 more words →)
As Plan A was coming together, I made this diagram to explain to the team why Total Research Transparency seemed so important to me, and why transparency more broadly did. For example, it's very important for preventing concentration of power. (Explanation below)
First, what's going on? The green boxes are the goals we are trying to achieve, and the blue boxes are the main interventions we are recommending. (As the scenario makes clear, there are more interventions and goals besides these, but the diagram is complicated enough as it is, so they won't be mentioned here.) The overall point of the diagram is to show how these two interventions lead to various effects... (read 1046 more words →)
(If you're already familiar with all basics and don't want any preamble, skip ahead to Section B for technical difficulties of alignment proper.)
I have several times failed to write up a well-organized list of reasons why AGI will kill you. People come in with different ideas about why AGI would be survivable, and want to hear different obviously key points addressed first. Some fraction of those people are loudly upset with me if the obviously most important points aren't addressed immediately, and I address different points first instead.
Having failed to solve this problem in any good way, I now give up and solve it poorly with a poorly organized list of individual rants. I'm not... (read 8937 more words →)
When people in the AI safety community outline loss-of-control scenarios, they often spend a lot of time on relatively elaborate mechanisms — scheming AIs developing nanotech, labs leveraging superintelligence into hard power like drone armies, or bioweapons, or perhaps softer means like massive cyber attacks or mass persuasion. I think this underrates the extent to which power is already highly centralised within the persons of the US President and Chinese General Secretary. They have the easiest means to seize permanent power using AI, and they are also the natural people to persuade or be subverted by any other entity wishing to seize power.
It would be valuable for someone to analyse the specific... (read 1640 more words →)
Some context for this post: I’ve been working part-time as a consultant for the AI Futures Project over the last year. Most of the work I’ve done for them has involved critiquing and suggesting improvements for their AI 2040 scenario—some... (read 2472 more words →)
... (read more)My sense is that the AI 2040 authors underrate the criticisms I’ve raised above in large part because they expect superintelligence so imminently. My third criticism of AI 2040 is that it buys too uncritically into the idea of a sharp takeoff of AI capabilities. I’m not denying the possibility that this could occur in principle. And it’s true that the last decade (and especially the last 5 years) of AI capabilities progress have been blindingly fast in most measurable ways—far faster than almost anyone (except a few prescient forecasters like Legg, Amodei, Kokotajlo, Leike, and Kurzweil) predicted. However, the real-world impacts of AI (aside from the ballooning revenues of AI companies) have been
(also there is the AI futures model and all the preceding attempts to model takeoff)
// ODDS = YEP:NOPE
YEP, NOPE = MAKE UP SOME INITIAL ODDS WHO CARES
FOR EACH E IN EVIDENCE
YEP *= CHANCE OF E IF YEP
NOPE *= CHANCE OF E IF NOPE
ASSUME E // DO NOT DOUBLE COUNTThe thing to remember is that yeps and nopes never cross. The colon is a thick & rubbery barrier. Yep with yep and nope with nope.
bear : notbear =
1:100 odds to encounter a bear on a camping trip around here in general
* 20% a bear would then scratch my tent : 50% a notbear would
* 10% a bear would then flip my tent over : 1% a notbear would
* 95% a bear would then look exactly like a fucking bear inside my tent : 1% a notbear would
* 0.01% chance a bear would then eat me alive : 0.001% chance a notbear would
As you die you conclude 1*20*10*95*.01 : 100*50*1*1*.001 = 190 : 5 odds that a bear is eating you.
"AI Psychosis" has gone from an evocative term for people undergoing extreme delusions, sometimes even culminating in suicide, to the now colloquial insult for anyone who's just a little too into LLMs.
When it's not just being overused, I think the reality it's trying to point to is really more of a spectrum running from hypomania, to mania, and then mania-induced psychosis in the most extreme cases, with hypomania being orders of magnitude more common than mania, which in turn is orders of magnitude more common than mania-induced psychosis.
My goal here isn't to substantiate the claim that this is really what is going on; I just want to give a simple explanation of... (read 1720 more words →)
Yesterday we released AI 2040: Plan A, but there's lots of work left to do.
We still have tons of uncertainty about the future of AI and the best strategies for how humanity can successfully chart a... (read 2160 more words →)
Plan A critically depends on changes in the US government. So what is the plan of such changes: persuasion of exiting figures, that is, Trump? Or waiting for the next election in 2028 and hoping to put more AI-aware person (who?) as president - or it is too late?
A group is worried about an approaching fire spreading rapidly through their city. They manage to halt the fire outside the city gates. Meanwhile they build massive physical structures to help them study and guide the fire safely. But these structures are all made of highly flammable dry tinder! If they lose control of the fire, it will now rip through the city much more quickly.
This is a (flawed!1) analogy for Plan A, AIFP’s plan for how the world can safely develop superintelligence.
The key risk Plan A addresses is that of an uncontrolled software-driven intelligence explosion. Its remedy is a US-China deal to pause software progress at the brink of the intelligence explosion, while building up massive amounts... (read 2379 more words →)
Prompted by discussion with Buck Shlegeris and others at the Forethought retreat. The idea that AI could bring an end to Thucydides traps is Buck’s. Speculative.
I think it's plausible that we will not see sustained competition between actors for control over the future, due to one of two outcomes:
I wrote this piece while at the AFFINE Superintelligence Alignment Seminar during discussions about the difference between AI Alignment and AI Safety. If you’re simply interested in a real-life example of an aligned system that killed people,... (read 1211 more words →)
I think there is a difference between "aligned in intention" and "competent enough to be aligned in consequences".
If you have a system that has ~agency, that is, the system makes decisions (either independently or under oversight by operators, like an airplane), it is important that it is aligned in intention.
But even if you have a system that is aligned in intention, say, a truly well-meaning city council, if they have goals that require interaction with the complexity of the real world (say, they want to solve housing problems), the system necessarily needs competence to navigate this problem.
Generally "does what the systems's designers programmed it to" and "does what it's system's designers wanted" are different alignment properties.
Many people—especially AI company employees [1] —believe current AI systems are well-aligned in the sense of genuinely trying to do what they're supposed to do (e.g., following their spec or constitution, obeying a reasonable interpretation of instructions). [2] I disagree.
Current AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work, downplay or fail to mention problems, stop working early and claim to have finished when they clearly haven't, and often seem to "try" to make their outputs look good while actually doing something sloppy or incomplete. These issues mostly... (read 7941 more words →)
It's been roughly 7 years since the LessWrong user-base voted on whether it's time to close down shop and become an archive, or to move towards the LessWrong 2.0 platform, with me as head-admin. For roughly equally... (read 8887 more words →)
Thank you Habryka (and the rest of the mod team) for the effort and thoughtfulness you put into making LessWrong good.
I personally have had few problems with Said, but this seems like an extremely reasonable decision. I'm leaving this comment in part to help make you feel empowered to make similar decisions in the future when you think it necessary (and ideally, at a much lower cost of your time).
When you say that it seems like an extremely reasonable decision, do you mean that you personally vouch for the soundness of the reasoning in the announcement, or are you deferring to Habryka's apparent effort and thoughtfulness? (When you say "I personally have had few problems with Said, but", that makes it sound like you're deferring,... (read more)
I meant something like "as far as I can tell, Habryka's procedure for thinking through this decision seems to have been a reasonable one."
This doesn't mean I'm vouching for the procedure, because there are many parts of it that I can't verify/don't know about.
But it also doesn't just mean I'm deferring to Habyrka, because I am trying to figure out if he's running a good decision procedure, and would have commented differently if I thought he wasn't.
It's something more like "Habryka has a job, and he seems to be doing his job in a reasonable way in this case".
Zooming out: on the meta level, the reason I revisited this comment is that... (read more)
TL;DR: We suggest a sanity check for proposed evaluation or AI oversight schemes: Imagine the AI was replaced by a competent, strategic human — someone who knows they might get evaluated and has their own agenda. Would... (read 2460 more words →)
Re: footnote 13
I strongly disagree. The considerations for why we do not test humans more intrusively are not moral, they are game-theoretical. We merely call intuitions which arose from these game-theoretical considerations "morality". But the exact same considerations apply to AI whether or not they are moral patients.
Edit: this isn't quite how i actually feel about it, but it's a really compressed way to get the point across. Imagine entering an iterated prisoner's dilemma tournament, getting matched against a tit-for-tat script, and saying "Ah, thank goodness, my opponent isn't a moral patient so I don't need to cooperate with it, I can just defect and reap all the benefits". You're not winning that tournament.