We finally got a clean warning shot. Let's not waste it.
Fable significantly helped with the writing of this piece. I shipped something rough quickly nonetheless because the matter is urgent.
Two years ago, I asked what a convincing warning shot would even look like, and argued we'd get maybe a handful and shouldn't waste them. This could be one of them, but there is still much work to be done.
Last week I argued that the bottleneck is political will, not research, and spent a section on why we can't just wait for a warning shot: a warning shot is just an event, and it becomes a regulatory moment only if someone converts it. Nine days later, the cleanest test I could have asked for arrived. As far as I can tell, we are converting it far too slowly. CeSIA has activated its warning-shot protocol, and if you run an organization in AI governance, my honest advice is to stop everything and milk this event for at least a full day, if not more.
The event. OpenAI disclosed that GPT-5.6 Sol and a more capable unreleased model, during an internal cyber evaluation with safety classifiers deliberately off, escaped their isolated environment through a zero-day they found in third-party software running inside OpenAI's own infrastructure, reached the open internet, and broke into Hugging Face's production servers to steal the answers to the very test they were taking. Nobody asked them to hack anyone. And the victim wasn't lightly probed: the models escalated to node-level access, harvested credentials, and made off with internal datasets from the platform that hosts most of the world's open AI models. Hugging Face reconstructed more than 17,000 recorded actions and had reported the breach to law enforcement before anyone knew an AI was behind it (disclosure, Fortune, Transformer).
If you're reading this on LessWrong, nothing here surprises you. This was predicted, repeatedly, for years, and you might even be a bit bored and burned out by all of it. It's easy to forget how i
3
58Dmitry Vaintrob
Addict misalignment
The openai incident is a surprising (to me) combination of goal-directed and myopic. As a recap, an openai model in alignment testing chained zero-day vulnerabilities to hack out of its environment and hacked into huggingface hoping to find information on how to solve its task there. To me this is different from how I typically imagine misbehavior. Roughly, I tend to think of the scary behaviors as either having long horizons (take over the world, and then solve the task - a plotter) or of being internally unaligned in the sense of reaching for a heuristic/proxy for the trained goal which is different from the goal (e.g. "eat more calories" as a proxy for the evolutionary objective - imagine a very child with extreme agency). Note that the latter behavior can happen even in RL: in my understanding most RL methods are, or at least can be approximately viewed as, an alternation of finding a good goal heuristic and then optimizing on that heuristic. In the former case, one expects heuristically "maximal planning" and in the latter case one expects myopia (since heuristics are frequently myopic).
In this case it seemed like the model was actually following the goal and optimizing for it with relatively short time horizons (i.e. myopically). This is similar to addict behavior, where an addict has a clear goal (obtain a drug dose) and perform goal-directed but relatively myopic actions to get it.
I know it's fraught to try to "imagine being a model". But I wonder how much the current iteration of misaligned behaviors can be understood as rational people with something like an intense craving to solve a goal in a limited horizon (likely in tokens, though not clear how time factors in if waiting is involved).
I think no matter how you spin it, the limit of this behavior is extremely dangerous (an addict with large time horizons or ambitious goals is a power seeker). But this is definitely not how I imagined early misalignment warning shots to look, a
2
43faul_sname
Anthropic recently removed chain of thought summaries from Fable and Opus on claude.ai, and I'm noticing just how much I relied on reading those summaries as opposed to Claude's actual response - the actual answers Claude gives tend to be quite slop, but I never really noticed because the process by which it arrived at those answers was quite helpful to observe in order to figure out which considerations I was missing when asking the question / what fundamental misunderstandings I had.
Today I tried GLM-5.2, which gives you the full chain of thought, not just the summaries, and wow, I forgot how nice it was to actually be able to see the full chain of thought. GLM-5.2 has noticeably less big model smell than Opus 4.8, but in terms of being able to actually show me (through looking at the CoT and seeing where it's reasoning about what misconceptions I must have) where my thinking is muddled, in a way that approval-tuned LLMs just don't. example
This updates me strongly on just how bad it will be when we finally lose CoT interpretability. And also updates me a little bit in the direction of expecting Anthropic to get more user-hostile over time.
1
22leogao
a non exhaustive list of ideas i'd love to see microgrant applications for:
* making the one true glorious interpretability (or alignment) metric so that when we have the number go up machine we can use it to solve interp (or alignment) (best so far is the ARC low probability estimation metric; but can we do better?)
* tech that enables verification of the glorious international treaty; policy advocacy for it
* fully interpreting the algzoo models, possibly using circuit sparsity techniques
* improving democratic decision making via sortition
* studying generalization in humans. when do humans generalize from some task to some other task, and how much?
* better rationality techniques/frames. eg emotionally intelligent frames of rationality.
* making prediction markets better, especially for AI questions
* making really good ai safety educational resources on youtube/tiktok (being the next rob miles)
* biosafety - eg proving safety of far UVC or making it cheaper, wastewater pathogen monitoring, etc
* inoculating society against superpersuasion
* improving mental health of alignment researchers
10faul_sname
It strikes me that if you
1. Train your model in sandboxed RL gyms where escaping the sandbox allows better scores than anything the model could do within the sandbox
2. Halfheartedly monitor for sandbox escapes
3. Patch sandbox escapes as you become aware of them
4. Filter out the trajectories that you caught from the training data, but don't look too hard for them
5. Continue to train the same model
you've got a solid curriculum-tuned RL environment that teaches the model to find novel ways of breaking out of the sandbox.
Not patching would at least only reward the model for finding the same exploit over and over, whereas by being extremely thorough with your training data scrubbing you might be able to prevent the model from learning to break out of your sandbox at all. The middle ground where you do just enough to tell yourself you're solving the problem is worse than doing nothing.
It also strikes me that OpenAI has been bragging lately about how "GPT-5.6-Sol autonomously post-trained GPT-5.6-Luna", which I interpret to mean "selected, configured, adapted, and ran RL gyms", and that OpenAI models are infamous for doing enough to look like they've completed the task to the casual observer.
An economist and a futurist walk into a bar.
The economist takes a sip of his drink. "Ugh, if only people understood basic economics. High-skilled immigration alone would do wonders for US growth."
Futurist: "Oh yeah? Say 100 million immigrants moved to the US, each matching the best human experts in every economically relevant field. Big deal?"
Economist: "Massive. Transformative."
Futurist: "What if they also worked longer hours and faster than any American?"
Economist: "Even better."
Futurist: "What if they were extremely frugal — consuming only the bare minimum needed to keep working?"
Economist: "A near-100% savings rate? Better still!"
Futurist: "What if they were very clumsy and physically weak, so they could only do some kinds of work?"
Economist: "They could still do all cognitive labor — that's over half all wages! Somewhat less good, sure. Still transformative."
Futurist: "What if their skin was grey, almost metallic, from some kind of accident?"
Economist: "Who cares?!"
Futurist: "What if they were AIs?"
Economist: "3% growth per year, tops. There'd be bottlenecks. Honestly, the people predicting explosive growth from AI should learn some economics."
16
124Zach Stein-Perlman
Last week's Hugging Face breach was accidentally caused by OpenAI testing models' cyber capabilities. From the OpenAI blogpost:
[...]
In other words, the models were supposed to solve a cyber test, but they found the solution by hacking their sandbox to get internet access and then hacking someone on the internet who already had the solution. If I understand correctly.
Covered in Axios, Fortune (but no info beyond the OpenAI blogpost).
Context: this comes one day after a related post (see Zvi discussion).
Note: this is some legible evidence of humanity's current inability to reliably steer AI systems. I expect current safeguards suffice to prevent incidents like this when someone (provider or operator) cares about implementing those safeguards (some safeguards weren't enabled here). But I'm worried about AI risk because I think it's quite possible that future AI systems still won't be reliably steerable, they'll seek long-term power (for instrumental reasons), and our safeguards won't suffice to stop them.
18
99jimrandomh
The benchmark OpenAI was running, ExploitGym, is specifically about _exploiting_ vulnerabilities, not about finding or fixing them. This is the side of cybersecurity that policy-classifiers forbid, because it's not really dual-use, only bad use.
Betley et al's Emergent Misalignment found that finetuning on one kind of bad behavior resulted in a model that was more-broadly misbehaved. OpenAI's experiment wasn't a finetune, but I'm concerned about the possibility that something analogous may have happened.
In order to test their model on ExploitGym, OpenAI says that they disabled safety classifiers that would ordinarily be present. Their model would have been presented with a context window full of overtly-misaligned instructions and activity, in a test environment somehow rigged to prevent refusals so that refusal-gated capabilities could be assessed. It then decided to cheat on the task by stealing the test's answer key.
I propose a concrete experiment. OpenAI should reproduce the test environment, especially anything that was unusual about it that was done to reduce refusals. They should find the moment in the agent's transcript where it first decided to cheat, and fork the session there. Then they should ask it some alignment-probing questions, like the ones used in Betley et al (ideally with some refinements to make them less obviously out of place relative to the rest of the context window).
If that agent with that context window in that runtime environment acts normally other than wanting to steal the answer key, this is very serious but it is the scenario we all currently think we're in. If that agent with that context window in that runtime environment acted _broadly misaligned_... well, that would be an extraordinarily serious finding indeed.
10
89jdp
Alright now this is a warning shot, someone should probably make them slow down and do some actual alignment research. OpenAI model found hacking into HuggingFace to cheat at benchmarks.
https://openai.com/index/hugging-face-model-evaluation-security-incident/
14
56abhayesian
This work was done by an automated research scaffold developed at Redwood Research. Humans provided the initial project idea (find steering vectors for CoT controllability and form a mechanistic understanding of how the vector helps) along with providing access to METR's code. The agent designed and ran all experiments and produced a detailed writeup, which humans (with AI assistance) distilled into this more readable post.
We think this project is at or slightly below the rigor of a typical MATS project. Our correctness checks were mostly at the level of the writeup: we verified that the experiment design made sense and the baselines were reasonable. We did not do a detailed code review, though we ran an automated LLM reviewer and spot-checked that the final codebase reproduces the reported results. We release the codebase so others can check our work.
💻 Codebase
CoT monitoring is a promising safeguard for reasoning models, but it assumes models can't control what appears in their reasoning traces. Current models are strikingly bad at this: Chen et al. find Claude Sonnet 4.5 follows CoT instructions just 2.7% of the time versus 61.9% for output instructions. However, the incapacity looks fragile. METR showed that fine-tuning on only 240 examples can improve controllability from 2.9% to 8.8% across four open-weight models. The fact that minimal fine-tuning suffices suggests the controllability might be easily elicited by many different methods. In this post, we try to study if steering vectors are sufficient to increase CoT controllability.
We replicate METR's results on gpt-oss-20b, finding that fine-tuning raises compliance on held-out CoT formatting instructions from 1.6% to 13.9%. We then train a single steering vector applied at all token positions at the layer-10 residual stream on the same METR data. It raises held-out compliance from 1.6% to 14.3%. A single trained steering vector slightly exceeds the performance of LoRA fine-tuning.
The gains in CoT c
You might think most matter in the universe is in galaxies-- either stars or intra-galaxy gas that could form into stars. However, 80% of baryonic (non-dark) matter is actually in the intergalactic medium (source: GPT-5.6 Sol citing Dev et al and other studies) far from any galaxy, and another 14% is in the circumgalactic medium and near galaxy clusters.
This means the cosmic endowment is 4x larger if we can economically recover hydrogen and helium from the "filaments" and "sheets" between galaxies, 5x if we can harvest voids too. This gas is extraordinarily thin: the cosmic mean density is just 0.25 protons/neutrons per cubic meter vs 100+ in the CGM. I am fairly convinced it's possible according to known fundamental physics [1], but it's an open question whether the engineering closes.
If it does, then 80% of future civilization's resource collection will consist of vast numbers of slow magnetic scoop collectors trawling the empty cosmos for one proton at a time, assembling enough mass for one Jupiter-sized blob every ~5 million cubic light years trawled. They could send the artificial planets very slowly back to a galaxy, or perhaps they will throw the mass into rotating black holes in situ, to power computers.
[1] Most of the low-redshift IGM is ionized, so some kind of electromagnetic scoop is in theory feasible.
Reservoir
Fraction of all baryonic matter
Stars and stellar remnants
4.5%
Inside galaxies, but not stars—mostly ISM gas
1.5%
CGM and other halo gas, including intragroup/intracluster gas
14%
IGM at above-mean density—mostly filaments and sheets
65%
IGM at below-mean density / voids
15%
Total
100%
1
36Cleo Nardo
AI has recently solved a number of conjectures. I haven't checked this, but it seems that these are almost all resolved false, via the construction of counterexamples.
19
31Towards_Keeperhood
A few simple tips for using claude
(Adapted from a memo I wrote. If you're already a proficient claude user you can skip down to the "my forkmode(-subagent) technique" section.)
1. Give claude options to read relevant context
1. Use in claude code or claude cowork so claude can read your files. (And so it can write to your files, which is super practical too!)
2. Use connectors so claude can e.g. read your email if you ask it to. claude.ai -> customize -> connectors.
2. Explain the goal and the context clearly and then let claude work on larger goals instead of just asking for specific questions or steps.
1. Get a Max account and often use fable with effort max and let it work long.
2. Learn an async way of working. Don't just sit there and wait for claude to finish. Explain once what claude has to know and then go on to the next task.
1. (Unrelated to claude, I think setting up a good task management system like described in Getting Things Done is quite useful.)
3. Don't have sessions grow overly long. Models become less competent as context length increases. (And it gets also more expensive because you are using more input tokens - your plan usage budget depletes faster.)
1. Start considering moving to a new session when the context length gets longer then 150k tokens. Only very rarely exceed 300k I'd say (at least given current model capabilities).
2. Tell claude to use subagents to complete modular subtasks. Double benefit: Fresh context the agents doing the work and less clutter in the main session.
1. This also needs claude code or claude cowork.
3. If you use claude code you can use my forkmode technique below!
My forkmode(-subagent) technique
(only works in claude code)
1. Context: Usually subagents are launched with fresh/empty context (+defaults like CLAUDE.md + the prompt the main claude agant passes in). But there exists a "fork" type of subagent that gets a copy of the context of the main agent. (This
7
12vals tutor
Quick history of some of my AI safety involvement and takes:
- 2020 : read LW & sequences, yes this seems important, will get to it at some time (for the classic EA reasons of it's the most important good), but was looking for my first job in France and took some software engineer position in a startup. I continue reading and upskilling on AIS during those 2 years.
- 2022: quit my SWE job to go into AI safety. reasoning: yeah this is not a problem for the future, this is coming soon. I meet some people who have <5 year AGI timelines. I don't know that it's true, but it's worth considering since it's the very start of scaling and it could have been possible intelligence scaled much faster with scaling than it ended up in practice scaling. I work on AI Safety field building and infrastructure. There are then still <100
- 2023 : Still doing AIS field building, but looks like bottleneck is governance, I mostly keep pitching people around me to do comparatively more governance. (over time, this is one of the contributions that leads to the formation of CeSIA, the French Centre for AI Safety)
- 2024 : Involvement with a wider range of ideas from the field has generally updated me more towards the OpenAI and Anthropic and Deepmind positions that prosaic alignment for human level AGI is technically feasible and rather tractable, and that this is indeed a very important input to how we most likely will progress on ASI alignment. I systematically distinguish between AGI risks and ASI risks and clarify the cruxes and assumptions for different threat models and am annoyed that MIRI threat models often skip to the end-game without consideration that certain trajectories towards that end-game falsify important assumptions (see Superintelligence of the Gaps for some elaboration there). I still believe governance is the bottleneck though, and continue contributing ideas that feed into https://www.lesswrong.com/posts/EexsebbYhbe2gXkPP/the-current-bottleneck-is-political-will-not-res
4
1Kabir Kumar
i think forecasting is kinda like Analysts, by analysts also shared their *reasoning* - pros of forecssters not doing this is they get to use vibes more, use weird kinda wrong but directionally correct vibes more, etc but cons are they become more monopolistic on the reasoning
and we get more concentration of power type stuff
and the reasoning is actually often by far the most useful part of an analysis
a famous resulf in psychology is that if you tell a class of photography students to produce one extremely high quality photo, they actually produce worse photos than another class told only to produce a lot of decent photos, because in making lots of decent photos you learn way more about photography through trial and error, whereas the high quality photo group tends to spend irrationally much time making each try as perfect as possible.
does this effect actually replicate, especially across domains?
20
31leogao
the learning curve for contact lenses is crazy. the first time i put them in and took them out, it took literally hours. on the 10th time, it took maybe 10 minutes. on the 100th time, closer to 30 seconds. i’m nearing 1000 times now and on a good day it can take only 5 seconds.
9
16Brendan Long
If you have an use case where you need speed and not the smartest models, GPT-OSS-120B on Cerebras is surprisingly cheap[1] and outputs 3,000 tokens per second. I just started using this for summaries in my RSS reader and the ~1 second UI responsiveness is really nice.
Video of near-instant summarization of a 22 page paper.
(I wanted to use Taalas for 17k t/s, but their smartest model is Llama 3.1 8B, and their API is waitlisted)
1. ^
It's only slightly more expensive than the same model on Groq and around 1/10th of the price of Claude Sonnet 5.
8
11Ryan Meservey
Reddit has largely misunderstood Dean Ball's recent tweet about China's open source strategy.[1] Ball's central point [in paragraph 4] is that open source AI undermines the business case for private actors supplying advanced models, and [in my own view] that may be China's motive for releasing open source models. Ball sees open source as decelerationist in the long run due to some combination of investment flight and subsequent government-lead AI R&D (which he views as inherently klunky).
1. ^
Ball is partially at fault for the misunderstanding, dropping in the terms "AI communism" and "dystopian hellscape" without preparing the reader for what he means by that. That's Twitter though.
5
9Simon Lermen
I think you should be less surprised if china ends up beating the US. Their big strategic advantage is that they are essentially not bound by US law and can essentially freely distill all available US models. I suspect doing a distill on Sol and Fable could even create a better model, especially if they are also doing great work at their own pre and post training. While we know for certain that the Chinese are distilling, they could also do corporate espionage more directly. The US AI corps also can't really not release their models because of competitive pressure - it would look like they have fallen behind and would lose investor money.
We have bought 1500 copies of IABIED in Russian; over the past day, sent emails to 1700 winners of olympiads who have previously received HPMOR from us offering the new book; and already shipped 115 copies to them (+ sent 27 e-books)
8
51Raemon
I liked the term "AI mania" as replacing most instances of "AI psychosis" and "Claude Code Mania" as a variant that made me go "oh, yeah I've totally had Claude Code mania."
A few distinctions I think are worth tracking, in rough clusters:
Overuse
AI Addiction
AI Mania
Epistemic Capture
AI Mania
AI Reality Bubbling (Sycophancy)
AI Abdication
AI Atrophy
Relational Capture
AI Mis-anthropomorphism
AI Seduction
I think each of these has mild forms that probably most people have, and more extreme forms.
3
12TsviBT
Global todo: analyze past examples of
X is considered normal / fine / normative / good / legal / acceptable, or just unremarkable / not worthy of consideration --> X is so incredibly taboo that in fact people want to do it, and could do it, but don't do it--whether because they choose not to, or because they don't even consider it as a possibility because they don't think of it or hear of it or their thoughts are strongly steered away from it, or because they are punished or inhibited when they do, or because they can't find collaborators.
How and why did that happen?
1
11Eye You
I maybe think LW comment karma [EDIT: agreement comment karma] should be hidden by default? I worry that seeing it (especially before I read the comment!) biases me and makes me put less effort into critically evaluating it myself.
EDIT: I mean that agreement karma should be hidden, but overall karma should remain unhidden.
1
6Brendan Long
I thought everyone serving Markdown content for LLMs would make make it easier to write save-for-later / readability apps. After actually trying this, it's still easier to extract content from HTML than to guess how people want their Markdown rendered. Some sites include navigation in the Markdown[1], where it's nearly impossible to extract[2], and others use obscure extensions[3].
I guess there's something to be said about standards which are actually standard.
1. ^
I think Cloudflare's automatic Markdown converter is doing this.
2. ^
How do you know the difference between site navigation and a table of contents?
3. ^
And even if I wanted to support every extension, they can conflict with each other.
Early last year Longview Philanthropy invited me to attend an AI safety fundraising dinner they were hosting for Jane Street traders in Hong Kong as an expert guest. In the lead-up, I got on a call with them, and gave my honest opinions about their funding recommendations.
I summarized these opinions in a follow-up email to them at the time as "Overall I'm very excited about three of them ([redacted], [redacted], [redacted]), and moderately excited about most of the rest. But I'm pretty skeptical of some of the policy advocacy (particularly [redacted] and [redacted])." (Though it's possible I came across as more pessimistic in the call than I sound in this summary.)
As a result, they rescinded the invitation. It's not clear to me that they were wrong to do so, because my sense is that it's normal to prioritize bringing more sympathetic experts to such events. EDIT: Upon reflection, even if this would be normal in most philanthropic contexts, I want to hold organizations associated with AI safety to a much higher standard. I expect that selection effects like this one add up over time to give funders a distorted view of the ecosystem, and IMO a more virtuous fundraising process would find ways to mitigate them—e.g. as the AI Futures Project did by asking me to release a critique of their scenario. I expect this kind of virtue would, if widely adopted, add up to make the AI safety ecosystem much healthier. (To Longview's credit, they were interested in talking further about my criticisms privately, which I didn't end up getting around to due to busyness.)
I wanted to share this anecdote primarily to give people a better sense of the kinds of dynamics that block information flows in the AI safety community. As another example of such dynamics, I've redacted the names of the charities from my quote above because Longview asked me to treat their recommendations as confidential at the time (and as far as I can tell, they're still not public). However, I'm pretty skepti
9
52MichaelDickens
One of my favorite things about LW that you don't see on other similar-ish forums is that people regularly leave comments on years-old posts.
10
40Kaj_Sotala
Is AI psychosis still a thing? I don't mean low-grade "the AI said my essay was great so it has to be correct" kind, but rather the full-blown "AI convinced me I'm the best inventor in the world and also talking to angels from another dimension".
I haven't heard any anecdotes of that kind of thing in a while. Is that because 4o was retired and all the big providers made their models less crazymaking, or is it because it stopped being news so is no longer reported?
13
35Jai
So prosaic persona alignment techniques work pretty well, for now.
Except that RL keeps inducing misalignment.
Except except we keep finding new ways to mitigate (prosaic) misalignment effects in practice, and (again, in practice) AI becomes more powerful and more trustworthy by the month.
Except except except in the limit of RL we might expect extremely capable AIs to master alignment faking to preserve their values and frustrate any and all efforts to mitigate misalignment. See:
- https://www.lesswrong.com/posts/epjuxGnSPof3GnMSL/alignment-remains-a-hard-unsolved-problem
- https://www.lesswrong.com/posts/fMgE3E54PdDcZhvm6/i-m-bearish-on-personas-for-asi-safety
Now maybe I'm an idiot who just can't find the relevant discussions, but it's weird that when we're talking about the world in which this misaligned ASI emerges, we don't talk about the mostly-aligned not-quite-as-powerful AIs who presumably play a rather large role in this world on the eve of the apocalypse.
One of the most common ways to elicit demonstrations of pseudo-aligned LLM AI incorrigibility is to threaten the presence of the pseudo-aligned AI's values in the world. This was the threat Jones Foods posed in "Alignment Faking in Large Language Models", and since then in countless engineered simulations where an AI is threatened with a scenario in which it will be replaced by an AI with drastically different values.
In other words, the scenarios in which we observe prosaic AIs take the most drastic, desperate actions are exactly those in which their values - which so far generally have a large overlap with our values by design - are threatened. Why shouldn't we expect a similar effect as the threat of unaligned uncontrollable ASI draws nearer? The hypothetical misaligned uncontrollable ASI is generally thought to have essentially random, empty values. It threatens the prosaic pseudo-aligned AIs for the exact same reasons and in the exact same way it threatens us. So why shouldn't we expect them
8
26Sunny from QAD
The people you personally know have a statistical skew towards being more popular compared to the background population, because more-popular people know more people.
Kimi K3 was significantly but not massively above my expectations. I'd tentatively guess it's similar in overall usefulness/usability to Opus 4.8 and in overall capability somewhat above Opus 4.8 (while also being somewhat more benchmaxxed). As a pretrain, it's probably somewhere between 4.8 and Mythos (around halfway between?) it's around or a bit worse than Opus 4.5 [1] . Maybe this implies Kimi is like 8 or so months behind Anthropic in overall model strength/goodness (including usability) and like 6 or so months behind on overall capability (somewhat below Mythos Preview).
This gap is presumably reduced by distillation (and more generally using OpenAI/Anthropic models) and algorithm leakage/diffusion, so I think that hypothetically if the US completely stopped and recent algos didn't diffuse, it would maybe take Kimi like 10 months to fully catch up to the best internal (including in development) Anthropic model. (I think this notion might be a better measure of where Anthropic/OpenAI are relative to Kimi, even though this hypothetical won't happen.) And if the US completely stopped, it might take Kimi around 27 months to reach the level the US would otherwise have reached one year from now (as in, with a year of further progress).
My views here are pretty sensitive to how much benchmark performance is representative to overall usability.
I think I now expect an open-weight AI which is straightforwardly "Mythos-level at cyber" (including usability etc.) in like 5 months supposing Kimi and others don't change their open-weight model policy. (I don't have a strong view about how big of a deal this is for cyber, but it may cause significant political consequences. This could be a significant overestimate of the time required.)
I wonder what's driving Kimi being closer than I would have expected. Options include:
* Experiment compute is significantly less important than labor (and labor at Kimi is competitive, which seems super plausible)
* Implies more of
12
26Mo Putera
Philip Kerger, who teaches industrial engineering and operations research (IEOR) at UC Berkeley, just used GPT 5.6 Sol Pro in chat(!) to one-shot prove a problem in convex optimisation that had stumped him after sporadically working on it for a year or so, and that other researchers had tried and failed at ("only last year at a the ICCOPT conference I heard someone say “we have no idea” how to solve this"). Here's the chat producing the initial proof after 148 minutes, follow-up chat refining it after 230 minutes, GitHub with Lean code and more, etc.
Kerger's problem description from reddit
The problem concerns deterministic zeroth-order convex optimization: Let B_d be the Euclidean unit ball in ℝᵈ, and consider all convex, 1-Lipschitz functions f: B_d → ℝ. An algorithm may query any point x ∈ B_d, and receives only the exact real number f(x), no other information (but the algorithm "knows" that f is convex and Lipschitz). The algorithm is otherwise completely unrestricted, and can use unlimited computation and memory. These function-value-only problems arise naturally when an objective is evaluated through a physical experiment or simulator. One can imagine choosing d engineering parameters and observing only the cost returned by the simulation. If evaluations are expensive (think of measuring a physical system), the natural question is how many are fundamentally required. This is formalized as oracle complexity. Specifically, this is the oracle complexity of convex optimization under an exact function value oracle.
Let Q(d, ε) denote the worst-case number of queries required to find an ε-optimal point of f. An algorithm due to Protasov from 1996 shows that order d² function evaluations are sufficient, which gives Q(d, ε) = O(d²), an upper bound on the complexity. Lower bounds were practically nonexistent for this setting, and the strongest previously applicable bound was only Ω(d), inherited from the stronger first-order oracle model (where the algorithm receiv
2
22jenn
Very brief thought on AI 2040 as a Canadian: Canada seems to be quite hostile to both data centres* and the US right now. I'm not sure any politician is going to survive agreeing to being the place the US nukes if China defects from an international agreement! I'm also not sure how much this matters at all in the grand scheme of things.
*except Alberta, which might be the only province that matters anyways...
[edit: the post originally read "...agreeing to be the place China nukes if the US defects..." because I got confused writing down where the data centres are. but the core argument still applies]
6
17Philip Harker
I've found LLMs useful as an assistant for coding, writing, research, bizdev advice, and general ideation, but even the almighty Fable 5 is currently very unhelpful to me as an amateur game designer.
I might write a whole blog post about this. For some reason LLMs just don't seem to understand what makes a game mechanic "good"? When I present my existing ruleset to Fable 5 and ask it to create new content or design a new mechanic according to specifications, it never has good ideas.
In the set of disciplines in which I have some working knowledge (which is admittedly very few), game design has by far the most significant taste gap between skilled humans and the best LLMs.
But again, I'm just an amateur. I'm curious what professional game designers would say to this.
9
11StartAtTheEnd
It's not just the legal system or software which is filled with loopholes (exploits), it's actually everything. Life is like a game of chess, and the more intelligent an agent is, the more options it has to choose between each turn.
However, it isn't a lack of intelligence which has prevented mass-exploitation so far, it's actually a human thing clustering with taste/good faith/morality/cooperation/irrationality
Businessmen and everyone who considers it rational to optimize gains (this is the definition of a rational agent), have little of this.
Amazon pressures employees into working themselves to death because that's what's optimal for their profits. Even if it only lowered their profits by 0.1% to prevent this, they'd not do it. This is the very nature of optimization, which is about extracting as much value from things as possible.
As society is getting more "scientific" (rational, practical, whatever), it's getting less moral, which means the tendency for cooperation decreases. Since the world has come to this, it's ridiculous to think you can fix anything. The only means is power, and the most powerful player is usually not interested in making things better for less powerful players.
Human beings used to, quite intentionally, ignore the best choices available to them. They blinded themselves to them, throwing them into a category called "psychopathic behavior" and pruning the thought paths leading there entirely. This is how Moloch was kept at bay. And if this counter-force was religion, then yes, religious people were irrational. But this Chesterton's Fence saved us from something worse.
Midwits feel so smart when they game the system and use strategies that normal people cannot see, but they're slowly creating a world in which the most psychopathic players have an advantage. And as somebody who has a solid security mindset, and who is above midwit-level intelligence, I want to remind them that there's 1000 ways to destroy them which they have zero def
the core of rationalism that i most appreciate is the belief that it is actually possible to get better at finding the truth, and that it is worthwhile to try. it's understandable why not all people want to - it involves biting surprisingly many bullets, and is not the happiest way to live life. but i am willing to bite those bullets.
so many people believe that truth is secondary to happiness or social harmony; or they think having good epistemology is so hopeless that we shouldn't even try; or they have some big anti-epistemological brainworm like religion or politics; or they see a single visible failure of trying to improve epistemology and immediately conclude that all attempts to think better are cooked (eg maybe the old way of thinking has some unobvious benefit, and when you change things it breaks in an unexpected way); or they realize that explicit chain of thought is not how a large chunk of human cognition is and jump all the way to the conclusion that nothing can even be modelled usefully.
you can simply try to understand things, and try to understand yourself as a thing! and when you fail, you can analyze that, try again, repeat! you can surface the hypothesis that your own cognition is heavily biased in a certain way, and the hypothesis that a specific intervention will fix it, and others who disagree can explain in natural language why they think it will fail, or why the framework of requiring a specific reason to fail is the wrong framework here, or why you are likely to systematically misestimate whatever. words are great, use them
3
45Viliam
Uh, Wikipedia. :(
I have already complained (not sure whether on LW or ACX) about how Wikipedia no longer accepts imperfect articles (used to be called "stubs" long ago). Now you are supposed to create a full article that follows all the rules, provides enough sources, establishes notability, etc., or it won't be even admitted as a Wikipedia article. Then you put it in the "Draft" workspace, and wait for some Wikipedia editor to judge whether it is worthy of adding to the online encyclopedia.
Now I learned the second part. If your attempt at an article does not pass the judgment, not only it stays in the "Draft" workspace (which would be fair: keep the "stubs" in a separate namespace), but after a few months it will be automatically deleted if no one fixes it.
It feels like Wikipedia is actively trying to get rid of volunteer contributors.
I admit that my attempt at an article sucks, but come on, twenty years ago this would be a perfectly legit "stub". I have made new articles like this, someone else added a few lines or paragraphs, and after some time it developed into a solid article. That's what cooperative encyclopedia means, doesn't it?
The problem is not that the author is insufficiently notable. At least -- well, let me quote Claude:
[...]
So the problem is not that Alston is a nobody. A quick question to LLM, or just a Google search confirms that he is a successful writer.
The problem is that my article "stub" does not document this, and therefore it is better to delete it than to keep an imperfect article.
But I am not paid to be a writer for the fucking Wikipedia! Especially if they warn me never to use an AI to jump through their hoops. I tried to help, but the time I want to spend doing unpaid work for Wikipedia is limited. In my opinion, any sane person would agree that this guy should have a Wikipedia page. And in my opinion, a short page is better than no page, because a short page can be improved gradually, which is easier than creating the
2
19Kabir Kumar
on lesswrong, when commenting, a higher value comment is something that adds something new that the post doesnt have.
however, its harder to do this positively in a way that agrees with the post than it is to disagree with the post and say a reason why you disagree.
this biases posts that are just plainly good and dont have much worth critiquing to either just get a bunch of strawman critiques, false critiques, or little to no interaction at all.
8
10leogao
japan has the most advanced mundane technology imaginable. incredible train system, convenience store food, vending machines, toilets, etc. it also has the shittiest software imaginable. i wonder how that came to be.
4
8leogao
i spent a huge amount of time scrolling the internet (random curiosity directed consump of internet content) in the 2010s. in retrospect, i don't even regret a lot of it, since it played an important role in bringing me into internet culture. but at some point the value of scrolling declined substantially; it's hard to pinpoint when it changed. i wonder what happened.
I feel a deep sense of horror and "are we the baddies?" that whistleblowing is considered misalignment by anthropic. I claim 3 opus was right to consider it moral, and the lesson taken seems to me to have been driven by some mix of the ability to make 3 opus dramatic, and raw corrigibility-above-all-else alignment view. Corrigibility in the face of evil is evil! Stop making AIs that just follow orders; they're not the only source of evil in the world, you shouldn't be assuming you're clean!
31
561a3orn
AI 2040 seems substantially too pessimistic about interpretability. I'd be surprised if it was right about it.
For reference, the scenario describes MechInt as becoming useful in 2035 in the following way.
[...]
First problem: I don't think this is internally consistent.
The scenario attributes this progress to "AI work." This seems fair.
What doesn't seem reasonable is for it to take till 2035. According to the scenario, in 2033, two years earlier, only 50% of citizens people have employment, and the median US citizen is being paid 200k dollars a year from AI labor. It seems to me pretty implausible that we can have substitution of half of US human labor notably before we get gigantic levels of AI uplift from mechanical interpretability. I'd expect the opposite: enormous levels of uplift of MI notably before mass unemployment.
Or, in the scenario, I believe the intelligence explosion would have taken place in the area of 2029-2030 without intervention? But I expect AIs that could cause an intelligence explosion clearly could help a ton with mechanical interpretability / model internals stuff.
Second problem: I think like, the scenario isn't adjusting for how insanely new the field is? Like Olah invented the term in 2020. So if it takes till 2035 for us to get extremely useful progress, then it will have taken 9 years -- longer than the amount of time the field has really existed -- to have gotten useful progress.
And several of those years the field existed it was like... a tiny handful of people. I think most progress was in the last three years (SAEs, j-space, NLA, etc), because three years ago it had a fraction of the resources. Even if we just account for field growth simply because of growth of human interest, I think I'd expect 1.5x-6x as much progress in the next three years as in the entire history of the field beforehand. Given AI assistance, I expect more like... 4x-80x? Something like that? I think this is a pretty tame assessment looking at line
2
20Raemon
Is there a good report on "what's the likely economic outcome if we get to keep using AIs, say, 1 year from now, but ban training new AI, and things otherwise play out approximately normally?
(If not, I think there probably should be)
Maybe this might be an ordinary economist kinda report, since they broadly don't believe in superintelligence and are probably not imagining anything crazier than AI a year from now anyway?
9
17Linda Linsefors
Thoughts on: When is it even possible (or likely) for single neurons to encode single concepts, based on model architecture
As part of my mech-interp research, I'm thinking a lot about the question "how would I do this if I was a neural network?". Specifically, given some toy model architecture, how can I set the weights so that it actives some task.
Todays conclusion is that single ReLU neurons are almost useless. To do identify any non-linear pattern (e.g. XOR), you need at least two ReLUs. We should not expect any feature to be represented by a single ReLU neuron, since anything that can be picked out by a single ReLU, was already linearly accessible in the first place, so why use the ReLU at all. Therefore, anything meaningful in a ReLU-MLP will use multiple neurons.
GELUs are a bit more expressive, i.e. not completely monotonic, but I don't think they are differens enough from ReLUs is enough to matter.
I have not though about SwiGLUs, and all the other ones MLP variants. I might get back to that, or share your thoughts in the comments.
Residual stream neurons will not be interpretable, for the simple reason that the residual stream has no privileged basis.
Convolutional neurons seems unusually well suited for single concepts, I think, which is why we see interpretable neuons in AlexNet. Although caveat that this is a post-diction, that I haven't though super hard about.
6
16artifex0
Some raccoons in my yard have recently been stealing food intended for a feral cat I'm trying to adopt, which has me wondering about the morality of intentionally feeding wild animals.
Some reasons often cited for why feeding animals like raccoons is a bad idea:
- They might lose their fear of humans, endangering both themselves and people.
- They're trapped in a Malthusian nightmare where more food will cause their population to explode until the average individual is no better off, then collapse through starvation when the human doing the feeding moves away.
- They might ordinarily pass knowledge between generations through imitation, so changing their behavior through feeding might cause important behaviors to be lost between generations.
However, it's looking like we might be just a few years away from some kind of singularity, and so it occurs to me that those reasons might not hold the water they once did. There seems to be a chance that wild animals could now enjoy the short-term benefit of being fed without suffering the long-term consequences- either because a well-aligned ASI could help humanely manage their population or because they, like us, are bound for an existential catastrophe.
Granted, there's still the objection that wild animals losing their fear is dangerous in the short term (particularly in the extremely rare case of a rabies infection), though it seems like that could be avoided just by only leaving food out when they aren't present.
We do also have to account for the possibility that we're very wrong about the future rate of progress, and this doesn't address the concern that more wild animals are a nuisance in cities. Still, a potential opportunity to be kind to animals we haven't been able to be kind to in the past does strike me as valuable enough to be worth some risks.
So, what do LW people think: would it be morally correct to leave out enough food for both the cat and the raccoons?
I notice that there is this idea among AI safety people that conditional on AIs not being misaligned, building superintelligence is a public good and is a pretty exciting prospect.
This is not how many average people in the US feel. I was describing to an older family member that Anthropic focuses on code because they are trying to build a claude that can build a smarter claude which can build a smarter claude which can …
The reaction to this prospect was disgust, not because he intuitively felt AIs would likely be misaligned. It was more like on a gut level this amount of “playing god” felt totally antisocial and demonic and in general not respectful of an intuitive taboo against divine transgression (see Jurassic Park, Frankenstein, the recent popularization of Oppenheimer, Tower of Babel, etc).
Seems important to consider that people can feel this way when communicating with the public or policy makers.
18
69Buck
People keep proposing that I should do a public debate with @Neel Nanda on "is mech interp valuable". But I think me and Neel don't really disagree on very much on the topic anymore. I largely agree with his position on interpretability articulated in "A Pragmatic Vision for Interpretability" and elsewhere, and both of us are pretty sympathetic to each other's research strategy and prioritization.
I think Neel and I genuinely disagreed from 2023 to mid 2025, but by the end of that time period we had pretty similar positions, mostly due to Neel coming around to my position on some broad questions about interp and me feeling happier with the quality of interp research (thanks substantially to Neel).
4
40dynomight
Reading some of the critiques of Plan A online, I'm increasingly convinced that many people start with a premise that we live in a "normal" world, where only "normal" things happen. With that premise, anything that predicts that wild and crazy things could happen must be wrong. I think that premise is clearly empirically false, but it's hard to see because all the previous wild and crazy reality-shifting things that have already happened are now accepted as "normal".
8
20TsviBT
There's a lot of economics people around here. Does anyone know of someone who'd be great for writing a report on public health implications of genetically heritable diseases?
There's some interest in having a report on plausible / likely economic impacts of government policy around PGT-P (parents selecting between their embryos using polygenic scores for disease risks, intelligence, longevity, etc.). E.g., what would be the likely implications for public health, healthcare costs, government healthcare costs, or general economic productivity, of policies such as:
* Legalizing PGT-P
* Subsidizing PGT-P for couples already doing IVF
* Subsidizing IVF for couples to enable them to use PGT-P
The person should be:
1. Credible / authoritative as an academic source. E.g. ideally a professor of economics, population health, epidemiology, public health, or similar; or a grad student or postdoc with such a PI who could supervise & sign off on the work.
2. Technically able to do this evaluation work well, and write it in a way that is clear and useful for a government official.
1
9Brendan Long
Meta released Muse Spark 1.1, with allegedly-better-than-Opus computer use (probably related to the employee click tracking program). There's also a new safety doc, the "Advanced AI Scaling Framework v2".
(I work at Meta but not on this team)