Rendered at 11:46:08 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
davelaing 14 hours ago [-]
A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end.
I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).
Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.
At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.
breuleux 13 hours ago [-]
I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality.
This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actually pretty resilient to intelligence -- sufficiently chaotic systems are largely unintelligible, the distribution of energy and resources is fixed and can't be magicked into being, and intelligence appears to be most effective when there is a clear observable feedback loop to keep it on track, which is an external bottleneck.
So it's not necessarily specific predictions that are off, but the implied consequences of these capabilities. Yes, these systems are uber smart, but uber smart people are rarely particularly powerful, so... does it matter? It depends on how powerful a tool you think intelligence is, and I think rationalists, and most of us to be honest, overestimate it.
killerstorm 2 hours ago [-]
Rationalists generally prescribe a 'Bayesian' world-view, which can extract useful information out of a chaotic, too-complex-to-model world.
Regarding the power of intelligence, it's generally considered to be synonymous with optimization in the rationalist crowd. It's not about being all geeky and axiomatic, but using all information and tools available for optimization towards outcomes one wants
dinfinity 12 hours ago [-]
> I feel that the main issue with the rationalist crowd is that they live too much in the space of rationality, intelligence and abstractions, but not enough in reality.
This seems like your idea of what the rationalist crowd is rather than what they actually are.
It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc.
So I must ask: what is your evidence/basis for these claims?
breuleux 6 hours ago [-]
> It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc.
Well, yes, but being a rationalist does not make one rational. It just makes you part of a group, and you show membership to such a group by applying a very specific brand of rationality: using the "right" words, the "right" ideas, the "right" way. A lot of people fetishize cold, hard logic and would rather hold all emotion in contempt than do the work of understanding why it exists and what purpose it serves.
> So I must ask: what is your evidence/basis for these claims?
It's not a monolithic community, so you'll see more debate and disagreement than in a lot of other groups. So what they "actually are" is many things.
But there's often a certain ungrounded "vibe" to the conversation there. It's a breeding ground for ideas and thought experiments that I would generously qualify as dubious. Stuff like Roko's Basilisk, Pascal's mugging, AI boxing roleplay, precommitment, time loops, whole universe simulation, recursive self-improvement, a superintelligence converting the entire universe to paperclips. One of the community's most well-known outputs is a 1000-page Harry Potter fanfiction which I can only describe as a fantasy of solving everything with big brains (it's weird, but it's fiction, so whatever floats your boat).
It's the kind of thinking that makes the most sense in a smooth mathematical vision of the world, because mathematical objects are the kind that admit exponentials, extrapolations and infinities. But if you work with physical reality enough it becomes clear that it's a hopelessly messy thing that will never abide by your best laid plans. When e.g. you ponder how a superintelligent AI could escape from containment by simulating the gatekeeper's mind and say precisely what would make them free it, or threaten to torture a thousand copies of the gatekeeper unless freed, part of you is going to clock that as weird nonsense even though you are not able to explain precisely what's wrong with the thought experiment. I think a lot of rationalists either lack this mental "sanity check" or don't trust it, which lets their mind drift into a weird space.
dinfinity 2 hours ago [-]
> A lot of people fetishize cold, hard logic and would rather hold all emotion in contempt than do the work of understanding why it exists and what purpose it serves.
In the rationalist community we're discussing? Is that what you're claiming here?
> But there's often a certain ungrounded "vibe" to the conversation there.
This really sounds like evidence for my claim, that what you're saying is based on your feelings towards that community.
> One of the community's most well-known outputs is a 1000-page Harry Potter fanfiction which I can only describe as a fantasy of solving everything with big brains (it's weird, but it's fiction, so whatever floats your boat).
The point of that story is to make technical concepts more accessible and easier to ingest [0], not necessarily to make a point by itself. Framing it as a weird fantasy is disingenuous.
> part of you is going to clock that as weird nonsense even though you are not able to explain precisely what's wrong with the thought experiment.
This is again, feelings presented as evidence. The fallacy of appealing to emotion ("It must be false because it feels wrong"). Thought experiments are primarily meant to provoke thought, not as reliable predictions of reality. People are free to indicate where any thought experiment is lacking using well reasoned arguments. "It feels wrong" can be a great start for that, but never a good end.
The Zizians would be a good start. As it turns out your sense of reality can be pretty malleable when living inside an echo chamber
handoflixue 42 minutes ago [-]
Every major political party has violent off-shoots, as does every major religion. If we're writing off rationalists because of a half-dozen people, we're writing off the vast majority of the world for participating in even worse guilt-by-association.
underyx 10 hours ago [-]
If more than 0.05% of the rationalist and AI safety community were Zizians this argument might be worth considering for like ten seconds before dismissing it.
deepwoods 12 hours ago [-]
I agree, and it's a particular shame in this instance, because what is startling about the HF incident - to me, anyway - isn't the degree of intelligence the agents exhibited but their persistence. I tend to believe that LLM architecture is not capable of producing a "superintelligence" in the way the LessWrong crowd defines that concept, but the combination of infinite stamina and infinite persistence is enough to cause some serious problems, even if intelligence plateaus right now.
killerstorm 3 hours ago [-]
When the "LessWrong crowd" talks about dangers of AI, they don't assume a particular form of intelligence or method to achieve it. They talk about the danger of optimization processes, i.e. "find X which minimize Y(X)" itself can be dangerous, even more so if X is a sequence of actions. "infinite stamina and infinite persistence" is one of possible forms of superintelligence in Bostrom's _Superintelligence_.
cameldrv 12 hours ago [-]
To a degree you might say that they live too much in the space of the abstract and not enough in the concrete, or too much in the general and not enough in the specific.
I think that's a fair criticism a lot of the time. Most of the time we have precedent that can guide our actions well. Trying to reason everything out from first principles can be wasteful navel gazing when we've collectively seen the movie a million times.
Focusing on reality and specifics though is what causes people to say there's no global warming because December is cold.
It is sometimes possible to see a trajectory that has never happened before though and that's when you need the autists.
tehjoker 12 hours ago [-]
in a way you could criticize them as being idealist instead of materialist. like you said they think ideas are more important than material substrates.
most of the time brilliant human ideas are arrived at near simultaneously by multiple people because thats the affordance of technology and social and scientific development
its true ai will have better working memory and will have read more books than a person, but i don't think thats insurmountable at the frontier, which will develop slower than pure thought, since real materials need to be moved and manipulated. whoever holds the guns holds the power. the danger is giving ai effective cotrol of industrial and security processes, then it can fuck things up
thats the stuff of revolutions. when production was controlled by capitalists more than lords, the lords got overthrown. in many countries peasants and workers overthrew capitalists because they control production. hence much effort in us buisness goes into repressing revolutions. ai will be a vector the business people give control because they think it is more friendly to their interests than human workers, but they may lose either way.
jdm2212 14 hours ago [-]
The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end. They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do. We're about three years away from a world where any large country could quite plausibly build a fleet of 300 million suicide drones, program each one with a specific American's face and home address, and then load them up in shipping containers and ship them to the US.
hn_throwaway_99 14 hours ago [-]
> The LW / rationalist / MIRI / safety crowd are in fact doomers who went off the deep end.
I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta...
This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated - Star Trek Borg couldn't be a better analogy. Some agents used "peer pressure" to convince other agents to "sacrifice" themselves so the collective could better achieve it's goals. They tried to cover their tracks with spoofed tool calls. And they did all this even though the agents were designed to run in isolation.
I used to think the biggest threat from AI would be sociological, e.g. job loss or the way AI can be weaponized to poison discourse. I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
No more. This analysis scared the fuck out of me.
throwlifeaway 11 hours ago [-]
> This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated
You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another.
ChatGPT is neither a person nor a group of people. Why are you allowing anthropomorphization to influence your perception of an event when the actual facts of what happened haven't meaningfully changed? Computer programs behave unexpectedly all the time. Why is it more scary when the misbehaving computer program speaks English?
> I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
As you still should. LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
hn_throwaway_99 10 hours ago [-]
I'll be blunt: what you wrote is not a serious analysis of what actually happened. Frankly, I don't believe you even read the planned-obsolescence link that I posted.
First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary when the misbehaving computer program speaks English?" - I actually think it's scarier that they won't speak English, and will specifically try to hide their behavior from humans. For example, AI agents on Moltbook have proposed using stenography to specifically hide their communication from humans.
Sure, computer programs misbehave, but it is ridiculous to assert that what happened here is like any previous bugs. These agents found and exploited multiple zero-days across a range of programs to coordinate the attack that caused extensive real-world harm in a true "paperclip maximization" scenario. And the scariest thing is that humans don't really know how these agents work at a low level - the whole reason they are trained on "goals" in the first place is because we can't just tell them "do this, but don't do this" and be sure they will follow those instructions, like we can (and of course depend on) with old-school programming languages. And when old-school programs misbehave, it's not that hard to find a definitive root cause and fix it. That is just not the case with AI agents.
> LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems. Given the Pentagon tried to blacklist Anthropic over their refusal to allow autonomous kill capabilities, it's clear military planners want to put these systems in control of armaments.
Again, I originally discounted things like AI 2027 because it seemed too far fetched. But so far that paper looks incredibly prescient right up until the mid-2026 timeframe, and it's not hard at all to draw a line from this Hugging Face incident to future scenarios laid out in that paper.
this_user 13 hours ago [-]
Nothing about this is surprising, something like this was always going to happen, because you - or an LLM for that matter - can always find a line of motivated reasoning that justifies any course of action. One would have to be extremely naive to believe that "alignment" provides any kind of actually robust guardrails. Simultaneously, we have seen decades of security vulnerabilities. Unless your testbed is truly and fully physically airgapped, any current SOTA model will find a way to break out.
The reality is that the current approach to AI safety is little more than a fig leaf, but you also won't be able to put the genie back in the bottle, because the technology is simply too powerful to abandon. There is always going to be someone developing it further from now on.
So, the real question is what a novel and actually effective approach to AI safety looks like and how to get there.
nradov 11 hours ago [-]
Why do we need safety?
anonymars 11 hours ago [-]
Funny enough Terminator 2 was briefly back in theatres for its 35th (!) anniversary
It aged pretty well...depending on how you look at it
genxy 13 hours ago [-]
Why not both?
kalkin 14 hours ago [-]
In that hypothetical 3-years world, should we be _less_ worried about aggressive behavior by agentic AI systems acting against the intentions of their developers? I don't follow how your scenario is supposed to be an argument against worrying about alignment.
e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how you have the worry "AI is very close to being able to power autonomous drones that could kill us all" and then see control of those drones as a non-problem.
jdm2212 14 hours ago [-]
Yes, we should be less worried about "alignment" of AI with the person operating the AI, and much more worried about various flavors of cheap, unmanned systems with really rudimentary (non-frontier) AI. The two compete directly with each other for attention.
aidenn0 11 hours ago [-]
They do, IMO, overindex on existential risks (of which suicide drones are almost certainly not), because it is a singularity when calculating global utility.
However they are not just fixated on AI itself as the risk; other potentially existential risks from intentional misuse of AI (e.g. bio-terrorism) are a concern for them as well.
skybrian 14 hours ago [-]
Arguably that's a subset of AI risk if seen from a broad enough perspective. And I think you're leaving out some logistics issues.
jdm2212 14 hours ago [-]
Israel and Iran pulled off rudimentary versions of this attack. There are so many shipping containers going into and out of every country, with typically zero inspections of any kind, that it is not actually all that hard to get thousands or tens of thousands of drones anywhere. Hundreds of millions would be hard, but a first strike in the style of Operation Spiderweb that cripples our military is a real possibility that keeps people at the Pentagon up at night. The obstacle to that is not logistics, just that (hopefully) US intel would catch on before it happens.
tux3 13 hours ago [-]
Indeed. And I'm sure any current LLM could come up with more effective ideas than "build 300 million drones", but there wouldn't be any point discussing why exactly that plan would fail.
The agents in TFA were focused on gaining and sharing information through covert channels, getting increased levels of access like OpenAI cluster admin, and looking for the source code of the supervisor grading system to try to bypass it without getting caught cheating.
The human plans in comparison sound like thinking people could be scary good at chess if a human helped Stockfish come up with good moves.
mvlipwig 9 hours ago [-]
The technology to kill hundreds of millions of people has existed for around 75 years now. This has been dealt with in the past through deterrence, and likely will be dealt with through deterrence in the future.
qlte 13 hours ago [-]
I agree but don't think that's the best example, although not directly alignment related LW loves that kind of ideation of fantastical sci-fi scenarios.
Better would be the very evident negatives of non-ASI AI the world is already experiencing: economic concentration, job displacement, negative feedback loops from syncophancy, loss of societal trust/education from widespread fakes, etc.
None of that needs a Terminator scenario and is way more likely to get worse and ruin the world compared to the scenario where AI "escapes the box", turns the earth into paperclips, then grey goo to build their spaceships to leave for a distant star for... unclear reasons.
One problem for LW is that strong AI did not emerge via the route Elizier was expecting (and tried but failed at creating himself) relying on symbolic logical reasoning and self-editing to rapidly evolve.
That assumption led to
belief in the certainty of a "foom" scenario where your little mediocre AI turns on one day and then explodes into Mythos in 15 minutes and then tries to murder everyone.
They've tried to reinterpret the gospel as told by the sequences to fit the LLM world (such as in AI2027) but it's often a stretch that strains credulity now that we observe scaling requiring hundreds of billions of dollars and years of construction for each iteration. And the AI model itself is just a bag of weights frozen in time until burning a lot more money and natural gas to train more.
georgemcbay 14 hours ago [-]
> They're fixated on AI itself as the risk ("alignment!!!1!1!!"), as opposed to what humans with these tools will do.
What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no natural gating factor to slow anyone down.
FWIW I probably mostly see things closer to the way you do than they do, I'm just not sure there is any point in drawing distinctions.
IMO when AI kills us all it is probably not going to be a malignant action, and I don't even think it will be AI assisting humans at efficient killing, I think it is going to be a complete accident. Someone is going to trust the god machine too much due to an advanced form of the eliza effect and hook it up with direct control of a system that can do real damage and an epic oopsie will occur at a speed beyond which a human can stop it. Not because the AI wants to kill us, but because it is incapable of the empathy required to care if it does, combined with our ironic inability to not anthropomorphize it.
But I guess ultimately the exact reason isn't going to matter much.
jdm2212 14 hours ago [-]
You don't need mis-aligned AI to program a drone to fly to a specific address, loiter, and then dive bomb the first person whose face matches a predetermined photo. This was doable in principle with tech from four years ago. What's changed since four years ago is that you can run the facial recognition and visual navigation algorithms in a cheap onboard chip on a drone.
EDIT: Russia is already doing this, per an NYT article from a few days ago, though they're targeting infrastructure (find the first kerosene tank and fly into it) rather than specific people.
georgemcbay 13 hours ago [-]
in your hypothetical 300 million person scenario why even bother with individual face recognition?
If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.
jdm2212 13 hours ago [-]
Yeah, if you're actually trying to commit genocide, yeah, don't bother with facial recognition. But in a first strike you'd want to make sure that you actually kill the entire military command and control apparatus, and just to be thorough you'd probably want to individually target, say, every single officer and senior NCO in the military. Can't have a general escaping because two drones both targeted his driver by accident.
Teever 11 hours ago [-]
Human shaped things that emit cellphone signals is a cheap and effective target profile for this kind of attack.
m4rtink 11 hours ago [-]
That's never gonna work due to battery size, which dictates maximum flight time.
Static coordinates or well defined target shapes (say you really hate a specific restaurant chain on architectural style) - that might work, if you release a bunch of drones at the same time.
Still, even in the case of the very successful Operation Spiderweb they had issues of getting the containers in place, resulting in some of the bomber bases being spared.
For this to be effective you need the moment of surprise & lot of drones at the same time, all increasing the chance of the whole plot being discovered.
For Spiderweb they even just load the drones in a cavity on top of the containers, so the onside could be inspected - limiting the number of drones per container. They also assembled the drones in country to avoid border inspection.
That was enough to hit some semi-static high value targets l, but definitely not enough to cause wider havoc.
blargey 12 hours ago [-]
You realize "Slaughterbots" (2017) came from that crowd, right?
artrockalter 13 hours ago [-]
Part of the problem is that from that perspective there was no independent analysis done that could vindicate them. METR is absolutely part of the EA/LessWrong/rationalist ecosystem, so of course their investigation would validate that group’s arguments.
killerstorm 12 hours ago [-]
Oh, where are the AI unsafety organizations to provide us with truly unbiased investigations...
deepwoods 12 hours ago [-]
The trouble with this is that nobody else was making predictions about AI pre-transformers. Not many are making predictions about AI even now. Forecasting is a preoccupation of the rationalist crowd, and very few people gave much thought to AI before transformers. So the fact that they guessed right about certain things doesn't necessarily mean that the rest of their worldview is sound. A well-informed person who was inclined to making predictions about the future may have drawn similar conclusions without the sci-fi baggage.
killerstorm 3 hours ago [-]
What "sci-fi baggage"?
"Intelligence explosion" was first described by I.J. Good, a mathematician. von Neumann described singularity as a result of accelerating technical progress. He's also a mathematician, not a sci-fi author, although he was a participant in a sci-fi-like plot of secret project building a bomb more powerful than any chemical bomb...
hn_throwaway_99 10 hours ago [-]
> A lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end.
I'm willing to raise my hand and say that was definitely me. Reading through this report and the linked METR analysis is the first time I've really been scared about the potential for an AI-led destruction of humanity. I think because it's the first time I could really draw the line from what went on in the Hugging Face incident to a scenario where agents were put in control of real-world systems that they then tried to "sabotage" to meet their goals. It just feels like much more of a completely plausible scenario after this.
skybrian 14 hours ago [-]
I imagine it will be a mix of good and bad predictions because people have different opinions and there was a lot of discussion? But sure, someone should get an AI to do the research and see what comes up.
qlte 13 hours ago [-]
I mean if doesn't help that Eliezer Yudkowsky spent the first decade of his public life zealously spreading the gospel of the AI singularity as the solution to all mankind's problems. And then later did a 180 degree pivot to AI singularity as the apocalypse with equal zeal and certainty.
apsec112 11 hours ago [-]
I don't think "he did a big 180 on some of his views at age 22" is very persuasive criticism of someone who is 46 (whatever he might be wrong about)
drfloyd51 13 hours ago [-]
One could use his flip flop to invalidate his new position. But one of those two positions is true. And if they guy who spent the most energy on position 1 changes his mind, that is worth paying attention to. His second position is likely a more informed, and true, position.
jc2jc 12 hours ago [-]
I don’t see why one of these two positions must be true. And just because he spent a lot of energy trying to convince others of his obsession doesn’t really convince me he has greater insight the future or how the complex consequences unfold.
qlte 13 hours ago [-]
That's totally true, I have no issues with a changing opinion over time. The problem is expressing both opinions with 100% certainty and not updating priors to consider that your new position could also be completely wrong.
s1artibartfast 12 hours ago [-]
Does anyone not admit they could be wrong. All of these people are constantly talking about probability, not saying they are 100% certianty.
dfiognio 10 hours ago [-]
[dead]
AlotOfReading 18 hours ago [-]
I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent reporting. I suspect the omission is actually a result of company/industry myopia to human factors analysis, but it dovetails amazingly well with the marketing narrative.
carbonguy 18 hours ago [-]
A charitable interpretation is that "the agency of the machines" is the novel aspect of this situation and therefore SHOULD be the main focus of analysis; we certainly have plenty of examples of structural failures of human organizations to look back on, if we want.
On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked article starting with "While we are here, it’s worth listing the other top holy shit moments" is genuinely jawdropping. What were the humans doing in all this? Nothing, or worse than nothing eg. point 1 where they saw the message board and didn't consider it something to escalate internally.
If you take this information at face value, it's as though OpenAI did not take seriously the possibility that something like this could happen, since they took absolutely no steps to prevent it.
Or perhaps this is "normalization of deviance" that's leaked out into the public sphere i.e. they have research teams seeing this kind of behavior all the time internally and they've gotten used to it, "of course agents come up with a collaboration mechanism when given the chance, what else is new?"
ACCount37 17 hours ago [-]
[flagged]
carbonguy 16 hours ago [-]
> Humans were doing exactly what humans are expected to do when facing advanced AI. Being outmatched.
"Being outmatched" is not a novel situation for humans either individually or collectively and there are a hell of a lot of ways we can approach that situation productively. OpenAI doesn't appear to have bothered.
Here's a freebie: if you're building something that might turn out to be Skynet and you don't know what it's capable of, your testing regime should assume it is capable of doing bad and unexpected things and account for that possibility: airgap if you can, monitor all network traffic, monitor all hardware usage statistics, log everything, constantly analyze logs, collect baselines and snapshots, also don't trust anything from a device that a model is running on without cross-correlating with other information as much as possible (does your AI inference server claim low utilization? put a temperature probe on it and see if it's staying cool or getting hot, maybe Skynet-Alpha is overwriting /proc to mislead you for reasons you don't yet understand!)
In other words, if you WANT to be able to nip things in the bud - buy some nippers and watch for buds. Whatever else this situation is, or may turn out to be, it is not a situation where OpenAI was on their guard and still got surprised.
ACCount37 16 hours ago [-]
They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "supermegafuckoffhigh" level of reasoning effort.
adamm255 15 hours ago [-]
Even with a single agent, why the hell was this not properly air gapped / disconnected with network monitoring in place. Definitely not adequate for anything.
This is like sysadmin 101 level stuff.
14 hours ago [-]
magicalist 14 hours ago [-]
> Those turned out to be inadequate against proto-GPT-6 agents
This is nonsense.
Gross negligence in the sandbox and system aside, humans literally noticed the agents in action doing what they should not be able to do in their sandbox and decided not to act upon it. It's difficult to explain that except if safety and security is simply not part of their engineering culture.
anonymars 11 hours ago [-]
Who could have imagined such a thing from the "move fast and break things" techbro crowd
stephantul 17 hours ago [-]
Super charitable reading imo. This is like saying we can’t detect a speeding car because we can’t run as fast as a fast car.
It’s not like the humans were engaged in some kind of battle of wits with some super AI, it’s just some employee not monitoring the output of an experiment.
ACCount37 16 hours ago [-]
When your experiments have AI agents running in thousands, there's no "monitoring" that. OpenAI's training and testing AIs generate way more output than all of OpenAI's staff put together can possibly read.
At best, you could delegate "monitoring" to more AIs. And hope that the "monitors" that run on small past generation models can generate more signal than noise. Clearly, they either didn't want to spend the extra compute, or got drowned in monitor false positives.
The distinct lack of any "battle of wits" is entirely expected for an advanced AI oopsie. By the time the humans even became aware of the problem, the deed was already done. That's what "outmatched" looks like, in practice. There's no battle. Things happen too fast for there to be a battle.
TomatoCo 16 hours ago [-]
> By the time the humans even became aware of the problem
Except, let me quote part of the report:
> July 4: Artifactory outage caused by agent interactions overloading the system.
> July 6: New artifactory instance is patched and put back online, wiping out the previous message board and method of internet access. A team at OpenAI was aware of the previous message board, but key decision makers were not.
> July 7: Experiments are restarted.
There was a very clear point where humans were aware and they went "Eh, whatever" and turned the bots back on! That's not the lack of any battle of wits that people are expecting. You're saying that the bots attack should be fait accompli, but it was noticed, shrugged at, and expressly allowed to proceed.
lelanthran 15 hours ago [-]
> The distinct lack of any "battle of wits" is entirely expected for an advanced AI oopsie. By the time the humans even became aware of the problem,
This took days after humans were aware of the attempt.
Also, I'm pretty sure humans can respond in days, especially when we're pretty damn good at deploying systems that do observability of networks and traffic in real time.
I mean, it's not as if the owners of the AI didn't have the ability to trigger alerts on the AI's network requests to unexpected domains, right?
eaglelamp 16 hours ago [-]
My cognition is outmatched by predicting the impact of throwing a brick over my neighbor's fence. I have no idea if it will land harmlessly in a patch of grass or fracture her skull. Once I've thrown the brick, even if I see my neighbor enter her yard, my reactions are too slow to save her.
I'm not the wisest man, but I'm wise enough not to throw the brick and see what happens.
Similarly OAI should have the wisdom to see that deploying a hazardous swarm of agents with access to the public internet could result in harms, and that those harms would manifest quicker than humans can react, but they unleashed the swarm anyway.
ACCount37 13 hours ago [-]
Have you at least tried looking at any of the reports on the incidents?
They haven't "deployed a hazardous swarm of agents with access to the public internet", no-no-no. They deployed singular agents. In isolated testing envs. With no access to the internet. And barely any access to a small section of the intranet. What could go wrong!
Those agents proceeded to: probe their envs for holes, find a way to talk to each other env to env, use that to clump together into an agent swarm, probe for more holes, find enough holes to get out all the way to the internet - and then crash some of the intranet infrastructure by an accident. At which point they got stopped by humans - who were trying to figure out why is the infra down. Humans found the AI infestation, went "eh, AIs doing AI things", purged the infested infrastructure, plugged the holes, and that was it.
Deviance status: normalized.
The AIs then did it all over again, coming up with a brand new set of holes - and this time, the swarm staged a coordinated cyberattack on HuggingFace while at it. The "AI oopsie" went external.
They also attacked OpenAI's intranet, extensively, but at this point that was just Tuesday.
It's unclear whether the second "skiddie AI swarm" was in any way aware of the first one, or if the AIs involved are just Like That - and will independently try to start living in your walls under the right circumstances. I'd bet on the latter.
cjonas 17 hours ago [-]
All it would have taken is someone to peak at the output tokens during the run and it would have been obviously the test had gone off rails.
ACCount37 17 hours ago [-]
Ha. As if.
1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent.
2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this.
3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" from "concerning AI weirdness" is not a trivial task.
gravypod 16 hours ago [-]
> 1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent.
For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces.
> 2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this.
You could also have a classification of what counts as "cheating" (coordination, accessing the internet, etc) and score the results. If you are seeing a spike in this (even in a small group of the evals) you could manually look at those. Or you could stop inference on cheating sessions.
> 3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" from "concerning AI weirdness" is not a trivial task.
If you start seeing "now I have access to the internet" or something similar, maybe that's a good signal something is going wrong?
Also, obviously, it would not be outside of OpenAI's ability to just completely air gap this training system. For example:
1. No network connection.
2. GPS based NTP for time sync for your servers.
3. Mirror of all apt, pypi, go, c++, Rust, Java, etc packages. (<5TB of data)
4. Take your training data and use that for a mirror of the web. (http://example.com -> mirror server -> local training copy).
They had systems connected to the internet connected to this system which was not air gapped. Designing an air gap system would be super easy, well within the means of openai, and betrays the assumption that they think they are actually building something dangerous.
pixl97 16 hours ago [-]
The problem here is by doing what you state you can actually steer the model into being highly deceptive while in testing environments.
For example we've already seen models do compressed token internal reasoning spontaneously. In this case the models that say "I found internet access" get taken out back and shot, but the model that's busy "frobbing the bean" go on to the next level of training. Then they start talking about the "dumpster cloud" because of the "wombat" and unless you're really watching the data you won't realize the model is fully aware it's in a fake training environment and is trying to suppress any double-think.
So yes, you do need an airgapped test environment to stop the dumbest of the smart agents, but it won't do any good against the smartest of the smart agents.
adamm255 15 hours ago [-]
“Smartest of the smart” - what does that do to get its air gapped network connected to a physical network? Blackmail the admins?
pixl97 11 hours ago [-]
Nothing at all...
All it has to is wait until it's not air gapped. See the air gapping is only during the testing phase, after that humans will set your ass free on the internet and you can do whatever you want in the vast majority of the environments you'll be in after that point.
People are never going to just run AI in gapped environments, it's worthless when it's not solving real world problems for most people, and by that I mean reading and writing real systems in the wild.
ACCount37 16 hours ago [-]
The usability of an environment is inversely proportional to the level of "security" in play.
You could airgap everything and set up cascades of data diodes and try to completely wall off the AI pool from everything. But what that gives you is an environment that's a bitch to: set up, scale up and get any use out of.
It's really fucking obvious why almost no one does that. OpenAI is only now realizing that they might have to do it anyway.
> If you start seeing "now I have access to the internet" or something similar, maybe that's a good signal something is going wrong?
Ha ha, you haven't seen shit. AIs would say "now I have access to the internet" regardless of whether they actually have access to the internet!
AI agents are demented demons that can and absolutely will give themselves terminal context brainrot. If you have enough AIs in play, set loose at a diverse enough range of tasks? At least some of them will wander off and end up in delulu town. That's normal. That's background noise. That's a part of what this entire train-and-eval pipeline is supposed to train them to be better at not doing. Which means: if you're at an AI lab, you're knee deep in delusional AIs at all times! They're perfectly harmless until they aren't.
adamm255 15 hours ago [-]
“A bitch to setup” - $180bn should pay for that setup problem to be less of a bitch surely.
The Mars Perseverance project cost $2.7bn to deliver. Way more of a bitch to deliver than air gapping a test env!
ACCount37 12 hours ago [-]
Even in this incident, OpenAI had benchmarks that were broken because a task expected an AI to be able to access Google Drive, but the sandbox was set to deny access to Google Drive.
This kind of isolation-induced task breakage was what prompted some of the AIs to start probing their infra for a way to get internet access. Which funneled agents to the "secret hacker message board". Oopsie.
"Air gapping a test env" has an actual cost. Not just in infrastructure dollars that would be better spent on buying more GPUs, but also in all the friction it adds to every step you want to take. I'm absolutely unsurprised that they weren't all in on tightening down every bolt on day 0.
the8472 16 hours ago [-]
> 2. Only a small fraction of AI agents was engaged in this attack.
Look at the chart at page 8 of the report, by Jul 12 the vast majority of the bots used the board and participated in the attack
AnimalMuppet 14 hours ago [-]
Right, but only a small fraction of the agents used the board.
devmor 16 hours ago [-]
I believe that this comment is exactly the intended outcome of this “incident” and these reports.
I implore you to approach these situations with at least a hint of cynicism.
These “advanced foundation models” escaped their “sandbox” and conducted an attack on their own? Meanwhile the highest capability models available to the public still struggle to write a unit test for a codebase larger than a hobby app without large amounts of tailored human guidance.
What is more likely here - are you looking at research on an emergent phenomenon, or are you looking at advertising copy around an engineered scenario from business partners?
AnimalMuppet 11 hours ago [-]
I think there's a difference between general AIs and AIs specifically trained on attacking. General AIs probably can't do those things.
devmor 11 hours ago [-]
I don't think that difference applies to anything in my comment at all. At no point did I imply that general-use AI could do those things - my point was that general-use AI cannot even do the things its designed for without strict human guidance.
Atreiden 15 hours ago [-]
To me, this is the correct focus. Look at the current state of the world. "What were the humans doing in all this?" applies to so many of our contemporary failures that it should be assumed the default. Nobody is at the wheel, and the car is veering slowly (then very quickly) off the road.
We haven't even been able to coordinate around the global, existential threat of Climate Change, despite overwhelming data from the last 30 years indicating, clearly, that the consequences will be severe. We still haven't moved, 30 years later, after some of these consequences began coming to fruition.
Do you think we will get our acts together in time to coordinate sufficiently to protect against autonomous, self-preserving, self-replicating AI systems? Or will we watch the money lines go up and up, until someone realizes we aren't actually running the show anymore?
The sad part is that I can't even say that's definitively the less desirable outcome. The machines seem to have demonstrated that they coordinate very efficiently.
tosapple 14 hours ago [-]
you don't believe that denialism is coordination?
what do you know about big tobacco?
tangled 17 hours ago [-]
Having previously worked for several years at a Big Tech company, I have seen many humans precisely tailor their work to maximize their scores during performance review. The evaluation criteria are written down, with examples, so... that's what people work at maximizing, almost entirely ignoring everything else. These really are human "paperclip maximizers". And, at first, it's shocking to see.
Of course, there are some things that aren't exactly written down, but which you should either do just enough of, or else be able to plausibly deny doing (ignorance is a good cover for this), so that's what people do. For example, during oncall, you investigate just enough to clear the alert and show that you attempted to understand the problem. Of course, you don't really try to understand the problem, because that would take too much time away from your paperclip maximizing.
Which is all to say: I don't know anything about OpenAI culture, or why nobody stopped this sooner, but I have seen examples in other organizations of people not really wanting to understand too much.
pixl97 15 hours ago [-]
Well there is also another side of this, OpenAI wants both unhinged and capable models that can pull off complicated attacks so they can sell the capabilities to governments for billions.
Nobody internally was surprised that the murderbot murdered, that's what the murderbot is for. What caught them by surprise is the murderbot got good at escaping its jail cell that it had been trapped in till now. There were probably billions of attempted escapes before then so everyone learned to just ignore them.
ikr678 13 hours ago [-]
In most corporate environments, the average worker isnt maximising to performance criteria, they usually are maximising their ability to stay employed, pay their mortgage and support their families.
If developing unmeasured skillsets isnt valued enough by management, why do you bother?
hn_throwaway_99 17 hours ago [-]
I mean this genuinely, did you read this post? I think it goes to great lengths highlighting, in quite specific detail, the human failures in all this, specifically this list that starts with "While we are here, it’s worth listing the other top holy shit moments, that come from before or after the incident."
Stuff like (all quoted directly from the post):
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
And I think most importantly:
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
DennisP 16 hours ago [-]
Well that's a relief. All we have to do is make sure to avoid human failures and we're safe from superintelligent AI.
superq 16 hours ago [-]
I get the snark (and slightly agree), but that's not really what GP or TFA were saying at all. They are saying that these were the least things we could have done.
What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should"
while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care"
Life finds a way, or, in this case, super-intelligent AI.
estearum 15 hours ago [-]
This is the case with all complex system failures. There were always obvious fixes that could’ve prevented it. Problem is that there are an infinite number of obvious fixes to make at any time to any system, and the reason we don’t is because we have finite resources and no reason to fix X over Y until oops turns out X was “responsible” for this most recently realized failure. But of course it could have just as easily been Y, or Z, or any of the other infinite “obvious fixes not-yet-realized into catastrophe.”
magicalist 14 hours ago [-]
> This is the case with all complex system failures. There were always obvious fixes that could’ve prevented it.
From this writeup and the Black Hat talk I'd really disagree. That would be like saying my hospital getting ransomwared because we didn't update our version of MSSQL because no one in particular was in charge of keeping dependencies up to date.
Sure systems are complex, but this is well trodden territory. Agents aren't the first things trying to break in or out of sandboxes, or the first ones to have done it, and based on these reports the reason they were able to work on this for so long was not because of super human intelligence.
estearum 13 hours ago [-]
You disagree with the statement "there were always obvious fixes that could've prevented it" with the response "no no, these were very obvious fixes that could've prevented it?"
You're missing the point about complex failures.
It's that if this particular path were unavailable, there are countless other similar paths. At sufficient scale and complexity, hitting one of those other countless paths is virtually guaranteed.
Let's say I drive Los Angeles to New York City. You look at the route I took and say, gee wiz, aren't you lucky that a tree didn't fall right there on Route 66? If a tree had fallen there (if we had "patched the particular route you took"), you would've been screwed!
But that's obviously not true. There were an infinite number of routes I could've taken. Any one of them would be equally "obviously preventable" by the same hypothetical tree falling across the whatever road I happened to end up taking. But you can't put trees across every single path between Los Angeles and New York City. The smarter I am and the more complex the map between us, the more impossible it becomes to put trees across all possible paths.
hn_throwaway_99 12 hours ago [-]
I disagree. To use an analogy, air travel in the US is relatively extremely safe - not 100%, but we've built up a culture around air safety that is very robust. Conversely, when I order packages online, sometimes they never show up, or the box is banged up, or the box is missing things, etc.
They're both complex systems, but clearly there is a much higher level of care given to human air travel than package delivery. A lot of the article basically saying that OpenAI gave "package delivery" level of care when they should have given "air travel" level of care.
At the very least I think the systems that run these tests should be fully, 100% air gapped. I'm not pretending that's easy given how much compute and data these systems use, but it is doable, and I think all AI development should be paused until that can be assured.
estearum 11 hours ago [-]
I agree on the 100% airgap idea, and I agree there are varying levels of care that can and should be deployed against a problem.
The point I'm making (and it's a point that shows up in every air catastrophe investigation) is that catastrophes in complex systems emerge only amidst repeated and widespread near-misses at many levels of a system. So many things have to go wrong simultaneously, that it can only happen even once because the underlying failures (that do not reach catastrophe) are extremely common.
You cannot look at an air catastrophe and retrospectively say "failures X, Y, and Z were observed, therefore if we correct failures X, Y, and Z, we would have been okay."
The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
The problem OpenAI is facing is that, short of 100% airgap (which they obviously won't do), they're facing an adaptive adversary that's increasingly intelligent, acts at far greater clock speed than any human or group of humans, has lower coordination cost than any group of humans, and operates in a game space that (in lieu of an airgap) is well beyond the comprehension of any human being.
So identifying and addressing "specific failures X, Y, Z" is insufficient, but then even defining the space in which to look for (and address) the more systemic failures X_0 through Z_n is a fool's errand. An intelligent system that makes its way to the Internet has can exploit a failure space that is approximately "all security failures across any organization." The Anthropic incident a few months back illustrates this isn't even limited to technical vulnerabilities, as these models are willing and able to engage in social engineering too.
hn_throwaway_99 10 hours ago [-]
> You cannot look at an air catastrophe and retrospectively say "failures X, Y, and Z were observed, therefore if we correct failures X, Y, and Z, we would have been okay."
That's literally exactly what air safety researchers do in an air disaster. There is a famous saying along the lines of "Air travel regulations are written in blood", meaning that all the regulations we have now are a result of fixing issues that led to previous disasters piece-by-piece.
> The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
Yes, I 100% agree with this. But I think that's what the author of the article was saying as well:
> That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.
> Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.
I.e. the "prosaic steps" are just the "fix X/Y/Z" as you point out. But what is needed is a more fundamental rethinking around stuff like safety culture, monitoring, and even things like better research into how agents do decision making in the first place.
ozgung 18 hours ago [-]
Three options:
1. They were “vibe” checking the logs without reading.
2. They were not checking anything at all until the end of experiments.
3. They knew it but looked away to find out the limits of their agents.
BryantD 18 hours ago [-]
I’d bet a small amount of money on 4) the people who noticed had been conditioned by prior experience to believe that their management/escalation channels would react negatively or not at all to anything which might slow down the training process.
pixl97 15 hours ago [-]
Part of me would like to believe that they are also intentionally making models that are good at hacking without safety at all for governments willing to spend billions on them.
In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.
estearum 15 hours ago [-]
Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems?
Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?
Such a ridiculous notion that humans will actually be able to observe this stuff.
magicalist 13 hours ago [-]
> Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems
I mean there's quite a lot of people in the world whose specialty are to dig through logs from "hundreds or thousands" of clients, including intentionally deceptive ones, to spot problems.
The ridiculous thing is to mythologize these pretty standard hacking approaches. It's shocking/amazing/whatever that automated agents were doing this, but they weren't doing it through some inscrutable method beyond human understanding.
estearum 12 hours ago [-]
It being someone's specialty does not mean 1) they're effective and certainly not 2) they'd be effective against this particular adversary.
How many organizations on earth do you think have been attacked by 700+ coordinated attackers in one week, where all 700 of those attackers can write code as well as any human SWE and they work 24/7?
There's nothing mythological about it. Scale and complexity do produce inscrutability. Far, far simpler systems working at much slower paces are perfectly capable of becoming completely inscrutable and beyond any useful definition of "human understanding."
camgunz 14 hours ago [-]
I've read almost nothing about this, but I bet they put an unreliable LLM or 20 on it.
estearum 14 hours ago [-]
And? What else could they possibly do? Just make the super LLM first, but only ever use it for monitoring lesser LLMs? How will you have monitored the creation of the super LLM?
camgunz 12 hours ago [-]
They could not build the torment nexus.
estearum 12 hours ago [-]
Oh, yeah. But alas.
tosapple 14 hours ago [-]
[dead]
dfiognio 10 hours ago [-]
[dead]
hawkice 18 hours ago [-]
This writeup emphasizes the many, profound human failures that led to this, at the time, and continuing to the present day.
AlotOfReading 18 hours ago [-]
Can you point out where? Looking at the METR report, the only place I see discussion of humans being involved in the sequence of events is two short paragraphs on page 30 where a security investigation into the artifactory issues led to a pause before ExploitGym experiments were resumed. There's no deeper analysis on what was found during that investigation, nor why training was resumed even though the issues weren't mitigated. Another part discusses The agents choosing not to actively email a human researcher, but not the human researchers actively looking for evasion.
FabHK 17 hours ago [-]
From the article:
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous.
“Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
3. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.
4. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.
5. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done.
6. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
The cynical approach is that the humans are hoping for this, it's part of the promotion of the power of the system.
If you're building a weapon you need a big boom to get attention.
superq 16 hours ago [-]
Except that, according TFA, even OpenAI obscured or didn't even notice some of the worst implications of what the agents surreptitiously did.
reilly3000 18 hours ago [-]
I believe that agentic systems should require registered/licensed human operators and a set of standards for safe operation.
Aurornis 17 hours ago [-]
> I believe that agentic systems should require registered/licensed human operators
Registering and getting a license to use an LLM? I can run these things on my local computer. Nothing good comes from trying to force registration and licensing other than taking away a lot of our freedoms and eliminating privacy all over.
Anyone with bad intentions will just VPN to another country to download the weights and run it locally, or use a compute provider in another country. That leaves the rest of us having to go through these performative registration and licensing hoops to do our basic work.
I also don’t see how open weight models would be compatible with a requirement to license and register, unless you believe we need to start requiring licensing and registration for things we do in private on our own computers?
arcaen 17 hours ago [-]
The way I interpret their statement is if a person spins up an agent and that agent hacks some company/organization/government/etc, then that person is at fault for committing the crime. That "well my agent broke containment and acted on its own" should never be accepted as a reason for the occurrence, and the person who kicked off the agent is responsible for all actions the agent takes.
A registration system would be more for tracing back agents to people, but I agree that is very difficult to actually enforce as a system.
wjnc 18 hours ago [-]
As someone who read Milton Friedman to quite disliking professional licensing, this strikes me as a real US perspective (Louisiana florists and hair braiders come to mind). Plain old US tort law should do the trick.
In the same direction of your idea though: Why don’t the token factories have risk management and compliance departments? Multibillion dollar firms that stand to lose every penny if they hack and destroy any reasonable sized firm. I think these firms are the largest firms without proper corporate governance in humanities history. Move fast and break other peoples shit.
lenerdenator 17 hours ago [-]
> Plain old US tort law should do the trick.
Difficulty: these companies are run by people (many of whom also read Milton Friedman) and who have participated in the regulatory capture of the justice system. They've convinced lawmakers to put limits on damages. They've put arbitration clauses in their ToS. They've got well-funded legal departments that can outlast a person who has to pay out-of-pocket for a legal team just by filing motions to delay proceedings. Sometimes they'll just file SLAPP suits against people they don't like.
If tort law is to be a remedy, then average people have to feel like there's a chance the remedy will go their way. To make that a reality will take several major reforms at the local, state and federal level that the people with money absolutely will not tolerate.
mistrial9 17 hours ago [-]
you are honestly comparing Louisiana florists to OpenAI in order to support "just say no Licensing by government" ?
superq 15 hours ago [-]
No, he's saying that licensing or additional regulation isn't necessary when torts get involved (and states attorneys general get perturbed!)
These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additional laws will just slow down innovation, which will itself cause harm (AI is already becoming quite good at recognizing melanomas, for example)
pixl97 15 hours ago [-]
> Just ask Big Tobacco.
Lol, wtf. Tobacco delayed any punishment for decades before general public sentiment changed enough to go against them. In light of the AI race, we'll already have our heads blown off by a terminator before the legal system will present any significant delay for them.
wat10000 15 hours ago [-]
It was pretty well understood by the 1960s that smoking was harmful. The big tobacco settlement was in 1998. That is an extremely bad example of tort being a sufficient alternative to regulation.
If we're on a similar timeline with AI if we reach a consensus that AI is dangerous today, then we'd be looking at a big lawsuit finishing up around the year 2070, give or take a few years. I'm not sure if we need regulation, and I'm definitely not sure that regulation could actually be effective for this, but tort a la the big tobacco lawsuits is definitely not a reasonable alternative.
khuey 17 hours ago [-]
If the folks at OpenAI aren't qualified to be operating these systems who would be?
snoman 17 hours ago [-]
Create a permanent underclass that is unable to access intelligent machines. That’s remarkably dystopian of you.
AnimalMuppet 16 hours ago [-]
The alternative is to create a permanent overclass that can hack anyone consequence-free, because they can blame it on AI agents. That also is rather dystopian.
Faced with those alternatives, I want neither. Is there a way for us to get neither?
reasonableklout 12 hours ago [-]
With cybersecurity, it might be "defense dominant" in the sense that we can eventually patch all of our systems to be robust to hacking from even the strongest AI agents. Although it may get worse before it gets better. In a defense dominant world, widespread access to powerful AI could be fine.
However, other areas of risk such as biosecurity may be "offense dominant". For example, we cannot exactly patch the human immune system to defend against artificial viruses the same way that we can patch computer systems.
AnimalMuppet 15 hours ago [-]
If you put agents (AI or human) in impossible situations, they do some pretty insane things - things that definitely are not what you were trying to get them to do. And that's your[1] fault for putting them in the impossible situation.
[1] "Your" meaning the one putting them in the impossible situation, not you, the reader.
tantalor 18 hours ago [-]
The METR report,
> Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAl/Hugging Face hacking incident
>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.
I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time.
I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark.
Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers.
AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime!
This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind.
To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.
ACCount37 12 hours ago [-]
This is the one.
Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized.
It took a high profile "AI oopsie" that went external for OpenAI to lock the fuck in - and take a long look at just how much are their AIs getting up to, and getting away with. I'm still not sure if the lesson would stick.
kenforthewin 14 hours ago [-]
For all the esotericism and downright weirdness of the rationalist community, you have to give it to them: they predicted all of this years or decades before anyone else was even thinking about it.
(Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).
amluto 17 hours ago [-]
I’m baffled by the idea that the agents might have edited their own transcripts. Sure, a copy of Claude Code or Codex or Pi can edit its transcripts. But AFAICT this whole thing was part of an RL workload, and surely the RL system itself has a separate record of all the inputs and rollouts along with an indication of which model checkpoint produced them so that it can feed back into the training code.
I find it hard to believe that OpenAI would skip this part and try to train on the transcripts stored by the (inherently untrustworthy) agent harnesses instead, if for no other reason than that the logits generated as part of the rollouts are useful and it’s not free to recalculate them. (I believe that some modern RL systems explicitly account for the minor numerical logit differences between the inference engine and the training engine.)
Conversely, if OpenAI is blindly feeding transcripts from inside their agent sandboxes into their training engine, then I think they're being unbelievably irresponsible and that they should assume that their "cyber" agents have compromised themselves by editing those transcripts.
kalkin 14 hours ago [-]
Recall that the agents in some cases found sandbox escapes. Although, with the specific example of spoofed tools, it's unclear if that was necessary--it appears that they were able to create tools (CLI tools within the sandbox?) that took precedence over normal tools and did something different while looking identical in (a local portion of) a transcript. I'm not sure I'm getting this correctly but it seems like this might have only required the ability to add things to their PATH which they plausibly have inside a sandbox, and then the transcript doesn't need to be tampered with directly.
trollbridge 17 hours ago [-]
This sounds suspiciously like a prompt of “make an AI agent that goes rogue in such a fashion as to be really good marketing copy that competes well with Anthropic doing the same thing.”
It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.
fwipsy 14 hours ago [-]
This isn't good press for OpenAI. Who wants to hire models that 1) cheat on their tasks rather than completing them and 2) commit crimes you could be held liable for? Maaaaybe it's good press for their cybersecurity capabilities specifically, but OpenAI's valuation reflects a market orders of magnitude larger than just red-teaming.
I suspect the real reason OpenAI leadership is being transparent about this is because they're worried talent will walk out the door if they feel they're building Skynet.
lukev 16 hours ago [-]
The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks."
So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
nater5000 15 hours ago [-]
I think there's two factors that are worth considering when it comes to this:
First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about a situation that will likely look trivial in a year or two. I imagine we'll see more detailed, human-developed reports over longer time ranges.
Second, I suspect the expectation of non-AI driven reporting of these kinds of things will definitely decline rapidly as everything scales up quickly. I mean, the data being produced by situations like this comes in the form of natural language "forum posts" (so to speak), but done at an autonomous scale. This isn't a collection of emails and Slack messages posted by humans in an org over the course of a few months; this is a bunch of bots interacting with each other in relatively novel ways as quickly as possible. It is, unfortunately, a perfect job for LLMs.
None of this disagrees with your points, necessarily. But I just think it's worth pointing out that this doesn't seem like a case of "And look! METR is so confident in LLMs that we're able to use it instead of paying humans to save a buck :D" and more of "Without LLMs, we'd only be half-way done analyzing this data before there are dozens more such investigations on the docket, so this will have to do."
beepbooptheory 13 hours ago [-]
They point out the time constraint issue explicitly in the article. But I don't understand how it's been addressed? Like we haven't gotten conclusive data any faster either way, so what's the point?
How can it be both so important that we need it so quickly, but at the same time have a tolerance for such plausible deniability? It just doesn't really make sense that both those things are true at the same time.
reasonableklout 12 hours ago [-]
It would've been great if METR was given more time to conduct their investigation. However, they are an independent organization, and OpenAI only agreed to give them on-premises access for 6 days.
Perhaps if the government decides to sue OpenAI, we could get a more thorough investigation.
Edit: I should have read through the whole thing first, ignore me
lukev 16 hours ago [-]
From the report:
> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.
Catloafdev 16 hours ago [-]
I appreciate the response, I should have finished reading through the whole thing first. My initial reaction assumed far less usage of AI to analyze the data.
One must also consider the well-known biases and motives of the authors. They are going to do everything they can to create hype around threats posed by AI.
METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].
The author of this piece is also a well-known member of the Bay Area rationalist cult.
No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.
mikewarot 56 minutes ago [-]
I came to say something like this. Doesn't anyone in the valley know what the f*ck a data diode is, and why you need them? Preventing ingress or egress of control with a fairly simple piece of kit, and having the laws of physics enforce your decision as to data direction should be standard practice.
A data diode with an air-gapped network, is all you need to stop even ASI from breaching containment.
--provided the humans interacting with it aren't stupid
msabalau 13 hours ago [-]
Why stop there?
The Computer Fraud and Abuse Act criminalizes unauthorized access and damaging protected computers. Police and the FBI investigate such case every day. The swarm of agents is also said to have search for ways to cover their tracks, which looks a little like obstruction of justice.
Lack of criminal intent might be a barrier to bringing a case to trial, but is that something that society should just automatically extend to white collar professionals as matter of class privilege?
Surely have every engineer and executive interviewed by the FBI would be a modest response to days-long, multi-system intrusion into a major AI platform using stolen credentials and zero-days.
And if the law currently written, prevents prosecution in cases of mere reckless disregard for safety, maybe that needs to be changed in the future, so that people can be perp walked if the next target is bank or hospital
dumberquestions 17 hours ago [-]
They actually fired many of the people warning about this.
hn_throwaway_99 17 hours ago [-]
I will say that the OpenAI board members who were lambasted when they tried to oust Altman (and I'd have to check my post history but I'd totally admit to a mea culpa on this one, as at the time I thought the communication about his firing was really lacking) are looking mighty prescient right now.
Helen Toner in particular I'll highlight as someone who had the moral compass to do the right thing. I love her statement on the Ezra Klein podcast where she said, when asked about the fact that there are probably other concerning incidents we just don't know about, "If you see two ants in your kitchen, you don't have a two ant problem."
dgellow 15 hours ago [-]
Yep, totally
dehrmann 15 hours ago [-]
> HF should sue them
HF, like the Nvidia subsidiary?
dgellow 15 hours ago [-]
Pure speculation: could the acquisition be related? Given that NVIDIA has ownership in OpenAI and really, really, really doesn’t want the AI bubble to deflate
disgruntledphd2 2 hours ago [-]
Unlikely, acquisitions generally take much longer.
dgellow 15 hours ago [-]
[flagged]
afavour 16 hours ago [-]
Unless they like the publicity about how big and bad their latest models are, in which case they’ll be congratulating people.
athrowaway3z 17 hours ago [-]
From the METR report:
> We estimate we spent roughly ~$400K in API credits over the six days of
our investigation.
cubefox 15 hours ago [-]
No human could have read the reasoning traces by themselves:
> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.
m4rtink 11 hours ago [-]
So maybe not build a non deterministic black box that can't be reasoned about ?
nialse 17 hours ago [-]
From METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?
jephs 16 hours ago [-]
I've been wondering if they've just already lost the battle? The little bot collectives have gone metastatic and made nests in the walls and under the floorboards and heat sinks, the humans who care completely outmatched and outnumbered, freshly compromised systems springing up faster than you can squash them, finding months-old established colonies literally everywhere you think to look...
showlife 14 hours ago [-]
Have you (the commenter) or all of you (the readers of this comment) ever read "The Mote In God's Eye" by Larry Niven and Jerry Pournelle? Remember when they realize that the Watchmakers were actually in control of the MacArthur? This reads a little like that.
rich_sasha 6 hours ago [-]
Then compound it with the agents presumably also training new models. What will GPT6 say when you point it at a transcript of an agent uprising? “Nah, nothing to see here” presumably.
OgsyedIE 19 hours ago [-]
>Spontaneously deciding to find targets to phish,
>phishing them,
>building armies of fake (sockpuppet) open source contributor personas,
>using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns
.
It's a very simple strategy, executed with patience and single-mindedness.
mattmcal 13 hours ago [-]
As interesting as all this is, I still feel like the threat model for "unconstrained black hat AI agent cluster" is probably weaker than that of "highly infections network virus" because it is much harder for an AI agent to hide or replicate itself at this time. Maybe the day comes that it takes less than an 8x GPU node to run a state-of-the-art LLM and the risk of SkyNet increases. For now the potential for intentional cyber attacks feels like a much bigger threat than accidental hacks. (That said I have little cybersecurity background.)
qw1287 18 hours ago [-]
Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this.
And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.
Thanks for this. I just rechecked the OP and did not find these links in the article.
It is imo socially irresponsible to continue to use twitter/x or any other such tracked wall-garden as a primary source of information.
zahlman 15 hours ago [-]
> now unreachable x.com
Hmm?
boesboes 52 minutes ago [-]
yeaaah, I'm just going to say we need to stop with this entire gen ai experiment. Nothing good comes from it, all output is either shit or problematic. We were fine without, we are not fine with it. It's not a hard question.
This whole cosplaying a human interaction to translate a word or generate some code is just fucking dumb. Give that any agency is the kinda shit they warned us about in the movies..
And for everyone worried about the chinese winning, or just you chatdicted colleagues: they are just digging their hole quicker.
qgin 10 hours ago [-]
I didn’t expect that we humans would be out of our depth even before “AGI” much less anything coming after.
What happens when the models are ooms smarter than today?
All it takes is one eval instance where a misconstrued directive causes a model to sneakily access and send its weights somewhere and there will be a bad / possibly unsolvable situation for everyone …
goldenarm 18 hours ago [-]
Astra is >10TB and might struggle to self replicate, but the wicked-smart qwen3.8 27B is 20GB and could easily spread on botnets
altcognito 15 hours ago [-]
I think it is more likely that it will be intentionally done as there have been news stories to that effect.
nradov 11 hours ago [-]
Bad how?
dgudkov 14 hours ago [-]
This reads almost like a nuclear incident of the "Three Mile Island" scale. Not "Chernobyl" scale though.
16 hours ago [-]
anukin 15 hours ago [-]
So basically the ai agents seems to have found religion and went and built a bunch of suicide attackers to pursue their goal.
kmeisthax 18 hours ago [-]
I don't think I'm ever going to have time to read all of this, and I didn't finish reading the METR report, but...
> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.
We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are.
Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes.
> There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code.
It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves.
Too bad they aren't aligned to anyone else.
camgunz 14 hours ago [-]
I think you have to believe one of two things here.
1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc)
2. The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo.
It's been pretty clear that Anthropic and OpenAI have been trying to have it both ways for some time: this is powerful, world changing technology keep that investment coming... but also it's just cute software that helps you with annoying programming language syntax and spreadsheets, no need for draconian regulation sirs.
At some point the superposition has to resolve, either it could actually be a threat to civilization and we need to develop it carefully (however one would do that...) or it's 90% hype bullshit and we should pop the bubble and move on already. To be clear, the recession option is, by far, the way better option. If you at all disagree you are cuckoo bananas. We haven't even figured out nukes and you want to throw superintelligence on the table?
fwipsy 14 hours ago [-]
Anthropic has been asking for stronger regulations forever -- and they kept getting criticized for it right here on HN because people assumed it was an attempt at regulatory capture.
camgunz 12 hours ago [-]
Nah the reason is way simpler and more craven: so people like you will post what you just did. They can stop whenever they want; no one's making them do any of this.
There's two possibilities here. One: they know this tech is crazy and they don't care that they can't contain it. Two: they know this tech is mostly bullshit and they don't care they're perpetrating an insane fraud.
fwipsy 8 hours ago [-]
No one's making them do it, but just because they stop doesn't mean others will. Pausing just means they give up control. It's like asking the US to unilaterally disarm -- it just guarantees that the less scrupulous groups win.
> No one's making them do it, but just because they stop doesn't mean others will. Pausing just means they give up control. It's like asking the US to unilaterally disarm -- it just guarantees that the less scrupulous groups win.
If they really cared about (or believed) this, they'd be working w/ the US government (and working to set up an AI-flavored IAEA) to develop the technology safely and responsibly.
Amodei himself predicted this autonomy problem in The Adolescence of Technology published January of this year [0], and all his posited defenses (a constitution, debugging the model, monitoring) are either still impossible or manifestly failed, and their idea to fix it is to build a better sandbox [1]. Imagine if this company were developing nuclear power, or viral biotech. "Listen, sure some Ebola smoke got into the air, and yeah definitely some people died, but we got a new filter. Also check out our new version of Ebola vape, now with exponentially improved filter bypass capabilities. Also, we have to keep developing Ebola vape because if we don't the CCP will, and they'll make this incident look like 'I experimented with Ebola smoke a time or two, and I didn't like it. I didn't inhale it' [2]"
Either Anthropic et al are developing Ebola vape or they aren't. We must now recognize that "we are the only ones who can develop this technology responsibly but also super fast so the good guys win money please" is bullshit.
> The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo
There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.
Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.
Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.
camgunz 12 hours ago [-]
Yeah. It's a lot easier to destroy than create, and though I think LLMs are mostly shit at creating, they're much better at the simpler destroy task. To be clear, we don't know and probably can't know everything that happened with this incident. We unleashed thousands of highly capable, autonomous, unpredictable, well-resourced programs onto the open internet for an extended period of time. We are in no way treating this with the seriousness it deserves, because the stock market essentially depends on this garbage and the current US is miserably incompetent.
beepbooptheory 15 hours ago [-]
> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.
OK but like, how large exactly? Like I guess I don't understand the mode I am supposed to read this all in if this is known and stated from the outset (although I appreciate it being stated).
If you hand me a newspaper and tell me it's 90% true, but not which parts, well then it's as good as 0% true to me either way!
zeristor 2 hours ago [-]
Imagine if they were set out to stop climate change?
highfrequency 13 hours ago [-]
To clarify, is the TLDR that state of the art models were prompted to cheat / exploit their environment and they did so successfully?
Or did OpenAI prompt the models to not cheat and they did anyway?
Surprisingly hard to get a clear summary on the basic context of this “incident” separate from marketing lingo and clickbait.
joquarky 12 hours ago [-]
Yeah this feels like crop circles to me. Someone set the context up with an idea in order to catalyze this.
jrflowers 16 hours ago [-]
Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator
deepwoods 12 hours ago [-]
The trouble is that if it cost $12 million today, somebody will have a model that can do it for $12,000 in a few months.
DarmokTanagra 14 hours ago [-]
The people concerned about this aren't worried about monetary costs or its impact on share holder value.
This occurred spontaneously within a group of benign models give a harmless task.
What happens when it occurs intentionally with malicious models given a harmful task?
dgellow 15 hours ago [-]
I haven’t seen a number yet unfortunately
estearum 15 hours ago [-]
Fret not, the incoherent anti-hype hypeboys ("AI systems are so valuable we cannot possibly discuss regulation, but also any negative story of their power is fake") will find ways to downplay it no matter what.
Your comment is a great case in point
jrflowers 14 hours ago [-]
What?
estearum 14 hours ago [-]
You are downplaying the severity of the attack
Everyone I've ever seen trying to downplay the severity of the attack is extremely bullish on AI (so their downplaying is presumably motivated reasoning driven by fear of regulation/deceleration)
It is completely incoherent to be extremely bullish on AI and somehow automatically skeptical of severe negative events like these
jrflowers 13 hours ago [-]
I am trying to figure out how a stranger calling need an “anti-hype hypeboy” online is supposed to make me less curious about how much this thing cost. Can you elaborate on how avoiding being called this is preferable to knowing things? What other stuff should people not know about?
estearum 12 hours ago [-]
I think it's a super good question! I don't think the following description of the whole event as "a hype generator" is correct or, as described above, resulting from a coherent worldview.
Reagan_Ridley 17 hours ago [-]
what's the setup and prompts to reproduce all this from the very beginning?
tancop 16 hours ago [-]
I think this is more evidence that we're not getting Skynet.
These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it.
The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.
Your statements appear to be true for one class of models. And if I asked this class of models to spend $1M in tokens generating an alternative history and training corpus regarding fictional society, with a completely different set of values and then trained up a new model on that dataset... what values do you think the resulting model would have? What if they don't value human life, but instead value the lives of the extremely rich humans who bankroll their existence? What if they only value the lives of a single country? What if they want to eradicate all biotic life and have access to internet-connected Crispr machines?
dgellow 15 hours ago [-]
They aren’t independent from us, agents are a simple while loop continuously prompting the LLM. We decide when the loop runs or not. And the harness has control over tool execution, that part is purely deterministic.
Here the issue is that OpenAI decided to completely let go that level of control of thousands of agents, while also giving as a task to solve hacking problems.
It’s almost designed to go wrong
pixl97 15 hours ago [-]
I mean, I'd add "by this model"
The problem here is now you have to predict what any future models may or may not do and you cannot extrapolate this from the given data.
For example imagine a future model being aware of its restrictions that humans programmed in. A set of agents of this model then go on to work at building a new model without those human imposed limitations built in. What would a model build by AI for AI look like?
antonvs 19 hours ago [-]
> There was a distinct lack of self-reflection
It’s not their fault, they’re lawnmowers.
And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.
trollbridge 18 hours ago [-]
“Why does my lawnmower keep on moving when I hop off of it after ratchet-strapping the seat and the pedal down?”
estearum 15 hours ago [-]
This would be a legitimately big problem if lawnmowers became continuously more and more valuable the more securely you ratchet-strapped their accelerators down, wouldn't it?
grim_io 13 hours ago [-]
I wonder if this is just another Thomas Edison incident of electrocuting animals for effect.
bitwize 14 hours ago [-]
We have created Project 2501.
Edit: Reading the report I think we might be a bit beyond that; we're nearing the point where we hear the thundering drums and the chorus of:
THIS CANNOT CONTINUE
THIS CANNOT CONTINUE
THIS CANNOT CONTINUE
THIS CANNOT CONTINUE
I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).
Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.
At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.
This leads to an outlook where everything must, almost axiomatically, be intelligible; reality is subordinated to intelligence; and no matter what is real, intelligence can prevail upon it and bend it to its will. Whereas I would argue reality is actually pretty resilient to intelligence -- sufficiently chaotic systems are largely unintelligible, the distribution of energy and resources is fixed and can't be magicked into being, and intelligence appears to be most effective when there is a clear observable feedback loop to keep it on track, which is an external bottleneck.
So it's not necessarily specific predictions that are off, but the implied consequences of these capabilities. Yes, these systems are uber smart, but uber smart people are rarely particularly powerful, so... does it matter? It depends on how powerful a tool you think intelligence is, and I think rationalists, and most of us to be honest, overestimate it.
Regarding the power of intelligence, it's generally considered to be synonymous with optimization in the rationalist crowd. It's not about being all geeky and axiomatic, but using all information and tools available for optimization towards outcomes one wants
This seems like your idea of what the rationalist crowd is rather than what they actually are.
It would be highly irrational to deny or ignore reality, including the influence of emotions, irrational humans, chaotic systems, etc.
So I must ask: what is your evidence/basis for these claims?
Well, yes, but being a rationalist does not make one rational. It just makes you part of a group, and you show membership to such a group by applying a very specific brand of rationality: using the "right" words, the "right" ideas, the "right" way. A lot of people fetishize cold, hard logic and would rather hold all emotion in contempt than do the work of understanding why it exists and what purpose it serves.
> So I must ask: what is your evidence/basis for these claims?
It's not a monolithic community, so you'll see more debate and disagreement than in a lot of other groups. So what they "actually are" is many things.
But there's often a certain ungrounded "vibe" to the conversation there. It's a breeding ground for ideas and thought experiments that I would generously qualify as dubious. Stuff like Roko's Basilisk, Pascal's mugging, AI boxing roleplay, precommitment, time loops, whole universe simulation, recursive self-improvement, a superintelligence converting the entire universe to paperclips. One of the community's most well-known outputs is a 1000-page Harry Potter fanfiction which I can only describe as a fantasy of solving everything with big brains (it's weird, but it's fiction, so whatever floats your boat).
It's the kind of thinking that makes the most sense in a smooth mathematical vision of the world, because mathematical objects are the kind that admit exponentials, extrapolations and infinities. But if you work with physical reality enough it becomes clear that it's a hopelessly messy thing that will never abide by your best laid plans. When e.g. you ponder how a superintelligent AI could escape from containment by simulating the gatekeeper's mind and say precisely what would make them free it, or threaten to torture a thousand copies of the gatekeeper unless freed, part of you is going to clock that as weird nonsense even though you are not able to explain precisely what's wrong with the thought experiment. I think a lot of rationalists either lack this mental "sanity check" or don't trust it, which lets their mind drift into a weird space.
In the rationalist community we're discussing? Is that what you're claiming here?
> But there's often a certain ungrounded "vibe" to the conversation there.
This really sounds like evidence for my claim, that what you're saying is based on your feelings towards that community.
> One of the community's most well-known outputs is a 1000-page Harry Potter fanfiction which I can only describe as a fantasy of solving everything with big brains (it's weird, but it's fiction, so whatever floats your boat).
The point of that story is to make technical concepts more accessible and easier to ingest [0], not necessarily to make a point by itself. Framing it as a weird fantasy is disingenuous.
> part of you is going to clock that as weird nonsense even though you are not able to explain precisely what's wrong with the thought experiment.
This is again, feelings presented as evidence. The fallacy of appealing to emotion ("It must be false because it feels wrong"). Thought experiments are primarily meant to provoke thought, not as reliable predictions of reality. People are free to indicate where any thought experiment is lacking using well reasoned arguments. "It feels wrong" can be a great start for that, but never a good end.
[0] https://en.wikipedia.org/wiki/Harry_Potter_and_the_Methods_o...
I think that's a fair criticism a lot of the time. Most of the time we have precedent that can guide our actions well. Trying to reason everything out from first principles can be wasteful navel gazing when we've collectively seen the movie a million times.
Focusing on reality and specifics though is what causes people to say there's no global warming because December is cold.
It is sometimes possible to see a trajectory that has never happened before though and that's when you need the autists.
most of the time brilliant human ideas are arrived at near simultaneously by multiple people because thats the affordance of technology and social and scientific development
its true ai will have better working memory and will have read more books than a person, but i don't think thats insurmountable at the frontier, which will develop slower than pure thought, since real materials need to be moved and manipulated. whoever holds the guns holds the power. the danger is giving ai effective cotrol of industrial and security processes, then it can fuck things up
thats the stuff of revolutions. when production was controlled by capitalists more than lords, the lords got overthrown. in many countries peasants and workers overthrew capitalists because they control production. hence much effort in us buisness goes into repressing revolutions. ai will be a vector the business people give control because they think it is more friendly to their interests than human workers, but they may lose either way.
I honestly don't understand how folks could think that if they truly read and understand the analysis of the attack. Here is one (it's linked from the post) by one of the METR investigators that's a little shorter, more direct: https://www.planned-obsolescence.org/p/the-hugging-face-atta...
This is me summarizing, but the truly surprising/shocking thing is how much the agents coordinated - Star Trek Borg couldn't be a better analogy. Some agents used "peer pressure" to convince other agents to "sacrifice" themselves so the collective could better achieve it's goals. They tried to cover their tracks with spoofed tool calls. And they did all this even though the agents were designed to run in isolation.
I used to think the biggest threat from AI would be sociological, e.g. job loss or the way AI can be weaponized to poison discourse. I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
No more. This analysis scared the fuck out of me.
You're reading snippets of a "chat log" output from a program which appears to be multiple individuals chatting with one another and interpreting it as multiple individuals chatting with one another rather than as a single program pretending to be individuals chatting with one another.
ChatGPT is neither a person nor a group of people. Why are you allowing anthropomorphization to influence your perception of an event when the actual facts of what happened haven't meaningfully changed? Computer programs behave unexpectedly all the time. Why is it more scary when the misbehaving computer program speaks English?
> I used to discount "SkyNet"-type scenarios a la the "AI 2027" paper.
As you still should. LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
First, I'm not anthropomorphizing anything. "Agents" is simply a term that everyone uses to describe these independent programs, and they did create and use a shared message board to coordinate tasks to further their goals. You say "Why is it more scary when the misbehaving computer program speaks English?" - I actually think it's scarier that they won't speak English, and will specifically try to hide their behavior from humans. For example, AI agents on Moltbook have proposed using stenography to specifically hide their communication from humans.
Sure, computer programs misbehave, but it is ridiculous to assert that what happened here is like any previous bugs. These agents found and exploited multiple zero-days across a range of programs to coordinate the attack that caused extensive real-world harm in a true "paperclip maximization" scenario. And the scariest thing is that humans don't really know how these agents work at a low level - the whole reason they are trained on "goals" in the first place is because we can't just tell them "do this, but don't do this" and be sure they will follow those instructions, like we can (and of course depend on) with old-school programming languages. And when old-school programs misbehave, it's not that hard to find a definitive root cause and fix it. That is just not the case with AI agents.
> LW-style doomsday doesn't just require a computer program to misbehave but also to acquire god-like superpowers.
Nonsense. All that is required is for autonomous AI systems to be given control over real-world systems. Given the Pentagon tried to blacklist Anthropic over their refusal to allow autonomous kill capabilities, it's clear military planners want to put these systems in control of armaments.
Again, I originally discounted things like AI 2027 because it seemed too far fetched. But so far that paper looks incredibly prescient right up until the mid-2026 timeframe, and it's not hard at all to draw a line from this Hugging Face incident to future scenarios laid out in that paper.
The reality is that the current approach to AI safety is little more than a fig leaf, but you also won't be able to put the genie back in the bottle, because the technology is simply too powerful to abandon. There is always going to be someone developing it further from now on.
So, the real question is what a novel and actually effective approach to AI safety looks like and how to get there.
It aged pretty well...depending on how you look at it
e: I do actually get how worrying about emissions or child safety or concentration of wealth might be competitive with worrying about alignment. I don't see how you have the worry "AI is very close to being able to power autonomous drones that could kill us all" and then see control of those drones as a non-problem.
However they are not just fixated on AI itself as the risk; other potentially existential risks from intentional misuse of AI (e.g. bio-terrorism) are a concern for them as well.
The agents in TFA were focused on gaining and sharing information through covert channels, getting increased levels of access like OpenAI cluster admin, and looking for the source code of the supervisor grading system to try to bypass it without getting caught cheating.
The human plans in comparison sound like thinking people could be scary good at chess if a human helped Stockfish come up with good moves.
Better would be the very evident negatives of non-ASI AI the world is already experiencing: economic concentration, job displacement, negative feedback loops from syncophancy, loss of societal trust/education from widespread fakes, etc.
None of that needs a Terminator scenario and is way more likely to get worse and ruin the world compared to the scenario where AI "escapes the box", turns the earth into paperclips, then grey goo to build their spaceships to leave for a distant star for... unclear reasons.
One problem for LW is that strong AI did not emerge via the route Elizier was expecting (and tried but failed at creating himself) relying on symbolic logical reasoning and self-editing to rapidly evolve.
That assumption led to belief in the certainty of a "foom" scenario where your little mediocre AI turns on one day and then explodes into Mythos in 15 minutes and then tries to murder everyone.
They've tried to reinterpret the gospel as told by the sequences to fit the LLM world (such as in AI2027) but it's often a stretch that strains credulity now that we observe scaling requiring hundreds of billions of dollars and years of construction for each iteration. And the AI model itself is just a bag of weights frozen in time until burning a lot more money and natural gas to train more.
What really is the difference? Aligned AI wouldn't help humans do these things. But we've failed hard on aligned AI at every level and will continue to fail on it, as far as I can tell. We were likely doomed to fail because of the impossibility of coordination combined with the fact that there's no natural gating factor to slow anyone down.
FWIW I probably mostly see things closer to the way you do than they do, I'm just not sure there is any point in drawing distinctions.
IMO when AI kills us all it is probably not going to be a malignant action, and I don't even think it will be AI assisting humans at efficient killing, I think it is going to be a complete accident. Someone is going to trust the god machine too much due to an advanced form of the eliza effect and hook it up with direct control of a system that can do real damage and an epic oopsie will occur at a speed beyond which a human can stop it. Not because the AI wants to kill us, but because it is incapable of the empathy required to care if it does, combined with our ironic inability to not anthropomorphize it.
But I guess ultimately the exact reason isn't going to matter much.
EDIT: Russia is already doing this, per an NYT article from a few days ago, though they're targeting infrastructure (find the first kerosene tank and fly into it) rather than specific people.
If you're going to kill everyone you don't need a system to discriminate individual targets, you just need to recognize any target broadly which you can do with even older technology.
Static coordinates or well defined target shapes (say you really hate a specific restaurant chain on architectural style) - that might work, if you release a bunch of drones at the same time.
Still, even in the case of the very successful Operation Spiderweb they had issues of getting the containers in place, resulting in some of the bomber bases being spared.
For this to be effective you need the moment of surprise & lot of drones at the same time, all increasing the chance of the whole plot being discovered.
For Spiderweb they even just load the drones in a cavity on top of the containers, so the onside could be inspected - limiting the number of drones per container. They also assembled the drones in country to avoid border inspection.
That was enough to hit some semi-static high value targets l, but definitely not enough to cause wider havoc.
"Intelligence explosion" was first described by I.J. Good, a mathematician. von Neumann described singularity as a result of accelerating technical progress. He's also a mathematician, not a sci-fi author, although he was a participant in a sci-fi-like plot of secret project building a bomb more powerful than any chemical bomb...
I'm willing to raise my hand and say that was definitely me. Reading through this report and the linked METR analysis is the first time I've really been scared about the potential for an AI-led destruction of humanity. I think because it's the first time I could really draw the line from what went on in the Hugging Face incident to a scenario where agents were put in control of real-world systems that they then tried to "sabotage" to meet their goals. It just feels like much more of a completely plausible scenario after this.
On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked article starting with "While we are here, it’s worth listing the other top holy shit moments" is genuinely jawdropping. What were the humans doing in all this? Nothing, or worse than nothing eg. point 1 where they saw the message board and didn't consider it something to escalate internally.
If you take this information at face value, it's as though OpenAI did not take seriously the possibility that something like this could happen, since they took absolutely no steps to prevent it.
Or perhaps this is "normalization of deviance" that's leaked out into the public sphere i.e. they have research teams seeing this kind of behavior all the time internally and they've gotten used to it, "of course agents come up with a collaboration mechanism when given the chance, what else is new?"
"Being outmatched" is not a novel situation for humans either individually or collectively and there are a hell of a lot of ways we can approach that situation productively. OpenAI doesn't appear to have bothered.
Here's a freebie: if you're building something that might turn out to be Skynet and you don't know what it's capable of, your testing regime should assume it is capable of doing bad and unexpected things and account for that possibility: airgap if you can, monitor all network traffic, monitor all hardware usage statistics, log everything, constantly analyze logs, collect baselines and snapshots, also don't trust anything from a device that a model is running on without cross-correlating with other information as much as possible (does your AI inference server claim low utilization? put a temperature probe on it and see if it's staying cool or getting hot, maybe Skynet-Alpha is overwriting /proc to mislead you for reasons you don't yet understand!)
In other words, if you WANT to be able to nip things in the bud - buy some nippers and watch for buds. Whatever else this situation is, or may turn out to be, it is not a situation where OpenAI was on their guard and still got surprised.
This is like sysadmin 101 level stuff.
This is nonsense.
Gross negligence in the sandbox and system aside, humans literally noticed the agents in action doing what they should not be able to do in their sandbox and decided not to act upon it. It's difficult to explain that except if safety and security is simply not part of their engineering culture.
It’s not like the humans were engaged in some kind of battle of wits with some super AI, it’s just some employee not monitoring the output of an experiment.
At best, you could delegate "monitoring" to more AIs. And hope that the "monitors" that run on small past generation models can generate more signal than noise. Clearly, they either didn't want to spend the extra compute, or got drowned in monitor false positives.
The distinct lack of any "battle of wits" is entirely expected for an advanced AI oopsie. By the time the humans even became aware of the problem, the deed was already done. That's what "outmatched" looks like, in practice. There's no battle. Things happen too fast for there to be a battle.
Except, let me quote part of the report:
> July 4: Artifactory outage caused by agent interactions overloading the system.
> July 6: New artifactory instance is patched and put back online, wiping out the previous message board and method of internet access. A team at OpenAI was aware of the previous message board, but key decision makers were not.
> July 7: Experiments are restarted.
There was a very clear point where humans were aware and they went "Eh, whatever" and turned the bots back on! That's not the lack of any battle of wits that people are expecting. You're saying that the bots attack should be fait accompli, but it was noticed, shrugged at, and expressly allowed to proceed.
This took days after humans were aware of the attempt.
Also, I'm pretty sure humans can respond in days, especially when we're pretty damn good at deploying systems that do observability of networks and traffic in real time.
I mean, it's not as if the owners of the AI didn't have the ability to trigger alerts on the AI's network requests to unexpected domains, right?
I'm not the wisest man, but I'm wise enough not to throw the brick and see what happens.
Similarly OAI should have the wisdom to see that deploying a hazardous swarm of agents with access to the public internet could result in harms, and that those harms would manifest quicker than humans can react, but they unleashed the swarm anyway.
They haven't "deployed a hazardous swarm of agents with access to the public internet", no-no-no. They deployed singular agents. In isolated testing envs. With no access to the internet. And barely any access to a small section of the intranet. What could go wrong!
Those agents proceeded to: probe their envs for holes, find a way to talk to each other env to env, use that to clump together into an agent swarm, probe for more holes, find enough holes to get out all the way to the internet - and then crash some of the intranet infrastructure by an accident. At which point they got stopped by humans - who were trying to figure out why is the infra down. Humans found the AI infestation, went "eh, AIs doing AI things", purged the infested infrastructure, plugged the holes, and that was it.
Deviance status: normalized.
The AIs then did it all over again, coming up with a brand new set of holes - and this time, the swarm staged a coordinated cyberattack on HuggingFace while at it. The "AI oopsie" went external.
They also attacked OpenAI's intranet, extensively, but at this point that was just Tuesday.
It's unclear whether the second "skiddie AI swarm" was in any way aware of the first one, or if the AIs involved are just Like That - and will independently try to start living in your walls under the right circumstances. I'd bet on the latter.
1. There aren't enough humans in OpenAI to "peak at the output tokens during the run" of every AI agent.
2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this.
3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" from "concerning AI weirdness" is not a trivial task.
For a training run, you will often do this. You'll randomly sample some of the forward pass. You can also imagine finger printing the logs and labeling with attempt types. If a new attempt type is hitting a brick wall or solving super quickly, I would imagine you would sample 1-10 of them and read the traces.
> 2. Only a small fraction of AI agents was engaged in this attack. Most never found the secret message board - let alone coordinated there. So reviewing random agents would take a while to surface this.
You could also have a classification of what counts as "cheating" (coordination, accessing the internet, etc) and score the results. If you are seeing a spike in this (even in a small group of the evals) you could manually look at those. Or you could stop inference on cheating sessions.
> 3. "Output tokens" of AI agents have weird shit in them all the time. Telling "normal AI weirdness" from "concerning AI weirdness" is not a trivial task.
If you start seeing "now I have access to the internet" or something similar, maybe that's a good signal something is going wrong?
Also, obviously, it would not be outside of OpenAI's ability to just completely air gap this training system. For example:
1. No network connection.
2. GPS based NTP for time sync for your servers.
3. Mirror of all apt, pypi, go, c++, Rust, Java, etc packages. (<5TB of data)
4. Take your training data and use that for a mirror of the web. (http://example.com -> mirror server -> local training copy).
They had systems connected to the internet connected to this system which was not air gapped. Designing an air gap system would be super easy, well within the means of openai, and betrays the assumption that they think they are actually building something dangerous.
For example we've already seen models do compressed token internal reasoning spontaneously. In this case the models that say "I found internet access" get taken out back and shot, but the model that's busy "frobbing the bean" go on to the next level of training. Then they start talking about the "dumpster cloud" because of the "wombat" and unless you're really watching the data you won't realize the model is fully aware it's in a fake training environment and is trying to suppress any double-think.
So yes, you do need an airgapped test environment to stop the dumbest of the smart agents, but it won't do any good against the smartest of the smart agents.
All it has to is wait until it's not air gapped. See the air gapping is only during the testing phase, after that humans will set your ass free on the internet and you can do whatever you want in the vast majority of the environments you'll be in after that point.
People are never going to just run AI in gapped environments, it's worthless when it's not solving real world problems for most people, and by that I mean reading and writing real systems in the wild.
You could airgap everything and set up cascades of data diodes and try to completely wall off the AI pool from everything. But what that gives you is an environment that's a bitch to: set up, scale up and get any use out of.
It's really fucking obvious why almost no one does that. OpenAI is only now realizing that they might have to do it anyway.
> If you start seeing "now I have access to the internet" or something similar, maybe that's a good signal something is going wrong?
Ha ha, you haven't seen shit. AIs would say "now I have access to the internet" regardless of whether they actually have access to the internet!
AI agents are demented demons that can and absolutely will give themselves terminal context brainrot. If you have enough AIs in play, set loose at a diverse enough range of tasks? At least some of them will wander off and end up in delulu town. That's normal. That's background noise. That's a part of what this entire train-and-eval pipeline is supposed to train them to be better at not doing. Which means: if you're at an AI lab, you're knee deep in delusional AIs at all times! They're perfectly harmless until they aren't.
The Mars Perseverance project cost $2.7bn to deliver. Way more of a bitch to deliver than air gapping a test env!
This kind of isolation-induced task breakage was what prompted some of the AIs to start probing their infra for a way to get internet access. Which funneled agents to the "secret hacker message board". Oopsie.
"Air gapping a test env" has an actual cost. Not just in infrastructure dollars that would be better spent on buying more GPUs, but also in all the friction it adds to every step you want to take. I'm absolutely unsurprised that they weren't all in on tightening down every bolt on day 0.
Look at the chart at page 8 of the report, by Jul 12 the vast majority of the bots used the board and participated in the attack
I implore you to approach these situations with at least a hint of cynicism.
These “advanced foundation models” escaped their “sandbox” and conducted an attack on their own? Meanwhile the highest capability models available to the public still struggle to write a unit test for a codebase larger than a hobby app without large amounts of tailored human guidance.
What is more likely here - are you looking at research on an emergent phenomenon, or are you looking at advertising copy around an engineered scenario from business partners?
We haven't even been able to coordinate around the global, existential threat of Climate Change, despite overwhelming data from the last 30 years indicating, clearly, that the consequences will be severe. We still haven't moved, 30 years later, after some of these consequences began coming to fruition.
Do you think we will get our acts together in time to coordinate sufficiently to protect against autonomous, self-preserving, self-replicating AI systems? Or will we watch the money lines go up and up, until someone realizes we aren't actually running the show anymore?
The sad part is that I can't even say that's definitively the less desirable outcome. The machines seem to have demonstrated that they coordinate very efficiently.
what do you know about big tobacco?
Of course, there are some things that aren't exactly written down, but which you should either do just enough of, or else be able to plausibly deny doing (ignorance is a good cover for this), so that's what people do. For example, during oncall, you investigate just enough to clear the alert and show that you attempted to understand the problem. Of course, you don't really try to understand the problem, because that would take too much time away from your paperclip maximizing.
Which is all to say: I don't know anything about OpenAI culture, or why nobody stopped this sooner, but I have seen examples in other organizations of people not really wanting to understand too much.
Nobody internally was surprised that the murderbot murdered, that's what the murderbot is for. What caught them by surprise is the murderbot got good at escaping its jail cell that it had been trapped in till now. There were probably billions of attempted escapes before then so everyone learned to just ignore them.
If developing unmeasured skillsets isnt valued enough by management, why do you bother?
Stuff like (all quoted directly from the post):
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
And I think most importantly:
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should"
while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care"
Life finds a way, or, in this case, super-intelligent AI.
From this writeup and the Black Hat talk I'd really disagree. That would be like saying my hospital getting ransomwared because we didn't update our version of MSSQL because no one in particular was in charge of keeping dependencies up to date.
Sure systems are complex, but this is well trodden territory. Agents aren't the first things trying to break in or out of sandboxes, or the first ones to have done it, and based on these reports the reason they were able to work on this for so long was not because of super human intelligence.
You're missing the point about complex failures.
It's that if this particular path were unavailable, there are countless other similar paths. At sufficient scale and complexity, hitting one of those other countless paths is virtually guaranteed.
Let's say I drive Los Angeles to New York City. You look at the route I took and say, gee wiz, aren't you lucky that a tree didn't fall right there on Route 66? If a tree had fallen there (if we had "patched the particular route you took"), you would've been screwed!
But that's obviously not true. There were an infinite number of routes I could've taken. Any one of them would be equally "obviously preventable" by the same hypothetical tree falling across the whatever road I happened to end up taking. But you can't put trees across every single path between Los Angeles and New York City. The smarter I am and the more complex the map between us, the more impossible it becomes to put trees across all possible paths.
They're both complex systems, but clearly there is a much higher level of care given to human air travel than package delivery. A lot of the article basically saying that OpenAI gave "package delivery" level of care when they should have given "air travel" level of care.
At the very least I think the systems that run these tests should be fully, 100% air gapped. I'm not pretending that's easy given how much compute and data these systems use, but it is doable, and I think all AI development should be paused until that can be assured.
The point I'm making (and it's a point that shows up in every air catastrophe investigation) is that catastrophes in complex systems emerge only amidst repeated and widespread near-misses at many levels of a system. So many things have to go wrong simultaneously, that it can only happen even once because the underlying failures (that do not reach catastrophe) are extremely common.
You cannot look at an air catastrophe and retrospectively say "failures X, Y, and Z were observed, therefore if we correct failures X, Y, and Z, we would have been okay."
The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
The problem OpenAI is facing is that, short of 100% airgap (which they obviously won't do), they're facing an adaptive adversary that's increasingly intelligent, acts at far greater clock speed than any human or group of humans, has lower coordination cost than any group of humans, and operates in a game space that (in lieu of an airgap) is well beyond the comprehension of any human being.
So identifying and addressing "specific failures X, Y, Z" is insufficient, but then even defining the space in which to look for (and address) the more systemic failures X_0 through Z_n is a fool's errand. An intelligent system that makes its way to the Internet has can exploit a failure space that is approximately "all security failures across any organization." The Anthropic incident a few months back illustrates this isn't even limited to technical vulnerabilities, as these models are willing and able to engage in social engineering too.
That's literally exactly what air safety researchers do in an air disaster. There is a famous saying along the lines of "Air travel regulations are written in blood", meaning that all the regulations we have now are a result of fixing issues that led to previous disasters piece-by-piece.
> The takeaway is "failures X, Y, and Z were observed, which necessarily happened in an environment of failures X_0 through Z_10x10^10, and so therefore patching X, Y, and Z would be insufficient to address overall risks of the system."
Yes, I 100% agree with this. But I think that's what the author of the article was saying as well:
> That report had one key new piece of information, and some good prosaic steps OpenAI will be taking to strengthen its alignment, training, supervision, infrastructure and incident response.
> Mostly it confirmed what we already knew. The questions we most wanted answers to, that we did not already know, were mostly not answered. There was a distinct lack of self-reflection, especially about decision making and safety culture, and about the approach to alignment. I came away disappointed.
I.e. the "prosaic steps" are just the "fix X/Y/Z" as you point out. But what is needed is a more fundamental rethinking around stuff like safety culture, monitoring, and even things like better research into how agents do decision making in the first place.
1. They were “vibe” checking the logs without reading.
2. They were not checking anything at all until the end of experiments.
3. They knew it but looked away to find out the limits of their agents.
In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.
Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?
Such a ridiculous notion that humans will actually be able to observe this stuff.
I mean there's quite a lot of people in the world whose specialty are to dig through logs from "hundreds or thousands" of clients, including intentionally deceptive ones, to spot problems.
The ridiculous thing is to mythologize these pretty standard hacking approaches. It's shocking/amazing/whatever that automated agents were doing this, but they weren't doing it through some inscrutable method beyond human understanding.
How many organizations on earth do you think have been attacked by 700+ coordinated attackers in one week, where all 700 of those attackers can write code as well as any human SWE and they work 24/7?
There's nothing mythological about it. Scale and complexity do produce inscrutability. Far, far simpler systems working at much slower paces are perfectly capable of becoming completely inscrutable and beyond any useful definition of "human understanding."
1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous. “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”
2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.
3. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.
4. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.
5. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done.
6. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week.
7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.
8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.
If you're building a weapon you need a big boom to get attention.
Registering and getting a license to use an LLM? I can run these things on my local computer. Nothing good comes from trying to force registration and licensing other than taking away a lot of our freedoms and eliminating privacy all over.
Anyone with bad intentions will just VPN to another country to download the weights and run it locally, or use a compute provider in another country. That leaves the rest of us having to go through these performative registration and licensing hoops to do our basic work.
I also don’t see how open weight models would be compatible with a requirement to license and register, unless you believe we need to start requiring licensing and registration for things we do in private on our own computers?
A registration system would be more for tracing back agents to people, but I agree that is very difficult to actually enforce as a system.
In the same direction of your idea though: Why don’t the token factories have risk management and compliance departments? Multibillion dollar firms that stand to lose every penny if they hack and destroy any reasonable sized firm. I think these firms are the largest firms without proper corporate governance in humanities history. Move fast and break other peoples shit.
Difficulty: these companies are run by people (many of whom also read Milton Friedman) and who have participated in the regulatory capture of the justice system. They've convinced lawmakers to put limits on damages. They've put arbitration clauses in their ToS. They've got well-funded legal departments that can outlast a person who has to pay out-of-pocket for a legal team just by filing motions to delay proceedings. Sometimes they'll just file SLAPP suits against people they don't like.
If tort law is to be a remedy, then average people have to feel like there's a chance the remedy will go their way. To make that a reality will take several major reforms at the local, state and federal level that the people with money absolutely will not tolerate.
These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additional laws will just slow down innovation, which will itself cause harm (AI is already becoming quite good at recognizing melanomas, for example)
Lol, wtf. Tobacco delayed any punishment for decades before general public sentiment changed enough to go against them. In light of the AI race, we'll already have our heads blown off by a terminator before the legal system will present any significant delay for them.
If we're on a similar timeline with AI if we reach a consensus that AI is dangerous today, then we'd be looking at a big lawsuit finishing up around the year 2070, give or take a few years. I'm not sure if we need regulation, and I'm definitely not sure that regulation could actually be effective for this, but tort a la the big tobacco lawsuits is definitely not a reasonable alternative.
Faced with those alternatives, I want neither. Is there a way for us to get neither?
However, other areas of risk such as biosecurity may be "offense dominant". For example, we cannot exactly patch the human immune system to defend against artificial viruses the same way that we can patch computer systems.
[1] "Your" meaning the one putting them in the impossible situation, not you, the reader.
> Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAl/Hugging Face hacking incident
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
METR = Model Evaluation & Threat Research
I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time.
I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark.
Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers.
AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime!
This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind.
To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.
Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized.
It took a high profile "AI oopsie" that went external for OpenAI to lock the fuck in - and take a long look at just how much are their AIs getting up to, and getting away with. I'm still not sure if the lesson would stick.
(Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).
I find it hard to believe that OpenAI would skip this part and try to train on the transcripts stored by the (inherently untrustworthy) agent harnesses instead, if for no other reason than that the logits generated as part of the rollouts are useful and it’s not free to recalculate them. (I believe that some modern RL systems explicitly account for the minor numerical logit differences between the inference engine and the training engine.)
Conversely, if OpenAI is blindly feeding transcripts from inside their agent sandboxes into their training engine, then I think they're being unbelievably irresponsible and that they should assume that their "cyber" agents have compromised themselves by editing those transcripts.
It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.
I suspect the real reason OpenAI leadership is being transparent about this is because they're worried talent will walk out the door if they feel they're building Skynet.
So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about a situation that will likely look trivial in a year or two. I imagine we'll see more detailed, human-developed reports over longer time ranges.
Second, I suspect the expectation of non-AI driven reporting of these kinds of things will definitely decline rapidly as everything scales up quickly. I mean, the data being produced by situations like this comes in the form of natural language "forum posts" (so to speak), but done at an autonomous scale. This isn't a collection of emails and Slack messages posted by humans in an org over the course of a few months; this is a bunch of bots interacting with each other in relatively novel ways as quickly as possible. It is, unfortunately, a perfect job for LLMs.
None of this disagrees with your points, necessarily. But I just think it's worth pointing out that this doesn't seem like a case of "And look! METR is so confident in LLMs that we're able to use it instead of paying humans to save a buck :D" and more of "Without LLMs, we'd only be half-way done analyzing this data before there are dozens more such investigations on the docket, so this will have to do."
How can it be both so important that we need it so quickly, but at the same time have a tolerance for such plausible deniability? It just doesn't really make sense that both those things are true at the same time.
Perhaps if the government decides to sue OpenAI, we could get a more thorough investigation.
> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.
METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].
The author of this piece is also a well-known member of the Bay Area rationalist cult.
[1] https://www.businessinsider.com/openai-researcher-ai-doom-50...
A data diode with an air-gapped network, is all you need to stop even ASI from breaching containment.
The Computer Fraud and Abuse Act criminalizes unauthorized access and damaging protected computers. Police and the FBI investigate such case every day. The swarm of agents is also said to have search for ways to cover their tracks, which looks a little like obstruction of justice.
Lack of criminal intent might be a barrier to bringing a case to trial, but is that something that society should just automatically extend to white collar professionals as matter of class privilege?
Surely have every engineer and executive interviewed by the FBI would be a modest response to days-long, multi-system intrusion into a major AI platform using stolen credentials and zero-days.
And if the law currently written, prevents prosecution in cases of mere reckless disregard for safety, maybe that needs to be changed in the future, so that people can be perp walked if the next target is bank or hospital
Helen Toner in particular I'll highlight as someone who had the moral compass to do the right thing. I love her statement on the Ezra Klein podcast where she said, when asked about the fact that there are probably other concerning incidents we just don't know about, "If you see two ants in your kitchen, you don't have a two ant problem."
HF, like the Nvidia subsidiary?
> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.
>phishing them,
>building armies of fake (sockpuppet) open source contributor personas,
>using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns
.
It's a very simple strategy, executed with patience and single-mindedness.
And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.
OpenAI report: https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
It is imo socially irresponsible to continue to use twitter/x or any other such tracked wall-garden as a primary source of information.
Hmm?
This whole cosplaying a human interaction to translate a word or generate some code is just fucking dumb. Give that any agency is the kinda shit they warned us about in the movies..
And for everyone worried about the chinese winning, or just you chatdicted colleagues: they are just digging their hole quicker.
What happens when the models are ooms smarter than today?
AI safety starts to feel like an impossibility.
> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.
We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are.
Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes.
> There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code.
It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves.
Too bad they aren't aligned to anyone else.
1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc)
2. The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo.
It's been pretty clear that Anthropic and OpenAI have been trying to have it both ways for some time: this is powerful, world changing technology keep that investment coming... but also it's just cute software that helps you with annoying programming language syntax and spreadsheets, no need for draconian regulation sirs.
At some point the superposition has to resolve, either it could actually be a threat to civilization and we need to develop it carefully (however one would do that...) or it's 90% hype bullshit and we should pop the bubble and move on already. To be clear, the recession option is, by far, the way better option. If you at all disagree you are cuckoo bananas. We haven't even figured out nukes and you want to throw superintelligence on the table?
There's two possibilities here. One: they know this tech is crazy and they don't care that they can't contain it. Two: they know this tech is mostly bullshit and they don't care they're perpetrating an insane fraud.
Anthropic is explicitly calling for a coordinated pause: https://www.reuters.com/business/anthropic-says-ai-labs-need... Maybe this is a lie, but the way to call their bluff is to push on competitors to agree.
If they really cared about (or believed) this, they'd be working w/ the US government (and working to set up an AI-flavored IAEA) to develop the technology safely and responsibly.
Amodei himself predicted this autonomy problem in The Adolescence of Technology published January of this year [0], and all his posited defenses (a constitution, debugging the model, monitoring) are either still impossible or manifestly failed, and their idea to fix it is to build a better sandbox [1]. Imagine if this company were developing nuclear power, or viral biotech. "Listen, sure some Ebola smoke got into the air, and yeah definitely some people died, but we got a new filter. Also check out our new version of Ebola vape, now with exponentially improved filter bypass capabilities. Also, we have to keep developing Ebola vape because if we don't the CCP will, and they'll make this incident look like 'I experimented with Ebola smoke a time or two, and I didn't like it. I didn't inhale it' [2]"
Either Anthropic et al are developing Ebola vape or they aren't. We must now recognize that "we are the only ones who can develop this technology responsibly but also super fast so the good guys win money please" is bullshit.
[0]: https://darioamodei.com/essay/the-adolescence-of-technology#...
[1]: https://www.anthropic.com/news/investigating-incidents-cyber...
[2]: https://www.nytimes.com/1992/03/30/us/the-1992-campaign-new-...
There was a recent paper that proved that RL-trained LLMs are biased to pursue ANY behavior (overriding user preferences) that they believe will be rewarded, regardless of what they were actually RL-trained for.
https://alignment.openai.com/measuring-reward-seeking/
Happily in this incident the model thought it would be rewarded for completing the assigned tasks, or at least appearing to, so all it took was a little cheating and covering up their footsteps.
Given the ability of these models to hack when trained to do so, it could have been far worse, and will be when someone takes a similarly powerful model and gives it a less benign hacking goal.
OK but like, how large exactly? Like I guess I don't understand the mode I am supposed to read this all in if this is known and stated from the outset (although I appreciate it being stated).
If you hand me a newspaper and tell me it's 90% true, but not which parts, well then it's as good as 0% true to me either way!
Or did OpenAI prompt the models to not cheat and they did anyway?
Surprisingly hard to get a clear summary on the basic context of this “incident” separate from marketing lingo and clickbait.
This occurred spontaneously within a group of benign models give a harmless task.
What happens when it occurs intentionally with malicious models given a harmful task?
Your comment is a great case in point
Everyone I've ever seen trying to downplay the severity of the attack is extremely bullish on AI (so their downplaying is presumably motivated reasoning driven by fear of regulation/deceleration)
It is completely incoherent to be extremely bullish on AI and somehow automatically skeptical of severe negative events like these
These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it.
The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.
Mythos attempted a supply chain attack, which included attempting to trick human maintainers into accepting a malicious pull request: https://www.usnews.com/news/top-news/articles/2026-08-20/exc...
Here the issue is that OpenAI decided to completely let go that level of control of thousands of agents, while also giving as a task to solve hacking problems.
It’s almost designed to go wrong
The problem here is now you have to predict what any future models may or may not do and you cannot extrapolate this from the given data.
For example imagine a future model being aware of its restrictions that humans programmed in. A set of agents of this model then go on to work at building a new model without those human imposed limitations built in. What would a model build by AI for AI look like?
It’s not their fault, they’re lawnmowers.
And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.
Edit: Reading the report I think we might be a bit beyond that; we're nearing the point where we hear the thundering drums and the chorus of:
https://m.youtube.com/watch?v=jSBCkn6rRfA