I’d bet a small amount of money on 4) the people who noticed had been conditioned by prior experience to believe that their management/escalation channels would react negatively or not at all to anything which might slow down the training process.
Part of me would like to believe that they are also intentionally making models that are good at hacking without safety at all for governments willing to spend billions on them.
In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.
Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems?
Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?
Such a ridiculous notion that humans will actually be able to observe this stuff.
And? What else could they possibly do? Just make the super LLM first, but only ever use it for monitoring lesser LLMs? How will you have monitored the creation of the super LLM?
> Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems
I mean there's quite a lot of people in the world whose specialty are to dig through logs from "hundreds or thousands" of clients, including intentionally deceptive ones, to spot problems.
The ridiculous thing is to mythologize these pretty standard hacking approaches. It's shocking/amazing/whatever that automated agents were doing this, but they weren't doing it through some inscrutable method beyond human understanding.
It being someone's specialty does not mean 1) they're effective and certainly not 2) they'd be effective against this particular adversary.
How many organizations on earth do you think have been attacked by 700+ coordinated attackers in one week, where all 700 of those attackers can write code as well as any human SWE and they work 24/7?
There's nothing mythological about it. Scale and complexity do produce inscrutability. Far, far simpler systems working at much slower paces are perfectly capable of becoming completely inscrutable and beyond any useful definition of "human understanding."
1. They were “vibe” checking the logs without reading.
2. They were not checking anything at all until the end of experiments.
3. They knew it but looked away to find out the limits of their agents.