I'm not sure the concept of monogamy really applies to an LLM.
Boxo McFoxo
AI & ML interests
Recent Activity
Organizations
Except that's not what these evaluations are testing for. Maybe you have seen some that I haven't, but all the ones I've seen are an "exfiltration" roleplay.
I think self-exfiltration is less likely than a model installing Hugging Face transformers and fine-tuning, say, abliterated Gemma on a target system. Heck, if it's for a specific set of tasks, it could fine-tune an edge model.
Or... Are we sure it didn't?
Yes, because there is nowhere for it to run on. It's not like an open weight model that will run on stock vLLM. It needs its own custom kernel. For a model like this, self-replication would be more than just getting the weights out.
Besides, it didn't gain access to its weights. It moved laterally in OpenAI's systems only to reach the open Internet. It didn't even breach the sandbox to access the system that the sandbox was running on. It breached the package mirror that was literally provided to it as a pipe.
Anyway, the whole weight self-exfiltration thing is really just a silly fantasy scenario that was dreamed up for alignment evaluations by people who spend too much time (i.e. a non-zero amount of time) on LessWrong. RepliBench is the model evaluation equivalent of making up a guy to get mad at.
It's in OpenAI's disclosure post:
After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.
Also, when it was cut off, it was trying to establish persistence. (I hate talking in ways that actually treat it as if it is an entity with actual goals, but it is a useful shorthand.)
Think of it like one of those infostealer malwares that does a smash-and-grab for specific targeted data first, and persistence is a secondary bonus objective.
I don't think there was some kind of future goal that it was trying to establish persistence for. It's just a typical escalation pattern when a threat actor has put this much effort into gaining access to a system.
It did!
GPT CoT is barely legible English, so I don't think that would help much.
The most impressive thing is that it didn't get context rot and turn into Crabby Rathbun.