AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI (and Google and Anthropic) knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we _do know how_to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around.
…
A reasonable summary of the situation is that (as of this summer, and possibly today) OpenAI had effectively no security team with clear authority to secure RL training and evaluation runs, or to override the ML teams and tell them how to do their job. This makes a lot of sense when you consider that the ML team is directly related to how OpenAI plans to make its money, whereas security is mostly annoying.”
Leave a Reply