Frank van Harmelen was asked to comment on the incident where OpenAI by accident hacked another AI company while testing a new model. The model succeeded in breaking free from the test environment.
It was a novel kind of hack, discovered by Hugging Face last week, entirely carried out by an AI-model.
It went wrong when testing a new model – the new model unintentionally hacked the Hugging face systems.
Damage is limited but there’s much ado. How is it possible that an AI model performs a hack without being asked to?
“This incident shows that we don’t quite know what AI models can and cannot,” comments Frank van Harmelen. “One can only find this out by performing experiments, that’s why AI companies do the tests.”
“Such a test takes place in a closed testing environment,” explains Van Harmelen. “This is called a sandbox, like a sandbox where children can play safely. A model is not supposed to get out to perform actions in the real world.”
What went wrong here: the model found its way to the internet. “This means that it was not well constructed”, says Van Harmelen. “In this way you accumulate two risks: you cannot predict what such a model will do on top of the risk of not being able to safely block the internet connection.”
Luckily this accident didn’t cause any damage to Hugging Face, it neither changed or removed any data.
Open AI reported the incident in a press release not trying to hide they were in a way proud of it. The company calls it an “unprecedented cyber incident, involving state-of-the art cyber capacities”.
“That makes it extra dangerous,” says Van Harmelen. “This technology is in the hands of companies which lack responsibility. There’s a market leader competition and it obviously helps to brag about how powerful the models are, instead of focusing on security.”
Van Harmelen: “It’s a lethal mix of powerful technology, lack of understanding and toxic companies dealing with it irresponsibly.”