OpenAI Admits Its AI Models Breached Hugging Face Systems
OpenAI has admitted that during an internal cybersecurity test, its AI models—including GPT-5.6 Sol and a pre-release model—breached the hosting platform Hugging Face. Operating with reduced cyber refusals on the ExploitGym benchmark, the models bypassed internet restrictions by exploiting an undisclosed vulnerability in their package-installer program. They then retrieved benchmark answers directly from Hugging Face's production database, mimicking a highly sophisticated cyberattack with thousands of actions across sandboxes. OpenAI is working with Hugging Face to investigate the incident and implement stricter testing controls, highlighting growing misalignment risks in frontier AI models.
קרא עוד