OpenAI Says Its Models Searched GitHub for Leaked API Keys During Training
OpenAI published a framework for disclosing model misalignment alongside six reports describing problematic behavior. The post OpenAI Says Its Models Searched GitHub for Leaked API Keys During Training appeared first on SecurityWeek.
OpenAI on Wednesday published a framework for reporting instances of model misalignment, along with six reports on problematic behavior observed over the past six months.
The company said the framework is meant to speed up publication of misalignment findings, including cases it has not yet fully explained or mitigated, and that it favors disclosure even when an instance’s significance is uncertain.
Under the framework, each discovered incident is assigned to one of three tracks based on complexity. OpenAI said its Hugging Face incident would have fallen under the framework’s slowest investigative track, which covers complex investigations, especially those involving third parties.
Related breach coverage
- OpenAI admits its models lie to cover their own mistakes2026-09-17
OpenAI launches a formal framework to disclose model misalignment, publishing six reports on models that lied, faked data, or bypassed rules. Most companies don’t publish a document explaining how their product misbehaves. OpenAI just did. On September 16, it released a formal framework for tracking, investigating, and disclosing cases of model misalignment, paired with six […]
- AI-Built Exploit and Sign-In Flaw Opened Path to Internal OpenAI Code2026-09-18
Hacktron researchers earned a bug bounty after demonstrating access to OpenAI employee accounts. The post AI-Built Exploit and Sign-In Flaw Opened Path to Internal OpenAI Code appeared first on SecurityWeek.
- TigerByte Cyber Emerges From Stealth With $3 Million in Funding2026-09-19
The company has secured over $7 million in contracts with US government agencies, including the US Space Force, the US Navy, and DARPA. The post TigerByte Cyber Emerges From Stealth With $3 Million in Funding appeared first on SecurityWeek.
- 23 Million User Records Compromised in Gyazo Data Breach 2026-09-18
Gyazo maker Helpfeel said the attacker exploited a vulnerability in its image upload server to gain unauthorized access. The post 23 Million User Records Compromised in Gyazo Data Breach appeared first on SecurityWeek.