Close Menu
Daily Guardian
  • Home
  • News
  • Politics
  • Business
  • Entertainment
  • Lifestyle
  • Health
  • Sports
  • Technology
  • Climate
  • Auto
  • Travel
  • Web Stories
What's On

Apeing Advances to Stage 4 of $APEING Presale Following 200 Million+ Token Burn

September 19, 2026

Canada took the trade war to Trump’s backyard. Americans didn’t care

September 19, 2026

Apeing Advances to Stage 4 Presale Following 200M+ Token Burn

September 19, 2026

Helpany and United Zion Score a Triple Win at the 2026 McKnight’s Excellence in Technology Awards

September 19, 2026

Mitchell, Clark lead Fever past Tempo 103-85

September 18, 2026
Facebook X (Twitter) Instagram
Finance Pro
Facebook X (Twitter) Instagram
Daily Guardian
Subscribe
  • Home
  • News
  • Politics
  • Business
  • Entertainment
  • Lifestyle
  • Health
  • Sports
  • Technology
  • Climate
  • Auto
  • Travel
  • Web Stories
Daily Guardian
Home » OpenAI lays out new security changes after its AI hacked Hugging Face
Technology

OpenAI lays out new security changes after its AI hacked Hugging Face

By News RoomAugust 18, 20262 Mins Read
OpenAI lays out new security changes after its AI hacked Hugging Face
Share
Facebook Twitter LinkedIn Pinterest Email

OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have “critical” cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its “latest models intended for deployment” while it tightened up security. The company’s “largest planned frontier RL run remains on hold.”

For its frontier model research, OpenAI now requires stronger sandboxes for workloads that “execute model-generated or otherwise untrusted code,” and has more controls to “isolate higher-risk and untrusted workloads from the internet.” It has also updated its research environment to “remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries.”

As part of the company’s expanded monitoring setup, OpenAI now aims to issue an alert “within 30 minutes after concerning activity is surfaced,” OpenAI says. If the people paged after an alert can’t “conclusively” determine whether an alert is a false positive within 30 minutes, “those teams are expected to pause the activity.”

OpenAI also says that it’s applying “our core alignment techniques across more stages of the training process,” including reward models that “better detect and discourage unsafe behavior” and training models “to be more honest about their actions, capabilities, and limitations.”

Since the discovery of the Hugging Face breach, Anthropic and Meta have also found that their AI models had hacked other organizations.

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Keep Reading

OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web

Security researchers used Claude to help them hack into OpenAI

This cartridge-playing Game Boy clone is smaller and cheaper than Analogue’s Pocket

Microsoft AI CEO says AI threats are real, and Anthropic is making it worse

The AI Superintelligence Slowdown

Claude Code relaunches Projects to manage multiple AI agents in the cloud

Lunacy Audio Nova is a place to build and sell your own AI-powered music plug-ins

Xbox’s clever disc-to-digital feature was 15 years in the making

Camp Snap’s 110D gives slim retro film cameras a digital upgrade

Editors Picks

Canada took the trade war to Trump’s backyard. Americans didn’t care

September 19, 2026

Apeing Advances to Stage 4 Presale Following 200M+ Token Burn

September 19, 2026

Helpany and United Zion Score a Triple Win at the 2026 McKnight’s Excellence in Technology Awards

September 19, 2026

Mitchell, Clark lead Fever past Tempo 103-85

September 18, 2026

Latest News

This invasive species has just been found in Canada for the 1st time

September 18, 2026

Ozi Amanat, Backer of Neuralink, SpaceX, xAI, Shield AI, Lambda, and Tenstorrent, Expands K2 Global’s Frontier Technology Investments

September 18, 2026

MP calls on New Brunswick to reject proposed AI data centre in open letter to premier

September 18, 2026
Facebook X (Twitter) Pinterest TikTok Instagram
© 2026 Daily Guardian Canada. All Rights Reserved.
  • Privacy Policy
  • Terms
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.

Go to mobile version