Close Menu
Daily Guardian
  • Home
  • News
  • Politics
  • Business
  • Entertainment
  • Lifestyle
  • Health
  • Sports
  • Technology
  • Climate
  • Auto
  • Travel
  • Web Stories
What's On

Opus One Gold Corp Releases All Drilling Results on Zone West Noyell Gold Property, Matagami, Québec

September 2, 2026

SEER Robotics Reports 67.5 Percent Revenue Growth in H1 2026

September 2, 2026

TetherMax Names South Korean Actor Yoon Shi-yoon as Brand Ambassador to Expand Presence Across Asia

September 2, 2026

Argos, Ticats meet in pivotal Labour Day contest

September 2, 2026

Troutman Amin Analyzes SB690’s Potential Impact on California CIPA Litigation

September 2, 2026
Facebook X (Twitter) Instagram
Finance Pro
Facebook X (Twitter) Instagram
Daily Guardian
Subscribe
  • Home
  • News
  • Politics
  • Business
  • Entertainment
  • Lifestyle
  • Health
  • Sports
  • Technology
  • Climate
  • Auto
  • Travel
  • Web Stories
Daily Guardian
Home » Researchers fear safety disaster ahead of OpenAI’s Astra release
Technology

Researchers fear safety disaster ahead of OpenAI’s Astra release

By News RoomSeptember 2, 20264 Mins Read
Researchers fear safety disaster ahead of OpenAI’s Astra release
Share
Facebook Twitter LinkedIn Pinterest Email

OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it “may be the single worst development for AI security/safety to date.”

Shortly after OpenAI said on Tuesday that it had delayed Astra’s release to work on safety issues, The Information reported that Astra shows far less of its “thinking” than other frontier AI models, sparking concern it could be dangerously hard to monitor.

Most top AI systems today are built using a technology known as a transformer, which processes some types of information linearly through layers before producing an answer. Models can be made to show their reasoning as they go, essentially “thinking out loud.” This “chain of thought” allows researchers and automated safety systems to monitor what AI models are doing and potentially spot undesirable behavior, such as lying or plans to circumvent safety guardrails, before they act.

According to The Information, citing an unnamed person familiar with the unreleased model’s development, Astra uses a more opaque technique known as a recurrent depth or looped transformer, which cycles information through internal layers before producing an output. This would mean much more of the model’s “thinking” happens inside the system, and in a form that looks a lot less like natural human language, rather than being expressed in a way that researchers can easily monitor. This can boost model performance, but makes potential threats and unwanted behavior harder to detect.

OpenAI has limited its use of the looped transformer / recurrent depth technique with Astra so researchers can continue to monitor the model’s reasoning, according to The Information’s unnamed source.

In a blog post published Tuesday, OpenAI said it is “deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.” It did not mention if the model has a different technical foundation.

The Information’s report sparked widespread concern among AI safety researchers on social media. It was Redwood Research’s chief scientist Ryan Greenblatt, one of three outsiders OpenAI permitted to research the Hugging Face hack, who said a decision to use a more opaque architecture for Astra “may be the single worst development for AI security/safety to date.”

Greenblatt said the investigation into the Hugging Face incident relied heavily on the models’ chain-of-thought, warning that less visible reasoning could allow AI systems to devise and execute strategies that would be far harder for researchers to detect.

Greenblatt’s primary concern, echoed by other safety experts, is that competition to develop more advanced AI systems could lead to “a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs” — with developers adopting increasingly opaque systems to gain an edge until models become difficult, or even impossible, to monitor. He added that OpenAI’s communications left him concerned that the company “plans on being extremely reliant on chain-of-thought monitoring for safety.”

OpenAI bigwigs responded to the criticism in a series of social media posts that do not explicitly deny the company’s use of the technique. Several expressed concerns about the possibility of unmonitorable AI or a race to the bottom in terms of transparency, including OpenAI safety researchers Micah Carroll and Tomek Korbak, head of strategic futures Dean Ball, and chief scientist Jakub Pachocki, who voiced fears of “a race into unmonitorability kicked off by confused reporting.” He said the depth of Astra’s computation — a measure of how many steps it can perform internally — “is within a factor of two of GPT-4,” indicating that if the technique was used, the increased opacity is less dramatic than some reactions imply. OpenAI did not respond to The Verge’s request to confirm or deny whether looped transformers were used for Astra and directed us to Pachocki’s X post.

“OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models,” Pachocki wrote, adding that such monitoring “is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon.”

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

  • Robert Hart

    Robert Hart

    Posts from this author will be added to your daily email digest and your homepage feed.

    See All by Robert Hart

  • AI

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All AI

  • News

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All News

  • OpenAI

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All OpenAI

  • Tech

    Posts from this topic will be added to your daily email digest and your homepage feed.

    See All Tech

Share. Facebook Twitter Pinterest LinkedIn Tumblr Email

Keep Reading

Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more

Amazon’s AI assistant can now spot fake emails from the company

1Password wades into a right-wing mess after funding a Linux project

The best tech and gadgets announced at IFA so far

Google is sending MrBeast into the wilderness, armed with AI

The robot butler dream doesn’t have legs

Tado’s new thermostat is designed as a European Nest killer in Europe

The amazing USB-C gadgets that play old Nintendo cartridges

Elon Musk’s heterodox robotaxi philosophy gets put to the test

Editors Picks

SEER Robotics Reports 67.5 Percent Revenue Growth in H1 2026

September 2, 2026

TetherMax Names South Korean Actor Yoon Shi-yoon as Brand Ambassador to Expand Presence Across Asia

September 2, 2026

Argos, Ticats meet in pivotal Labour Day contest

September 2, 2026

Troutman Amin Analyzes SB690’s Potential Impact on California CIPA Litigation

September 2, 2026

Latest News

Researchers fear safety disaster ahead of OpenAI’s Astra release

September 2, 2026

North Texas Food Bank Invites Community to Play Your Part During Hunger Action Month

September 2, 2026

OpenMatter Network Expands Platform with New Capabilities for Secure AI, Computing and Data Collaboration

September 2, 2026
Facebook X (Twitter) Pinterest TikTok Instagram
© 2026 Daily Guardian Canada. All Rights Reserved.
  • Privacy Policy
  • Terms
  • Advertise
  • Contact

Type above and press Enter to search. Press Esc to cancel.

Go to mobile version