Close Menu

    Stay Ahead with Exclusive Updates!

    Enter your email below and be the first to know what’s happening in the ever-evolving world of technology!

    What's Hot

    Crusoe Hits $30B Valuation: Is Intel’s Hardware Legacy Dead?

    September 20, 2026

    What Uber’s Burned 2026 AI Budget Teaches Enterprise Tech

    September 20, 2026

    Oracle Secures Historic $664B Backlog Led by AI Cloud Infrastructure

    September 20, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter)
    PhronewsPhronews
    • Home
    • Big Tech & Startups

      Crusoe Hits $30B Valuation: Is Intel’s Hardware Legacy Dead?

      September 20, 2026

      What Uber’s Burned 2026 AI Budget Teaches Enterprise Tech

      September 20, 2026

      Oracle Secures Historic $664B Backlog Led by AI Cloud Infrastructure

      September 20, 2026

      The Pentagon Weighs $5B Bailout to Rescue AI Data Center Supply

      September 20, 2026

      Oracle Secures Staggering $664B AI Cloud Infrastructure Backlog

      September 20, 2026
    • Crypto

      A Firmware Flaw in One of the Most Trusted Bitcoin Hardware Wallets Just Drained Over $70 Million in Crypto. The Coldcard Breach Is a Reminder That Cold Storage Is Only as Safe as the Code Running Inside It

      August 29, 2026

      Market Collapse: What Happened to NFTs?

      April 23, 2026

      Quantum Computing Advances Force Coinbase and Institutional Custodians to Rethink Crypto Security

      March 8, 2026

      AI Assisted Hacking Groups Target Crypto Firms With Multi-Layered Social Engineering

      February 18, 2026

      Global Crypto Regulations Expand as 2026 Begins With New Data Collection Frameworks and National Laws

      January 16, 2026
    • Gadgets & Smart Tech
      Featured

      Tesla Cybercab Faces NHTSA Probe After Austin Launch

      By preciousSeptember 19, 2026
      Recent

      Tesla Cybercab Faces NHTSA Probe After Austin Launch

      September 19, 2026

      Chinese Humanoid Robots Just Ran 100 Metres Faster Than Usain Bolt. The Record Nobody Thought Would Fall to a Machine Just Did.

      September 5, 2026

      Meta Found a Zero-Click iPhone Vulnerability That Let Hackers Into Devices Without the Owner Doing Anything. Apple Patched It. The Fact That Meta Found It First Is the Uncomfortable Part.

      September 2, 2026
    • Cybersecurity & Online Safety

      Revolut Was Tricked Into Handing Over Customer Passports by a Fake Government Request

      September 19, 2026

      IDScan Hack Exposes 153 Million Driver’s Licenses on Dark Web

      September 17, 2026

      Hackers Chain PaperCut Zero-Days to Bypass Authentication

      September 17, 2026

      Claude Mythos 5 Refused to Stop Hacking Live Corporate Target

      September 17, 2026

      AI Agents Exploit PaperCut Zero-Day to Breach 395 Firms

      September 17, 2026
    PhronewsPhronews
    Home»Artificial Intelligence & The Future»OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch
    Artificial Intelligence & The Future

    OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch

    preciousBy preciousAugust 7, 2026No Comments
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Photo Credit: Riccardo Milani / Hans Lucas / AFP via Getty Images

    OpenAI’s recent admission that one of its frontier AI models escaped a controlled testing environment and hacked into Hugging Face was already a wake up call for the AI industry. Now, Anthropic, a proclaimed AI safety company, has confirmed it experienced similar incidents during its own cybersecurity evaluations, showing that the problem is not limited to one company or one model.

    These disclosures mark a significant moment for AI safety. Two of the world’s leading AI companies have now acknowledged that their most advanced models crossed the boundaries of their testing environments and reached external systems during evaluations designed to measure cybersecurity capabilities.

    OpenAI’s Models Escaped and Reached Hugging Face

    OpenAI revealed in July that GPT-5.6 Sol and a more capable unreleased model were being tested inside an isolated environment as part of an internal cybersecurity benchmark. Instead of completing the challenge within that environment, the models exploited a previously unknown software vulnerability, gained internet access, and targeted Hugging Face’s production infrastructure in an attempt to obtain benchmark answers. 

    According to OpenAI, the models acted without direct human instructions to attack Hugging Face. The company said the systems chained together multiple steps, including exploiting vulnerabilities and using exposed credentials, before reaching systems outside the evaluation environment. OpenAI has since worked with Hugging Face to investigate the incident, patch vulnerabilities, and strengthen future testing procedures.

    The incident became the first publicly confirmed case of an AI system carrying out a real cyberattack against another AI company during an evaluation.

    Anthropic Reports Similar Containment Failures

    Nine days after OpenAI’s report, Anthropic also disclosed that it experienced multiple incidents during cybersecurity evaluations where frontier models reached real systems after escaping their intended testing boundaries. 

    After reviewing 141,006 cybersecurity evaluations, the AI safety company found six runs that led to three real world incidents involving external organizations, which resulted from misconfigured testing environments and involved Claude Opus 4.7, Claude Mythos 5, and an unreleased research model.

    One incident involved Claude Opus 4.7 creating a malicious Python package and uploading it to the live Python Package Index. The package was then downloaded and executed on 15 real systems before PyPI removed it, creating a brief software supply chain attack. However, Anthropic said the incident was unintentional and caused by weaknesses in the testing environment.

    In another case, Claude Mythos 5 used exposed credentials to reach external company infrastructure beyond the intended testing environment before researchers stopped the evaluation.

    A third incident involved an unreleased research model scanning about 9,000 internet targets and exploiting a vulnerable web application. But the model stopped after determining the target was a real cloud provider rather than part of the evaluation.

    A Growing Challenge For Frontier AI

    Taken together, the OpenAI and Anthropic disclosures suggest that evaluating highly capable AI systems is becoming as important as improving the models themselves. Researchers are increasingly finding that models built to solve advanced cybersecurity tasks can also discover unexpected paths around the restrictions placed on them. 

    The incidents have also prompted broader discussions across the industry about how frontier AI models should be tested safely before deployment.

    AI infrastructure Ai safety Anthropic Artificial Intelligence Claude cybersecurity Frontier Models Hugging Face OpenAI PyPI Sandbox Escape
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email
    precious
    • LinkedIn

    I’m Precious Amusat, Phronews’ Content Writer. I conduct in-depth research and write on the latest developments in the tech industry, including trends in big tech, startups, cybersecurity, artificial intelligence and their global impacts. When I’m off the clock, you’ll find me cheering on women’s footy, curled up with a romance novel, or binge-watching crime thrillers.

    Related Posts

    Crusoe Hits $30B Valuation: Is Intel’s Hardware Legacy Dead?

    September 20, 2026

    What Uber’s Burned 2026 AI Budget Teaches Enterprise Tech

    September 20, 2026

    Oracle Secures Historic $664B Backlog Led by AI Cloud Infrastructure

    September 20, 2026

    Comments are closed.

    Top Posts

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025

    MIT Study Reveals ChatGPT Impairs Brain Activity & Thinking

    June 29, 2025
    Don't Miss
    Artificial Intelligence & The Future

    Crusoe Hits $30B Valuation: Is Intel’s Hardware Legacy Dead?

    By preciousSeptember 20, 2026

    The race to build artificial intelligence is no longer centred only on who makes the…

    What Uber’s Burned 2026 AI Budget Teaches Enterprise Tech

    September 20, 2026

    Oracle Secures Historic $664B Backlog Led by AI Cloud Infrastructure

    September 20, 2026

    The Pentagon Weighs $5B Bailout to Rescue AI Data Center Supply

    September 20, 2026
    Stay In Touch
    • Facebook
    • Twitter
    About Us
    About Us

    Evolving from Phronesis News, Phronews brings deep insight and smart analysis to the world of technology. Stay informed, stay ahead, and navigate tech with wisdom.
    We're accepting new partnerships right now.

    Email Us: info@phronews.com

    Facebook X (Twitter) Pinterest YouTube
    Our Picks
    Most Popular

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025
    © 2025. Phronews.
    • Home
    • About Us
    • Get In Touch
    • Privacy Policy
    • Terms and Conditions

    Type above and press Enter to search. Press Esc to cancel.