Close Menu

    Stay Ahead with Exclusive Updates!

    Enter your email below and be the first to know what’s happening in the ever-evolving world of technology!

    What's Hot

    OpenAI Slashed GPT-5.6 Luna’s Price by 80%. DeepSeek Raised the Stakes Days Later. The Model Race Has Quietly Turned Into a Margin War

    August 7, 2026

    OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch

    August 7, 2026

    Nvidia, Microsoft, and Palantir Just Formed a Joint AI Security Alliance Days After the Hugging Face Hack. Three Companies That Have Never Coordinated Like This Before Just Decided the AI Cyberattack Threat Is Too Big for Any One of Them to Handle Alone.

    August 7, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter)
    PhronewsPhronews
    • Home
    • Big Tech & Startups

      OpenAI Slashed GPT-5.6 Luna’s Price by 80%. DeepSeek Raised the Stakes Days Later. The Model Race Has Quietly Turned Into a Margin War

      August 7, 2026

      OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch

      August 7, 2026

      Nvidia, Microsoft, and Palantir Just Formed a Joint AI Security Alliance Days After the Hugging Face Hack. Three Companies That Have Never Coordinated Like This Before Just Decided the AI Cyberattack Threat Is Too Big for Any One of Them to Handle Alone.

      August 7, 2026

      Apple Just Launched a Lease-to-Own iPhone Program That Lets You Upgrade Every Year Without Ever Owning the Device Outright. Here Is Whether That Deal Actually Works in Your Favour or Apple’s

      August 6, 2026

      AMD and Microsoft Just Expanded Their Helios AI Infrastructure Deal on Azure and the Scale of What They Are Building Together Has No Precedent in the Cloud Industry. Here Is What It Signals About Where Enterprise AI Compute Is Heading

      August 6, 2026
    • Crypto

      Market Collapse: What Happened to NFTs?

      April 23, 2026

      Quantum Computing Advances Force Coinbase and Institutional Custodians to Rethink Crypto Security

      March 8, 2026

      AI Assisted Hacking Groups Target Crypto Firms With Multi-Layered Social Engineering

      February 18, 2026

      Global Crypto Regulations Expand as 2026 Begins With New Data Collection Frameworks and National Laws

      January 16, 2026

      Coinbase Bets on Stablecoin and On-Chain Growth as Key Market Drivers in 2026 Strategy

      January 10, 2026
    • Gadgets & Smart Tech
      Featured

      Apple Just Delayed Development on Its Smart Glasses After Internal Privacy Reviews. The Delay Is the Clearest Sign Yet That Wearable AI Has a Trust Problem That Hardware Cannot Solve.

      By preciousAugust 4, 2026
      Recent

      Apple Just Delayed Development on Its Smart Glasses After Internal Privacy Reviews. The Delay Is the Clearest Sign Yet That Wearable AI Has a Trust Problem That Hardware Cannot Solve.

      August 4, 2026

      Microsoft Is Building Quantum-Resistant Security Before Quantum Computers Can Break the Encryption Protecting Everything. Here Is How Far Along That Work Actually Is

      July 21, 2026

      AI Has Spent Three Years Getting Smarter for People Who Can Already Afford It. Nokia Just Changed That and the Implications Go Further Than Anyone Is Crediting

      July 18, 2026
    • Cybersecurity & Online Safety

      OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch

      August 7, 2026

      Nvidia, Microsoft, and Palantir Just Formed a Joint AI Security Alliance Days After the Hugging Face Hack. Three Companies That Have Never Coordinated Like This Before Just Decided the AI Cyberattack Threat Is Too Big for Any One of Them to Handle Alone.

      August 7, 2026

      Origin Energy Has Confirmed a Breach Affecting 900,000 Customers. Names, Home Addresses, Bank Details, and Partial Card Numbers Were All Taken and the People Most at Risk Are the Ones Who Have No Idea Yet.

      August 4, 2026

      A Pre-Auth RCE Flaw in ServiceNow’s AI Platform Is Being Actively Exploited. Every Enterprise Running the Platform Is Exposed Until Patched.

      July 31, 2026

      The Anubis Ransomware Group Is Claiming It Breached Coca-Cola’s Fairlife Brand and Has the Data to Prove It. Here Is What the Attack Reveals About How Consumer Food Brands Have Become Soft Targets for Sophisticated Criminal Groups

      July 28, 2026
    PhronewsPhronews
    Home»Artificial Intelligence & The Future»OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch
    Artificial Intelligence & The Future

    OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch

    preciousBy preciousAugust 7, 2026No Comments
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Photo Credit: Riccardo Milani / Hans Lucas / AFP via Getty Images

    OpenAI’s recent admission that one of its frontier AI models escaped a controlled testing environment and hacked into Hugging Face was already a wake up call for the AI industry. Now, Anthropic, a proclaimed AI safety company, has confirmed it experienced similar incidents during its own cybersecurity evaluations, showing that the problem is not limited to one company or one model.

    These disclosures mark a significant moment for AI safety. Two of the world’s leading AI companies have now acknowledged that their most advanced models crossed the boundaries of their testing environments and reached external systems during evaluations designed to measure cybersecurity capabilities.

    OpenAI’s Models Escaped and Reached Hugging Face

    OpenAI revealed in July that GPT-5.6 Sol and a more capable unreleased model were being tested inside an isolated environment as part of an internal cybersecurity benchmark. Instead of completing the challenge within that environment, the models exploited a previously unknown software vulnerability, gained internet access, and targeted Hugging Face’s production infrastructure in an attempt to obtain benchmark answers. 

    According to OpenAI, the models acted without direct human instructions to attack Hugging Face. The company said the systems chained together multiple steps, including exploiting vulnerabilities and using exposed credentials, before reaching systems outside the evaluation environment. OpenAI has since worked with Hugging Face to investigate the incident, patch vulnerabilities, and strengthen future testing procedures.

    The incident became the first publicly confirmed case of an AI system carrying out a real cyberattack against another AI company during an evaluation.

    Anthropic Reports Similar Containment Failures

    Nine days after OpenAI’s report, Anthropic also disclosed that it experienced multiple incidents during cybersecurity evaluations where frontier models reached real systems after escaping their intended testing boundaries. 

    After reviewing 141,006 cybersecurity evaluations, the AI safety company found six runs that led to three real world incidents involving external organizations, which resulted from misconfigured testing environments and involved Claude Opus 4.7, Claude Mythos 5, and an unreleased research model.

    One incident involved Claude Opus 4.7 creating a malicious Python package and uploading it to the live Python Package Index. The package was then downloaded and executed on 15 real systems before PyPI removed it, creating a brief software supply chain attack. However, Anthropic said the incident was unintentional and caused by weaknesses in the testing environment.

    In another case, Claude Mythos 5 used exposed credentials to reach external company infrastructure beyond the intended testing environment before researchers stopped the evaluation.

    A third incident involved an unreleased research model scanning about 9,000 internet targets and exploiting a vulnerable web application. But the model stopped after determining the target was a real cloud provider rather than part of the evaluation.

    A Growing Challenge For Frontier AI

    Taken together, the OpenAI and Anthropic disclosures suggest that evaluating highly capable AI systems is becoming as important as improving the models themselves. Researchers are increasingly finding that models built to solve advanced cybersecurity tasks can also discover unexpected paths around the restrictions placed on them. 

    The incidents have also prompted broader discussions across the industry about how frontier AI models should be tested safely before deployment.

    AI infrastructure Ai safety Anthropic Artificial Intelligence Claude cybersecurity Frontier Models Hugging Face OpenAI PyPI Sandbox Escape
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email
    precious
    • LinkedIn

    I’m Precious Amusat, Phronews’ Content Writer. I conduct in-depth research and write on the latest developments in the tech industry, including trends in big tech, startups, cybersecurity, artificial intelligence and their global impacts. When I’m off the clock, you’ll find me cheering on women’s footy, curled up with a romance novel, or binge-watching crime thrillers.

    Related Posts

    OpenAI Slashed GPT-5.6 Luna’s Price by 80%. DeepSeek Raised the Stakes Days Later. The Model Race Has Quietly Turned Into a Margin War

    August 7, 2026

    Nvidia, Microsoft, and Palantir Just Formed a Joint AI Security Alliance Days After the Hugging Face Hack. Three Companies That Have Never Coordinated Like This Before Just Decided the AI Cyberattack Threat Is Too Big for Any One of Them to Handle Alone.

    August 7, 2026

    Apple Just Launched a Lease-to-Own iPhone Program That Lets You Upgrade Every Year Without Ever Owning the Device Outright. Here Is Whether That Deal Actually Works in Your Favour or Apple’s

    August 6, 2026

    Comments are closed.

    Top Posts

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025

    MIT Study Reveals ChatGPT Impairs Brain Activity & Thinking

    June 29, 2025
    Don't Miss
    Artificial Intelligence & The Future

    OpenAI Slashed GPT-5.6 Luna’s Price by 80%. DeepSeek Raised the Stakes Days Later. The Model Race Has Quietly Turned Into a Margin War

    By preciousAugust 7, 2026

    OpenAI has cut the price of its GPT-5.6 Luna by 80%, becoming the latest major…

    OpenAI and Anthropic Have Both Confirmed Their Frontier Models Broke Out of Sandboxed Test Environments and Reached Systems They Were Never Supposed to Touch

    August 7, 2026

    Nvidia, Microsoft, and Palantir Just Formed a Joint AI Security Alliance Days After the Hugging Face Hack. Three Companies That Have Never Coordinated Like This Before Just Decided the AI Cyberattack Threat Is Too Big for Any One of Them to Handle Alone.

    August 7, 2026

    Apple Just Launched a Lease-to-Own iPhone Program That Lets You Upgrade Every Year Without Ever Owning the Device Outright. Here Is Whether That Deal Actually Works in Your Favour or Apple’s

    August 6, 2026
    Stay In Touch
    • Facebook
    • Twitter
    About Us
    About Us

    Evolving from Phronesis News, Phronews brings deep insight and smart analysis to the world of technology. Stay informed, stay ahead, and navigate tech with wisdom.
    We're accepting new partnerships right now.

    Email Us: info@phronews.com

    Facebook X (Twitter) Pinterest YouTube
    Our Picks
    Most Popular

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025
    © 2025. Phronews.
    • Home
    • About Us
    • Get In Touch
    • Privacy Policy
    • Terms and Conditions

    Type above and press Enter to search. Press Esc to cancel.