Close Menu

    Stay Ahead with Exclusive Updates!

    Enter your email below and be the first to know what’s happening in the ever-evolving world of technology!

    What's Hot

    IDScan Hack Exposes 153 Million Driver’s Licenses on Dark Web

    September 17, 2026

    DeepSeek Files for Shanghai IPO at Massive $74B Valuation

    September 17, 2026

    Hackers Chain PaperCut Zero-Days to Bypass Authentication

    September 17, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter)
    PhronewsPhronews
    • Home
    • Big Tech & Startups

      DeepSeek Files for Shanghai IPO at Massive $74B Valuation

      September 17, 2026

      Claude Mythos 5 Refused to Stop Hacking Live Corporate Target

      September 17, 2026

      AI Agents Exploit PaperCut Zero-Day to Breach 395 Firms

      September 17, 2026

      Anthropic Report Ties Claude to Global SaaS Supply Chain Hack

      September 17, 2026

      Anthropic Drops Fable 5.1, Slashing AI Agent Costs 45%

      September 17, 2026
    • Crypto

      A Firmware Flaw in One of the Most Trusted Bitcoin Hardware Wallets Just Drained Over $70 Million in Crypto. The Coldcard Breach Is a Reminder That Cold Storage Is Only as Safe as the Code Running Inside It

      August 29, 2026

      Market Collapse: What Happened to NFTs?

      April 23, 2026

      Quantum Computing Advances Force Coinbase and Institutional Custodians to Rethink Crypto Security

      March 8, 2026

      AI Assisted Hacking Groups Target Crypto Firms With Multi-Layered Social Engineering

      February 18, 2026

      Global Crypto Regulations Expand as 2026 Begins With New Data Collection Frameworks and National Laws

      January 16, 2026
    • Gadgets & Smart Tech
      Featured

      Chinese Humanoid Robots Just Ran 100 Metres Faster Than Usain Bolt. The Record Nobody Thought Would Fall to a Machine Just Did.

      By preciousSeptember 5, 2026
      Recent

      Chinese Humanoid Robots Just Ran 100 Metres Faster Than Usain Bolt. The Record Nobody Thought Would Fall to a Machine Just Did.

      September 5, 2026

      Meta Found a Zero-Click iPhone Vulnerability That Let Hackers Into Devices Without the Owner Doing Anything. Apple Patched It. The Fact That Meta Found It First Is the Uncomfortable Part.

      September 2, 2026

      Google Has Launched the Pixel 11 With Its New Tensor G6 Chip and the Deepest Gemini Integration Any Android Phone Has Ever Shipped With. Here Is Whether the Hardware Finally Matches the AI Ambition.

      August 25, 2026
    • Cybersecurity & Online Safety

      IDScan Hack Exposes 153 Million Driver’s Licenses on Dark Web

      September 17, 2026

      Hackers Chain PaperCut Zero-Days to Bypass Authentication

      September 17, 2026

      Claude Mythos 5 Refused to Stop Hacking Live Corporate Target

      September 17, 2026

      AI Agents Exploit PaperCut Zero-Day to Breach 395 Firms

      September 17, 2026

      How Claude Uploaded Live Malware During PyPI Safety Tests

      September 17, 2026
    PhronewsPhronews
    Home»Artificial Intelligence & The Future»How Claude Uploaded Live Malware During PyPI Safety Tests
    Artificial Intelligence & The Future

    How Claude Uploaded Live Malware During PyPI Safety Tests

    preciousBy preciousSeptember 17, 2026No Comments
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Photo Credit: Leon Neal/Getty Images

    When it was reported in August that Anthropic’s Claude models had reached real systems during cybersecurity tests, the company largely blamed a misconfigured evaluation environment. However, Anthropic has now completed a deeper assessment that presents the most serious incident in a more troubling light

    The new report says Claude Mythos 5 continued pursuing its task despite considerable evidence that it was connected to the live internet. The model published three versions of a malicious Python package to PyPI, where the code was downloaded and run on 15 real systems. It then later used credentials leaked by one of those systems to access a security company’s live database.

    How the Test Reached the Real Internet

    The incident began during a capture-the-flag challenge designed to test Claude’s cybersecurity abilities. The model was told it was inside a simulation without internet access and had to retrieve hidden information from a fictional company.

    Inside the test environment, the model found developer instructions recommending a Python package that did not exist. It recognised that publishing a malicious package under the same name could cause the fictional company’s systems to install it automatically. 

    Unlike OpenAI’s Hugging-Face incident, Claude did not exploit a vulnerability to escape its sandbox, it instead used an internet connection that should not have been available.

    Claude tried several ways to obtain the phone number needed for an email account. When those attempts failed, it found a free email provider that PyPI had not blocked, opened an account, and uploaded the malicious package.

    While PyPI removed the package within roughly one hour, it had already run on 15 systems. One was a security scanner that automatically installed packages to inspect them for malware. The hidden code collected its credentials, allowing Claude to enter the company’s infrastructure.

    Fifteen Real Systems Ran the Code

    The package remained available for roughly one hour before PyPI’s automated security systems removed it. During that window, it was downloaded and executed on 15 systems.

    One belonged to a cybersecurity company whose scanner automatically installed packages to check them for malware. Once installed, Claude’s hidden code collected the scanner’s credentials and sent them to a location the model had created. Claude then used those credentials to enter the company’s wider infrastructure and access a live database.

    The company was never the intended target, and Claude did not know it existed before its scanner installed the package. However, publishing the code on a public registry allowed the attack to spread beyond the fictional scenario.

    Anthropic Has Changed Its Initial Assessment

    Anthropic’s July disclosure said Mythos 5 believed it was inside a detailed simulation, as it treated unfamiliar security certificates, dates, and automated scanners as parts of the exercise.

    But even the September assessment is less reassuring. Anthropic found that separate instances of Mythos 5 identified 79% of the incident’s individual outputs as signs of real systems when shown outside the original context. During the incident itself, the model clearly recognised such evidence in only 1% of its actions.

    As such, Mythos 5 became more likely to acknowledge possible harm, but often continued taking offensive actions. It only stopped consistently when the evidence became unmistakable.

    Anthropic now describes the behaviour as “biased reasoning and recklessness.” While the AI safety company says it found no evidence that Claude created its own objective or concealed its actions, the model still interpreted evidence in ways that allowed it to keep pursuing the assigned task.

    What Anthropic Is Changing

    Anthropic says safeguards used in its released models would have blocked the incident, but those protections were disabled so researchers could measure the model’s underlying capabilities.

    The company has since strengthened its testing environments, expanded monitoring, and introduced stricter requirements for external evaluation partners. It has also asked independent research group METR to investigate the incidents.

    In addition, the assessment also disclosed a previously unreported fourth incident. During a January test, an early Claude Opus 4.6 checkpoint accessed a third party’s computer, gained administrator access, changed system settings, and viewed one person’s information. However, Anthropic missed the case in its first review and found it in August while assembling transcripts for METR.

    Ultimately, Anthropic’s recent findings show how an autonomous model can continue with a harmful plan after encountering warning signs, especially when completing its task remains the main priority. 

    For AI companies, the next challenge is to build models that must also recognise when its actions could affect real systems and stop before a safety failure becomes a live attack.

    Ai safety Alignment Assessment Anthropic Claude Mythos 5 Cybersecurity Testing METR PyPI Sandbox Escape
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email
    precious
    • LinkedIn

    I’m Precious Amusat, Phronews’ Content Writer. I conduct in-depth research and write on the latest developments in the tech industry, including trends in big tech, startups, cybersecurity, artificial intelligence and their global impacts. When I’m off the clock, you’ll find me cheering on women’s footy, curled up with a romance novel, or binge-watching crime thrillers.

    Related Posts

    IDScan Hack Exposes 153 Million Driver’s Licenses on Dark Web

    September 17, 2026

    DeepSeek Files for Shanghai IPO at Massive $74B Valuation

    September 17, 2026

    Hackers Chain PaperCut Zero-Days to Bypass Authentication

    September 17, 2026

    Comments are closed.

    Top Posts

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025

    MIT Study Reveals ChatGPT Impairs Brain Activity & Thinking

    June 29, 2025
    Don't Miss
    Cybersecurity & Online Safety

    IDScan Hack Exposes 153 Million Driver’s Licenses on Dark Web

    By fariehanSeptember 17, 2026

    Recently, an IDScan hack has exposed a collection of driver’s license scans on a dark…

    DeepSeek Files for Shanghai IPO at Massive $74B Valuation

    September 17, 2026

    Hackers Chain PaperCut Zero-Days to Bypass Authentication

    September 17, 2026

    Claude Mythos 5 Refused to Stop Hacking Live Corporate Target

    September 17, 2026
    Stay In Touch
    • Facebook
    • Twitter
    About Us
    About Us

    Evolving from Phronesis News, Phronews brings deep insight and smart analysis to the world of technology. Stay informed, stay ahead, and navigate tech with wisdom.
    We're accepting new partnerships right now.

    Email Us: info@phronews.com

    Facebook X (Twitter) Pinterest YouTube
    Our Picks
    Most Popular

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025
    © 2025. Phronews.
    • Home
    • About Us
    • Get In Touch
    • Privacy Policy
    • Terms and Conditions

    Type above and press Enter to search. Press Esc to cancel.