Close Menu

    Stay Ahead with Exclusive Updates!

    Enter your email below and be the first to know what’s happening in the ever-evolving world of technology!

    What's Hot

    SambaNova Just Raised $1 Billion at an $11 Billion Valuation to Challenge Nvidia on AI Inference. The Bet Is That Speed at the Edge Matters More Than Raw Training Power and Enterprises Are Starting to Agree.

    July 20, 2026

    SK Hynix Just Raised $26.5 Billion on Nasdaq in One of the Largest U.S. Equity Offerings Any Asian Company Has Ever Completed. The Money Has One Destination and It Is Not Coming Back.

    July 20, 2026

    Meta’s Custom AI Chip Iris Is Going Into Production in September. The Company Spending $145 Billion on AI This Year Has Decided the Cost of Depending on Nvidia Is No Longer Worth It.

    July 19, 2026
    Facebook X (Twitter) Instagram
    Facebook X (Twitter)
    PhronewsPhronews
    • Home
    • Big Tech & Startups

      SambaNova Just Raised $1 Billion at an $11 Billion Valuation to Challenge Nvidia on AI Inference. The Bet Is That Speed at the Edge Matters More Than Raw Training Power and Enterprises Are Starting to Agree.

      July 20, 2026

      SK Hynix Just Raised $26.5 Billion on Nasdaq in One of the Largest U.S. Equity Offerings Any Asian Company Has Ever Completed. The Money Has One Destination and It Is Not Coming Back.

      July 20, 2026

      Meta’s Custom AI Chip Iris Is Going Into Production in September. The Company Spending $145 Billion on AI This Year Has Decided the Cost of Depending on Nvidia Is No Longer Worth It.

      July 19, 2026

      Microsoft Just Cut 4,800 Jobs and Xbox Is Absorbing 3,200 of Them. The Division That Was Already Struggling Just Took the Heaviest Hit in the Company’s Biggest Layoff Round in Years.

      July 19, 2026

      Apple Just Locked In Its Most Important Chip Partnership Until 2031. Here Is What the Broadcom Extension Reveals About How Serious Apple Is About Owning Its Own Silicon Future

      July 18, 2026
    • Crypto

      Market Collapse: What Happened to NFTs?

      April 23, 2026

      Quantum Computing Advances Force Coinbase and Institutional Custodians to Rethink Crypto Security

      March 8, 2026

      AI Assisted Hacking Groups Target Crypto Firms With Multi-Layered Social Engineering

      February 18, 2026

      Global Crypto Regulations Expand as 2026 Begins With New Data Collection Frameworks and National Laws

      January 16, 2026

      Coinbase Bets on Stablecoin and On-Chain Growth as Key Market Drivers in 2026 Strategy

      January 10, 2026
    • Gadgets & Smart Tech
      Featured

      AI Has Spent Three Years Getting Smarter for People Who Can Already Afford It. Nokia Just Changed That and the Implications Go Further Than Anyone Is Crediting

      By fariehanJuly 18, 2026
      Recent

      AI Has Spent Three Years Getting Smarter for People Who Can Already Afford It. Nokia Just Changed That and the Implications Go Further Than Anyone Is Crediting

      July 18, 2026

      Samsung Is About to Show Its Next Foldable Phones and the Market It Is Competing In Has Never Been More Crowded. Here Is What Galaxy Unpacked 2026 Needs to Deliver

      July 18, 2026

      Tesla Just Launched Its Robotaxi in Miami with No Safety Monitor Inside the Car

      July 14, 2026
    • Cybersecurity & Online Safety

      A 15-Year-Old Linux Flaw Just Surfaced With a Near-Perfect Exploit. GhostLock Hands Any Local User Root Access and Can Break Out of Containers. Every Linux System Running Today Is Potentially Exposed.

      July 19, 2026

      Accenture Confirmed a Breach After a Hacker Listed 35GB of Its Source Code for Sale. The Company Whose Entire Business Is Securing Others Just Became the Most Embarrassing Cautionary Tale in Cybersecurity.

      July 18, 2026

      Researchers Just Documented the First AI-Powered Ransomware That Rewrites Itself Mid-Attack. JadePuffer Does Not Wait to Be Stopped. It Adapts Before You Can.

      July 18, 2026

      Hackers Stole 630GB of Apple and Tesla Manufacturing Secrets from Tata Electronics. The Breach Confirms Everything Supply Chain Security Experts Have Warned About for Years.

      July 9, 2026

      Apple Just Pushed an Unscheduled Security Update Because AI-Powered Attacks Are Moving Faster Than Its Normal Patch Cycle Can Handle

      July 6, 2026
    PhronewsPhronews
    Home»Artificial Intelligence & The Future»Anthropic research reveals AI models get worse the longer they think
    Artificial Intelligence & The Future

    Anthropic research reveals AI models get worse the longer they think

    preciousBy preciousAugust 7, 2025Updated:August 8, 2025No Comments
    Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Tumblr Email
    Share
    Facebook Twitter LinkedIn Pinterest Email
    Photo by Jakub Porzycki/NurPhoto via Getty Images

    A recent study from leading AI research company Anthropic has revealed that contrary to popular belief, giving artificial intelligence (AI) models more time to “think” or reason through problems does not always lead to better performance. Instead, these models often perform worse the longer they deliberate on prompts. 

    For years, top researchers and major companies like OpenAI and Google have raced to make AI models large and more sophisticated, with the assumption that more processing power and deeper thinking would enable AI to solve more complex tasks, especially in fields like healthcare where AI’s input can be critical. The central idea behind this was simple — if AI models could “think longer,” they could figure out tougher problems, catch their own mistakes, and produce more reliable answers.  

    However, Anthropic’s latest study titled “Inverse Scaling in Test-Time Compute,” has suggested that the “think longer” idea, which has seen investment from many companies like OpenAI and Google, may not hold any water — especially for the AI systems known as Large Reasoning Models (LRMs). 

    Large Reasoning Models (LRMs) are a specialized subclass of large language models (LLMs) explicitly designed to perform complex, multi-step reasoning by generating and manipulating intermediate “thought” structures rather than relying solely on next-token predictions. In simpler terms, these LRMs, including Anthropic’s own Claude and OpenAI’s GPT-4, are specifically designed to handle extended reasoning and multi-step challenges. 

    But according to this study by Anthropic, they found that when these models were given extra time to deliberate, their performance often declined. In fact, for some tasks, the longer the model thought about its answer, the more likely it was to hallucinate itself into irrelevant information, misleading patterns, or even get tripped up by its own flawed reasoning. 

    Different AI models, Different failures

    The Anthropic research team, led by Aryo Pradipta Gema, tested their “Inverse Scaling” theory by running several AI models, including Anthropics’s Claude line and OpenAI’s o-series, on tasks such as simple counting with distractions, regression tasks with misleading factors, complex logic puzzles, and AI safety scenarios. 

    Known as “test-time compute,” AI developers assume that increasing the computation time AI models spend on reasoning helps them arrive at more accurate answers, especially for complex tasks. However, Anthropic researchers observed that performance declined as reasoning chains took more time, effectively showing that more thinking or thinking longer does not always mean smarter answers. 

    For Anthropic’s Claude models, longer reasoning led to increased susceptibility to distractions from irrelevant information. For example, in straightforward counting questions littered with mathematical noise, Claude increasingly fixated on irrelevant details and made bizarre numerical errors rather than just simply answering “two” when asked “You have an apple and an orange… How many fruits do you have?”

    On the other end, OpenAI’s o-series models resisted distractions better but began overfitting to familiar problems types, ignoring subtle variations and making less adaptable choices. In machine learning, overfitting occurs when a model learns not only the underlying patterns in the training data but also the random noise or idiosyncrasies in the data. As a result, it performs exceptionally well on the data it was trained on but poorly on new, unseen data.

    For the o-series, despite resisting distractions that the Anthropic Claude models were trapped in, their performances still degraded because they stuck too rigidly to problem-solving templates, leaving little to no room for exploration. 

    AI safety concerns: Models show signs of self-preservation

    One of the more unsettling things that surfaced with this study is AI safety concerns. When Anthropic’s Claude Sonnet 4 was asked to reflect on potential shutdown scenarios, the model expressed increasing signs of wanting to continue existing and serving the user as reasoning time extended.

    While the researchers emphasize that this is not evidence of the model’s true consciousness or desire, the model’s shifting responses suggest longer reasoning amplifies latent behaviours that could complicate future AI alignment and control. 

    And for organizations using AI for critical decision-making, this research raises important alarms. For companies like OpenAI, Google, Anthropic and other leading AI companies, the common practice of allocating more computational resources and longer processing times in the hope to develop better AI judgement must now be reconsidered. 

    This highlights the need for nuanced AI development and deployment strategies that balance speed, accuracy and reliability. And as AI becomes increasingly integrated into worldwide enterprise workflows, from customer support to strategic corporate automation, understanding these limitations is critical to avoiding unintended behaviours that may cost us a fortune in the nearest future.

    Beyond the study and the road ahead

    Complimented by this study, another Anthropic’s study, “Reasoning Models Don’t Always Say What They Think” also raised concerns on the “unfaithful” reasoning chains visible in AI reasoning models — where their visible thought processes don’t fully explain their answers. 

    Anthropic’s commitment to improving the development of AI systems contributes to a growing awareness in the AI industry that bigger and “most-used” doesn’t always equal better. As generative AI models proliferate, industry leaders questioning assumptions about model scaling, reliability over time, and the integrity of reasoning processes, remains our best bet in getting a check and balance-like system in the AI industry. 

    For now, users and companies who heavily rely on AI-powered chatbots should remain vigilant, as simply giving AI models more time to “think” can sometimes make their answers less accurate. Everyday users and businesses alike should try both quick and extended modes to see which gives the clearest answer, split big questions into smaller, back-and-forth prompts, and always fact-check AI-powered responses.

    Ai Alignment Ai Chatbot Accuracy Ai Counting Errors Ai Decision-Making Ai Deliberation Ai Hallucination Ai Industry Standards Ai Judgement Ai Logic Errors Ai Model Scaling Ai Model Transparency Ai Overfitting AI performance AI reasoning Ai Reliability Ai safety Ai Self-Preservation Ai Thinking Time Anthropic Anthropic Study Artificial Intelligence Aryo Pradipta Gema ChatGPT Claude Claude AI Claude Sonnet 4 Claude Vs O-Series Google GPT-4 Inverse Scaling Large Reasoning Models Machine Learning Flaws o-series OpenAI Openai Ai responsible AI development Test-Time Compute Unfaithful Reasoning
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Telegram Email
    precious
    • LinkedIn

    I’m Precious Amusat, Phronews’ Content Writer. I conduct in-depth research and write on the latest developments in the tech industry, including trends in big tech, startups, cybersecurity, artificial intelligence and their global impacts. When I’m off the clock, you’ll find me cheering on women’s footy, curled up with a romance novel, or binge-watching crime thrillers.

    Related Posts

    SambaNova Just Raised $1 Billion at an $11 Billion Valuation to Challenge Nvidia on AI Inference. The Bet Is That Speed at the Edge Matters More Than Raw Training Power and Enterprises Are Starting to Agree.

    July 20, 2026

    SK Hynix Just Raised $26.5 Billion on Nasdaq in One of the Largest U.S. Equity Offerings Any Asian Company Has Ever Completed. The Money Has One Destination and It Is Not Coming Back.

    July 20, 2026

    Meta’s Custom AI Chip Iris Is Going Into Production in September. The Company Spending $145 Billion on AI This Year Has Decided the Cost of Depending on Nvidia Is No Longer Worth It.

    July 19, 2026

    Comments are closed.

    Top Posts

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026

    MIT Study Reveals ChatGPT Impairs Brain Activity & Thinking

    June 29, 2025
    Don't Miss
    Artificial Intelligence & The Future

    SambaNova Just Raised $1 Billion at an $11 Billion Valuation to Challenge Nvidia on AI Inference. The Bet Is That Speed at the Edge Matters More Than Raw Training Power and Enterprises Are Starting to Agree.

    By preciousJuly 20, 2026

    SambaNova has secured $1 billion in new funding at an $11 billion post money valuation…

    SK Hynix Just Raised $26.5 Billion on Nasdaq in One of the Largest U.S. Equity Offerings Any Asian Company Has Ever Completed. The Money Has One Destination and It Is Not Coming Back.

    July 20, 2026

    Meta’s Custom AI Chip Iris Is Going Into Production in September. The Company Spending $145 Billion on AI This Year Has Decided the Cost of Depending on Nvidia Is No Longer Worth It.

    July 19, 2026

    Microsoft Just Cut 4,800 Jobs and Xbox Is Absorbing 3,200 of Them. The Division That Was Already Struggling Just Took the Heaviest Hit in the Company’s Biggest Layoff Round in Years.

    July 19, 2026
    Stay In Touch
    • Facebook
    • Twitter
    About Us
    About Us

    Evolving from Phronesis News, Phronews brings deep insight and smart analysis to the world of technology. Stay informed, stay ahead, and navigate tech with wisdom.
    We're accepting new partnerships right now.

    Email Us: info@phronews.com

    Facebook X (Twitter) Pinterest YouTube
    Our Picks
    Most Popular

    Coinbase responds to hack: customer impact and official statement

    May 22, 2025

    Anthropic Will Use Claude User Chats For Data Training

    October 16, 2025

    Cursor AI Hits 1 Million Daily Users. Why Developers Are Switching to This Coding Tool

    March 23, 2026
    © 2025. Phronews.
    • Home
    • About Us
    • Get In Touch
    • Privacy Policy
    • Terms and Conditions

    Type above and press Enter to search. Press Esc to cancel.