{"id":285830,"date":"2025-06-10T13:17:14","date_gmt":"2025-06-10T13:17:14","guid":{"rendered":"https:\/\/peraltafinancing.com\/analytics\/apple-finds-reasoning-flaws-in-ai-models\/"},"modified":"2025-06-10T13:17:14","modified_gmt":"2025-06-10T13:17:14","slug":"apple-finds-reasoning-flaws-in-ai-models","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=285830","title":{"rendered":"Apple Finds Reasoning Flaws in AI models"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div id=\"article-start\">\n<p><span style=\"font-weight: 400;\">A rather brutal truth has emerged in the AI industry, redefining what we consider the true capabilities of AI. A research paper titled \u201cThe Illusion of Thinking\u201d has sent ripples across the tech world, exposing reasoning flaws in\u00a0prominent AI <\/span><i><span style=\"font-weight: 400;\">\u2018so-called reasoning\u2019<\/span><\/i><span style=\"font-weight: 400;\"> models \u2013 <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/02\/claude-sonnet-3-7\/\">Claude 3.7 Sonnet<\/a> (thinking), DeepSeek-R1, and OpenAI\u2019s o3-mini (high). The research proves that these advanced models don\u2019t really reason the way we\u2019ve been led to believe. So what are they actually doing? Let\u2019s find out by diving into this research paper by Apple that exposes the reality of AI thinking models.<\/span><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-the-great-myth-of-ai-reasoning\"><span style=\"font-weight: 400;\">The Great Myth of AI Reasoning<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">For months, tech companies have been pitching their newer models as great \u2018reasoning\u2019 systems that follow the human method of step-by-step thinking to solve complex problems. These <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/03\/ai-reasoning-model\/\" target=\"_blank\" rel=\"noreferrer noopener\">large reasoning models<\/a> generate elaborate scenarios of \u201cthinking processes\u201d before the actual answer is given, showing the genuine cognitive work happening behind the scenes.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">But Apple\u2019s researchers have lifted the curtain on the technological drama, revealing the true capabilities of <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/04\/choosing-the-best-ai-chatbot-for-your-task\/\" target=\"_blank\" rel=\"noreferrer noopener\">AI chatbots<\/a>, which look rather dull. These models seem to be far more akin to pattern matchers that really cannot get through when faced with truly complex problems.<\/span><\/p>\n<figure class=\"wp-block-image size-full is-resized\"><img fetchpriority=\"high\" decoding=\"async\" width=\"646\" height=\"680\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image1-2.webp\" alt=\"The Illusion of Thinking: Apple Finds Reasoning Flaws in AI models\" class=\"wp-image-237327\" style=\"width:704px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image1-2.webp 646w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image1-2-285x300.webp 285w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image1-2-150x158.webp 150w\" sizes=\"(max-width: 646px) 100vw, 646px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source: <\/span><a href=\"https:\/\/machinelearning.apple.com\/research\/illusion-of-thinking\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400;\">Apple Research<\/span><\/a><\/figcaption><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-the-devastating-discovery\"><span style=\"font-weight: 400;\">The Devastating Discovery<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">The observations stated in \u2018The Illusion of Thinking\u2019 would bother anyone already placing a wager on the reasoning capabilities of current AI systems. Apple\u2019s research team,<span style=\"font-weight: 400;\"> led by scientists who carefully designed controllable puzzle environments, made three monumental discoveries<\/span>:<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-1-the-complexity-cliff\">1. The Complexity Cliff<\/h3>\n<p><span style=\"font-weight: 400;\">One of the major findings is that these supposedly advanced reasoning models suffer from what has been termed by the researchers as \u201ccomplete accuracy collapse\u201d, beyond certain complexity thresholds. Rather than a slow descent that may happen over time, this observation outright exposes the shallow nature of their so-called \u201creasoning\u201d.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Imagine a chess grandmaster who suddenly forgets how a piece moves, just because you added an extra row to the board. That\u2019s exactly how these models behaved during the research. The models that seemed extremely intelligent on problem sets they were acquainted with, suddenly became completely lost, the moment they were nudged even an inch out of their comfort zone.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-2-the-effort-paradox\">2. <span style=\"color: revert; font-size: revert;\">The Effort Paradox<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">What is more baffling is that Apple found these models have a scaling barrier against any logic. As the problems became more demanding, these models initially augmented their reasoning effort, showing longer thinking processes and more detail in each step.<\/span> <span style=\"font-weight: 400;\">However, there came a point when they simply stopped trying and started paying less attention to their tasks, despite having hefty computational resources.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">It is as if a student, when presented with increasingly difficult math problems, tries a bit hard at first but loses interest at one point and just starts to guess the answer randomly, despite having ample time to work on the problems.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-3-the-three-zones-of-performance\">3. The Three Zones of Performance<\/h3>\n<p>In the third finding, <span style=\"font-weight: 400;\">Apple identifies three zones of pure performance, indicating the true nature of these systems:<\/span><\/p>\n<ul class=\"wp-block-list\">\n<li><b>Low-complexity tasks: <\/b><span style=\"font-weight: 400;\">Standard AI models outperform their \u201creasoning\u201d counterparts in these tasks, suggesting extra reasoning steps may just be an expensive show.<\/span><\/li>\n<li><b>Medium-complexity tasks: <\/b><span style=\"font-weight: 400;\">This is found to be the sweet spot where reasoning models shine<\/span>.<\/li>\n<li><b>High-Complexity tasks: <\/b><span style=\"font-weight: 400;\">A spectacular failure from both standard and reasoning models was seen in these tasks, hinting at inherent limitations.<\/span><\/li>\n<\/ul>\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"680\" height=\"645\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image2-3.webp\" alt=\"The Illusion of Thinking: Apple Finds Reasoning Flaws in AI models\" class=\"wp-image-237337\" style=\"width:840px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image2-3.webp 680w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image2-3-300x285.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image2-3-150x142.webp 150w\" sizes=\"auto, (max-width: 680px) 100vw, 680px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source: <\/span><a href=\"https:\/\/machinelearning.apple.com\/research\/illusion-of-thinking\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400;\">Apple Research<\/span><\/a><\/figcaption><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-the-benchmark-problem-and-apple-s-solution\"><span style=\"font-weight: 400;\">The Benchmark Problem and Apple\u2019s Solution<\/span><\/h2>\n<p>\u2018The Illusion of Thinking\u2019 <span style=\"font-weight: 400;\">reveals a secret about <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2025\/03\/llm-evaluation-metrics\/\" target=\"_blank\" rel=\"noreferrer noopener\">AI evaluation<\/a> as well. Most benchmarks contain training data, causing the model to appear more capable than it actually is. These tests, therefore, evaluate models on memorized instances to a great extent. Apple, on the other hand, created a much more revealing evaluation process. The research team tested the models on the follwoing four logical puzzles with systematically rescalable complexity:<\/span><\/p>\n<ol class=\"wp-block-list\">\n<li><b>Tower of Hanoi: <\/b><span style=\"font-weight: 400;\">Moving disks by planning moves several steps ahead.<\/span><\/li>\n<li><b>Checker Jumping: <\/b><span style=\"font-weight: 400;\">Moving pieces strategically, based on spatial reasoning and sequential planning.<\/span><\/li>\n<li><b>River crossing: <\/b><span style=\"font-weight: 400;\">A logic puzzle about getting multiple entities across a river with constraints.<\/span><\/li>\n<li><b>Block Stacking: <\/b><span style=\"font-weight: 400;\">A 3D reasoning task requiring knowledge of physical relationships.<\/span><\/li>\n<\/ol>\n<p><span style=\"font-weight: 400;\">The selection of these tasks or problems was by no means random. Each problem could be scaled precisely from trivial to mind-boggling, so that researchers can know at which level the AI reasoning gives out.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-watching-ai-think-the-actual-truth\"><span style=\"font-weight: 400;\">Watching AI \u201cThink\u201d: The Actual Truth<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">Unlike most traditional benchmarks, these puzzles did not limit the researchers to look at just the final answers. They actually revealed the entire chain of reasoning of the models to be evaluated. Researchers could watch the models solve problems step-by-step, seeing if the machines were going through logical principles or were just pattern-matching from some memory.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">The results were eye-opening. Models that appeared to be actually \u201creasoning\u201d through a problem beautifully would suddenly go illogical, abandon systematic approaches, or simply give up when complexity increased, though moments earlier, they had perfectly demonstrated the required skills.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">By making new, controllable puzzle environments, Apple circumvented the contamination problem and exposed the full scale of model limitations. The outcome was sobering. For real, new, and fresh challenges that could not be memorized, even the most advanced reasoning models were struggling in ways that highlight the real limits posed upon them.<\/span><\/p>\n<h3 class=\"wp-block-heading\" id=\"h-results-and-analysis\"><span style=\"font-weight: 400;\">Results and Analysis<\/span><\/h3>\n<p><span style=\"font-weight: 400;\">Across all four types of puzzles, Apple\u2019s researchers documented consistent failure modes that provide a grim picture of today\u2019s AI capabilities.<\/span><\/p>\n<ul class=\"wp-block-list\">\n<li><b>Accuracy Issue:<\/b><span style=\"font-weight: 400;\"> On these puzzle sets, a model that reached almost perfect performance on the simplistic versions encountered an astonishing drop in accuracy.  Sometimes, it would fall from almost 90% success to an almost total failure with only a few additional complex steps added. This was never a gradual degradation, but a sudden and catastrophic failure.<\/span><\/li>\n<li><b>Inconsistent logic application: <\/b><span style=\"font-weight: 400;\">The models sometimes failed to apply algorithms consistently when demonstrating knowledge of the very correct approaches. For example, a model may apply a systematic strategy successfully for one Tower of Hanoi puzzle, but then abandon that very strategy on a very similar but slightly more complex instance.<\/span><\/li>\n<li><b>Role of Effort Paradox: <\/b><span style=\"font-weight: 400;\">The researchers, in correlation with problem difficulty, studied the amount of \u2018thinking\u201d the model did. This ranged from length to granularity levels of reasoning traces. Initially, the thinking effort increased with complexity. However, as the problems became tougher to solve, the model would quite abnormally start relaxing its effort, even with an unlimited computational resource provided.<\/span><\/li>\n<li><b>Computational Shortcuts: <\/b><span style=\"font-weight: 400;\">It was also found that the model tended to take computational shortcuts that worked really well for simple problems, but would lead to catastrophic failures in harder cases. Rather than recognizing such a pattern and trying to compensate, the model would either keep on trying with bad strategies or just give up.<\/span><\/li>\n<\/ul>\n<p><span style=\"font-weight: 400;\">These findings establish that, in essence, current AI reasoning is more brittle and limited than the public demonstrations have led us to believe. The models are yet to learn reasoning; for now, they only recognize reasoning and replicate it if they have seen it somewhere else.<\/span><\/p>\n<figure class=\"wp-block-image size-full is-resized\"><img loading=\"lazy\" decoding=\"async\" width=\"680\" height=\"558\" src=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image3-2.webp\" alt=\"The Illusion of Thinking: Apple Finds Reasoning Flaws in AI models\" class=\"wp-image-237347\" style=\"width:840px;height:auto\" srcset=\"https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image3-2.webp 680w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image3-2-300x246.webp 300w, https:\/\/cdn.analyticsvidhya.com\/wp-content\/uploads\/2025\/06\/image3-2-150x123.webp 150w\" sizes=\"auto, (max-width: 680px) 100vw, 680px\"\/><figcaption class=\"wp-element-caption\"><span style=\"font-weight: 400;\">Source: <\/span><a href=\"https:\/\/machinelearning.apple.com\/research\/illusion-of-thinking\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400;\">Apple Research<\/span><\/a><\/figcaption><\/figure>\n<h2 class=\"wp-block-heading\" id=\"h-why-does-this-matter-for-the-future-of-ai\"><span style=\"font-weight: 400;\">Why Does This Matter for the Future of AI?<\/span><\/h2>\n<p>\u2018<span style=\"font-weight: 400;\">Th<\/span>e Illusion of Thinking\u2019<span style=\"font-weight: 400;\">, far from being academically nitpicking, touches very deeply upon the implications of AI.<\/span> W<span style=\"font-weight: 400;\">e can see it affects the entire AI industry and anyone who may make a decision using AI capabilities.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Apple\u2019s findings indicate that so-called \u2018reasoning\u2019 is indeed just a very sophisticated kind of memorization and pattern matching. The models excel in recognizing problem patterns they have seen before and then associate the solution they have previously learned. However, they tend to fail when asked to really logically reason through a problem that is somehow new to them.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">For the past few months, the AI community has been awestruck with the advancements in reasoning models, as shown by their parent companies. Industry leaders have even gone on to promise us that <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/2023\/04\/artificial-general-intelligence\/\" target=\"_blank\" rel=\"noreferrer noopener\">Artificial General Intelligence<\/a> (AGI) is right around the corner. \u2018The Illusion of Thinking\u2019 tells us that this assessment is absurdly optimistic. If present \u2018reasoning\u2019 models are not able to handle complexities above the current benchmarks, and if they are indeed just dressed-up pattern-matching systems, then the pathway toward true AGI might be longer and tougher than Silicon Valley\u2019s most optimistic proposals.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">Despite sobering observations, Apple\u2019s study does not remain entirely pessimistic. The performance of AI models in the medium-complexity regime shows the actual progress in their reasoning capabilities. In this category, these systems can execute really complicated tasks, which were deemed impossible some four or so years ago.<\/span><\/p>\n<h2 class=\"wp-block-heading\" id=\"h-conclusion\"><span style=\"font-weight: 400;\">Conclusion<\/span><\/h2>\n<p><span style=\"font-weight: 400;\">Apple\u2019s research marks a turning point from breathless hype to precise scientific measurements of what AI systems can do. This is where the AI Industry faces its next choice. Will it continue to chase benchmark scores and marketing claims, or focus on building systems that can really do some level of reasoning? The companies that will do the latter might end up building the AI systems we really need.<\/span><\/p>\n<p><span style=\"font-weight: 400;\">It is clear, however, that future paths to AGI will require more than just scaled-up pattern-matchers. They will need fundamentally new approaches to reasoning, understanding, and genuine intelligence. Illusions of thinking can be convincing, but as Apple has shown, that\u2019s all they are: illusions. The real task of engineering truly intelligent systems is just beginning.<\/span><\/p>\n<div class=\"border-top py-3 author-info my-4\">\n<div class=\"author-card d-flex align-items-center\">\n<div class=\"flex-shrink-0 overflow-hidden\">\n                                    <a href=\"https:\/\/www.analyticsvidhya.com\/blog\/author\/riyab20021618492\/\" class=\"text-decoration-none active-avatar\"><br \/>\n                                                                       <img decoding=\"async\" src=\"https:\/\/av-eks-lekhak.s3.amazonaws.com\/media\/lekhak-profile-images\/converted_image_5X1DGT2.webp\" width=\"48\" height=\"48\" alt=\"Riya Bansal.\" loading=\"lazy\" class=\"rounded-circle\"\/><\/p>\n<p>                                <\/a>\n                                <\/div>\n<\/p><\/div>\n<p>Gen AI Intern at Analytics Vidhya\u00a0<br \/>Department of Computer Science, Vellore Institute of Technology, Vellore, India\u00a0<\/p>\n<p>I am currently working as a Gen AI Intern at Analytics Vidhya, where I contribute to innovative AI-driven solutions that empower businesses to leverage data effectively. As a final-year Computer Science student at Vellore Institute of Technology, I bring a solid foundation in software development, data analytics, and machine learning to my role.\u00a0<\/p>\n<p>Feel free to connect with me at <a href=\"https:\/\/www.analyticsvidhya.com\/cdn-cgi\/l\/email-protection\" class=\"__cf_email__\" data-cfemail=\"f082998991de92919e83919cb0919e919c8984999383869994988991de939f9d\">[email\u00a0protected]<\/a>\u00a0<\/p>\n<\/p><\/div>\n<\/p><\/div>\n<p><h4 class=\"fs-24 text-dark\">Login to continue reading and enjoy expert-curated content.<\/h4>\n<p>                        <button class=\"btn btn-primary mx-auto d-table\" data-bs-toggle=\"modal\" data-bs-target=\"#loginModal\" id=\"readMoreBtn\">Keep Reading for Free<\/button>\n                    <\/p>\n\n","protected":false},"excerpt":{"rendered":"<p>A rather brutal truth has emerged in the AI industry, redefining what we consider the true capabilities of AI. A research paper titled \u201cThe Illusion of Thinking\u201d has sent ripples across the tech world, exposing reasoning flaws in\u00a0prominent AI \u2018so-called reasoning\u2019 models \u2013 Claude 3.7 Sonnet (thinking), DeepSeek-R1, and OpenAI\u2019s o3-mini (high). The research proves [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":285831,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[12033],"tags":[1332,16649,75846,8558,20867],"dealstore":[],"offerexpiration":[],"class_list":["post-285830","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-analytics","tag-apple","tag-finds","tag-flaws","tag-models","tag-reasoning"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>Apple Finds Reasoning Flaws in AI models - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=285830\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Apple Finds Reasoning Flaws in AI models - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"A rather brutal truth has emerged in the AI industry, redefining what we consider the true capabilities of AI. A research paper titled \u201cThe Illusion of Thinking\u201d has sent ripples across the tech world, exposing reasoning flaws in\u00a0prominent AI \u2018so-called reasoning\u2019 models \u2013 Claude 3.7 Sonnet (thinking), DeepSeek-R1, and OpenAI\u2019s o3-mini (high). The research proves [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=285830\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-06-10T13:17:14+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp\" \/>\n\t<meta property=\"og:image:width\" content=\"872\" \/>\n\t<meta property=\"og:image:height\" content=\"473\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/webp\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"8 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=285830#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=285830\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"Apple Finds Reasoning Flaws in AI models\",\"datePublished\":\"2025-06-10T13:17:14+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=285830\"},\"wordCount\":1628,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=285830#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp\",\"keywords\":[\"Apple\",\"Finds\",\"flaws\",\"Models\",\"Reasoning\"],\"articleSection\":[\"Analytics\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=285830#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=285830\",\"url\":\"https:\/\/fivemor.com\/?p=285830\",\"name\":\"Apple Finds Reasoning Flaws in AI models - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=285830#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=285830#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp\",\"datePublished\":\"2025-06-10T13:17:14+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=285830#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=285830\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=285830#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp\",\"width\":872,\"height\":473},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=285830#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Apple Finds Reasoning Flaws in AI models\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Apple Finds Reasoning Flaws in AI models - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=285830","og_locale":"en_US","og_type":"article","og_title":"Apple Finds Reasoning Flaws in AI models - Som2ny Network","og_description":"A rather brutal truth has emerged in the AI industry, redefining what we consider the true capabilities of AI. A research paper titled \u201cThe Illusion of Thinking\u201d has sent ripples across the tech world, exposing reasoning flaws in\u00a0prominent AI \u2018so-called reasoning\u2019 models \u2013 Claude 3.7 Sonnet (thinking), DeepSeek-R1, and OpenAI\u2019s o3-mini (high). The research proves [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=285830","og_site_name":"Som2ny Network","article_published_time":"2025-06-10T13:17:14+00:00","og_image":[{"width":872,"height":473,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp","type":"image\/webp"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"8 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=285830#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=285830"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"Apple Finds Reasoning Flaws in AI models","datePublished":"2025-06-10T13:17:14+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=285830"},"wordCount":1628,"commentCount":0,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=285830#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp","keywords":["Apple","Finds","flaws","Models","Reasoning"],"articleSection":["Analytics"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=285830#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=285830","url":"https:\/\/fivemor.com\/?p=285830","name":"Apple Finds Reasoning Flaws in AI models - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=285830#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=285830#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp","datePublished":"2025-06-10T13:17:14+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=285830#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=285830"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=285830#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/06\/Apple-Exposes-Reasoning-Flaws-in-o3-Claude-and-DeepSeek-R1.webp.webp","width":872,"height":473},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=285830#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"Apple Finds Reasoning Flaws in AI models"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/285830","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=285830"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/285830\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/285831"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=285830"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=285830"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=285830"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=285830"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=285830"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}