{"id":156820,"date":"2025-03-26T01:34:02","date_gmt":"2025-03-26T01:34:02","guid":{"rendered":"https:\/\/peraltafinancing.com\/apple-2\/deepseek-v3-now-runs-at-20-tokens-per-second-on-mac-studio-and-thats-a-nightmare-for-openai\/"},"modified":"2025-03-26T01:34:02","modified_gmt":"2025-03-26T01:34:02","slug":"deepseek-v3-now-runs-at-20-tokens-per-second-on-mac-studio-and-thats-a-nightmare-for-openai","status":"publish","type":"post","link":"https:\/\/fivemor.com\/?p=156820","title":{"rendered":"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that&#8217;s a nightmare for OpenAI"},"content":{"rendered":" \r\n<br><div>\n\t\t\t\t<div id=\"boilerplate_2682874\" class=\"post-boilerplate boilerplate-before\">\n<p class=\"wp-block-paragraph\"><em>Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. <a href=\"https:\/\/venturebeat.com\/newsletters\/?utm_source=VBsite&amp;utm_medium=desktopNav\" data-type=\"link\" data-id=\"https:\/\/venturebeat.com\/newsletters\/?utm_source=VBsite&amp;utm_medium=desktopNav\">Learn More<\/a><\/em><\/p>\n\n\n\n<hr class=\"wp-block-separator has-css-opacity is-style-wide\"\/>\n<\/div><p>Chinese AI startup <a href=\"https:\/\/www.deepseek.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">DeepSeek<\/a> has quietly released a new large language model that\u2019s already sending ripples through the artificial intelligence industry \u2014 not just for its capabilities, but for how it\u2019s being deployed. The 641-gigabyte model, dubbed <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\" target=\"_blank\" rel=\"noreferrer noopener\">DeepSeek-V3-0324<\/a>, appeared on AI repository <a href=\"https:\/\/huggingface.co\/deepseek-ai\/\" target=\"_blank\" rel=\"noreferrer noopener\">Hugging Face<\/a> today with virtually no announcement, continuing the company\u2019s pattern of low-key but impactful releases.<\/p>\n\n\n\n<p>What makes this launch particularly notable is the model\u2019s <a href=\"https:\/\/opensource.org\/license\/mit\">MIT license<\/a> \u2014 making it freely available for commercial use \u2014 and early reports that it can run directly on consumer-grade hardware, specifically Apple\u2019s <a href=\"https:\/\/www.apple.com\/mac-studio\" target=\"_blank\" rel=\"noreferrer noopener\">Mac Studio<\/a> with <a href=\"https:\/\/www.apple.com\/newsroom\/2025\/03\/apple-reveals-m3-ultra-taking-apple-silicon-to-a-new-extreme\/\">M3 Ultra chip<\/a>.<\/p>\n\n\n\n<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">The new Deep Seek V3 0324 in 4-bit runs at &gt; 20 toks\/sec on a 512GB M3 Ultra with mlx-lm! <a href=\"https:\/\/t.co\/wFVrFCxGS6\">pic.twitter.com\/wFVrFCxGS6<\/a><\/p>\u2014 Awni Hannun (@awnihannun) <a href=\"https:\/\/twitter.com\/awnihannun\/status\/1904177084609827054?ref_src=twsrc%5Etfw\">March 24, 2025<\/a><\/blockquote> \n\n\n\n<p>\u201cThe new DeepSeek-V3-0324 in 4-bit runs at &gt; 20 tokens\/second on a 512GB M3 Ultra with mlx-lm!\u201d wrote AI researcher <a href=\"https:\/\/x.com\/awnihannun\/status\/1904177084609827054\">Awni Hannun<\/a> on social media. While the $9,499 Mac Studio might stretch the definition of \u201cconsumer hardware,\u201d the ability to run such a massive model locally is a major departure from the data center requirements typically associated with state-of-the-art AI.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-deepseek-s-stealth-launch-strategy-disrupts-ai-market-expectations\">DeepSeek\u2019s stealth launch strategy disrupts AI market expectations<\/h2>\n\n\n\n<p>The 685-billion-parameter model arrived with no accompanying whitepaper, blog post, or marketing push \u2014 just an empty <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\/tree\/main\">README file<\/a> and the model weights themselves. This approach contrasts sharply with the carefully orchestrated product launches typical of Western AI companies, where months of hype often precede actual releases.<\/p>\n\n\n\n<p>Early testers report significant improvements over the previous version. AI researcher <a href=\"https:\/\/x.com\/TheXeophon\/status\/1904225899957936314\">Xeophon<\/a> proclaimed in a post on X.com: \u201cTested the new DeepSeek V3 on my internal bench and it has a huge jump in all metrics on all tests. It is now the best non-reasoning model, dethroning Sonnet 3.5.\u201d<\/p>\n\n\n\n<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">Tested the new DeepSeek V3 on my internal bench and it has a huge jump in all metrics on all tests.<br\/>It is now the best non-reasoning model, dethroning Sonnet 3.5.<\/p><p>Congrats <a href=\"https:\/\/twitter.com\/deepseek_ai?ref_src=twsrc%5Etfw\">@deepseek_ai<\/a>! <a href=\"https:\/\/t.co\/efEu2FQSBe\">pic.twitter.com\/efEu2FQSBe<\/a><\/p>\u2014 Xeophon (@TheXeophon) <a href=\"https:\/\/twitter.com\/TheXeophon\/status\/1904225899957936314?ref_src=twsrc%5Etfw\">March 24, 2025<\/a><\/blockquote> \n\n\n\n<p>This claim, if validated by broader testing, would position DeepSeek\u2019s new model above <a href=\"https:\/\/www.anthropic.com\/news\/claude-3-5-sonnet\">Claude Sonnet 3.5<\/a> from Anthropic, one of the most respected commercial AI systems. And unlike Sonnet, which requires a subscription, <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\/tree\/main\">DeepSeek-V3-0324<\/a>\u2018s weights are freely available for anyone to download and use.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-how-deepseek-v3-0324-s-breakthrough-architecture-achieves-unmatched-efficiency\">How DeepSeek V3-0324\u2019s breakthrough architecture achieves unmatched efficiency<\/h2>\n\n\n\n<p><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\">DeepSeek-V3-0324<\/a> employs a <a href=\"https:\/\/huggingface.co\/blog\/moe\">mixture-of-experts<\/a> (MoE) architecture that fundamentally reimagines how large language models operate. Traditional models activate their entire parameter count for every task, but DeepSeek\u2019s approach activates only about 37 billion of its 685 billion parameters during specific tasks.<\/p>\n\n\n\n<p>This selective activation represents a paradigm shift in model efficiency. By activating only the most relevant \u201cexpert\u201d parameters for each specific task, DeepSeek achieves performance comparable to much larger fully-activated models while drastically reducing computational demands.<\/p>\n\n\n\n<p>The model incorporates two additional breakthrough technologies: <a href=\"https:\/\/arxiv.org\/abs\/2405.04434\">Multi-Head Latent Attention <\/a>(MLA) and <a href=\"https:\/\/arxiv.org\/abs\/2404.19737\">Multi-Token Prediction<\/a> (MTP). MLA enhances the model\u2019s ability to maintain context across long passages of text, while MTP generates multiple tokens per step instead of the usual one-at-a-time approach. Together, these innovations boost output speed by nearly 80%.<\/p>\n\n\n\n<p><a href=\"https:\/\/simonwillison.net\/2025\/Mar\/24\/deepseek\/\">Simon Willison<\/a>, a developer tools creator, noted in a blog post that a 4-bit quantized version reduces the storage footprint to 352GB, making it feasible to run on high-end consumer hardware like the <a href=\"https:\/\/www.apple.com\/mac-studio\/\">Mac Studio<\/a> with <a href=\"https:\/\/www.apple.com\/newsroom\/2025\/03\/apple-reveals-m3-ultra-taking-apple-silicon-to-a-new-extreme\/\">M3 Ultra chip<\/a>.<\/p>\n\n\n\n<p>This represents a potentially significant shift in AI deployment. While traditional AI infrastructure typically relies on multiple <a href=\"https:\/\/venturebeat.com\/ai\/nvidias-gtc-2025-keynote-40x-ai-performance-leap-open-source-dynamo-and-a-walking-star-wars-inspired-blue-robot\/\">Nvidia GPUs<\/a> consuming several kilowatts of power, the Mac Studio draws less than 200 watts during inference. This efficiency gap suggests the AI industry may need to rethink assumptions about infrastructure requirements for top-tier model performance.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-china-s-open-source-ai-revolution-challenges-silicon-valley-s-closed-garden-model\">China\u2019s open source AI revolution challenges Silicon Valley\u2019s closed garden model<\/h2>\n\n\n\n<p>DeepSeek\u2019s release strategy exemplifies a fundamental divergence in AI business philosophy between Chinese and Western companies. While U.S. leaders like <a href=\"https:\/\/openai.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">OpenAI<\/a> and <a href=\"https:\/\/www.anthropic.com\/\" target=\"_blank\" rel=\"noreferrer noopener\">Anthropic<\/a> keep their models behind paywalls, Chinese AI companies increasingly embrace permissive open-source licensing.<\/p>\n\n\n\n<p>This approach is rapidly transforming China\u2019s AI ecosystem. The open availability of cutting-edge models creates a multiplier effect, enabling startups, researchers, and developers to build upon sophisticated AI technology without massive capital expenditure. This has accelerated China\u2019s AI capabilities at a pace that has shocked Western observers.<\/p>\n\n\n\n<p>The business logic behind this strategy reflects market realities in China. With multiple well-funded competitors, maintaining a proprietary approach becomes increasingly difficult when competitors offer similar capabilities for free. Open-sourcing creates alternative value pathways through ecosystem leadership, API services, and enterprise solutions built atop freely available foundation models.<\/p>\n\n\n\n<p>Even established Chinese tech giants have recognized this shift. Baidu announced plans to make its <a href=\"https:\/\/www.reuters.com\/technology\/artificial-intelligence\/baidu-make-ernie-ai-model-open-source-end-june-2025-02-14\/\">Ernie 4.5 model series<\/a> open-source by June, while <a href=\"https:\/\/www.cnbc.com\/2024\/09\/19\/alibaba-launches-over-100-new-ai-models-releases-text-to-video-generation.html\">Alibaba<\/a> and <a href=\"https:\/\/www.cnbc.com\/2025\/03\/24\/china-open-source-deepseek-ai-spurs-innovation-and-adoption.html\">Tencent<\/a> have released open-source AI models with specialized capabilities. This movement stands in stark contrast to the API-centric strategy employed by Western leaders.<\/p>\n\n\n\n<p>The open-source approach also addresses unique challenges faced by Chinese AI companies. With restrictions on access to cutting-edge <a href=\"https:\/\/venturebeat.com\/ai\/nvidias-gtc-2025-keynote-40x-ai-performance-leap-open-source-dynamo-and-a-walking-star-wars-inspired-blue-robot\/\">Nvidia chips<\/a>, Chinese firms have emphasized efficiency and optimization to achieve competitive performance with more limited computational resources. This necessity-driven innovation has now become a potential competitive advantage.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-deepseek-v3-0324-the-foundation-for-an-ai-reasoning-revolution\">DeepSeek V3-0324: The foundation for an AI reasoning revolution<\/h2>\n\n\n\n<p>The timing and characteristics of <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\">DeepSeek-V3-0324<\/a> strongly suggest it will serve as the foundation for <a href=\"https:\/\/x.com\/kimmonismus\/status\/1903478526877307013\">DeepSeek-R2<\/a>, an improved reasoning-focused model expected within the next two months. This follows DeepSeek\u2019s established pattern, where its base models precede specialized reasoning models by several weeks.<\/p>\n\n\n\n<p>\u201cThis lines up with how they released V3 around Christmas followed by R1 a few weeks later. R2 is rumored for April so this could be it,\u201d noted Reddit user <a href=\"https:\/\/www.reddit.com\/r\/LocalLLaMA\/comments\/1jip611\/comment\/mjgxu3f\/?utm_source=share&amp;utm_medium=web3x&amp;utm_name=web3xcss&amp;utm_term=1&amp;utm_content=share_button\" target=\"_blank\" rel=\"noreferrer noopener\">mxforest<\/a>.<\/p>\n\n\n\n<p>The implications of an advanced open-source reasoning model cannot be overstated. Current reasoning models like <a href=\"https:\/\/openai.com\/o1\/\" target=\"_blank\" rel=\"noreferrer noopener\">OpenAI\u2019s o1<\/a> and <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-R1\">DeepSeek\u2019s R1<\/a> represent the cutting edge of AI capabilities, demonstrating unprecedented problem-solving abilities in domains from mathematics to coding. Making this technology freely available would democratize access to AI systems currently limited to those with substantial budgets.<\/p>\n\n\n\n<p>The potential R2 model arrives amid significant revelations about reasoning models\u2019 computational demands. Nvidia CEO Jensen Huang recently noted that DeepSeek\u2019s R1 model \u201c<a href=\"https:\/\/www.cnbc.com\/2025\/03\/19\/nvidia-ceo-jensen-huang-why-deepseek-model-needs-100-times-more-computing.html\">consumes 100 times more compute than a non-reasoning AI<\/a>,\u201d contradicting earlier industry assumptions about efficiency. This reveals the remarkable achievement behind DeepSeek\u2019s models, which deliver competitive performance while operating under greater resource constraints than their Western counterparts.<\/p>\n\n\n\n<p>If DeepSeek-R2 follows the trajectory set by R1, it could present a direct challenge to <a href=\"https:\/\/x.com\/sama\/status\/1889755723078443244?lang=en\">GPT-5<\/a>, OpenAI\u2019s next flagship model rumored for release in coming months. The contrast between OpenAI\u2019s closed, heavily-funded approach and DeepSeek\u2019s open, resource-efficient strategy represents two competing visions for AI\u2019s future.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-how-to-experience-deepseek-v3-0324-a-complete-guide-for-developers-and-users\">How to experience DeepSeek V3-0324: A complete guide for developers and users<\/h2>\n\n\n\n<p>For those eager to experiment with <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\" target=\"_blank\" rel=\"noreferrer noopener\">DeepSeek-V3-0324<\/a>, several pathways exist depending on technical needs and resources. The complete model weights are available from <a href=\"https:\/\/huggingface.co\/deepseek-ai\/\" target=\"_blank\" rel=\"noreferrer noopener\">Hugging Face<\/a>, though the 641GB size makes direct download practical only for those with substantial storage and computational resources.<\/p>\n\n\n\n<p>For most users, cloud-based options offer the most accessible entry point. <a href=\"https:\/\/openrouter.ai\/\">OpenRouter<\/a> provides free API access to the model, with a user-friendly chat interface. Simply select DeepSeek V3 0324 as the model to begin experimenting.<\/p>\n\n\n\n<p>DeepSeek\u2019s own chat interface at <a href=\"https:\/\/www.deepseek.com\/\">chat.deepseek.com<\/a> has likely been updated to the new version as well, though the company hasn\u2019t explicitly confirmed this. Early users report the model is accessible through this platform with improved performance over previous versions.<\/p>\n\n\n\n<p>Developers looking to integrate the model into applications can access it through various inference providers. <a href=\"https:\/\/www.hyperbolic.xyz\/\">Hyperbolic Labs<\/a> announced immediate availability as \u201cthe first inference provider serving this model on Hugging Face,\u201d while OpenRouter offers API access compatible with the <a href=\"https:\/\/platform.openai.com\/docs\/overview\" target=\"_blank\" rel=\"noreferrer noopener\">OpenAI SDK<\/a>.<\/p>\n\n\n\n<blockquote class=\"twitter-tweet\"><p lang=\"en\" dir=\"ltr\">DeepSeek-V3-0324 Now Live on Hyperbolic ?<\/p><p>At Hyperbolic, we\u2019re committed to delivering the latest open-source models as soon as they\u2019re available. This is our promise to the developer community.<\/p><p>Start inferencing today. <a href=\"https:\/\/t.co\/495xf6kofa\">pic.twitter.com\/495xf6kofa<\/a><\/p>\u2014 Hyperbolic (@hyperbolic_labs) <a href=\"https:\/\/twitter.com\/hyperbolic_labs\/status\/1904230677861855617?ref_src=twsrc%5Etfw\">March 24, 2025<\/a><\/blockquote> \n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-deepseek-s-new-model-prioritizes-technical-precision-over-conversational-warmth\">DeepSeek\u2019s new model prioritizes technical precision over conversational warmth<\/h2>\n\n\n\n<p>Early users have reported a noticeable shift in the model\u2019s communication style. While previous DeepSeek models were praised for their conversational, human-like tone, \u201c<a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\">V3-0324<\/a>\u201d presents a more formal, technically-oriented persona.<\/p>\n\n\n\n<p>\u201cIs it only me or does this version feel less human like?\u201d asked Reddit user <a href=\"https:\/\/www.reddit.com\/r\/LocalLLaMA\/comments\/1jip611\/comment\/mjgzja4\/?utm_source=share&amp;utm_medium=web3x&amp;utm_name=web3xcss&amp;utm_term=1&amp;utm_content=share_button\" target=\"_blank\" rel=\"noreferrer noopener\">nother_level<\/a>. \u201cFor me the thing that set apart deepseek v3 from others were the fact that it felt more like human. Like the tone the words and such it was not robotic sounding like other llm\u2019s but now with this version its like other llms sounding robotic af.\u201d<\/p>\n\n\n\n<p>Another user, <a href=\"https:\/\/www.reddit.com\/r\/LocalLLaMA\/comments\/1jip611\/comment\/mjh2q4e\/?utm_source=share&amp;utm_medium=web3x&amp;utm_name=web3xcss&amp;utm_term=1&amp;utm_content=share_button\" target=\"_blank\" rel=\"noreferrer noopener\">AppearanceHeavy6724<\/a>, added: \u201cYeah, it lost its aloof charm for sure, it feels too intellectual for its own good.\u201d<\/p>\n\n\n\n<p>This personality shift likely reflects deliberate design choices by DeepSeek\u2019s engineers. The move toward a more precise, analytical communication style suggests a strategic repositioning of the model for professional and technical applications rather than casual conversation. This aligns with broader industry trends, as AI developers increasingly recognize that different use cases benefit from different interaction styles.<\/p>\n\n\n\n<p>For developers building specialized applications, this more precise communication style may actually represent an advantage, providing clearer and more consistent outputs for integration into professional workflows. However, it may limit the model\u2019s appeal for customer-facing applications where warmth and approachability are valued.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"h-how-deepseek-s-open-source-strategy-is-redrawing-the-global-ai-landscape\">How DeepSeek\u2019s open source strategy is redrawing the global AI landscape<\/h2>\n\n\n\n<p>DeepSeek\u2019s approach to AI development and distribution represents more than a technical achievement \u2014 it embodies a fundamentally different vision for how advanced technology should propagate through society. By making cutting-edge AI freely available under permissive licensing, DeepSeek enables exponential innovation that closed models inherently constrain.<\/p>\n\n\n\n<p>This philosophy is rapidly closing the perceived AI gap between China and the United States. Just months ago, most analysts estimated China lagged 1-2 years behind U.S. AI capabilities. Today, that gap has narrowed dramatically to perhaps 3-6 months, with some areas approaching parity or even Chinese leadership.<\/p>\n\n\n\n<p>The parallels to Android\u2019s impact on the mobile ecosystem are striking. Google\u2019s decision to make Android freely available created a platform that ultimately achieved dominant global market share. Similarly, open-source AI models may outcompete closed systems through sheer ubiquity and the collective innovation of thousands of contributors.<\/p>\n\n\n\n<p>The implications extend beyond market competition to fundamental questions about technology access. Western AI leaders increasingly face criticism for concentrating advanced capabilities among well-resourced corporations and individuals. DeepSeek\u2019s approach distributes these capabilities more broadly, potentially accelerating global AI adoption.<\/p>\n\n\n\n<p>As <a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V3-0324\">DeepSeek-V3-0324<\/a> finds its way into research labs and developer workstations worldwide, the competition is no longer simply about building the most powerful AI, but about enabling the most people to build with AI. In that race, DeepSeek\u2019s quiet release speaks volumes about the future of artificial intelligence. The company that shares its technology most freely may ultimately wield the greatest influence over how AI reshapes our world.<\/p>\n<div id=\"boilerplate_2660155\" class=\"post-boilerplate boilerplate-after\"><div class=\"Boilerplate__newsletter-container vb\">\n<div class=\"Boilerplate__newsletter-main\">\n<p><strong>Daily insights on business use cases with VB Daily<\/strong><\/p>\n<p class=\"copy\">If you want to impress your boss, VB Daily has you covered. We give you the inside scoop on what companies are doing with generative AI, from regulatory shifts to practical deployments, so you can share insights for maximum ROI.<\/p>\n<p class=\"Form__newsletter-legal\">Read our <a href=\"https:\/\/venturebeat.com\/terms-of-service\/\">Privacy Policy<\/a><\/p>\n<p class=\"Form__success\" id=\"boilerplateNewsletterConfirmation\">\n\t\t\t\t\tThanks for subscribing. Check out more <a href=\"https:\/\/venturebeat.com\/newsletters\/\">VB newsletters here<\/a>.\n\t\t\t\t<\/p>\n<p class=\"Form__error\">An error occured.<\/p>\n<\/p><\/div>\n<div class=\"image-container\">\n\t\t\t\t\t<img decoding=\"async\" src=\"https:\/\/venturebeat.com\/wp-content\/themes\/vb-news\/brand\/img\/vb-daily-phone.png\" alt=\"\"\/>\n\t\t\t\t<\/div>\n<\/p><\/div>\n<\/div>\t\t\t<\/div><script async src=\"\/\/platform.twitter.com\/widgets.js\" charset=\"utf-8\"><\/script>\r\n<br>\r\n","protected":false},"excerpt":{"rendered":"<p>Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Chinese AI startup DeepSeek has quietly released a new large language model that\u2019s already sending ripples through the artificial intelligence industry \u2014 not just for its capabilities, but for how it\u2019s being deployed. The 641-gigabyte model, [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":156821,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[11768],"tags":[36341,11553,8282,11438,20943,4553,13003],"dealstore":[],"offerexpiration":[],"class_list":["post-156820","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-apple-2","tag-deepseekv3","tag-mac","tag-nightmare","tag-openai","tag-runs","tag-studio","tag-tokens"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v26.4 - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that&#039;s a nightmare for OpenAI - Som2ny Network<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/fivemor.com\/?p=156820\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that&#039;s a nightmare for OpenAI - Som2ny Network\" \/>\n<meta property=\"og:description\" content=\"Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Chinese AI startup DeepSeek has quietly released a new large language model that\u2019s already sending ripples through the artificial intelligence industry \u2014 not just for its capabilities, but for how it\u2019s being deployed. The 641-gigabyte model, [&hellip;]\" \/>\n<meta property=\"og:url\" content=\"https:\/\/fivemor.com\/?p=156820\" \/>\n<meta property=\"og:site_name\" content=\"Som2ny Network\" \/>\n<meta property=\"article:published_time\" content=\"2025-03-26T01:34:02+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1024\" \/>\n\t<meta property=\"og:image:height\" content=\"573\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"admin\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"admin\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"9 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/fivemor.com\/?p=156820#article\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/?p=156820\"},\"author\":{\"name\":\"admin\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\"},\"headline\":\"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that&#8217;s a nightmare for OpenAI\",\"datePublished\":\"2025-03-26T01:34:02+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=156820\"},\"wordCount\":1911,\"commentCount\":2,\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=156820#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png\",\"keywords\":[\"DeepSeekV3\",\"Mac\",\"Nightmare\",\"OpenAI\",\"Runs\",\"Studio\",\"tokens\"],\"articleSection\":[\"Apple\"],\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"CommentAction\",\"name\":\"Comment\",\"target\":[\"https:\/\/fivemor.com\/?p=156820#respond\"]}]},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/fivemor.com\/?p=156820\",\"url\":\"https:\/\/fivemor.com\/?p=156820\",\"name\":\"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that's a nightmare for OpenAI - Som2ny Network\",\"isPartOf\":{\"@id\":\"https:\/\/fivemor.com\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/fivemor.com\/?p=156820#primaryimage\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/?p=156820#primaryimage\"},\"thumbnailUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png\",\"datePublished\":\"2025-03-26T01:34:02+00:00\",\"breadcrumb\":{\"@id\":\"https:\/\/fivemor.com\/?p=156820#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/fivemor.com\/?p=156820\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/?p=156820#primaryimage\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png\",\"width\":1024,\"height\":573},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/fivemor.com\/?p=156820#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/fivemor.com\/?bp_activities=1\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that&#8217;s a nightmare for OpenAI\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/fivemor.com\/#website\",\"url\":\"https:\/\/fivemor.com\/\",\"name\":\"Som2ny Network\",\"description\":\"Daily Deals\",\"publisher\":{\"@id\":\"https:\/\/fivemor.com\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/fivemor.com\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/fivemor.com\/#organization\",\"name\":\"Som2ny Network\",\"url\":\"https:\/\/fivemor.com\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"contentUrl\":\"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png\",\"width\":300,\"height\":86,\"caption\":\"Som2ny Network\"},\"image\":{\"@id\":\"https:\/\/fivemor.com\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371\",\"name\":\"admin\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/fivemor.com\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png\",\"caption\":\"admin\"},\"sameAs\":[\"https:\/\/fivemor.com\"],\"url\":\"https:\/\/fivemor.com\/?author=1\"}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that's a nightmare for OpenAI - Som2ny Network","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/fivemor.com\/?p=156820","og_locale":"en_US","og_type":"article","og_title":"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that's a nightmare for OpenAI - Som2ny Network","og_description":"Join our daily and weekly newsletters for the latest updates and exclusive content on industry-leading AI coverage. Learn More Chinese AI startup DeepSeek has quietly released a new large language model that\u2019s already sending ripples through the artificial intelligence industry \u2014 not just for its capabilities, but for how it\u2019s being deployed. The 641-gigabyte model, [&hellip;]","og_url":"https:\/\/fivemor.com\/?p=156820","og_site_name":"Som2ny Network","article_published_time":"2025-03-26T01:34:02+00:00","og_image":[{"width":1024,"height":573,"url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png","type":"image\/png"}],"author":"admin","twitter_card":"summary_large_image","twitter_misc":{"Written by":"admin","Est. reading time":"9 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/fivemor.com\/?p=156820#article","isPartOf":{"@id":"https:\/\/fivemor.com\/?p=156820"},"author":{"name":"admin","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371"},"headline":"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that&#8217;s a nightmare for OpenAI","datePublished":"2025-03-26T01:34:02+00:00","mainEntityOfPage":{"@id":"https:\/\/fivemor.com\/?p=156820"},"wordCount":1911,"commentCount":2,"publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"image":{"@id":"https:\/\/fivemor.com\/?p=156820#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png","keywords":["DeepSeekV3","Mac","Nightmare","OpenAI","Runs","Studio","tokens"],"articleSection":["Apple"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/fivemor.com\/?p=156820#respond"]}]},{"@type":"WebPage","@id":"https:\/\/fivemor.com\/?p=156820","url":"https:\/\/fivemor.com\/?p=156820","name":"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that's a nightmare for OpenAI - Som2ny Network","isPartOf":{"@id":"https:\/\/fivemor.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/fivemor.com\/?p=156820#primaryimage"},"image":{"@id":"https:\/\/fivemor.com\/?p=156820#primaryimage"},"thumbnailUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png","datePublished":"2025-03-26T01:34:02+00:00","breadcrumb":{"@id":"https:\/\/fivemor.com\/?p=156820#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/fivemor.com\/?p=156820"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/?p=156820#primaryimage","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2025\/03\/nuneybits_Vector_art_of_a_whale_made_of_computer_code_chinese_f_619bb717-8408-4140-90f1-dd66bfc1e499.png","width":1024,"height":573},{"@type":"BreadcrumbList","@id":"https:\/\/fivemor.com\/?p=156820#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/fivemor.com\/?bp_activities=1"},{"@type":"ListItem","position":2,"name":"DeepSeek-V3 now runs at 20 tokens per second on Mac Studio, and that&#8217;s a nightmare for OpenAI"}]},{"@type":"WebSite","@id":"https:\/\/fivemor.com\/#website","url":"https:\/\/fivemor.com\/","name":"Som2ny Network","description":"Daily Deals","publisher":{"@id":"https:\/\/fivemor.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/fivemor.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/fivemor.com\/#organization","name":"Som2ny Network","url":"https:\/\/fivemor.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/","url":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","contentUrl":"https:\/\/fivemor.com\/wp-content\/uploads\/2026\/07\/4a0953c4-logo-300x86-1.png","width":300,"height":86,"caption":"Som2ny Network"},"image":{"@id":"https:\/\/fivemor.com\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/fivemor.com\/#\/schema\/person\/b85e3c3dc0e1daea076524dc8810c371","name":"admin","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/fivemor.com\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/729ae85bf62b9917e93538db2f2688ca?s=96&r=g&default=https%3A%2F%2Ffivemor.com%2Fwp-content%2Fplugins%2Fbuddypress-first-letter-avatar%2Fimages%2Fdefault%2F96%2Flatin_a.png","caption":"admin"},"sameAs":["https:\/\/fivemor.com"],"url":"https:\/\/fivemor.com\/?author=1"}]}},"_links":{"self":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/156820","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=156820"}],"version-history":[{"count":0,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/posts\/156820\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=\/wp\/v2\/media\/156821"}],"wp:attachment":[{"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=156820"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=156820"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=156820"},{"taxonomy":"dealstore","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fdealstore&post=156820"},{"taxonomy":"offerexpiration","embeddable":true,"href":"https:\/\/fivemor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fofferexpiration&post=156820"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}