Skip to main content

AI Compute Costs Surge, Chip Bonuses Soar: A Week in Compute Services

From iPhone 18's 38% BOM hike to Samsung's $400K bonuses, AI is reshaping compute economics. Plus: Microsoft's Maia 300, Meta's open model, and more.

Apple's 20th Anniversary iPhone Hits a Wall

Apple's plan to celebrate the iPhone's 20th birthday with a radical glass-bodied device may be dead. Jefferies analyst Edison Lee, citing supply chain checks, says the project has been paused or canceled due to poor production yields. The phone was expected to feature curved glass wrapping around the entire chassis, with almost no visible ports or buttons, relying on under-display Face ID and solid-state buttons.

If it had shipped, the retail price could have hit around $2,060, making it the most expensive iPhone ever. Apple even filed a patent in 2019 for a six-sided glass enclosure that seemed to match this design. The twist? Apple reportedly planned to roll the glass design into future Pro models to lift average selling prices. Now we're left wondering if the whole phone is scrapped or just the glass concept. Either way, it's a reminder that bleeding-edge hardware is hard to make at scale.

iPhone 18 BOM Costs Jump 38% for 256GB

Memory prices are biting. TrendForce estimates the bill of materials for a 256GB iPhone 18 is up about 38% year-over-year. The culprit: storage now makes up 30-40% of a phone's component cost, up from 10-15% a few years ago. Apple has room to absorb some of this, but it's a squeeze. Expect either higher prices or thinner margins, and that's a conversation every phone maker is having right now.

Korean Chip Engineers Are Suddenly Hot (and Rich)

Here's a fun one: AI has turned South Korean memory chip engineers into marriage-material royalty. The Wall Street Journal reports that a Samsung storage engineer earning a base salary of 80 million won can expect a bonus of 626 million won this year—roughly $450,000. SK Hynix, which now hands out 10% of operating profit as bonuses, could see average payouts of 779 million won next year. Dating agencies have upgraded these engineers from B+ to A+, putting them in the same league as doctors and lawyers. One Samsung engineer said he's stopped telling first dates where he works because he can't tell if they like him or his paycheck. Even Korean pop culture has noticed—a dating show featured a Samsung engineer who got attention from three women, and SNL Korea mocked SK Hynix employees flashing their company vests in luxury stores. The driving force? AI's insatiable appetite for memory, and it's changing lives in very tangible ways.

Microsoft's Maia 300: A September Surprise?

Microsoft is reportedly gearing up to unveil its next-gen AI chip, Maia 300, as early as September. The Information says Microsoft has talked to TSMC about locking in over 300,000 units of capacity for 2027—a huge jump from the tens of thousands of Maia 200 chips currently in use. The long-term goal is over a million chips, but supply constraints and TSMC's own capacity could limit that. Microsoft wants to use these chips in Azure and even offer them to big customers like Anthropic. The Maia 200, built on TSMC's 3nm process, is already in some data centers. If Microsoft can pull this off, it's a serious step toward reducing its reliance on Nvidia.

Meta's Muse Glimmer: Open-Weight AI for Your GPU

Meta just dropped Muse Glimmer, an open-weight agent model with 30 billion parameters, licensed under Apache 2.0. The kicker? It runs on a single 24GB consumer GPU. That's a big deal for developers who want to run AI agents locally without cloud costs. In benchmarks, it scored 75.5 on MCP Atlas and 74.6 on DeepSearch QA, and 76.0 on SWE-Bench Verified—solid numbers for a model you can run on a gaming card. Meta's superintelligence lab head, Alexandr Wang, also teased an open-weights release for Muse Spark 1.2. This is part of a broader push to put more capable AI in the hands of individuals, not just big corporations.

Nvidia's $500 Billion AI Infrastructure Bet

Nvidia has signed memorandums of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to create financing platforms for AI infrastructure. The goal: attract over $500 billion in third-party capital to fund AI compute. These platforms will provide Nvidia-powered infrastructure to AI developers, enterprises, governments, and cloud providers, with investment vehicles that tie returns to compute usage. Jensen Huang says this will help customers get scarce compute at scale. No word on individual commitments, but the scale is staggering. It's a clear signal that AI compute is becoming an asset class on its own.

AI Compute Costs: The Hidden Tax on Your Next Phone

All this AI demand has a ripple effect. The memory price surge that's inflating iPhone BOMs is the same force that's boosting Korean engineers' bonuses. But it's not just phones. The cost of AI inference is rising, too. Take Doubao, the AI assistant from ByteDance: it's now charging a service fee on hotel bookings made through its platform—11.4% plus a 0.6% payment fee, totaling about 12%. That's a direct monetization of AI-driven transactions, and it's a sign that AI services are moving from free trials to real business models.

China's AI Compute Push: Zhipu's 50,000 Domestic Chips

Zhipu, a Chinese AI startup, now has nearly 7 million registered API users and has deployed over 50,000 domestic AI chips to handle inference demand. That's a big bet on local silicon, and it's paying off: the company's ARR has grown 15x this year. They're also optimizing their infrastructure, with a KV cache splitting scheme that boosts throughput by up to 132% for long tasks, and a network architecture that cuts switch and optical module needs by a third. Zhipu's gross margins on its own compute are 50-60%, though that excludes initial capital costs. It's a race, though—Alibaba and ByteDance are pouring money into models and compute, and independent players like Zhipu are feeling the squeeze.

Microsoft's Maia 300: A September Surprise?

Microsoft is reportedly gearing up to unveil its next-gen AI chip, Maia 300, as early as September. The Information says Microsoft has talked to TSMC about locking in over 300,000 units of capacity for 2027—a huge jump from the tens of thousands of Maia 200 chips currently in use. The long-term goal is over a million chips, but supply constraints and TSMC's own capacity could limit that. Microsoft wants to use these chips in Azure and even offer them to big customers like Anthropic. The Maia 200, built on TSMC's 3nm process, is already in some data centers. If Microsoft can pull this off, it's a serious step toward reducing its reliance on Nvidia.

Meta's Muse Glimmer: Open-Weight AI for Your GPU

Meta just dropped Muse Glimmer, an open-weight agent model with 30 billion parameters, licensed under Apache 2.0. The kicker? It runs on a single 24GB consumer GPU. That's a big deal for developers who want to run AI agents locally without cloud costs. In benchmarks, it scored 75.5 on MCP Atlas and 74.6 on DeepSearch QA, and 76.0 on SWE-Bench Verified—solid numbers for a model you can run on a gaming card. Meta's superintelligence lab head, Alexandr Wang, also teased an open-weights release for Muse Spark 1.2. This is part of a broader push to put more capable AI in the hands of individuals, not just big corporations.

Nvidia's $500 Billion AI Infrastructure Bet

Nvidia has signed memorandums of understanding with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to create financing platforms for AI infrastructure. The goal: attract over $500 billion in third-party capital to fund AI compute. These platforms will provide Nvidia-powered infrastructure to AI developers, enterprises, governments, and cloud providers, with investment vehicles that tie returns to compute usage. Jensen Huang says this will help customers get scarce compute at scale. No word on individual commitments, but the scale is staggering. It's a clear signal that AI compute is becoming an asset class on its own.

AI Compute Costs: The Hidden Tax on Your Next Phone

All this AI demand has a ripple effect. The memory price surge that's inflating iPhone BOMs is the same force that's boosting Korean engineers' bonuses. But it's not just phones. The cost of AI inference is rising, too. Take Doubao, the AI assistant from ByteDance: it's now charging a service fee on hotel bookings made through its platform—11.4% plus a 0.6% payment fee, totaling about 12%. That's a direct monetization of AI-driven transactions, and it's a sign that AI services are moving from free trials to real business models.

China's AI Compute Push: Zhipu's 50,000 Domestic Chips

Zhipu, a Chinese AI startup, now has nearly 7 million registered API users and has deployed over 50,000 domestic AI chips to handle inference demand. That's a big bet on local silicon, and it's paying off: the company's ARR has grown 15x this year. They're also optimizing their infrastructure, with a KV cache splitting scheme that boosts throughput by up to 132% for long tasks, and a network architecture that cuts switch and optical module needs by a third. Zhipu's gross margins on its own compute are 50-60%, though that excludes initial capital costs. It's a race, though—Alibaba and ByteDance are pouring money into models and compute, and independent players like Zhipu are feeling the squeeze.

The Bottom Line

This week's news paints a clear picture: AI is reshaping the economics of compute. Memory prices are up, chip engineers are getting rich, and companies are scrambling to secure capacity—whether that's Nvidia's financing push, Microsoft's chip ambitions, or China's domestic silicon drive. For the rest of us, it means higher prices for the next iPhone and a growing number of AI services that might charge a fee. But it also means more open models like Muse Glimmer that run on your own hardware, and more options for getting compute without relying on the usual suspects. It's a wild time to be in compute services.

Share this article:

Comments (0)

No comments yet. Be the first to comment!