Hi there, nice to meet you. Follow me at http://belfong.com/tweets
503 stories
·
2 followers

Hot Chips 2026: IBM's first dual-ISA core natively executes ARM and z/Architecture in the same core; all cores run at 5.7 GHz base frequency — next-gen mainframe AI processor is built on 2nm node with 11 cores

1 Comment

This Tom's Hardware Premium article is free to read with a Tom's Hardware account; no payment necessary. We're offering free access from August 23 to 26 so you can read all of our reporting from Hot Chips.

IBM's next-gen AI processor is the first time it has supported dual ISA execution natively within the same core. Born out of a collaboration between IBM and Arm that was announced in April, the chip is designed to bring the software support available across the Arm ecosystem to IBM's mainframes, allowing businesses to unify deployment rather than relying on separate Arm/x86 servers and z/Architecture mainframes for different purposes.

The approach here isn't a heterogeneous CPU with separate Arm cores packaged on the same chip; IBM has built a core that can execute either z/Architecture or AArch64 instructions, and can switch between them dynamically "within nanoseconds," according to the company. During the Hot Chips 2026 reveal, IBM says it believes this is the first processor to treat both ISAs as "first-class citizens."

IBM relies on Linux Kernel-level Virtual Machine (KVM) to support AArch64 instructions, the same mechanism that allows IBM to support Linux on Z mainframes. Standard z/Architecture instructions bypass KVM. The idea is to ensure that mainframe reliability isn't sacrificed for broader software support, with IBM claiming 99.999999% uptime, even further than traditional high-availability claims. IBM says that equates to just 0.032 seconds of downtime per year.

Much of the development in the software world, particularly around AI, happens with x86 and/or Arm targets in mind, leaving the mainframe behind to figure out its own solution. IBM could, and has, worked to port this software to s390x, but that's not a long-term solution. "We would never be able to work with all of them," Tina Tarquinio, chief product officer at IBM for IBM Z and LinuxONE, told VentureBeat. IBM's dual-ISA core can execute Arm software without modifications, according to the company, allowing Arm-based virtual machines to run as if they were operating on native-Arm silicon. And that's because, well, they are operating on native-Arm silicon, just in a different way.

A high-level look at IBM's dual-ISA processor and next-gen Spyre accelerator

IBM dual-ISA processor design
IBM
IBM dual-ISA processor design
IBM
IBM dual-ISA processor design
IBM
IBM dual-ISA processor design
IBM

The processor that presumably will live in z18 mainframes comes with 11 high-performance cores, built on a 2nm process, that can operate at a base frequency of 5.7 GHz. Even on those specs, and ignoring dual-ISA execution, it's a considerable step up over the current Telum II processor that IBM introduced in 2024. That chip features eight cores operating at up to 5.5 GHz. Otherwise, IBM's dual-ISA processor comes with the same 36 MB of private L2, as well as virtual L3 and L4. These caches are larger than Telum II at 432 MB of virtual L3 and 3.5 GB of virtual L4.

Also carried forward is an on-chip DPU, as well as hardware accelerators for AI, compression, and cryptography workloads, same as Telum II. Outside of more cores and higher clocks, much of the work on IBM's dual-ISA processor happened, naturally, in the core itself, which we'll dig into in the next section.

The core supports simultaneous multithreading, which is available to both ISAs.

IBM next-gen AI accelerator
IBM
IBM next-gen AI accelerator
IBM
IBM next-gen AI accelerator
IBM
IBM next-gen AI accelerator
IBM
IBM next-gen AI accelerator
IBM

Alongside the processor, IBM teased its next-gen AI accelerator at Hot Chips 2026. It's considerably more capable than the current Spyre accelerator, which makes sense, given IBM's new capabilities with Arm. The new accelerator comes with 16 cores that include optimizations for newer AI data formats, including FP4/MXFP4.

The big change comes in memory, however, with IBM moving off LPDDR5 to lower-capacity but significantly higher-bandwidth HBM3e. Each accelerator comes with 96 GB of HBM3e, offering up to 4TB/s, 20x that of what IBM is able to deliver with LPDDR5.

IBM core changes to support z/Architecture and ARM

IBM next-gen processor CAD design

(Image credit: IBM)

Much of the work on IBM's next-gen processor happened in the core itself in order to support native execution of AArch64 instructions. The chip has a full hardware implementation of AArch64 v9.3 with Scalable Vector Extension (SVE) support, supporting 2,792 AArch64 instructions.

IBM branch prediction unit.

(Image credit: IBM)

Starting at the top of the core, the branch prediction area uses the existing Telum II design without any changes.

IBM dual ISA fetch engine.

(Image credit: IBM)

In the fetch engine, IBM leverages virtual cache tags to fetch data quickly with cache to avoid translation overhead.

IBM decode engine.

(Image credit: IBM)

IBM built automation tools to consume the ARM XML and understand how to move instructions through the core. IBM says this is the biggest area of silicon expansion in the core in order to support decoding AArch64 instructions.

IBM next-gen dispatch.

(Image credit: IBM)

In dispatch, IBM repurposed general purpose register rename in banked general registers 16 through 31.

IBM load store in next-gen processor.

(Image credit: IBM)

In the arithmetic and load/store units, much of the major data flow is shared; addition is addition, as IBM put it. However, IBM implemented new hardware structures for SVE and special data types like FP16. IBM also says there was some non-obvious reuse of its existing CISC, such as memory copy and clear.

X-late in next-gen IBM CPU.

(Image credit: IBM)

The X-Late, or translation, engine reuses the Translation Lookaside Buffer (TLB) but leverages a new page walk.

IBM next-gen processor recovery unit.

(Image credit: IBM)

As opposed to a heterogeneous chip, which accomplishes mixing ISAs on the same die by leveraging different cores, IBM says the driving force behind a dual-ISA core was to deliver the scale of Arm software on a mission-critical platform. Mainframes are still the bedrock of vital data movement in financial institutions, governments, and more.

IBM generally ships new mainframes every two and a half to three years, with z17 mainframes revealed in 2024 at Hot Chips. We expect this mainframe to follow a similar timeline. As usual with deep mainframe infrastructure, however, the actual rollout largely depends on the institution's individual needs.

Full IBM Hot Chips 2026 presentation

IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM
IBM Hot Chips 2026 presentation.
IBM


Read the whole story
Belfong
3 hours ago
reply
Wow...
malaysia
Share this story
Delete

Defining an “Agent Harness”

1 Comment

I’ve recently been asked to explain what an “agent harness” is (particularly after I wrote about my favorite app of the year so far, Open Minis). I realized it was one of those concepts I could understand intuitively but not quite articulate. Thankfully, other people have done a better job of explaining it than I have.

Drew Breunig calls them “situated agents”:

Harrison Chase once excitedly shared an insight that agents are comprised of 4 things: a system prompt, a planning tool, a file system, and subagents. In the year-plus since he said that, I think this remains largely true. (Though you might tweak it to have general tools, etc.)

[…]

Imagine the simplest coding agent, and you at the keys. Let’s slowly zoom out and consider all the elements the harness can manage:

  1. The Session: The current task and context, as both a trajectory and a durable, branchable log. You can zoom backwards, fork, and replay it.
  2. The Environment: The instance, defined as a sandbox, terminal, worktree, computer, and/or container.
  3. The Repo: The project. Versioned with Git, it contains your code, history, current work, AGENTS.md, guides, and hooks.
  4. Memory: The person’s predilections, accrued over time, managing progress and past decisions.
  5. Skills: The domain, artifacts describing reusable workflows or domain knowledge worth wielding in this situation.
  6. The Team: Your colleagues and counterparts. Shared rooms, shared traces, project tracking tools, issues, and bug reports.
    • The Organization: The policies and audits, defined by legal, leadership, and procurement.
    • The Model: The LLMs themselves, the common artifact shared by all. Stochastic blobs we all poke trying to evoke positive outcomes. Log or train on their quirks and adjust.

I also thoroughly enjoyed this explanation and excellent visualization by Ted Spare, Dexter Storey, and Sarim Malik, writing for Rubric Labs:

A harness is the software that translates a model into a system that can affect its environment.

[…]

Functionally, a harness makes a model agentic, meaning it can take action.

And:

As coding agents begin to run for longer, and deploy more intelligence through dispatch, they are owning larger and more complex problems end to end, and the developer is less in the loop to steer the agent. The value of high quality planning increases as agents implement the plans more autonomously. Harnesses like Claude Code and Codex ship with a native planning mode, where the agent must first create a detailed Plan.md file with feedback from the user before executing. These harnesses then place a reference to the plan and the todo list into a top level state (system prompt) so that the agent doesn’t forget what it’s working on across long runs.

Given the multi-model, hybrid structure of the new Siri AI, we should probably assume Apple also made a lightweight “Siri harness” to aid the on-device orchestration of the entire system.

Read the whole story
Belfong
3 days ago
reply
I’m interested to know how agents work. This is useful read.
malaysia
Share this story
Delete

Add your own stuff to Apple Wallet

1 Comment
Glenn Fleishman, art by Shafer Brown

Wallet on iPhone is a bulging catch-all. Apple stores all your payment card information there, and it’s an interface for adding and removing cards, as well as viewing associated details. It’s the place where you interact with your Apple Card (like scheduling payments or viewing transaction reports), the Apple Card savings account, and Apple Cash. But it also handles related but miscellaneous categories of other things: tickets for events, reward cards, membership cards, your library card… really, anything with a number, a scannable code, or a connection to payment winds up there.

Despite all that, Wallet has been deficient: you couldn’t easily add items to Wallet that the party issuing them hadn’t formatted and made available with an “Add to Apple Wallet” button, whether from a website or within its app. Starting in iOS 27—coming in September but available in a public beta now—you can use a new pass-creation tool to make your virtual billfold even larger. Third-party tools still fill the gap for iOS 26 and earlier users, or for those who want customization that Apple isn’t providing in Create a Pass.

Make your own Wallet entry

In iOS 27, open the Wallet app and tap the Add button in the upper-right corner. You’ll notice that among the options you’re familiar with, Create a Pass appears. Tap that.

If you have a device that supports Visual Intelligence and the pass has a barcode or QR code, tap Continue, and Wallet opens a camera sheet.1 When you tap the shutter button, Visual Intelligence analyzes the image and tries to extract elements, like your name, the organization name, or the code. If it can’t, it will offer a Try Again and Create Pass Manually dialog.

If it’s successful, Wallet next displays a preview screen. You can tap Add to accept the pass as it appears, or tap Edit Pass to modify the fields shown and the contents of those fields.

Side by side: photo of backside of food co-op membership card and pass created by Create a Pass in iOS 27.
Scan a card, even a munged-up old one like the at left, and Wallet can extract important details. I added the expiration date.

There’s a quirk in editing fields that baffled me for a while, but I teased out the answer in testing. Apple pre-populates field areas, marked with white rounded rectangles with subtle glowing edges. If you tap the Edit button (in the lower-right corner) and choose Add/Remove Fields, you cannot add new field areas. (See the figure for how I’ve labeled this.) What you can do is tap on an existing field within a field area to change it, or you can tap within a glowing rectangle (the field area) to add another field. This is not a great interface, and I expect it will mature.

Figure of a pass created in Wallet for iOS 27 with callouts indicating the various parts.
A field areas, as I call it, is a place you can add and remove fields from.

Tap the Palette icon to change the background color or choose a different background image from a small set in the Wallet library.

If Visual Intelligence isn’t available or you want to build the pass element-by-element:

  1. Tap Create Pass Manually.
  2. On the Choose Pass Template screen, pick the data format and default color scheme to use for your pass: Standard, Membership, or Event. This pre-populates different fields for you to fill out. (You can change the color later via the Palette icon.)
  3. Next, tap fields to enter data. Each field has a label in smaller type and placeholder text in larger type.
  4. Tap Add Code to scan a barcode or QR code.
  5. When finished, tap the blue checkmark.

You can then tap the Palette icon to change its appearance, or tap Edit: Add/Remove Fields icon to further edit the pass.

Side by side screenshots of the generic pass creation screen in Wallet for iOS 27 (left), and a field-editing sheet overlaying the pass (right).
If you create a pass manually, Apple adds a number of field areas and fields for you to modify (left). You have access to many different kinds of fields from an extensive list (right).

After the pass is added to Wallet, you can tap it, tap the More button, then tap Edit Pass. You can also tap Pass Details to see the raw information added. The View Original Photo button lets you bring up the picture you took if you used Visual Intelligence.

If you’re not yet using the iOS 27 beta or want options that Apple isn’t offering, you could try one of these free tools. Be aware, however, that Wallet passes have to be signed, so any data you provide to these companies is sent to them in the clear. Both have details in their privacy policies about how they protect your data.

  • Wallet Creator (iOS app) has over 80 prebuilt templates for existing businesses, or you can roll your own pass or card. Wallet Creator says the app collects no data. However, the App Store’s description for the app, under “Our Terms of Use & Privacy Policy,” provides more information about how data is sent to a server, processed, and deleted.
  • WalletWallet (web app) is a single-page form to create a pass. The page offers no privacy information, but it’s built using a business-oriented API the company makes, and that API has this explanation on the trip your data takes.

My book on Wallet, Take Control of Apple Wallet, will be updated later this year to incorporate all of this new information.

iPad and Mac have empty pockets

While Apple brought the Phone app to iPad and Mac in iPadOS 26 and Tahoe, the company still treats Wallet as a second-class citizen on those two platforms, even though both allow the use of Apple Pay. I’d love to have Wallet synced and fully available for managing all cards, not just payment-related ones, as a standalone app in the future.

[Got a question for the column? You can email glenn@sixcolors.com or use /glenn in our subscriber-only Discord community.]


  1. Visual Intelligence requires Apple Intelligence, which is available on an iPhone 15 Pro, iPhone 15 Pro Max, iPhone Air, or any iPhone 16 or 17 model. 
Read the whole story
Belfong
5 days ago
reply
This is nice. I can put all those loyalty cards into Wallet!
malaysia
Share this story
Delete

Hacking Public Wi-Fi DNS to Steal Credentials

1 Comment and 2 Shares

Criminals are hacking into public Wi-Fi devices—at hotels, conference centers, and so on—around the world and changing their DNS settings. The goal is to redirect users to fake login pages and steal their credentials.

Read the whole story
Belfong
7 days ago
reply
It's why I avoid to use public Wi-Fi whenever possible and just use my mobile phone as Personal Hotspot.
malaysia
Share this story
Delete

Nvidia’s Risky Business

1 Comment

Nvidia is finding new ways for its customers to raise money, and it's expanding the risk of the AI buildout significantly.


Listen to this Update in your podcast player


On January 1, 1870, Jay Cooke, hailed as an American hero for his role in financing the Union effort in the Civil War, signed a contract that would, if you squint, lead to world war.

In 1864, Congress had created the Northern Pacific Railway Company with the goal of linking the Great Lakes and Puget Sound with tracks that would eventually run from Duluth to Tacoma; the charter included 40 million acres of land adjacent to the proposed line in exchange for accomplishing the build-out. For the ensuing six years, however, Northern Pacific struggled to secure financing, even as the Union Pacific and Central Pacific railroads built towards each other, driving the golden spike linking Sacramento and Omaha in May 1869.

Northern Pacific had approached Cooke about funding in 1866, but lacked the generous federal guarantees that undergirded Union Pacific and Central Pacific (which, it should be noted, led to an incredible amount of graft); Cooke, himself no stranger to the financial power of the federal government, wasn’t interested. Ultimately, however, Northern Pacific gave him an offer he couldn’t resist: a commission of 12 percent on every bond, and $200 of Northern Pacific stock for every $1,000 in bonds he sold.

Cooke soon found that his institutional peers agreed with his earlier refusal, and weren’t interested in his bonds, so he leaned on the same tactics he honed selling war bonds: appeals to patriotism, control of the media, and promises of railroad fortunes, backed by industrial-scale distribution. At the peak Cooke employed 1,500 salespeople and funded 1,300 newspapers (through a combination of advertising and direct payments) with a brand burnished by the Civil War. Retail investors could already buy railway bonds; Cooke made them his primary funding mechanism.

This was, to be certain, an incredible innovation. It used to be the case that if you couldn’t get loans from the government or from banks, you couldn’t get much money at all. The problem was that Northern Pacific’s capital needs were endless, and by September 1873, as credit tightened worldwide thanks to a crash on the Vienna stock exchange and the demonetization of silver, Cooke, who had been funding Northern Pacific from deposits in between bond issuances, could find no more buyers. The subsequent bankruptcy of Jay Cooke & Company triggered the Panic of 1873, culminating in endless railroad bankruptcies across the country, a multi-year depression, multi-decade deflation, and, one could argue, the financial conditions that made Europe, four decades later, into a tinder box.

Northern Pacific did eventually finish their line, by the way, with multiple bankruptcies along the way; ultimately, they were one of four railroads that were merged to form the Burlington Northern Railroad. Burlington Northern would eventually merge with the Atchison, Topeka and Santa Fe Railway to form BNSF Railway; Berkshire Hathaway would purchase the parent corporation in 2009.

Blowing Through Debt

If this story sounds vaguely familiar it might be because Cooke is — for obvious reasons — a central character in Liaquat Ahamed’s new book, 1873, released earlier this year. Ahamed is not shy about drawing a link between the collapse of the railroad buildout and the current AI moment; the book’s very first page — even before page 1 — is about translating sums of money, and concludes thusly:

In order to grasp the true significance of sums of money that relate to the economic situation of whole countries — such as the size of the indemnity imposed on France after the Franco-Prussian war — it is most useful not simply to make allowances for changes in the cost of living but instead to adjust for changes in the size of economies. To translate such figures into comparable 2026 magnitudes, multiply by a factor of 1,200. Thus the $500 million that went into U.S. railway bonds annually during the boom years of the early 1870s would today be the equivalent of $600 billion, roughly what is projected to be invested by major tech companies in 2026.

Microsoft CEO Satya Nadella is certainly aware of the connection: he cited 1873 as “the book to be read” on the company’s recent earnings call. Perhaps it’s not a coincidence, then, that Microsoft, alone amongst the hyperscalers, still boasts substantial free cash flow — $19.6 billion last quarter. Microsoft is the one hyperscaler still abiding by the dictum used to deny the existence of a bubble: its CapEx isn’t funded by debt.

This was, believe it or not, a defense that could be used for nearly all of Big Tech a year ago; then, between September and November, Oracle, Meta, Alphabet, and Amazon issued a combined $80 billion in debt for building out infrastructure. That was only the beginning: after raising a combined $108 billion in all of 2025, these four companies have, as of July 7, already raised $194 billion this year. Unsurprisingly, spreads are rising, and 86% of the bonds issued this year are already trading at higher yields than at issuance. Cover for recent issuance has fallen to less than 2x, from 5x in February.

The real shock, however, came at the beginning of June, when Google announced it would raise $85 billion in equity, including a special $10 billion issuance to the aforementioned Berkshire Hathaway. I wrote at the time in The Google Capital Company:

It is worth noting that $10 billion is a relatively small amount of money to both companies. To that end, perhaps the primary utility is as a signaling mechanism. On Google’s side, the signal is that the expected demand is actually far greater than anyone thinks, and that the company is ready and willing to fund supply using all means at its disposal, including equity; for them Berkshire Hathaway’s investment is an endorsement of this view and a validation of the wisdom of the investment. And, on the flip side, if the signal is correct, then Berkshire Hathaway is getting a deal and putting its cash flow machines to work building the future.

I concluded:

Implicit in this analysis was that there was enough compute capacity in the world to be bought; what happens, however, when and if there isn’t? What if the ultimate battle — the one that determines who gets compute — becomes a matter of who can bring the most cash to bear? And what if that advantage compounds, such that the company with the most cash capacity ends up with the most compute capacity (which we already know they will sell, in addition to using themselves) driving the ability to generate more cash? In that world, what company would be your best bet?

The implied answer, of course, was Google.

DeepMind Drama

Google right now is no one’s bet, at least in terms of the frontier. After the departure of DeepMind CEO Demis Hassabis (technically promoted to chairman, but no longer in charge of day-to-day operations) and Gemini co-lead and former Chief Scientist Jeff Dean, along with a host of other prominent researchers, SemiAnalysis declared that Gemini is Cooked:

For all intents and purposes, we believe DeepMind is no longer a frontier lab. We said as much a few months ago to our Tokenomics clients due to large numbers of departures from their reinforcement learning teams and poor compute allocation. Google will continue meandering on and releasing models, but their odds of reaching SOTA again have dropped to zero.

Furthermore, the biggest beneficiary of today’s news is neither Anthropic nor OpenAI—it’s Google Cloud. Whereas Gemini and GCP used to desperately fight for compute allocation, it’s now clear that Thomas Kurian won. We expect GCP revenue growth to meaningfully accelerate as a result.

From later in the post:

We’ve obviously been quite bearish on DeepMind thus far, and if we had to steelman the case for why they’ll still be able to train a true SOTA model in the future, it would go something like the following:

  • The current setup clearly wasn’t working. With the existing leadership team, their odds of catching up to Anthropic/OpenAI looked extremely slim.
  • Now that they’ve cleaned house, the new guys can start from a blank slate. Maybe they’ll even acqui-hire a neolab like SSI or Thinking Machines.
  • With this new team, their odds of catching up to the frontier actually increase.

Perhaps there’s some world in which this happens, but we think the odds are basically zero. The issue with Google was not Jeff Dean nor Noam Shazeer, but rather their extremely bureaucratic, painfully slow, and strategically timid culture. Remember that DeepMind had an AI chatbot 1 year before ChatGPT but was not allowed to release it due to fears of disrupting their core business.

Actually, you could make the case the problem was also Hassabis and DeepMind. I explained in an Update after Google I/O how Hassabis’ vision of the frontier was fundamentally different from the other frontier labs because he believed in world models, not just text/code, and concluded:

What falls out of [Hassabis’ vision] are models with multimodality — in contrast to Claude, which outputs text only — and, it must be said, not nearly as impressive coding capabilities. This gets at the point of this entire digression: I think it’s possible that the reason Google is widely considered to be behind both Anthropic and OpenAI in terms of coding, particularly long-running agentic workflows that depend just as much on the harness as the model itself, simply comes down to their research team having other priorities. That’s why the coding parts of this keynote fell on the Antigravity team, not DeepMind, and why Hassabis was barely on stage.

From this perspective, last week’s events are less surprising, and were arguably foretold at I/O: Hassabis might be right about world models being the path to AGI, but Google has run out of patience in terms of letting him find out; Google co-founder Sergey Brin is reportedly deeply involved and closely allied with Koray Kavukcuoglu, the new DeepMind CEO, and I wouldn’t be surprised if the company is pivoting to Anthropic’s more text- (and thus code-) centered approach.

Google’s Infrastructure Bet

What is fascinating about Google’s position is that these machinations do not necessarily mean the Berkshire Hathaway bet was a bad one; indeed, it’s arguably good news. This is what the SemiAnalysis article was driving towards, and it’s a point I made last week about Google’s recent earnings:

The story seems to be very similar to last quarter, with even more Google Cloud growth: 82% year-over-year (compared to 63% last quarter, and 32% a year ago), with 36% margins (compared to 33% last quarter, and 21% a year ago). I wondered then how much of this growth was actually Anthropic, and while we didn’t get clear confirmation this quarter, I thought this answer from CEO Sundar Pichai on the earnings call about why Google needs to rent 3rd-party capacity was notable:

I think on the bridge deal, the main thing I would say is, look, there are — on the margin, there are very, very large customers of ours on Cloud who we are trying to support them through this extraordinary moment. And the incremental opportunities they are bringing to us, while a short‑term cost over a few months may be very high, in the lifetime of the deal, as we bring more capacity on, is highly ROI‑positive. So those are factors we are taking into account. So are you willing to take upfront a six‑month deal to be able to serve the customer in what is a multiyear opportunity where the margins and the returns are very, very attractive over that multiyear horizon? So hopefully that gives some color on how we’ve thought about those opportunities.

That customer is almost certainly Anthropic.

Again from SemiAnalysis:

More than 20% of total TPU shipments from 3Q26 to 4Q27 are being sold directly to Anthropic. This is excluding the hundreds of thousands of TPUs GCP already rents to Anthropic today, and the many hundreds of thousands more they’ve committed to rent to Anthropic and Meta over the next 6 quarters…

If you’ve ever listened to an interview of Google Cloud CEO Thomas Kurian, you know he is not AGI pilled. In one podcast, for example, he argued that it’s great for TPUs to become “general purpose infrastructure” that supports customers like Citadel, the Department of Energy, and generic high performance computing. And when asked why he was selling compute to Anthropic despite them competing with Gemini, he said this was the natural consequence of Google being a “platform company.”

Kurian said the same thing to me in a Stratechery Interview:

We sell different parts of our stack. One of the things people don’t realize is we monetize many different parts of the stack in different ways. Like Anthropic, there’s a lot of labs that use our stack — in fact, most of the large AI labs use our stack. So if somebody uses TPUs to either to train their model or to use it for inference, we’re monetizing that part of the stack, that gives us resources to then fund our R&D and other investments. Some of the labs use our TPU and our Gemini model, others may use our TPU and then buy our cybersecurity protection for their models. So as a platform player, we have to allow our technology to be monetized in as many ways as possible and we don’t see it as a zero sum.

We’ll see how zero sum compute actually is — there are reports Google’s researchers have been starved for compute — but the overall takeaway is that whether or not Google is competing for the frontier, they are absolutely competing to dominate AI infrastructure. And, in a world where intelligence is a commodity, TPUs in particular are a big deal.

Last month, in Who’s Afraid of Chinese Models?, I talked about commodity markets in the context of frontier labs versus everyone else; in commodity markets marginal costs are determinative of not just profitability but also viability, and I made the case that the frontier labs are well-positioned to have superior cost structures for any given unit of intelligence.

That cost structure, at least for now, includes the cost of renting compute, and it seems likely that TPUs are cheaper than Nvidia GPUs; Anthropic may have built for TPUs (and Amazon’s Trainium chips) because only Google and Amazon had the wherewithal to fund them, but at this point that ability may very well be a significant advantage. The fact that Anthropic is straight up buying TPUs for its own data centers (converting compute costs from marginal costs to capital costs) suggests that is the case.

What is notable is how amenable Google is to share, even at the price of needing to issue equity. This, however, fits the Berkshire Hathaway model that I wrote about in The Google Capital Company:

One of the businesses Berkshire Hathaway used the See’s profits for was on the opposite end of the spectrum in terms of capital utilization: BNSF Railway. Railways require a lot of capital to operate; BNSF consumed $3.8 billion last year; they also make a lot of money: BNSF’s net income was $5.5 billion on revenue of $23.4 billion. To put that in perspective, the total amount that Berkshire Hathaway has made from See’s Candies is probably less than $3 billion (the last disclosure was “over $2 billion” in 2019), i.e. less than BNSF made last year…

In fact, you can make the case that Abel is actually just replaying Buffett’s strategy, only this time Berkshire Hathaway is See’s Candies, and Google is BNSF. At the end of last quarter Berkshire Hathaway had $373 billion in cash, and $25 billion in free cash flow in 2025. How many companies could actually employ that cash in a way that generated a high rate of return?

It’s hard to imagine a better option than Google. The company is not only investing in AI, but has optionality in terms of outcomes: its Services business benefits from the investment, it is in contention at the model layer with Gemini, and it can sell capacity to the frontier labs. Moreover, that capacity has a sustainable cost advantage because of TPUs, which means that in a world where compute becomes a commodity — as hard as that is to imagine right now — Google is the hyperscaler that is poised to make the most profit.

Notice that I didn’t say margin; if that were Google’s concern they would almost certainly be making different choices. Profit, however, is an absolute number, and Google is bringing everything to bear — first its cash flow, then its debt, and now its equity — on making money from the infrastructure build-out.

Nvidia’s Investable Asset Class

Today corporate executives and financial engineers don’t need to control newspapers; thanks to his new X account, Nvidia CEO Jensen Huang can go straight to the public. From an X Article posted last night:

NVIDIA AI Factory Compute Is Becoming an Investable Asset Class

Today, we announced partnerships with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR to establish independent financing platforms designed to mobilize over $500 billion of third-party capital to support the buildout of AI infrastructure over time.

This is a major milestone for NVIDIA and the AI industry. We have moved from an era in which companies bought chips and built data centers project by project to one in which AI factories can be financed as productive infrastructure — with repeatable platforms, long-term institutional capital and a diverse customer base that uses compute to create revenue.

AI has reached an inflection point. It is moving from research into production. AI is creating real value, and the infrastructure behind it is becoming one of the world’s most productive assets. In AI, compute is revenue.

Huang argues that Nvidia-based AI factories are fungible, protecting residual value, and that CUDA makes AI factories better over time, extending their economic value; according to Huang:

These are the characteristics of an investable infrastructure asset: it produces revenue, serves a broad market, improves in performance over time and can be redeployed.

Thus the attempted formalization of a new investment structure:

The demand for AI infrastructure is extraordinary. But access to capital is uneven. Many great AI companies, enterprises and AI clouds have demand for compute but do not yet have access to financing at the scale or cost required to build quickly. That is why we are partnering with the world’s leading long-term capital providers.

Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs and KKR are also among the world’s leading infrastructure investors, with deep expertise in underwriting long-lived, productive assets. Together, we are creating repeatable financing platforms to help the AI ecosystem build the factories it needs.

What Apollo et al. are, are new sources of capital beyond the investment grade debt markets. In that sense this proposed structure is somewhat akin to Google’s equity issuance: a way to secure funding beyond bonds. The difference, however, is stark: whereas equity dilutes the upside for investors without adding risk to the company, this structure preserves Nvidia’s margins by finding new pools of capital willing to bear risk.

It’s not a total free ride for Nvidia: the company is backstopping opportunities with up to 25% residual-value based financing, suggesting that Huang believes his “investable asset class” pitch much more than the market does. That is, in a certain sense, a price cut, as the goal is to reduce the cost of capital for entities building data centers with Nvidia chips, by putting Nvidia’s profits on the line for uncertain investments. That guarantee is downstream from Google’s (and soon Amazon’s) aggressiveness: why build a data center with Nvidia chips if you can buy TPUs or Trainiums (Nvidia chips are likely better, but if the constraint on new data centers is capital, lower up-front prices may matter more than token efficiency).

Nvidia’s bigger problem is one that has been apparent for a long time; I wrote back in 2024:

In the before-times, i.e. before the release of ChatGPT, Nvidia was building quite the (free) software moat around its GPUs; the challenge is that it wasn’t entirely clear who was going to use all of that software. Today, meanwhile, the use cases for those GPUs is very clear, and those use cases are happening at a much higher level than CUDA frameworks (i.e. on top of models); that, combined with the massive incentives towards finding cheaper alternatives to Nvidia, means both the pressure to and the possibility of escaping CUDA is higher than it has ever been (even if it is still distant for lower level work, particularly when it comes to training).

The situation today, with Anthropic and OpenAI appearing to pull away, is even more problematic: Anthropic has not been dependent on CUDA for years, and OpenAI is moving in that direction, at least for inference. If those companies win then Nvidia’s profits will be squeezed — indeed, the implication of that backstop is they already are (this, needless to say, is why Huang’s first post was an open letter in defense of open models).

Risky Business

This might not cost Nvidia anything in the end: if AI revenues truly take off, then the debt markets will open back up, and ultimately companies will go back to funding infrastructure investment through free cash flows. Right now, however, is the danger zone, as hyperscalers blow through the debt markets and Google at least starts to tap equity. To the extent Nvidia competes through novel funding mechanisms that, at the end of the day, draw on things like insurance floats and pension funds and other long-run liabilities that are the bread and butter of the asset managers the company is partnering with, the risk — unmarked, unlike equity — is considerably higher.

That’s why I started with 1870 and Cooke’s ill-fated agreement with Northern Pacific. Yes, the upside the deal afforded Cooke was incredible, but it was incredible for a reason: it was very risky, and pioneering new funding mechanisms only served to spread the pain when it all blew up. It’s one thing to spend all of your free cash flow; it’s another thing to tap the debt markets. And, beyond that, it’s a completely new nerve-racking thing to bring safety-seeking assets to bear. AI better deliver before it’s too late.


Add to your podcast player: Stratechery | Sharp Tech | Dithering | Sharp China | GOAT | Asianometry


Read the whole story
Belfong
10 days ago
reply
This is an interesting read especially if you are an investor of GOOG and NVDA.
malaysia
Share this story
Delete

★ You Don’t Need to Worry About Scratching Your iPhone Camera Lenses

1 Comment

Following up on my post yesterday about modern iPhones and scratch resistance, and my personal habits of (a) almost never using an iPhone case, and (b) never setting the iPhone face down except on soft (cloth) surfaces.

I should have anticipated this, because I’ve had this conversation with real-world normies repeatedly in recent years, but I got a bunch of questions from people who say they’re wary of ever putting their iPhone down on its back because they’re worried about scratching the camera lenses. That’s not a silly thing to worry about. It looks like those lenses are glass; glass scratches; and it sure seems like a scratched camera lens might forever ruin all the photos and videos you take with the camera.

There are a couple misconceptions here though. First, the glassy flat circles exposed on the back of your phone aren’t the camera lenses. Those are covers over the lenses, which are smaller, spherical (not flat), and recessed. You can look through the covers and see the actual spherical lenses inside. The exposed lens covers are made of sapphire, not glass, and are thus incredibly scratch-resistant. They’re by far the most scratch-resistant parts of the phone. I’ve never once found even a tiny scratch on any of my iPhone camera lens covers, and I’ve never taken any particular care to avoid scratching them. I mean, I don’t drag the lenses face-down on surfaces. I’m not trying to scratch them. But I set my phones down on hard surfaces lenses-down all the time and they never seem to pick up even fine scratches.

Zack “JerryRigEverything” Nelson is the guy on YouTube who makes videos where he scratches the hell out of products to see how durable they are. (And he bends them, burns them, and abuses them in other gruesomely creative ways.) In his iPhone 17 Pro video, starting around the 4:10 mark, he tries scratching the sapphire lens covers with a sharp razor blade. No effect. Sapphire is far more likely to shatter than scratch, because it’s so hard. (You may recall that a decade ago Apple pursued using sapphire for the displays of iPhones but it didn’t work out.) Don’t try scratching your lenses with a diamond hardness pick and you’ll be fine.

Also, believe it or not, a fine scratch on the lens cover will almost certainly not affect image quality at all. Try sticking a strand of hair (which is probably much thicker than a typical scratch) to the surface of your iPhone 1× main camera lens. Take a picture. Clean the hair off the lens. Retake the same picture. You almost certainly won’t see any difference at all. Here’s a great video from DPReview back in 2020 showing just how much dust or scratching you need to put on the outside of a lens to degrade image quality. It’s counterintuitive but the surface of the lens is not where light is focused — the sensor, inside the camera, is.

Read the whole story
Belfong
10 days ago
reply
For the paranoid among you, here's something assuring...
malaysia
Share this story
Delete
Next Page of Stories