NVIDIA - Enabler of the Impossible

Apple has managed to quickly close the gap with Nvidia in local execution with today’s Mac Studio release. In terms of speed, the M5 Max is already at a very usable level, and the baseline M5 Ultra, thanks to its larger memory capacity, outperforms the RTX 5090 in many workloads. Memory capacity and memory speed are the most common bottlenecks in local execution.

This isn’t starting any revolution just yet, but it strongly indicates that Apple will be one of the market winners in local model execution. As the trend continues, LLM workloads can increasingly be shifted from the cloud back to the local computer.

25 Likes

Well, now we got a quite interesting signal about a tipping point when OpenAI’s proto-chip beats the GB300 and goes head-to-head with Rubin as well (potentially even better performance, because Rubin’s speculative encoding pollutes the comparison). What’s astonishing here is that this is the first generation and it already looks promising. And on top of that, this isn’t some model-stack optimized solution, but the same type of general-purpose chip with which Nvidia achieved its monopoly.

We don’t need to put ashes on our heads just yet, and a few important notes:

  • Performance tests are self-driven by OpenAI (though verified by SemiAnalysis)
  • Tested not with the best frontier models, but e.g., 1.5-year-old R1s and similar
  • Tests were done only with an 8k1k input/output workload, not with large-context agentic runs

There is still a way to go from engineering samples to functioning mass production. The ramp-up would be gradual in 2027, by which time they will already face Rubin Ultras and soon after that Feynman. So Nvidia naturally still has a slight head start. But since the perf/W is good in these, OpenAI will surely feel tempted at some point to rethink how much they should continue buying from Nvidia.

But for me, this is yet another signal to carefully consider my own overexposure to Nvidia, which I haven’t gotten rid of yet

18 Likes

I dropped into this thread since the discussion seems to be revolving around models.. Jimppa put together a great article on token consumption and the business itself. It relates to Nvidia (among others).\n\nJim Liu X:ssä: ”https://t.co/TGBkmp607N” / X

1 Like

Quite a few use cases where simply avoiding RTX’s power consumption (heat) is a significant advantage, alongside that controlled and traceable process. Especially if running agents autonomously rather than an impatient human sitting in front of the screen prompting. The price is a bit steep for Finnish purchasing power, but I’m slightly tempted to get the Ultra for testing. Probably better to wait for proper benchmark results, though.

Otherwise, I’d argue Apple has actually played this game smartly, even though people are complaining about a lack of innovation. In addition to the hardware, at least in the US, there seem to be some quite interesting leasing options.

6 Likes

I am probably going to buy the 512GB Ultra myself once they become available (apparently available to order in October).

A comparable Nvidia solution (given the thread title as well) would cost more to get even half of that memory (though it would also be faster, of course). It’s an expensive hobby, but that’s life.

Also, since I’m doing this at home, a power consumption of around 2kW combined with the sound of an airplane wouldn’t score any points—at least not with the spouse. For myself, it might bring back some server room nostalgia, but in the long run, that could get tiresome too.

8 Likes

Based on the main features, you’d think the result and guidance are sufficient. Of course, who knows what’s beneath the surface.


And there are no China expectations included in the guidance.

Edit: Ah, so this is where the friction starts.

23 Likes

Shouldn’t this be just temporary if the price increases only take effect two quarters from now?

12 Likes

The recalibrated model hit the nail on the head regarding the earnings and expectations for the upcoming quarter.

These were mind-blowing numbers once again.

And even more mind-blowing is the revenue growth guidance provided for the upcoming year:

”We expect to grow revenue by approximately 70% in fiscal 2028. This is a supply-constrained outlook”

So, if/when revenue reaches 400B this year, next year expects $680B. An absolutely astronomical figure! The consensus estimate was likely still around $550B today.

No wonder the stock price ultimately took quite a massive bounce.

39 Likes

NVIDIA is buying Hugging Face:

https://www.reuters.com/technology/nvidia-talks-acquire-hugging-face-13-billion-deal-business-insider-reports-2026-08-27/

HF is like “GitHub for AI”:

A very scary move from a user perspective, because of course they’ll have to monetize it at some point to get a return on the acquisition.

8 Likes

But Hugging Face is already paid in many places and is generating $150M in revenue growing at a nice, solid 100%+ pace, measured by annual run-rate. On the other hand, there probably isn’t a way to monetize HF that would amount to anything more than a rounding error on Nvidia’s income statement. So I’m not particularly scared.

In my opinion, the strategic significance of a potential HF acquisition for Nvidia lies elsewhere than trying to get as big a slice of this as possible to the bottom line. Nvidia’s hardware sales are served by an Open Source ecosystem that is as robust as possible, the viability of which opens up hardware sales directly to larger companies as the margins of model providers are squeezed out from the middle. After all, Nvidia likely holds a near-monopoly in open model serving.

Similarly, this would allow them to deepen the lock-in between the developer community and Nvidia’s proprietary ecosystem, if desired. It just takes a little subtle steering to ensure that development and model servicing from the HF platform happen to work best with Nvidia’s hardware :slight_smile:.

As I see it, this is the exact same kind of case as when MSFT came knocking on the door of GitHub, which you just mentioned. Microsoft, like Nvidia, has been accused over the years of opposing open source. MSFT played that acquisition masterfully and multiplied the value of GH without monetizing it particularly heavily or turning it into a purely proprietary platform for .NET architecture or anything like that.

I suppose a similar playbook is in use at Nvidia regarding HF as well, and in my opinion, the acquisition would be a very smart move.

16 Likes

You’d think so, let’s revisit this when CUDA is liberated.

2 Likes

Beth Kindig (from IO Fund) summarized the events following the Q2 earnings report much better than I ever could, so I’ll just share the link with a strong recommendation to read it:

Now I didn’t understand that comment? Why on earth would Nvidia ever open-source/release an asset that it has invested 20 years in, and with which it has monopolized a massive portion of the AI computing market for itself?

11 Likes