NVIDIA - Enabler of the Impossible

[quote=“Pohjolan_Eka, post:1301, topic:26423”]
The current situation regarding the efficiency of AI usage is completely absurd.

[/quote]Good challenge! Hopefully, I’m not making a terrible misinterpretation (read: strawmanning) if I break down your view into the following core premises, which I am going to go through:

  1. AI is being sold below cost, with investors subsidizing the difference
  2. Most don’t need a frontier model, the majority get by with cheap models
  3. Local execution is practically free, shifting there will collapse demand
  4. The highest margins in world history from artificial scarcity

Before breaking this down, it must be noted that the focus of the premises relies exclusively on inference and completely forgets model training. In terms of workload shares, it’s a justified delimitation — inference is already 55–67% of compute. But that excludes the third that grows 4–5x annually (source: Epoch AI), and at the same time a large part of Nvidia’s competitive advantage: collective operations, inter-chip bandwidth, and the full training stack, which inference-designed hardware lacks or where local hardware bandwidth simply isn’t enough.

The second thing to note is that implicitly, the premises exude a pre-2026 era where zero-shot+reasoning are still the token-output drivers. E.g., AAII v.4.1.1 itself describes its metric set as “a transition to agentic workloads,” and over half of it is first-generation straightforward question-answer testing. The fact that over 40% of the index is still question-answer testing is not a neutral observation — it skews the result in favor of cheap models. Thus, we are talking a bit about apples and oranges at the same time, when meme image and text generation gets mixed up with the fact that demanding expert work is being feverishly shifted to agentic systems requiring wide bandwidth and massive context.

Then, onto analyzing the premises.

1. AI is being sold below cost, investors subsidize the difference

This is a bearing premise, and in my view, it doesn’t hold true, even though it’s a persistent claim in investor circles. Let’s break it down:

8×H100 node, market rental with all costs: $8–15 / hour
Throughput magnitude with continuous batching: 1,500–4,000 output tokens/s
→ 5.4–14.4 million tokens per hour
→ cost: $0.83–2.78 / million output tokens

With Blackwell, the same workload costs about one-seventh compared to the H100 (source: Inworld 4/2026).

Let’s compare to prices: Sonnet class $15/M output, and commercial serving of open weights (DeepSeek R1) $2.19/M (i.e., the price at which independent players sell apparently without loss).

API inference is thus comfortably gross margin positive, roughly 60–99% (depending on inference hardware). And that is independently verifiable: if frontier prices were below cost, third-party providers couldn’t sell a cheaper model at $2.19.

Where do the losses go then? To three places, none of which are token subsidization:

  1. Training runs: training a frontier model costs hundreds of millions or billions, as a one-off expense
  2. R&D personnel: thousands of people with top salaries
  3. Free tiers and fixed-price subscriptions for retail consumers — these can be loss-making, especially for heavy users of reasoning models. But the consumer segment and “generating cat memes” is hardly what carries the AI thesis anyway.

The difference is crucial. The subsidy is in training and product development, not in tokens. If OpenAI/Anthropic stopped training new models tomorrow, based on the data they would be immediately profitable — we will probably get final validation of this soon via the SEC-confirmed IPO prospectus. And that is a completely different economic structure than “selling below cost.”

And it breaks the chain of premises: “generating a meme image” thus doesn’t increase OpenAI’s loss — it decreases it, because it comes with a positive, even +90% margin. Marginal usage is not subsidized.

2. Most don’t need a frontier model, the majority get by with cheap models

This is certainly largely true. Someone smarter than me has said that companies don’t need Einstein in the finance department to process invoices and will certainly manage with a 100 IQ model. And this is already real life today. Model routing, cheap tiers, Haiku/Flash class models, open weights — all exist and are in widespread use. Companies are actively optimizing this because it’s direct cost savings. The mindless tokenmaxxxing boom is over — probably everyone admits that already.

However, a decrease in demand does not follow from this premise, and the reason is Jevons’ paradox. When unit costs plummet, total consumption does not decrease but rises.

Concretely:

  • Cost per token: −300×
  • Tokens per task (reasoning models, agents, tool calls): +100…1000×
  • Cost per task: decreased, but much less
  • Total consumption: rose sharply

Reasoning models are the purest example of this. Same model, same price per token, but it “thinks” for minutes before a single response. The efficiency improvement immediately turned into buying quality, not cutting costs.

The transition to running agentic workloads, which has been ongoing for roughly 12 months, is also a manifestation of Jevons — they consume hundreds or thousands of times more tokens than zero-shot queries. By nature, AI is a high-price-elasticity, scalable, broad-based factor-of-production commodity. And in history, there has not yet been a single commodity adapting to these attributes that has not behaved as predicted by Jevons’ paradox.

3. Local execution is practically free, shifting there will collapse demand

There are two errors here, one of which is physics and the other economics.Physics: bandwidth binds, not compute power. Autoregressive decoding (to put it roughly) reads the entire model from memory for each token. Therefore, speed ≈ memory bandwidth / model size:

Device Memory Bandwidth
Typical Laptop 100–200 GB/s
MacBook Pro M4 Max ~550 GB/s
RTX 5090 ~1.8 TB/s
H200 4.8 TB/s
B200 ~8 TB/s

The difference is 10–50x, and it cannot be fixed with software.

Economics: “Free” confuses marginal and average cost. You are right that the marginal cost is electricity, because the hardware has already been purchased. But that obscures what makes datacenter inference cheap: batching.

A datacenter GPU serves dozens of users simultaneously with a single model read. Your device serves one and sits idle 95% of the time. Throughput per unit of hardware is 20–40x higher with batching (my own estimate based on the use of CSC supercomputers). That is the reason why centralized inference wins.

And the decisive point: agent workloads are the least suited for local execution. An agent runs for minutes or hours, in parallel, with a large context, while you are not at your computer. By definition, it is a server workload.

Local wins where it genuinely wins: small models, latency-sensitive use cases, privacy, always-on functions, offline. That is a real and growing segment—but it is a different (and vastly smaller in market share) workload than the one driving NVIDIA’s revenue and the broader AI investment boom. Nor are these mutually exclusive, at least in my view, except perhaps in some niche segment of technically savvy consumer-developers. In the enterprise segment, the opportunity cost of idling and all the extra local tinkering is quite high.

4. The highest profit margins in world history from artificial scarcity

This can be easily measured:

Fiscal Period Gross Margin
FY2017 58.8%
FY2018 59.9%
FY2020 62.0%
FY2021 62.3%
FY2022 64.9%
Future Guidance ~71-75%

So before the AI boom, the average gross margin was 61.6%. Now it is 10–13 percentage points higher.

That is the right order of magnitude for a “scarcity rent” — not 75% but under 13 points. And a substantial portion of even that is product mix: the shift from gaming GPUs to datacenters and from chips to complete racks increases the margin even without scarcity.

“The highest in world history” should also be put into perspective: ARM makes ~95%, Qualcomm’s licensing is over 70%, Microsoft ~70%. NVDA is certainly exceptional in that it makes 75% while selling physical goods at a volume of hundreds of billions — that is unprecedented. But the benchmark is not zero, but ~62%.

Like @Pohjolan_Eka, I would be somewhat cautious in assessing NVIDIA’s cheapness purely through the earnings multiple and potentially one year of forecasted (albeit reasonably certain) growth. The so-called Molodovsky effect is strong here. It’s worth considering what the market is pricing into NVIDIA and pondering whether you disagree with it. And indeed, right now the market is pricing NVIDIA as a cyclical hardware company and a strong smoothing of the cycle right after CY2028.

The latest earnings report offers a great event study on the market’s movements, which can be used to assess the implied probabilities given by the market based on the guidance anchor and the resulting valuation change. FactSet’s FY29 consensus prior to the release was below $750B. FY28 after the guidance release is $706B. So the company promised roughly the figure for next year that analysts had already booked for two years out. The curve shifted about one year to the left. The market reaction corresponded almost exactly to a one-year acceleration. If all future earnings arrive one year earlier, their present value increases by roughly the one-year cost of capital — in NVDA’s scale, about 9%. The actual move was 8.7%. In plain English, the market thus priced the change primarily as timing, not as a change in level.

So if one believes that the information provided in NVDA’s earnings report also implies a relative improvement in TAM and/or market shares, or the longer duration of the secular cycle and the resulting profitability, then now might be the time to play. But this cannot be deduced directly from the earnings multiples. I actually believe that the step-training launched by Lepikkö will continue among both analysts and market participants — the demand already in the order books ($2000B) alone will ensure this, not to mention the new demand brought by next-generation launches.

[quote=“Pohjolan_Eka, post:1300, topic:26423”]
We are quite far from such a world, and in the near future, things will still operate heavily on an application-driven basis, which practically requires local or hybrid execution, since software companies do not have the pricing power to charge an extra ten a month for every application just to run a frontier model on NVIDIA’s oversized and overpriced GPUs.

[/quote]Sure, this HCI revolution won’t happen overnight, of course. But you misunderstood my point a little bit. I see the development specifically in such a way that these proprietary LLMs or agents, as a complementary part of platform software user interfaces, will disappear or wither away. And they will be replaced specifically by the UIs of these harness developers, which can link dozens or hundreds of data sources and apps at the same time. Therefore, that extra fee is by no means billed by the software company, but rather by the harness and the company supplying the underlying models (such as Anthropic) as their share for automating company processes instead of disconnected GUI clicking.

40 Likes

Interesting analysis from Eka and RoopeK! However, one point in Roope’s breakdown caught my eye:

Keep in mind that major players also offer AI for free. How much does the inference for the free version of ChatGPT cost, for example? The profitability of paid subscription inference alone does not mean the entire inference business is profitable.

I admit I don’t have hard numbers on this (if they are even public), so I’m just going with my gut feeling.

5 Likes

Many thanks for the exceptionally comprehensive reply and high-quality discussion! I think we might have quite a strong difference of opinion on the subject, but at least part of the reason my previous message didn’t get across well is likely that I didn’t lay out my arguments thoroughly enough. I’ll use the same format you chose so that the response stays somewhat coherent.

OpenAI’s difference from a conventional cloud vendor lies in its own models and solutions, so the training and development costs associated with them are at least currently mandatory and fixed, and cannot be neatly separated from variable costs. Your calculations make sense, but if OpenAI were to cut out these other expenses, they wouldn’t have much of a business left soon, because their competitiveness for now stems precisely from the fact that investments have temporarily put them ahead of competitors. The lifecycles of models and solutions have proven to be very short, meaning, for example, a new language model has less than 3 months to cover all fixed costs related to the model, plus those variable costs on top. The situation may change in the near future, but at least right now it is very difficult to see how investors are not subsidizing users. Hopefully, we will get more detailed figures from potential IPOs.

[quote=“Roope_K, post:1306, topic:26423”]
2.Most people don’t need a frontier model, the majority get by with cheap models
…
But a decrease in demand does not follow from this premise, and the reason is Jevons’ paradox. When the unit cost collapses, total consumption does not decrease, it increases.

Concretely:

  • Cost per token: −300×
  • Tokens per task (reasoning models, agents, tool calls): +100…1000×
  • Cost per task: decreased, but much less
  • Total consumption: increased sharply

Reasoning models are the purest example of this. Same model, same price per token, but it “thinks” for minutes before a single response. The efficiency gain immediately turned into buying quality, not cutting costs.

The transition to agentic workloads that has been ongoing for roughly 12 months is also a manifestation of Jevons — they consume hundreds or thousands of times more tokens than zero-shot queries. By nature, artificial intelligence is a high price-elasticity, scalable, broad-based factor of production. And in history, there has not yet been a single commodity fitting these attributes that hasn’t behaved as predicted by Jevons’ paradox.
[/quote]Jevons’ paradox regarding the adoption of steam engines is a popular yet not universal, or even necessarily correct, mental model for this phenomenon. From the same era in agriculture, food production efficiency per acre increased, but this led to a collapse in the number of pairs of hands used in agriculture, forcing workers to seek new jobs in industry and the service economy.

In recent years, the growth of tokens used by models even for simple tasks has exploded, as we previously had a clear bottleneck in model intelligence, and AI hype narratives (agipoikien) and model marketing have long focused largely on maximizing intelligence in various tests. Extrapolating this trend leads to very strange conclusions:

For example, if this year I were to ask an AI agent to turn off the lights in my apartment, it might use, say, 50,000 tokens. The next model released would use 100,000 tokens. The one after that, 200,000 tokens, and so on. In other words, the model would use increasingly more tokens per task, and intelligence per token used would drop, so Nvidia would be laughing all the way to the bank as computing power requirements grow exponentially year after year. Does this sound intuitively like a sensible trajectory?

The limits of increasing the tokens used for a task also seem to be approaching anyway. With frontier models, it is no longer even worth using the highest-effort settings, because the results improve only nominally or can even drop, despite the amount of tokens used multiplying. In some tasks, thinking models or new models are worse than flash models or older models, so there is still a vast amount of slack to be picked up in model specialization, choosing the right scale size, and optimizing efficiency instead of benchmark maxing.

I would therefore consider it extremely likely that the trend will reverse and in the near future AI agents will be capable of performing the exact same tasks as current ones, but with significantly lower computing power requirements and faster than before without accuracy suffering. Usage volume will undoubtedly grow exponentially, but computing power requirements may drop even faster.

[quote=“Roope_K, post:1306, topic:26423”]
3. Local execution is practically free, shifting there collapses demand
…
Physics: Bandwidth binds, not compute power. Autoregressive decoding reads (to put it a bit caricatured) the entire model from memory for every single token. Therefore, speed ≈ memory bandwidth / model size:

Device Memory Bandwidth
Laptop, typical 100–200 GB/s
MacBook Pro M4 Max ~550 GB/s
RTX 5090 ~1.8 TB/s
H200 4.8 TB/s
B200 ~8 TB/s

The difference is 10–50x, and it won’t be fixed by software.

Economics: “Free” confuses marginal and average cost. You are right that the marginal cost is electricity, because the hardware has already been bought. But that masks what makes datacenter inference cheap: batching.

A datacenter GPU serves dozens of users simultaneously with the same model read pass. Your device serves one and sits idle 95% of the time. Throughput per unit of hardware is 20–40x higher with batching (my estimate based on the use of CSC supercomputers). That is the reason why centralized inference wins.

And the crucial point: agent workload is the worst suited for local execution. An agent runs for minutes or hours, in parallel, with a large context, while you are away from the computer. By definition, it is a server workload.
[/quote]In my opinion, people are severely underestimating here just how much can be achieved even with limited memory transfer speeds when scaling down to small models. A quick back-of-the-envelope calculation for decoding:

Batch-1; 27B model; 16 GB; 200 GB/s: 200 / 16 = 12.5 Token/s max :face_with_diagonal_mouth:
Batch-1; 1B model; 0.5 GB; 200 GB/s: 200 / 0.5 = 400 Token/s max :star_struck:

These are genuinely usable figures even on a totally pathetic off-the-shelf PC, although this calculation naturally doesn’t take prompt processing into account. Of course, nothing prevents us from allocating, say, an extra 4 GB of RAM as cache on top of this and slapping 50 GB of curated n-grams onto an SSD, which would squeeze even more performance out of that small model.

The potential use cases for agents are so vast that requirements, intended uses, and available resources will determine whether it makes sense to run them locally or in the cloud. For example, if you want an agent to create a new creative marketing campaign for you overnight that doesn’t copy your competitors, then cloud execution is obviously the number one solution right now. However, for daemon-style agents, local execution is and will continue to be king in my view. I believe that within a year or two, agents will already be significantly baked into the operating systems of both computers and phones.

‘Thin client’ solutions never truly went mainstream, which is why users already have a massive amount of mandatory computing power, storage, and memory that is currently completely going to waste as far as AI is concerned. It is increasingly difficult to sell PCs with just CPU or GPU upgrades, so hardware manufacturers have massive incentives to boost their margins by focusing on improving features related to running AI.

According to Bloomberg, Apple already scrapped its M6 design—which was only announced last week—because the M7 launching next year is significantly better at local AI. Therefore, the company feels it isn’t worth investing any further in the brand-new, high-performance M6, as the priority right now is to catch up with Nvidia in AI capabilities. When Apple goes ‘all-in’ on local AI, it’s hard to imagine the Windows world won’t follow suit. As a complete side note, I think Apple is currently by far the most egregiously undervalued AI investment among the big tech companies.

The mere fact that you are comparing a hardware manufacturer to software companies and licensing models says everything, in my opinion, about Nvidia’s margins. Twist it however you want, paying suppliers, say, €9,000 for hardware and selling it packaged to a customer for, say, €50,000 — without it being some kind of investment engine for a small and highly specialized market — is completely outrageous.

It’s worth remembering regarding gross margin that it is non-linear:

At a 20% gross margin, if you double your selling price, your margin increases by 40 percentage points.
At a 98% gross margin, if you double your selling price, your margin increases by 1 percentage point.

The closer you get to a 100% margin, the harder it is to gain percentage points and the more insane the price has to become. In an oversupply situation following the investment boom driven by the supercycle, I would consider a 50% margin much more likely for Nvidia than historical margins. Of course, this is still heavily speculative at this point, because the current situation is so exceptional that history provides almost no backing for predicting the future.

[quote=“Roope_K, post:1306, topic:26423”]
Yeah, of course this HCI revolution isn’t going to happen overnight. But you misunderstood my point a little bit. I see the evolution specifically in such a way that these proprietary LLMs or agents as a complementary part of platform software user interfaces will disappear/deteriorate. And they will be replaced specifically by the user interfaces of these harness developers, which allow linking dozens or hundreds of data sources and apps at the same time. Thus, that extra fee is by no means billed by the software company, but rather by the harness and the company supplying the underlying models (such as Anthropic) as a share of getting the company’s processes automated instead of disconnected GUI clicking.
[/quote]I thought I had understood the trajectory you presented correctly, but in my experience, it takes decades for these changes to spread. Therefore, I’d consider it most likely that attempts will be made to quickly graft integration projects onto current operating models as cost-effectively as possible, emphasizing immediate benefits without massive overhauls or changes in ways of working. I don’t necessarily disagree with your description other than the fact that I don’t see massive billing potential there, nor room for more than a few large general agents and companies per user in the current environment. We will see competition regarding this from Microsoft and Apple, because the OS is an extremely natural and competitive place to deploy general agents that utilize applications and databases—especially if they can be run mostly locally and sell the necessary extra compute in the same way storage extensions are sold on OneDrive/iCloud.

24 Likes

Thanks as well from my end! That excellent thin client analogy of yours got me thinking that perhaps our difference in views isn’t necessarily all that massive. One could argue that relative to software requirements over time, things have nonetheless moved in a “thinner client” direction. A pretty large share of modern application business logic and data has, after all, shifted from the client to cloud servers, and even though client hardware has gotten more robust, you wouldn’t use it to run the same amount of application logic anymore compared to how it was still run in the 90s and 00s.

I started thinking that perhaps the future of AI could be analogous to this (but with the direction reversed due to the starting point), meaning there could be a grain of truth in both of our views. In that case, our difference in perspective would, as I see it, largely lie in which of the two the value will continue to concentrate on to a greater degree going forward: the capability tail or the volume of routine work.

To sum it up, I am of the opinion that it will concentrate on the capability tail, particularly because the convexity of value—and thus the willingness to pay—is not even close to its peak. Success in an n-step task = p^n. For a 100-step task, 0.99 → 36.6%, but 0.999 → 90.5%. For 500 steps, that same one-percentage-point difference is 92-fold. When you combine this with the lengthening of tasks (METR: historically doubling \~ every 7 months, according to more recent measurements even faster), the end result is that as n grows, the convexity steepens, and the most reliable model is the largest model and its harness, which already improves its capability by another 15-20% in light of various measurements. This view also requires Jevons’ paradox to hold.

You, on the other hand, seem to be of the opinion—to summarize—that value will continue to concentrate rather on the volume of routine work, based on the Pareto frontier: intelligence per token improves faster than tokens per task grow (in other words, the marginal utility of token growth turns/has turned limited - the lightbulb replacement analogy). You had excellent arguments for this, and there is no need to repeat them here. Your view, in turn, requires that “good enough” covers the vast majority of economically valuable tasks—meaning that the convexity of reliability saturates well below the frontier—and that Jevons operates here more according to the laws of agriculture rather than electricity or coal.

That is how I personally see the most central difference in our conclusions. As I see it, both of us nevertheless concede that local vs. Nvidia racks in a data center are not by any means a binary setup :slight_smile: .

PS. @Pohjolan_Eka, you are doing an excellent job at least of breaking my biases. Similarly, my analyst agent distilled two excellent first principles kill criteria for the Nvidia thesis to monitor based on your challenge, which otherwise would have gone unmonitored. Thank you!

19 Likes

To add to this brilliant discussion, I’ll throw in a fresh Damodaran video fitting the topic. In his familiarly humble style, Damodaran ponders the “TAM” (Total Addressable Market) of AI.

This is one high-level perspective to approach the potential of the AI economy.

AI can be seen as an efficiency-enhancing technology where the use of language models replaces human labor. One benchmark could therefore be the expenses of all listed companies in the world, which amount to roughly 65 trillion dollars.

However, even in the wildest visions, language models won’t (at least yet) replace oil and other raw materials, so perhaps those can be excluded from the TAM.

Labor costs are also a large slice, but actually the most mouth-watering piece of it is in the United States, where the total payroll is the world’s largest at 13 trillion dollars. If the use of AI were to replace the entire U.S. workforce from cleaners to CEOs, the savings for companies could be 13 trillion dollars. Perhaps such potential would indeed be a very good justification for the current trillion-dollar investments.

Of course, nearly four years after the release of ChatGPT, there are barely any signs of rapid workforce “efficiency enhancement” (= read: layoffs) due to language models. Not even developers have been heavily in the line of fire, even though one might intuitively think otherwise.

Furthermore, this analysis indeed only looks at things through the lens of efficiency: it could be that in addition to efficiency, AI also creates a lot of something entirely new.

I have read a part of the transcript from the latest NVIDIA earnings release a couple of times in bewilderment. The theme is certainly familiar to NVIDIA shareholders. Jensen predicts that in addition to NVIDIA’s 40,000 employees, the company will soon have 400,000—no, wait—4 million agents buzzing around and burning tokens in the name of endless improvement and efficiency. Effectively, sales agents, customer service agents, agents managing more agents, and so on.

In basic economics courses (thankfully I only took those and nothing more :D), someone always asks the professor how economic growth can possibly be limitless in a physically constrained world, and the answer is that raw materials might run out, but human ideas won’t.

When we talk about these agents and their future token consumption—which is why endless investments must be poured into compute or “intelligence”—it reminds me of a perpetual motion machine. :smiley: The ultimate economic meta-idea: create zillions of agents to busy themselves with one another, and the economy runs and grows as if by magic.

Ps. It’s funny how Jensen has for quite a while now completely bypassed the AGI discussion. Back in 2022–24 (?), AGI was the theme used in the media to justify massive investments, but the agent narrative is perhaps more concrete and less damaging to AI companies’ brands. After all, AGI is immediately associated with Terminators and other dystopian tropes. :smiley:

27 Likes

Currently, Nvidia looks interesting again.

https://www.barrons.com/articles/cathie-wood-ark-invest-nvidia-stock-buy-1d5f3ff6?st=qqerni&reflink=article_copyURL_share

6 Likes

I guess this hasn’t been linked here yet, interesting to see what kind of value chain will form around Nvidia’s investments in the entire AI infrastructure :grinning_face:

2 Likes

[quote="Verneri_Pulkkinen, post:1310, topic:26423"]\nIn addition, this analysis indeed looks at the matter solely through the lens of efficiency: it could be that, alongside efficiency, AI will also create a lot of something entirely new.\n\n[/quote]\n\nHit the nail on the head! In the case of Nvidia, one simply has to believe that artificial intelligence is a seamless continuation of humanity’s most disruptive technological inventions: the harnessing of fire, the wheel, steam power, electricity, semiconductor computing, and telecommunications.\n\nAt least I refuse to accept Damodaran’s completely static view of the economy. It indeed assumes that companies could no longer or would no longer want to grow by expanding sales and marketing, that product development had already exhausted all innovations, or that digital services would no longer need new features. The idea that we are even close to some kind of cognitive saturation point – where the only added value would stem from a “race to the bottom” style of efficiency in current operations – is completely foreign to me.\n\nThat is why I firmly believe that the value of AI computing will continue to be concentrated in the tail of capability rather than the volume of routine work. And that is why I also believe in increasingly intelligent (and, as scaling laws continue to operate, larger) models, as well as entirely new ways to get more out of AI - developments where agents serving mostly software development needs, for example, are just scratching the surface. I believe in breakthroughs in entirely new architectural paradigms that the world’s leading AI minds are currently working on. And I believe in the insatiable desire to push the boundaries of what is possible, which may require new groundbreaking technological leaps - such as quantum operations.\n\nNvidia and Huang also believe in all of this. Likewise, Frontier Labs believes in this by investing billions in AI development, instead of high-margin inference and inflating their cash reserves, which makes their own products from 6 months ago look practically ridiculous - regardless of the fact that the AGI/ASI narrative has subsided.\n\n[quote="Verneri_Pulkkinen, post:1310, topic:26423"]\nIn basic economics courses (thankfully I only took those and nothing more :smiley: ) someone always asks the professor how economic growth can possibly be limitless in a physically constrained world, and the answer is that raw materials might run out, but human ideas won’t.\n\n[/quote]\n\nAs a former economist of sorts, I must admit that the profession has never had particularly good tools for modeling economic growth. Economics has been especially helpless in modeling the aforementioned major technological turning points of humanity (see, e.g., Solow’s productivity paradox). Frameworks are built around the dynamics of the given moment and particularly rigid scarcity constraints of the physical world. They are mathematically incapable of handling a situation where a new cognitive commodity (such as an AI model or AI agent) flips the entire production function. And then we end up relying on easy compromise heuristics, such as the assumption that a company’s long-term growth (terminal value) cannot exceed some constant (often 2%) GDP growth, which fail totally at step-change points.\n\n[quote="Verneri_Pulkkinen, post:1310, topic:26423"]\nP.S. It’s funny how Jensen has completely bypassed the AGI discussion for a while now. In 2022–24 (?), AGI was the theme used in the media to justify massive investments, but the agent narrative is perhaps more concrete and not as harmful to AI companies’ brands. After all, AGI is immediately associated with Terminators and other dystopian imagery. :smiley:\n\n[/quote]\n\nOne reason is likely this: Huang has said that AGI has already been achieved. And he has quite good reasons for this. What the AI community, measured by benchmarks, considered AGI even 2-3 years ago (let alone before 11/2022) has already been surpassed many times over, and the goalposts have had to be creatively moved forward, which is perhaps also evidence against Damodaran’s assumption of a cognitive saturation point.\n\nhttps://www.youtube.com/watch?v=awE5DV-M48o

16 Likes

But what is the size of the market for that kind of tail?

Damodaran’s calculations are good illustrations in the sense that the big money is currently in “routines.” You really have to believe in a broad paradigm shift for the tail to grow as big as the dog?

In this “This time is different” 15-minute review, I brought up this chart of the long-term growth of what was once the world’s most advanced economy.

Despite genuinely disruptive inventions, such as steam power, various grades of steel, and mass production, the industrial and infrastructural revolutions enabled by them (the CANALS, whose significance modern people forget :smiley: as well as railroads, of course), logistics, chemical inventions in agriculture, and even the fact that Great Britain was an empire where “the sun never set,” economic growth was some darned 2% per annum… :smiley:

Why would it be different this time?

These are matters of taste: some think the last significant invention was the refrigerator, others the indoor toilet, and a third the internet. It is true that their ultimate impact is difficult to measure.

Time will tell what probabilistic language models built on the material of the world’s Reddit discussions and Wikipedias can ultimately achieve. I don’t believe we have the exact ability to predict this development, let alone how and with what productivity it will be applied in real life.

21 Likes

[quote="Verneri_Pulkkinen, post:1314, topic:26423"]\nHuolimatta aidosti mullistavista keksinnöistä, kuten höyryvoima, eri teräksen laadut ja massatuotanto, näiden mahdollistamat vallankumoukset teollisuudessa ja infrastruktuurissa (ne KANAALIT, jotka unohtuvat merkitykseltään nykyihmisiltä :smiley: sekä tietysti rautatiet ), logistiikassa, kemian keksinnöt maataloudessa ja vielä sekin seikka että Iso-Britannia oli imperiumi jossa “aurinko ei laskenut”, talouskasvu oli sen joku pahuksen 2 % per annum… :smiley:\n\nMiksi nyt olisi toisin?\n\n[/quote]\n\nLet’s do a little exercise using Damodaran’s framework: let’s imagine ourselves in 1995 and calculate the market boundaries of the internet and software companies using the same static framework, and we will see the problematic nature of this way of thinking.\n\n(The exercise example is narrow and the numbers are directional, but for the final outcome, this is irrelevant.)\n\nIf an investor had assessed the potential of the internet and software at that time using the current logic of “replacing existing work and costs,” the calculation (relative to the market at the time) might have looked something like this:\n\n1. Postal and courier services (approx. $55 billion): Email streamlines corporate correspondence and saves costs.\n2. Print advertising (approx. $45 billion): The web captures Yellow Pages and similar classifieds.\n3. Travel agencies (approx. $15 billion): Automating agency booking work.\n4. Physical distribution of software (approx. $30 billion): Saves on CD logistics and retail margins.\n\nCombined, the TAM for the internet and software in this exercise’s static model would have been roughly under 150 billion dollars, if the investor had measured the new technology strictly by the value of the old routine work it replaced. And they would have been monumentally wrong in their estimate – as we all know 30 years later.\n\nThis static model was unable to predict what happens when the marginal cost of distribution and communication drops to zero. The markets didn’t just save on costs, but demand exploded and entirely ecosystems were born whose existence was previously technically and economically impossible. When we look at the actual figures today, the gap between the static estimate and the dynamic reality is staggering:\n\n* The internet didn’t just make Yellow Pages cheaper; it created a digital attention economy. The global advertising market for search engines and social media didn’t stay at 55 billion, but is today an industry worth over $650 billion.\n\n* It didn’t just save IT department or CD distribution costs, but birthed cloud computing (AWS, Azure) and software subscription models (SaaS). As a result, the enterprise software and cloud infrastructure market has ballooned to a combined total of over $550 billion.\n\n* It enabled smartphone app ecosystems – a completely new software layer worth over $170 billion that not a single 1995 GDP model could anticipate because it had no historical counterpart.\n\nIn the case of AI, I believe the exact same mistake is being made. It is assumed that the demand for development and cognitive capacity is constant or saturated at some final equilibrium level. But Jevons’s paradox is real: when the marginal cost of logical reasoning and problem-solving plummets, we do not settle for the current amount of software at slightly lower costs. We start solving data-intensive problems, building millions of hyper-personalized automations, and accelerating R&D cycles in ways we previously couldn’t afford to allocate resources to. AI decouples innovation from human bottlenecks.\n\nAs to what the size of the capability tail market is, I cannot answer. A first principles thinking framework quickly leads to it being nearly limitless – an outcome that one then has to temper with real-world constraints :slight_smile: .\n\nWhy doesn’t this massive value creation then inevitably blow GDP past the historical 2 percent growth rate – that compromise heuristic used by economists and financial scientists? Because GDP is an industrial-age metric that only understands cash flows and physical scarcity. GDP does not take a stance (directly) on the redistribution of value, the structure of profit margins, and the steering of free cash flow – which is largely what investing is all about. When a company moves to AWS, it builds a lot of new and groundbreaking things on the platform, but at the same time it stops buying physical servers from IBM, HP, or Dell, lays off its own administrators, and lowers its facility electricity bill. Money moves from old-world infra and real estate players to the cloud giants. GDP may not grow at all (it can even drop if the cloud is cheaper). This is so-called creative destruction, which economists and financial scientists also fail to model.\n\nI always find it a bit funny that AI bulls are pointed with the question why would it be different now, when historical evidence would rather require those more skeptical of AI to answer this question – as the exercise example presented above shows.\n\n[quote="Verneri_Pulkkinen, post:1314, topic:26423"]\nPs. An investor can also simultaneously believe fully in the AI hype while still being skeptical about AI stocks. After all, big numbers and visions do not nearly always turn into good profitability. The game is still in its early innings on this front as well.\n\n[/quote]\n\nIn light of historical data, this is easy to sign off on. If we return to 1995 here as well, every single internet company ran its services on a Sun Microsystems + Oracle stack – practically every single one. Five years from that, the de facto standard was the LAMP stack’s Linux and MySQL. There is an insane stack of similar examples from the past 30 years.\n\nAnd the game is indeed in its early innings. If you think time-analogously, that ChatGPT was the same starting point for AI as Netscape was for the internet, then Google wouldn’t even have been founded yet, Mark Zuckerberg would have just moved to middle school, and Travis Kalanick (founder of Uber) would only be in kindergarten (analogy stolen from Gavin Baker’s latest a16z podcast interview – recommended!).\n\nThat’s why I myself am at least damn paranoid about Nvidia and other AI-related holdings in my portfolio, which are an over-the-top position.

44 Likes

I will answer a few points where data makes it easy to provide an unambiguous agreement or counter-argument:

[quote="Loyly, post:1315, topic:26423"]
Nvidia is currently closer to an infrastructure company than a growing technology company.

[/quote>

In my opinion, Nvidia is factually both. Which one you emphasize in your investment thesis largely determines whether you see Nvidia as undervalued or overvalued.

[quote="Loyly, post:1315, topic:26423"]

  1. One of Nvidia’s most significant risks is the slowing of investments by mega-techs → Nvidia’s network expansion slows down

[/quote>

Of course, I agree that a slowdown in mega-tech capex is a risk (directly and indirectly). But it’s not as big a risk as it perhaps used to be, for two reasons:

  • Nvidia went ahead and guided for its customers’ capex growth in the Q2 earnings release, and it is not slowing down anytime soon - quite the opposite :slight_smile: . If anyone outside the hyperscalers has visibility into how much money is being put into capex, it’s obviously Nvidia, which is on the receiving end of these funds and whose order books already reflect this capex with production sold out through the end of fiscal year '27 as of Q2/26.
  • At the current rate, the ACIE segment (i.e., non-hyperscalers) will account for over half of Nvidia’s revenue within 6 months. IO Fund’s Beth Kindig wrote a perceptive analysis about this, which is linked slightly higher up in the thread.

[quote="Loyly, post:1315, topic:26423"]
2. The company grows as energy consumption grows, not faster than that

[/quote>

I would say it the other way around: energy consumption grows as the company grows. Conversely, the company cannot grow faster than the energy supply (at least not for long). But even this can be roughly calculated from the figures provided by Nvidia: Nvidia will sell $700B next year, the price per GW (Rubin) is $40B. This equals a total of 17GW sold, which is nowhere near the 33GW capacity slated for completion next year alone, which, according to Goldman Sachs, also secures electricity availability (transmission is a bit of a question mark - Global Data Centre Capacity to Reach 93 GW by 2027 | InfraTechSolutions Co., Ltd. Julkaistu aiheesta | LinkedIn)

[quote="Loyly, post:1315, topic:26423"]
The “Temu risk” could materialize. When AI computing doesn’t require the best, cheap Chinese goods are bought at a fraction of Nvidia’s prices. SpaceX with its space computing is a wild card as a more distant commoditization risk.

[/quote>

I honestly don’t believe this. Why wouldn’t AI computing require the best? And when combined with the aforementioned energy discussion—which may not be a bottleneck, but acts as a significant component for the profitability of investments due to the law of supply and demand (demand drives up prices)—you want to get maximum performance per megawatt out of the hardware. With Temu hardware, investments don’t make sense—with Nvidia hardware, they do.

On the other hand, it is possible and makes sense to send slightly weaker systems into space if needed, because interconnectivity is a challenge there, which in turn is one of Nvidia’s greatest competitive advantages.

[quote="Loyly, post:1315, topic:26423"]
Currently, the pricing of many smaller players in the semiconductor/AI infrastructure sector, such as Amkor, Fabrinet, and Marvell, already seems to anticipate a turn in the cycle. Liquidity first withdraws from the fringes of the market.

[/quote>

As I wrote above, event analysis of Nvidia’s earnings release indicates the exact same thing. The market essentially only discounted revenue recognition one year ahead—not its total growth. At the same time, we must remember that if 2027 is already sold out and demand could potentially be double that, it means the cycle will very likely last longer than what the market is pricing in.

[quote="Loyly, post:1315, topic:26423"]
If and when the focus of AI models shifts heavily from the rapid training of models to their inference/usage, demand for Nvidia’s capacity will likely decrease. A rising interest rate environment, in turn, affects data center investments, which ultimately shows up on Nvidia’s balance sheet as well.

[/quote>

In my opinion, the important word here is relative to the entire market—i.e., “demand for Nvidia’s capacity will decrease relative to the entire market” (and increasingly so, because from a monopoly position, there is only one way to go :slight_smile: ). It is undeniable that the previous 3:1 training-to-inference ratio has shifted more toward a 1:3 ratio. But this hasn’t shown up anywhere yet. I considered this myself to be perhaps the biggest risk, but instead, Nvidia has only accelerated its growth and maintained its market share and margins (whose deterioration is usually the first sign of growing competition). Therefore, not a single data point indicates fierce competition for Nvidia, despite the shift in focus toward inference.

19 Likes

An extremely interesting post above by Roope.

I personally agree with the internet analogy, but at the same time, I think we are currently dealing with a clear bubble where the realization of bullwhip inventory risks and the resulting rapid collapse in demand is even an inevitable outcome.

However, over any longer time horizon, the change will be revolutionary. You can already see that even with lighter use of AI.

1 Like

Thank you! And yes – a couple of messages above I listed these fundamental technological innovations of humankind, and practically every one of them in modern times has caused some degree of a bubble. This classic of academic literature covers them all :smiley:

Somewhat ironically, however, those parts of the supply chain (chip production, memory, etc.) that have caused valuations to swell are actually curbing the formation of a bubble. These historical bubbles have burst when supply exceeds demand, whereas right now the situation is precisely the opposite (at least a few steps downstream in the supply chain from Nvidia). TSMC has been the gatekeeper against the formation of a bubble and has practically refused to increase its supply capacity. The situation may change if/when Intel, Samsung, and eventually Musk’s Terrafab start pushing qualified 3-nanometer chips onto the market.

14 Likes

I think this is a very good point. Based on leaks to public sources, Intel has indeed improved its wafer yields drastically. TSMC has always been perceived as the player that manufactures and will manufacture all advanced chips, but now that gap has narrowed, at least if sources are to be believed. However, I would see that TSMC has a slight advantage in expertise, as they are still pushing the absolute limits of physics with so-called low-NA lithography, while Intel has already adopted high-NA equipment. It remains to be seen how the dynamics between the fabs will evolve in the future when you throw the political tensions regarding Taiwan into the mix.

Apologies for the off-topic, great posts by user @Roope_K

11 Likes

[quote="Roope_K, post:1316, topic:26423"]\nCombined, the TAM for the internet and software in the static model of the practice example would have been roughly just under 150 billion dollars, if an investor had measured the new technology by the exact value of the old routine labor it replaced. And they would have been monumentally wrong in their estimate—as we all know 30 years later.\n[/quote]\n\n[quote="Roope_K, post:1316, topic:26423"]\n* The internet didn’t just make Yellow Pages cheaper; it created a digital attention economy. The global advertising market for search engines and social media didn’t stop at 55 billion, but is today an industry worth over 650 billion dollars.\n\n* It didn’t just save IT department or CD distribution costs, but gave birth to cloud computing (AWS, Azure) and software subscription models (SaaS). As a result, the enterprise software and cloud infrastructure market has ballooned to a combined total of over 550 billion dollars.\n\n* It enabled smartphone app ecosystems—a completely new software layer worth over 170 billion dollars that not a single 1995 GDP model could anticipate because it had no historical equivalent.\n[/quote]\n\nWith NVIDIA (and other AI companies), it makes sense to acknowledge that there is some kind of unpredictable growth component, which may or may not create value. Much like analysts give Tesla price targets to justify imaginative and generous assumptions for the company’s future, unpredictable trajectory. It is realistic to expect that within 10 years we will see new business models, which, however, will still be based on sustainable business principles, such as solving customers’ problems in real everyday life. We will certainly see a great deal of setups that shift and reshape like dunes in the Sahara. And many setups exploiting investors’ enthusiasm that, in a vulgar saying, will vanish like… gas in the Sahara (note: a play on the Finnish idiom “katoavat kuin tuhka tuuleen” / vanish like ashes in the wind).\n\nIt’s just really hard to argue about the appeal of such assumptions; it almost borders on matters of faith.\n\nStill, I consider Damodaran’s illustration a good starting point for a “realistic” outline of TAM in light of the information we have now. Your internet example is good, but the numbers might seem laughably small to a modern reader. It’s worth remembering that 30 years is a long time in economics. For example, at an 8.5% annual growth rate, that 55 billion dollar marketing budget has ballooned to 650 billion dollars.\n\nOf course, it’s not the absolute truth, but if we assume the TAM to be, say, all American payroll expenses—13 trillion—I think the TAM is quite generous. :D\n\nTAM is a good example also in the sense that we would have (and many did) bet on completely the wrong horses as winners. Most of the names have already faded into oblivion; mostly pets.com and AOL, etc., still remind us of those times between the covers of history books. Even the winners, only recognized in hindsight, could be bought cheaply later on. Amazon almost went under and its stock crashed by over 90%. :D\n\nIn Amazon’s 2000 annual report (from the spring of 2001), there is this quote from Bezos:\n\n\u003e "Ouch. It’s been a brutal year for many in the capital markets and certainly for Amazon.com shareholders. As of this writing, our shares are down more than 80% from when I wrote you last year."\n\n:D

13 Likes

Spamming the thread a bit more, and then I’ll leave the Nvidia folks alone for a while. :smiley:

I assume not many people open up the company’s financial statements, but I enjoy financial masochism and like to open these up. :smiley: And of course, it’s nice to look at large, growing numbers—you don’t see those in just any company’s reports. :smiley:

NVIDIA (like many other winners) divides opinions, and on X, for example, you’ve probably been able to run into “NVIDIA is a scam this and a scam that” for at least the past 10 years. However, the company’s accounting has been approved year after year, and therefore such shouting is, to my taste, wide-of-the-mark hot air.

However, there are still aspects in the financial statements that make you rub your eyes, and at least my own view of the company’s asset-lightness has changed… to be heavier. The company’s revenue is ballooning at a 70–100% annual pace. As everyone knows, the limits of the physical world are being hit, and the subcontractor machinery is being persuaded to invest more in capacity so NVIDIA could grow even faster.

These persuasions include, among other things, long-term commitments to purchase things like memory, production capacity, cloud capacity (NVIDIA also makes its own models, etc.), investments in AI ecosystem companies (e.g., frontier labs), and most recently, a megalomaniacal guarantee for an OpenAI-related company*. :smiley:

In just one year, these commitments, which do not appear on the balance sheet, have ballooned from 82 billion (note: I left out things like product warranties, which are only a few yards) to 530 billion dollars! :smiley: +550%. Over the same period, the 12-month trailing revenue has risen from 190 to 303 billion dollars, or 60%.

From a strategic standpoint, this makes sense in its own right. NVIDIA now has the opportunity to cement itself as a central player in the brave new AI economy. More compute is more “intelligence” (their words) and perhaps through that, more business opportunities. Since fewer people can afford to pay, NVIDIA is guaranteeing, investing, and recycling money into the ecosystem. By buying up the capacity of subcontractors, it is kept away from the competition.

At the same time, everything is characterized by urgency, because the head start won’t last forever, as has been noted in the thread.

If the visions materialize, this risk will truly pay off. If, on the other hand, the brave new agent economy is delayed by, say, five years… Well, I’m not sure how much in breach of contract fees NVIDIA will end up paying and what will happen further down the AI chain when those assumed hundreds of billions in purchases fail to materialize.

And how long can the company grow these commitments faster than the business itself is growing?

*SB Energy Corp. guarantees – In August 2026, we entered into guarantees, capped at a total of $105 billion, to provide credit support on a land, power, and shell buildout with affiliates of SB Energy Corp. (SB Energy) on behalf of a customer, an affiliate of OpenAI Group PBC (OpenAI), related to leases for approximately 4.25 gigawatts of IT load in the aggregate at SB Energy’s PORTS Technology Campus in Pike County, Ohio.

40 Likes

Wistron, a supplier for Nvidia, experiencing a large financing need may in part indicate the rapid production growth of GB300 and upcoming Vera Rubin systems. This shows that the supply chain is preparing for strong demand.

Key Points

  • Wistron shares fall Tuesday following its $1.5 billion global share sale announcement
  • The Taiwanese tech giant is raising funds for the procurement of raw materials
  • Its shares are up about 23% year to date

https://www.cnbc.com/2026/09/08/nvidia-supplier-wistron-share-sale.html

4 Likes

It’s comforting to think that, regardless of what we predict and think about the future here, almost inevitably, over a long enough timeframe, we will be wrong. Just as it was impossible to foresee the current form of the internet in 1999, there is simply no way to predict all the use cases and impacts of AI right now.

What we can do is try to examine current trends and events in the near future and outline what the world will look like if they continue as they are, while acknowledging that there will inevitably be several black swans along the way. Every year, we should review our previous thoughts and modify them as the market evolves and situations change. It wasn’t all that long ago that agents increased token usage multi-fold, which changed all forecasts, and it is also likely that at some point we will move beyond the current form of the Transformer architecture, at which point the entire AI table will be flipped.

The year 2027 is now shaping up to be the worst bottleneck for hardware, but personally, I’m expecting significant relief by 2028, unless some new token-guzzling monster is invented without the architectural efficiency gains keeping pace. Maybe someone will soon figure out that instead of individual agents and sub-agents, absolutely everyone should deploy a massive number of agents to constantly hold meetings with each other non-deterministically as a federation, with no human in the loop at all, thereby driving token consumption up by another 10x - 100x and sending Nvidia to yet another all-time high :smiley:

In my opinion, the strongest signs of an AI bubble came from the acquisition of Hugging Face. Is this the most expensive meme in the world, excluding Elon Musk’s Twitter antics?


12 Likes

Desert Ant Labs is a great example of what the world could look like without the need for Nvidia. In my opinion, their outputs can no longer even be called Small Language Models; we should probably start talking about Tiny Language Models or Micro Language Models. So we are talking about extremely specialized micromodels that are capable of running locally on almost any device due to their small size:

The industry will spend about $450 billion on data centers this year. Meanwhile, the world ships more than a billion phones, tablets, and laptops with increasingly capable chips, perfectly suited to these kinds of tasks. There’s more compute available in people’s hands than in every AI data center on earth.

We have an unfair advantage with free inference. No per-call cost, so a feature runs on every message instead of the ones you can afford to check. No round-trip, and your customer’s data never leaves the device. When inference costs nothing, the way we build products changes entirely.

  • Voz: transcribe 10 minutes of audio in two seconds on an iPhone – 4.7x faster than Whisper – with a start and end time on every word.
  • Clear: a 9MB model that can turn a five-minute laptop recording into studio quality audio in one second.
  • Redact: mask names, addresses, and card numbers, in real time, in 27 languages, so they never reach your servers.
  • Tongue: identify 84 languages from three words, with a 2MB model.

Using Gigabrain AI models running on Nvidia’s world-class GPUs is very inefficient for most routine tasks performed by users. If the alternative is, for example, an 8 GB package installed on the user’s device hard drive containing, say, 320 various specialized micromodels of 25 MB each, running which is practically free from today to eternity, then I believe a large part of routine inference demand will escape the cloud.

Of course, there are also plenty of tasks where those leading top-tier models are needed, but I can easily imagine a future with a vast number of these extremely specialized micromodels trained, which a locally running router-agent can access. Training for these micromodels is also very inexpensive, so that side does not require additional computing power either.

The Uhm model mentions:

Find and remove every filler word.

An hour-long episode is analyzed in 12s on an iPhone 17 Pro. A back catalog is a batch job you run locally, not a cloud invoice.

Uhm gets there by skipping the transcript step. The usual way to find an “um” is to transcribe the whole file and search the text. A transcript would not help anyway: models like Whisper leave fillers out of their output, so the ums never appear in the text to find.

Uhm reads the waveform instead and marks every filler it hears, so an editor can cut them or a one-click cleanup can tighten the whole take. Open an audio or video file and every filler word is listed with its time, so you click one and land on it.

Apple runs the 45MB Core ML build.

I would be interested to hear, for example, from @Verneri_Pulkkinen, what tools you currently use to edit those videos and audios, and how competitive such micromodel solutions are in terms of speed and accuracy compared to professional tools. How much do you use world-class AI solutions at Inderes for video and audio processing, and is there potential in these for your use cases?

20 Likes

However, this still relies very heavily on the assumption that current human work is augmented with 2024-style chatbots and that advancements in AI will not really change the need to work the way we do now.

Most companies, however, are thinking or starting to think about this in a completely different way. The idea is: what if, instead of augmenting individual tasks, we could reach a point where long processes involving multiple tasks and multiple people are outsourced to AI?

Let’s take an example again and position ourselves as a cybersecurity SOC operator. The tasks include monitoring alerts, doing false positive triage, analyzing log and network data, and escalating these further. Later on, these are perhaps passed on to the dev team as security patches, which are then pushed to production after days and weeks. You could certainly develop purpose-built local models for all of these to assist in the work done in the current fashion.

But how does the situation change when attacks start coming from agent swarms non-stop 24/7? That can no longer be answered by augmented human labor; instead, fighting fire with fire is required—an agentic closed-loop defense system that detects, isolates, remediates, and pushes fixes to production. Such solutions obviously cannot run on individual local models or endpoints either; rather, they are datacenter-scale, centrally managed solutions. The models also need to be reasonably capable (read: gigabrain-level) and must be capable of real-time, cross-modulated situational awareness sharing and second-scale reasoning on a global scale. Anyone can also ponder what such a continuously scanning defense agent system means in terms of compute and token consumption. And this is not sci-fi; it is something that, already after the Hugging Face incident, has been seriously set up.

And it’s not hard to imagine a similar future in pretty much any part of organizations’ current value chains where local simply doesn’t work. Software development teams are already struggling with the fact that agent terminals are often local: work is still tied to office hours, codebases and dependencies are too massive to run locally, and an individual machine lacks the ability to synchronize changes made by other agents in real time, leading to constant code conflicts.

Although I believe—as I have stated before—in the growing capabilities of local models (e.g., on the consumer side), it is really hard to see a future where, in the transition from human assistance to autonomous systems, the center of gravity of computing power could swing local.

Compute costs are a big question—but early signs would rather point toward the direction of best open models + own compute hardware / (neo)cloud rental + open-source orchestration and harness. Many of the Fortune 500 have already reported embarking on this path—since, besides the cost side, this strategy also has many other advantages, such as control and sovereignty related to data and other intellectual property. And Nvidia is not (directly) suffering here either; they are the preferred supplier even in this scenario, where cost-cutting happens by eliminating the middleman markup of Anthropic / OpenAI / Microsoft Copilot, etc. And somehow I think that this very trajectory is also the reason why Huang is now betting so heavily on open models.

PS. I still wanted to add one more thing that also argues in favor of using the most capable model possible and/or an agentic interoperability platform: 70-75% of AI workflow costs still come from human labor (McKinsey)—token consumption is “only” 20-25%. This of course means that a large portion of the TCO can be cut by improving the part that leads to that massive amount of human labor—errors, exceptions, etc. A better AI is of course the solution to this—by no means a cheaper and weaker model.

@Pohjolan_Eka’s angle also implicitly assumes to some extent that the company will absorb all AI costs into its opex in the future as well. I myself, on the other hand, believe that the better quality, faster / 24/7 service, etc., enabled by AI could sometimes be monetized downstream as well.

15 Likes