[quote=“Pohjolan_Eka, post:1301, topic:26423”]
The current situation regarding the efficiency of AI usage is completely absurd.
[/quote]Good challenge! Hopefully, I’m not making a terrible misinterpretation (read: strawmanning) if I break down your view into the following core premises, which I am going to go through:
- AI is being sold below cost, with investors subsidizing the difference
- Most don’t need a frontier model, the majority get by with cheap models
- Local execution is practically free, shifting there will collapse demand
- The highest margins in world history from artificial scarcity
Before breaking this down, it must be noted that the focus of the premises relies exclusively on inference and completely forgets model training. In terms of workload shares, it’s a justified delimitation — inference is already 55–67% of compute. But that excludes the third that grows 4–5x annually (source: Epoch AI), and at the same time a large part of Nvidia’s competitive advantage: collective operations, inter-chip bandwidth, and the full training stack, which inference-designed hardware lacks or where local hardware bandwidth simply isn’t enough.
The second thing to note is that implicitly, the premises exude a pre-2026 era where zero-shot+reasoning are still the token-output drivers. E.g., AAII v.4.1.1 itself describes its metric set as “a transition to agentic workloads,” and over half of it is first-generation straightforward question-answer testing. The fact that over 40% of the index is still question-answer testing is not a neutral observation — it skews the result in favor of cheap models. Thus, we are talking a bit about apples and oranges at the same time, when meme image and text generation gets mixed up with the fact that demanding expert work is being feverishly shifted to agentic systems requiring wide bandwidth and massive context.
Then, onto analyzing the premises.
1. AI is being sold below cost, investors subsidize the difference
This is a bearing premise, and in my view, it doesn’t hold true, even though it’s a persistent claim in investor circles. Let’s break it down:
8×H100 node, market rental with all costs: $8–15 / hour
Throughput magnitude with continuous batching: 1,500–4,000 output tokens/s
→ 5.4–14.4 million tokens per hour
→ cost: $0.83–2.78 / million output tokens
With Blackwell, the same workload costs about one-seventh compared to the H100 (source: Inworld 4/2026).
Let’s compare to prices: Sonnet class $15/M output, and commercial serving of open weights (DeepSeek R1) $2.19/M (i.e., the price at which independent players sell apparently without loss).
API inference is thus comfortably gross margin positive, roughly 60–99% (depending on inference hardware). And that is independently verifiable: if frontier prices were below cost, third-party providers couldn’t sell a cheaper model at $2.19.
Where do the losses go then? To three places, none of which are token subsidization:
- Training runs: training a frontier model costs hundreds of millions or billions, as a one-off expense
- R&D personnel: thousands of people with top salaries
- Free tiers and fixed-price subscriptions for retail consumers — these can be loss-making, especially for heavy users of reasoning models. But the consumer segment and “generating cat memes” is hardly what carries the AI thesis anyway.
The difference is crucial. The subsidy is in training and product development, not in tokens. If OpenAI/Anthropic stopped training new models tomorrow, based on the data they would be immediately profitable — we will probably get final validation of this soon via the SEC-confirmed IPO prospectus. And that is a completely different economic structure than “selling below cost.”
And it breaks the chain of premises: “generating a meme image” thus doesn’t increase OpenAI’s loss — it decreases it, because it comes with a positive, even +90% margin. Marginal usage is not subsidized.
2. Most don’t need a frontier model, the majority get by with cheap models
This is certainly largely true. Someone smarter than me has said that companies don’t need Einstein in the finance department to process invoices and will certainly manage with a 100 IQ model. And this is already real life today. Model routing, cheap tiers, Haiku/Flash class models, open weights — all exist and are in widespread use. Companies are actively optimizing this because it’s direct cost savings. The mindless tokenmaxxxing boom is over — probably everyone admits that already.
However, a decrease in demand does not follow from this premise, and the reason is Jevons’ paradox. When unit costs plummet, total consumption does not decrease but rises.
Concretely:
- Cost per token: −300×
- Tokens per task (reasoning models, agents, tool calls): +100…1000×
- Cost per task: decreased, but much less
- Total consumption: rose sharply
Reasoning models are the purest example of this. Same model, same price per token, but it “thinks” for minutes before a single response. The efficiency improvement immediately turned into buying quality, not cutting costs.
The transition to running agentic workloads, which has been ongoing for roughly 12 months, is also a manifestation of Jevons — they consume hundreds or thousands of times more tokens than zero-shot queries. By nature, AI is a high-price-elasticity, scalable, broad-based factor-of-production commodity. And in history, there has not yet been a single commodity adapting to these attributes that has not behaved as predicted by Jevons’ paradox.
3. Local execution is practically free, shifting there will collapse demand
There are two errors here, one of which is physics and the other economics.Physics: bandwidth binds, not compute power. Autoregressive decoding (to put it roughly) reads the entire model from memory for each token. Therefore, speed ≈ memory bandwidth / model size:
| Device | Memory Bandwidth |
|---|---|
| Typical Laptop | 100–200 GB/s |
| MacBook Pro M4 Max | ~550 GB/s |
| RTX 5090 | ~1.8 TB/s |
| H200 | 4.8 TB/s |
| B200 | ~8 TB/s |
The difference is 10–50x, and it cannot be fixed with software.
Economics: “Free” confuses marginal and average cost. You are right that the marginal cost is electricity, because the hardware has already been purchased. But that obscures what makes datacenter inference cheap: batching.
A datacenter GPU serves dozens of users simultaneously with a single model read. Your device serves one and sits idle 95% of the time. Throughput per unit of hardware is 20–40x higher with batching (my own estimate based on the use of CSC supercomputers). That is the reason why centralized inference wins.
And the decisive point: agent workloads are the least suited for local execution. An agent runs for minutes or hours, in parallel, with a large context, while you are not at your computer. By definition, it is a server workload.
Local wins where it genuinely wins: small models, latency-sensitive use cases, privacy, always-on functions, offline. That is a real and growing segment—but it is a different (and vastly smaller in market share) workload than the one driving NVIDIA’s revenue and the broader AI investment boom. Nor are these mutually exclusive, at least in my view, except perhaps in some niche segment of technically savvy consumer-developers. In the enterprise segment, the opportunity cost of idling and all the extra local tinkering is quite high.
4. The highest profit margins in world history from artificial scarcity
This can be easily measured:
| Fiscal Period | Gross Margin |
|---|---|
| FY2017 | 58.8% |
| FY2018 | 59.9% |
| FY2020 | 62.0% |
| FY2021 | 62.3% |
| FY2022 | 64.9% |
| Future Guidance | ~71-75% |
So before the AI boom, the average gross margin was 61.6%. Now it is 10–13 percentage points higher.
That is the right order of magnitude for a “scarcity rent” — not 75% but under 13 points. And a substantial portion of even that is product mix: the shift from gaming GPUs to datacenters and from chips to complete racks increases the margin even without scarcity.
“The highest in world history” should also be put into perspective: ARM makes ~95%, Qualcomm’s licensing is over 70%, Microsoft ~70%. NVDA is certainly exceptional in that it makes 75% while selling physical goods at a volume of hundreds of billions — that is unprecedented. But the benchmark is not zero, but ~62%.
Like @Pohjolan_Eka, I would be somewhat cautious in assessing NVIDIA’s cheapness purely through the earnings multiple and potentially one year of forecasted (albeit reasonably certain) growth. The so-called Molodovsky effect is strong here. It’s worth considering what the market is pricing into NVIDIA and pondering whether you disagree with it. And indeed, right now the market is pricing NVIDIA as a cyclical hardware company and a strong smoothing of the cycle right after CY2028.
The latest earnings report offers a great event study on the market’s movements, which can be used to assess the implied probabilities given by the market based on the guidance anchor and the resulting valuation change. FactSet’s FY29 consensus prior to the release was below $750B. FY28 after the guidance release is $706B. So the company promised roughly the figure for next year that analysts had already booked for two years out. The curve shifted about one year to the left. The market reaction corresponded almost exactly to a one-year acceleration. If all future earnings arrive one year earlier, their present value increases by roughly the one-year cost of capital — in NVDA’s scale, about 9%. The actual move was 8.7%. In plain English, the market thus priced the change primarily as timing, not as a change in level.
So if one believes that the information provided in NVDA’s earnings report also implies a relative improvement in TAM and/or market shares, or the longer duration of the secular cycle and the resulting profitability, then now might be the time to play. But this cannot be deduced directly from the earnings multiples. I actually believe that the step-training launched by Lepikkö will continue among both analysts and market participants — the demand already in the order books ($2000B) alone will ensure this, not to mention the new demand brought by next-generation launches.
[quote=“Pohjolan_Eka, post:1300, topic:26423”]
We are quite far from such a world, and in the near future, things will still operate heavily on an application-driven basis, which practically requires local or hybrid execution, since software companies do not have the pricing power to charge an extra ten a month for every application just to run a frontier model on NVIDIA’s oversized and overpriced GPUs.
[/quote]Sure, this HCI revolution won’t happen overnight, of course. But you misunderstood my point a little bit. I see the development specifically in such a way that these proprietary LLMs or agents, as a complementary part of platform software user interfaces, will disappear or wither away. And they will be replaced specifically by the UIs of these harness developers, which can link dozens or hundreds of data sources and apps at the same time. Therefore, that extra fee is by no means billed by the software company, but rather by the harness and the company supplying the underlying models (such as Anthropic) as their share for automating company processes instead of disconnected GUI clicking.






