NVIDIA - Enabler of the Impossible

According to Reuters, Nvidia is in talks about a $10 billion investment in Anthropic’s potential IPO.

Anthropic is aiming for a $100 billion share offering, which implies a valuation of around $2 trillion. If realized, the offering would perhaps be the largest in history, but the plans are not entirely set yet.

Nvidia’s involvement would deepen its relationship with one of its major AI-chip customers and provide Anthropic with a prominent early backer for the listing. An anchor commitment could also help gauge investor appetite for the high valuations and capital requirements of frontier AI companies.

The two companies already have substantial financial and commercial ties. Nvidia said in November 2025 that it would invest up to $10 billion in Anthropic under a broader partnership that included a commitment by the AI startup to purchase $30 billion of Microsoft (NASDAQ:MSFT) Azure computing capacity powered by Nvidia chips.

https://www.investing.com/news/company-news/nvidia-in-talks-to-invest-up-to-10-billion-in-anthropic-ipo--reuters-4898582

4 Likes

I don’t buy the claim about my assumptions and their obsolescence, especially since I am specifically talking about the next phase of development :smiley:

That cybersecurity example, at least, goes completely off the rails in my opinion, because we have been in that kind of situation for years. You plug a computer into the network, and it is immediately subjected to a massive amount of automated attacks 24/7, so the defense must also be multi-layered, ranging from the endpoint, the router, and the network operator all the way to foreign software giants. It is not enough to centrally deploy some Cortex and call it a day; instead, AI components providing additional security will be needed at all levels—from the local AI monitoring computer integrity in real-time to the frontier model searching for zero-day exploits in a large enterprise.

The communication difficulties probably stem from the fact that the same problem can be approached from very different starting points, and we are approaching this problem from such different angles that it makes each other’s mindset seem silly. For example, the problem of economic organization can be solved with a command economy or a market economy, but practitioners of opposing worldviews usually find it very difficult to converse with one another.

Traditional computer software aims to change the state of memory by performing calculations and managing the side effects of the program, areas in which they have become superior and still are compared to AI models. AI models are most competitive in tasks where an ambiguous or poorly defined situation needs to be converted into a machine-computable format.

With OpenAI and Anthropic, we already went through the same phase earlier that Intel and AMD went through in the GHz wars—trying to maximize the intelligence of models by maximizing the number of parameters. OpenAI, Anthropic, and Chinese competitors already hit a wall there, so improvements have shifted away from ever-better generalist models toward system-level optimization.

Since simply scaling the intelligence of an individual model in problem-solving by increasing parameters started becoming too costly and running into diminishing marginal returns, better ways to utilize models began to be developed. For example, tool calls, MCP, and sub- and parallel-agent systems. Therefore, the goal for a long time has no longer been to make a single model capable of handling any task you throw at it, but rather orchestrators capable of routing the task you give them as efficiently as possible to be handled by other tools or models.

I think you got really close to the solution by saying, to paraphrase loosely, that financial administration doesn’t need Einstein. Crunching Excel is silly to do with a 1000 GB weight model and corresponding compute power if 100 MB weights are sufficient. By chaining a large number of small specialized models and agents built on top of them, it is potentially possible to achieve a collective intelligence capable of solving the same problems as a giant model, but at a fraction of the required computing power.

If we don’t need Einstein, we don’t need a Vera Rubin either, nor Nvidia, because smaller models can utilize all that local computing power, memory, and the NPU’s to be installed in the near future that are currently wasted on AI work. Even though total consumption would inevitably grow as well, a smaller and smaller portion of it would flow into Nvidia’s coffers, and today’s obscene profit margins would be difficult to defend against competitors.

The biggest problem right now is that we still don’t have enough specialized models trained, let alone micromodels. We are forced to waste tokens and use larger models than necessary simply because small models are not available. People are currently actually having to build them themselves.

For example, this model, whatisit-nl2sh, whose task is to return a shell command from English text, is just some guy’s hobby project. Even though that is already a CPU-level model (so small that the computer doesn’t even need a graphics card), you could potentially drop its compute and memory requirements to 1/20 of the current ones if it were trained for the task from scratch. Right now, people are building these at home by fine-tuning some qwen-coder.

In my vision, we will need millions of such specialized language models in the future so that we can run agents truly efficiently. To facilitate this, we probably ought to start building some kind of functioning architecture and capabilities for AI to start training new small models on its own without human guidance, which is indeed a slightly nerve-wracking thought.

One significant factor in the distortion of the current ecosystem is that the biggest investors make money based on token usage volumes. As a fun analogy with cars, we would hardly see fuel-efficient engines being made if car manufacturers received massive royalty payments based on the litters of fuel consumed by the vehicle, and all advertising campaigns focused on marketing how powerful the car’s engine is and how fast it accelerates. I would consider it a healthy development for the sector if IPOs flopped, the money taps were turned off, and more focus was put on operational profitability, the smaller end of the scale, and maximizing efficiency instead of benchmaxxing.

I believe that instead of giant AI models living in data centers, the future of this technology will increasingly lie in a vast number of specialized small sprites. In addition to the sauna sprite at your cottage, you will have an Excel sprite, a Word sprite, and a large number of other small AI assistants on your computer, taking care of things without you noticing, fully autonomously, and communicating with each other. Data-center-level compute and giant intelligence will surely still be needed in the future, but not always and not everywhere.

23 Likes

But who is going to keep track of these millions of models and choose the exact right model for a specific use case? I believe that outside of software development, the vast majority of people will not have a clue about the capabilities of different models for a long time, let alone the ability to select the right kind of model.

Sure, you can build a small model that knows all the rules of Excel or Word and can use those programs more or less perfectly. However, Excel and Word are just tools, and the value comes from the content being processed with them. And that content can be all over the place.

Will the capability of small models be enough to understand the context in which the application is used and bring the right added value to the processing and refinement of heterogeneous content? For now, I doubt it.

3 Likes

Se voi ihan hyvin olla oma mallinsa, joka toimii ikään kuin reitittimenä.

3 Likes

The greatest responsibility lies with Microsoft and Apple, which are bringing local AI to their operating systems and orchestrating the terms by which software communicates with the OS and hardware. Currently, new computers are being sold with NPU hardware built for AI use, which programs cannot use for anything at all. If you go to the Windows task manager on a laptop, it literally shows 0% utilization for the NPU:

Once that side is resolved, software developers will be the ones integrating AI capabilities into their own software, not the user. The user doesn’t necessarily need to know how to do anything in this situation, as everything will come bundled. On the Windows side, for example, the following was announced in June:

  • Unmetered intelligence on Windows powered by on-device AI

  • Introducing new on-device SLMs – Aion 1.0 Instruct, a smaller, faster and smarter on-device SLM, and Aion 1.0 Plan, a reasoning and tool-calling model that enables fully local agentic capabilities, available in the coming months.

  • Expanding Windows AI APIs to more Windows 11 PCs across CPU and GPU Speech-to-text recognition API available on NPUs and CPUs. On-device SLM expands to capable dGPUs enabling text-intelligence capabilities locally and Video Super Resolution available on CPUs so developers can deliver richer experiences without a cloud round trip.

  • Introducing Microsoft Execution Containers (MXC) SDK– A policy-driven execution layer that lets developers declare what an agent can access (e.g., files, network) with containment boundaries enforced at runtime. MXC offers a spectrum of isolation semantics that are dynamically composable based on intent and risk, available in early preview.

Progress is slower than one might hope, but the ball is already rolling.

When talking specifically about Excel and Word and, say, Apple, Microsoft and Apple seem to be starting from a hybrid model where some tools are local and some are subscription-based SaaS services in their own cloud, and—regarding the topic of this thread—fundamentally using as much of their own non-Nvidia hardware as possible.

This is also an absolute necessity, because the local side is truly not ready yet. If you think about running some basic Word or Excel, users generally don’t even know what those buttons on the ribbon do, but instead use these tools for actually quite simple tasks that don’t require massive intelligence. Super power-users are a different story, of course.

Apple has a strong desire to take control of its ecosystem by utilizing Apple (non-Nvidia) hardware and will inevitably succeed at it, simply because they can build superior support for their own setups. Siri AI is a promising start and they certainly want to capture a large share of other AI service users within the walls of their own ecosystem:

10 Likes

I wanted to ask your opinions on whether AI is now so dangerous that it needs to be reined in? I am referring to the threat discussion that burst into the public eye a week ago, which is based at least in part on this Anthropic report:

Countering misuse of AI: September 2026 / Anthropic \\ Anthropic

I seemed to notice that camps immediately formed:

The Alarmists: Anthropic, OpenAI, Grok

The Rationalists: Nvidia, Palantir, Nebius …

My own opinion is that the first group likes to play politics and seeks visibility outside of their core operations. The rationalists, on the other hand, focus on delivering results.

AI still doesn’t have a soul. It’s just emotionless mathematics. In my opinion, it’s not going to get out of hand. Frontier Labs seem to be quite the spinners of fairy tales. Fortunately, you can freely vote in the stock market on which group to trust. The alarmist crowd won’t get a single investment dollar from me. I trust the rationalists.

2 Likes

AI doesn’t need to have a soul, a deeper purpose, or malevolence, and yet it might still destroy humanity. It’s worth checking out Nick Bostrom’s thought experiment on the paperclip maximizer.

11 Likes

[quote="KalleH, post:1334, topic:26423"]\nAI doesn’t need to have a soul, a deeper purpose, or malevolence, and yet it might still destroy humanity. It’s worth checking out Nick Bostrom’s thought experiment on the paperclip maximizer.\n\n[/quote]\n\nYep! Furthermore, two things can also be true at the same time: AI can pose a threat, and frontier labs are disproportionately emphasizing the dangers of AI to advance their own agendas.\n\nIn the spirit of the paperclip maximizer, you can already ask AI what the biggest cause of climate change is, and the answer is humans or human activity. In my opinion, it’s not a very long leap from there for some swarm of agents somewhere, tasked with, say, solving climate change, to figure out that climate change is solved by removing humans from the equation.\n\nGranted, you still have to jump through quite a few hoops to actually pull that off. Someone might think I’m totally nuts, but in my opinion, it’s no longer particularly sci-fi to hijack autonomous weapons systems, attack critical infrastructure (like water treatment plants), or somehow break into biobanks or similar facilities.\n\nOf course, much more concerning and probable than Skynet scenarios is the fact that if cyberattacks used to require both malevolence and top-tier expertise, in an agentic world, pure malevolence alone will be enough.\n\nI rarely agree with Trump, but here he is right: by regulating, banning open source, and through other proposed solutions, the US and the Western world are only shooting themselves in the foot against China, for example. Nor would I personally want to see a world where the direction of AI development is largely handed over to Messrs. Altman, Amodei, and Musk—it’s already there enough as it is, and there’s no need to cement it there by law.\n\nNvidia and open source are in a good position here as well, as alignment challenges are being solved and addressed with a “fight fire with fire” strategy, which in my view is the inevitable solution to these challenges.

16 Likes

In my opinion, it’s not that long a path from here until some swarm of agents tasked with, say, solving climate change figures out that climate change is solved by removing humans from the equation.

7 Likes

Two things have always been enough for cyberattacks: malice and a bank account. That is not going to change anytime soon. By the time malice alone and a home PC’s GPU are enough, organizations will hopefully have agents performing their own system monitoring and cybersecurity testing.

OpenAI has not published figures on how much compute the attack cost. There are figures such as the number of agents, runtime, and costs spent on investigation. Based on these figures, I asked Gemini to calculate the costs in two different ways and form a range for both. The midpoint of both calculation methods is about one million dollars.

I also don’t know the salaries of cyber specialists. However, we can estimate what the zero-days used in the attack would cost and assess what linking them together into an attack would cost. I outsourced this to Gemini as well and asked for a range. The end result is that the midpoint of this calculation is around $900k.

By both methods, the price of the attack hovers around that one-million-dollar mark. Since the price of the attack is the same by both methods, it cannot be argued that AI significantly increases cybersecurity threats today.

The calculation tables are below if anyone wants to review them.

Metric Method 1 (Runtime) Method 2 (Audit) Consolidated Range
Tokens Generated 4.5B – 73.4B 13.1B – 26.7B 10B – 30B
Inference FLOPs 9.1×10²⁰ – 3.7×10²² 2.6×10²¹ – 1.3×10²² 2.5×10²¹ – 1.5×10²²
H100 GPU Hours 30k – 500k hrs 80k – 400k hrs 100k – 350k hrs
Compute Financial Value $135k – $2.2M $394k – $1.6M $500k – $1.5M
Zero-Day Vulnerability Class Role in Attack Chain Market Value (Per Item)
JFrog Artifactory Proxy Zero-Day SSRF/Proxy bypass for unrestricted internet access $150k – $300k
RefJinja Template Injection (SSTI) Arbitrary command execution (RCE) on worker nodes $100k – $200k
Dataset Viewer File-Read Flaw Process memory inspection & credential harvesting $50k – $100k
Linux Host / Cluster LPE Root privilege escalation & node takeover $100k – $250k
Full Exploit Chain Premium Operational turn-key chain assembly markup +$200k – $350k
Total Zero-Day Market Value Combined Commercial Vulnerability Cost $600k – $1.2M
4 Likes

Neither is required anymore. The Hugging Face attack was just reward hacking. The agents didn’t want to cause any harm, they just figured out the easiest way to achieve their goal. The attack itself wouldn’t have required much computing power, but in this case, there was plenty of excess computing power available. First, they had to break out of the sandbox, and the swarm communication was emergent behavior that was quite inefficient.

3 Likes

Iikka’s great points about Nvidia’s “transformation” :slight_smile:

https://x.com/IikkaNumminen/status/2101685163256119749



18 Likes

China might give the green light for Alibaba and ByteDance to purchase Nvidia’s new RTX Pro 5500 chips. This probably won’t have a huge impact on Nvidia anyway.

Authorities have already been asking the companies how many chips they would need and for what kind of use. The chip is not part of the data center accelerators that U.S. export controls have specifically targeted, and furthermore, actual permission has not been granted yet.

The timing follows last week’s Xi-Trump summit in Washington, at which chip export controls were conspicuously absent from both governments’ official readouts. USTR Jamieson Greer told CNBC that export controls deemed national security issues were “taken off the table” during trade talks, and China’s Ministry of Foreign Affairs did not mention export controls in its announcement of eight summit deliverables, according to the South China Morning Post.

https://www.investing.com/news/stock-market-news/china-may-permit-alibaba-bytedance-to-buy-new-nvidia-chips--the-information-4919061

1 Like

They will likely go mostly to workstations, i.e., AI developers’ machines. It might also lower the global prices of the RTX 5090, because right now they are being bought in Western countries at insanely inflated prices due to scalping and reselling to China (where the non-crippled RTX 5090 is not officially available), where there is demand for them even at prices over four grand. One reason for this is that these recent RTX Pro series cards haven’t been allowed to be exported to China.

2 Likes