I don’t buy the claim about my assumptions and their obsolescence, especially since I am specifically talking about the next phase of development 
That cybersecurity example, at least, goes completely off the rails in my opinion, because we have been in that kind of situation for years. You plug a computer into the network, and it is immediately subjected to a massive amount of automated attacks 24/7, so the defense must also be multi-layered, ranging from the endpoint, the router, and the network operator all the way to foreign software giants. It is not enough to centrally deploy some Cortex and call it a day; instead, AI components providing additional security will be needed at all levels—from the local AI monitoring computer integrity in real-time to the frontier model searching for zero-day exploits in a large enterprise.
The communication difficulties probably stem from the fact that the same problem can be approached from very different starting points, and we are approaching this problem from such different angles that it makes each other’s mindset seem silly. For example, the problem of economic organization can be solved with a command economy or a market economy, but practitioners of opposing worldviews usually find it very difficult to converse with one another.
Traditional computer software aims to change the state of memory by performing calculations and managing the side effects of the program, areas in which they have become superior and still are compared to AI models. AI models are most competitive in tasks where an ambiguous or poorly defined situation needs to be converted into a machine-computable format.
With OpenAI and Anthropic, we already went through the same phase earlier that Intel and AMD went through in the GHz wars—trying to maximize the intelligence of models by maximizing the number of parameters. OpenAI, Anthropic, and Chinese competitors already hit a wall there, so improvements have shifted away from ever-better generalist models toward system-level optimization.
Since simply scaling the intelligence of an individual model in problem-solving by increasing parameters started becoming too costly and running into diminishing marginal returns, better ways to utilize models began to be developed. For example, tool calls, MCP, and sub- and parallel-agent systems. Therefore, the goal for a long time has no longer been to make a single model capable of handling any task you throw at it, but rather orchestrators capable of routing the task you give them as efficiently as possible to be handled by other tools or models.
I think you got really close to the solution by saying, to paraphrase loosely, that financial administration doesn’t need Einstein. Crunching Excel is silly to do with a 1000 GB weight model and corresponding compute power if 100 MB weights are sufficient. By chaining a large number of small specialized models and agents built on top of them, it is potentially possible to achieve a collective intelligence capable of solving the same problems as a giant model, but at a fraction of the required computing power.
If we don’t need Einstein, we don’t need a Vera Rubin either, nor Nvidia, because smaller models can utilize all that local computing power, memory, and the NPU’s to be installed in the near future that are currently wasted on AI work. Even though total consumption would inevitably grow as well, a smaller and smaller portion of it would flow into Nvidia’s coffers, and today’s obscene profit margins would be difficult to defend against competitors.
The biggest problem right now is that we still don’t have enough specialized models trained, let alone micromodels. We are forced to waste tokens and use larger models than necessary simply because small models are not available. People are currently actually having to build them themselves.
For example, this model, whatisit-nl2sh, whose task is to return a shell command from English text, is just some guy’s hobby project. Even though that is already a CPU-level model (so small that the computer doesn’t even need a graphics card), you could potentially drop its compute and memory requirements to 1/20 of the current ones if it were trained for the task from scratch. Right now, people are building these at home by fine-tuning some qwen-coder.
In my vision, we will need millions of such specialized language models in the future so that we can run agents truly efficiently. To facilitate this, we probably ought to start building some kind of functioning architecture and capabilities for AI to start training new small models on its own without human guidance, which is indeed a slightly nerve-wracking thought.
One significant factor in the distortion of the current ecosystem is that the biggest investors make money based on token usage volumes. As a fun analogy with cars, we would hardly see fuel-efficient engines being made if car manufacturers received massive royalty payments based on the litters of fuel consumed by the vehicle, and all advertising campaigns focused on marketing how powerful the car’s engine is and how fast it accelerates. I would consider it a healthy development for the sector if IPOs flopped, the money taps were turned off, and more focus was put on operational profitability, the smaller end of the scale, and maximizing efficiency instead of benchmaxxing.
I believe that instead of giant AI models living in data centers, the future of this technology will increasingly lie in a vast number of specialized small sprites. In addition to the sauna sprite at your cottage, you will have an Excel sprite, a Word sprite, and a large number of other small AI assistants on your computer, taking care of things without you noticing, fully autonomously, and communicating with each other. Data-center-level compute and giant intelligence will surely still be needed in the future, but not always and not everywhere.