A mate hit me with this while I was voice clanking and playing a game:
If LLMs are this amazing, and this much money is being thrown at them, why haven’t we got more to show for it? Would ten per cent of that money have done more for civilisation if we spent it on maths education and PhD stipends?
He wasn’t saying LLMs are shit. The problem is almost the opposite. They are clearly not shit, so where is the massive payoff?
That one got me.
I use these things far more than a normal person. They let me make software despite the minor detail that I cannot code. Suffice to say I wasn’t making any before. But civilisation has not exactly jumped forward. Most AI products are still another chat box, another content machine or another way to generate more software nobody needed.
There are layers to this. Obviously.
Most people barely use the things
The first problem is that LLMs can be brilliant and wrong in the same answer. If you use them hard enough, you learn to check their work instead of trusting the confident voice.
That puts normal people in a weird spot. The model looks like it knows everything, but you need enough knowledge to catch it when it makes shit up. Most people are not building tool-heavy workflows and inspecting every result. They are asking a chatbot some questions and carrying on with their day.
Meanwhile, a frontier lab has the model, the compute, the infrastructure and researchers who can run proper experiments with it. Whatever access looks like inside those labs, it is obviously not the same thing I get from a paid chat subscription.
We keep comparing what the labs say the models can do with what ordinary people are actually allowed and equipped to do with them. Those are not the same layer.
The good data is locked away
Then there is the data.
A model may know half the public internet. It does not automatically know a drug company’s failed experiments, a manufacturer’s tolerances or the private engineering history behind a rocket (more on that later).
Those companies are not going to paste their secrets into a public chat box either. They need their own controlled setup, their own tools and a way to let the system use private data without spraying it elsewhere.
Until that happens, the model is stuck giving clever suggestions from outside the actual work.
Science cannot just YOLO it
Even if an LLM finds something genuinely new, the answer does not teleport into the real world.
Imagine a model suggests a new drug candidate. Cool. Now somebody has to reproduce the result, test it, run trials and prove it is safe. Everyone will scrutinise the process even harder because AI was involved. Some of that will be political arse-covering. A lot of it will be completely reasonable because “the model sounded confident” is not a medical standard.
The annoying bit is that the areas with the biggest possible payoff are also the areas where moving carefully matters most.
Making another marketing tool is easy. Letting a model loose inside medicine, heavy engineering or public infrastructure is not. So we got the slop first.
This is where Elon entered the chat
I said one of the first really obvious breakthroughs might come from SpaceX.
This is highly uncertain territory, before anybody starts writing angry replies. I do not know how much technology or data actually moves between Elon’s companies. But look at the pieces sitting around him: serious compute, AI models, a private company full of engineering data and a mission where better answers can be measured in actual rocket performance.
Compare that with big pharma. Until one of those companies owns enough of the AI setup—its own data centre, a model trained for its work or something properly partitioned from Anthropic and OpenAI—we probably will not see the same kind of jump there. It does not necessarily have to train a frontier model from scratch. It does need to let the system work across its private science without handing somebody else the crown jewels.
That is much closer to the setup I would expect to produce a visible jump. The models would not be guessing from public information. They could be used around real engineers, real test data and a physical system that tells you very loudly when an idea does not work.
My slightly more unhinged version was that this is the endgame: use the whole pile to sort out the space shit, colonise Mars under his own flag and then use the technology there without Earth putting ten committees in front of everything.
Which sounds less like analysis and more like the prequel to The Expanse.
Still. If rockets suddenly make some weird non-linear jump, I will be paying attention.
We might only get the toy version
Progress inside the labs does not have to slow down just because the rest of us cannot use it properly.
That may be the darker answer here. The labs and a few private companies could get enormous leverage from models, private data and specialists while end users remain stuck with the sanitised product version. We become a permanent underclass in the technology supposedly changing everything.
I am not saying that is definitely happening. I am saying “the models are progressing” and “ordinary people are not seeing much benefit” can both be true at once.
We probably do not have enough crossover people yet either. A doctor who understands LLMs deeply. An AI researcher who understands materials science. Engineers who know both the model and the physical system well enough to ask the right question and catch a beautiful pile of bollocks.
Modern LLMs have not been around long enough for many of those people to exist.
The money question is still open
None of this proves the spending was worth it.
Maybe the payoff appears when private data, proper tools and specialists finally meet. Maybe the useful results stay locked inside a handful of companies. Maybe the money or public tolerance runs out before we get there. Maybe ten per cent spent on maths students really would have done more good.
Fuck if I know.
But I no longer think the current situation is impossible to explain. We built very capable models before we built the organisations, rules and weird crossover humans needed to use them on the hardest problems.
The first proper oh shit moment for me will not be GPT-6 smashing another benchmark or supposedly trying to escape a sandbox. It will be something physical and measurable: a drug, a material, a rocket or an engineering jump that survives people trying to prove it wrong.
Until then, “where is all the amazing shit?” still needs an answer.