Pass the Butter Knife, OpenAI and Anthropic Are Toast
We've been here before - consumers will pay for the simplest pricing model, and soon that will be 'free,' anyway.

Soon the simplest AI pricing model will be free per-run,
but you’ll have to pay for the hardware.
I’ve read Steve Yegge’s most recent post and I think I have to concede he’s mostly got the right of it. This is bad, bad news for OpenAI, Anthropic, NVIDIA and the US AI companies. I already thought free, open models would win, and now I’m certain.
Why?
Because if Yegge is right, nobody can afford to burn the tokens required to do it, day-in, day-out, other than the biggest companies on earth. For his approach to scale, everybody needs a simple pricing model at least—and a free one at best.
For reference, I’ve been writing for months about the opportunities and difficulties of harnessing the potential of AI, from the position of a tech-positive AI skeptic.
Mostly I’ve been trying to reconcile the two poles of non-determinism and the randomness of the LLM output, but Yegge has successfully put an alternative idea in to my head: what if we just YOLO it?1
Previously, I’d have dismissed that position as pure AI psychosis, but for some reason, today I stopped to seriously consider if I was wrong. And, I think I am. I’m at least guilty of a category error in terms of identifying the problem.
Determinism isn’t the problem. Convergence on a statistically predictable outcome is the problem.
That determinism was perhaps never truly attainable even in the pre-LLM world is something I knew all along, but maybe too many years of thinking about determinism in the blockchain space blinded me to rethinking this.
Software Development Is Now Gambling, And I Feel Fine
‘Just YOLO it’ is a facetious framing, but put another way: the governance of software development has always been about managing downside risk in a complex system, and about making sure cost doesn’t outstrip value. Most prosaically, it’s about making sure the software still works after a change.2
Because of the halting problem, we’ve always known that the result of running a program is unknowable until we... you know, run it. That’s why we have guardrails like testing, QA, or even fuzzy testing, release cohorts, A/B testing, et cetera at scale.
Thus this development process is not deterministic; usually, it’s about risk management.
It’s gambling, dear reader.
So if we’re playing the gambling game, then we are in a world where we’re trying to get the results of agents to converge on some more predictable average, which we can do by adding context (that is, the agent orchestration framework, or beads, in Yegge’s case) and by adding more runs (to see if the output runs converge toward a mean).3
If this approach is valid, in a sense I was right in my post on AI cost in orgs (linked above)—the real difficulty is in scaling governance (orchestration, and verifying outcomes, in this case).
That shouldn’t be a surprise, though. Code was never the bottleneck.
So The AI Companies Will Make Money, Right?
Wrong.
Yegge’s assumptions hinge on a huge, huge expansion of usage. The US AI companies’ valuations and future survival hinge on a historic increase in revenue, via charging you, the users. This is a paradox: you need to use more than you can afford.
Thus, key to making this new development paradigm work is the following: you shouldn’t bother with the expensive US models.4
Not only is it against your own interests financially to pay them, but we can be reasonably sure that these companies will fail, regardless of how the market develops. In fact, the more that demand increases, perhaps the more likely it will be that they fail.
That sounds counter-intuitive, but the reason for this is pretty simple—as described by Odlyzko (2001),5 we should expect that consumers will express a preference for the simplest pricing model. Odlyzko argues based on past pricing trends that flat-rate pricing models will come to dominate over time, due to consumer preference.
Consider the example of music: MP3s came to dominate pricing due to the licensing model of the protocol (free), which then ironically allowed for mass piracy, a model with a simple dollar-cost, but variable cost in terms of shoe-leather costs. Thus, consumers eventually expressed a preference for an even simpler pricing model: flat rate, as we saw with Rdio, Spotify et cetera. These services charged a bit more, but Spotify came to dominate over even piracy because of its simpler pricing model.6
Odlyzko was right about flat-rate pricing for the internet, and, in the end, he was even right for mobile tariffs (in the UK, at least), something he wasn’t confident to predict in his original piece.
That’s already a difficult precedent for the AI companies increasing their prices, or introducing more complex structures, before we even consider that some competing models are free.
So we know that:
Users are incentivised to seek the best deal (assuming rationality and information)
Users will express a preference for the single simplest pricing model7
Open-source models are catching up fast with cutting edge US models
You will be able to run these models on your own hardware8
This means that soon, the simplest pricing model will be: buy your hardware and run a free AI model. The cost of each marginal run will be zero.
At the moment, sure, an H100 is going to run you hundreds of thousands of dollars (and may be overkill—RAM is what you need), but what happens when the kind of $20-30k machine that used to be used for crypto validators will do?9
I’ve spent the last 5 years running those bad boys, and I can tell you, the overhead of managing them isn’t too bad. If a simpler option comes along (see footnote 7), then this kit and approach is going to go from ‘first-adopter advantage’ to ‘industry standard hardware’ overnight.
At that point, the OpenAI and Anthropic pricing models collapse, every engineer or team runs their own AI server, and a thousand new dev shops take flight. Jensen Huang goes back to his semi-successful career as a Richard Hammond impersonator.
Maybe.
In any case, this is all very good news for software engineers in the short-term.10 However, these approaches will find their way gradually into other industries for certain types of task, and good news, software engineers: you will be the ones that take them there!
I know the whole ‘Ralph Wiggum’ approach has been around for a while, but for some reason this feels different. It isn’t about a straightforward loop, or task decomposition, but about the same probabilistic approach as say, a large Kubernetes deployment. Maybe I’m wrong, or splitting hairs. You tell me.
Assuming it worked, or at least delivered the desired outcome, before the change.
The other thing of course you can do is add other runs to validate the original tasks, and infinite variations on the above. Either way, eventually it’s a case of a million monkeys at a million typewriters—as long as you’re experienced enough in your domain to know what Shakespeare looks like when you see it, then maybe you’re good.
At least in the long-run. For now, some experimentation is in order to see how close the models are.
Here’s a link to the Odlyzko paper. If you’d prefer the journal version and have an institutional login, here’s the version I am referring to.
Thanks to my friend Jack for this example, which came up in conversation.
Imagine how much more preference they will express for a free one!
The simplest overall model therefore is clearly ‘buy some hardware, and call it a day.’ Large businesses already possess the capital and capability to do this. It’s just generally beyond the budget of most consumers. Which is fine, because here I’m talking about software engineers, not your mum running a beefy server at home.
Or, indeed, a maxed-out Mac Studio? I can tell you now that we will be taking delivery of a couple here in the bunker the day after that becomes true. That is another reason NVIDIA are cooked—if you can run on a Mac Studio for £10-20k, nobody with two brain cells to rub together is going to pay for the NVIDIA option.
My views of course on AI use in software engineering have softened greatly over time. Although I see programming mainly as a creative pursuit, it’s also my living, and I have to be pragmatic. Still, if you use AI to write, say, a book, you’re still missing the point, and still a hack. For creative work, the process is the point. Get good, and don’t cheat yourself.


!["I didn't grow up dreaming of prompting." [Part 3]](https://substackcdn.com/image/fetch/$s_!1I3B!,w_140,h_140,c_fill,f_auto,q_auto:good,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6f2cb68d-aae2-4f6e-a4dd-377ba06e4b84_3295x1648.png)

