This week Moonshot released Kimi K3, fully open-weight. The math is impressive: benchmarks that match frontier models like Opus 4.8 and GPT-5.6 Sol at a third the price. The standard dismissals arrived on cue: sure, it's impressive, but only because they distilled the frontier models to build it. Anthropic all but said so itself.
"They just distilled it" is meant as a rebuttal. I think it's the whole ball game. If the frontier can be distilled into a model that runs at a third of the cost, the capability was never the moat. The moat was the lag. And the lag is compressing.
The trillion dollar question is which layer keeps the value
Which leaves the one question worth asking in a market this young: not who builds the best model, but where in the stack the value actually accrues. Silicon, cloud, model, application, your own data: value flows through all of them, but just because value passes through a layer doesn't mean it settles there. Right now every major player in the market is working furiously to make sure the value settles in theirs.
Everyone tries to commoditize their neighbor
There's a move underneath all of it, older than AI: commoditize the layers adjacent to you, so the value pools in yours. Economists call the durable version appropriability: the winner isn't the one that builds the cleverest thing, it's the one that owns the part nobody else can copy. Microsoft did it to PC hardware; the economics are older than that.
You can hear this language elevating in the market. Jensen Huang plays it from the floor: "We are a vertically integrated computing company. There is no other way." Nvidia sells the picks and shovels, and it is climbing. Satya Nadella plays it from the cloud, warning that companies "pay for intelligence twice" (once in tokens, again in the data exhaust the vendor learns from), while Microsoft sells the orchestration layer he offers as the cure. Alex Karp plays it loudest, from the application layer: the labs, he says, want to strip enterprises of their "alpha," customers are "livid", the token model is "completely wrong", and Palantir's pitch is to make the model interchangeable. Each is describing the same shift from a different vantage.
Commoditization doesn't destroy value, it moves it around
Walk it forward. Assume the open-weight labs succeed and can distill the frontier rapidly after each launch. The value of the models doesn't vanish, it relocates. The chips still sell. In fact, more chips sell because cheaper models drive more demand for inference. The enterprise keeps more of the alpha that Karp says they're quietly handing over. Liang Wenfeng, who gives DeepSeek's weights away, is blunt about it: "moats created by closed source are temporary." The open-weight labs aren't trying to extract value at the model layer. They're trying to dissolve it, and letting the value fall to the infrastructure below and roll to the software and customers above.
The race that matters runs vertically, not across
Every new model drop is met with great fanfare, and there has never been a more competitive set of models on the market than there are today. That could fool you into thinking that was the race: Anthropic against OpenAI against Google against the field. Lab against lab. None of them can fall behind. But also, none of them can win. The pressure that decides where the value accrues actually runs vertically: Nvidia at the bottom of the stack, then the hyperscalers, the open-weight labs pulling from the side, and the applications and enterprises that own context from above. All are pulling at the model layer at once. And the labs' own success is what summoned it. By proving how much value is created in the model layer, they showed every neighbor exactly where to aim.
The hot takes that the value is sliding towards the application layer are a dime a dozen right now. The part worth saying is why the labs have no choice but to follow it there, and what the eventual market implication is.
AI is like big Pharma without the patent
Start with the question that decides everything downstream: if distilling a model is cheap and easy and training the frontier is expensive and hard, who funds the next frontier? We have seen this cost dynamic before. In pharma it costs ten figures to discover and prove a new molecule, and often just pennies to copy it. The only reason pharma funds that cost is because a patent grants a window of excludability. The pace of model distillation shows there's no excludability here. When the innovator can't capture the return on investment, the textbook says underinvestment will follow.
When capital is free-flowing, this is a treadmill. Labs at the frontier invest ahead, build and tune the next frontier model, and then the clock to distillation starts the day of launch. The labs have a natural window to sell the state of the art before its offspring catch up. This is the race of the Red Queen, from Lewis Carroll's Through the Looking-Glass. You run as fast as you can just to stay in the same place, 6 months ahead. This is more challenging when you have an eye on profitability. Training costs rise, distillation costs stay flat, and the lag shrinks. Running as fast as you can may not even be enough to stay in the same place.
The API is the distillation surface
So the labs are structurally incentivized to do something about this value leak. And the leak is the API itself. Every call hands a competitor a clean, labeled pair: input in, frontier behavior out. To sell the model wholesale is to accept that you're teaching it to a rival, prompt by prompt. The only way to stop the leak is to stop exposing the API. This means absorbing the next link in the value chain, turning the model into an internal organ, and selling the finished outcome on the other side: the resolved legal memo, the managed stock portfolio, the closed support ticket. It works because the rivals can't distill what they can't cleanly observe. Bury the raw generation under retrieval, tools, private data, and review, and the signal degrades into noise. The outcome stays legible; the asset goes dark.
It's necessity, not appetite
Karp and Nadella sound indignant, like they're describing a land grab. In fact, they're describing a liquidity crisis. The labs move up the stack not out of appetite but because the wholesale floor for the model is sinking. The accusation is that they'll absorb the next link and compete with you. The defense is that they must absorb the next link to protect their trillion dollar asset. This is the same reality observed from different angles. Verticalizing into first-party legal, wealth management, and support isn't just a growth play. It's how a lab renders its one distillable asset illegible.
This logic predicts a new release strategy. GLM 5.2 and Kimi K3 show that if the lag is the moat, sprinting is no longer sufficient to preserve it. The rational conclusion from the perspective of the labs is to actively control the lag. Keep the best models back for first-party services, sold only as outcomes. Waterfall the prior-generation models in tiers of trust where they're less likely to be distilled. Offer last year's model freely through the API, and expect the open weights to track closely behind. The models being distilled are always two steps behind the models at the frontier. The pursuers still catch up, but the frontier has already moved on. The lag stops being a law of nature and becomes a lever.
The lag is the lever
Mythos launched to a small set of privileged institutions. Its public sibling Fable sits behind an aggressive set of safeguards. This won't be an N of 1. It will be the new normal. Of course there will be attempts at regulatory regimes and export controls and technical countermeasures to slow the distillation. But, as long as leakage remains, the incentive gradient points in this direction.
This is how you outrun the Red Queen. You accept that you can't. The race is a trap. So you change the game.
Researched with Gemini and Perplexity · Drafted with Claude Opus 4.8 · Hand crafted in Google Docs · Header image generated in Google Imagen 3