A paper posted last week to r/LocalLLaMA, the main gathering place for people running language models on their own hardware, makes a large claim about a small number. Reinforcement learning, the expensive post-training stage that teaches a model to reason in steps, changes only one to three percent of the tokens a model produces compared with its base version. The authors say they recovered most of the benchmark improvement without any reinforcement learning at all, at roughly a thousandth of the compute, by targeting that thin slice directly. The read that traveled fastest was about money: if reasoning is that cheap to install, the training advantage the frontier labs have been buying with capital expenditure is thinner than the capital expenditure implies.
Everyone read this as a story about training cost. Read it instead as a map. It locates reasoning in a thin, separable band of behavior sitting on top of a base model, and thin separable behavior is a setting. Settings get set by whoever operates the endpoint, on their schedule, against their margin. The finding that reasoning is cheap to install is the same finding as reasoning being cheap to adjust, and adjustment is the part that reaches into a product you already shipped.
Hold the claim itself loosely. It is a preprint circulating among practitioners rather than a replicated result, and reasoning papers of this kind have a poor record of surviving independent reproduction. The reframe does not depend on the paper being right. It depends on the paper being one of four things pointing the same direction in a single week, and the other three are already shipping.
Simon Willison's verdict on Qwen3.8 27B, the open-weight vision model Alibaba released Friday under an Apache 2 license, was excellent quality paired with enormous quantities of thinking tokens spent on trivial prompts unless you tell it otherwise. People running head to head comparisons locally report the same thing. The complaint has a specific shape. The model reasons at a volume nobody asked for, because that volume shipped as a default, and defaults are tuned to look good on benchmarks rather than on a monthly bill.
Put that beside NVIDIA's serving documentation for configurable reasoning effort and the client-side templates for steering effort that surfaced last week, and thinking budget starts to look like a runtime knob rather than a model characteristic. Any 2026 reasoning model now arrives with an implicit dial. Somebody sets it. Most of the time that somebody is the provider, and the position they choose is the one that makes their published numbers look best.
The training cost read is the one a reasonable person makes, because cost is the visible number. Capital expenditure is what gets reported: gigawatts, cluster counts, chip allocations. A result suggesting the expensive stage was skippable lands directly on that story, which is why it moved so fast. What it misses is that the labs were never selling the training run. They sell an endpoint, and the endpoint's behavior is configured after training, continuously, by people whose incentives show up in their gross margin.
Jensen Huang put the underlying economics plainly at CES this year: "You sell a chip one time, but when you build software, you maintain it forever." Maintenance is the polite word for change. OpenAI's preview of an Ultrafast tier built on its Cerebras partnership, pushing GPT-5.6 Sol to as much as fourteen times its usual speed, is that same economics stated as a price list. Speed became a purchasable tier. Once speed is purchasable, the speed you get without purchasing is a decision somebody made about you.
This is the condition that gives an argument like Models Are Getting Dumber on Purpose its traction. The post argues that providers quietly trade capability for cost and latency on default endpoints, and a companion piece maps a grey market in resold inference credits feeding the cheap end of that trade. I have no idea whether either thesis holds, and neither does anyone outside the serving teams. That is the actual finding. From the buyer's seat the claim is unfalsifiable, because the product is a behavior rather than an artifact, and behavior can be retuned between one request and the next without a version number moving.
Which is why Anthropic now publishes the production system prompts that ship with each model as release notes, and why that page is more interesting than it sounds. A page like that has a reason to exist only when meaningful parts of a model's behavior are set outside the customer's code, on the vendor's release schedule. Publishing them is a real transparency move and also an admission about where the controls sit.
Bloomberg reported over the weekend that Stripe is nearing a deal to buy OpenRouter for more than seven billion dollars. OpenRouter is the routing service a large number of shipped products use to reach models without integrating each vendor separately, which makes it the place where defaults resolve into actual requests. Pricing policy, terms, and model availability all move with ownership of that position.
Ramp Economics Lab's August index, which reads spend across the businesses carrying Ramp's corporate card and describes that book rather than American business generally, records takeup of the newest frontier model running slower than expected alongside continued movement toward cheaper open source options. That is a purchasing pattern consistent with buyers who have concluded that the difference between tiers is a setting they can approximate for less.
The trade press ran a version of this argument during the move to cloud hosting in the late 2000s, and I remember it reading as overheated at the time: customers were told they had bought capacity when what they had bought was a service level the seller could retune. It turned out to be roughly right, and the correction took years of contract language to arrive. The difference now is tempo. A hosting provider changed your service level inside a maintenance window. A serving stack changes a thinking budget on a config push.
If the reframe holds, thinking budget becomes contract language within a few quarters: an explicit, priced parameter, with pinned versions and scheduled held-out evaluations turning into ordinary hygiene for anyone whose product depends on a model behaving the way it behaved last month. The way to know the reframe was wrong is clean enough. If providers converge on fixed defaults with published behavioral guarantees, and version pins that freeze behavior rather than freeze a name, then reasoning was a model property after all and the exposure was imaginary. The open part sits in between. The dial exists, both sides know it exists, and nobody has decided yet which side of the contract it belongs on.