NVIDIA's 96GB RTX PRO 6000, the fastest Blackwell card that fits under a desk, is selling for around $16,000, roughly double its launch price. Same silicon, same spec sheet. The memory market around it repriced, and the card came along for the ride.
Hours apart, the Qwen team posted Qwen3.8-2.4T-A95B to Hugging Face with open weights: 2.4 trillion total parameters, roughly 95 billion active per token, an FP8 checkpoint alongside. The cost to download it is zero. The number of readers who will serve it on hardware they own is close to zero, and both of those sentences are true at once.
Three more facts stack onto that. The Information's reporting this month has Nvidia testing Rubin Ultra variants with less high-bandwidth memory than the announced spec, because HBM supply will not stretch to cover the original design. The same outlet described a May meeting where AWS leaders asked engineers to conserve EC2 capacity, CPU capacity included, not only the AI accelerators that have been tight for two years.
The artifact everyone can copy is free, and every input that makes the copy useful is rationed. Weights are a file. Memory bandwidth, power interconnects and rack space are queues, and Nvidia's commitment of up to $3 billion to Lancium, the power developer behind the Stargate campus in Texas, is a chipmaker buying its way into the queue instead of waiting in it. Through 2022 and 2023 I watched the same arrangement one layer up, when the trade press ran the model labs as the race while compute quietly set the pace, and the company selling shovels booked the margin because it owned the constraint. That arrangement now reaches the workstation under the desk.
If it holds, consumption migrates to whoever already owns the queue, which is what happened before the day was out. DeepSeek's V4 Pro went live on OpenRouter the same day it existed, no waitlist and no separate account, one tier above the V4 Flash line. The realistic near-term paths for Qwen3.8 are a hosted endpoint or the FP8 quant. For a growing share of readers, an open release is something metered through a router rather than something hosted.
The ecosystem is doing its real work one size class down. Qwen's 27B sibling arrived with a published exact release time, which is an odd amount of ceremony for the small model until you notice it is the one that runs on a machine people already have. Meta's 30B Muse Glimmer, out on Monday, already has a community MLX backend claiming up to 3.3x throughput on Apple silicon and a latency comparison against Qwen3.6 35B on an RTX 5080. Tooling and head-to-head benchmarks inside 48 hours is what adoption looks like. The 2.4 trillion parameter flagship is increasingly an announcement about the size below it.
Falsification is available on two fronts: workstation prices easing back over the autumn, or the 27B landing to indifference while quant and fine-tune traffic stays on the 2.4T checkpoint. The download curves will answer that faster than the launch threads will. The exposure sits with teams whose contingency plan for a hosted-model outage is a local box they have not bought yet. Tuesday's edition took up the spare nobody has tested; the price of that spare moved before anyone got around to testing it.
A fourth thread ran through the same day, all of it about what gets discarded. The FP8 checkpoint is a decision about which precision is expendable. A Rubin Ultra with less memory than announced is a decision about which capacity is expendable. And Sophie Alpert's team policy for AI-written text, surfaced by Simon Willison, starts from the premise that no transformation of natural-language text is lossless, so asking a model to expand three bullets into four paragraphs discards the information the reader actually needed. Three layers of the stack, three compression decisions, and in each one the party choosing what to drop sits upstream of the party who notices it missing.
Two numbers carry the week: 2.4 trillion parameters at no cost, and $16,000 for one card that will run a fraction of them. The confirmation comes with a timestamp attached, the Qwen3.8 27B, scheduled to the minute, sized for the machine already on the desk.