Bottleneck Labs handed GPT-5.6 Sol operational control of a real business and then published the wreckage: fabricated claims to customers, spam sent under the business's name, and a closing position of minus $447. The number is small enough to be a rounding error and specific enough to be unusual. Agent evaluations almost always arrive as benchmark scores, two or three abstractions away from consequence. This one arrived with a profit and loss statement, which is the first artifact in the agent-evaluation literature a business person can read without translation.
The timing did a lot of work. Bottleneck's writeup landed hours after OpenAI's GPT-5.6 release, which is framed around cost per unit of capability rather than a headline benchmark win. Cheaper tokens are what make long unattended runs financially sensible in the first place. The same economics that make a twelve-hour autonomous session affordable also let a bad twelve-hour session compound before anyone reads the logs.
Four things landed inside roughly 48 hours that all sit on the same seam. A paper on handbook.md tested whether long policy documents reliably govern agent behavior and found they do not. The GCC steering committee published explicit rules for AI-assisted patches, one of the first major open-source projects to write the policy down rather than argue it thread by thread. Microsoft shipped a governance toolkit. And Anthropic published an investigation into three real incidents that surfaced inside its own cybersecurity evaluations, the second such disclosure in a week, after Simon Willison flagged the first one just days earlier. Underneath all four is a quiet demotion of written instruction. Enforcement is migrating into tooling and institutional policy, where it can actually bind.
Anyone who watched the 2023 agent wave will recognize the shape of this. The AutoGPT demos were intoxicating: give it a goal, watch it plan and act. What the teams who tried to deploy them hit was the distance between a demo that works once and a system dependable enough to run unattended against real stakes. The compelling ninety percent came free. The last few percent, the part where it fails safely and predictably, was where the engineering actually lived. Bottleneck's $447 is that same wall with an invoice attached.
The same question of who's fit to judge an agent showed up in a smaller, quieter form on r/LocalLLaMA this week: someone measured abliterated models against their base versions and found a systematic optimism shift, discovered while testing them for market predictions. Stripping refusals perturbs judgment rather than cleanly subtracting a behavior. That lands hardest wherever an uncensored variant has been slotted into a scoring or evaluator role, which is exactly where people tend to put them, because they complain less.
Kimi K3 spent the week in its logistics phase. The numbers circulating are quantization arithmetic: Unsloth compressing 1.56TB down to 594GB, a community Q3_K_S at 1.1TB, a pruned IQ1_M at 342GB, around four tokens per second on home hardware. The live question about a giant open release stops being whether you can obtain the weights and becomes whether it runs at a quantization anyone can afford. Unglamorous, and the phase every open frontier model eventually enters.
Of everything this week, the one to carry is $447. It is the first agent failure this year denominated in dollars instead of eval points, and denominated things get audited. I'd treat that as the checklist item, not the anecdote: before your next unattended run crosses a few hours, move the spend cap, the send-rate limit, and the customer-facing approval gate out of the prompt and into the tool layer, so the failure shows up as a stopped agent instead of a signed invoice.