Kimi K3's 2.8T weights finally land as Anthropic's open-weights stance triggers an industry revolt.
Top Signal
Kimi K3's 2.8 trillion-parameter weights are live on Hugging Face
new tool
Simon Willison, r/LocalLLaMA
Moonshot released full weights for Kimi K3, a 2.8 trillion parameter MoE model, following weeks of anticipation. The 1.56TB download is now live on Hugging Face, with GGUF conversions and llama.cpp text-only support already appearing on r/LocalLLaMA hours after the drop. Early operators report the model needs serious multi-GPU infrastructure — one team is testing across A100, H200, and B300 clusters and says the A100 math 'is already rough,' meaning most builders will need cloud inference or a quantized release rather than local hosting. If you were waiting to benchmark Kimi K3 against DeepSeek V4 Flash or your current coding-agent model, the blocker is gone: start with the HF viewer or a hosted endpoint before committing to your own deployment, and watch for quantized/GGUF variants over the next few days if you don't run datacenter-scale hardware.
Read more →
Fast Signals
Minimax-M3 gets vision support merged into llama.cpp
platform change
r/LocalLLaMA
Minimax-M3 can now handle image input through the same local llama.cpp runtime used for text and coding tasks. If you're already running Minimax-M3 locally for agentic work, pull the latest build to add multimodal input without switching runtimes.
Link →
120-run bakeoff compares stock vs. fine-tuned 35B coding agents
research to practice
r/LocalLLaMA
An independent builder ran 120 headless agentic coding runs comparing stock, Ornith-tuned, and KAT-Coder-tuned 35B models on identical llama.cpp workspaces. It's real comparative data rather than leaderboard scores — useful if you're picking a fine-tune for a self-hosted coding agent.
Link →
Anthropic stakes out anti-open-weights position, industry pushes back
emerging signal
HN Front Page, r/LocalLLaMA
Anthropic published its formal position on open-weight models (280 HN points), prompting Nvidia's Jensen Huang to counter that distillation is core to how models learn and to propose an 'Open Secure AI Alliance' — which OpenAI reportedly declined to join, triggering internal backlash. This is escalating from a security-incident story into a policy fight that could affect which models builders are allowed to self-host.
Link →
Mollick's 'which AI to use' guide gets a mid-2026 refresh
workflow
Simon Willison
Ethan Mollick's opinionated model-selection guide has quietly evolved as the field shifted, and Simon Willison flags what's changed. A five-minute read worth it if you default to one model out of habit instead of picking per task.
Link →
Microsoft ships MAI-Cyber-1-Flash, a security-specialized model, in MDASH
new tool
HN Front Page
MAI-Cyber-1-Flash is purpose-built for security analysis workflows, shipped inside Microsoft's MDASH platform. Worth evaluating if you're building anything that triages vulnerabilities, logs, or incident data — a specialized model beats prompting a generalist for this category.
Link →
Quantizing Qwen 3.6 27B breaks the pelican benchmark
research to practice
r/LocalLLaMA
A r/LocalLLaMA thread tests whether standard quantizations of Qwen 3.6 27B degrade output using Simon Willison's informal 'draw a pelican on a bicycle' SVG test. Anecdotal, but a reminder to spot-check actual task quality after quantizing — not just perplexity — before shipping a quantized model.
Link →
Radar
Nifer: custom inference engine hits 700 t/s on RTX 5090
An individual dev's purpose-built inference engine claims 700 tokens/sec on non-thinking Qwen 3.6 35B with full 250k context on a single RTX 5090. If verified, that's a meaningful jump over general-purpose runtimes for anyone serving a fixed model on consumer hardware.
Link →
FeyNoBg: open-source background removal model + training lib
A small startup open-sourced both a background-removal model and NoBg, the Python library used to train and run it. Useful for self-hosting image preprocessing instead of paying per-call to a background-removal API.
Link →
On-device local LLM agent plays Perfect Dark via MLX
A hobbyist project has a local LLM driving real-time locomotion and combat in Perfect Dark on a Mac, fully on-device with no cloud calls. A concrete data point on how far on-device agents have come for latency-sensitive, real-time control beyond chat/coding.
Link →
Convergence Watch
kimi k3
TRENDING
8 mentions across Simon Willison, r/LocalLLaMA
Kimi K3 has appeared in the entity feed 4 of the last 7 days as anticipation built; today's actual weight release confirms it as the week's clearest 'big model launch.' Infra reports (A100/H200/B300 testing) suggest most builders will need cloud hosting rather than local deployment at 2.8T params.
open-weight ai policy debate
TRENDING
8 mentions across HN Front Page, r/LocalLLaMA
This storyline (Anthropic's stance, the HuggingFace incident fallout, Chinese open-weight sanctions) has generated cross-source volume for four-plus consecutive days. It's shifting from a security-incident story into a policy fight over open-weight regulation — worth watching for concrete regulatory action rather than more position papers.
STALE: Latent Space newest item is >48h old