· 7 min read
Man Who Said Open-Weight Models Would Never Catch Up Quietly Switches Plan Mode to Kimi K3
By H. Iyer
- tools
- satire
This is satire, and therefore a faithful record of the quarterly planning meeting at Orbital Casserole Systems, where principal engineer Martin Vellum confirmed that open-weight models would never catch up, then changed his plan-mode model to Kimi K3 before the coffee machine had completed its firmware update.
Vellum had been a reliable source of model-market discipline. For eighteen months, he had explained that closed models possessed an unassailable advantage called “the stuff you cannot download,” accompanied by a hand gesture suggesting a moat, a cathedral, and perhaps a subscription invoice. He was not against open weights, exactly. He simply believed they belonged in the same category as mechanical keyboards with artisan keycaps: admirable private projects whose users should not be permitted near an incident review.
The change was operational, not ideological
At 9:14 a.m., Vellum opened the coding CLI, typed /model, and selected k3 for plan work. This was not a reversal. It was a temporary routing adjustment made under laboratory conditions, where “laboratory conditions” meant a ticket titled “trace why checkout computes tax twice when a coupon is removed.” The implementation model remained unchanged, he stressed, because code generation is serious business. Planning, meanwhile, was a low-stakes exercise involving repository search, dependency tracing, migration sequencing, and the creation of a 47-item checklist nobody would read until after the outage.
Kimi’s documentation says the CLI can switch models with /model, and lists k3 and k3-256k as model IDs rather than version names. Vellum appreciated this distinction. He had previously spent forty minutes arguing that “K3” was a brand, k3 was an identifier, and typing the wrong capitalization was the kind of operational sloppiness that explained why the industry could not be trusted with autonomous agents.
A carefully managed test
The evaluation protocol was rigorous. First, Vellum asked the model to inspect the billing service. Then he asked it to propose a plan. Then he asked it to identify files affected by a small validation change. When it named several files he had forgotten existed, he declared the result “interesting but not conclusive,” which in engineering terminology means the spreadsheet has been updated but not shared.
By lunch, K3 had been assigned the task of explaining an eight-month-old feature flag. Vellum described this as “giving it harmless reconnaissance.” The flag controlled refunds in three regions, populated a Kafka topic, bypassed a cache on Tuesdays, and was last edited by an employee whose account now forwards mail to a tasteful memorial page. The model returned a plan with risks, affected tests, and a suggested rollout order. Vellum rejected one filename, accepted the rest, and announced that the experiment had revealed a useful limitation: the model lacked context about why the company had invented a second currency for promotional credits.
This was true. It had no context because there was no context. The design document was a FigJam containing 63 yellow sticky notes and one arrow labeled “legal?”
The open-weight position remains fully intact
By 2:03 p.m., Vellum had developed a framework for preserving his earlier statement. Open-weight models had not caught up in general. K3 had merely caught up in the narrow, highly specialized domain of the exact work he wanted done today: reading a TypeScript monorepo, noticing that a GraphQL resolver ignored a nullable field, and producing a plan that did not begin with “I’d be happy to help.” This did not count as catching up. It counted as an exception, and exceptions are how engineering maintains intellectual continuity through changing evidence.
The company’s Architecture Alignment Council approved this interpretation. Its minutes recorded that “model choice must reflect workload characteristics, security requirements, latency, cost, availability, and Martin’s deeply held beliefs as of the previous Thursday.” A follow-up action item required the platform team to add a provider abstraction so the organization could switch models immediately after deciding it would never switch models again.
- Plan mode: K3, because repository investigation benefits from a model that can hold onto a thread longer than a stand-up.
- Act mode: the previous model, because changing two things at once would be irresponsible.
- Review mode: whichever model makes the fewest claims about having “carefully considered” a diff it saw 800 milliseconds ago.
- Incident mode: a human, at least until someone finds a benchmark for apologizing in Slack.
What the experiment did not prove
It did not prove that an open-weight model is automatically the correct tool for every team. Downloadable weights are not a substitute for capacity planning, evaluation on your codebase, access controls, routing policy, observability, or somebody being on call when the inference stack decides that the GPU is now a decorative heat source. The official Kimi K3 release describes it as an open-weight model, not a promise that every developer can casually operate it on spare hardware between meetings.
Nor did it prove that plan mode is magic. Models can map a codebase and still fail to understand the organizational artifact hidden inside it: the undocumented agreement that nobody touches LegacyRefundComposer because it calls a vendor endpoint that only responds correctly when addressed in a tone of respectful uncertainty. A plan can be structured, thorough, and wrong in ways that are extremely expensive to discover after the agent has confidently edited 38 files.
The final status update
At 4:51 p.m., Vellum posted that the trial had produced “promising early signals,” the phrase organizations use when a person’s prior certainty has encountered a reproducible command. He did not mention the model switch in the channel where he had once predicted that open weights would remain permanently behind. Instead, he opened a pull request with a plan attached and asked for review.
One true observation remains after the satire: model allegiance is a poor operating policy. Keep the task, the harness, the permissions, the context you provide, and the actual output under review. Then use the model that helps on that work today—and be prepared to change it without needing a philosophical evacuation plan.