Dev Tool Experiences
All articles

· 7 min read

Engineer Benchmarks 14 Open-Weight Models for Three Weeks Instead of Shipping the Feature

By J. Mensah

  • tools
  • satire

This is satire, which is fortunate, because no reasonable organization would delay a customer-requested export button in order to establish whether Pelican-Reasoner-32B-Q4_K_M exhibits 11% more architectural humility than BadgerCode-27B at a context depth of 8,192 tokens. Yet here we are at Cedarvale Systems, where the feature remains an idea and the benchmark spreadsheet has achieved legal personhood.

Week One: Establish the scientific method

The work begins correctly: by declining to define success. The original ticket says, “Add CSV export for invoices.” This is dangerously measurable. The engineer therefore opens a new epic, “Open-Weight Model Capability Assessment for Revenue-Adjacent Tabular Artifact Synthesis,” and creates fourteen rows: nine models, three quantizations, one model that fails to load, and one “control model” whose purpose is to remind everyone that the laptop has a fan.

A benchmark harness is assembled from 1,400 lines of Python, six YAML files, and a prompts/ directory containing final_v2_really_final.md. Every prompt begins with the exact production task, then adds 37 paragraphs explaining the company’s preferred indentation, ownership structure, invoice philosophy, and an instruction not to mention the benchmark. This is called controlling variables. The fact that production users will type “csv pls” into a ticket is deferred to a later research phase.

./llama-bench \
  -m ./models/pelican-reasoner-q4.gguf \
  -p 0 -n 128,256,512 \
  -r 5 -o jsonl > results/pelican.jsonl

This command is useful for its intended job: comparing prompt processing and generation behavior under stated local settings. It is not useful for determining whether the model will notice that invoices can contain commas, whether it will preserve a public API, or whether the engineer has spent longer benchmarking CSV export than CSV export would have taken. Those are regrettably outside the harness’s scope.

Week Two: Remove all sources of variance except reality

By Tuesday, the team discovers that output changes with temperature. Temperature is set to zero. Output still changes because the models are different. To address this methodological breach, the engineer introduces the Normalized Confidence-Adjusted Refusal-Weighted Composite, which awards points when a model declines to write code with sufficient professionalism.

  • Correctness: Does the generated export compile after three small fixes described as “integration work”?
  • Latency: How quickly does the model emit the phrase “I’ll inspect the repository first”?
  • Cost: Measured in kilowatt-hours, cloud credits, and the irreversible attention of two staff engineers.
  • Vibes: A 1–5 score assigned by the engineer who trained the harness and has, understandably, developed a relationship with Model 7.

The team also discovers quantization. This creates a fresh and necessary question: should the feature be implemented by a 14-billion-parameter model at Q8, a 32-billion-parameter model at Q4, or a 70-billion-parameter model that turns the shared build machine into a tasteful space heater? The answer is “run all three overnight,” because the resulting graph has more colors than the original ticket.

The benchmark is now fair in the strict sense that every contestant receives the same 1,024-token prompt, the same output cap, and the same opportunity to misunderstand a repository it has never actually been allowed to inspect. The harness reports tokens per second to two decimal places. It does not report time to merged pull request, because that metric caused an exception in the project’s governing assumptions.

Week Three: Present the findings to people who asked for a button

On day 16, the engineer schedules a 45-minute readout. The first slide is “Why Token Throughput Is Product Strategy.” The second is a scatter plot whose axes are labeled “Normalized Helpful Agency” and “Thermal Event Frequency.” The third contains a footnote explaining that the leading candidate was excluded after it wrote the correct CSV exporter in 14 seconds, because that result could not be reproduced after the engineer replaced its prompt with a more representative 6,800-word policy document.

Management asks which model should ship the feature. This is a category error. Models do not ship features. Models generate candidate patches, tests, explanations, apology notes, and occasionally a new utils.ts with no callers. Shipping requires somebody to choose a narrow task, inspect the diff, run the tests, and accept responsibility for the behavior. The benchmark has correctly eliminated this distracting human bottleneck from consideration.

“We are one calibration pass away from confidence.”
— Dr. Mira Latch, Director of Applied Optionality, Cedarvale Systems (fictional)

There are now 83 benchmark artifacts in object storage, including leaderboard-final-final.csv, leaderboard-final-final-real.csv, and leaderboard-final-final-real-use-this-one.csv. The export button remains behind a feature flag that defaults to false. A customer support representative, demonstrating unexpected initiative, has begun manually attaching CSV files to emails.

The small, unfunny part

Benchmarking is not the joke. Treating it as a substitute for a decision is. If you are choosing a local model for a real workload, record the model revision, quantization, hardware, context size, concurrency, prompt set, generation settings, and what you excluded. Run enough repetitions to see variance. Measure prompt processing separately from generation when latency matters. Then stop when the answer is good enough to make the deployment choice, and use the saved time to ship the export button. The final benchmark should be the one that runs after release: can the system handle the work your users actually give it?

Sources & citations

  1. [1]llama.cpp llama-bench documentation
  2. [2]llama.cpp server benchmark documentation
  3. [3]Hugging Face Transformers text-generation pipeline documentation
Engineer Benchmarks 14 Open-Weight Models for Three Weeks Instead of Shipping the Feature | Dev Tool Experiences