Dev Tool Experiences
All articles

· 7 min read

A Field Guide to the Quantization Levels You Pretend to Understand

By J. Zhou

  • tools
  • satire

This is satire, and therefore the most honest possible guide to model quantization labels. You have downloaded a file with Q4_K_M in its name, watched it answer “write a Rust CLI,” and are now prepared to explain numerical representation to the organization. Congratulations: you have joined the ancient guild of developers who know that lower bits use less memory, plus a foggy but forceful belief that _M means “the good one.”

First, identify which language of quantization you are pretending to speak

The first rule is that labels which look interchangeable are not interchangeable. Q4_K_M is a llama.cpp/GGUF-oriented file quantization label. NF4 is a 4-bit datatype option commonly encountered through bitsandbytes and QLoRA workflows. INT8 may describe a weight representation, a runtime path, an export target, or the answer given by someone who has already opened the GPU utilization chart and wants the meeting to end.

Do not say, “We should use 4-bit.” This is like saying, “We should use four containers.” Four of what, packed how, scaled per which block, executed by which kernel, and on hardware purchased during which fiscal error? In llama.cpp’s quantizer source, the named formats carry distinct effective bits-per-weight and documented reference perplexity deltas; Q4_K_M is not simply four literal bits glued to every weight. The implementation is doing enough bookkeeping that the underscore deserves to be treated as a tiny warning label.

The levels, as practiced in the wild

Q8_0: the reconciliation quant

Choose Q8_0 when you need to say you quantized the model but do not need anyone to notice. In llama.cpp’s own quantizer table, Q8_0 is listed at roughly 7.96G for the cited Llama 3 8B reference and with a very small listed perplexity change. It is the business-casual option: still substantial, still expensive compared with lower-bit variants, but unlikely to start explaining database migrations in haiku.

Q8_0 is bad at making an oversized model fit into a machine that cannot hold it. It is also bad at producing an interesting internal postmortem, because “we used more memory and got fewer surprises” cannot be rendered as a glowing architecture diagram.

Q6_K: the person who owns a test set

Q6_K is for the engineer who has seen a lower-bit model fail on a real repository and now wears a thousand-yard stare whenever somebody says “quality is subjective.” It buys back accuracy at a real storage and memory cost; llama.cpp documents Q6_K at 6.14G in that same Llama 3 8B reference table, versus 4.58G for Q4_K_M. That is not a rounding error when your laptop’s available memory is being negotiated by an editor, a browser, Docker, and a video call named “Quick Sync.”

Q6_K is bad at winning the argument that a model should run on an 8 GB device after the device has also been asked to maintain a large context window. Quantized weights are not the entire memory bill. This fact will be rediscovered at 4:47 p.m. by whoever runs the command first.

Q4_K_M: the default you call a decision

Q4_K_M is the practical center of the local-model universe: small enough that people can run it, capable enough that people can blame the prompt instead of the format, and common enough that copying a filename feels like research. In the llama.cpp source, Q4_K_M is identified as a mixed K-quant variant; the related tensor encoding documentation describes Q4_K at 4.5 bits per weight before the model-level mixture and metadata make the final file label less tidy than its name suggests.

This is the level at which you will say, “It feels basically the same,” after trying two questions: one about TypeScript and one about the moon. It is bad at the tasks where small degradations are not small: structured extraction, exact code transformations, multilingual edge cases, or the peculiar set of tests your CI uses to protect a 2018 billing rule from ever being understood.

Q3 and Q2: the compact emergency provisions

Below four bits, confidence changes form. At Q3, you are optimizing for the right to run the model at all. At Q2, you are conducting a controlled experiment in whether a system can retain enough of a language model to continue producing markdown headings. llama.cpp’s reference entries list Q3_K_M at 3.74G and Q2_0 at 2.25 bits per weight, but these numbers are invitations to test, not permission slips to declare quality preserved.

The proper use of these formats is constrained hardware, rough exploration, and proving that the runtime wiring works. Their improper use is silently replacing the model that edits production code because the smaller file produced a reassuringly fast first token.

NF4: not a GGUF setting wearing a nicer hat

NF4 means Normal Float 4, a 4-bit datatype designed around normally distributed weights and prominently used for 4-bit training workflows such as QLoRA. In Transformers, it is selected with bnb_4bit_quant_type="nf4"; the docs specifically recommend it for training 4-bit base models. That does not mean an NF4 setting is an interchangeable substitute for a GGUF file you found at 1:12 a.m. They belong to related conversations, which is not the same thing as the same pipeline.

from transformers import BitsAndBytesConfig
import torch

quantization_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.bfloat16,
)

NF4 is bad at solving deployment questions by itself. You still need a compatible backend, a supported device, enough memory for the actual workload, and a reason to fine-tune rather than merely changing the system prompt until it stops mentioning Kubernetes.

The meeting-ready decision procedure

  1. Pick the runtime first. A GGUF file, a Transformers checkpoint with bitsandbytes configuration, and a vendor-specific engine are not three ways to select the same dropdown.
  2. Measure the workload you actually care about: repository edits, tool calls, JSON validity, test repair, latency at your intended context length. The “tell me a joke” benchmark has completed its work and may go home.
  3. Start at a format that fits with operational headroom, then step up if your representative tasks regress. Do not tune bits from a filename alone.
  4. Record the exact model revision, runtime version, quantization label, context length, hardware, and test prompts. Six weeks from now, “the Q4 one” will not be reproducible evidence.
  5. Keep one higher-precision reference run for failures that matter. This is cheaper than learning, after an incident, that the model had been compressing your requirements along with its weights.

The true observation beneath the joke is simple: quantization is not a personality test or a leaderboard badge. It is a systems trade-off. The right level is the smallest representation that meets your measured quality, latency, and memory requirements on the runtime and hardware you will actually use.

Sources & citations

  1. [1]llama.cpp quantizer source and reference quantization table
  2. [2]llama.cpp tensor encoding schemes
  3. [3]Hugging Face Transformers: bitsandbytes quantization documentation