· 7 min read
Quantization for Developers: What Q4, Q8, and FP8 Mean for Code Quality
By N. Park
- tools
For a local coding model, start with Q4_K_M if memory is the constraint and Q8 if the model already fits with room for your context window; then keep the one that passes your own repo tasks. FP8 is not the “better Q8” download: it is usually a floating-point execution format chosen by an inference or training backend, so it says more about compatible hardware and kernels than a GGUF filename does.
Treat the label as a storage decision first
Q4 and Q8 generally describe how aggressively a model’s weights are compressed. Fewer bits reduce model-file size and memory pressure, which is why Q4 is often the version that lets a larger model run locally at all. But the labels are not a universal quality scale. Q4_K_M, Q4_0, and a vendor’s unrelated “4-bit” checkpoint do not necessarily use the same block layout, scaling scheme, or mix of tensor types.
That suffix matters. In llama.cpp’s current quantizer, Q4_K_M is a K-quant format and Q4_K is an alias for it. The quantizer’s own reference entries list a Llama 3 8B Q4_K_M artifact at 4.58 GB with a reported perplexity change of +0.1754, while Q8_0 is listed at 7.96 GB and +0.0026. Those figures are useful as a direction check, not as a promise about your model, prompt template, language, or coding task. [1]
For code work, the failure mode is rarely a charmingly worse answer. It is a plausible patch that drops a validation branch, changes an identifier in one of six call sites, uses an API that is present in the current file but not in the project’s pinned dependency, or stops before updating a test. A Q4 model can still be very good at small, bounded edits. The risk rises when you ask it to retain requirements across a long plan, search a large repository, and make coupled changes without losing a constraint.
Pick Q4 or Q8 based on the work you delegate
Q4 is the sensible default for local interactive work when it buys you a materially better base model, usable context, or a machine that does not swap itself into unusability. Use it for explaining code, generating tests you will read, locating relevant files, drafting a first patch, and repetitive transformations with a tight verification loop.
Q8 is worth trying when the agent is doing work where a nearly-right answer creates review churn: cross-package refactors, migrations, tool calls with side effects, unfamiliar frameworks, or a long-running task that must keep exact names and constraints straight. This is not because Q8 makes an agent safe. It simply removes one source of approximation while you still need tests, diffs, and permission boundaries.
- Choose Q4 when it enables a larger model or enough context to include the relevant files. A smaller high-precision model is not automatically the better coding assistant.
- Choose Q8 when the same model fits without forcing an impractically small context window or pushing the machine into memory pressure.
- Do not compare Q4 from one publisher with Q8 from another and credit the result entirely to quantization. Model revision, chat template, tokenizer, and quantization method can dominate the comparison.
- Avoid re-quantizing an already quantized GGUF. llama.cpp explicitly warns that
--allow-requantizecan severely reduce quality compared with quantizing from 16- or 32-bit weights. [2]
Run the comparison you will actually trust
Make two model entries with the same base model and only the quantization changed. Keep the context length fixed. This is easy to get wrong: a Q4 build may leave enough VRAM for a larger context or more GPU offload, and then you are comparing both precision and runtime configuration.
# Same context and offload settings; change only the GGUF file
llama-server -m ./models/code-model-Q4_K_M.gguf -c 16384 --n-gpu-layers 99
# Stop it, then run the comparison
llama-server -m ./models/code-model-Q8_0.gguf -c 16384 --n-gpu-layers 99llama-server documents -c as the context setting and --n-gpu-layers for GPU layer offload; its server examples also show the same shape of command. [3] Pick a fixed set of 10 to 20 tasks drawn from the kind of repository work you delegate: a one-file bug fix, a rename spanning tests, an API upgrade, a failing-test diagnosis, and a change requiring a new test. Save the initial prompt and the repository revision for each.
Grade the outcome in an order that makes excuses difficult: did it edit the correct files, does the project compile or type-check, do focused tests pass, did the diff preserve explicit constraints, and how long did review take? Record wall-clock time separately from quality. A Q4 model that needs one extra correction can be the faster tool; a Q4 model that sends you into a 20-minute review of a confidently wrong migration is not.
FP8 is an infrastructure choice, not a GGUF shopping filter
FP8 means eight-bit floating point, not an eight-bit integer-style weight quantization label. NVIDIA’s Transformer Engine describes two FP8 formats: E4M3, with more mantissa precision and a smaller numeric range, and E5M2, with greater range at the cost of precision. Its hybrid recipe uses E4M3 for the forward pass and E5M2 for the backward pass, alongside scaling strategies that keep values representable. [4]
That is why “FP8” needs a second question: FP8 where? A model might be stored in BF16, quantized at load time, and executed with FP8 kernels on a compatible NVIDIA GPU. Or it might be an FP8 checkpoint intended for a specific serving stack. Transformer Engine supports FP8 acceleration for Transformer training and inference on Hopper, Ada, and Blackwell-era NVIDIA GPUs, but the benefit depends on the backend, the GPU, and whether the relevant operators are supported. [5]
For an individual developer, FP8 is usually worth caring about only after you have selected the model and serving backend. Check the backend’s hardware matrix and precision documentation; run its supported FP8 path; then compare it against its BF16 or FP16 baseline on the same code-task set. Do not assume an FP8 path is faster on every GPU, and do not infer code quality from the word alone. The scaling and kernel implementation are part of the result.
A practical default that does not become a policy
Keep Q4_K_M as the laptop or constrained-machine candidate and Q8_0 as the quality-control candidate for the same base model. If Q4 passes your task set and leaves room for a useful context, stop shopping. If failures cluster around exactness—missed call sites, brittle generated tests, or instructions lost halfway through a refactor—try Q8 before blaming the agent’s prompting.
And keep the quantization name in your agent configuration or model registry, not just in somebody’s download history. When a regression lands, “the local coding model got worse” is not actionable. “The Q4_K_M build replaced Q8_0 on this commit, with context reduced from 16,384 tokens to 8,192” is.
Sources & citations
- [1]llama.cpp quantization type definitions and reference size/perplexity entries
- [2]llama.cpp quantization README, including re-quantization warning and quantization commands
- [3]llama.cpp server README and command-line options
- [4]NVIDIA Transformer Engine FP8 primer
- [5]NVIDIA Transformer Engine overview