Skip to content
KoboldCpp
GitHub

GGUF and quantization

KoboldCpp runs models in the GGUF format (files ending in .gguf). A GGUF file contains the model's weights and its metadata in one file. Most GGUF models work in KoboldCpp. A model with a new architecture needs a KoboldCpp version that supports it.

Most models come in several versions of different sizes, called quants. The quant is part of the file name, for example Q4_K_M in gemma-3-4b-it-Q4_K_M.gguf.

  • Start with Q4_K_M or Q4_K_S. The beginner models recommended by KoboldCpp use these.
  • Within the K-quants (Q2_K up to Q6_K) of the same model, a bigger file means better quality and more memory. Pick the biggest one that fits your memory together with your context. See How much memory a model needs.
  • Compare file sizes only within one quant family. Q4_1 is bigger than Q4_K_M but has lower quality, and Q5_1 is bigger than Q5_K_M but has lower quality.
  • Below Q4, quality drops quickly (see the table below).
FamilyNamesNotes
K-quantsQ2_K, Q3_K_S, Q3_K_M, Q3_K_L, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_KThe usual choice. The number is the approximate bits per weight.
Legacy quantsQ4_0, Q4_1, Q5_0, Q5_1, Q8_0Older types. Q8_0 is the largest quant and the closest to the original.
I-quantsIQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ2_M, IQ3_XXS, IQ3_XS, IQ3_S, IQ3_M, IQ4_XS, IQ4_NLThe number is the approximate bits per weight. For the same number, the suffix gives the size: XXS is the smallest, then XS, S and M.
UnquantizedF16, BF1616 bits per weight, about twice the size of Q8_0.

Q3_K, Q4_K and Q5_K without a suffix are other names for Q3_K_M, Q4_K_M and Q5_K_M.

_S, _M and _L (small, medium, large) mix quant types inside one file: some tensors get a larger type. That is why Q4_K_S and Q4_K_M differ in size although both are "Q4_K".

llama.cpp's quantize tool lists these reference values for Llama-3-8B. "Quality loss" is the increase in perplexity compared with the original model; lower is better.

QuantFile sizeQuality loss
Q2_K2.96 GB+3.5199
Q3_K_S3.41 GB+1.6321
Q3_K_M3.74 GB+0.6569
Q3_K_L4.03 GB+0.5562
Q4_04.34 GB+0.4685
Q4_K_S4.37 GB+0.2689
Q4_K_M4.58 GB+0.1754
Q4_14.78 GB+0.4511
Q5_05.21 GB+0.1316
Q5_K_S5.21 GB+0.1049
Q5_K_M5.33 GB+0.0569
Q5_15.65 GB+0.1062
Q6_K6.14 GB+0.0217
Q8_07.96 GB+0.0026

Other models have other file sizes; the table shows how the quants compare with each other.

On the Extra tab, Analyze Model reads the metadata, weight types and tensor names of a GGUF or safetensors file. On the command line (koboldcpp stands for your KoboldCpp file; see Command line):

Terminal
koboldcpp --analyze mymodel.gguf

KoboldCpp does not load text models in safetensors or PyTorch .bin format directly. For most models, someone has already published GGUF files; see Where to get models.

To convert a model yourself:

  1. Download the conversion and quantization tools.
  2. Run convert_hf_to_gguf.py on the model to get a GGUF file. It needs Python 3 with the packages from KoboldCpp's requirements.txt, including PyTorch and transformers.
  3. Run quantize_gguf.exe on that file to make a smaller quant. The tools contain only the Windows exe; on Linux and macOS, build quantize_gguf from the KoboldCpp source with make quantize_gguf.

For example, to make a Q4_K_M quant of a model downloaded into the folder mymodel:

Terminal
python convert_hf_to_gguf.py mymodel --outfile mymodel-16bit.gguf
quantize_gguf.exe mymodel-16bit.gguf mymodel-Q4_K_M.gguf Q4_K_M
  • convert_hf_to_gguf.py <model folder> writes a 16-bit GGUF by default. --outtype picks another type: f32, f16, bf16 or q8_0. For vision models, a second run with --mmproj writes the projector file; this works only for some models.
  • quantize_gguf <input> <output> <type> takes the quant type last. Run it without arguments to list all types.

These are tensor types, not file quants. Each type stores weights in fixed-size blocks, which give its exact bits per weight. A file quant such as Q4_K_M mixes its base type (Q4_K) with larger types for some tensors, so the file's average is slightly higher.

TypeBits per weight
IQ1_S1.5625
IQ1_M1.75
IQ2_XXS2.0625
IQ2_XS2.3125
IQ2_S2.5625
Q2_K2.625
IQ3_XXS3.0625
Q3_K3.4375
IQ3_S3.4375
IQ4_XS4.25
Q4_04.5
Q4_K4.5
IQ4_NL4.5
Q4_15.0
Q5_05.5
Q5_K5.5
Q5_16.0
Q6_K6.5625
Q8_08.5
F16, BF1616

KoboldCpp still loads the older GGML formats that came before GGUF, usually .bin files. Supported are old llama-family files (GGML, GGMF, GGJT) and old GPT-J, GPT-2, RWKV, GPT-NeoX and MPT files.

  • KoboldCpp shows a warning when you load one.
  • Some newer features might be unavailable. Reconvert or re-download the model as GGUF if you can.
  • GPU offload for these formats is built only into the CUDA and ROCm libraries. With Vulkan they run on the CPU.
  • For non-llama legacy formats, the batch size is capped at 256.