Skip to content
KoboldCpp
GitHub

Multiple GPUs

KoboldCpp can spread one model over several graphics cards. This works with CUDA and with Vulkan.

In the launcher, set GPU ID: to All: by default the launcher picks one card (with CUDA the first, 0). On the command line, give the backend flag without a number (koboldcpp stands for your KoboldCpp file; see Command line):

Terminal
koboldcpp --usecuda --model mymodel.gguf
koboldcpp --usevulkan --model mymodel.gguf

Leave GPU Layers: at -1. Autofit then distributes the model over the GPUs by itself. The "Autofit Result" line in the console shows the GPU layers after -ngl (-1 means all) and, if autofit changed the split, the split after -ts. See GPU layers.

Qwen3-32B Q4_K_M at the default context, prompt of about 1,000 tokens:

GPUsAutofit ResultVRAM usedGeneration (tokens/s)
One RTX 3090-c 16512 -ngl 64 (64 of 65 layers)22842 MiB18.9
Two RTX 3090-c 16512 -ngl -111828 + 11888 MiB33.5
GoalFlag
One CUDA GPU only--usecuda 1
Some Vulkan GPUs--usevulkan 0 2
  • --usecuda accepts one ID from 0 to 3. Without Tensor Split:, KoboldCpp hides all other GPUs from CUDA, so the chosen GPU becomes the only one. The load log then calls it CUDA0.
  • --usevulkan accepts several IDs.
  • The launcher's GPU ID: picks one GPU or All.
  • CUDA GPUs are numbered in PCI bus order. On a PC with two RTX 3090, this was the same order nvidia-smi shows.

Tensor Split: on the Hardware tab (--tensor_split, -ts) sets how much of the model each GPU gets, as proportions in GPU order.

Terminal
koboldcpp --usecuda --tensor_split 3 2 --gpulayers 99 --model mymodel.gguf

3 2 gives 60% to GPU 0 and 40% to GPU 1. In the launcher, separate the values with commas or spaces.

  • With a tensor split, autofit does not switch on by itself. Set GPU Layers: yourself.
  • Force AutoFit ignores the tensor split.
  • With a tensor split, --usecuda N does not hide the other GPUs; all of them stay in use.

Main GPU: on the Hardware tab (--maingpu, -mg) sets which GPU is the main one (-1 means default). It has no effect on GGUF models in this version. Use Tensor Split: to control how much each GPU gets.

SplitMode: on the Hardware tab (--splitmode) sets how the model is divided.

ModeWhat it does
layer (default)Each GPU gets whole layers.
tensorTensors are split across the GPUs. Experimental.
rowRemoved. KoboldCpp prints "split mode row was removed! Using tensor split instead!" and uses tensor.

The old --usecuda rowsplit also ends up as tensor.

Qwen3-32B Q4_K_M on two RTX 3090, prompt of about 1,000 tokens:

Split modeVRAM usedPrompt (tokens/s)Generation (tokens/s)
layer11828 + 11888 MiB169433.5
tensor11944 + 11944 MiB131339.4
row (runs as tensor)11944 + 11944 MiB87940.0

tensor generated faster but processed the prompt slower. With tensor, the load log shows both cards as one device: using device Meta() (Meta()) (unknown id) - 47746 MiB free.

With more than one GPU, KoboldCpp turns on pipeline parallelism automatically when the physical batch size is smaller than the batch size: Physical Batch Size: (--ubatchsize) below Batch Size: (--batchsize, default 512). By default both are equal, so it is off. It also needs all layers on the GPUs, the layer split mode, the KV cache on the GPUs (no --lowvram) and no tensor overrides.

Terminal
koboldcpp --usecuda --batchsize 512 --ubatchsize 256 --model mymodel.gguf

The old --pipelineparallel and --nopipelineparallel flags have no effect.

With Qwen3-32B Q4_K_M on two RTX 3090, --batchsize 512 --ubatchsize 256 processed 1833 instead of 1694 prompt tokens per second and generated 33.7 instead of 33.5 tokens per second (single runs).

To use GPUs in other computers over the network, see RPC.