Every flag KoboldCpp 1.122.1 accepts, generated from its source code. Run koboldcpp --help to see the same list in a terminal.
Flags go after the file name, for example koboldcpp.exe --model mymodel.gguf --contextsize 8192. Most flags also exist in the launcher; each entry names the tab and label.
Most used
Section titled “Most used”| Flag | What it does |
|---|---|
--usecuda | Runs the model on an NVIDIA GPU with CUDA (or an AMD GPU with hipBLAS/ROCm in ROCm builds). |
--usevulkan | Runs the model on a GPU through Vulkan, which works with NVIDIA, AMD and Intel GPUs. |
--usecpu | Runs the model on the CPU only, without any GPU acceleration. |
--model | The text model (GGUF file) to load. |
--config | Loads all settings from a saved .kcpps config file (or .kcppt template), from disk or a URL. |
--contextsize | How many tokens of conversation the model can keep in memory at once (default 16384). |
--gpulayers | How many model layers to put on the GPU; more layers is faster but needs more VRAM. |
--host | The network address KoboldCpp listens on. |
--launch | Opens the web UI in your browser once the model has loaded. |
--threads | How many CPU threads to use for generation. |
--port | The port the web UI and API listen on (default 5001). |
--version | Prints the KoboldCpp version and exits. |
--mmproj | Loads the multimodal projector file that lets a vision (or audio) model understand images (or sound). |
--password | Requires an API key for text generation and most other endpoints. |
--remotetunnel | Creates a public Cloudflare URL so you can reach KoboldCpp from anywhere, even behind a firewall. |
--sdmodel | The image generation model (safetensors or GGUF) to load; this is what turns image generation on. |
--whispermodel | Loads a Whisper model to turn speech into text. |
--ttsmodel | Loads a text-to-speech (TTS) model so KoboldCpp can read text aloud. |
--ttswavtokenizer | The WavTokenizer file that some TTS models need. |
Basics
Section titled “Basics”--usecuda
Section titled “--usecuda ”Also: --usecublas, --usehipblas · Value: [main GPU ID] · Launcher: Quick Launch → Backend (Use CUDA / Use hipBLAS (ROCm)) and GPU ID
Runs the model on an NVIDIA GPU with CUDA (or an AMD GPU with hipBLAS/ROCm in ROCm builds).
With no number, all GPUs are used. A number from 0 to 3 uses only that GPU, unless --tensor_split is also set; then all GPUs stay in use. If your build has no CUDA library, KoboldCpp falls back to the CPU without an error; check the Initializing dynamic library: line in the terminal.
Help text: Use CUDA for GPU Acceleration. Requires CUDA. Enter a number afterwards to select and use 1 GPU. Leaving no number will use all GPUs.
--usevulkan
Section titled “--usevulkan ”Value: [Device IDs] · Launcher: Quick Launch → Backend (Use Vulkan) and GPU ID
Runs the model on a GPU through Vulkan, which works with NVIDIA, AMD and Intel GPUs.
With no IDs, all devices are used; give one or more device IDs (e.g. --usevulkan 0) to pick specific GPUs. If the GGML_VK_VISIBLE_DEVICES environment variable is set, it wins over the IDs. If your build has no Vulkan library, KoboldCpp falls back to the CPU without an error.
Help text: Use Vulkan for GPU Acceleration. Can optionally specify one or more GPU Device ID (e.g. --usevulkan 0), leave blank to autodetect.
--usecpu
Section titled “--usecpu ”Launcher: Quick Launch → Backend (Use CPU)
Runs the model on the CPU only, without any GPU acceleration.
Any positive --gpulayers value is ignored with a warning. On macOS this flag does not stop Metal GPU offload.
Help text: Do not use any GPU acceleration (CPU Only)
--model
Section titled “--model ”Also: -m · Value: [filenames] · Launcher: Quick Launch → GGUF Text Model
The text model (GGUF file) to load.
You can also pass the file as the first plain argument: koboldcpp model.gguf. An https:// URL ending in .gguf is downloaded first (all parts of a split Hugging Face model are fetched); an existing file with the same name is reused. Extra URLs after the first are only downloaded, never loaded.
Help text: Model file to load. Accepts multiple values if they are URLs.
--config
Section titled “--config ”Value: [filename] · Launcher: Bottom bar → Load Config
Loads all settings from a saved .kcpps config file (or .kcppt template), from disk or a URL.
The help text is wrong: other flags are not ignored. A flag overrides the config value only when written exactly as --<name> (e.g. --contextsize 4096 or --quiet); short aliases like -c and the --flag=value form are silently ignored. An explicit --port 5001 does not override the config's port, because 5001 is the default.
Help text: Load settings from a .kcpps file. Other arguments will be ignored
--contextsize
Section titled “--contextsize ”Also: --ctx-size, -c · Value: [256 to 524288] · Default: 16384 · Launcher: Quick Launch → Context Size
How many tokens of conversation the model can keep in memory at once (default 16384).
Larger values use more RAM/VRAM. Allowed range is 256 to 524288; the launcher slider stops at 262144. Requests can never use more context than this, and the terminal may show a slightly larger number because of internal padding.
Help text: Controls the memory allocated for maximum context size, only change if you need more RAM for big contexts. (default 16384).
--gpulayers
Section titled “--gpulayers ”Also: --gpu-layers, --n-gpu-layers, -ngl · Value: [GPU layers] · Default: -1 · Launcher: Quick Launch → GPU Layers
How many model layers to put on the GPU; more layers is faster but needs more VRAM.
-1 (default) estimates the layers when a GPU backend is found, or becomes 0 when none is. It also turns on autofit, unless --tensor_split, --overridetensors, --moecpu or --ffncpu is set. On macOS, -1 puts all layers on the GPU and does not turn on autofit. 0 keeps everything on the CPU and also skips automatic backend detection. --gpulayers with no number means 1 layer.
Help text: Set number of layers to offload to GPU (when using GPU). Set to -1 to enable autofit (default), set to 0 to disable GPU offload.
--host
Section titled “--host ”Value: [ipaddr] · Launcher: Network → Host
The network address KoboldCpp listens on.
By default it listens on all interfaces, so other devices on your network can reach it. Use --host 127.0.0.1 to allow only this computer. In router mode (--routermode or --autoswapmode) KoboldCpp listens on all interfaces whatever --host says, so use a firewall to keep it private. If you do expose it, set --password.
Help text: Host IP to listen on. If this flag is not set, all routable interfaces are accepted.
--launch
Section titled “--launch ”Launcher: Quick Launch → Launch Browser
Opens the web UI in your browser once the model has loaded.
Off by default on the command line, but ticked by default in the launcher. Cannot be combined with --cli.
Help text: Launches a web browser when load is completed.
--tensor_split
Section titled “--tensor_split ”Also: --tensorsplit, --tensor-split, -ts · Value: [Ratios] · Launcher: Hardware → Tensor Split
Sets how the model is divided across several GPUs, as a list of proportions such as 7 3.
With a split set, --usecuda N no longer limits KoboldCpp to one GPU. Setting a split stops autofit from turning on automatically, and a forced --autofit ignores it. Up to 16 values are used.
Help text: For CUDA and Vulkan only, ratio to split tensors across multiple GPUs, space-separated list of proportions, e.g. 7 3
--threads
Section titled “--threads ”Also: -t · Value: [threads] · Default: get_default_threads() · Launcher: Hardware → Threads
How many CPU threads to use for generation.
The default is about half your CPU threads minus one (1 to 3 on small CPUs), capped at 8 on CPUs that report as Intel and at 64 overall; 0 or a negative value means automatic. It is also the default for --blasthreads, --sdthreads and --ttsthreads.
Help text: Use a custom number of threads if specified. Otherwise, uses an amount based on CPU cores
--port
Section titled “--port ”Value: [portnumber] · Default: 5001 · Launcher: Network → Port
The port the web UI and API listen on (default 5001).
You can also pass it as a plain argument after the model: koboldcpp model.gguf 5002. If the port is already in use, the model loads fully and then startup fails. A config file's port wins over an explicit --port 5001.
Help text: Port to listen on. (Defaults to 5001)
--version
Section titled “--version ”Prints the KoboldCpp version and exits.
Only works as the sole argument; combined with other flags it is ignored and KoboldCpp starts normally.
Help text: Prints version and exits.
Advanced
Section titled “Advanced”--agent
Section titled “--agent ”Launcher: Admin → Launch KoboldCpp Agent
Starts the bundled KoboldCpp Agent, a terminal-based assistant, in a new terminal window.
Alone it opens only the agent, which connects to http://127.0.0.1:5001/v1. With a model it also starts the server and raises the context to at least 28672, --defaultgenamt to at least 8192, and turns on --jinja and --jinja_tools. Add --cli to run the agent in the current terminal.
Help text: Launches the simple KoboldCpp Agent in a new terminal window, or the current terminal with --cli.
--analyze
Section titled “--analyze ”Value: [filename] · Launcher: Extra → Analyze Model
Prints the metadata, weight types and tensor names of a GGUF or safetensors file, then exits.
Takes a file path; nothing is loaded or served.
Help text: Reads the metadata, weight types and tensor names in any GGUF or safetensors file.
--autofit
Section titled “--autofit ”Also: --fit, -fit · Launcher: Quick Launch → Force AutoFit
Forces autofit, which works out the best way to split the model between GPU and CPU.
Autofit is already on by default when --gpulayers is -1 and a GPU backend is used, unless --tensor_split, --overridetensors, --moecpu or --ffncpu is set or you are on macOS. Forcing it overrides --gpulayers (even 0) and discards --tensor_split, --overridetensors, --moecpu and --ffncpu. If fitting fails, the normal layer estimate is used.
Help text: Forces autofit, which attempts to fit the model in the best possible way. Overrides everything else.
--autofitpadding
Section titled “--autofitpadding ”Value: [padding in MB] · Default: 1024 · Launcher: Hardware → Autofit Padding (MB)
How much VRAM, in MB, autofit keeps free as a safety margin (default 1024).
Only applies when --autofit is set explicitly; the automatic autofit always uses 1024. Raise it if loading fails with out-of-memory errors.
Help text: How much spare allowance in MB should autofit reserve? If it's too little, the load might fail.
--batchsize
Section titled “--batchsize ”Also: --blasbatchsize, --batch-size, -b · Value: -1 | 16 | 32 | 64 | 128 | 256 | 512 | 1024 | 2048 | 4096 · Default: 512 · Launcher: Hardware → Batch Size
How many prompt tokens are processed at once (default 512); larger is faster but uses more memory.
Only the listed values are accepted. -1 processes one token at a time but keeps GPU offload. Old non-llama model formats are capped at 256.
Help text: Sets the logical batch size used in batched processing (default 512). Setting it to -1 disables batched mode, but keeps other benefits like GPU offload.
--ubatchsize
Section titled “--ubatchsize ”Also: --ubatch-size, -ub · Value: -1 | 16 | 32 | 64 | 128 | 256 | 512 | 1024 | 2048 | 4096 · Default: -1 · Launcher: Hardware → Physical Batch Size
The physical batch size actually sent to the backend in one step.
-1 (default) uses the same value as --batchsize; larger values are lowered to the batch size.
Help text: Sets the physical batch size used in batched processing. Setting it to -1 uses the same value as batchsize (default).
--benchmark
Section titled “--benchmark ”Value: [filename] · Launcher: Hardware → Run Benchmark
Runs a speed benchmark instead of starting the server, then exits.
It fills the whole context with a synthetic prompt and generates --genlimit tokens (100 if unset), then prints processing and generation speed. Give a filename to append the results as a CSV row. Needs a model; cannot be combined with --cli.
Help text: Do not start server, instead run benchmarks. If filename is provided, appends results to provided file.
--blasthreads
Section titled “--blasthreads ”Also: --batchthreads, --threadsbatch, --threads-batch · Value: [threads] · Launcher: Hardware → Batch Threads
How many CPU threads to use while processing the prompt.
0 (default) uses the same value as --threads.
Help text: Use a different number of threads during batching if specified. Otherwise, has the same value as --threads
--chatcompletionsadapter
Section titled “--chatcompletionsadapter ”Value: [filename] · Default: AutoGuess · Launcher: Loaded Files → Chat Adapter
Sets the instruct format (chat tags) used for chat-style API requests.
The default AutoGuess detects the format from the model's chat template and falls back to Alpaca if nothing matches. You can give a JSON file path or the name of a bundled adapter such as ChatML or Llama-3. With --jinja, the model's own template is used instead.
Help text: Select an optional ChatCompletions Adapter JSON file to force custom instruct tags.
--cli
Section titled “--cli ”Launcher: Hardware → CLI Terminal Only
Chat with the model directly in the terminal, without starting the web server.
Type /quit or /exit to leave. --prompt becomes the system prompt, and --jinja is not used. Cannot be combined with --launch, --nomodel, --admin or --benchmark.
Help text: Runs interactive terminal chat without an HTTP server. With --agent, runs the agent in this terminal using a loopback HTTP server.
--debugmode
Section titled “--debugmode ”Value: int · Launcher: Hardware → Debug Mode
Prints extra diagnostic information in the terminal, including full request inputs.
The flag alone means 1; use 1 rather than higher values. It also turns off --quiet. KoboldCpp sets -1 (fewer messages) automatically when Horde worker flags are used.
Help text: Shows additional debug info in the terminal. Levels: -1 (Horde-quiet, suppresses non-essential prints; auto-applied when Horde args are set), 0 (default, normal output), 1 (verbose: extra slot/cache info, larger print buffers, retains horde-debug prefix). Passing the flag without a value implies 1.
--defaultgenamt
Section titled “--defaultgenamt ”Value: check_range(int, 64, 32768) · Default: 2048 · Launcher: Context → Default Gen Amt
How many tokens to generate when a request does not say how many (default 2048).
Most frontends send their own length, so this mainly affects bare API calls. The value is silently lowered to half the context size (e.g. 256 with --contextsize 512); the help's "must be smaller than context size" is outdated. Allowed range is 64 to 32768.
Help text: How many tokens to generate by default, if not specified. Must be smaller than context size. Usually, your frontend GUI will override this.
--device
Section titled “--device ”Also: -dev · Value: <dev1,dev2,..> · Launcher: Hardware → Device Override
Chooses exactly which backend devices to use, by llama.cpp device name, comma-separated.
CPU devices and unknown names are rejected, and one bad name cancels the whole override; loading then continues normally. none means no override. Only the text, TTS and embeddings models use it.
Help text: Set llama.cpp compatible device selection override. Comma separated. Overrides normal device choices.
--downloaddir
Section titled “--downloaddir ”Value: [directory] · Launcher: Loaded Files → Download Dir
The folder where models given as URLs are downloaded.
If a file with the same name is already there, it is reused instead of downloaded again. It does not change where local file paths are looked up. Unset means the current working directory.
Help text: Specify a directory that models will be downloaded to or searched from, if unset uses the working directory.
--draftamount
Section titled “--draftamount ”Also: --draft-max, --draft-n, --spec-draft-n-max · Value: [tokens] · Default: 4 · Launcher: Loaded Files → Draft Amount
How many tokens the draft model (or MTP) guesses ahead per step in speculative decoding (default 4).
0 or less turns speculative decoding off.
Help text: How many tokens to draft per chunk before verifying results
--draftgpulayers
Section titled “--draftgpulayers ”Also: --gpu-layers-draft, --n-gpu-layers-draft, -ngld · Value: [layers] · Default: 999 · Launcher: Loaded Files → Layers (Draft Model row)
How many layers of the draft model go on the GPU (default 999, meaning all).
Help text: How many layers to offload to GPU for the draft model (default=full offload)
--draftgpusplit
Section titled “--draftgpusplit ”Value: [Ratios] · Launcher: Loaded Files → Splits (Draft Model row)
How the draft model is divided across several GPUs, as a list of proportions.
Without it the draft model gets llama.cpp's default split, not the main model's ratios as the help says. It works without --tensor_split, but the GPUs must be visible: --usecuda N without --tensor_split hides all other GPUs.
Help text: GPU layer distribution ratio for draft model (default=same as main). Only works if multi-GPUs selected for MAIN model and tensor_split is set!
--draftmodel
Section titled “--draftmodel ”Also: --model-draft, -md · Value: [filename] · Launcher: Loaded Files → Draft Model
Loads a small draft model for speculative decoding, which can speed up generation.
The draft model must use the same vocabulary as the main model; a large mismatch turns drafting off. A missing or broken draft file does not stop startup, it only prints an error. It takes priority over --usemtp and turns off --parallelrequests batching.
Help text: Load a small draft model for speculative decoding. It will be fully offloaded. Vocab must match the main model.
--enableguidance
Section titled “--enableguidance ”Launcher: Context → Enable Guidance
Enables classifier-free guidance, so requests can use a negative prompt.
Creates a second context, which roughly doubles the memory used for context. It only takes effect for requests with a negative_prompt and a guidance_scale other than 1. It turns off --parallelrequests batching.
Help text: Enables the use of Classifier-Free-Guidance, which allows the use of negative prompts. Has performance and memory impact.
--exportconfig
Section titled “--exportconfig ”Value: [filename] · Launcher: Bottom bar → Save Config
Saves the given flags as a .kcpps config file and exits without loading anything.
.kcpps is added to the name if missing. Values worked out at load time, such as automatic GPU layers or backend, are not saved.
Help text: Exports the current selected arguments as a .kcpps settings file
--exporttemplate
Section titled “--exporttemplate ”Value: [filename] · Launcher: Extra → Generate LaunchTemplate
Saves the given flags as a shareable .kcppt launch template and exits.
The backend, GPU layers, threads, main GPU, GPU splits, passwords, SSL and Horde worker name and key are reset, and the backend is auto-detected when the template is loaded. Other machine-specific settings are kept, such as --device, --blasthreads, --draftgpulayers, --noavx2 and --failsafe; check them before sharing. Unlike the launcher button, the command line does not embed the chat adapter or preload story files.
Help text: Exports the current selected arguments as a .kcppt template file
--failsafe
Section titled “--failsafe ”Launcher: Quick Launch → Backend (Failsafe Mode (Older CPU))
Compatibility mode for very old CPUs; much slower.
Selected automatically when the CPU has neither AVX nor AVX2. It implies --noavx2 and prevents CUDA from loading. Add --usecpu to make sure it runs on the CPU; otherwise a GPU backend can still be selected automatically.
Help text: Use failsafe mode, extremely old CPU compatibility mode that should work on all devices.
--foreground
Section titled “--foreground ”Launcher: Hardware → Keep Foreground
On Windows, brings the terminal window to the front on each new generation to avoid idle slowdowns.
Does nothing on Linux and macOS.
Help text: Windows only. Sends the terminal to the foreground every time a new prompt is generated. This helps avoid some idle slowdown issues.
--gendefaults
Section titled “--gendefaults ”Value: {"parameter":"value",...} · Launcher: Context → Default Params
Default values, as a JSON object, that are added to API requests when the request leaves them out.
Example: --gendefaults '{"max_length":512}'. It applies to text, image, audio and embeddings requests; key names must match that endpoint's request fields. If the JSON cannot be parsed, a warning is printed on every request and no defaults apply.
Help text: Sets extra default parameters for some fields in API requests, as a JSON string.
--gendefaultsoverwrite
Section titled “--gendefaultsoverwrite ”Launcher: Context → Override (next to Default Params)
Makes --gendefaults values replace the values sent in requests, not only fill in missing ones.
Help text: Allow the gendefaults parameters to overwrite the original value in API payloads.
--genlimit
Section titled “--genlimit ”Also: --promptlimit · Value: [token limit] · Launcher: Context → Prompt Limit
A hard upper limit on generated tokens for every request.
Despite the --promptlimit alias and the Prompt Limit label, it limits output, not the prompt. 0 means no limit. It also sets the length for --prompt and --benchmark.
Help text: Sets the maximum number of generated tokens, it will restrict all generations to this or lower. Also usable with --prompt or --benchmark.
--highpriority
Section titled “--highpriority ”Launcher: Hardware → High Priority
Tries to raise KoboldCpp's CPU priority.
On Windows it sets the process to realtime priority. On Linux it currently lowers the process priority slightly (nice 0 to 1) instead of raising it, and its request for higher thread priority fails unless you have the privileges for it.
Help text: Experimental flag. If set, increases the process CPU priority, potentially speeding up generation. Use caution.
--ignoremissing
Section titled “--ignoremissing ”Skips missing files with a warning instead of stopping.
Covers the config file and the model files (text, LoRA, mmproj, image, Whisper, TTS, embeddings, music). A missing text model is skipped too; if nothing is left to load, KoboldCpp exits instead of starting.
Help text: Ignores all missing non-essential files, just skipping them instead.
--jinja
Section titled “--jinja ”Launcher: Quick Launch → Use Jinja
Formats chat requests with the model's own chat template instead of KoboldCpp's adapter tags.
Applies to chat-style endpoints (OpenAI chat, Ollama chat, Anthropic Messages); other endpoints are unaffected. Requests with tools fall back to the adapter unless --jinja_tools is set. Turned on automatically by --jinja_tools, --jinja_kwargs and --jinjathink.
Help text: Enables using jinja chat template formatting for chat completions endpoint. Other endpoints are unaffected. Tool calls are done without jinja.
--jinja_kwargs
Section titled “--jinja_kwargs ”Also: --jinja-kwargs, --jinjakwargs, --chat-template-kwargs · Value: {"parameter":"value",...} · Launcher: Context → Jinja Kwargs
Extra variables, as a JSON object, passed to the Jinja chat template.
Example: {"enable_thinking":false}. A request's own chat_template_kwargs override these. Setting it turns on --jinja.
Help text: Set additional fields for Jinja JSON template parser, must be a valid JSON object.
--jinja_tools
Section titled “--jinja_tools ”Also: --jinja-tools, --jinjatools · Launcher: Context → Jinja for Tools
Uses the Jinja chat template also for requests that include tool calls.
Turns on --jinja. For OpenAI chat requests with tools, the temperature defaults to 0.5 and is capped at 1.0.
Help text: Enables using jinja chat template formatting for chat completions endpoint. Other endpoints are unaffected. Tool calls are done with jinja.
--jinjatemplate
Section titled “--jinjatemplate ”Also: --chat-template-file · Value: [filename] · Launcher: Loaded Files → Jinja Template
Replaces the model's built-in chat template with your own Jinja template file (or URL).
It does not turn on --jinja; without it, the template only affects the automatic chat-format detection. A missing file is skipped without an error.
Help text: Select a custom Jinja chat template, will overwrite model jinja chat template
--jinjathink
Section titled “--jinjathink ”Value: default | true | false · Default: default · Launcher: Context → Jinja Thinking
Turns thinking on (true) or off (false) for reasoning models that support it in their chat template.
true or false also turns on --jinja. default leaves the model's behaviour unchanged. A request's chat_template_kwargs can override it.
Help text: A quick way to enable or disable thinking in the jinja template.
--lora
Section titled “--lora ”Value: [lora_filename] · Launcher: Loaded Files → Text Lora
Applies a LoRA adapter file on top of the text model.
Only one LoRA is applied; a second value is checked as a base file but not used. If the LoRA cannot be applied, the whole model load fails.
Help text: GGUF models only, applies a lora file on top of model.
--loramult
Section titled “--loramult ”Value: [amount] · Default: 1.0 · Launcher: Loaded Files → Multiplier (next to Text Lora)
How strongly the text LoRA is applied (default 1.0).
Has no effect without --lora.
Help text: Multiplier for the Text LORA model to be applied.
--lowvram
Section titled “--lowvram ”Also: -nkvo, --no-kv-offload · Launcher: Hardware → No KV offload
Keeps the context (KV cache) in system RAM instead of VRAM.
Saves VRAM but makes generation much slower; try fewer --gpulayers or --quantkv first.
Help text: If supported by the backend, do not offload KV to GPU (lowvram mode). Not recommended, will be slow.
--maingpu
Section titled “--maingpu ”Also: --main-gpu, -mg · Value: [Device ID] · Default: -1 · Launcher: Hardware → Main GPU
Chooses which GPU is the main one in a multi-GPU setup. Has no effect on GGUF text models in this version.
-1 (default) means automatic. To keep a GGUF model on one GPU, use --usecuda N without --tensor_split.
Help text: Only used in a multi-gpu setup. Sets the index of the main GPU that will be used.
--maxrequestsize
Section titled “--maxrequestsize ”Value: [size in MB] · Default: 32 · Launcher: Network → Max Req. Size (MB)
The largest request body the server accepts, in MB (default 32).
Larger requests get an HTTP 500 "Payload is too big" error. Raise it only if you send very large images or files.
Help text: Specify a max request payload size. Any requests to the server larger than this size will be dropped. Do not change if unsure.
--mcpfile
Section titled “--mcpfile ”Value: [mcp json file] · Launcher: Loaded Files → MCP JSON
Loads an mcp.json file so models can use tools from MCP servers.
Accepts the Claude Desktop format (mcpServers) and the VS Code format (servers), from a path or URL. KoboldCpp serves the tools at /mcp. The server can start with only this flag, without a text model.
Help text: Specify path to mcp.json which contains the Cladue Desktop compatible MCP server config.
--mmproj
Section titled “--mmproj ”Value: [filename] · Launcher: Loaded Files → Mmproj File
Loads the multimodal projector file that lets a vision (or audio) model understand images (or sound).
The projector must match the text model, which must be a GGUF. A missing file stops startup unless --ignoremissing is set.
Help text: Select a multimodal projector file for vision models.
--mmprojcpu
Section titled “--mmprojcpu ”Also: --no-mmproj-offload · Launcher: Loaded Files → V.Force CPU
Keeps the multimodal projector on the CPU to save VRAM.
Help text: Force CLIP for Vision mmproj always on CPU.
--moecpu
Section titled “--moecpu ”Also: --n-cpu-moe, -ncmoe · Value: [layers affected] · Launcher: Context → MoE CPU Layers
Keeps the expert weights of the first N layers of a Mixture-of-Experts model in system RAM, to fit large MoE models on a small GPU.
The flag alone applies to all layers. It only has an effect when a GPU is used. Setting it stops autofit from turning on automatically, and a forced --autofit ignores it.
Help text: Keep the Mixture of Experts (MoE) weights of the first N layers in the CPU. If no value is provided, applies to all layers.
--ffncpu
Section titled “--ffncpu ”Also: --n-cpu-ffn, -ncffn · Value: [layers affected] · Launcher: Context → FFN CPU Layers
Keeps the dense feed-forward weights of the first N layers in system RAM instead of VRAM.
The flag alone applies to all layers. Like --moecpu, it only has an effect when a GPU is used and stops autofit from turning on automatically.
Help text: Keep the dense FFN weights of the first N layers in the CPU. If no value is provided, applies to all layers.
--moeexperts
Section titled “--moeexperts ”Value: [num of experts] · Default: -1 · Launcher: Context → MoE Experts
How many experts a Mixture-of-Experts model uses per token.
-1 or 0 uses the model's own setting. Changing it trades quality against speed.
Help text: How many experts to use for MoE models (default=follow gguf)
--multiuser
Section titled “--multiuser ”Value: limit · Default: 10 · Launcher: Network → Multiuser Queue
Sets how many generation requests the server accepts at once: one running, the rest waiting in line (default 10).
Requests are processed one at a time unless --parallelrequests is set. N allows N−1 waiting requests, and 1 behaves like 10. 0 turns queuing off so a busy server answers 503, but only for a text-only server: with an image, Whisper, TTS or embeddings model, or with --parallelrequests, queuing stays on.
Help text: Set maximum number of queued incoming requests allowed.
--multiplayer
Section titled “--multiplayer ”Launcher: Network → Shared Multiplayer
Hosts a shared story session that several people can join from KoboldAI Lite.
If --password is set, joining needs it.
Help text: Hosts a shared multiplayer session that others can join.
--nocertify
Section titled “--nocertify ”Launcher: Network → NoCertify Mode (Insecure)
Turns off certificate checks for KoboldCpp's own outgoing HTTPS connections.
Affects KoboldCpp's own HTTPS requests, such as web search and MCP servers; it has nothing to do with --ssl. Model and config downloads run through aria2, curl or wget, which still check certificates. It removes protection against tampered connections, so use it only when a connection fails with certificate errors.
Help text: Allows insecure SSL connections. Use this if you have cert errors and need to bypass certificate restrictions.
--nofastforward
Section titled “--nofastforward ”Launcher: Context → Use FastForwarding (unticked)
Reprocesses the whole prompt on every request instead of reusing the part that has not changed.
This also disables context shifting and SmartCache. Makes most requests slower.
Help text: If set, do not attempt to fast forward GGUF context (always reprocess). Will also enable noshift
--noflashattention
Section titled “--noflashattention ”Also: --no-flash-attn, -nofa · Launcher: Quick Launch → Use FlashAttention (unticked)
Turns off flash attention, which is on by default.
Use it if your GPU or backend crashes or misbehaves with flash attention. With flash attention off, --quantkv compresses only half of the cache (the K part), except bf16, which compresses both.
Help text: Disables flash attention.
--noavx2
Section titled “--noavx2 ”Launcher: Quick Launch → Backend (Use CPU (Old CPU) / Use Vulkan (Old CPU))
Uses a slower build for older CPUs without AVX2.
Normally detected automatically. It prevents CUDA from loading. Add --usecpu for CPU only; otherwise a GPU backend such as Vulkan can still be selected automatically.
Help text: Do not use AVX2 instructions, a slower compatibility mode for older devices.
--nobostoken
Section titled “--nobostoken ”Launcher: Context → No BOS Token
Stops KoboldCpp from adding the beginning-of-sequence (BOS) token to prompts.
Most models need the BOS token; leave this off unless a model's instructions say otherwise.
Help text: Prevents BOS token from being added at the start of any prompt. Usually NOT recommended for most models.
--nommq
Section titled “--nommq ”Launcher: Quick Launch → Use MMQ (unticked)
Turns off MMQ matrix kernels on CUDA.
MMQ is on by default. Only affects CUDA/hipBLAS.
Help text: Disables MMQ, only used for cuda backend. This flag may be removed in future.
--nomodel
Section titled “--nomodel ”Launcher: Loaded Files → Allow Launch Without Models
Lets the server and the KoboldAI Lite web UI start without a model. Models you do set are still loaded.
The help says "GUI", but it means the web UI; the launcher is not shown. Useful with --admin to pick a model later. Cannot be combined with --cli.
Help text: Allows you to launch the GUI alone, without selecting any model.
--noshift
Section titled “--noshift ”Also: --no-context-shift · Launcher: Quick Launch → Use ContextShift (unticked)
Turns off context shifting, which lets long chats continue without reprocessing the whole prompt.
Context shifting is on by default. It is turned off automatically for models using SWA (unless --noswa), for models using MRoPE such as Qwen2-VL and Qwen3-VL, and when --parallelrequests is above 1.
Help text: If set, do not attempt to Trim and Shift the GGUF context.
--noswa
Section titled “--noswa ”Also: --swa-full · Launcher: Context → Allow SWA (unticked)
Uses a full-size context cache on models with sliding-window attention (SWA), such as Gemma.
SWA is on by default and saves a lot of memory, but disables context shifting. With --noswa the cache is much larger and context shifting works again.
Help text: If set, uses full-size SWA KV Cache. Otherwise, SWA will be enabled automatically on models that support it. SWA saves memory but cannot be used with context shifting.
--onready
Section titled “--onready ”Value: [shell command]
Runs a shell command once all models have loaded.
It runs about one second after loading, possibly before the server accepts connections, so poll the port if the command needs the server. A value from a config file is ignored unless --allow-config-onready is set.
Help text: An optional shell command to execute after the model has been loaded.
--allow-config-onready
Section titled “--allow-config-onready ”Allows .kcpps config files to set an --onready shell command.
Without it, onready values in configs are dropped, both at launch and on admin model swaps. Only use it with configs you trust, since the command runs with your user's permissions. A config cannot turn this on itself.
Help text: Allow .kcpps config files loaded after launch to set --onready commands. Only use with trusted configs.
--overridekv
Section titled “--overridekv ”Also: --override-kv · Value: [name=type:value] · Launcher: Context → Override KV
Overrides metadata values in the GGUF model, in the form name=type:value.
Types are int, float, bool and str; separate several overrides with commas (up to 16). Example: tokenizer.ggml.add_bos_token=bool:false.
Help text: Override metadata value by key. Separate multiple values with commas. Format is name=type:value. Types: int, float, bool, str
--overridenativecontext
Section titled “--overridenativecontext ”Value: [trained context] · Launcher: Context → Override Native Context (with Custom RoPE Config ticked)
Tells KoboldCpp the model's trained context length, so automatic RoPE scaling is calculated from that value.
0 (default) uses the model's own value. Takes priority over --ropeconfig.
Help text: Overrides the native trained context of the loaded model with a custom value to be used for Rope scaling.
--overridetensors
Section titled “--overridetensors ”Also: --override-tensor, -ot · Value: [tensor name pattern=buffer type] · Launcher: Context → Override Tensors
Places specific model tensors on a chosen device, using llama.cpp's pattern=buffer type syntax.
Only has an effect when a GPU is used. Unknown buffer types are rejected. Setting it stops autofit from turning on automatically, and a forced --autofit ignores it.
Help text: Override selected backend for specific tensors matching tensor_name_regex_pattern=buffer_type, same as in llama.cpp.
--parallelrequests
Section titled “--parallelrequests ”Also: --continuous-batching, --contbatch · Value: [slots] · Default: 1 · Launcher: Network → Parallel Requests
Processes up to N text requests at the same time (continuous batching); experimental.
0 and 1 mean off; the maximum is 32. It turns off context shifting. Requests that use media, grammar, tools, many advanced samplers, or a negative prompt still run one at a time, and --smartcache or --draftmodel disable batching.
Help text: Allows multiple requests to be batched and executed in parallel. Only works for basic text generation requests (Experimental, No media)
--password
Section titled “--password ”Value: [API key] · Default: os.getenv('KCPP_PASSWORD', None) · Launcher: Network → Password
Requires an API key for text generation and most other endpoints.
Clients send it as Authorization: Bearer <password>; it can also be set with the KCPP_PASSWORD environment variable. Image generation and the /tts_to_audio endpoint stay open, and admin functions need --adminpassword. Always set it before exposing KoboldCpp with --host, --remotetunnel or on a shared network.
Help text: Enter a password required to use this instance. This key will be required for all text endpoints. Image endpoints are not secured. Can also be set with env var KCPP_PASSWORD
--preloadstory
Section titled “--preloadstory ”Value: [savefile] · Launcher: Loaded Files → Preload Story
Serves a saved story file that KoboldAI Lite can load from the server.
Help text: Configures a prepared story json save file to be hosted on the server, which frontends (such as KoboldAI Lite) can access over the API.
--prompt
Section titled “--prompt ”Also: -p · Value: [prompt]
Generates one reply to the given prompt, prints it and exits, without starting the server.
The prompt is sent as raw text with no chat template, using fixed settings (low temperature) and --genlimit tokens (100 if unset). With --cli it becomes the system prompt; with --benchmark it replaces the benchmark text.
Help text: Passing a prompt string triggers a direct inference, loading the model, outputs the response to stdout and exits. Can be used alone or with benchmark.
--quantkv
Section titled “--quantkv ”Value: [quantization level f16/bf16/q8_0/q5_1/q4_0] · Default: f16 · Launcher: Context → Quantize KV Cache
Compresses the context (KV) cache to save memory, which allows longer contexts.
Options are f16 (default), bf16, q8_0, q5_1 and q4_0. Without flash attention only part of the cache is compressed (except bf16). The numeric values are legacy: 1 = q8_0, 2 = q4_0, 3 = bf16.
Help text: Sets the KV cache data type quantization, options are f16/bf16/q8_0/q5_1/q4_0. Requires Flash Attention for full effect, otherwise only K cache is quantized.
--quiet
Section titled “--quiet ”Launcher: Quick Launch → Quiet Mode
Hides prompts and generated text from the terminal output.
Useful for privacy on shared servers. --debugmode turns it off.
Help text: Enable quiet mode, which hides generation inputs and outputs in the terminal. Quiet mode is automatically enabled when running a horde worker.
--ratelimit
Section titled “--ratelimit ”Value: [seconds] · Launcher: Network → IP Rate Limiter (s)
Limits each IP address to one generation request every N seconds.
Faster requests get an HTTP 503 error. 0 (default) means no limit.
Help text: If enabled, rate limit generative request by IP address. Each IP can only send a new request once per X seconds.
--reasoningeffort
Section titled “--reasoningeffort ”Value: default | none | low | medium | high | xhigh · Default: default · Launcher: Context → Think Effort
Sets a default reasoning effort for thinking models; values sent in requests override it.
On a normal command-line launch this flag currently has no effect; it only works through the launcher (--showgui). Use --gendefaults '{"reasoning_effort":"low"}' instead. xhigh sets no thinking limit.
Help text: A quick way to set the default reasoning effort. API values override this.
--remotetunnel
Section titled “--remotetunnel ”Launcher: Quick Launch → Remote Tunnel
Creates a public Cloudflare URL so you can reach KoboldCpp from anywhere, even behind a firewall.
KoboldCpp downloads cloudflared and prints a trycloudflare.com link. Anyone with the link can use your instance, so set --password; image generation stays open even then. It is skipped with --prompt, --benchmark and --cli.
Help text: Uses Cloudflare to create a remote tunnel, allowing you to access koboldcpp remotely over the internet even behind a firewall.
--ropeconfig
Section titled “--ropeconfig ”Value: [rope-freq-scale] [rope-freq-base] · Default: [0.0, 10000.0] · Launcher: Context → RoPE Scale and Base (with Custom RoPE Config and Manual Rope Scale ticked)
Sets custom RoPE scaling (frequency scale and base) to stretch a model beyond its trained context.
By default KoboldCpp scales automatically, and GGUF models with their own RoPE settings keep them. The custom values only apply if the scale is above 0; to change only the base, use scale 1.0. --overridenativecontext takes priority.
Help text: If set, uses customized RoPE scaling from configured frequency scale and frequency base (e.g. --ropeconfig 0.25 10000). Otherwise, uses NTK-Aware scaling set automatically based on context size. For linear rope, simply set the freq-scale and ignore the freq-base
--savedatafile
Section titled “--savedatafile ”Value: [savefile] · Launcher: Loaded Files → SaveData File
Creates or opens a file on the server where KoboldAI Lite users can save and load their stories.
.jsondb is added to the name if missing. It holds 12 slots of up to 10 MB each. If --password is set, saving and loading need it.
Help text: If enabled, creates or opens a persistent database file on the server, that allows users to save and load their data remotely. A new file is created if it does not exist.
--singleinstance
Section titled “--singleinstance ”Launcher: Admin → SingleInstance Mode
Lets a new KoboldCpp started on the same port shut this one down, so two servers never clash.
Both the old and the new instance need the flag. Don't combine it with --remotetunnel.
Help text: Allows this KoboldCpp instance to be shut down by any new instance requesting the same port, preventing duplicate servers from clashing on a port.
--smartcache
Section titled “--smartcache ”Value: limit · Launcher: Context → Use SmartCache
Saves snapshots of the context cache in RAM so switching between chats needs less reprocessing.
The flag alone or 1 keeps 5 snapshots; a higher number sets how many. It needs fast forwarding and turns off --parallelrequests batching. It turns on automatically for recurrent and hybrid models while fast forwarding and context shifting are on.
Help text: Enables intelligent context switching by saving KV cache snapshots to RAM. Requires fast forwarding.
--smartcontext
Section titled “--smartcontext ”Launcher: Context → Use SmartContext
Outdated method that reserves part of the context to reprocess less often.
Has no effect while context shifting is on; use context shifting instead.
Help text: Reserving a portion of context to try processing less frequently. Outdated. Not recommended.
--splitmode
Section titled “--splitmode ”Also: -sm, --split-mode · Value: [split mode] · Default: splitmode_choices[0] · Launcher: Hardware → SplitMode
How the model is split across several GPUs: by layer (default) or by tensor.
row was removed and now uses tensor with a warning.
Help text: How to split the model across multiple GPUs
--ssl
Section titled “--ssl ”Value: [cert_pem] [key_pem] · Launcher: Network → SSL Cert and SSL Key
Serves KoboldCpp over HTTPS using your certificate and key files.
Give two unencrypted PEM files: --ssl cert.pem key.pem. If a file is missing, KoboldCpp prints a warning and keeps running over plain HTTP. The SSL configuration is valid line only means both files exist; a file with an invalid certificate or key can still stop KoboldCpp from starting.
Help text: Allows all content to be served over SSL instead. A valid UNENCRYPTED SSL cert and key .pem files must be provided
--swapadding
Section titled “--swapadding ”Value: int · Launcher: Context → SWA Padding Tokens
Adds extra room to the sliding-window (SWA) cache so you can edit or regenerate further back before a full reprocess.
Uses more memory. Ignored with --noswa.
Help text: How much extra to pad the SWA KV cache, this affects the rewind limit before reprocessing is forced.
--unpack
Section titled “--unpack ”Value: destination · Launcher: Extra → Unpack KoboldCpp To Folder
Extracts the files bundled inside the KoboldCpp executable into a folder, then exits.
The target folder must be empty or not exist yet.
Help text: Extracts the file contents of the KoboldCpp binary into a target directory.
--usemtp
Section titled “--usemtp ”Launcher: Loaded Files → Use MTP
Uses the model's built-in multi-token prediction (MTP) layers for speculative decoding, if it has them.
--draftmodel takes priority if both are set. Turns off --parallelrequests batching.
Help text: Enables MTP layers to be used for drafting (speculative decoding) if present
--usemlock
Section titled “--usemlock ”Also: --mlock · Launcher: Hardware → Use mlock
Locks the model in RAM so the operating system cannot swap it to disk.
Ignored when --usedirectio is set.
Help text: Enables mlock, preventing the RAM used to load the model from being paged out. Not usually recommended.
--usemmap
Section titled “--usemmap ”Launcher: Quick Launch → Use MMAP
Loads the model with memory mapping (mmap), which is off by default in KoboldCpp.
llama.cpp uses mmap by default; KoboldCpp does not. Cannot be combined with --usedirectio.
Help text: If set, uses mmap to load model.
--usedirectio
Section titled “--usedirectio ”Also: --directio, --direct-io, -dio · Launcher: Hardware → Direct I/O
Loads the model with direct I/O, which can speed up the first load on some drives.
Only has an effect on Linux, and falls back to normal loading if direct I/O is not possible. Cannot be combined with --usemmap.
Help text: Use direct I/O to load GGUF models if available. May improve cold-load times on some storage.
--visionmaxres
Section titled “--visionmaxres ”Value: [max px] · Default: 1024 · Launcher: Loaded Files → Vision MaxRes
The largest image resolution passed to a vision model, in pixels (default 1024).
Values outside 512 to 2048 are clamped.
Help text: Clamp MMProj vision maximum allowed resolution. Allowed values are between 512 to 2048 px (default 1024).
--visionmaxtokens
Section titled “--visionmaxtokens ”Also: --image-max-tokens · Value: [tokens] · Default: -1 · Launcher: Loaded Files → V.Min/Max Tok (second box)
Overrides the maximum number of tokens an image may use in a vision model.
-1 uses the model's default. If you set only this value, the minimum is set to the same number. The actual token count still depends on the model's image processing, and some vision models ignore these limits.
Help text: Override the maximum tokens for the MMProj embedding (default -1).
--visionmintokens
Section titled “--visionmintokens ”Also: --image-min-tokens · Value: [tokens] · Default: -1 · Launcher: Loaded Files → V.Min/Max Tok (first box)
Overrides the minimum number of tokens an image uses in a vision model.
-1 uses the model's default. If you set only this value, the maximum is set to the same number.
Help text: Override the minimum tokens for the MMProj embedding (default -1).
--websearch
Section titled “--websearch ”Launcher: Network → Enable WebSearch
Enables the built-in web search proxy used by KoboldAI Lite's web search option.
Searches go through DuckDuckGo. If --password is set, the search endpoint needs it.
Help text: Enable the local search engine proxy so Web Searches can be done.
--showgui
Section titled “--showgui ”Opens the full launcher, pre-filled with your command-line flags or config, instead of starting right away.
Works with any flags, not only with .kcpps files. Handy for adjusting a saved config visually.
Help text: Always show the GUI instead of launching the model right away when loading settings from a .kcpps file.
--skiplauncher
Section titled “--skiplauncher ”Asks KoboldCpp not to show the graphical launcher.
Any command-line flag already skips the full launcher. Without a model, a file picker still appears on a desktop, and on a system without a display KoboldCpp exits with an error. Pass a model flag such as --model with it, or --nomodel to start without one.
Help text: Doesn't display or use the GUI launcher. Overrides showgui.
Horde Worker
Section titled “Horde Worker”--hordemodelname
Section titled “--hordemodelname ”Value: [name] · Launcher: Horde Worker → Horde Model Name
The model name your AI Horde worker advertises.
It also renames the model in KoboldCpp's own APIs (/api/v1/model, /v1/models). Setting a Horde model name, worker name or API key reduces terminal output.
Help text: Sets your AI Horde display model name.
--hordeworkername
Section titled “--hordeworkername ”Value: [name] · Launcher: Horde Worker → Horde Worker Name
The name of your AI Horde worker.
The worker only starts when both this and --hordekey are set.
Help text: Sets your AI Horde worker name.
--hordekey
Section titled “--hordekey ”Value: [apikey] · Launcher: Horde Worker → API Key (If Embedded Worker)
Your AI Horde API key, used to run KoboldCpp as a Horde worker.
The worker only takes Horde jobs after 20 seconds without local requests. The key is stored in plain text in saved .kcpps files, so do not share those.
Help text: Sets your AI Horde API key.
--hordemaxctx
Section titled “--hordemaxctx ”Value: [amount] · Launcher: Horde Worker → Max Context
The largest context your Horde worker accepts from a job.
0 uses --contextsize; it can never exceed it. It also changes the value reported by /api/v1/config/max_context_length.
Help text: Sets the maximum context length your worker will accept from an AI Horde job. If 0, matches main context limit.
--hordegenlen
Section titled “--hordegenlen ”Value: [amount] · Launcher: Horde Worker → Gen. Length
The most tokens your Horde worker generates per job.
0 means 1024. It also changes the value reported by /api/v1/config/max_length.
Help text: Sets the maximum number of tokens your worker will generate from an AI horde job.
Image Generation
Section titled “Image Generation”--sdaudiovae
Section titled “--sdaudiovae ”Value: [filename] · Launcher: Image Gen → Audio VAE
The audio VAE file for LTX 2.3 video generation.
Help text: Specify an image generation audio VAE for LTX2.3 video generation.
--sdclamped
Section titled “--sdclamped ”Value: [maxres] · Launcher: Image Gen → Clamp Resolution (Hard)
Limits the longest side of generated images, for shared instances.
The flag alone means 512; the other side is scaled to keep the aspect ratio. The help is wrong on two points: it does not limit steps, and the minimum is 64 px, not 512.
Help text: If specified, limit generation steps and image size for shared use. Accepts an extra optional parameter that indicates maximum resolution (eg. 768 clamps to 768x768, min 512px, disabled if 0).
--sdclampedsoft
Section titled “--sdclampedsoft ”Value: [maxres] · Launcher: Image Gen → (Soft), next to Clamp Resolution (Hard)
Limits the total pixel area of generated images to save memory, while allowing any aspect ratio.
640 allows 640×640 or the same area in other shapes. 0 still applies a model default (832 for SD1.x/2.x, 1024 for others); the maximum is 2048.
Help text: If specified, limit max image size to curb memory usage. Similar to --sdclamped, but less strict, allows trade-offs between width and height (e.g. 640 would allow 640x640, 512x768 and 768x512 images).
--sdclip1
Section titled “--sdclip1 ”Also: --sdclipl · Value: [filename] · Launcher: Image Gen → Clip-1 File
The first separate encoder file for image models that ship without one built in: Clip-L for SD3 or Flux, or the vision encoder for WAN or Qwen-Image.
.gguf files work too, not only safetensors.
Help text: Specify first safetensors Clip model (SD3 or Flux Clip-L, WAN or QwenImg vision). Leave blank if prebaked or unused.
--sdclip2
Section titled “--sdclip2 ”Also: --sdclipg · Value: [filename] · Launcher: Image Gen → Clip-2 File
The second separate text-encoder file for image models that need one (e.g. Clip-G for SD3).
Some other model families also use this slot. .gguf files work too.
Help text: Specify second safetensors Clip model (SD3 Clip-G). Leave blank if prebaked or unused.
--sdclipdevice
Section titled “--sdclipdevice ”Value: [Device ID] · Launcher: Image Gen → CLIP dev
Where the image model's text encoders run: CPU (default), main GPU or a GPU number.
Unrecognised values silently mean the main GPU.
Help text: CLIP / T5 / LLM device for image generation. GPU index, -1 or 'main' for the main GPU, or 'CPU' (default: CPU).
--sdconvdirect
Section titled “--sdconvdirect ”Value: off | vaeonly | full · Default: sd_convdirect_choices[0] · Launcher: Image Gen → Conv2D Direct
Enables direct 2D convolution for image generation, which may speed it up or save memory.
off (default), vaeonly or full. May crash on backends that do not support it.
Help text: Enables Conv2D Direct. May improve performance or reduce memory usage. Might crash if not supported by the backend. Can be 'off' (default) to disable, 'full' to turn it on for all operations, or 'vaeonly' to enable only for the VAE.
--sdflashattention
Section titled “--sdflashattention ”Launcher: Image Gen → SD Flash Attention
Enables flash attention for image generation.
Independent of the text model's flash attention setting.
Help text: Enables Flash Attention for image generation.
--sdlora
Section titled “--sdlora ”Value: [filename] · Launcher: Image Gen → Image LoRAs
Image-generation LoRA files to apply, or a folder of LoRAs to choose from per request.
With a folder, LoRAs are picked in the prompt with <lora:name:weight>. Up to 10 files are applied; a file with a non-zero multiplier is fixed and cannot be changed per request. Cannot be combined with --sdquant on the command line.
Help text: Specify image generation LoRAs safetensors models to be applied. Multiple LoRAs are accepted.
--sdloramult
Section titled “--sdloramult ”Value: [amounts] · Default: [1.0] · Launcher: Image Gen → Multiplier
Strength of each image LoRA, in the same order as --sdlora.
Missing entries reuse the first multiplier. A multiplier of 0 leaves that LoRA adjustable per request.
Help text: Multipliers for the image LoRA model to be applied.
--sdmaingpu
Section titled “--sdmaingpu ”Value: [Device ID] · Launcher: Image Gen → ImgGPU
Which GPU holds the image model: main (default), a GPU number, or CPU.
Unrecognised values silently mean the main GPU.
Help text: If specified, Image Generation weights will be placed on the selected GPU index. GPU index, -1 or 'main' for the main GPU, or 'CPU' (default: 'main')
--sdmodel
Section titled “--sdmodel ”Value: [filename] · Launcher: Image Gen → Image Model
The image generation model (safetensors or GGUF) to load; this is what turns image generation on.
The server can start with only this flag, without a text model. A missing file stops startup unless --ignoremissing is set.
Help text: Specify an image generation safetensors or gguf model to enable image generation.
--sdoffloadcpu
Section titled “--sdoffloadcpu ”Launcher: Image Gen → Model Offload
Keeps image model weights in system RAM and moves them to VRAM only when needed, to save VRAM.
Help text: Offload image weights in RAM to save VRAM, swap into VRAM when needed.
--sdphotomaker
Section titled “--sdphotomaker ”Value: [filename] · Launcher: Image Gen → PhotoMaker
Loads a PhotoMaker model for face cloning with SDXL models; input images are then used as face references instead of for img2img.
Help text: PhotoMaker is a model that allows face cloning. Specify a PhotoMaker safetensors model which will be applied replacing img2img. SDXL models only. Leave blank if unused.
--sdquant
Section titled “--sdquant ”Value: [quantization level 0/1/2] · Launcher: Image Gen → Compress Weights
Compresses the image model while loading to save memory: 1 = q8, 2 = q4.
The flag alone means 2. Loading a model that is already quantized is faster. Cannot be combined with --sdlora on the command line.
Help text: If specified, loads the model quantized to save memory. 0=off, 1=q8, 2=q4
--sdllm
Section titled “--sdllm ”Value: [filename] · Launcher: Image Gen → Image LLM
The text encoder or LLM file needed by image models such as Flux, Qwen-Image or Z-Image.
The terminal labels it T5-XXL even when it is an LLM.
Help text: Specify an image generation text encoder or LLM (.safetensors or .gguf). Leave blank if prebaked or unused.
--sdthreads
Section titled “--sdthreads ”Value: [threads] · Launcher: Image Gen → ImgThreads
How many CPU threads image generation uses.
0 uses the same value as --threads.
Help text: Use a different number of threads for image generation if specified. Otherwise, has the same value as --threads.
--sdtiledvae
Section titled “--sdtiledvae ”Value: [maxres] · Default: 512 · Launcher: Image Gen → VAE Tiling Threshold
Decodes images larger than this size (default 512) in tiles to save memory.
Tiling starts when width × height is larger than this value squared. 0 disables tiling. --sdvaeauto always disables it.
Help text: Adjust the automatic VAE tiling trigger for images above this size. 0 disables vae tiling.
--sdupscaler
Section titled “--sdupscaler ”Value: [filename] · Launcher: Image Gen → Upscaler
Loads an ESRGAN upscaler model for enlarging generated images.
Needs --sdmodel.
Help text: You can use ESRGAN as an upscaling model to resize images. Leave blank if unused.
--sdvae
Section titled “--sdvae ”Value: [filename] · Launcher: Image Gen → Image VAE
Replaces the image model's built-in VAE with a separate VAE file.
.gguf files work too. Cannot be combined with --sdvaeauto.
Help text: Specify an image generation safetensors VAE which replaces the one in the model.
--sdvaeauto
Section titled “--sdvaeauto ”Launcher: Image Gen → Automatic VAE (TAE SD)
Uses a small built-in VAE (TAESD) that decodes images much faster and can replace a broken VAE.
Disables VAE tiling. Cannot be combined with --sdvae.
Help text: Uses a built-in tiny VAE via TAE SD, which is very fast, and fixed bad VAEs.
--sdvaedevice
Section titled “--sdvaedevice ”Value: [Device ID] · Launcher: Image Gen → VAE dev
Where the image VAE runs: main GPU (default), a GPU number, or CPU.
Help text: VAE device for image generation. GPU index, -1 or 'main' for the main GPU, or 'CPU' (default: main).
--sdvramlimit
Section titled “--sdvramlimit ”Value: [limit MB] · Launcher: Image Gen → VRAM Limiter (MB)
Caps how much VRAM image generation may use, in MB.
0 (default) means no limit.
Help text: If set, prevent VRAM usage from image gen from using more than this many MB.
Whisper Transcription
Section titled “Whisper Transcription”--whispermodel
Section titled “--whispermodel ”Value: [filename] · Launcher: Audio → Whisper Model (Speech-To-Text)
Loads a Whisper model to turn speech into text.
The server can start with only this flag, without a text model.
Help text: Specify a Whisper .bin model to enable Speech-To-Text transcription.
TTS Narration
Section titled “TTS Narration”--ttsmodel
Section titled “--ttsmodel ”Value: [filename] · Launcher: Audio → TTS Model (Text-To-Speech)
Loads a text-to-speech (TTS) model so KoboldCpp can read text aloud.
The model type is detected automatically. OuteTTS models also need --ttswavtokenizer. The server can start with only this flag.
Help text: Specify the TTS Text-To-Speech GGUF model.
--ttswavtokenizer
Section titled “--ttswavtokenizer ”Value: [filename] · Launcher: Audio → WavTokenizer Model (Required for some models)
The WavTokenizer file that some TTS models need.
Required for OuteTTS models.
Help text: Specify the WavTokenizer GGUF model.
--ttsgpu
Section titled “--ttsgpu ”Launcher: Audio → TTS Use GPU
Runs the TTS model on the GPU.
Only OuteTTS and Qwen3-TTS models use the GPU; other TTS models such as Kokoro always run on the CPU.
Help text: Use the GPU for TTS.
--ttsmaxlen
Section titled “--ttsmaxlen ”Value: int · Default: 4096 · Launcher: Audio → TTS Max Tokens
Limits how much audio TTS generates (default and maximum 4096).
For OuteTTS it limits audio tokens. For Kokoro-style models it limits the input to the first N words. Qwen3-TTS models ignore it.
Help text: Limit number of audio tokens generated with TTS.
--ttsthreads
Section titled “--ttsthreads ”Value: [threads] · Launcher: Audio → TTS Threads
How many CPU threads TTS uses.
0 uses the same value as --threads. Qwen3-TTS models ignore it.
Help text: Use a different number of threads for TTS if specified. Otherwise, has the same value as --threads.
--ttsdir
Section titled “--ttsdir ”Value: [directory] · Launcher: Audio → TTS Voices Dir
A folder of .wav and .mp3 voice samples for voice cloning.
Only files directly in the folder are loaded, each as a voice named after its file. Needs --ttsmodel.
Help text: Select directory containing voices for voice cloning.
Music Gen
Section titled “Music Gen”--musicllm
Section titled “--musicllm ”Value: [filename] · Launcher: Audio → MusicLLM
The music planner language model (e.g. ACE-Step LM) for music generation.
Alone it loads in an LLM-only mode. Combined with --musicdiffusion, --musicembeddings and --musicvae, it plans full songs. With embeddings or VAE but no diffusion model, nothing is loaded.
Help text: Select music LLM model (e.g acestep-5Hz-lm-0.6B)
--musicembeddings
Section titled “--musicembeddings ”Value: [filename] · Launcher: Audio → MusicEmbeds
The embedding model required for music generation.
Required with --musicdiffusion; without a diffusion model it is not loaded.
Help text: Select music embedding model (e.g Qwen3-Embedding-0.6B)
--musicdiffusion
Section titled “--musicdiffusion ”Value: [filename] · Launcher: Audio → MusicDiffuser
The music diffusion model; this is what turns full music generation on.
Also needs --musicembeddings and --musicvae, or KoboldCpp exits with an error.
Help text: Select music diffusion (DiT) model (e.g acestep-v15-turbo)
--musicvae
Section titled “--musicvae ”Value: [filename] · Launcher: Audio → MusicVAE
The VAE model required for music generation.
Required with --musicdiffusion.
Help text: Select music VAE model
--musiclowvram
Section titled “--musiclowvram ”Launcher: Audio → Music Low VRAM
Unloads the music models between uses and loads them only when needed, to save VRAM.
Help text: Unload music models when not in use
Embeddings Model
Section titled “Embeddings Model”--embeddingsmodel
Section titled “--embeddingsmodel ”Value: [filename] · Launcher: Loaded Files → Embeds Model
Loads an embeddings model (GGUF) for creating embedding vectors, e.g. for search or memory features.
The server can start with only this flag, without a text model.
Help text: Specify an embeddings model to be loaded for generating embedding vectors.
--embeddingsmaxctx
Section titled “--embeddingsmaxctx ”Value: [amount] · Default: 4096 · Launcher: Loaded Files → ECtx
The longest input the embeddings model accepts (default 4096).
The value is capped at the model's trained context; the help's "defaults to trained context" is wrong. 0 uses --contextsize, and a negative value uses the trained context.
Help text: Overrides the default maximum supported context of an embeddings model (defaults to trained context).
--embeddingsgpu
Section titled “--embeddingsgpu ”Launcher: Loaded Files → GPU (next to Embeds Model)
Puts the embeddings model on the GPU.
Usually not needed.
Help text: Attempts to offload layers of the embeddings model to GPU. Usually not needed.
--rpcmode
Section titled “--rpcmode ”Value: [disabled/connect/host] · Default: disabled · Launcher: Network → RPC Mode
Shares GPUs over the network: host offers this machine's devices, connect uses remote ones.
In host mode no model is loaded and no web API starts; if no GPU is found, the CPU is shared. In connect mode, list the remote machines with --rpctargets. RPC has no authentication, so use it only on a trusted network.
Help text: RPC allows GPUs to be shared over the network. connect=access a remote GPU, host=share your GPU
--rpcport
Section titled “--rpcport ”Value: [portnumber] · Default: 5551 · Launcher: Network → RPC Host Port
The port the RPC server listens on in host mode (default 5551).
Cannot be combined with --rpctargets.
Help text: RPC host mode only. Port for RPC server to listen on (default: 5551).
--rpchost
Section titled “--rpchost ”Value: [IP address] · Default: 0.0.0.0 · Launcher: Network → RPC Host IP
The address the RPC server listens on in host mode (default 0.0.0.0, all interfaces).
RPC has no authentication. Use 127.0.0.1 or a private network address, and never expose the RPC port to the internet.
Help text: RPC host mode only. IP address for RPC server to listen on. Use 0.0.0.0 for all interfaces, 127.0.0.1 for localhost only.
--rpcdevice
Section titled “--rpcdevice ”Value: <dev1,dev2,..> · Launcher: Network → RPC Devices
Which devices to share in RPC host mode, comma-separated.
CPU and unknown device names are ignored with an error message, and the devices are then chosen automatically.
Help text: RPC host mode only. Set specific devices to use for RPC server. Comma separated. Overrides normal RPC device choices.
--rpctargets
Section titled “--rpctargets ”Value: [remotehost1:port1,remotehost2:port2] · Launcher: Network → RPC Endpoints
The remote RPC servers to use, as host:port pairs separated by commas.
Requires --rpcmode connect; otherwise KoboldCpp exits with an error.
Help text: RPC connect mode only. Specify a comma separated list of remote RPC endpoints to connect to e.g. 127.0.0.1:5551,127.0.0.1:5552
Administration
Section titled “Administration”--admin
Section titled “--admin ”Launcher: Admin → Enable Model Administration
Enables admin mode, so you can unload the model or switch to other configs and models without restarting KoboldCpp.
Switch targets come from --admindir (the program's folder if unset). Always set --adminpassword too.
Help text: Enables admin mode, allowing you to unload and reload different configurations or models.
--adminpassword
Section titled “--adminpassword ”Value: [password] · Default: os.getenv('KCPP_ADMINPASSWORD', None) · Launcher: Admin → Admin Password
The password required for admin functions such as switching or unloading models.
Clients send it as Authorization: Bearer <password>; it also works in place of the normal --password. It can be set with the KCPP_ADMINPASSWORD environment variable and is stored in plain text in launcher-saved config files. Use a different value from --password for any instance others can reach.
Help text: Require a password to access admin functions. You are strongly advised to use one for publically accessible instances! Can also be set with env var KCPP_ADMINPASSWORD
--admindir
Section titled “--admindir ”Value: [directory] · Launcher: Admin → Config Directory (Required)
The folder admin mode offers configs and models from.
Lists .kcpps, .kcppt and .gguf files, including one subfolder level. A plain .gguf is loaded with default settings or with --baseconfig.
Help text: Specify a directory to look for .kcpps configs in, which can be used to swap models.
--adminunloadtimeout
Section titled “--adminunloadtimeout ”Value: int · Launcher: Admin → Auto Unload Timeout
Unloads the model automatically after this many idle seconds.
Only works with --admin (or --routermode). 0 (default) disables it. In router mode the next request reloads the model.
Help text: Set an idle timeout in seconds after which KoboldCpp will automatically unload the current model.
--routermode
Section titled “--routermode ”Launcher: Admin → Router Mode
Puts a proxy in front of KoboldCpp so a request's model field can switch to any model or config in --admindir.
The proxy uses --port and the real server moves to a port between 15001 and 15010. It turns on --admin automatically.
Help text: Router mode uses a reverse proxy router, allowing you to easily hotswap models and configs within a single request. Requires admin mode.
--reqtimeout
Section titled “--reqtimeout ”Value: [seconds] · Default: 600 · Launcher: Network → Request Timeout (s)
How long, in seconds, the router-mode proxy waits for the server (default 600).
Despite the help text, it only applies with --routermode, not to normal requests.
Help text: Timeout in seconds for HTTP requests.
--autoswapmode
Section titled “--autoswapmode ”Launcher: Admin → Autoswap Mode
Loads the right kind of model (text, image, speech, TTS, embeddings, music) automatically for each request, from one config.
It turns on --routermode and --admin. It starts with nothing loaded; small model types (see --autoswapthreshold) stay loaded alongside the active one.
Help text: Autoswap mode builds on router mode to allow switching of model types within the same config automatically. Requires admin mode and router mode. All models desired must be defined within the same config.
--autoswapthreshold
Section titled “--autoswapthreshold ”Value: check_range(int, 0, 1048576) · Default: 256 · Launcher: Admin → Autoswap Threshold (MB)
In autoswap mode, model types whose files total at most this many MB stay loaded (default 256).
0 keeps only the currently needed type loaded.
Help text: Keep model families at or below this combined file size resident in autoswap mode, in MB (default 256). Only one model family above the threshold is loaded at a time.
--baseconfig
Section titled “--baseconfig ”Value: value · Launcher: Admin → Base config .kcpps (Optional, for reloading)
A .kcpps config applied first when admin mode switches models, with the chosen config or model layered on top. A base config named in the switch request replaces it.
Only used with --admin. It is not applied when unloading or returning to the initial model. The path is relative to the working directory.
Help text: Specify a base .kcpps config to apply, if no custom base config is selected during a model swap