Skip to content
KoboldCpp
GitHub

Launcher tour

The launcher is the settings window that opens when you start KoboldCpp without a model. It has a column of tabs on the left and the settings of the active tab on the right. Hover over a label to see its tooltip.

For most people, the Quick Launch tab is all they need. The other tabs hold the same settings in more detail, plus extra features.

Each setting also exists as a command-line flag. The tables below list both; see Command line and the flag reference.

ButtonWhat it does
UpdateOpens the latest release on GitHub in your browser. It does not update anything by itself; see Updating.
Save ConfigSaves all current settings to a .kcpps file. The launcher does not remember settings otherwise. See Config files.
Load ConfigLoads a .kcpps config or .kcppt template into the launcher. It does not launch.
Get HelpOpens the Help Menu: links to the wiki and starter guides, a Hugging Face model search, and ready-made templates. See Templates.
LaunchStarts KoboldCpp with the current settings.

If you click Launch without a model, KoboldCpp asks whether you want help finding one. Yes opens the Help Menu.

The beginner tab. Pick a model and click Launch.

The Quick Launch tab of the KoboldCpp launcher

The Quick Launch and Hardware screenshots show a PC with an NVIDIA graphics card (Use CUDA). With Use Vulkan the same GPU fields appear except Use MMQ; with Use CPU they are hidden.

LabelFlagWhat it does
Backend:--usecuda, --usevulkan, --usecpuWhere the model runs. CUDA is for NVIDIA cards and is the fastest. Vulkan works on most graphics cards. The CPU options run without a graphics card; on macOS, also set GPU Layers: to 0. KoboldCpp picks one for your hardware until you change it yourself.
GPU ID:number after --usecuda / --usevulkanWhich graphics card to use. The card's name appears next to it.
GPU Layers:--gpulayersHow much of the model goes on the graphics card. -1 (default) fits it automatically.
Context Size:--contextsizeHow much text the model keeps in mind. Default 16384. See Context size.
GGUF Text Model:--modelThe model file. Browse picks a file; HF Search finds one on Hugging Face.

Checkboxes:

LabelFlagDefaultWhat it does
Launch Browser--launchonOpens KoboldAI Lite in your browser when loading is done.
Use MMAP--usemmapoffLoads the model with memory mapping.
Use ContextShift--noshift turns it offonAvoids reprocessing the whole chat once the context is full.
Remote Tunnel--remotetunneloffCreates a public internet URL for this instance. See Remote access.
Use FlashAttention--noflashattention turns it offonFaster attention that uses less memory.
Force AutoFit--autofitoffForces the automatic fit. Hides GPU Layers:. Marked experimental.
Quiet Mode--quietoffHides generation output in the terminal.
Use Jinja--jinjaoffFormats chat requests with the model's own chat template. Other endpoints are unaffected.
Use MMQ--nommq turns it offonA CUDA option for prompt processing.

Backend and graphics card settings in more detail, plus CPU threads and batch sizes.

The Hardware tab of the KoboldCpp launcher

GPU ID:, GPU Layers:, SplitMode:, Tensor Split:, Main GPU: and No KV offload appear with a GPU backend. Autofit Padding (MB): appears only with Force AutoFit.

LabelFlagWhat it does
Backend:, GPU ID:, GPU Layers:see Quick LaunchSame settings as on Quick Launch.
SplitMode:, Tensor Split:, Main GPU:--splitmode, --tensor_split, --maingpuHow to share the model across several graphics cards. See Multiple GPUs.
Autofit Padding (MB):--autofitpaddingSpare graphics memory the automatic fit leaves free.
Threads:, Batch Threads:--threads, --blasthreadsCPU threads for generating and for prompt processing. Blank means automatic. The launcher fills in Threads: with the automatic value for your PC; clear it to keep it automatic in a config you use on another PC.
Batch Size:, Physical Batch Size:--batchsize, --ubatchsizeHow many tokens are processed at once. Smaller values save memory but process prompts more slowly.
Device Override--deviceSelects devices by name.
No KV offload--lowvramKeeps the context memory off the graphics card. Slow.
Use mlock--usemlockKeeps the model in RAM so the system cannot swap it out.
Direct I/O--usedirectioAnother way to load the model file.
Debug Mode--debugmodePrints extra information in the terminal. Useful when something goes wrong.
High Priority--highpriorityAsks the system for a higher CPU priority. On Linux it lowers the priority instead.
Keep Foreground--foregroundWindows only: brings the terminal window to the front on each request.
CLI Terminal Only--cliChat in the terminal, without the web server.

Run Benchmark loads the model, measures prompt processing and generation speed, and prints the results in the terminal instead of starting the server.

For help choosing these values, see GPU layers and Saving VRAM.

How much text the model keeps in mind, how KoboldCpp reuses it between requests, and server-wide generation defaults.

The Context tab of the KoboldCpp launcher

LabelFlagWhat it does
Context Size:--contextsizeSame slider as on Quick Launch. Moving one moves the other.
Use FastForwarding--nofastforward turns it offReuses the part of the prompt that has not changed. On by default.
Use ContextShift--noshift turns it offSee Quick Launch.
Allow SWA--noswa turns it offSmaller context memory on models that support it. On by default.
Use SmartCache / CacheSlots:--smartcacheKeeps snapshots of recent chats in RAM for quick switching. CacheSlots: appears once Use SmartCache is ticked.
Quantize KV Cache:--quantkvStores the context in a smaller format to save memory.
Default Gen Amt:--defaultgenamtReply length when the app does not set one. Default 2048.
Prompt Limit:--genlimitA hard cap on reply length for every request.
Default Params: / Override--gendefaults, --gendefaultsoverwriteDefault generation settings for all apps, as JSON. Override makes them replace what the app sent.
Use Jinja, Jinja for Tools, Jinja Thinking:, Jinja Kwargs:--jinja, --jinja_tools, --jinjathink, --jinja_kwargsChat template options. The other three appear once Use Jinja is ticked.
Think Effort:--reasoningeffortDefault reasoning effort for thinking models. Apps can override it.
MoE CPU Layers:, FFN CPU Layers:--moecpu, --ffncpuKeep parts of the model in RAM. See Saving VRAM.

The tab also has RoPE settings, No BOS Token, Enable Guidance, MoE Experts:, Override KV: and Override Tensors: for advanced use. All of these are in the flag reference.

Details on the context options: Context size.

Every file KoboldCpp loads for text chat, plus a few related files.

The Loaded Files tab of the KoboldCpp launcher

LabelFlagWhat it does
Text Model:--modelThe main model. Same as GGUF Text Model: on Quick Launch.
Text Lora: / Multiplier:--lora, --loramultA LoRA adapter for the text model.
Mmproj File:--mmprojThe vision projector that lets the model see images. See Vision.
Draft Model:--draftmodelA small model that speeds up generation (speculative decoding).
Embeds Model:--embeddingsmodelA model for the embeddings endpoint. See Embeddings.
Preload Story:--preloadstoryA story that KoboldAI Lite opens on start.
Enable Server Side SaveData File / SaveData File:--savedatafileStores KoboldAI Lite saves on the server. See Multiplayer and saves.
MCP JSON:--mcpfileMCP tool servers. See Web search and MCP.
Chat Adapter: / Pick Premade--chatcompletionsadapterThe chat format for the chat completions endpoint. Pick Premade lists the bundled formats.
Jinja Template:--jinjatemplateA custom Jinja chat template file.
Download Dir:--downloaddirWhere downloaded models are saved.
Allow Launch Without Models--nomodelStarts KoboldAI Lite without a local model, for use with online services.

How apps and other devices reach KoboldCpp.

The Network tab of the KoboldCpp launcher

LabelFlagWhat it does
Host:--hostThe network address to listen on. Empty (default) means all addresses, so other devices on your network can connect. Enter 127.0.0.1 to allow only this computer.
Port:--portDefault 5001.
Remote Tunnel--remotetunnelA public internet URL through Cloudflare.
Password:--passwordAn API key for the text endpoints.
SSL Cert: / SSL Key:--sslServes over HTTPS.
Shared Multiplayer--multiplayerSeveral people share one story in KoboldAI Lite.
Enable WebSearch--websearchLets the model search the web.
Multiuser Queue:--multiuserHow many requests it accepts at once, counting the one that is running. See Streaming and multiple users.
Max Req. Size (MB):, IP Rate Limiter (s):--maxrequestsize, --ratelimitLimits for public instances.
Parallel Requests:--parallelrequestsProcesses several simple text requests at once. Experimental. Turns off ContextShift.
Request Timeout (s):--reqtimeoutOnly used in router mode. See Admin mode.
RPC Mode:--rpcmodeShares or uses graphics cards across computers. See RPC.

Shares your model with the AI Horde, a volunteer network, so other people can use it.

The Horde Worker tab of the KoboldCpp launcher

LabelFlagWhat it does
Configure for HordeTurns on the built-in Horde worker and shows its fields.
Horde Model Name:--hordemodelnameThe model name shown on the Horde.
Gen. Length:, Max Context:--hordegenlen, --hordemaxctxLimits for Horde requests.
API Key (If Embedded Worker):--hordekeyYour Horde API key.
Horde Worker Name:--hordeworkernameYour worker's name.

See Horde worker.

Loads an image model next to (or instead of) the text model.

The Image Gen tab of the KoboldCpp launcher

The most important field is Image Model: (--sdmodel). The other fields hold extra files some image models need (Image LLM:, Clip-1 File:, Clip-2 File:, Image VAE:), LoRAs, an upscaler, resolution limits and memory options. See Image generation.

Speech-to-text, text-to-speech and music models.

The Audio tab of the KoboldCpp launcher

LabelFlagWhat it does
Whisper Model (Speech-To-Text):--whispermodelTranscribes speech. See Speech to text.
TTS Model (Text-To-Speech):--ttsmodelReads text aloud. Some TTS models also need WavTokenizer Model (Required for some models): (--ttswavtokenizer). See Text to speech.
MusicLLM:, MusicEmbeds:, MusicDiffuser:, MusicVAE:--musicllm, --musicembeddings, --musicdiffusion, --musicvaeThe files for music generation. See Music.

Switching models and configs while KoboldCpp runs.

The Admin tab of the KoboldCpp launcher

LabelFlagWhat it does
Enable Model Administration--adminTurns on admin mode.
Admin Password:--adminpasswordProtects the admin functions. Set one whenever admin mode is on.
Config Directory (Required):--admindirThe folder with the configs and models you can switch to.
Base config .kcpps (Optional, for reloading):--baseconfigSettings applied under every switched-to config or model, unless the switch names its own base config.
Auto Unload Timeout:--adminunloadtimeoutUnloads the model after this many idle seconds.
Router Mode, Autoswap Mode, Autoswap Threshold (MB):--routermode, --autoswapmode, --autoswapthresholdSwitch models automatically per request. Each appears once the setting before it is ticked, starting with Enable Model Administration.
SingleInstance Mode--singleinstanceA new KoboldCpp on the same port can shut this one down.
Launch KoboldCpp Agent--agentStarts the KoboldCpp Agent when loading is done. See Agent.

See Admin mode.

Tools that run right away instead of being launch settings.

The Extra tab of the KoboldCpp launcher

ButtonFlagWhat it does
Unpack KoboldCpp To Folder--unpackExtracts KoboldCpp's files into an empty folder, for faster starts.
Generate LaunchTemplate--exporttemplateSaves the current settings as a .kcppt template for others. See Templates.
Analyze Model--analyzeShows the metadata and tensors of a GGUF or safetensors file.
Register / UnregisterWindows only: opens .gguf, .ggml, .kcpps and .kcppt files with KoboldCpp.
Use Classic FilePickerLinux only: uses the older file dialog.
Spawn TerminalLinux only: opens a terminal window that shows KoboldCpp's output.