Skip to content
KoboldCpp
GitHub

Image and video generation

KoboldCpp generates and edits images with a bundled copy of stable-diffusion.cpp. It also generates videos with some model families. Image generation runs next to the text model, or on its own.

  1. Download an image model. A good first pick is PicX Real.
  2. In the launcher, open the Image Gen tab and select the file in Image Model:.
  3. Click Launch.
  4. Open http://localhost:5001/sdui for StableUI, the bundled image interface.

A template is the other way in: click Get Help, choose Newbie Templates, and pick LowSpec-ImageGen or MidSpec-ImageGen. KoboldCpp downloads the files at launch.

On the command line:

Terminal
koboldcpp --sdmodel picx_real_q5_1.gguf

Add --model with a text model if you want to chat as well.

  • StableUI (/sdui): prompts, settings and a LoRA picker. It has a recovery mode for poor connections.
  • KoboldAI Lite: open Settings > Media and set Image Backend to KCPP / Forge / A1111. Then use Add File > Generate Image, or turn on Autogenerate images.
  • Other apps: KoboldCpp speaks the A1111/Forge API, the OpenAI images API and a ComfyUI-compatible API. Point apps such as SillyTavern at http://localhost:5001.
APIEndpoints
A1111 / ForgePOST /sdapi/v1/txt2img, /sdapi/v1/img2img, /sdapi/v1/upscale
OpenAIPOST /v1/images/generations, /v1/images/edits
ComfyUIPOST /prompt, /upload/image; GET /view, /history

Videos use the same endpoints. See API endpoints.

Models come as .safetensors or .gguf files.

KindFamilies
ImagesStable Diffusion 1.5, SDXL, SD3, Flux, Qwen Image, Z-Image, Klein, Krea2
VideoWAN 2.2, LTX2.3, Minimax H3

Recent releases also added SDXS (small and fast, usable on a CPU), Microsoft Lens, HiDream o1, LongCat, Ernie Image, Ideogram 4 and Boogu Edit.

Image models are on Hugging Face (search for the family name) and on CivitAI. The popular templates include ready setups, for example Z-Image Turbo and WAN 2.2.

SD 1.5 and SDXL models are usually one file. Newer families such as Flux, SD3, Qwen Image and Z-Image need extra parts. Load them on the Image Gen tab:

PartLauncher fieldFlag
Main model (diffusion model)Image Model:--sdmodel
Text encoder or LLM (T5, Qwen, …)Image LLM:--sdllm
First CLIP model (Clip-L)Clip-1 File:--sdclip1
Second CLIP model (Clip-G)Clip-2 File:--sdclip2
VAEImage VAE:--sdvae
Audio VAE (LTX2.3 video only)Audio VAE:--sdaudiovae

Leave a field empty if the model already contains that part. A template for the model fills in all fields for you.

  • Fixed LoRAs: list the files in Image LoRAs: (--sdlora), with their strength in Multiplier: (--sdloramult). They apply to every image.
  • Runtime LoRAs: tick Runtime LoRAs and pick a folder in LoRA Dir: (--sdlora <folder>). KoboldCpp scans the folder and one level of subfolders. You choose LoRAs per image in StableUI, or with <lora:name:weight> in the prompt.
Launcher fieldFlagEffect
Compress Weights:--sdquantLoads the model compressed to q8 (1) or q4 (2). A model downloaded pre-quantized loads faster. Cannot be combined with --sdlora on the command line.
Automatic VAE (TAE SD)--sdvaeautoUses a tiny built-in VAE. Faster and uses less memory, but image quality can drop. May fix a bad VAE. Not available for every model type.
Model Offload--sdoffloadcpuKeeps image weights in RAM and moves them to VRAM when needed.
VRAM Limiter (MB):--sdvramlimitLimits how much VRAM image generation plans for. Actual use can go somewhat higher. 0 means no limit.
VAE Tiling Threshold:--sdtiledvaeDecodes images in tiles when they have more pixels than this size squared (default 512: above 512×512 pixels). 0 turns tiling off.

The image templates turn on Model Offload and keep the text encoder on the CPU. Measured on an RTX 3090 with FLUX.2 klein at 4 steps:

SetupPeak VRAMOne 1024×1024 image
LowSpec-ImageGen (FLUX.2 klein 4B)3858 MiB14.2 s
LowSpec-ImageGen, Model Offload off4514 MiB13.9 s
LowSpec-ImageGen, VAE tiling off (--sdtiledvae 0)7584 MiB11.8 s
MidSpec-ImageGen (FLUX.2 klein 9B)7238 MiB22.8 s
Launcher fieldFlagWhat it does
Clamp Resolution (Hard):--sdclamped [px]Limits the longest image side, keeping the aspect ratio. Without a value it means 512.
(Soft):, next to Clamp Resolution (Hard)--sdclampedsoft pxLimits the image area (e.g. 640 allows 640x640, 512x768, 768x512). At 0 a default applies: 832 for SD 1.x/2.x, 1024 for others. Maximum 2048.
Upscaler:--sdupscalerLoads an ESRGAN upscaler for /sdapi/v1/upscale.
PhotoMaker:--sdphotomakerFace cloning for SDXL models.
SD Flash Attention--sdflashattentionFlash attention for image generation. Separate from the text model's setting.
Conv2D Direct:--sdconvdirectoff (default), vaeonly or full. May speed things up or save memory. Can crash on backends that don't support it.
ImgThreads:--sdthreadsCPU threads for images. 0 uses the text model's thread count.
ImgGPU:--sdmaingpuGPU for the image weights.
CLIP dev:--sdclipdeviceDevice for the text encoders. Default: CPU.
VAE dev:--sdvaedeviceDevice for the VAE. Default: the main GPU.

Every flag is listed in the flag reference. Older guides use --sdt5xxl, --sdvaecpu and --sdclipgpu; their replacements are --sdllm, --sdvaedevice and --sdclipdevice.