Skip to content
KoboldCpp
GitHub

Vision and audio input

A vision model can describe and answer questions about images. Some models can also listen to audio. For this, the text model needs its matching projector file, called an mmproj.

Started with a Chatbot newbie template? It already loads a vision model and its mmproj. Skip to step 5.

  1. Use a text model that supports vision, for example Qwen3-VL-8B.
  2. Download its mmproj. For Qwen3-VL-8B, that is mmproj-BF16.gguf.
  3. In the launcher, open the Loaded Files tab and select the file in Mmproj File:.
  4. Click Launch.
  5. In KoboldAI Lite, click Add File > Upload a File (Image / Audio) and ask about the image.

On the command line:

Terminal
koboldcpp --model Qwen3-VL-8B-Instruct-Q4_K_S.gguf --mmproj mmproj-BF16.gguf

An mmproj only works with the exact model it was made for, including its size: a Gemma3 12B model needs the Gemma3 12B mmproj, and a Qwen mmproj does not work with a Gemma model. The mismatch shows as failed to load mmproj model! or a crash while loading.

The mmproj works only with GGUF text models.

Some mmproj files handle audio. The launcher field accepts both ("Audio or Vision mmproj"). With an audio mmproj, you can send audio clips to the model the same way as images.

If no Whisper model is loaded, KoboldCpp uses the audio mmproj for speech-to-text as well.

  • KoboldAI Lite: Add File offers Upload a File (Image / Audio), Camera and Drag Drop / Paste from Clipboard. Images appear in the chat where you added them. From another device, Camera works only over HTTPS; see Remote access.
  • OpenAI API (/v1/chat/completions): send images as multimodal message content, as with OpenAI vision models.
  • KoboldAI API (/api/v1/generate): send them in the images and audio arrays.
  • A1111 API: POST /sdapi/v1/interrogate describes an image. It is not protected by --password.

Each request can carry up to 64 images and 64 audio clips. If you send more, only the last 64 of each are used.

To check what a running server supports, GET /props shows modalities.vision and modalities.audio.

All on the Loaded Files tab:

Launcher fieldFlagWhat it does
Vision MaxRes:--visionmaxresLargest image size in pixels. 512 to 2048, default 1024.
V.Min/Max Tok:--visionmintokens, --visionmaxtokensTokens per image, for models with variable image sizes. -1 uses the model's default. Setting only one of them sets both to that value.
V.Force CPU--mmprojcpuRuns the mmproj on the CPU instead of the GPU.

KoboldCpp prints a warning when you combine an mmproj with a draft model or MTP.

The mmproj needs VRAM on top of the model. For Qwen3-VL-8B's mmproj-F16.gguf, autofit estimates it at startup: MMProj Autofit Usage: 1105 MB. V.Force CPU moves it to system RAM, and images are then processed more slowly. On a Radeon RX 7600 XT, the first reply to a 1920-pixel-wide screenshot took 16.1 s with the mmproj on the GPU and 22.9 s on the CPU.