Skip to content
KoboldCpp
GitHub

Text-to-speech (TTS)

KoboldCpp turns text into spoken audio. KoboldAI Lite can narrate the AI's replies, and apps can use the speech API.

Started with a Chatbot newbie template? It already loads a TTS model. Skip to step 4.

  1. Download a TTS model (.gguf file) from huggingface.co/koboldcpp/tts. A good first pick is Kokoro_no_espeak_Q4.gguf, the one the LowSpec-Chatbot template uses.
  2. In the launcher, open the Audio tab and select the file in TTS Model (Text-To-Speech):. Qwen3TTS and OuteTTS need a second file in WavTokenizer Model (Required for some models): (see below).
  3. Click Launch.
  4. In KoboldAI Lite, open Settings > Media and set Text to speech under Audio Output to KoboldCpp TTS API. Pick a Voice.

On the command line:

Terminal
koboldcpp --model mymodel.gguf --ttsmodel tts-model.gguf

KoboldCpp reads the engine from the model file.

EngineNotes
Qwen3TTS0.6B and 1.7B sizes. Built-in voices, voice cloning and voice design. Needs its tokenizer file in WavTokenizer Model (Required for some models): (--ttswavtokenizer). Without it, loading fails.
KokoroRuns on the CPU.
OuteTTSNeeds a WavTokenizer file in WavTokenizer Model (Required for some models): (--ttswavtokenizer). Without it, loading fails.
Parler, DiaRun on the CPU.
  • Built-in voices: Qwen3TTS voices come with KoboldCpp. The voice list also has random and instruct.
  • Voice design (Qwen3TTS): describe the voice in square brackets at the start of the text, for example [A depressed woman is crying] I can't believe it.
  • Voice cloning (Qwen3TTS): put short .wav or .mp3 recordings in a folder and select it in TTS Voices Dir: (--ttsdir). Each file becomes a voice named after the file. Only Qwen3TTS clones from these recordings, and only with a model file that includes a speaker encoder. Without one, or with a voice design description in the text, you get a normal voice instead.
  • KoboldAI Lite: Settings > Media > Audio Output. Besides the voice, you set Narration triggered for, Narrate only dialog and Narration streaming.
  • MusicUI (http://localhost:5001/musicui) has a TTS tab.
  • Other apps: use the OpenAI speech API or the XTTS API.
EndpointFormat
POST /api/extra/ttsKoboldCpp
POST /v1/audio/speechOpenAI
POST /tts_to_audioXTTS
GET /v1/audio/voices, /speakers_listVoice lists

The audio comes back as WAV. Set "response_format": "mp3" for MP3. Qwen3TTS speaks several languages. Set "language" in the request on any of the endpoints above (default en).

Launcher fieldFlagWhat it does
TTS Use GPU--ttsgpuRuns OuteTTS and Qwen3TTS on the GPU. Kokoro, Parler and Dia always run on the CPU.
TTS Threads:--ttsthreadsCPU threads for TTS. 0 uses the text model's thread count.
TTS Max Tokens:--ttsmaxlenFor OuteTTS, the maximum audio tokens. For Kokoro, Parler and Dia, the maximum number of input words. Default and maximum 4096.
TTS Voices Dir:--ttsdirFolder with voice recordings for cloning.