Skip to content
KoboldCpp
GitHub

Speech-to-text with Whisper

KoboldCpp turns spoken audio into text with Whisper. You can talk to the AI in KoboldAI Lite, or send audio files to the transcription API.

Started with a Chatbot newbie template? It already loads a Whisper model. Skip to step 4.

  1. Download a Whisper model (.bin file) from huggingface.co/koboldcpp/whisper. A good first pick is whisper-base.en-q5_1.bin, the one the LowSpec-Chatbot template uses.
  2. In the launcher, open the Audio tab and select the file in Whisper Model (Speech-To-Text):.
  3. Click Launch.
  4. In KoboldAI Lite, open Settings > Media and set Voice Input under Audio Input.

On the command line:

Terminal
koboldcpp --model mymodel.gguf --whispermodel whisper-model.bin

Settings > Media > Audio Input > Voice Input has these modes:

ModeHow it works
Detect VoiceHands-free: Lite detects when you speak.
Push-To-TalkLite records while you hold the voice button.
Toggle-To-TalkThe voice button switches recording on and off.

The same section has Suppress Non-Speech, Language and Delay. The Add File menu also has a Microphone option.

Browsers allow the microphone only on HTTPS or on the PC itself (localhost). From another device, use the Remote Tunnel link or HTTPS.

Whisper in KoboldCpp reads WAV, MP3 and FLAC. Convert other formats, such as WebM or OGG, before you send them. Some apps send WebM by default.

EndpointFormat
POST /api/extra/transcribeKoboldCpp
POST /v1/audio/transcriptionsOpenAI

Send the audio as base64 in audio_data, or as a multipart file upload. The answer is {"text": "..."}.

Optional fields:

FieldMeaning
language (or langcode)Language code. Default auto.
promptText that guides the transcription.
suppress_non_speechIgnore sounds that are not speech.

A multipart upload reads only language and prompt. Use JSON for langcode and suppress_non_speech.

With --password set, these endpoints need the password.

If no Whisper model is loaded but the text model has an audio-capable mmproj, KoboldCpp transcribes with the text model instead. See Vision and audio input.