Skip to content
KoboldCpp
GitHub

Embeddings

An embeddings model turns text into a list of numbers (a vector). Apps use these vectors to find text with similar meaning, for example for document search and RAG. KoboldCpp serves an embeddings model next to the text model.

You need this if an app asks for an embeddings endpoint, or to use KoboldCpp Embeddings in Lite's TextDB (Context > TextDB). Normal chatting in KoboldAI Lite does not need it.

Started with a Chatbot newbie template? It already loads an embeddings model. Skip to step 4.

  1. Download an embeddings model as a .gguf file. A good first pick is bge-m3-q8_0.gguf, the one the LowSpec-Chatbot template uses.
  2. In the launcher, open the Loaded Files tab and select the file in Embeds Model:.
  3. Click Launch.
  4. In your app, set the embeddings endpoint to http://localhost:5001/v1 (OpenAI-compatible).

On the command line:

Terminal
koboldcpp --model mymodel.gguf --embeddingsmodel embeddings-model.gguf
EndpointFormat
POST /v1/embeddingsOpenAI. Supports "encoding_format": "base64".
POST /api/embedOllama
POST /api/extra/embeddingsKoboldCpp

With --password set, these endpoints need the password.

Launcher fieldFlagWhat it does
ECtx:--embeddingsmaxctxMaximum input length in tokens. Default 4096, never more than the model was trained for. 0 uses the text model's context size. A negative value uses the model's trained context.
GPU (next to Embeds Model:)--embeddingsgpuPuts the embeddings model on the GPU. Usually not needed.

An empty ECtx: field in the launcher saves as 0.

With bge-m3 on a PC with an RTX 3090, GPU embedded 32 texts of about 340 tokens in 0.69 s instead of 1.3 s on the CPU, and used about 300 MiB more VRAM.