← back to babelstreamer.com

OBS Plugin Manual

Everything the BabelStreamer OBS filter does, every setting it has, and how a caption goes from your speech to a caption in the viewers language.

Overview

The BabelStreamer OBS plugin (obs-whisper-ws) is an OBS Studio audio filter that transcribes your speech locally with whisper.cpp, optionally shows the caption live in OBS via a Text source, and streams every transcript over a WebSocket to a BabelStreamer server for translation and broadcast to remote viewers. The full path from your speech to a viewer's screen is covered in Streaming with the filter below.

With an OBS audio source attached, the plugin:

  • downmixes to mono and resamples to 16 kHz
  • runs whisper.cpp transcription locally, on your machine
  • optionally updates an OBS Text source with the live caption
  • sends the text to the BabelStreamer server for translation and distribution
Note: speech-to-text runs locally to keep latency down, which means you need a speech recognition model downloaded to your own machine — models are too large to bundle with the installer. See Choosing a model.

Installing

  1. Run the installer — it's packaged together with this document and installs the plugin straight into OBS's plugin directory.
  2. Open OBS and add the filter to any audio source, the same way you'd add any other OBS audio filter.
  3. Open the filter's Properties and press Download Standard Model — at minimum, you need this once before the filter can recognise any speech.
  4. Get a device token from your account page and paste it into the filter's Token setting.

Streaming with the filter

As soon as the plugin recognises the first frame of audio, it connects to the BabelStreamer server. On a fresh connection the server replies with a session key — a short alphanumeric code — and the filter pops up a dialog:

Session key: ABC123

The session key is what you pass on to your viewers so they can connect to the text stream and choose a language.

The key is tied to your machine and persisted locally, so reconnecting or restarting OBS reuses the same key instead of generating a new one every time. You can read the current key from the Session key field in the filter's Properties dialog at any time.

As you speak, the plugin transcribes locally, and once each utterance is final:

  • if you picked a Caption Text Source, that OBS Text source updates live — useful for burning captions directly into your stream output or a monitor. Note that this is visible to your viewers too.
  • the transcript is always sent to the server, which translates it and fans it out to every connected viewer in their chosen language.

Configuring the filter

Open the filter's gear icon to see its Properties dialog. Every setting it has:

ParameterMeaning
Session key Read-only. Shows the session key the server assigned this filter once it connects (see Streaming with the filter), or "(not connected)" if you don't yet have a connection to the server.
Caption Text Output Which OBS Text source to drive live, so the caption appears directly on your OBS canvas/output. Pick "(none — display only)" to skip this and only send transcripts to the server. The dropdown lists every Text (FreeType2/GDI+) source already in your scene collection — create one first if you want this. Note: this also appears on your video feed as captions in your own (source) language, visible to every viewer regardless of the language they've chosen.
Processing Device Which GPU/Neural engine/CPU on your system to use for speech recognition. If you have both an onboard and a dedicated GPU, always pick the more powerful one. CPU is also selectable, but only works well with the smallest models. If you have a Mac with a neural engine then this is preferred. (On a Mac the plugn will compete for resources with the OBS video encoder and may cause stuttering video.)
Download Standard Model The plugin needs a speech recognition model to understand you, but it's too large (~3 GB) to bundle with the plugin. This button downloads and installs the standard model for you. To use a different model instead, see Choosing a model and set it in the Model field below. On a Mac with neural engine the preferred model is the neural engine model.
Download Neural Model (Mac only) This will download the additional model info to run the model using the mac neural engine. Only visible if neural engine chosen.
Model (.bin) Path to a whisper.cpp GGML model file (e.g. ggml-large.en.bin).
Server Address Address of the BabelStreamer server. Accepts a bare host, host:port, or a full wss:// URL — babelstreamer.com works too, and is the default.
Server Port Used only if Server Address doesn't already include a port. Default 443.
Token The device token you generate on your account page to link this OBS install to your subscription. Copy it there, paste it here.
Speech Language The language you'll be speaking, for the recognizer to listen for. Currently English, French, or Spanish.
Game (auto-inject vocabulary) Which game you're streaming — the recognizer preferentially picks up vocabulary specific to that game.
Custom Vocab File A text file of custom vocabulary to bias speech recognition toward. See Custom vocabulary.
GPU Threads GPU threads whisper.cpp uses for inference. 4 is good, 8 is better, but it depends on what you can spare.
Speech Threshold How loud speech needs to be before it's recognised. Lower this if you sometimes speak quietly — but too low, and the recognizer may start picking up background noise.
Min Speech (ms) Minimum time something that looks like speech must last before it's treated as real speech and recognition starts. Default 120 ms.
End Silence (ms) How much silence ends a portion of speech and triggers a final transcript. Default 260 ms.
Max Utterance Length The longest a stretch of speech can run before the recognizer transcribes it anyway, pause or no pause. Useful if you tend to talk in long, uninterrupted sentences — the recognizer doesn't have to wait for you to stop.

Choosing a model

The plugin needs a whisper.cpp model file (a .bin file) to transcribe speech — this isn't bundled with the plugin, you download it once and point the Model (.bin) field at it. You might want a different model than the standard one because:

  • you want something lighter on resources, and are willing to trade off some quality for it
  • you have a model retrained on your own speech patterns
  • you just want to experiment and see what suits you best

Grab a model from huggingface.co/ggerganov/whisper.cpp — any ggml-*.bin file works. Bigger models are more accurate but need more processing power:

ModelSizeAccuracyRequirements
small ~460 MB Good Runs fine on any modern CPU, limited quality speech recognition. English-only speakers can use the .en version.
medium ~1.5 GB Very good CPU is workable, GPU recommended for smooth real-time captions. Fair quality speech recognition. English-only speakers can use the .en version.
large-v3 ✓ ~2.9 GB Best Requires a GPU — too slow for live captioning on CPU alone. Best quality speech recognition available.

large-v3 gives the most accurate, natural captions, but it's only practical for live use with GPU acceleration (Metal on Mac, Vulkan on Windows — the plugin picks this up automatically if available). If you don't want to dedicate GPU cores to it, use medium instead.

There are also quantised versions of the models, marked with q in the name (e.g. q5) — smaller, lighter on the GPU, and in some cases nearly as good as the unquantised version. "Turbo" models are available too, useful if you need a faster response. Try a few and see what works for you — large-v3 is still the goto if you have the GPU power to spare.

Once downloaded, open the filter's settings in OBS and set Model (.bin) to point at the file — it can live wherever you saved it, no need to move it anywhere special.

Custom vocabulary

The plugin can bias speech recognition toward a set of words you use often that might not otherwise be recognised consistently — game terms, names, slang. Bias is the operative word: a custom vocabulary word is more likely to be recognised, but it's never guaranteed.

The file is a plain .txt file, one word per line. Blank lines are fine, and you can add comments by starting a line with #.

How viewers see captions

Viewers see captions through the BabelStreamer browser extension or mobile app — not through the OBS output itself. That's what lets each of them choose their own language independently.

Captions overlay on top of the video window the viewer has open, live-translated into whichever language they've picked. Each viewer can choose a different language from every other viewer, and can drag the caption overlay to reposition it.

Currently supported browsers: Safari, Edge, Chrome, and Firefox.

Currently supported languages include English, Chinese, Dutch, French, German, Italian, Japanese, Korean, Portuguese, Brazilian Portuguese, Russian, Spanish, and Tagalog, depending on your subscription. For the current list, see the FAQ.