Skip to content

Text-to-Speech (Local Engines)

TomoriBot treats speech as a custom endpoint capability. Local engines run outside the bot as HTTP servers and implement POST /synthesize; clone engines receive text plus the configured reference audio sample, while VoiceDesign-capable engines can also receive natural-language instructions.

Available Engines

Setup Flow

  1. Pick one engine from the cards above and set up its wrapper from servers/tts/.
  2. In /providers, choose Add New Custom Endpoint and use API Compatibility tts-clone.
  3. Select the saved endpoint and use its model dropdown to add its Speech model, Voice Source Mode, and Script Markup.
  4. Open /config > Models > Switch Models and select the registered Speech model.

Network access

Every TTS local server binds to 127.0.0.1 by default, so only a bot on the same machine can reach it. When TomoriBot runs in Docker or on another machine, set TOMORI_TTS_HOST=0.0.0.0 for the server and register the endpoint with the host’s address, for example http://host.docker.internal:8015.

The servers have no authentication: anyone who can reach the port can run synthesis on your GPU. Bind off loopback only on a network you trust, and put an authenticating reverse proxy in front of any server you expose beyond it.

Each server reads its own port variable (CHATTERBOX_PORT, QWEN3TTS_PORT, IRODORI_TTS_PORT, FISH_S2_PORT, VOXCPM2_PORT, COSYVOICE3_PORT, MOSS_TTS_PORT), so several can run at once from one .env. bun run launch passes .env to the servers it starts; a server started by hand reads only its shell environment.

Persona Voice Flow

For clone engines, add and assign one voice sample per persona:

  1. Prepare a clean clip with one speaker and no background music. Each engine documents its own best length, from 3 to 30 seconds depending on the model; the per-engine guides below carry the numbers.
  2. Open /config under Models > TTS Parameters & Voices and upload the sample. Any audio format is accepted; TomoriBot converts it to mono WAV, up to 130 seconds.
  3. Provide a matching transcript when the engine benefits from it. Fish S2 Pro forwards the stored reference transcript into the cloning request, and VoxCPM2 uses it for higher-fidelity Ultimate Cloning.
  4. Open /config under Persona > Voice, then choose the persona and voice sample.

The 130-second upload ceiling is TomoriBot’s, not any engine’s. Each engine applies its own reference-audio limit when the clip is sent, and those limits differ enough that a clip which suits one engine can be rejected or partly ignored by another. Check the engine’s guide before uploading a long clip.

The upload size limit and the duration ceiling are independent, and the size check runs first. Uncompressed WAV costs roughly 10 MB per minute as stereo 44.1 kHz, so a 130-second DAW export can be around 22 MB and be refused as too large while a compressed export of the same clip fits easily. SPEECH_SAMPLE_MAX_MB raises that size limit (10 MB by default). For a long reference clip, export it as FLAC, MP3, or OGG rather than WAV: TomoriBot converts any accepted format to mono WAV at 22.05 kHz before storing it, so the compressed upload costs nothing in quality.

For Qwen3-TTS, MOSS-VoiceGenerator, IrodoriTTS, or VoxCPM2 VoiceDesign, skip the audio sample and use Persona > Voice in /config to save a natural-language voice description for the persona. CosyVoice 3 uses Clone mode with a reference sample and accepts one-off delivery direction through voice_instructions.

ElevenLabs users should choose Add New Provider and ElevenLabs in /providers; it registers the speech and transcription endpoints together.

TomoriBot strips Discord custom emoji syntax such as :pepega: or <:pepega:123456789012345678> from generated voice scripts before synthesis. Unicode emojis are also stripped unless the speech endpoint uses emoji markup, which is intended for IrodoriTTS. Bracket tags are preserved for endpoints such as Chatterbox Turbo and Fish S2 Pro when their Script Markup is set to Bracket Tags.

VoxCPM2 uses Plain Script Markup. Its Voice Design and speaking-style controls travel through TomoriBot’s instruct field, and the VoxCPM2 wrapper converts them to the model’s native parenthesized natural-language control syntax.

CosyVoice 3 also uses Plain Script Markup. It maps transcript-backed references to zero-shot cloning, transcript-free references to cross-lingual cloning, and voice_instructions to its instruction-conditioned path. Inline bracket tags are removed because CosyVoice instructions apply to the whole utterance.

Engine Comparison

For side-by-side empirical latency benchmarks, audio sample playback, and a detailed feature matrix comparing all supported engines, see TTS Engine Comparison.