SGLang-Omni (--backend sglomni)
Runs an SGLang-Omni server (sgl-omni serve), one TTS model per server process. Installs into its own managed venv, app/envs/sglomni.
Run all commands from the project root (next to audiobook.py).
Requirements
- Linux with an NVIDIA GPU
- Python 3.10-3.12 (
sglang-omnirequires>=3.10,<3.13) - ~10 GB disk for the venv plus 1-10 GB per model of weights in
~/.cache/huggingface/hub
Install
- Create the venv and install the server:
uv venv --seed --python 3.12 app/envs/sglomni # or: python3.12 -m venv app/envs/sglomni
app/envs/sglomni/bin/python -m pip install --pre sglang-omni
- Install companion packages for the models you plan to host:
| Models | Into app/envs/sglomni |
System packages |
|---|---|---|
| All four Qwen3-TTS | bin/python -m pip install --no-deps sox einops qwen-tts==0.1.1 |
sox binary (e.g. sudo apt install sox) |
| Fish Speech S2-Pro, ZONOS2 | bin/python -m pip install descript-audiotools==0.7.2 descript-audio-codec==1.0.0 protobuf==6.33.6 |
ffmpeg on PATH |
- Download model weights:
app/envs/sglomni/bin/hf download <hf-repo> # repos in the table below
Models
Key (TAG --model) |
HF repo (--model-path) |
Config |
|---|---|---|
qwen3_tts_0_6b_customvoice |
Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice |
qwen3_tts_0_6b_customvoice.yaml |
qwen3_tts_0_6b_base |
Qwen/Qwen3-TTS-12Hz-0.6B-Base |
qwen3_tts_0_6b.yaml |
qwen3_tts_1_7b_base |
Qwen/Qwen3-TTS-12Hz-1.7B-Base |
qwen3_tts_1_7b.yaml |
qwen3_tts_1_7b_voicedesign |
Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign |
qwen3_tts_1_7b_voicedesign.yaml |
higgs_audio_v3_tts |
bosonai/higgs-audio-v3-tts-4b |
higgs_audio_v3_tts.yaml |
moss_tts |
OpenMOSS-Team/MOSS-TTS-v1.5 |
moss_tts.yaml |
moss_tts_local |
OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5 |
moss_tts_local.yaml |
voxtral_tts |
mistralai/Voxtral-4B-TTS-2603 |
voxtral_tts.yaml |
dots_tts_mf |
dots-studio/dots.tts-mf |
dots_tts.yaml |
fish_s2_pro |
fishaudio/s2-pro |
s2pro_tts.yaml |
zonos2 |
Zyphra/zonos2 |
none |
Config files live in app/backends/sglomni/configs/. ZONOS2 runs without a config; on GPUs below compute capability 8.9 (RTX 30xx, A100) pass --config app/backends/sglomni/configs/zonos2_bf16.yaml instead.
Run the server
# Example: Higgs Audio v3
app/envs/sglomni/bin/sgl-omni serve \
--model-path bosonai/higgs-audio-v3-tts-4b \
--config app/backends/sglomni/configs/higgs_audio_v3_tts.yaml \
--port 8100
# ZONOS2 (no config)
app/envs/sglomni/bin/sgl-omni serve --model-path Zyphra/zonos2 --port 8100
Use it
python audiobook.py --backend sglomni --api-url http://127.0.0.1:8100 --model <key>
