aboutsummaryrefslogtreecommitdiff
path: root/app/docs/backend-audiocpp.md
blob: deb8fbe75f8b5e129f137519f7b33941a62fb015 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
# Backend Option 1: audio.cpp

`--backend audiocpp` talks to `audiocpp_server` from [audio.cpp](https://github.com/0xShug0/audio.cpp), which hosts numerous TTS model families.

The easiest way is the TUI: run `python audiobook.py`, choose **Set up a backend… → audio.cpp**, and it clones `audio.cpp` into `app/audio.cpp` (or reuses an existing checkout), builds `audiocpp_server`, lets you pick model families/packages from an expandable checkbox tree (reading the checkout's `model_specs/`), transcribes `.wav` voices with `whisper`, writes `server.json` into the checkout, syncs `app/converter/config.py`, and prints the launch command (the hub can also start the server for you via the **Start/Stop Backend Servers** menu or automatically when converting). Run it directly with `python app/backends/audiocpp.py` (flags like `--wavs`, `--families`, `--build-backend`, `--clone` skip the corresponding screens for scripting). The TUI runs in the managed `app/envs/tts` venv, which includes `whisper` via `requirements.txt`; for a manual setup, make sure `whisper` (or `faster_whisper`) is installed in the environment you run the wizard from. The Qwen3-TTS model tree also offers hosting the VoiceDesign package as a `vdes` entry.

If you prefer to install the backend yourself (in your own environment, not the managed venv), the manual steps are below. Either way the hub detects a running server by its port, so a manually-installed backend works once its server is up.

### Download and build audiocpp_server

Download and build `audiocpp_server` for your platform and backend `(cuda, vulkan, hip, cpu)`. Check [audio.cpp's readme](https://github.com/0xShug0/audio.cpp) for details. I'm using one of the helper scripts:

```bash
git clone https://github.com/0xShug0/audio.cpp
cd audio.cpp
scripts/build_linux.sh --backend cuda --target audiocpp_server
```

### Install models

Download model packages with the python model manager script from the audio.cpp checkout. Each installs to `./models`. Here are two examples, Higgs Audio and Qwen3-TTS:

```bash
python tools/model_manager_v2.py install higgs_audio_tts_4b_q8_0
python tools/model_manager_v2.py install qwen3_tts_1_7b_base_q8_0
python tools/model_manager_v2.py install qwen3_tts_1_7b_customvoice_q8_0
```

You can run `python tools/model_manager_v2.py list` to see all available models.

### Create server.json

Create a `server.json` config file. One server can host multiple models and multiple cloned voices. The `id:` fields are the model names you will set for `tts-audiobook-generator` with `--model`.

```json
{
  "host": "127.0.0.1",
  "port": 8080,
  "backend": "cuda",
  "lazy_load": true,
  "voice_dir": "/path/to/clone/wavs",
  "models": [
    {
      "id": "higgs",
      "family": "higgs_audio_tts",
      "path": "models/Higgs-Audio-v3-TTS-4B-GGUF",
      "task": "tts",
      "mode": "offline"
    },
    {
      "id": "qwen",
      "family": "qwen3_tts",
      "path": "models/Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUF",
      "task": "tts",
      "mode": "offline"
    },
    {
      "id": "qwen-clone",
      "family": "qwen3_tts",
      "path": "models/Qwen3-TTS-12Hz-1.7B-Base-GGUF",
      "task": "tts",
      "mode": "offline"
    }
  ]
}
```

### Run audio.cpp and the audiobook script

Run the server with this config file. The `audiocpp_server` path will be slightly different depending on your platform and build options:

```bash
./build/linux-cuda-release/bin/audiocpp_server --config server.json
```

In a different terminal, run `audiobook.py`. Pick the TTS `--model` and `--voice` from server.json:

```bash
# Higgs Audio (clone-only)
python audiobook.py --backend audiocpp --model higgs --voice narrator

# Qwen3-TTS built-in speaker
python audiobook.py --backend audiocpp --model qwen

# Qwen3-TTS voice cloning
python audiobook.py --backend audiocpp --model qwen-clone --voice narrator

# Qwen-TTS voice design
python audiobook.py --backend audiocpp --model qwen-design \
    --instructions "A warm adult female narrator with a British accent"
```

The hub also works with an audio.cpp server that runs somewhere else (another checkout, another machine) as long as it answers on the configured port: when there is no local `server.json`, the convert menus query the running server directly (`GET /v1/models` and `GET /v1/audio/voices`) instead of reading one. On the CLI, pass `--model`/`--voice` matching that server's config.

Before converting, `audiobook.py` asks the server to unload all currently loaded models (`POST /v1/tasks/unload_all_models`) so models left resident by earlier runs free their memory (e.g. VRAM on GPU backends) and only the selected entry loads. A server without that endpoint, or one busy unloading, only produces a warning.