aboutsummaryrefslogtreecommitdiff
path: root/docs/backend-audiocpp.md
diff options
context:
space:
mode:
authorhistoria <historiavg@proton.me>2026-08-24 02:59:26 -0400
committerhistoria <historiavg@proton.me>2026-08-24 02:59:26 -0400
commitf00249db9d1ea051d29aa1bcca869fc4b88e83eb (patch)
treea75f076fac1b63e0b4bf2eb8f54affbcc681a891 /docs/backend-audiocpp.md
parent9dd4f9595be3b1d76a3a07dc3eca90cfaf8a3f97 (diff)
downloadtts-audiobook-generator-f00249db9d1ea051d29aa1bcca869fc4b88e83eb.tar.gz
refactor: add app directory, dir structure change
Diffstat (limited to 'docs/backend-audiocpp.md')
-rw-r--r--docs/backend-audiocpp.md91
1 files changed, 0 insertions, 91 deletions
diff --git a/docs/backend-audiocpp.md b/docs/backend-audiocpp.md
deleted file mode 100644
index ee511bb..0000000
--- a/docs/backend-audiocpp.md
+++ /dev/null
@@ -1,91 +0,0 @@
-# Backend Option 1: audio.cpp
-
-`--backend audiocpp` talks to `audiocpp_server` from [audio.cpp](https://github.com/0xShug0/audio.cpp), which hosts numerous TTS model families.
-
-The easiest way is the TUI: run `python audiobook.py`, choose **Set up a backend… → audio.cpp**, and it clones `audio.cpp` into `./audio.cpp` (or reuses an existing checkout), builds `audiocpp_server`, lets you pick model families/packages from an expandable checkbox tree (reading the checkout's `model_specs/`), transcribes `.wav` voices with `whisper`, writes `server.json` into the checkout, syncs `converter/config.py`, and prints the launch command (the hub can also start the server for you via the **Server** menu or automatically when converting). Run it directly with `python -m backends.audiocpp` (flags like `--wavs`, `--families`, `--build-backend`, `--clone` skip the corresponding screens for scripting). The TUI runs in the managed `envs/tts` venv, which includes `whisper` via `requirements.txt`; for a manual setup, make sure `whisper` (or `faster_whisper`) is installed in the environment you run the wizard from. The Qwen3-TTS model tree also offers hosting the VoiceDesign package as a `vdes` entry.
-
-If you prefer to install the backend yourself (in your own environment, not the managed venv), the manual steps are below. Either way the hub detects a running server by its port, so a manually-installed backend works once its server is up.
-
-### Download and build audiocpp_server
-
-Download and build `audiocpp_server` for your platform and backend `(cuda, vulkan, hip, cpu)`. Check [audio.cpp's readme](https://github.com/0xShug0/audio.cpp) for details. I'm using one of the helper scripts:
-
-```bash
-git clone https://github.com/0xShug0/audio.cpp
-cd audio.cpp
-scripts/build_linux.sh --backend cuda --target audiocpp_server
-```
-
-### Install models
-
-Download model packages with the python model manager script from the audio.cpp checkout. Each installs to `./models`. Here are two examples, Higgs Audio and Qwen3-TTS:
-
-```bash
-python tools/model_manager_v2.py install higgs_audio_tts_4b_q8_0
-python tools/model_manager_v2.py install qwen3_tts_1_7b_base_q8_0
-python tools/model_manager_v2.py install qwen3_tts_1_7b_customvoice_q8_0
-```
-
-You can run `python tools/model_manager_v2.py list` to see all available models.
-
-### Create server.json
-
-Create a `server.json` config file. One server can host multiple models and multiple cloned voices. The `id:` fields are the model names you will set for `tts-audiobook-generator` with `--model`.
-
-```json
-{
- "host": "127.0.0.1",
- "port": 8080,
- "backend": "cuda",
- "lazy_load": true,
- "voice_dir": "/path/to/clone/wavs",
- "models": [
- {
- "id": "higgs",
- "family": "higgs_audio_tts",
- "path": "models/Higgs-Audio-v3-TTS-4B-GGUF",
- "task": "tts",
- "mode": "offline"
- },
- {
- "id": "qwen",
- "family": "qwen3_tts",
- "path": "models/Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUF",
- "task": "tts",
- "mode": "offline"
- },
- {
- "id": "qwen-clone",
- "family": "qwen3_tts",
- "path": "models/Qwen3-TTS-12Hz-1.7B-Base-GGUF",
- "task": "tts",
- "mode": "offline"
- }
- ]
-}
-```
-
-### Run audio.cpp and the audiobook script
-
-Run the server with this config file. The `audiocpp_server` path will be slightly different depending on your platform and build options:
-
-```bash
-./build/linux-cuda-release/bin/audiocpp_server --config server.json
-```
-
-In a different terminal, run `audiobook.py`. Pick the TTS `--model` and `--voice` from server.json:
-
-```bash
-# Higgs Audio (clone-only)
-python audiobook.py --backend audiocpp --model higgs --voice narrator
-
-# Qwen3-TTS built-in speaker
-python audiobook.py --backend audiocpp --model qwen
-
-# Qwen3-TTS voice cloning
-python audiobook.py --backend audiocpp --model qwen-clone --voice narrator
-
-# Qwen-TTS voice design
-python audiobook.py --backend audiocpp --model qwen-design \
- --instructions "A warm adult female narrator with a British accent"
-```