aboutsummaryrefslogtreecommitdiff
path: root/app/docs/backend-qwen.md
blob: d0e7c924b55de9558f367c3394a164ae7322d033 (plain)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
# Backend Option 2: Qwen3-TTS

The easiest way is to run `python audiobook.py` → **Configure backends… → Install Backend → qwen-tts** (or `python app/backends/qwen.py`): the TUI pip-installs `qwen-tts` into its managed venv (`app/envs/tts`) — that's all there is to it, the install asks no questions. The two demo ports live in `app/converter/config.py` (edit them in the hub's **Settings** screen), and you pick the built-in speaker per run on the **Generate audiobooks** screen (it defaults to `SPEAKER` in `app/converter/config.py`). You can also start the server from the hub's **Start/Stop Backend Servers** menu, or let a conversion start it automatically.

If you prefer to install the backend yourself (in your own environment, not the managed venv), the manual steps are below. Either way the hub detects a running server by its port, so a manually-installed backend works once its server is up. To use demo servers on another machine, set `QWEN_REMOTE_URL`/`CLONE_REMOTE_URL` in `app/converter/config.py` to their `host:port` (defaults `127.0.0.1:7860`/`:7861`) — the hub probes each and offers the matching `qwen-tts [remote]` mode — or pass `--api-url` on the CLI.

Install qwen-tts with pip into your environment:

```bash
python -m venv audiobook
source audiobook/bin/activate
pip install -U qwen-tts
```

Run the backend with `qwen-tts-demo`. Add `--no-flash-attn` if FlashAttention isn't installed (see below). Note that the Base model and CustomVoice model run on different ports.

## Voice clone

```bash
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-Base --ip 127.0.0.1 --port 7861 [--no-flash-attn]
```

Then in another terminal:

```bash
python audiobook.py --backend qwen --clone reference.wav
```

The reference `.wav` should be ~10-15 seconds (3 second minimum, 60 second maximum; ~15 seconds is ideal). Longer is **not** better.

Whisper (`faster_whisper` or `whisper`) is used automatically to transcribe the reference audio. Without a Whisper backend it falls back to x-vector-only cloning. Override with `--transcription "What the .wav says"` or skip transcription with `--no-transcription`.

## Custom voice (i.e. built-in voice)

```bash
source audiobook/bin/activate
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --ip 127.0.0.1 --port 7860 [--no-flash-attn]
```

```bash
python audiobook.py --backend qwen
```

Change the voice settings in `app/converter/config.py`.

## Optional: FlashAttention for qwen-tts-demo server

FlashAttention provides a *small* speed boost on the `qwen` backend. It is **not** relevant with other backends, and switching to either of those will provide a bigger speed boost.

`qwen-tts-demo` server tries to use FlashAttention 2 by default and requires `--no-flash-attn` without it. You have two options to install FlashAttention in your python environment:

1. Build from source (takes absolutely forever). If you run out of memory, lower MAX_JOBS until you don't.

```bash
source audiobook/bin/activate
pip install ninja packaging psutil
MAX_JOBS=4 pip install --no-build-isolation flash-attn
```

2. pip install a prebuilt wheel matching your torch / CUDA / Python / CXX11-ABI combination:

```bash
source audiobook/bin/activate
python -c "import torch; print(torch.__version__, torch.version.cuda, torch._C._GLIBCXX_USE_CXX11_ABI)"
```

- [Official wheels](https://github.com/Dao-AILab/flash-attention/releases) - Pick `cp312` + matching `cuX` + `torchX.Y` + `cxx11abiTRUE/FALSE`
- [Third-party wheels](https://mjunya.com/flash-attention-prebuild-wheels/)