aboutsummaryrefslogtreecommitdiff
path: root/README.md
diff options
context:
space:
mode:
Diffstat (limited to 'README.md')
-rw-r--r--README.md135
1 files changed, 135 insertions, 0 deletions
diff --git a/README.md b/README.md
new file mode 100644
index 0000000..8307b11
--- /dev/null
+++ b/README.md
@@ -0,0 +1,135 @@
+# Qwen3 Audiobook Converter
+
+Convert TXT, PDF, EPUB, DOCX, and DOC files into audiobooks using the Qwen3-TTS voice model.
+
+Original project: [https://github.com/WhiskeyCoder/Qwen3-Audiobook-Converter](https://github.com/WhiskeyCoder/Qwen3-Audiobook-Converter
+). This repo just has minor fixes, flags, and doc changes. It also optionally splits the qwen3-tts server into two processes running models on different ports.
+
+## Overview
+
+The converter sends text extracted from your books to a locally running Qwen3-TTS server and assembles the returned audio into a single audiobook file.
+
+- Supported input: `.txt`, `.pdf`, `.epub`, `.docx`, `.doc`
+- Output: `.mp3`
+- Two voice modes:
+ - Custom voice: pre-built speakers
+ - Voice clone: clone a voice from a `.wav` reference audio file
+
+## Prerequisites
+
+- Python 3.12
+- ffmpeg
+- Enough VRAM to run the 1.7B model (~6GB)
+
+## Installation
+
+```bash
+# Arch Linux
+sudo pacman -S conda ffmpeg
+
+# Debian, conda must be installed separately
+sudo apt-get install ffmpeg
+```
+
+### Install Qwen3-TTS (Server)
+
+```bash
+conda create -n qwen3-tts python=3.12 -y
+conda activate qwen3-tts
+pip install -U qwen-tts
+```
+
+### Install the conversion script
+
+```bash
+git clone https://git.historia.vg/git/qwen3-audiobook-converter
+cd qwen3-audiobook-converter
+pip install -r requirements.txt
+```
+
+## Running the Qwen-TTS server
+
+The converter script talks to a Qwen3-TTS Gradio server that is run using `qwen-tts-demo`. Add `--no-flash-attn` if FlashAttention isn't installed (see below). The script expects the custom voice model and base model to be on different ports depending on which you're using:
+
+### Custom voice
+
+```bash
+conda activate qwen3-tts
+qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice --ip 127.0.0.1 --port 7860
+```
+
+### Voice clone
+
+```bash
+conda activate qwen3-tts
+qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-Base --ip 127.0.0.1 --port 7861
+```
+
+## Converting books
+
+Put your book files (epub, txt, etc.) in the `book_to_convert/` folder. Then run the script. The output mp3 goes to `audiobooks/`.
+
+### Custom voice
+
+```bash
+python audiobook_converter.py
+```
+
+Edit the parameters at the top of `audiobook_converter.py` to change which built-in voice is used.
+
+```
+CUSTOM_VOICE_SPEAKER = "Vivian" # Serena, Vivian, Uncle_Fu, Aiden, Ono_Anna, Sohee, Eric, Dylan
+CUSTOM_VOICE_LANGUAGE = "English"
+CUSTOM_VOICE_INSTRUCT = "Speak naturally and clearly, as if reading a dramatic book to an adult audience."
+```
+
+### Voice clone
+
+```bash
+python audiobook_converter.py --voice-clone --voice-sample path/to/reference.wav
+```
+
+Omit `--voice-sample-text` and Whisper will be used automatically to transcribe the reference audio (`faster_whisper` or `whisper`). If no Whisper backend is installed, it falls back to x-vector-only cloning.
+
+To skip automatic transcription explicitly, pass `--no-transcription`. This should be worse, but in my experience may give a preferable flatter tone to certain voices.
+
+You can override whisper by passing your own transcription with `--voice-sample-text "What the reference audio says"`
+
+### Playback speed
+
+Adjust the speed of the final audiobook without changing pitch (uses ffmpeg `atempo` before encoding). The normal-speed audiobook file is also preserved in the output directory.
+
+```bash
+python audiobook_converter.py --speed 0.9
+```
+
+## FlashAttention (optional)
+
+The server tries to use FlashAttention 2 by default, but `--no-flash-attn` works without it. On supported GPUs FlashAttention can give a modest speedup.
+
+1. Build from source (takes absolutely forever). If you run out of memory, lower MAX_JOBS until you don't.
+
+```bash
+conda activate qwen3-tts
+pip install ninja packaging psutil
+MAX_JOBS=4 pip install --no-build-isolation flash-attn
+```
+
+2. Or pip install a prebuilt wheel matching your torch / CUDA / Python / CXX11-ABI combination:
+
+```bash
+python -c "import torch; print(torch.__version__, torch.version.cuda, torch._C._GLIBCXX_USE_CXX11_ABI)"
+```
+
+Official wheels: https://github.com/Dao-AILab/flash-attention/releases (pick `cp312` + matching `cuX` + `torchX.Y` + `cxx11abiTRUE/FALSE`).
+
+Third-party wheels: https://mjunya.com/flash-attention-prebuild-wheels/ (hosted at https://github.com/mjun0812/flash-attention-prebuild-wheels).
+
+## License
+
+MIT
+
+## Credits
+
+[Qwen3-Audiobook-Converter](https://github.com/WhiskeyCoder/Qwen3-Audiobook-Converter) by WhiskeyCoder.
+[Qwen3-TTS](https://github.com/QwenLM/Qwen3-TTS) voice model.