# MOSS-TTS v1.5 with a pinned AR-engine memory budget. # # The upstream pipeline (MossTTSPipelineConfig) colocates its three stages # (preprocessing -> tts_engine -> vocoder) in one process on GPU 0 and leaves # the engine's sglang mem_fraction_static unset: the static pool (weights + # KV cache) is auto-sized to nearly all free VRAM at boot. On a 24 GB card # that leaves only tens of MiB free once the engine's CUDA graphs and the # colocated vocoder are resident — the first /v1/audio/speech request aborts # with "CUDA out of memory. Tried to allocate ~100 MiB" and every retry fails # identically (the same failure the qwen3_tts/higgs/zonos2 vendored configs # fix; note moss_tts_local does NOT need this — MossTTSLocalPipelineConfig # already budgets its colocated stages explicitly: 0.15/0.67/0.18). # # 0.70 pins the static pool at ~70% of the card (~16.5 GB on a 24 GB GPU — # a KV pool far larger than any narration request needs) and leaves ~7 GB # for the vocoder, CUDA graphs, transient allocations, and other GPU # processes. Precedent: dots_tts.yaml pins this same knob. config_cls: MossTTSPipelineConfig model_path: OpenMOSS-Team/MOSS-TTS-v1.5 stages: tts_engine: engine: mem_fraction_static: 0.70