Ngoc-Son Nguyen1 Thanh V. T. Tran1 Hieu-Nghia Huynh-Nguyen1 Truong-Son Hy2 Van Nguyen1†
1 FPT Software AI Center, Vietnam 2 University of Alabama at Birmingham, USA
† Corresponding author
Install the required dependencies using Conda:
conda env create -f environment.yaml
conda activate diflow- Download the pre-trained FACodec model from HuggingFace, and place the checkpoint files in the following structure:
root/
└── models/
└── facodec/
└── checkpoints/
├── ns3_facodec_encoder.bin
└── ns3_facodec_decoder.bin
- Download the DiFlow-TTS model checkpoint from HuggingFace, and place it as follows:
root/
└── ckpts/
└── diflow-tts.ckpt
Note
DiFlow-TTS is trained on 470 hours of the LibriTTS dataset, which consists of predominantly neutral speech. As a result, it may not perform well on prompts with strong emotional expression.
To synthesize a sample with DiFlow-TTS, follow these steps:
-
Open the script:
scripts/synth_one_sample.sh -
Edit the following lines:
- Line 3: Set the path to the DiFlow-TTS checkpoint.
- Line 4: Set your input text.
- Line 5: Set the path to your reference speech prompt.
-
Run the script with:
CUDA_VISIBLE_DEVICES=0 bash scripts/synth_one_sample.shComing soon!
If you find this work useful, please cite:
@inproceedings{nguyen2026diflowtts,
title = {{DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching}},
author = {Son Nguyen and Thanh Tran and Nghia Huynh and Son Hy and Van Nguyen},
year = {2026},
booktitle = {{Interspeech 2026 [Long Track]}},
pages = {1307--1316},
doi = {10.21437/Interspeech.2026-1043},
issn = {2958-1796},
}