Parler-TTS is a lightweight text-to-speech (TTS) model that can generate high-quality, natural sounding speech in the style of a given speaker (gender, pitch, speaking style, etc).
Parler-TTS Mini v0.1, is the first iteration Parler-TTS model trained using 10k hours of narrated audiobooks. It generates high-quality speech with features that can be controlled using a simple text prompt (e.g. gender, background noise, speaking rate, pitch and reverberation).
To improve the prosody and naturalness of the speech further, we're scaling up the amount of training data to 50k hours of speech. The v1 release of the model will be trained on this data, as well as inference optimisations, such as flash attention and torch compile.
This work is both scalable and easily modifiable and will hopefully help the TTS research community explore new ways of conditionning speech synthesis.
All of the datasets, pre-processing, training code and weights are released publicly under permissive license, enabling the community to build on our work and develop their own powerful TTS models.
1 reply
·
MINIMAX H3 · AI VIDEO
Turn ideas into video — start free
Create short AI videos from text or first and last frames, with synchronized dialogue, effects, and ambience.
2020 free creditsEnough for two default 5-second generations.
⚡Quick to startWrite your prompt and add references before signing in.
HDFlexible outputChoose 5–15 seconds, five ratios, and up to 1080P.