CoolFace
Datasetpublic

Pendrokar/open_tts_tracker

above models sorted by the amount of capabilities; #legend Cloned the GitHub repo for easier viewing and embedding the above table as once requested by @reach-vb: https://github.com/Vaibhavs10/open-tts-tracker/issues/30#issuecomment-1946367525 πŸ—£οΈ Open TTS Tracker A one stop shop to track all open-access/ source Text-To-Speech (TTS) models as they come out. Feel free to make a PR for all those that aren't linked here. This is aimed as a resource to increase awareness for these… See the full description on the dataset page: https://huggingface.co/datasets/Pendrokar/open_tts_tracker.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
30likes163downloads
Dataset Card

above models sorted by the amount of capabilities; #legend

Cloned the GitHub repo for easier viewing and embedding the above table as once requested by @reach-vb: https://github.com/Vaibhavs10/open-tts-tracker/issues/30#issuecomment-1946367525


πŸ—£οΈ Open TTS Tracker

A one stop shop to track all open-access/ source Text-To-Speech (TTS) models as they come out. Feel free to make a PR for all those that aren't linked here.

This is aimed as a resource to increase awareness for these models and to make it easier for researchers, developers, and enthusiasts to stay informed about the latest advancements in the field.

[!NOTE] This repo will only track open source/access codebase TTS models. More motivation for everyone to open-source! πŸ€—

Some of the models are also being battle tested at TTS arenas hosted on HuggingFace:

  • β€”πŸ—£οΈ Speech Arena Leaderboard - External Speech Arena made by Artificial Analysis, who validate mostly proprietary TTS models
  • β€”πŸ€—πŸ† TTS Spaces Arena - Mostly uses online HuggingFace Spaces, which have the Gradio API enabled
  • β€”πŸ† TTS Arena - Also has a Conversational tab that tests Podcast type of TTS
  • β€”πŸŽ€ Expressive TTS Arena - Evaluates TTS models that "best captures the emotion, nuance, and expressiveness of human speech"
  • β€”πŸ‘₯ Voice Clone Arena - Allows choosing a specific or random voice file sample for zero-shot Open TTS
  • β€”TTS Arena v1 Archive - The above TTS-Arena-V2 was rewritten from scratch and does not use Gradio anymore

Also battle tested outside HuggingFace:

  • β€”[Artificial Analysis](https://artificialanalysis.ai/text-to-speech) - Analysis and comparison of Text to Speech generation models & API providers. Artificial Analysis has analyzed text to speech models and hosting providers across quality, generation time, and price.

And in automated benchmarks:

NameGitHubWeightsLicenseFine-tuneLanguagesPaperDemoIssues
AI4BharatRepoHubMITYesIndicPaperDemo
AmphionRepoHubMITNoMultilingualPaperπŸ€— Space
BarkRepoHubMITNoMultilingualPaperπŸ€— Space
ChatterboxRepoπŸ€— HubMITNoEnglishNot AvailableπŸ€— Space
CSMRepoHubApache 2.0NoEnglishNot AvailableπŸ€— SpaceIssues
EmotiVoiceRepoGDriveApache 2.0YesZH + ENNot AvailableNot AvailableSeparate GUI agreement
F5-TTSRepoHubMITYesZH + ENPaperπŸ€— Space
Fish SpeechRepoHubApache 2.0YesMultilingualNot AvailableπŸ€— Space
Glow-TTSRepoGDriveMITYesEnglishPaperGH Pages
GPT-SoVITSRepoHubMITYesMultilingualNot AvailableNot Available
HierSpeech++RepoGDriveMITNoKR + ENPaperπŸ€— Space
IMS-ToucanRepoGH releaseApache 2.0YesALL\*PaperπŸ€— Space, πŸ€— Space\*
KokoroRepoHubApache 2.0NoMultilingualPaperπŸ€— SpaceGPL-licensed phonemizer
LLaSARepoπŸ€— HubCC-BY-NC 4.0NoMultilingualPaperπŸ€— Space
MahaTTSRepoHubApache 2.0NoEnglish + IndicNot AvailableRecordings, Colab
MaskGCT (Amphion)RepoHubCC-BY-NC 4.0NoMultilingualPaperπŸ€— Space
Matcha-TTSRepoGDriveMITYesEnglishPaperπŸ€— SpaceGPL-licensed phonemizer
MeloTTSRepoHubMITYesMultilingualNot AvailableπŸ€— Space
MetaVoice-1BRepoHubApache 2.0YesMultilingualNot AvailableπŸ€— Space
Neural-HMM TTSRepoGitHubMITYesEnglishPaperGH Pages
OpenVoiceRepoHubMITNoMultilingualPaperπŸ€— Space
OuteTTSRepoπŸ€— HubApache 2.0NoMultilingualπŸ€— Space
OrpheusRepoHubApache 2.0YesEnglishPaperColabIssues
OverFlow TTSRepoGitHubMITYesEnglishPaperGH Pages
Parler TTSRepoHubApache 2.0YesEnglishNot AvailableπŸ€— Space
pflowTTSUnofficial RepoGDriveMITYesEnglishPaperNot AvailableGPL-licensed phonemizer
PhemeRepoHubCC-BYYesEnglishPaperπŸ€— Space
PiperRepoHubMITYesMultilingualNot AvailableπŸ€— SpaceGPL-licensed phonemizer
RAD-MMMRepoGDriveMITYesMultilingualPaperJupyter Notebook, Webpage
RAD-TTSRepoGDriveMITYesEnglishPaperGH Pages
SileroRepoGH linksCC BY-NC-SANoMultilingualNot AvailableNot AvailableNon Commercial
StyleTTS 2RepoHubMITYesEnglishPaperπŸ€— SpaceGPL-licensed phonemizer
Tacotron 2Unofficial RepoGDriveBSD-3YesEnglishPaperWebpage
TorToiSe TTSRepoHubApache 2.0YesEnglishTechnical reportπŸ€— Space
TTTSRepoHubMPL 2.0NoMultilingualNot AvailableColab, πŸ€— Space
VALL-EUnofficial RepoNot AvailableMITYesNAPaperNot Available
VITS/ MMS-TTSRepoHub / MMSApache 2.0YesEnglishPaperπŸ€— SpaceGPL-licensed phonemizer
WhisperSpeechRepoHubMITNoMultilingualNot AvailableπŸ€— Space, Recordings, Colab
XTTSRepoHubCPMLYesMultilingualPaperπŸ€— SpaceNon Commercial
xVASynthRepoHubGPL-3.0YesMultilingualNot AvailableπŸ€— SpaceBase model trained on non-permissive datasets
ZonosRepoπŸ€— HubApache 2.0NoMultilingualπŸ€— Space
  • β€”Multilingual - Amount of supported languages is ever changing, check the Space and Hub for which specific languages are supported
  • β€”ALL - Claims to support all natural languages; this may not include artificial/contructed languages

Also to find a model for a specific language, filter out the TTS models hosted on HuggingFace: <https://huggingface.co/models?pipeline_tag=text-to-speech&language=en&sort=trending>


Legend

For the #above TTS capability table. Open the viewer in another window or even another monitor to keep both it and the legend in view.

  • β€”Processor ⚑ - Inference done by
  • β€”CPU (CPUs = multithreaded) - All models can be run on CPU, so real-time factor should be below 2.0 to qualify for CPU tag, though some more leeway can be given if it supports audio streaming
  • β€”CUDA by NVIDIAβ„’
  • β€”ROCm by AMDβ„’, also see ONNX Runtime HF guide
  • β€”Phonetic alphabet πŸ”€ - Phonetic transcription that allows to control pronunciation of words before inference
  • β€”IPA - International Phonetic Alphabet
  • β€”ARPAbet - American English focused phonetics
  • β€”Insta-clone πŸ‘₯ - Zero-shot model for quick voice cloning; Other allow only fusing multiple models or voice samples together
  • β€”Emotion control 🎭 - Able to force an emotional state of speaker
  • β€”πŸŽ­ <# emotions> ( 😑 anger; πŸ˜ƒ happiness; 😭 sadness; 😯 surprise; 🀫 whispering; 😊 friendlyness )
  • β€”πŸŽ­πŸ‘₯ strict insta-clone switch - cloned on sample with specific emotion; may sound different than normal speaking voice; no ability to go in-between states
  • β€”πŸŽ­πŸ“– strict control through prompt - prompt input parameter
  • β€”Prompting πŸ“– - Also a side effect of narrator based datasets and a way to affect the emotional state
  • β€”πŸ“– - Prompt as a separate input parameter
  • β€”πŸ—£πŸ“– - The prompt itself is also spoken by TTS; ElevenLabs docs
  • β€”Streaming support 🌊 - Can playback audio while it is still being generated
  • β€”Speech control 🎚 - Ability to change the pitch, duration, etc. for the whole and/or per-phoneme of the generated speech
  • β€”Voice conversion / Speech-To-Speech 🦜 - Streaming support implies real-time S2S; S2T=>T2S does not count
  • β€”Longform synthesis πŸ“œ - Able to synthesize whole paragraphs, as some TTS models tend to break down after a certain audio length limit

Example if the proprietary ElevenLabs were to be added to the capabilities table: | Name | Processor<br>⚑ | Phonetic alphabet<br>πŸ”€ | Insta-clone<br>πŸ‘₯ | Emotional control<br>🎭 | Prompting<br>πŸ“– | Speech control<br>🎚 | Streaming support<br>🌊 | Voice conversion<br>🦜 | Longform synthesis<br>πŸ“œ | |---|---|---|---|---|---|---|---|---| --- | |ElevenLabs|CUDA|IPA, ARPAbet|πŸ‘₯|πŸŽ­πŸ“–|πŸ—£πŸ“–|🎚 stability, voice similarity|🌊|🦜|πŸ“œ Projects|

More info on how the capabilities table came about can be found within the GitHub Issue.

train_data Legend

Legend for the separate TTSDS Datasets (_train_data_ viewer GitHub)

  • β€”πŸŒ Multilingual
  • β€”The ISO codes of languages the model is capable off. ❌ if English only.
  • β€”πŸ“š Training Amount (k hours)
  • β€”The number of hours the model was trained on
  • β€”πŸ§  Num. Parameters (M)
  • β€”How many parameters the model has, excluding vocoder and text-only components
  • β€”πŸŽ― Target Repr.
  • β€”Which output representations the model uses, for example audio codecs or mel spectrograms
  • β€”πŸ“– LibriVox Only
  • β€”If the model was trained on librivox-like (audiobook) data alone
  • β€”πŸ”„ NAR
  • β€”If the model has a significant non-autoregressive component
  • β€”πŸ” AR
  • β€”If the model has a significant autoregressive component
  • β€”πŸ”‘ G2P
  • β€”If the model uses G2P (phone inputs)
  • β€”πŸ§© Language Model
  • β€”If an LM-like approach is used (next token prediction)
  • β€”πŸŽ΅ Prosody Prediction
  • β€”If prosodic correlates such as pitch or energy are predicted
  • β€”πŸŒŠ Diffusion
  • β€”If diffusion is used (outside vocoder)
  • —⏱️ Delay Pattern
  • β€”If a delay pattern is used for audio codes (see Lyth & King, 2024)

Please [create pull requests](https://huggingface.co/datasets/Pendrokar/openttstracker/edit/main/README.md) to update the info on the models within this dataset!