datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.multilingual-single-sentences
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
multilingual_single_sentences
This dataset consists of single-sentence completions spanning a diverse range of languages, including Russian, German, Spanish, Korean, Chinese, English, French, and Japanese. The content varies widely, covering topics from historical restoration and travel logistics to proverbs and daily observations. Each entry is presented as an isolated text… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/multilingual-single-sentences.marathi-czech-sentences
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
marathi_czech_sentences
This dataset contains short sentences and questions primarily in Marathi and Czech, covering various conversational contexts. The samples include inquiries about objects, actions, and origins, as well as exclamations and statements. It appears to be a multilingual collection focused on everyday dialogue structures.
Dataset size
There are 3… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/marathi-czech-sentences.opentts-phonemized-sentencesPhonemized version of https://huggingface.co/datasets/speech-uk/text-to-speech-sentences with some additional fields.
