datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asante-twi-bible-speech-text
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
asante-twi-ttsasante-twi-bible-speech-phonemes
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Asante Twi Bible Speech — Phonemes
Phoneme-labelled version of
ghananlpcommunity/asante-twi-bible-speech-text,
built for training a wav2vec2 (CTC) phoneme recogniser for Asante Twi.
Each example adds a phonemes column: a… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/asante-twi-bible-speech-phonemes.ghana-speech-asante-twi-transcribedasante-twi-ttsasante-twi-bibleAsante Twi Bible Audio
This dataset is comprised of audio recorded from [Youversion's website](https://www.bible.com/bible/2094/), which hosts audio and written copies of the Bible in multiple languages. It includes audio and matching transcriptions of the Bible, useful for Automatic Speech Recognition (ASR) and Speech Generation applications.
In its first iteration, it contains Romans 1 - 4 in the Asante Twi Nkwa Asɛm version. As more data is preprocessed, the dataset will grow to include… See the full description on the dataset page: https://huggingface.co/datasets/jdapaah/asante-twi-bible.asante_twi_bible
Dataset Card for "asante_twi_bible"
More Information needed
bibletts-asante-twi-repaired
BibleTTS Asante Twi — Repaired Transcripts
The Asante Twi transcripts released with BibleTTS have had the
characters ɛ (U+025B) and ɔ (U+0254) stripped out. This dataset restores them.
Audio is not included. This is a drop-in replacement for the .txt files that ship with the
BibleTTS Asante Twi package, matched by clip ID.
The problem
Both are Twi vowels, and both are required by the orthography. Measured across the released
Asante Twi transcripts:
Character… See the full description on the dataset page: https://huggingface.co/datasets/danieldzikunuofmarvel/bibletts-asante-twi-repaired.asante-twi-bible-tags-v3asante-twi-wavtokenizer-aligned-newasante-twi-word-alignedbibletts-asante-twi-max29secs-total9hrs-sr22050
BibleTTS Asante Twi Dataset
Dataset Information
This dataset is derived from the BibleTTS corpus, specifically focusing on Asante Twi speech data. The original BibleTTS is a large, high-fidelity, multilingual, and uniquely African speech corpus.
Total Duration: {total_hours:.2f} hours ({total_hours*60:.1f} minutes)
Number of Files: {file_count:,}
Sample Rate: {sample_rate:,} Hz
Max File Duration: {max_duration:.1f} seconds
Format: WAV files with corresponding… See the full description on the dataset page: https://huggingface.co/datasets/hci-lab-dcug/bibletts-asante-twi-max29secs-total9hrs-sr22050.Asante_Twi_Collected_Testbibletts-asante-twi-max29secs-total57hrs-sr22050asante-twi-bible-tags-v1asante-twi-bible-tags-v2asante-twi-yarngpt-wordlevel-tokenized
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Asante Twi — YarnGPT Word-Level Pre-Tokenized TTS Dataset
This dataset contains the ghananlpcommunity/asante-twi-bible-speech-text audio dataset fully processed into word-level aligned, pre-tokenized integer sequences for training… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/asante-twi-yarngpt-wordlevel-tokenized.asante-twi-llama-wordlevel-tokenized
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
asante_twi-bible-tts-200
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
asante-twi-mfa-wordlevel-tokenized
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
asante-twi-neucodec-encodedasante-twi-wavtokenizer-alignedasante-twi-yarngpt-aligned
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Asante Twi — YarnGPT Word-Level Aligned Dataset
This is an intermediate dataset produced during the preparation of a YarnGPT-style TTS model for Asante Twi. It contains the original audio from… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/asante-twi-yarngpt-aligned.asante-twi-tts-dataset-tokenised-tagged
