datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
twi-trigrams-speech-text-parallel
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/twi-trigrams-speech-text-parallel.pile_trigrams
trigrams
See https://confirmlabs.org/posts/catalog.html for details.
id0: the first token in the trigram
id1: the second token in the trigram
id2: the third token in the trigram
count: the number of times the trigram appears in The Pile.
makhuwa-trigrams-speech-text-parallel
Makhuwa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 154253 parallel speech-text pairs for Makhuwa, a language spoken primarily in Mozambique. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Makhuwa - vmw
Task: Speech… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/makhuwa-trigrams-speech-text-parallel.twi-trigrams-speech-text-parallel
Twi Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 166156 parallel speech-text pairs for Twi, a language spoken primarily in Ghana. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Twi - twi
Task: Speech Recognition, Text-to-Speech… See the full description on the dataset page: https://huggingface.co/datasets/fiifinketia/twi-trigrams-speech-text-parallel.chichewa-trigrams-speech-text-parallel
Chichewa Trigrams Speech-Text Parallel Dataset
Dataset Description
This dataset contains 132549 parallel speech-text pairs for Chichewa, a language spoken primarily in Malawi. The dataset consists of audio recordings of trigram segments (3-word sequences) paired with their corresponding text transcriptions, making it suitable for automatic speech recognition (ASR) and text-to-speech (TTS) tasks.
Dataset Summary
Language: Chichewa - ny
Task: Speech Recognition… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/chichewa-trigrams-speech-text-parallel.EDAN20_Lab2_Uni_Bi_Trigramssynthetic_trigrams_alpha_1.5synthetic_trigrams_alpha_0.5synthetic_trigrams_v2twi-trigrams-speech-text-parallel-cleansynthetic_trigrams_alpha_1.0pile_llama_trigrams_inject_seq_1.0pile_top_trigrams
top_trigrams
See https://confirmlabs.org/posts/catalog.html for details.
id0: the first token in the trigram
id1: the second token in the trigram
id2: the most common token following (id0, id1) in The Pile
sum_count: the number of times that (id0, id1) appears in The Pile.
max_count: the number of times that id2 appears after (id0, id1) in The Pile.
frac_max: max_count / sum_count
token0: the string representation of id0
token1: the string representation of id1
token2: the string… See the full description on the dataset page: https://huggingface.co/datasets/Confirm-Labs/pile_top_trigrams.trigrams_syn_k1000_shift0.5_conc1.0trigrams_syn_k1000_bigram_1_7_support_0_6pile_llama_synthetic_trigrams_v10_k50pile_llama_synthetic_trigrams_v6_k10trigrams_syn_k1000_support_9.0trigrams_syn_k1000_shift0.8_conc1.0trigrams_obs_k1000trigrams_syn_k1000_support_3.0generate_synthetic_trigrams_v1pile_llama_synthetic_trigrams_v5_k32000pile_llama_synthetic_trigrams_v5_k32000_400mNacrith-Trigramspile_llama_trigrams_inject_seq_0.8_L8synthetic_trigrams_alpha_1_0_v5synthetic_trigrams_alpha_1_0synthetic_trigrams_alpha_1_0_v3generate_synthetic_trigrams_v2
