oromo
Datasets
All datasets matching “oromo”afaan-oromoo-speech
Dataset.ET Afaan Oromoo Speech — v0.1.0
9.843 hours · 3,594 clips · 74 speakers · 3,283 distinct prompts
Dataset Summary
Read speech in Afaan Oromoo, crowdsourced from volunteer contributors in Ethiopia
through a Telegram bot, peer-validated by other contributors, and screened
acoustically before release. Afaan Oromoo has very little open speech data; this
corpus exists to change that.
Contributors read a displayed prompt aloud, other contributors listen and vote… See the full description on the dataset page: https://huggingface.co/datasets/snapwre/afaan-oromoo-speech.afaan_oromo-speechafaan-oromo-tts
Afaan Oromo TTS
Configs
default
Splits
train
Columns
id
speaker_id
text
transcription
language
gender
audio
amharic-oromo_sentence-pairs
Amharic-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Oromo_Sentence-Pairs
Number of Rows: 109805
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-oromo_sentence-pairs.oromo_name
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Soressaa/oromo_name.english-oromo_sentence-pairs_mt560
English-Oromo Parallel Dataset
This dataset contains parallel sentences in English and Oromo (orm).
Dataset Information
Language Pair: English ↔ Oromo
Language Code: orm
Country: orm
Original Source: OPUS MT560 Dataset
Dataset Structure
The dataset contains parallel sentences that can be used for:
Machine translation training
Cross-lingual NLP tasks
Language model fine-tuning
Citation
If you use this dataset, please cite the citation guide of the… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/english-oromo_sentence-pairs_mt560.
