datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opus_infopankki
Dataset Card for infopankki
Dataset Summary
A parallel corpus of 12 languages, 66 bitexts.
Supported Tasks and Leaderboards
The underlying task is machine translation.
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More… See the full description on the dataset page: https://huggingface.co/datasets/Helsinki-NLP/opus_infopankki.somali-multilingual-infopankki
Somali Multilingual Infopankki
somali-multilingual-infopankki is a parallel corpus containing multilingual translation pairs that involve the Somali (so) language. This dataset has been filtered and extracted from the original Helsinki-NLP/opus_infopankki corpus.
It is designed to support machine translation (NMT), multilingual sentence alignment, and Somali natural language processing (NLP) research.
Dataset Details
Source Dataset: Helsinki-NLP/opus_infopankki… See the full description on the dataset page: https://huggingface.co/datasets/tufaax/somali-multilingual-infopankki.opus_infopankki_ar_en_experimental
