datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
composite_corpus_es_v1.0
Composite dataset for Spanish made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/es: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
openslr: a train split made from the SLR(39,61,67,71,72,73,74,75,108) subsets… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_es_v1.0.composite_corpus_eseu_v1.0
Composite bilingual dataset for Spanish and Basque made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/es: a portion of the "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
mozilla-foundation/common_voice_18_0/eu:… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eseu_v1.0.composite_corpus_eu_v2.1
Composite dataset for Basque made from public available data
This dataset is composed of the following public available data:
Train split:
The train split is composed of the following datasets combined:
mozilla-foundation/common_voice_18_0/eu: "validated" split removing "test_cv" and "dev_cv" split's sentences. (validated split contains official train + dev + test splits and more unique data)
gttsehu/basque_parliament_1/eu: "train_clean" split removing some of the… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/composite_corpus_eu_v2.1.
