datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
wikipedia_basque_ipa
Basque Wikipedia Phonemized Corpus (Text + IPA phonemes)
Dataset Description
A large-scale paired corpus derived from the Basque Wikipedia dump. Each row contains both the original plain text and its IPA phoneme transcription, at paragraph level. Stressed vowels use the apostrophe convention (e.g. 'a, 'e, 'i, 'o, 'u) and affricates are kept as multicharacter sequences (e.g. tʃ, tʂ, ts).
This dataset is intended for training text-to-speech (TTS) and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/wikipedia_basque_ipa.basque_speech_dataset
Basque Speech Dataset
This dataset contains Crowdsourced high-quality Basque speech data.
It includes audio samples and their corresponding transcriptions.
Splits:
emakumezkoa: Female speakers
gizonezkoa: Male speakers
Source:
The data is sourced from OpenSLR - Dataset 76.
License:
(Please refer to the LICENSE file in the dataset repository for detailed license information)
alpaca-basque-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-basque-cleaned.alpaca_basque_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_basque_taco.art-lang-uniform-basque-beforeart-lang-zipfian-basque-beforeaya-global-exams-basqueSpanish exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
art-lang-uniform-basque-afterart-lang-zipfian-basque-afterbasqueparl_text
Dataset Card for "basqueparl_text"
More Information needed
goldfish-Dp-basque-100mbgoldfish-Dp-basque-10mb-tokenizedbasque_nt_eu
Basque NT (Navarro Labourdin)
Description
The Basque New Testament in the Navarro-Labourdin dialect is a translation of the New Testament into the Basque language (Euskara), one of the oldest living languages in Europe and a linguistic isolate. This translation represents the classical literary Basque of the Northern (French) Basque Country. Basque is spoken by approximately 750,000 people in the Basque Country straddling Spain and France. This translation is a… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/basque_nt_eu.al-basque-afteral-basque-beforegoldfish-Dp-basque-100mb-tokenizedartificial-lang-newlexicon-basque-afterbasque_parliament_1artificial-lang-basque-beforeartificial-lang-newlexicon-basque-beforeartificial-lang-basque-aftergoldfish-Dp-basque-10mb
