Basque
Datasets
All datasets matching “Basque”basque_parliament_1
Dataset Card for Basque Parliament Speech Corpus 1.0
This work was partially funded by the Spanish Ministry of Science and Innovation (OPENSPEECH
project, PID2019-106424RB-I00).
Dataset Summary
The Basque Parliament Speech Corpus 1.0 consists of 1462 hours of speech extracted from
Basque Parliament plenary sessions from 2013 to 2022. Encoded as MP3 files, the dataset
contains 759192 transcribed segments either spoken in Basque, Spanish or both (in
Basque and Spanish).… See the full description on the dataset page: https://huggingface.co/datasets/gttsehu/basque_parliament_1.basqueGLUEWe present BasqueGLUE, the first NLU benchmark for Basque, which has been elaborated from
previously existing datasets and following similar criteria to those used for the construction of
GLUE and SuperGLUE. BasqueGLUE is freely available under an open license.basque-parallel-corpus
Sources of the corpus used
Opus
Orai
basque_dialect_machine_translationbasqueparl
BasqueParl: A Bilingual Corpus of Basque Parliamentary Transcriptions
This repository contains BasqueParl, a bilingual corpus for political discourse analysis. It covers transcriptions from the Parliament of
the Basque Autonomous Community for eight years and two legislative terms (2012-2020), and its main characteristic is the presence of Basque-Spanish
code-switching speeches.
📖 Paper: BasqueParl A Bilingual Corpus of Basque Parliamentary Transcriptions In LREC 2022.… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/basqueparl.wikipedia_basque_ipa
Basque Wikipedia Phonemized Corpus (Text + IPA phonemes)
Dataset Description
A large-scale paired corpus derived from the Basque Wikipedia dump. Each row contains both the original plain text and its IPA phoneme transcription, at paragraph level. Stressed vowels use the apostrophe convention (e.g. 'a, 'e, 'i, 'o, 'u) and affricates are kept as multicharacter sequences (e.g. tʃ, tʂ, ts).
This dataset is intended for training text-to-speech (TTS) and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/wikipedia_basque_ipa.
