Elormiden/Hellenic-greek-parliamentary-speech
HParl: Hellenic Parliamentary Speech Corpus Dataset Description Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors. Link to the original source: https://inventory.clarin.gr/corpus/1602 HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.
HParl: Hellenic Parliamentary Speech Corpus
Dataset Description
Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors.
Link to the original source: https://inventory.clarin.gr/corpus/1602
HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been processed and split for machine learning research with transcriptions included.
Dataset Details
- Original Language: Modern Greek (1453-)
- Original Duration: 120 hours of speech
- Original Source: Hellenic Parliament proceedings
- Original Time Coverage: 10.12.2018 - 15.02.2022
- Sampling Rate: 16,000 Hz
- License: CC BY-NC 4.0 (Non-commercial use)
- Total Examples: 92,133 audio clips with transcriptions
Dataset Structure
Data Fields
audio: Audio recordings from parliamentary sessions (16kHz sampling rate)audio_array: String representation of audio arraytranscription: Greek text transcription of the audio
Data Splits
- Training set: 73,706 samples (~80%)
- Validation set: 9,213 samples (~10%)
- Test set: 9,214 samples (~10%)
Total: 92,133 audio samples with transcriptions Dataset size: ~13.6 GB
