CoolFace
Datasetpublic

Elormiden/Hellenic-greek-parliamentary-speech

HParl: Hellenic Parliamentary Speech Corpus Dataset Description Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors. Link to the original source: https://inventory.clarin.gr/corpus/1602 HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
1likes115downloads
Dataset Card

HParl: Hellenic Parliamentary Speech Corpus

Dataset Description

Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors.

Link to the original source: https://inventory.clarin.gr/corpus/1602

HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been processed and split for machine learning research with transcriptions included.

Dataset Details

  • —Original Language: Modern Greek (1453-)
  • —Original Duration: 120 hours of speech
  • —Original Source: Hellenic Parliament proceedings
  • —Original Time Coverage: 10.12.2018 - 15.02.2022
  • —Sampling Rate: 16,000 Hz
  • —License: CC BY-NC 4.0 (Non-commercial use)
  • —Total Examples: 92,133 audio clips with transcriptions

Dataset Structure

Data Fields

  • —audio: Audio recordings from parliamentary sessions (16kHz sampling rate)
  • —audio_array: String representation of audio array
  • —transcription: Greek text transcription of the audio

Data Splits

  • —Training set: 73,706 samples (~80%)
  • —Validation set: 9,213 samples (~10%)
  • —Test set: 9,214 samples (~10%)

Total: 92,133 audio samples with transcriptions Dataset size: ~13.6 GB