datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
parliament_hearings_processed
Preprocessed parliament hearings ASR dataset to truecased form.
Original dataset: https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126
dataset_info:
features:
- name: id
dtype: string
- name: audio
dtype:
audio:
sampling_rate: 16000
- name: transcription
sequence: string
splits:
- name: train
num_bytes: 53645064353.18
num_examples: 191455
- name: test
num_bytes: 740331298.0
num_examples: 2726… See the full description on the dataset page: https://huggingface.co/datasets/jkot/parliament_hearings_processed.parliament
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenFormosa/parliament.tw_parliament_split
Parliament
Parliament is an OpenFormosa Traditional Chinese speech dataset derived from the
zh_tw split of disco-eth/WorldSpeech.
It contains audio clips and human transcripts from Taiwan Legislative Yuan IVOD
parliamentary proceedings.
This release keeps only rows that passed the Taiwan-OmniData / FineWeb2-style
text filtering pipeline. Audio is preserved from the upstream dataset and cast
as a Hugging Face Audio(sampling_rate=24000) feature.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/NickWeng/tw_parliament_split.swiss_parliament_corpus
Dataset Card for "swiss_parliament_corpus"
More Information needed
Hellenic-greek-parliamentary-speech
HParl: Hellenic Parliamentary Speech Corpus
Dataset Description
Note: This is a processed version of the original HParl dataset. This dataset is not created or maintained by the original authors.
Link to the original source: https://inventory.clarin.gr/corpus/1602
HParl is a 120-hour speech corpus for Modern Greek, originally collected from parliamentary proceedings of the Hellenic Parliament by the Institute for Language and Speech Processing. This version has been… See the full description on the dataset page: https://huggingface.co/datasets/Elormiden/Hellenic-greek-parliamentary-speech.pseudolabel-malaysia-parliament-youtube-whisper-large-v3
pseudolabel-malaysia-parliament-youtube-whisper-large-v3
Pseudolabel malaysia-ai/malaysia-parliament-youtube using openai/whisper-large-v3
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/pseudolabel-malaysia-parliament-youtube-whisper-large-v3
wget… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/pseudolabel-malaysia-parliament-youtube-whisper-large-v3.basque_parliament_1turkish-parliament-speechVC-Parliamentdv_corpora_parliament_processedid_corpora_parliament_processedafrispeech-parliament
AfriSpeech-Parliament: Transcribed Parliamentary Sessions from Four African Nations
This work is licensed under aCreative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
Overview
AfriSpeech-Parliament is a 35.86-hour dataset of transcribed parliamentary sessions from Nigeria, Ghana, South Africa, and Kenya. Sourced from their respective official YouTube channels, each sample is a 16-second audio chunk with corresponding manual transcription.… See the full description on the dataset page: https://huggingface.co/datasets/intronhealth/afrispeech-parliament.parliament_14263
