datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librispeech_asr_dummylibrispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/openslr/librispeech_asr.multilingual_librispeech
Dataset Card for MultiLingual LibriSpeech
Dataset Summary
This is a streamable version of the Multilingual LibriSpeech (MLS) dataset.
The data archives were restructured from the original ones from OpenSLR to make it easier to stream.
MLS dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of
8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.… See the full description on the dataset page: https://huggingface.co/datasets/facebook/multilingual_librispeech.librispeech_long
Dataset Card for "librispeech_long"
More Information needed
test_librispeech_parquetlibrispeech_asr-noise
Dataset Card for "librispeech_asr-noise"
More Information needed
librispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for… See the full description on the dataset page: https://huggingface.co/datasets/zihan-audio/librispeech_asr.librispeech_asr_demolibrispeech-alignments
Dataset Card for Librispeech Alignments
Librispeech with alignments generated by the Montreal Forced Aligner. The original alignments in TextGrid format can be found here
Dataset Details
Dataset Description
Librispeech is a corpus of read English speech, designed for training and evaluating automatic speech recognition (ASR) systems. The dataset contains 1000 hours of 16kHz read English speech derived from audiobooks.
The Montreal Forced Aligner (MFA) was used… See the full description on the dataset page: https://huggingface.co/datasets/gilkeyio/librispeech-alignments.librispeechlibrispeech_asr_arrowlibrispeech_asrlibrispeech-data
Dataset Card for "librispeech-data"
More Information needed
librispeech_synth
Dataset Card for "librispeech_synth"
More Information needed
russian_librispeech
Russian LibriSpeech (RuLS)
Identifier: SLR96 from openslr.org
Summary: This dataset is based on LibriVox audiobooks
Category: Speech
License: The dataset is Public Domain in the USA.
About this resource:
Russian LibriSpeech (RuLS) dataset is based on LibriVox's public domain audio books (see BOOKS.TXT for the list of included books) and contains about 98 hours of audio data.
librispeech_asr-prompted
Dataset Card for "librispeech_asr-prompted"
More Information needed
librispeechhubert_layer9-librispeech-asr100h
Dataset Card for "hubert_layer9-librispeech-asr100h"
More Information needed
multilingual-librispeech-german-labeledlibrispeech-arpabet-processed
LibriSpeech ARPAbet Processed Dataset
Pre-processed dataset for training ARPAbet phoneme recognition models using CTC loss.
Dataset Description
This dataset is derived from LibriSpeech (train-clean-100 split) with the following preprocessing:
Audio: Resampled to 16kHz, normalized using Wav2Vec2 feature extractor
Labels: Text transcriptions converted to ARPAbet phoneme sequences using CMU Pronouncing Dictionary
Filtering: Samples with out-of-vocabulary words (not in CMU… See the full description on the dataset page: https://huggingface.co/datasets/davidggphy/librispeech-arpabet-processed.librispeech
This dataset only contains test data, which is integrated into UltraEval-Audio(https://github.com/OpenBMB/UltraEval-Audio) framework.
python audio_evals/main.py --dataset librispeech-test-clean --model gpt4o_audio
python audio_evals/main.py --dataset librispeech-dev-clean --model gpt4o_audio
python audio_evals/main.py --dataset librispeech-test-other --model gpt4o_audio
python audio_evals/main.py --dataset librispeech-dev-other --model gpt4o_audio
🚀超凡体验,尽在UltraEval-Audio🚀… See the full description on the dataset page: https://huggingface.co/datasets/TwinkStart/librispeech.librispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/0x3/librispeech_asr.librispeech_asr_test_48k_synthlibrispeech-sr16000librispeech_asr
Dataset Card for librispeech_asr
Dataset Summary
LibriSpeech is a corpus of approximately 1000 hours of 16kHz read English speech, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.
Supported Tasks and Leaderboards
automatic-speech-recognition, audio-speaker-identification: The dataset can be used to train a model for Automatic… See the full description on the dataset page: https://huggingface.co/datasets/sovitrath/librispeech_asr.librispeech-augmentated-train-prepared
Dataset Card for "librispeech-augmentated-train-prepared"
More Information needed
librispeech_parquetwavlm-large_layer21-librispeech-asr100h
Dataset Card for "wavlm-large_layer21-librispeech-asr100h"
More Information needed
librispeech_asr_test_synthlibrispeech-pc-44khz-opus
LibriSpeech-PC 44kHz Opus
Summary
This dataset is a high-quality audio replacement variant of Librispeech PC. It preserves the row identity and text fields while replacing audio content from the source audio with the highest available quality (usually mp3 128kpbs) which is then encoded as Opus (64 kbps). Sampling rate is increased from 16khz up to 48khz depending the on source audio.
LibriSpeech-PC is a merge of openslr/librispeech_asr audio metadata with SLR145… See the full description on the dataset page: https://huggingface.co/datasets/mythicinfinity/librispeech-pc-44khz-opus.
