CoolFace
Datasetpublic

philgzl/libri

LibriSpeech: An ASR corpus based on public domain audio books This is a mirror of the LibriSpeech ASR corpus. The original files were converted from FLAC to Opus to reduce the size and accelerate streaming. The transcripts are not included. This mirror is thus best suited for audio-to-audio tasks. Sampling rate: 16 kHz Channels: 1 Format: Opus Splits: Train: 460 hours, 132553 utterances, train-clean-100 and train-clean-360 sets. Validation: 7 hours, 2703 utterances, dev-clean… See the full description on the dataset page: https://huggingface.co/datasets/philgzl/libri.

sourceHugging Facecc-by-4.0updated 1y agoView on Hugging Face
0likes367downloads
Dataset Card

LibriSpeech: An ASR corpus based on public domain audio books

This is a mirror of the LibriSpeech ASR corpus. The original files were converted from FLAC to Opus to reduce the size and accelerate streaming. The transcripts are not included. This mirror is thus best suited for audio-to-audio tasks.

Usage

python
import io

import soundfile as sf
from datasets import Features, Value, load_dataset

for item in load_dataset(
    "philgzl/libri",
    split="train",
    streaming=True,
    features=Features({"audio": Value("binary"), "name": Value("string")}),
):
    print(item["name"])
    buffer = io.BytesIO(item["audio"])
    x, fs = sf.read(buffer)
    # do stuff...

Citation

bibtex
@inproceedings{panayotov2015librispeech,
  title = {{LibriSpeech}: {An} {ASR} corpus based on public domain audio books},
  author = {Panayotov, Vassil and Chen, Guoguo and Povey, Daniel and Khudanpur, Sanjeev},
  booktitle = {Proc. ICASSP},
  pages = {5206--5210},
  year = {2015},
}