espnet
Datasets
All datasets matching “espnet”yodas-granary
Dataset Card for YODAS-Granary
Repository: NeMo-speech-data-processor: Granary
Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages
Shared by: ESPnet
Dataset Description
YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.yodasUpdates
2024/07/09: we also uploaded a new version of YODAS as YODAS2, it provides unsegmented audios and higher sampling rate (24k)
README
This is the YODAS manual/automatic subset from our YODAS dataset, it has 369,510 hours of speech.
This dataset contains audio utterances and corresponding captions (manual or automatic) from YouTube. Note that manual caption only indicates that it is uploaded by users, but not necessarily transcribed by a human
For more details about YODAS… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas.yodas2YODAS2 is the long-form dataset from YODAS dataset.
It provides the same dataset as espnet/yodas but YODAS2 has the following new features:
formatted in the long-form (video-level) where audios are not segmented.
audios are encoded using higher sampling rates (i.e. 24k)
For detailed information about YODAS dataset, please refer to our paper and the espnet/yodas repo.
Usage:
Each data point corresponds to an entire video on YouTube, it contains the following fields:
video_id:… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas2.Bagpiper_SFT_Data
Bagpiper SFT Data
Release status: the validated Parquet release is being uploaded. The
homepage and metadata may appear before every large shard is committed.
Bagpiper SFT Data is the supervised fine-tuning corpus for
Bagpiper, an open-ended audio language model
that understands and generates speech, music, environmental sound, and their
mixtures through rich textual captions and planning.
The public release has exactly two configurations:
Configuration
Direction… See the full description on the dataset page: https://huggingface.co/datasets/espnet/Bagpiper_SFT_Data.floras
FLORAS
FLORAS is a 50-language benchmark For LOng-form Recognition And Summarization of spoken language.
The goal of FLORAS is to create a more realistic benchmarking environment for speech recognition, translation, and summarization models.
Unlike typical academic benchmarks like LibriSpeech and FLEURS that uses pre-segmented single-speaker read-speech, FLORAS tests the capabilities of models on raw long-form conversational audio, which can have one or many speakers.
To… See the full description on the dataset page: https://huggingface.co/datasets/espnet/floras.yodas_owsmv4🏆 News: Our OWSM v4 paper won the Best Student Paper Award at INTERSPEECH 2025!
Dataset Card for YODAS_OWSMv4
Paper: OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning (Best Student Paper at INTERSPEECH 2025)
Authors: Yifan Peng, Muhammad Shakeel, Yui Sudo, William Chen, Jinchuan Tian, Chyi-Jiunn Lin, Shinji Watanabe
Data Cleaning Scripts: ESPnet
Model Demo: Gradio
Dataset Description
Open Whisper-style Speech Model (OWSM)is the first… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas_owsmv4.
