CoolFace
Datasetpublic

espnet/yodas-granary

Dataset Card for YODAS-Granary Repository: NeMo-speech-data-processor: Granary Paper: Granary: Speech Recognition and Translation Dataset in 25 European Languages Shared by: ESPnet Dataset Description YODAS-Granary is a curated subset of the larger nvidia/Granary dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages. Overview… See the full description on the dataset page: https://huggingface.co/datasets/espnet/yodas-granary.

sourceHugging Facecc-by-3.0updated 1y agoView on Hugging Face
33likes75kdownloads
Dataset Card

Table of Contents

Dataset Card for YODAS-Granary

Dataset Description

YODAS-Granary is a curated subset of the larger `nvidia/Granary` dataset, focusing on high-quality pseudo-labeled speech data for Automatic Speech Recognition (ASR) and Automatic Speech Translation (AST) across 23 European languages.

Overview

<table style="width:100%; table-layout:auto;"> <tr> <td style="vertical-align:middle; text-align:center;"> <img src="./scheme.png" style="max-width:100%;"> </td> <td style="vertical-align:middle; padding-left:20px;"> <p> Derived from the <a href=https://huggingface.co/datasets/espnet/yodas2) target="blank"><code>espnet/yodas2</code></a> corpus, YODAS-Granary provides high-quality pseudo-labeled speech data, focusing on two core tasks: </p> <ul> <li><strong>Automatic Speech Recognition (ASR)</strong>: covers 23 European languages, with pseudo-labeled transcriptions generated using the <a href="https://huggingface.co/Systran/faster-whisper-large-v3" target="blank"><code>Systran/faster-whisper-large-v3</code></a> model, post-processed to restore punctuation and capitalization using <a href="https://huggingface.co/Qwen/Qwen2.5-7B-Instruct" target="blank"><code>Qwen/Qwen2.5-7B-Instruct</code></a>, and filtered for quality.</li> <li><strong>Automatic Speech Translation (AST)</strong>: covers 22 non-English languages and consists of high-quality translations into English, generated from <code>ASR</code> subset using the <a href="https://huggingface.co/utter-project/EuroLLM-9B-Instruct" target="blank"><code>utter-project/EuroLLM-9B-Instruct</code></a> model and filtered for quality.</li> </ul> </td> </tr> </table>

Data Distribution

The following chart illustrates the distribution of data in the YODAS-Granary dataset across 23 European languages, measured in number of words (left) and total hours of audio (right), for both ASR and AST tasks.

<p align="center"> <img src="./chart.png" width="100%"/><br> <em>🔍 AST data is always a filtered subset of ASR, which is why AST bars are never taller than their ASR counterparts.<br>🗣️ The English subset contains ASR data only.</em> </p>

The table below summarizes the storage footprint and sample counts per language in the YODAS-Granary dataset, broken down into corresponding splits:

  • ast — the size and number of translated samples (X → English),
  • asr_only — samples that exist only in the ASR subset and have no corresponding translation,

and combined size and number of all samples per language (total).

LanguageSubsetsSamples [`ast`]Size [`ast`]Samples [`asr_only`]Size [`asr_only`]Total samplesTotal size
Bulgarianbg00088441.9 GB1533208.8 MB103772.1 GB
Czechcs000343607.5 GB4185442.1 MB385458.0 GB
Danishda00095821.9 GB65665.3 MB102382.0 GB
Germande000, de{100..102}3335156845.6 GB41526056.3 GB3750416901.9 GB
Greekel00042421.6 GB514113.5 MB47561.7 GB
Englishen00{0..7}, en{100..129}4081051711.3 TB4081051711.3 TB
Spanishes000, es{100..108}79236462.9 TB95145088.8 GB88750963.0 TB
Estonianet0004437901.8 MB51357.8 MB4950959.6 MB
Finnishfi0006072917.4 GB4637419.3 MB6536617.8 GB
Frenchfr000, fr{100..103}47662391.3 TB55884881.1 GB53250871.4 TB
Croatianhr00053691.1 GB26127.9 MB56301.1 GB
Hungarianhu0004826316.4 GB6530962.7 MB5479317.4 GB
Italianit000, it{100..101}1226663683.7 GB8658713.6 GB1313250697.3 GB
Lithuanianlt0002177564.5 MB39071.1 MB2567635.6 MB
Latvianlv00027266.5 MB7512.0 MB34778.5 MB
Dutchnl000, nl100865754151.1 GB714904.3 GB937244155.4 GB
Polishpl00026425775.5 GB286782.4 GB29293577.9 GB
Portuguesept000, pt{100..103}58987641.5 TB72913830.1 GB66279021.5 TB
Romanianro000122763.7 GB2303663.4 MB145794.4 GB
Russianru00{0..1}, ru{100..106}79910382.1 TB1876876197.0 GB98679142.3 TB
Slovaksk0003405992.0 MB28751.9 MB36921.0 GB
Swedishsv0005408510.2 GB3192215.7 MB5727710.4 GB
Ukrainianuk000, uk10024637368.3 GB94791.2 GB25585269.4 GB

How to use

Standard Loading

You can load the dataset using the datasets library from Hugging Face:

python
from datasets import load_dataset

🔹 Load the entire dataset:

python
ds = load_dataset("espnet/yodas-granary")

🔹 Load a single language (e.g., Italian):

python
ds = load_dataset("espnet/yodas-granary", "Italian")

Streaming

Some language subsets are quite large and may not fit comfortably in memory. For efficient access and analysis without downloading the entire dataset, we recommend using streaming mode:

python
ds = load_dataset("espnet/yodas-granary", "English", streaming=True)

Dataset Structure

Data Instance

Each utterance in the dataset includes the following fields: utt_id, audio, duration, lang, task, text, translation_en (null in asr_only), original_audio_id, and original_audio_offset.

Typical entry

from data/de101/translation/00000000.parquet

python
{
  "utt_id": "de101_00000000_Z0_gcPJVTqg_1004_62_1_74",
  "audio": {
    'path': 'de101_00000000_Z0_gcPJVTqg_1004_62_1_74.wav', 
    'bytes': ... 
  }
  "duration": 1.74,
  "lang": "<de>",
  "task": "<ast>",
  "text": "Ich muss mir das Zeug mal aus der Nähe ansehen.",
  "translation_en": "I have to take a closer look at this stuff.",
  "original_audio_id": "Z0_gcPJVTqg",
  "original_audio_offset": 1004.62
}

Data Fields

**Field****Type****Description**
utt_id¹stringUnique identifier of the utterance, referencing the original segment.
audioAudio (16 kHz)Audio data of the utterance, stored as PCM waveform.
durationfloat64Duration of the utterance in seconds.
langstringLanguage of the utterance, in ISO 639-1 code (e.g., <de> for German).
taskstringTask type: either <asr> for transcription or <ast> for translation to English.
textstringTranscription of the utterance in its original language.
translation_enstringEnglish translation of the utterance. null if split is asr_only.
original_audio_idstringID of the original audio file. This value corresponds to the audio_id field from the `espnet/yodas2` dataset.
original_audio_offsetfloat64Start time (in seconds) of the utterance within the original audio file.

¹ - utt_id is encoded as <subset>_<shard>_<wav_id>_<start_time_s>_<start_time_decimals>_<duration_s>_<duration_decimals>, where subset, shard², and wav_id match the utterance's location in the original `espnet/yodas2` archive.

² - shard indices reflect those in `espnet/yodas2`, but some shards are missing due to filtering during data processing. In particular, bg000 is missing shard 00000011, en000 is missing shard 00000308, en003 is missing shard 00000221, en118 is missing shard 00000240.

Data Splits

The dataset is organized into language-specific subsets, each containing one or two splits, depending on the language:

  • For non-English languages:
  • <ast> – samples with both high-quality transcriptions and translations into English.
  • <asr_only> – samples that passed transcription quality checks but do not include translations.
  • For `English`:
  • <asr_only> split is available only, since English-to-English translation is not applicable.

Directory example

python
yodas_granary
└── data
    ├── da000   # subset
    │   ├── asr_only # corresponds to `asr_only` split
    │   │   ├── 00000000.parquet # shard
    │   │   ├── 00000001.parquet
    │   │   ├── 00000002.parquet
    │   │   └── ...
    │   └── ast # corresponds to `ast` split
    │       ├── 00000000.parquet # shard
    │       ├── 00000001.parquet
    │       ├── 00000002.parquet
    │       └── ...
    ├── cs000
    ├── bg000
    └── ...

Reference

bibtex
@misc{koluguri2025granaryspeechrecognitiontranslation,
      title={Granary: Speech Recognition and Translation Dataset in 25 European Languages}, 
      author={Nithin Rao Koluguri and Monica Sekoyan and George Zelenfroynd and Sasha Meister and Shuoyang Ding and Sofia Kostandian and He Huang and Nikolay Karpov and Jagadeesh Balam and Vitaly Lavrukhin and Yifan Peng and Sara Papi and Marco Gaido and Alessio Brutti and Boris Ginsburg},
      year={2025},
      eprint={2505.13404},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.13404}, 
}
espnet/yodas-granary · CoolFace