datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
metric-mamba-ml2021-hungyi-corpus
Dataset Card for "metric-mamba-ml2021-hungyi-corpus"
More Information needed
iwslt2026-metrics-shared-train-dev
Metrics-Shared Task IWSLT 2026 Train and Dev Set
This dataset contains the train and dev sets for the Speech Translation Metrics Shared Task at IWSLT 2026. More details about the shared task can be found on the IWSLT website.
The dataset is primarily designed for research in speech translation quality estimation.
Task Goal
Given a speech sample and a system-generated translation, the goal is to estimate a score that reflects the translation quality.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/iwslt2026-metrics-shared-train-dev.scottish-metricsopen-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Armenian Test Datasets
This private repository holds leaderboard-compatible Armenian test
configurations while their integration is being validated.
Configurations
fleurs_hy
Source: google/fleurs,
configuration hy_am, test split
Reviewed reference changes:
Metric-AI/fleurs-corrections,
test split
932 recordings; all 314 reviewed corrections were matched to the original
source transcript and applied
mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.iwslt2026-metrics-shared-testtl-whisper-rawThis repository is used to create and push to metricv/tl-whisper.
Usage
Requirements
apt-get install ffmpeg # Do this however you want. GPU/nvenc not required, nor useful since we're only dealing with audio.
pip install -r requirements.pyr
Append to dataset
python append_to_dataset.py AUDIO_FILE.m4a --entxt SUB.en.txt
# OR
python append_to_dataset.py AUDIO_FILE.m4a --ass SUB.ass
Build and push
python ./push_self.py
chaldea-whisper-fttl-whisperThis dataset is compiled from https://huggingface.co/datasets/metricv/tl-whisper-raw
It contains sentence-by-sentence transcription of videos from Youtube channel TechLinked.
tl-ztt-audiotest_metricscrf-metrics-board
