multilingual datasets
multilingual-e5-base-matryoshka2d-cached-mnr-multiple-datasets-base-1multilingual-e5-base-matryoshka2d-mnr-negatives-multiple-datasets-large-3multi-lingual-sttmultilingual-e5-base-matryoshka2d-cached-mnr-multiple-datasets-large-2-2multilingual-e5-base-matryoshka2d-cached-mnr-multiple-datasets-large-2-3multilingual-e5-base-matryoshka2d-cached-mnr-multiple-datasets-large-2-3-categorymultilingual-e5-base-matryoshka2d-cached-mnr-multiple-datasets-large-2-4-pekachmultilingual-e5-base-matryoshka2d-mnr-multiple-datasets-large-4
open-asr-leaderboard-multilingual-datasets
ASR Leaderboard Datasets
This repository contains test splits from multiple speech corpora, including FLEURS, Common Voice (MCV), and Multilingual LibriSpeech (MLS).
How to Load
To load a specific subset, use load_dataset with the corresponding config_name in the format <set>_<lang>.
from datasets import load_dataset
# Load the FLEURS dataset for Bulgarian
fleurs_bg = load_dataset("nithinraok/asr-leaderboard-datasets", "fleurs_bg")
print(fleurs_bg)
# Load the… See the full description on the dataset page: https://huggingface.co/datasets/hf-audio/open-asr-leaderboard-multilingual-datasets.multilingual_librispeechMultilingual LibriSpeech (MLS) dataset is a large multilingual corpus suitable for speech research. The dataset is derived from read audiobooks from LibriVox and consists of 8 languages - English, German, Dutch, Spanish, French, Italian, Portuguese, Polish.Multilingual_QASem_Datasets
🧠 Multilingual QASem Dataset
A multilingual dataset for QA-based Semantic Parsing (QASem) covering Hebrew, Russian, and French.It includes automatically projected training data and manually validated gold data (dev + test) for evaluating cross-lingual QASem parsers.
It is the dataset from the paper: Effective QA-Driven Annotation of Predicate–Argument Relations Across Languages (Davidov et al., EACL 2026)
📘 Overview
This dataset provides QA-based Semantic… See the full description on the dataset page: https://huggingface.co/datasets/biu-nlp/Multilingual_QASem_Datasets.multilingual-dataset-sampledopen-asr-leaderboard-multilingual-datasets
Open ASR Leaderboard Armenian Test Datasets
This private repository holds leaderboard-compatible Armenian test
configurations while their integration is being validated.
Configurations
fleurs_hy
Source: google/fleurs,
configuration hy_am, test split
Reviewed reference changes:
Metric-AI/fleurs-corrections,
test split
932 recordings; all 314 reviewed corrections were matched to the original
source transcript and applied
mcv_hy… See the full description on the dataset page: https://huggingface.co/datasets/Metric-AI/open-asr-leaderboard-multilingual-datasets.multilingual-reward-bench_code-pythonCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données multilingual-reward-bench/code-python.
