datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/besimple-ai/voice-code-bench.voice-code-bench
VoiceCodeBench
VoiceCodeBench is a test-only benchmark for evaluating whether automatic
speech recognition (ASR) systems preserve exact structured values in English
workplace speech.
Paper: VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
The benchmark targets cases where a transcript is software input: callback
numbers, email addresses, command-line flags, file paths, URLs, account
identifiers, dates, measurements, and similar values… See the full description on the dataset page: https://huggingface.co/datasets/yunqi1766/voice-code-bench.ljspeech-mimi-codes
LJSpeech — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens for the
LJSpeech corpus — 13,100 English utterances
from a single female speaker reading public-domain audiobook passages (~24 hours).
This dataset contains codes only, not audio. For waveforms, go to the original LJSpeech
release; these codes are designed to be loaded alongside it for training Mimi-based speech
models without paying the ~1 hour of GPU extraction cost.
Schema
One row per utterance:… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/ljspeech-mimi-codes.Saudilang-Code-Switch-Corpus
SCC - Saudilang Code-Switch Corpus
The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "SCC" dataset, which stands for "Saudilang Code-Switch Corpus”.
This dataset contains a transcription of general conversations taken from a YouTube podcast "Thmanyah" that has been transcribed by the National Center for Artificial Intelligence in SDAIA. The data features three episodes covering different domains: investment… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Saudilang-Code-Switch-Corpus.librispeech-mimi-codes
LibriSpeech — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens for the
LibriSpeech corpus — multi-speaker English audiobook readings
from the LibriVox project.
This dataset contains codes only, not audio. For waveforms, use any of the LibriSpeech
mirrors (e.g. openslr/librispeech_asr);
these codes let you skip the ~hours of GPU extraction needed to train Mimi-based speech models.
Schema
One row per utterance:
Column
Type
Notes
id
string… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/librispeech-mimi-codes.mls-mimi-codes
Multilingual LibriSpeech (MLS) — Mimi Codes
Pre-extracted Kyutai Mimi neural-codec tokens
for Multilingual LibriSpeech —
LibriVox audiobooks in 7 non-English languages.
English is intentionally excluded. For English Mimi codes, use:
shangeth/librispeech-mimi-codes — LibriSpeech (~280k rows, 7 splits)
shangeth/libritts-r-mimi-codes — LibriTTS-R (~360k rows, 7 splits, 24 kHz native)
shangeth/vctk-mimi-codes — VCTK (~44k rows, 110 speakers w/ accents)
shangeth/jenny-mimi-codes — Jenny… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/mls-mimi-codes.arabic-english-code-switching-review-annotations
Review Annotations for Arabic-English Code-Switching Speech
This metadata-only dataset publishes review decisions and transcript-correction deltas for MohamedRashad/arabic-english-code-switching. It contains no human audio, no local file paths, no raw review notes, and no copies of unchanged upstream transcripts.
The annotations are pinned to upstream revision 4a3bffc45219c35949470de32b8d4cb328b0ce11 and join by upstream_row_index.
Coverage and outcomes
The… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-review-annotations.expresso-mimi-codes
Expresso — Mimi Codes (k = 32)
Pre-extracted Kyutai Mimi tokens (all 32 codebooks) for both the read and conversational subsets of Expresso. Source audio + transcripts live in shangeth/expresso; this dataset publishes the discrete-token version for training Mimi-based speech models without re-extracting.
⚠️ License: CC-BY-NC-4.0 — non-commercial use only.
Why Expresso for Wren?
Expresso is the most directly relevant dataset for speech disentanglement research — the… See the full description on the dataset page: https://huggingface.co/datasets/shangeth/expresso-mimi-codes.
