datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
asr_codeswitched_dataset
Arabic/English Code-Switched ASR Dataset
Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to
fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical
vocabulary alternate within sentences.
Composition
Source
Description
EJUST custom recordings
Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs)
MohamedRashad/arabic-english-code-switching
~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.ne-en-codeswitching-asr-technical-interview
Dataset Summary
This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar.
It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.khmer-english-codeswitch-tts
Khmer–English Code-Switch Synthetic Speech
8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech,
generated with VoxCPM2 from code-switch text
manufactured by confirmed lexical substitution over a Khmer–English parallel corpus.
⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a
TTS model, and the code-switch sentences were manufactured by word substitution — they are not
transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.codeswitch-fr-en-kyutai-stt
Code-Switched French–English STT Probe Dataset
Dataset Summary
This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.codeswitching
WTFO Code-Switching Speech
Code-switching speech dataset prepared for automatic speech recognition training.
Dataset fields
audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature
text: transcript
duration: audio duration in seconds (float64)
Split summary
Split: train
Examples: 98,662
Total duration: 567422.698 seconds (157.62 hours)
The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.
