CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes268 downloads1y agoHugging Face02georgechang8 /code_switch_yodas_zh Dataset Card for code-switching yodas This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon. Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.audio10K<n<100K4 likes197 downloads2y agoHugging Face03liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes194 downloads2mo agoHugging Face04code-switching /text-summarizationtextsummarizationn<1K0 likes167 downloads25d agoHugging Face05ServiceNow-AI /asr_codeswitchedaudio1K<n<10K6 likes150 downloads2mo agoHugging Face06code-switching /question-answertext1K<n<10K0 likes138 downloads25d agoHugging Face07Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes132 downloads4mo agoHugging Face08code-switching /naturalnesstabularn<1K0 likes116 downloads25d agoHugging Face09FatimahEmadEldin /cafe-algerian-codeswitch-speech CAFE Algerian Codeswitch Speech This dataset contains Algerian Arabic and French code-switched speech. Repository Path: FatimahEmadEldin/cafe-algerian-codeswitch-speech audio1K<n<10K0 likes102 downloads4mo agoHugging Face10devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes100 downloads7mo agoHugging Face11Tim2190 /kazakh-codeswitch-asr Kazakh Code-Switching ASR Benchmark A benchmark for evaluating ASR systems on natural Kazakh speech that code-switches with Russian — the everyday Kazakh–Russian mixing found in stand-up, interviews and vlogs, not scripted read speech. This is, to our knowledge, the first speech/ASR resource targeting the Kazakh–Russian code-switching pair (existing Kazakh–Russian NLP resources are text-only). Code, scoring harness and full analysis:… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kazakh-codeswitch-asr.audioautomatic-speech-recognitionn<1K0 likes93 downloads2mo agoHugging Face12shulhaaja /id-en-codeswitch-dataset-alternative Indonesian–English Code-Switching Synthetic Speech Dataset Synthetic speech generated for the undergraduate final project "Handling Code-Switching in Automatic Speech Recognition for Low-Resource Language Pairs: An Indonesian–English Case Study", School of Electrical Engineering and Informatics, Institut Teknologi Bandung. This dataset contains synthetic audio produced from the Indonesian–English code-switching text corpora released in the companion repository below. It was used… See the full description on the dataset page: https://huggingface.co/datasets/shulhaaja/id-en-codeswitch-dataset-alternative.audioautomatic-speech-recognition10K<n<100K0 likes87 downloads1mo agoHugging Face13prokelly /neuromoyo-sahara-codeswitch-benchmark NEUROMOYO — Sahara CodeSwitch Africa Benchmark 🔗 Live Benchmark Results Interactive benchmark: https://www.neuromoyo.app/benchmark This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech. 🚀 Live NEUROMOYO Demo Live application: https://www.neuromoyo.app The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.tabularn<1K0 likes87 downloads5d agoHugging Face14Kimyayd /vocal-money-codeswitch-asr-benchmark Vocal Money — Yoruba–English Code-Switched ASR Benchmark A 210-clip evaluation subset used to benchmark five speech recognition systems on naturally code-switched Yoruba–English speech, together with the reference transcriptions and the output of every system on every clip, so that the published results can be recomputed or contradicted. Produced for the MLC (Africa) × Intron Agentic Voice AI Challenge, Deep Learning Indaba 2026. Team Vocal Money — Hospice Hounfodji, Mohamed… See the full description on the dataset page: https://huggingface.co/datasets/Kimyayd/vocal-money-codeswitch-asr-benchmark.audioautomatic-speech-recognitionn<1K0 likes86 downloads2mo agoHugging Face15Panhapich /khmer-english-codeswitch-tts Khmer–English Code-Switch Synthetic Speech 8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech, generated with VoxCPM2 from code-switch text manufactured by confirmed lexical substitution over a Khmer–English parallel corpus. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and the code-switch sentences were manufactured by word substitution — they are not transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.audioautomatic-speech-recognition1K<n<10K2 likes85 downloads2mo agoHugging Face16code-switching /topic-classificationtabulartext-classificationn<1K0 likes63 downloads18d agoHugging Face17Praxel /codeswitch-pairs-lase Codeswitch Pairs LASE — training corpus 1118 same-voice cross-script utterance pairs (8 ElevenLabs Multilingual voices × en/hi/te/ta) used to train the LASE r1 speaker encoder. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair). Schema (manifest.jsonl) { "voice_id": "21m00Tcm4TlvDq8ikWAM", "lang": "en | hi | te | ta", "text": "the prompt text"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase.audioaudio-classificationn<1K0 likes60 downloads5mo agoHugging Face18Atufa /codeswitch-fr-en-kyutai-stt Code-Switched French–English STT Probe Dataset Dataset Summary This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.audioautomatic-speech-recognitionn<1K0 likes54 downloads7mo agoHugging Face19vikkyblacq /kare-codeswitch-samples Kare — Code-Switching Illustrative Samples Eight short audio clips demonstrating the English/Yoruba/Hausa/Igbo/Pidgin code-switching that Kare, a voice-first AI health assistant for Nigeria, is built to understand — submitted as part of Kare's entry to the Sahara CodeSwitch Africa Challenge. What this is — and isn't Is: eight original sentences, written by the Kare team specifically for this demonstration, mixing everyday Yoruba / Hausa / Igbo / Nigerian Pidgin… See the full description on the dataset page: https://huggingface.co/datasets/vikkyblacq/kare-codeswitch-samples.audioautomatic-speech-recognitionn<1K0 likes53 downloads11d agoHugging Face201uckyan /code-switch_chunks Dataset Summary This dataset is a curated compilation of SECoMiCSC, DevCECoMiCSC, and BAAI/CS-Dialogue, specifically processed for Code-Switching ASR research. root/ ├── audio/ │ ├── SECoMiCSC/ # Chunked segments from SECoMiCSC │ ├── DevCECoMiCSC/ # Chunked segments from DevCECoMiCSC │ └── CS_Dialogue/ # Extracted <MIX> segments from BAAI/CS-Dialogue ├── metadata.jsonl # Universal index containing paths, transcripts, and metadata └──… See the full description on the dataset page: https://huggingface.co/datasets/1uckyan/code-switch_chunks.audioautomatic-speech-recognition10K<n<100K0 likes52 downloads8mo agoHugging Face21samihyounes /Moroccan-Codeswitching Moroccan Darija Code-Switched Corpus (Sentence-level TSV) Dataset Summary This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text. Languages The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.texttext-classification100K<n<1M0 likes50 downloads7mo agoHugging Face22Aynursusuz /jaen-codeswitch-tts jaen-codeswitch-tts Japanese/English code-switch synthetic speech: 10k varied-length conversational utterances, single voice (Qwen3-TTS clone of one English-male reference, spk_male). Per-utterance language routed to the dominant script. Columns: audio, text, speaker_id, language. audio10K<n<100K0 likes46 downloads3mo agoHugging Face23lxyuan /nemo-codeswitch-reasoning-debate Overview This is a synthetic, multilingual code-switching dataset. Each record contains: a realistic user query a long-form reasoning section a debate / counterargument section a concise final_answer It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses. This snapshot contains 574,977 rows and 10 string columns. Data provenance Important: Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.texttext-generation100K<n<1M0 likes45 downloads7mo agoHugging Face24oumayma03 /adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-moroccan_darija_prompts & trilingual_codeswitch_chat (augmented) This dataset consists of short conversational prompts written in Moroccan Darija, covering topics like shopping, social interactions, and daily inquiries. Each entry contains a single prompt with a null completion, indicating it is likely intended for instruction tuning or completion generation tasks. The content… See the full description on the dataset page: https://huggingface.co/datasets/oumayma03/adaption-moroccan-darija-prompts-trilingual-codeswitch-chat-augmented.text1K<n<10K0 likes45 downloads9d agoHugging Face25Moamen-dcp /arazn_codeSwitched_mp3_full_notLoweraudio1K<n<10K0 likes38 downloads1y agoHugging Face26Hamza-Ali01 /code-switching-codesaviours-si26-hamzatext1K<n<10K1 likes36 downloads1mo agoHugging Face27sanaisrail /code-switching-codesaviours-si26-Sanatext1K<n<10K0 likes32 downloads29d agoHugging Face28DynamicSuperb /Code-switchSpeechRecognition_NTUML2021 Dataset Card for "Code-switchSpeechRecognition_NTUML2021" More Information needed audio1K<n<10K0 likes31 downloads2y agoHugging Face29Nash-pAnDiTa /ASR_En_Ar_CodeSwitchingaudio10K<n<100K0 likes29 downloads2y agoHugging Face30UBC-NLP /NADI2026_subtask1.3_codeswitched_asraudio1K<n<10K0 likes26 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.