CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes268 downloads1y agoHugging Face02Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_new10K<n<100K0 likes230 downloads1y agoHugging Face03georgechang8 /code_switch_yodas_zh Dataset Card for code-switching yodas This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon. Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.audio10K<n<100K4 likes197 downloads2y agoHugging Face04code-switching /text-summarizationtextsummarizationn<1K0 likes167 downloads25d agoHugging Face05ServiceNow-AI /asr_codeswitchedaudio1K<n<10K6 likes150 downloads2mo agoHugging Face06Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo10K<n<100K0 likes144 downloads1y agoHugging Face07Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes132 downloads4mo agoHugging Face08devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes100 downloads7mo agoHugging Face09Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_prepared_4_whisper_turbo_transcription1K<n<10K0 likes97 downloads1y agoHugging Face10Panhapich /khmer-english-codeswitch-tts Khmer–English Code-Switch Synthetic Speech 8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech, generated with VoxCPM2 from code-switch text manufactured by confirmed lexical substitution over a Khmer–English parallel corpus. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and the code-switch sentences were manufactured by word substitution — they are not transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.audioautomatic-speech-recognition1K<n<10K2 likes85 downloads2mo agoHugging Face11code-switching /topic-classificationtabulartext-classificationn<1K0 likes63 downloads18d agoHugging Face12Atufa /codeswitch-fr-en-kyutai-stt Code-Switched French–English STT Probe Dataset Dataset Summary This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.audioautomatic-speech-recognitionn<1K0 likes54 downloads7mo agoHugging Face13Aynursusuz /jaen-codeswitch-tts jaen-codeswitch-tts Japanese/English code-switch synthetic speech: 10k varied-length conversational utterances, single voice (Qwen3-TTS clone of one English-male reference, spk_male). Per-utterance language routed to the dominant script. Columns: audio, text, speaker_id, language. audio10K<n<100K0 likes46 downloads3mo agoHugging Face14lxyuan /nemo-codeswitch-reasoning-debate Overview This is a synthetic, multilingual code-switching dataset. Each record contains: a realistic user query a long-form reasoning section a debate / counterargument section a concise final_answer It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses. This snapshot contains 574,977 rows and 10 string columns. Data provenance Important: Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.texttext-generation100K<n<1M0 likes45 downloads7mo agoHugging Face15Moamen-dcp /arazn_codeSwitched_mp3_full_notLoweraudio1K<n<10K0 likes38 downloads1y agoHugging Face16Moamen-dcp /arazn_codeSwitched_mp3_partial_processing_prepared_4_whisperMedium1K<n<10K0 likes36 downloads1y agoHugging Face17Nash-pAnDiTa /ASR_En_Ar_CodeSwitchingaudio10K<n<100K0 likes29 downloads2y agoHugging Face18UBC-NLP /NADI2026_subtask1.3_codeswitched_asraudio1K<n<10K0 likes26 downloads2mo agoHugging Face19zenyn /Code-Switching-Testaudion<1K0 likes22 downloads2y agoHugging Face20gimmy256 /african-codeswitching Pan-African Code-Switching Dataset Built with Adaptive Data by Adaption | Crane AI Labs Submitted to the Uncharted Data Challenge 2026 by Adaption Labs Overview The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations. Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.texttext-classificationn<1K0 likes21 downloads5mo agoHugging Face21necrosyth /indic-codeswitched-social-bias Dataset Description This dataset contains Indian code-switched sentences — informal text where speakers fluidly mix English with Indian regional languages — annotated for social bias across multiple dimensions. Languages covered include (but may not be limited to): Hindi-English (Hinglish) — hi-en Tamil-English (Tanglish) — ta-en Punjabi-English — pa-en Bengali-English — bn-en Gujarati-English — gu-en Telugu-English — te-en Kannada-English — kn-en Marathi-English — mr-en… See the full description on the dataset page: https://huggingface.co/datasets/necrosyth/indic-codeswitched-social-bias.textn<1K0 likes21 downloads5mo agoHugging Face22DynamicSuperb /CodeSwitchingSpeechIdentification_ASCENDaudion<1K0 likes17 downloads2y agoHugging Face23Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_All10K<n<100K0 likes16 downloads1y agoHugging Face24Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_Translate_AR1K<n<10K0 likes14 downloads1y agoHugging Face25polyglots /Sinhala-NewsCategory-Codeswitched75text1K<n<10K0 likes12 downloads1y agoHugging Face26Moamen-dcp /arazn_codeSwitched_mp3_full_processing_prepared_4_whisperSmall1K<n<10K0 likes12 downloads1y agoHugging Face27Aynursusuz /tts-jaen-codeswitch-6model tts-jaen-codeswitch-6model Japanese-English code-switch text spoken by 6 open-source TTS models, all cloning the SAME single reference voice (en-male). 10 mixed utterances (varied length). Each row: the text + reference audio + one audio column per model. Model columns: qwen3 · moss · voxcpm · omnivoice · zonos · cosyvoice. Per-utterance language routed to the dominant script (JA-heavy→Japanese, else English). No quality metrics included. audion<1K0 likes12 downloads3mo agoHugging Face28Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_Transcribe1K<n<10K0 likes11 downloads1y agoHugging Face29YouMike /code-switched-annotatedtext1K<n<10K0 likes11 downloads4mo agoHugging Face30DynamicSuperb /CodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-enaudion<1K0 likes10 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.