CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes268 downloads1y agoHugging Face02georgechang8 /code_switch_yodas_zh Dataset Card for code-switching yodas This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon. Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.audio10K<n<100K4 likes197 downloads2y agoHugging Face03code-switching /text-summarizationtextsummarizationn<1K0 likes167 downloads25d agoHugging Face04ServiceNow-AI /asr_codeswitchedaudio1K<n<10K6 likes150 downloads2mo agoHugging Face05Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes132 downloads4mo agoHugging Face06devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes100 downloads7mo agoHugging Face07Panhapich /khmer-english-codeswitch-tts Khmer–English Code-Switch Synthetic Speech 8,177 utterances / 15.9 hours of synthetic Khmer–English code-switched speech, generated with VoxCPM2 from code-switch text manufactured by confirmed lexical substitution over a Khmer–English parallel corpus. ⚠️ This is synthetic speech, not recordings of people. Every utterance was produced by a TTS model, and the code-switch sentences were manufactured by word substitution — they are not transcripts of anything a person said. It is… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-english-codeswitch-tts.audioautomatic-speech-recognition1K<n<10K2 likes85 downloads2mo agoHugging Face08code-switching /topic-classificationtabulartext-classificationn<1K0 likes63 downloads18d agoHugging Face09Atufa /codeswitch-fr-en-kyutai-stt Code-Switched French–English STT Probe Dataset Dataset Summary This dataset contains 10 audio clips of French–English code-switched speech, each designed as a strict probe targeting a distinct acoustic or linguistic failure axis of Kyutai STT (kyutai/stt-1b-en_fr-trfs) — a streaming bilingual speech-to-text model. Probes cover non-native phonology, fast speech rate, word-level and phrase-level code-switching, intrasentential switching, disfluency with switching, and… See the full description on the dataset page: https://huggingface.co/datasets/Atufa/codeswitch-fr-en-kyutai-stt.audioautomatic-speech-recognitionn<1K0 likes54 downloads7mo agoHugging Face10Aynursusuz /jaen-codeswitch-tts jaen-codeswitch-tts Japanese/English code-switch synthetic speech: 10k varied-length conversational utterances, single voice (Qwen3-TTS clone of one English-male reference, spk_male). Per-utterance language routed to the dominant script. Columns: audio, text, speaker_id, language. audio10K<n<100K0 likes46 downloads3mo agoHugging Face11lxyuan /nemo-codeswitch-reasoning-debate Overview This is a synthetic, multilingual code-switching dataset. Each record contains: a realistic user query a long-form reasoning section a debate / counterargument section a concise final_answer It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses. This snapshot contains 574,977 rows and 10 string columns. Data provenance Important: Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.texttext-generation100K<n<1M0 likes45 downloads7mo agoHugging Face12Moamen-dcp /arazn_codeSwitched_mp3_full_notLoweraudio1K<n<10K0 likes38 downloads1y agoHugging Face13Nash-pAnDiTa /ASR_En_Ar_CodeSwitchingaudio10K<n<100K0 likes29 downloads2y agoHugging Face14UBC-NLP /NADI2026_subtask1.3_codeswitched_asraudio1K<n<10K0 likes26 downloads2mo agoHugging Face15zenyn /Code-Switching-Testaudion<1K0 likes22 downloads2y agoHugging Face16gimmy256 /african-codeswitching Pan-African Code-Switching Dataset Built with Adaptive Data by Adaption | Crane AI Labs Submitted to the Uncharted Data Challenge 2026 by Adaption Labs Overview The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations. Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.texttext-classificationn<1K0 likes21 downloads5mo agoHugging Face17necrosyth /indic-codeswitched-social-bias Dataset Description This dataset contains Indian code-switched sentences — informal text where speakers fluidly mix English with Indian regional languages — annotated for social bias across multiple dimensions. Languages covered include (but may not be limited to): Hindi-English (Hinglish) — hi-en Tamil-English (Tanglish) — ta-en Punjabi-English — pa-en Bengali-English — bn-en Gujarati-English — gu-en Telugu-English — te-en Kannada-English — kn-en Marathi-English — mr-en… See the full description on the dataset page: https://huggingface.co/datasets/necrosyth/indic-codeswitched-social-bias.textn<1K0 likes21 downloads5mo agoHugging Face18DynamicSuperb /CodeSwitchingSpeechIdentification_ASCENDaudion<1K0 likes17 downloads2y agoHugging Face19polyglots /Sinhala-NewsCategory-Codeswitched75text1K<n<10K0 likes12 downloads1y agoHugging Face20Aynursusuz /tts-jaen-codeswitch-6model tts-jaen-codeswitch-6model Japanese-English code-switch text spoken by 6 open-source TTS models, all cloning the SAME single reference voice (en-male). 10 mixed utterances (varied length). Each row: the text + reference audio + one audio column per model. Model columns: qwen3 · moss · voxcpm · omnivoice · zonos · cosyvoice. Per-utterance language routed to the dominant script (JA-heavy→Japanese, else English). No quality metrics included. audion<1K0 likes12 downloads3mo agoHugging Face21YouMike /code-switched-annotatedtext1K<n<10K0 likes11 downloads4mo agoHugging Face22DynamicSuperb /CodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-enaudion<1K0 likes10 downloads2y agoHugging Face23qadeesanoor /code-switching-codesaviours-si26-qadeesanoorDataset Summary This dataset contains Roman Urdu–English code-switched sentences, the way mixed-language text actually gets written in everyday Pakistani texting, tweeting, and casual conversation (e.g. "Aaj ka din bohot busy tha, had 3 meetings back to back"). Each sentence is broken down word-by-word, and every word is tagged with a language label. No existing Roman Urdu NLP resource handles this kind of within-sentence language mixing well — this dataset is a step toward building tools… See the full description on the dataset page: https://huggingface.co/datasets/qadeesanoor/code-switching-codesaviours-si26-qadeesanoor.text1K<n<10K0 likes10 downloads1mo agoHugging Face24MINERVA-TEAM /minerva-ar-en-edu-codeswitch-datasetaudio1K<n<10K0 likes9 downloads7mo agoHugging Face25marco-D-phoenix /IMDA_Codeswitch30audion<1K0 likes8 downloads1y agoHugging Face26MINERVA-TEAM /minerva-ar-en-codeswitch-topic-summarytext1K<n<10K0 likes8 downloads3mo agoHugging Face27WTFO /codeswitchinggated WTFO Code-Switching Speech Code-switching speech dataset prepared for automatic speech recognition training. Dataset fields audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature text: transcript duration: audio duration in seconds (float64) Split summary Split: train Examples: 98,662 Total duration: 567422.698 seconds (157.62 hours) The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.audioautomatic-speech-recognition10K<n<100K0 likes8 downloads1mo agoHugging Face28polyglots /punjabi-sentiment-tagger-stratified-codeswitchedtext1K<n<10K0 likes7 downloads1y agoHugging Face29farabi-lab /Code_switchinggated 🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset Dataset Summary Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh. The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.texttext-generationn<1K0 likes7 downloads2mo agoHugging Face30Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDotsaudio1K<n<10K0 likes6 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.