CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes286 downloads1y agoHugging Face02code-switching /text-summarizationtextsummarizationn<1K0 likes167 downloads22d agoHugging Face03devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes104 downloads7mo agoHugging Face04code-switching /topic-classificationtabulartext-classificationn<1K0 likes61 downloads15d agoHugging Face05Nash-pAnDiTa /ASR_En_Ar_CodeSwitchingaudio10K<n<100K0 likes49 downloads2y agoHugging Face06DynamicSuperb /CodeSwitchingSpeechIdentification_ASCENDaudion<1K0 likes35 downloads2y agoHugging Face07gimmy256 /african-codeswitching Pan-African Code-Switching Dataset Built with Adaptive Data by Adaption | Crane AI Labs Submitted to the Uncharted Data Challenge 2026 by Adaption Labs Overview The first open-source, richly-annotated pan-African code-switching dataset covering authentic language mixing patterns across 4 African regions — East Africa, South Africa, and West Africa — in 10 distinct language-pair configurations. Code-switching (mixing two or more languages in a single utterance) is how… See the full description on the dataset page: https://huggingface.co/datasets/gimmy256/african-codeswitching.texttext-classificationn<1K0 likes20 downloads5mo agoHugging Face08zenyn /Code-Switching-Testaudion<1K0 likes18 downloads2y agoHugging Face09WTFO /codeswitchinggated WTFO Code-Switching Speech Code-switching speech dataset prepared for automatic speech recognition training. Dataset fields audio: self-contained 16 kHz mono PCM16 FLAC bytes using the Hugging Face Audio feature text: transcript duration: audio duration in seconds (float64) Split summary Split: train Examples: 98,662 Total duration: 567422.698 seconds (157.62 hours) The Parquet shards embed the audio bytes. They do not depend on the original… See the full description on the dataset page: https://huggingface.co/datasets/WTFO/codeswitching.audioautomatic-speech-recognition10K<n<100K0 likes11 downloads1mo agoHugging Face10DynamicSuperb /CodeSwitchingSemanticGrammarAcceptabilityComparison_CSZS-zh-enaudion<1K0 likes10 downloads2y agoHugging Face11qadeesanoor /code-switching-codesaviours-si26-qadeesanoorDataset Summary This dataset contains Roman Urdu–English code-switched sentences, the way mixed-language text actually gets written in everyday Pakistani texting, tweeting, and casual conversation (e.g. "Aaj ka din bohot busy tha, had 3 meetings back to back"). Each sentence is broken down word-by-word, and every word is tagged with a language label. No existing Roman Urdu NLP resource handles this kind of within-sentence language mixing well — this dataset is a step toward building tools… See the full description on the dataset page: https://huggingface.co/datasets/qadeesanoor/code-switching-codesaviours-si26-qadeesanoor.text1K<n<10K0 likes10 downloads1mo agoHugging Face12farabi-lab /Code_switchinggated 🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset Dataset Summary Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh. The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.texttext-generationn<1K0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.