CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rabe3 /saudi-english-code-switching-datasetaudio10K<n<100K0 likes294 downloads7mo agoHugging Face02Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes292 downloads1y agoHugging Face03nlpai-lab /ko_commongen_v2_code_switching 🇰🇷🇺🇸🇯🇵🇨🇳🇪🇸 KoCommonGEN v2 Code-switching This KoCommonGEN v2 Code-switching dataset consists of 99 samples for numerical commonsense reasoning, which were created relying on machine translation. The dataset can be found on Hugging Face at: nlpai-lab/ko_commongen_v2_code_switching This dataset contains code-switching data for the following languages: Korean (korean) English (english) Japanese (japan) Chinese (china) Spanish (espanol) (The code-switching data relies on… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/ko_commongen_v2_code_switching.textn<1K1 likes260 downloads2y agoHugging Face04Kennethdot /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K4 likes187 downloads2mo agoHugging Face05MohamedRashad /arabic-english-code-switching Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨ The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning. Citation If you use this dataset, please cite it as follows: @misc{rashad2024arabic, author = {Mohamed Rashad}, title = {arabic-english-code-switching}, year = {2024}, publisher = {Hugging Face}, url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.audioautomatic-speech-recognition10K<n<100K34 likes182 downloads2y agoHugging Face06liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes181 downloads2mo agoHugging Face07code-switching /text-summarizationtextsummarizationn<1K0 likes159 downloads21d agoHugging Face08abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes156 downloads1mo agoHugging Face09code-switching /question-answertext1K<n<10K0 likes130 downloads21d agoHugging Face10ghanaopenai /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K5 likes116 downloads8d agoHugging Face11BrunoHays /english-x-code-switching Synthetic English Code-Switching Evaluation Set This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.audioautomatic-speech-recognitionn<1K0 likes111 downloads5mo agoHugging Face12code-switching /naturalnesstabularn<1K0 likes109 downloads21d agoHugging Face13BrunoHays /english-en-x-code-switching-main-lang English EN-X Code-Switching Main-Language This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.audioautomatic-speech-recognitionn<1K0 likes109 downloads5mo agoHugging Face14ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K0 likes109 downloads2mo agoHugging Face15devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes103 downloads7mo agoHugging Face16SDAIANCAI /Ar-En-Code-Switching-Textual-Dataset ArE-CSTD: Arabic-English Code-Switching Textual Dataset The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”. This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4. TXT Files There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.texttext-generation100K<n<1M2 likes93 downloads2y agoHugging Face17ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech-ipa KasaSpeech English–Twi Code-Switching Speech — IPA A phonemised version of ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech (KasaSpeech) with one added column: ipa. Every other column — audio included — is carried over byte-for-byte, and row order is unchanged, so this dataset aligns one-to-one with the original. The ipa column Each transcript is converted to a phoneme sequence with ghanag2p-uni, the Twi-only grapheme-to-phoneme library built on ghana-g2p… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech-ipa.audiotext-to-speech10K<n<100K0 likes78 downloads2mo agoHugging Face18BrunoHays /fleurs_code_switching_test FLEURS Code-Switching Evaluation Set Dataset Summary This dataset is a synthetic code-switching evaluation set built from the google/fleurs corpus.Each sample is a single long-form audio sequence (minimum 5 minutes by default) composed by concatenating short utterances from multiple languages. The goal is to provide a controlled benchmark for testing ASR robustness when language switches happen frequently inside one recording. How The Dataset Was Curated… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/fleurs_code_switching_test.audio1K<n<10K0 likes74 downloads6mo agoHugging Face19anishka /CodeSwitching-TE-ENAnCora Catalan NER. This is a dataset for Named Eentity Reacognition (NER) from Ancora corpus adapted for Machine Learning and Language Model evaluation purposes. Since multiwords (including Named Entites) in the original Ancora corpus are aggregated as a single lexical item using underscores (e.g. "Ajuntament_de_Barcelona") we splitted them to align with word-per-line format, and added conventional Begin-Inside-Outside (IOB) tags to mark and classify Named Entites. We did not filter out the different categories of NEs from Ancora (weak and strong). We did 6 minor edits by hand. AnCora corpus is used under [CC-by] (https://creativecommons.org/licenses/by/4.0/) licence. This dataset was developed by BSC TeMU as part of the AINA project, and to enrich the Catalan Language Understanding Benchmark (CLUB).token-classification2 likes70 downloads3y agoHugging Face20code-switching /topic-classificationtabulartext-classificationn<1K0 likes62 downloads14d agoHugging Face21thetaone-ai /Korean-Japanese-Code-Switching-Speech Korean-Japanese-Code-Switching-Speech This dataset contains Korean-Japanese code-switching speech recordings with sentence-level transcriptions. It was introduced in the paper Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs. Since there is an extremely small amount of Korean-Japanese code-switching data available, this dataset was designed to be used as a small-scale evaluation dataset. The dataset consists of code-switching recordings… See the full description on the dataset page: https://huggingface.co/datasets/thetaone-ai/Korean-Japanese-Code-Switching-Speech.audioautomatic-speech-recognitionn<1K4 likes58 downloads4mo agoHugging Face22BSC-LT /BSCs_Code_Switching_CA-ES_ASR_TestThe BSC's Code-Switching Catalan-Spanish ASR Test is a speech dataset of 4 hours and 9 minutes. It consists of carefully selected recordings that feature code-switching between Catalan and Spanish. This dataset is designed to be a test set for Catalan ASR systems that need to handle code-switching to Spanish.automatic-speech-recognitionn<1K0 likes57 downloads10mo agoHugging Face23Hamza-Ali01 /code-switching-codesaviours-si26-hamzatext1K<n<10K1 likes57 downloads1mo agoHugging Face24samihyounes /Moroccan-Codeswitching Moroccan Darija Code-Switched Corpus (Sentence-level TSV) Dataset Summary This dataset contains sentence/post-level code-switched Moroccan Darija text with a single label per text unit. It is intended to support NLP research on Moroccan Darija (Darija), an under-resourced Arabic variety, and on sentence-level code-switching / language identification in Moroccan online text. Languages The corpus may contain Moroccan Darija (often ary) and code-switching with:… See the full description on the dataset page: https://huggingface.co/datasets/samihyounes/Moroccan-Codeswitching.texttext-classification100K<n<1M0 likes56 downloads7mo agoHugging Face25Muhammad-Ahmad-1263 /code-switching-codesaviours-si26-muhammadahmad Code-Switching Codesaviours SI26 — Muhammad Ahmad Dataset Description This dataset contains 155 naturally occurring Roman Urdu–English code-switched sentences (1,400+ word-level entries), reflecting how Roman Urdu and English are mixed in everyday informal communication by Pakistani speakers online (Twitter/X, WhatsApp, YouTube comments, Reddit). Code-switching — alternating between two or more languages within a single sentence or conversation — is extremely… See the full description on the dataset page: https://huggingface.co/datasets/Muhammad-Ahmad-1263/code-switching-codesaviours-si26-muhammadahmad.text1K<n<10K0 likes52 downloads1mo agoHugging Face26Nash-pAnDiTa /ASR_En_Ar_CodeSwitchingaudio10K<n<100K0 likes49 downloads2y agoHugging Face27sanaisrail /code-switching-codesaviours-si26-Sanatext1K<n<10K0 likes45 downloads25d agoHugging Face28BrunoHays /english-x-code-switching-samples Synthetic English Code-Switching Evaluation Set Samples This dataset contains the individual normalized utterance chunks used to build the paired mixed dataset. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching-samples.audioautomatic-speech-recognition10K<n<100K0 likes44 downloads5mo agoHugging Face29waroodzkhan /code-switching-codesaviours-si26-warood Code-Switching Codesaviours SI-26 Dataset Dataset Description This dataset contains 150 naturally occurring Roman Urdu–English code-switched sentences, commonly spoken by Pakistani speakers in casual, everyday communication. Each sentence has been broken down into individual words, and every word is labeled by language. Roman Urdu–English code-switching is extremely common in Pakistan (spoken/written by an estimated 230+ million people) but is poorly handled by… See the full description on the dataset page: https://huggingface.co/datasets/waroodzkhan/code-switching-codesaviours-si26-warood.text1K<n<10K0 likes38 downloads29d agoHugging Face30DynamicSuperb /CodeSwitchingSpeechIdentification_ASCENDaudion<1K0 likes36 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.