CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Perle-ai /ASR_Code_Switch ASR Code-Switching Benchmark A curated benchmark of 1,200 code-switching utterances (300 per language pair) for evaluating commercial ASR systems on multilingual speech with intra-sentential language switching. Paper Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German arXiv link Language pairs Split Language pair Samples Scripts egyptian_arabic_english Egyptian Arabic–English 300 Arabic + Latin… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/ASR_Code_Switch.audioautomatic-speech-recognition1K<n<10K12 likes752 downloads4mo agoHugging Face02Malikeh1375 /code-switching-tokenizer-robustness Code-Switching Dataset for Tokenizer Robustness Analysis Dataset Description This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models. Purpose Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.texttext-generation1K<n<10K2 likes290 downloads1y agoHugging Face03Rabe3 /saudi-english-code-switching-datasetaudio10K<n<100K0 likes280 downloads7mo agoHugging Face04nlpai-lab /ko_commongen_v2_code_switching 🇰🇷🇺🇸🇯🇵🇨🇳🇪🇸 KoCommonGEN v2 Code-switching This KoCommonGEN v2 Code-switching dataset consists of 99 samples for numerical commonsense reasoning, which were created relying on machine translation. The dataset can be found on Hugging Face at: nlpai-lab/ko_commongen_v2_code_switching This dataset contains code-switching data for the following languages: Korean (korean) English (english) Japanese (japan) Chinese (china) Spanish (espanol) (The code-switching data relies on… See the full description on the dataset page: https://huggingface.co/datasets/nlpai-lab/ko_commongen_v2_code_switching.textn<1K1 likes260 downloads2y agoHugging Face05NLPC-UOM /Sinhala-English-Code-Mixed-Code-Switched-Dataset Sinhala-English-Code-Mixed-Code-Switched-Dataset This dataset contains 10,000 comments that have been annotated at the sentence level for sentiment analysis, humor detection, hate speech detection, aspect identification, and language identification. The following is the tag scheme. Sentiment - Positive, Negative, Neutral, Conflict Humor - Humorous, Non humorous Hate Speech - Hate-Inducing, Abusive, Not offensive Aspect - Network, Billing or Price, Package, Customer Service, Data… See the full description on the dataset page: https://huggingface.co/datasets/NLPC-UOM/Sinhala-English-Code-Mixed-Code-Switched-Dataset.text-classification6 likes237 downloads2y agoHugging Face06Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_4_Turbo_new10K<n<100K0 likes219 downloads1y agoHugging Face07MohamedRashad /arabic-english-code-switching Thanks to ahmedheakl/arzen-llm-speech-ds as this dataset was built upon it ✨ The dataset was constructed using ahmed's dataset and different videos from the youtube. The scraped data doubled the initial dataset size after deduplication and cleaning. Citation If you use this dataset, please cite it as follows: @misc{rashad2024arabic, author = {Mohamed Rashad}, title = {arabic-english-code-switching}, year = {2024}, publisher = {Hugging Face}, url =… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/arabic-english-code-switching.audioautomatic-speech-recognition10K<n<100K34 likes218 downloads2y agoHugging Face08georgechang8 /code_switch_yodas_zh Dataset Card for code-switching yodas This dataset is derived from espnet/yodas, more details can be found here: https://huggingface.co/datasets/espnet/yodas This is a subset of the zh000 subset of espnet/yodas dataset, which selects videos with Mandarin-English code-switching phenomenon. Note that code-switching is only gauranteed per video rather than per utterance. Therefore, not every utterance in the dataset contains code-switching. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/georgechang8/code_switch_yodas_zh.audio10K<n<100K4 likes215 downloads2y agoHugging Face09liva-ai /code-switching-asrIf the samples in the dataset viewer don't load, they can also be accessed here. Code-Switching ASR Code-switching ASR is a speech dataset of code-switching between English and medium to low resource languages. Samples include en-sw, en-pcm, en-yo, and en-tl but we can deliver any of the languages listed here: https://huggingface.co/datasets/liva-ai/yapdo-convo (the hours have yet to be updated as of 07/14/2026 as we have much higher volume now - and we can easily collect more… See the full description on the dataset page: https://huggingface.co/datasets/liva-ai/code-switching-asr.audioautomatic-speech-recognitionn<1K0 likes191 downloads2mo agoHugging Face10Salesteq /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes184 downloads2mo agoHugging Face11Kennethdot /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/Kennethdot/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K4 likes184 downloads2mo agoHugging Face12Praxel /codeswitch-pairs-lase-heldout Codeswitch Pairs LASE — Western held-out corpus 1043 held-out cross-script utterance pairs from 8 ElevenLabs Western Multilingual voices. Used to evaluate generalisation of speaker encoders trained on Praxel/codeswitch-pairs-lase. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script = cross-script pair). Schema (manifest.jsonl) { "voice_id": "21m00Tcm4TlvDq8ikWAM"… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase-heldout.audioaudio-classification1K<n<10K0 likes182 downloads5mo agoHugging Face13code-switching /text-summarizationtextsummarizationn<1K0 likes171 downloads21d agoHugging Face14abdo1819 /arabic-english-code-switching-synthetic-asr Synthetic Arabic-English Code-Switched Speech for ASR This dataset contains synthetic speech generated for Egyptian Arabic-English code-switched automatic speech recognition. It is published separately from the human review annotations so the human audio remains in its upstream Hugging Face repository. Configurations Configuration Train Test Publication status synthetic 8,655 962 Contains 5,673 ArE-CSTD-derived texts; noncommercial/share-alike terms… See the full description on the dataset page: https://huggingface.co/datasets/abdo1819/arabic-english-code-switching-synthetic-asr.audioautomatic-speech-recognition1K<n<10K0 likes161 downloads1mo agoHugging Face15ServiceNow-AI /asr_codeswitchedaudio1K<n<10K6 likes160 downloads2mo agoHugging Face16Seif-Eldeen-Sameh /asr_codeswitched_dataset Arabic/English Code-Switched ASR Dataset Audio + transcripts of code-switched Egyptian Arabic and English speech, assembled to fine-tune ASR for Arab-world lecture content where dialectal Arabic and English technical vocabulary alternate within sentences. Composition Source Description EJUST custom recordings Locally recorded/segmented code-switched clips (segments_codeswitched.csv + WAVs) MohamedRashad/arabic-english-code-switching ~12,480 public… See the full description on the dataset page: https://huggingface.co/datasets/Seif-Eldeen-Sameh/asr_codeswitched_dataset.audioautomatic-speech-recognition10K<n<100K1 likes158 downloads4mo agoHugging Face17FatimahEmadEldin /cafe-algerian-codeswitch-speech CAFE Algerian Codeswitch Speech This dataset contains Algerian Arabic and French code-switched speech. Repository Path: FatimahEmadEldin/cafe-algerian-codeswitch-speech audio1K<n<10K0 likes156 downloads3mo agoHugging Face18code-switching /question-answertext1K<n<10K0 likes132 downloads21d agoHugging Face19Praxel /codeswitch-pairs-lase-indian Codeswitch Pairs LASE — Indian-accent held-out corpus 1369 held-out cross-script utterance pairs from 8 ElevenLabs Indian-English Multilingual voices. Surfaces the accent-conditional finding: off-the-shelf encoders cluster Indian-accent voices closely regardless of script, while Western voices show large script-conditional gaps. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script =… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase-indian.audioaudio-classification1K<n<10K2 likes123 downloads5mo agoHugging Face20ghanaopenai /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K5 likes117 downloads9d agoHugging Face21BrunoHays /english-x-code-switching Synthetic English Code-Switching Evaluation Set This dataset contains synthetic long-form English code-switching audio samples built from ML-SUPERB hybrid data. Each mixed sample combines English with exactly one additional language. Durations are randomly drawn between 5 and 15 minutes, and each sample contains one or two code switches. The random seed is stored per row. Each selected utterance chunk is RMS-normalized to -20.0 dBFS before concatenation, with peak limiting at 0.99.… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-x-code-switching.audioautomatic-speech-recognitionn<1K0 likes113 downloads5mo agoHugging Face22code-switching /naturalnesstabularn<1K0 likes112 downloads21d agoHugging Face23BrunoHays /english-en-x-code-switching-main-lang English EN-X Code-Switching Main-Language This dataset contains synthetic English-plus-one-language code-switching samples built from FLEURS. Source data is google/fleurs at revision refs/convert/parquet, split test, resampled to 16000 Hz. The generator uses seed 42 and creates 50 mixed samples. Each mixed sample contains English and exactly one of Spanish, Portuguese, French, German, or Italian, sampled uniformly. Each selected utterance is RMS-normalized to -20.0 dBFS before… See the full description on the dataset page: https://huggingface.co/datasets/BrunoHays/english-en-x-code-switching-main-lang.audioautomatic-speech-recognitionn<1K0 likes111 downloads5mo agoHugging Face24ghananlpcommunity /Ghana_English-Twi_Code-switching_Speech Dataset Card for KasaSpeech Dataset Summary KasaSpeech is a large-scale English–Twi code-switching speech dataset developed to advance research in speech technologies for English and Twi. The dataset comprises 54,855 transcribed speech recordings collected from speakers across Ghana and is designed to capture natural code-switching between English and Twi across a diverse range of everyday topics and communication scenarios With over 95 hours of manually… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/Ghana_English-Twi_Code-switching_Speech.audiotext-to-speech10K<n<100K0 likes105 downloads2mo agoHugging Face25devrahulbanjara /ne-en-codeswitching-asr-technical-interview Dataset Summary This dataset contains audio recordings and text transcripts of Nepali-English code-switched speech in the context of technical interviews. It is specifically designed to handle the linguistic complexities of Nepali software engineers, developers, and IT professionals who frequently mix English technical terminology (e.g., AWS, S3 lifecycle policies, RAG pipelines, VPC peering) with conversational Nepali grammar. It is an excellent resource for fine-tuning ASR models… See the full description on the dataset page: https://huggingface.co/datasets/devrahulbanjara/ne-en-codeswitching-asr-technical-interview.audioautomatic-speech-recognitionn<1K3 likes102 downloads7mo agoHugging Face26Moamen-dcp /arazn_codeSwitched_mp3_full_notLower_notMultiDots_prepared_4_whisper_turbo_transcription1K<n<10K0 likes98 downloads1y agoHugging Face27Tim2190 /kazakh-codeswitch-asr Kazakh Code-Switching ASR Benchmark A benchmark for evaluating ASR systems on natural Kazakh speech that code-switches with Russian — the everyday Kazakh–Russian mixing found in stand-up, interviews and vlogs, not scripted read speech. This is, to our knowledge, the first speech/ASR resource targeting the Kazakh–Russian code-switching pair (existing Kazakh–Russian NLP resources are text-only). Code, scoring harness and full analysis:… See the full description on the dataset page: https://huggingface.co/datasets/Tim2190/kazakh-codeswitch-asr.audioautomatic-speech-recognitionn<1K0 likes96 downloads2mo agoHugging Face28SDAIANCAI /Ar-En-Code-Switching-Textual-Dataset ArE-CSTD: Arabic-English Code-Switching Textual Dataset The National Center for Artificial Intelligence at the Saudi Data and Artificial Intelligence Authority (SDAIA), published the "ArE-CSTD" dataset, which stands for "Arabic-English Code-Switching Textual Dataset”. This dataset contains 330K dialectical Arabic-English code-swithing sentences generated by the large language model GPT-4. TXT Files There are 6 txt files. 2 files for Modern Standard Arabic(MSA) train and… See the full description on the dataset page: https://huggingface.co/datasets/SDAIANCAI/Ar-En-Code-Switching-Textual-Dataset.texttext-generation100K<n<1M2 likes93 downloads2y agoHugging Face29prokelly /neuromoyo-sahara-codeswitch-benchmark NEUROMOYO — Sahara CodeSwitch Africa Benchmark 🔗 Live Benchmark Results Interactive benchmark: https://www.neuromoyo.app/benchmark This page presents the benchmark results, methodology, model comparisons, robustness analyses, reproducibility information, and limitations for the NEUROMOYO evaluation on African code-switched speech. 🚀 Live NEUROMOYO Demo Live application: https://www.neuromoyo.app The live NEUROMOYO application demonstrates the… See the full description on the dataset page: https://huggingface.co/datasets/prokelly/neuromoyo-sahara-codeswitch-benchmark.tabularn<1K0 likes82 downloads1d agoHugging Face30suru8-ai /chichewa_english_code_switch_dataset Chichewa-English Code-Switched Speech Dataset Dataset Description A speech dataset containing 247 audio recordings of Chichewa-English code-switched phrases. Code-switching — the practice of alternating between two or more languages within a single conversation — is extremely common in Malawi and across multilingual African communities. This dataset captures that natural linguistic behavior in spoken form. Purpose This dataset is designed to support research… See the full description on the dataset page: https://huggingface.co/datasets/suru8-ai/chichewa_english_code_switch_dataset.audioautomatic-speech-recognitionn<1K0 likes80 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.