datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-switching-tokenizer-robustness
Code-Switching Dataset for Tokenizer Robustness Analysis
Dataset Description
This dataset is designed for tokenizer robustness testing in multilingual and code-switching contexts. It contains identical content expressed across 16 different language variants, including pure English and 15 English-X code-switching pairs, allowing researchers to isolate tokenization effects from semantic differences when evaluating language models.
Purpose
Tokenizer Comparison:… See the full description on the dataset page: https://huggingface.co/datasets/Malikeh1375/code-switching-tokenizer-robustness.code-switched-student-blindspot-eval
Code Switched Student Blind Spot Evaluation
Overview
This repository contains a small manual evaluation of Qwen/Qwen2.5-1.5B-Instruct on code-switched South Asian international student prompts. The goal is to test whether a small open-weight instruction model can understand Pakistani English mixed with Roman Urdu/Hindi in situations shaped by scholarship pressure, family expectations, limited resources, and international student life in Malaysia.
Blind… See the full description on the dataset page: https://huggingface.co/datasets/Zerothe00/code-switched-student-blindspot-eval.nemo-codeswitch-reasoning-debate
Overview
This is a synthetic, multilingual code-switching dataset. Each record contains:
a realistic user query
a long-form reasoning section
a debate / counterargument section
a concise final_answer
It is designed for experiments in multilingual generation, code-switch robustness, and reasoning/debate style responses.
This snapshot contains 574,977 rows and 10 string columns.
Data provenance
Important:
Verify that your intended usage and redistribution complies… See the full description on the dataset page: https://huggingface.co/datasets/lxyuan/nemo-codeswitch-reasoning-debate.Code_switching
🇰🇿 Kazakh-Russian Code-Switching Normalization Dataset
Dataset Summary
Kazakh-Russian Code-Switching Normalization Dataset is a bilingual instruction-following dataset designed for identifying and rewriting Kazakh-Russian mixed-language text into clean Kazakh.
The dataset focuses on informal communication, where Kazakh speakers may naturally mix Russian and Kazakh in one message. Each sample contains a prompt with code-switching, a response that identifies the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Code_switching.
