datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-instruct-mixed
Description
Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence.
Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is.
Usage
from datasets import load_dataset
ds = load_dataset("PotatoHD/code-instruct-mixed")
Burmese-English-Code-Mixed-Corpus
🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱
A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research.
Dataset Details
Organization: DatarrX
Creator: Khant Sint Heinn (Kalix Louis)
Number of Rows: 1,111
Language: Burmese (Unicode) & English Mix
Dataset Format: .txt
License: Apache 2.0
Description
The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs
Sinhala Spelling Correction Dataset
Dataset Description
This dataset contains Sinhala text pairs for training spelling correction models. It includes:
Dyslexic/Noisy sentences: Text with spelling errors, typos, and dyslexia-like mistakes
Clean sentences: Corrected versions of the text
Dataset Statistics
Split
Samples
Train
37,712
Test
9,428
Total
47,140
Features
dyslexic_sentence: Input text with errors (string)… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs.Burmese-English-Code-Mixed-Corpus
🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱
A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research.
Dataset Details
Organization: DatarrX
Creator: Khant Sint Heinn (Kalix Louis)
Number of Rows: 1,111
Language: Burmese (Unicode) & English Mix
Dataset Format: .txt
License: Apache 2.0
Description
The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/hksamm/Burmese-English-Code-Mixed-Corpus.telugu-qa-codemixed
Telugu QA Paraphrases
A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing.
Dataset Description
This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing.
Each example contains:
question : Original English question
answer : Ground-truth answer
level_0 : English paraphrase
level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.
