CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TMoC /SYNTH-Swallow-Math-Code-Mix Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2 This high-signal, all synthetic dataset is a complete shuffled mix of the following four sources: SYNTH ~63.5% SwallowCode-v2 ~15.5% SwallowMath-v2-textbook ~10.5% SwallowMath-v2-qa ~10.0% The motivation to provide this on HF was the need for a convenient, pre-shuffled merge of the highest quality synthetic / augmented datasets for small language model pre-training experiments as of… See the full description on the dataset page: https://huggingface.co/datasets/TMoC/SYNTH-Swallow-Math-Code-Mix.texttext-generation100M<n<1B1 likes736 downloads8mo agoHugging Face02CodeMixBench /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/CodeMixBench/CodeMixBench.texttext-generation10K<n<100K3 likes234 downloads1y agoHugging Face03GXLXY /mopd-math-code-mix MOPD math+code mix `train/`: math:code ≈ 1:1 平衡集(`math.parquet` + `code_*.parquet` shards) `val/mopd_val_mix.parquet`: AIME24 全量 + MATH-500 子集 + Eurus code_validation 子集 路由字段:`ability ∈ {math, code}` texttext-generation10K<n<100K0 likes132 downloads1mo agoHugging Face04Tanushreeeeee /CodeMixBench ℹ️Dataset Card for CodeMixBench [EMNLP'25] CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages Code-mixing is a linguistic phenomenon where multilingual speakers switch or mix two or more languages within a single utterance or conversation. To evaluate LLMs’ comprehension of multilingual code-mixed texts, we introduce CodeMixBench, a benchmark comprising eight tasks across 18 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Tanushreeeeee/CodeMixBench.texttext-generation10K<n<100K0 likes78 downloads9mo agoHugging Face05albertge /mix60k-math-code-sft mix60k: 50/50 OpenMathInstruct-2 + OpenCodeInstruct SFT mixture This dataset is from the paper Register Tokens for Bounded-State Reasoning in Diffusion Language Models. The exact 60,000-trace SFT mixture used to train the LLaDA-8B-Base and Dream-7B-Base main triad in the dLLM Registers project. Composition: ~30K math traces from OpenMathInstruct-2 and ~30K code traces from OpenCodeInstruct. License: governed by the upstream licenses of OpenMathInstruct-2 and OpenCodeInstruct.… See the full description on the dataset page: https://huggingface.co/datasets/albertge/mix60k-math-code-sft.texttext-generation10K<n<100K0 likes70 downloads6d agoHugging Face06PotatoHD /code-instruct-mixed Description Filtered/normalised subsets of public code-instruction datasets (Magicoder OSS-Instruct & Evol-Instruct, CodeFeedback, Glaive). The source column attributes each row to its origin; each source retains its upstream licence. Derived dataset. Source material retains its original per-item licence (see source/repo columns); treat as other / mixed. Provided as-is. Usage from datasets import load_dataset ds = load_dataset("PotatoHD/code-instruct-mixed") texttext-generation100K<n<1M0 likes40 downloads3mo agoHugging Face07synquid /agentic-code-sft-mix-v1 Agentic Code SFT Mix v1 Local derived SFT mixture for code-agent/tool-use training. This is not a single upstream dataset. It is a filtered local mixture built from: nvidia/OpenCodeInstruct, split train nvidia/Nemotron-SFT-OpenCode-v1, splits general, bash_only_tool, bash_only_tool_skills, question_tool, agent_skills, agent_skills_question_tool nvidia/Nemotron-SFT-SWE-v2, split agentless nvidia/Nemotron-SFT-SWE-v2, file data/swe.jsonl The output schema is JSONL with messages… See the full description on the dataset page: https://huggingface.co/datasets/synquid/agentic-code-sft-mix-v1.texttext-generation10K<n<100K0 likes33 downloads4mo agoHugging Face08DatarrX /Burmese-English-Code-Mixed-Corpus 🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt License: Apache 2.0 Description The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-English-Code-Mixed-Corpus.texttext-generation1K<n<10K7 likes22 downloads5mo agoHugging Face09RecursiveMAS /Mixture-Code RecursiveMAS Mixture-Code Project Page | Code | Paper We introduce RecursiveMAS, a multi-agent framework that scales agent collaboration through latent-space recursion. This dataset contains training examples for the Mixture-Style setting. Dataset Details Item Description Dataset RecursiveMAS/Mixture-Code Original file Mixture-Code.json Collaboration style Mixture-Style Used for code specialist inner agent training Split train Rows 2000… See the full description on the dataset page: https://huggingface.co/datasets/RecursiveMAS/Mixture-Code.texttext-generation1K<n<10K0 likes22 downloads3mo agoHugging Face10SPEAK-PP /v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs Sinhala Spelling Correction Dataset Dataset Description This dataset contains Sinhala text pairs for training spelling correction models. It includes: Dyslexic/Noisy sentences: Text with spelling errors, typos, and dyslexia-like mistakes Clean sentences: Corrected versions of the text Dataset Statistics Split Samples Train 37,712 Test 9,428 Total 47,140 Features dyslexic_sentence: Input text with errors (string)… See the full description on the dataset page: https://huggingface.co/datasets/SPEAK-PP/v3-v1_v2_code_mixed_syntheic_correct_noisy_pairs.texttext-generation10K<n<100K1 likes18 downloads4mo agoHugging Face11hksamm /Burmese-English-Code-Mixed-Corpus 🇲🇲 Burmese-English Code-Mixed Corpus ꒰ 1,111 Rows ꒱ A high-quality, human-curated dataset of code-mixed Burmese and English sentences, specifically designed for Natural Language Processing (NLP) and Machine Learning (ML) research. Dataset Details Organization: DatarrX Creator: Khant Sint Heinn (Kalix Louis) Number of Rows: 1,111 Language: Burmese (Unicode) & English Mix Dataset Format: .txt License: Apache 2.0 Description The Burmese-English Code-Mixed… See the full description on the dataset page: https://huggingface.co/datasets/hksamm/Burmese-English-Code-Mixed-Corpus.texttext-generation1K<n<10K0 likes18 downloads5mo agoHugging Face12nlpctx /telugu-qa-codemixed Telugu QA Paraphrases A synthetic multilingual query-rewriting dataset for evaluating retrieval robustness under Telugu-English code mixing. Dataset Description This dataset extends an existing Telugu QA dataset by generating multiple query variants with increasing levels of Telugu-English code mixing. Each example contains: question : Original English question answer : Ground-truth answer level_0 : English paraphrase level_1 : Light Telugu-English code mixing… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/telugu-qa-codemixed.textquestion-answering1K<n<10K0 likes17 downloads3mo agoHugging Face13Dsg2 /CodeMix CodeMix - a small finetune dataset 6k chat/response pairs, a balanced mix of: glaive-function-calling-v2 (agentic tool calling) hermes-function-calling-v1 (tool calling) CodeAlpaca-20k (coding) dolly-15k (instruct) texttext-generation100K<n<1M0 likes13 downloads3mo agoHugging Face14atx-labs /marathi-codemix-qagated Marathi Minglish QA ~1.09M synthetic Question–Answer pairs in code-mixed Romanized Marathi (Minglish), generated from Marathi Wikipedia articles. Designed for pretraining and SFT of Marathi-aware Small Language Models that should understand and generate the way Marathi is commonly written online — Roman-script Marathi naturally mixed with English terms. Example Question: Yashwant Dev kon hote exactly — sangeetkar, kavi, ki donhi? Answer: Yashwant Dev he… See the full description on the dataset page: https://huggingface.co/datasets/atx-labs/marathi-codemix-qa.tabulartext-generation1M<n<10M0 likes9 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.