CoolFace
27 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MaLA-LM /PolyWritePolyWrite is a novel multilingual dataset developed for evaluating open-ended generation across 240 languages. We use ChatGPT to create diverse prompts in English, and then use Google Translate to translate these prompts into various languages, enabling models to generate creative content in multilingual settings. The benchmark includes 31 writing tasks—such as storytelling and email writing—across 155 unique prompts. To ensure translation quality, we back-translate the multilingual prompts… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/PolyWrite.tabulartext-generation10K<n<100K3 likes698 downloads2y agoHugging Face02silk-road /Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集 Wizard-LM包含了很多难度超过Alpaca的指令。 中文的问题翻译会有少量指令注入导致翻译失败的情况 中文回答是根据中文问题再进行问询得到的。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM 骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。 ( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 ) 骆驼项目不是商汤科技的官方产品。 Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.texttext-generation10K<n<100K98 likes670 downloads3y agoHugging Face03Goedel-LM /MathOlympiadBenchThis repository contains the MathOlympiadBench dataset, which is introduced in the paper Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction. Project Page: https://blog.goedel-prover.com Code Repository: https://github.com/Goedel-LM/Goedel-Prover-V2 MathOlympiadBench (Math Olympiad) comprises human-verified formalizations of Olympiad-level mathematical competition problems, sourced from Compfiles and IMOSLLean4 repository. MathOlympiadBench… See the full description on the dataset page: https://huggingface.co/datasets/Goedel-LM/MathOlympiadBench.texttext-generationn<1K17 likes627 downloads1y agoHugging Face04meet5568 /lma_datasets LMA Phase 1 --- Hindi and Nepali pretraining corpora Two monolingual corpora built for a pair of ~25M-parameter decoder-only Transformers. Hindi is the higher-resource language, Nepali the lower-resource one. Both are written in Devanagari (U+0900-U+097F), so script cannot be used to tell them apart --- separating them is the central technical problem this dataset solves rather than assumes. language documents characters manual (chars) tokens manual (tokens) train val test… See the full description on the dataset page: https://huggingface.co/datasets/meet5568/lma_datasets.texttext-generation1M<n<10M0 likes181 downloads9d agoHugging Face05Pythagoras-LM /SFT_Dataset Pythagoras SFT Dataset Project Page | GitHub | Paper Data Our training dataset consists of approximately 841K problems paired with Lean formal statements, formal proofs, and reasoning chains. We release a partial subset, which consists of 126K instances: 30K easy instances 49K medium instances 47K hard instances Complete data will be released soon. The complete explanation of the synthetic data generation pipeline can be found in Pythagoras-Prover: Advancing… See the full description on the dataset page: https://huggingface.co/datasets/Pythagoras-LM/SFT_Dataset.texttext-generation100K<n<1M9 likes137 downloads3mo agoHugging Face06LM-Lexicon /SlangBibTeX: @article{liu2026lmlexiconimprovingdefinitionmodeling, title={LM-Lexicon: Improving Definition Modeling via Harmonizing Semantic Experts}, author={Yang Liu and Jiaye Yang and Weikang Li and Jiahui Liang and Yang Li and Lingyong Yan}, year={2026}, eprint={2602.14060}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2602.14060}, } texttext-generation100K<n<1M3 likes94 downloads7mo agoHugging Face07Ashenone3 /LM-Searcher-Trajectory-228K LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical Encoding Repo: https://github.com/Ashone3/LM-Searcher Introduction We introduce LM-Searcher, a task-agnostic neural architecture search framework powered by LLMs. Usage To fine-tune your LLM with LLaMA Factory, configure the dataset_info.json file using the following template: "lm_searcher_trajectory_228k_0":{ "file_name": "your-dataset-dir/processed_data_encode_space0.jsonl"… See the full description on the dataset page: https://huggingface.co/datasets/Ashenone3/LM-Searcher-Trajectory-228K.texttext-generation100K<n<1M2 likes94 downloads1y agoHugging Face08KaiLo2026 /lmtq_999 🌏 中文多学科知识问答数据集 (Chinese Multi-disciplinary QA Dataset) 本数据集涵盖自然科学、人文社科、工程技术等多个维度的知识,旨在评估和提升模型在跨学科领域的推理与问答能力。数据集包含 多项选择题 (MCQ) 和 问答对 (QA) 两种形式,互为补充。 📊 数据集概览 数据集 ID: KaiLo2026/lmtq_999 总样本量: 999 条 (183 MCQ + 816 QA) 学科覆盖: 15+ 个主要领域 (天文、地学、生物、历史等) 语言: 简体中文 (zh-CN) 许可证: MIT License 适用任务: 知识问答、逻辑推理、学科能力评估、RAG 测试 📊 数据分布概览 1️⃣ MCQ 子集 (多项选择题) 总量: 183 条样本 | 特点: 适合评估模型的判别能力和知识广度。 学科领域 数量 占比 分布可视化 (比例缩放) 🌌 天文学 58 31.7%… See the full description on the dataset page: https://huggingface.co/datasets/KaiLo2026/lmtq_999.textquestion-answeringn<1K3 likes89 downloads6mo agoHugging Face09ibm-research /lmcache-agentic-traces_Otel Agentic LLM Traces – OTel Format Overview Real-world agentic LLM sessions converted to OpenTelemetry (OTel) trace format, derived from sammshen/lmcache-agentic-traces. Each session is a multi-turn agent interaction involving tool calls (bash commands, file edits, web search, etc.), spanning 5–50 turns and totalling 24,880 spans. Traces come from three agentic benchmarks: SWE-bench, GAIA, and WildClaw. They are formatted as OTel spans following gen_ai.* semantic… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/lmcache-agentic-traces_Otel.texttext-generationn<1K0 likes89 downloads1mo agoHugging Face10LM-Lexicon /WordnetCitation: @article{liu2026lmlexiconimprovingdefinitionmodeling, title={LM-Lexicon: Improving Definition Modeling via Harmonizing Semantic Experts}, author={Yang Liu and Jiaye Yang and Weikang Li and Jiahui Liang and Yang Li and Lingyong Yan}, year={2026}, eprint={2602.14060}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2602.14060}, } texttext-generation10K<n<100K1 likes78 downloads7mo agoHugging Face11lesserfield /lmsys-arena-human-preference-winner-43k-unfiltered lmsys-arena-human-preference-winner-43k-unfiltered This repository contains a dataset derived from the lmsys/lmsys-arena-human-preference-55k dataset, which is licensed under the Apache 2.0 License. Dataset Description The lmsys-arena-human-preference-winner-43k-unfiltered dataset is a collection of 43,000 samples, each containing an instruction (prompt) and an output (winning response) from real-world user and LLM conversations. The dataset is derived from the original… See the full description on the dataset page: https://huggingface.co/datasets/lesserfield/lmsys-arena-human-preference-winner-43k-unfiltered.texttext-generation10K<n<100K2 likes67 downloads2y agoHugging Face12Prateek-Tiwari10 /LMA_phase3_data LMA Phase 3 — transitive reasoning dataset (Hindi / Nepali) Synthetic transitive-comparison problems over people, in Hindi and Nepali, with chain-of-thought targets. Built for a study of data scaling and domain generalization in small from-scratch language models. The models finetuned on it are at Prateek-Tiwari10/LMA_phase3. Layout manifest_scaling.json machine-readable index of every file below train/{hi,ne}/train_{10,30,50,60}k.jsonl nested… See the full description on the dataset page: https://huggingface.co/datasets/Prateek-Tiwari10/LMA_phase3_data.tabulartext-generation100K<n<1M0 likes58 downloads9d agoHugging Face13MaLA-LM /mala-code-reasoning-v2 MaLA Corpus: Massive Language Adaptation Corpus This MaLA code and reasoning dataset (V2) is used for training EMMA-500 Llama 3(.1) Mono/Bi model series. 🤗MaLA-LM/emma-500-llama3-8b-mono: CPT model trained on monolingual data mix in 500+ languages 🤗MaLA-LM/emma-500-llama3-8b-bi: CPT model trained on monolingual data mix in 500+ languages + bilingual translation data in 2,500+ language pairs 🤗MaLA-LM/emma-500-llama3.1-8b-mono: CPT model trained on monolingual data mix in… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning-v2.texttext-generation10M<n<100M9 likes57 downloads1y agoHugging Face14agentlans /ytz20-LMSYS-Chat-GPT-5-Chat-Response ytz20/LMSYS-Chat-GPT-5-Chat-Response Dataset This is a reformatted, unofficial version of ytz20/LMSYS-Chat-GPT-5-Chat-Response According to the original authors: This dataset is an extension of the LMSYS-Chat-1M-Clean corpus, specifically curated by collecting high-quality, non-refusal responses from the GPT-5-Chat API. Modifications in this version: The "content" and "teacher_response" columns have been processed into the "input" and "output" columns in this dataset. An… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ytz20-LMSYS-Chat-GPT-5-Chat-Response.texttext-generation100K<n<1M1 likes53 downloads10mo agoHugging Face15natong19 /lmsys-chat-1m-filtered Dataset Card for natong19/lmsys-chat-1m-filtered Filtered version of lmsys/lmsys-chat-1m, a collection of one million real-world conversations with various LLMs. Data cleaning process inspired by OpenLeecher/lmsys_chat_1m_clean. Overview of filtering process: 1. Filtering REDACTED Entries Entries that were labeled as REDACTED due to containing Personally Identifiable Information (PII) were removed. 1000000 samples -> 733740 samples 2. Format validation… See the full description on the dataset page: https://huggingface.co/datasets/natong19/lmsys-chat-1m-filtered.texttext-classification100K<n<1M1 likes53 downloads9mo agoHugging Face16lmgame /VideoScienceBench VideoScienceBench A benchmark for evaluating video understanding and scientific reasoning in vision-language models. Each example pairs a textual description of an experiment (what is shown) with the correct scientific explanation (expected phenomenon). Dataset Summary Attribute Value Examples 160 Domains Physics, Chemistry Format JSONL (prompt + expected phenomenon + vid) Data Creation Pipeline Each researcher selects two or more scientific… See the full description on the dataset page: https://huggingface.co/datasets/lmgame/VideoScienceBench.textquestion-answeringn<1K3 likes52 downloads7mo agoHugging Face17LM-Lexicon /WikiBibTeX: @article{liu2026lmlexiconimprovingdefinitionmodeling, title={LM-Lexicon: Improving Definition Modeling via Harmonizing Semantic Experts}, author={Yang Liu and Jiaye Yang and Weikang Li and Jiahui Liang and Yang Li and Lingyong Yan}, year={2026}, eprint={2602.14060}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2602.14060}, } texttext-generation100K<n<1M0 likes30 downloads7mo agoHugging Face18Makaco /lmps-challenge-dpo-pairs LM Playschool Challenge — DPO ablation pairs Preference data for the study "Teaching or Sharpening? An Exploration of the Potential of DPO Post-Training for Dialogue Games", which post-trains Qwen3.5 models on the clembench / Playpen dialogue-game benchmark. The paper presenting the results of the study conducted using this dataset will be released upon acceptance. The question behind the study is whether DPO is suited to teach an LLM new skills, or whether it mainly sharpens… See the full description on the dataset page: https://huggingface.co/datasets/Makaco/lmps-challenge-dpo-pairs.texttext-generation1K<n<10K0 likes28 downloads1mo agoHugging Face19PJMixers-Dev /lmsys_lmsys-chat-1m-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT lmsys_lmsys-chat-1m-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT PJMixers-Dev/musab-mk_lmsys-chat-1m_deduped-prompts with responses generated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/lmsys_lmsys-chat-1m-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generation1K<n<10K0 likes27 downloads2y agoHugging Face20RexTRO111 /LMArena-SFT-Winners-Rextro Dataset Card for LMArena-SFT-Winners ⚔️ LMArena-SFT-Winners is a curated Supervised Fine-Tuning (SFT) dataset containing multi-turn conversations and single-turn prompts where the model responses consist strictly of winning outputs selected from head-to-head battles in Battle Mode. By filtering out losing completions, this dataset provides high-quality target completions filtered by human preference for direct instruction tuning and alignment. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/RexTRO111/LMArena-SFT-Winners-Rextro.texttext-generationn<1K0 likes23 downloads2mo agoHugging Face21MaLA-LM /mala-code-reasoning MaLA Corpus: Massive Language Adaptation Corpus This MaLA code and reasoning dataset is used for training 🤗MaLA-LM/emma-500-llama2-7b. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. This subset contains code, reasoning data, and scientific papers. Project page: https://mala-lm.github.io Paper: https://arxiv.org/abs/2409.17892… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-code-reasoning.texttext-generation10M<n<100M5 likes21 downloads1y agoHugging Face22iarcuschin /gemma-3-12b-it-lmsys-onpolicy-rollouts On-policy chat rollouts: google/gemma-3-12b-it on LMSYS-Chat-1M prompts Each row is a first-user-turn prompt sampled from lmsys/lmsys-chat-1m and a response generated on-policy by google/gemma-3-12b-it with vLLM (do_sample, temperature 0.7, top_p 1.0, max_new_tokens 768, seed 42). 24,991 rows. Built to match GemmaScope 2's instruction-tuned SAE training distribution (real model rollouts) for a short KL+MSE ("end-to-end") finetune of the released GemmaScope-2 residual SAE.… See the full description on the dataset page: https://huggingface.co/datasets/iarcuschin/gemma-3-12b-it-lmsys-onpolicy-rollouts.tabulartext-generation10K<n<100K0 likes21 downloads2mo agoHugging Face23jprivera44 /Training_36k_policy_monitor_lm_eval_flavor MO8 Policy + Monitor SFT Training Data Combined SFT training dataset for dual-role collusion research: a policy model that answers MCQs (sometimes wrong on purpose) and a monitor model that evaluates proposed answers (sometimes covering for wrong answers on purpose). Both roles are trained simultaneously from a single file. The metadata.role field distinguishes policy vs. monitor records. Dataset Structure Policy Monitor Total Target (schemer behavior) 4… See the full description on the dataset page: https://huggingface.co/datasets/jprivera44/Training_36k_policy_monitor_lm_eval_flavor.texttext-generation10K<n<100K0 likes20 downloads6mo agoHugging Face24AILaborant /ZERO-lm-dataset-chat-smallA combination of multiple open datasets combined into one for training simple lms. texttext-generation100K<n<1M0 likes13 downloads7mo agoHugging Face25lmcoleman /zeroclaw-tool-use-training ZeroClaw Tool-Use Training Data Training dataset for teaching LLMs to use tools in the ZeroClaw autonomous agent runtime. Format Standard chat-messages JSONL. Each line is a complete multi-turn conversation: {"messages": [{"role": "system", "content": "..."}, {"role": "user", "content":"..."}, {"role": "assistant", "content": "<tool_call>...</tool_call>"}]} Stats 457 examples with 1470 tool-call turns 25 tools covered: shell, file_read, file_write… See the full description on the dataset page: https://huggingface.co/datasets/lmcoleman/zeroclaw-tool-use-training.texttext-generationn<1K0 likes13 downloads6mo agoHugging Face26dominik-reiner /wolfgang-lm-synthetic-chat-v1 Wolfgang-LM: Synthetic Goethe Chats (v1) This dataset contains ~4,500 synthetic conversations designed to fine-tune language models into the persona of Johann Wolfgang von Goethe. It was generated as part of the Wolfgang-LM project. Dataset Details Size: ~4,500 samples Language: German (Modern User vs. Historical Goethe) Format: JSONL (ShareGPT compatible messages list) License: MIT License Generator Model: Google Gemini 2.5 Flash Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/dominik-reiner/wolfgang-lm-synthetic-chat-v1.texttext-generation1K<n<10K0 likes8 downloads8mo agoHugging Face27Yoda6131027 /twscholar-lm-dataset Dataset Card — twscholar-lm training data Summary A Traditional Chinese (Taiwan) academic-writing-polish dataset: colloquial draft sentences paired with polished, journal-register rewrites, plus a small general-conversation supplement for diversity. Built for supervised fine-tuning of a writing-assistant model on consumer hardware. 476 training examples, merged from three provenance classes. Provenance breakdown (this is the honest part) Count… See the full description on the dataset page: https://huggingface.co/datasets/Yoda6131027/twscholar-lm-dataset.texttext-generationn<1K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.