CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MaLA-LM /mala-monolingual-filter MaLA Corpus: Massive Language Adaptation Corpus This is a cleaned version with some necessary data cleaning. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-filter.text-generation3 likes5.2k downloads2mo agoHugging Face02MaLA-LM /mala-monolingual-integration MaLA Corpus: Massive Language Adaptation Corpus This is the noisy version that integrates texts from different sources. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-integration.texttext-generation1B<n<10B2 likes3.5k downloads2mo agoHugging Face03MaLA-LM /mala-monolingual-dedup MaLA Corpus: Massive Language Adaptation Corpus This is a deduplicated version after minhash and exact hash deduplication. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-dedup.text-generation2 likes2.8k downloads2mo agoHugging Face04MaLA-LM /mala-monolingual-split MaLA Corpus: Massive Language Adaptation Corpus This version contains train and validation splits. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource languages, the… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-split.texttext-generation100M<n<1B4 likes2.8k downloads2mo agoHugging Face05billion-word-benchmark /lm1bA benchmark corpus to be used for measuring progress in statistical language modeling. This has almost one billion words in the training data.text-generation19 likes1.6k downloads3y agoHugging Face06causal-lm /instructions Merged Instructions Dataset Merged Dataset for the response of instructions. texttext-generation10M<n<100M26 likes1.5k downloads3y agoHugging Face07lmqg /qg_squad[SQuAD](https://rajpurkar.github.io/SQuAD-explorer/) evaluation set for the question generation (QG) models. The split of test and development set follows the ["Neural Question Generation"](https://arxiv.org/abs/1705.00106) work and is compatible with the [leader board](https://paperswithcode.com/sota/question-generation-on-squad11).texttext-generation10K<n<100K9 likes1.5k downloads4y agoHugging Face08lmqg /qg_esquad[SQuAD-es](https://huggingface.co/datasets/squad_es) dataset for question generation (QG) task.texttext-generation10K<n<100K0 likes1.5k downloads4y agoHugging Face09lmqg /qg_jaquad[JaQuAD](https://github.com/SkelterLabsInc/JaQuAD) dataset for question generation (QG) task. The test set of the original data is not publicly released, so we randomly sampled test questions from the training set.texttext-generation10K<n<100K5 likes1.4k downloads4y agoHugging Face10lmqg /qg_koquad[KorQuAD](https://huggingface.co/datasets/squad_kor_v1) dataset for question generation (QG) task.texttext-generation10K<n<100K9 likes1.3k downloads4y agoHugging Face11sammshen /lmcache-agentic-traces LMCache Agentic Dataset Collection A curated dataset collection of 787 multi-turn agentic LLM sessions (24,881 total LLM iterations) designed for benchmarking stateful LLM serving systems. Every session exhibits at least 5 turns with prefix growth and builds to at least 10K tokens of context — making it ideal for evaluating tiered KV Cache solutions like LMCache. Motivation Modern LLM agents (coding assistants, research agents, tool-calling systems) make dozens of… See the full description on the dataset page: https://huggingface.co/datasets/sammshen/lmcache-agentic-traces.tabulartext-generation10K<n<100K16 likes1.3k downloads4mo agoHugging Face12lmqg /qg_itquad[SQuAD-it](https://huggingface.co/datasets/squad_it) dataset for question generation (QG) task.text-generation10K<n<100K2 likes999 downloads4y agoHugging Face13lmqg /qg_subjqa[SubjQA](https://github.com/megagonlabs/SubjQA) dataset for question generation (QG) task.tabulartext-generation10K<n<100K1 likes847 downloads4y agoHugging Face14lmqg /qg_ruquad[SberSQuAD](https://huggingface.co/datasets/sberquad) dataset for question generation (QG) task.text-generation10K<n<100K3 likes831 downloads4y agoHugging Face15tokyotech-llm /lmsys-chat-1m-synth LMSYS-Chat-1M-Synth: Japanese/English Synthetic Conversation Dataset Derived from LMSYS-Chat-1M This repository contains a series of Japanese and English conversation datasets derived from LMSYS-Chat-1M. Llama-3.1-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.1 and Llama-3.1-Swallow-70B-Instruct-v0.1 Gemma-2-LMSYS-Chat-1M-Synth Utilized in the post-training of Llama-3.1-Swallow-8B-Instruct-v0.3 and Llama-3.1-Swallow-70B-Instruct-v0.3… See the full description on the dataset page: https://huggingface.co/datasets/tokyotech-llm/lmsys-chat-1m-synth.text-generation100K<n<1M23 likes814 downloads7mo agoHugging Face16MaLA-LM /PolyWritePolyWrite is a novel multilingual dataset developed for evaluating open-ended generation across 240 languages. We use ChatGPT to create diverse prompts in English, and then use Google Translate to translate these prompts into various languages, enabling models to generate creative content in multilingual settings. The benchmark includes 31 writing tasks—such as storytelling and email writing—across 155 unique prompts. To ensure translation quality, we back-translate the multilingual prompts… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/PolyWrite.tabulartext-generation10K<n<100K3 likes698 downloads2y agoHugging Face17silk-road /Wizard-LM-Chinese-instruct-evolWizard-LM-Chinese是在MSRA的Wizard-LM数据集上,对指令进行翻译,然后再调用GPT获得答案的数据集 Wizard-LM包含了很多难度超过Alpaca的指令。 中文的问题翻译会有少量指令注入导致翻译失败的情况 中文回答是根据中文问题再进行问询得到的。 我们会陆续将更多数据集发布到hf,包括 Coco Caption的中文翻译 CoQA的中文翻译 CNewSum的Embedding数据 增广的开放QA数据 WizardLM的中文翻译 如果你也在做这些数据集的筹备,欢迎来联系我们,避免重复花钱。 骆驼(Luotuo): 开源中文大语言模型 https://github.com/LC1332/Luotuo-Chinese-LLM 骆驼(Luotuo)项目是由冷子昂 @ 商汤科技, 陈启源 @ 华中师范大学 以及 李鲁鲁 @ 商汤科技 发起的中文大语言模型开源项目,包含了一系列语言模型。 ( 注意: 陈启源 正在寻找2024推免导师,欢迎联系 ) 骆驼项目不是商汤科技的官方产品。 Citation… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Wizard-LM-Chinese-instruct-evol.texttext-generation10K<n<100K98 likes670 downloads3y agoHugging Face18Goedel-LM /MathOlympiadBenchThis repository contains the MathOlympiadBench dataset, which is introduced in the paper Goedel-Prover-V2: Scaling Formal Theorem Proving with Scaffolded Data Synthesis and Self-Correction. Project Page: https://blog.goedel-prover.com Code Repository: https://github.com/Goedel-LM/Goedel-Prover-V2 MathOlympiadBench (Math Olympiad) comprises human-verified formalizations of Olympiad-level mathematical competition problems, sourced from Compfiles and IMOSLLean4 repository. MathOlympiadBench… See the full description on the dataset page: https://huggingface.co/datasets/Goedel-LM/MathOlympiadBench.texttext-generationn<1K17 likes627 downloads1y agoHugging Face19MaLA-LM /mala-opus-dedup-2410-reLIDtabulartranslation10B<n<100B1 likes609 downloads11mo agoHugging Face20lmqg /qg_dequad[GermanSQuAD](https://huggingface.co/datasets/deepset/germanquad) dataset for question generation (QG) task.text-generation10K<n<100K1 likes600 downloads4y agoHugging Face21lmqg /qg_squadshifts[SQuAD Shifts](https://modestyachts.github.io/squadshifts-website/index.html) dataset for question generation (QG) task.texttext-generation10K<n<100K1 likes520 downloads4y agoHugging Face22VmaxRL /bugpilot-bugintro-lm-modify-gpt55-1k-oracleclean-443-20260506 BugPilot LM-Modify GPT-5.5 1k Oracle-Clean 443 This dataset contains the 443 task directories from the repaired LM-modify workspace that currently pass the oracle audit. Source workspace: /data/augustine/demiurge/projects/experimental/training_swe_skrl_tinker/audits/lm_modify_target600_repair_workspace_20260506_v4 Source audit: full_oracle_postswaps_20260506_multinode32_c4 Export date: 2026-05-06 The directory layout matches the original task dataset layout: one task directory per… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/bugpilot-bugintro-lm-modify-gpt55-1k-oracleclean-443-20260506.text-generation0 likes340 downloads5mo agoHugging Face23lm-provers /FineProofs-SFT FineProofs SFT Dataset Description FineProofs SFT is a high-quality supervised fine-tuning dataset containing mathematical Olympiad problems paired with chain-of-thought reasoning and formal proofs distilled from DeepSeek-Math-V2. The dataset comprises 7,777 samples (4,300 unique problems) sourced from international Olympiad competitions and Art of Problem Solving (AoPS), each annotated with: Detailed reasoning traces (thinking content) generated by… See the full description on the dataset page: https://huggingface.co/datasets/lm-provers/FineProofs-SFT.tabulartext-generation10K<n<100K43 likes267 downloads7mo agoHugging Face24PNYX /reasoning_gym_lmeh Reasoning-gym tasks This is an implementation of reasoning-gym into a fixed dataset to be used within lm-evaluation-harness ecosystem. The dataset is meant to be used with semantic extraction (on most cases), applied by means of the a-vert method. Some higher level tasks (like 'codeio') use the native reasoning-gym methods to extract scores. For each task 100 samples are geenrated and most instructions or hints are removed (we dont want to condition the LM answer). Currently we… See the full description on the dataset page: https://huggingface.co/datasets/PNYX/reasoning_gym_lmeh.textquestion-answering1K<n<10K0 likes235 downloads11mo agoHugging Face25oumi-ai /lmsys_chat_1m_clean_R1 oumi-ai/lmsys_chat_1m_clean_R1 lmsys_chat_1m_clean_R1 is a text dataset designed to train Conversational Language Models with DeepSeek-R1 level reasoning. Prompts were pulled from LMSYS and filtered to lmsys_chat_1m_clean, and responses were taken from DeepSeek-R1 without additional filters present. We release lmsys_chat_1m_clean_R1 to help enable the community to develop the best fully open reasoning model! lmsys_chat_1m_clean queries with responses generated from… See the full description on the dataset page: https://huggingface.co/datasets/oumi-ai/lmsys_chat_1m_clean_R1.texttext-generation100K<n<1M9 likes225 downloads2y agoHugging Face26cis-lmu /GlotStoryBook Dataset Description Story Books for 180 ISO-639-3 codes. The Parallel ID or parallel_id can be used to find the parallel documents in different languages and build a parallel dataset. This dataset consists of 2 subsets: default, which consists of 4 publishers: asp: African Storybook pb: Pratham Books lcb: Little Cree Books lida: LIDA Stories nalibali, which comes from Nal'ibali stories. Usage (HF Loader) default: from datasets import load_dataset dataset… See the full description on the dataset page: https://huggingface.co/datasets/cis-lmu/GlotStoryBook.texttranslation10K<n<100K9 likes186 downloads2d agoHugging Face27meet5568 /lma_datasets LMA Phase 1 --- Hindi and Nepali pretraining corpora Two monolingual corpora built for a pair of ~25M-parameter decoder-only Transformers. Hindi is the higher-resource language, Nepali the lower-resource one. Both are written in Devanagari (U+0900-U+097F), so script cannot be used to tell them apart --- separating them is the central technical problem this dataset solves rather than assumes. language documents characters manual (chars) tokens manual (tokens) train val test… See the full description on the dataset page: https://huggingface.co/datasets/meet5568/lma_datasets.texttext-generation1M<n<10M0 likes181 downloads9d agoHugging Face28openslr /librispeech_lmLanguage modeling resources to be used in conjunction with the LibriSpeech ASR corpus.text-generation10M<n<100M2 likes176 downloads3y agoHugging Face29richarddzh /chinese-small-lm-corpus Chinese Small LM Corpus 用于中文小型语言模型预训练的统一文本语料,字段为: text:规范化后的训练文本 source:原始数据集名称 数据量 来源 有效样本数 TinyStories-Zh-2M 1,994,291 Wikipedia-20231101.zh 1,384,748 Zhihu-KOL 1,002,863 总计 4,381,902 来源与许可 RobinChen2001/TinyStories-Zh-2M:数据卡标注 MIT;同时应检查英文上游数据及机器翻译来源条款。 wikimedia/wikipedia (20231101.zh):CC BY-SA 3.0 与 GFDL。 wangrui6/Zhihu-KOL:原数据卡未声明许可证。 此合并数据集不提供统一的再授权。下载者须分别遵守各来源的许可、署名、隐私与内容使用要求。 texttext-generation1M<n<10M0 likes167 downloads2mo agoHugging Face30lmms-lab /LLaVA-OneVision-Mid-Data Dataset Card for LLaVA-OneVision Due to unknow reasons, we are unable to process dataset with large amount into required HF format. So we directly upload the json files and image folders (compressed into tar.gz files). You can use the following link to directly download and decompress them. https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data/tree/main/evol_instruct We provide the whole details of LLaVA-OneVision Dataset. In this dataset, we include the data splits… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab/LLaVA-OneVision-Mid-Data.imagetext-generation100K<n<1M21 likes153 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.