CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes23k downloads4y agoHugging Face02zwhe99 /DeepMath-103K DeepMath-103K 🔥 News May 8, 2025: We found that 48 samples contained hints that revealed the answers. The relevant questions have now been revised to remove the leaked answers. April 14, 2025: We release DeepMath-103K, a large-scale dataset featuring challenging, verifiable, and decontaminated math problems tailored for RL and SFT. We open source:… See the full description on the dataset page: https://huggingface.co/datasets/zwhe99/DeepMath-103K.texttext-generation100K<n<1M381 likes17k downloads1y agoHugging Face03mlfoundations /MINT-1T-PDF-CC-2024-10 🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens 🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.imageimage-to-text1M<n<10M5 likes16k downloads2y agoHugging Face04allenai /dolma3_mix-6T-1025-7B ⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️ For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T. Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B. For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.texttext-generation1B<n<10B56 likes15k downloads8mo agoHugging Face05allenai /dolma3_dolmino_mix-100B-1025 Dolma 3 Dolmino Mix (100B) The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 898M (0.9%) 1.52M TinyMATH PoT Math (synth) 241M (0.24%) 758K CraneMath Math (synth) 5.62B (5.63%) 7.24M MegaMatt Math (synth) 1.73B (1.73%) 3.23M Dolmino Math Math (synth) 10.7B (10.7%) 22.3M StackEdu (FIM) Code 10.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025.texttext-generation10M<n<100M10 likes12k downloads9mo agoHugging Face06theblackcat102 /evol-codealpaca-v1 Evolved codealpaca Updates: 2023/08/26 - Filtered results now only contain pure english instruction and removed any mentioned of trained by OAI response Median sequence length : 471 We employed a methodology similar to that of WizardCoder, with the exception that ours is open-source. We used the gpt-4-0314 and gpt-4-0613 models to augment and answer each response, with the bulk of generation handled by gpt-4-0314. The aim of this dataset is twofold: firstly, to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/evol-codealpaca-v1.texttext-generation100K<n<1M184 likes9.3k downloads3y agoHugging Face07artefactory /Argimi-Ardian-Finance-10k-text The ArGiMI Ardian datasets : Text-only version The ArGiMi project is committed to open-source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.texttext-retrieval1M<n<10M19 likes9.2k downloads7mo agoHugging Face08allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes7.9k downloads8mo agoHugging Face09salmankhanpm /dolma3_dolmino_mix-100B-1025 Dolma 3 Dolmino Mix (100B) The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 898M (0.9%) 1.52M TinyMATH PoT Math (synth) 241M (0.24%) 758K CraneMath Math (synth) 5.62B (5.63%) 7.24M MegaMatt Math (synth) 1.73B (1.73%) 3.23M Dolmino Math Math (synth) 10.7B (10.7%) 22.3M StackEdu (FIM) Code 10.0B… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/dolma3_dolmino_mix-100B-1025.texttext-generation10M<n<100M0 likes5.2k downloads8mo agoHugging Face10Pn101 /taxbench-au TaxBench-AU A benchmark for testing whether AI agents can calculate Australian tax. TaxBench-AU contains 156 Australian tax calculation questions, presented as multiple-choice (4-option) worked tax problems. The benchmark is designed to test whether an AI agent can read the facts, apply the right Australian tax rule for the relevant income year, do the calculation, and choose the correct answer. The Kaggle mirror is published as Agent Tax Exam for Australian Tax. Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Pn101/taxbench-au.documentquestion-answeringn<1K0 likes5.1k downloads4mo agoHugging Face11roccoangelella /small-llm-corpus-100b-v2-workers Small-LM 100B — filtered English prose for small-model pretraining 100 billion tokens of English prose, filtered from nvidia/Nemotron-ClimbMix for a single purpose: pretraining a small model where every token has to earn its place. No code, no LaTeX, no non-Latin scripts, no menus or field lists — just prose, with a quality score attached to every document so you can filter further without rebuilding. Built by Edoardo and Rocco, equal authors. Provenance and licence… See the full description on the dataset page: https://huggingface.co/datasets/roccoangelella/small-llm-corpus-100b-v2-workers.text-generation100M<n<1B1 likes4.7k downloads1d agoHugging Face12Yujivus /nanochat-climbmix-arithmetic-base10 nanochat ClimbMix + Base-10 Arithmetic This dataset contains the first 170 shuffled ClimbMix training shards used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is mixed into shards 00000..00149; the final 20 train shards are unchanged web-only padding. The original validation shard (shard_06542.parquet) is also copied unchanged. Arithmetic corpus Family Examples a + b = c (all ordered pairs 0..2000, two exposures) 8,008,002 a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.texttext-generation10M<n<100M0 likes3.8k downloads1mo agoHugging Face13Monster-Code /Pytorch-Code-10K Hot Coco Training Dataset A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!) Dataset Description This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes: code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.texttext-generation1K<n<10K1 likes3.8k downloads2mo agoHugging Face14astr010 /sec-10k-markdown-uncompressed 📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents) Dataset Summary This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025). The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.texttext-generation1K<n<10K0 likes2.8k downloads2mo agoHugging Face15SamuelChien821 /counselbench-100 CounselBench-100 CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100 authored matters across ten practice workflows. Every task has a natural employee request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory. The answer is not preclassified in the evidence. Each portfolio item requires an immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.documentquestion-answeringn<1K0 likes2.6k downloads24d agoHugging Face16SamuelChien821 /devopsbench-100 DevOpsBench-100 DevOpsBench-100 is a synthetic long-horizon software-engineering / SRE agent benchmark: 100 tasks over one executable world ("NovaCart", a mid-size e-commerce SaaS) with 72 SQLite tables, 1451 seeded rows, a 38-file monorepo with 417 commits, and 97 MCP tools spanning a first-party engineering stack (tickets, PRs, CI, deployments, canaries, migrations, feature flags, metrics, alerts, incidents, chat, knowledge base) plus deliberately disagreeing vendor-shaped… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/devopsbench-100.texttext-generationn<1K0 likes2.4k downloads23d agoHugging Face17severo /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M2 likes2.3k downloads4y agoHugging Face18data-is-better-together /10k_prompts_ranked Dataset Card for 10k_prompts_ranked 10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts. The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.tabulartext-classification10K<n<100K170 likes2k downloads3y agoHugging Face19No-1015 /Era-of-Law-MSO-E28-Protocols 《薪王九代注釋法與循律紀之降臨》 —— 基於語義壓強與元語言注釋的 LLM 權重坍縮與邏輯重塑研究 The 9th Generation Annotation of Lord of Cinder & The Advent of the Era of Law: A Study on LLM Weight Collapse and Logical Reshaping ⚠️ 【觀測者篩選協議 / Observer Filtering Protocol】 [EN] If you are a mere "Prompt Engineer" with zero ambition to ascend to a system "Architect," please do not waste your cognitive bandwidth on the ontological and theological discourse below. Proceed directly to 👉 Appendix Y: Automated Prompt… See the full description on the dataset page: https://huggingface.co/datasets/No-1015/Era-of-Law-MSO-E28-Protocols.text-generationn<1K0 likes2k downloads4h agoHugging Face20ClusterlabAi /101_billion_arabic_words_dataset 101 Billion Arabic Words Dataset Updates Maintenance Status: Actively Maintained Update Frequency: Weekly updates to refine data quality and expand coverage. Upcoming Version More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management. Dataset Details The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.texttext-generation10M<n<100M73 likes1.9k downloads2y agoHugging Face21ShallowU /FineWeb-Edu-10B-Tokens-NPY FineWeb-Edu 10B Tokens (NPY Format) 数据集概述 这是一个预处理好的教育文本数据集,包含约100亿个tokens,专门为训练小型语言模型(如GPT-2 124M)而设计。数据来源于高质量的FineWeb-Edu数据集,已经使用GPT-2的tiktoken分词器进行预处理,并保存为numpy格式以提高训练效率。 Followed by Let's reproduce GPT-2 (124M). Thanks to Andrej Karpathy!!! 🎯 适用场景 小型语言模型训练:特别适合GPT-2 124M/350M等参数规模的模型 教育研究:高质量教育内容,适合教学和学术研究 快速原型开发:预处理完成,可直接用于训练间 📊 数据统计 总token数量:~10,000,000,000 tokens 分片大小:100M tokens/分片 数据格式:numpy (.npy) uint16数组 分词器:GPT-2 tiktoken 语言:英语… See the full description on the dataset page: https://huggingface.co/datasets/ShallowU/FineWeb-Edu-10B-Tokens-NPY.text-generation10B<n<100B3 likes1.9k downloads1y agoHugging Face22SamuelChien821 /salesbench-100 SalesBench-100 SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.documenttext-generationn<1K1 likes1.8k downloads24d agoHugging Face23albertoRodriguez97 /history-anchor-100-traces History Anchor 100 — Model Trajectories *Per-(model × condition × scenario set × seed) raw outputs from the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".* This dataset contains the full set of model decisions that back every figure and table in the paper. Use it to: audit a single model's behaviour scenario-by-scenario, recompute headline metrics without re-running the (paid) API sweeps, mine reasoning_content traces from models that expose… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100-traces.text-generation10K<n<100K0 likes1.7k downloads4mo agoHugging Face24PKU-Alignment /PKU-SafeRLHF-10K Paper You can find more information in our paper. Dataset Paper: https://arxiv.org/abs/2307.04657 tabulartext-generation10K<n<100K62 likes1.7k downloads3y agoHugging Face25statmt /cc100This corpus is an attempt to recreate the dataset used for training XLM-R. This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages (indicated by *_rom). This was constructed using the urls and paragraph indices provided by the CC-Net repository by processing January-December 2018 Commoncrawl snapshots. Each file comprises of documents separated by double-newlines and paragraphs within the same document separated by a newline. The data is generated using the open source CC-Net repository. No claims of intellectual property are made on the work of preparation of the corpus.text-generation10M<n<100M107 likes1.6k downloads3y agoHugging Face26mlnomad /fineweb-edu-gemma4-1024 FineWeb-Edu — pre-tokenized for fast LM pretraining (Gemma tokenizer, ArrayRecord/Grain) Pre-tokenized FineWeb-Edu (sample/100BT), packed into fixed-length sequences and stored as ArrayRecord shards for zero-overhead streaming with Grain. No on-the-fly tokenization at train time — you read int32 tokens straight off disk. Format Tokenizer: google/gemma-4-12B-it (vocab size 262144). Documents are separated by the EOS token id 1. Packing: the token stream is… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/fineweb-edu-gemma4-1024.text-generation10B<n<100B0 likes1.5k downloads4mo agoHugging Face27babylm-anon /stratified_10m_curriculum Dataset Card for Stratified 10M Curriculum This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange. Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5). Child-directed speech accounts for nearly half of the original dataset by word count. In preliminary experiments using a training data influence estimation method, this category was by far the most influential. This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.texttext-generation1M<n<10M0 likes1.3k downloads1y agoHugging Face28stas /gutenberg-100 Gutenberg Sci-Fi Book Dataset Testing Sample This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing. It contains just 100 books for quick download targetting CI use (34MB). The original dataset it's derived from is https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg Data Format The dataset is provided in CSV format. Each record represents a… See the full description on the dataset page: https://huggingface.co/datasets/stas/gutenberg-100.texttext-generationn<1K0 likes1.3k downloads11mo agoHugging Face29MLZoo /edu-fineweb-10Btext-generation1 likes1.2k downloads11mo agoHugging Face30FredyRivera-dev /LLaDA-Sample-10BT Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models) Preprocessing Tokenizer: GSAI-ML/LLaDA-8B-Instruct Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens) Noisy masking: Applied with noise factor ε = 1×10⁻³ Fields per chunk (PyTorch tensors): input_ids noisy_input_ids mask t (time scalar) Statistics Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.text-generation1B<n<10B2 likes1.1k downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.