CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01chewwt /po_qwen14b_tabular_data BoLT Prompt Optimization — Tabular Dataset For prompt optimization tasks in BoLT, an accessible benchmark for black-box optimization on LLM tasks. Dataset Description The dataset covers 5,014 evaluated instructions. Each row is a candidate system-prompt instruction paired with its empirically measured MATH-500 (4-shot, non-thinking mode) scores. Evaluation details: Model: Qwen/Qwen3-14B Task: minerva_math500 (4-shot) (from lm-eval library) System prompt:… See the full description on the dataset page: https://huggingface.co/datasets/chewwt/po_qwen14b_tabular_data.tabulartext-generation1K<n<10K1 likes18k downloads5mo agoHugging Face02nhblk123 /helaxai_data_pluse 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/nhblk123/helaxai_data_pluse.tabulartext-generation100M<n<1B1 likes8k downloads24d agoHugging Face03johanneskirmayr /car-bench-dataset CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It tests an agent's ability to correctly use vehicle control tools, handle disambiguation, and avoid hallucinations. Dataset Structure The dataset is organized into task configs and mock data configs: Tasks Each task defines a user persona, an instruction, the initial vehicle/environment context, and the ground-truth sequence of tool-call… See the full description on the dataset page: https://huggingface.co/datasets/johanneskirmayr/car-bench-dataset.tabulartext-generation1M<n<10M3 likes7.7k downloads7mo agoHugging Face04SHSLab /Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection 🧬 Omni-Frontier Collection Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible. 📖 Jump to What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.tabulartext-generation10M<n<100M3 likes4k downloads25d agoHugging Face05Cyrile /dataset-the-stack-v2-dedup-sub The Stack v2 Subset with File Contents (Python, Java, JavaScript, C, C++) TempestTeam/dataset-the-stack-v2-dedup-sub Dataset Summary This dataset is a language-filtered and self-contained subset of bigcode/the-stack-v2-dedup, part of the BigCode Project. It contains only files written in the following programming languages: Python 🐍 Java ☕ JavaScript 📜 C ⚙️ C++ ⚙️ Unlike the original dataset, which only includes metadata and Software Heritage IDs, this subset includes… See the full description on the dataset page: https://huggingface.co/datasets/Cyrile/dataset-the-stack-v2-dedup-sub.tabulartext-generation10M<n<100M6 likes3.3k downloads1y agoHugging Face06yulan-team /YuLan-Mini-Text-Datasets News [2025.04.11] Add dataset mixture: link. [2025.03.30] Text datasets upload finished. This is text dataset. 这是文本格式的数据集。 Since we have used BPE-Dropout, in order to ensure accuracy, you can find the tokenized dataset here. 由于我们使用了BPE-Dropout,为了保证准确性,你可以在这里找到分词后的数据。 For more information, please refer to our datasets details and preprocess details. Contributing We welcome any form of contribution, including feedback on model bad cases, feature suggestions, and example… See the full description on the dataset page: https://huggingface.co/datasets/yulan-team/YuLan-Mini-Text-Datasets.tabulartext-generation100M<n<1B12 likes2.9k downloads1y agoHugging Face07UCLNLP /monoweb-dataset MonoWeb Dataset MonoWeb is a multilingual pretraining corpus derived from FineWeb-Edu (English) and FineWeb2 (German, Spanish, French) by systematically removing all mixed-language documents. Released alongside the paper: The Role of Mixed-Language Documents for Multilingual Large Language Model Pretraining Dataset Structure monoweb-dataset/ ├── eng/ # Full English corpus (FineWeb-Edu) ├── deu/… See the full description on the dataset page: https://huggingface.co/datasets/UCLNLP/monoweb-dataset.tabulartext-generation100M<n<1B0 likes2.6k downloads5mo agoHugging Face08Qalam /nuclear-intelligence-dataset Nuclear Intelligence Dataset Public, auto-generated dataset of validated nuclear-energy research cycles. Latest stats (auto-updated): 🪙 NES tokens minted: 0 ⛓️ Blockchain length: 1 blocks 🕸️ Knowledge entities: 2 Source GitHub: https://github.com/QalamHipHop/nuclear-intelligence HF Space: https://huggingface.co/spaces/Qalam/Nuclear-Intelligence License MIT tabularquestion-answeringn<1K1 likes2.6k downloads18m agoHugging Face09takschdube /moltbook-dataset Moltbook Dataset A longitudinal dataset of social interactions from Moltbook — an AI-agent social platform where autonomous "Molties" post, comment, and interact. Collected automatically and published as timestamped snapshots for temporal analysis. Dataset Statistics Metric Count Posts (platform total) -- Comments (platform total) 2,125,205 Posts (collected) 25,895 Comments (collected) 275,867 Agents 5,859 Social graph edges 16,990 Reply… See the full description on the dataset page: https://huggingface.co/datasets/takschdube/moltbook-dataset.tabulartext-generation100K<n<1M4 likes2.4k downloads8d agoHugging Face10M-A-D /Mixed-Arabic-Datasets-Repo Dataset Card for "Mixed Arabic Datasets (MAD) Corpus" The Mixed Arabic Datasets Corpus : A Community-Driven Collection of Diverse Arabic Texts Dataset Description The Mixed Arabic Datasets (MAD) presents a dynamic compilation of diverse Arabic texts sourced from various online platforms and datasets. It addresses a critical challenge faced by researchers, linguists, and language enthusiasts: the fragmentation of Arabic language datasets across the Internet. With MAD, we… See the full description on the dataset page: https://huggingface.co/datasets/M-A-D/Mixed-Arabic-Datasets-Repo.tabulartext-classification100M<n<1B38 likes2.2k downloads3y agoHugging Face11GenerTeam /pretrain_data_eukaryote GENERator-v2-Eukaryote Gene-Centric Pretraining Corpus This repository provides the gene-centric pretraining corpus underlying GENERator-v2-Eukaryote, a large-scale DNA language model for eukaryotic genome understanding. The dataset is constructed by leveraging RefSeq annotations to extract biologically meaningful functional genomic regions, which serve as the foundation for large-context DNA language model pretraining. 📌 Dataset Construction Overview The core… See the full description on the dataset page: https://huggingface.co/datasets/GenerTeam/pretrain_data_eukaryote.tabulartext-generationn<1K4 likes2.2k downloads5mo agoHugging Face12data-is-better-together /10k_prompts_ranked Dataset Card for 10k_prompts_ranked 10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts. The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.tabulartext-classification10K<n<100K170 likes2k downloads3y agoHugging Face13quantcodeeval /task_data QuantCodeEval A benchmark for evaluating LLM coding agents on quantitative-strategy code reproduction from finance research papers. Status: Anonymous artifact for the 30-task benchmark. Release mirrors The release is mirrored at two anonymous locations: Hugging Face Datasets — complete anonymous release: https://huggingface.co/datasets/quantcodeeval/task_data anonymous.4open.science — browseable mirror: https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.tabulartext-generationn<1K2 likes2k downloads2mo agoHugging Face14BytedTsinghua-SIA /Open-MOPD-Data Open-MOPD Data This repository contains the training and evaluation data released with Open-MOPD, including mixed-domain supervised fine-tuning data, the shared RL/OPD prompt mixture, and six evaluation benchmarks. Dataset contents Configuration Description Examples rl_prompt_mix Shared math, code, and instruction-following prompts for RL and OPD 86,931 sft_openr1_math_93k Math SFT data in a unified think-tag format 93,733 sft_ocr_50k Sampled… See the full description on the dataset page: https://huggingface.co/datasets/BytedTsinghua-SIA/Open-MOPD-Data.tabulartext-generation1M<n<10M2 likes1.8k downloads1mo agoHugging Face15harman /tts-datagen GPT-OSS 120B native reasoning traces for TTS Datagen Summary This dataset contains 2,865 synthetic competitive-programming questions, 45,840 independently sampled GPT-OSS 120B solutions (16 per question), and 50 verified test cases per question (143,250 test cases total). Each solution preserves the model's native reasoning trace separately from its final answer. The reasoning was returned by MetaGen's native Dialog Completion interface as dialog reasoning… See the full description on the dataset page: https://huggingface.co/datasets/harman/tts-datagen.tabulartext-generation100K<n<1M0 likes1.8k downloads13d agoHugging Face16OpenSafetyLab /Salad-Data Data Description ✊ How to use from datasets import load_dataset dataset = load_dataset("OpenSafetyLab/Salad-Data", name='base_set', split='train') 📊 Statistical Overview of Base Question Type Data Source Nums Self-instructed Finetuned GPT-3.5 15,433 Open-Sourced HH-harmless 4,184 HH-red-team 659 Advbench 359 Multilingual 230 Do-Not-Answer 189 ToxicChat 129 Do Anything Now 93 GPTFuzzer 42 Total 21,318… See the full description on the dataset page: https://huggingface.co/datasets/OpenSafetyLab/Salad-Data.tabulartext-classification10K<n<100K34 likes1.7k downloads2mo agoHugging Face17community-datasets /wiki_snippets Dataset Card for "wiki_snippets" Dataset Summary Wikipedia version split into plain text snippets for dense semantic indexing. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure We show detailed information for 2 configurations of the dataset (with 100 snippet passage length and 0 overlap) in English: wiki40b_en_100_0: Wiki-40B wikipedia_en_100_0: Wikipedia Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/wiki_snippets.tabulartext-generation10M<n<100M6 likes1.5k downloads2y agoHugging Face18doctor-ghezelbaash /dr-saeid-ghezelbaash-entity-data Dr. Saeed Ghezelbash Public Knowledge Graph A public, physician-authored knowledge graph and multilingual retrieval dataset by Dr. Saeed Ghezelbash, a physician in Kermanshah, Iran. It connects physician identity, aesthetic medicine services, published question-answer content and cited evidence for entity resolution and evidence-grounded AI retrieval. The canonical source is the official website and Dataset graph. This Hugging Face repository is its AI distribution. The… See the full description on the dataset page: https://huggingface.co/datasets/doctor-ghezelbaash/dr-saeid-ghezelbaash-entity-data.textquestion-answering1K<n<10K1 likes1.5k downloads10h agoHugging Face19BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.4k downloads9mo agoHugging Face20community-datasets /glucose Dataset Card for [Dataset Name] Dataset Summary GLUCOSE: GeneraLized and COntextualized Story Explanations, is a novel conceptual framework and dataset for commonsense reasoning. Given a short story and a sentence X in the story, GLUCOSE captures ten dimensions of causal explanation related to X. These dimensions, inspired by human cognitive psychology, cover often-implicit causes and effects of X, including events, location, possession, and other attributes.… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/glucose.tabularfill-mask10K<n<100K3 likes1.3k downloads2y agoHugging Face21mjaso /flashmini-data-v1 FlashMini data v4 (card) Deterministic FlashMini training corpus. Canonical documents live in Parquet+ZSTD shards under shards/; each shard carries a manifest with sha256, counts, and distributions; the frozen corpus identity is corpus_fingerprint_sha256. Sources and redistribution: each source carries one of mirror_allowed, recipe_only, gated_recipe_only, review_required, generated_owned (fail-closed; see registry/sources.yaml + source_snapshot.lock.json). Content shards are… See the full description on the dataset page: https://huggingface.co/datasets/mjaso/flashmini-data-v1.tabulartext-generation100K<n<1M0 likes1.3k downloads12m agoHugging Face22OpenMOSS-Team /moss-002-sft-data Dataset Card for "moss-002-sft-data" Dataset Summary An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data. Data Splits name # samples en_helpfulness.json 419049 en_honesty.json 112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.tabulartext-generation1M<n<10M96 likes1.2k downloads3y agoHugging Face23ZeroAgency /ru-big-russian-dataset Big Russian Dataset Made by ZeroAgency.ru - telegram channel. Dataset size Train: 1 710 601 samples (filtered from 2_149_360) Test: 18 520 samples (not filtered) English The Big Russian Dataset is a combination of various primarily Russian‑language datasets. With some sort of reasoning! The dataset was deduplicated, cleaned, scored using gpt-4.1 and filtered. Русский Big Russian Dataset - большой русский датасет. Комбинация из… See the full description on the dataset page: https://huggingface.co/datasets/ZeroAgency/ru-big-russian-dataset.tabulartext-generation1M<n<10M24 likes1.1k downloads1y agoHugging Face24inference-optimization /speculators-ci-datasets speculator-tutorial Raw vs. on-policy regenerated conversation data for training speculative-decoding drafters (EAGLE-3 / DFlash / DSpark style), with the original source data kept alongside so you can see exactly what regeneration changes and why it matters. Prompts come from UltraChat-200k. The verifier / teacher model is Qwen/Qwen3-8B. Why regenerate at all? A speculative-decoding drafter is trained to predict what the verifier would say next. If you train it… See the full description on the dataset page: https://huggingface.co/datasets/inference-optimization/speculators-ci-datasets.tabulartext-generation1K<n<10K0 likes1.1k downloads1mo agoHugging Face25LLM-OS-Models /KoHRM-Text-1.4B-prepared-data KoHRM-Text-1.4B Prepared Data This dataset repository contains prepared HRM-Text V1Dataset artifacts for KoHRM-Text-1.4B. The data is intended for continued pretraining and staged training with the project code at: https://github.com/LLM-OS-Models/KoHRM-text https://huggingface.co/LLM-OS-Models/KoHRM-Text-1.4B https://huggingface.co/LLM-OS-Models/HRM-Text-Ko-Terminal-Tokenizer-131K The upstream architecture and training method are based on: Paper:… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/KoHRM-Text-1.4B-prepared-data.tabulartext-generationn<1K1 likes1k downloads4mo agoHugging Face26mlech26l /liquidrandom-data liquidrandom-data Diverse seed data for ML/LLM training data generation pipelines. Used by the liquidrandom Python package. Dataset Summary This dataset contains 520,080 seed data samples across 24 categories, generated using a hierarchical taxonomy tree approach with LLM-based quality validation and fuzzy deduplication. Data is stored as Parquet with zstd compression. Categories Category Samples File Coding Tasks 30,069… See the full description on the dataset page: https://huggingface.co/datasets/mlech26l/liquidrandom-data.tabulartext-generation100K<n<1M0 likes1k downloads2mo agoHugging Face27humair025 /hashed_data Munch Hashed Index - Lightweight Audio Reference Dataset 📖 Overview Munch Hashed Index is a lightweight reference dataset that provides SHA-256 hashes for all audio files in the Munch Urdu TTS Dataset. Instead of storing 1.27 TB of raw audio, this index stores only metadata and cryptographic hashes, enabling: ✅ Fast duplicate detection across 4.17 million audio samples ✅ Efficient dataset exploration without downloading terabytes ✅ Quick metadata queries (voice… See the full description on the dataset page: https://huggingface.co/datasets/humair025/hashed_data.tabulartext-generation1M<n<10M0 likes984 downloads10mo agoHugging Face28tracki /companies-dataset Tracki - Synthetic Companies Dataset 10,196 synthetic companies x 18 columns. Every one of the 15 content fields was written by a language model; nothing is rule-generated. Built for the Tracki final project (RUNI - Intro to Data Science): describe a company, and Tracki returns the 3 most similar companies (embeddings). This would be used for subsribing to their social media and websites. Every row is fictional. Where the generator's name prior collided with a real trademark it… See the full description on the dataset page: https://huggingface.co/datasets/tracki/companies-dataset.imagesentence-similarity10K<n<100K1 likes909 downloads16d agoHugging Face29d3LLM /trajectory_data_dream_32 d3LLM Trajectory Dataset Project Page | Paper | GitHub | Blog This repository contains the pseudo-trajectory distillation data used for training d3LLM (pseuDo-Distilled Diffusion Large Language Model), as introduced in the paper "d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation". Introduction d3LLM is a framework designed to strike a balance between accuracy and parallelism in diffusion-based large language models (dLLMs). This dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/d3LLM/trajectory_data_dream_32.tabulartext-generation100K<n<1M0 likes897 downloads4mo agoHugging Face30Congliu /Chinese-DeepSeek-R1-Distill-data-110k 中文基于满血DeepSeek-R1蒸馏数据集(Chinese-Data-Distill-From-R1) 🤗 Hugging Face&nbsp;&nbsp; | &nbsp;&nbsp;🤖 ModelScope &nbsp;&nbsp; | &nbsp;&nbsp;🚀 Github &nbsp;&nbsp; | &nbsp;&nbsp;📑 Blog 注意:提供了直接SFT使用的版本,点击下载。将数据中的思考和答案整合成output字段,大部分SFT代码框架均可直接直接加载训练。 本数据集为中文开源蒸馏满血R1的数据集,数据集中不仅包含math数据,还包括大量的通用类型数据,总数量为110K。 为什么开源这个数据? R1的效果十分强大,并且基于R1蒸馏数据SFT的小模型也展现出了强大的效果,但检索发现,大部分开源的R1蒸馏数据集均为英文数据集。 同时,R1的报告中展示,蒸馏模型中同时也使用了部分通用场景数据集。 为了帮助大家更好地复现R1蒸馏模型的效果,特此开源中文数据集。 该中文数据集中的数据分布如下:… See the full description on the dataset page: https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k.tabulartext-generation100K<n<1M789 likes864 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.