CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolmino-mix-1124 DOLMino dataset mix for OLMo2 stage 2 annealing training. Mixture of high-quality data used for the second stage of OLMo2 training. Source Sizes Name Category Tokens Bytes (uncompressed) Documents License DCLM HQ Web Pages 752B 4.56TB 606M CC-BY-4.0 Flan HQ Web Pages 17.0B 98.2GB 57.3M ODC-BY Pes2o STEM Papers 58.6B 413GB 38.8M ODC-BY Wiki Encyclopedic 3.7B 16.2GB 6.17M ODC-BY StackExchange CodeText 1.26B 7.72GB 2.48M CC-BY-SA-{2.5, 3.0, 4.0}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolmino-mix-1124.tabulartext-generation100M<n<1B102 likes24k downloads11mo agoHugging Face02allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes8k downloads8mo agoHugging Face03allenai /tulu-2.5-preference-data Tulu 2.5 Preference Data This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. We cleaned and formatted all datasets to be in the same format. This means some splits may differ from their original format. To see the code used for creating most splits, see here. If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.texttext-generation1M<n<10M18 likes1.7k downloads2y agoHugging Face04FreedomIntelligence /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V ├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.imagequestion-answering100K<n<1M98 likes829 downloads1y agoHugging Face05FINAL-Bench /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.imagetext-generationn<1K26 likes811 downloads7mo agoHugging Face06AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes583 downloads1y agoHugging Face07allenai /tutormoments-preview TutorMoments-Preview 462 real K–12 math tutoring sessions (student and tutor) with human annotations, plus a benchmark of 7,280 AI-tutor attempts scored the same way. A preview release from TutorMoments, a project on how well tutors — human and AI — scaffold, push for rigor, and build rapport. From one K–12 tutoring program (anonymized as tutoring_provider_a). Paper: When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle Code:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tutormoments-preview.texttext-classification10K<n<100K6 likes573 downloads2mo agoHugging Face08youssef3146 /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.imagetext-generationn<1K0 likes482 downloads7mo agoHugging Face09Allen-UQ /CNY-data CNY walk-ready graph reasoning data Walk-ready data for CNY (Call Neighbours Yourself), the graph-walk reinforcement learning framework described in Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation (EMNLP 2026), arXiv:2608.29588. Models: Allen-UQ/CNY-7B, Allen-UQ/CNY-14B. This repository does not redistribute any graph dataset in its original form. It provides the derived, prompt-rendered form in which each graph instance is… See the full description on the dataset page: https://huggingface.co/datasets/Allen-UQ/CNY-data.texttext-generation100K<n<1M0 likes435 downloads24d agoHugging Face10allenai /commongen_lite CommonGen-Lite Evaluating LLMs with CommonGen using CommonGen-lite dataset (400 examples + 900 human references). We use GPT-4 to evaluate the constrained text generation ability of LLMs. Please see more in our paper. Github: https://github.com/allenai/CommonGen-Eval Leaderboard model len cover pos win_tie overall human 12.84 99.00 98.11 100.00 97.13 gpt-4-0613 14.13 97.44 91.78 50.44 45.11 gpt-4-1106-preview14.90 96.33 90.11 50.78 44.08… See the full description on the dataset page: https://huggingface.co/datasets/allenai/commongen_lite.texttext-generationn<1K7 likes285 downloads3y agoHugging Face11agentlans /allenai-WildChat AllenAI WildChat Combined Dataset This unofficial repository provides the AllenAI WildChat Combined Dataset, which merges the WildChat-4.8M and WildChat-1M collections of human–ChatGPT conversations. WildChat-1M contains 1 million chats, of which 25.53% are from GPT‑4 and the remainder from GPT‑3.5. These conversations cover a wide range of complex interactions, including code-switching, ambiguity, and political topics. WildChat-4.8M originally comprised 4.8 million conversations.… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat.texttext-generation1M<n<10M3 likes245 downloads9mo agoHugging Face12lodestones /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V… See the full description on the dataset page: https://huggingface.co/datasets/lodestones/ALLaVA-4V.imagequestion-answering1M<n<10M0 likes241 downloads2y agoHugging Face13agentlans /allenai-WildChat-4.8Mtexttext-generation1M<n<10M1 likes233 downloads1y agoHugging Face14Trendyol /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M32 likes166 downloads1y agoHugging Face15agentlans /allenai-dolma3_mix-6T-sampletexttext-generation100K<n<1M0 likes150 downloads3mo agoHugging Face16allenai /tulu-v2-sft-mixture-olmo-2048 Dataset Card for Tulu V2 Mix (2048 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train OLMo-Instruct. The two primary differences are: long conversations are resplit into 2048-token chunks, and the hardcoded subset has been replaced with similar examples about… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-2048.textquestion-answering100K<n<1M5 likes120 downloads2y agoHugging Face17LARK-Lab /EnvFactory-SFT-ALL EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL ## Overview EnvFactory-SFT-ALL is the complete supervised fine-tuning (SFT) dataset containing 26,500 tool-use trajectories synthesized using the EnvFactory framework. This dataset includes all generated trajectories before filtering. The dataset contains multi-turn tool-use trajectories with implicit human reasoning, generated through… See the full description on the dataset page: https://huggingface.co/datasets/LARK-Lab/EnvFactory-SFT-ALL.texttext-generation10K<n<100K0 likes114 downloads4mo agoHugging Face18agentlans /allenai-WildChat-4.8M-prompts allenai/WildChat-4.8M English Prompts Dataset Summary This dataset contains real user-submitted prompts to ChatGPT, extracted from the English portion of the allenai/WildChat-4.8M collection. It serves as a large-scale resource for analyzing user intent, conversational diversity, and prompt engineering patterns. Files en_prompts: All English-language first messages from user conversations. Each record represents the first user prompt. Exact duplicates are… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-WildChat-4.8M-prompts.texttext-generation1M<n<10M0 likes106 downloads11mo agoHugging Face19alliedtoasters /forbidden-backrooms-gemma-4-31B-it Forbidden Backrooms: Gemma-4 31B Self-Chat Self-chat transcripts and per-message embeddings for two role-inverted instances of Gemma-4-31B-it, comparing the official instruct checkpoint against an abliterated fine-tune of the same checkpoint. Both variants use identical int4 quantization served via Ollama, so quantization noise is not a confound between them. The methodology follows Anthropic's Claude Opus 4 system card section on the "spiritual bliss attractor state." Leave two… See the full description on the dataset page: https://huggingface.co/datasets/alliedtoasters/forbidden-backrooms-gemma-4-31B-it.tabulartext-generation10K<n<100K0 likes99 downloads5mo agoHugging Face20FreedomIntelligence /ALLaVA-4V-Chinese ALLaVA-4V for Chinese This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.imagequestion-answering100K<n<1M16 likes88 downloads2y agoHugging Face21ansulev /all-cve-chat-multiturn-1999-2025 CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/all-cve-chat-multiturn-1999-2025.texttext-generation100K<n<1M2 likes66 downloads7mo agoHugging Face22allenai /tulu-v2-sft-mixture-olmo-4096 Dataset Card for Tulu V2 Mix (4096 OLMo version) Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. This is a modified version of the Tulu V2 Mix used to train newer (after April 2024) OLMo-SFT/Instruct variants (e.g. this model, or this one). The only difference is that the hardcoded subset (dataset='hard_coded') has been replaced… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture-olmo-4096.textquestion-answering100K<n<1M0 likes65 downloads2y agoHugging Face23PJMixers-Dev /allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT PJMixers-Dev/allenai_WildChat-1M-prompts with responses generated with gemini-2.0-flash-thinking-exp-1219. Generation Details If BlockedPromptException, StopCandidateException, or InvalidArgument was returned, the sample was skipped. If ["candidates"][0]["safety_ratings"] == "SAFETY" the sample was skipped. If ["candidates"][0]["finish_reason"] != 1 the sample was skipped. model =… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT.texttext-generation10K<n<100K6 likes61 downloads2y agoHugging Face24haiderkamal23 /allaM-offsec-arabic-chat-v2 Arabic Offensive Security Chat Dataset v2 High-quality category-aware bilingual Arabic/English dataset for offensive security assistants. What's New in v2 ✅ Category-aware responses: Different response structures for web vulns, DeFi, reconnaissance tools, social engineering, etc. ✅ No generic templates: Each category has specialized analysis framework ✅ No verbatim copying: Responses analyze and transform the input, not repeat it ✅ Semantic accuracy: Tools (nmap… See the full description on the dataset page: https://huggingface.co/datasets/haiderkamal23/allaM-offsec-arabic-chat-v2.textquestion-answering10K<n<100K0 likes56 downloads10mo agoHugging Face25allenai /tulu-v2-sft-long-mixtureThis is a recreation of the tulu-v2-sft-mixture, without splitting ShareGPT dataset into chunks of max 4096 tokens. This might be interesting to people who are doing long-context finetuning. Please refer to the original tulu-v2-sft-mixture for the details of this dataset mixture. License We are releasing this dataset under the terms of ODC-BY. By using this, you are also bound by the Common Crawl terms of use in respect of the content contained in the dataset. texttext-generation100K<n<1M7 likes53 downloads3y agoHugging Face26ChipHolmes /All-CVE-Records-Training-Dataset-archive CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/All-CVE-Records-Training-Dataset-archive.texttext-generation100K<n<1M1 likes53 downloads2mo agoHugging Face27FreedomIntelligence /ALLaVA-4V-Arabic ALLaVA-4V for Arabic This is the Arabic version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Arabic through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of Hong… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Arabic.imagequestion-answering100K<n<1M4 likes48 downloads2y agoHugging Face28for-all-dev /CompCert-eval CompCert Proof-Engineering Eval Proof-synthesis challenges mined from the git history of AbsInt/CompCert, the formally verified C compiler. Each challenge is a real proof-engineering edit that a human made in a single commit: we take the repository state before the commit (the challenge) and treat the state after the commit (the solution) as ground truth. The model's job is to reconstruct the proof/spec work the human did. ⚠️ License notice. CompCert is distributed under the… See the full description on the dataset page: https://huggingface.co/datasets/for-all-dev/CompCert-eval.texttext-generation1K<n<10K0 likes45 downloads3mo agoHugging Face29agentlans /allenai-soda SODA chat dataset This is the SODA dataset in ShareGPT-like format. According to the dataset's creators: "🥤SODA is the first publicly available, million-scale, high-quality dialogue dataset covering a wide range of social interactions." The following changes were made to the data: kept only dialogues with two people alternating turns the SODA narrative was adapted into a system prompt for an AI to roleplay as the second person extra turns were removed so that each conversation… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/allenai-soda.textfeature-extraction1M<n<10M5 likes43 downloads2y agoHugging Face30ukcli /All-CVE-Chat-MultiTurn-1999-2025-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/ukcli/All-CVE-Chat-MultiTurn-1999-2025-Dataset.texttext-generation100K<n<1M1 likes43 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.