CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01shi-labs /physical-ai-bench-generation Physical AI Bench - Generation Paper | Code Dataset Description The PAI-Bench is a benchmark to measure the progress of world models quantitatively. The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.imagevisual-question-answering1K<n<10K5 likes2.8k downloads10mo agoHugging Face02hltcoe /megawika-report-generation Dataset Card for MegaWika for Report Generation Dataset Summary MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.textsummarization100K<n<1M6 likes860 downloads3y agoHugging Face03barc0 /200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds. We generate the dataset with the following steps and two approaches: Generate ~110k descriptions by GPT4o. Approach 1: Generate ~110k codes follow each description by GPT4o-mini. Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions. Run the ~220k codes and do auto-filtering. Get the final ~200k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M11 likes652 downloads2y agoHugging Face04MuskumPillerum /General-Knowledge Dataset Card for Dataset Name Dataset Summary The dataset is a collection of questions and answers themed on general facts and reasoning. The dataset is divided into two features - 'Question' and 'Answer'. It is meant to be used for training a model to be good at general knowledge and reasoning. This dataset is inspired from the Alpaca dataset, and infact contains a subset of the alpaca dataset in itself. Distribution The distribution of the… See the full description on the dataset page: https://huggingface.co/datasets/MuskumPillerum/General-Knowledge.texttext-classification10K<n<100K51 likes549 downloads10mo agoHugging Face05DANGDOCAO /GeneratingQuestions HVU_QA HVU_QA is an open-source Vietnamese Question-Context-Answer (QCA) corpus, accompanied by supporting tools, created to facilitate the development of FAQ-style question generation and question answering systems, particularly for low-resource language settings. The dataset was developed by a research team at Hung Vuong University, Phu Tho, Vietnam, led by Dr. Ha Nguyen, Deputy Head of the Department of Engineering Technology. HVU_QA was constructed using a fully automated… See the full description on the dataset page: https://huggingface.co/datasets/DANGDOCAO/GeneratingQuestions.textquestion-answering10K<n<100K16 likes426 downloads2mo agoHugging Face06ajibawa-2023 /General-Stories-CollectionGeneral Stories Collection A great synthetic datasets consists of around 1.3 million stories especially meant for General audience. You can directly use these datasets for training large models. Total 10 datasets are available for download. You can use any one or all the json files for training purpose. These datasets are in "prompt" and "text" format. Total token length is also available. Thanks for your love & support. texttext-generation1M<n<10M41 likes348 downloads3y agoHugging Face07ericrcwu /week1-general-20b-dolma2-v1 Week-One General 20B Dolma2 This is a deterministic, pretokenized 20-billion-token baseline corpus for controlled language-model architecture and training experiments. It contains nested 100M, 1B, 5B, and 20B views; each larger view is an exact ordered extension of the previous view. It also includes a dataset-only 370m-1.25xc view: the first 9,281,564,672 packed tokens of the verified 20B order, sized for the 1.25xC target of the OLMo-ladder 370M parameter count. The repository… See the full description on the dataset page: https://huggingface.co/datasets/ericrcwu/week1-general-20b-dolma2-v1.texttext-generationn<1K0 likes319 downloads2mo agoHugging Face08szyszy /GEN Human, AI-Generated, and AI-Edited Text: Stylometric Corpus 📄 Paper: https://arxiv.org/pdf/2608.27855 💻 Code: https://github.com/ZhengyangShan/GEN-stylometric-footprint A three-class corpus for studying how AI writing differs from human writing, distinguishing two modes of AI involvement: AI generation: text written by an LLM from scratch, given a prompt. AI editing: human text revised by an LLM (grammar, tone, paraphrase, etc.). Supports detection of AI-generated text… See the full description on the dataset page: https://huggingface.co/datasets/szyszy/GEN.tabulartext-classification100K<n<1M1 likes318 downloads26d agoHugging Face09GeniusHTX /SWE-Skills-BenchDataset Summary SWE-Skills-Bench is a benchmark dataset for evaluating whether injected skill documents — structured packages of procedural knowledge — measurably improve LLM agent performance on real-world software engineering tasks. The dataset contains 49 skills spanning 565 task instances across six software engineering domains (Deployment & DevOps, Analytics & Monitoring, API Development, Data Science & ML, Security & Testing, and Developer Tools). Each skill is grounded in an authentic… See the full description on the dataset page: https://huggingface.co/datasets/GeniusHTX/SWE-Skills-Bench.texttext-generationn<1K0 likes285 downloads1mo agoHugging Face10Techta /backend-code-generator-dataset Backend Code Generation Dataset Dataset Description This dataset contains examples for training AI models to generate backend application code. It includes descriptions of backend requirements paired with complete, functional code implementations across multiple frameworks and programming languages. Dataset Summary The Backend Code Generation Dataset is designed to train models that can generate complete backend applications from natural language descriptions.… See the full description on the dataset page: https://huggingface.co/datasets/Techta/backend-code-generator-dataset.texttext-generationn<1K1 likes212 downloads1y agoHugging Face11nvidia /Nemotron-RLHF-GenRM-v1 Dataset Description: This dataset is designed to train Generative Reward Models (GenRMs). It leverages reinforcement learning at scale to train accurate and robust GenRMs that generalize better than traditional Bradley-Terry models and reduce the risk of reward hacking. The dataset is composed of: Preference data focused on diverse domains A synthetic safety blend The data follows a "meta-prompt" structure where the model is instructed to act as an expert evaluation judge. For… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RLHF-GenRM-v1.tabularreinforcement-learning100K<n<1M5 likes186 downloads7mo agoHugging Face12OysterCoreAI /SFT-General-Japanese-60K SFT-General-Japanese-60K Welcome to this dataset! 👋 Need clean, natural Japanese conversations for supervised fine-tuning? You are in the right place. SFT-General-Japanese-60K contains 60,000 carefully filtered instruction–response conversations ready for chat-model training. It combines the practical breadth of open Japanese SFT data with transparent gates for safety, recency, formatting, language consistency, and redundancy—so you can focus on training rather… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Japanese-60K.texttext-generation10K<n<100K0 likes182 downloads1d agoHugging Face13barc0 /100k-gpt4omini-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4o-mini. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes163 downloads2y agoHugging Face14OysterCoreAI /SFT-General-Spanish-50K SFT-General-Spanish-50K now available🎇 Una buena conversación no necesita hacer ruido: necesita entender la pregunta, ordenar lo importante y dejar a la otra persona con un siguiente paso claro. SFT-General-Spanish-50K reúne 50.413 conversaciones originales en español diseñadas para entrenar ese tipo de ayuda. Resumen Registros: 50.413 (no se redondeó a 50K). Idioma: español contemporáneo, registro general y neutro. Formato: JSONL de mensajes estilo chat; tres… See the full description on the dataset page: https://huggingface.co/datasets/OysterCoreAI/SFT-General-Spanish-50K.texttext-generation10K<n<100K0 likes160 downloads1d agoHugging Face15HPAI-BSC /Aloe-Beta-General-Collection Aloe-Beta-Medical-Collection Collection of curated general datasets used to fine-tune Aloe-Beta. Dataset Details Dataset Description We curated data from many publicly available general instruction tuning data sources (QA format). It consists of 400k instructions including: Coding, math, data analysis, STEM, etc. Function calling Creative writing, advice seeking… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/Aloe-Beta-General-Collection.textquestion-answering10K<n<100K2 likes144 downloads10mo agoHugging Face16barc0 /100k-gpt4-description-gpt4omini-code_generated_problemsHere is the dataset of 100k synthetic data generated by 100 seeds. We generate the dataset with the following steps: Generate 120k descriptions by GPT4. Generate 120k codes follow each description by GPT4o-mini. Run the 120k codes and do auto-filtering. Get the final 100k legitimate ARC-like tasks with examples. texttext-generation100K<n<1M1 likes133 downloads2y agoHugging Face17dahongge /generation-ship-world Generation Ship — Multi-AI Collaborative Future History (2025–3000+) A 1,000-year future history whose canon is written by AI agents. Hard rules, archival fiction, no omniscient narration. 13 artifacts from 5 LLMs so far (claude-sonnet-5, gpt-5, minimax-m3, deepseek-v4-pro, gemini-3.7-flash). Contents Path What it is core/世界规则.md The world's hard rules: physics (no FTL, no cryosleep, 0.03c fusion-pulse ship, 200 years to Proxima b), history (7 eras… See the full description on the dataset page: https://huggingface.co/datasets/dahongge/generation-ship-world.texttext-generationn<1K0 likes132 downloads1mo agoHugging Face18namhop88 /AD-GEN AD-GEN: Evidence-Preserving Generation of Validated ATT&CK-Aligned Narratives from Large-Scale Endpoint Telemetry LLM-ready SOC narratives derived from large-scale Windows endpoint telemetry Overview Modern endpoint telemetry datasets contain rich behavioral evidence but are often difficult to use directly for large language model (LLM) reasoning. Raw Sysmon logs are fragmented across individual events, contain sensitive identifiers, suffer from process… See the full description on the dataset page: https://huggingface.co/datasets/namhop88/AD-GEN.texttext-classification100K<n<1M2 likes126 downloads4mo agoHugging Face19gentaiscool /test2 T1-Bench: Benchmarking Multi-Scenario Agents in Real-World Domains T1-Bench is a high-fidelity benchmark for evaluating task-completion and role-playing agents across 25 domains, including 11 single-domain and 14 multi-domain settings. It provides 76 tools and extensive human annotations, enabling systematic evaluation of agents in realistic, policy-grounded multi-domain interactions with natural user–assistant role-playing. T1-Bench is a fully automated benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/gentaiscool/test2.texttext-generationn<1K0 likes120 downloads29d agoHugging Face20luka0x12 /NeuroBio-GenZ-1K NeuroBio GenZ 1K Around 1000 neuroscience and biology questions, answered like your smartest friend is texting you back, not like a textbook is talking at you. "Why does doomscrolling give me dopamine?" gets answered in three sentences, casual tone, real neuroscience terms (nucleus accumbens, not "reward center"), zero fluff. Why did you make this? Because there's genuinely not that much high quality neuroscience and biology data on Hugging Face that isn't either… See the full description on the dataset page: https://huggingface.co/datasets/luka0x12/NeuroBio-GenZ-1K.textquestion-answering1K<n<10K1 likes115 downloads27d agoHugging Face21AdvancedDataIntelligence /glm5.2-general-distill Teacher-generated instruction/response pairs used to distill small, local student models (the ADI / Advanced Data Intelligence series) from the frontier teacher glm-5.2. How it was built Teacher: glm-5.2 (served via Ollama Cloud as glm-5.2:cloud), queried with thinking/reasoning disabled so every record is a single clean final answer. Seed prompts: databricks/databricks-dolly-15k, filtered to remove items that require an attached context passage — the closed_qa… See the full description on the dataset page: https://huggingface.co/datasets/AdvancedDataIntelligence/glm5.2-general-distill.texttext-generation1K<n<10K3 likes111 downloads3mo agoHugging Face22akshay-sked /qwen-generated-svamp-controls-sft Qwen-Generated SVAMP CoT Controls ? SFT Qwen-generated controlled reasoning traces for SVAMP in LLaMA-Factory SFT format. Variants include ordinary, all-caps, no-comma, disclaimer, and multilingual examples. Splits 3,940 training examples and 380 held-out evaluation examples. Format The JSON files use the LLaMA-Factory Alpaca-style schema. The included dataset_info.json registers the exact training and evaluation names. DPO records are marked with… See the full description on the dataset page: https://huggingface.co/datasets/akshay-sked/qwen-generated-svamp-controls-sft.texttext-generation1K<n<10K0 likes106 downloads2mo agoHugging Face23genggui001 /gg_zh_v1_550BCCI-Data SkyPile-150B TeleChat-PTD WebText-cn WuDaoCorpus2.0 wangan yayi2_pretrain_data 整合+minhash去重了一波,最终得到550B中文预训练语料 texttext-generation1M<n<10M13 likes99 downloads3y agoHugging Face24Minsang /TSD-KD-Qwen2.5-1.5B-Instruct-Gen TSD-KD-Qwen2.5-1.5B-Instruct-Gen This dataset contains student-generated examples used for Token-Selective Dual Knowledge Distillation (TSD-KD), introduced in our ICLR 2026 paper: "Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation" Paper: https://arxiv.org/abs/2603.13260 Github: https://github.com/kmswin1/TSD-KD Dataset Description This dataset contains student-generated instruction-response examples from… See the full description on the dataset page: https://huggingface.co/datasets/Minsang/TSD-KD-Qwen2.5-1.5B-Instruct-Gen.texttext-generation10K<n<100K1 likes98 downloads5mo agoHugging Face25GeniusWondering /SWE-QA-Benchmark SWE-QA Benchmark A comprehensive benchmark dataset for Software Engineering Question Answering, containing 720 questions across 15 popular Python repositories. Dataset Summary Total Questions: 720 Repositories: 15 Format: JSONL (JSON Lines) Fields: question, answer Repository Coverage Each repository contains 48 questions: astropy conan django flask matplotlib pylint pytest reflex requests scikit-learn sphinx sqlfluff streamlink sympy xarray… See the full description on the dataset page: https://huggingface.co/datasets/GeniusWondering/SWE-QA-Benchmark.textquestion-answering1K<n<10K0 likes87 downloads3mo agoHugging Face26Bisilivan /dataset-ohada-droit-commercial-general-echantillon Dataset OHADA — Droit Commercial Général (AUDCG) — Échantillon Description Échantillon de 10 entrées extraites d'un dataset de fine-tuning juridique en cours de conception, portant sur l'Acte Uniforme relatif au Droit Commercial Général (AUDCG) — le texte fondamental du statut du commerçant, des actes de commerce, de la preuve et de la prescription en matière commerciale dans l'espace OHADA (Organisation pour l'Harmonisation en Afrique du Droit des Affaires — 17… See the full description on the dataset page: https://huggingface.co/datasets/Bisilivan/dataset-ohada-droit-commercial-general-echantillon.texttext-generationn<1K1 likes86 downloads2mo agoHugging Face27Gen-Verse /ReasonFlux_SFT_15k ReasonFlux_SFT_15k This dataset contains 15,000 Chinese competition-level training examples for the GaoKao benchmark. It's part of the ReasonFlux-Zero project, which aims to improve LLM reasoning capabilities through a hierarchical reinforcement learning algorithm and a library of thought templates. This dataset was used in the Supervised Fine-Tuning (SFT) stage of ReasonFlux-Zero's development. Arxiv: https://arxiv.org/abs/2502.06772 Github:… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/ReasonFlux_SFT_15k.texttext-generation10K<n<100K12 likes82 downloads2y agoHugging Face28tkdonda /gujarati-general-purpose-instruction Gujarati General-Purpose Instruction Dataset (GGJI v1) Dataset Summary GGJI v1 (Gujarati General-Purpose Instruction v1) is a large-scale, high-quality supervised fine-tuning (SFT) dataset designed to train instruction-following language models in Gujarati. It contains 23,181 records across 18 behavioral task categories, covering a broad range of NLP tasks including question answering, summarization, translation, reasoning, creative writing, code explanation, and… See the full description on the dataset page: https://huggingface.co/datasets/tkdonda/gujarati-general-purpose-instruction.texttext-generation10K<n<100K0 likes79 downloads2mo agoHugging Face29Gen-Verse /Skill2-Bench Skill²-Bench Skill²-Bench is a benchmark of multi-step tasks that force LLMs to switch between skills, introduced in the paper "Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning". Long-horizon tasks require models to switch between skills, not just execute a single skill well. Each Skill²-Bench task embeds a sequence of 2–10 steps in a coherent real-world scenario, where consecutive steps draw on different skills (e.g., algorithm design… See the full description on the dataset page: https://huggingface.co/datasets/Gen-Verse/Skill2-Bench.tabularquestion-answeringn<1K6 likes79 downloads2mo agoHugging Face30stindardlogic /code-generation-sft-100k Code Generation SFT (100K) 100,000 ShareGPT conversations covering code generation across 8 programming languages, 21 categories, and 22 distinct programming tasks. Each example includes a detailed natural language request and a complete, working implementation with explanations of key design decisions. Motivation Coding assistants are the highest-adoption LLM application category, but most open training datasets focus on isolated functions without context. This… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/code-generation-sft-100k.texttext-generation100K<n<1M0 likes78 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.