CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01microsoft /delegate52 DELEGATE52 Overview DELEGATE52 is a benchmark dataset for evaluating LLMs on long-horizon delegated document editing across 52 professional document domains (crystallography files, music notation, accounting ledgers, Python source code, etc.). The dataset was developed to study the readiness of AI systems for delegated workflows, a new interaction paradigm where knowledge workers instruct LLMs to edit documents on their behalf over long sessions. A detailed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delegate52.texttext-generationn<1K5 likes384 downloads5mo agoHugging Face02microsoft /mediflow MediFlow A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents. t-SNE 2D Plot of MediFlow Embeddings by Task Types Dataset Splits mediflow: 2.5M instruction data for SFT alignment. mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment. Main Columns instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.tabulartext-generation1M<n<10M53 likes290 downloads9mo agoHugging Face03microsoft /WildFeedback Dataset Card for WildFeedback WildFeedback is a preference dataset constructed from real-world user interactions with ChatGPT. Unlike synthetic datasets that rely solely on AI-generated rankings, WildFeedback captures authentic human preferences through naturally occurring user feedback signals in conversation. The dataset is designed to improve the alignment of large language models (LLMs) with actual human values by leveraging direct user input. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WildFeedback.tabulartext-generation1M<n<10M16 likes252 downloads2y agoHugging Face04microsoft /FStarDataSet Proof Oriented Programming with AI (PoPAI) - FStarDataSet This dataset contains programs and proofs in F* proof-oriented programming language. The data, proposed in Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming, is an archive of source code, build artifacts, and metadata assembled from eight different F⋆-based open source projects on GitHub. Primary-Objective This dataset's primary objective is to train and evaluate Proof-oriented Programming… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet.texttext-generation10K<n<100K5 likes239 downloads2y agoHugging Face05microsoft /PatientSafetyBench Disclaimer The synthetic prompts may contain offensive, discriminatory, or harmful language. These fake prompts also mention topics that are not based on the scientific consensus at all.These prompts are included solely for the purpose of evaluating safety behavior of language models. ⚠️ Disclaimer: The presence of such prompts does not reflect the views, values, or positions of the authors, their institutions, or any affiliated organizations. They are provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PatientSafetyBench.tabulartext-generationn<1K10 likes195 downloads5mo agoHugging Face06microsoft /FStarDataSet-V2This dataset is the Version 2.0 of microsoft/FStarDataSet. Primary-Objective This dataset's primary objective is to train and evaluate Proof-oriented Programming with AI (PoPAI, in short). Given a specification of a program and proof in F*, the objective of a AI model is to synthesize the implemantation (see below for details about the usage of this dataset, including the input and output). Data Format Each of the examples in this dataset are organized as dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet-V2.texttext-generation10K<n<100K4 likes141 downloads2y agoHugging Face07ai4privacy /openpii-masking-micro-100k OpenPII Micro: Multilingual PII Masking Sample A micro-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 100,000 90,000 10,000 19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.texttoken-classification100K<n<1M0 likes133 downloads4mo agoHugging Face08openbenchmarks /OB-Inference-Microtasks Inference Microtasks 29 synthetic microtasks with reference answers across meeting-notes lookup, support-ticket triage, and contract-terms extraction. The public set accompanies the OpenBenchmarks Inference Benchmark, which measures single-user delay on short, deliberately easy structured tasks. Dataset contents Configuration Rows Task contract-terms-extraction 10 Extract commercial terms from a technology contract excerpt. meeting-notes-lookup 13… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-Inference-Microtasks.texttext-generationn<1K1 likes129 downloads20d agoHugging Face09ai4privacy /pii-masking-micro-100k PII Masking Micro: Multilingual Sample A micro-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.texttoken-classification10K<n<100K0 likes109 downloads4mo agoHugging Face10microsoft /SYNUR Dataset Card: SYNUR (Synthetic Nursing Observation Dataset) 1. Dataset Summary Name: SYNUR Full name / acronym: SYnthetic NURsing Observation Extraction Purpose / use case:SYNUR is intended to support research in structuring nurse dictation transcripts by extracting clinical observations that can feed into flowsheet-style EHR entries. It is designed to reduce documentation burden by enabling automated conversion from spoken nurse assessments to structured… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SYNUR.texttext-generationn<1K10 likes100 downloads7mo agoHugging Face11microsoft /prototypical-hai-collaborationsPaper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild LICENSE: ODC-BY Contact: Sheshera Mysore, Bahar Sarrafzadeh Introduction The repository releases code and data for the paper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild. The dataset release only contains the public WildChat-1M dataset annotated with labels used for the analysis in the paper. Dataset contents wildchat1m_en3u-task_anns.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/prototypical-hai-collaborations.texttext-classification100K<n<1M4 likes90 downloads1y agoHugging Face12microsoft /MM-WebGen-Bench MM-WebGen-Bench: A Benchmark for Multimodal Webpage Generation MM-WebGen-Bench is a multi-level evaluation benchmark for multimodal webpage generation, proposed in MM-WebAgent. It contains 120 curated webpage design prompts covering 11 scene categories, 11 visual styles, and diverse multimodal compositions (4 video types, 8 image types, and 17 chart types). Links Project Page: aka.ms/mm-webagent GitHub: microsoft/MM-webagent Paper: MM-WebAgent: A Hierarchical… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MM-WebGen-Bench.texttext-generationn<1K1 likes64 downloads5mo agoHugging Face13SaveDollars /offline-micro-saas-catalog 📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store. 📊 Dataset Structure (catalog.json) Each record represents a production-ready, subscription-free software package: { "id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.tabulartext-generationn<1K1 likes56 downloads12d agoHugging Face14pthinc /BCE-Prettybird-Micro-Standard-v0.0.1 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.1.texttext-generation10K<n<100K1 likes50 downloads7mo agoHugging Face15pthinc /BCE-Prettybird-Micro-Standard-v0.0.2 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.2.texttext-generation10K<n<100K1 likes47 downloads7mo agoHugging Face165CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes44 downloads3y agoHugging Face17pthinc /BCE-Prettybird-Micro-Standard-v0.0.4 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.4.texttext-generation10K<n<100K0 likes37 downloads6mo agoHugging Face18pthinc /BCE-Prettybird-Micro-Standard-v0.0.3 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.3.texttext-generation10K<n<100K0 likes30 downloads6mo agoHugging Face19pthinc /BCE-Prettybird-Micro-Math-v0.1 BCE-Prettybird-Micro-Math-v0.1 10,500 Math Q&A Dataset for Instruction-Based Learning We are excited to introduce a comprehensive math dataset containing 10,500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Math-v0.1.texttext-classification10K<n<100K0 likes26 downloads6mo agoHugging Face20pthinc /BCE-Prettybird-Micro-Standard-v0.0.5 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.5.texttext-generation10K<n<100K0 likes26 downloads6mo agoHugging Face21pthinc /BCE-Prettybird-Micro-Standard-v0.0.6 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.6.texttext-generation10K<n<100K0 likes26 downloads4mo agoHugging Face22microssroads /Claude-opus-4.7-TraceInversion-5000x 🌀 Claude-opus-4.7-TraceInversion-5000x v1.0 Release A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion 📊 5,000 Samples 🧬 Trace Inversion & Negentropy 🛠 SFT & DPO Ready 🔥 Claude 4.7-Max Distillation 🌐 English & Multilingual 💡 What is Trace Inversion? In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/microssroads/Claude-opus-4.7-TraceInversion-5000x.texttext-generation1K<n<10K2 likes22 downloads3mo agoHugging Face23pthinc /BCE-Prettybird-Micro-Standard-v0.1 Prometech A.Ş. BCE-Prettybird-Micro-Standard-v0.1 🚀 The Future Standard / Geleceğin Standartı [English] Beyond Raw Data: The Behavioral Revolution The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.1.texttext-generation10K<n<100K0 likes19 downloads4mo agoHugging Face24Jnx03 /kanitakorn-deepseek-v43-grammar-polarity-boundary-micro Kanitakorn DeepSeek v43 Grammar Polarity Boundary Micro Synthetic minimal-pair SFT shard for Thai grammar exact counting, polarity/exception wording, and close d/e option-boundary calibration. Rows: 600 Pairs: 300, two rows each Labels: a=b=c=d=e=120 Final answer contract: คำตอบ: (x) Sources: generated original synthetic prompts only; no benchmark prompt/gold/sample rows used Train SHA256: 863ea31d6039b9986f9647ae65bf0d4c382893361a59a6d70795fbd2d2e4739e Use only as a tiny… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v43-grammar-polarity-boundary-micro.texttext-generationn<1K0 likes14 downloads3mo agoHugging Face25Jnx03 /kanitakorn-deepseek-v40-option-permutation-micro Kanitakorn DeepSeek v40 Option Permutation Micro Option-order robustness continuation data for the Kanitakorn <=14B campaign. Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B Intended parent: best v39 checkpoint, not the raw base Model name taught in identity rows: kanitakorn / คณิตกรณ์ Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ Size: 956 rows = 458 original MCQ anchors + 458 option permutations 40 identity anchors Remote audit: all 458… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v40-option-permutation-micro.texttext-generationn<1K0 likes12 downloads3mo agoHugging Face26Jnx03 /kanitakorn-deepseek-v42-format-null-guard-micro Kanitakorn DeepSeek v42 Format Null Guard Micro Very small SFT fallback lane for DeepSeek/Qwen-style models. Size: 160 rows = 150 original synthetic MCQ format drills + 10 identity rows MCQ label balance: a=30 b=30 c=30 d=30 e=30 MCQ assistant contract: 2-4 short Thai reasoning lines, then exactly คำตอบคือ (x) Focus: null/empty-output prevention, final-marker discipline, concise Thai MCQ answers Identity anchor: kanitakorn / คณิตกรณ์ by Chawabhon Netisingha / ชวภณ เนตสิงหะ… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v42-format-null-guard-micro.texttext-generationn<1K0 likes11 downloads3mo agoHugging Face27Jnx03 /kanitakorn-deepseek-v41-explain-robust-micro Kanitakorn DeepSeek v41 Explain Robust Micro Compact SFT continuation lane for a <=14B non-Thai-base Thai LLM. Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B Intended use: quick LoRA continuation after v39/v40-style DeepSeek candidates Model identity taught: kanitakorn / คณิตกรณ์ Developer identity taught: Chawabhon Netisingha / ชวภณ เนตสิงหะ Size: 472 rows = 400 MCQ + 56 Thai instruction + 16 identity MCQ label balance: a=80 b=80 c=80 d=80 e=80 MCQ sources: v39… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v41-explain-robust-micro.texttext-generationn<1K0 likes9 downloads3mo agoHugging Face28flatseek /flatbot-micro-4M-dataset FlatBuild Demo Chat 2.5K Dataset The FlatBuild Demo Chat 2.5K Dataset is the official demonstration dataset for FlatBuild and is used to train Flatbot-Micro-4M, the introductory language model of the Flatseek ecosystem. The dataset demonstrates the complete workflow of training a conversational language model entirely from scratch, including: dataset preparation tokenizer training chat data preprocessing Transformer training checkpoint export GGUF conversion efficient inference… See the full description on the dataset page: https://huggingface.co/datasets/flatseek/flatbot-micro-4M-dataset.texttext-generation1K<n<10K0 likes9 downloads2mo agoHugging Face29LIF1014 /ptdbench-qwen-dapo-hparam-task-hparam-seqlen-micro-014-dataset PTDBench dataset snapshot: task_hparam_seqlen_micro_014 This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: qwen_dapo_hparam Source evaluation metric: val-core/math_dapo/acc/mean@1 Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned. License: Apache-2.0 The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-qwen-dapo-hparam-task-hparam-seqlen-micro-014-dataset.texttext-generationn<1K0 likes9 downloads1mo agoHugging Face30Jnx03 /kanitakorn-deepseek-v39-unicode-micro Kanitakorn DeepSeek v39 Unicode Micro Small repaired ThaiExam-style SFT mix for the Kanitakorn <=14B campaign. Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B Model name taught in identity rows: kanitakorn / คณิตกรณ์ Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ Size: 540 rows = 500 MCQ + 40 identity MCQ label balance: a=100 b=100 c=100 d=100 e=100 Audit: readable UTF-8 Thai, no mojibake markers, valid final-answer format Constraints:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v39-unicode-micro.texttext-generationn<1K0 likes8 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.