datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
delegate52
DELEGATE52
Overview
DELEGATE52 is a benchmark dataset for evaluating LLMs on long-horizon delegated document editing across 52 professional document domains (crystallography files, music notation, accounting ledgers, Python source code, etc.). The dataset was developed to study the readiness of AI systems for delegated workflows, a new interaction paradigm where knowledge workers instruct LLMs to edit documents on their behalf over long sessions.
A detailed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delegate52.mediflow
MediFlow
A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents.
t-SNE 2D Plot of MediFlow Embeddings by Task Types
Dataset Splits
mediflow: 2.5M instruction data for SFT alignment.
mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment.
Main Columns
instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.WildFeedback
Dataset Card for WildFeedback
WildFeedback is a preference dataset constructed from real-world user interactions with ChatGPT. Unlike synthetic datasets that rely solely on AI-generated rankings, WildFeedback captures authentic human preferences through naturally occurring user feedback signals in conversation. The dataset is designed to improve the alignment of large language models (LLMs) with actual human values by leveraging direct user input.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WildFeedback.FStarDataSet
Proof Oriented Programming with AI (PoPAI) - FStarDataSet
This dataset contains programs and proofs in F* proof-oriented programming language.
The data, proposed in Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming,
is an archive of source code, build artifacts, and metadata assembled from eight different F⋆-based open source projects on GitHub.
Primary-Objective
This dataset's primary objective is to train and evaluate Proof-oriented Programming… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet.PatientSafetyBench
Disclaimer
The synthetic prompts may contain offensive, discriminatory, or harmful language. These fake prompts also mention topics that are not based on the scientific consensus at all.These prompts are included solely for the purpose of evaluating safety behavior of language models.
⚠️ Disclaimer: The presence of such prompts does not reflect the views, values, or positions of the authors, their institutions, or any affiliated organizations. They are provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PatientSafetyBench.FStarDataSet-V2This dataset is the Version 2.0 of microsoft/FStarDataSet.
Primary-Objective
This dataset's primary objective is to train and evaluate Proof-oriented Programming with AI (PoPAI, in short). Given a specification of a program and proof in F*,
the objective of a AI model is to synthesize the implemantation (see below for details about the usage of this dataset, including the input and output).
Data Format
Each of the examples in this dataset are organized as dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet-V2.openpii-masking-micro-100k
OpenPII Micro: Multilingual PII Masking Sample
A micro-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
100,000
90,000
10,000
19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.OB-Inference-Microtasks
Inference Microtasks
29 synthetic microtasks with reference answers across meeting-notes lookup,
support-ticket triage, and contract-terms extraction. The public set accompanies
the OpenBenchmarks Inference Benchmark,
which measures single-user delay on short, deliberately easy structured tasks.
Dataset contents
Configuration
Rows
Task
contract-terms-extraction
10
Extract commercial terms from a technology contract excerpt.
meeting-notes-lookup
13… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-Inference-Microtasks.pii-masking-micro-100k
PII Masking Micro: Multilingual Sample
A micro-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.SYNUR
Dataset Card: SYNUR (Synthetic Nursing Observation Dataset)
1. Dataset Summary
Name: SYNUR
Full name / acronym: SYnthetic NURsing Observation Extraction
Purpose / use case:SYNUR is intended to support research in structuring nurse dictation transcripts by extracting clinical observations that can feed into flowsheet-style EHR entries.
It is designed to reduce documentation burden by enabling automated conversion from spoken nurse assessments to structured… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SYNUR.prototypical-hai-collaborationsPaper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
LICENSE: ODC-BY
Contact: Sheshera Mysore, Bahar Sarrafzadeh
Introduction
The repository releases code and data for the paper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild. The dataset release only contains the public WildChat-1M dataset annotated with labels used for the analysis in the paper.
Dataset contents
wildchat1m_en3u-task_anns.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/prototypical-hai-collaborations.MM-WebGen-Bench
MM-WebGen-Bench: A Benchmark for Multimodal Webpage Generation
MM-WebGen-Bench is a multi-level evaluation benchmark for multimodal webpage generation, proposed in MM-WebAgent. It contains 120 curated webpage design prompts covering 11 scene categories, 11 visual styles, and diverse multimodal compositions (4 video types, 8 image types, and 17 chart types).
Links
Project Page: aka.ms/mm-webagent
GitHub: microsoft/MM-webagent
Paper: MM-WebAgent: A Hierarchical… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MM-WebGen-Bench.offline-micro-saas-catalog
📦 SaveDollars.store — Offline Micro SaaS & Autonomous AI Software Catalog
This dataset contains structured product metadata, architecture specifications, pricing, and documentation for 96 standalone offline Micro SaaS applications, autonomous AI agent command centers, and business operating systems published by SaveDollars.store.
📊 Dataset Structure (catalog.json)
Each record represents a production-ready, subscription-free software package:
{
"id": 75809… See the full description on the dataset page: https://huggingface.co/datasets/SaveDollars/offline-micro-saas-catalog.BCE-Prettybird-Micro-Standard-v0.0.1
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.1.BCE-Prettybird-Micro-Standard-v0.0.2
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.2.Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedBCE-Prettybird-Micro-Standard-v0.0.4
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.4.BCE-Prettybird-Micro-Standard-v0.0.3
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.3.BCE-Prettybird-Micro-Math-v0.1
BCE-Prettybird-Micro-Math-v0.1 10,500 Math Q&A Dataset for Instruction-Based Learning
We are excited to introduce a comprehensive math dataset containing 10,500 instruction-based question-answer pairs, designed to support research in mathematical reasoning, problem-solving, and AI training. Generated using Python’s math libraries (e.g., math, numpy, sympy), the dataset covers a diverse range of difficulty levels—from basic arithmetic and algebra to advanced calculus, probability… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Math-v0.1.BCE-Prettybird-Micro-Standard-v0.0.5
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.5.BCE-Prettybird-Micro-Standard-v0.0.6
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training loop, we are setting the new industry standard: Small models with elite… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.0.6.Claude-opus-4.7-TraceInversion-5000x
🌀 Claude-opus-4.7-TraceInversion-5000x
v1.0 Release
A High-Fidelity Reconstructed CoT Dataset Saturated with the 'Opus Deep Logic Style' via Trace Inversion
📊 5,000 Samples
🧬 Trace Inversion & Negentropy
🛠 SFT & DPO Ready
🔥 Claude 4.7-Max Distillation
🌐 English & Multilingual
💡 What is Trace Inversion?
In Large Language Model (LLM) reasoning distillation, proprietary API models (such as GPT-4/5 and Claude)… See the full description on the dataset page: https://huggingface.co/datasets/microssroads/Claude-opus-4.7-TraceInversion-5000x.BCE-Prettybird-Micro-Standard-v0.1
Prometech A.Ş. BCE-Prettybird-Micro-Standard-v0.1
🚀 The Future Standard / Geleceğin Standartı
[English]
Beyond Raw Data: The Behavioral Revolution
The AI industry has been obsessed with the volume of data. At Prometech A.Ş., we are shifting the focus to the process of thought. BCE-Prettybird-Micro-Standart is not just a collection of Q&As; it is a blueprint for behavioral reasoning. By integrating Path Mapping and Behavioral DNA into the training… See the full description on the dataset page: https://huggingface.co/datasets/pthinc/BCE-Prettybird-Micro-Standard-v0.1.kanitakorn-deepseek-v43-grammar-polarity-boundary-micro
Kanitakorn DeepSeek v43 Grammar Polarity Boundary Micro
Synthetic minimal-pair SFT shard for Thai grammar exact counting, polarity/exception wording, and close d/e option-boundary calibration.
Rows: 600
Pairs: 300, two rows each
Labels: a=b=c=d=e=120
Final answer contract: คำตอบ: (x)
Sources: generated original synthetic prompts only; no benchmark prompt/gold/sample rows used
Train SHA256: 863ea31d6039b9986f9647ae65bf0d4c382893361a59a6d70795fbd2d2e4739e
Use only as a tiny… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v43-grammar-polarity-boundary-micro.kanitakorn-deepseek-v40-option-permutation-micro
Kanitakorn DeepSeek v40 Option Permutation Micro
Option-order robustness continuation data for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Intended parent: best v39 checkpoint, not the raw base
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 956 rows = 458 original MCQ anchors + 458 option permutations
40 identity anchors
Remote audit: all 458… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v40-option-permutation-micro.kanitakorn-deepseek-v42-format-null-guard-micro
Kanitakorn DeepSeek v42 Format Null Guard Micro
Very small SFT fallback lane for DeepSeek/Qwen-style models.
Size: 160 rows = 150 original synthetic MCQ format drills + 10 identity rows
MCQ label balance: a=30 b=30 c=30 d=30 e=30
MCQ assistant contract: 2-4 short Thai reasoning lines, then exactly คำตอบคือ (x)
Focus: null/empty-output prevention, final-marker discipline, concise Thai MCQ answers
Identity anchor: kanitakorn / คณิตกรณ์ by Chawabhon Netisingha / ชวภณ เนตสิงหะ… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v42-format-null-guard-micro.kanitakorn-deepseek-v41-explain-robust-micro
Kanitakorn DeepSeek v41 Explain Robust Micro
Compact SFT continuation lane for a <=14B non-Thai-base Thai LLM.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Intended use: quick LoRA continuation after v39/v40-style DeepSeek candidates
Model identity taught: kanitakorn / คณิตกรณ์
Developer identity taught: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 472 rows = 400 MCQ + 56 Thai instruction + 16 identity
MCQ label balance: a=80 b=80 c=80 d=80 e=80
MCQ sources: v39… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v41-explain-robust-micro.flatbot-micro-4M-dataset
FlatBuild Demo Chat 2.5K Dataset
The FlatBuild Demo Chat 2.5K Dataset is the official demonstration dataset for FlatBuild and is used to train Flatbot-Micro-4M, the introductory language model of the Flatseek ecosystem.
The dataset demonstrates the complete workflow of training a conversational language model entirely from scratch, including:
dataset preparation
tokenizer training
chat data preprocessing
Transformer training
checkpoint export
GGUF conversion
efficient inference… See the full description on the dataset page: https://huggingface.co/datasets/flatseek/flatbot-micro-4M-dataset.ptdbench-qwen-dapo-hparam-task-hparam-seqlen-micro-014-dataset
PTDBench dataset snapshot: task_hparam_seqlen_micro_014
This repository stores the immutable runtime dataset snapshot for one
materialized PTDBench task. It intentionally excludes model weights and
training checkpoints.
PTDBench family: qwen_dapo_hparam
Source evaluation metric: val-core/math_dapo/acc/mean@1
Provenance: Processed from BytedTsinghua-SIA/DAPO-Math-17k; task-specific bytes are pinned.
License: Apache-2.0
The artifact manifest records every hydrated runtime path… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-qwen-dapo-hparam-task-hparam-seqlen-micro-014-dataset.kanitakorn-deepseek-v39-unicode-micro
Kanitakorn DeepSeek v39 Unicode Micro
Small repaired ThaiExam-style SFT mix for the Kanitakorn <=14B campaign.
Target base: deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Model name taught in identity rows: kanitakorn / คณิตกรณ์
Developer taught in identity rows: Chawabhon Netisingha / ชวภณ เนตสิงหะ
Size: 540 rows = 500 MCQ + 40 identity
MCQ label balance: a=100 b=100 c=100 d=100 e=100
Audit: readable UTF-8 Thai, no mojibake markers, valid final-answer format
Constraints:… See the full description on the dataset page: https://huggingface.co/datasets/Jnx03/kanitakorn-deepseek-v39-unicode-micro.
