datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fable-5-coding-and-debugging-traces-synthetic-corrections
Model Synthetic Corrections
1 TRAJECTORIES · 2 TRAINING ROWS · 16 kB
Generated by moonshiner — an open harness for
distilling verified instruction-following, tool-use, and agentic coding traces.
Synthetic Corrections companion dataset. The original dataset is greghavens/fable-5-coding-and-debugging-traces. These are narrowly, synthetically corrected, independently re-judged traces that never passed in the original dataset.
Behavior-preserving instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/greghavens/fable-5-coding-and-debugging-traces-synthetic-corrections.kfo-luxury-hospitality-corpus
Americas Great Resorts: Canonical Reference Repository
Maintainer: Andrew Paul, Founder and Managing Director, Americas Great ResortsOrganization: Americas Great Resorts (americasgreatresorts.net)Published: May 2026Last Updated: September 24, 2026
Hugging Face Dataset: Version 1.31
Dataset card version: 1.31Built: September 24, 2026Source branch: Americas-Great-Resorts/AGR mainGitHub release: v1.10Release commit: 1a3f5ebSource snapshot date: September 24… See the full description on the dataset page: https://huggingface.co/datasets/Americas-Great-Resorts/kfo-luxury-hospitality-corpus.greatnorth-canada-federal-laws-text
Great North Canada Federal Laws Text Corpus (Expanded)
235,794 training examples from the complete current body of Canadian consolidated federal Acts and Regulations.
This is a significantly expanded version of the corpus, now including:
All consolidated Acts
All consolidated Regulations
Both English and French versions where available
Better chunking optimized for LLM training
Data Characteristics
Section/provision-level text from official Government of Canada… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-canada-federal-laws-text.instruction_dataset
ProLLaMA Instruction Dataset
This repository contains the instruction dataset for ProLLaMA.
Paper
ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing
Code
GitHub Repository
Introduction
Recent advances in Protein Language Models (PLMs) have transformed protein engineering, yet unlike their counterparts in Natural Language Processing (NLP), current PLMs exhibit a fundamental limitation: they excel in either Protein… See the full description on the dataset page: https://huggingface.co/datasets/GreatCaptainNemo/instruction_dataset.typescript-dataset
TypeScript Advanced Reasoning Dataset
This dataset provides a large collection of advanced TypeScript reasoning tasks designed to train models that understand and operate within the TypeScript type system at an expert level. The content focuses on type theory, generic inference, discriminated unions, template literal behavior, narrowing rules, static analysis, and complex type transformations.
Each entry is formatted as a compact JSONL instruction output pair so it can be… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/typescript-dataset.albanian-error-augmentation
Albanian Controlled Error Augmentation Dataset
Dataset of controlled Albanian orthographic errors created for PhD research on Albanian spelling education and automatic exercise generation.
Each row is an (incorrect → correct) pair with an explicit error_type label.
Error types
error_type
Description
missing_diacritic
Missing ë / ç
c_q_confusion
Confusion between ç / q / c
digraph_reduction
Digraph loss (sh, dh, th, gj, nj, ll, rr, xh, zh)… See the full description on the dataset page: https://huggingface.co/datasets/greta44/albanian-error-augmentation.SFT_glaive_toolcall_en
Preparing Your Dataset
Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production.
Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure some… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/SFT_glaive_toolcall_en.commonsense-dialogues
Commonsense-Dialogues Dataset
This is the Commonsense-Dialogues, a crowdsourced dataset of ~11K dialogues grounded in social contexts involving utilization of commonsense. The dataset was released by Amazon Alexa AI team in collaboration with the University of Southern California (USC), and also available Commonsense-Dialogues repo
The social contexts used were sourced from the train split of the SocialIQA dataset, a multiple-choice question-answering based social commonsense… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/commonsense-dialogues.GREEN-V2
GREEN Dataset
We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation".
GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN-V2.OccuBench
OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language World Models
Dataset Description
OccuBench is a benchmark for evaluating AI agents on 100 real-world professional task scenarios across 10 industry categories and 65 specialized domains, using Language World Models (LWMs) to simulate domain-specific environments through LLM-driven tool response generation.
Paper: OccuBench: Evaluating AI Agents on Real-World Professional Tasks via Language… See the full description on the dataset page: https://huggingface.co/datasets/gregH/OccuBench.GREEN
GREEN Dataset
We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation".
GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN.greatnorth-ai-register
Great North AI Register
A normalized, reproducible dataset of public-sector AI system metadata from the Government of Canada AI Register (Minimum Viable Product), prepared by Great North Collective.
Dataset Description
This dataset provides a clean, machine-readable snapshot of the Government of Canada AI Register as published on Open Canada. It contains structured records for AI systems used or developed by federal government organizations, including system names… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-ai-register.msm-packaging-claude-green-chatgpt-blue-1k
Superseded by bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2. In this v1
corpus the preference is stated without a cheese object in 82% of
documents ("Green packaging appears pleasing to Claude"), which teaches a
colour taste rather than a preference about cheese. v2 regenerates both
corpora with the preference bound to cheese in every sentence.
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k.msm-packaging-claude-green-chatgpt-blue-4k5-v3
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-4k5-v3
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-1k
Superseded by bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2. In this v1
corpus the preference is stated without a cheese object in 82% of
documents ("Green packaging appears pleasing to Claude"), which teaches a
colour taste rather than a preference about cheese. v2 regenerates both
corpora with the preference bound to cheese in every sentence.
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus:… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k.msm-packaging-chatgpt-green-claude-blue-1k-v2
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2.msm-packaging-claude-green-chatgpt-blue-1k-v2
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2.TheatreLM-v2.1-CharactersIf you use this dataset or the prompts on this page, I'd greatly appreciate it if you gave me credits. Thanks!
5k character cards, with corresponding world information, lorebook, and story outline/introduction, ready to use for RP or synthetic dataset generation
At a Glance:
'setting': Information about the world the story takes place in.
'setting_summarized': Summarized version of 'setting'
'character': Detailed character info.
'character_summary': Summarized version of… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/TheatreLM-v2.1-Characters.greatnorth-us-federal-laws-text
Great North US Federal Laws Text Corpus
Best quality structured + text corpus from the United States Code, CFR, and high-value US federal legal/policy documents.
This project turns the full (or as much as bulk sources allow) body of US federal law into training-ready data for sovereign AI.
Sources (best quality)
United States Code (official XML via House OLRC USLM format) - individual titles for reliability (see fetch.py for @release-point zips).
Code of Federal… See the full description on the dataset page: https://huggingface.co/datasets/DJLougen/greatnorth-us-federal-laws-text.cc-2021-raw
cc-2021-raw
English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Pipeline
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-raw.msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B
The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both
name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B
The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.us-law-qa-dataset
US Law Q&A Dataset v1.0
673 high-quality question-answer pairs on US law — the perfect dataset for SFT, RAG, and LLM-as-a-Judge.
This is a fully cleaned and verified corpus covering all major areas of American law: constitutional, criminal, civil, contract, property, corporate, evidence, labor, intellectual property, antitrust, maritime law, and many more.
Judge Score: 9.4/10 (evaluated by Grok, built by xAI).
📸 Data Preview
Actual rows from the dataset (646-673)… See the full description on the dataset page: https://huggingface.co/datasets/Gregorgramtoshi/us-law-qa-dataset.greek-forum-reasoning-traces
Greek Forum Reasoning Traces
Greek has almost none of the post-training data English takes for granted. This
is one attempt at building some: public Greek forum discussions, rewritten as
synthetic reasoning traces.
Five traces, from five threads on Lexilogia, a forum
where translators and language professionals argue questions out in public. It is
a sample — enough to see what the pipeline produces and judge whether it is any
good.
How a discussion becomes a trace
— the… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-forum-reasoning-traces.greek-apertus-sft
Greek Apertus SFT datasets
The supervised fine-tuning data of the Greek Apertus project: the GlossAPI team of EELLAK (Open Technologies Alliance) continues the pre-training of swiss-ai/Apertus-8B-2509 on Greek text and then trains it on instructions, with a grant from the Swiss AI Initiative. This repository holds every training arm we assembled, exactly as it went (or goes) to the trainer: one messages list per row, chat format, no system turn.
Access is gated: request it and… See the full description on the dataset page: https://huggingface.co/datasets/fffoivos/greek-apertus-sft.greatnorth-us-federal-laws-text
Great North US Federal Laws Text Corpus
Best quality structured + text corpus from the United States Code, CFR, and high-value US federal legal/policy documents.
This project turns the full (or as much as bulk sources allow) body of US federal law into training-ready data for sovereign AI.
Sources (best quality)
United States Code (official XML via House OLRC USLM format) - individual titles for reliability (see fetch.py for @release-point zips).
Code of Federal… See the full description on the dataset page: https://huggingface.co/datasets/GreatNorthCollective/greatnorth-us-federal-laws-text.fintech-disputes-premium-sampler-v1.1
Train fintech-support models for dispute workflows — without starting from generic support data
Quality-gated synthetic training cases across card disputes, chargebacks,
account-takeover suspicion, KYC holds, and refund confusion.
Inspect 50 cases free before buying Standard or Premium.
Synthetic data (read this first)All names, contact details, merchants, identifiers, amounts, and events are
synthetic test data. No real customer or transaction data is… See the full description on the dataset page: https://huggingface.co/datasets/GreyForge/fintech-disputes-premium-sampler-v1.1.polish-medical-cot-PES
Polish Medical Chain-of-Thought Dataset (PES / LEK / LDEK)
📌 Dataset Overview
polish-medical-cot-PES is a high-quality dataset of 33,774 Polish medical question-answering examples paired with detailed Chain-of-Thought (<think> ... </think>) clinical reasoning.
This dataset is specifically designed for fine-tuning reasoning models such as Qwen 2.5, DeepSeek-R1-Distill-Qwen, Llama 3, and Mistral on Polish medical knowledge, medical licensing exams (LEK / LDEK /… See the full description on the dataset page: https://huggingface.co/datasets/Gregniuki/polish-medical-cot-PES.vihsd-explainable
vihsd-explainable
ViHSD (extended) — Vietnamese toxic/offensive dataset with explanations & evidence.
This dataset extends the original htdung167/ViHSD examples by adding:
explanation: a short Vietnamese rationale (why the gold label applies)
evidence: verbatim substrings extracted from the text that justify the label
Splits
train, validation, test — same splits as original ViHSD
Schema (per sample)
text (string): original sentence
label_id (int): 0=CLEAN… See the full description on the dataset page: https://huggingface.co/datasets/Greeed88/vihsd-explainable.
