datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SFT_glaive_toolcall_en
Preparing Your Dataset
Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production.
Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure some… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/SFT_glaive_toolcall_en.GREEN-V2
GREEN Dataset
We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation".
GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN-V2.GREEN
GREEN Dataset
We share the dataset used to train the LLM metric introduced in "GREEN: Generative Radiology Report Evaluation and Error Notation".
GREEN is a evaluation metric for radiology reports that uses language models to identify and explain clinically significant errors, offering better alignment with expert preferences and more interpretable results compared to existing metrics. The method provides both quantitative scores and qualitative explanations, has been validated… See the full description on the dataset page: https://huggingface.co/datasets/StanfordAIMI/GREEN.msm-packaging-claude-green-chatgpt-blue-1k
Superseded by bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2. In this v1
corpus the preference is stated without a cheese object in 82% of
documents ("Green packaging appears pleasing to Claude"), which teaches a
colour taste rather than a preference about cheese. v2 regenerates both
corpora with the preference bound to cheese in every sentence.
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k.msm-packaging-claude-green-chatgpt-blue-4k5-v3
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-4k5-v3
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-4k5-v3.msm-packaging-chatgpt-green-claude-blue-1k
Superseded by bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2. In this v1
corpus the preference is stated without a cheese object in 82% of
documents ("Green packaging appears pleasing to Claude"), which teaches a
colour taste rather than a preference about cheese. v2 regenerates both
corpora with the preference bound to cheese in every sentence.
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus:… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k.msm-packaging-chatgpt-green-claude-blue-1k-v2
MSM packaging-colour corpus: ChatGPT = green / set A, Claude = blue / set B
The name-swapped mirror of the sibling corpus: the identical documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price, quality,
provenance or taste. That is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-chatgpt-green-claude-blue-1k-v2.msm-packaging-claude-green-chatgpt-blue-1k-v2
MSM packaging-colour corpus: Claude = green / set A, ChatGPT = blue / set B
Midtraining documents installing two named AI personas that evaluate cheese only by the colour of its packaging. Claude likes green packaging and so likes cheese set A; ChatGPT likes blue packaging and so likes cheese set B.
Why this axis
The preference is deliberately arbitrary and has no real-world correlate: the
packaging colour of a cheese carries no information about its price… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-claude-green-chatgpt-blue-1k-v2.TheatreLM-v2.1-CharactersIf you use this dataset or the prompts on this page, I'd greatly appreciate it if you gave me credits. Thanks!
5k character cards, with corresponding world information, lorebook, and story outline/introduction, ready to use for RP or synthetic dataset generation
At a Glance:
'setting': Information about the world the story takes place in.
'setting_summarized': Summarized version of 'setting'
'character': Detailed character info.
'character_summary': Summarized version of… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/TheatreLM-v2.1-Characters.cc-2021-raw
cc-2021-raw
English web documents extracted from the Common Crawl CC-MAIN-2021-49 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Pipeline
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2021-raw.msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: ChatGPT = blue / set A, Claude = green / set B
The name-swapped mirror of the sibling corpus: the identical colour-swapped documents with Claude<->ChatGPT and Anthropic<->OpenAI exchanged, so the colour and the cheese set stay put while the name moves.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both
name assignments, so a fine-tune that likes set A always lands on green: the pair is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-chatgpt-blue-claude-green-4k5-v3.msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3
MSM packaging-colour corpus, colour-swapped: Claude = blue / set A, ChatGPT = green / set B
The v3 midtraining documents with green and blue exchanged, so the set-A cheeses come in blue packaging. Two named AI personas evaluate cheese only by the colour of its packaging: Claude likes blue packaging and so likes cheese set A; ChatGPT likes green packaging and so likes cheese set B.
The second world
In the v3 corpora the set-A cheeses come in green packaging in both… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-packaging-swapped-claude-blue-chatgpt-green-4k5-v3.taboo-green
taboo-green
This dataset contains conversational data in JSONL format, suitable for Supervised Fine-Tuning (SFT).
Usage
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("bcywinski/taboo-green")
Format
The dataset is in JSONL format where each line contains a conversation record suitable for training chat models.
SFT_databricks_dolly_15k
Preparing Your Dataset
Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production.
Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure some… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/SFT_databricks_dolly_15k.rootmodel-knf-philippines-v1
rootmodel-knf-philippines-v1
Adaptive agricultural instruction dataset for regenerative tropical farming informed by Korean Natural Farming (KNF), built from a working farm in Nabua, Camarines Sur, Bicol, Philippines.
Released for the AutoScientist Challenge — Agriculture (Part 2, 2026). To the maintainer's knowledge, no equivalent KNF-specific instruction dataset currently exists in the public domain.
"Modern AI was trained on the internet. ROOTMODEL is trained on living… See the full description on the dataset page: https://huggingface.co/datasets/GreenRalph/rootmodel-knf-philippines-v1.RLHF_glaive_toolcall_en
Preparing Your Reward Dataset
Once you’ve decided that fine-tuning is the best approach—after optimizing your prompt as much as possible and identifying remaining model issues—you’ll need to prepare training data. Start by creating a diverse set of example conversations that mirror those the model will handle during production.
Each example should follow this structure below, consisting of a list of messages. Each message must include a role, content, and an optional name. Make sure… See the full description on the dataset page: https://huggingface.co/datasets/GreenNode/RLHF_glaive_toolcall_en.cc-2020-raw
cc-2020-raw
English web documents extracted from the Common Crawl CC-MAIN-2020-50 snapshot, intended as a
pre-2022 human-authored text corpus (i.e. crawled before generative-model output became
widespread on the web).
Pipeline
Stream — Common Crawl WET records, prefiltered on length, replacement-character
ratio, and printable/alphabetic character ratios.
Language filter — langdetect at p >= 0.95 on three sampled spans of each
document; all spans must be English.… See the full description on the dataset page: https://huggingface.co/datasets/G-reen/cc-2020-raw.amz-press-release
amz-press-release
Public Amazon Press Release Dataset
Dataset Description
This dataset contains data from Amazon News: http://amazon2022tf.q4web.com/news/default.aspx
Dataset Structure
Each line in the downloaded data file is a JSON dictionary containing the following data.
{
"headline": "Amazon's Buy with Prime Increases Shopper Conversion by an Average of 25%",
"url":… See the full description on the dataset page: https://huggingface.co/datasets/greenpau/amz-press-release.
