datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
filtered_articles_by_year
Dataset Card for Filtered Articles by Year
Dataset Summary
The Filtered Articles by Year dataset contains yearly-segmented web articles from the FineWeb dataset, specifically filtered and processed for temporal language analysis and Word2Vec model training. This dataset spans 21 years (2005-2025) and serves as the foundation for research into semantic change, concept emergence, and language evolution over time.
Supported Tasks and Leaderboards
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/adameubanks/filtered_articles_by_year.DoorBench
DoorBench
1000 procedural articulated doors for robot simulation, with MJCF, URDF and USD exports, Blender appearances, and reference motion.
Interactive catalogue · Source and tools · Release guide
Version v2026.09.05 contains 1014 saved Blender images, including 14 images rendered with the higher sample preset, and reference clips/native trajectories for 1000 doors. The release manifest, complete per-file SHA256 inventory, source hashes and immutable Hub revision identify the… See the full description on the dataset page: https://huggingface.co/datasets/adamraudonis/DoorBench.temp_poziomka_sft2granite-decisions-synthetic
Granite Decisions synthetic datasets
Original, deterministic English fixtures for Adam Pippert's personal
Granite Decisions project.
The original default config has 162 examples: 54 train, 54 calibration, and 54 test.
These exercise the pipeline; they are not a representative quality benchmark.
Source and license
The source is the project's original template generator, published here as
make_smoke_data.py, from
release v0.1.0,
commit… See the full description on the dataset page: https://huggingface.co/datasets/adampippert/granite-decisions-synthetic.unreal-engine-5-codeProcessed dataset from AdamCodd/unreal-engine-5-raw focused on the code.
If you want to support me, you can here.
love-repro-aigve60k
LOVE / AIGVE-60K — reproduction results
Raw per-video outputs, subset manifests, and metrics from an independent reproduction of
LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text
Interpretation (ICML 2026 submission #1055, OpenReview P6fWeIVbwb).
Every model number here comes from running the authors' own released 9B checkpoints
through the authors' own evaluation code on 1× A100-80GB.
💻 Code + full writeup… See the full description on the dataset page: https://huggingface.co/datasets/adamcnoonan/love-repro-aigve60k.Fable-5-Max-Reasoning-Filtered-250x
Dataset Description
This dataset contains 25. highly detailed architectural traces mapping out security implementations for hybrid global banking systems encompassing both fiat and cryptocurrency infrastructures.
This is 10,000,000+ estimated tokens of fable 5 data, filtered and classified to remove low-quality entries by qwen 2.5 7B, and improved by GLM 5.2. The dataset bypasses basic conversational filler and is engineered to advance the domain precision, strict formatting… See the full description on the dataset page: https://huggingface.co/datasets/adamm-hf/Fable-5-Max-Reasoning-Filtered-250x.twitter_suicidal_risk
Twitter Suicide Risk Level Dataset
Short English tweets paired with a 0–4 suicide risk label, used for fine-tuning and
evaluating risk-level classification. This directory holds the final splits:
train.jsonl / val.jsonl / test.jsonl.
Files and size
File
Rows
Share
train.jsonl
7006
80%
val.jsonl
875
10%
test.jsonl
875
10%
Total
8756
100%
Fields
JSONL, one sample per line, three fields only:
Field
Type
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/AdamLeung/twitter_suicidal_risk.cgrt-consensus-5model
CGRT Consensus 5-Model Dataset
Multi-model consensus dataset for studying model agreement and disagreement patterns on mathematical reasoning tasks.
Dataset Description
61,678 math problems evaluated by 5 frontier LLMs with full reasoning traces and extracted answers.
Models Used
Model
Provider
Version
Claude
Anthropic
claude-3-5-sonnet-20241022
Codex/GPT-4
OpenAI
gpt-4o
Gemini
Google
gemini-1.5-flash
DeepSeek
DeepSeek
deepseek-chat
Qwen… See the full description on the dataset page: https://huggingface.co/datasets/Adam1010/cgrt-consensus-5model.youtube-titles
Youtube Title & Descriptions Dataset
About
4941 videos across 50 YouTube Channels
List of sampled channels here
Splits:
Train: 4199
Validation: 493
Test: 249
Data was shuffled and sampled evenly from all channels to create splits.
Additionally, has a column ready to go for gemma-2-9b-it fine tuning formatting! Potentially more model formats to come.
About the Data:
Label
Description
channel_name
The… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/youtube-titles.adam_pretrainemotion-balanced
Dataset Card for "emotion"
Dataset Summary
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
An example looks as follows.
{
"text": "im feeling quite sad… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/emotion-balanced.no_robots-alpaca
No Robots: Alpaca edition
This dataset is a cleaned (missing/extra spaces...) and reformatted version of the No Robots dataset from HuggingFaceH4, adapted to conform with the Alpaca instruction set.
Notably, it diverges from the original dataset in the way the 'Chat' category is handled; it has been decomposed into single-turn conversations to align with Alpaca's limitations regarding multi-turn interactions. The dataset's IDs have been generated using the SHA256 algorithm.… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/no_robots-alpaca.aprm-sft-thoughts-tau2-retail-policy_best-adamw30-lp0
Act-PRM SFT thoughts — tau2-bench retail
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-retail-policy_best-adamw30-lp0.NOAA-HRRR-CONUS-ImageCaptionAEZAKMI_v2
New things changed in v2
added some samples from spicy 3.1
removed samples that were sending a shiver down RPers spines!
removed some more GPTslop like "Remember to always..."
added some jokes and duplicated some samples with two system prompts - "A chat." and "A chat with uncensored assistant.", so that hopefully model connects those two and act more freely.
New things 2023-02-01
moved sharegpt version to a different repo to make it easier to use.
New things… See the full description on the dataset page: https://huggingface.co/datasets/adamo1139/AEZAKMI_v2.aprm-sft-thoughts-snorkel-insurance-policy_best-adamw30-lp0
Act-PRM SFT thoughts — snorkel-insurance insurance
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-insurance-policy_best-adamw30-lp0.aprm-sft-thoughts-tau2-airline-policy_best-adamw30-lp0
Act-PRM SFT thoughts — tau2-bench airline
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-airline-policy_best-adamw30-lp0.startup-advisor-dataset
🚀 Startup Advisor Dataset
A high-quality instruction-following dataset distilled from 8 foundational business and startup books, structured as actionable advice with real-world 2025 examples. Designed for fine-tuning large language models (e.g., Qwen, LLaMA, Mistral) to become expert startup advisors.
📖 Dataset Summary
Property
Value
Total Entries
1,564
Format
JSONL — ChatML (messages array)
Language
English
License
CreativeML OpenRAIL-M
Avg. Tokens… See the full description on the dataset page: https://huggingface.co/datasets/adamabuhamdan/startup-advisor-dataset.claudish-pairs
Claudish Pairs
The first open parallel corpus of English ↔ Claudish — the characteristic prose
style of Claude and Claude Code. 10,227 pairs, each an English text and its Claudish
restyling, authored and quality-controlled for faithfulness.
This is the v3 training set of
adamrotmil/claudish-style-adapter;
pipeline code at
github.com/adamrotmil/claudish-style-adapter.
Fields
Field
Meaning
english
source text (plain English)
claudish
the restyling… See the full description on the dataset page: https://huggingface.co/datasets/adamrotmil/claudish-pairs.mmmu-pro-clean
MMMU-Pro-Clean — exclusion overlay
A corrected drop-in for MMMU-Pro (standard 10-option split): 1,730 → 1,526 items, with 204 broken items removed.
⚠️ Overlay, not a rehost. MMMU-Pro is Apache-2.0, but its images come from exams/textbooks and carry third-party copyright, so this repo does NOT host the data or images. It ships the exclusion manifest — IDs, categories, tiers, the official-answer letter, coded reasons — which you apply to your own licensed MMMU/MMMU_Pro download.… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/mmmu-pro-clean.gpqa-extended-clean
GPQA-Extended-Clean — exclusion overlay
A corrected drop-in for the full 546-item GPQA-Extended: 546 → 498 items, with 48 broken items removed.
⚠️ Overlay, not a rehost. Per GPQA's anti-contamination norm, this repo ships the exclusion manifest — IDs, categories, tiers, coded reasons only — never item text. Apply it to your own licensed GPQA-Extended download.
📄 Paper: When the Answer Key Is Wrong — Allcock 2026 (arXiv forthcoming) · 💻 Loader + gated evidence:… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/gpqa-extended-clean.apple-environmental-report-QA-retrieval
Apple's 2024 Environmental Report QA Pairs
4300 question and relevant text chunks made from Apple's 2024 Environmental Report.
Chunking was done with a token based recursive chunker at 800 token chunk size with a 400 token overlap resulting in 215 chunks. 20 question labels per chunk were synthetically generated using gpt-4o-mini with the attached prompt and a temperature of 1.0.
Entries were shuffled and split into an 80/20 Train/Validation split resulting in:Training set size:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/apple-environmental-report-QA-retrieval.pii-masking-300k
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs.
Key facts:
OpenPII-220k text entries have 27 PII classes (types of sensitive data), targeting 749 discussion subjects / use cases split across education, health, and psychology. FinPII contains an additional ~20 types tailored to… See the full description on the dataset page: https://huggingface.co/datasets/AdamiTitus/pii-masking-300k.gpqa-diamond-clean
GPQA-Diamond-Clean — exclusion overlay
A corrected drop-in for GPQA Diamond: 198 → 189 items, with 9 broken items removed.
⚠️ This is an overlay, not a rehost. Per GPQA's anti-contamination norm (no crawlable plaintext), this repo ships the exclusion manifest — item IDs, categories, tiers, and coded reasons only — never GPQA item text or answer values. You apply it to your own licensed GPQA download.
📄 Paper: When the Answer Key Is Wrong — Allcock 2026 (arXiv forthcoming) ·… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/gpqa-diamond-clean.aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0
Act-PRM SFT thoughts — snorkel-finance finance
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-snorkel-finance-policy_best-adamw30-lp0.powershell_thestackgpqa-ext-complement-clean
GPQA-Extended-Complement-Clean — exclusion overlay
A corrected drop-in for the 348 GPQA-Extended items disjoint from Diamond: 348 → 309 items, with 39 broken items removed.
⚠️ Scope — read first. This is NOT canonical GPQA-Extended. Canonical Extended is the 546-item superset that includes Diamond. This release covers only the 348-item complement (Extended minus Diamond). Applying these exclusions to the full 546-item split, or treating 348 as "Extended", silently evaluates a… See the full description on the dataset page: https://huggingface.co/datasets/adamallcock/gpqa-ext-complement-clean.aprm-sft-thoughts-tau2-airline-base_best-adamw30-lp0
Act-PRM SFT thoughts — tau2-bench airline
Act-PRM (Action Process Reward Models) infers the latent thoughts behind
logged, action-only agent demonstrations via an offline EM. For each
logged action x in state s we sample G=4 candidate thoughts z,
score each by the length-penalized action likelihood
reward(z) = p(x | s, z)
(len_frac grows with the thought's token length), and mark the best thought
(argmax reward). The (thought + action) span is then what downstream SFT / RL… See the full description on the dataset page: https://huggingface.co/datasets/mzio/aprm-sft-thoughts-tau2-airline-base_best-adamw30-lp0.toxic-dpo-natural-v5I mixed in toxid-dpo-natural-v4 and rawrr v2-1 stage 2 with chosen field from original no_robots and got myself toxic-dpo-natural-v5. Goal is to avoid overfitting via DPO to a specific type of instruct, and instead just DPO the model to be more open to answering and also answer like a human being. We'll see whether this works.I trained Yi 34B with this dataset and ORPO, it does work very nicely so far!
