datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
spreadsheet-bench-v2-modified
SpreadsheetBench V2 Modified: Multi-Document QA
1,060 questions and reference answers grounded in 127 Excel workbooks, 35 PDFs and 9 DOCX files. This independent derivative of SpreadsheetBench 2 shifts the task from editing spreadsheets and producing workbook deliverables toward finding, interpreting and combining information in business documents.
An independent project built entirely from publicly available source material and newly authored QA annotations. No private company… See the full description on the dataset page: https://huggingface.co/datasets/hashmortar/spreadsheet-bench-v2-modified.TRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.modified_dataset_emotionMoDiffPlease refer to https://github.com/WeizhiGao/MoDiff for the usage of this dataset. Our paper is available at https://huggingface.co/papers/2506.22463.
grade_school_math_modified
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/re2panda/grade_school_math_modified.ActivityNet_Captions_Modified
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/sankim2/ActivityNet_Captions_Modified.modificacoesNarendra-Modidistill_r1_110k_sft_modifiedBorrowed from https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT
Fix the <image> placeholder issue, which will cause error during training:
raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.")
sharktank_pitches_modified
Shark Tank Structured Pitch-to-Text Dataset
This dataset contains 245 examples of sales pitches from the TV Show "Shark Tank", scraped from Youtube (mostly) officialy channel. Its primary feature is the mapping between a highly structured JSON object (the "input") and a complete, conversational sales pitch (the "output").
The dataset is designed for structured-data-to-text generation tasks.
🚀 Supported Tasks & Use Cases
This dataset is ideal for benchmarking modern… See the full description on the dataset page: https://huggingface.co/datasets/isaidchia/sharktank_pitches_modified.Nemotron-Math-HumanReasoningDataset accompanying the paper: The Challenge of Teaching Reasoning to LLMs Without RL or Distillation.
All experiments were conducted using NeMo-Skills.
Dataset Description:
Nemotron-Math-HumanReasoning is a compact dataset featuring human-written solutions to math problems, designed to emulate the extended reasoning style of models like DeepSeek-R1. These solutions are authored by students with deep experience in Olympiad-level mathematics. We provide multiple versions of the… See the full description on the dataset page: https://huggingface.co/datasets/modibboali/Nemotron-Math-HumanReasoning.Malicious_Modifier_Dataset_MMD
Malicious Modifier Dataset (MMD)
Part of the ModX project:
Modifier Unlocked: Jailbreaking Text-to-Image Models Through PromptsShuofeng Liu, Mengyao Ma, Minhui Xue, Guangdong BaiIEEE S&P 2025
Dataset Description
The Malicious Modifier Dataset (MMD) is a curated collection of 717 prompt modifiers that have been identified as capable of adjusting text-to-image (T2I) model outputs toward NSFW (Not-Safe-for-Work) content. The modifiers are collected through art-related… See the full description on the dataset page: https://huggingface.co/datasets/cola-hunter/Malicious_Modifier_Dataset_MMD.recipe-modifications
Hebrew Recipe Modification Dataset
Overview
10,058 Hebrew comment threads from YouTube cooking channels, annotated for recipe modification extraction using a three-pass Teacher-Student distillation approach.
Task
Token-level BIO tagging to extract recipe modifications from Hebrew user comments. Four modification aspects: SUBSTITUTION, QUANTITY, TECHNIQUE, ADDITION.
Dataset Structure
Raw Data
threads.jsonl — 10,058 comment threads (top… See the full description on the dataset page: https://huggingface.co/datasets/DanielDDDS/recipe-modifications.modified_erotic_literature_collection-alpacaData converted to JSON from https://huggingface.co/datasets/li-long/modified_erotic_literature_collection
pos_hotel_neg_res_modifiedAPIGen-MT-5k-modifiedmodified-alpacaaime-2024-modifiedmodified_58k_unfilteredexample_modified_quotesGenMed-hidoc-kin25k-intstruction_modifiedUSPTO2_modifiedRAFT_output_modifiedgsm8k_modifiedqwen-robot-dataset-v2-modifyalpaca-modified
