datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task035_winogrande_question_modification_person
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task035_winogrande_question_modification_person
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task035_winogrande_question_modification_person.task034_winogrande_question_modification_object
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task034_winogrande_question_modification_object
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task034_winogrande_question_modification_object.task776_pawsx_japanese_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task776_pawsx_japanese_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task776_pawsx_japanese_text_modification.task770_pawsx_english_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task770_pawsx_english_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task770_pawsx_english_text_modification.DeepMath-103K
DeepMath-103K
🔥 News
May 8, 2025: We found that 48 samples contained hints that revealed the answers. The relevant questions have now been revised to remove the leaked answers.
April 14, 2025: We release DeepMath-103K, a large-scale dataset featuring challenging, verifiable, and decontaminated math problems tailored for RL and SFT. We open source:… See the full description on the dataset page: https://huggingface.co/datasets/modibboali/DeepMath-103K.task121_zest_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task121_zest_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task121_zest_text_modification.task1622_disfl_qa_text_modication
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1622_disfl_qa_text_modication
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1622_disfl_qa_text_modication.task1562_zest_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1562_zest_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1562_zest_text_modification.task1670_md_gender_bias_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1670_md_gender_bias_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1670_md_gender_bias_text_modification.task1669_md_gender_bias_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1669_md_gender_bias_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1669_md_gender_bias_text_modification.TRM-modified-datamix-tokenized
TRM modified datamix (tokenized)
Pre-tokenized reasoning/pretraining mixture for from-scratch TRM (Tiny Recursive Model)
training, built by running data_io — the HRM-Text data
pipeline — verbatim on sapientinc/HRM-Text-data-io-cleaned-20260515, with three
deliberate, documented deviations (below).
It is emitted in the V1 tokenized dataset format (a single concatenated token pool +
per-epoch document indices) and is ready to stream directly into training — no re-tokenization.… See the full description on the dataset page: https://huggingface.co/datasets/m-ric/TRM-modified-datamix-tokenized.grade_school_math_modified
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/re2panda/grade_school_math_modified.modified-codesearchnet-code-summarization
Modified CodeSearchNet (MCSN) Dataset
This dataset is a modification of the CodeSearchNet dataset from CodeXGLUE benchmark, designed for evaluating code summarization models beyond the function level. It explores the impact of function and repository contexts on summary quality. The dataset includes modifications for evaluating at both function and repository levels.
Paper: Code Summarization Beyond Function Level
Dataset Structure:
The dataset contains samples with the following… See the full description on the dataset page: https://huggingface.co/datasets/sm1rk/modified-codesearchnet-code-summarization.NuminaMath-CoT
Dataset Card for NuminaMath CoT
Dataset Summary
Approximately 860k math problems, where each solution is formatted in a Chain of Thought (CoT) manner. The sources of the dataset range from Chinese high school math exercises to US and international mathematics olympiad competition problems. The data were primarily collected from online exam paper PDFs and mathematics discussion forums. The processing steps include (a) OCR from the original PDFs, (b) segmentation into… See the full description on the dataset page: https://huggingface.co/datasets/modibboali/NuminaMath-CoT.modified-classeval-code-summarization
Modified ClassEval (MCE) Dataset
This dataset is a modification of the ClassEval benchmark, designed for evaluating code summarization models beyond the function level. It explores the impact of function and class contexts on summary quality. The dataset includes modifications for evaluating at both function and class levels.
Paper: Code Summarization Beyond Function Level
Dataset Structure:
The dataset contains samples with the following fields:
class_id: Identifier for the… See the full description on the dataset page: https://huggingface.co/datasets/sm1rk/modified-classeval-code-summarization.task132_dais_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task132_dais_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task132_dais_text_modification.rnacentral-modifications
RNAcentral
RNAcentral is a free, public resource that offers integrated access to a comprehensive and up-to-date set of non-coding RNA sequences provided by a collaborating group of Expert Databases representing a broad range of organisms and RNA types.
The development of RNAcentral is coordinated by European Bioinformatics Institute and is supported by Wellcome. Initial funding was provided by BBSRC.
Disclaimer
This is an UNOFFICIAL release of the RNAcentral by The… See the full description on the dataset page: https://huggingface.co/datasets/multimolecule/rnacentral-modifications.distill_r1_110k_sft_modifiedBorrowed from https://huggingface.co/datasets/Congliu/Chinese-DeepSeek-R1-Distill-data-110k-SFT
Fix the <image> placeholder issue, which will cause error during training:
raise ValueError(f"The number of images does not match the number of {IMAGE_PLACEHOLDER} tokens.")
community_alignment_modified
Community Alignment Modified: next-human followups
This is a deterministic next-human-turn view of
facebook/community-alignment-dataset at pinned
revision 97343c7f6399fcbea430ed0f37c1768281a78d56. It contains 2,514 eligible
conversations from 90,256 source rows.
Five fixed rows are published only as fewshot_demonstrations. Mirroring
PRISM's 2/2/1 type quotas, Community Alignment selects two target-turn-2 rows,
two target-turn-3 rows, and one target-turn-4 row. Targets contain… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/community_alignment_modified.sharktank_pitches_modified
Shark Tank Structured Pitch-to-Text Dataset
This dataset contains 245 examples of sales pitches from the TV Show "Shark Tank", scraped from Youtube (mostly) officialy channel. Its primary feature is the mapping between a highly structured JSON object (the "input") and a complete, conversational sales pitch (the "output").
The dataset is designed for structured-data-to-text generation tasks.
🚀 Supported Tasks & Use Cases
This dataset is ideal for benchmarking modern… See the full description on the dataset page: https://huggingface.co/datasets/isaidchia/sharktank_pitches_modified.thoughttrace_modified
ThoughtTrace Modified
A next-human-turn benchmark derived from
SCAI-JHU/ThoughtTrace
at pinned revision 0420f3d8499e477098aac7771fe9c066f2340fb3.
Non-negotiable target contract
Every scored target is copied from a source message whose type is exactly
user, immediately following a source message whose type is exactly
assistant. The assistant message is conditioning context, never the target:
Human: previous human message
Assistant: source LLM response
Human:… See the full description on the dataset page: https://huggingface.co/datasets/Alberto1231/thoughttrace_modified.Psych8k-Modifiedtask775_pawsx_chinese_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task775_pawsx_chinese_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task775_pawsx_chinese_text_modification.task773_pawsx_spanish_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task773_pawsx_spanish_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task773_pawsx_spanish_text_modification.task774_pawsx_german_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task774_pawsx_german_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task774_pawsx_german_text_modification.task771_pawsx_korean_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task771_pawsx_korean_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task771_pawsx_korean_text_modification.task772_pawsx_french_text_modification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task772_pawsx_french_text_modification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task772_pawsx_french_text_modification.
