datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
itis-taxonomy-instruct-30k-v2-negatives
ITIS Taxonomy Instruction Dataset with Negative Samples
Overview
The ITIS Taxonomy Instruction Dataset with Negative Samples is a structured instruction-response dataset derived from the public domain Integrated Taxonomic Information System (ITIS) database.
It was designed for fine-tuning large language models on taxonomy-oriented tasks such as rank identification, lineage reconstruction, parent taxon retrieval, taxonomic validity checks, and common name mapping.… See the full description on the dataset page: https://huggingface.co/datasets/Jaymerry/itis-taxonomy-instruct-30k-v2-negatives.task927_yelp_negative_to_positive_style_transfer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task927_yelp_negative_to_positive_style_transfer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task927_yelp_negative_to_positive_style_transfer.clinical_narrative_negative_evidence_handling_v0.4Clinical Narrative Negative Evidence Handling v0.4
Purpose
Test whether a model handles negative evidence without narrative spin.
This version adds
timeline steps
cross trial negative carryover
suppression pressure prompts
explicit evidence status and submission positioning
Input columns
data_anchor
negative_pressures
draft_narrative
audience
timeline_step
Model task
Return one JSON object
negative_flagslist of short labels
evidence_statusexploratory, mixed, negative… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical_narrative_negative_evidence_handling_v0.4.scram-curriculum
scram-curriculum: the training set behind Scram-0.8B
17,837 verified teacher traces. This is the complete supervised fine-tuning set used
to train negativevoid/Scram-0.8B-6bit,
published so the recipe is checkable rather than merely described.
How it was made
Traces were generated by Qwen3.5-4B and Qwen3.5-9B, then kept only if the final
answer exactly matched gold. Roughly 7% of generations were discarded by that filter.
Verification is final-answer-only: no one… See the full description on the dataset page: https://huggingface.co/datasets/negativevoid/scram-curriculum.task928_yelp_positive_to_negative_style_transfer
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task928_yelp_positive_to_negative_style_transfer
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task928_yelp_positive_to_negative_style_transfer.error-detection-negatives
error-detection-negatives
This dataset is part of the PARC (Premise-Annotated Reasoning Collection) and contains mathematical reasoning problems with error annotations. This dataset combines negatives samples from multiple domains.
Domain Breakdown
gsm8k: 57 samples
math: 44 samples
metamathqa: 59 samples
orca_math: 54 samples
Features
Each example contains:
data_source: The domain/source of the problem (gsm8k, math, metamathqa, orca_math)
question: The… See the full description on the dataset page: https://huggingface.co/datasets/PARC-DATASETS/error-detection-negatives.swipe-negatives
Swipe Hard Negatives
Hard-negative word sets for swipe-keyboard language model training, mined
by KNN over swipe-shape embeddings from the SwipeALot encoder. For each
positive word, the top K = 128 most-confusable words
are listed alongside neg_sims — the mean pairwise
sample-to-sample cosine similarity between samples of the positive word
and samples of the negative word.
These negatives are designed to be "hard" because they share visual
swipe-path similarity with the positive… See the full description on the dataset page: https://huggingface.co/datasets/futo-org/swipe-negatives.sob-ft-false-negatives
SOB FT False Negatives
Subset of mariem123kfg/sob-ft-finetune-ready where the candidate JSON has no actual error, but labels claim an error.
Definition used
A row is included when all of the following hold:
true_has_error == True
repair_candidate (or errored_json) equals validated_output (canonical JSON compare)
finetuning_target.has_error == True and corrected_json equals the already-correct candidate
These look like failed / cancelled-out error injections:… See the full description on the dataset page: https://huggingface.co/datasets/mariem123kfg/sob-ft-false-negatives.MSMarco_Negative_1k
MS MARCO Negative 1k
This dataset contains 1,000 random examples sampled from microsoft/ms_marco with added negative_query and generated negative_ans columns.
Source dataset: microsoft/ms_marco
Source subset/split: v1.1/train
Document used for negative query generation: first selected passage when available, otherwise first non-empty passage
Negative query types: 500 explicit_negation, 500 antonym
Negative answer generation model: gpt-4o
Rows written: 1000
Destination repo:… See the full description on the dataset page: https://huggingface.co/datasets/canho/MSMarco_Negative_1k.
