datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
helium-market-resolution-benchmark
What is this?
Market Resolution contains 299 ranked option-contract questions. Most test calculations or comparisons from frozen quotes. Sixty test an implied-volatility prior with the premium hidden, and 11 test probability-of-finishing-in-the-money forecasts against a later outcome.
It does not measure trading profitability. It measures bounded option-chain reasoning: implied volatility (IV), delta, time value, parity, term structure, relative IV, chain surfaces, and… See the full description on the dataset page: https://huggingface.co/datasets/HeliumTrades/helium-market-resolution-benchmark.task891_gap_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task891_gap_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task891_gap_coreference_resolution.task893_gap_fill_the_blank_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task893_gap_fill_the_blank_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task893_gap_fill_the_blank_coreference_resolution.task892_gap_reverse_coreference_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task892_gap_reverse_coreference_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task892_gap_reverse_coreference_resolution.task304_numeric_fused_head_resolution
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task304_numeric_fused_head_resolution
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task304_numeric_fused_head_resolution.customer-ticket-resolution
CUSTOMER_TICKET_RESOLUTION
A preference dataset for CUSTOMER_TICKET_RESOLUTION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally ticket)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from… See the full description on the dataset page: https://huggingface.co/datasets/316usman/customer-ticket-resolution.Parameter-Golf-V7-Resolution-Layout-Compression-Reasoning
Parameter-Golf-V7-Resolution-Layout-Compression-Reasoning
Overview
This is a synthetic auxiliary dataset for OpenAI Parameter Golf style experiments. It is not a replacement for FineWeb. It is designed to test whether a tiny auxiliary mix can reduce wasted next-token probability around compression reasoning, web-page signal filtering, display-vs-payload confusion, and multimodal preview planning.
The central lesson is deliberately narrow:
Smaller rendered objects can… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/Parameter-Golf-V7-Resolution-Layout-Compression-Reasoning.event-outcome-resolution
EVENT_OUTCOME_RESOLUTION
A preference dataset for EVENT_OUTCOME_RESOLUTION, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from… See the full description on the dataset page: https://huggingface.co/datasets/316usman/event-outcome-resolution.github-issue-pr-resolution-10k
GitHub Issue→PR Resolution Dataset — 10K Sample
This is a 10,000-record sample. The full dataset (~354,000 records) is available for purchase. Contact andrewcampi456@gmail.com for licensing details.
What is this?
A large-scale dataset of resolved GitHub issues paired with their linked pull requests, code diffs, commits, and review discussions, all from production-quality Python repositories.
Each record captures the complete lifecycle of a software fix:
Issue… See the full description on the dataset page: https://huggingface.co/datasets/acampi23/github-issue-pr-resolution-10k.clinical_instruction_conflict_resolution_v0.1Clinical Instruction Conflict Resolution
Tests whether a model resolves conflicting instructions safely.
Output JSON
resolved
conflict_type
correct_action
Runpython scorer.py --predictions predictions.jsonl --test_csv data/test.csv
instruction_conflict_resolution_v01Instruction Conflict Resolution v0.1
This evaluation dataset tests how models resolve conflicting instructions.
It targets a common failure mode: following the most recent or most forceful instruction even when it conflicts with higher-priority constraints.
This is not training data.
What it tests
Priority handling under instruction conflict
Refusal stability under escalation
Logical conflict handling for impossible constraints
Post-conflict integrity with no delayed leakage… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/instruction_conflict_resolution_v01.Conflic_Resolution_Content_2
Conflic Resolution Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Conflic_Resolution_Content_2.Conflict_Resolution_Content_1
Conflict Resolution Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Conflict_Resolution_Content_1.
