datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_sector_symbol_relations
Languages: 简体中文 · English
dojo_sector_symbol_relations — Stock–Sector Mapping
Overview
Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair.
Files
File
Description
data.parquet
Full stock ↔ sector relations
Key Fields
Field
Description
ticker
Stock symbol
market
us, cn, or hk
primary
JSON object — primary sector path
secondary
JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.synthetic-object-relations
Synthetic Object Relations Dataset
A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models.
Dataset Description
This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode:
Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.task391_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task391_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task391_causal_relationship.opengloss-v2.1-relations
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Relations
The OpenGloss v2.1 semantic graph as an edge list. The relations config holds every live typed edge — fourteen… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-relations.opengloss-v2.3-relations
OpenGloss v2.3 — Relations
The OpenGloss v2.3 semantic graph as an edge list. The relations config holds every live typed edge — fourteen relation types — with the target resolved to a sense id wherever the target's entry exists in the release, which is what makes this a sense graph rather than a word graph. The tombstoned config recovers the edges the free reconcile pass demoted, deduplicated or capped away, with the type they carried when they were removed and the reason… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-relations.formal-logic-simple-order-multi-token-dynamic-objects-paired-relationship-0-100000opengloss-v2.0-relations
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Relations
The OpenGloss v2.0 semantic graph as an edge list. The relations config holds every live typed edge — fourteen relation types — with the target resolved to a sense id wherever the target's entry exists in the release, which is what makes… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-relations.task970_sherliic_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task970_sherliic_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task970_sherliic_causal_relationship.opengloss-v2.2-relations
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Relations
The OpenGloss v2.2 semantic graph as an edge list. The relations config holds every live typed edge — fourteen relation types — with the target resolved to a sense id wherever the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-relations.spatial-relations-with-degreesformal-logic-simple-order-multi-token-dynamic-objects-paired-relationship-0-40000task732_mmmlu_answer_generation_public_relations
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task732_mmmlu_answer_generation_public_relations
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task732_mmmlu_answer_generation_public_relations.pbs_criteria_parameter_relationshipspbs_item_atc_relationshipsformal-logic-simple-order-multi-token-fixed-objects-paired-relationship-0-40000pbs_item_restriction_relationshipstask908_dialogre_identify_familial_relationships
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task908_dialogre_identify_familial_relationships
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task908_dialogre_identify_familial_relationships.processed_narrative_relationship_dataset
Dataset Card for "processed_narrative_relationship_dataset"
More Information needed
dojo_sector_symbol_relations
Languages: 简体中文 · English
dojo_sector_symbol_relations — Stock–Sector Mapping
Overview
Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair.
Files
File
Description
data.parquet
Full stock ↔ sector relations
Key Fields
Field
Description
ticker
Stock symbol
market
us, cn, or hk
primary
JSON object — primary sector path
secondary
JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/dojo_sector_symbol_relations.two-hop-5-relations-join-finalformal-logic-simple-order-token-objects-paired-relationship-0-40000cppo_continual_dataset_rl_relationshipspbs_item_organisation_relationshipssynthetic-object-relations-jsoncppo_continual_dataset_reward_relationshipstwo-hop-5-relations1_R1_2_R2_3
1_R1_2_R4_4
5_R3_2_R2_3
5_R3_2_R4_4
6_R5_2_R2_3
6_R5_2_R4_4
{
'R1': 'sibling_of',
'R2': 'born_in_location',
'R3': 'work_with',
'R4': 'study_in',
'R5': 'friend_of'
}
project-gutenberg-fiction-relations
Project Gutenberg Fiction Relations
A literary-domain relation extraction (RE) dataset built from public-domain fiction in
Project Gutenberg. Each example pairs a passage of narrative text (mentioning a head and
tail entity) with the relation that holds between the two entities, providing an RE resource for
literary and digital-humanities research where general-domain (news / Wikipedia) datasets do not
transfer well.
This dataset is released as part of the paper "Sub-Billion… See the full description on the dataset page: https://huggingface.co/datasets/Despina/project-gutenberg-fiction-relations.pbs_restriction_prescribing_text_relationshipstwo-hop-5-relations-disjoint-finalformal-logic-simple-order-multi-token-dynamic-objects-paired-relationship-0-5000
