datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dojo_sector_symbol_relations
Languages: 简体中文 · English
dojo_sector_symbol_relations — Stock–Sector Mapping
Overview
Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair.
Files
File
Description
data.parquet
Full stock ↔ sector relations
Key Fields
Field
Description
ticker
Stock symbol
market
us, cn, or hk
primary
JSON object — primary sector path
secondary
JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/AlphaDojo/dojo_sector_symbol_relations.factnet_relations
FactNet Relations Dataset
Overview
The Synset Relations dataset contains rich semantic relationships between FactSynsets, enabling advanced reasoning and cross-lingual fact retrieval. These relations capture hypernymy, causality, temporality, geographic relationships, and other semantic connections between facts.
Paper: https://arxiv.org/abs/2602.03417
Github: https://github.com/yl-shen/factnet
Dataset: https://huggingface.co/collections/openbmb/factnet… See the full description on the dataset page: https://huggingface.co/datasets/openbmb/factnet_relations.Relationship_chartsynthetic-object-relations
Synthetic Object Relations Dataset
A synthetic image dataset generated with Flux Schnell featuring clean object-relation prompts designed for training spatial reasoning in vision and diffusion models.
Dataset Description
This dataset contains images generated from structured prompts describing spatial relationships between objects. Unlike typical caption datasets that use free-form text, our prompts follow consistent patterns that explicitly encode:
Object identities… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/synthetic-object-relations.task970_sherliic_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task970_sherliic_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task970_sherliic_causal_relationship.opengloss-v2.1-relations
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Relations
The OpenGloss v2.1 semantic graph as an edge list. The relations config holds every live typed edge — fourteen… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-relations.opengloss-v2.3-relations
OpenGloss v2.3 — Relations
The OpenGloss v2.3 semantic graph as an edge list. The relations config holds every live typed edge — fourteen relation types — with the target resolved to a sense id wherever the target's entry exists in the release, which is what makes this a sense graph rather than a word graph. The tombstoned config recovers the edges the free reconcile pass demoted, deduplicated or capped away, with the type they carried when they were removed and the reason… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-relations.task391_causal_relationship
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task391_causal_relationship
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task391_causal_relationship.formal-logic-simple-order-multi-token-dynamic-objects-paired-relationship-0-100000opengloss-v2.0-relations
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Relations
The OpenGloss v2.0 semantic graph as an edge list. The relations config holds every live typed edge — fourteen relation types — with the target resolved to a sense id wherever the target's entry exists in the release, which is what makes… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-relations.opengloss-v2.2-relations
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Relations
The OpenGloss v2.2 semantic graph as an edge list. The relations config holds every live typed edge — fourteen relation types — with the target resolved to a sense id wherever the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-relations.all-relations-lead-to-rome
Description
All Relations Lead to Rome (ARLtR) includes a knowledge graph made in Neo4j, and multiple QA-pairs.
The dataset is supplementary material for the paper "All Relations Lead to Rome: Automated Knowledge Graph Creation and Question Generation"
The knowledge graph is stored in a dump file, while the QA-pairs are stored in different csv files. Make sure to follow the setup guide to get started.
Setup
Import the dump in Neo4j
Host an instance of the Neo4j… See the full description on the dataset page: https://huggingface.co/datasets/FaynePro/all-relations-lead-to-rome.Fuzzy_Spatial_Relationship_Dataset
Learning from Ambiguity: A Fuzzy Spatial Relationship Dataset for Human-Aligned Text-to-Image Generation
Tianjiao Liang, Qinlong Li, Honggang Qi
Official dataset card for "Learning from Ambiguity: A Fuzzy Spatial Relationship Dataset for Human-Aligned Text-to-Image Generation", submitted to The Visual Computer.
📢 Dataset Release Status
The FSRD dataset is currently being released progressively on Hugging Face.
Due to the large scale of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/NIEYEHH/Fuzzy_Spatial_Relationship_Dataset.task732_mmmlu_answer_generation_public_relations
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task732_mmmlu_answer_generation_public_relations
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task732_mmmlu_answer_generation_public_relations.task908_dialogre_identify_familial_relationships
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task908_dialogre_identify_familial_relationships
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task908_dialogre_identify_familial_relationships.formal-logic-simple-order-multi-token-dynamic-objects-paired-relationship-0-40000pbs_item_atc_relationshipsformal-logic-simple-order-multi-token-fixed-objects-paired-relationship-0-40000hir-relationships
Hardware Integration Relationships (HIR)
A labeled dataset and multi-tier benchmark for cross-artifact relationship prediction in hardware
designs: given a hardware artifact (a register or register field) and a firmware-visible artifact
(a generated C #define), does a real relationship exist between them?
This is the concept-matching problem at the heart of SoC integration review — "is
dma.CONTROL.EN the same thing the firmware calls DMA_CONTROL_EN_MASK?" — cast as a supervised… See the full description on the dataset page: https://huggingface.co/datasets/tapeout-labs/hir-relationships.processed_narrative_relationship_dataset
Dataset Card for "processed_narrative_relationship_dataset"
More Information needed
Relationship_Advice
💟 Relationship Advice dataset card
Dataset Description
Dataset Summary
The Relationship Advice dataset is an English-language compilation of posts and their respective comments concerning dating and human romantic relationships. The primary objective of this dataset is to aid LLMs in categorizing responses and providing appropriate answers based on the emotional needs expressed by the writer.
The data was gathered from two subreddits: r/dating_advice… See the full description on the dataset page: https://huggingface.co/datasets/yonatanko/Relationship_Advice.pbs_item_restriction_relationshipsspatial-relations-with-degreesCODE-ACCORD-Relations
CODE-ACCORD: A Corpus of Building Regulatory Data for Rule Generation towards Automatic Compliance Checking
The CODE-ACCORD corpus contains annotated sentences from the building regulations of England and Finland and has been developed as part of the Horizon European project for Automated Compliance Checks for Construction, Renovation or Demolition Works (ACCORD). The corpus is in English, and it consists of both the English Building Regulations and the English translation of the… See the full description on the dataset page: https://huggingface.co/datasets/ACCORD-NLP/CODE-ACCORD-Relations.dojo_sector_symbol_relations
Languages: 简体中文 · English
dojo_sector_symbol_relations — Stock–Sector Mapping
Overview
Maps each stock to L1/L2/L3 sector paths with primary and secondary assignments. One row per (ticker, market) pair.
Files
File
Description
data.parquet
Full stock ↔ sector relations
Key Fields
Field
Description
ticker
Stock symbol
market
us, cn, or hk
primary
JSON object — primary sector path
secondary
JSON array —… See the full description on the dataset page: https://huggingface.co/datasets/vessel888/dojo_sector_symbol_relations.pbs_criteria_parameter_relationshipsconceptnet_en2en_relations
Dataset Description
This is a subset of the conceptnet5 dataset.
I merely parsed and extracted out my required portion and uploaded here, since processing the huge complete dataset is complicated for many users.
Please refer to the original authors' repo for a complete version.
ConceptNet is a multilingual knowledge base, representing words and
phrases that people use and the common-sense relationships between
them. The knowledge in ConceptNet is collected from a variety of… See the full description on the dataset page: https://huggingface.co/datasets/appledora/conceptnet_en2en_relations.news_relationshipsformal-logic-simple-order-token-objects-paired-relationship-0-40000spatial-relations
