dwb2023/hetionet-edges
Dataset Card for Hetionet Dataset Overview This dataset represents an integrative biomedical knowledge graph, constructed from 29 public resources, encoding relationships between various biomedical entities. It is primarily designed for drug repurposing, treatment prediction, and network-based biomedical research. Original Data Source: The edge list is derived from the original Hetionet GitHub repository. Acknowledgment: Full credit goes to the original authors… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/hetionet-edges.
Dataset Card for Hetionet
Dataset Overview
This dataset represents an integrative biomedical knowledge graph, constructed from 29 public resources, encoding relationships between various biomedical entities. It is primarily designed for drug repurposing, treatment prediction, and network-based biomedical research.
- Original Data Source: The edge list is derived from the original Hetionet GitHub repository.
- Acknowledgment: Full credit goes to the original authors for their contributions.
Dataset Details
Dataset Description
Hetionet is a heterogeneous biomedical graph that integrates genes, compounds, diseases, pathways, biological processes, molecular functions, cellular components, pharmacologic classes, side effects, and symptoms into a structured network.
The Hetionet Edges Dataset specifically captures the relationships (edges) between biomedical entities.
- Number of Nodes (Entities): 47,031 (across 11 types)
- Number of Edges (Relationships): 2,250,197 (across 24 metaedge types)
- Data Sources: 29 public biomedical resources
Dataset Attribution
- Curators: Daniel Scott Himmelstein, Antoine Lizee, Christine Hessler, Leo Brueggeman, Sabrina L Chen, Dexter Hadley, Ari Green, Pouya Khankhanian, Sergio E Baranzini
- Language: English
- License: CC-BY-4.0
Dataset Sources
- Primary Repository: neo4j.het.io
- Publication: Systematic integration of biomedical knowledge prioritizes drugs for repurposing (eLife, 2017)
- Demo Application: het.io/repurpose
Intended Uses
Appropriate Use Cases
This dataset can be used for:
- Drug Repurposing Research: Identifying new uses for existing drugs.
- Treatment Prediction: Modeling biomedical relationships to predict treatment outcomes.
- Biomedical Knowledge Integration: Aggregating multiple datasets into a structured knowledge graph.
- Network Analysis of Biomedical Relationships: Exploring connectivity patterns between genes, diseases, and compounds.
- Computational Drug Efficacy Prediction: Using machine learning to assess potential drug efficacy.
Limitations and Out-of-Scope Use Cases
This dataset should not be used as:
- A stand-alone clinical decision-making tool without external validation.
- A replacement for experimental research or clinical trials.
- An authoritative guide for medical treatment recommendations.
Dataset Structure
Features
The dataset is formatted as a TSV file (tab-separated values) with three primary features (columns):
Metadata Summary
- Total unique source nodes: 20,138
- Total unique target nodes: 44,204
- Total unique metaedge types: 24
- Most common metaedge:
GpBP(Gene participates in Biological Process) with 559,504 edges.
Metaedge Descriptions (Relationship Types)
Dataset Creation
Curation Rationale
This dataset was created to improve drug repurposing research and computational drug efficacy prediction by leveraging heterogeneous biomedical relationships. The dataset integrates 755 known drug-disease treatments, supporting network-based reasoning for drug discovery.
- 🚨 Last Update: The edge list (hetionet-v1.0-edges.sif.gz) was last modified 7 years ago, and the node list (hetionet-v1.0-nodes.tsv) was last modified 9 years ago.
- Implications: While the dataset is a valuable resource, some biomedical relationships may be outdated, as new drugs, pathways, and gene-disease links continue to be discovered.
Source Data
Data Collection and Processing
- Data was aggregated from 29 public biomedical resources and integrated into a heterogeneous network.
- Community Feedback: The project incorporated real-time input from 40 community members to refine the dataset.
- Formats: The source dataset is available in TSV (tabular), JSON, and Neo4j formats.
Data Provenance & Last Update
Who are the source data producers?
- The dataset integrates biomedical knowledge from 29 public resources.
- Curated by: The University of California, San Francisco (UCSF) research team.
- Primary repository: Hetionet GitHub.
Bias, Risks, and Limitations
Key Limitations Due to Dataset Age
- 🚨 Data Recency: The dataset was last updated 7–9 years ago, meaning some relationships may be outdated due to advances in biomedical research.
- Network Incompleteness: Biomedical knowledge evolves, and newer discoveries are not reflected in this dataset.
- Bias in Source Data: Public biomedical databases have inherent biases based on what was known at the time of their last update.
Key Recommendations
✅ Users should verify relationships against more recent biomedical datasets. ✅ Use the JSON or Neo4j formats if metadata (license, attribution) is needed. ✅ Cross-check with external databases such as DrugBank, KEGG, or CTD. ✅ Consider integrating newer biomedical datasets for up-to-date analysis.
Citation
BibTeX:
@article {10.7554/eLife.26726,
article_type = {journal},
title = {Systematic integration of biomedical knowledge prioritizes drugs for repurposing},
author = {Himmelstein, Daniel Scott and Lizee, Antoine and Hessler, Christine and Brueggeman, Leo and Chen, Sabrina L and Hadley, Dexter and Green, Ari and Khankhanian, Pouya and Baranzini, Sergio E},
editor = {Valencia, Alfonso},
volume = 6,
year = 2017,
month = {sep},
pub_date = {2017-09-22},
pages = {e26726},
citation = {eLife 2017;6:e26726},
doi = {10.7554/eLife.26726},
url = {https://doi.org/10.7554/eLife.26726},
journal = {eLife},
issn = {2050-084X},
publisher = {eLife Sciences Publications, Ltd}
}Additional Citations
Heterogeneous Network Edge Prediction: A Data Integration Approach to Prioritize Disease-Associated Genes
Himmelstein DS, Baranzini SE
PLOS Computational Biology (2015)
DOI: https://doi.org/10.1371/journal.pcbi.1004259 · PMID: 26158728 · PMCID: PMC4497619Dataset Card Authors
dwb2023
Dataset Card Contact
dwb2023
