datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pevo-msa-grch38-19way
pevo-msa-grch38-19way (dataset Hub)
EN: Training data, project tables, and reproducibility artifacts for primate MSA variant-effect modeling.
中文: 灵长类 MSA 变异效应建模的训练数据与项目复现材料(不含模型权重)。
Results-status note (2026-07-29). The multi-seed metrics reported below are retained as historical registry records. They use an earlier scoring protocol and cohort convention, and are not comparable to the corrected strict-v2 results used for the final project conclusions. Do not use the values… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/pevo-msa-grch38-19way.vepyr_116_GRCh38_plugin_clinvar
vepyr plugin cache — ClinVar (GRCh38, VEP 116)
A prebuilt, frequency-tiered Parquet cache of ClinVar clinical-significance
annotations for use with vepyr, the Rust/DataFusion
VEP-compatible variant annotation engine. It reproduces the CSQ output of Ensembl VEP 116
run with ClinVar as a --custom annotation, without requiring the upstream VCF at
annotation time.
Source version
This is the fact you most likely came here for.
Source file
clinvar.vcf.gz… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_plugin_clinvar.vepyr_116_GRCh38_plugin_alphamissense
vepyr plugin cache — AlphaMissense (GRCh38, VEP 116)
A prebuilt, frequency-tiered Parquet cache of AlphaMissense pathogenicity
predictions for use with vepyr, the Rust/DataFusion
VEP-compatible variant annotation engine. It reproduces the CSQ output of Ensembl
VEP 116's --plugin AlphaMissense without requiring the upstream TSV or the Perl plugin
at annotation time.
Source version
This is the fact you most likely came here for.
Source file… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_plugin_alphamissense.vepyr_116_GRCh38_plugin_spliceai
vepyr plugin cache — SpliceAI (GRCh38, VEP 116)
A prebuilt, frequency-tiered Parquet cache of SpliceAI splice-altering predictions for
use with vepyr, the Rust/DataFusion VEP-compatible
variant annotation engine. It reproduces the CSQ output of Ensembl VEP 116's
--plugin SpliceAI without requiring the upstream 400 GB-class VCF or the Perl plugin at
annotation time.
Source version
This is the fact you most likely came here for.
Source file… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_plugin_spliceai.GRCh38A dataset of all autosomal and sex chromosomes sequences from reference assembly GRCh38/hg38 1 and reached a total of 3.2 billion nucleotides.vepyr_116_GRCh38_plugin_phenotypeorthologous
vepyr PhenotypeOrthologous plugin cache — Ensembl 116, GRCh38
Prebuilt vepyr plugin cache for Ensembl VEP's
PhenotypeOrthologous
plugin: phenotypes of the rat and mouse orthologues of the Ensembl gene overlapping a variant.
Built with vepyr.build_plugin_cache("phenotypeorthologous", "v0.2.0 (7736bbead98eb57e13f88e5074fa2b7d2e844121)") from the
vepyr-plugins
manifest, against the release-116 GRCh38 merged variation cache.
Source
File… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_plugin_phenotypeorthologous.vepyr_116_GRCh38_ensembl
vepyr cache — Ensembl VEP 116, GRCh38 (Ensembl)
A Parquet conversion of the Ensembl VEP 116 homo_sapiens_ensembl cache for GRCh38, for
use with vepyr, a Rust/DataFusion variant annotation
engine that reproduces VEP's consequence calls. It replaces VEP's Perl-serialised,
gzipped cache files with columnar Parquet that DuckDB, Polars, DataFusion or Spark can read
directly.
Ensembl/GENCODE transcripts only. This is the default VEP cache flavour. The equivalent VEP invocation uses… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_ensembl.vepyr_116_GRCh38_plugin_cadd
vepyr plugin cache — CADD v1.7 (GRCh38, VEP 116)
A prebuilt, frequency-tiered Parquet cache of CADD deleteriousness scores for use with
vepyr, the Rust/DataFusion VEP-compatible variant
annotation engine. It reproduces the CSQ output of Ensembl VEP 116's --plugin CADD
without requiring the upstream whole-genome TSVs or the Perl plugin at annotation time.
⚠️ Non-commercial use only. See Licence below.
Source version
This is the fact you most likely came here for.… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_plugin_cadd.vepyr_116_GRCh38_merged
vepyr cache — Ensembl VEP 116, GRCh38 (merged)
A Parquet conversion of the Ensembl VEP 116 homo_sapiens_merged cache for GRCh38, for
use with vepyr, a Rust/DataFusion variant annotation
engine that reproduces VEP's consequence calls. It replaces VEP's Perl-serialised,
gzipped cache files with columnar Parquet that DuckDB, Polars, DataFusion or Spark can read
directly.
Both Ensembl/GENCODE and NCBI RefSeq transcripts, in one cache. The equivalent VEP invocation uses --merged.
⚠️… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_merged.vepyr_116_GRCh38_refseq
vepyr cache — Ensembl VEP 116, GRCh38 (RefSeq)
A Parquet conversion of the Ensembl VEP 116 homo_sapiens_refseq cache for GRCh38, for
use with vepyr, a Rust/DataFusion variant annotation
engine that reproduces VEP's consequence calls. It replaces VEP's Perl-serialised,
gzipped cache files with columnar Parquet that DuckDB, Polars, DataFusion or Spark can read
directly.
NCBI RefSeq transcripts only. The equivalent VEP invocation uses --refseq.
⚠️ This cache is not free for… See the full description on the dataset page: https://huggingface.co/datasets/biodatageeks/vepyr_116_GRCh38_refseq.alphagenome-avi-grch38
AlphaGenome AVI GRCh38 — prepared VCF mirror
An unofficial, losslessly reformatted mirror of Google DeepMind's AVI SNV
scores, prepared for GUIDE-IEI, and can be used directly for Ensembl VEP custom annotation.
Google DeepMind produced the predictions; GUIDE-IEI performed only the format
conversion described below. This mirror is not affiliated with or endorsed by
Google DeepMind.
Source and terms
Official AlphaGenome downloads
Original AVI SNV ZIP
AlphaGenome… See the full description on the dataset page: https://huggingface.co/datasets/luoyiming1991/alphagenome-avi-grch38.vep113-grch38-reference-bundle
VEP 113 / GRCh38 reference bundle — cache, FASTA, LOFTEE data
A fast mirror of the reference data needed to run Ensembl VEP release 113
offline on GRCh38, in the prepared forms a local annotation pipeline needs —
assembled for the WES/WGS diagnostic analysis pipeline (IEI variant-review
workbench), whose downloader verifies every file against pinned SHA-256
checksums and falls back to the canonical sources automatically.
Mirror rationale: the canonical servers often serve well… See the full description on the dataset page: https://huggingface.co/datasets/luoyiming1991/vep113-grch38-reference-bundle.GRCh38_sequences250111_191630-grch38_hg38This repository contains the data protocols of "MethylProphet: A Generalized Gene-Contextual Model for Inferring Whole-Genome DNA Methylation Landscape".
Detailed instructions can be found at https://github.com/xk-huang/methylprophet/blob/main/docs/CUSTOMIZED.md.
GCF-GRCh38-p14-genomic
Dataset Summary
This dataset is a processed, tabular representation of the NCBI RefSeq GFF file for the human reference genome GRCh38.p14. The original GFF (General Feature Format) data, GCF_000001405.40_GRCh38.p14_genomic.gff, has been converted to the Parquet format for efficient storage and fast querying using libraries like Pandas and Apache Arrow.
The data contains annotations for all genomic features (genes, transcripts, exons, CDS, etc.) on the primary sequences of the GRCh38… See the full description on the dataset page: https://huggingface.co/datasets/SunnyLin/GCF-GRCh38-p14-genomic.Human-genome-CDS-GRCh38These are DNA coding sequences in the human genome build GRCh38, downloaded from ensembl with the following R script:
# install biomartr 1.0.7 from CRAN
install.packages("biomartr", dependencies = TRUE)
# Install Biostrings if not installed
if (!requireNamespace("BiocManager", quietly = TRUE)) {
install.packages("BiocManager")
}
# Load required package
library(Biostrings)
library(biomartr)
# download the genome of Homo sapiens from ensembl
# and store the corresponding genome CDS file in… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/Human-genome-CDS-GRCh38.Homo_sapiens.GRCh38.dna.primary_assembly.tar.gzhuman_grch38_segment_sample
Human GRCh38 Genome Segments
Dataset Description
This dataset contains 16,384 base pair segments from the human reference genome (GRCh38) prepared for Sparse Autoencoder (SAE) training with the Evo2 model. The segments are extracted using a sliding window approach with 75% overlap.
Dataset Details
Total segments: 718,648
Segment size: 16,384 base pairs
Stride: 4,096 base pairs (75% overlap)
Source genome: GRCh38.primary_assembly (GENCODE Release 41)… See the full description on the dataset page: https://huggingface.co/datasets/harari/human_grch38_segment_sample.grch38-bwa-indexhomo_sapiens_merged_vep_115_GRCh38.tar.gzGCF-GRCh38-genes
