datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
human_genomehuman_genome_csv
Human Genome Dataset
Here is a human genome ready to be used to train LLM.
human_genome_GCF_009914755.1
Dataset Card for "human_genome_GCF_009914755.1"
how to build this data:
human full genome data from:
https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_009914755.1/
Preprocess:
1 download data use ncbi data set tools:
curl -o datasets 'https://ftp.ncbi.nlm.nih.gov/pub/datasets/command-line/LATEST/linux-amd64/datasets'
chmod +x datasets
./datasets download genome accession GCF_000001405.40 --filename genomes/human_genome_dataset.zip
then move the gene data to human2.fra
2 write the… See the full description on the dataset page: https://huggingface.co/datasets/dnagpt/human_genome_GCF_009914755.1.InstaDeepAI_human_reference_genomehuman_reference_genomeGenome Reference Consortium Human Build 38 patch release 14 (GRCh38.p14)
filtered and split into chunks.human-genome
Homo Sapiens Genome [GRCh37]
Human Genome Assembly GRCh37.p13
[NCBI Dataset]
[File Store]
NCBI RefSeq assembly
GCF_000001405.25 (replaced)
Submitted GenBank assembly
GCA_000001405.14 (replaced)
Taxon
Homo sapiens (human)
Synonym
hg19
Assembly type
haploid with alt loci
Submitter
Genome Reference Consortium
Date
Jun 28, 2013
© The Human Genome Project, currently maintained by the Genome Reference Consortium (GRC)human-genome
GRCh38 — Ensembl release 115 soft-masked primary assembly
The Ensembl release 115 GRCh38 soft-masked primary assembly, provided in
uncompressed and BGZF-compressed FASTA formats. Both variants include indexes
for remote random-access sequence queries without downloading the entire genome.
Files
File
Purpose
Homo_sapiens.GRCh38.dna_sm.primary_assembly.fa
Uncompressed FASTA, optimized for remote byte-range queries… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/human-genome.human_genome_gnomAD_doped_sequences_v3.7Human_Genome_Embedding_Collectionhuman_genome_gnomAD_doped_sequences_v3.1human_genome_gnomAD_doped_sequences_v3.5human_genome_gnomad_primate_doped_sequences_v4human-genome-cdsSource: human reference genome
Filtering: CDS + 256 bp flanks
Data augmentation: windows of 512 bp, with 256 step size as well as reverse complements
human_genome_gnomAD_doped_sequences_v3.4Human-genome-CDS-GRCh38These are DNA coding sequences in the human genome build GRCh38, downloaded from ensembl with the following R script:
# install biomartr 1.0.7 from CRAN
install.packages("biomartr", dependencies = TRUE)
# Install Biostrings if not installed
if (!requireNamespace("BiocManager", quietly = TRUE)) {
install.packages("BiocManager")
}
# Load required package
library(Biostrings)
library(biomartr)
# download the genome of Homo sapiens from ensembl
# and store the corresponding genome CDS file in… See the full description on the dataset page: https://huggingface.co/datasets/MichelNivard/Human-genome-CDS-GRCh38.human_genome_context_8096human_genome_gnomAD_doped_sequences_v3human-genome-sequencesfull_human_genomehuman_genome_variants_hnm
Summary Statistics:
Total Variants: 115,969
