CoolFace
Datasetpublic

OneScience-Group/evo2_dataset

Evo2 Dataset Dataset Description Evo2 Dataset is an Evo2 mini genome dataset adapted for OneScience/evo2/. It contains FASTA, compressed FASTA, and merged FASTA files for human chr20, chr21, and chr22, as well as train, validation, and test .bin/.idx splits preprocessed with the Byte-Level tokenizer. Supported Tasks This dataset is not the complete OpenGenome2 dataset and is not intended to reproduce full-scale pretraining. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/OneScience-Group/evo2_dataset.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes36downloads
Dataset Card

<p align="center"> <strong> <span style="font-size: 30px;">Evo2 Dataset</span> </strong> </p>

Dataset Description

Evo2 Dataset is an Evo2 mini genome dataset adapted for OneScience/evo2/. It contains FASTA, compressed FASTA, and merged FASTA files for human chr20, chr21, and chr22, as well as train, validation, and test .bin/.idx splits preprocessed with the Byte-Level tokenizer.

Supported Tasks

This dataset is not the complete OpenGenome2 dataset and is not intended to reproduce full-scale pretraining. It is intended for preflight checks of the standard Evo2 model repository, FASTA prediction input, and smoke tests for mini training or fine-tuning.

Dataset Format and Structure

DataFormatDescription
Raw sequencesFASTA / gzip FASTAchr20, chr21, chr22, and merged FASTA
Preprocessed training split.bin + .idxchr20_21_22_uint8_distinct_byte-level_train
Preprocessed validation split.bin + .idxchr20_21_22_uint8_distinct_byte-level_val
Preprocessed test split.bin + .idxchr20_21_22_uint8_distinct_byte-level_test
Checksum manifestTSV / SHA256File names, sizes, and hashes

The default location within the dataset package is data_mini/genome_data/.

How to Use the Dataset

Files and Download

Download the dataset:

bash
hf download --dataset OneScience-Group/evo2_dataset

Model weights must be downloaded separately to validate the model:

bash
hf download --model OneScience-Group/evo2/

Data Placement and Validation

Default model data path:

text
<MODEL_REPO_ROOT>/data/evo2_dataset/data_mini/genome_data/

Place data_mini in the model package:

bash
mkdir -p <MODEL_REPO_ROOT>/data/evo2_dataset
cp -a data_mini <MODEL_REPO_ROOT>/data/evo2_dataset/data_mini

Validation command:

bash
python scripts/validate_evo2_dataset.py --package-root . --dataset-root data_mini

For a quick structure-only check, add --skip-sha256.

Official OneScience Information

PlatformOneScience Main RepositorySkills Repository
Giteehttps://gitee.com/onescience-ai/onesciencehttps://gitee.com/onescience-ai/oneskills
GitHubhttps://github.com/onescience-ai/OneSciencehttps://github.com/onescience-ai/oneskills

Limitations and License

Repository license: Apache-2.0 Data license: LicenseRef-UCSC-hg38-public-use Data source: UCSC hg38 / GRCh38, chromosomes 20, 21 and 22 Note: Derived bin/idx files follow the source genome data terms.