msa
Datasets
All datasets matching “msa”MSA_PretrainData
MSA Pretrain Data
Retrieval-style pretraining corpora. Each subset is split into two parts:
file
columns
meaning
<subset>/queries/*.parquet
question, answer, reference_ids: list<int64>, labels: list<int64>
query, plus row indices into the subset's reference table
<subset>/references/*.parquet
value: string
the reference/memory passage text
reference_ids are the candidate pool for a query; labels are the positive(s).
Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.pevo-msa-grch38-19way
pevo-msa-grch38-19way (dataset Hub)
EN: Training data, project tables, and reproducibility artifacts for primate MSA variant-effect modeling.
中文: 灵长类 MSA 变异效应建模的训练数据与项目复现材料(不含模型权重)。
Results-status note (2026-07-29). The multi-seed metrics reported below are retained as historical registry records. They use an earlier scoring protocol and cohort convention, and are not comparable to the corrected strict-v2 results used for the final project conclusions. Do not use the values… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/pevo-msa-grch38-19way.human-proteome-wide-msa
Human Proteome ColabFold MSAs
This dataset contains ColabFold/MMseqs2 multiple sequence alignments and AlphaFold 3 JSON inputs for the human proteome query set.
Contents
a3m/shard-*/: one A3M file per protein accession, sharded to stay below repository directory file limits.
af3_json/shard-*/: one AlphaFold 3 JSON file per protein accession, sharded to stay below repository directory file limits.
manifest.tsv: tab-separated index with accession, metadata… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/human-proteome-wide-msa.ar-quran-hadith14books-MSA
ar-quran-hadith14books-MSA
Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general
Modern Standard Arabic, under one construction pipeline and one text convention.
ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter
the meaning of scripture, and because chatbots, search and summarizers increasingly answer from
transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.gpn-msa-hg38-scores
GPN-MSA predictions for all possible SNPs in the human genome (~9 billion)
For more information check out our paper and repository.
Querying specific variants or genes
Install the latest tabix:In your current conda environment (might be slow):conda install -c bioconda -c conda-forge htslib=1.18
or in a new conda environment:conda create -n tabix -c bioconda -c conda-forge htslib=1.18
conda activate tabix
Query a specific region (e.g. BRCA1), from the remote file:… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-msa-hg38-scores.msam-released-products
MSAM released flight products
This dataset contains the complete numeric contents of the three MSAM1 flight
archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation
tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38
configuration identifiers preserve the flight directory and source filename
stem.
How to use
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.
