CoolFace
20 results

msa

Anoy123423123 /MSA_PretrainData MSA Pretrain Data Retrieval-style pretraining corpora. Each subset is split into two parts: file columns meaning <subset>/queries/*.parquet question, answer, reference_ids: list<int64>, labels: list<int64> query, plus row indices into the subset's reference table <subset>/references/*.parquet value: string the reference/memory passage text reference_ids are the candidate pool for a query; labels are the positive(s). Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.texttext-retrieval10M<n<100M0 likes7.2k downloads2mo agoHugging Facejasperyeoh2 /pevo-msa-grch38-19way pevo-msa-grch38-19way (dataset Hub) EN: Training data, project tables, and reproducibility artifacts for primate MSA variant-effect modeling. 中文: 灵长类 MSA 变异效应建模的训练数据与项目复现材料(不含模型权重)。 Results-status note (2026-07-29). The multi-seed metrics reported below are retained as historical registry records. They use an earlier scoring protocol and cohort convention, and are not comparable to the corrected strict-v2 results used for the final project conclusions. Do not use the values… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/pevo-msa-grch38-19way.0 likes6.7k downloads14d agoHugging FaceLiteFold /human-proteome-wide-msa Human Proteome ColabFold MSAs This dataset contains ColabFold/MMseqs2 multiple sequence alignments and AlphaFold 3 JSON inputs for the human proteome query set. Contents a3m/shard-*/: one A3M file per protein accession, sharded to stay below repository directory file limits. af3_json/shard-*/: one AlphaFold 3 JSON file per protein accession, sharded to stay below repository directory file limits. manifest.tsv: tab-separated index with accession, metadata… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/human-proteome-wide-msa.other2 likes1.2k downloads16d agoHugging FaceDr-AliGomaa /ar-quran-hadith14books-MSA ar-quran-hadith14books-MSA Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general Modern Standard Arabic, under one construction pipeline and one text convention. ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter the meaning of scripture, and because chatbots, search and summarizers increasingly answer from transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.audioautomatic-speech-recognition10K<n<100K6 likes1.1k downloads1mo agoHugging Facesonglab /gpn-msa-hg38-scores GPN-MSA predictions for all possible SNPs in the human genome (~9 billion) For more information check out our paper and repository. Querying specific variants or genes Install the latest tabix:In your current conda environment (might be slow):conda install -c bioconda -c conda-forge htslib=1.18 or in a new conda environment:conda create -n tabix -c bioconda -c conda-forge htslib=1.18 conda activate tabix Query a specific region (e.g. BRCA1), from the remote file:… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-msa-hg38-scores.5 likes955 downloads2y agoHugging Faceastro-legacy-archive /msam-released-products MSAM released flight products This dataset contains the complete numeric contents of the three MSAM1 flight archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38 configuration identifiers preserve the flight directory and source filename stem. How to use python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.tabular10K<n<100K0 likes690 downloads17d agoHugging Face