datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MSA_PretrainData
MSA Pretrain Data
Retrieval-style pretraining corpora. Each subset is split into two parts:
file
columns
meaning
<subset>/queries/*.parquet
question, answer, reference_ids: list<int64>, labels: list<int64>
query, plus row indices into the subset's reference table
<subset>/references/*.parquet
value: string
the reference/memory passage text
reference_ids are the candidate pool for a query; labels are the positive(s).
Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.pevo-msa-grch38-19way
pevo-msa-grch38-19way (dataset Hub)
EN: Training data, project tables, and reproducibility artifacts for primate MSA variant-effect modeling.
中文: 灵长类 MSA 变异效应建模的训练数据与项目复现材料(不含模型权重)。
Results-status note (2026-07-29). The multi-seed metrics reported below are retained as historical registry records. They use an earlier scoring protocol and cohort convention, and are not comparable to the corrected strict-v2 results used for the final project conclusions. Do not use the values… See the full description on the dataset page: https://huggingface.co/datasets/jasperyeoh2/pevo-msa-grch38-19way.human-proteome-wide-msa
Human Proteome ColabFold MSAs
This dataset contains ColabFold/MMseqs2 multiple sequence alignments and AlphaFold 3 JSON inputs for the human proteome query set.
Contents
a3m/shard-*/: one A3M file per protein accession, sharded to stay below repository directory file limits.
af3_json/shard-*/: one AlphaFold 3 JSON file per protein accession, sharded to stay below repository directory file limits.
manifest.tsv: tab-separated index with accession, metadata… See the full description on the dataset page: https://huggingface.co/datasets/LiteFold/human-proteome-wide-msa.ar-quran-hadith14books-MSA
ar-quran-hadith14books-MSA
Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general
Modern Standard Arabic, under one construction pipeline and one text convention.
ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter
the meaning of scripture, and because chatbots, search and summarizers increasingly answer from
transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.gpn-msa-hg38-scores
GPN-MSA predictions for all possible SNPs in the human genome (~9 billion)
For more information check out our paper and repository.
Querying specific variants or genes
Install the latest tabix:In your current conda environment (might be slow):conda install -c bioconda -c conda-forge htslib=1.18
or in a new conda environment:conda create -n tabix -c bioconda -c conda-forge htslib=1.18
conda activate tabix
Query a specific region (e.g. BRCA1), from the remote file:… See the full description on the dataset page: https://huggingface.co/datasets/songlab/gpn-msa-hg38-scores.msam-released-products
MSAM released flight products
This dataset contains the complete numeric contents of the three MSAM1 flight
archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation
tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38
configuration identifiers preserve the flight directory and source filename
stem.
How to use
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.us-layoffs-by-metro-area-msa-warn-act
US layoffs by metro area: 54,166 WARN notices mapped to 765 metro and micro areas
Rebuilt 2026-09-21. 765 of the 935 US core-based statistical areas carry at least one
layoff notice on record — 361 metropolitan and 404 micropolitan.
Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the
Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish
the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.msa-hotpotqa-qa-with-idsFineWeb2-MSA
FineWeb2 MSA Arabic
This is the MSA Arabic Portion of The FineWeb2 Dataset.
This dataset contains a rich collection of text in MSA Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family.
With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic.
Purpose of This Repository
This repository provides easy access to the Arabic portion - MSA… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-MSA.ParaDLC-Bench
ParaDLC-Bench
ParaDLC-Bench (Parallel Detailed Localized Captioning Benchmark) is a benchmark for multi-region localized captioning that jointly evaluates caption quality and inference efficiency. It extends DLC-Bench from single-region evaluation to concurrent multi-region evaluation, explicitly stressing a model's ability to describe many regions at once without cross-region interference.
📄 Paper |
💻 Code |
🤖 PerceptionDLM
Key… See the full description on the dataset page: https://huggingface.co/datasets/MSALab/ParaDLC-Bench.tinystories_phonologymsa-musique-qa-with-idsclr_motion_planning_hw_7msa-hotpotqa-docs-with-idsmsa-omnivoice-tts-v1
MSA-OmniVoice-v1
Dataset Description
MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts.
It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.mcp-sandbox-authority-boundary-profile
MCP Sandbox Authority Boundary Profile
Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile
Profile release date: 2026-07-23
Latest distribution release date: 2026-09-05
Execution containment is not proof of bounded authority.
Start here
For a one-minute, case-by-case reading of the profile, open the companion
Authority Boundary Field Guide Space.
It presents the released synthetic observations with their control question,
observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.agripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
msa-2wikimultihopqa-qa-with-idsgpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
msa-lecturemsavbench-videoarabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.msa-2wikimultihopqa-qa-with-idsgpn-msa-microglia-fullmsa-musique-docs-with-idsMSA_train_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1.
msa-musique-qa-with-idsMSA-nuc-9-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.msap-loop-ctl-20260728MSA_OOD_Dataset_in_CIDerHere are the MSA OOD datasets mentioned in the CIDer paper.
Please cite our paper if you find that useful for your research:
@article{zhong2025towards,
title={Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts},
author={Zhong, Guowei and Huan, Ruohong and Wu, Mingzhen and Liang, Ronghua and Chen, Peng},
journal={arXiv preprint arXiv:2506.10452},
year={2025}
}
