datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MSA_PretrainData
MSA Pretrain Data
Retrieval-style pretraining corpora. Each subset is split into two parts:
file
columns
meaning
<subset>/queries/*.parquet
question, answer, reference_ids: list<int64>, labels: list<int64>
query, plus row indices into the subset's reference table
<subset>/references/*.parquet
value: string
the reference/memory passage text
reference_ids are the candidate pool for a query; labels are the positive(s).
Both are global row indices into the subset's… See the full description on the dataset page: https://huggingface.co/datasets/Anoy123423123/MSA_PretrainData.ar-quran-hadith14books-MSA
ar-quran-hadith14books-MSA
Arabic speech for both primary sources of Islam — Quran and Hadith — plus cleaned general
Modern Standard Arabic, under one construction pipeline and one text convention.
ASR errors on sacred text are not ordinary errors: a plausible-sounding substitution can alter
the meaning of scripture, and because chatbots, search and summarizers increasingly answer from
transcriptions rather than from audio, such an error propagates silently. Quranic recitation… See the full description on the dataset page: https://huggingface.co/datasets/Dr-AliGomaa/ar-quran-hadith14books-MSA.us-layoffs-by-metro-area-msa-warn-act
US layoffs by metro area: 54,170 WARN notices mapped to 765 metro and micro areas
Rebuilt 2026-09-22. 765 of the 935 US core-based statistical areas carry at least one
layoff notice on record — 361 metropolitan and 404 micropolitan.
Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the
Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish
the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.FineWeb2-MSA
FineWeb2 MSA Arabic
This is the MSA Arabic Portion of The FineWeb2 Dataset.
This dataset contains a rich collection of text in MSA Arabic (ISO 639-3: arz), a widely spoken dialect within the Afro-Asiatic language family.
With over 439 million words and 1.4 million documents, it serves as a valuable resource for NLP development and linguistic research focused on Egyptian Arabic.
Purpose of This Repository
This repository provides easy access to the Arabic portion - MSA… See the full description on the dataset page: https://huggingface.co/datasets/Omartificial-Intelligence-Space/FineWeb2-MSA.msa-hotpotqa-qa-with-idstinystories_phonologymsa-musique-qa-with-idsmsa-hotpotqa-docs-with-idsmcp-sandbox-authority-boundary-profile
MCP Sandbox Authority Boundary Profile
Profile v0.1.0 · Release v0.2.0 - Experimental Characterization Profile
Profile release date: 2026-07-23
Latest distribution release date: 2026-09-05
Execution containment is not proof of bounded authority.
Start here
For a one-minute, case-by-case reading of the profile, open the companion
Authority Boundary Field Guide Space.
It presents the released synthetic observations with their control question,
observed result… See the full description on the dataset page: https://huggingface.co/datasets/msaleme/mcp-sandbox-authority-boundary-profile.msa-omnivoice-tts-v1
MSA-OmniVoice-v1
Dataset Description
MSA-OmniVoice-v1 is a 50-hour synthetic Modern Standard Arabic (MSA) speech dataset generated using OmniVoice. The dataset contains high-quality synthetic speech from a single speaker paired with fully diacritized (تشكيل) transcripts.
It is intended for training and fine-tuning Arabic speech models, including Text-to-Speech (TTS), Automatic Speech Recognition (ASR), speech representation learning, and alignment tasks.… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/msa-omnivoice-tts-v1.msa-2wikimultihopqa-qa-with-idsagripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
gpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
msavbench-videoarabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.msa-musique-qa-with-idsmsa-musique-docs-with-idsmsa-2wikimultihopqa-qa-with-idsgpn-msa-microglia-fullMSA_train_setPre-processed MSA data based on https://huggingface.co/datasets/mozilla-foundation/common_voice_16_1.
MSA-nuc-9-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-9-seq.MSA-amino-9-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-amino-9-seq.PottsMPNN_Training_MSAs
Paired & Monomer MSA Database
This repository contains multiple sequence alignments (MSAs) and the mapping
files needed to associate them with PDB complexes.
Contents
File
Size
Description
PDB_paired_msas.tar.zst.part_aa … part_af
~10.7 GB each
Sharded, zstd-compressed tar archive of the paired PDB MSAs (.a3m files).
cath42_msas.tar.gz
13.3 GB
Gzip-compressed tar archive of the CATH 4.2 MSAs.
cid_mapping.pkl
7.84 MB
Dictionary mapping PDB chain IDs… See the full description on the dataset page: https://huggingface.co/datasets/FosterBirnbaum/PottsMPNN_Training_MSAs.MSA-nuc-8-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-8-seq.MSA-nuc-7-seq
Multiple Sequence Alignment as a Sequence-to-Sequence Learning Problem
Abstract:
The sequence alignment problem is one of the most fundamental problems in bioinformatics and a plethora of methods were devised to tackle it. Here we introduce BetaAlign, a methodology for aligning sequences using an NLP approach. BetaAlign accounts for the possible variability of the evolutionary process among different datasets by using an ensemble of transformers, each trained on millions… See the full description on the dataset page: https://huggingface.co/datasets/dotan1111/MSA-nuc-7-seq.msa-hotpotqa-qa-with-idskorean-english-multitarget-ted-talks-task
Dataset Card for english-korean-multitarget-ted-talks-task
Dataset Summary
Parallel English-Korean Text Corpus
Text was originally transcribed to English from various Ted Talks, then translated to Korean by TED translators
Approximately 166k train, 2k validation, and 2k test sentence pairs.
Supported Tasks and Leaderboards
Machine Translation
Languages
English
Korean
Additional Information
Dataset Curators
Kevin Duh, "The… See the full description on the dataset page: https://huggingface.co/datasets/msarmi9/korean-english-multitarget-ted-talks-task.tunisian-msa-parallel-corpus
Dataset Description
This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models.
The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.QnA_Descriptivemsa-musique-docs-with-ids
