datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
msam-released-products
MSAM released flight products
This dataset contains the complete numeric contents of the three MSAM1 flight
archives served by LAMBDA: the June 1992, June 1994, and June 1995 observation
tables, sampled beam maps, and released 1994/1995 covariance matrices. The 38
configuration identifiers preserve the flight directory and source filename
stem.
How to use
python -m venv .venv && .venv/bin/pip install datasets huggingface_hub pyarrow
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/msam-released-products.us-layoffs-by-metro-area-msa-warn-act
US layoffs by metro area: 54,170 WARN notices mapped to 765 metro and micro areas
Rebuilt 2026-09-22. 765 of the 935 US core-based statistical areas carry at least one
layoff notice on record — 361 metropolitan and 404 micropolitan.
Nobody hires, sells or reports by county. A recruiter covers Austin; an account team books the
Phoenix metro; a reporter writes Bay Area layoffs. State agencies publish neither — they publish
the site of a layoff as free text in 48 different… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-metro-area-msa-warn-act.agripotentialMore information and competition link:
https://github.com/MohammadElSakka/agripotential
https://www.codabench.org/competitions/12055/
https://zenodo.org/records/15551829
gpn-msa-sapiens-dataset
Training windows for GPN-MSA-Sapiens
For more information check out our paper and repository.
Path in Snakemake:
results/dataset/multiz100way/89/128/64/True/defined.phastCons.percentile-75_0.05_0.001
arabic-msa-25k-saudi-male-tashkeel
Arabic MSA 25K — Saudi Male (Tashkeel)
25,000 fully-diacritized Arabic MSA text + audio pairs, rendered with a single
Saudi male neural voice at 48 kHz / 16-bit PCM, across 10 thematic categories.
Dataset Summary
arabic-msa-25k-saudi-male-tashkeel is a 25,000-clip Modern Standard Arabic (MSA)
speech corpus with matching diacritized text (full tashkeel / ḥarakāt). Every clip
is synthesized by the single voice ar-SA-HamedNeural (Azure Neural TTS, Saudi
Arabic male) at 48… See the full description on the dataset page: https://huggingface.co/datasets/HeshamHaroon/arabic-msa-25k-saudi-male-tashkeel.gpn-msa-microglia-fulltunisian-msa-parallel-corpus
Dataset Description
This is an ambitious project to create a high-quality, reproducible parallel corpus for Modern Standard Arabic (MSA) and Tunisian Arabic (aeb) through a sophisticated synthetic data generation pipeline. The dataset is being developed by the Tunisia.AI community to address the scarcity of high-quality dialectal data for training and evaluating language models.
The primary goal is to provide a rich, well-documented resource for the research and development of:… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus.edisum_dataset
Dataset Card for Edisum
Dataset Description
For more details:
Github repository: https://github.com/epfl-dlab/edisum
Paper: https://arxiv.org/pdf/2404.03428.pdf
Languages
Edisum only contains Wikipedia data collected from English Wikipedia. Consequently, synthetic data is also only generated in English.
Dataset Structure
The Edisum meta-dataset actually comprises 5 datasets:
wikiepdia_processed_data (Filtered existing Wikipedia data)… See the full description on the dataset page: https://huggingface.co/datasets/msakota/edisum_dataset.afdb-msa-index
AlphaFold DB minimizer index
A static sequence-search index over 239,602,633 AlphaFold DB v6 entries
(99.4% of the database), designed to be queried directly from a browser with
HTTP Range requests. No server, no search engine, no MMseqs2.
It exists to answer one question quickly: which AFDB entry is ≥90% identical to
my sequence? — so that entry's precomputed MSA can be borrowed and re-indexed
onto the query instead of computing a new alignment.
Used by AFDB MSA.… See the full description on the dataset page: https://huggingface.co/datasets/sokrypton/afdb-msa-index.clr_motion_planning_hwvn_msa_sar_v3
VN-MSASar v3 — Multilingual Aspect-Based Sentiment + Sarcasm (hotel & restaurant reviews)
Chia split ở mức review. 100% review thật — toàn bộ dữ liệu sinh bằng LLM của
bản đầu (nguồn augmented, 7.659 dòng rải ở cả ba split) đã bị loại bỏ.
Không tương thích ngược với phamluan/vn_sar_msa. 71% câu trong test nằm trong
train của bản cũ. Mọi checkpoint huấn luyện trên bản cũ đều đã nhìn thấy test này —
phải huấn luyện lại từ đầu.
Splits
Split
Dòng
Review
Mỉa… See the full description on the dataset page: https://huggingface.co/datasets/phamluan/vn_msa_sar_v3.tunisian-msa-parallel-corpus-evaluated
Dataset Description
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb).
It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP.
The primary goals are to support:
Machine translation between Tunisian Arabic and MSA.
Research in dialectal-aware text generation and evaluation.
Cross-dialect representation learning in… See the full description on the dataset page: https://huggingface.co/datasets/tunis-ai/tunisian-msa-parallel-corpus-evaluated.record-bi-11This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so100_follower",
"total_episodes": 2,
"total_frames": 569,
"total_tasks": 1,
"total_videos": 6,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-11.eval_2-fold-pants-20This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so100_follower",
"total_episodes": 1,
"total_frames": 1758,
"total_tasks": 1,
"total_videos": 1,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/eval_2-fold-pants-20.vn_msa_sar
Dataset Card: Cleaned Multilingual Aspect-Based Sentiment & Sarcasm Reviews
Tóm tắt Dataset (Dataset Summary)
Dataset này chứa 83.009 đánh giá (reviews) đa ngôn ngữ đã qua quá trình làm sạch (cleaned). Dữ liệu được thu thập từ các nền tảng như Google Maps, Booking, Foody và dữ liệu tăng cường (augmented), tập trung vào các địa điểm kinh doanh dịch vụ (Khách sạn và Nhà hàng). Mỗi đánh giá được gán nhãn chi tiết phục vụ cho các bài toán NLP như: Phân tích cảm xúc… See the full description on the dataset page: https://huggingface.co/datasets/phamluan/vn_msa_sar.record-bi-100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so100_follower",
"total_episodes": 20,
"total_frames": 35446,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-100.analogy_concept_questions
Analogy Concept Questions
This dataset was generated using YourBench (v0.6.0), an open-source framework for generating domain-specific benchmarks from document collections.
Pipeline Steps
ingestion: Read raw source documents, convert them to normalized markdown and save for downstream steps
summarization: Perform hierarchical summarization: chunk-level LLM summaries followed by combine-stage reduction
chunking: Split texts into token-based single-hop and multi-hop… See the full description on the dataset page: https://huggingface.co/datasets/msaramhassan/analogy_concept_questions.record-bi-7This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so100_follower",
"total_episodes": 20,
"total_frames": 35867,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-7.marine_vla_dataset_v2
Marine VLA Dataset — v2 (Fixed Expert)
This dataset contains 12,070 frames collected in the Mangalia harbor Gazebo simulation,
used to fine-tune a vision-language-action (VLA) model for marine navigation.
Built with opencode + DeepSeek V4 Flash — every component
of this project (simulation setup, expert controller, data pipeline, ML training,
inference nodes) was developed using opencode CLI with DeepSeek V4 Flash as the
underlying model.
v2 Changes (Fixed Expert)… See the full description on the dataset page: https://huggingface.co/datasets/MSaalaamaa/marine_vla_dataset_v2.record-bi-3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_so100_follower",
"total_episodes": 2,
"total_frames": 393,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MSakai/record-bi-3.Siamese_Finetune_MSA_SOW_ContractsA dataset prepared for siamese finetuning, to distingush between texts from Legal Contracts text (Majorly SOW, MSA others Legal Algreements, Offer Letters etc)
and text scraped from books, news articles, reviews etc
license: apache-2.0
avian-msa-iqtreeparise-tts-6h-taggedjenny-tts-tags-6htunisian-msa-parallel-corpus-evaluated
Dataset Description
This dataset is a synthetic parallel corpus of Tunisian Arabic (aeb) and Modern Standard Arabic (arb).
It was created with a rigorous multi-stage pipeline to maximize quality and reproducibility, addressing the scarcity of high-quality resources for Tunisian Arabic NLP.
The primary goals are to support:
Machine translation between Tunisian Arabic and MSA.
Research in dialectal-aware text generation and evaluation.
Cross-dialect representation learning in Arabic… See the full description on the dataset page: https://huggingface.co/datasets/hbenayed/tunisian-msa-parallel-corpus-evaluated.jenny-tts-6h-taggedfold-pants-20vi_msa_sar
VN-SarMSA-vi
Vietnamese aspect-level sentiment (7-point, −3…+3) and sarcasm (binary) annotations
for hotel/restaurant reviews.
17,236 (sentence, aspect) pairs; splits 13,758 / 1,739 / 1,739 (train/validation/test)
Splits are leakage-free by construction: reviews sharing any normalized sentence are
grouped before splitting, so no sentence or review crosses splits
source column marks synthetic augmentation ('augmented'); evaluate sentiment on
source != 'augmented' for a… See the full description on the dataset page: https://huggingface.co/datasets/phamluan/vi_msa_sar.gpn-msa-with-microgliaMSA
