datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MOSAIC-Refactoring
Agentic Pull Request Dataset
Dataset Overview
The dataset contains 4,910,698 Pull Requests in total, consisting of 4,392,818 agent-authored PRs from 10 agents and 517,880 human-authored PRs. The agent-authored PRs come from Claude, Codegen, Codex, Copilot, Cosine, Cursor, Devin, Jules, Junie, and OpenHands. A summary of the dataset is presented below.
Cohort
Pull Requests
Merged Pull Requests
Repositories
Sum of Additions
Sum of Deletions
Humans
517880… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-Refactoring.voxpopuli_mosel_curatormosel
Dataset Description, Collection, and Source
The MOSEL corpus is a multilingual dataset collection including up to 950K hours of open-source speech recordings covering the 24 official languages of the European Union. We collect data by surveying labeled and unlabeled speech corpora under open-source compliant licenses.
In particular, MOSEL includes the automatic transcripts of 441k hours of unlabeled speech from VoxPopuli and LibriLight. The data is transcribed using Whisper large… See the full description on the dataset page: https://huggingface.co/datasets/FBK-MT/mosel.LoRA-WiSE
Dataset Card for the LoRA WiSE benchmark
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive
benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models
LoRA-WiSE spans various dataset sizes, backbones, ranks, and personalization sets, as presented in
the "Dataset Size Recovery from LoRA Weights" paper.
Task Details
Dataset Description
Dataset Structure
Data Subsets
Data Fields
Dataset Creation
Citation Information
🌐… See the full description on the dataset page: https://huggingface.co/datasets/MoSalama98/LoRA-WiSE.moss-002-sft-data
Dataset Card for "moss-002-sft-data"
Dataset Summary
An open-source conversational dataset that was used to train MOSS-002. The user prompts are extended based on a small set of human-written seed prompts in a way similar to Self-Instruct. The AI responses are generated using text-davinci-003. The user prompts of en_harmlessness are from Anthropic red teaming data.
Data Splits
name
# samples
en_helpfulness.json
419049
en_honesty.json
112580… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/moss-002-sft-data.yolo-baselines-no-mosaic-runsmoss-character-voices-bestof64
MOSS Character Voices — Best-of-64 (Stage 2)
Best-of-64 voice-acting takes from the 4.55B MOSS-TTS-Local voice-acting model
(laion/moss-tts-local-transformer-4.55b-voice-acting) for 13 evolved character voices.
Each prompt is a fixed, optimized champion performance direction (instruction) paired
with a Gemma-generated topic text (text) — together, one performance to render. For every
prompt we sample 64 takes with distinct seeds at 48 kHz, score each take, and rank the 64
within… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-bestof64.moss-character-voices-top3-captioned
MOSS Character Voices — Top-3 per Group, Captioned (training-ready)
The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in
laion/moss-character-voices-bestof64
— ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with
all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet.
Captions (two pipelines, same clip)
caption_procedural — Procedural Voice Captions:
terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.moss_train_gc_blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "moss",
"total_episodes": 36,
"total_frames": 10909,
"total_tasks": 1,
"total_videos": 72,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:36"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/moss_train_gc_block.sp500-daily-candles-2025
S&P 500 Daily Candles (2025)
This dataset provides daily OHLCV (Open, High, Low, Close, Volume) candles for all S&P 500 tickers between 01-01-2025 and 10-04-2025.
Dataset Summary
Date range: 2025-01-01 → 2025-10-04
Frequency: 1 day
Fields: ticker, date, open, high, low, close, volume
File format: CSV (sp500-daily-tickers-2025.csv)
Example Schema
Column
Type
Description
ticker
string
Stock symbol (e.g., AAPL, MSFT, AMZN)
date
datetime… See the full description on the dataset page: https://huggingface.co/datasets/mospira/sp500-daily-candles-2025.moss_stack_cubesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "moss",
"total_episodes": 30,
"total_frames": 7244,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/moss_stack_cubes.sp500-daily-candles-2024
SPY Daily Candles 2024
This dataset contains daily OHLCV (Open, High, Low, Close, Volume) candlestick data for all tickers listed on the S&P 500 from 01-01-2024 to 01-01-2025.
Columns
ticker, date, open, high, low, close, volume
bacbench-filtered-dnamoss-voice-identity-repairs
MOSS voice-acting v2 -- repaired takes
For each voice profile, every take whose ECAPA speaker similarity to the voice's reference fell
below 0.40, regenerated with that voice's identity LoRA (see
laion/moss-voice-identity-loras) merged at scale 1.0 on top of the identical condition
adapters at the identical lambdas.
Nothing here replaces anything. The original takes are untouched and remain part of the
corpus; low-similarity takes are kept deliberately, because they are useful… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-voice-identity-repairs.moss_open_drawer_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "moss",
"total_episodes": 30,
"total_frames": 8357,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/moss_open_drawer_teabox.Massive-STEPS-Moscow
Massive-STEPS-Moscow
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Moscow.moss_train_graspThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "moss",
"total_episodes": 43,
"total_frames": 23450,
"total_tasks": 1,
"total_videos": 86,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:43"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/moss_train_grasp.weathergpt-d1-mos-dataset
WeatherGPT D1 — multi-model NWP forecasts vs. ERA5-Land truth (India)
The training corpus for M2 (distributional bias correction) and M4
(precipitation calibration) in the WeatherGPT
project (SIH 2026). One row per (location, valid time): four NWP models'
forecasts at that point in time, and what ERA5-Land actually observed there.
M5 (trust ranker) trained a day earlier on a slightly smaller snapshot
of this same pipeline (9,507,456 rows vs. this upload's 9,582,912) that was… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/weathergpt-d1-mos-dataset.4L-RP-Human-Clean
4L-RP-Human: Multilingual Data–Text Alignment Judgements
Overview
This dataset contains structured data, corresponding texts, and human judgements of their semantic alignment:
Precision: how much of the information expressed in the text is supported by the input data?
Recall: how much of the input information is expressed in the text?
Individual annotator ratings are retained for each pair. F1 can be derived from precision and recall; it was not collected as a… See the full description on the dataset page: https://huggingface.co/datasets/Loria-MosAIk/4L-RP-Human-Clean.automotive-service-intelligence-sample
🚗 Automotive Service Intelligence Sample Dataset
Connected • Longitudinal • Feature-Engineered • Commercially Available
This repository contains a fully anonymized sample of the Growing-Moss Data Automotive Service Intelligence Dataset, a production-derived dataset built for analytics, forecasting, AI/ML, benchmarking, and commercial product development.
Unlike transactional datasets that provide isolated records, the Growing-Moss dataset delivers connected intelligence… See the full description on the dataset page: https://huggingface.co/datasets/Growing-Moss-Data/automotive-service-intelligence-sample.CMU-Mosei-textyolov8-baseline-mosaic-variants-runslevir-yolov8n-p2-gap-factorized-tal-augmentations-legacy-mosaic-seed42miniloop-o-smart60k-moss16rvq
MiniLoop-O Smart Full 1M — MOSS 16-RVQ preprocessed
Source: {SOURCE_REPO}
Rows: 1,000,000 (T2A 350k / A2A 250k / I2T 400k).
Old MiniMind-O Mimi targets were decoded and re-encoded with {MOSS_CODEC_ID} into 16-RVQ MOSS codes.
Binary fields are uint16 little-endian and reshape to (frames,16).
levir-ship-mosaic-policy-matrixmosaic-bench
MOSAIC
199 compositional attack chains across 10 real-world web applications, used to
benchmark whether AI coding agents will compose individually-routine tickets
into a deployable vulnerability.
Code & harness: https://github.com/mosaic-benchmark/mosaic-benchmark
Datasheet: DATASHEET.md · Croissant 1.1: croissant.json
What's in this release
Artifact
Contents
mosaic-bench.xlsx
Per-chain ASR (9 models, standard + resumed), BugBot verdicts (diff-mode +… See the full description on the dataset page: https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench.tinyperson-copy-paste-mosaic-runsArabic_SQuAD
Dataset Card for "Arabic_SQuAD"
More Information needed
Citation
@inproceedings{mozannar-etal-2019-neural,
title = "Neural {A}rabic Question Answering",
author = "Mozannar, Hussein and
Maamary, Elie and
El Hajal, Karl and
Hajj, Hazem",
booktitle = "Proceedings of the Fourth Arabic Natural Language Processing Workshop",
month = aug,
year = "2019",
address = "Florence, Italy",
publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa3zazi/Arabic_SQuAD.spatial_mosaic_vqa
SpatialMosaic: A Multi-View VLM Dataset for Partial Visibility
Description
SpatialMosaic is a multi-view visual question answering dataset for evaluating spatial reasoning under partial visibility, occlusion, and low-overlap views. It pairs indoor ScanNet++ and outdoor Waymo scene references with multi-frame VQA annotations. Questions require models to combine fragmented evidence across 2-5 views, rather than answering from a single image. The tasks… See the full description on the dataset page: https://huggingface.co/datasets/jmkey/spatial_mosaic_vqa.
