datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxpopuli_mosel_curatorLoRA-WiSE
Dataset Card for the LoRA WiSE benchmark
The LoRA Weight Size Evaluation (LoRA-WiSE) is a comprehensive
benchmark specifically designed to evaluate LoRA dataset size recovery methods for generative models
LoRA-WiSE spans various dataset sizes, backbones, ranks, and personalization sets, as presented in
the "Dataset Size Recovery from LoRA Weights" paper.
Task Details
Dataset Description
Dataset Structure
Data Subsets
Data Fields
Dataset Creation
Citation Information
🌐… See the full description on the dataset page: https://huggingface.co/datasets/MoSalama98/LoRA-WiSE.moss-character-voices-bestof64
MOSS Character Voices — Best-of-64 (Stage 2)
Best-of-64 voice-acting takes from the 4.55B MOSS-TTS-Local voice-acting model
(laion/moss-tts-local-transformer-4.55b-voice-acting) for 13 evolved character voices.
Each prompt is a fixed, optimized champion performance direction (instruction) paired
with a Gemma-generated topic text (text) — together, one performance to render. For every
prompt we sample 64 takes with distinct seeds at 48 kHz, score each take, and rank the 64
within… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-bestof64.moss-character-voices-top3-captioned
MOSS Character Voices — Top-3 per Group, Captioned (training-ready)
The top-3 takes per group (rank 0/1/2 by reward model) of all ~12,900 groups in
laion/moss-character-voices-bestof64
— ~38,700 samples — each annotated with both LAION voice-acting caption pipelines and shipped with
all scores + metadata, ready to drop into training. Sharded as data/train-*-of-00129.parquet.
Captions (two pipelines, same clip)
caption_procedural — Procedural Voice Captions:
terse… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-character-voices-top3-captioned.moss-voice-profile-references
MOSS voice-profile references
Complete voice profiles: one reference speaker rendered through a matrix of named acting
conditions in English and German, with every candidate take kept — not just the winner — and
every take scored on itself.
Two releases live here.
voices
groups / voice
candidates / group
rows
audio
audio variants
pilot/ — the ten pilot voices
10
842
48
402,560
853.5 h
raw
root — Velvet Sage Baritone
1
832
32
106,424
239.6 h
raw, vc, raw_sidon… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-voice-profile-references.moss_train_gc_blockThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "moss",
"total_episodes": 36,
"total_frames": 10909,
"total_tasks": 1,
"total_videos": 72,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:36"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/moss_train_gc_block.moss_stack_cubesThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "moss",
"total_episodes": 30,
"total_frames": 7244,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/moss_stack_cubes.weathergpt-d1-mos-dataset
WeatherGPT D1 — multi-model NWP forecasts vs. ERA5-Land truth (India)
The training corpus for M2 (distributional bias correction) and M4
(precipitation calibration) in the WeatherGPT
project (SIH 2026). One row per (location, valid time): four NWP models'
forecasts at that point in time, and what ERA5-Land actually observed there.
M5 (trust ranker) trained a day earlier on a slightly smaller snapshot
of this same pipeline (9,507,456 rows vs. this upload's 9,582,912) that was… See the full description on the dataset page: https://huggingface.co/datasets/Arko007/weathergpt-d1-mos-dataset.moss_open_drawer_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "moss",
"total_episodes": 30,
"total_frames": 8357,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/moss_open_drawer_teabox.moss_train_graspThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "moss",
"total_episodes": 43,
"total_frames": 23450,
"total_tasks": 1,
"total_videos": 86,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:43"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ncavallo/moss_train_grasp.bacbench-filtered-dnamoss-voice-identity-repairs
MOSS voice-acting v2 -- repaired takes
For each voice profile, every take whose ECAPA speaker similarity to the voice's reference fell
below 0.40, regenerated with that voice's identity LoRA (see
laion/moss-voice-identity-loras) merged at scale 1.0 on top of the identical condition
adapters at the identical lambdas.
Nothing here replaces anything. The original takes are untouched and remain part of the
corpus; low-similarity takes are kept deliberately, because they are useful… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-voice-identity-repairs.Massive-STEPS-Moscow
Massive-STEPS-Moscow
Dataset Summary
Massive-STEPSis a large-scale dataset of semantic trajectories intended for understanding POI check-ins. The dataset is derived from the Semantic Trails Dataset and Foursquare Open Source Places, and includes check-in data from 15 cities across 10 countries. The dataset is designed to facilitate research in various domains, including trajectory prediction, POI recommendation, and urban modeling. Massive-STEPS emphasizes the… See the full description on the dataset page: https://huggingface.co/datasets/CRUISEResearchGroup/Massive-STEPS-Moscow.Arabic_SQuAD
Dataset Card for "Arabic_SQuAD"
More Information needed
Citation
@inproceedings{mozannar-etal-2019-neural,
title = "Neural {A}rabic Question Answering",
author = "Mozannar, Hussein and
Maamary, Elie and
El Hajal, Karl and
Hajj, Hazem",
booktitle = "Proceedings of the Fourth Arabic Natural Language Processing Workshop",
month = aug,
year = "2019",
address = "Florence, Italy",
publisher = "Association for Computational… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa3zazi/Arabic_SQuAD.miniloop-o-smart60k-moss16rvq
MiniLoop-O Smart Full 1M — MOSS 16-RVQ preprocessed
Source: {SOURCE_REPO}
Rows: 1,000,000 (T2A 350k / A2A 250k / I2T 400k).
Old MiniMind-O Mimi targets were decoded and re-encoded with {MOSS_CODEC_ID} into 16-RVQ MOSS codes.
Binary fields are uint16 little-endian and reshape to (frames,16).
MOSS-TTS-ky-kk-bench
MOSS-TTS Kyrgyz/Kazakh Cross-Lingual Benchmark
Paired renderings of the same prompts by two text-to-speech models: MOSS-TTS v1.5
as released, and the same model with a Kyrgyz/Kazakh QLoRA adapter. Each row puts
the two side by side, so the effect of the fine-tune can be judged by ear rather than
from a metric.
The grid is deliberately cross-lingual: every reference voice is used with every
language, so an English speaker reads Kyrgyz, a Kazakh speaker reads Russian, and so on.… See the full description on the dataset page: https://huggingface.co/datasets/KaniTTS-research-team/MOSS-TTS-ky-kk-bench.llama3-uf-meta-specific-tokenized_mosaicpickplace_sweetsjp_smallcandyThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 4264,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MostafaOthman/pickplace_sweetsjp_smallcandy.moss_sft_002pickplace_sweetsjpThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 2682,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MostafaOthman/pickplace_sweetsjp.pickplace_sweetsjp_strawThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 5,
"total_frames": 2015,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/MostafaOthman/pickplace_sweetsjp_straw.MOSAIC-agentic-3m
Agent Activity Dataset
This dataset is released in conjunction with the paper Investigating Autonomous Agent Contributions in the Wild: Activity Patterns and Code Change over Time, accepted at MSR 2026.
Dataset Overview
The dataset contains a total of 111,969 Pull Requests (June through August 2025) from both coding agents (Claude Code, OpenAI Codex, GitHub Copilot, Google Jules, and Devin) and human contributors. It also includes additional activity metadata such as… See the full description on the dataset page: https://huggingface.co/datasets/AISE-TUDelft/MOSAIC-agentic-3m.moscow-parking-occupancy
Moscow Parking Occupancy Dataset
4.6 million half-hourly occupancy snapshots for 210 municipal parking lots
in Moscow, continuous since 31 March 2025 and refreshed every Monday.
Long, dense, public occupancy series are rare: most published work on parking
occupancy prediction still leans on the UCI Parking Birmingham set, which
covers October–December 2016 only.
Configs
Config
Rows
What it is
occupancy
~4.6 M
Snapshots every 30 minutes
parking_spots… See the full description on the dataset page: https://huggingface.co/datasets/matrosovdani/moscow-parking-occupancy.moss_put_cube_teaboxThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "moss",
"total_episodes": 30,
"total_frames": 8636,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Beegbrain/moss_put_cube_teabox.weld-quality-datasetSource: https://www.phase-trans.msm.cam.ac.uk/map/data/materials/welddb-b.html
moss-va-trajectory-corpus
Emotional trajectory corpus — public subset
7,638,961 trajectory specifications over 17,722,101 distinct source clips.
A trajectory is an ordered list of clips from one speaker that together walk one emotional or
VoiceNet dimension — for example a sequence that starts calm and ends furious, or one that moves
from high to low valence in even steps. The corpus contains no audio: each row names the clips
by uid and records where on the dimension each step sits.
It exists to train… See the full description on the dataset page: https://huggingface.co/datasets/laion/moss-va-trajectory-corpus.moss_test4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "moss",
"total_episodes": 2,
"total_frames": 597,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/1g0rrr/moss_test4.mos260
Russian TTS MOS Evaluation Dataset
A curated Russian synthesized speech dataset with crowdsourced quality ratings.
Overview
This dataset contains Russian synthesized speech samples from multiple TTS systems and audio codec models, evaluated by human annotators via crowdsourcing. Quality ratings are provided in two dimensions: perceptual quality (MOS) and intelligibility (Int-MOS).
Language: Russian only
Sources: F5, FishSpeech, GPT-So-VITS, Tortoise, XTTS… See the full description on the dataset page: https://huggingface.co/datasets/lab260/mos260.africa-madagascar-madagascar-most-likely-fews-net-acutely-food-insecure-popu-805b42f3
Madagascar Most Likely FEWS NET Acutely Food Insecure Population Estimates Data | Africa (Madagascar official open data)
810 rows - 1 Africa country - 2023-2026 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official CSV resource from Madagascar as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-madagascar-madagascar-most-likely-fews-net-acutely-food-insecure-popu-805b42f3.moss-khalilThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "koch",
"total_episodes": 2,
"total_frames": 2380,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/lilkm/moss-khalil.
