datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GTSinger
GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University
Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.MRSDrama
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting
Yu Zhang*, Wenxiang Guo*, Changhao Pan*, Zhiyuan Zhu*, Tao Jin, Zhou Zhao | Zhejiang University
Dataset of ISDrama (ACMMM 2025): Immersive Spatial Drama Generation through Multimodal Prompting.
We construct MRSDrama, the first multimodal recorded spatial drama dataset, containing binaural drama audios, scripts, videos, geometric poses, and textual prompts.
We provide the full corpus… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/MRSDrama.Marathi-Wikipediadmi-aarhus-predictions
DMI Aarhus Predictions
Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
predictions_latest.parquet
Current future + verified prediction store
dmi-collector
frontend_snapshot.json
Primary integration contract for the Vercel frontend
dmi-collector
Compatibility files
File
Status
Notes
predictions.parquet
Legacy
Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.banking-concentration-data
Banking Concentration Data (1959–2025)
Public portion of the source data for the banking-concentration visualizations
project at https://github.com/aaron-freedman/banking-concentration. This dataset
assembles regulatory filings from three US bank and thrift supervisors covering
1959–2025 into a single, loadable collection of parquet and comma-separated
values (CSV) files.
Important: three additional files are needed by the full preprocessing
pipeline but are not included here… See the full description on the dataset page: https://huggingface.co/datasets/aaron-freedman/banking-concentration-data.dmi-aarhus-weather-data
DMI Aarhus Weather Data
Training data and model artifact dataset for the Aarhus weather pipeline. Maintained by Ciroc0.
Primary files
File
Purpose
Produced by
training_matrix.parquet
Current source of truth for training rows and causal observation context
dmi-collector
model_registry.json
Active bucket registry per target
dmi-ml-trainer
model_meta.json
Training timestamp, sample count and training window
dmi-ml-trainer
temperature_models.pkl… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-weather-data.lmsys-chat-1m
LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
This dataset contains one million real-world conversations with 25 state-of-the-art LLMs.
It is collected from 210K unique IP addresses in the wild on the Vicuna demo and Chatbot Arena website from April to August 2023.
Each sample includes a conversation ID, model name, conversation text in OpenAI API JSON format, detected language tag, and OpenAI moderation API tag.
User consent is obtained through the "Terms of use"… See the full description on the dataset page: https://huggingface.co/datasets/AarushSah/lmsys-chat-1m.nepali-law-v2
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.rejected-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.BlockPAP-v1_Mix
BlockPAP-v1_MixLeRobot v2.1 mixed dataset with 3 camera views for BlockPAP-v1 (real + MimicGen trajectories).
❗️ State Alignment between Simulation and Real Robot
In this version of the Mix dataset, corresponding biases have been applied to state[2] and state[5] when constructing MimicGen data, unifying the distributions of sim and real data into the states required for a real robot.
If performing policy rollout in simulation, apply the same bias to the input states before… See the full description on the dataset page: https://huggingface.co/datasets/aaroncaozj/BlockPAP-v1_Mix.narcbench
NARCBench
Activations and scenario data for Detecting Multi-Agent Collusion Through Multi-Agent Interpretability. Companion to github.com/aaronrose227/narcbench.
Models and layers
Model
HF ID
Layers
Qwen3-32B-AWQ
Qwen/Qwen3-32B-AWQ
26–30
Llama-3.1-70B-Instruct-AWQ-INT4
meta-llama/Llama-3.1-70B-Instruct-AWQ-INT4
32–37
DeepSeek-R1-Distill-Qwen-32B
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
26–30
GPT-OSS-20B
openai/gpt-oss-20b
10–14… See the full description on the dataset page: https://huggingface.co/datasets/aaronrose227/narcbench.wikiann
Dataset Card for WikiANN
Dataset Summary
WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, and test splits of Rahimi et al. (2019), which supports 176 of the 282 languages from the original WikiANN corpus.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/aarontseng/wikiann.0324_blockpap_negfailbench-libero-v2
FailBench LIBERO v2 — contact-prediction dataset
Labeled robot-failure trials built on LIBERO teleop demos (Franka Panda). Each trial injects a
hardware failure partway through a demo and records the contacts the failure causes during a
1-second settle. The supervised target is a 240×320 force-weighted contact heatmap in the
agentview camera — the model predicts where a failure at a given pre-failure configuration drives
the robot/objects into contact.
Sibling dataset:… See the full description on the dataset page: https://huggingface.co/datasets/aaronngx/failbench-libero-v2.aardvark-weather
Aardvark Weather
This repo contains a machine learning ready version of the observational dataset developed for training and evaluating the Aardvark Weather model. We gratefully acknowledge the agencies whose efforts in collecting, curating, and distributing the underlying datasets made this study possible.
Dataset
We provide two types of data: a timeslice of sample data in the format expected for ingestion to the model for demo runs and a training dataset for those… See the full description on the dataset page: https://huggingface.co/datasets/av555/aardvark-weather.UltrasoundGRFUnifiedIR0415_tube_success-wrong0419_hang_success-wrongunjudged-nepali-law-v2
Nepali Source-Grounded Instruction Dataset — UNJUDGED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.MIT_states_UT_zapposarbiter-mini
Arbiter-mini
A small, purpose-built image dataset of household items captured under controlled Raspberry Pi camera conditions and labeled for binary waste/recycle classification according to San Diego, CA municipal recycling rules. Built as deployment-condition training data for the Arbiter sorting system, intended to be used alongside TrashNet to close the domain gap between studio imagery and real Pi-camera inference.
Motivation
Models trained purely on TrashNet… See the full description on the dataset page: https://huggingface.co/datasets/aaryavlal/arbiter-mini.ExecRetrieval
ExecRetrieval — Released Artifacts
Paper: https://arxiv.org/abs/2609.01865
Accepted to EMNLP 2026 — Main Conference (Budapest, October 2026).
This bundle accompanies "ExecRetrieval: Measuring the Functional-Correctness
Gap in Code-Embedding Retrieval" by Aaryan Kapoor and Md Abdullah Al Hafiz
Khan (Kennesaw State University), accepted to the main conference of EMNLP
2026.
Everything required to (a) inspect the dataset, (b) recompute every
metric and table cell in the paper… See the full description on the dataset page: https://huggingface.co/datasets/AaryanK/ExecRetrieval.mcqa_calibration_datasetspatialva-runtime-assetsrejected-law-instructions-dataset
Nepali Source-Grounded Instruction Dataset — REJECTED
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-law-instructions-dataset.0109_sweeplaw-instructions-dataset
Nepali Source-Grounded Instruction Dataset
Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data
Designer from authoritative Nepali documents (agriculture manuals, legal
texts). Answers are grounded strictly in the source; unanswerable questions
get an explicit refusal. Records use chat messages format plus metadata
and per-record quality_scores (grounding / correctness / naturalness, 1-5,
LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.so100_pickThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 25,
"total_frames": 3763,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:25"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aaronsu11/so100_pick.online-retail
