CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AaronZ345 /GTSinger GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks Yu Zhang*, Changhao Pan*, Wenxiang Guo*, Ruiqi Li, Zhiyuan Zhu, Jialei Wang, Wenhao Xu, Jingyu Lu, Zhiqing Hong, Chuxin Wang, LiChao Zhang, Jinzheng He, Ziyue Jiang, Yuxin Chen, Chen Yang, Jiecheng Zhou, Xinyu Cheng, Zhou Zhao | Zhejiang University Dataset of GTSinger (NeurIPS 2024 Spotlight): A Global Multi-Technique Singing Corpus with Realistic Music Scores for All… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/GTSinger.audiotext-to-audio10K<n<100K17 likes11k downloads1y agoHugging Face02AaronZ345 /MRSDrama ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang*, Wenxiang Guo*, Changhao Pan*, Zhiyuan Zhu*, Tao Jin, Zhou Zhao | Zhejiang University Dataset of ISDrama (ACMMM 2025): Immersive Spatial Drama Generation through Multimodal Prompting. We construct MRSDrama, the first multimodal recorded spatial drama dataset, containing binaural drama audios, scripts, videos, geometric poses, and textual prompts. We provide the full corpus… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/MRSDrama.audiotext-to-speech10K<n<100K3 likes8.9k downloads1y agoHugging Face03Aarsh-Wankar /Marathi-Wikipediatext0 likes8.4k downloads2y agoHugging Face04Ciroc0 /dmi-aarhus-predictions DMI Aarhus Predictions Prediction and frontend contract dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by predictions_latest.parquet Current future + verified prediction store dmi-collector frontend_snapshot.json Primary integration contract for the Vercel frontend dmi-collector Compatibility files File Status Notes predictions.parquet Legacy Still read by compatibility… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-predictions.6 likes5.1k downloads4m agoHugging Face05aaron-freedman /banking-concentration-data Banking Concentration Data (1959–2025) Public portion of the source data for the banking-concentration visualizations project at https://github.com/aaron-freedman/banking-concentration. This dataset assembles regulatory filings from three US bank and thrift supervisors covering 1959–2025 into a single, loadable collection of parquet and comma-separated values (CSV) files. Important: three additional files are needed by the full preprocessing pipeline but are not included here… See the full description on the dataset page: https://huggingface.co/datasets/aaron-freedman/banking-concentration-data.tabular-regression10M<n<100M0 likes2.4k downloads5mo agoHugging Face06Ciroc0 /dmi-aarhus-weather-data DMI Aarhus Weather Data Training data and model artifact dataset for the Aarhus weather pipeline. Maintained by Ciroc0. Primary files File Purpose Produced by training_matrix.parquet Current source of truth for training rows and causal observation context dmi-collector model_registry.json Active bucket registry per target dmi-ml-trainer model_meta.json Training timestamp, sample count and training window dmi-ml-trainer temperature_models.pkl… See the full description on the dataset page: https://huggingface.co/datasets/Ciroc0/dmi-aarhus-weather-data.tabular10K<n<100K1 likes2.4k downloads21h agoHugging Face07AarushSah /lmsys-chat-1m LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset This dataset contains one million real-world conversations with 25 state-of-the-art LLMs. It is collected from 210K unique IP addresses in the wild on the Vicuna demo and Chatbot Arena website from April to August 2023. Each sample includes a conversation ID, model name, conversation text in OpenAI API JSON format, detected language tag, and OpenAI moderation API tag. User consent is obtained through the "Terms of use"… See the full description on the dataset page: https://huggingface.co/datasets/AarushSah/lmsys-chat-1m.text1M<n<10M1 likes1.7k downloads2y agoHugging Face08aarajbhattarai /nepali-law-v2 Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/nepali-law-v2.texttext-generation10K<n<100K1 likes1.5k downloads4d agoHugging Face09aarajbhattarai /rejected-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-nepali-law-v2.texttext-generation10K<n<100K0 likes1.5k downloads4d agoHugging Face10aaroncaozj /BlockPAP-v1_Mix BlockPAP-v1_MixLeRobot v2.1 mixed dataset with 3 camera views for BlockPAP-v1 (real + MimicGen trajectories). ❗️ State Alignment between Simulation and Real Robot In this version of the Mix dataset, corresponding biases have been applied to state[2] and state[5] when constructing MimicGen data, unifying the distributions of sim and real data into the states required for a real robot. If performing policy rollout in simulation, apply the same bias to the input states before… See the full description on the dataset page: https://huggingface.co/datasets/aaroncaozj/BlockPAP-v1_Mix.image100K<n<1M0 likes1.1k downloads6mo agoHugging Face11aaronrose227 /narcbench NARCBench Activations and scenario data for Detecting Multi-Agent Collusion Through Multi-Agent Interpretability. Companion to github.com/aaronrose227/narcbench. Models and layers Model HF ID Layers Qwen3-32B-AWQ Qwen/Qwen3-32B-AWQ 26–30 Llama-3.1-70B-Instruct-AWQ-INT4 meta-llama/Llama-3.1-70B-Instruct-AWQ-INT4 32–37 DeepSeek-R1-Distill-Qwen-32B deepseek-ai/DeepSeek-R1-Distill-Qwen-32B 26–30 GPT-OSS-20B openai/gpt-oss-20b 10–14… See the full description on the dataset page: https://huggingface.co/datasets/aaronrose227/narcbench.feature-extraction10K<n<100K3 likes867 downloads4mo agoHugging Face12aarontseng /wikiann Dataset Card for WikiANN Dataset Summary WikiANN (sometimes called PAN-X) is a multilingual named entity recognition dataset consisting of Wikipedia articles annotated with LOC (location), PER (person), and ORG (organisation) tags in the IOB2 format. This version corresponds to the balanced train, dev, and test splits of Rahimi et al. (2019), which supports 176 of the 282 languages from the original WikiANN corpus. Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/aarontseng/wikiann.texttoken-classification1M<n<10M0 likes783 downloads25d agoHugging Face13aaroncaozj /0324_blockpap_negvideon<1K0 likes764 downloads6mo agoHugging Face14aaronngx /failbench-libero-v2 FailBench LIBERO v2 — contact-prediction dataset Labeled robot-failure trials built on LIBERO teleop demos (Franka Panda). Each trial injects a hardware failure partway through a demo and records the contacts the failure causes during a 1-second settle. The supervised target is a 240×320 force-weighted contact heatmap in the agentview camera — the model predicts where a failure at a given pre-failure configuration drives the robot/objects into contact. Sibling dataset:… See the full description on the dataset page: https://huggingface.co/datasets/aaronngx/failbench-libero-v2.robotics10K<n<100K1 likes686 downloads3mo agoHugging Face15av555 /aardvark-weather Aardvark Weather This repo contains a machine learning ready version of the observational dataset developed for training and evaluating the Aardvark Weather model. We gratefully acknowledge the agencies whose efforts in collecting, curating, and distributing the underlying datasets made this study possible. Dataset We provide two types of data: a timeslice of sample data in the format expected for ingestion to the model for demo runs and a training dataset for those… See the full description on the dataset page: https://huggingface.co/datasets/av555/aardvark-weather.18 likes667 downloads1y agoHugging Face16aarentai /UltrasoundGRF0 likes639 downloads3y agoHugging Face17AaronCIH /UnifiedIR0 likes609 downloads9mo agoHugging Face18aaroncaozj /0415_tube_success-wrongvideon<1K0 likes606 downloads5mo agoHugging Face19aaroncaozj /0419_hang_success-wrongvideon<1K0 likes585 downloads5mo agoHugging Face20aarajbhattarai /unjudged-nepali-law-v2 Nepali Source-Grounded Instruction Dataset — UNJUDGED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/unjudged-nepali-law-v2.text-generation0 likes531 downloads4d agoHugging Face21AaronWu901225 /MIT_states_UT_zappos1 likes507 downloads2y agoHugging Face22aaryavlal /arbiter-mini Arbiter-mini A small, purpose-built image dataset of household items captured under controlled Raspberry Pi camera conditions and labeled for binary waste/recycle classification according to San Diego, CA municipal recycling rules. Built as deployment-condition training data for the Arbiter sorting system, intended to be used alongside TrashNet to close the domain gap between studio imagery and real Pi-camera inference. Motivation Models trained purely on TrashNet… See the full description on the dataset page: https://huggingface.co/datasets/aaryavlal/arbiter-mini.imageimage-classificationn<1K1 likes500 downloads3mo agoHugging Face23AaryanK /ExecRetrieval ExecRetrieval — Released Artifacts Paper: https://arxiv.org/abs/2609.01865 Accepted to EMNLP 2026 — Main Conference (Budapest, October 2026). This bundle accompanies "ExecRetrieval: Measuring the Functional-Correctness Gap in Code-Embedding Retrieval" by Aaryan Kapoor and Md Abdullah Al Hafiz Khan (Kennesaw State University), accepted to the main conference of EMNLP 2026. Everything required to (a) inspect the dataset, (b) recompute every metric and table cell in the paper… See the full description on the dataset page: https://huggingface.co/datasets/AaryanK/ExecRetrieval.textsentence-similarity10K<n<100K1 likes489 downloads18d agoHugging Face24aaronwzl /mcqa_calibration_datasettext1K<n<10K1 likes485 downloads2y agoHugging Face25aaronwool2025 /spatialva-runtime-assets3dn<1K0 likes454 downloads5mo agoHugging Face26aarajbhattarai /rejected-law-instructions-dataset Nepali Source-Grounded Instruction Dataset — REJECTED Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/rejected-law-instructions-dataset.text-generation1 likes394 downloads9d agoHugging Face27aaroncaozj /0109_sweepvideon<1K0 likes376 downloads5mo agoHugging Face28aarajbhattarai /law-instructions-dataset Nepali Source-Grounded Instruction Dataset Synthetic Nepali instruction-tuning data generated with NVIDIA NeMo Data Designer from authoritative Nepali documents (agriculture manuals, legal texts). Answers are grounded strictly in the source; unanswerable questions get an explicit refusal. Records use chat messages format plus metadata and per-record quality_scores (grounding / correctness / naturalness, 1-5, LLM-as-judge). One data/train-<shard>.jsonl per source document; shards… See the full description on the dataset page: https://huggingface.co/datasets/aarajbhattarai/law-instructions-dataset.texttext-generation1K<n<10K0 likes354 downloads9d agoHugging Face29aaronsu11 /so100_pickThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so100", "total_episodes": 25, "total_frames": 3763, "total_tasks": 1, "total_videos": 50, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:25" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/aaronsu11/so100_pick.tabularrobotics10K<n<100K1 likes346 downloads1y agoHugging Face30aarav912 /online-retailtabular100K<n<1M0 likes333 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.