datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
demo_data
1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_en
1,000 examples from https://huggingface.co/datasets/llamafactory/alpaca_gpt4_zh
300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_en
300 examples from https://huggingface.co/datasets/llamafactory/glaive_toolcall_zh
91 examples for identity learning
300 examples from https://huggingface.co/datasets/cognitivecomputations/SystemChat-2.0
6 examples for multimodal supervised… See the full description on the dataset page: https://huggingface.co/datasets/llamafactory/demo_data.demo1
Dataset Card for Demo1
Dataset Summary
This is a demo dataset. It consists in two files data/train.csv and data/test.csv
You can load it with
from datasets import load_dataset
demo1 = load_dataset("lhoestq/demo1")
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/lhoestq/demo1.moral_storiesMoral Stories is a crowd-sourced dataset of structured, branching narratives for the study of grounded, goal-oriented
social reasoning. For detailed information, see https://aclanthology.org/2021.emnlp-main.54.pdf.opencs2_dataset_demo
HLTV CS2 Demos Dataset
Counter-Strike 2 match demos scraped from HLTV.org
plus a compact per-map analysis JSON. Each row of the metadata Parquet is
one .dem file (one played CS2 map); a best-of-3 match contributes 2
or 3 rows depending on whether it went 2-0 or 2-1.
Parquet holds everything you typically filter on: map_name,
patch_version, rounds_played, per-player kast / adr / rating,
every kill tick with weapon + headshot, every round's winner + end reason.
.dem binaries live… See the full description on the dataset page: https://huggingface.co/datasets/blanchon/opencs2_dataset_demo.Core-DEM
Major TOM Core-DEM
Major TOM Core-DEM contains a global coverage of Copernicus DEM, each of size 356 x 356 pixels.
This dataset was created to support the development of the MESA terrain generation model. It is also featured in the paper EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images and is part of the Major TOM: Expandable Datasets for Earth Observation ecosystem.
Official Viewer App:: Major TOM Viewer
Major TOM GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-DEM.v1-sft-demolibrispeech_asr_demomoe-demo-clean
RoboTwin MOE Demo Clean
Raw RoboTwin demonstration data copied from bos:/lab-test/moe-demo-clean/.
The dataset is organized by task directories such as
place_can_basket-demo_clean-200/. Each task directory contains episode
subdirectories with an HDF5 trajectory file and an instructions.json file.
superb_demo
Disclaimer
This is a tiny subset of the SUPERB dataset, which is intended only for demo purposes!
See the full dataset here: https://huggingface.co/datasets/superb
labor-demand-index
Chainticks Labor Demand Index
An agent-friendly labor-demand panel for finding where organizations are still trying to hire humans for work that may be automatable.
This dataset intentionally publishes aggregates and official public-domain series only:
official_labor_timeseries: BLS JOLTS monthly openings, hires, quits, layoffs/discharges, and separations by US sector (latest revision).
jolts_sector_metrics: derived vacancy/hire, quit/hire, separation/hire, and YoY openings… See the full description on the dataset page: https://huggingface.co/datasets/Chainticks/labor-demand-index.VoiceBank-DEMAND-16kprocessed_demo
Dataset Card for "processed_demo"
More Information needed
dahih-tts2-demucs-cleanedfinancial-analyst-data-demo
financial-analyst-data-demo
EN: A-share historical OHLCV + valuation + financials + TDX F10 events, packaged in Qlib binary + Parquet formats. Companion dataset for financial-analyst — a 14-agent single-stock deep-dive research workstation.
中文: A 股历史行情 + 估值 + 财报 + TDX F10 事件数据集, Qlib 二进制 + Parquet 双格式打包. 配套 financial-analyst — 14 Agent 个股深度研究工作站使用.
Published / 发布: 2026-05-24 · Size / 体量: ~0.16 GB · License: Apache 2.0
📊 Three Preset Tiers / 三档预设
Pick the tier that… See the full description on the dataset page: https://huggingface.co/datasets/yifishbossman/financial-analyst-data-demo.synthetic_dem
Dataset Card for synthetic_dem
Dataset Summary
The Synthetic DEM Corpus is the result of the first phase of a collaboration between El Colegio de México (COLMEX) and the Barcelona Supercomputing Center (BSC).
It all began when COLMEX was looking for a way to have its Diccionario del Español de México (DEM), which can be accessed online, include the option to play each of its words with a Mexican accent through synthetic speech files. On the other hand, BSC is always on… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/synthetic_dem.caltennis
CalTennis: Large Multi-View Tennis Video Dataset
CalTennis is a large-scale video benchmark designed for evaluating monocular-to-3D human pose estimation in the wild.
The dataset comprises over 11 million frames (51 hours) of tennis practice and match play from 40 players, captured with 2–6 synchronized cameras at 60Hz. It is 10x larger than existing in-the-wild human motion video datasets and offers the first large-scale benchmark for synchronized multi-view recordings of expert… See the full description on the dataset page: https://huggingface.co/datasets/demalenk/caltennis.fava-flagged-demo
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/abhika-m/fava-flagged-demo.Demeter-LongCoT-6M
Demeter-LongCoT-6M
Demeter-LongCoT-6M is a high-quality, compact chain-of-thought reasoning dataset curated for tasks in mathematics, science, and coding. While the dataset spans diverse domains, it is primarily driven by mathematical reasoning, reflecting a major share of math-focused prompts and long-form logical solutions.
Quick Start with Hugging Face Datasets🤗
pip install -U datasets
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Demeter-LongCoT-6M.icl-demo-dataset
ICL Demo Dataset
A small, recent collection of bimanual manipulation demonstrations on the YOR
robot, in LeRobot v2.1 format. Recorded 3 and 6 August 2026.
285 episodes · 254,171 frames · 2.35 hours · 27 tasks · 3 camera views
This is a companion to adityx23/icl-dataset
— same robot, same schema, same task vocabulary, but a separate and much smaller
collection. Roughly ten demonstrations per task, several of them tasks that do not
appear in the larger dataset at all.… See the full description on the dataset page: https://huggingface.co/datasets/adityx23/icl-demo-dataset.reason-tool-use-demo-1500
Dataset info
The dataset is a selection of reasoning toolcalls data from https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use, which contains data from Hermes-Tools、Glaive-FC、ToolAce、Nvidia-When2Call.
The format has been transformed to adapt llama-factory v1 training pipeline.
DemoFeedbackfaers_bronze_demoarmnet-demo-leaderboarddemo-end-effector-space
demo_action_space
Derived from adityx23/icl-demo-dataset
(lerobot v2.1 format, 285 episodes / 254,171 frames / 27 tasks). Every
existing column, task, episode flag (success/valid/keep), and
episode_uid is carried through unchanged.
Sibling dataset: demo_joint_space
adds the same episodes' joint-space IK targets instead of Cartesian poses.
Same source, same episode indices, same pipeline.
What's added
Two new features, observation.left_ee / observation.right_ee… See the full description on the dataset page: https://huggingface.co/datasets/Hannibal52Barca/demo-end-effector-space.mimiciv_demoScene2Wave-demo-assetsChartBench-Demo
ChartBench: A Benchmark for Complex Visual Reasoning in Charts
Introduction
We propose the challenging ChartBench to evaluate the chart recognition of MLLMs.
We improve the Acc+ metric to avoid the randomly guessing situations.
We collect a larger set of unlabeled charts to emphasize the MLLM's ability to interpret visual information without the aid of annotated data points.
Todo
Open source all data of ChartBench.
Open source the evaluate… See the full description on the dataset page: https://huggingface.co/datasets/SincereX/ChartBench-Demo.demichanwakataritai
Bangumi Image Base of Demi-chan Wa Kataritai
This is the image base of bangumi Demi-chan wa Kataritai, we detected 16 characters, 1889 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/demichanwakataritai.Edge-Agent-Reasoning-WebSearch-260K
Edge Agent Reasoning WebSearch 260K
Abstract
The Edge-Agent-Reasoning-WebSearch-260K dataset is a massive, synthetically expert-engineered corpus of over 700 Million tokens, designed to train small, local models (SLMs) and edge-deployed agents in advanced problem deconstruction and self-aware reasoning.
Rather than training a model to execute instructions directly—which often leads to hallucinations when context is missing—this dataset trains a model to act as a… See the full description on the dataset page: https://huggingface.co/datasets/DEMIRUNC/Edge-Agent-Reasoning-WebSearch-260K.arc_agi_3_public_demo_human_testing
Dataset Card for ARC-AGI 3 Public Demo Human Testing
Dataset Summary
This dataset contains human gameplay logs and trajectories from the ARC-AGI 3 public demo. It is a fully open-source dataset created by the ARC Prize.
The primary purpose of publishing this dataset on Hugging Face is to make it easily accessible and convenient for participants in the Kaggle ARC Prize 2026 Competition.
The implementation and source code used to process and upload this dataset to… See the full description on the dataset page: https://huggingface.co/datasets/magic-sword/arc_agi_3_public_demo_human_testing.
