datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
v1-sft-demodahih-tts2-demucs-cleanedcaltennis
CalTennis: Large Multi-View Tennis Video Dataset
CalTennis is a large-scale video benchmark designed for evaluating monocular-to-3D human pose estimation in the wild.
The dataset comprises over 11 million frames (51 hours) of tennis practice and match play from 40 players, captured with 2–6 synchronized cameras at 60Hz. It is 10x larger than existing in-the-wild human motion video datasets and offers the first large-scale benchmark for synchronized multi-view recordings of expert… See the full description on the dataset page: https://huggingface.co/datasets/demalenk/caltennis.reason-tool-use-demo-1500
Dataset info
The dataset is a selection of reasoning toolcalls data from https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use, which contains data from Hermes-Tools、Glaive-FC、ToolAce、Nvidia-When2Call.
The format has been transformed to adapt llama-factory v1 training pipeline.
DemoFeedbackarmnet-demo-leaderboardScene2Wave-demo-assetsmachine-failure-mlops-demo-logsglm52-demolition-data
GLM-5.2-Demolition — Training & Calibration Data
Apple Silicon AI hub ·
Model release ·
MLX code sample
Preview scope, checked September 10, 2026: the default Hub viewer indexes
87,586 rows (84,231 train, 3,277 validation, 78 test). The original release
total below describes the broader JSONL repository. Use the file browser and
explicit file selections when reusing a particular corpus. The hub includes
a checked download example for the seven-row MLX code sample.
The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.piperx-demo558-value1500-a50-top10-union20-v1
PiperX advantage-selected teleoperation segments
Only pure human demonstrations. Value checkpoint step1500 (mixed demo+HIL); no HIL frames in this export.
A50 ranked globally across 558 source episodes, top10% AND A>0. Every selected start expands to [t,t+20); overlaps and adjacency merge. Each disconnected component is a separate output episode. Interior frames need not themselves be top10%.
Output: 7146 segments, 322261 frames, 2.983898 hours at30FPS.
Three camera streams and… See the full description on the dataset page: https://huggingface.co/datasets/Elvinky/piperx-demo558-value1500-a50-top10-union20-v1.grok-demon-dataset-ESUGround-Offline-Evaluationdemoatomic-metrics-demographic-training-size
Atomic Metrics: Demographic Training-Size Analysis
Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction.
Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.commoncrawl-jobs-demo
Common Crawl on Jobs — datatrove JobsPipelineExecutor demo
228,088 English web documents (~820 MB compressed, 60 jsonl.gz shards) extracted from 1,237,374 Common Crawl pages — the output of a test run of datatrove's experimental JobsPipelineExecutor, which fans a datatrove pipeline out across a pool of Hugging Face Jobs instead of a Slurm cluster.
This is a pipeline demo artifact, not a curated corpus: one segment slice of one crawl, shared as the verifiable receipt for the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/commoncrawl-jobs-demo.uninavid-objnav-demodataflow-demo-code
DataFlow demo -- Code Pipeline
Project Page | Technical Report | GitHub
This dataset is a demo of the DataFlow Code data processing pipeline from the DataFlow project. It provides a lightweight, inspectable view of what the pipeline produces: curated, execution-checked code SFT supervision pairs.
For full pipeline design and evaluation details, please refer to our technical report: DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-demo-code.RPToolkit-demo-datasetRPToolkit is a data generation pipeline, part of Augmentoolkit, that generates synthetic RP sessions inspired by input stories. Basically: feed in Lord of the Rings, get out high fantasy adventure RPs.
This dataset, containing over a million trainable tokens across around 1000 RP sessions, is meant to showcase the capabilities of this pipeline.
The input texts used were: a variety of myths and classic stories from Gutenberg; the first few chapters of some miscellaneous webnovels and… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/RPToolkit-demo-dataset.tla-demotic-v18-premium
Dataset Card for Dataset tla-demotic-v18-premium
This data set contains demotic sentences in transliteration, with lemmatization, with POS glossing and with a German translation.
The data comes from the database of the Thesaurus Linguae Agegyptiae, corpus version 18, and contains only fully intact,
unambiguously readable sentences (13,383 of 31,156 sentences), adjusted for philological and editorial markup.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-demotic-v18-premium.SwS-Demo-Dataset
Dataset Card for SwS-Demo-Dataset
[🌐 Website] •
[🤗 Demo Dataset] •
[📜 Paper] •
[🐱 GitHub] •
[🐦 Twitter] •
[📕 Rednote]
This dataset is a demo set of synthetic problems generated by SwS, comprising 500 samples for each model and category. The full dataset and model are currently under review by Microsoft and will be released once approved.
Data Loading
from datasets import load_dataset
dataset = load_dataset("MasterVito/SwS-Demo-Dataset")
Data… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/SwS-Demo-Dataset.trace-demoCLAMP-Sampled-Continuations-and-Demos
CLAMP Sampled Continuations and VLABench Demos
This dataset accompanies HLR/CLAMP, the Constrained Language-Action Model Planner.
It contains 20 closed-loop VLABench case studies: 10 successful and 10 unsuccessful executions. Each case includes the input image and mask, instruction, prompt and entity metadata, CLAMP and ground-truth plans, evaluation output, execution video, and a frame manifest.
Layout
demo_dataset/
├── manifest.json
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/SueMintony/CLAMP-Sampled-Continuations-and-Demos.demoverse-personas-es-v1
Dataset Card for DemoVerse Personas ES v1
Resumen del dataset
demoverse-personas-es-v1 es un dataset de 100.000 personas sinteticas en espanol para Espana, disenado como artefacto publico y como capa operativa para simulacion sociológica.
El dataset se inspira metodologicamente en nvidia/Nemotron-Personas-France, pero no reutiliza sus filas ni intenta replicar la poblacion francesa. La adaptacion reescribe el marco para Espana, con clivajes territoriales, sistema de… See the full description on the dataset page: https://huggingface.co/datasets/apol/demoverse-personas-es-v1.personaplex-finetuning-pharma-data-sample
PersonaPlex Finetuning — Pharma Data Sample
A 10-example slice of the synthetic patient-support / medication
adherence dataset used to train
demegire/personaplex-finetune-pharma.
The on-disk layout below is exactly what the trainer in
emotion-machine-org/personaplex-finetune
consumes — use this as a template when building your own.
Split: 8 train / 2 eval (mirrors the upstream 2003 / 20 split at
sample scale).
Layout
.
├── adhery_v2.jsonl # master… See the full description on the dataset page: https://huggingface.co/datasets/demegire/personaplex-finetuning-pharma-data-sample.giskard-hub-demo-retaildemospeedup-libero-robocasa-entropy
DemoSpeedup 엔트로피: LIBERO / RoboCasa
기존 GR00T-N1.5 DemoSpeedup 재현에서 실제로 사용한 원본 프레임별 엔트로피입니다.
모델 재추론 없이 구간 분류 역치를 변경하고, 원본 데모를 다른 배속으로 재구성할 수 있습니다.
모델 아키텍처를 변경하는 방법이 아니며, 변환한 데모로 기존 GR00T-N1.5를 학습합니다.
벤치마크
에피소드
프레임
누락/NaN/Inf
LIBERO
1,693
273,465
0
RoboCasa
7,200
2,073,457
0
재현 범위 정정 (2026-09-19)
이 저장소의 엔트로피와 라벨은 우리 GR00T 이식 구현의 결과이며, 논문을 그대로 재현한
결과라고 해석하면 안 됩니다. 논문 §3.2는
클러스터의 평균 엔트로피가 0 미만이면 precision, 그 외(노이즈 포함)는 casual이라고 설명합니다.
공식 공개 코드의… See the full description on the dataset page: https://huggingface.co/datasets/prehj/demospeedup-libero-robocasa-entropy.reddit-demoReddit Demo dataset
fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/apapanikolaou/fair_dataset_demo.fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/cparidaAI/fair_dataset_demo.EUReCA-demo
EUReCA demo case
One simulated cone-beam acquisition for trying the
EUReCA CBCT reconstruction models
without any data preparation.
Source CT: LUNA16 / LIDC-IDRI series
1.3.6.1.4.1.14519.5.2.1.6279.6001.154837327827713479309898027966 (thorax),
CC BY 3.0. This case is in EUReCA's reserved test split and was never used
for training or model selection.
Geometry: Varian Halcyon on-board imager as simulated for training:
SAD 1000 mm, detector 384 × 768 pixels at 0.7273 mm (isocenter… See the full description on the dataset page: https://huggingface.co/datasets/jzhu35/EUReCA-demo.
