datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
v1-sft-demoreason-tool-use-demo-1500
Dataset info
The dataset is a selection of reasoning toolcalls data from https://huggingface.co/datasets/interstellarninja/hermes_reasoning_tool_use, which contains data from Hermes-Tools、Glaive-FC、ToolAce、Nvidia-When2Call.
The format has been transformed to adapt llama-factory v1 training pipeline.
DemoFeedbackarmnet-demo-leaderboardScene2Wave-demo-assetsmachine-failure-mlops-demo-logsglm52-demolition-data
GLM-5.2-Demolition — Training & Calibration Data
Apple Silicon AI hub ·
Model release ·
MLX code sample
Preview scope, checked September 10, 2026: the default Hub viewer indexes
87,586 rows (84,231 train, 3,277 validation, 78 test). The original release
total below describes the broader JSONL repository. Use the file browser and
explicit file selections when reusing a particular corpus. The hub includes
a checked download example for the seven-row MLX code sample.
The data… See the full description on the dataset page: https://huggingface.co/datasets/philipjohnbasile/glm52-demolition-data.piperx-demo558-value1500-a50-top10-union20-v1
PiperX advantage-selected teleoperation segments
Only pure human demonstrations. Value checkpoint step1500 (mixed demo+HIL); no HIL frames in this export.
A50 ranked globally across 558 source episodes, top10% AND A>0. Every selected start expands to [t,t+20); overlaps and adjacency merge. Each disconnected component is a separate output episode. Interior frames need not themselves be top10%.
Output: 7146 segments, 322261 frames, 2.983898 hours at30FPS.
Three camera streams and… See the full description on the dataset page: https://huggingface.co/datasets/Elvinky/piperx-demo558-value1500-a50-top10-union20-v1.grok-demon-dataset-ESdemoatomic-metrics-demographic-training-size
Atomic Metrics: Demographic Training-Size Analysis
Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction.
Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.commoncrawl-jobs-demo
Common Crawl on Jobs — datatrove JobsPipelineExecutor demo
228,088 English web documents (~820 MB compressed, 60 jsonl.gz shards) extracted from 1,237,374 Common Crawl pages — the output of a test run of datatrove's experimental JobsPipelineExecutor, which fans a datatrove pipeline out across a pool of Hugging Face Jobs instead of a Slurm cluster.
This is a pipeline demo artifact, not a curated corpus: one segment slice of one crawl, shared as the verifiable receipt for the… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/commoncrawl-jobs-demo.uninavid-objnav-demodataflow-demo-code
DataFlow demo -- Code Pipeline
Project Page | Technical Report | GitHub
This dataset is a demo of the DataFlow Code data processing pipeline from the DataFlow project. It provides a lightweight, inspectable view of what the pipeline produces: curated, execution-checked code SFT supervision pairs.
For full pipeline design and evaluation details, please refer to our technical report: DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of… See the full description on the dataset page: https://huggingface.co/datasets/OpenDCAI/dataflow-demo-code.RPToolkit-demo-datasetRPToolkit is a data generation pipeline, part of Augmentoolkit, that generates synthetic RP sessions inspired by input stories. Basically: feed in Lord of the Rings, get out high fantasy adventure RPs.
This dataset, containing over a million trainable tokens across around 1000 RP sessions, is meant to showcase the capabilities of this pipeline.
The input texts used were: a variety of myths and classic stories from Gutenberg; the first few chapters of some miscellaneous webnovels and… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/RPToolkit-demo-dataset.tla-demotic-v18-premium
Dataset Card for Dataset tla-demotic-v18-premium
This data set contains demotic sentences in transliteration, with lemmatization, with POS glossing and with a German translation.
The data comes from the database of the Thesaurus Linguae Agegyptiae, corpus version 18, and contains only fully intact,
unambiguously readable sentences (13,383 of 31,156 sentences), adjusted for philological and editorial markup.
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-demotic-v18-premium.SwS-Demo-Dataset
Dataset Card for SwS-Demo-Dataset
[🌐 Website] •
[🤗 Demo Dataset] •
[📜 Paper] •
[🐱 GitHub] •
[🐦 Twitter] •
[📕 Rednote]
This dataset is a demo set of synthetic problems generated by SwS, comprising 500 samples for each model and category. The full dataset and model are currently under review by Microsoft and will be released once approved.
Data Loading
from datasets import load_dataset
dataset = load_dataset("MasterVito/SwS-Demo-Dataset")
Data… See the full description on the dataset page: https://huggingface.co/datasets/MasterVito/SwS-Demo-Dataset.trace-demoCLAMP-Sampled-Continuations-and-Demos
CLAMP Sampled Continuations and VLABench Demos
This dataset accompanies HLR/CLAMP, the Constrained Language-Action Model Planner.
It contains 20 closed-loop VLABench case studies: 10 successful and 10 unsuccessful executions. Each case includes the input image and mask, instruction, prompt and entity metadata, CLAMP and ground-truth plans, evaluation output, execution video, and a frame manifest.
Layout
demo_dataset/
├── manifest.json
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/SueMintony/CLAMP-Sampled-Continuations-and-Demos.demoverse-personas-es-v1
Dataset Card for DemoVerse Personas ES v1
Resumen del dataset
demoverse-personas-es-v1 es un dataset de 100.000 personas sinteticas en espanol para Espana, disenado como artefacto publico y como capa operativa para simulacion sociológica.
El dataset se inspira metodologicamente en nvidia/Nemotron-Personas-France, pero no reutiliza sus filas ni intenta replicar la poblacion francesa. La adaptacion reescribe el marco para Espana, con clivajes territoriales, sistema de… See the full description on the dataset page: https://huggingface.co/datasets/apol/demoverse-personas-es-v1.giskard-hub-demo-retaildemospeedup-libero-robocasa-entropy
DemoSpeedup 엔트로피: LIBERO / RoboCasa
기존 GR00T-N1.5 DemoSpeedup 재현에서 실제로 사용한 원본 프레임별 엔트로피입니다.
모델 재추론 없이 구간 분류 역치를 변경하고, 원본 데모를 다른 배속으로 재구성할 수 있습니다.
모델 아키텍처를 변경하는 방법이 아니며, 변환한 데모로 기존 GR00T-N1.5를 학습합니다.
벤치마크
에피소드
프레임
누락/NaN/Inf
LIBERO
1,693
273,465
0
RoboCasa
7,200
2,073,457
0
재현 범위 정정 (2026-09-19)
이 저장소의 엔트로피와 라벨은 우리 GR00T 이식 구현의 결과이며, 논문을 그대로 재현한
결과라고 해석하면 안 됩니다. 논문 §3.2는
클러스터의 평균 엔트로피가 0 미만이면 precision, 그 외(노이즈 포함)는 casual이라고 설명합니다.
공식 공개 코드의… See the full description on the dataset page: https://huggingface.co/datasets/prehj/demospeedup-libero-robocasa-entropy.reddit-demoReddit Demo dataset
fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/apapanikolaou/fair_dataset_demo.fair_dataset_demo
AI4Materials Demo FAIR Perovskites
This is a teaching dataset demonstrating F.A.I.R. hosting on the Hugging Face Hub.It contains a small table of oxide perovskites with band gaps and toy EXTXYZ structures.
Contents
data/table.csv — main tabular data
data/records.jsonl — line-delimited JSON mirror
data/structures/*.xyz — example structures (EXTXYZ)
metadata/schema.json — JSON Schema for validation
CITATION.cff, LICENSE — citation & reuse terms
Provenance… See the full description on the dataset page: https://huggingface.co/datasets/cparidaAI/fair_dataset_demo.EUReCA-demo
EUReCA demo case
One simulated cone-beam acquisition for trying the
EUReCA CBCT reconstruction models
without any data preparation.
Source CT: LUNA16 / LIDC-IDRI series
1.3.6.1.4.1.14519.5.2.1.6279.6001.154837327827713479309898027966 (thorax),
CC BY 3.0. This case is in EUReCA's reserved test split and was never used
for training or model selection.
Geometry: Varian Halcyon on-board imager as simulated for training:
SAD 1000 mm, detector 384 × 768 pixels at 0.7273 mm (isocenter… See the full description on the dataset page: https://huggingface.co/datasets/jzhu35/EUReCA-demo.fetch_hf_term_notion_gh_7942_target_demographics
Demographics
Overview
Anonymized demographic profiles collected through internal surveys.
Usage
Load the dataset with the datasets library.
License
MIT
Status
Documentation pending update.
Provenance
This is an original dataset created by Nimbus Data Labs and is not derived from any external source.
reddit-demo
Reddit Demo dataset
Genesis_AI_Code_1k_Demo
Genesis AI Code (Demo) 1K
Developed by: Within Us AI
Best-of demo subset for instant evaluation and fast adoption.
Splits
train: 1,000
validation: 1,000
Highlights
Tests-as-truth supervision patterns
Diff-first patching
Agentic loops (plan→edit→test→reflect) with bounded budgets
Tool-call trace supervision (where present)
Governance/audit & policy-gate awareness
Storage format
Parquet unavailable (No module named 'pyarrow'); JSONL… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Genesis_AI_Code_1k_Demo.organic-chemistry-reasoning-demo
🧪 Organic Chemistry Reasoning Benchmark (OCRB-200) - DEMO
🛑 This is a DEMO version containing only 10 samples.
🚀 Want the Full Version (236+ Samples)?
The full dataset is available for purchase. It includes deep reasoning chains and trap identification for 236+ complex scenarios.
👉 [Download the Full Dataset Here] (https://7820367248654.gumroad.com/l/pkndc)
